Brain-computer dynamic continuous emotion recognition method, system and device based on channel selection
The optimal channel combination is selected through EEG-Net and non-dominant sorting genetic algorithm, combining short-term and long-term feature extraction networks and attention mechanism layers, and the accuracy problem of dynamic emotions recognition is solved, achieving more efficient emotions recognition effect.
Patent Information
- Application Number
- CN202510314851.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-18
AI Technical Summary
It is difficult for the prior art to accurately monitor and identify dynamic continuous emotions in humans, especially emotional changes under the influence of external environment and social interactions.
EEG-Net and non-dominant sorting genetic algorithm are used to determine the optimal channel combination, combined with short-term flow feature extraction network, long-term flow feature extraction network and attention mechanism layer, emotional recognition is performed through EEG timing signals, and long-term time dependence and global features are captured.
It improves the accuracy of dynamic continuous emotion recognition, can identify emotional changes more accurately, and enhances the stability and feature extraction ability of emotion recognition.
Smart Images

Figure CN119837531B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electroencephalogram (EEG) emotion recognition, and particularly to a method, system, and device for dynamic continuous EEG emotion recognition based on channel selection. Background Art
[0002] Human emotions are internal manifestations integrated from personal behaviors, thinking patterns, and feelings towards objective things. In real life, human emotions are dynamic and diverse, affected by external environments, social interactions, and individual internal states. Emotions exhibit continuous characteristics, having a profound impact on an individual's subsequent behaviors and psychological motivations, and playing a crucial role in the interaction process between subjective feelings and external environments. Under normal conditions, people usually maintain stable emotions, which is one of the most common states. However, when subjected to positive external stimuli, this stable emotion may be enhanced, manifested as positive emotions such as joy. On the contrary, when suffering from negative stimuli, emotions may turn negative, manifested as negative emotions such as sadness and tension. Therefore, it is necessary to accurately monitor the dynamic continuous emotions of humans. Summary of the Invention
[0003] The purpose of this application is to provide a method, system, and device for dynamic continuous EEG emotion recognition based on channel selection, which can capture the long-term time dependence and global characteristics of EEG time series signals, thereby improving the accuracy of dynamic continuous emotion recognition.
[0004] To achieve the above purpose, this application provides the following solutions:
[0005] In the first aspect, this application provides a method for dynamic continuous EEG emotion recognition based on channel selection, including:
[0006] Collect multi-channel EEG time series signals of a subject in a dynamic scenario;
[0007] According to the multi-channel EEG time series signals, use EEG-Net and non-dominated sorting genetic algorithm to determine the optimal channel combination;
[0008] According to the EEG time series signals corresponding to the optimal channel combination, use a pre-trained emotion recognition model to determine the continuous dynamic emotion category of the subject in the dynamic scenario;
[0009] Among them, the emotion recognition model includes a short-term flow feature extraction network, a long-term flow feature extraction network, an attention mechanism layer, and a classifier; the short-term flow feature extraction network is used to extract local features of the EEG time series signal by using a temporal convolutional module; the long-term flow feature extraction network is used to extract global features of the EEG time series signal by using a Transformer module; the attention mechanism layer is used to fuse the local features and the global features to obtain fused features; the classifier is used to determine continuous dynamic emotion categories according to the fused features.
[0010] In a second aspect, the present application provides an EEG dynamic continuous emotion recognition system based on channel selection, including:
[0011] A signal acquisition unit, configured to acquire multi-channel EEG time series signals of a subject in a dynamic scenario;
[0012] A channel selection unit, configured to determine an optimal channel combination according to the multi-channel EEG time series signals by using EEG-Net and non-dominated sorting genetic algorithm;
[0013] An emotion recognition unit, configured to determine continuous dynamic emotion categories of the subject in the dynamic scenario according to the EEG time series signals corresponding to the optimal channel combination by using a pre-trained emotion recognition model.
[0014] Among them, the emotion recognition model includes a short-term flow feature extraction network, a long-term flow feature extraction network, an attention mechanism layer, and a classifier; the short-term flow feature extraction network is used to extract local features of the EEG time series signal by using a temporal convolutional module; the long-term flow feature extraction network is used to extract global features of the EEG time series signal by using a Transformer module; the attention mechanism layer is used to fuse the local features and the global features to obtain fused features; the classifier is used to determine continuous dynamic emotion categories according to the fused features.
[0015] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned EEG dynamic continuous emotion recognition method based on channel selection.
[0016] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0017] The present application provides a method, system and device for EEG-based dynamic continuous emotion recognition using channel selection. First, EEG-Net and non-dominated sorting genetic algorithm are used to determine the optimal channel combination. Then, according to the EEG time series signals corresponding to the optimal channel combination, a short-term flow feature extraction network is used to extract the local features of the EEG time series signals corresponding to the optimal channel combination, and a long-term flow feature extraction network is used to extract the global features of the EEG time series signals corresponding to the optimal channel combination, effectively capturing the long-term time dependence and global features of the EEG time series signals. An attention mechanism layer is used to fuse the local features and global features, and finally, the continuous dynamic emotion category is determined based on the fused features, improving the recognition accuracy of dynamic continuous emotions. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is an application environment diagram of a method for EEG-based dynamic continuous emotion recognition using channel selection according to an embodiment of the present application.
[0020] Figure 2 It is a schematic flowchart of a method for EEG-based dynamic continuous emotion recognition using channel selection according to an embodiment of the present application.
[0021] Figure 3 It is a schematic diagram of the principle of a method for EEG-based dynamic continuous emotion recognition using channel selection according to an embodiment of the present application.
[0022] Figure 4 It is a schematic diagram of the acquisition process of EEG time series signals according to an embodiment of the present application.
[0023] Figure 5 It is a schematic diagram of the structure of a two-stream network according to an embodiment of the present application.
[0024] Figure 6 It is a schematic diagram of the functional modules of a system for EEG-based dynamic continuous emotion recognition using channel selection according to an embodiment of the present application.
[0025] Figure 7 It is a schematic diagram of the structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0027] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] The EEG dynamic continuous emotion recognition method based on channel selection provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the multi-channel EEG time series signals of the subject in a dynamic scenario to the server 104. After receiving the multi-channel EEG time series signals, the server 104 determines the optimal channel combination according to the multi-channel EEG time series signals by using EEG-Net and non-dominated sorting genetic algorithm, and determines the continuous dynamic emotion category of the subject in the dynamic scenario according to the EEG time series signals corresponding to the optimal channel combination by using a pre-trained emotion recognition model. The server 104 can feedback the continuous dynamic emotion category of the subject in the dynamic scenario to the terminal 102. In addition, in some embodiments, the EEG dynamic continuous emotion recognition method based on channel selection can also be implemented separately by the server 104 or the terminal 102.
[0029] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0030] In an exemplary embodiment, as Figure 2 and Figure 3 shown, a method for EEG dynamic continuous emotion recognition based on channel selection is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, this method is applied to Figure 1Taking server 104 in [the context] as an example for illustration, it includes the following steps 201 to 204.
[0031] Step 201, collect multi-channel EEG time series signals of the subject in a dynamic scenario.
[0032] In an exemplary embodiment, create an EEG acquisition paradigm in a dynamic scenario. Based on the EEG acquisition paradigm, use a 64-channel electrode cap to collect multi-channel EEG time series signals of the subject under target stimuli. When training EEG-Net, simultaneously collect dynamic emotion labels corresponding to the multi-channel EEG time series signals, and perform normalization processing on the dynamic emotion labels. The dynamic emotion labels include arousal and pleasure scores of six dynamic emotions.
[0033] Specifically, organize experts to analyze the application scenarios and possible situations, and design a relatively complete paradigm that can accurately reflect the brain activities in this cognitive state. Under the isolation of external environmental interference, play multiple dynamic emotion video clips that induce different dynamic emotions for the subject. During the process of the subject watching the dynamic emotion video clips, collect the discharge signals of the cerebral cortex of the subject under different cognitive activities, and after watching a dynamic emotion video clip, perform pleasure and arousal scoring on the dynamic emotion video clip. After the scoring is completed until the EEG signal of the subject is in a stable state, then watch the next dynamic emotion video clip until all dynamic emotion video clips are watched and scored. Comprehensive evaluation of the state of the subject at the acquisition site and the quality of the collected EEG signals, and eliminate the EEG signals with poor performance and those that cannot accurately reflect the dynamic emotion process.
[0034] As a specific implementation manner, the acquisition process of the multi-channel EEG time series signals is as follows.
[0035] 1) Select 5 experts to participate in the selection of dynamic emotion video clips. All of them have received professional film training and have experience in rehearsing films; select 24 healthy subjects to participate in the experiment, including 15 males and 9 females, with an average age of (20 ± 1.5) years.
[0036] 2) Wear an EEG signal acquisition device for the subject in a quiet room with appropriate temperature and brightness.
[0037] 3) The subject conducts the experiment in a dedicated environment that can shield external interference. During the experiment, always set the brightness at a comfortable level so that the subject can clearly see the content of the dynamic emotion video clips. To ensure the smooth progress of the experiment, before the official start of the experiment, require each subject to fill out a mental health assessment questionnaire, and train each subject with additional dynamic emotion video clips.
[0038] 4) AsFigure 4 As shown, at the beginning of the experiment, a prompt and precautions were played first. The experiment was divided into 3 sections. For each section, a dynamic emotion video that induced two emotions was played for the subjects. There were a total of six dynamic emotion videos that induced different dynamic emotions. Each dynamic emotion video had a duration of approximately 180 seconds, with each emotion lasting 90 seconds. Each dynamic emotion video was segmented into 5-second units, and each dynamic emotion video had a total of 36 segmented clips. The presentation order of each dynamic emotion video was fixed. The subjects first adjusted their own emotions to a relatively relaxed state. When the brainwaves of the subjects were stable, the dynamic emotion video clips were played.
[0039] 5) Emotion assessment: The subjects rated the arousal and pleasure levels of the two different dynamic emotion videos in the dynamic emotion video, and also rated the arousal and pleasure levels of the 36 5-second segmented clips of the dynamic emotion video. The rating included 9 levels. First, the arousal level was rated, with the numbers 1 to 9 corresponding to "very relaxed" to "very intense" respectively, and the number 5 representing "calm". Then, the pleasure level was rated, with the numbers 1 to 9 corresponding to "very negative" to "very positive" respectively, and the number 5 representing "normal". The subjects filled in the ratings according to their true feelings.
[0040] 6) Set a rest time. After each dynamic emotion video was rated, the subjects rested for 15 to 30 seconds to keep their brainwaves in a stable state. Then, the above process was continuously cycled until the subjects completed all processes of the entire paradigm and the experiment ended.
[0041] This application can fully reflect the expression and change of emotions in a dynamic environment by obtaining the electroencephalogram (EEG) time-series signals and dynamic emotion labels of the subjects in a dynamic scenario, thereby providing rich emotion samples for the training of EEG-Net and emotion recognition models, greatly improving the practicality of emotion recognition, and filling the gap in the field of EEG emotion recognition in the direction of dynamic emotion recognition.
[0042] Step 202: Perform filtering, artifact removal, and resampling processing on the multi-channel EEG time-series signals in sequence. Further perform differential operation on the multi-channel EEG time-series signals to obtain differential entropy features.
[0043] Step 203: According to the multi-channel EEG time-series signals, use EEG-Net and Non-dominated Sorting Genetic Algorithm (NSGA) to determine the optimal channel combination.
[0044] Among them, EEG-Net is obtained by pre-training with a training sample set. The training sample set includes preprocessed multi-channel EEG time-series sample signals and corresponding dynamic emotion labels. The non-dominated sorting genetic algorithm uses NSGA-II.
[0045] In an exemplary embodiment, the EEG time-series signal of each channel is input into EEG-Net, and the scores of each channel (including classification accuracy and computational complexity) are sorted according to the loss score. The channels with scores higher than the average value of all channels are screened out, and the non-dominated sorting genetic algorithm II is used to find the optimal channel combination. Step 203 includes the following steps 31 to step 35.
[0046] Step 31, determining an initial population according to the multi-channel EEG time-series signal. An individual in the initial population represents a channel combination.
[0047] Step 32, for the m th iteration, using EEG-Net to classify the EEG time-series signal of each channel in the population at the m th iteration respectively, and determining the classification accuracy and computational complexity of each channel; m >0; the population at the first iteration is the initial population.
[0048] The EEG-Net includes a two-dimensional convolutional layer, a depth convolutional layer, a channel attention layer, and a separable convolutional layer connected in sequence.
[0049] (1) The two-dimensional convolutional layer is used to extract the time-frequency features of the EEG time-series signal, mainly focusing on feature extraction in the channel dimension and time dimension. In the two-dimensional convolutional layer, a two-dimensional convolutional kernel with a size of (1, 64) is used for feature extraction, indicating that the feature extraction range is between 2 Hz and 128 Hz. The formula is as follows:
[0050] ;
[0051] Among them, represents the time-frequency feature of the n th sample at the th channel at the th time point, represents the signal value of the n th sample at the th channel at the th time point, represents the size of the convolutional kernel of the two-dimensional convolutional layer, represents the sample serial number, and each sample is a signal extracted from the EEG activity data in a time period or a certain experimental task, f represents the serial number of the frequency component extracted after the convolutional operation, Indicates the channel number. Channels represent different electrode positions in an EEG recording, and each channel corresponds to an electrode at a specific position. Indicates a time point. Indicates the size of the convolutional kernel in the channel dimension. Indicates the size of the convolutional kernel in the time dimension. c Indicates the index in the channel dimension. Indicates the index in the time dimension.
[0052] (2) In the depth convolutional layer, depth convolution operations are performed separately on the time-frequency features of each channel output by the two-dimensional convolutional layer, so as to extract higher-level spatial features from the relationships between different channels, in order to better understand the differences between channels. The formula for the depth convolutional layer is as follows:
[0053] ;
[0054] Where, Indicates the n th sample at the th channel at the th time point, Indicates the size of the convolutional kernel of the depth convolutional layer.
[0055] (3) The channel attention layer introduces a channel attention mechanism to perform global average pooling on each channel, obtaining the global features of each channel and generating a scalar value to measure the importance of each channel, and then assigning an attention weight to each channel, so as to strengthen the attention of EEG-Net to important channel features. The formula for the channel attention mechanism is as follows:
[0056] ;
[0057] Where, Indicates the n th sample at the th channel.
[0058] (4) The separable convolutional layer is used to integrate the global features of each channel output by the channel attention layer. The formula is as follows:
[0059] ;
[0060] Where, Indicates the n th sample at the th channel at the th time point of the final feature.
[0061] Further classify the EEG time series signals of each channel based on the final features output by the separable convolutional layer, and determine the classification accuracy and computational complexity of each channel in combination with the true labels.
[0062] Step 33: According to the classification accuracy and computational complexity of each channel, perform non-dominated sorting on the population at the m -th iteration, and divide the individuals in the population at the m -th iteration into multiple non-dominated fronts.
[0063] Step 34: Calculate the crowding distance of the individuals in each non-dominated front.
[0064] Step 35: According to the crowding distance of the individuals in each non-dominated front, determine the population at the m +1-th iteration, and perform the m +1-th iteration. Until the preset termination condition is met, determine the optimal channel combination according to the population when the preset termination condition is met.
[0065] Specifically, a new population is selected based on non-dominated sorting and crowding distance, and then a new offspring population is generated through crossover and mutation operations.
[0066] In order to further optimize the performance of the EEG-Net model and reduce the computational complexity, this application uses the non-dominated sorting genetic algorithm II for multi-objective channel screening, and determines a set of optimal channel combinations that perform well in terms of classification accuracy and computational complexity through NSGA-II.
[0067] Step 204: According to the EEG time series signals corresponding to the optimal channel combination, use a pre-trained emotion recognition model to determine the continuous dynamic emotion categories of the subject in a dynamic scenario. The EEG time series signals input to the emotion recognition model are the signals after preliminary feature extraction.
[0068] The continuous dynamic emotion categories include joy to calm, calm to joy, sadness to calm, calm to sadness, tension to calm, and calm to tension.
[0069] As Figure 5 shown, the emotion recognition model is based on the Transformer and Convolutional Neural Network (Transformer-based Convolutional Neural Network, TCNN) algorithms. The EEG time series signals corresponding to the optimal channel combination are used to extract local features and global features through the short-term stream and the long-term stream respectively, and then the attention mechanism is used for feature fusion. The emotion recognition model includes a short-term stream feature extraction network, a long-term stream feature extraction network, an attention mechanism layer, and a classifier.
[0070] (1) The short-term flow feature extraction network is used to extract local features of EEG time series signals by using a temporal convolutional module. The main function of the short-term flow feature extraction network is to extract local features within a short time span.
[0071] In an exemplary embodiment, the temporal convolutional module includes a causal convolutional layer, a dilated convolutional layer, a residual connection, a ReLU activation function, and a normalization layer.
[0072] ① The input of the causal convolutional layer is the EEG time series signal. Through causal convolution, convolution operations are performed on the time series data to ensure that the output of each time step depends only on the current time step and previous inputs, and the output is the short-term emotion feature. The causal convolutional layer can ensure the causality of the signal, that is, the convolution operation does not introduce future information, thus avoiding data leakage and adapting to the characteristics of time series data. The formula of the causal convolutional layer is as follows:
[0073] ;
[0074] Where, represents the output of the causal convolutional layer at the t -th time step, represents the input data of the causal convolutional layer at the t - j -th time step, represents the convolutional kernel weight of the causal convolutional layer, represents the size of the convolutional kernel of the causal convolutional layer, represents the bias term of the causal convolutional layer.
[0075] ② The input of the dilated convolutional layer is the short-term emotion feature output by the causal convolutional layer. Through dilated convolution, the receptive field of the convolutional kernel is expanded to capture features with long-term dependencies, and the output is the dilated short-term emotion feature. The dilated convolutional layer introduces a dilation rate, adds gaps between convolutional kernels to expand the receptive field, and at the same time does not increase the number of parameters. The formula is as follows:
[0076] ;
[0077] Where, represents the output of the dilated convolutional layer at the t -th time step, represents the input signal of the dilated convolutional layer at the -th time step, d 1 represents the dilation rate, represents the convolutional kernel weight of the dilated convolutional layer, represents the bias term of the dilated convolutional layer.
[0078] ③ The input of the residual connection is the short-term emotion features output by the causal convolutional layer and the dilated short-term emotion features. The short-term emotion features output by the causal convolutional layer are added to the dilated short-term emotion features, and the output is the features after the residual connection.
[0079] In this application, a residual connection is introduced between the input and the convolutional output to help information flow in the network, solve the problem of vanishing gradients in deep networks, and stabilize the training process. The formula is as follows:
[0080] ;
[0081] where, represents the output after the residual connection at the t -th time step, represents the input signal at the t -th time step, represents the feature map extracted through the convolution operation, and represents the convolutional kernel parameters.
[0082] ④ The input of the ReLU activation function is the features after the residual connection. Through the ReLU activation function, the values less than zero in the features after the residual connection are set to zero, and the values greater than zero remain unchanged. The output is the features after ReLU activation. The formula of the ReLU activation function is as follows:
[0083] ;
[0084] where, represents the output of the activation function at the t -th time step.
[0085] ⑤ The input of the normalization layer is the features after ReLU activation. The normalization layer normalizes the features after ReLU activation, and the output is the local features. The normalization layer ensures the stability of the feature distribution and accelerates the training. The formula is as follows:
[0086] ;
[0087] where, represents the output after normalization at the t -th time step, represents the mean, represents the variance, and represent the scaling parameters, represents the translation parameter, and represents the smoothing term to prevent the denominator from being zero.
[0088] In an exemplary embodiment, during the training process of the short-term flow feature extraction network, L2 regularization is used in the loss function to prevent overfitting. The formula is as follows:
[0089] ;
[0090] Where, represents the regularization loss term, H represents the number of weights in the short-term flow feature extraction network, represents the h th weight in the short-term flow feature extraction network, represents the regularization coefficient, which controls the intensity of the penalty term.
[0091] During the training process of the short-term flow feature extraction network, Dropout is further used to randomly "drop out" a part of the neurons in the neural network to reduce the complexity of the short-term flow feature extraction network, thereby improving the generalization ability. The formula is as follows:
[0092] ;
[0093] Where, represents the output after Dropout at the t th time step, represents a random mask matrix with the same dimension as , represents that the neuron is retained with probability .
[0094] (2) The long-term flow feature extraction network is used to extract the global features of the EEG time series signal by adopting a Transformer module. The main function of the long-term flow feature extraction network is to capture the long-term time dependence of the emotion data.
[0095] In an exemplary embodiment, the Transformer module includes a feature encoding layer, a multi-head self-attention layer, and a feed-forward neural network connected in sequence.
[0096] The process of the long-term flow feature extraction network adopting the Transformer module to extract the global features of the EEG time series signal includes the following steps 46 to 48.
[0097] Step 46: Add positional encoding to the EEG time series signal through the feature encoding layer to obtain the encoded EEG time series signal. Since the Transformer module itself does not have the ability to process sequence order information, it is necessary to add order information to the EEG signal at each time step through positional encoding so that the Transformer module can recognize the temporal dependence relationship in the EEG time series signal. Positional encoding is achieved by adding a vector representing the position of the time step to the EEG signal at each time step.
[0098] The feature encoding layer adds positional encoding to the EEG time series signal using the following formula:
[0099] ;
[0100] ;
[0101] ;
[0102] where, is the encoded EEG time series signal, X is the EEG time series signal, PE is the positional encoding, is the time step, is the feature dimension, is the index of the feature dimension, is the th positional encoding of the th feature dimension at the th time step. is the positional encoding of the 2 +1th feature dimension at the
[0103] Step 47: Extract the global relationship of the encoded EEG time series signal through the multi-head self-attention layer and generate multi-head attention features.
[0104] Through the parallel calculation of the multi-head self-attention layer, it is possible to simultaneously focus on multiple attention heads, thereby capturing information from different parts. The calculation formula for each attention head is as follows:
[0105] ;
[0106] ;
[0107] ;
[0108] where, represents the attention feature output by the th attention head, represents the The attention scores output by each attention head, denotes the query vector of the th attention head, denotes the th key vector of the attention head, denotes the th value vector of the attention head, denotes the scaling factor, denotes the query weight matrix, denotes the key weight matrix,
[0109] Step 48, transform the multi-head attention features through the feed-forward neural network to obtain global features.
[0110] The feed-forward neural network consists of two linear layers and a non-linear activation function, and is used to further transform the features at each time step. The formula is as follows:
[0111] ;
[0112] where, L denotes the global feature, denotes the feed-forward neural network, denotes the multi-head attention feature, and are weight matrices, and are bias terms.
[0113] (3)The attention mechanism layer is used to fuse the local features and the global features to obtain fused features.
[0114] This application uses the attention mechanism layer to dynamically generate fusion weights, and specifically calculates the attention weights of the local features and the global features using the following formula:
[0115] ;
[0116] ;
[0117] where, denotes the attention weight of the local feature, denotes the attention weight of the global feature, is the bias of the global feature, is the bias term of the local feature, S denotes the local feature, L denotes the global feature, denotes the weight matrix of the local feature, denotes the weight matrix of the global feature.
[0118] (4) The classifier is used to determine continuous dynamic emotion categories based on the fused features.
[0119] In this application, by constructing a two-stream network structure, the long-term stream feature extraction network ensures that the input sequence can correctly express the temporal information through feature encoding, then captures the global dependencies through the multi-head self-attention mechanism, enhances the attention ability of the emotion recognition model to different features, and finally performs non-linear transformation through the feed-forward neural network to generate high-level feature representations, effectively capturing the long-term temporal dependencies and global features of the EEG temporal signals. Furthermore, it overcomes the deficiencies of the TCNN in difficultly capturing long-term emotion changes and lacking global information integration, has higher feature extraction accuracy and stronger ability to capture emotion changes, and thus improves the accuracy and stability of emotion recognition.
[0120] Based on the same inventive concept, the embodiment of this application also provides an EEG dynamic continuous emotion recognition system for implementing the above-mentioned EEG dynamic continuous emotion recognition method. The implementation solutions provided by this system to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the EEG dynamic continuous emotion recognition system provided below can refer to the limitations on the EEG dynamic continuous emotion recognition method in the above text, and will not be elaborated here.
[0121] In an exemplary embodiment, as Figure 6 shown, an EEG dynamic continuous emotion recognition system is provided, including: a signal acquisition unit 601, a channel selection unit 602, and an emotion recognition unit 603.
[0122] The signal acquisition unit 601 is used to acquire multi-channel EEG temporal signals of a subject in a dynamic scenario.
[0123] The channel selection unit 602 is used to determine an optimal channel combination according to the multi-channel EEG temporal signals by using EEG-Net and the non-dominated sorting genetic algorithm.
[0124] The emotion recognition unit 603 is used to determine the continuous dynamic emotion categories of the subject in the dynamic scenario according to the EEG temporal signals corresponding to the optimal channel combination by using a pre-trained emotion recognition model.
[0125] Among them, the emotion recognition model includes a short-term stream feature extraction network, a long-term stream feature extraction network, an attention mechanism layer, and a classifier. The short-term stream feature extraction network is used to extract local features of the EEG temporal signals by using a temporal convolutional module. The long-term stream feature extraction network is used to extract global features of the EEG temporal signals by using a Transformer module. The attention mechanism layer is used to fuse the local features and global features to obtain fused features; the classifier is used to determine continuous dynamic emotion categories according to the fused features.
[0126] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as shown in Figure 7 the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multi-channel EEG time series signals of the subject in a dynamic scenario. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for EEG-based dynamic continuous emotion recognition based on channel selection.
[0127] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0128] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0129] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0131] In the present application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.
[0132] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0133] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0134] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0135] In this text, specific examples are used to illustrate the principles and implementation modes of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for dynamic continuous emotion recognition of electroencephalogram based on channel selection, characterized in that The brain-computer electroencephalogram (EEG) dynamic continuous emotion recognition method based on channel selection includes the following steps: Collect multi-channel EEG time series signals of a subject in a dynamic scenario; Based on the multi-channel EEG time-series signal, use EEG-Net and non-dominated sorting genetic algorithm to determine the optimal channel combination; based on the multi-channel EEG time-series signal, use EEG-Net and non-dominated sorting genetic algorithm to determine the optimal channel combination, specifically including: determining the initial population according to the multi-channel EEG time-series signal; an individual in the initial population represents a channel combination; for the m th iteration, use EEG-Net to classify the EEG time-series signal of each channel in the population at the m th iteration respectively, and determine the classification accuracy and computational complexity of each channel; m >0; the population at the first iteration is the initial population; according to the classification accuracy and computational complexity of each channel, perform non-dominated sorting on the population at the m th iteration, and divide the individuals in the population at the m th iteration into multiple non-dominated fronts; calculate the crowding distance of the individuals in each non-dominated front; according to the crowding distance of the individuals in each non-dominated front, determine the population at the m +1th iteration, and perform the m +1th iteration until the preset termination condition is satisfied, and determine the optimal channel combination according to the population when the preset termination condition is satisfied; in order to further optimize the performance of the EEG-Net model and reduce the computational complexity, use the non-dominated sorting genetic algorithm II for multi-objective channel selection, and determine a set of optimal channel combinations that perform well in classification accuracy and computational complexity through NSGA-II; Based on the EEG time series signals corresponding to the optimal channel combination, use a pre-trained emotion recognition model to determine the continuous dynamic emotion categories of the subject in the dynamic scenario. By constructing a two-stream network structure, the long-term stream feature extraction network ensures that the input sequence can correctly express the time series information through feature encoding, then captures the global dependencies through multi-head self-attention, enhances the attention ability of the emotion recognition model to different features, and finally performs a non-linear transformation through a feed-forward neural network to generate high-level feature representations, effectively capturing the long-term time dependencies and global features of the EEG time series signals; Among them, the emotion recognition model includes a short-term stream feature extraction network, a long-term stream feature extraction network, an attention mechanism layer, and a classifier. The short-term stream feature extraction network is used to extract local features of the EEG time series signals using a temporal convolutional module. The long-term stream feature extraction network is used to extract global features of the EEG time series signals using a Transformer module. The attention mechanism layer is used to fuse the local features and the global features to obtain fused features. The classifier is used to determine the continuous dynamic emotion categories based on the fused features.
2. The method for dynamically and continuously recognizing brain - electrical - signal - based emotions by channel selection according to claim 1, wherein, The EEG-Net includes a two-dimensional convolutional layer, a depth convolutional layer, a channel attention layer, and a separable convolutional layer connected in sequence.
3. The method for dynamically and continuously recognizing brain-computer based on channel selection according to claim 1, characterized in that, The temporal convolutional module includes a causal convolutional layer, a dilated convolutional layer, a residual connection, a ReLU activation function, and a normalization layer; The input of the causal convolutional layer is the EEG time series signal. Convolution operations are performed on the time series data through causal convolution to ensure that the output of each time step depends only on the current time step and previous inputs, and the output is short-term emotion features; The input of the dilated convolutional layer is the short-term emotion features output by the causal convolutional layer. The receptive field of the convolutional kernel is expanded through dilated convolution to capture features with long-term dependencies, and the output is the dilated short-term emotion features; The input of the residual connection is the short-term emotion features output by the causal convolutional layer and the dilated short-term emotion features. The short-term emotion features output by the causal convolutional layer are added to the dilated short-term emotion features, and the output is the features after the residual connection; The input of the ReLU activation function is the features after the residual connection. The values less than zero in the features after the residual connection are set to zero through the ReLU activation function, and the values greater than zero remain unchanged. The output is the features after ReLU activation; The input of the normalization layer is the features after ReLU activation. The features after ReLU activation are normalized through the normalization layer, and the output is local features.
4. The method for dynamically and continuously recognizing brain-computer based on channel selection according to claim 1, characterized in that The Transformer module includes a feature encoding layer, a multi-head self-attention layer, and a feed-forward neural network connected in sequence; The process of the long-term stream feature extraction network using the Transformer module to extract global features of the EEG time series signals includes: Adding positional encoding to the EEG time series signal through the feature encoding layer to obtain the encoded EEG time series signal; Extract the global relationship of the encoded EEG time series signal through the multi-head self-attention layer and generate multi-head attention features; Transform the multi-head attention features through the feed-forward neural network to obtain global features.
5. The method for dynamically and continuously recognizing brain-computer emotion based on channel selection according to claim 4, characterized in that The feature encoding layer adds position encoding to the EEG time series signal using the following formula: ; ; ; Among them, is the encoded EEG time series signal, X is the EEG time series signal, PE is the position encoding, is the time step, is the feature dimension, is the index of the feature dimension, is the th time step and the position encoding of the 2nd th feature dimension, is the th time step and the position encoding of the 2nd +1th feature dimension.
6. The method for electroencephalogram-based dynamic continuous emotion recognition based on channel selection according to claim 1, wherein The continuous dynamic emotion categories include joy to calm, calm to joy, sadness to calm, calm to sadness, tension to calm, and calm to tension.
7. The method for dynamically and continuously recognizing EEG-based emotions based on channel selection according to claim 1, wherein The EEG dynamic continuous emotion recognition method based on channel selection further includes: Perform filtering, artifact removal, and resampling processing on the multi-channel EEG time series signal in sequence.
8. A brain-computer electroencephalogram dynamic continuous emotion recognition system based on channel selection, which is applied to the brain-computer electroencephalogram dynamic continuous emotion recognition method according to any one of claims 1-7, and is characterized in that, The EEG dynamic continuous emotion recognition system based on channel selection includes: A signal acquisition unit for acquiring multi-channel EEG time series signals of a subject in a dynamic scenario; A channel selection unit, configured to determine an optimal channel combination according to the multi-channel EEG time series signals by using EEG-Net and the non-dominated sorting genetic algorithm; determining an optimal channel combination according to the multi-channel EEG time series signals by using EEG-Net and the non-dominated sorting genetic algorithm, specifically including: determining an initial population according to the multi-channel EEG time series signals; an individual in the initial population represents a channel combination; for the m th iteration, using EEG-Net to classify the EEG time series signals of each channel in the population at the m th iteration respectively, and determining the classification accuracy and computational complexity of each channel; m >0; the population at the first iteration is the initial population; according to the classification accuracy and computational complexity of each channel, performing non-dominated sorting on the population at the m th iteration, and dividing the individuals in the population at the m th iteration into multiple non-dominated fronts; calculating the crowding distance of the individuals in each non-dominated front; according to the crowding distance of the individuals in each non-dominated front, determining the population at the m +1th iteration, and performing the m +1th iteration until the preset termination condition is satisfied, and determining the optimal channel combination according to the population when the preset termination condition is satisfied; in order to further optimize the performance of the EEG-Net model and reduce the computational complexity, the non-dominated sorting genetic algorithm II is used for multi-objective channel screening, and a group of optimal channel combinations that perform well in terms of classification accuracy and computational complexity are determined through NSGA-II; An emotion recognition unit for determining the continuous dynamic emotion category of the subject in the dynamic scenario according to the EEG time series signal corresponding to the optimal channel combination, using a pre-trained emotion recognition model; by constructing a two-stream network structure, the long-term stream feature extraction network ensures that the input sequence can correctly express the time series information through feature encoding, then captures the global dependence relationship through multi-head self-attention, enhances the attention ability of the emotion recognition model to different features, and finally performs non-linear transformation through a feed-forward neural network to generate a high-level feature representation, effectively capturing the long-term time dependence and global features of the EEG time series signal; Among them, the emotion recognition model includes a short-term stream feature extraction network, a long-term stream feature extraction network, an attention mechanism layer, and a classifier; the short-term stream feature extraction network is used to extract local features of the EEG time series signal using a temporal convolutional module; the long-term stream feature extraction network is used to extract global features of the EEG time series signal using a Transformer module; the attention mechanism layer is used to fuse local features and global features to obtain fused features; the classifier is used to determine the continuous dynamic emotion category according to the fused features.
9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the EEG dynamic continuous emotion recognition method based on channel selection according to any one of claims 1-7.
Citation Information
Patent Citations
Emotional state regulation and control method and system based on EEG signal
CN113208626A
Electroencephalogram signal recognition model training method, recognition method, device and equipment
CN114418026A
Model compression method of electroencephalogram deep neural network
CN114580629A
Long-time electroencephalogram emotion recognition method based on electroencephalogram micro-state
CN117609863A
Electroencephalogram emotion recognition method and device under dynamic situation, medium and product
CN118490231A