Emotion recognition method and system based on electroencephalogram eye movement multi-mode cross-attention feature fusion
Through the multimodal cross-attention feature fusion method of EEG eye movement, the problems of single-modal feature incomplete information and poor multimodal interaction are solved, and high accuracy and high real-time emotion recognition is achieved, which is suitable for multimodal physiological signal interaction systems.
Patent Information
- Application Number
- CN202510324236.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
AI Technical Summary
The existing deep learning-based emotion recognition methods have problems such as incomplete single-modal feature information, low recognition accuracy, poor adaptability of multimodal feature extraction, and poor interactiveness of multimodal physiological signal interaction systems.
The emotional recognition method based on the fusion of multimodal cross-attention feature of EEG eye movement is adopted. Through the dynamic graph EEG signal feature extraction module and the multidimensional eye movement signal feature extraction module, combined with the multimodal cross-attention feature fusion module, a deep learning model is built to realize efficient fusion of multimodal data and emotional state recognition.
On the public data set SEED, a high accuracy recognition with an average accuracy of 97.94%, and an identification accuracy of 87.4% was achieved in the multimodal emotional interaction system, enhancing the interactivity and real-timeness of emotional recognition.
Smart Images

Figure CN120296550A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a method and system for emotion recognition based on cross-attention feature fusion of electroencephalogram and eye movement multimodality. Background Art
[0002] Emotion recognition is a key research area for accurately understanding and identifying human emotions. Emotions play a very important role in daily life, not only in interpersonal communication, but also in the perception of the world, and in human decision-making and behavior. The mental state of a person is relatively complex, with a relatively short duration, and is accompanied by corresponding physiological or expressive responses, which may include language, behavior, neural mechanisms, etc., such as being reflected in the changes of their own electroencephalogram signals and eye movement signals.
[0003] With the development of computer technology and artificial intelligence technology, human-computer interaction, especially emotional interaction, has become particularly important. Emotional computing simply means integrating technology and emotions into human-computer interaction, simulating emotional interaction between humans and computers by measuring the emotional state of users, and calculating the emotional state of people's hearts. At the same time, electroencephalogram technology and eye movement technology have gradually matured and have received extensive attention from researchers due to their non-invasive, fast, and easy-to-use characteristics. By searching relevant materials, papers, and research results, it is confirmed that the feasibility of using electroencephalogram signals and eye movement signals for mental state detection.
[0004] As the core technology of multimodal emotion recognition, deep learning methods can automatically learn and extract emotion-related features from multi-source data. It can effectively fuse data of different modalities by constructing complex machine learning models and establish a deep association between emotions and data. At the same time, through continuous training and optimization, deep learning algorithms have significantly improved the generalization ability of the model, making it perform well in various emotion recognition tasks. However, multimodal data usually has high dimensions and complex structures, and extracting useful emotion-related features from these data is still a major challenge. To address this challenge, representation learning between features becomes particularly crucial. By learning effective representations, the correlation between different modalities and their unique descriptions of emotional states can be better captured. In addition, since different modalities have different contributions and importance in emotion recognition tasks, how to reasonably weigh and determine the weights of each modality, make full use of their respective information, and avoid the negative impact of unimportant or redundant information on emotion recognition has become a key problem to be solved urgently. This requires comprehensively considering the characteristics of each modality and the complex relationships between data in model design to achieve more efficient and accurate emotion recognition.
[0005] On the other hand, an emotion interaction system may become an important intervention means for treating mental diseases such as depression and monitoring the state of patients. The interaction system can adjust the interaction interface in real time and provide a personalized user experience, thereby enhancing user satisfaction and efficiency. At the same time, in the field of mental health intervention, through multimodal emotion recognition, more reasonable solutions can be provided to researchers, and a closed-loop emotion interaction system can also directly provide real-time feedback to the subjects to help them better evaluate their own emotional states. Summary of the Invention
[0006] To solve the problems of incomplete single-modal feature information, low recognition accuracy, and poor adaptability of multi-modal feature extraction in existing deep learning-based emotion recognition methods, the present invention proposes a multi-modal cross-attention feature fusion emotion recognition method based on electroencephalogram and electrooculogram. Through five steps of data acquisition, data processing, model construction, method testing, and method verification, high-accuracy recognition and high-real-time recognition of emotional states are achieved. The performance of this method is evaluated and experimentally verified on the internationally public multi-modal emotion dataset SEED, and the average accuracy rate reaches 97.94%, which is better than other existing optimal methods. To solve the problem of poor interactivity of existing interaction systems based on multi-modal physiological signals, the present invention develops a set of emotion recognition interaction systems based on multi-modal cross-attention feature fusion of electroencephalogram and electrooculogram, forming a complete operation platform with real-time display and warning functions. The real-time effectiveness of this method is verified on the movement imagination data of 8 subjects collected, and the average accuracy rate reaches 87.4%, meeting the basic application conditions.
[0007] First, preprocess the obtained multi-modal physiological signal data and establish corresponding training sets and test sets; secondly, construct a multi-modal cross-attention feature fusion emotion recognition model based on electroencephalogram and electrooculogram; then input the preprocessed training set data into the emotion recognition model for model training, update the model parameters and save the model. Next, input the preprocessed test set into the model for performance testing of the method, and finally output the results; finally, build a multi-modal emotion interaction system, collect the multi-modal physiological signals of 8 subjects, including electroencephalogram signals and electrooculogram signals, and conduct effectiveness verification of the method.
[0008] A multi-modal cross-attention feature fusion emotion recognition method based on electroencephalogram and electrooculogram according to an embodiment of the present invention includes the following steps:
[0009] Step 1: Obtain the publicly available emotion recognition dataset SEED;
[0010] Step 2: Preprocess the obtained publicly available dataset and establish corresponding training sets and test sets;
[0011] Step 3: Use the Pytorch deep learning framework to construct an emotion recognition model based on cross-attention feature fusion of EEG and eye movement multimodality;
[0012] Step 4: Input the preprocessed training set in Step 2 into the emotion recognition model for model training; then input the preprocessed test set in Step 2 into the trained emotion recognition model for performance testing of the method, and conduct a performance comparison with the existing state-of-the-art methods;
[0013] Step 5: Build a multimodal emotion interaction system, use the system to collect the emotion recognition EEG signals and eye movement signals of 8 subjects, and verify the effectiveness of the method by recording the recognition accuracy during the interaction process; where:
[0014] In Step 2, data preprocessing includes operations such as data screening, band-pass filtering, baseline correction, re-referencing, and bad channel rejection;
[0015] In Step 3, the emotion recognition model based on cross-attention feature fusion of EEG and eye movement multimodality includes: a dynamic graph EEG signal feature extraction module, a multi-dimensional eye movement signal feature extraction module, a multimodal cross-attention feature fusion module, and an emotion state recognition and classification module, where: the dynamic graph EEG signal feature extraction module is used to extract the deep spatio-temporal domain features of EEG signals, the multi-dimensional eye movement signal feature extraction module is used to extract the deep time domain features of eye movement signals, and then the high-level features of the two modalities are input into the multimodal attention feature fusion module for fusing EEG signal features and eye movement signal features. Finally, the fusion result is input into the emotion state classification block to output the result, and the emotion recognition result includes discrete emotion states such as positive, negative, and neutral;
[0016] In Step 4, use the training set to train the model by the method of five-fold cross-validation, then save the model training parameters, and use the test set to conduct performance verification such as accuracy testing of the model.
[0017] In Step 5, build an online emotion recognition system for EEG and eye movement multimodality. Use the system to first collect the scalp EEG data and eye movement data of the 8 subjects; secondly, preprocess the data according to the method in Step 2, and establish a training set and a test set for the collected data; then input the preprocessed training set into the emotion recognition model constructed in Step 3, and train the emotion recognition model according to the method described in Step 4; then, use the system to collect the scalp EEG data and eye movement signals of the 8 subjects again, and verify the effectiveness of the method by real-time recording of the system recognition accuracy during the physiological signal data, data preprocessing, emotion recognition, and interaction feedback processes.
[0018] The main advantages of an emotion recognition method and system based on cross-attention feature fusion of EEG and eye movement multimodality proposed by the present invention include:
[0019] 1. The dynamic EEG signal feature extraction module designed by the present invention combines the EEG time-domain features with the positional features between EEG channels to dynamically and deeply extract the spatio-temporal features of EEG information; the multi-dimensional eye movement signal feature extraction module uses the attention mechanism to assign different weights to different eye movement features and deeply extracts the time-domain features of eye movement information. The combination of the two overcomes the problems of lack of modality targeting and self-adaptability in multi-modal feature extraction, and effectively improves the accuracy of emotion recognition;
[0020] 2. The present invention utilizes a multi-modal cross-attention feature fusion module, which deeply fuses multi-modal physiological signals through two parts: a low-level cross-attention module and a high-level cross-attention module, solves the problems of insufficient single-modal feature extraction and low recognition accuracy, effectively improves the recognition accuracy, and validates the effectiveness of this method on the publicly available dataset SEED dataset. The average accuracy of intention reaches 97.94%, which is better than other existing optimal methods;
[0021] 3. The multi-modal emotion interaction system of the present invention integrates the brain-computer interface, eye movement device and real-time feedback warning in an integrated manner, enhances the interactivity of emotion recognition, and validates the online effectiveness of this method and this system on the multi-modal physiological data of 8 subjects collected. The average accuracy of emotion recognition reaches 87.4%, meeting the requirements of the practicality of the emotion recognition interaction system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 FIG. is a flowchart of a method for emotion recognition based on multi-modal cross-attention feature fusion of EEG and eye movement according to an embodiment of the present invention.
[0023] Figure 2 FIG. is a structural diagram of a neural network for emotion recognition based on multi-modal cross-attention feature fusion of EEG and eye movement according to an embodiment of the present invention.
[0024] Figure 3 FIG. is a schematic diagram of the composition of an emotion recognition interaction system based on multi-modal cross-attention feature fusion of EEG and eye movement according to an embodiment of the present invention.
[0025] Figures 4A - 4C FIG. is an interface diagram of an emotion recognition system based on multi-modal cross-attention feature fusion of EEG and eye movement. DETAILED DESCRIPTION OF THE INVENTION
[0026] The overall process of the method for emotion recognition based on multi-modal cross-attention feature fusion of EEG and eye movement according to an embodiment of the present invention is as Figure 1 shown, and it includes:
[0027] Step S1: Obtain the publicly available emotion recognition dataset SEED;
[0028] Step S2: Preprocess the obtained public dataset and establish a training set and a test set for the public data;
[0029] This step specifically includes:
[0030] The specific steps for data preprocessing of the public dataset include:
[0031] Step S2.1: Data screening. For EEG signals, select the data of all channels for use; for eye movement signals, select four types of indicators, namely eye movement trajectory, eye movement pupil diameter, eye movement fixation point, and eye movement saccade, to capture the eye movement changes of the subject;
[0032] Step S2.2: Band-pass filtering. For both the public emotion dataset SEED and the data collected by the system, select an appropriate band-pass filter for processing. For EEG signals, first perform 256Hz resampling, and then use a 1 - 50Hz band-pass filter to adapt to the frequency band range where emotion information exists in the EEG signal data. For eye movement signals, select a corresponding band-pass filter to perform noise reduction processing according to the frequency characteristics of eye movement;
[0033] Step S2.3: Baseline calibration. In order to avoid the baseline drift phenomenon existing in long-term EEG acquisition, it is necessary to perform baseline correction on the EEG, set the same reference starting point, such as selecting the data mean value in a period of time before the task event occurs as the baseline, and then calculate the relative value relative to the baseline after the task event occurs as the new EEG data value;
[0034] Step S2.4: Re-reference. Select an appropriate reference electrode according to the need to re-reference the EEG signal. In the SEED dataset, select the central electrode on the top of the head (Cz) as the reference electrode to observe the changes of all brain electrodes during the task. The re-reference process can remove the reference electrode information in the signal to reduce the possible interference caused by the reference. A similar method can also be used for eye movement signals;
[0035] Step S2.5: Reject bad channels. Reject the bad segments of EEG data exceeding plus or minus 100 microvolts. These bad segments may be caused by noise or bad electrodes, such as sensor failures, poor electrode contacts, etc.;
[0036] Step S3: Use the Pytorch deep learning framework to construct the structure of an emotion recognition model based on cross-attention feature fusion of EEG and eye movement multimodality. The structure is as Figure 2 shown. This model includes: a dynamic graph EEG signal feature extraction module, a multi-dimensional eye movement signal feature extraction module, a multimodal cross-attention feature fusion module, and an emotion state recognition and classification module.
[0037] Specifically, the dynamic graph electroencephalogram (EEG) signal feature extraction module and the multi-dimensional eye movement signal feature extraction module are used to extract high-level features of the signals. Next, the EEG and eye movement signals, as high-level features of the two modalities, are input into the multi-modal cross-attention feature fusion module. Finally, the fusion result is input into the emotional state classification block to output the result;
[0038] Specifically, the dynamic graph EEG signal feature extraction module is used to extract the depth spatio-temporal domain features of the EEG signal, combining the EEG information with the channel information; the multi-dimensional eye movement signal feature extraction module is used to extract the depth time domain features of the eye movement signal, capturing the real-time changes in eye movement features; the multi-modal cross-attention feature fusion module is used to fuse the EEG signal features and the eye movement signal features; the emotional state recognition and classification module is used to output the emotional recognition result, and the emotional recognition result includes discrete emotional states such as positive, negative, and neutral.
[0039] The modeling method includes:
[0040] (a) Construct a dynamic graph EEG signal feature extraction module, which is used to extract the high-level features in the time domain and space domain of the EEG data, including:
[0041] At the time domain level, first, multiple (preferably three) convolutional layer units with different convolutional kernel sizes are used to initially extract the time domain features from the EEG signal.
[0042] Then, by calculating the time-frequency graph of the EEG signal to generate significant features in different frequency bands, the different frequency bands include: delta band [0 - 4Hz], theta band [4 - 8Hz], alpha band [8 - 12Hz], beta band [12 - 30Hz], and gamma band [30 - 50Hz]. The primary EEG data with a feature dimension of (batchsize, C, 5) (C is the number of EEG channels) is obtained, and batchsize is taken as 32.
[0043] Subsequently, depthwise separable convolution is used to extract the high-level EEG features in sub-bands. Compared with the conventional convolution method, depthwise separable convolution first uses depthwise convolutional kernels to group and analyze the emotional recognition features of the five frequency bands, and then uses pointwise convolutional kernels to fuse the global full-band EEG features, which is beneficial to capturing the most discriminative emotional features, reducing the number of model parameters, and preventing overfitting. The above process obtains the high-level EEG feature F g , where ∈ represents belonging to, R represents the set of real numbers, and C×D g represents the feature dimension of this feature, where C is the number of EEG channels, and D g is the number of frequency bands of the high-level EEG feature, and here the value is 5.
[0044] Again, a dynamic adjacency matrix of EEG channel weights is used to learn the connections between different EEG channels, including:
[0045] First, randomly initialize the adjacency matrix \(A\in\mathbb{R}\) C×C , where the \((i, j)\)-th element represents the coupling strength between the \(i\)-th and \(j\)-th EEG channels, including the direction and strength between the channels.
[0046] Then, use the ELU and Tanh non-linear activation functions to further sparsify the adjacency matrix, denoted by \(\delta(\cdot)\) and \(\sigma(\cdot)\) respectively, and at the same time perform a fully connected transformation using the weight matrices \(W_1\) and \(W_2\) to obtain the dense adjacency matrix \(A\) T in vectorized form Thereby reducing the problems of feature smoothing and noise information redundancy during the graph convolution process. The calculation formula is as follows:
[0047]
[0048] where
[0049] is obtained by vectorizing \(A\), \(A\) T is the dense adjacency matrix, is the vectorized form of the dense connection matrix \(A\) T ,
[0050] and represent the weight matrices of the fully connected layers,
[0051] \(\delta(\cdot)\) and \(\sigma(\cdot)\) represent the ELU and Tanh activation functions respectively,
[0052] \(r\) represents the decay rate hyperparameter.
[0053] After that, by transforming into \(\mathbb{R}\) C×C , the dense adjacency matrix \(A\) T \(\in\mathbb{R}\) C×C is obtained, where the \((i, j)\)-th element represents the dynamically updatable parameter, reflecting the directional correlation between the \(i\)-th and \(j\)-th EEG channels.
[0054] To dynamically incorporate the weight information of EEG channels into EEG features, two steps are taken. First, combine the optimized adjacency matrix \(A\) T with the high-level EEG features to form the graph feature \(G\) containing the associated node features g , where is the vectorized form of \(F\) g , \(v\) represents the vertex set \(|v| = C\), from the feature The connections between nodes are determined by the adjacency matrix \(A\) T . The calculation formula is as follows:
[0055]
[0056] Among them,
[0057] represents the high-level features obtained from the time-domain branch and the frequency-domain branch,
[0058] D S represents the degree matrix of A T Each element in it can be specifically expressed as
[0059] and represents the weight matrix,
[0060] D g ′ represents an adjustable hyperparameter,
[0061] δ1(·) and δ2(·) represent the ELU non-linear activation function.
[0062] Finally, the high-level EEG signal features G g fused with channel weight information are used as the input of the next dynamic channel weight attention module, and the calculation formula is as follows:
[0063] H g = MHSA g (LN(G g )) + G g (3),
[0064] where LN(·) represents layer normalization, and MHSA g represents the graph self-attention mechanism and the fully connected layer. The specific calculation formula is:
[0065]
[0066] where Q g , K g and V g represent the corresponding query matrix, key matrix, and value matrix respectively, and K g T represents the transpose matrix of K g , and D Kg represents the set hyperparameter. Thus, the spatio-temporal depth features H g of the EEG signal are obtained as the output of this module, and its feature dimension is (batchsize, C, 5);
[0067] (b) Construct a multi-dimensional eye movement signal feature extraction block,
[0068] Select multi-dimensional eye movement signal characterization metrics: eye movement trajectory, eye movement pupil diameter, eye movement fixation points, and eye movement saccades to extract the key features of the eye movement signal. Specifically:
[0069] The eye movement trajectory records the movement trajectory of the subject's eyes during viewing. These trajectories are represented in the form of two-dimensional coordinates and can provide information about the subject's fixation path and visual saccade behavior.
[0070] The eye movement pupil diameter reflects the change in the pupil size of the subject during viewing. This change can be used to evaluate the subject's cognitive load and emotional state.
[0071] The eye movement fixation points record the specific areas and time periods where the subject stays on the screen. These fixation points are usually related to the target of interest or a specific task and can provide information about the subject's visual attention and attention allocation.
[0072] Eye movement saccades refer to the time and position of the subject's rapid eye movement between different targets. This saccade behavior can provide information about the subject's visual search and attention transfer.
[0073] According to an embodiment of the present invention, all eye movement signal features are used for feature extraction, that is, multi-dimensional eye movement signal features. Regarding the problem that the number of dimensions of the multi-dimensional eye movement signal features is small and overfitting is likely to occur during the deep learning process:
[0074] First, use an encoder to expand the dimension of the initial eye movement signal features obtained from each segment where N y is the dimension of the eye movement feature,
[0075] Then, use a convolutional network with a residual structure to preliminarily process the encoded and dimension-expanded eye movement features to obtain the initial eye movement features where D y is the number of expanded dimensions. Here, the value is taken as 5, which makes the dimension-expanded eye movement data more closely related to the original eye movement data and contains more sufficient information.
[0076] Next, for the same type of eye movement feature, use the multi-head self-attention mechanism to capture the internal connection of the eye movement feature, and at the same time use multi-layer convolution to deepen the feature. The two are parallel, and high-level feature extraction is achieved on a single feature scale. The calculation formula is as follows:
[0077] F ydd = MHSA y (LN(F yd )) + Conv(F yd ) (5),
[0078] where LN(·) represents layer normalization, MHSAy (·) mainly represents the self-attention mechanism and the fully connected layer. Conv(·) represents the multi-layer convolutional layer, and F ydd represents the advanced eye movement features obtained after processing.
[0079] Finally, the initial eye movement feature F yd and F ydd are subjected to residual stacking on the feature scale to obtain the depth features of the eye movement signal while preventing overfitting and reducing the loss of data content of the eye movement information, where H y belongs to the real number set R, N y ×D y represents the feature dimension, N y is the dimension of the eye movement feature, and D y is the number of expanded dimensions, and the value is taken as 5 here;
[0080] (c) Construct a multi-modal attention feature fusion module:
[0081] The depth features H g of the electroencephalogram signal and the depth features H y of the eye movement signal are obtained from the above modules. These two parts are used as the input of the adaptive multi-level attention feature fusion module to fuse the features of the two different modalities and realize the final extraction and classification of the emotion features.
[0082] This operation / module is divided into two parts, including low-level attention feature fusion and high-level attention feature fusion. Research has shown that low-level feature fusion can better ensure the integrity of the feature content, while high-level feature fusion can better ensure the feature pertinence of the target task. In an embodiment according to the present invention, the calculation formula of the low-level attention feature fusion part is:
[0083] R L = MHSA L (LN(H g , H y )) (6),
[0084] The calculation formula of the high-level attention feature fusion part is:
[0085] R H = MHSA H (LN(H g , H y )) (7),
[0086] where:
[0087] MHSA L and MHSA H respectively represent the low-level and high-level attention mechanisms,
[0088] LN(·) represents layer normalization;
[0089] MHSA L The specific calculation steps are as follows:
[0090] H1 = Concat(H g , H y ) (8),
[0091]
[0092] where Concat(·) represents matrix concatenation, softmax(·) represents the normalization function, and among the three vectors Q L , K L and V L represent the query matrix, key matrix, and value matrix respectively, and they each correspond to three weight matrices. H1 is multiplied by these three weight matrices for linear transformation and is mapped to the corresponding Q L , K L and V L ,, D KL represents the set hyperparameter.
[0093] MHSA H The specific calculation steps are as follows:
[0094]
[0095] R H = Concat(R H1 , R H2 ) (12),
[0096] The three vectors Q H1 , K H1 and V H1 represent the query matrix, key matrix, and value matrix respectively, and they each correspond to three weight matrices. H g is multiplied by these three weight matrices for linear transformation and is mapped to the corresponding Q H1 , K H1 and V H1 ; The three vectors Q H2 , K H2 and V H2 represent the query matrix, key matrix, and value matrix respectively, and they each correspond to three weight matrices. H y is multiplied by these three weight matrices for linear transformation and is mapped to the corresponding Q H2 , K H2 and V H2 , D KH1 and D KH represent hyperparameters.
[0097] (d) Construct an emotional state recognition and classification module for outputting multi-class emotion recognition results:
[0098] For the two fusion features R obtained from the above process L and R H Flatten them into one-dimensional vectors respectively and input them into two fully connected layers. Then, use the SoftMax function to calculate the classification probability of the output vector, and perform decision fusion on the two sets of probabilities to obtain the final classification probability, where the maximum value is defined as the classification result;
[0099] Step S4: Input the preprocessed training set in Step S2 into the emotion recognition model for training, update the model parameters and save the model, specifically:
[0100] First, train by minimizing the cross-entropy loss J between the model prediction and the label, and the calculation formula is as follows:
[0101]
[0102] where p i is the i-th conditional probability generated by the emotion recognition model, l i is the i-th category of the label set, ω(·) is the indicator function, Θ represents the learnable parameters in the intent recognition model, ‖·‖ is the regularization term for alleviating the overfitting problem, λ is the regularization weight trade-off, which can be taken as 0.01, M represents the size of the Batchsize, and set M to 32.
[0103] During the process of training the model, set the batchsize size to 32, set the model learning rate to 0.0001, use the Adam optimizer to optimize the model parameters, stop training after 30 iterations of training or when the accuracy starts to decline, and save the model parameters.
[0104] To verify the effectiveness of the proposed method, performance tests are carried out on the public dataset SEED.
[0105] The classification results can be divided into four categories: True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN). The effectiveness of the proposed method is evaluated by accuracy (Acc), AUC, and F1-Score, where the definitions of Acc and F1 metrics are as follows:
[0106]
[0107] To compare the performance with the proposed method, in the same experimental environment, three baseline models are used, and these models are retested on the SEED dataset. A brief introduction to the three baseline models is as follows:
[0108] 1) DCCA (W. Liu et al., Multimodal Emotion Recognition Using Deep Canonical Correlation Analysis, 2018): DCCA calculates the maximum correlation between vectors, solves the linearized feature vectors of different modalities through a deep network, and inputs the extracted deep features into a support vector machine for classification. DCCA-AM combines the attention mechanism with DCCA, which is more conducive to multimodal-related classification and improves the classification efficiency.
[0109] 2) DGCNN (T. Song et al., EEG emotion recognition using dynamical graph convolutional neural networks, 2018): This method uses a dynamic adjacency matrix to simulate the channel relationship of shallow DE features, which has good results in EEG feature extraction and performs well in emotion recognition tasks.
[0110] 3) ETF (Y. Wang et al., Emotion transformer fusion: Complementary representation properties of EEG and eye movements on recognizing anger and surprise, 2021): ETF inputs data of different modalities into corresponding Transformer encoders, and then fuses them into a joint representation space through a fusion layer based on the attention mechanism. It is the basis for some multi-feature fusion methods.
[0111] Table 1 Performance comparison of emotion recognition methods based on EEG and eye movement multimodal cross-attention feature fusion
[0112]
[0113] As can be seen from Table 1, the average recognition accuracy of the emotion recognition method based on EEG and eye movement multimodal cross-attention feature fusion proposed in the present invention reaches 97.94% on SEED, which is 21.95%, 8.28% and 2.37% higher than the baseline methods respectively. The significantly improved accuracy of the method of the present invention indicates that the method of the present invention can better fuse multi-source information features, and at the same time indicates its robustness to high-level feature extraction, can provide better classification performance, and meets the requirements of human-computer emotion interaction for recognition accuracy.
[0114] Step S5: Build an emotion recognition interaction system, use the system to collect the electroencephalogram (EEG) signals and eye movement signals of 8 subjects, and verify the effectiveness of the method by recording the recognition accuracy and interaction duration during the brain-computer interaction process, as follows:
[0115] When using a scalp electroencephalogram acquisition device to collect EEG signals, a Smarting wireless scalp electroencephalogram acquisition device is adopted. The electrode positions adopt the international standard 10-20 electrode lead positioning, the reference electrode is set in the central area of the head, the amplifier sampling frequency is 1000 Hz, and the acquisition channels are EEG signals of 24 leads, including: Fp1, Fp2, AFZ, F7, F3, Fz, F4, F8, T7, C3, Cz, C4, T8, CPZ, M1, M2, P7, P3, Pz, P4, P8, PoZ, O1, O2, and the collected EEG signals are uploaded in real time through a data cable; in the part of eye movement signal acquisition, a Tobii Pro Glasses 3 eye tracker is selected to collect four types of data: the eye movement trajectory, eye movement pupil diameter, eye movement fixation point, and eye movement saccade of the subject, and the collected eye movement signals are uploaded in real time through a wifi module;
[0116] In the actual operation process, first construct an emotional stimulation scenario, and then use a scalp electroencephalogram acquisition device to collect EEG signals and an eye movement acquisition device to collect eye movement signals. When starting to construct the emotional stimulation scenario, the prompt "Experiment starts" appears on the screen to let the subject concentrate. The subject generates corresponding emotions according to the video clips played, specifically: there is a 20-second rest time before each trial, a 3-second video category prompt time, the duration of each video clip is about 3-4 minutes, and there is a 30-second self-evaluation time after each clip to evaluate the stimulation effect of this clip on the subject. The presentation order of the video clips is arranged randomly, but two movie clips targeting the same emotion will not be displayed continuously;
[0117] When verifying the effectiveness of the method, conduct an emotion recognition experiment on 8 subjects. The verification process is mainly divided into two parts: The first part is to collect all multimodal physiological data of 8 subjects through an EEG and eye movement multimodal online emotion recognition system. Each subject has 18 trials. The collected data is saved and used for model training to obtain an emotion recognition model for each subject respectively; The second part first deploys the emotion recognition model obtained in the above process to the interaction system, and then conducts a second collection on 8 subjects. The collected multimodal physiological signals are directly input into the emotion interaction system, and the recognition accuracy and interaction duration of the system during the emotion recognition process are recorded to verify the effectiveness of the method. The specific experimental results are shown in Table 2. It can also be seen from Table 2 that this system
[0118] Table 2 Subject information and emotion recognition accuracy
[0119]
[0120] The system can obtain a relatively high recognition accuracy rate (average recognition accuracy rate > 87%) with a relatively fast response speed (average < 1 s), realizing the rapid response of the emotion recognition system and meeting the application conditions.
[0121] According to another aspect of the present invention, there is provided a multi-modal cross-attention feature fusion emotion recognition interaction system based on electroencephalogram and electrooculogram. The schematic diagram of its constituent modules is as Figure 3 shown and specifically includes:
[0122] Control module: This module is used to record the information of the test subject and the selection of video clips, and control the start and stop of the test process;
[0123] Multi-modal data acquisition module and visualization display module: Control the electroencephalogram acquisition device and the eye tracker to complete the acquisition of multi-modal data signals, and send the acquired data into the mental state training model in real time. The acquired data is visually updated and displayed in real time on the UI interface. The electroencephalogram signals are displayed with electroencephalogram waveforms and topographic maps, and the eye movement coordinate information is output in tabular form to display the eye movement signals;
[0124] Among them, the devices used for collecting online electroencephalogram data include: Smarting wireless electroencephalogram electrode cap, with the electrode positions located by the international standard 10-20 electrode lead system, and the reference electrode is set in the central area of the head; electroencephalogram amplifier, with the sampling frequency set to 1000 Hz, and the acquisition channels are 24-lead electroencephalogram signals, including: Fp1, Fp2, AFZ, F7, F3, Fz, F4, F8, T7, C3, Cz, C4, T8, CPZ, M1, M2, P7, P3, Pz, P4, P8, PoZ, O1, O2, and the acquired electroencephalogram signals are uploaded in real time through the data line to the built-in lsl module. The devices used for collecting online eye movement data include: Tobii Pro Glasses 3 eye tracker data recording unit and head-mounted acquisition unit. The data recording unit acts like a small computer and is used to control the head-mounted acquisition unit. It can record and store the eye movement data, sound, and video of the scene camera on a removable SD card; the head-mounted acquisition unit is the core component of this eye tracker. It is a highly complex measuring device composed of multiple extremely sensitive sensors. These include infrared sensors, high-resolution cameras, and eye movement tracking cameras, etc. The coordinated work of these sensors makes the acquisition of eye movement data possible; these two units work together to provide a powerful eye movement data acquisition function.
[0125] Multi-modal data preprocessing module: For the collected multi-modal signals, data preprocessing is performed separately, including filtering and artifact removal of EEG signals; removal of invalid data and resampling of eye movement signals; both types of signals need to maintain a data window size of 1 s for further identification;
[0126] Detection model loading module: Embed the feature extraction and fusion decoding network of EEG and eye movement signals, realize the feature extraction, analysis and state recognition of multi-modal signals, and import the model library into the system software platform to perform model calculations on the online collected data and output the recognition results;
[0127] Recognition result display module: Visualize the interface and display the psychological state recognition results, which can realize the display of recognition duration, recognition status, recognition results and abnormal warnings, etc., and can compare the recognition results, that is, the training results of the model with the labels of the corresponding videos to output the stage accuracy;
[0128] Abnormal state warning module: The result output by the model is added with a threshold judgment to determine the degree of abnormal results. When the abnormal threshold result is greater than 0.1 and less than 0.4, the model judges that the user's emotional abnormality degree is "relatively abnormal". At this time, a pop-up window is displayed on the interface, and the small window shows "The current user state is relatively abnormal". The duration of the pop-up window is 2 s. You can click the "cancel" button or close the pop-up window in the upper right corner. At the same time, the system emits an alarm sound with a sound frequency of 300 Hz and a sound duration of 1 s; when the abnormal threshold result is greater than 0.4, the model judges that the user's emotional abnormality degree is "very abnormal". At this time, a pop-up window is displayed on the interface, and the small window shows "The current user state is very abnormal". The duration of the pop-up window is 2 s. You can click the "cancel" button or close the pop-up window in the upper right corner. At the same time, the system emits an alarm sound with a sound frequency of 500 Hz and a sound duration of 1 s.
[0129] The overall system interface diagram and the pop-up window interface diagram of the abnormal state warning module are shown in Figures 4A-4C.
[0130] The above has described in detail the brain-computer interaction intention recognition method and system based on the lightweight gradient boosting decision tree provided by the present invention. However, it is obvious that the scope of the present invention is not limited thereto. Without departing from the protection scope defined by the appended claims, various changes to the above embodiments are within the scope of the present invention.
Claims
1. A modeling method for an emotion recognition model based on cross-attention features fusion of electroencephalogram and electrooculogram multimodality for emotion recognition based on cross-attention features fusion of electroencephalogram and electrooculogram multimodality, characterized in that including; A) Construct a dynamic graph electroencephalogram (EEG) signal feature extraction module for extracting high-level features in the time domain and spatial domain of EEG data, and combining EEG information with channel information; B) Construct a multi-dimensional eye movement signal feature extraction module for extracting the deep time-domain features of eye movement signals and capturing real-time changes in eye movement features; C) Construct a multi-modal cross-attention feature fusion module for fusing EEG signal features and eye movement signal features; D) Construct an emotion state recognition and classification module for outputting emotion recognition results, where the emotion recognition results include discrete emotion states such as positive, negative, and neutral. where: The step A) includes: A1) At the time domain level, first use multiple convolutional layer units with different convolutional kernel sizes to preliminarily extract time-domain features from EEG signals; A2) Then, generate significant features in different frequency bands by calculating the time-frequency graph of EEG signals. The different frequency bands include: delta band [0 - 4Hz], theta band [4 - 8Hz], alpha band [8 - 12Hz], beta band [12 - 30Hz], and gamma band [30 - 50Hz], obtaining primary EEG data with a feature dimension of (batchsize, C, 5), where batchsize is taken as 32 and C is the number of EEG channels. A3) Subsequently, depthwise separable convolution is used to divide the frequency bands into groups, and high-level EEG features are deeply extracted, including: First, the emotion recognition features of the five frequency bands are analyzed using depthwise convolutional kernels grouped by depth, and then pointwise convolutional kernels are used to fuse the global full-band EEG features, which is beneficial to capturing the most discriminative emotion features, while reducing the number of model parameters and preventing overfitting, thereby obtaining high-level EEG features where ∈ represents belonging to, R represents the set of real numbers, X×D g represents the feature dimension of the feature, where is the number of EEG channels, D g is the number of frequency bands of the high-level EEG features A4) Again, use a dynamic adjacency matrix of EEG channel weights to learn the connections between different EEG channels, including: A41) First, randomly initialize the adjacency matrix \(A\in\mathbb{R}\) C×C , where the \((i, j)\)-th element represents the coupling strength between the \(i\)-th and \(j\)-th EEG channels, including the direction and strength between channels. A42) Then, the ELU and Tanh non-linear activation functions are used to further sparsely encode the adjacency matrix, denoted by δ(·) and σ(·) respectively, and at the same time, the weight matrices W1 and W2 are used for fully connected transformation to obtain the dense adjacency matrix A T Vectorized representation of Thus, the problems of feature smoothing and noise information redundancy in the graph convolution process are alleviated, and the formula is as follows: where, Obtained by vectorizing A, where A T is a dense adjacency matrix, is the vectorized representation of the dense adjacency matrix A T and is the vectorized representation of and represent the weight matrix of the fully connected layer, δ(·) and σ(·) represent the ELU and Tanh activation functions respectively, r represents the decay rate hyperparameter, After A43), by transforming into R C×C , a dense adjacency matrix A T ∈R C×C is obtained, where the (i, j)-th element represents a dynamically updatable parameter, reflecting the directional correlation between the i-th and j-th EEG channels. Among them, in order to dynamically incorporate the weight information of EEG channels into EEG features, two steps are taken. First, the optimized adjacency matrix A T is combined with the high-level EEG features to form the graph feature G that contains the associated node features g , where is the vectorized representation of F g , υ represents the vertex set |υ| = C, and the connections between nodes from the feature are determined by the adjacency matrix A T , and the calculation formula is as follows: where, denotes the high-level features obtained from the time-domain branch and the frequency-domain branch, D S represents A T 's degree matrix, and each element therein can be specifically expressed as and represent the weight matrix, D g ′ represents an adjustable hyperparameter, δ1(·), δ2(·) represent the ELU non-linear activation function, A5) Finally, the advanced EEG signal feature G that incorporates channel weight information g is used as the input to the next part, the dynamic channel weight attention module. The calculation formula is as follows: H g = MHSA g (LN(G g )) + G g (3) where: H g is the depth feature of the electroencephalogram signal, LN(·) represents layer normalization, MHSA g Represents the graph self-attention mechanism and the fully connected layer The specific calculation formula is: where: Q g , K g and V g represent the corresponding query matrix, key matrix, and value matrix, respectively. K g T represents K g is the transposed matrix of D Kg Indicates the set hyperparameters Thus, the spatio-temporal depth feature H of the EEG signal is obtained. g As the output of this module, its feature dimension is (batchsize, C, 5); The step B) includes: B1) Select multi-dimensional eye movement signal characterization metrics, including: eye movement trajectory, eye movement pupil diameter, eye movement fixation point, and eye movement saccade, to extract key features of eye movement signals, where: The eye movement trajectory records the movement trajectory of the subject's eyes during viewing, and these trajectories are represented in the form of two-dimensional coordinates, providing information about the subject's fixation path and visual saccade behavior; The eye movement pupil diameter reflects the change in the pupil size of the subject during viewing, and this change can be used to evaluate the subject's cognitive load and emotional state; The eye movement fixation point records the specific areas and time periods where the subject stays on the screen, and these fixation points are usually related to the target of interest or a specific task, and can provide information about the subject's visual attention and attention allocation; The eye movement saccade refers to the time and position of the subject's rapid eye movement between different targets, and this saccade behavior can provide information about the subject's visual search and attention transfer. B2) Next, for the same eye movement feature, use the multi-head self-attention mechanism to capture the internal connection of eye movement features, and at the same time use multiple convolutional layers to deepen the features, achieving high-level feature extraction on a single feature scale. The calculation formula is as follows: F ydd = MHSA y (LN(F yd )) + Conv(F yd ) (5) where: LN(·) represents layer normalization, MHSA y (·) mainly represents the self-attention mechanism and the fully connected layer, Conv(·) represents multiple convolutional layers, F ydd represents the advanced eye movement features obtained after processing B3) Finally, the initial eye movement feature F yd and F ydd are subjected to residual superposition on the feature scale to obtain the depth feature of the eye movement signal so as to reduce the loss of data content of eye movement information while preventing overfitting, where: H y belongs to the set of real numbers R N y ×D y represents the feature dimension N y is the dimension of the eye movement feature D y For the number of expanded dimensions, the value is taken as 5; Step C) includes taking the depth feature H of the electroencephalogram signal g and the depth feature H of the eye movement signal y , fusing these two parts of features to achieve the final extraction and classification of emotional features The step C) includes low-level attention feature fusion and high-level attention feature fusion, where: The calculation formula for low-level attention feature fusion is as follows: R L = MHSA L (LN(H g ,H y )) (6) The calculation formula for high-level attention feature fusion part is as follows: R H = MHSA H (LN(H g ,H y )) (7), Where: MHSA L and MHSA H represent the low - level and high - level attention mechanisms respectively, LN(·) represents layer normalization; MHSA L The specific calculation steps include: H1 = Concat(H g , H y ) (8), Where: Concat(·) represents matrix concatenation, softmax(·) represents the normalization function, Three vectors Q L , K L and V L represent the query matrix, the key matrix, and the value matrix respectively, each of which corresponds to three weight matrices. H1 is linearly transformed by multiplying with these three weight matrices and is mapped to the corresponding Q L , K L and V L , D KL represents the set hyperparameters MHSA H The specific calculation steps include: R H = Concat(R H1 , R H2 ) (13), Where: Three vectors Q H1 , K H1 and V H1 represent the query matrix, the key matrix, and the value matrix respectively, each corresponding to three weight matrices. H g is multiplied by these three weight matrices for linear transformation and is mapped to the corresponding Q H1 , K H1 and V H1 , Three vectors Q H2 , K H2 and V H2 represent the query matrix, the key matrix, and the value matrix respectively, each corresponding to three weight matrices, H y is multiplied by these three weight matrices for linear transformation and is mapped to the corresponding Q H2 , K H2 and V H2 , D KH1 and D KH2 represent hyperparameters, Step D) includes: Flatten the two fusion features R L and R H into one-dimensional vectors respectively, and then input them into two fully connected layers respectively. Then, use the SoftMax function to calculate the classification probability for the output vector, and perform decision fusion on the two sets of probabilities to obtain the final classification probability, where the maximum value is defined as the classification result.
2. The modeling method of the fusion emotion recognition model according to claim 1, characterized in that; Use the Pytorch deep learning framework for modeling.
3. The modeling method of the fusion emotion recognition model according to claim 1, characterized in that; D g takes a value of 5.
4. The modeling method of the fusion emotion recognition model according to claim 1, characterized in that Step B1 includes using all eye movement signal features for feature extraction, that is, multi-dimensional eye movement signal features, including: First, the encoder is used to expand the dimension of the initial eye movement signal feature F obtained from each segment y ∈R Ny where N y is the dimension of the eye movement feature Then, a convolutional network with a residual structure is used to preliminarily process the eye movement features after encoding and dimension expansion to obtain the initial eye movement feature F yd ∈R Ny×Dy , where D y is the number of expanded dimensions. Here, the value is taken as 5, so that the eye movement data after dimension expansion is more closely related to the original eye movement data and contains more sufficient information.
5. A fusion emotion recognition method based on cross-attention features of electroencephalogram and electrooculogram multimodality, characterized in that Including: Step S1: Obtain multi-modal physiological data in the publicly available emotion recognition dataset SEED, including electroencephalogram (EEG) data and eye movement data; Step S2: Preprocess the obtained publicly available dataset and establish a training set and a test set; Step S3: Execute the modeling method of the fusion emotion recognition model according to one of claims 1-4; Step S4: Input the preprocessed multi-modal data training set in Step S2 into the emotion recognition model for model training; Then input the preprocessed test set in Step S2 into the trained model for performance testing of the method, Where: In the said Step S2, data preprocessing includes data screening, band-pass filtering, baseline correction, re-referencing, and bad channel rejection; In the said Step S3, the emotion recognition model based on cross-attention feature fusion of EEG and eye movement multi-modalities includes: a dynamic graph EEG signal feature extraction module, a multi-dimensional eye movement signal feature extraction module, a multi-modal cross-attention feature fusion module, and an emotion state recognition and classification module, where: The dynamic graph EEG signal feature extraction module is used to extract high-level EEG features, and the multi-dimensional eye movement signal feature extraction module is used to extract high-level eye movement features. Next, the high-level features of the two modalities are input into the multi-modal cross-attention feature fusion module, and finally the fusion result is input into the emotion state classification block to output the result. Specifically; The EEG signal feature extraction module is used to extract the deep spatio-temporal domain features of EEG signals; The eye movement signal feature extraction module is used to extract the deep time domain features of eye movement signals; The multi-modal attention feature fusion module is used to fuse EEG signal features and eye movement signal features; the emotion state recognition and classification module is used to output emotion recognition results, and the emotion recognition results include discrete emotion states such as positive, negative, and neutral; In the said Step S4: Adopt the method of five-fold cross-validation to train the model using the training set, then save the trained parameters of the model, and use the test set to perform performance verification such as accuracy testing on the model.
6. The fusion emotion recognition method according to claim 5, wherein Further includes: Step S5: Build an online emotion recognition system for EEG and eye movement multi-modalities, use the system to collect emotion recognition EEG and eye movement signals of 8 subjects, and verify the effectiveness of the method by recording the recognition accuracy and interaction duration during the interaction process, Including: Build an online emotion recognition system for EEG and eye movement multi-modalities, and use the system to first collect scalp layer EEG data and eye movement data of the multiple subjects; Secondly, preprocess the data according to the method of step S2, and establish a training set and a test set for the collected data; Then, input the preprocessed training set into the emotion recognition model constructed in step S3, and train the emotion recognition model according to the method described in step S4; Then, use the system to collect the scalp electroencephalogram data and eye movement signals of the multiple subjects again, and verify the effectiveness of the method through the system recognition accuracy rate in the processes of real-time recording of physiological signal data, data preprocessing, emotion recognition, and interactive feedback.
7. The fused emotion recognition method according to claim 5 or 6, wherein: The preprocessing in step S2 includes: Step S2.1: Data screening. For electroencephalogram signals, select the data of all channels for use; for eye movement signals, select four types of indicators, namely, eye movement trajectory, eye movement pupil diameter, eye movement fixation point, and eye movement saccade, to capture the eye movement changes of the subject; Step S2.2: Band-pass filtering. Select appropriate band-pass filters to process the data collected from the public emotion dataset SEED and the electroencephalogram-eye movement multimodal online emotion recognition system. For electroencephalogram signals, first perform 256Hz resampling, and then use a 1-50Hz band-pass filter to adapt to the frequency band range where emotion information exists in the electroencephalogram signal data. For eye movement signals, select the corresponding band-pass filter to perform noise reduction processing on the signals according to the frequency characteristics of eye movement; Step S2.3: Baseline calibration. In order to avoid the baseline drift phenomenon existing in the long-term electroencephalogram acquisition, it is necessary to perform baseline correction on the electroencephalogram. Set the same reference starting point, such as selecting the data mean value in a period of time before the task event occurs as the baseline, and then calculate the relative value relative to the baseline after the task event occurs as the new electroencephalogram data value; Step S2.4: Re-reference. Select appropriate reference electrodes to re-reference the electroencephalogram signals according to needs. In the SEED dataset, select the central electrode Cz on the top of the head as the reference electrode to observe the changes of all-brain electrodes when the task occurs. The re-reference process can remove the reference electrode information in the signal to reduce the possible interference caused by the reference. A similar method can also be used for eye movement signals; Step S2.5: Reject bad channels. Reject the bad segments of electroencephalogram data exceeding plus or minus 100 microvolts. These bad segments may be caused by noise or bad electrodes.
8. The fused emotion recognition method according to claim 5 or 6, wherein: The label is the label of each segment. Step S4 includes inputting the training set preprocessed in step S2 into the constructed electroencephalogram-eye movement multimodal cross-attention feature fusion emotion recognition network model respectively, updating the model parameters, and generating the final training model, including: First, train by minimizing the cross-entropy loss J between the model prediction and the label. The calculation formula is as follows: where p i is the i-th conditional probability generated by the emotion recognition model, l i is the i-th category of the label set, ω(·) is the indicator function, Θ represents the learnable parameters in the intention recognition model, ‖·‖ is the regularization term used to alleviate the overfitting problem, λ is the regularization weight trade-off, M represents the size of the Batchsize, set M to 32, set the model learning rate to 0.0001, use the Adam optimizer to optimize the model parameters, stop training after 30 iterations of training or when the accuracy starts to decline, save the model parameters, Then, input the test set preprocessed in step S2 into the trained electroencephalogram-eye movement multimodal cross-attention feature fusion emotion recognition network model respectively, obtain the multimodal emotion recognition classification performance indicators, and evaluate the effectiveness of the proposed method through the accuracy rate Acc, AUC, and F1-Score.
9. The fused emotion recognition method according to claim 5 or 6, characterized in that: In the step S5, when using the electroencephalogram and electrooculogram multimodal online emotion recognition system to collect electroencephalogram and electrooculogram, first construct an emotional stimulation scenario, and then use the scalp electroencephalogram acquisition device to collect electroencephalogram signals and the electrooculogram acquisition device to collect electrooculogram signals, where: When constructing the emotional stimulation scenario starts, the screen shows "Experiment starts" to prompt the subject to concentrate. The subject generates corresponding emotions according to the video clips played. Specifically: there is a 20-second rest time before each trial, a 3-second video category prompt time, the duration of each film clip is about 3 - 4 minutes, and there is a 30-second self-evaluation time after each clip to evaluate the stimulation effect of this film on the subject. The arrangement order of the video clips is random, but two movie clips targeting the same emotion will not be displayed continuously; When using the scalp electroencephalogram acquisition device to collect electroencephalogram signals, use the Smarting wireless scalp electroencephalogram acquisition device. The electrode positions adopt the international standard 10 - 20 electrode lead positioning. The reference electrode is set in the central area of the head. The amplifier sampling frequency is 1000 Hz, and the acquisition channels are electroencephalogram signals of 24 leads, including: Fp1, Fp2, AFZ, F7, F3, Fz, F4, F8, T7, C3, Cz, C4, T8, CPZ, M1, M2, P7, P3, Pz, P4, P8, PoZ, O1, O2, and the collected electroencephalogram signals are uploaded in real time through the data cable; in the part of electrooculogram signal acquisition, select the Tobii Pro Glasses 3 eye tracker to collect four types of data of the subject's eye movement trajectory, electrooculogram pupil diameter, electrooculogram fixation point and electrooculogram saccade, and upload the collected electrooculogram signals in real time through the wifi module; When verifying the effectiveness of the method, conduct an emotion recognition experiment on 8 subjects. The verification process is mainly divided into two parts: the first part is to collect all multimodal physiological data of 8 subjects through the electroencephalogram and electrooculogram multimodal online emotion recognition system. Each subject has 18 trials. The collected data is saved and used for model training to obtain the emotion recognition model of each subject respectively; the second part first deploys the emotion recognition model obtained in the above process to the interactive system, and then conducts a second collection on 8 subjects. The collected multimodal physiological signals are directly input into the emotion interaction system, and the recognition accuracy and interaction duration of the system during the emotion recognition process are recorded to verify the effectiveness of the method.
10. A computer-readable storage medium storing a computer-executable program, and the computer-executable program can enable a processor to execute the method according to any one of claims 1 - 9.
Citation Information
Cited By
Multi-modal signal fusion emotion recognition method based on attention mechanism
CN120477781A
Continuous monitoring method and system for multi-modal physiological parameter fusion
CN120753668A
Emotion recognition method based on multi-mode collaborative optimization
CN120892820A
Motion imagination coupling system and method based on emotion prediction
CN121365240A
A motion-image coupling system and method based on emotion prediction
CN121365240B