Quick emotion assessment method and device under variable gravity condition

By collecting and preprocessing EEG and video data in microgravity environments, using single-modal and multi-modal emotion recognition models, the problem of insufficient accuracy and stability of emotion recognition in variable gravity environments is solved, and more efficient emotion recognition effect is achieved.

CN120323973APending Publication Date: 2025-07-18SCI RES TRAINING CENT FOR CHINESE ASTRONAUTS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104958.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing emotion recognition technology lacks generalization ability of facial expression recognition models under changing gravity environments. Physiological signals are affected by gravity changes and cannot be accurately analyzed. There is a lack of customized design for changing gravity environments, resulting in a decrease in recognition accuracy and insufficient stability.

Method used

EEG data and video data in microgravity environments are collected, and after data preprocessing is performed, single-modal and multimodal basic emotion recognition models are used for emotion recognition, including single-modal basic emotion recognition models and multimodal basic emotion recognition models, and are trained and recognized through classifiers such as convolutional neural networks.

Benefits of technology

It improves the accuracy and real-time emotion recognition in variable gravity environments, and enhances the stability and ease of use of the system in extreme environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120323973A_ABST
    Figure CN120323973A_ABST
Patent Text Reader

Abstract

The invention provides a quick emotion assessment method and device under a variable gravity condition. The method comprises the following steps: acquiring electroencephalogram data and video data under a microgravity environment; performing data preprocessing on the electroencephalogram data and the video data to obtain electroencephalogram features and video features; substituting the electroencephalogram features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result; wherein the basic emotion recognition model comprises a single-mode basic emotion recognition model and a multi-mode basic emotion recognition model; the single-mode basic emotion recognition model and the multi-mode basic emotion recognition model are obtained by training a classifier based on features extracted from the video data and the electroencephalogram data and corresponding emotion categories. According to the invention, the accuracy and real-time performance of emotion recognition in a variable gravity environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emotion recognition in variable gravity environments, and specifically to a method and device for rapid emotion assessment under variable gravity conditions. Background Art

[0002] Emotion recognition technology has made remarkable progress in the field of artificial intelligence in recent years, and is widely used in daily life, such as in fields like social media interaction, virtual reality, and mental health monitoring. However, when faced with special physical environments, especially variable gravity conditions, emotion recognition technology faces severe challenges. Traditional ground-based emotion recognition technologies mainly rely on changes in facial expressions, speech tones, and physiological signals. However, in spaceflight or other simulated microgravity environment conditions, these features may change significantly, such as restricted facial muscle activity, changes in sound propagation characteristics, and atypical physiological responses, which directly reduce the reliability and effectiveness of current emotion recognition systems.

[0003] Most existing emotion recognition technologies are based on research results in the normal gravity environment on the ground. For example, facial expression recognition and speech emotion analysis are carried out using video and audio signals, or physiological data such as heart rate and skin conductance are collected with the help of physiological sensors. However, when applied to variable gravity environments, due to expression variations, physiological index fluctuations caused by changes in the force state, and the influence of complex environmental factors, conventional emotion recognition methods often cannot accurately capture and interpret the emotional changes of astronauts.

[0004] The current emotion recognition technology performs relatively maturely in normal environments. However, under variable gravity conditions, such as in manned spaceflight missions, due to changes in the facial expression change rules, increased interference of physiological signals, and increased influence of environmental factors on emotion expression, there are significant deficiencies in the accuracy, real-time performance, and stability of existing emotion recognition methods and devices. The present invention aims at these problems, and focuses on solving the problems of how to accurately and quickly identify and evaluate the basic emotions and compound emotion states of individuals in variable gravity environments, and how to effectively capture and warn of abnormal emotions, while also taking into account the usability, safety, and scalability of the system in variable gravity environments.

[0005] Combined with the above background, the main technical defects of the existing technology are as follows:

[0006] In variable gravity environments, the generalization ability and robustness of facial expression recognition models are insufficient, which may lead to a decrease in recognition accuracy;

[0007] Physiological signals are affected by gravity changes, and the original models are difficult to correctly analyze and correlate physiological indicators and emotion states under different gravity conditions;

[0008] The current emotion recognition system lacks customized design for variable gravity environments and cannot effectively handle the noise interference and data distortion problems brought about by extreme environments;

[0009] The system lacks robustness in abnormal situations (such as power interruption, data transmission delay), which may lead to data loss or unstable evaluation results;

[0010] There is a lack of training data for the multi-modal fusion model under variable gravity conditions, and there is a lack of targeted model adjustment and optimization. Summary of the Invention

[0011] In order to solve the problem that the generalization ability and robustness of the facial expression recognition model are insufficient in the variable gravity environment in the prior art, which may lead to a decrease in recognition accuracy, the present invention proposes a method for rapid emotion assessment under variable gravity conditions, including:

[0012] Collect electroencephalogram data and video data in the microgravity environment;

[0013] Perform data preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features;

[0014] Substitute the electroencephalogram features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result;

[0015] Among them, the basic emotion recognition model includes a single-modal basic emotion recognition model and a multi-modal basic emotion recognition model;

[0016] The single-modal basic emotion recognition model and the multi-modal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and electroencephalogram data, and the corresponding emotion categories.

[0017] Optionally, the training of the single-modal basic emotion recognition model includes:

[0018] Obtain electroencephalogram data and video data, and the emotion categories corresponding to the electroencephalogram data and the video data;

[0019] Perform preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features;

[0020] Construct a sample set from the electroencephalogram features, video features, and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0021] Use the electroencephalogram features and video features in the training set as inputs, and the corresponding emotion categories as outputs to train the classifier to obtain a preliminarily trained single-modal basic emotion recognition model;

[0022] Input the EEG features and video features in the test set into the preliminarily trained single-modal basic emotion recognition model to obtain the emotion recognition results;

[0023] Determine whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained single-modal basic emotion recognition model is used as the trained single-modal basic emotion recognition model; otherwise, continue to train the preliminarily trained single-modal basic emotion recognition model based on the training set.

[0024] Optionally, the training of the multi-modal basic emotion recognition model includes:

[0025] Obtain EEG data and video data, as well as the corresponding emotion categories of the EEG data and the video data;

[0026] Preprocess the EEG data and video data to obtain EEG features and video features;

[0027] Fuse the EEG features and video features to obtain multi-modal features;

[0028] Construct a sample set from the multi-modal features and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0029] Use the multi-modal features in the training set as the input and the corresponding emotion categories as the output to train the classifier to obtain a preliminarily trained multi-modal basic emotion recognition model;

[0030] Input the multi-modal features in the test set into the preliminarily trained multi-modal basic emotion recognition model to obtain the emotion recognition results;

[0031] Determine whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained multi-modal basic emotion recognition model is used as the trained multi-modal basic emotion recognition model; otherwise, continue to train the preliminarily trained multi-modal basic emotion recognition model based on the training set.

[0032] Optionally, the preprocessing of the EEG data and video data to obtain EEG features and video features includes:

[0033] Fill in the missing data in the EEG data, remove the high-frequency components and artifacts to obtain the artifact-free EEG data, and use the modal decomposition method to extract features from the artifact-free EEG data to obtain EEG features;

[0034] Perform face detection and localization using a face detection algorithm based on video data, perform key point detection on key positions of the face, extract facial expression features based on the key points, and use the facial expression features as video features.

[0035] Optionally, filling in the missing data in the EEG data, and removing high-frequency components and artifacts to obtain artifact-free EEG data, includes:

[0036] Filling in the missing data in the EEG data using an interpolation method; using a rereferencing method to change the reference baseline of the EEG signal;

[0037] Removing high-frequency components in the EEG data through low-pass filtering, extracting slow waves or oscillations at specific frequencies in the EEG data, to obtain EEG data with high-frequency components removed;

[0038] Using an artifact removal technique to remove artifacts in the EEG data with high-frequency components removed, to obtain artifact-free EEG data.

[0039] Optionally, the classifier includes: a convolutional neural network, a Yolo-v5 network, or a VGG network.

[0040] On the other hand, the present invention also provides an apparatus for rapid emotion assessment under variable gravity conditions, including:

[0041] An experimental data collection layer for collecting EEG data and video data in a microgravity environment;

[0042] A data processing layer for preprocessing the EEG data and video data to obtain EEG features and video features;

[0043] An emotion recognition layer for substituting the EEG features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result;

[0044] Wherein, the basic emotion recognition model includes a unimodal basic emotion recognition model and a multimodal basic emotion recognition model;

[0045] The unimodal basic emotion recognition model and the multimodal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and EEG data, and the corresponding emotion categories.

[0046] Optionally, it further includes a unimodal training module including:

[0047] Obtain EEG data and video data, and the emotion categories corresponding to the EEG data and the video data;

[0048] Preprocess the EEG data and video data to obtain EEG features and video features;

[0049] Construct a sample set from the EEG features, video features, and corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0050] Use the EEG features and video features in the training set as inputs, and the corresponding emotion categories as outputs to train the classifier, obtaining a preliminarily trained single-modal basic emotion recognition model;

[0051] Input the EEG features and video features in the test set into the preliminarily trained single-modal basic emotion recognition model to obtain an emotion recognition result;

[0052] Judge whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, use the preliminarily trained single-modal basic emotion recognition model as the trained single-modal basic emotion recognition model; otherwise, continue to train the preliminarily trained single-modal basic emotion recognition model based on the training set.

[0053] Optionally, it further includes a multi-modal training module for:

[0054] Obtain EEG data and video data, as well as the emotion categories corresponding to the EEG data and the video data;

[0055] Preprocess the EEG data and video data to obtain EEG features and video features;

[0056] Fuse the EEG features and video features to obtain multi-modal features;

[0057] Construct a sample set from the multi-modal features and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0058] Use the multi-modal features in the training set as inputs, and the corresponding emotion categories as outputs to train the classifier, obtaining a preliminarily trained multi-modal basic emotion recognition model;

[0059] Input the multi-modal features in the test set into the preliminarily trained multi-modal basic emotion recognition model to obtain an emotion recognition result;

[0060] Judge whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, use the preliminarily trained multi-modal basic emotion recognition model as the trained multi-modal basic emotion recognition model; otherwise, continue to train the preliminarily trained multi-modal basic emotion recognition model based on the training set.

[0061] Optionally, the specific implementation steps of preprocessing the EEG data and video data to obtain EEG features and video features in the single-modal training module and the multi-modal training module include:

[0062] Fill in the missing data in the EEG data, remove the high-frequency components and artifacts to obtain artifact-free EEG data, and use the modal decomposition method to extract features from the artifact-free EEG data to obtain EEG features;

[0063] Based on the video data, use a face detection algorithm to detect and locate the face, detect the key points at the key positions of the face, extract the facial expression features based on the key points, and use the facial expression features as video features.

[0064] Optionally, the steps of filling in the missing data in the EEG data and removing the high-frequency components and artifacts to obtain artifact-free EEG data in the single-modal training module and the multi-modal training module specifically include:

[0065] Fill in the missing data in the EEG data using an interpolation method; use a re-reference method to change the reference baseline of the EEG signal;

[0066] Remove the high-frequency components in the EEG data through low-pass filtering, extract the slow waves or oscillations at specific frequencies in the EEG data to obtain artifact-free EEG data after removing the high-frequency components;

[0067] Use an artifact removal technique to remove the artifacts in the EEG data after removing the high-frequency components to obtain artifact-free EEG data.

[0068] Optionally, the classifier includes: a convolutional neural network, a Yolo-v5 network, or a VGG network.

[0069] On the other hand, the present application also provides a computing device, including: at least one processor and a memory;

[0070] The memory is used to store one or more programs;

[0071] When the one or more programs are executed by the at least one processor, an emotion rapid assessment method under variable gravity conditions as described above is implemented.

[0072] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, an emotion rapid assessment method under variable gravity conditions as described above is implemented.

[0073] Compared with the prior art, the beneficial effects of the present invention are:

[0074] The present invention provides a method for rapid emotion assessment under variable gravity conditions, including: collecting electroencephalogram (EEG) data and video data in a microgravity environment; preprocessing the EEG data and video data to obtain EEG features and video features; substituting the EEG features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result; wherein, the basic emotion recognition model includes a single-modal basic emotion recognition model and a multi-modal basic emotion recognition model; the single-modal basic emotion recognition model and the multi-modal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and EEG data, and the corresponding emotion categories. The present invention improves the accuracy and real-time performance of emotion recognition under variable gravity conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a flowchart of a method for rapid emotion assessment under variable gravity conditions according to the present invention;

[0076] Figure 2 It is a flowchart of preprocessing video data and EEG data according to the present invention;

[0077] Figure 3 It is a process diagram of data annotation according to the present invention;

[0078] Figure 4 It is a statistical chart of the training set according to the present invention;

[0079] Figure 5 It is a diagram of electrode polarization according to the present invention;

[0080] Figure 6(a) 、 6(b) They are respectively analysis diagrams of electrooculogram artifacts presenting red and blue on the left front side of the eyes according to the present invention;

[0081] Figure 6(c) 、 6(d) They are respectively analysis diagrams of electrooculogram artifacts presenting red on the right front side of the eyes, and presenting red and blue on the front side of the eyes according to the present invention;

[0082] Figure 7 It is a comparison diagram of ICA analysis components according to the present invention;

[0083] Figure 8 It is a data comparison diagram before and after removing artifacts by ICA according to the present invention;

[0084] Figure 9 It is a frequency domain division diagram according to the present invention;

[0085] Figure 10 It is a schematic diagram of the kernel function according to the present invention;

[0086] Figure 11Schematic diagram of the design structure of the single-modal basic emotion recognition module of the present invention;

[0087] Figure 12 Flowchart of the operation of the single-modal basic emotion recognition module of the present invention;

[0088] Figure 13 Schematic diagram of the design structure of the multi-modal basic emotion recognition module of the present invention;

[0089] Figure 14 Flowchart of the operation of the multi-modal basic emotion recognition module of the present invention;

[0090] Figure 15 Schematic diagram of an emotion rapid assessment device under variable gravity conditions of the present invention;

[0091] Figure 16 Schematic diagram of a computer device of the present invention. Detailed implementation manners

[0092] The present invention provides an emotion rapid assessment method under variable gravity conditions, which can effectively identify and evaluate basic emotions and compound emotion states through a pre-trained basic emotion recognition model, and improve the accuracy and real-time performance of emotion recognition in a variable gravity environment.

[0093] Example 1:

[0094] An emotion rapid assessment method under variable gravity conditions, as Figure 1 shown, includes:

[0095] Step S1: Collect electroencephalogram data and video data in a microgravity environment;

[0096] Step S2: Perform data preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features;

[0097] Step S3: Substitute the electroencephalogram features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result;

[0098] Among them, the basic emotion recognition model includes a single-modal basic emotion recognition model and a multi-modal basic emotion recognition model;

[0099] The single-modal basic emotion recognition model and the multi-modal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and electroencephalogram data, and the corresponding emotion categories.

[0100] Regarding the basic emotion recognition model involved in the present application, this model includes: a single-modal basic emotion recognition model and a multi-modal basic emotion recognition model. The training processes of these two models are introduced below:

[0101] The training of the unimodal basic emotion recognition model includes:

[0102] Obtain electroencephalogram (EEG) data and video data, as well as the emotion categories corresponding to the EEG data and the video data;

[0103] Preprocess the EEG data and video data to obtain EEG features and video features;

[0104] Construct a sample set from the EEG features, video features, and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0105] Use the EEG features and video features in the training set as inputs, and the corresponding emotion categories as outputs to train a classifier, obtaining a preliminarily trained unimodal basic emotion recognition model;

[0106] Input the EEG features and video features in the test set into the preliminarily trained unimodal basic emotion recognition model to obtain emotion recognition results;

[0107] Judge whether the emotion recognition results are consistent with the corresponding emotion categories in the test set. If they are consistent, use the preliminarily trained unimodal basic emotion recognition model as the trained unimodal basic emotion recognition model; otherwise, continue to train the preliminarily trained unimodal basic emotion recognition model based on the training set.

[0108] The training of the multimodal basic emotion recognition model includes:

[0109] Obtain electroencephalogram (EEG) data and video data, as well as the emotion categories corresponding to the EEG data and the video data;

[0110] Preprocess the EEG data and video data to obtain EEG features and video features;

[0111] Fuse the EEG features and video features to obtain multimodal features;

[0112] Construct a sample set from the multimodal features and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0113] Use the multimodal features in the training set as inputs, and the corresponding emotion categories as outputs to train a classifier, obtaining a preliminarily trained multimodal basic emotion recognition model;

[0114] Input the multimodal features in the test set into the preliminarily trained multimodal basic emotion recognition model to obtain emotion recognition results;

[0115] Determine whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained multi-modal basic emotion recognition model is used as the trained multi-modal basic emotion recognition model; otherwise, the preliminarily trained multi-modal basic emotion recognition model is continuously trained based on the training set.

[0116] The single-modal basic emotion recognition model and the multi-modal basic emotion recognition model are introduced as follows:

[0117] The single-modal basic emotion recognition module aims to perform basic emotion recognition on EEG data, picture data, and video data. The main emotion types to be recognized include neutral, sad, afraid, disgusted, angry, surprised, and pleasant. Before training the single-modal emotion model, customized data preprocessing operations need to be performed to adapt to different modal types.

[0118] In the process of emotion recognition of EEG data, picture data, and video type data, as Figure 2 shown, extract and analyze the key features of specific emotion categories to train an accurate single-modal basic emotion model.

[0119] For basic emotion recognition, since no other multi-modal information such as audio and text is involved, the input can only be video images. For face expression recognition of video images, the basic steps are divided into four steps, namely: data collection, data preprocessing, feature extraction, and expression classification.

[0120] (1) The data set mainly includes two aspects. On the one hand, it is the facial videos recorded through microgravity simulation environment experiments, including two situations of head-high and head-low positions. On the other hand, it is an open-source basic emotion facial data set used to enhance the model classification effect.

[0121] (2) Data preprocessing includes: ① Face detection and localization: Use algorithms similar to MTCNN to detect and frame the facial area. ② Facial key point detection: Detect key positions such as eyebrows, eyes, nose, and mouth. ③ Expression feature extraction: Based on the key points, extract facial expression features such as eyebrow shape and eye state. ④ Expression classification: Input the extracted expression features into pre-trained convolutional neural network and other models to complete expression classification.

[0122] Data Collection and Annotation

[0123] Through experiments in the microgravity simulation environment, we used a video acquisition module to collect the facial videos of the subjects in each experiment, and cut them according to the time nodes of each emotion induction based on the facial videos, so as to construct a preliminary labeled emotion picture data set under the microgravity simulation environment. Collect a picture data set containing emotion information to ensure that there are various emotion expressions in the data set. This part is accurately operated manually.

[0124] After preliminary data processing, it is necessary to further annotate the pictures, marking the position of each face and the corresponding emotion category. This system uses annotation tools such as Labelme to complete this step, and the specific operation process is shown in Figure 3 as follows:

[0125] Model training

[0126] The classifier includes: convolutional neural network, Yolo-v5 network or VGG network. In this embodiment, taking the training of Yolo-v5 and VGG models as examples, the model training process will be introduced in detail.

[0127] ①Yolo-v5

[0128] This system uses the official provided Yolo-v5: According to the official documentation of Yolo-v5, install Yolo-v5 and its dependencies.

[0129] Prepare training data: Divide the annotated dataset into training set, test set and validation set, with a ratio of 7:2:1.

[0130] For the emotion recognition task, modify the configuration file of Yolo-v5, set the number of classes to the number of emotion categories in the dataset. The number of basic emotions to be recognized in this experiment is 7. Finally, adjust other hyperparameters of the model, such as learning rate, batch size, etc.

[0131] Statistics such as Figure 4 Histogram of the number of each category in the training set data shown in (upper left), set the x and y center values of all boxes at the same position to view the length and width of each label box in the training set data (upper right), draw a histogram of x and y variables to show the distribution of the dataset (lower left), draw a histogram of width and height variables to show the distribution of the dataset (lower right). In the figure, instances are examples, Anger is anger, Disgust is disgust, Fear is fear, Happy is happy, Neutral is neutral, Sad is sad, Surprise is surprise.

[0132] Summarize the labels of the training set data and draw a relationship diagram (linear or non-linear, with or without obvious correlation) between the four variables of x, y, width, and height of the training set data labels.

[0133] The following are the training set pictures and their labels used by the model in the first three batches of the first training epoch.

[0134] ①VGG

[0135] For the VGG model, the CrossEntropy loss function is set to calculate between the input and the label. The optimizer selects the momentum optimizer paddle.optimizer.Momentum() in the paddle api. The specific parameters are as follows:

[0136] learning_rate (float|_LRScheduler, optional) - Learning rate, used for the calculation of parameter updates. It can be a floating-point value or a _LRScheduler class, and the default value is 0.001

[0137] momentum (float, optional) - Momentum factor.

[0138] parameters (list, optional) - Specify the parameters that the optimizer needs to optimize. This parameter must be provided in dynamic graph mode; in static graph mode, the default value is None, and all parameters will be optimized at this time.

[0139] use_nesterov (bool, optional) - Enable Nesterov momentum, default value is False.

[0140] weight_decay (float|Tensor, optional) - Weight decay coefficient, which is a float type or a shape type, and a Tensor type with the data type of float32. The default value is 0.01

[0141] grad_clip (GradientClipBase, optional) – Gradient clipping strategy, supporting three clipping strategies: paddle.nn.ClipGradByGlobalNorm, paddle.nn.ClipGradByNorm, paddle.nn.ClipGradByValue. The default value is None, and no gradient clipping will be performed at this time.

[0142] name (str, optional) - This parameter is used by developers when printing debug information. For specific usage, please refer to Name, and the default value is None for the cross-entropy loss.

[0143] The basic EEG-based basic emotion recognition process:

[0144] In the field of emotion analysis, the specificity of electroencephalogram (EEG) activity refers to the degree of association between EEG signals and specific emotional states. Electroencephalogram is a neurophysiological technique that records the electrical activity of the brain by placing electrodes on the scalp. When a person experiences different emotional states, their brain electrical activity changes, and these changes can be detected in the EEG. Emotional specificity means that the EEG signals contain specific patterns or features related to specific emotional states. Researchers attempt to determine which patterns or frequency components in the EEG signals correspond to different emotional states (such as anger, happiness, sadness, etc.). For example, the enhancement or weakening of certain frequency bands may be related to anger, while changes in other frequency bands may be related to happiness.

[0145] The EEG signals of this system are collected by a wired measurement device, and the basic emotions of the user can be identified by analyzing the recorded EEG data. Since these EEG signals contain very complex vibration patterns, the analysis of the original EEG data may be very inefficient. Traditional methods require data preprocessing for the obtained original EEG data. Here, the data preprocessing mainly eliminates signals such as electrooculogram and electrocardiogram. On this basis, further feature extraction guided by domain knowledge (signal processing, sequence pattern mining, affective psychology, etc.) is carried out to extract as many emotion-related features as possible. The refined features are called data features in machine learning and are further organized to construct training samples. Further, feature selection methods can screen out features that are closely related to the emotional state and contribute to improving the emotion recognition performance from all candidate feature sets. Finally, we try to select various statistical machine learning models or deep learning models to build the selected features and iteratively evaluate their performance by calculating evaluation metrics until the goal that the model output is close to the true value is achieved.

[0146] (1) Data preprocessing

[0147] ① Event alignment

[0148] The EEG instrument in this experiment does not involve the event recording function, so the collected EEG signals need to be event-aligned and cut.

[0149] ② Electrode positioning 10 - 20

[0150] Electrode positioning in EEG data processing is a crucial step. It involves placing electrodes on the scalp to record the electrical activity of the brain. The purpose of electrode positioning is to ensure that the recorded signals can accurately reflect the activity of specific brain regions and can provide accurate information about brain function. According to the research purpose and the region of interest in the brain, the electrode layout is determined. The selection and placement of electrodes usually rely on the international standard 10 - 20 system or other similar systems. These systems divide the scalp into specific regions and specify the relative positions of placing electrodes in these regions, such as Figure 5As shown in the figure, where 1. Frontal Region, midline electrodes: FPZ, FZ, left electrodes: FP1, AF3, F3, F5, F7, FC5, FC3, right electrodes: FP2, AF4, F4, F6, F8, FC6, FC4; 2. Central Region, midline electrode: CZ, left electrodes: C3, C5, right electrodes: C4, C6; 3. Parietal Region, midline electrode: PZ, left electrodes: P3, P5, P7, right electrodes: P4, P6, P8; 4. Occipital Region, midline electrode: OZ, left electrode: O1, right electrode: O2; 5. Temporal Region, left electrodes: T7, TP7, right electrodes: T8, TP8; 6. Posterior Parietal Region, left electrodes: PO3, PO5, PO7, right electrodes: PO4, PO6, PO8; 7. Fronto-Central Region, midline electrodes: AFZ, FCZ, left electrodes: AF3, FC3, FC5, right electrodes: AF4, FC4, FC6; 8. Other Regions, occipital bottom electrodes: CB1 (left), CB2 (right).

[0151] ③ Bad Channel Interpolation and Re-referencing

[0152] Bad Channel Interpolation and Re-referencing are common operations in EEG data processing, used to improve signal quality and accuracy.

[0153] For Bad Channel Interpolation: When data is missing from a certain electrode due to a fault or other reasons, in order to maintain data integrity, interpolation methods can be used to fill in the data of these bad channels. Bad Channel Interpolation can also be used to repair signal anomalies caused by electrode movement or other external interferences. Generally, linear interpolation is used: by using the data of adjacent electrodes, the data on the bad channel electrode is estimated according to a linear relationship. The calculation formula (linear interpolation) is as follows:

[0154]

[0155] Where Y i is the estimated value of the bad channel electrode, X i is the position of the bad channel single electrode, Y i+1 and Y i-1 are the values of adjacent electrodes, X i+1 and X i-1 are the positions of adjacent electrodes.

[0156] Re-referencing is to change the reference baseline of the electroencephalogram (EEG) signal to reduce external interference and make the data easier to interpret and analyze.

[0157] Common re-referencing methods include average reference, ear reference, etc., to ensure that the recorded signal reflects brain activity rather than some global changes on the scalp surface. The commonly used methods are as follows:

[0158] Average reference: The signals of all electrodes are averaged to obtain a global average, and then the signal of each electrode is subtracted by this global average.

[0159] Ear reference: Use the electrodes on the ears as a reference because there is less tissue under the scalp and it is closer to the brain surface.

[0160] The calculation formula for average reference is as follows:

[0161]

[0162] X new is the signal after the new reference, X original is the original signal, and N is the number of electrodes.

[0163] ④ Low-pass filtering / High-pass filtering

[0164] Low-pass filtering makes the low-frequency components in the signal more prominent by removing high-frequency components. This is very useful for extracting slow waves or oscillations at specific frequencies in the EEG signal. Low-pass filtering can also help remove high-frequency noise and improve the signal-to-noise ratio of the signal. The data reliability of physiological signals is determined by the signal-to-noise ratio. The better the noise detection method, the higher the reliability and accuracy. The commonly used method: Butterworth Filter: This is a commonly used filter type that can achieve a smooth frequency cut-off. The calculation formula (Butterworth low-pass filter) H(s) is as follows:

[0165]

[0166] where s is the frequency variable, w c is the cut-off frequency, and n is the order of the filter.

[0167] High-pass filtering highlights the high-frequency information in the signal by removing low-frequency components. This is very useful for analyzing and studying the rapid changes in the EEG signal, such as the analysis of event-related potentials (ERPs). High-pass filtering can also help remove low-frequency noise and improve the signal-to-noise ratio of the signal. The Butterworth filter can also be used to implement high-pass filtering by adjusting the order and cut-off frequency of the filter to control the filtering effect.

[0168] The specific formula is as follows:

[0169] In these two filters, f is the complex frequency, where f = σ + jω, σ represents the damping term, ω represents the angular frequency, and j is the imaginary unit. The cut-off frequency ω c determines the position where the signal is cut off in the frequency domain, while the order n affects the steepness of the filter.

[0170] ④ ICA artifact removal

[0171] Generally speaking, electroencephalogram (EEG) signals are measured by installing sensors on the scalp. However, during this process, noise or various artifacts, usually including eye movements, muscle movements, heartbeats, sweat, and tongue movements, are mixed with the signals. It is impossible to completely eliminate these noises and artifacts; however, they can be minimized using techniques such as principal component analysis (PCA), regression analysis, adaptive filtering, independent component analysis (ICA), and wavelet transform (WT).

[0172] Preprocessing the collected EEG data mainly involves removing EEG artifacts with a frequency less than 4 Hz caused by blinking, electrocardiogram (ECG) artifacts around 1.2 Hz, electromyogram (EMG) artifacts with a frequency greater than 30 Hz, power frequency artifacts in the environment, etc. Frequencies between 50 and 60 Hz, and so on. Independent component analysis (ICA), discrete wavelets, or band-pass filters (such as Butterworth filters) can be used to remove these above-mentioned artifacts to retain the rhythm components related to emotional activities. The filtering process cannot eliminate all artifacts in the EEG signal, so additional processing is required. This system uses ICA to remove artifacts. The basic idea of ICA is to regard the observed signal as a linear combination of independent components and find a set of weights to maximize the statistical independence of the mixed signal. In EEG signal processing, ICA attempts to separate the EEG signal and artifact signals into independent components. There are the following main techniques for distinguishing artifacts through the power spectral density map:

[0173] Eye components describe eye movements. Each retina (the part of the eye that records incident light) generates an electric field that can be effectively modeled as an "equivalent current dipole" (infant development). Generally, eye movements are divided into two parts: vertical movements and horizontal movements. Other eye components can also be found, such as diagonal directions, but they are rare and depend on the experiment. For all electrooculogram (EOG) artifacts, the power spectrum varies depending on the experiment and the person, but generally most of the power will reside at frequencies below 5 Hz because people generally do not move their eyes faster. As shown in Figures 6(a), 6(b), 6(c), and 6(d), they can all be considered EOG artifacts, where, Component Time Series: Displays the time series of independent components (ICA components), reflecting the dynamic changes of the signal on the time axis. Can be used to analyze whether the artifact is associated with a specific time event. Activity Power Spectrum: Shows the distribution of the power of the signal as a function of frequency. Helps identify the activity in specific frequency bands and judge whether there are artifacts (such as EOG artifacts usually having more low-frequency components). Frequency: The abscissa represents the frequency range, used to analyze the spectral characteristics of the signal. When analyzing EOG artifacts, the low-frequency band (such as 0 - 4 Hz) is the focus. Continuous Data: Refers to the original unsegmented signal data. Suitable for observing the persistence of the artifact and its influence range. Dipole: Displays the source model of the ICA component, marking the source point with a dipole. Used to judge whether the artifact is caused by a specific brain region or non-brain region (such as the eye). Epoched Data: The data displayed after cutting the signal by time windows. Facilitates alignment analysis with stimulus events, especially before and after artifact correction.

[0174] The remaining artifacts can be manually removed by observing the specific power spectral density feature map. Using the ICA module provided in Python-MNE, the number of independent components to be extracted in this system is specified as 63. In electroencephalogram (EEG) signal processing, it is usually desirable to select a sufficient number of components to ensure that all possible signal sources are taken into account, but attention should also be paid not to select too many to avoid overfitting. When using the ICA algorithm, plot the time-domain or spatial-domain representation of the independent components (ICs), as Figure 7 shown, where, ICA Components: Each independent component represents a relatively independent component in the EEG signal, which may be a brain signal, an artifact, or noise. ICA000 to ICA059: Represent the 0th to 59th independent components decomposed from the original EEG signal. By comparing the time series, power spectrum, and spatial distribution of these components, it can be judged which components may be artifacts (such as EOG artifacts, electromyogram (EMG) artifacts).

[0175] This figure can help visually view the time / spatial characteristics of each independent component, thus enabling better artifact recognition and selection of components to be removed. As shown below, we can select the parts we need to remove based on the characteristics of different artifacts.

[0176] After removing artifacts through ICA and viewing the comparison data before and after, there are obvious differences. For example, Figure 8 the visualization effect of EEG data at the same time point is shown as follows:

[0177] (2) Feature extraction

[0178] EEG feature extraction is the process of extracting relevant information about brain activities from EEG signals. These features can be used for analyzing, classifying, recognizing, or understanding different brain states and activities. The following are some common EEG feature extraction methods and features:

[0179] EEG time-domain features: The most direct way to extract EEG signals is to calculate statistics such as mean, variance, skewness, kurtosis, and peak-to-peak interval. These statistics can characterize the time-domain characteristics of EEG signals.

[0180] EEG frequency-domain features: Another way to describe the signal, which has good anti-interference ability against noise and can reflect the details of each component of the signal. For example, the power spectral density (PSD) is usually obtained through the fast Fourier transform (FFT). The original EEG signal can be decomposed into several different frequency bands, such as Delta (1 - 4Hz), Theta (4 - 8Hz), Alpha (8 - 13Hz), Beta (13 - 30Hz), and Gamma (>30Hz). Then calculate the average power of a specific frequency band and use this feature.

[0181] Time-frequency composition features: A single time domain or frequency domain cannot fully characterize the features of the signal. For non-stationary signals, their frequency components often change with the changes of brain nerves and cognitive processes. It is necessary to transition from static frequency component analysis to dynamic time-frequency joint analysis. Two methods can be used for the transition to dynamic frequency joint analysis:

[0182] (1) Using the short-time Fourier transform (STFT), which is only suitable for analyzing stable signals and it is difficult to obtain high resolution in both the frequency domain and the time domain based on the Heisenberg uncertainty principle.

[0183] (2) Wavelet analysis (such as discrete wavelet transform), which is good at showing subtle local features in both signal domains and is widely used in processing non-stationary neural signals, especially electroencephalograms.

[0184] Specific operation process: The original EEG signal is decomposed by basis functions obtained by scaling and translating the mother wavelet.

[0185] The discrete wavelet transform (DWT) decomposes a signal into approximation coefficients and detail coefficients. The approximation coefficients describe the high-scale, low-frequency part of the signal, while the detail coefficients describe the low-scale, high-frequency part of the signal. The decomposition process is an iterative process, i.e., the approximation coefficients at one scale can be further decomposed into finer-grained coefficients corresponding to the next scale. This process is iterative, generating a series of approximation coefficients and detail coefficients belonging to different scales, as Figure 9 shown, where Level 1 to Level 7 represent the decomposition levels of the wavelet decomposition. Each level corresponds to a different frequency range: Level 1: 0.5 - 128 Hz, Level 2: 0.5 - 64 Hz, Level 3: 0.5 - 32 Hz, Level 4: 0.5 - 16 Hz, Level 5: 0.5 - 8 Hz, Level 6: 0.5 - 4.2 Hz, Level 7: 1.5 - 2.4 Hz, a1 to a7 represent the approximation coefficients (Approximation Coefficients) after each level of decomposition, reflecting the low-frequency components of the signal. d1 (band1) to d7 (band7) represent the detail coefficients (Detail Coefficients) of each level of decomposition, corresponding to the high-frequency components in different frequency bands: d1: 64 - 128 Hz, d2: 32 - 64 Hz, d3: 16 - 32 Hz, d4: 8 - 16 Hz, d5: 4.3 - 8 Hz, d6: 2.4 - 4.2 Hz, d7: 1.5 - 2.4 Hz. In addition, there are many modal decomposition methods, such as empirical mode decomposition (EMD), multi-scale empirical mode decomposition (MEMD), and variational mode decomposition (VMD), which are also very suitable for extracting non-linear and non-stationary EEG features. Among them, multi-channel electroencephalogram is decomposed into multiple intrinsic mode functions (IMFs), based on which more representative features can be extracted.

[0186] (3) Fitting of multiple machine learning models

[0187] The emotion recognition task usually involves mapping EEG data to emotion categories. In deep learning, a common method is to use a neural network model. Figure 10 is a basic model fitting process route, which includes data preprocessing, construction of multiple models, training and evaluation. Finally, the prediction results of multiple models are combined using ensemble learning to obtain a more robust and accurate classification result.

[0188] ① Random forest

[0189] Random Forest is an ensemble learning algorithm based on decision trees. It improves the prediction performance of the model by constructing multiple decision trees and aggregating the prediction results of these trees. During the construction of each tree, Random Forest adopts two random strategies: sampling with replacement (Bootstrap Aggregating, i.e., Bagging) and random feature selection. Sampling with replacement ensures the diversity of training samples, while random feature selection ensures that only part of the features are considered when splitting nodes in each tree, thus further enhancing the generalization ability of the model.

[0190] The structure of Random Forest can be summarized as the following parts:

[0191] (1) Multiple decision trees: Random Forest consists of multiple decision trees, and each tree is trained and predicted independently.

[0192] (2) Node splitting: During the construction of the tree, node splitting adopts random feature selection to find the optimal splitting feature from part of the features.

[0193] (3) Leaf nodes: The leaf nodes of the tree are used to output the prediction results. For classification problems, voting is usually adopted; for regression problems, the average value is used.

[0194] For classification problems, the prediction formula of Random Forest is as follows, where represents the prediction result of Random Forest, T represents the number of trees, represents the prediction result of the t-th tree:

[0195]

[0196] ② Support Vector Machine

[0197] Support Vector Machine is a machine learning algorithm for classification and regression. Its basic idea is to map the data into a high-dimensional space and then find an optimal hyperplane in this space to separate data of different classes. SVM achieves this goal by maximizing the margin, that is, maximizing the distance between two classes of data while ensuring that the distance from each class of data to the hyperplane is greater than 1.

[0198] The structure of SVM can be summarized as the following parts:

[0199] Hyperplane: SVM separates data of different classes by finding an optimal hyperplane. The schematic diagram is as follows. The hyperplane can be represented by the classification function f(x) = W Tx + b indicates that when the function value is greater than 0, it is located on the hyperplane; otherwise, it is below the hyperplane. Here, f(x) is the classification function, W is the weight vector representing the normal vector of the hyperplane, which determines the direction of the hyperplane, x is the sample feature vector, and b is the bias term, which determines the position of the hyperplane.

[0200] (2) Support vectors: The nearest data points on both sides of the hyperplane are called support vectors, and they have an important impact on the position of the hyperplane.

[0201] (3) Kernel function: When the data cannot be directly mapped to a high-dimensional space, SVM maps the data to a high-dimensional space through a kernel function. In the scenario of emotion recognition in a microgravity simulation environment, a kernel function is needed to map EEG data to a high-dimensional space. In the system, we use a Gaussian kernel, and its formula is as follows:

[0202]

[0203] In the formula, k(x1, x2) is the kernel function used to map the samples to a high-dimensional space. x1 and x2 are different sample points respectively, and σ is the width parameter of the Gaussian kernel function.

[0204] The Gaussian kernel will map the original space to a function of an infinite-dimensional space. If σ is chosen to be very large, the weights on high-order features actually decay very quickly, so in fact (numerically approximated), it is equivalent to a low-dimensional subspace; conversely, if σ is chosen to be very small, any data can be mapped to be linearly separable, but the resulting problem may be very serious overfitting. Generally speaking, by adjusting the parameter σ, the Gaussian kernel actually has quite high flexibility and is also one of the most widely used kernel functions. Figure 10 The example shown maps low-dimensional linearly inseparable data to a high-dimensional space through a Gaussian kernel function. Here, Classification model represents the classification model used to divide the input data into multiple categories according to the input data. In EEG data, the classification model can be used to identify emotional states. Theta wave is a range of brain wave frequencies, usually appearing between 4 - 7 Hz. Alpha wave is a range of brain wave frequencies, usually appearing between 8 - 12 Hz. Gamma wave is a high-frequency brain wave, usually appearing between 30 - 100 Hz. Beta wave is a range of brain wave frequencies, usually appearing between 13 - 30 Hz. EEG data is the scalp potential change recorded by an electroencephalogram device, used to reflect the activity state of the brain.

[0205] ③ k-Nearest Neighbors algorithm (KNN)

[0206] The K-Nearest Neighbors (KNN) algorithm is an instance-based learning method. Its basic idea is that for a sample with an unknown class, KNN will find the K nearest neighbors of this sample in the training set, and then predict the class of this sample through majority voting. In KNN, the number of neighbors K is a parameter that needs to be adjusted, and it has an important impact on the prediction performance of the model.

[0207] The structure of KNN can be summarized into the following parts:

[0208] (1) Training set: KNN first requires a set of training data, which is used to find the nearest neighbors during the prediction process.

[0209] (2) Unknown sample: During prediction, KNN will compare the features of the unknown sample with the data in the training set to find the K nearest neighbors.

[0210] (3) Majority voting: By performing majority voting on the labels of the neighbors, KNN predicts the class of the unknown sample.

[0211] The prediction formula of KNN is as follows:

[0212]

[0213] In the formula, is the predicted value of the KNN model, representing the prediction result of the class or attribute value of the target sample. K is the number of nearest neighbors, referring to the number of the nearest samples used for prediction, and y neighbork is the true value of the k-th neighbor sample.

[0214] ④ Voting mechanism

[0215] The voting mechanism is a method of Ensemble Learning, which obtains a more robust and accurate classification result by combining the prediction results of multiple models. The voting mechanism used in this system uses a combination of soft voting and weighted voting.

[0216] Soft Voting:

[0217] Soft voting takes into account the prediction probabilities of each model, and the final class is the weighted average of the prediction probabilities of each model. For models with probability outputs, such as support vector machines and some ensemble models, soft voting can make better use of the confidence information of the models.

[0218] Weighted Voting:

[0219] Assign a weight to each model, and determine the weight according to its performance or confidence level. This method can be used in both hard voting and soft voting.

[0220] Voting mechanisms are commonly used to combine models of different types, with different parameter settings, or using different feature subsets to improve overall performance. This can help reduce overfitting, improve the robustness of the model, and in some cases, improve accuracy. Through this mechanism, the single-modal EEG emotion recognition model can be made more perfect.

[0221] Based on the neural network model of deep learning theory, the design of the single-modal basic emotion recognition module includes EEG data preprocessing and feature extraction, as well as subject face detection and feature localization, neutral emotion recognition, happy emotion recognition, sad emotion recognition, fear emotion recognition, disgust emotion recognition, anger emotion recognition, surprise emotion recognition, etc. functions, as Figure 11 shown.

[0222] Table 1 Module and Function Description Table

[0223]

[0224]

[0225] In the single-modal basic emotion recognition module, the user enters the system through the GUI interface design of each function of the module. First, log in through the login page. After successful login, the system will automatically jump to the main page. The user can select the single-modal basic emotion recognition function, and can analyze for picture and video modalities, or can analyze for EEG data. The EEG modality data can be analyzed for continuous segments and single segments. When the recognition is completed, the recognition result will be returned in the form of a percentage. After the recognition function is completed, the system can be exited. The specific working process of the single-modal basic emotion recognition module is as Figure 12 shown.

[0226] The multi-modal basic emotion recognition module aims to fuse EEG data and picture data and then recognize basic emotions. The main emotion types to be recognized include neutral, sad, fear, disgust, anger, surprise, and happy. The fusion algorithms of multi-modal learning can generally be divided into three levels: data fusion, feature fusion, and decision fusion.

[0227] Early feature-based multi-modal fusion methods perform shallow fusion after early feature extraction. Fusing features of different modalities at the shallow layer of the model is equivalent to unifying the features of different single-modalities into the same parameter space. Due to the differences in information between different modalities, the features usually contain a large amount of redundant information, and dimensionality reduction methods are often needed to remove the redundant information. The dimensionality-reduced features are input into the model to complete feature extraction and prediction.

[0228] The mid-term model-based multimodal fusion method inputs multimodal data into a network, and the intermediate layer of the model performs feature fusion between modalities. The model-based modality fusion method can select the location of modality feature fusion to achieve intermediate interaction. Model-based fusion usually uses multi-kernel learning, neural networks, graph models, and alternative methods.

[0229] The late decision-level fusion method is used to fuse information from different modalities. Decision-level fusion means training models separately for data from different modalities and combining the outputs of different modalities into the final decision. Multimodal models based on decision fusion usually use methods such as averaging, majority voting, weighted, and learnable models to fuse modalities. Such models are usually lightweight and flexible. When any modality is missing, decisions can be made using the remaining modalities.

[0230] In the process of fusing EEG data and picture data, multimodal data is often large-capacity, high-rate, and diverse in data forms, including structured, semi-structured, etc. data, and these data are located in different spaces. Direct fusion not only has technical difficulties but also brings problems of dimensionality disaster and model convergence. On the one hand, when using decision fusion, since people are redundant when expressing emotional information, decision fusion often brings classification errors; on the other hand, the independence assumption of EEG data and image information loses the mutual information between modalities.

[0231] Figure 13 and Figure 14 are the main steps and functional flowcharts of the multimodal basic emotion recognition module under both strategies.

[0232] Feature-based fusion method:

[0233] To further improve the accuracy of emotion recognition, an emotion recognition method based on multimodal fusion is introduced. The multimodal fusion model can obtain emotion recognition results by fusing different physiological signal features.

[0234] Aiming at the problems of single modality and low accuracy, this system uses a multimodal emotion recognition method based on expression-EEG interaction with a deep autoencoder, that is, fusing EEG signals and facial expression signals through a bimodal deep autoencoder (BDAE), and then using a supervised learning method with a classifier for emotion recognition.

[0235] Deeply exploring the capabilities of EEG signals and facial expression signals is to distinguish and characterize different emotions. Combining EEG signals with facial expression signals through different model fusion strategies (including deep neural networks) is to establish a multimodal emotion recognition model that combines internal neural models and external subconscious behaviors.

[0236] To study the stability of the emotion expression ability of EEG and facial expression signals over time, BDAE was used for model fusion. After fusing the EEG signals and facial expression signals, the accuracy of the emotion recognition model was effectively improved. After the training process, the middle layer (i.e., the third layer of BDAE, with a total of five layers) was used as the extracted feature, and then it was sent to the classifier for supervised learning training to obtain the emotion classification model. In this way, the emotion recognition ability of the model will be significantly improved, which is beneficial to the high-order features related to emotions in the two signals extracted by the deep neural network.

[0237] ②EEG feature extraction

[0238] The extraction of EEG features in the multimodal basic emotion model recognition module is different from that in the single-modal basic emotion recognition module. Here, we propose a method for selecting EEG signal characteristics based on decision trees, which can not only achieve the objectivity of feature selection but also reach a relatively high classification accuracy.

[0239] Principle of decision tree Decision tree is one of the most classic and commonly used algorithms in data mining. Compared with other data mining algorithms, decision tree has three advantages:

[0240] (1) Decision tree is an algorithm that is very easy to understand;

[0241] (2) During the process of training the decision tree, researchers do not need to know the relevant background knowledge of the training data;

[0242] (3) The classification accuracy of the decision tree is relatively high. Considering the three major advantages of the decision tree, the decision tree algorithm is usually used for data classification.

[0243] The decision tree here uses the greedy algorithm to output a tree-like structure. EEG signal is a non-stationary, non-linear, and high-dimensional weak physiological signal, so it is very suitable to use the decision tree for extracting EEG features. The specific steps include:

[0244] The first step: According to the traditional method of manually extracting EEG features, time-domain or frequency-domain features and other features related to the nature of EEG can be extracted. After the feature extraction operation, the EEG signal feature vector is input.

[0245] The second step: Use the decision tree c4.5 algorithm to divide the input EEG signal feature vector into dominant features and non-dominant features.

[0246] The third step: Discard the non-dominant EEG signals from the input EEG signals, reorganize the dominant signals, and finally obtain a new vector, and then output the reorganized vector to represent the feature vector re-extracted by the decision tree.

[0247] In the process of constructing a decision tree, the attribute with the largest information gain ratio of the training set is selected as the splitting attribute, and then the branches are divided according to different values, and each branch is recursively operated. When pruning a complete tree, the post-pruning method is selected to form a simple decision tree. The splitting attributes of all nodes included in the simple decision tree are defined as the dominant attributes.

[0248] After removing the dominant features from the initial features, the remaining features are non-dominant features. For the initial feature vector X, it is reorganized according to the dominant features to form a new feature vector. This reorganized feature vector is selected after applying the decision-tree-based feature selection method.

[0249] ③ Facial feature extraction

[0250] Facial expressions are mirrors of visual information and a key part of conveying emotions. Human faces and facial emotions can reflect different emotional states of people. For example, when people are in a happy state, their lips often open, the corners of their mouths turn up, and their eyes become smaller. When people are in an angry state, they often open their eyes, frown, furrow their eyebrows, and twitch the zygomatic muscles. Identifying these states by computer is facial expression recognition.

[0251] Sparse representation is a method for feature extraction, often applied to computer vision tasks such as facial emotion recognition. This method captures the important features of the data by finding a set of sparse representations in high-dimensional data, that is, representations with only a few non-zero elements. In facial emotion recognition, sparse representation can be used to extract the key features of facial expressions.

[0252] The vector coefficients determine the facial expression category of the test sample. The specific implementation process is as follows:

[0253] Step 1: Preprocessing stage. Input seven types of facial expression samples, g represents a certain type of emotion, u g,1 represents a picture sample in class g, B g represents the matrix formed by this sample, and each specific sample is represented as b g,h , and h is a grayscale image with a size of m*n pixels.

[0254] Step 2: Stack each column of the sample b g to form a column vector u g,h , and the new matrix B g = [u g,1 , u g,2 , u g,3 ,...,] ∈ R m×n , where R m×n is the two-dimensional matrix into which each image sample in the training sample set is converted, m is the number of rows of the image, and n is the number of columns of the image.

[0255] Each column represents the facial expression training sample of the g-th object, and a total of n training samples of type k form the total training sample set matrix B.

[0256]

[0257] In the formula, B g represents the emotion training sample of the g-th object, B k represents the matrix composed of the training samples of the i-th class, u g,1 , is the feature vector of the g-th object in the training sample set.

[0258] Step 3: After obtaining the matrix B g For the test sample, first convert it into the form of a column vector, and any g-th object can be approximated by a linear combination of the training samples of this class:

[0259] y = a g,1 u g,1 a g,2 u g,2 ,..., a g,h u g,h

[0260] In the formula, x is the expansion coefficient vector of the test sample y with respect to the total training sample set B:

[0261]

[0262] In the formula, y1, y2, y g are the feature vectors representing the training samples from the 1st to the g-th respectively, a g,1 , a g,2 , a g,h are the coefficients related to the test sample y, usually used for linear combination, reflecting the projection weight of the test sample on the training samples, u g,h is the feature vector representing the g-th class of training samples.

[0263] Step 4: Find the approximate sparse solution of y = Ax is to obtain the emotional states of different facial expressions. The number of variables in the equation system is greater than the number of equations. Therefore, the solution of y is not unique. Since sparsity is defined by the 0-norm, the L0-norm minimization method can be used to solve it as follows: Subject to y = Ax;

[0264] In the formula, || ||0 is the L0-norm of the vector, representing the number of non-zero elements in the vector Ax. When the coefficient vector x is sparse enough, the L1-norm approximation method can be used to solve it as follows: Constrained by y = Ax, where || ||1 in the formula is the L1 norm of the vector, representing the sum of the absolute values in vector x, V t is a linear transformation matrix used to constrain x, representing the prediction matrix of the model or a certain prior information matrix. It plays a constraining role in the solution of the sparse vector x.

[0265] ③ Modal fusion

[0266] Using a deep neural network to fuse two signals can effectively improve the accuracy of emotion recognition and has better performance than traditional model fusion methods. We use a deep neural network to fuse EEG signals and facial expression signals to further improve the accuracy of emotion recognition. The deep neural network consists of two parts: encoding and decoding. In the encoding part, two RBM (Restricted Boltzmann Machine) models are trained using the features of EEG signals and eye movement signals.

[0267] After training, the hidden layers h EEG 、h Face and weights w1, w2 of the two RBMs can be obtained. Combine h EEG and h Face into the visible layer of another new RBM, and then train the RBM to obtain the corresponding weights.

[0268] In the decoding part, two-layer RBMs are implemented to reconstruct the input features, thus forming a deep encoder. The weights of each layer of network connection are W1, W2, W1 T , W2 T .

[0269] Finally, unsupervised training is first used, mainly for feature extraction. Then, the reconstructed features obtained are used as the input of the classifier, and supervised learning is used to finally obtain our multi-modal emotion recognition model.

[0270] Considering that the multi-modal emotion recognition model based on feature fusion requires more computing resources, even after the model is trained, we still need to extract and fuse the data of EEG and picture modalities first every time we perform an emotion recognition task. The process of reshaping multi-modal data into a shape acceptable to the emotion recognition model will take a long time. At this time, drawing on the idea of ensemble learning, using a decision-based fusion method, only the information of multiple features needs to be combined through certain decision rules or strategies to obtain the final fusion result. The algorithms for EEG and picture modality emotion recognition are the same as those of the single-modal basic emotion recognition module.

[0271] (1) Decision rule

[0272] Decision-based fusion methods involve defining a set of rules or conditions that describe how to combine different features or information sources. These rules can be set manually or learned. In this system, we use manually set rules as follows:

[0273] Use EEG data for emotion classification to obtain a probability result P1; use facial images for emotion classification to obtain a probability result P2. If the predictions of P1 and P2 are consistent, then output that category as the final result. If P1 and P2 are inconsistent, then compare the probability values of the two results and output the category with the higher probability as the final result.

[0274] (2) Weight assignment

[0275] In the decision-making process, different features or information sources may have different importance. Therefore, decision rules usually assign weights to each feature to reflect its relative contribution. These weights can be fixed or learnable parameters.

[0276] Since EEG directly measures brain activity, emotion recognition is more accurate. Give a greater weight to the EEG result, set as w1 = 0.7, and give a smaller weight to the facial image result, set as w2 = 0.3.

[0277] (1) Decision-making process

[0278] In the process of feature fusion, according to the decision rules and weight assignment, the system or model makes a final decision or generates a fusion result. This process can be a simple weighted sum, or more complex decision trees, rule sets, or other decision models. For simplicity and speed, we use the following decision-making process:

[0279] Calculate the weighted average result: P = w1P1 + w2P2, and select the category with the highest probability in P as the final output.

[0280] (2) Adaptive adjustment

[0281] Some decision-based feature fusion methods may be adaptive, that is, they can adjust decision rules or weight assignment according to real-time or dynamic situations to adapt to different working conditions. For multi-modal emotion recognition in a microgravity simulation environment, we use the following strategy for adjustment:

[0282] We have collected enough training data, evaluated the accuracy of the two modal results in real time, and dynamically adjusted w1 and w2 to reflect the latest accuracy situation.

[0283] Through the above design, the complementary information of the two modalities is combined, and the priority of the EEG results is also considered, realizing intuitive and interpretable feature fusion. When more data is obtained, the decision rules and weight allocation can be continuously improved through automatic optimization to improve the classification performance.

[0284] The design of the multi-modal basic emotion recognition module includes a multi-modal model based on feature fusion and a multi-modal model based on decision fusion. Both methods can have the following functions: neutral emotion recognition, pleasant emotion recognition, sad emotion recognition, fear emotion recognition, disgust emotion recognition, anger emotion recognition, surprise emotion recognition, etc., as shown in Table 2.

[0285] Table 2 Module and Function Description Table

[0286]

[0287] In the multi-modal emotion recognition module, the user enters the system through the GUI interface design of each function of the module. First, log in through the login page. After successful login, the system will automatically jump to the main page. The user can select the multi-modal basic emotion recognition function. When using the multi-modal basic emotion recognition, it is necessary to pay attention to uploading a matching pair of EEG and pictures. Both modalities of data must be uploaded for correct operation. When the recognition is completed, the recognition result will be returned in the form of a percentage. After the recognition function is completed, the system can be exited. The specific working process of the multi-modal basic emotion recognition module is as Figure 14 shown.

[0288] Embodiment 2:

[0289] Based on the same inventive concept, the present invention also provides an emotion rapid assessment device under variable gravity conditions, as Figure 15 shown, including:

[0290] An experimental data collection layer for collecting EEG data and video data in a microgravity environment;

[0291] A data processing layer for preprocessing the EEG data and video data to obtain EEG features and video features;

[0292] An emotion recognition layer for substituting the EEG features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result;

[0293] Wherein, the basic emotion recognition model includes a single-modal basic emotion recognition model and a multi-modal basic emotion recognition model;

[0294] The unimodal basic emotion recognition model and the multimodal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and EEG data, as well as the corresponding emotion categories.

[0295] Optionally, it further includes a unimodal training module, which includes:

[0296] Obtain EEG data and video data, as well as the emotion categories corresponding to the EEG data and the video data;

[0297] Preprocess the EEG data and video data to obtain EEG features and video features;

[0298] Construct a sample set from the EEG features, video features, and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0299] Use the EEG features and video features in the training set as inputs, and the corresponding emotion categories as outputs to train the classifier, obtaining a preliminarily trained unimodal basic emotion recognition model;

[0300] Input the EEG features and video features in the test set into the preliminarily trained unimodal basic emotion recognition model to obtain an emotion recognition result;

[0301] Judge whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained unimodal basic emotion recognition model is used as the trained unimodal basic emotion recognition model; otherwise, continue to train the preliminarily trained unimodal basic emotion recognition model based on the training set.

[0302] Optionally, it further includes a multimodal training module for:

[0303] Obtain EEG data and video data, as well as the emotion categories corresponding to the EEG data and the video data;

[0304] Preprocess the EEG data and video data to obtain EEG features and video features;

[0305] Fuse the EEG features and video features to obtain multimodal features;

[0306] Construct a sample set from the multimodal features and the corresponding emotion categories; and divide the sample set into a training set and a test set according to a set ratio;

[0307] Use the multimodal features in the training set as inputs, and the corresponding emotion categories as outputs to train the classifier, obtaining a preliminarily trained multimodal basic emotion recognition model;

[0308] Input the multi-modal features in the test set into the preliminarily trained multi-modal basic emotion recognition model to obtain the emotion recognition result;

[0309] Determine whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained multi-modal basic emotion recognition model is used as the trained multi-modal basic emotion recognition model; otherwise, continue to train the preliminarily trained multi-modal basic emotion recognition model based on the training set.

[0310] Optionally, the specific implementation steps of preprocessing the EEG data and video data in the unimodal training module and the multi-modal training module to obtain EEG features and video features include:

[0311] Fill in the missing data in the EEG data, remove the high-frequency components and artifacts to obtain the artifact-free EEG data, and use the modal decomposition method to extract features from the artifact-free EEG data to obtain EEG features;

[0312] Based on the video data, use the face detection algorithm to detect and locate the face, detect the key points of the face at the key positions, extract the facial expression features based on the key points, and use the facial expression features as the video features.

[0313] Optionally, the specific implementation steps of filling in the missing data in the EEG data in the unimodal training module and the multi-modal training module, and removing the high-frequency components and artifacts to obtain the artifact-free EEG data include:

[0314] Fill in the missing data in the EEG data using the interpolation method; use the re-reference method to change the reference baseline of the EEG signal;

[0315] Remove the high-frequency components in the EEG data through low-pass filtering, extract the slow waves or oscillations of specific frequencies in the EEG data to obtain the EEG data with high-frequency components removed;

[0316] Use the artifact removal technique to remove the artifacts in the EEG data with high-frequency components removed to obtain the artifact-free EEG data.

[0317] Embodiment 3:

[0318] As Figure 16 shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and the data can be called and / or modified when the instructions are executed.

[0319] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for rapid emotion assessment under variable gravity conditions in the above embodiments.

[0320] Embodiment 4

[0321] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device, and of course, can also include the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a method for rapid emotion assessment under variable gravity conditions in the above embodiments can be implemented.

[0322] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0323] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0324] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0325] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0326] The above are only embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.

Claims

1. A method for rapid emotional assessment under variable gravity conditions, characterized in that, Including: Collecting electroencephalogram data and video data in a microgravity environment; Performing data preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features; Substituting the electroencephalogram features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result; Wherein, the basic emotion recognition model includes a unimodal basic emotion recognition model and a multimodal basic emotion recognition model; The unimodal basic emotion recognition model and the multimodal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and electroencephalogram data, and the corresponding emotion categories.

2. The method according to claim 1, wherein The training of the unimodal basic emotion recognition model includes: Obtaining electroencephalogram data and video data, and the emotion categories corresponding to the electroencephalogram data and the video data; Performing preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features; Constructing a sample set from the electroencephalogram features, video features, and the corresponding emotion categories; and dividing the sample set into a training set and a test set according to a set ratio; Using the electroencephalogram features and video features in the training set as inputs and the corresponding emotion categories as outputs to train a classifier to obtain a preliminarily trained unimodal basic emotion recognition model; Inputting the electroencephalogram features and video features in the test set into the preliminarily trained unimodal basic emotion recognition model to obtain an emotion recognition result; Judging whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained unimodal basic emotion recognition model is used as the trained unimodal basic emotion recognition model. Otherwise, the preliminarily trained unimodal basic emotion recognition model is continuously trained based on the training set.

3. The method according to claim 1, wherein The training of the multimodal basic emotion recognition model includes: Obtaining electroencephalogram data and video data, and the emotion categories corresponding to the electroencephalogram data and the video data; Performing preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features; Fusing the electroencephalogram features and video features to obtain multimodal features; Constructing a sample set from the multimodal features and the corresponding emotion categories; and dividing the sample set into a training set and a test set according to a set ratio; Using the multimodal features in the training set as inputs and the corresponding emotion categories as outputs to train a classifier to obtain a preliminarily trained multimodal basic emotion recognition model; Inputting the multimodal features in the test set into the preliminarily trained multimodal basic emotion recognition model to obtain an emotion recognition result; Judging whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained multimodal basic emotion recognition model is used as the trained multimodal basic emotion recognition model. Otherwise, the preliminarily trained multimodal basic emotion recognition model is continuously trained based on the training set.

4. The method according to claim 2 or 3, characterized in that, The performing preprocessing on the electroencephalogram data and video data to obtain electroencephalogram features and video features includes: Fill in the missing data in the EEG data, remove high-frequency components and artifacts to obtain artifact-free EEG data, and use a modal decomposition method to extract features from the artifact-free EEG data to obtain EEG features; Based on the video data, use a face detection algorithm to detect and locate human faces, detect key points at key positions of the human face, extract facial expression features based on the key points, and use the facial expression features as video features.

5. The method according to claim 4, wherein The step of filling in the missing data in the EEG data, removing high-frequency components and artifacts to obtain artifact-free EEG data includes: Fill in the missing data in the EEG data using an interpolation method; use a rereferencing method to change the reference baseline of the EEG signal; Remove high-frequency components in the EEG data through low-pass filtering, extract slow waves or oscillations at specific frequencies in the EEG data to obtain EEG data with high-frequency components removed; Use an artifact removal technique to remove artifacts in the EEG data with high-frequency components removed to obtain artifact-free EEG data.

6. The method according to claim 2 or 3, characterized in that, The classifier includes: a convolutional neural network, a Yolo-v5 network, or a VGG network.

7. An emotional rapid assessment device under variable gravity conditions, characterized in that It includes: An experimental data collection layer for collecting EEG data and video data in a microgravity environment; A data processing layer for preprocessing the EEG data and video data to obtain EEG features and video features; An emotion recognition layer for substituting the EEG features and the video features into a pre-trained basic emotion recognition model to obtain an emotion recognition result; Among them, the basic emotion recognition model includes a unimodal basic emotion recognition model and a multimodal basic emotion recognition model; The unimodal basic emotion recognition model and the multimodal basic emotion recognition model are obtained by training a classifier based on the features extracted from video data and EEG data, and the corresponding emotion categories.

8. The system according to claim 7, wherein It also includes a unimodal training module for: Obtaining EEG data and video data, and the emotion categories corresponding to the EEG data and the video data; Preprocessing the EEG data and video data to obtain EEG features and video features; Constructing a sample set from the EEG features, video features, and the corresponding emotion categories; and dividing the sample set into a training set and a test set according to a set ratio; Using the EEG features and video features in the training set as inputs and the corresponding emotion categories as outputs to train the classifier to obtain a preliminarily trained unimodal basic emotion recognition model; Inputting the EEG features and video features in the test set into the preliminarily trained unimodal basic emotion recognition model to obtain an emotion recognition result; Judging whether the emotion recognition result is consistent with the corresponding emotion category in the test set. If it is consistent, the preliminarily trained unimodal basic emotion recognition model is used as the trained unimodal basic emotion recognition model. Otherwise, continue to train the preliminarily trained unimodal basic emotion recognition model based on the training set.

9. An electronic device, characterized in that, It includes: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a method for rapid emotion assessment under variable gravity conditions as described in any one of claims 1 to 6 is implemented.

10. A readable storage medium, characterized in that, A program for execution is stored thereon, and when the program for execution is executed, a method for rapid emotion assessment under variable gravity conditions as described in any one of claims 1 to 6 is implemented.