A multi-modal emotion recognition system based on global-local spatio-temporal semantic alignment
The multimodal emotion recognition system with global-local spatiotemporal semantic alignment solves the problems of neglecting the diversity of physiological signal data and the semantic correlation characteristics of brain regions in multimodal emotion recognition. It achieves spatiotemporal semantic alignment of physiological signals, thereby improving the accuracy of emotion recognition and the robustness of the model.
Patent Information
- Application Number
- CN202411466424.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-10-21
AI Technical Summary
Existing multimodal emotion recognition methods neglect the frequency, phase integrity, and semantic association characteristics of physiological signals when processing diverse and variable physiological signal data, resulting in information loss and low model robustness, and failing to effectively perform spatiotemporal semantic alignment.
A multimodal emotion recognition system employing global-local spatiotemporal semantic alignment achieves spatiotemporal semantic alignment of EEG and other modal signals through data acquisition and feature extraction, temporal information encoding, local temporal semantic alignment, local-global spatial semantic alignment, and global fusion modules, thereby enhancing the interpretability and robustness of the model.
It effectively solves the problems of inconsistent physiological signal length and modal response delay, optimizes emotion recognition results, and improves the accuracy and interpretability of the model.
Smart Images

Figure CN119538086B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of depression auxiliary diagnosis and treatment, and particularly relates to a multi-modal emotion recognition system based on global-local spatio-temporal semantic alignment. BACKGROUND
[0002] Emotion recognition, as a hot research field of artificial intelligence, aims to automatically identify the emotional state of individuals by collecting and analyzing various data. Although humans can easily perceive emotions in daily communication, it still requires further research and exploration to enable machines to accurately recognize and simulate human emotional states. Similar to humans having five senses, machines acquire data through various sensors, which can be categorized into different modalities. Generally, the modalities that respond to human emotions can be divided into two categories: (1) overt behaviors, such as voice, posture, gestures, facial expressions, etc. (2) physiological signals, including Electroencephalography (EEG), Electrocardiography (ECG), Galvanic Skin Reaction (GSR), Electromyography (EMG), Eye Movements (EM), etc. Overt behaviors are easy to obtain, but can be disguised or concealed, which leads to certain unreliability in reflecting the true emotional state. In contrast, physiological signals directly reflect the activity of the nervous system and are difficult to manipulate and conceal, thus more objectively reflecting the emotional state changes caused by stimuli.
[0003] Among various physiological signals, Electroencephalography (EEG) is a physiological signal that records the electrical activity of the brain cortex. Compared with other physiological signals, EEG provides direct information about brain activity, and the changes in frequency and amplitude can reflect the synchronous and asynchronous activity of neurons, effectively revealing the cognitive and emotional processing of individuals, which is crucial for understanding different cognitive and emotional states and has significant advantages in emotion recognition tasks. With the continuous development of multi-modal learning, multi-modal emotion recognition based on psychophysiological data has become a research hotspot. Multi-modal emotion recognition perceives emotions from multiple perspectives, combining subjective and objective factors, making the features more comprehensive, and the information of each modality complementary. EEG provides direct information about brain activity and can effectively reveal the cognitive and emotional processing of individuals. Eye movement signals provide additional insights into visual attention and emotional responses. Electromyography (EMG), Galvanic Skin Response (GSR), and other physiological signals provide physiological indicators related to autonomic nervous system activity and can also be used to indirectly reflect emotional states. By combining these signals, researchers can more accurately understand the response patterns of individuals to different emotional stimuli, thereby improving the accuracy and reliability of emotion recognition.
[0004] Although the related research on multi-modal representation and fusion provides rich and in-depth theoretical basis, there is still a problem to be solved in the field of multi-modal emotion recognition based on electroencephalogram and other modalities, that is, how to explore the response time delay and brain region semantic correlation characteristics between different emotional modalities from diverse and variable data.
[0005] Firstly, the reaction time and duration of individuals to different emotions vary from person to person, which leads to the diversity and variability of data. Existing methods usually divide the signal into equal-length small pieces and then input them into deep neural networks. However, this approach ignores the basic characteristics of physiological signals as a special carrier carrying valuable information, such as the integrity of frequency, phase and the time fluctuation of electroencephalogram components, which may lead to the loss of information. In addition, the emotion recognition task relies on understanding and interpreting the whole context, and this approach may destroy the coherence of the context. Therefore, the diversity and variability of such data need to be fully considered.
[0006] Secondly, there is a response time delay between different emotional modalities, and different physiological signals at the same time may express completely different emotional semantic information, which makes not all time periods in the whole stimulus source effective. After receiving the stimulus, there is a certain information transmission lag between different body parts, which can be explained by the temporal alignment relationship between emotional semantic information. Unfortunately, current research generally ignores this phenomenon, resulting in low robustness and interpretability of multi-modal models. Therefore, it is particularly necessary to explore the activation sequence relationship of multi-modal signals in time and align the semantic information in the time dimension.
[0007] Finally, different regions of the brain have their own unique functions, and the semantic correlation characteristics of electroencephalogram signals in each brain region and other physiological signals and their importance are different. In emotion recognition, brain regions related to emotions and other physiological signals play a crucial role, which can be understood as semantic alignment in space. Although some studies have explored from the perspective of brain region space, they only focus on a single modality of electroencephalogram signals, resulting in redundant information in brain region space, which will have a negative impact on emotion recognition results. Therefore, it is also crucial to align semantic information in the spatial dimension to enhance the semantic information of important brain regions.
[0008] In summary, the present invention aims to propose a multi-modal emotion recognition model that aligns multi-modal signals in time and space to solve the above scientific problems. SUMMARY
[0009] Therefore, the purpose of the present invention is to provide a multi-modal emotion recognition system based on global-local spatiotemporal semantic alignment.
[0010] A multimodal emotion recognition system based on global-local spatiotemporal semantic alignment, comprising:
[0011] (1) Data acquisition and feature extraction unit, including:
[0012] 1) A data acquisition module, configured to acquire EEG data and at least one other modality of physiological signals, referred to as modality b physiological signals;
[0013] 2) Feature extraction module, used to extract features from EEG data and physiological signals of modality b to obtain EEG feature data X EEG and the characteristic data X of mode b B ;
[0014] 3) A brain region division module, which is used to divide the cerebral cortex into multiple specific regions and divide the EEG feature data according to the specific regions;
[0015] The global-local multimodal spatiotemporal semantic alignment unit includes:
[0016] 1) Time information encoding module, used to convert EEG feature data X EEG and feature data X B Input 1D-CNN network to extract deep features and obtain feature X′ EEG and X′ B , the feature X′ B Do R sublinear mapping to get ″′ with R subsets
[0017] Feature X B , where R represents the number of brain regions; EEG and X B Add the respective position encoding information to obtain the deep feature sequence with position information. and
[0018] 2) Local temporal semantic alignment module, used to:
[0019] Use multi-brain cross-modal attention mechanism to analyze deep feature sequences and Obtain the representation of the two queries separately through linear transformation and Key representation and And the value representation and Brain regions correspond to different attention heads;
[0020] Next, we calculate the cross-modal attention of multiple brain regions and obtain the attention weights in two directions as follows:
[0021]
[0022] where M B denotes the mask of modality b; M EEG denotes the mask of EEG data; the subscripts B→EEG and EEG→B represent different attention directions;
[0023] To only keep the correct alignment information, the alignment matrix of the rth brain region in two directions is obtained and
[0024]
[0025] where, and denote the elements of the i-th row and j-th column in the alignment matrix A and denote the elements of the i-th row and j-th column in the matrix A
[0026] The output of each brain region of the two modalities is calculated using the alignment matrix, i.e., the weighted sum of values:
[0027]
[0028] where r represents the rth brain region, r∈(1,R);
[0029] Next, the deep features of the target modality are added to the output of the linear layer through the residual connection operation, so that the alignment of each direction contains the data of the two modalities; finally, one-dimensional global average pooling is used to summarize in the time dimension, and the obtained result is the one-way alignment representation of each brain region of the two modalities, which is defined as follows:
[0030]
[0031] where GAP() represents one-dimensional global average pooling;
[0032] Finally, all brain regions are summarized to obtain the one-way time alignment representation of the two modalities, which is represented as follows:
[0033]
[0034] 3) Local-global spatial semantic alignment module, used for:
[0035] A, first, learn the dynamic attention weight of each brain region: the model learns the local time alignment relationship between the two modalities (EEG modality and modality b), and obtains the one-way time alignment representation and The dynamic attention weights of each brain region are learned by two linear layers. The weight parameters of the first linear layer are The bias parameters are used to integrate the unidirectional time alignment representation and use the tanh function for nonlinear activation; the second linear layer calculates the spatial alignment score, and the weight parameters are The bias parameters are Finally, the spatial alignment weight is obtained by using the Softmax function for normalization:
[0036]
[0037] where the sum of the weights of all brain regions is 1;
[0038] Similarly, a set of spatial alignment weights reflecting modality b is obtained
[0039]
[0040] B. Secondly, integrate the unidirectional time alignment representation: use the learned spatial alignment weight to scale the unidirectional time alignment representation, that is:
[0041]
[0042]
[0043] where the symbol ⊙ represents Hadamard product, and and are spliced to obtain
[0044] C. Then, add the output to the original representation to realize the residual connection, concatenate the output of each brain region in order to form a higher-dimensional representation; finally, process through the linear layer and perform layer normalization operation (LN()), and the final result is the unidirectional space-time alignment representation:
[0045]
[0046] where concat() represents splicing processing; Linear() represents linear layer processing; LN() represents layer normalization operation;
[0047] 4) Global fusion module, used to:
[0048] First, concatenate the unidirectional space-time alignment representation to obtain the fusion representation:
[0049]
[0050] Then, a multi-layer perceptron is used to output the score logits of the current sample belonging to C different emotion categories:
[0051] logits=MLP(O fusion )=(ReLU( fusion W1+b1))W2+b2
[0052] Where MLP stands for multi-layer perceptron; W1 and W2 are the weight parameters of the multi-layer perceptron, b1 and b2 are the bias parameters of the multi-layer perceptron, and the ReLU activation function is used for nonlinear mapping;
[0053] 5) Optimization module: After obtaining the output logits, the Softmax function is further used to convert it into a predicted probability distribution P.
[0054] P = Softmax(logits)
[0055] The sum of the probabilities of all emotion categories is equal to 1;
[0056] Set the loss function and train and optimize the entire multimodal emotion recognition system by minimizing the value of the loss function:
[0057] (3) The emotion recognition unit uses the trained multimodal emotion recognition system to perform emotion recognition on the input data to obtain the corresponding emotion category.
[0058] Furthermore, the data acquisition and feature extraction unit also includes a data preprocessing module for downsampling and filtering the data.
[0059] Preferably, the mask of mode b is expressed as express; Indicates the padding mask used to make the EEG data consistent with the length of modality b data; represents the future mask of modality b data;
[0060] The mask representation of EEG data is express, Indicates the padding mask used to make the EEG data consistent with the length of modality b data. Represents the future mask of EEG data.
[0061] Preferably, the process of constructing the loss function includes: using a unidirectional spatiotemporal alignment representation and The L2 loss between them is used to constrain the collaborative representation of the two modalities, that is, to calculate the square difference of each corresponding element between them and take the average value. The specific calculation is as follows:
[0062]
[0063] wherein Rd represents the dimension of the deep feature;
[0064] The cross-entropy loss and L2 regularization are added again, assuming that the total number of samples is N, and the total optimization target is calculated as:
[0065]
[0066] wherein W is a trainable weight parameter, and lambda is the weight decay of regularization.
[0067] Preferably, the emotion recognition unit further comprises an evaluation test of the trained multi-modal emotion recognition system, wherein leave-one-subject-out cross-validation (LOSO-CV) is selected as the evaluation method.
[0068] Preferably, the modal b data comprises at least one of electroencephalogram, electrocardiogram, and eye movement.
[0069] Preferably, when the feature extraction module performs feature extraction, it includes calculating differential entropy for each frequency band and each lead of the electroencephalogram signal.
[0070] Preferably, when the feature extraction module performs feature extraction, it includes performing z-score normalization processing on the electroencephalogram signal and other physiological signals.
[0071] The present application has the following beneficial effects:
[0072] 1. In the existing multi-modal emotion recognition method, when solving the problem of inconsistent physiological signal length, the physiological signal is usually divided into equal-length small segments, and then input into a deep neural network, which may cause loss of key information and damage to the context coherence. In the present application, after the multi-modal physiological signal is divided into multiple single-modal physiological signals, all single-modal physiological signals are divided into multiple batches, the maximum sequence length of the feature sequence of the single-modal physiological signal segment in each batch is obtained, and the sequence length of the single-modal physiological signal segment in each batch is padded to the maximum sequence length with the maximum sequence length in each batch as the reference value, so that the length of the feature sequence of the single-modal physiological signal segment in each batch is the same, solving the problem of inconsistent physiological signal data length and key information loss.
[0073] 2. The existing multi-modal emotion recognition method generally ignores the response delay between different modalities, so that not all time periods in the entire stimulus source are effective. The present application proposes a local time semantic alignment module to align each modality in the time dimension, enhancing the interpretability and robustness of the model.
[0074] 3、The prior art often ignores the semantic correlation characteristics and importance of the brain signals in each brain region and other physiological signals, resulting in redundant information in the brain region space, which has a negative impact on the emotion recognition result. The present application aligns the semantic information of each modality in the spatial dimension, thereby optimizing the emotion recognition process. BRIEF DESCRIPTION OF DRAWINGS
[0075] Figure 1 A schematic diagram in which 62 electrodes are divided into 16 brain regions in the embodiment of the present application;
[0076] Figure 2 A structural diagram of a local-time semantic alignment module;
[0077] Figure 3 A workflow diagram of the system of the present application;
[0078] Figure 4 A structural diagram of a local-global spatial semantic alignment module;
[0079] Figure 5 A flowchart of a leave-one-cross-validation method. DETAILED DESCRIPTION
[0080] The present application will be described in detail below with reference to the accompanying drawings and embodiments.
[0081] The present application provides a multi-modal emotion recognition system based on global-local spatio-temporal semantic alignment, comprising:
[0082] (1) A data acquisition and feature extraction unit: physiological signals such as electroencephalogram, electrocardiogram, and eye movement of different subjects are collected and preprocessed, and features of multiple single modal physiological signals are extracted respectively;
[0083] (2) A global-local multi-modal spatio-temporal semantic alignment unit: the electroencephalogram features divided into multiple brain regions and other modal features are calculated for local inter-modal temporal semantic alignment, and the activation degree of each brain region is calculated from a global perspective to achieve spatial semantic alignment between modalities, obtaining the representation of global fusion spatio-temporal semantic alignment.
[0084] (3) An emotion recognition unit: after extracting the spatio-temporal semantic features of different data, using leave-one-subject cross-validation (LOSO-CV) as an evaluation method, the effectiveness of the model proposed in the present application in the emotion recognition task is evaluated and its generalization performance is verified.
[0085] (1) The data acquisition and feature extraction unit comprises the following modules:
[0086] 1) Data acquisition module: the types of physiological signals are different, so the acquisition methods are different. The types of physiological signals can be classified in various forms such as collecting device information and interface information. For example, when collecting the EEG information of a target person, an EEG acquisition device is used. The connection interface of the EEG acquisition device and the terminal device can be determined, and the source of the EEG data transmitted through the connection interface can also be determined. Therefore, the multi-modal physiological signals can be divided into multiple single-modal physiological signals through the information of the connection interface.
[0087] 2) Data preprocessing module: the types of physiological signals are different, and the preprocessing methods are different. For example, for EEG signals, the sampling frequency of the signals is reduced from 1000 Hz to 200 Hz to reduce the storage space and computational load, improve the subsequent processing speed, and use a band-pass filter of 1-75 Hz to process the EEG signals, further removing low-frequency baseline drift and high-frequency noise. Then, a notch filter is used to filter the 50 Hz power frequency signal to obtain a more pure EEG signal.
[0088] 3) Feature extraction module: the types of physiological signals are different, and the feature extraction methods are different. For example, for EEG signals, DE is the generalization of Shannon entropy on continuous variables, which is often used to measure the total amount of information of continuous random signals, thereby statistically analyzing the uncertainty of the probability density distribution of continuous signals. The differential entropy is calculated for each frequency band and lead of the EEG signal:
[0089]
[0090] Finally, in order to eliminate the influence of individual differences, the z-score normalization processing is performed on all physiological features of each subject:
[0091]
[0092] where X is the original feature sequence of a subject on a test day, μ is the mean, σ is the standard deviation, and X' is the feature sequence after z-score normalization processing.
[0093] 4) Brain region division module (such as the attached Figure 1):According to the knowledge of the anatomy and function of the cerebral cortex, the cerebral cortex is divided into multiple specific regions, such as frontal lobe, temporal lobe, parietal lobe and occipital lobe, which have unique functions and tasks, such as cognitive function, emotion regulation, perception and motor control, etc. In order to more accurately study the brain activity and function, we use the international standard 10-20 system to arrange the electrodes on the brain, which ensures the uniform distribution of the electrodes on the scalp surface to capture comprehensive information of the brain electrical activity. According to the standard electrode position of the 10-20 system, the brain is divided into 16 specific regions, each region contains a specific set of electroencephalogram electrodes, and the position of these electrodes accurately corresponds to the specific functional area of the brain. The electroencephalogram modality features are divided according to the different 16 brain regions.
[0094] (2) The global-local multi-modal spatio-temporal semantic alignment unit includes the following modules:
[0095] 1) Time information encoding module: In this module, the EEG feature sequence of multiple brain regions and the feature sequence of another modality (hereinafter referred to as modality b, such as eye movement, skin electricity, etc.) are input into the 1D-CNN network after padding to extract deep features, and position encoding information is added. This enables the present application (such as the attached Figure 3 ) to learn the changes of time information in each modality, thereby solving the problem of variable data length. Before input, the original electroencephalogram signal needs to be divided into multiple small subsets according to the brain region, and the modality b signal is also segmented into several segments in time, and the aforementioned features are extracted to generate the feature sequence.
[0096] Assuming that a total of R brain regions are divided, the EEG (EEG) and modality b (B) feature sequences of a certain sample are as follows:
[0097]
[0098]
[0099] wherein the feature sequence of a single brain region is is the length of the EEG feature sequence (the number of segments), is the feature dimension of the current brain region, r≤R. is the length of the modality b feature sequence (the number of segments), d b is the feature dimension of modality b.
[0100] Next, find the maximum sequence length in the batch b where the sample x is located. Fill all samples in batch b with a specific marker PAD to match the maximum sequence length, and get the padding mask
[0101] After padding, the feature sequence of each brain region and the modal b feature sequence are sent into the 1D-CNN network for point-by-point convolution processing to ensure that they have the same feature dimension. Since there is a difference in the feature dimension of each brain region, the 1D-CNN of each brain region is initialized using different parameters:
[0102]
[0103] Therefore, there are the following deep feature sequences of electroencephalogram and modal b:
[0104]
[0105] wherein, is the maximum sequence length in the batch, and d is the deep feature dimension. In order to make the deep feature of modal b have more semantic information and match each brain region in the subsequent temporal semantic alignment, the deep feature sequence X' of modal b is mapped to R groups of data by R times of linear mapping, and the deep feature dimension thereof is set to Rd. B
[0106] After that, the deep electroencephalogram feature sequence of each brain region is spliced together to form a Rd-dimensional feature consistent with the data of modal b. In order to ensure that the sequence can carry time information, the present application also introduces position embedding into the electroencephalogram and modal b feature sequence. The process can be represented as:
[0107]
[0108] Finally, the and are connected to obtain the deep feature with position information as the output of the module.
[0109] 2) Local temporal semantic alignment module: in this module, multi-region cross-modal attention (MRCMA) is used to ensure that the correct alignment information of each brain region is obtained, and the corresponding sparse alignment matrix is generated, specifically:
[0110] A, first, the deep feature is divided back into according to the number of brain regions R. The representations of queries, keys and values are obtained through linear transformation, wherein the linear transformation parameters of all brain regions are not shared, and the brain regions correspond to different attention heads:
[0111]
[0112] where d is the dimension of each head. Each attention head is responsible for learning the attention pattern from the query, key and value of the corresponding brain region, so as to help the model better capture the inter-sequence association information of the local brain region.
[0113] B. Similarly, the application also uses future masks and Ensure that the future time information of the source modality is not used in the dot product calculation. Add the padding mask and the future mask, that is, two direction mask matrices can be obtained, and the calculation is as follows:
[0114]
[0115] C. Next, the multi-brain region cross-modal attention is calculated, and the attention weights in two directions can be obtained, as follows:
[0116]
[0117] Where B→EEG and EEG→B represent different attention directions. By adding the scaled dot product and the mask matrix, after the Softmax function, the masked part in each row, that is, the negative infinity position, becomes zero, thereby preventing it from participating in subsequent calculations. After masking, the attention weights The actual meaningful part in the matrix forms a lower triangular matrix , which means that only the current time and past time information of the source modality is retained in the attention calculation, and the future time information is not included.
[0118] D. In order to only retain the correct alignment information, an adaptive threshold is set on the attention weight of each brain region, and the following two direction alignment matrices and
[0119]
[0120]
[0121] Where, and respectively represent the elements of the i-th row and the j-th column in the alignment matrix the alignment matrix ; and respectively represent the elements of the i-th row and the j-th column in the matrix ;
[0122] The threshold is set as the current row number, so that the threshold size can be adaptively determined according to the sequence length, and only the attention weights higher than the threshold are retained, which represents the thresholded one-way alignment information. After thresholding, the alignment matrix of each brain area is sparse and different.
[0123] E, the output of each brain area is calculated using the alignment matrix, that is, the weighted sum of values:
[0124]
[0125] wherein, Each row vector in the represents a segment in the current brain area electroencephalogram sequence, which is obtained by weighting according to the alignment information of all segments in the modal b sequence. This means that, given a target electroencephalogram segment in the current brain area, the alignment matrix determines which modal b segments are the most relevant and important. On the other hand, Then, from the perspective of modal b, each row vector represents a segment in the modal b sequence, also obtained by weighting according to the alignment information of all electroencephalogram segments in the current brain area. This means that the alignment matrix of the current brain area determines which electroencephalogram segments are the most relevant and important to the given target segment in modal b.
[0126] F, next, the deep features of the target modal are added to the output of the linear layer through the residual connection operation, so that the alignment in each direction contains data of both modalities. Finally, one-dimensional global average pooling (GAP) is used to summarize in the time dimension. The result obtained in this way is the one-way alignment representation of each brain area, defined as follows:
[0127]
[0128] Finally, all brain areas are summarized, and the one-way time alignment representation obtained can be represented as follows:
[0129]
[0130] The whole process of the local-global spatial semantic alignment module can be summarized in the attached Figure 2 .
[0131] 3) Local-global spatial semantic alignment module:
[0132] This module is used to integrate the time alignment information from different brain areas into a unified semantic space, realizing the transition from local to global. The brain area level spatial attention is used on the time alignment representation, learning the dynamic attention weight of each brain area, and through the weighting operation, the model pays more attention to the more important brain area information, realizing the spatial semantic alignment.
[0133] A、Firstly, the module learns the dynamic attention weights of each brain region. The model learns the local temporal alignment relationship between the two modalities (EEG modality and modality b), obtaining a one-way temporal alignment representation and The dynamic attention weights of each brain region are learned through two linear layers. The weight parameters of the first linear layer are The bias parameters are used to integrate the one-way temporal alignment representation and use the tanh function for nonlinear activation. The second linear layer calculates the spatial alignment score, and the weight parameters are The bias parameters are Finally, the spatial alignment weights are obtained by using the Softmax function for normalization:
[0134]
[0135] where the sum of the weights of all brain regions is 1;
[0136] Similarly, a set of spatial alignment weights reflecting modality b can be obtained
[0137]
[0138] B、Secondly, the module integrates the one-way temporal alignment representation. The learned spatial alignment weights are used to scale the one-way temporal alignment representation. Here, the Hadamard product (element-wise multiplication) is used to achieve this, that is:
[0139]
[0140] where the symbol ⊙ represents the Hadamard product, and and are concatenated to obtain
[0141] C、Then, the module adds the output to the original representation to implement residual connection. The module adds the scaled representation to the original representation. This residual connection helps to preserve the original information and makes the model easier to train and converge.
[0142] D、Next, the module concatenates the brain region outputs in order. The module concatenates the outputs of each brain region in order to form a higher-dimensional representation.
[0143] E、Finally, the module processes through a linear layer (Linear()) and performs layer normalization (LN()). The module processes the concatenated representation through a linear layer to abstract and integrate the information of all brain regions. Then, the result is subjected to layer normalization to improve the stability and training effect of the model.
[0144] By the above steps, the final result is the one-way spatio-temporal alignment representation:
[0145]
[0146]
[0147] These representations can reflect the dynamic characteristics of different brain regions and have more rich and complex representation capabilities. This process can be summarized in the attached Figure 4 .
[0148] 4) Global fusion module: After the above three modules, the one-way spatio-temporal alignment representation and has learned the spatio-temporal alignment relationship between the two modalities, this module uses this spatio-temporal alignment information to facilitate emotion recognition.
[0149] First, they are spliced to obtain a fusion representation
[0150]
[0151] Using a multi-layer perceptron to get the output This is the score of the current sample belonging to C different emotion categories. For the SEED-IV dataset, C=4, and for the SEED-V dataset, C=5. This step can be represented as:
[0152] logits=MLP(O fusion )=(ReLU(O fusion W1+b1))W2+b2
[0153] Where W1 and W2 are the weight parameters of the multi-layer perceptron, b1 and b2 are the bias parameters of the multi-layer perceptron, and the ReLU activation function is used for non-linear mapping.
[0154] 5) Optimization module: After obtaining the output logits, further use the Softmax function to convert it to the predicted probability distribution P, where the sum of the probabilities of all emotion categories is equal to 1:
[0155] P=Softmax(logits)
[0156] Using the L2 loss between the one-way spatio-temporal alignment representation and to constrain the collaborative representation of the two modalities, that is, to calculate the square difference of each corresponding element between them and take the average, the specific calculation is as follows:
[0157]
[0158] By adding the cross-entropy loss and L2 regularization, assuming the total number of samples is N, the total optimization objective of the application can be calculated as:
[0159]
[0160] Where W is the trainable weight parameter, and lambda is the weight decay of regularization.
[0161] By minimizing the total loss , the whole multimodal emotion recognition system is trained and optimized:
[0162]
[0163] Where W * is the weight parameter that minimizes the total loss. By the above formula, the proposed optimization objective can be effectively solved, which not only accurately identifies different emotion categories, but also narrows the distance between the two modalities, thereby achieving better spatio-temporal alignment effect.
[0164] (3) The emotion recognition unit includes evaluation test of the trained multimodal emotion recognition system and recognition of the input data, specifically:
[0165] The application selects Leave One Subject Out Cross Validation (LOSO-CV) as the evaluation method. The process is shown in the accompanying Figure 5 .
[0166] In each cycle of the experiment, one subject is selected as the test set, and the remaining subjects are used as the training set to ensure that each subject is used as the test set at least once. The purpose of this strategy is to test whether individual differences will affect model evaluation.
[0167] The average and standard deviation of the test accuracy of all cycles are taken as evaluation indicators. The accuracy of the i-th fold of cross-validation can be expressed as:
[0168]
[0169] Then the expression of the average accuracy and standard deviation is:
[0170]
[0171] Where K is the number of subjects, which is also the number of cross-validation folds.
[0172] The EEG and at least one other physiological signal of the subject are input into the trained multimodal emotion recognition system for emotion recognition.
[0173] To sum up, the above is only the preferred embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A multi-modal emotion recognition system based on global-local spatio-temporal semantic alignment, characterized in that, Comprise: (1) Data acquisition and feature extraction unit, comprising: 1) Data acquisition module, for collecting electroencephalogram data and at least one other modality of physiological signals, denoted as physiological signals of modality b; 2) a feature extraction module for extracting features from the electroencephalogram data and the physiological signals of modality b to obtain electroencephalogram feature data X EEG and the feature data X of modality b B ; 3) Brain region division module, for dividing the cerebral cortex into multiple regions, and dividing the electroencephalogram feature data according to the regions; (2) Global-local multi-modal spatio-temporal semantic alignment unit comprises: 1) a time information encoding module, configured to: input brain electrical feature data X EEG and feature data X B into a 1D-CNN network respectively to extract corresponding features X' EEG and X' B , perform R times linear mapping on the feature X' B to obtain a feature X' B with R subsets, where R represents the number of brain regions; and add position encoding information to the features X' EEG and X' B respectively to obtain deep feature sequences with position information and 2) Local time semantic alignment module, for: Using multi-brain region cross-modal attention mechanism on deep feature sequence and Obtaining representations of queries of both by linear transformation and representation of keys and and representation of values and Brain regions correspond to different attention heads; Next, calculate the multi-brain region cross-modal attention, get the attention weight in two directions, as follows: where M B denotes the mask of modality b; M EEG denotes the mask of electroencephalography data; subscripts B→EEG and EEG→B denote different attention directions; To keep only the correct alignment information, get the alignment matrix of the rth brain region in both directions and wherein and denote the i-th row, j-th column element of the alignment matrix and the alignment matrix respectively. and denote the i-th row, j-th column element of the matrix and the matrix respectively. Use the alignment matrix to calculate the output of each brain region of the two modalities, that is, the weighted sum of values: Wherein, r represents the rth brain region, r∈(1,R); Next, through the residual connection operation, add the deep features of the target modality and the output of the linear layer, so that the alignment of each direction contains the data of the two modalities; Finally, use one-dimensional global average pooling to summarize in the time dimension, and the obtained result is the one-way alignment representation of the two modalities of each brain region, defined as follows: Wherein, GAP() represents one-dimensional global average pooling; Finally, all brain regions are summarized, and the obtained one-way time alignment representation of the two modalities is represented as follows: 3) Local-global spatial semantic alignment module, for: A、First, learn the dynamic attention weights of each brain region: the model learns the local temporal alignment relationship between the two modalities and obtains a unidirectional temporal alignment representation and The dynamic attention weights of each brain region are learned through two linear layers; the weight parameters of the first linear layer are The bias parameters are used to integrate the unidirectional temporal alignment representation and use the tanh function for nonlinear activation; the second linear layer calculates the spatial alignment score, and the weight parameters are The bias parameters are Finally, the spatial alignment weights are obtained by using the Softmax function for normalization: Wherein, the sum of the weights of all brain regions is 1; Similarly, a set of spatial alignment weights reflecting modality b is obtained B、Secondly, integrate the one-way time alignment representation: use the learned spatial alignment weight to scale the one-way time alignment representation, that is: wherein the symbol denotes Hadamard product, and and are concatenated to obtain C、Then, add the output to the original representation to realize the residual connection, concatenate the output of each brain region in order to form a higher-dimensional representation; Finally, through linear layer processing and layer normalization operation, the final result is the one-way spatio-temporal alignment representation: Wherein, concat() represents concatenation processing; Linear() represents linear layer processing; LN() represents layer normalization operation; 4) Global fusion module, for: First, concatenate the one-way spatio-temporal alignment representation to obtain the fusion representation: Then, use a multi-layer perceptron to output the score logits of the current sample belonging to C different emotion categories: logits = MLP(O fusion ) = (ReLU(O fusion W1 + b1)) W2 + b2 Wherein, MLP represents a multi-layer perceptron; W1 and W2 are weight parameters of the multi-layer perceptron, b1 and b2 are bias parameters of the multi-layer perceptron, and ReLU activation function is used for nonlinear mapping; 5) Optimization module: after obtaining the output logits, further convert it to a prediction probability distribution P using the Softmax function P=Softmax(logits) Wherein, the sum of the probabilities of all emotion categories is equal to 1; Set the loss function, minimize the value of the loss function, so as to train and optimize the whole multi-modal emotion recognition system: (3) The emotion recognition unit uses the trained multi-modal emotion recognition system to perform emotion recognition on the input data, and obtains the corresponding emotion category.
2. A global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1, wherein, The data acquisition and feature extraction unit further comprises a data preprocessing module for downsampling and filtering the data.
3. A global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1 wherein, The mask for modality b is represented as is represented as is represented as a padding mask used to make the brain electrical data consistent with the modality b data length; is represented as a future mask for modality b data; Masked representation of electroencephalography data is represented as represents, represents a padding mask used to make the electroencephalography data consistent with the length of the modality b data, represents a future mask of the electroencephalography data.
4. The global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1, wherein, The process of constructing the loss function includes: using the one-way spatio-temporal alignment representation and L2 loss between them to constrain the collaborative representation of the two modalities, that is, to calculate the square difference of each corresponding element between them and take the average, which is calculated as follows: Wherein, Rd represents the dimension of the deep feature; Then add the cross-entropy loss and L2 regularization, assuming that the total number of samples is N, and the total optimization target is calculated as: Wherein, W is a trainable weight parameter, and λ is a weight decay of regularization.
5. The global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1, wherein, The emotion recognition unit further comprises an evaluation test of the trained multi-modal emotion recognition system, wherein a leave-one-subject-out cross-validation (LOSO-CV) is selected as the evaluation method.
6. A global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1 wherein, The modal b data comprises at least one of electroencephalogram, electrocardiogram and eye movement.
7. A global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1 wherein, When the feature extraction module performs feature extraction, the feature extraction module comprises calculating differential entropy for each frequency band and each lead of the electroencephalogram signal.
8. A global-local spatio-temporal semantic alignment based multi-modal emotion recognition system as claimed in claim 1 wherein, When the feature extraction module performs feature extraction, the feature extraction module comprises performing z-score normalization processing on the electroencephalogram signal and other physiological signals.
Citation Information
Patent Citations
Multi-modal physiological signal semantic alignment method and system for emotion recognition
CN116561634A
Depression intensity identification method based on spatial-temporal feature integration and global-local feature fusion
CN118506421A