Easy-to-acquire physiological signal-based emotion recognition method
Through the ACST-Net model of multi-scale convolution and multi-head self-attention mechanism combined with knowledge distillation technology, the problem of easy-to-acquire physiological signal recognition is solved, high-precision emotion recognition is achieved, and the application of technology in wearable devices and mental health monitoring is promoted.
Patent Information
- Application Number
- CN202510750906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-11
AI Technical Summary
The existing emotional recognition technology based on electroencephalogram signals (EEG) is costly and complicated to collect, while the recognition accuracy based on easy-to-acquire physiological signals (such as GSR, ECG, etc.) is low, and there is a lack of high-precision recognition methods.
The emotion recognition method based on easy-to-acquire physiological signals is adopted, and features are extracted through ACST-Net model with multi-scale convolution and multi-head self-attention mechanism, and combined with knowledge distillation technology, emotional knowledge is transmitted from EEG signals to improve recognition accuracy.
It improves the emotion recognition accuracy of easy access to physiological signals, promotes the practical application of this technology in scenarios such as wearable devices and mental health monitoring, and enhances the universality and wide applicability of the technology.
Smart Images

Figure CN120296568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotion analysis, and particularly to a method for emotion recognition based on easily obtainable physiological signals. Background Art
[0002] Existing emotion recognition technologies are mostly based on electroencephalogram (EEG) signals. However, their high acquisition cost and complex process limit their practical applications. Although easily obtainable physiological signals (such as GSR, ECG, etc.) can be collected simply, the recognition accuracy is relatively low. Therefore, there is an urgent need for a method for emotion recognition of easily obtainable physiological signals with high accuracy. Summary of the Invention
[0003] The purpose of the present invention is to solve the above problems, and a method for emotion recognition based on easily obtainable physiological signals is designed.
[0004] To achieve the above purpose, the technical solution of the present invention is a method for emotion recognition based on easily obtainable physiological signals, including the following steps:
[0005] Preprocess the original physiological signals to obtain preprocessed physiological signal data;
[0006] Parallelly extract temporal features using three convolutional kernels with different temporal kernel sizes, and concatenate them in the feature dimension;
[0007] Add positional information, emotion cues, and category information to the temporal features;
[0008] Use a parallel spatio-temporal attention mechanism to extract the most discriminative temporal features and channel features at different time points and channels;
[0009] Use a soft attention mechanism to fuse the temporal features and channel features;
[0010] Input the fused features into a classifier composed of a linear layer and a ReLU layer for emotion recognition.
[0011] The lengths of the convolutional kernels of the three different temporal kernel sizes are set to specific ratios of the physiological signal sampling rate, and the expression is ;
[0012] , where represents the th convolution, is the ratio,
[0013] The extraction of the temporal features includes:
[0014] Perform convolution and non-linear activation operations on each sample;
[0015] Use average pooling to reduce the dimensionality of the features;
[0016] The features extracted by different convolution kernels are concatenated in the feature dimension and normalized through the BatchNorm layer.
[0017] The parallel spatiotemporal attention mechanism includes a temporal attention block and a channel attention block:
[0018] The temporal attention block calculates the correlation between different temporal features through a multi-head self-attention mechanism;
[0019] The channel attention block calculates the correlation between different channel features through a multi-head self-attention mechanism.
[0020] The specific steps of the soft attention mechanism fusing time features and channel features include:
[0021] Normalize and linearly embed the temporal features and channel features respectively;
[0022] Calculate the weight score for each feature;
[0023] The features are weightedly fused according to the weight scores.
[0024] A method for easily accessible physiological signal emotion recognition based on knowledge distillation, comprising the following steps:
[0025] Preprocessing the original physiological signals;
[0026] The ACST-Net model is used to extract the features of EEG signals and easily accessible physiological signals;
[0027] The heterogeneous features of EEG signals and easily accessible physiological signals are projected into a common space through a linear layer, and the InfoNCE loss function is used to achieve knowledge distillation at the feature level.
[0028] The heterogeneous features of EEG signals and easily accessible physiological signals are fed into the softmax layer to generate probability distribution, and the KL loss function is used to achieve knowledge distillation at the label level.
[0029] The features of easily accessible physiological signals are fed into the classifier for emotion classification, and the cross entropy loss function is used to guide the classification;
[0030] Combine the loss functions at the feature level, label level, and classification level to generate the final loss function and perform multiple rounds of training;
[0031] The trained feature extraction model is used to realize emotion recognition based on easily accessible physiological signals.
[0032] The feature-level knowledge distillation shortens the distance between positive sample pairs and increases the distance between negative sample pairs through contrastive learning. The specific loss function is:
[0033] ;
[0034] Among them and represent the feature extraction and projection of the teacher model and the student model is a temperature parameter used to control the sharpness or smoothness of the distribution.
[0035] The knowledge distillation at the label level calculates the difference between the output probability distributions of the teacher model and the student model through KL divergence. The specific loss function is:
[0036] ;
[0037] The final loss function is .
[0038] A high-precision emotion recognition system based on easily obtainable physiological signals, comprising:
[0039] A signal preprocessing module for preprocessing the original physiological signals;
[0040] A feature extraction module for extracting multi-scale temporal features and channel features using the ACST-Net model;
[0041] An attention mechanism module for extracting discriminative features through a parallel spatio-temporal attention mechanism;
[0042] A feature fusion module for fusing temporal and channel features through a soft attention mechanism;
[0043] A classification module for inputting the fused features into a classifier for emotion recognition;
[0044] A knowledge distillation module for improving the emotion recognition accuracy of easily obtainable physiological signals through knowledge distillation at the feature level and the label level.
[0045] The method for emotion recognition based on easily obtainable physiological signals fabricated using the technical solution of the present invention transfers the rich emotional knowledge of EEG signals to modalities with poor performance by using the method of knowledge distillation, improves the emotion recognition accuracy of easily obtainable physiological signals, overcomes the disadvantages of high cost and low accuracy in emotion recognition based on physiological signals in the past, and provides effective support for the practical application of emotion recognition technology based on physiological signals. Description of the Drawings
[0046] Figure 1 is a flowchart of the cross-modal distillation EEG-EOPS-DM method based on easily obtainable physiological signals according to the present invention.
[0047] Figure 2It is the structure diagram of the ACST-Net model of the method for emotion recognition based on easily obtainable physiological signals according to the present invention.
[0048] Figure 3 It is the ablation experiment diagram of the ACST-Net of the method for emotion recognition based on easily obtainable physiological signals according to the present invention.
[0049] Figure 4 It is the comparison diagram of the number of convolutional kernels of the ACST-Net of the method for emotion recognition based on easily obtainable physiological signals according to the present invention on the EEG signals of the DEAP dataset.
[0050] Figure 5 It is the comparison diagram of different feature fusion methods of the ACST-Net of the method for emotion recognition based on easily obtainable physiological signals according to the present invention on the EEG signals of the DEAP dataset.
[0051] Figure 6 It is the framework diagram of the EEG-EOPS-DM method for emotion recognition based on easily obtainable physiological signals according to the present invention.
[0052] Figure 7 It is the ablation experiment diagram of the distillation method for emotion recognition based on easily obtainable physiological signals according to the present invention on the GSR signals of the AMIGOS dataset.
[0053] Figure 8 It is the step diagram of the method for emotion recognition based on easily obtainable physiological signals according to the present invention. Detailed implementation manners
[0054] The present invention will be specifically described below with reference to the accompanying drawings. As Figure 1-8 shown, the method for emotion recognition based on easily obtainable physiological signals includes the following steps:
[0055] Preprocess the original physiological signals to obtain preprocessed physiological signal data;
[0056] Use three convolutional layers with different-sized temporal kernels to extract temporal features in parallel and concatenate them in the feature dimension;
[0057] Add positional information, emotional cues, and category information to the temporal features;
[0058] Use a parallel spatio-temporal attention mechanism to extract the most discriminative temporal and channel features at different time points and channels;
[0059] Use a soft attention mechanism to fuse the temporal and channel features;
[0060] Input the fused features into a classifier composed of a linear layer and a ReLU layer for emotion recognition.
[0061] The convolution kernel lengths of the three different-sized temporal kernels are set to a specific ratio of the physiological signal sampling rate, and the expression is ;
[0062] , where represents the nth convolution, is the ratio, is the sampling rate.
[0063] The extraction of the temporal features includes:
[0064] Performing convolution and non-linear activation operations on each sample;
[0065] Using average pooling to reduce the dimensionality of the features;
[0066] Concatenating the features extracted by different convolution kernels in the feature dimension and normalizing through the BatchNorm layer.
[0067] The parallel spatio-temporal attention mechanism includes a temporal attention block and a channel attention block:
[0068] The temporal attention block calculates the correlation between different temporal features through the multi-head self-attention mechanism;
[0069] The channel attention block calculates the correlation between different channel features through the multi-head self-attention mechanism.
[0070] The specific steps for the soft attention mechanism to fuse temporal features and channel features include:
[0071] Normalizing and linearly embedding the temporal features and channel features respectively;
[0072] Calculating the weight scores for each feature;
[0073] Weightedly fusing the features according to the weight scores.
[0074] A method for easily accessible physiological signal emotion recognition based on knowledge distillation, comprising the following steps:
[0075] Preprocessing the original physiological signal;
[0076] Using the ACST-Net model to extract the features of electroencephalogram signals and easily accessible physiological signals;
[0077] Projecting the heterogeneous features of EEG signals and easily accessible physiological signals into a common space through a linear layer, and using the InfoNCE loss function to achieve knowledge distillation at the feature level;
[0078] The heterogeneous features of EEG signals and easily accessible physiological signals are fed into the softmax layer to generate probability distribution, and the KL loss function is used to achieve knowledge distillation at the label level.
[0079] The features of easily accessible physiological signals are fed into the classifier for emotion classification, and the cross entropy loss function is used to guide the classification;
[0080] Combine the loss functions at the feature level, label level, and classification level to generate the final loss function and perform multiple rounds of training;
[0081] The trained feature extraction model is used to realize emotion recognition based on easily accessible physiological signals.
[0082] The feature-level knowledge distillation shortens the distance between positive sample pairs and increases the distance between negative sample pairs through contrastive learning. The specific loss function is:
[0083] ;
[0084] in and represents the feature extraction and projection of the teacher model and the student model, is a temperature parameter that controls how sharp or smooth the distribution is.
[0085] The knowledge distillation at the label level calculates the difference in probability distribution output by the teacher model and the student model through KL divergence. The specific loss function is:
[0086] ;
[0087] The final loss function is .
[0088] A high-precision emotion recognition system based on easily accessible physiological signals, comprising:
[0089] A signal preprocessing module, used for preprocessing the original physiological signal;
[0090] Feature extraction module, used to extract multi-scale temporal features and channel features using the ACST-Net model;
[0091] An attention mechanism module for extracting discriminative features through a parallel spatiotemporal attention mechanism;
[0092] Feature fusion module, used to fuse temporal and channel features through soft attention mechanism;
[0093] The classification module is used to input the fused features into the classifier for emotion recognition;
[0094] A knowledge distillation module for improving the emotion recognition accuracy of easily accessible physiological signals through knowledge distillation at the feature level and label level.
[0095] The characteristics of this implementation are as follows: A combined multi-scale convolution and multi-head self-attention mechanism is proposed to design an ACST-Net model suitable for emotion recognition of easily accessible physiological signals. This model not only performs well on easily accessible physiological signals and EEG signals, but also combines the proposed cross-modal knowledge distillation framework EEG-EOPS-DM to achieve knowledge transfer at the feature and label levels, fully improving the performance of easily accessible physiological signals. Experimental results show that the method in this paper has achieved good performance on two widely recognized physiological signal datasets, DEAP and AMIGOS. Compared with the advanced deep learning method TimeMixer, the ACST-Net model has superior performance. Finally, the research in this paper makes up for the lack of research in the field of emotion recognition based on easily accessible physiological signals. The proposed method improves the universality and wide applicability of emotion recognition technology based on physiological signals. The method proposed in this paper not only promotes the development of emotion recognition technology based on physiological signals, but also promotes the process of bringing high-performance emotion recognition technology from the laboratory to real life, further accelerating the implementation and application of emotion recognition technology in scenarios such as wearable devices and mental health monitoring.
[0096] In this implementation, the architecture of the emotion recognition model ACST-Net is as Figure 2 shown. ACST-Net includes parallel convolutions for extracting fine-grained temporal features, two multi-head attention mechanisms that focus on the dependencies between temporal features and between channels, followed by a module that fuses the features passing through the two attention mechanisms, and finally a classifier for the ACST-Net model.
[0097] Given n data samples of a physiological signal , where represents the corresponding emotion label. For each , the dimension of each sample is , where C represents the number of electrode channels and T represents the number of time points.
[0098] Temporal Feature: In the ACST-Net model, three different temporal kernels are set to extract features of different granularities from physiological signal data in parallel, and then the features of different scales are integrated. Specifically, the ACST-Net model chooses to concatenate in the temporal dimension for input into the subsequent attention module. This parallel multi-scale convolution operation enables the network to capture details and abstract features at different levels. Another reason for choosing the feature extraction method of multi-scale parallel convolution is to fully extract the temporal features of easily obtainable physiological signals with fewer input channels to ensure an adequate number of features. Referring to EEGNet and TSception , the ACST-Net model sets the length of the convolution kernel to a specific ratio of the physiological signal sampling rate . These ratios are defined as , where represents the nth convolution. The convolution kernel expression is as follows:
[0099] ;
[0100] From the frequency perspective, EEGNet and TSception point out that different convolution kernel sizes can learn rich frequency features related to emotions. Longer temporal kernels can learn different representations of low frequencies, while short kernels can extract high-frequency representations. From the temporal perspective, using convolution to capture multi-scale temporal features along the temporal dimension, time series of different scales exhibit different advantages, where fine scales mainly focus on detailed fluctuations and coarse scales highlight macroscopic changes . The multi-scale view can essentially unravel the complex changes of multiple components, thus facilitating modeling based on temporal changes . In summary, temporal receptive fields of different sizes can extract signal change features within different temporal scales, thereby generating more diverse features and capturing more emotion-related information. represents the output of the th convolutional layer. For each sample , after convolution, a non-linearity is applied to the calculation. In the ACST-Net model, the activation function is used, and average pooling is used to perform dimensionality reduction on the features. As follows:
[0101] ;
[0102] where represents the convolution kernel, is the input physiological signal sample. is the average pooling operation, is the convolution operation, the convolution kernel size is , the stride is 1×1, denote The activation function. The input of the two-dimensional convolution in the ACST-Net model is , in order to enrich the features, the output channels of the convolution are set to 6, and the convolution output is . The final output is adjusted to .
[0103] Finally, the features extracted by different convolutional kernels are concatenated in the feature dimension. The final temporal feature The expression is as follows:
[0104] ;
[0105] is layer, which is used to normalize variables to reduce the internal covariate shift problem in the neural network. At the same time, it enhances the stability of the training process. denotes concatenation in the feature dimension.
[0106] Due to the heterogeneity between different physiological signals, the ACST-Net model uses different temporal kernels and pooling sizes for different physiological signals.
[0107] Attention block: In order to extract the relationship between each temporal feature and channel in the easily accessible physiological signal data, the ACST-Net model combines the features obtained in the previous step with the self-attention mechanism. The attention block consists of two different attention mechanisms: one focuses on temporal dynamics and the other focuses on information specific to a particular channel.
[0108] In the temporal attention block, the different-scale temporal features extracted by three parallel convolutions are input into the temporal feature attention mechanism. By calculating the correlation between different features, the local features highly correlated with the emotional state are found, and a higher weight is assigned to this feature. Although the concatenated temporal features may disrupt the temporal continuity, they provide more comprehensive scale features, enabling the ACST-Net model to no longer be limited to the local temporal features of a single convolution. Under different emotional states, the model's focus can be concentrated on the most representative features at different time scales.
[0109] Before using the attention mechanism, the ACST-Net model first uses a learnable class token as a prefix, whose main function is to aggregate information from the entire sequence and then be used for emotion classification. In order to make the model adapt to all physiological signals for emotion recognition, an emotion prompt is added, inspired by MAET . At the same time, in order to fuse positional information, a learnable positional embedding is added, and the final expression is as follows:
[0110] ;
[0111] Among them . The ACST-Net model uses the multi-head self-attention mechanism in Transformer to transform the input into Key ( ), Query ( ), and Value ( ) through three linear layers. The self-attention can calculate the expression as follows:
[0112] ;
[0113] One attention head is used in the ACST-Net model for calculation. To ensure the stability and effectiveness of training, residual connections are used.
[0114] The final temporal feature The expression is as follows:
[0115] ;
[0116] Among them represents the transformation matrix, and represents the layer normalization operation.
[0117] In the channel attention block, due to the certain dependence relationship between different channels of physiological signals, especially EEG signals, different brain activation regions correspond to different emotional states. Here, similar to the temporal attention, the multi-head self-attention mechanism in Transformer is used, and the attention mechanism is used to help find out which channels are more related to the emotional state. The input is adjusted to , and other operations are the same as those in the temporal attention block. Finally, the channel attention feature is obtained.
[0118] Before performing the classification task finally, the temporal attention feature and the channel attention feature should be fused.
[0119] Fusion classification layer: In order to fuse the temporal attention feature and the channel attention feature and make them complementary according to different contribution degrees. The ACST-Net model is adjusted according to the modality attention fusion strategy and the attention-based hybrid spatio-temporal feature fusion strategy . First, the two attention features are normalized, then the two features are embedded using three linear layers, and then the soft attention mechanism is used to calculate the relevant weights of each feature embedding. The weight scores of the features The calculation process is as follows:
[0120] ;
[0121] Where and are learnable parameters, and represent the time attention features embedded by the linear layer and the channel attention features , represents the concatenation operation.
[0122] Then, the two features are weighted and fused using the weights of the features. The expression is as follows:
[0123] ;
[0124] Where and are the feature weights corresponding to the time attention features and the channel attention features , represents the concatenation operation, is the layer normalization process.
[0125] Finally, the classifier consists of two fully connected layers and layers. The final expression is as follows:
[0126] ;
[0127] Where represents the fully connected layer, represents the activation function.
[0128] The DEAP dataset and the Amigos dataset are used to verify the superiority of the proposed method.
[0129] The DEAP dataset is a multimodal dataset for analyzing human emotional states. It records the physiological signals of 32 subjects during the process of watching 40 one-minute-long music videos, including 32-channel EEG signals and 8-channel peripheral physiological signals. In this paper, the 8-channel peripheral physiological signals are used as easily accessible physiological signals, including EOG, EMG, GSR, Respiration belt, BVP, and Temperature. Each recording lasts for 1 minute, and 3 seconds of baseline signals are recorded before watching the music video. The subjects rate each video from 1 to 9 according to four emotional dimensions, including arousal, valence, liking, and dominance.
[0130] The Amigos dataset is a dataset for studying emotion recognition in different social environments. The experimental phase is divided into two parts: watching short videos in a personal solitude environment and watching long videos in a group environment to record physiological signals. In the experiments in this paper, the data collected in the personal environment was used. After watching the videos, the subjects rated arousal, valence, likeability, and dominance on a scale of 1 to 9. 40 subjects watched 16 short videos to record 14-channel EEG signals and 3-channel peripheral physiological signals. Peripheral physiological signals, as easily accessible physiological signals, include ECG and GSR. Each experiment included 5 seconds of baseline signals.
[0131] In the experiments in this paper, the DEAP and Amigos datasets used a threshold of 5 to distinguish high and low ratings for the four emotional dimensions. All signals were downsampled to 128 Hz and segmented using non-overlapping 1-s time windows.
[0132] Data preprocessing: In this paper, a preprocessing method of baseline filtering was performed on all physiological signal data. The formula is as follows:
[0133] ;
[0134] where L is the number of time windows of the baseline signal.
[0135] Experimental settings: The experiments were conducted using the Pytorch framework and on an NVDIA 3080Ti GPU device. To ensure fair comparison, all methods used the same data processing and experimental settings. For each classification task, ten-fold cross-validation was used in this chapter, and then the final average results were reported. The Adam optimizer was selected with a learning rate of 0.001 and a batch size of 128.
[0136] Emotion Recognition Performance on Easily Obtainable Physiological Signals: The ACST-Net model is a general emotion recognition model for easily obtainable physiological signals. The easily obtainable physiological signals in the DEAP dataset include EOG, EMG, GSR, Respiration belt, BVP, and Temperature. Additionally, the easily obtainable physiological signals in the AMIGOS dataset, ECG and GSR, were tested on all subjects (across subjects). This section presents the average accuracy of emotion classification in four dimensions: valence, arousal, dominance, and liking. On the DEAP dataset, the recognition accuracy of EOG and EMG is close to 1. Except for GSR and Temperature, the accuracy of the ACST-Net model is higher than 90%. The results are shown in Table 3.1. On the AMIGOS dataset, the accuracy of ECG is close to 1, and the accuracy of GSR is close to 90%. The results are shown in Table 3.2. Since the existing research on emotion recognition based on easily obtainable physiological signals is an under-explored research area, this paper uses the attention-based temporal convolutional network ATCNet model as a comparison method. The ATCNet model combines temporal convolution and multi-head attention mechanisms and mainly focuses on the temporal features of physiological signals. This paper also uses the TimeMixer, a time series prediction classification model with decomposable multi-scale fusion as a comparison method. This model decouples the information of multi-scale time series to achieve excellent performance and efficiency in long-term and short-term classification tasks. At the same time, this model has achieved state-of-the-art performance in multiple long-term and short-term prediction tasks and demonstrated excellent efficiency in all experiments.
[0137] The ACST-Net model is 10.74% higher than the baseline model TimeMixer in the four emotion dimensions of the DEAP dataset. It is 3.74% higher than the baseline model TimeMixer in the four emotion dimensions of the AMIGOS dataset.
[0138] From the above results, it can be seen that for the generalization of heterogeneous modalities and the emotion recognition performance of easily obtainable physiological signals, the proposed ACST-Net model in this paper is a more robust and powerful model.
[0139] Table 3.1 Average Precision (Acc%) of the ACST-Net Model on the Easily Obtainable Physiological Signals of All Subjects in the DEAP Dataset:
[0140]
[0141] Continued Table 3.1:
[0142]
[0143] Table 3.2 Average accuracy (Acc%) of the ACST-Net model on easily accessible physiological signals of all subjects in the AMIGOS dataset:
[0144]
[0145] Continued Table 3.2:
[0146]
[0147] Emotion recognition performance on EEG signals: In this paper, the effectiveness of the ACST-Net model was further verified on EEG signals, and a performance comparison was made with previous methods. The average accuracy of emotion classification on EEG signals of the DEAP dataset and the AMIGOS dataset in terms of valence, arousal, dominance, and liking dimensions is reported in this section. As shown in Table 3.3, the ACST-Net model outperforms other methods on the EEG signals of the DEAP dataset and the AMIGOS dataset. Specifically, in the subject-dependent results on the DEAP dataset, the performance of the ACST-Net model on the four emotion dimensions is on average 0.785% higher than that of the existing optimal ASTDF-Net work. In the cross-subject case, the ACST-Net model also achieved a high accuracy, and the results are shown in Table 3.4. It is on average 3.14% higher than MEEG-Transformer on two emotion dimensions, and the accuracy in the cross-subject case is also close to 1, demonstrating the strong generalization ability of the ACST-Net model.
[0148] On the AMIGOS dataset, the subject-dependent performance can also be close to 1, and the results are shown in Table 3.5. Although the cross-subject performance is not close to 1, it is all higher than 97%. The results are shown in Table 3.6. Due to the different acquisition environments and the number of EEG channels of the DEAP dataset and the AMIGOS dataset, there is a slight difference in the results. However, the ACST-Net model shows strong generalization ability on the EEG data of both datasets.
[0149] Table 3.3 Average accuracy and standard deviation (Acc±Std%) of the ACST-Net model on EEG signals of the DEAP dataset subject-dependent:
[0150]
[0151] Table 3.4 Average accuracy (Acc%) of the ACST-Net model on EEG signals of the DEAP dataset cross-subject:
[0152]
[0153] Table 3.5 Mean accuracy and standard deviation (Acc±Std%) of the ACST-Net model on EEG signals of the AMIGOS dataset, dependent on the subject:
[0154]
[0155] Table 3.6 Mean accuracy (Acc%) of the ACST-Net model across subjects on EEG signals of the AMIGOS dataset:
[0156]
[0157] To verify the effectiveness of each part of the ACST-Net model, since some physiological signals in easily accessible physiological signals have only one channel and cannot demonstrate the role of channel attention of the ACST-Net model. Therefore, an ablation study was conducted on the EEG signals with sufficient number of channels in the DEAP dataset, and the comparison results are as Figure 3 shown. Specifically, several different models were designed in this paper.
[0158] Temporal and channel attention: To study the impact of temporal and channel attention on the model performance, three independent models were trained in this paper: without both temporal and channel attention, without temporal attention, and without channel attention. As can be seen from the figure such as Figure 3 shown, the accuracy decreased significantly when both temporal and channel attention were removed simultaneously. The average decrease was 2.8%, 2.72%, 2.7%, and 2.73% in the four emotional dimensions of valence, arousal, dominance, and liking. When removing channel and temporal attention separately, it was found that removing temporal attention decreased by an average of 1.16% in the four emotional dimensions, and removing channel attention decreased by an average of 1.03%. These findings indicate that in the emotion recognition task based on physiological signals, temporal information plays a more important role than channel information and spatial information.
[0159] Fusion method: To evaluate the effectiveness of the fusion method of the ACST-Net model, the fusion method of the ACST-Net model was compared with a simple concatenation operation in this paper.
[0160] The weight-based attention fusion method, which assigns weights according to the importance of temporal and channel information, improved by 2.88%, 2.11%, 5.75%, and 2.31% respectively in the four emotional dimensions of valence, arousal, dominance, and liking compared with the simple concatenation operation. These comparisons demonstrated the superiority of the attention-based feature fusion method in the emotion recognition task based on physiological signals.
[0161] Number of convolutional kernels: In this paper, the impact of the number of parallel convolutional kernels on performance was evaluated, and the number of convolutional kernels was discussed. In this paper, the ACST-Net model with 1, 2, and 4 convolutional kernels was trained on the EEG signals of all subjects in the DEAP dataset. The results are as Figure 4 shown. The ACST-Net model showed the highest accuracy when performing three parallel convolutions. When setting 1, 2, and 4 parallel convolutions to extract time features, the accuracy was always lower than 99%. There were slight overfitting or insufficient time feature extraction drawbacks when setting more or fewer convolutional kernels.
[0162] Comparative analysis of fusion methods: When designing the model in this paper, the contribution rates of time features and channel features were assumed to be different, and the situation where the contribution rates were assumed to be the same was not considered. Therefore, a comparison was made on the EEG signals of all subjects in the DEAP dataset using the feature concatenation method Concatenate and the method of directly adding two features. The results are as Figure 5 shown. It was found from the results that if the contribution rates of time and channel features were assumed to be the same, the results of the ACST-Net model decreased by 3.26% and 2.84% on average.
[0163] Therefore, as a time series signal, physiological signals contain more time information than channel information. Therefore, the time and channel feature fusion method based on attention weights in the ACST-Net model is the preferred solution.
[0164] The designed cross-modal knowledge distillation method transfers the rich emotional knowledge in EEG to physiological signals such as GSR and Temperature with poor emotion recognition performance. The distillation structure is as Figure 6 shown. The distillation method includes knowledge transfer at both the feature and label levels. This method promotes the application of emotion recognition using physiological signals in actual scenarios.
[0165] To improve the emotion recognition performance of physiological signals such as GSR and Temperature with poor performance, based on the ACST-Net model as a feature extraction method, a knowledge distillation method was used to transfer the rich emotional knowledge of EEG signals to modalities with poor performance. A cross-modal knowledge distillation method was proposed, and knowledge distillation was performed at the feature and label levels.
[0166] Feature-level knowledge distillation: At the feature level, a feature distillation method based on contrastive learning was designed in this paper to perform interactions between modalities at the feature level, guiding the GSR and Temperature modalities to be more aligned with EEG at the feature level.
[0167] First, project the heterogeneous features before inputting them into the classifiers of the teacher model and the student model into a common space, and then conduct feature guidance to align the feature structure of the student model with that of the teacher model. Use the features of the teacher model as the supervision signal to enable the student model to learn the feature structure of the teacher model. This method reduces the distance between positive samples that describe the same target emotional state and increases the distance from negative samples that are inconsistent with the target emotional state. Specifically, the positive sample pairs of the teacher model and the student model are There are K negative samples The negative sample pairs are .
[0168] The loss function of feature-level knowledge distillation is as follows:
[0169] ;
[0170] where and represent the feature extraction and projection of the teacher model and the student model, and is a temperature parameter used to control the sharpness or smoothness of the distribution.
[0171] Label-level knowledge distillation: The knowledge transfer at the label level. In this paper, the teacher model plus layers are used to generate the probability distribution, i.e., soft labels, for the input data instead of the final emotional category labels. These soft labels reflect the relative probabilities between different categories and the boundary ambiguity between them. The student model learns based on these soft labels to imitate the decision-making process of the teacher model, and finally makes the outputs of the teacher model and the student model more consistent. The method of knowledge transfer at the label level uses the Kullback-Leibler (KL) divergence to calculate. The output probabilities of the EEG sample data after passing through the teacher model and layers are , and the output probabilities of the GSR or Temperature sample data after passing through the student model and layers are . is the number of categories for the final classification to be predicted. The loss can be expressed as in Equation (3.11):
[0172] ;
[0173] Optimization: To optimize the model, this paper adds a cross-entropy loss function to guide the classification accuracy of the student model. The one-hot label distribution probability of the target emotion is , where M is the number of categories for the final classification to be predicted. Its expression is as follows:
[0174] ;
[0175] Finally, this loss function is combined with the loss functions in feature-level and label-level knowledge distillation. The expression of the final loss function of the student model is as shown in Equation (3.13):
[0176] ;
[0177] where 、 and are used to adjust the weights of different loss terms.
[0178] During the previous process of emotion recognition using ACST-Net, it was found that the recognition performance of GSR and Temperature in easily accessible physiological signals was relatively poor compared to other easily accessible physiological signals. Therefore, this paper uses cross-modal knowledge distillation to improve the performance. This method is applied to GSR and Temperature, which perform poorly in the DEAP dataset. The results are shown in Table 3.7. The GSR in the AMIGOS dataset is also verified. The results are shown in Table 3.8. Through the distillation method proposed in this paper on the four emotion dimensions of the DEAP dataset, the average improvement of GSR is 15.3%, and the emotion recognition accuracy is improved from an average of 77% to 93%. The average improvement of Temperature is 1.14%, from an average of 68% to 70%. On the four dimensions of the AMIGOS dataset, the average improvement of GSR is 5.53%, and the emotion recognition accuracy is improved from an average of 86.5% to 92.3%.
[0179] Table 3.7 Performance of the distillation method proposed in this paper on GSR and Temperature in the DEAP dataset (Acc%):
[0180]
[0181] Table 3.8 Performance of the distillation method proposed in this paper on GSR in the AMIGOS dataset (Acc%):
[0182]
[0183] Analysis of cross-modal knowledge distillation: To verify the effectiveness of the proposed distillation method in both the label level and the feature level, this paper conducted ablation experiments on the GSR signals of the AMIGOS dataset. Specifically, the knowledge transfer of the label layer and the knowledge transfer of the feature layer were separately removed from the overall cross-modal knowledge distillation results.
[0184] Finally, the results before and after distillation were compared. The results are as Figure 7As shown, the results verified the effectiveness of knowledge transfer at two different levels. When knowledge transfer at the label - missing level was absent, the accuracies in the four emotional dimensions of valence, arousal, dominance, and liking decreased by 0.26%, 0.93%, 1.35%, and 0.61% respectively. When knowledge transfer at the feature - missing level was absent, the accuracies in the four emotional dimensions of valence, arousal, dominance, and liking decreased by 2.28%, 1.44%, 3.35%, and 1.44% respectively. Overall, the knowledge transfer process at the feature level had a greater impact on the entire distillation process.
[0185] The above - mentioned technical solutions only reflect the preferred technical solutions of the technical solutions of the present invention. Some changes that those skilled in the art of this technology may make to some parts thereof all reflect the principles of the present invention and fall within the protection scope of the present invention.
Claims
1. A method for emotion recognition based on easily obtainable physiological signals, characterized in that, The following steps are involved: Preprocessing the original physiological signal to obtain preprocessed physiological signal data; The time features are extracted in parallel using convolutions of three different-sized time kernels and concatenated in the feature dimension; Add location information, emotional cues, and category information to temporal features; Utilize the parallel spatiotemporal attention mechanism to extract the most discriminative temporal and channel features at different time points and channels; Use soft attention mechanism to fuse temporal features and channel features; The fused features are input into a classifier consisting of a linear layer and a ReLU layer for emotion recognition.
2. The method for emotion recognition based on easily obtainable physiological signals according to claim 1, characterized in that The convolution kernel lengths of the three different-sized temporal kernels are set to specific ratios of the physiological signal sampling rate, and the expression is ; , where represents the nth convolution, is the ratio, is the sampling rate.
3. The method for emotion recognition based on easily obtainable physiological signals according to claim 1, wherein The extraction of the time feature comprises: Perform convolution and nonlinear activation operations on each sample; Use average pooling to reduce the dimension of features; The features extracted by different convolution kernels are concatenated in the feature dimension and normalized through the BatchNorm layer.
4. The method for emotion recognition based on easily obtainable physiological signals according to claim 1, wherein, The parallel spatiotemporal attention mechanism includes a temporal attention block and a channel attention block: The temporal attention block calculates the correlation between different temporal features through a multi-head self-attention mechanism; The channel attention block calculates the correlation between different channel features through a multi-head self-attention mechanism.
5. The method for emotion recognition based on easily obtainable physiological signals according to claim 1, wherein The specific steps of the soft attention mechanism fusing time features and channel features include: Normalize and linearly embed the temporal features and channel features respectively; Calculate the weight score for each feature; The features are weighted and fused according to the weight scores.
6. A method for easily obtaining emotional recognition of physiological signals based on knowledge distillation, characterized in that, The following steps are involved: Preprocessing the original physiological signals; The ACST-Net model is used to extract the features of EEG signals and easily accessible physiological signals; The heterogeneous features of EEG signals and easily accessible physiological signals are projected into a common space through a linear layer, and the InfoNCE loss function is used to achieve knowledge distillation at the feature level. The heterogeneous features of EEG signals and easily accessible physiological signals are fed into the softmax layer to generate probability distribution, and the KL loss function is used to achieve knowledge distillation at the label level. The features of easily accessible physiological signals are fed into the classifier for emotion classification, and the cross entropy loss function is used to guide the classification; Combine the loss functions at the feature level, label level, and classification level to generate the final loss function and perform multiple rounds of training; The trained feature extraction model is used to realize emotion recognition based on easily accessible physiological signals.
7. The method for easily obtaining physiological signal emotion recognition based on knowledge distillation according to claim 6, characterized in that The feature-level knowledge distillation shortens the distance between positive sample pairs and increases the distance between negative sample pairs through contrastive learning. The specific loss function is: ; where and represent the feature extraction and projection of the teacher model and the student model, is a temperature parameter used to control the sharpness or smoothness of the distribution.
8. The method for easily obtaining physiological signal emotion recognition based on knowledge distillation according to claim 6, wherein The knowledge distillation at the label level calculates the difference in probability distribution output by the teacher model and the student model through KL divergence. The specific loss function is: 。 9. The method for easily obtaining physiological signal emotion recognition based on knowledge distillation according to claim 6, characterized in that, The final loss function is .
10. A high-precision emotion recognition system based on easily obtainable physiological signals, characterized in that, include: A signal preprocessing module, used for preprocessing the original physiological signal; Feature extraction module, used to extract multi-scale temporal features and channel features using the ACST-Net model; An attention mechanism module for extracting discriminative features through a parallel spatiotemporal attention mechanism; Feature fusion module, used to fuse temporal and channel features through soft attention mechanism; The classification module is used to input the fused features into the classifier for emotion recognition; The knowledge distillation module is used to improve the emotion recognition accuracy of easily accessible physiological signals through knowledge distillation at the feature level and label level.
Citation Information
Cited By
Image emotion prediction method based on double attention and diversified knowledge distillation
CN120932025A
Student emotion recognition method and system based on LLaVA multi-modal model
CN121438377A