Multi-modal physiological signal emotion recognition method and device, storage medium and electronic equipment
Through the multimodal physiological signal emotion recognition method, the VGG network and channel attention mechanism are used to capture the emotional-related characteristics in the physiological signal, solving the problem of subjectivity and privacy leakage of identification results in the prior art, and achieving high accuracy and privacy protection emotions recognition effects.
Patent Information
- Application Number
- CN202510205415.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
Existing emotions recognition technologies rely on text or images and lack directly related physiological information, resulting in a high degree of subjectivity of the recognition results and a low accuracy rate.
The multimodal physiological signal emotion recognition method is adopted, and the VGG network feature extraction and channel attention mechanism are used to capture the emotional-related features in electrocardiogram, skin electrical activity, electromyography and respiratory signals, and a one-dimensional visual geometric group VGG network model based on channel attention mechanism is constructed.
It significantly improves the accuracy of emotional recognition, avoids the leakage of user personal privacy, and enhances the objectivity and reliability of identification results.
Smart Images

Figure CN120203580A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of emotion recognition, and particularly to a multi-modal physiological signal emotion recognition method, device, storage medium and electronic device. Background Art
[0002] Different from emotion recognition which usually uses large language models for opinion mining, emotion recognition emphasizes real-time understanding of human mental states and behaviors, and is an important part of realizing human-computer interaction, with wide applications in education, traffic safety, medical assistance and other aspects.
[0003] In related technologies, research on emotion recognition is still limited to traditional recognition methods such as text emotion recognition and image emotion recognition. Although these emotion recognition methods can reflect human emotions to a certain extent, they have no direct relationship with physiological information related to emotional expression. Therefore, their recognition results are highly subjective, vulnerable to external noise and situational ambiguity, and also prone to revealing user personal privacy.
[0004] Based on this, there is an urgent need for an emotion recognition method to improve the accuracy of emotion recognition and avoid the leakage of user personal privacy. Summary of the Invention
[0005] The purpose of this application is to provide a multi-modal physiological signal emotion recognition method, device, storage medium and electronic device, which utilize the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features to capture important features related to emotions in physiological signals, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy.
[0006] This application provides a multi-modal physiological signal emotion recognition method, including: Obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the squeeze module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features.
[0007] Optionally, the normalization processing of the physiological signal to be processed to obtain the information to be recognized includes: normalizing different physiological signals in the physiological signal to be processed to the same magnitude; and combining the normalized physiological signals by data splicing to obtain the information to be recognized.
[0008] Optionally, the inputting the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized includes: using the compression module to compress the spatial dimension of the input feature map, and using the excitation module to learn the channel dimension of the input feature map to obtain channel weights for each channel; multiplying the channel weights of each channel by the input feature map to obtain a finally output feature map, and obtaining the emotion recognition result based on the finally output feature map; wherein, the input feature map is the input of the last multiple convolutional modules in the emotion recognition model, and the finally output feature map is the output of the multiple convolutional modules; each convolutional module in the multiple convolutional modules includes: a convolutional layer, a SENet module, and a pooling layer; in each convolutional module, the SENet module is arranged before the pooling layer, and the convolutional layer is arranged before the SENet module.
[0009] Optionally, the emotion recognition model is trained based on the following steps: obtaining a multi-modal physiological signal dataset, and dividing the physiological signal dataset into a training set and a test set according to a preset ratio; the physiological signal dataset includes physiological signals collected from different emotional states and different body parts; using the training set to train a one-dimensional Visual Geometry Group (VGG) network model based on a channel attention mechanism, and using the test set to verify the trained VGG network model after training to obtain the emotion recognition model.
[0010] Optionally, the using the training set to train a one-dimensional VGG network model based on a channel attention mechanism includes: updating the parameters of the model by using the gradient calculated by the cross-entropy function through a Ranger optimizer to obtain a minimized loss function, and ending the training of the VGG network model when the loss value of the cross-entropy function no longer decreases and continues for a preset number of times.
[0011] Optionally, the emotion recognition model includes five modules and multiple fully connected layers, namely: the first module, the second module, the third module, the fourth module, and the fifth module; each of the five modules contains two one-dimensional convolutional layers, two one-dimensional normalization layers, two rectified linear unit functions, and a one-dimensional max pooling layer; the first module includes: two convolutional layers with an input channel number of 1 and an output channel number of the first channel number; the second module includes: two convolutional layers with an output channel number of the second channel number; the third module includes: three convolutional layers with an output channel number of the third channel number; the fourth module and the fifth module each include: three convolutional layers with an output channel number of the fourth channel number; the fourth channel number is twice the third channel number; the third channel number is twice the second channel number; the second channel number is twice the first channel number; the output category number of the multiple fully connected layers is the probability distribution of the target value; the multiple convolutional modules include: the third module, the fourth module, and the fifth module.
[0012] The present application also provides a multi-modal physiological signal emotion recognition device, including: A signal processing module, configured to obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; an emotion recognition module, configured to input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the squeeze module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features.
[0013] Optionally, the signal processing module is specifically configured to perform normalization processing on different physiological signals in the physiological signal to be processed, and normalize the physiological signals with different characteristics to the same magnitude; the signal processing module is specifically further configured to merge the normalized physiological signals by means of data splicing to obtain the information to be recognized.
[0014] Optionally, the SENet module includes: a compression module and an excitation module; the emotion recognition module is specifically configured to use the compression module to compress the spatial dimension of the input feature map, use the excitation module to learn the channel dimension of the input feature map, and obtain the channel weights of each channel; the emotion recognition module is further specifically configured to multiply the channel weights of each channel by the input feature map to obtain the finally output feature map, and obtain the emotion recognition result based on the finally output feature map; wherein, the input feature map is the input of the last multiple convolutional modules in the emotion recognition model, and the finally output feature map is the output of the multiple convolutional modules; each convolutional module in the multiple convolutional modules includes: a convolutional layer, a SENet module, and a pooling layer; in each convolutional module, the SENet module is arranged before the pooling layer, and the convolutional layer is arranged before the SENet module.
[0015] Optionally, the device further includes: a data acquisition module and a model training module; the data acquisition module is configured to acquire a multi-modal physiological signal dataset, and divide the physiological signal dataset into a training set and a test set according to a preset ratio; the physiological signal dataset includes physiological signals collected from different emotional states and different body parts; the model training module is configured to use the training set to train a one-dimensional Visual Geometry Group (VGG) network model based on a channel attention mechanism, and use the test set to verify the trained VGG network model after training to obtain the emotion recognition model.
[0016] Optionally, the model training module is specifically configured to update the parameters of the model by using the gradient calculated by the cross-entropy function through a Ranger optimizer to obtain a minimized loss function, and end the training of the VGG network model when the loss value of the cross-entropy function no longer decreases and lasts for a preset number of times.
[0017] Optionally, the emotion recognition model includes five modules and multiple fully-connected layers, namely: the first module, the second module, the third module, the fourth module, and the fifth module; each of the five modules contains two one-dimensional convolutional layers, two one-dimensional normalization layers, two rectified linear units, and a one-dimensional max pooling layer; the first module includes: two convolutional layers with an input channel number of 1 and an output channel number of the first channel number; the second module includes: two convolutional layers with an output channel number of the second channel number; the third module includes: three convolutional layers with an output channel number of the third channel number; the fourth module and the fifth module each include: three convolutional layers with an output channel number of the fourth channel number; the fourth channel number is twice the third channel number; the third channel number is twice the second channel number; the second channel number is twice the first channel number; the output category number of the multiple fully-connected layers is the probability distribution of the target value; the multiple convolutional modules include: the third module, the fourth module, and the fifth module.
[0018] The present application also provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the multi-modal physiological signal emotion recognition method as described in any one of the above.
[0019] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the multi-modal physiological signal emotion recognition method as described in any one of the above.
[0020] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the multi-modal physiological signal emotion recognition method as described in any one of the above.
[0021] The multi-modal physiological signal emotion recognition method, device, storage medium and electronic device provided by the present application. First, obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; then, input the information to be recognized into the emotion recognition model to obtain the emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Networks (SENet) SENet module; the compression module is used to retain and strengthen important information in the input feature map; the excitation module is used to dynamically adjust the weights of different channels to highlight the emotion information features. In this way, by using the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features, important features related to emotions in physiological signals are captured, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 is a schematic flow chart of the multi-modal physiological signal emotion recognition method provided by the present application; Figure 2 is a schematic structural diagram of the emotion recognition model provided by the present application; Figure 3 is a schematic flow chart of the emotion recognition model training method provided by the present application; Figure 4 is a schematic structural diagram of the multi-modal physiological signal emotion recognition device provided by the present application; Figure 5 is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the scope of protection of the present application.
[0025] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0026] Emotion recognition has a wide range of applications in many aspects such as education, traffic safety, and medical assistance. For example, in the field of distance education, emotion recognition can be used to evaluate the teaching status of teachers in order to monitor the teaching quality. In the field of urban planning, by studying the impact of traffic noise, etc. on the emotions of residents, it is conducive to the formulation of relevant policies and the treatment of noise pollution. In the medical field, many patients with mental illnesses usually have defects in emotion recognition. By comparing the emotion recognition accuracy at the first onset and the later stage of the disease, it can effectively help doctors conduct phased treatment of relevant patients and respond in a timely manner to possible brain damage of the patients. However, many studies are still limited to traditional recognition methods such as text emotion recognition and image emotion recognition. Although they can reflect human emotions to a certain extent, they have no direct relationship with the physiological information related to emotional expression. Therefore, the results are highly subjective and are easily affected by the leakage of personal privacy and external noise and contextual ambiguity. In addition, using a single physiological signal is affected by noise and individual differences, it is difficult to obtain important features, easily leads to a low accuracy rate, the recognition effect fluctuates greatly, and it cannot comprehensively reveal the comprehensive role of physiological signals in emotional expression.
[0027] In view of the above technical problems existing in the related art, the embodiments of the present application provide an emotion recognition method. While avoiding the leakage of personal privacy information and improving the comprehensive understanding of the emotion state by relevant researchers, this method can achieve a relatively high emotion recognition accuracy. Its essence is as follows: ①. Combining the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features to capture important features related to emotions in physiological signals; ②. This method focuses on how to extract features related to emotions from physiological signals and uses a VGG network model with high scalability to improve the accuracy of emotion recognition. The application and improvement of this method can promote the in-depth understanding of emotions by relevant researchers, and embedding this method into wearable devices will also drive the further development of industries such as medical diagnosis and health monitoring.
[0028] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the multi-modal physiological signal emotion recognition method provided by the embodiments of the present application.
[0029] As Figure 1 shown, a multi-modal physiological signal emotion recognition method provided by the embodiments of the present application may include the following steps 101 and 102: Step 101: Obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized.
[0030] Among them, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal. The above normalization processing is used to normalize different physiological signals to the same magnitude.
[0031] Specifically, the above step 101 may further include the following steps 101a1 and 101a2: Step 101a1: Perform normalization processing on different physiological signals in the physiological signal to be processed to normalize physiological signals with different characteristics to the same magnitude.
[0032] Step 101a2: Combine the normalized physiological signals by means of data splicing to obtain the information to be recognized.
[0033] Exemplarily, after obtaining physiological signals with different characteristics, they need to be fused. When fusing different physiological signals, a normalization method can be used to normalize the selected physiological signals with different characteristics to the same magnitude, and these physiological signals are combined into the information to be recognized through data splicing.
[0034] Step 102: Input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized.
[0035] Among them, the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the squeeze module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features.
[0036] Exemplarily, after normalizing the physiological signals with different features obtained and obtaining the information to be recognized, the information to be recognized can be input into the emotion recognition model for emotion recognition to obtain the emotion recognition result for the object to be recognized.
[0037] Exemplarily, the emotion recognition model includes five modules and multiple fully connected layers, namely: the first module, the second module, the third module, the fourth module, and the fifth module; each of the five modules contains two one-dimensional convolutional layers, two one-dimensional normalization layers, two rectified linear unit functions, and a one-dimensional max pooling layer; the first module includes: two convolutional layers with an input channel number of 1 and an output channel number of the first channel number; the second module includes: two convolutional layers with an output channel number of the second channel number; the third module includes: three convolutional layers with an output channel number of the third channel number; the fourth module and the fifth module each include: three convolutional layers with an output channel number of the fourth channel number; the fourth channel number is twice the third channel number; the third channel number is twice the second channel number; the second channel number is twice the first channel number; the output category number of the multiple fully connected layers is the probability distribution of the target value; the multiple convolutional modules include: the third module, the fourth module, and the fifth module.
[0038] Exemplarily, as Figure 2 shown, it is a schematic structural diagram of the emotion recognition model provided by the embodiment of the present application. The emotion recognition module is a VGG16-SENet model. VGG16-SENet is constructed by combining a feature extraction module and an attention mechanism module. The backbone model adopted by the feature extraction module is a one-dimensional VGG16 network, and the introduced attention mechanism module is a SENet module. One-dimensional VGG16 extracts initial features from the input multimodal physiological signals. The SENet module is introduced in the last three convolutional modules of the one-dimensional VGG16 model, which can dynamically adjust the weights of different channels to further highlight the emotion information features.
[0039] Specifically, the above step 102 may further include the following steps 102a1 and 102a2: Step 102a1: Use the compression module to compress the spatial dimension of the input feature map, and use the excitation module to learn the channel dimension of the input feature map to obtain the channel weights of each channel.
[0040] Step 102a2: Multiply the channel weights of each channel with the input feature map to obtain the finally output feature map, and obtain the emotion recognition result based on the finally output feature map.
[0041] Among them, the input feature map is the input of the last multiple convolutional modules in the emotion recognition model, and the finally output feature map is the output of the multiple convolutional modules. Each convolutional module in the multiple convolutional modules includes: a convolutional layer, a SENet module, and a pooling layer; in each convolutional module, the SENet module is arranged before the pooling layer, and the convolutional layer is arranged before the SENet module.
[0042] As Figure 2 shown, the one-dimensional VGG16 network can be divided into five modules in total (i.e., the above-mentioned first module, second module, third module, fourth module, and fifth module). Each module includes two one-dimensional convolutional layers, two one-dimensional normalization layers, two ReLU functions (i.e., the above-mentioned rectified linear unit functions), and one one-dimensional max pooling layer. The first module includes the 1st and 2nd convolutional layers, which use 64 3×3 convolutional kernels respectively. The number of input channels is 1, and the number of output channels is 64 (i.e., the above-mentioned first number of channels); the second module includes the 3rd and 4th convolutional layers, which use 128 3×3 convolutional kernels respectively, and the number of output channels is 128 (i.e., the above-mentioned second number of channels); the third module includes the 5th, 6th, and 7th convolutional layers, which use 256 3×3 convolutional kernels respectively, and the number of output channels is 256 (i.e., the above-mentioned third number of channels); the last two modules each contain three convolutional layers, which use 512 3×3 convolutional kernels respectively, and the number of output channels of each module is 512 (i.e., the above-mentioned fourth number of channels). Next, three fully connected layers are used to output the probability distribution with the number of categories being 3 (i.e., the above-mentioned target value), and the dropout function is introduced in the first two fully connected layers to prevent overfitting.
[0043] As Figure 2 shown, to optimize the model performance, SENet is introduced into the last three convolutional modules of the one-dimensional VGG16 network, aiming to improve the discriminability of feature representation. The SENet module can be divided into a compression module and an excitation module. The compression module compresses the spatial dimension of the input feature map, while the excitation module learns the channel dimension of the input feature map, multiplies the obtained weights of each channel with the input feature map to obtain the finally output feature map.
[0044] As Figure 2As shown, since there is a pooling layer in the base model, which is responsible for downsampling and information compression of the feature map, in order to reduce the loss of key features and provide higher-quality input features for the subsequent convolutional layer and fully connected layer, an attention mechanism module is added before the pooling layer of the last three convolutional modules (i.e., the above-mentioned third module, fourth module, and fifth module) of the one-dimensional VGG16 network to ensure that the important information of the feature map is fully retained and strengthened.
[0045] Optionally, in the embodiments of the present application, physiological signals with different features extracted from the WESAD dataset can be used to train the model.
[0046] Exemplarily, the embodiments of the present application also provide a training method for the above-mentioned emotion recognition module. The training method may include the following steps 201 and 202: Step 201, obtain a multi-modal physiological signal dataset, and divide the physiological signal dataset into a training set and a test set according to a preset ratio.
[0047] Wherein, the physiological signal dataset contains physiological signals collected from different emotional states and different body parts.
[0048] Exemplarily, as Figure 3 shown, it is a schematic flowchart of the emotion recognition model training method provided by the embodiments of the present application. As Figure 3 shown, electrocardiogram, skin electroactivity, electromyogram, and respiratory signal are extracted from the WESAD dataset covering physiological signal data collected from different emotional states and different body parts and fused. When fusing different physiological signals, the normalization method is used to normalize the physiological signals with selected different features to the same magnitude, and these physiological signals are combined into a dataset through data splicing. Then, the training set and the test set are divided according to a ratio of 80:20 (i.e., the above-mentioned preset ratio), and the dataset is divided into ten mutually exclusive subsets of equal size for ten-fold cross-validation.
[0049] Step 202, use the training set to train a one-dimensional Visual Geometry Group (VGG) network model based on the channel attention mechanism, and use the test set to verify the trained VGG network model after training to obtain the emotion recognition model.
[0050] Exemplarily, after the training set, the test set, and the emotion recognition model are all available, the VGG16-SENet is input into the training set for training to obtain the optimal hyperparameters.
[0051] Specifically, the above step 202 may further include the following step 202a: Step 202a: Update the parameters of the model using the gradients calculated by the Ranger optimizer with the cross-entropy function to obtain a minimized loss function, and end the training of the VGG network model when the loss value of the cross-entropy function no longer decreases and persists for a preset number of times.
[0052] Exemplarily, as Figure 3 shown, during model training, the parameters of the model are updated using the gradients calculated by the Ranger optimizer with the cross-entropy function to find the minimized loss function and improve the accuracy of the model. The early stopping strategy is used during training. When the loss value obtained by the cross-entropy function no longer decreases in ten consecutive trainings, it can be considered that the model accuracy no longer improves, and the training is stopped. Then, the VGG16-SENet is input into the test set to test and evaluate the emotion recognition effect.
[0053] The multi-modal physiological signal emotion recognition method provided by the embodiments of the present application first obtains the physiological signals to be processed of the object to be recognized, and performs normalization processing on the physiological signals to be processed to obtain the information to be recognized. Then, the information to be recognized is input into the emotion recognition model to obtain the emotion recognition result for the object to be recognized. Among them, the physiological signals to be processed include at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, and respiratory signal. The emotion recognition model is constructed based on the one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism. The emotion recognition model includes a one-dimensional VGG16 network model and an attention mechanism module. The attention mechanism module is a Squeeze-and-Excitation Network (SENet) module. The squeeze module in the SENet module is used to retain and strengthen the important information in the input feature map. The excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features. In this way, by utilizing the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features, the important features related to emotions in physiological signals are captured, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy.
[0054] It should be noted that for the multi-modal physiological signal emotion recognition method provided by the embodiments of the present application, the execution subject can be a multi-modal physiological signal emotion recognition device, or a control module in the multi-modal physiological signal emotion recognition device for executing the multi-modal physiological signal emotion recognition method. In the embodiments of the present application, the multi-modal physiological signal emotion recognition method is executed by the multi-modal physiological signal emotion recognition device as an example to illustrate the multi-modal physiological signal emotion recognition device provided by the embodiments of the present application.
[0055] It should be noted that in the embodiments of the present application, the multi-modal physiological signal emotion recognition methods shown in the above-mentioned various method drawings are all exemplarily described by taking one drawing in the embodiments of the present application as an example. Specifically, when implemented, the multi-modal physiological signal emotion recognition methods shown in the above-mentioned various method drawings can also be implemented in combination with any other combinable drawings shown in the above embodiments, which will not be elaborated here.
[0056] The multi-modal physiological signal emotion recognition device provided by the present application will be described below, and the following description can be mutually corresponding and referred to the multi-modal physiological signal emotion recognition method described above.
[0057] Figure 4 It is a schematic structural diagram of the multi-modal physiological signal emotion recognition device provided by the embodiments of the present application, as Figure 4 shown, specifically including: A signal processing module 401, configured to obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; an emotion recognition module 402, configured to input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the squeeze module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features.
[0058] Optionally, the signal processing module 401 is specifically configured to perform normalization processing on different physiological signals in the physiological signal to be processed, and normalize the physiological signals with different characteristics to the same magnitude; the signal processing module 401 is specifically further configured to merge the normalized physiological signals by means of data splicing to obtain the information to be recognized.
[0059] Optionally, the emotion recognition module 402 is specifically configured to compress the spatial dimension of the input feature map by using the compression module, and learn the channel dimension of the input feature map by using the excitation module to obtain the channel weights of each channel; the emotion recognition module 402 is specifically applied to multiply the channel weights of each channel by the input feature map to obtain the finally output feature map, and obtain the emotion recognition result based on the finally output feature map; wherein, the input feature map is the input of the last multiple convolutional modules in the emotion recognition model, and the finally output feature map is the output of the multiple convolutional modules; each convolutional module in the multiple convolutional modules includes: a convolutional layer, a SENet module, and a pooling layer; in each convolutional module, the SENet module is arranged before the pooling layer, and the convolutional layer is arranged before the SENet module.
[0060] Optionally, the device further includes: a data acquisition module and a model training module; the data acquisition module is configured to acquire a multi-modal physiological signal dataset, and divide the physiological signal dataset into a training set and a test set according to a preset ratio; the physiological signal dataset includes physiological signals collected from different emotional states and different body parts; the model training module is configured to train a one-dimensional Visual Geometry Group (VGG) network model based on a channel attention mechanism by using the training set, and verify the trained VGG network model by using the test set after the training is completed to obtain the emotion recognition model.
[0061] Optionally, the model training module is specifically configured to update the parameters of the model by using the gradient calculated by the cross-entropy function through a Ranger optimizer to obtain a minimized loss function, and end the training of the VGG network model when the loss value of the cross-entropy function no longer decreases and lasts for a preset number of times.
[0062] Optionally, the emotion recognition model includes five modules and multiple fully connected layers, namely: the first module, the second module, the third module, the fourth module, and the fifth module; each of the five modules contains two one-dimensional convolutional layers, two one-dimensional normalization layers, two rectified linear unit functions, and a one-dimensional max pooling layer; the first module includes: two convolutional layers with an input channel number of 1 and an output channel number of the first channel number; the second module includes: two convolutional layers with an output channel number of the second channel number; the third module includes: three convolutional layers with an output channel number of the third channel number; the fourth module and the fifth module each include: three convolutional layers with an output channel number of the fourth channel number; the fourth channel number is twice the third channel number; the third channel number is twice the second channel number; the second channel number is twice the first channel number; the output category number of the multiple fully connected layers is the probability distribution of the target value; the multiple convolutional modules include: the third module, the fourth module, and the fifth module.
[0063] The multi-modal physiological signal emotion recognition device provided by the present application first obtains the physiological signal to be processed of the object to be recognized, and performs normalization processing on the physiological signal to be processed to obtain the information to be recognized; then, the information to be recognized is input into the emotion recognition model to obtain the emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on the one-dimensional Visual Geometry Group (VGG) network model with channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the squeeze module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features. In this way, by utilizing the advantages of VGG network feature extraction and the ability of channel attention mechanism to improve the attention to important features, the important features related to emotions in physiological signals are captured, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy.
[0064] Figure 5 An example of the physical structure diagram of an electronic device is as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a multi-modal physiological signal emotion recognition method, which includes: First, obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; After that, input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; Wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; The emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; The emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; The attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; The compression module in the SENet module is used to retain and strengthen important information in the input feature map; The excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features. In this way, by using the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features, important features related to emotions in physiological signals are captured, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy.
[0065] In addition, when the logic instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. And the foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0066] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the multi-modal physiological signal emotion recognition method provided by the above-mentioned various methods. The method includes: First, obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; Then, input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the compression module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features. In this way, by utilizing the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features, important features related to emotions in physiological signals are captured, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy.
[0067] On another aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the execution of the multi-modal physiological signal emotion recognition method provided by the above-mentioned various methods. The method includes: First, obtain the physiological signal to be processed of the object to be recognized, and perform normalization processing on the physiological signal to be processed to obtain the information to be recognized; Then, input the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; wherein, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, respiratory signal; the emotion recognition model is constructed based on a one-dimensional Visual Geometry Group (VGG) network model with a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a Squeeze-and-Excitation Network (SENet) module; the compression module in the SENet module is used to retain and strengthen important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the emotion information features. In this way, by utilizing the advantages of VGG network feature extraction and the ability of the channel attention mechanism to improve the attention to important features, important features related to emotions in physiological signals are captured, greatly improving the accuracy of emotion recognition and being able to avoid the leakage of user personal privacy.
[0068] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A multimodal physiological signal emotion recognition method, characterized in that: include: Acquiring a physiological signal to be processed of the object to be identified, and performing normalization processing on the physiological signal to be processed to obtain information to be identified; Inputting the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; Among them, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, and breathing signal; the emotion recognition model is constructed based on a one-dimensional visual geometry group VGG network model based on a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a compressed excitation network SENet module; the compression module in the SENet module is used to retain and enhance important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the characteristics of emotional information.
2. The method according to claim 1, characterized in that The step of normalizing the physiological signal to be processed to obtain information to be identified includes: Normalizing different physiological signals among the physiological signals to be processed, and normalizing physiological signals with different characteristics to the same magnitude; The normalized physiological signals are combined by data splicing to obtain the information to be identified.
3. The method according to claim 1 or 2, characterized in that: The step of inputting the information to be identified into an emotion recognition model to obtain an emotion recognition result for the object to be identified includes: The compression module is used to compress the spatial dimension of the input feature map, and the excitation module is used to learn the channel dimension of the input feature map to obtain the channel weight of each channel; After multiplying the channel weight of each channel by the input feature map, a final output feature map is obtained, and the emotion recognition result is obtained based on the final output feature map; Among them, the input feature map is the input of the last multiple convolution modules in the emotion recognition model, and the final output feature map is the output of the multiple convolution modules.
4. The method according to claim 1 or 3, characterized in that: The emotion recognition model is trained based on the following steps: Acquire a multimodal physiological signal data set, and divide the physiological signal data set into a training set and a test set according to a preset ratio; the physiological signal data set includes physiological signals collected from different emotional states and different body parts; The training set is used to train a one-dimensional visual geometry group VGG network model based on a channel attention mechanism, and after the training is completed, the trained VGG network model is verified using the test set to obtain the emotion recognition model.
5. The method according to claim 4, characterized in that The method of using the training set to train a one-dimensional visual geometry group VGG network model based on a channel attention mechanism includes: The gradients calculated by the cross entropy function are used by the Ranger optimizer to update the model parameters to minimize the loss function, and the training of the VGG network model is terminated when the loss value of the cross entropy function no longer decreases and continues for a preset number of times.
6. The method according to claim 3, characterized in that The emotion recognition model includes five modules and multiple fully connected layers, namely: a first module, a second module, a third module, a fourth module and a fifth module; each of the five modules includes two one-dimensional convolutional layers, two one-dimensional normalization layers, two linear rectification functions and a one-dimensional maximum pooling layer; The first module includes: two convolutional layers with an input channel number of 1 and an output channel number of the first channel number; the second module includes: two convolutional layers with an output channel number of the second channel number; the third module includes: three convolutional layers with an output channel number of the third channel number; the fourth module and the fifth module respectively include: three convolutional layers with an output channel number of the fourth channel number; the fourth channel number is twice the third channel number; the third channel number is twice the second channel number; the second channel number is twice the first channel number; the number of output categories of the multiple fully connected layers is a probability distribution of a target value; the multiple convolutional modules include: the third module, the fourth module and the fifth module.
7. A multimodal physiological signal emotion recognition device, characterized in that: The device comprises: A signal processing module, used for acquiring a physiological signal to be processed of the object to be identified, and performing normalization processing on the physiological signal to be processed to obtain information to be identified; An emotion recognition module, used for inputting the information to be recognized into an emotion recognition model to obtain an emotion recognition result for the object to be recognized; Among them, the physiological signal to be processed includes at least one of the following: electrocardiogram signal, skin electrical activity, electromyogram signal, and breathing signal; the emotion recognition model is constructed based on a one-dimensional visual geometry group VGG network model based on a channel attention mechanism; the emotion recognition model includes: a one-dimensional VGG16 network model and an attention mechanism module; the attention mechanism module is a compressed excitation network SENet module; the compression module in the SENet module is used to retain and enhance important information in the input feature map; the excitation module in the SENet module is used to dynamically adjust the weights of different channels to highlight the characteristics of emotional information.
8. The device according to claim 7, characterized in that The signal processing module is specifically used to perform normalization processing on different physiological signals in the physiological signals to be processed, and normalize physiological signals with different characteristics to the same magnitude; The signal processing module is specifically used to merge the normalized physiological signals by data splicing to obtain the information to be identified.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the multimodal physiological signal emotion recognition method as claimed in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the multimodal physiological signal emotion recognition method as claimed in any one of claims 1 to 6 are implemented.