Dual-Modal Emotion Recognition Method and System Based on Multi-Source Signals and Neural Networks

By combining the dual-mode emotion recognition method with radar and video, using neural networks and multimodal compact bilinear pooling algorithms, the discomfort and environmental impact problems of contact physiological signal acquisition are solved, and high-accurate emotion recognition is achieved.

CN114707530BActive Publication Date: 2025-07-22NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011492594.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-17
Publication Date
2025-07-22
Estimated Expiration
2040-12-17

AI Technical Summary

Technical Problem

The existing contact physiological signal collection methods are prone to cause discomfort to the human body, and are greatly affected by environmental factors in actual application scenarios, resulting in low accuracy and poor robustness in emotional recognition.

Method used

A dual-mode emotion recognition method combining vital sign monitoring radar and video is used to extract breathing signals through radar echo signals, facial expression changes are extracted through video, and feature fusion is used for neural network and multimodal compact bilinear pooling algorithm, and emotion recognition is performed by combining attention mechanisms.

Benefits of technology

The model calculation volume is reduced, the heartbeat signal characteristic information is restored, information loss is reduced, the emotion recognition accuracy is improved, and robustness is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114707530B_ABST
    Figure CN114707530B_ABST
Patent Text Reader

Abstract

The present invention discloses a dual-modal emotion recognition method and system based on multi-source signals and neural networks. First, the respiration signal is extracted from the radar echo signal, the PPG signal is extracted from the video face cheek area, and the heartbeat signal is extracted from the PPG signal. The one-dimensional convolutional neural network is used to extract the features of the physiological signals. Secondly, the continuous picture frames of the eye and mouth areas are extracted from the video, and the two-dimensional convolutional neural network and the long short-term memory network are used to extract their features. Then, feature fusion is performed based on the multi-modal compact bilinear pooling algorithm, and the attention mechanism is used to assign different weights to each dimension of the fused features. Finally, emotion recognition is carried out through the classification layer. The present invention uses a dual-modal sensor combined with a compact bilinear pooling feature fusion algorithm to achieve emotion recognition. Compared with the traditional single-modal sensor and the feature splicing-based feature fusion, it effectively reduces the feature dimension, avoids the dimension explosion, and at the same time improves the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of radar and multi-sensor fusion, and particularly relates to a dual-modal emotion recognition method and system based on multi-source signals and neural networks. Background Art

[0002] Emotion recognition is an important research content in the fields of psychology, cognitive science, computer science, etc.

[0003] Emotion recognition originated from facial expressions at the earliest. Due to its straightforwardness and individual differences, facial expressions have the situation of small expression changes. Coupled with the influence of factors such as ambient light intensity in actual application scenarios, the distinguishability of facial expressions corresponding to different emotional states is insufficient, and in some specific scenarios, people's facial expressions are easy to disguise, which brings subjective influence to emotion recognition to a certain extent, resulting in the actual emotion recognition accuracy often being lower than the results measured in the laboratory.

[0004] Physiological signals are more objective for emotion recognition, and emotion recognition based on physiological signals has gradually become one of the important research directions of emotion recognition. However, most of the existing physiological signal acquisitions are contact-based. This contact-based acquisition method is likely to cause discomfort to the human body and has a certain impact on the excitation of human emotions. Moreover, in actual application scenarios, the acquisition of physiological signals is easily affected by various environmental factors, resulting in a large degree of distortion and poor robustness, so emotion recognition based on physiological signals also has certain limitations. Summary of the Invention

[0005] The purpose of the present invention is to establish a dual-modal emotion recognition model combining an attention mechanism and a multi-modal compact bilinear pooling algorithm with the help of a vital sign monitoring radar and video to achieve emotion recognition in view of the deficiencies of contact-based physiological signal acquisition and single-modal emotion recognition models.

[0006] The technical solution for realizing the purpose of the present invention is: A dual-modal emotion recognition method based on multi-source signals and neural networks, the method comprising the following steps:

[0007] Step 1, using a camera to collect a video of the facial expression changes of a subject, and at the same time using a vital sign monitoring radar to collect a radar echo signal containing the chest and abdomen movement information of the subject, and performing arctangent demodulation and band-pass filtering on the radar echo signal to obtain a respiration signal;

[0008] Step 2, performing frame-by-frame segmentation on the video and extracting the face region to obtain continuous face picture frames, extracting a heartbeat signal from the face picture frames, and reconstructing the heartbeat signal;

[0009] Step 3: Segment the face image frame to obtain continuous image frames of the eye and mouth regions containing emotional information;

[0010] Step 4: Preprocess the data of the two modalities obtained in Steps 1 to 3, and then use a neural network to extract the physiological signal and emotion-related features in the continuous image frames respectively;

[0011] Step 5: Fuse the features of the two modalities extracted in Step 4. The fused features are processed by an attention mechanism to assign different weights to each dimension of the fused features, and then connected to a classification layer to construct a complete dual-modal emotion recognition model;

[0012] Step 6: Use the collected physiological signals and continuous image frame data to train the dual-modal emotion recognition model, and then use the trained model to predict the unknown emotional state to achieve emotion recognition.

[0013] Further, the process of extracting the heartbeat signal from the face image frame and reconstructing the heartbeat signal in Step 2 specifically includes:

[0014] Step 21-1: Extract the PPG signal from the cheek region in the continuous face image frames;

[0015] Step 21-2: Perform band-pass filtering on the PPG signal to obtain the heartbeat signal;

[0016] Step 21-3: Reconstruct the heartbeat signal using the phase tracking algorithm.

[0017] Further, the process of reconstructing the heartbeat signal in Step 2 specifically includes:

[0018] Step 22-1: Use the detrending method based on L1 trend filtering to detrend the distorted PPG signal;

[0019] Step 22-2: Perform independent component analysis on the detrended PPG signal. The component with the highest frequency domain peak among the obtained independent component components is the distorted heartbeat signal, denoted as S(t), where t represents the current moment;

[0020] Step 22-3: Perform Hilbert transform on the distorted heartbeat signal S(t) to obtain the phase phase(t) and amplitude mag(t) of the distorted heartbeat signal S(t);

[0021] Step 22-4: Initialize the observation matrix A and the prediction matrix Y for predicting the phase of the distorted signal, specifically:

[0022] Y = [1, w]

[0023] Where \(w\) is the number of sampling points for predicting the phase, and \(T\) is the matrix transpose symbol;

[0024] The least squares solution \(P\) for phase linear prediction obtained from the observation matrix \(A\) and the prediction matrix \(Y\) is:

[0025] \(P = Y\times(A T A) -1 A T

[0026] Where \(\times\) is the matrix outer product symbol, and \((A T A) -1 is the generalized inverse matrix of \(A T A\);

[0027] Step 22 - 5, use the phases of the first \(w\) distorted heartbeat signal sampling points at the current distortion time \(t\) to linearly predict the phase at the current distortion time \(t\), denoted as \(phase\_predict(t)\):

[0028] \(phase\_predict(t)=P\times phase(t)\)

[0029] Step 22 - 6, transform the value range of \(phase\_predict(t)\) to \((-\pi,\pi)\), record the difference \(diff(N\times2\pi)\) before and after the transformation of \(phase\_predict(t)\), and calculate the residual \(phase(t)-phase\_predict(t)\) between the phase \(phase(t)\) of the distorted signal at the current distortion time \(t\) and the predicted phase \(phase\_predict(t)\), denoted as \(res\);

[0030] Step 22 - 7, transform the value range of the residual \(res\) to \((-\pi,\pi)\), and use \(res\) and \(phase\_predict(t)\) to predict the phase \(phase\_reconstruct(t)\) of the reconstructed heartbeat signal:

[0031] \(phase\_reconstruct(t)=phase\_predict(t)+\alpha\times res + N\times2\pi\)

[0032] Where \(0\lt\alpha\leq1\) is the tolerance coefficient of phase change, and \(N = diff / 2\pi\);

[0033] Step 22 - 8, reconstruct the distorted heartbeat signal:

[0034] \(S(t)=mag(t).*phase\_reconstruct(t)\)

[0035] Where.\(*\) represents element - by - element multiplication.

[0036] Further, the preprocessing of the two modalities of data obtained in Steps 1 to 3 in Step 4 specifically includes:

[0037] Step 41-1: Downsample the sampling rate of the respiration signal to the same as that of the heartbeat signal, and save the downsampled respiration signal and heartbeat signal in different signal channels of the same packet of data;

[0038] Step 41-2: Store the consecutive picture frames of the eye and mouth regions extracted from the video in different channels of the same packet of data in the same storage manner as the physiological signals;

[0039] Step 41-3: Align the physiological signals and the consecutive picture frames in time, and locate the body movement time period of the subject through the collected video, and delete the physiological signal data and consecutive picture frame data containing the body movement time period to obtain physiological signal data and consecutive picture frame data not interfered by body movement.

[0040] Further, the use of neural networks to separately extract emotion-related features from physiological signals and consecutive picture frames in Step 4 specifically is: Construct a physiological signal feature extraction model using a one-dimensional convolutional neural network to extract the feature information of the respiration signal and heartbeat signal, and construct a model using a two-dimensional convolutional neural network and a long short-term memory network to extract the feature information of the consecutive picture frames of the eyes and mouth; The specific process includes:

[0041] Step 42-1: Build a physiological signal feature extraction model, specifically: Establish a neural network with the main body being a J-layer one-dimensional convolutional layer. After each layer of convolution, connect a max pooling layer for one-dimensional pooling, and use a flattening layer to flatten the temporal features extracted by the convolutional layer, and then output n-dimensional physiological signal features through a fully connected layer;

[0042] Step 42-2: Input the physiological signal data obtained in Step 41-3 into the physiological signal feature extraction model to extract the emotion-related features in the physiological signals;

[0043] Step 42-3: Build a consecutive picture frame feature extraction model, specifically: Establish a two-dimensional convolutional neural network with K layers to extract the emotion-related features in each frame of the picture, and then pass through a bidirectional long short-term memory network with the number of hidden neurons being S to capture the temporal information of the emotion changes in the consecutive picture frames, and finally output m-dimensional consecutive picture frame features through a fully connected layer;

[0044] Step 42-4: Input the consecutive picture frame data obtained in Step 41-3 into the consecutive picture frame feature extraction model to extract the emotion-related features in the consecutive picture frames.

[0045] Further, in step 5, the features of the two modalities extracted in step 4 are fused, which is implemented by using the multi-modal compact bilinear pooling algorithm; the specific process of step 5 includes:

[0046] Step 5-1, assume that the dimension of the feature vector f extracted from consecutive picture frames of the eyes and mouth is m, and the dimension of the output bimodal fusion feature vector is d, where d << m. Randomly initialize the vectors according to the following formula:

[0047]

[0048]

[0049] In the formula, h, s ∈ R m , and d is the dimension of the fused feature;

[0050] Through the above formula, s is randomly initialized as a vector of length m composed of -1 or 1, and h is randomly initialized as a vector of length m composed of any integer between 1 and d;

[0051] Step 5-2, perform dimensionality reduction on the consecutive picture frame features. Specifically: Initialize the feature y_video after dimensionality reduction ∈ 0 d , for each element f[i] in the input consecutive picture frame feature f, use h[i] as the index of the output feature after dimensionality reduction, and add f[i] * s[i] to the corresponding index element of the output feature after dimensionality reduction. Specifically, the formula is as follows:

[0052] y_video[h[i]] = y_video[h[i]] + f[i] * s[i]

[0053] In the formula, h[i], f[i], and s[i] respectively represent the element values of the vectors h, f, and s at the index position i;

[0054] Obtain the consecutive picture frame feature y_video with dimension d after dimensionality reduction;

[0055] Step 5-3, repeat steps 5-1 and 5-2 for the features extracted from the physiological signals to obtain the physiological signal feature y_physio with dimension d after dimensionality reduction;

[0056] Step 5-4, perform the fast Fourier transform on the dimensionality-reduced features y_video and y_physio simultaneously, multiply the transformed features element by element, and then perform the inverse fast Fourier transform on the multiplied result to obtain the final fused feature vector y;

[0057] Step 5-5: Pass the fused output feature vector y through a fully connected layer with the softmax activation function to output the fused feature weights feature_weights with a value range of (0, 1).

[0058] Step 5-6: Multiply the fused feature vector y by its corresponding feature weights feature_weights, and adaptively update the size of the feature weights through the training loss of the bimodal emotion recognition model to achieve the focused attention of the model on some features and filter out redundant features.

[0059] Step 5-7: The features selected in Step 5-6 pass through the classification layer to output the prediction probabilities of multiple emotional states, thus constructing a complete bimodal emotion recognition model.

[0060] A bimodal emotion recognition system based on multi-source signals and neural networks, the system includes:

[0061] The first physiological signal acquisition module is used to collect the video of the facial expression changes of the subject by using a camera, and at the same time collect the radar echo signal containing the chest and abdomen movement information of the subject by using a vital sign monitoring radar, and perform arctangent demodulation and band-pass filtering on the radar echo signal to obtain the respiration signal.

[0062] The second physiological signal acquisition module is used to segment the video frame by frame and extract the face region to obtain continuous face picture frames, extract the heartbeat signal from the face picture frames, and reconstruct the heartbeat signal.

[0063] The continuous picture frame data acquisition module is used to segment the face picture frames to obtain continuous picture frames of the eye and mouth regions containing emotion information.

[0064] The feature extraction module is used to preprocess the data of the two modalities obtained by the above modules, and then use a neural network to extract the emotion-related features in the physiological signals and continuous picture frames respectively.

[0065] The bimodal emotion recognition model construction module is used to fuse the features of the two modalities extracted by the feature extraction module. The fused features are processed by an attention mechanism to assign different weights to each dimension of the fused features, and then connected to the classification layer to construct a complete bimodal emotion recognition model.

[0066] The emotion recognition module is used to train the bimodal emotion recognition model by using the collected physiological signals and continuous picture frame data, and then use the trained model to predict the unknown emotional state to achieve emotion recognition.

[0067] Compared with the prior art, the significant advantages of the present invention are as follows: 1) Instead of directly using facial expressions for emotion recognition, continuous picture frames of the eyes and mouth are extracted from the face and input into the model, which reduces the computational load of the model while retaining the main emotion information of the face; 2) Based on the phase tracking algorithm, the distorted heartbeat signal is reconstructed, which can recover the heartbeat signal to the greatest extent and retain the relevant feature information in the signal; 3) A bimodal emotion recognition model is constructed using a neural network, eliminating the need for manual feature extraction, reducing information loss caused by feature extraction, and simplifying the emotion recognition process; 4) The multi-modal compact bilinear pooling algorithm is used for feature fusion, which greatly reduces the feature dimension and the computational load of the model while retaining the original feature information; 5) Two types of modal information are used to recognize different emotional states, which improves the accuracy of emotion recognition and enhances the robustness of the model.

[0068] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0069] Figure 1 It is a flow block diagram of a bimodal emotion recognition method based on multi-source signals and a neural network.

[0070] Figure 2 It is an illustration of the distortion of the PPG signal in the case of a sudden change in light intensity in one embodiment.

[0071] Figure 3 It is a comparison diagram of the PPG signals based on L1 trend filtering and L2 trend filtering in one embodiment.

[0072] Figure 4 It is a comparison diagram of the heartbeat signals before and after using the phase tracking algorithm in one embodiment.

[0073] Figure 5 It is a detailed comparison diagram of the distorted parts before and after the recovery of the heartbeat signal in one embodiment.

[0074] Figure 6 It is a diagram of a bimodal emotion recognition model combining the attention mechanism and the multi-modal compact bilinear pooling algorithm in one embodiment. Detailed Embodiments

[0075] In order to make the objectives, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0076] In addition, if the description in the present invention involves "first", "second", etc., the description of "first", "second", etc. is only for descriptive purposes and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0077] Combined Figure 1 , the present invention provides a dual-modal emotion recognition method based on multi-source signals and neural networks, and the method includes the following steps:

[0078] Step 1, use a camera to collect a video of the facial expression changes of the subject, and at the same time use a vital sign monitoring radar to collect a radar echo signal containing the chest and abdomen movement information of the subject, and perform arctangent demodulation and band-pass filtering on the radar echo signal to obtain a respiration signal;

[0079] Step 2, perform frame-by-frame segmentation on the video and extract the face region to obtain continuous face picture frames, extract the heartbeat signal from the face picture frames, and reconstruct the heartbeat signal; the specific process includes:

[0080] Step 21-1, extract the PPG signal from the cheek region in the continuous face picture frames of the face;

[0081] Step 21-2, perform band-pass filtering on the PPG signal to obtain a heartbeat signal;

[0082] Step 21-3, use the phase tracking algorithm to reconstruct the heartbeat signal (solving the problem of heartbeat signal distortion caused by sudden changes in light intensity in the actual application scenario). Specifically, it includes:

[0083] Step 22-1, use the detrending method based on L1 trend filtering to detrend the distorted PPG signal and reduce the signal distortion degree affected by sudden changes in light intensity;

[0084] Step 22-2, perform independent component analysis processing on the detrended PPG signal, and the component with the highest frequency domain peak in the obtained independent component components is the distorted heartbeat signal, denoted as S(t), where t represents the current moment;

[0085] Step 22-3, perform Hilbert transform on the distorted heartbeat signal S(t) to obtain the phase phase(t) and amplitude mag(t) of the distorted heartbeat signal S(t);

[0086] Step 22-4, initialize the observation matrix A and the prediction matrix Y for predicting the phase of the distorted signal, specifically:

[0087] Y = [1, w]

[0088] where w is the number of sampling points for predicting the phase, and T is the matrix transpose symbol;

[0089] Obtain the least-squares solution P for phase linear prediction from the observation matrix A and the prediction matrix Y as:

[0090] P = Y * (A T A) -1 A T

[0091] where * is the matrix outer product symbol, and (A T A) -1 is the generalized inverse matrix of A T A;

[0092] Step 22-5, linearly predict the phase at the current distortion time t using the phases of the previous w (w ≤ 90) distorted heartbeat signal sampling points at the current distortion time t, denoted as phase_predict(t):

[0093] phase_predict(t) = P * phase(t)

[0094] Step 22-6, transform the value range of phase_predict(t) to (-π, π), record the difference diff(N×2π) before and after the transformation of phase_predict(t), and calculate the residual phase(t) - phase_predict(t) between the phase phase(t) of the distorted signal at the current distortion time t and the predicted phase phase_predict(t), denoted as res;

[0095] Step 22-7, transform the value range of the residual res to (-π, π), and use res and phase_predict(t) to predict the phase phase_reconstruct(t) of the reconstructed heartbeat signal:

[0096] phase_reconstruct(t) = phase_predict(t) + α × res + N×2π

[0097] where 0 < α ≤ 1 is the tolerance coefficient for phase change, and N is diff / 2π;

[0098] Step 22-8, reconstruct the distorted heartbeat signal:

[0099] S(t) = mag(t).*phase_reconstruct(t)

[0100] In the formula,.* represents the product of corresponding elements.

[0101] Step 3: Segment the face picture frame to obtain continuous picture frames of the eye and mouth regions containing emotion information.

[0102] Step 4: Preprocess the data of the two modalities obtained in Steps 1 to 3, and then use a neural network to extract emotion-related features from the physiological signals and continuous picture frames respectively. The preprocessing specifically includes:

[0103] Step 41-1: Downsample the sampling rate of the respiratory signal (100 Hz) to the same as that of the heartbeat signal (30 Hz) to ensure that the number of sampling points of the respiratory signal and the heartbeat signal in each packet of data is the same, and save the downsampled respiratory signal and heartbeat signal in different signal channels of the same packet of data.

[0104] Step 41-2: Store the continuous picture frames of the eye and mouth regions extracted from the video in different channels of the same packet of data in the same storage manner as the physiological signals.

[0105] Step 41-3: Align the physiological signals and continuous picture frames in time, and locate the body movement time period of the subject through the collected video, and delete the physiological signal data and continuous picture frame data containing the body movement time period to obtain physiological signal data and continuous picture frame data not affected by body movement interference.

[0106] Among them, the use of a neural network to extract emotion-related features from the physiological signals and continuous picture frames respectively is specifically as follows: A one-dimensional convolutional neural network is used to construct a physiological signal feature extraction model to extract the feature information of the respiratory signal and the heartbeat signal, and a two-dimensional convolutional neural network and a long short-term memory network are used to construct a model to extract the feature information of the continuous picture frames of the eyes and mouth. The specific process includes:

[0107] Step 42-1: Build a physiological signal feature extraction model, specifically: Establish a neural network with the main body being J layers of one-dimensional convolutional layers. After each layer of convolution, a max pooling layer is connected for one-dimensional pooling, and a flattening layer is used to flatten the temporal features extracted by the convolutional layer, and then an n-dimensional physiological signal feature is output through a fully connected layer. According to the characteristics and sample size of the physiological signal dataset in the present invention, the optimal model parameters selected are as follows: The size of the input physiological signal is (1800, 2), the number of one-dimensional convolutional layers is 3, the convolutional kernel sizes are 1×5, 1×3, 1×5 respectively, the number of convolutional layer filters are 16, 32, 32 respectively, the one-dimensional max pooling sizes are 5, 5, 4 respectively, and the output feature dimension is 256.

[0108] Step 42-2: Input the physiological signal data obtained in Step 41-3 into the physiological signal feature extraction model to extract the emotion-related features in the physiological signal;

[0109] Step 42-3: Build a continuous picture frame feature extraction model, specifically: establish a two-dimensional convolutional neural network with K layers, extract the emotion-related features in each frame of the picture, and then pass through a bidirectional long short-term memory network with S hidden neurons to capture the temporal information of the emotion changes in the continuous picture frames. Finally, output the m-dimensional continuous picture frame features through a fully connected layer; according to the size and sample size of the continuous picture frame dataset in the present invention, the selected optimal model parameters are as follows: the size of the input continuous picture frames is (60, 32, 64, 2), the number of two-dimensional convolutional layers is 2, the convolutional kernel sizes are 5×5 and 5×5 respectively, the number of convolutional layer filters are 20 and 50 respectively, the two-dimensional maximum pooling sizes are 4×4 and 2×2 respectively, the number of hidden neurons in the LSTM unit is 64, and the output feature dimension is 512;

[0110] Step 42-4: Input the continuous picture frame data obtained in Step 41-3 into the continuous picture frame feature extraction model to extract the emotion-related features in the continuous picture frames.

[0111] Step 5: Use the multi-modal compact bilinear pooling algorithm to fuse the features of the two modalities extracted in Step 4. The fused features are processed by an attention mechanism to assign different weights to each dimension of the fused features, and then connected to the classification layer to construct a complete bimodal emotion recognition model; the specific process includes:

[0112] Step 5-1: Assume that the dimension of the feature vector f extracted from the continuous picture frames of the eyes and mouth is m, and the dimension of the output bimodal fusion feature vector is d, where d << m. Randomly initialize the vectors according to the following formula:

[0113]

[0114]

[0115] In the formula, h, s ∈ R m , d is the dimension of the fused features;

[0116] Through the above formula, s is randomly initialized into a vector of length m composed of -1 or 1, and h is randomly initialized into a vector of length m composed of any integer between 1 and d;

[0117] Step 5-2: Reduce the dimension of the continuous picture frame features, specifically: initialize the dimension-reduced feature y_video ∈ 0 d, for each element f[i] in the input continuous picture frame feature f, use h[i] as the output feature index after dimensionality reduction, and add the corresponding index element of the output feature after dimensionality reduction with f[i]*s[i]. Specifically, the formula is as follows:

[0118] y_video[h[i]] = y_video[h[i]] + f[i]*s[i]

[0119] In the formula, h[i], f[i], and s[i] respectively represent the element values of vectors h, f, and s at the index position i;

[0120] Obtain the continuous picture frame feature y_video with the dimensionality of d after dimensionality reduction;

[0121] Step 5-3, repeat Steps 5-1 and 5-2 for the features extracted from the physiological signals to obtain the physiological signal feature y_physio with the dimensionality of d after dimensionality reduction;

[0122] Step 5-4, perform the fast Fourier transform on the features y_video and y_physio after dimensionality reduction simultaneously, multiply the transformed features element by element, and then perform the inverse fast Fourier transform on the multiplied result to obtain the final fused feature vector y, realizing the fast calculation of feature fusion;

[0123] Step 5-5, pass the fused output feature vector y through a fully connected layer with the activation function softmax to output the fused feature weight feature_weights with the value range of (0,1);

[0124] Step 5-6, multiply the fused feature vector y by its corresponding feature weight feature_weights, and adaptively update the size of the feature weight through the training loss of the bimodal emotion recognition model to realize the focused attention of the model on some features and screen out redundant features;

[0125] Step 5-7, the features after feature selection in Step 5-6 pass through the classification layer to output the prediction probabilities of multiple emotion states (4 kinds, such as calm, happy, sad, fear), and thus a complete bimodal emotion recognition model is constructed.

[0126] Step 6, use the collected physiological signals and continuous picture frame data to train the bimodal emotion recognition model, and then use the trained model to predict the unknown emotion state to realize emotion recognition.

[0127] The present invention provides a bimodal emotion recognition system based on multi-source signals and neural networks, and the system includes:

[0128] The first physiological signal acquisition module is used to collect the video of the facial expression changes of the subject by using a camera, and at the same time collect the radar echo signal containing the chest and abdomen movement information of the subject by using a vital sign monitoring radar, and perform arctangent demodulation and band-pass filtering on the radar echo signal to obtain a respiration signal;

[0129] The second physiological signal acquisition module is used to perform frame-by-frame segmentation on the video and extract the face region to obtain continuous face picture frames, extract the heartbeat signal from the face picture frames, and reconstruct the heartbeat signal; this module includes the following steps executed in sequence:

[0130] The PPG signal extraction unit is used to extract the PPG signal from the cheek region in the continuous face picture frames of the face;

[0131] The heartbeat signal acquisition unit is used to perform band-pass filtering on the PPG signal to obtain the heartbeat signal;

[0132] The heartbeat signal reconstruction unit is used to reconstruct the heartbeat signal by using the phase tracking algorithm; this unit includes the following steps executed in sequence:

[0133] The detrending subunit is used to perform detrending on the distorted PPG signal by using the detrending method based on L1 trend filtering;

[0134] The first processing subunit is used to perform independent component analysis on the detrended PPG signal, and the component with the highest peak in the frequency domain among the obtained independent component components is the distorted heartbeat signal, denoted as S(t), where t represents the current moment;

[0135] The second processing subunit is used to perform Hilbert transform on the distorted heartbeat signal S(t) to obtain the phase phase(t) and amplitude mag(t) of the distorted heartbeat signal S(t);

[0136] The phase prediction parameter initialization subunit is used to initialize the observation matrix A and the prediction matrix Y for predicting the phase of the distorted signal, specifically:

[0137] Y = [1, w]

[0138] In the formula, w is the number of sampling points for predicting the phase, and T is the matrix transpose symbol;

[0139] The least squares solution P of the phase linear prediction obtained from the observation matrix A and the prediction matrix Y is:

[0140] P = Y * (A T A) -1 A T

[0141] In the formula, * is the matrix outer product symbol, (A T A) -1For A T The generalized inverse matrix of A;

[0142] A linear predictor unit, which is used to linearly predict the phase at the current distortion time t by using the phases of the first w sampled points of the distorted heartbeat signals at the previous w distortion times, denoted as phase_predict(t):

[0143] phase_predict(t) = P * phase(t)

[0144] A residual calculation unit, which is used to transform the value range of phase_predict(t) to (-π, π), record the difference diff(N×2π) before and after the transformation of phase_predict(t), and calculate the residual phase(t) - phase_predict(t) between the phase of the distorted signal phase(t) and the predicted phase phase_predict(t) at the current distortion time t, denoted as res;

[0145] A phase prediction unit, which is used to transform the value range of the residual res to (-π, π) and predict the phase phase_reconstruct(t) of the reconstructed heartbeat signal by using res and phase_predict(t):

[0146] phase_reconstruct(t) = phase_predict(t) + α × res + N×2π

[0147] In the formula, 0 < α ≤ 1 is the tolerance coefficient of phase change, and N is diff / 2π;

[0148] A reconstruction unit, which is used to reconstruct the distorted heartbeat signal:

[0149] S(t) = mag(t).*phase_reconstruct(t)

[0150] In the formula,.* represents the product of corresponding elements.

[0151] A continuous picture frame data acquisition module, which is used to segment the face picture frame to obtain continuous picture frames of the eye and mouth regions containing emotion information;

[0152] A feature extraction module, which is used to preprocess the data of the two modalities obtained by the above module, and then use a neural network to extract emotion-related features in the physiological signal and the continuous picture frame respectively; this module includes the following steps executed in sequence:

[0153] A preprocessing unit, which is used to downsample the sampling rate of the respiration signal to the same as that of the heartbeat signal and save the downsampled respiration signal and the heartbeat signal in different signal channels of the same packet of data;

[0154] The continuous picture frames of the eye and mouth regions extracted from the video are stored in different channels of the same packet of data in the same storage manner as the physiological signals;

[0155] Align the physiological signals and the continuous picture frames in time, and locate the body movement time period of the subject through the collected video. Delete the physiological signal data and continuous picture frame data containing the body movement time period to obtain physiological signal data and continuous picture frame data that are not interfered by body movement.

[0156] The first model construction unit is used to build a physiological signal feature extraction model. Specifically: establish a neural network with a main body of J one-dimensional convolutional layers. After each layer of convolution, connect a max-pooling layer for one-dimensional pooling, and use a flattening layer to flatten the temporal features extracted by the convolutional layer, and then output n-dimensional physiological signal features through a fully connected layer;

[0157] The first feature extraction unit is used to input the physiological signal data into the physiological signal feature extraction model to extract the emotion-related features in the physiological signals;

[0158] The second model construction unit is used to build a continuous picture frame feature extraction model. Specifically: establish a two-dimensional convolutional neural network with K layers to extract the emotion-related features in each frame of the picture, and then pass through a bidirectional long short-term memory network with S hidden neurons to capture the temporal information of the emotion changes in the continuous picture frames, and finally output m-dimensional continuous picture frame features through a fully connected layer;

[0159] The second feature extraction unit is used to input the continuous picture frame data into the continuous picture frame feature extraction model to extract the emotion-related features in the continuous picture frames.

[0160] The dual-modal emotion recognition model construction module is used to fuse the features of the two modalities extracted by the feature extraction module. The fused features are processed by an attention mechanism to assign different weights to each dimension of the fused features, and then connected to a classification layer to build a complete dual-modal emotion recognition model; this module includes the following steps executed in sequence:

[0161] The vector initialization unit assumes that the dimension of the feature vector f extracted from the continuous picture frames of the eyes and mouth is m, and the dimension of the output dual-modal fusion feature vector is d, where d << m. Randomly initialize the vector according to the following formula:

[0162]

[0163]

[0164] where h, s ∈ R m , d is the dimension of the fused feature;

[0165] Through the above formula, s is randomly initialized as a vector of length m consisting of -1 or 1, and h is randomly initialized as a vector of length m consisting of any integer between 1 and d;

[0166] A dimensionality reduction processing unit, which is used to perform dimensionality reduction on the continuous picture frame features. Specifically: Initialize the dimensionality-reduced feature y_video ∈ 0 d For each element f[i] in the input continuous picture frame feature f, use h[i] as the index of the output feature after dimensionality reduction, and add f[i]*s[i] to the corresponding index element of the output feature after dimensionality reduction. Specifically, the formula is as follows:

[0167] y_video[h[i]] = y_video[h[i]] + f[i]*s[i]

[0168] In the formula, h[i], f[i], and s[i] respectively represent the element values of the vectors h, f, and s at the index position i;

[0169] Obtain the continuous picture frame feature y_video with dimension d after dimensionality reduction;

[0170] A physiological signal feature acquisition unit, which is used to repeat the vector initialization unit and the dimensionality reduction processing unit for the features extracted from the physiological signal, and obtain the physiological signal feature y_physio with dimension d after dimensionality reduction;

[0171] A feature vector fusion unit, which is used to perform fast Fourier transform on the dimensionality-reduced features y_video and y_physio simultaneously, multiply the transformed features element by element, and then perform inverse fast Fourier transform on the multiplied result to obtain the final fused feature vector y;

[0172] A feature weight acquisition unit, which is used to pass the fused output feature vector y through a fully connected layer with the activation function softmax, and output the fused feature weight feature_weights with a value range of (0,1);

[0173] A feature weight update unit, which is used to multiply the fused feature vector y by its corresponding feature weight feature_weights, and adaptively update the size of the feature weight through the training loss of the bimodal emotion recognition model, so as to realize the focused attention of the model on some features and filter out redundant features;

[0174] A bimodal emotion recognition model construction unit, which is used to pass the features after the above feature selection through a classification layer, and output the prediction probabilities of multiple emotion states, thus constructing a complete bimodal emotion recognition model;

[0175] An emotion recognition module, which is used to train a bimodal emotion recognition model by using the collected physiological signals and continuous picture frame data, and then use the trained model to predict the unknown emotion state to achieve emotion recognition.

[0176] For the specific limitations of the bimodal emotion recognition system based on multi-source signals and neural networks, reference can be made to the limitations of the bimodal emotion recognition method based on multi-source signals and neural networks in the above text, which will not be elaborated here. Each module in the above-mentioned bimodal emotion recognition system based on multi-source signals and neural networks can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of the computer device in the form of hardware or independent of it, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0177] Neural networks can achieve adaptive learning of various modal data, automatically map the complex information deep in the data to a relatively simple feature representation through non-linear mapping, and have strong generalization ability and high fault tolerance. The bimodal emotion recognition combining neural networks and multi-source signals can make full use of the emotion-related information in the two modal data for mutual supplementation, while simplifying the complexity of the model, enhancing the robustness of the emotion recognition system, and effectively improving the accuracy of emotion recognition.

[0178] The following further describes the present invention in detail with reference to embodiments.

[0179] Embodiment

[0180] Combined with Figure 1 , the bimodal emotion recognition method based on multi-source signals and neural networks of the present invention includes the following steps:

[0181] Step 1, use the mobile phone camera to record the video of the facial expression changes of the subject, and at the same time use the vital sign monitoring radar to collect the radar echo signal containing the chest and abdomen movement information of the subject, and perform arctangent demodulation and band-pass filtering on the radar echo signal to obtain the respiration signal;

[0182] Step 2, segment the video frame by frame and extract the face region to obtain continuous face picture frames, extract the PPG signal from the cheek region in the continuous face picture frames, perform band-pass filtering on the PPG signal to obtain the heartbeat signal. In the actual application scenario, there is a problem of distortion of the PPG signal and the heartbeat signal when the light intensity suddenly changes. For example Figure 2 , the detrended PPG signal is severely distorted. After selecting the detrending method based on L1 trend filtering, the signal distortion degree is alleviated. For example Figure 3 shown; adopting an algorithm of phase tracking can relatively accurately recover the heartbeat signal. For example Figure 4 shown, and the heartbeat signal at the distorted part can be well recovered. For exampleFigure 5 as shown

[0183] Step 3: Segment the face picture frames extracted in Step 2 to obtain continuous picture frames of the eye and mouth regions containing emotion information;

[0184] Step 4: Preprocess the data of the two modalities obtained in Steps 1 - 3, and use a one-dimensional convolutional neural network to construct a physiological signal feature extraction model to extract the feature information of the respiration signal and the heartbeat signal. The specific structural parameters of the model are shown in Table 1.

[0185] Table 1 Physiological Signal Feature Extraction Model Parameter Table

[0186] Model parameters Set value Input layer size (1800,2) Number of one-dimensional convolutional layers 3 Convolution kernel size 1×5、1×3、1×5 Number of convolutional layer filters 16、32、32 One-dimensional max pooling size 5、5、4 Output feature dimension 256

[0187] Use a two-dimensional convolutional neural network and a long short-term memory network to construct a model to extract the feature information of the continuous picture frames of the eyes and mouth. The specific structural parameters of the model are shown in Table 2.

[0188] Table 2 Continuous Picture Frame Feature Extraction Model Parameter Table

[0189] Model parameters Set value Input layer size (60,32,64,2) Number of two-dimensional convolutional layers 2 Convolution kernel size 5×5、5×5 Number of convolutional layer filters 20、50 Two-dimensional max pooling size 4×4、2×2 Number of hidden neurons in LSTM cells 64 Output feature dimension 512

[0190] Step 5: Perform feature fusion through the multi-modal compact bilinear pooling algorithm. The dimension of the fused output feature is 128. The fused features are processed by an attention mechanism to assign different weights to each dimension of the feature, and then connected to the classification layer to construct a complete bimodal emotion recognition model. The specific model is as Figure 6 as shown

[0191] Step 6: Train the bimodal emotion recognition model with the physiological signals and continuous picture frame data collected in Steps 1 - 3. According to the sample size and data characteristics of the dataset of the present invention, select the optimal parameters for model training as: batch size of 200, learning rate of 4e-3, train for 80 epochs, and then use the trained model to predict the unknown emotion state to achieve emotion recognition.

[0192] In summary, the present invention uses a bimodal sensor combined with a compact bilinear pooling feature fusion algorithm to achieve emotion recognition. Compared with traditional single-modal sensors and feature splicing-based feature fusion, it effectively reduces the feature dimension, avoids dimensional explosion, and at the same time improves the accuracy of emotion recognition.

[0193] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A dual-modal emotion recognition method based on multi-source signals and neural networks, characterized in that, The method includes the following steps: Step 1, using a camera to collect a video of the facial expression changes of the subject, and at the same time using a vital sign monitoring radar to collect a radar echo signal containing the chest and abdomen movement information of the subject, and performing arctangent demodulation and band-pass filtering on the radar echo signal to obtain a respiration signal; Step 2, segmenting the video frame by frame and extracting the face region to obtain a continuous sequence of face picture frames, extracting a heartbeat signal from the face picture frames, and reconstructing the heartbeat signal; Step 3, segmenting the face picture frames to obtain continuous picture frames of the eye and mouth regions containing emotion information; Step 4, preprocessing the data of the two modalities obtained in Steps 1 to 3, and then using a neural network to extract the emotion-related features in the physiological signals and the continuous picture frames respectively; the physiological signals include a respiration signal and a heartbeat signal; Step 5, fusing the features of the two modalities extracted in Step 4, processing the fused features through an attention mechanism, assigning different weights to each dimension of the fused features, and then connecting a classification layer to construct a complete dual-modal emotion recognition model; Step 6, training the dual-modal emotion recognition model using the collected physiological signals and continuous picture frame data, and then using the trained model to predict an unknown emotion state to achieve emotion recognition; The process of extracting a heartbeat signal from the face picture frames and reconstructing the heartbeat signal in Step 2 specifically includes: Step 21-1, extracting a PPG signal from the cheek region in the continuous face picture frames; Step 21-2, performing band-pass filtering on the PPG signal to obtain a heartbeat signal; Step 21-3, reconstructing the heartbeat signal using a phase tracking algorithm; The process of reconstructing the heartbeat signal in Step 2 specifically includes: Step 22-1, using a detrending method based on L1 trend filtering to detrend the distorted PPG signal; Step 22-2, performing independent component analysis on the detrended PPG signal, and the component with the highest frequency domain peak in the obtained independent component components is the distorted heartbeat signal, denoted as S(t), where t represents the current moment; Step 22-3, performing a Hilbert transform on the distorted heartbeat signal S(t) to obtain the phase phase(t) and amplitude mag(t) of the distorted heartbeat signal S(t); Step 22-4, initializing an observation matrix A and a prediction matrix Y for predicting the phase of the distorted signal, specifically: In the formula, w is the number of sampling points for predicting the phase, and T is the matrix transpose symbol; The least squares solution P for linear prediction of the phase obtained from the observation matrix A and the prediction matrix Y is: P = Y * (A T A) -1 A T where * is the symbol of matrix outer product, (A T A) -1 is the generalized inverse matrix of A T ; Step 22-5, linearly predicting the phase at the current distortion moment t using the phases of the first w sampling points of the distorted heartbeat signal at the current distortion moment t, denoted as phase_predict(t): phase_predict(t) = P * phase(t) Step 22-6: Transform the value range of phase_predict(t) to (-π, π), record the difference diff(N×2π) before and after the transformation of phase_predict(t), and calculate the residual phase(t) - phase_predict(t) between the distorted signal phase phase(t) and the predicted phase phase_predict(t) at the current distortion moment t, denoted as res; Step 22-7: Transform the value range of the residual res to (-π, π), and use res and phase_predict(t) to predict the phase phase_reconstruct(t) of the reconstructed heartbeat signal: phase_reconstruct(t) = phase_predict(t) + α×res + N×2π where 0 < α ≤ 1 is the tolerance coefficient of phase change, and N is diff / 2π; Step 22-8: Reconstruct the distorted heartbeat signal: S(t) = mag(t).*phase_reconstruct(t) where.* represents element-wise multiplication; The preprocessing of the two modalities of data obtained in Step 4 for Steps 1 to 3 specifically includes: Step 41-1: Downsample the sampling rate of the respiratory signal to the same as that of the heartbeat signal, and save the downsampled respiratory signal and heartbeat signal in different signal channels of the same packet of data; Step 41-2: Store the consecutive picture frames of the eye and mouth regions extracted from the video in different channels of the same packet of data in the same storage manner as the physiological signal; Step 41-3: Align the physiological signal and the consecutive picture frames in time, and locate the body movement time period of the subject through the collected video, and delete the physiological signal data and consecutive picture frame data containing the body movement time period to obtain physiological signal data and consecutive picture frame data not affected by body movement interference.

2. The dual-modal emotion recognition method based on multi-source signals and neural networks according to claim 1, characterized in that, The use of neural networks in Step 4 to separately extract emotion-related features from physiological signals and consecutive picture frames is specifically as follows: A one-dimensional convolutional neural network is used to construct a physiological signal feature extraction model to extract the feature information of the respiratory signal and heartbeat signal, and a two-dimensional convolutional neural network and a long short-term memory network are used to construct a model to extract the feature information of the consecutive picture frames of the eye and mouth; The specific process includes: Step 42-1: Build a physiological signal feature extraction model, specifically: Establish a neural network with a main body of J one-dimensional convolutional layers. After each layer of convolution, a max pooling layer is connected for one-dimensional pooling, and a flattening layer is used to flatten the temporal features extracted by the convolutional layer, and then an n-dimensional physiological signal feature is output through a fully connected layer; Step 42-2: Input the physiological signal data obtained in Step 41-3 into the physiological signal feature extraction model to extract the emotion-related features in the physiological signal; Step 42-3: Build a continuous picture frame feature extraction model, specifically: establish a two-dimensional convolutional neural network with K layers to extract emotion-related features in each frame of the picture, then capture the temporal information of emotion changes in continuous picture frames through a bidirectional long short-term memory network with S hidden neurons, and finally output m-dimensional continuous picture frame features through a fully connected layer; Step 42-4: Input the continuous picture frame data obtained in Step 41-3 into the continuous picture frame feature extraction model to extract the emotion-related features in the continuous picture frames.

3. The dual-modal emotion recognition method based on multi-source signals and neural networks according to claim 2, characterized in that In Step 5, the features of the two modalities extracted in Step 4 are fused, which is realized by using the multi-modal compact bilinear pooling algorithm; the specific process of Step 5 includes: Step 5-1: Assume that the dimension of the feature vector f extracted from the continuous picture frames of the eyes and mouth is m, and the dimension of the output bimodal fusion feature vector is d, where d = m. Randomly initialize the vectors according to the following formula: where \(h, s\in R\) m , and \(d\) is the dimension of the fused feature; Through the above formula, s is randomly initialized into a vector of length m composed of -1 or 1, and h is randomly initialized into a vector of length m composed of any integer between 1 and d; Step 5-2: Dimension reduction is performed on the continuous picture frame features, specifically: Initialize the dimension-reduced feature y_video ∈ 0 d , for each element f[i] in the input continuous picture frame feature f, use h[i] as the output feature index after dimension reduction, and add f[i]*s[i] to the corresponding index element of the output feature after dimension reduction. Specifically, the formula is as follows: y_video[h[i]] = y_video[h[i]] + f[i] * s[i] In the formula, h[i], f[i], and s[i] respectively represent the element values of the vectors h, f, and s at the index position i; Obtain the continuous picture frame feature y_video with the dimension of d after dimensionality reduction; Step 5-3: Repeat Steps 5-1 and 5-2 for the features extracted from the physiological signals to obtain the physiological signal feature y_physio with the dimension of d after dimensionality reduction; Step 5-4: Perform fast Fourier transform on the features y_video and y_physio after dimensionality reduction at the same time, multiply the transformed features element by element, and then perform inverse fast Fourier transform on the multiplied result to obtain the final fusion feature vector y; Step 5-5: Pass the fused output feature vector y through a fully connected layer with the activation function softmax to output the fusion feature weight feature_weights with the value range of (0, 1); Step 5-6: Multiply the fused feature vector y by its corresponding feature weight feature_weights, and adaptively update the size of the feature weight through the training loss of the bimodal emotion recognition model to achieve the focused attention of the model on some features and filter out redundant features; Step 5-7: The features selected after Step 5-6 pass through the classification layer to output the prediction probabilities of multiple emotion states, thus constructing a complete bimodal emotion recognition model.

4. A dual-modal emotion recognition system based on multi-source signals and neural networks according to the method described in any one of claims 1 to 3, characterized in that, The system includes: The first physiological signal acquisition module is used to collect the video of the subject's facial expression changes by using a camera, and at the same time collect the radar echo signal containing the chest and abdomen movement information of the subject by using a vital sign monitoring radar, and perform arctangent demodulation and band-pass filtering on the radar echo signal to obtain the respiration signal; The second physiological signal acquisition module is used to segment the video frame by frame and extract the face area to obtain continuous face picture frames, extract the heartbeat signal from the face picture frames, and reconstruct the heartbeat signal; A continuous picture frame data acquisition module, which is used to segment the face picture frames to obtain continuous picture frames of the eye and mouth regions containing emotion information; A feature extraction module, which is used to preprocess the data of the two modalities obtained by the above module, and then use a neural network to extract emotion-related features in physiological signals and continuous picture frames respectively; A dual-modal emotion recognition model construction module, which is used to fuse the features of the two modalities extracted by the feature extraction module. The fused features are processed by an attention mechanism to assign different weights to each dimension of the fused features, and then connected to a classification layer to construct a complete dual-modal emotion recognition model; An emotion recognition module, which is used to train the dual-modal emotion recognition model with the collected physiological signals and continuous picture frame data, and then use the trained model to predict the unknown emotion state to achieve emotion recognition.

5. The dual-modal emotion recognition system based on multi-source signals and neural networks according to claim 4, wherein The second physiological signal acquisition module includes the following steps executed in sequence: A PPG signal extraction unit, which is used to extract the PPG signal from the cheek region in the face continuous picture frames; A heart rate signal acquisition unit, which is used to perform band-pass filtering on the PPG signal to obtain the heart rate signal; A heart rate signal reconstruction unit, which is used to reconstruct the heart rate signal by using a phase tracking algorithm.

6. The dual-modal emotion recognition system based on multi-source signals and neural network according to claim 5, characterized in that The heart rate signal reconstruction unit includes the following steps executed in sequence: A detrending subunit, which is used to perform detrending on the distorted PPG signal by using a detrending method based on L1 trend filtering; A first processing subunit, which is used to perform independent component analysis on the detrended PPG signal. The component with the highest frequency domain peak in the obtained independent component components is the distorted heart rate signal, denoted as S(t), where t represents the current moment; A second processing subunit, which is used to perform Hilbert transform on the distorted heart rate signal S(t) to obtain the phase phase(t) and amplitude mag(t) of the distorted heart rate signal S(t); A phase prediction parameter initialization subunit, which is used to initialize the observation matrix A and the prediction matrix Y for predicting the phase of the distorted signal, specifically: Y = [1, w] where w is the number of sampling points for predicting the phase, and T is the matrix transpose symbol; The least squares solution P for phase linear prediction obtained from the observation matrix A and the prediction matrix Y is: P = Y * (A T A) -1 A T where * is the symbol of matrix outer product, (A T A) -1 is the generalized inverse matrix of A T ; A linear prediction subunit, which is used to linearly predict the phase at the current distortion moment t by using the phases of the first w sampling points of the distorted heart rate signal at the current distortion moment t, denoted as phase_predict(t): phase_predict(t) = P * phase(t) A residual calculation subunit, which is used to transform the value range of phase_predict(t) to (-π, π), record the difference diff(N×2π) before and after the transformation of phase_predict(t), and calculate the residual phase(t) - phase_predict(t) between the phase phase(t) of the distorted signal at the current distortion moment t and the predicted phase phase_predict(t), denoted as res; A phase prediction sub-unit, which is used to transform the value range of the residual res to (-π, π), and predict the phase phase_reconstruct(t) of the reconstructed heartbeat signal by using res and phase_predict(t): phase_reconstruct(t) = phase_predict(t) + α × res + N × 2π In the formula, 0 < α ≤ 1, which is the tolerance coefficient of phase change, and N is diff / 2π; A reconstruction sub-unit, which is used to reconstruct the distorted heartbeat signal: S(t) = mag(t).*phase_reconstruct(t) In the formula,.* represents element-wise multiplication.

7. The dual-modal emotion recognition system based on multi-source signals and neural network according to claim 6, characterized in that The feature extraction module includes the following steps executed in sequence: A preprocessing unit, which is used to downsample the sampling rate of the respiratory signal to the same as that of the heartbeat signal, and save the downsampled respiratory signal and heartbeat signal in different signal channels of the same packet of data; Store the continuous picture frames of the eye and mouth regions extracted from the video in different channels of the same packet of data in the same storage manner as the physiological signal; Align the physiological signal and the continuous picture frames in time, and locate the body movement time period of the subject through the collected video, and delete the physiological signal data and continuous picture frame data containing the body movement time period to obtain the physiological signal data and continuous picture frame data not disturbed by body movement; A first model construction unit, which is used to build a physiological signal feature extraction model. Specifically: establish a neural network with a main body of J one-dimensional convolutional layers. After each layer of convolution, connect a max-pooling layer for one-dimensional pooling, and use a flattening layer to flatten the temporal features extracted by the convolutional layer, and then output n-dimensional physiological signal features through a fully connected layer; A first feature extraction unit, which is used to input the physiological signal data into the physiological signal feature extraction model to extract the emotion-related features in the physiological signal; A second model construction unit, which is used to build a continuous picture frame feature extraction model. Specifically: establish a two-dimensional convolutional neural network with K layers, extract the emotion-related features in each frame of the picture, and then pass through a bidirectional long short-term memory network with a hidden neuron number of S to capture the temporal information of the emotion changes in the continuous picture frames, and finally output m-dimensional continuous picture frame features through a fully connected layer; A second feature extraction unit, which is used to input the continuous picture frame data into the continuous picture frame feature extraction model to extract the emotion-related features in the continuous picture frames.

Citation Information

Patent Citations

  • Multi-mode intelligent emotion sensing system

    CN107220591A

  • Multi-mode based emotion recognition method

    CN108805089A