A method for judging emotional state of disabled people based on multi-modal fusion
Through a multimodal fusion method for judging the emotional state of people with disabilities, combining audio, video and physiological signal data, and adopting modal decoupling and attention mechanisms, the problems of low wearing comfort and recognition accuracy of equipment in existing technologies are solved, and real-time emotional companionship and accurate emotion recognition are achieved.
Patent Information
- Application Number
- CN202511130970.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing technologies have many shortcomings in emotion recognition and mental health support for people with disabilities, including insufficient mental health support system, poor wearing comfort of equipment, large size of multimodal emotion recognition system and difficulty in data synchronization, low accuracy of single modality recognition, and insufficient depth of modal fusion strategy.
A multimodal fusion method for judging the emotional state of people with disabilities is adopted. Audio, video and physiological signal data are obtained through sensors, and feature extraction and fusion are performed in combination with modal decoupling strategy and attention mechanism to achieve learning of modal consistency and heterogeneity, and finally perform emotion recognition.
It improves the accuracy and robustness of emotion recognition, provides real-time emotional companionship, realizes early warning and proactive prevention, alleviates the pressure of traditional manual care, and meets the multi-dimensional emotional needs of people with disabilities.
Smart Images

Figure CN120616533B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a multi-modal fusion-based disabled person emotional state judgment method and belongs to the technical field of multi-modal fusion. BACKGROUND
[0002] Emotion, as a close link between people and the world, is closely related to human life. Therefore, research on human emotion has become a focus of interdisciplinary research - linguistics deconstructs emotion expression mechanism, biology explores emotion physiological basis, and computer science is committed to building emotion computing framework. Today, when material needs are basically met, human spiritual needs are growing exponentially. This social change has promoted the progress of human-computer interaction: scholars gradually realize that the emotional factor plays a key role in decision-making, social maintenance, etc. This discovery has prompted them to explore the integration of emotional factors into computing systems in order to enable computers to recognize, understand and simulate human emotions. Emotion recognition is an important part of this exploration process, which aims to enlighten machines to capture the subtle connection between external emotional expression and internal emotional state, thereby achieving accurate interpretation of human emotional state. In recent years, with the iterative upgrade of deep learning algorithms and the exponential growth of computing resources, emotion recognition research has broken through the boundaries of the laboratory and gradually penetrated into various application scenarios.
[0003] The information technology revolution has driven the development of human-computer interaction and opened up new possibilities for socially disadvantaged groups. In the face of the dual dilemma of physiological limitations and social prejudice faced by the disabled for a long time, emotion recognition technology shows important value: by monitoring physical and mental state in real time and providing intelligent care, it not only can improve the quality of life, but also is helpful to guide the disabled to be self-esteem, self-confidence, self-reliance and self-reliance, which has significant social value.
[0004] Human perception of others' emotional state is a complex process that integrates multiple types of information. In terms of image information alone, it covers rich details in multiple dimensions such as facial expression, body posture, and limb movement. However, in the field of emotion recognition, early related research mainly focuses on the processing of single modal data, which often leads to the inability to fully and deeply capture the complexity of emotional information, and is particularly vulnerable to noise interference, with obvious limitations. In fact, different modalities exhibit different sensitivity and specificity in reflecting emotional state. Therefore, compared to a single modality, multi-modal analysis can more accurately and comprehensively reflect the emotional state of individuals, becoming an effective way to improve the accuracy and robustness of emotion recognition.
[0005] Multimodal emotion recognition technology is a technology that combines multiple sensors and input methods, aiming to recognize the emotional state of individuals by analyzing their speech, facial expressions, physiological signals, text, and other information. In terms of hardware, the components involved in this technology mainly include sensors, data acquisition devices, data processing units, etc.
[0006] Similar implementation schemes include:
[0007] 1. Integrated wearable device: This device integrates heart rate monitors, skin conductance sensors, and microphones into a wearable unit, such as a smartwatch or chest strap.
[0008] 2. High-resolution camera: used to capture facial expressions and body movements, usually connected to a data processing unit for real-time analysis of expression changes.
[0009] 3. Data synchronization technology: uses wireless communication technology (such as Bluetooth or Wi-Fi) to achieve synchronous transmission between sensor data and data processing units.
[0010] 4. Data processing and analysis software: This software can process data from different sensors and use existing algorithms to identify and classify emotional states.
[0011] Through in-depth analysis of similar implementation schemes in the prior art, it is found that each scheme often has breakthroughs in a single dimension, but it is difficult to balance the overall performance optimization of the system. In the data acquisition and preprocessing stage, although current research has adopted head-mounted integrated devices to achieve synchronous acquisition of multi-modal data such as speech and facial video, there are still many challenges in practical applications. In the feature extraction stage, existing schemes often face a dilemma: one type of scheme performs well in single-modal feature extraction and analysis, but these schemes perform poorly in multi-modal feature fusion, making it difficult to achieve cross-modal optimization. Another type of scheme designs a complex multi-modal fusion architecture, but in the pursuit of fusion results, it ignores the characteristics of each single modality, resulting in limited overall performance improvement. In addition, in the most critical modality fusion stage, although some advanced schemes try to integrate convolutional networks and attention mechanisms to improve performance, their core fusion strategy still remains at the level of simple feature concatenation, neither fully considering the dynamic interaction between different modalities nor fully preserving modality-specific characteristics, thus failing to effectively distinguish similar emotions, resulting in a sharp drop in model robustness when facing complex scenarios.
[0012] The existing technology has several key defects in the field of emotion recognition and mental health support for people with disabilities, which need to be solved urgently. This invention proposes a systematic solution to these problems:
[0013] First, there are significant deficiencies in the service system and technical implementation level. China's large population of disabled people faces serious mental health challenges, and psychological problems such as depression and anxiety are prevalent. However, the existing mental health support system has obvious deficiencies: it mainly relies on passive mental health lectures, which cannot meet the urgent needs of the disabled population with low social participation for emotional companionship; caregiver services often focus on physical care and ignore psychological care, and professional psychological support is costly; more importantly, the current system mainly intervenes after the emergence of psychological problems, lacking effective early warning and real-time monitoring mechanisms. In terms of technical support, existing solutions have many limitations: the wearing comfort of physiological monitoring devices is not high and they are not contactless, making it difficult to meet long-term use; multi-modal emotion recognition systems usually require multiple independent hardware devices to work together, which not only has a large volume and is inconvenient to use, but also has problems such as data synchronization difficulties, bringing many inconveniences to users' daily life. The present invention not only provides real-time intelligent emotional companionship for the disabled in terms of technology, meeting users' emotional needs through human-computer interaction, effectively alleviating the pressure on traditional manual care; it can also detect emotions and use emotion recognition technology to capture emotional fluctuations, alert and notify relevant personnel when emotions are abnormal, and achieve a shift from passive response to proactive prevention. In addition, in terms of data, it ensures the accurate alignment and effective fusion of multi-modal data, providing a more accurate analysis basis for emotion recognition.
[0014] Second, in terms of emotion recognition technology, existing technology has modality limitations. Traditional solutions are mostly limited to relying on single-modal data (such as facial expression analysis only for vision or emotion recognition only for speech), which is inconsistent with the complexity and diversity of human emotional expression. In fact, emotion recognition based on real-world scenarios requires integrating multi-dimensional information: in terms of behavior, facial expressions, body language, and speech intonation all carry rich emotional information; in terms of physiology, autonomic nervous system indicators such as skin conductance provide objective reflections of internal emotional states. Single-modal systems not only cannot capture this multi-level emotional expression, but also are vulnerable when faced with information gaps in real-world scenarios (such as face occlusion, environmental noise), with a sharp decline in recognition accuracy and a clear lack of anti-interference ability. In pursuit of overall performance excellence in practical applications, single-modal recognition strategies have been unable to meet actual needs. Compared to single-modal data, multi-modal data can more comprehensively reflect an individual's emotional state and provide more rich and accurate information. Therefore, the present invention combines image, audio, and physiological signal data for emotion recognition, which can improve emotion recognition accuracy and provide reliable technical support for real-world application scenarios.
[0015] Thirdly, in terms of modal fusion, the prior art has obvious modeling limitations. The current mainstream multi-modal emotion recognition technology often adopts simple feature concatenation or weighted average fusion strategy, and fails to deeply mine the consistency and heterogeneity features contained in multi-modal data. This technical limitation is manifested in two aspects: on the one hand, some methods excessively pursue the common feature extraction among multiple modalities, but ignore the unique advantages of different single modalities in expressing emotions, such as facial expressions are good at capturing instantaneous emotional changes, voice features can reflect the intensity changes of emotions, and physiological signals objectively reflect the continuous emotional state; on the other hand, the existing technology lacks the modeling ability of the complex interaction between modalities, resulting in that when distinguishing similar emotions (such as anger and fear), the emotion recognition model can only capture the surface features and cannot recognize the deep differences, causing high recognition confusion rate. In view of the above key challenge, the present application proposes a multi-modal decoupling strategy, which integrates the consistency and heterogeneity of multi-modal data into a unified framework to realize the learning of modal consistency and heterogeneity. This innovation not only solves the problem of ignoring the capture of emotions by single modalities in the existing multi-modal emotion recognition methods, but also provides a better technical path for fine-grained emotion recognition in complex scenarios.
[0016] In summary, the present application aims to overcome the shortcomings of the prior art and provide a more comprehensive, efficient, real-time and high-accuracy multi-modal emotion recognition system, method and device to better meet the psychological health needs of the disabled and other special groups with emotional companion needs. SUMMARY
[0017] The present application designs and develops a multi-modal fusion-based disabled person emotional state judgment method, which judges the emotional state based on multi-modal information to improve the emotion recognition accuracy.
[0018] The technical scheme provided by the present application is as follows:
[0019] A multi-modal fusion-based disabled person emotional state judgment method, comprising:
[0020] Step one, acquiring audio, video and physiological signal raw data of the user through a sensor; preprocessing the audio, video and physiological signal data;
[0021] Step two, identifying the raw data to extract audio feature data, visual feature data and physiological signal feature data;
[0022] Among them, the audio feature includes the time information of the voice, the visual feature data includes the face image, and the physiological signal data includes the heart rate and skin electric reaction data;
[0023] Step three, applying the audio feature data, visual feature data, physiological signal feature data to the modal decoupling strategy to realize the learning of modal consistency and heterogeneity;
[0024] Step four, fusing the features of different modalities through attention mechanism, performing emotion recognition based on multi-modal features, and finally obtaining the emotion change recognition result;
[0025] The emotion is divided into eight emotions: neutral, calm, happy, sad, angry, fearful, disgusted, and surprised.
[0026] Preferably, the step one comprises:
[0027] For audio data, the preprocessing process includes dividing each audio descriptor into t time periods to meet the input requirements of the convolutional neural network model;
[0028] For image data, the preprocessing process includes dividing the video into t parts and randomly sampling k consecutive frames of short clips from each segment, and taking the t short clips as the data representation of the entire video to meet the input requirements of the three-dimensional convolutional neural network model;
[0029] Each segment has k consecutive frames.
[0030] For ECG data, the preprocessing process includes ECG filtering using wavelet decomposition and reconstruction method, and filtering out R-R heart rate signals through R peak positioning;
[0031] For skin electricity data, the preprocessing process includes filtering out noise by wavelet denoising; and batch normalizing the physiological signals.
[0032] Preferably, the step two comprises:
[0033] Based on the 2D-CNN neural network model, the audio features are extracted by combining the attention mechanism;
[0034] The audio features mainly use mel-frequency cepstral coefficients;
[0035] Based on the 3D-CNN neural network model, the video frame features are extracted from the input video frames;
[0036] The extracted audio features enter the time attention module, and the extracted video frame features enter the spatial attention module, the channel attention module and the time attention module:
[0037] In the spatial attention module, the input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the spatial dimension. The attention weights are normalized by the softmax function and applied to the original spatial features to generate weighted spatial features.
[0038] In the channel attention module, the video frame feature matrix is transposed into a channel feature matrix, where the spatial dimension is interchanged with the channel dimension.
[0039] The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the channel dimension; the attention weights are normalized by the softmax function and applied to the original channel features to generate weighted channel features;
[0040] The video features output by the channel attention module are first spatially averaged through the pooling layer to obtain a temporal feature sequence;
[0041] The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the time dimension. The features are activated by the ReLU function to calculate the temporal attention weights, which are then weighted and fused with the original features to obtain the time-weighted features.
[0042] Perform average pooling to reduce feature dimensions and smooth feature information;
[0043] Extract features of input physiological signals based on 1D-CNN and LSTM;
[0044] ;
[0045] ;
[0046] in, is the original feature of the physiological signal at the tth time step, R is the 1D-CNN receptive field, that is, the number of time steps covered by the convolution kernel, is the 1D-CNN convolution kernel parameter, b is the bias, For 1D-CNN at the tth time step, the activation function After the output, is the hidden state of LSTM at the tth time step, is the hidden state at the t-1th time step, which is used to capture the temporal dependency of physiological signals;
[0047] In each time attention module, the input features are first extracted local temporal patterns by a Conv1d convolution layer, then the correlation features in time dimension are generated by a fully connected layer, and the time attention weights are calculated by a ReLU activation function, and then the original features are weighted and fused to generate time-weighted feature representation;
[0048] Average pooling is performed to reduce the feature dimension while smoothing the feature information;
[0049] Cascade fusion of features from different physiological signals to form joint feature representation;
[0050] The joint features are input to a fully connected layer for further processing to generate the final physiological emotion features.
[0051] Preferably, the step three comprises:
[0052] Using a modal decoupling module for modal decoupling strategy;
[0053] The modal decoupling module comprises a modal shared encoder and three modal private encoders.
[0054] The extracted audio features , image features , physiological signal features are projected into the modal shared encoder of the modal shared space to extract the shared features of the modal, including:
[0055] ;
[0056] ;
[0057] ;
[0058] In the formula, is the audio shared feature, is the image shared feature, is the physiological signal shared feature
[0059] The modal shared encoder is used to extract common features of different modalities. The trainable parameters of the shared encoder are the same set of parameters for all modalities, learning the consistency of the modalities.
[0060] The extracted audio, image, and physiological signal features are projected into the modal private encoder of the respective modal private space to extract specific features of the modal, including:
[0061] ;
[0062] ;
[0063] ;
[0064] wherein, is an audio modality specific feature, is an image modality specific feature, is a physiological signal modality specific feature: an audio modality private editor, is a video modality private editor, is a physiological signal modality private encoder; , , are trainable parameters of the private encoder.
[0065] Preferably, the step four comprises:
[0066] concatenating the audio, visual, physiological modality specific features , , into a unified multimodal feature sequence :
[0067] ;
[0068] learning inter-modality correlation with K heads of attention, concatenating the output features of each head, and then performing linear transformation to obtain cross-modality fusion features :
[0069] ;
[0070] obtaining the final representation of the specific features of each modality through an attention mechanism, inputting into a classifier, outputting a probability distribution of emotion, and obtaining a classification result.
[0071] The beneficial effects of the present application are:
[0072] 1. The present application performs emotion recognition by fusing multimodal emotion features, effectively making up for the possible deficiencies of single modality in emotion feature expression. The interaction and complementation between different modalities enable the present application to extract more comprehensive and accurate emotion features, thereby improving the accuracy and reliability of emotion recognition.
[0073] 2. The present application adopts a multimodal information fusion strategy based on an attention mechanism, which realizes deep mining of emotion features through the synergistic effect of multiple channels.
[0074] 3. The present invention includes a modal decoupling module for distinguishing similar emotions. By capturing the consistency and heterogeneity of different modalities, the modal shared features and modal specific features of each modality are extracted, effectively improving the accuracy of emotion recognition.
[0075] 4. This invention utilizes a data acquisition device integrated with multiple sensors to monitor multiple health indicators of the target subject in real time, adding the ability to extract physiological signals and detect emotional changes. Furthermore, this invention possesses data-driven self-adjustment and optimization capabilities, leveraging data uploaded by the acquisition device to continuously refine the prediction model, thereby gradually improving the accuracy and adaptability of the predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 This is a flowchart of the method for judging the emotional state of disabled people based on multimodal fusion described in the present invention.
[0077] Figure 2 Schematic diagram of the audio feature extraction model described in the present invention.
[0078] Figure 3 Schematic diagram of the video face image feature extraction model described in the present invention.
[0079] Figure 4 Schematic diagram of the physiological emotion feature extraction model described in the present invention.
[0080] Figure 5 This is a flow chart of the modal decoupling strategy described in the present invention.
[0081] Figure 6 This is a comparison chart of the accuracy and F1 score results described in the present invention. DETAILED DESCRIPTION
[0082] The present invention will be described in further detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.
[0083] like Figures 1-6 As shown, the present invention provides a method for determining the emotional state of disabled people based on multimodal fusion. This method is suitable for users with disabilities who have limited mobility due to leg disabilities or other reasons, but whose facial expressions and voice functions are normal. Multimodal information is used to determine their emotional state, thereby improving the accuracy of emotion recognition. Because this method requires the collection of information such as facial expressions and voice intonation, it is more suitable for users with limited mobility but whose physiological functions are not significantly affected. It can also be expanded to other users with similar characteristics. Specifically, it includes:
[0084] Step 1: Obtain the user's audio, video and physiological signal raw data through sensors; pre-process the audio, video and physiological signal data;
[0085] For audio data, the preprocessing process includes: dividing each audio descriptor into t time periods to meet the input requirements of the convolutional neural network model; for image data, the preprocessing process includes: dividing the video into t parts and randomly sampling k consecutive frames of short segments from each segment, taking t short segments as the data representation of the entire video to meet the input requirements of the three-dimensional convolutional neural network model; wherein each segment has k consecutive frames;
[0086] For ECG data, the preprocessing process includes: filtering the ECG using wavelet decomposition and reconstruction method, and filtering out the R-R heart rate signal through R peak positioning;
[0087] For skin electricity data, the preprocessing process includes: filtering out noise by wavelet denoising; batch normalizing physiological signals.
[0088] Step two, identify the original data, extract audio feature data, visual feature data, and physiological signal feature data;
[0089] Among them, the audio feature contains the time information of the speech, the visual feature data contains the face image, and the physiological signal data contains the heart rate and skin electricity response data;
[0090] Based on the 2D-CNN neural network model, the audio features are extracted by combining the attention mechanism;
[0091] Among them,
[0092] The audio feature mainly uses the mel frequency cepstral coefficient;
[0093] Based on the 3D-CNN neural network model, the video frame features are extracted from the input video frames, and the feature matrix after extraction is as follows:
[0094] ;
[0095] Among them, t is the number of video segments, corresponding to the total number of segments obtained after dividing the video according to certain rules; m is determined by the spatial size (height h and width w flattened, i.e. m = h x w) of the video frame feature map, representing the number of positions after flattening the features in the spatial dimension; n is the number of feature channels, for any vector in the feature matrix , which is the visual feature vector corresponding to the jth spatial position in the ith segment.
[0096] The extracted audio features enter the time attention module, and the extracted video frame features enter three attention mechanism modules:
[0097] The spatial attention module is used to extract the importance of the features in the spatial dimension;
[0098] In the spatial attention module, the input features are first processed by a Conv1d convolution layer and then by a fully connected layer to generate the correlation features in the spatial dimension. The attention weights are normalized by a softmax function and applied to the original spatial features to generate weighted spatial features.
[0099] ;
[0100] ;
[0101] ;
[0102] wherein, 、 is the trainable parameter matrix of the spatial attention module, which is used to linearly transform the input i-th video segment feature (taken from the corresponding segment); is the spatial attention intermediate feature obtained by transformation; is the spatial attention weight obtained by applying the Softmax function to ; is the weighted spatial feature obtained by element-wise multiplication of the attention weight and the original segment feature ; is the matrix transpose.
[0103] Channel attention module: used to extract the importance of features in the channel dimension;
[0104] First, transpose to , i.e., the channel feature matrix of the video frame:
[0105] ;
[0106] wherein, for any vector in the matrix, is the feature vector corresponding to the j-th channel position in the i-th segment.
[0107] Input features are first processed by a Conv1d convolution layer and then by a fully connected layer to generate the correlation features in the channel dimension. The attention weights are normalized by a softmax function and applied to the original channel features to generate weighted channel features.
[0108] ;
[0109] ;
[0110] ;
[0111] wherein, , is a trainable parameter matrix of the channel attention module, used to transform the channel features of the i-th segment (taken from the corresponding segment); is the channel attention intermediate feature; is the channel attention weight obtained after applying the Softmax function to normalize, used to measure the importance of different channel features within the i-th video segment; is the weighted channel feature, obtained by element-wise multiplication of the attention weight and the original channel feature .
[0112] Temporal attention module: used to extract the importance of features in the time dimension.
[0113] The video features output by the channel attention module are first subjected to spatial average pooling by the pooling layer to obtain a time sequence of features (t is the number of video segments, and n is the feature dimension), which is used as the input of the temporal attention;
[0114] The input features are first subjected to a Conv1d convolution layer, and then processed by a fully connected layer to generate the correlation features in the time dimension. The features are calculated by the ReLU activation function to obtain the temporal attention weight, and then fused with the original features by weighting to obtain the time-weighted features.
[0115] ;
[0116] ;
[0117] ;
[0118] wherein, is a time sequence of features of the audio or video, is the feature vector of the j-th time sequence segment; , is a trainable parameter matrix of the temporal attention module; is the temporal attention intermediate feature; is the temporal attention weight obtained after the ReLU activation function is applied to , used to measure the importance of different time sequence segment features; is the time-weighted feature obtained by weighting and summing the features of each time sequence segment according to the attention weight.
[0119] Average pooling is performed to reduce the feature dimension and smooth the feature information;
[0120] The physiological signal input is extracted based on the 1D-CNN and the LSTM;
[0121] ;
[0122] ;
[0123] wherein, is the original feature of the physiological signal at the t-th time step; R is the 1D-CNN receptive field, that is, the number of time steps covered by the convolution kernel; is the 1D-CNN convolution kernel parameter, and b is the bias; is the output of the 1D-CNN at the t-th time step after the activation function ; is the hidden state of the LSTM at the t-th time step, is the hidden state at the t-1-th time step, which is used to capture the time sequence dependence of the physiological signal.
[0124] In the respective time attention module: the input feature is first subjected to a Conv1d convolution layer to extract local time sequence patterns, then subjected to a fully connected layer to generate correlation features in the time dimension, then subjected to a ReLU activation function to calculate time attention weights, and then subjected to weighting fusion with the original feature to generate a time-weighted feature representation;
[0125] average pooling is performed to reduce the feature dimension and smooth the feature information;
[0126] features from different physiological signals are concatenated and fused, that is, the features of the electrocardiogram signal and the features of the electrodermal signal are spliced according to a specific dimension to form a joint feature representation;
[0127] The joint feature is input into a fully connected layer for further processing to generate the final physiological emotional feature.
[0128] Step three, applying a modal decoupling strategy to the audio feature data, visual feature data, and physiological signal feature data to learn the modal consistency and heterogeneity;
[0129] The modal decoupling strategy is performed using a modal decoupling module, and the modal decoupling module includes: one modal shared encoder and three modal private encoders;
[0130] The audio feature , image feature , and physiological signal feature extracted in step two are projected into the modal shared encoder of the modal shared space to extract shared features , , of the modal
[0131] ;
[0132] ;
[0133] ;
[0134] wherein, is an audio sharing feature, is an image sharing feature, is a physiological signal sharing feature, is a modal sharing encoder for extracting common features of different modalities; is a trainable parameter of the sharing encoder, all modalities share the same set of parameters, learning the consistency of modalities.
[0135] The extracted audio, image, and physiological signal features are projected into the modal private encoder of the respective modal private space to extract specific features of the modal 、 、 :
[0136] ;
[0137] ;
[0138] ;
[0139] wherein, is an audio modal specific feature, is an image modal specific feature, is a physiological signal modal specific feature: an audio modal private editor, is a video modal private editor, is a physiological signal modal private encoder; 、 、 is a trainable parameter of the private encoder, each modality is independently trained, learning the heterogeneity of the modalities;
[0140] Step four, the features of different modalities are fused through a cross-modal attention mechanism, and emotion recognition is performed based on the multi-modal features, to obtain the emotion change recognition result;
[0141] The modal specific features of audio, vision, and physiology 、 、 are spliced into a unified multi-modal feature sequence :
[0142] ;
[0143] Learn inter-modal correlation with K-head attention, concatenate the output of each head, and then linearly transform to get cross-modal fusion features :
[0144] ;
[0145] The specific features of each modality are obtained through attention mechanism to get the final representation, and input into the classifier to output the probability distribution of emotion, and get the classification result.
[0146] As shown in Figure 6 , the comparison of the benchmark method and the present method in accuracy (Accuracy) and F1 score (F1-Score) is shown, and it can be seen from the figure that the present method is superior to the benchmark method in accuracy and F1 score through the fusion of multi-modal information for emotion recognition.
[0147] Although the embodiments of the present application have been disclosed as above, it is not limited to the application listed in the specification and embodiments, and can be fully applied to various fields suitable for the present application, and other modifications can be easily realized by those skilled in the art, and therefore the present application is not limited to specific details and the figures shown and described herein, without departing from the general concept defined by the claims and the equivalent scope.
Claims
1. A method for judging the emotional state of disabled people based on multimodal fusion, characterized in that: include: Step 1: Obtain the user's audio, video and physiological signal raw data through sensors; Preprocess audio, video and physiological signal data; The step one comprises: For audio data, the preprocessing process includes: dividing each audio descriptor into t time segments to make it meet the input requirements of the convolutional neural network model; For image data, the preprocessing process includes: dividing the video into t parts and randomly sampling short segments of k consecutive frames from each segment, and using these t short segments as the data representation of the entire video to make it meet the input requirements of the 3D convolutional neural network model; Among them, each segment has k consecutive frames; For ECG data, the preprocessing process includes: using wavelet decomposition and reconstruction method to perform ECG filtering on the original data, and then filtering out the RR heart rate signal by R peak positioning; For the electrodermal data, the preprocessing process includes: filtering out noise from the raw data through wavelet denoising; normalizing the physiological signals in batches; Step 2: Identify the original data and extract audio feature data, visual feature data, and physiological signal feature data; Among them, audio features include the time information of speech, visual feature data includes facial images, and physiological signal data includes heart rate and skin electrical response data; The second step includes: Based on the 2D-CNN neural network model, audio features are extracted in combination with the attention mechanism; The audio features are mainly based on Mel frequency cepstral coefficients; Extract video frame features from the input video frames based on the 3D-CNN neural network model; The extracted audio features enter the temporal attention module, and the extracted video frame features enter the spatial attention module, channel attention module, and temporal attention module: In the spatial attention module, the input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the spatial dimension. The attention weights are normalized by the softmax function and applied to the original spatial features to generate weighted spatial features. In the channel attention module, the video frame feature matrix is transposed into the channel feature matrix, where the spatial dimension is interchanged with the channel dimension; The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the channel dimension; the attention weights are normalized by the softmax function, and weighted channel features are generated based on the original channel features; The video features output by the channel attention module are first spatially averaged through the pooling layer to obtain a temporal feature sequence; The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the time dimension. The features are activated by the ReLU function to calculate the temporal attention weights, which are then weighted and fused with the original features to obtain the time-weighted features. Perform average pooling to reduce feature dimensions and smooth feature information; Extract features of input physiological signals based on 1D-CNN and LSTM; ; ; in, is the original feature of the physiological signal at the tth time step, R is the 1D-CNN receptive field, that is, the number of time steps covered by the convolution kernel, is the 1D-CNN convolution kernel parameter, b is the bias, For 1D-CNN at the tth time step, the activation function After the output, is the hidden state of LSTM at the tth time step, is the hidden state at the t-1th time step, which is used to capture the temporal dependency of physiological signals; In each temporal attention module, the input features are passed through the Conv1d convolution layer to extract local temporal patterns, and then the correlation features in the time dimension are generated through the fully connected layer. The temporal attention weights are then calculated using the ReLU activation function and weightedly fused with the original features to generate a time-weighted feature representation. Perform average pooling to reduce feature dimensions and smooth feature information; Cascade the features from different physiological signals to form a joint feature representation; The combined features are input to the fully connected layer for further processing to generate the final physiological emotion features; Step 3: applying a modal decoupling strategy to the audio feature data, visual feature data, and physiological signal feature data to achieve learning of modal consistency and heterogeneity; Step 4: The features of different modalities are integrated through the attention mechanism, and emotion recognition is performed based on multimodal features to finally obtain the emotion change recognition result; Emotions are divided into eight categories: neutral, calm, happy, sad, angry, fearful, disgusted and surprised.
2. The method for judging the emotional state of disabled people based on multimodal fusion according to claim 1 is characterized in that: The step three includes: Use the modal decoupling module to perform modal decoupling strategy; Among them, the modal decoupling module includes: a modal shared encoder and three modal private encoders; The extracted audio features , image features , physiological signal characteristics Projected into the modality-sharing encoder of the modality-sharing space, the shared features of the modalities are extracted, including: ; ; ; Where, For audio sharing features, Shared features for images, For physiological signal sharing features, Shared encoder for modalities; are the trainable parameters of the shared encoder; The extracted audio, image, and physiological signal features are projected into the modality-specific encoders of their respective modality-specific spaces to extract modality-specific features, including: ; ; ; Where, is the audio modality specific feature, is the image modality specific feature, Specific features for physiological signal modalities: Audio modal private editor, Private editor for video modal, Private encoder for physiological signal modality; 、 、 are the trainable parameters of the private encoder.
3. The method for judging the emotional state of disabled people based on multimodal fusion according to claim 2 is characterized in that: The fourth step includes: Incorporating modality-specific features of audio, visual, and physiological 、 、 Splicing into a unified multimodal feature sequence : ; Use K-head attention to learn the correlation between modalities, concatenate the output features of each head, and then perform linear transformation to obtain cross-modal fusion features. : ; The specific features of each modality are finally represented through the attention mechanism and input into the classifier, which outputs the probability distribution of emotions and obtains the classification results.
Citation Information
Patent Citations
Multi-modal emotion recognition method for medical care robot
CN114724224A
Multimodal data-based method and system for recognizing cognitive engagement in classroom
US20250022314A1