Emotional state judgment method for disabled person based on multi-modal fusion
Through multimodal fusion technology, combining image, audio and physiological signal data, and adopting modal decoupling and attention mechanism, the problems of modal limitations and modeling limitations in emotion recognition for people with disabilities are solved, the accuracy of emotion recognition and the portability of equipment are improved, and real-time emotional companionship and early warning are achieved.
Patent Information
- Application Number
- CN202511130970.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing technologies have modal limitations, modeling limitations and equipment wearing comfort issues in emotion recognition for people with disabilities, making it difficult to achieve precise alignment and effective fusion of multimodal data, resulting in low emotion recognition accuracy and insufficient anti-interference ability.
A multimodal fusion strategy is adopted, combining image, audio and physiological signal data, feature extraction and fusion are performed through modal decoupling and attention mechanism, and integrated sensors are used for real-time monitoring to achieve accurate judgment of emotional state.
It improves the accuracy and robustness of emotion recognition, provides real-time emotional companionship, reduces the size and inconvenience of equipment, and achieves early warning and proactive prevention.
Smart Images

Figure CN120616533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for judging the emotional state of a disabled person based on multimodal fusion, and belongs to the technical field of multimodal fusion. Background Art
[0002] Emotion, as the bond that connects people to the world, is inextricably linked to human life. Therefore, the study of human emotion has become a focal area of interdisciplinary research: linguistics deconstructs the mechanisms of emotional expression, biology explores the physiological basis of emotion, and computer science strives to build frameworks for affective computing. While material needs are largely met, human spiritual needs are growing exponentially. This societal shift is driving progress in human-computer interaction. Scholars are increasingly recognizing the critical role that emotion plays in decision-making and maintaining social connections. This discovery has prompted them to explore integrating emotion into computing systems, aiming to enable computers to recognize, understand, and simulate human emotions. Emotion recognition is a crucial component of this quest, aiming to enable machines to capture the subtle connections between external emotional representations and internal emotional states, thereby accurately interpreting human emotions. In recent years, with the iterative advancements of deep learning algorithms and the exponential growth of computing resources, emotion recognition research has expanded beyond the laboratory and gradually permeated various application scenarios.
[0003] While the information technology revolution is driving the development of human-computer interaction, it is also opening up new possibilities for vulnerable groups in society. The 20th National Congress of the Communist Party of China proposed the important deployment of "improving the social security system and care service system for people with disabilities and promoting the comprehensive development of their cause," emphasizing the use of technology to support people with disabilities as a key measure of social progress. Facing the dual dilemma of physical limitations and social prejudice that people with disabilities have long faced, emotion recognition technology has demonstrated its significant value. By monitoring their physical and mental state in real time and providing intelligent care, it not only improves their quality of life but also helps them develop self-esteem, confidence, self-reliance, and independence, possessing significant social value.
[0004] Human perception of another person's emotional state is a complex process that integrates multiple types of information. Image information alone encompasses rich details across multiple dimensions, including facial expressions, body posture, and body movements. However, in the field of emotion recognition, early research focused on processing unimodal data, which often resulted in an inability to fully and deeply capture the complexity of emotional information. Furthermore, it was particularly vulnerable to noise interference and exhibited significant limitations. In fact, different modalities exhibit varying sensitivities and specificities when reflecting emotional states. Therefore, compared to single modal analysis, multimodal analysis can more accurately and comprehensively reflect an individual's emotional state, becoming an effective approach to improving the accuracy and robustness of emotion recognition.
[0005] Multimodal emotion recognition technology combines multiple sensors and input methods to identify an individual's emotional state by analyzing voice, facial expressions, physiological signals, text, and other information. The hardware components involved primarily include sensors, data acquisition devices, and data processing units.
[0006] Similar implementations include: 1. Integrated wearable device: This device integrates a heart rate monitor, skin conductance sensor, and a micro microphone into a single wearable unit, such as a smartwatch or chest strap.
[0007] 2. High-resolution camera: used to capture facial expressions and body movements, usually connected to a data processing unit to analyze facial expressions in real time.
[0008] 3. Data synchronization technology: Use wireless communication technology (such as Bluetooth or Wi-Fi) to achieve synchronous transmission between sensor data and data processing units.
[0009] 4. Data processing and analysis software: This software is capable of processing data from different sensors and using the derived algorithms to identify and classify emotional states.
[0010] Through an in-depth analysis of the implementation schemes similar to the present invention in the prior art, it is found that each scheme often makes breakthroughs in a single dimension, but it is difficult to take into account the optimization of the overall performance of the system. At the level of data acquisition and preprocessing, although there are studies that have used head-mounted integrated devices to achieve the synchronous acquisition of multimodal data such as voice, facial video, etc., it still faces many challenges in practical applications. In the feature extraction stage, existing solutions often fall into a dilemma: one type of solution performs well in single-modal feature extraction and analysis, but these solutions perform poorly in multimodal feature fusion, and it is difficult to achieve cross-modal collaborative optimization; the other type of solution designs a complex multimodal fusion architecture, but ignores the feature mining of each single modality in the pursuit of fusion effect, resulting in limited improvement in overall performance. In addition, in the most critical modal fusion stage, although some advanced solutions try to integrate convolutional networks and attention mechanisms to improve performance, their core fusion strategies still remain at the simple feature splicing level. They neither fully consider the dynamic interaction relationship between different modalities nor fully retain the modal specificity. As a result, they fail to effectively distinguish similar emotions, resulting in a sharp drop in the robustness of the model when facing complex scenarios.
[0011] Existing technologies for emotion recognition and mental health support for people with disabilities have several key deficiencies that need to be addressed. This paper proposes systematic solutions to these problems: First, there are significant deficiencies in the service system and technical implementation. my country's large population of people with disabilities faces severe mental health challenges, with depression, anxiety, and other psychological issues prevalent. However, the existing mental health support system has significant shortcomings: it primarily relies on passive mental health lectures, which cannot meet the urgent need for emotional companionship among people with disabilities who have low social participation; nursing services often focus only on physical care and neglect psychological care, and the cost of professional psychological support is high; more importantly, the current system often intervenes only after psychological problems manifest, lacking effective early warning and real-time monitoring mechanisms. Regarding technical support, existing solutions have multiple limitations: physiological monitoring devices are not very comfortable to wear and are not yet touchless, making them difficult to use over the long term; multimodal emotion recognition systems typically require the coordination of multiple independent hardware devices, which are not only bulky and inconvenient to use, but also face problems such as data synchronization, causing significant inconvenience to users' daily lives. This invention not only provides real-time intelligent emotional companionship for people with disabilities, meeting users' emotional needs through human-computer interaction and effectively alleviating the human resource pressure of traditional manual care, but also enables emotion detection, using emotion recognition technology to keenly capture emotional fluctuations, issuing warnings and notifying relevant personnel when emotions are abnormal, achieving a shift from passive response to proactive prevention. Furthermore, the system ensures the precise alignment and effective integration of multimodal data, providing a more accurate analytical foundation for emotion recognition.
[0012] Second, existing emotion recognition technologies suffer from modality limitations. Traditional solutions are often limited to relying on single-modal data (e.g., visual facial expression analysis or speech-only emotion recognition). This simplistic approach is inconsistent with the complexity and diversity of human emotional expression. In reality, emotion recognition in real-world scenarios requires integrating multidimensional information: at the behavioral level, facial expressions, body language, and voice intonation all convey rich emotional information; at the physiological level, autonomic nervous system indicators such as galvanic skin response provide an objective reflection of internal emotional states. Single-modal systems not only fail to capture this multi-layered emotional expression but are also vulnerable to information loss in real-world scenarios (e.g., facial occlusion and ambient noise), resulting in a sharp drop in recognition accuracy and significantly insufficient interference immunity. In practical applications that demand superior overall performance, single-modal recognition strategies are no longer sufficient. Compared to single-modal data, multimodal data can more comprehensively reflect an individual's emotional state, providing richer and more accurate information. Therefore, the present invention combines multimodal data such as images, audio, and physiological signals for emotion recognition, which can improve the accuracy of emotion recognition and provide reliable technical support for practical application scenarios.
[0013] Third, in terms of modal fusion, existing technologies suffer from significant modeling limitations. Current mainstream multimodal emotion recognition technologies often employ simple fusion strategies such as feature concatenation or weighted averaging, failing to fully exploit the consistency and heterogeneity inherent in multimodal data. These limitations manifest themselves in two key areas: First, some methods overly focus on extracting common features across multiple modalities while ignoring the unique emotional expression strengths of individual modalities. For example, facial expressions excel at capturing transient emotional changes, voice features can reflect changes in emotional intensity, and physiological signals more objectively reflect ongoing emotional states. Second, existing technologies lack the ability to model complex interactions between modalities. Consequently, when distinguishing similar emotions (such as anger and fear), emotion recognition models often only capture surface features and fail to identify underlying differences, resulting in high recognition confusion rates. To address this key challenge, this paper proposes a multimodal decoupling strategy that integrates the consistency and heterogeneity of multimodal data into a unified framework, enabling learning of both modal consistency and heterogeneity. This innovation not only solves the problem of existing multimodal emotion recognition methods ignoring the capture of emotions by a single modality, but also provides a better technical path for fine-grained emotion recognition in more complex scenarios.
[0014] In summary, the present invention aims to address the shortcomings of the existing technology and propose a more comprehensive, efficient, real-time and highly accurate multimodal emotion recognition system, method and device to better meet the mental health needs of the disabled and other special groups with emotional companionship needs. Summary of the Invention
[0015] The present invention designs and develops a method for judging the emotional state of disabled people based on multimodal fusion, which judges the emotional state based on multimodal information and improves the accuracy of emotion recognition.
[0016] The technical solution provided by the present invention is: A method for judging the emotional state of disabled people based on multimodal fusion, comprising: Step 1: Obtain the user's audio, video and physiological signal raw data through sensors; pre-process the audio, video and physiological signal data; Step 2: Identify the original data and extract audio feature data, visual feature data, and physiological signal feature data; Among them, audio features include the time information of speech, visual feature data includes facial images, and physiological signal data includes heart rate and skin electrical response data; Step 3: applying a modal decoupling strategy to the audio feature data, visual feature data, and physiological signal feature data to achieve learning of modal consistency and heterogeneity; Step 4: The features of different modalities are integrated through the attention mechanism, and emotion recognition is performed based on multimodal features to finally obtain the emotion change recognition result; Emotions are divided into eight categories: neutral, calm, happy, sad, angry, fearful, disgusted and surprised.
[0017] Preferably, the step 1 includes: For audio data, the preprocessing process includes: dividing each audio descriptor into t time segments to make it meet the input requirements of the convolutional neural network model; For image data, the preprocessing process includes: dividing the video into t parts and randomly sampling short segments of k consecutive frames from each segment, and using these t short segments as the data representation of the entire video to make it meet the input requirements of the 3D convolutional neural network model; Among them, each segment has k consecutive frames; For ECG data, the preprocessing process includes: using wavelet decomposition and reconstruction method to perform ECG filtering on the original data, and then filtering out the RR heart rate signal by R peak positioning; For skin electrical data, the preprocessing process includes: filtering out noise from the original data through wavelet denoising; and normalizing the physiological signals in batches.
[0018] Preferably, the step 2 includes: Based on the 2D-CNN neural network model, audio features are extracted in combination with the attention mechanism; The audio features are mainly based on Mel frequency cepstral coefficients; Extract video frame features from the input video frames based on the 3D-CNN neural network model; The extracted audio features enter the temporal attention module, and the extracted video frame features enter the spatial attention module, channel attention module, and temporal attention module: In the spatial attention module, the input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the spatial dimension. The attention weights are normalized by the softmax function and applied to the original spatial features to generate weighted spatial features. In the channel attention module, the video frame feature matrix is transposed into a channel feature matrix, where the spatial dimension is interchanged with the channel dimension.
[0019] The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the channel dimension; the attention weights are normalized by the softmax function and applied to the original channel features to generate weighted channel features; The video features output by the channel attention module are first spatially averaged through the pooling layer to obtain a temporal feature sequence; The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the time dimension. The features are activated by the ReLU function to calculate the temporal attention weights, which are then weighted and fused with the original features to obtain the time-weighted features. Perform average pooling to reduce feature dimensions and smooth feature information; Extract features of input physiological signals based on 1D-CNN and LSTM; ; ; in, is the original feature of the physiological signal at the tth time step, R is the 1D-CNN receptive field, that is, the number of time steps covered by the convolution kernel, is the 1D-CNN convolution kernel parameter, b is the bias, For 1D-CNN at the tth time step, the activation function After the output, is the hidden state of LSTM at the tth time step, is the hidden state at the t-1th time step, which is used to capture the temporal dependency of physiological signals; In each temporal attention module, the input features are passed through the Conv1d convolution layer to extract local temporal patterns, and then the correlation features in the time dimension are generated through the fully connected layer. The temporal attention weights are then calculated using the ReLU activation function and weightedly fused with the original features to generate a time-weighted feature representation. Perform average pooling to reduce feature dimensions and smooth feature information; Cascade the features from different physiological signals to form a joint feature representation; The joint features are input into the fully connected layer for further processing to generate the final physiological emotion features.
[0020] Preferably, the step three includes: Use the modal decoupling module to perform modal decoupling strategy; Among them, the modal decoupling module includes: a modal shared encoder and three modal private encoders; The extracted audio features , image features , physiological signal characteristics Projected into the modality-sharing encoder of the modality-sharing space, the shared features of the modalities are extracted, including: ; ; ; Where, For audio sharing features, Shared features for images, Sharing features for physiological signals It is a modality-sharing encoder used to extract common features of different modalities; To share the trainable parameters of the encoder, all modalities share the same set of parameters and learn the consistency of the modalities; The extracted audio, image, and physiological signal features are projected into the modality-specific encoders of their respective modality-specific spaces to extract modality-specific features, including: ; ; ; in, is the audio modality specific feature, is the image modality specific feature, Specific features for physiological signal modalities: Audio modal private editor, Private editor for video modal, Private encoder for physiological signal modality; 、 、 are the trainable parameters of the private encoder.
[0021] Preferably, the step 4 includes: Integrate audio, visual, and physiological modality-specific features 、 、 Splicing into a unified multimodal feature sequence : ; Use K-head attention to learn the correlation between modalities, concatenate the output features of each head, and then perform linear transformation to obtain cross-modal fusion features. : ; The specific features of each modality are finally represented through the attention mechanism and input into the classifier, which outputs the probability distribution of emotions and obtains the classification results.
[0022] The beneficial effects of the present invention are: 1. This invention uses multimodal emotion recognition to effectively overcome the potential shortcomings of a single modality in expressing emotion. The interaction and complementarity between different modalities enables the invention to extract more comprehensive and accurate emotion features, thereby improving the accuracy and reliability of emotion recognition.
[0023] 2. This paper adopts a multimodal information fusion strategy based on the attention mechanism, which realizes the deep mining of emotional features through the synergy of multiple channels.
[0024] 3. The present invention includes a modal decoupling module for distinguishing similar emotions. By capturing the consistency and heterogeneity of different modalities, the modal shared features and modal specific features of each modality are extracted, effectively improving the accuracy of emotion recognition.
[0025] 4. This invention utilizes a data acquisition device integrated with multiple sensors to monitor multiple health indicators of the target subject in real time, adding the ability to extract physiological signals and detect emotional changes. Furthermore, this invention possesses data-driven self-adjustment and optimization capabilities, leveraging data uploaded by the acquisition device to continuously refine the prediction model, thereby gradually improving the accuracy and adaptability of the predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a flowchart of the method for judging the emotional state of disabled people based on multimodal fusion described in the present invention.
[0027] Figure 2 Schematic diagram of the audio feature extraction model described in the present invention.
[0028] Figure 3 Schematic diagram of the video face image feature extraction model described in the present invention.
[0029] Figure 4 Schematic diagram of the physiological emotion feature extraction model described in the present invention.
[0030] Figure 5 This is a flow chart of the modal decoupling strategy described in the present invention.
[0031] Figure 6 This is a comparison chart of the accuracy and F1 score results described in the present invention. DETAILED DESCRIPTION
[0032] The present invention will be described in further detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.
[0033] like Figure 1-6As shown, the present invention provides a method for determining the emotional state of disabled people based on multimodal fusion. This method is suitable for users with disabilities who have limited mobility due to leg disabilities or other reasons, but whose facial expressions and voice functions are normal. Multimodal information is used to determine their emotional state, thereby improving the accuracy of emotion recognition. Because this method requires the collection of information such as facial expressions and voice intonation, it is more suitable for users with limited mobility but whose physiological functions are not significantly affected. It can also be expanded to other users with similar characteristics. Specifically, it includes: Step 1: Obtain the user's audio, video and physiological signal raw data through sensors; pre-process the audio, video and physiological signal data; For audio data, the preprocessing process includes: dividing each audio descriptor into t time segments to meet the input requirements of the convolutional neural network model; for image data, the preprocessing process includes: dividing the video into t parts and randomly sampling short segments of k consecutive frames from each segment, and using these t short segments as the data representation of the entire video to meet the input requirements of the 3D convolutional neural network model; where each segment has k consecutive frames; For ECG data, the preprocessing process includes: using wavelet decomposition and reconstruction method to perform ECG filtering on the original data, and then filtering out the RR heart rate signal by R peak positioning; For skin electrical data, the preprocessing process includes: filtering out noise from the original data through wavelet denoising; and normalizing the physiological signals in batches.
[0034] Step 2: Identify the original data and extract audio feature data, visual feature data, and physiological signal feature data; Among them, audio features include the time information of speech, visual feature data includes facial images, and physiological signal data includes heart rate and skin electrical response data; Based on the 2D-CNN neural network model, audio features are extracted in combination with the attention mechanism; in, The audio features are mainly based on Mel frequency cepstral coefficients; The input video frame is subjected to video frame feature extraction based on the 3D-CNN neural network model, and the feature matrix in the following form is obtained after extraction: ; Among them, t is the number of video clips, which corresponds to the total number of clips obtained after the video is divided according to certain rules; m is determined by the spatial size of the video frame feature map (after flattening the height h and width w, that is, m=h×w), which represents the number of locations of the feature after flattening the spatial dimension; n is the number of feature channels. For any vector in the feature matrix , which is the visual feature vector corresponding to the j-th spatial position in the i-th segment.
[0035] The extracted audio features enter the temporal attention module, and the extracted video frame features enter the three attention mechanism modules: Spatial attention module, used to extract the importance of features in spatial dimensions; In the spatial attention module, the input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the spatial dimension. The attention weights are normalized by the softmax function and applied to the original spatial features to generate weighted spatial features. ; ; ; in, 、 is the trainable parameter matrix of the spatial attention module, which is used to input the i-th video clip features (Taken from Perform linear transformation on the corresponding fragments); is the intermediate feature of spatial attention obtained by transformation; Yes The spatial attention weight obtained after applying the Softmax function normalization is used to measure the importance of different spatial position features in the i-th video clip; To pass the attention weight and the original segment features Multiply element by element to obtain the weighted spatial features, Transpose the matrix.
[0036] Channel attention module: used to extract the importance of features in the channel dimension; First, Transpose , that is, the channel feature matrix of the video frame: ; For any vector in the matrix , which is the feature vector corresponding to the j-th channel position in the i-th segment.
[0037] The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the channel dimension. The attention weights are normalized by the softmax function and applied to the original channel features to generate weighted channel features. ; ; ; in, 、 is the trainable parameter matrix of the channel attention module, which is used to transform the channel features of the i-th segment (Taken from corresponding fragment); is the intermediate feature of channel attention; Yes The channel attention weight obtained after applying the Softmax function normalization is used to measure the importance of different channel features in the i-th video clip; is the weighted channel feature, which is the combination of attention weight and original channel feature Element-wise multiplication.
[0038] Temporal attention module: used to extract the importance of features in the temporal dimension.
[0039] The video features output by the channel attention module are first spatially averaged through the pooling layer to obtain a temporal feature sequence. (t is the number of video clips, n is the feature dimension), which is used as the input of temporal attention; The input features first pass through the Conv1d convolution layer and then through the fully connected layer to generate correlation features in the time dimension. The features are activated by the ReLU function to calculate the temporal attention weights, and are weighted and fused with the original features to obtain the time-weighted features.
[0040] ; ; ; in, is the temporal feature sequence of audio or video, is the feature vector of the jth time series segment; 、 is the trainable parameter matrix of the temporal attention module; is the intermediate feature of temporal attention; Yes The temporal attention weight obtained after processing by the ReLU activation function is used to measure the importance of features of different temporal segments; It is the time-weighted feature obtained by summing the features of each time sequence segment according to the attention weight.
[0041] Perform average pooling to reduce feature dimensions and smooth feature information; Extract features of input physiological signals based on 1D-CNN and LSTM; ; ; in, is the original feature of the physiological signal at the tth time step; R is the 1D-CNN receptive field, that is, the number of time steps covered by the convolution kernel; is the 1D-CNN convolution kernel parameter, b is the bias; For 1D-CNN at the tth time step, the activation function The output after is the hidden state of LSTM at the tth time step, is the hidden state at the t-1th time step, which is used to capture the temporal dependencies of physiological signals.
[0042] In each temporal attention module: the input features first pass through the Conv1d convolutional layer to extract local temporal patterns, then pass through the fully connected layer to generate correlation features in the time dimension, and then use the ReLU activation function to calculate the temporal attention weights and weighted fusion with the original features to generate the time-weighted feature representation; Perform average pooling to reduce feature dimensions and smooth feature information; Cascade fusion of features from different physiological signals, that is, splicing the features of ECG signals and skin electrical signals according to specific dimensions to form a joint feature representation; The joint features are input into the fully connected layer for further processing to generate the final physiological emotion features.
[0043] Step 3: applying a modal decoupling strategy to the audio feature data, visual feature data, and physiological signal feature data to achieve learning of modal consistency and heterogeneity; A modal decoupling module is used to implement the modal decoupling strategy. The modal decoupling module includes: a modal shared encoder and three modal private encoders; The audio features extracted in step 2 , image features , physiological signal characteristics Projected into the modality shared encoder of the modality shared space to extract the shared features of the modalities 、 、 : ; ; ; Where, For audio sharing features, Shared features for images, For physiological signal sharing features, It is a modality-sharing encoder used to extract common features of different modalities; To share the trainable parameters of the encoder, all modalities share the same set of parameters and learn the consistency of the modalities.
[0044] Project the extracted audio, image, and physiological signal features into the modality-private encoder of each modality-private space to extract modality-specific features 、 、 : ; ; ; Where, is the audio modality specific feature, is the image modality specific feature, Specific features for physiological signal modalities: Audio modal private editor, Private editor for video modal, Private encoder for physiological signal modality; 、 、 It is a trainable parameter of the private encoder, which is trained independently for each modality to learn the heterogeneity of the modality; Step 4: The features of different modalities are integrated through the cross-modal attention mechanism, and emotion recognition is performed based on multimodal features to finally obtain the emotion change recognition result; Integrate audio, visual, and physiological modality-specific features 、 、 Splicing into a unified multimodal feature sequence : ; Use K-head attention to learn the correlation between modalities, concatenate the output features of each head, and then perform linear transformation to obtain cross-modal fusion features. : ; The specific features of each modality are finally represented through the attention mechanism and input into the classifier, which outputs the probability distribution of emotions and obtains the classification results.
[0045] like Figure 6 As shown in the figure, the accuracy and F1-Score comparison between the baseline method and this method are shown. It can be seen from the figure that by fusing multimodal information for emotion recognition, this method outperforms the baseline method in both accuracy and F1-Score.
[0046] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A method for judging the emotional state of disabled people based on multimodal fusion, characterized in that: include: Step 1: Obtain the user's audio, video and physiological signal raw data through sensors; pre-process the audio, video and physiological signal data; Step 2: Identify the original data and extract audio feature data, visual feature data, and physiological signal feature data; Among them, audio features include time information of speech, visual feature data include face images, and physiological signal data include heart rate and skin electrical response data; Step 3: applying a modal decoupling strategy to the audio feature data, visual feature data, and physiological signal feature data to achieve learning of modal consistency and heterogeneity; Step 4: The features of different modalities are integrated through the attention mechanism, and emotion recognition is performed based on multimodal features to finally obtain the emotion change recognition result; Emotions are divided into eight categories: neutral, calm, happy, sad, angry, fearful, disgusted and surprised.
2. The method for judging the emotional state of disabled people based on multimodal fusion according to claim 1 is characterized in that: The step one comprises: For audio data, the preprocessing process includes: dividing each audio descriptor into t time segments to make it meet the input requirements of the convolutional neural network model; For image data, the preprocessing process includes: dividing the video into t parts and randomly sampling short segments of k consecutive frames from each segment, and using these t short segments as the data representation of the entire video to make it meet the input requirements of the 3D convolutional neural network model; Among them, each segment has k consecutive frames; For ECG data, the preprocessing process includes: using wavelet decomposition and reconstruction method to perform ECG filtering on the original data, and then filtering out the RR heart rate signal by R peak positioning; For skin electrical data, the preprocessing process includes: filtering out noise from the original data through wavelet denoising; and normalizing the physiological signals in batches.
3. The method for judging the emotional state of disabled people based on multimodal fusion according to claim 2 is characterized in that: The second step includes: Based on the 2D-CNN neural network model, audio features are extracted in combination with the attention mechanism; The audio features are mainly based on Mel frequency cepstral coefficients; Extract video frame features from the input video frames based on the 3D-CNN neural network model; The extracted audio features enter the temporal attention module, and the extracted video frame features enter the spatial attention module, channel attention module, and temporal attention module: In the spatial attention module, the input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the spatial dimension. The attention weights are normalized by the softmax function and applied to the original spatial features to generate weighted spatial features. In the channel attention module, the video frame feature matrix is transposed into the channel feature matrix, where the spatial dimension is interchanged with the channel dimension; The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the channel dimension; the attention weights are normalized by the softmax function, and weighted channel features are generated based on the original channel features; The video features output by the channel attention module are first spatially averaged through the pooling layer to obtain a temporal feature sequence; The input features are first processed by the Conv1d convolution layer and then by the fully connected layer to generate correlation features in the time dimension. The features are activated by the ReLU function to calculate the temporal attention weights, which are then weighted and fused with the original features to obtain the time-weighted features. Perform average pooling to reduce feature dimensions and smooth feature information; Extract features of input physiological signals based on 1D-CNN and LSTM; ; ; in, is the original feature of the physiological signal at the tth time step, R is the 1D-CNN receptive field, that is, the number of time steps covered by the convolution kernel, is the 1D-CNN convolution kernel parameter, b is the bias, For 1D-CNN at the tth time step, the activation function After the output, is the hidden state of LSTM at the tth time step, is the hidden state at the t-1th time step, which is used to capture the temporal dependency of physiological signals; In each temporal attention module, the input features are passed through the Conv1d convolution layer to extract local temporal patterns, and then the correlation features in the time dimension are generated through the fully connected layer. The temporal attention weights are then calculated using the ReLU activation function and weightedly fused with the original features to generate a time-weighted feature representation. Perform average pooling to reduce feature dimensions and smooth feature information; Cascade the features from different physiological signals to form a joint feature representation; The joint features are input into the fully connected layer for further processing to generate the final physiological emotion features.
4. The method for judging the emotional state of disabled people based on multimodal fusion according to claim 3 is characterized in that: The step three includes: Use the modal decoupling module to perform modal decoupling strategy; Among them, the modal decoupling module includes: a modal shared encoder and three modal private encoders; The extracted audio features , image features , physiological signal characteristics Projected into the modality-sharing encoder of the modality-sharing space, the shared features of the modalities are extracted, including: ; ; ; Where, For audio sharing features, Shared features for images, For physiological signal sharing features, Shared encoder for modalities; are the trainable parameters of the shared encoder; The extracted audio, image, and physiological signal features are projected into the modality-specific encoders of their respective modality-specific spaces to extract modality-specific features, including: ; ; ; Where, is the audio modality specific feature, is the image modality specific feature, Specific features for physiological signal modalities: Audio modal private editor, Private editor for video modal, Private encoder for physiological signal modality; 、 、 are the trainable parameters of the private encoder.
5. The method for judging the emotional state of disabled people based on multimodal fusion according to claim 4 is characterized in that: The fourth step includes: Incorporating modality-specific features of audio, visual, and physiological 、 、 Splicing into a unified multimodal feature sequence : ; Use K-head attention to learn the correlation between modalities, concatenate the output features of each head, and then perform linear transformation to obtain cross-modal fusion features. : ; The specific features of each modality are finally represented through the attention mechanism and input into the classifier, which outputs the probability distribution of emotions and obtains the classification results.
Citation Information
Patent Citations
Multi-modal emotion recognition method for medical care robot
CN114724224A
Multi-modal sentiment analysis method combining pre-training model and self-attention block
CN118898046A
Multimodal data-based method and system for recognizing cognitive engagement in classroom
US20250022314A1
Method and system for early diagnosis of parkinson's disease based on multimodal deep learning
US20250213174A1