Emotional state evaluation method and system based on different space-time multi-modal information fusion

By integrating hetero-spatial multimodal information and using neural networks and psychological models, accurate monitoring and analysis of the emotional state of key monitoring objects is solved, and the problems of hetero-spatial data processing and continuous emotional changes in multi-modal emotion recognition are improved, and the comprehensiveness and robustness of emotional recognition are improved.

CN120337026AInactive Publication Date: 2025-07-18NANTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510357369.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multimodal emotion recognition technology has limitations in processing hetero-spatial-temporal data and describing continuous changes in emotions, and it is difficult to comprehensively and accurately capture and analyze multimodal information, especially the neglect of body posture modals and insufficient processing of hetero-spatial-temporal data.

Method used

The expression and posture features in the video information are extracted by a dual-channel convolutional neural network for first-level fusion, combined with the Gaussian hybrid model and deep neural network for emotional classification of audio information, and weighted through the selective attention weighting mechanism, and used a full connection layer to achieve secondary fusion in different time and space. The PAD model is used for multi-dimensional description, and an emotional change trend curve is generated through the dynamic time bending method.

Benefits of technology

It realizes accurate monitoring and analysis of the emotional state of key monitoring objects, can process multimodal information at different times and places, improves the robustness and continuity of emotional recognition, enhances the accuracy of abnormal behavior detection and the intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337026A_ABST
    Figure CN120337026A_ABST
Patent Text Reader

Abstract

The invention discloses an emotional state evaluation method and system based on different space-time multi-modal information fusion, and belongs to the technical field of emotion recognition and monitoring. According to the technical scheme, the method comprises the steps that S1, expression and posture features in video information are extracted through a two-channel convolutional neural network, and primary fusion is carried out; s2, performing sentiment classification on voice signals in the audio information by using a Gaussian mixture model and a deep neural network; s3, inputting the audio emotion level as a bias item into a full connection layer, and realizing second-level fusion of multi-modal information in a different time space; s4, performing multi-dimensional description on the emotion by adopting a PAD emotion model to generate an emotion description curve; and S5, processing the emotional description curve of each dimension by using a dynamic time bending method to generate an emotional change trend curve. The beneficial effect of the invention is that accurate monitoring and analysis of the emotional state of a key monitoring object can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition and monitoring, and in particular to an emotion state assessment method and system based on heterotemporal and spatial multimodal information fusion. Background Art

[0002] Emotion recognition technology is to infer the emotional state of an individual by analyzing human voice, facial expressions, body posture and other modal information. In recent years, with the rapid development of artificial intelligence and machine learning technology, emotion recognition technology has been widely used in many fields, such as human-computer interaction, mental health monitoring, security monitoring, etc.

[0003] Early emotion recognition was mainly based on single-modal information, such as speech, facial expressions, or body postures. However, single-modal emotion recognition has obvious limitations. For example, speech emotion recognition is easily disturbed by environmental noise, facial expression recognition may fail due to lighting conditions or occlusion, and body posture recognition has difficulty capturing subtle emotional changes. In order to overcome the limitations of single-modal emotion recognition, researchers began to explore multimodal emotion recognition technology, that is, to improve the accuracy and robustness of emotion recognition by integrating information from multiple modalities (such as speech, facial expressions, and body postures). The core idea of multimodal emotion recognition is to use the complementarity between different modalities to make up for the shortcomings of a single modality.

[0004] Although multimodal emotion recognition has advantages in theory, it still faces many challenges in practical applications: 1) Data of different modalities have different feature representations and spatiotemporal characteristics. How to effectively fuse these heterogeneous data is an important research issue; 2) In practical applications, multimodal data may come from different times and spaces. This "different time and space" problem increases the difficulty of multimodal information fusion; 3) Traditional emotion recognition methods usually divide emotions into discrete categories (such as anger, happiness, sadness, etc.), ignoring the continuity and dynamic changes of emotions. Such continuous changes are difficult to accurately describe through discrete categories.

[0005] At present, the research on multimodal emotion recognition mainly focuses on the fusion of audio and video data in the same space and time, and is mostly limited to the fusion of two modes (such as speech and facial expressions). However, these methods have the following limitations: 1) Body posture is one of the important modes of emotional expression, but most existing studies have ignored this mode, resulting in insufficient comprehensiveness of emotion recognition; 2) Existing methods usually assume that multimodal data comes from the same space and time, and it is difficult to process multimodal information from different times and places. This kind of "different space and time" data is very common in practical applications, such as multi-camera data in surveillance systems; 3) Most existing methods use discrete emotion category division, which is difficult to describe the continuous change and temporal nature of emotions.

[0006] Although multi-modal emotion recognition faces many challenges, some research has attempted to address these issues in recent years: 1) Researchers have proposed various multi-modal information fusion methods, such as early fusion, mid-term fusion, and late fusion; 2) Deep learning technologies (such as convolutional neural networks and recurrent neural networks) have been widely applied in multi-modal emotion recognition; 3) To describe the continuous change of emotions, researchers have proposed various continuous emotion representation models, such as the PAD (Pleasure-Arousal-Dominance) model, which describes the emotional state through three dimensions (pleasure, arousal, dominance) and can depict the continuous change of emotions more precisely.

[0007] Emotion recognition technology has evolved from single-modal to multi-modal. Although certain progress has been made, it still faces challenges such as inter-modal heterogeneity, spatio-temporal inconsistency, and discrete emotion representation.

[0008] How to solve the above technical problems is the subject of this invention. Summary of the Invention

[0009] The object of the present invention is to provide an emotion state evaluation method and system based on cross-time and space multi-modal information fusion, which can achieve precise monitoring and analysis of the emotion states of key monitored objects; the present invention realizes dynamic and continuous evaluation of emotion states by fusing multi-modal information from different times and different locations, combining neural networks and psychological models.

[0010] The inventive concept of the present invention is as follows: The present invention provides an emotion state evaluation method and system based on cross-time and space multi-modal information fusion. First, an expression and gesture feature in video information is extracted through a dual-channel convolutional neural network and subjected to primary fusion. Then, a Gaussian mixture model and a deep neural network are used to perform emotion classification on the speech signal in audio information, and the emotion classification result is weighted through a selective attention weighting mechanism to obtain an audio emotion level. After that, the audio emotion level is input as a bias term into a fully connected layer to achieve secondary fusion of multi-modal information across time and space. Subsequently, the PAD emotion model is used to describe the emotion in multiple dimensions to generate an emotion description curve. Finally, the dynamic time warping method is used to process each dimension of the emotion description curve to generate an emotion change trend curve.

[0011] To achieve the above object of the invention, the technical solution adopted by the present invention is specifically as follows: An emotion state evaluation method based on cross-time and space multi-modal information fusion, comprising the following steps:

[0012] 1.1 Extract the expression and gesture features in video information through a dual-channel convolutional neural network and perform primary fusion;

[0013] 1.2. Use the Gaussian mixture model and the deep neural network to perform sentiment classification on the speech signals in the audio information, and weight the sentiment classification results through a selective attention weighting mechanism to obtain the audio sentiment level;

[0014] 1.3. Take the audio sentiment level result as a bias term and input it into the fully connected layer to achieve the secondary fusion of multi-modal information in different time and space;

[0015] 1.4. Adopt the PAD (Pleasure-Arousal-Dominance) model to describe the sentiment in multiple dimensions and generate the sentiment description curves in three dimensions;

[0016] 1.5. Use the dynamic time warping method (DTW) to process the sentiment description curves in each dimension, generate the sentiment change trend curve, and realize the evaluation of the emotional state of the key monitoring object.

[0017] Further, the step 1.1 specifically includes the following steps:

[0018] 2.1. Use a 3×3 convolution kernel, a 1×1 convolution kernel, and max pooling to extract the facial expression features in the video information;

[0019] 2.2. Use a 3×3 convolution kernel, a 1×1 convolution kernel, and max pooling to extract the posture features in the video information;

[0020] 2.3. Use the weighted average method to perform the primary fusion on the extracted facial expression and posture features.

[0021] Further, the step 1.2 specifically includes the following steps:

[0022] 3.1. Use the Gaussian mixture model to model the multi-resolution spectral features of the enhanced speech signal;

[0023] 3.2. Based on the prosodic features of the enhanced speech signal, classify the speech signal from negative emotion to positive emotion into M levels;

[0024] 3.3. Calculate the syllable state transition probability matrix for the segmented audio segments, and use it as the input of the bidirectional long short-term memory network to complete the sensitive word detection;

[0025] 3.4. When a sensitive word appears, set different weights according to the different sensitive words, and positively weight the speech signals at M levels to obtain the audio sentiment level.

[0026] Further, the step 1.4 specifically includes the following steps:

[0027] 4.1. Combine the time change, use the time t as the horizontal axis, and generate the sentiment description curve in the pleasure dimension, and its formula is f P =(P1, P2, PN );

[0028] 4.2. Generate an emotional description curve for the arousal dimension with time t as the horizontal axis in combination with the change in time. Its formula is f A = (A1, A2,, A N );

[0029] 4.3. Generate an emotional description curve for the dominance dimension with time t as the horizontal axis in combination with the change in time. Its formula is f D = (D1, D2,, D N ).

[0030] Further, step 1.5 specifically includes the following steps:

[0031] 5.1. For the pleasure dimension sequence f P = (P1, P2, P N ), the arousal dimension sequence f A = (A1, A2,, A N ), construct an N×N distance matrix Each element in the matrix represents the Euclidean distance between P i and A j ;

[0032] 5.2. Construct an accumulated distance matrix where each element represents the minimum accumulated distance from the starting point (1, 1) to the point (i, j). The calculation formula for the accumulated distance matrix is

[0033]

[0034] 5.3. Starting from the lower right corner of the accumulated distance matrix , backtrack to the upper left corner according to the principle of selecting the minimum adjacent point of the current point, and find the shortest curved path between f P and f A , denoted as f S = (S1, S2, S N );

[0035] 5.4. For the shortest curved path f S = (S1, S2, S N ), the dominance dimension sequence f D = (D1, D2,, D N ), construct an N×N distance matrix Each element in the matrix represents the Euclidean distance between S i and D j ;

[0036] 5.5. Construct a cumulative distance matrix where each element represents the minimum cumulative distance from the starting point (1, 1) to the point (i, j). The calculation formula for the cumulative distance matrix is

[0037]

[0038] 5.6. Starting from the lower right corner of the cumulative distance matrix and backtracking to the upper left corner according to the principle of selecting the minimum adjacent point of the current point, find the shortest bending path between f S and f D and denote it as f B =(B1, B2, B N );

[0039] 5.7. Finally, output an emotional change trend curve with time as the horizontal axis and the shortest bending path as the vertical axis.

[0040] Compared with the prior art, the present invention has significant advantages in terms of the comprehensiveness and robustness of multimodal information fusion, the continuity and dynamics of emotional representation, the accuracy of abnormal behavior detection, and the intelligent level of the system. These advantages enable the present invention to have a wide range of application prospects in the fields of security monitoring, mental health monitoring, intelligent interaction systems, etc. Specifically, the beneficial effects of the present invention are as follows:

[0041] (1) By fusing information of multiple modalities such as voice, facial expressions, and body postures, the present invention can capture the emotional state of the monitored object more comprehensively and accurately; the present invention can process multimodal information from different times and different locations, overcoming the limitation that traditional methods can only process data in the same time and space, and improving the robustness of emotional recognition;

[0042] (2) The present invention uses the PAD (Pleasure - Arousal - Dominance) model to describe emotions in multiple dimensions, and can more finely depict the continuous change and time series of emotions; the present invention processes the emotional description curves of each dimension through the DTW method to generate an emotional change trend curve, which can reflect the change of the emotional state of the monitored object in real time (such as sudden emotional fluctuations, being in a negative emotion for a long time, etc.);

[0043] (3) The present invention introduces a selective attention weighting mechanism in speech emotion recognition, which can effectively detect sensitive words and perform weighted processing on the emotion recognition results, avoiding the monitored object from deliberately modifying the speech emotion and improving the accuracy of abnormal behavior detection;

[0044] (4) The present invention uses deep learning technologies (such as convolutional neural networks, deep neural networks) to automatically extract the features of multimodal data and realizes emotion recognition through end - to - end learning, reducing manual intervention and improving the intelligent level of the system. Brief Description of the Drawings

[0045] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.

[0046] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0047] Figure 2 It is a network diagram of the expression and gesture feature extraction and primary fusion process in video information;

[0048] Figure 3 It is a framework diagram of the emotional level classification process in audio information;

[0049] Figure 4 It is a network diagram of the secondary fusion process of multi-modal information;

[0050] Figure 5 It is a schematic diagram of using the PAD model and the DTW method to evaluate the emotional state;

[0051] Figure 6 It is the experimental effect diagram of the present invention;

[0052] Figure 7 It is a schematic diagram of the system structure of the present invention. Detailed Embodiment

[0053] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. Of course, the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0054] Embodiment 1

[0055] Refer to Figures 1 to 5 , this embodiment provides its technical solution as an emotional state evaluation method based on cross-time and space multi-modal information fusion, including the following steps:

[0056] 1. Use a two-channel convolutional neural network to extract the expression and gesture features in video information and perform primary fusion;

[0057] 2. Use a Gaussian mixture model and a deep neural network to perform emotional classification on the speech signal in audio information, and weight the emotional classification result through a selective attention weighting mechanism to obtain the audio emotional level;

[0058] 3. Input the audio emotional level result as a bias term into the fully connected layer to achieve secondary fusion of multi-modal information in cross-time and space;

[0059] 4. Use the PAD (Pleasure-Arousal-Dominance) model to describe emotions in multiple dimensions and generate emotional description curves in three dimensions;

[0060] 5. Use the Dynamic Time Warping (DTW) method to process the emotional description curves in each dimension, generate emotional change trend curves, and realize the evaluation of the emotional state of the key monitoring object.

[0061] Specifically, in step 1, referring to Figure 2 , use a two-channel convolutional neural network to extract the expression and pose features in the video information and perform primary fusion, including the following steps:

[0062] 1) Use a 3×3 convolutional kernel, a 1×1 convolutional kernel, and max pooling to extract the expression features in the video information;

[0063] 2) Use a 3×3 convolutional kernel, a 1×1 convolutional kernel, and max pooling to extract the pose features in the video information;

[0064] 3) Concatenate the expression and pose features, and use the weighted average method to perform primary fusion on the extracted expression and pose features.

[0065] Specifically, in step 2, referring to Figure 3 , use the Gaussian mixture model and the deep neural network to perform emotional classification on the speech signal in the audio information, and weight the emotional classification results through a selective attention weighting mechanism to obtain the audio emotional level, including the following steps:

[0066] 1) Use the Gaussian mixture model to model the multi-resolution spectral features of the enhanced speech signal;

[0067] 2) Based on the prosodic features of the enhanced speech signal, classify the speech signal from negative emotions to positive emotions into M levels;

[0068] 3) Calculate the syllable state transition probability matrix for the segmented audio segments, and use it as the input of the bidirectional long short-term memory network to complete the sensitive word detection;

[0069] 4) When a sensitive word appears, set different weights according to the different sensitive words and positively weight the speech signals at M levels to obtain the audio emotional level.

[0070] Specifically, in step 3, referring to Figure 4 , use the audio emotional level result as the bias term to input into the fully connected layer to realize the secondary fusion of multi-modal information in different time and space.

[0071] Specifically, in step 4, referring to Figure 5 , use the PAD (Pleasure-Arousal-Dominance) model to describe emotions in multiple dimensions and generate emotional description curves in three dimensions, including the following steps:

[0072] 1) Combine with the change of time, taking time t as the horizontal axis, generate the emotional description curve of the pleasure dimension, and its formula is f P =(P1, P2, P N );

[0073] 2) Combine with the change of time, taking time t as the horizontal axis, generate the emotional description curve of the arousal dimension, and its formula is f A =(A1, A2,, A N );

[0074] 3) Combine with the change of time, taking time t as the horizontal axis, generate the emotional description curve of the dominance dimension, and its formula is f D =(D1, D2,, D N ).

[0075] Specifically, in step 5, refer to Figure 5 , use the dynamic time warping method (DTW) to process the emotional description curves of each dimension, generate the emotional change trend curve, and realize the evaluation of the emotional state of the key monitoring object, including the following steps:

[0076] 1) For the pleasure dimension sequence f P =(P1, P2, P N ), the arousal dimension sequence f A =(A1, A2,, A N ), construct an N×N distance matrix Each element in the matrix represents the Euclidean distance between P i and A j ;

[0077] 2) Construct the cumulative distance matrix where each element represents the minimum cumulative distance from the starting point (1, 1) to the point (i, j). The calculation formula of the cumulative distance matrix is

[0078] 3) Starting from the lower right corner of the cumulative distance matrix , backtrack to the upper left corner according to the principle of selecting the minimum adjacent point of the current point, and find the shortest warping path between f P and f A , denoted as f S =(S1, S2, S N );

[0079] 4) For the shortest warping path f S =(S1, S2, S N ), the dominance dimension sequence f D =(D1, D2,, DN ) Construct an N×N distance matrix Each element in the matrix represents the Euclidean distance between S i and D j ;

[0080] 5) Construct a cumulative distance matrix where each element represents the minimum cumulative distance from the starting point (1,1) to the point (i,j). The calculation formula for the cumulative distance matrix is

[0081] 6) Starting from the lower right corner of the cumulative distance matrix and following the principle of selecting the minimum adjacent point of the current point, backtrack to the upper left corner to find the shortest curved path between f S and f D , denoted as f B =(B1,B2,B N );

[0082] 7) Finally, output an emotional change trend curve with time as the horizontal axis and the shortest curved path as the vertical axis.

[0083] The experiment uses the commonly used dataset CMU-MOSEI and divides it into a training set, a validation set, and a test set according to a ratio of 6:2:2. To verify the superiority of this embodiment, three types of methods, namely existing single-modal emotion recognition, dual-modal emotion recognition, and multi-modal emotion recognition, are selected as comparison methods. This embodiment is evaluated using evaluation metrics such as accuracy, F1 score, AUC (Area Under Curve), mean squared error MSE, and DTW similarity. As Figure 6 shown, the prediction results of this embodiment are better than those of single-modal, dual-modal, and existing multi-modal methods.

[0084] Embodiment 2

[0085] As Figure 7 shown, to achieve the above object, this embodiment discloses an emotional state evaluation system for cross-time and space multi-modal information fusion, including:

[0086] 1. A video information input module 11, which is used to extract expression and posture features. Through a dual-channel convolutional neural network, the expression and posture features in the video information are respectively extracted and subjected to primary fusion;

[0087] 2. The audio information input module 12 is used for emotional classification and weighted processing of voice signals. It uses a Gaussian mixture model to perform multi-resolution spectral feature modeling on voice signals, uses a deep neural network to perform emotional classification on voice signals, and performs weighted processing on the emotional recognition results through a selective attention weighting mechanism;

[0088] 3. The multi-modal information fusion module 13 is used for the fusion of multi-modal information. It inputs the audio emotional information as a bias term into the fully connected layer to achieve the secondary fusion of multi-modal information in different time and space;

[0089] 4. The emotional state evaluation module 14 is used for generating emotional description curves and emotional change trend curves. It uses the PAD model to perform multi-dimensional description of emotions, generates emotional description curves in three dimensions of pleasure, arousal, and dominance, and uses the dynamic time warping method (DTW) to process the emotional description curves in each dimension to generate emotional change trend curves.

[0090] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for evaluating emotional states based on cross - space - time multimodal information fusion, characterized in that, It includes the following steps: S1. Use a dual-channel convolutional neural network to extract the expression and pose features in the video information and perform primary fusion; S2. Use a Gaussian mixture model and a deep neural network to perform emotion classification on the speech signal in the audio information, and weight the emotion classification results through a selective attention weighting mechanism to obtain the audio emotion level; S3. Input the audio emotion level result as a bias term into the fully connected layer to achieve secondary fusion of multi-modal information in different time and space; S4. Adopt the PAD model to describe the emotion in multiple dimensions and generate emotion description curves in three dimensions; S5. Use the dynamic time warping method DTW to process each dimension of the emotion description curve to generate an emotion change trend curve and realize the evaluation of the emotion state of the key monitoring object.

2. The emotional state evaluation method based on cross-time and space multimodal information fusion according to claim 1, wherein The step S1 includes the following steps: S21. Use a 3×3 convolutional kernel, a 1×1 convolutional kernel, and max pooling to extract the expression features in the video information; S22. Use a 3×3 convolutional kernel, a 1×1 convolutional kernel, and max pooling to extract the pose features in the video information; S23. Use the weighted average method to perform primary fusion on the extracted expression and pose features.

3. The emotional state evaluation method based on cross-time and space multi-modal information fusion according to claim 1, wherein, The step S2 includes the following steps: S31. Use a Gaussian mixture model to model the multi-resolution spectral features of the enhanced speech signal; S32. Based on the prosodic features of the enhanced speech signal, classify the speech signal from negative emotion to positive emotion into M levels; S33. Obtain the syllable state transition probability matrix for the segmented audio segments as the input of the bidirectional long short-term memory network to complete the sensitive word detection; S34. When a sensitive word appears, set different weights according to the different sensitive words and positively weight the speech signals at M levels to obtain the audio emotion level.

4. The emotional state assessment method based on cross-time and space multimodal information fusion according to claim 1, characterized in that The step S4 includes the following steps: S41. Combine with the time change, take time t as the horizontal axis, and generate an emotional description curve of the pleasure dimension, whose formula is f P =(P1, P2, … P N ), where P N represents the positive and negative emotional experiences, ranging from unpleasant to pleasant, and the quantitative value range is [0, 1]; S42. Combine with the time change, use time t as the horizontal axis to generate an emotional description curve of the arousal dimension, and its formula is f A =(A1, A2, …, A N ), where A N represents the psychological activation level, ranging from calm to excited, and the quantitative value range is [0, 1]; S43. Combine with the time change, take time t as the horizontal axis, and generate an emotional description curve of the dominant dimension, whose formula is f D =(D1, D2, …, D N ), D N represents the sense of control over the situation, ranging from compliance to dominance, and the quantitative value range is [0, 1].

5. The emotional state evaluation method based on cross-time and space multi-modal information fusion according to claim 1, wherein The step S5 includes the following steps: S51. For the pleasure dimension sequence f P =(P1, P2, … P N ), and the arousal dimension sequence f A =(A1, A2, …, A N ), construct an N×N distance matrix Each element in the matrix represents the Euclidean distance between P i and A j ; S52. Construct a cumulative distance matrix where each element represents the minimum cumulative distance from the starting point (1, 1) to the point (i, j). The calculation formula for the cumulative distance matrix is S53. Starting from the lower right corner of the cumulative distance matrix , backtrack to the upper left corner according to the principle of selecting the minimum adjacent point of the current point to find the shortest curved path between f P and f A , denoted as f S =(S1, S2, … S N ); S54. For the shortest bending path f S =(S1, S2, … S N ), the dominant dimension sequence f D =(D1, D2, …, D N ), construct an N×N distance matrix Each element in the matrix represents the Euclidean distance between S i and D j ; S55. Construct the cumulative distance matrix where each element represents the minimum cumulative distance from the starting point (1, 1) to the point (i, j), and the calculation formula for the cumulative distance matrix is S56. Starting from the lower right corner of the cumulative distance matrix , backtrack to the upper left corner according to the principle of selecting the minimum adjacent point of the current point to find the shortest curved path between f S and f D , denoted as f B = (B1, B2, … B N ); S57. Finally, output an emotion change trend curve with time as the horizontal axis and the shortest bending path as the vertical axis.

6. An emotional state assessment system based on cross-time and space multi-modal information fusion, characterized in that, It includes: S61. A video information input module for extracting expression and pose emotion features; S62. An audio information input module for emotion classification and weighting processing of the speech signal; S63. A multi-modal information fusion module for fusing multi-modal information; S64. An emotion state evaluation module for generating emotion description curves and emotion change trend curves.

7. The system according to claim 6, wherein The video information input module uses a dual-channel convolutional neural network for feature extraction and fusion.

8. The system according to claim 6, wherein The audio information input module uses a Gaussian mixture model and a deep neural network for emotion classification.

9. The system according to claim 6, wherein The multi-modal information fusion module inputs the audio emotion information as a bias term into the fully connected layer.

10. The system according to claim 6, characterized in that, The emotion state evaluation module uses the PAD model and the dynamic time warping method DTW for emotion state evaluation.

Citation Information

Cited By

  • Case analysis and prediction method for secondary ear symptoms caused by maxillofacial joint disorder

    CN122024987A

  • A method, system, and medium for identifying psychological states based on the PAD emotion model.

    CN122556990A