Multi-modal emotion recognition method and system based on deep learning
By analyzing facial expression richness, audio complexity and text emotional consistency, multimodal emotion recognition features are weighted, and combined with deep learning models, the problem of modal bias in multimodal emotion recognition is solved, and the accuracy of emotion recognition is improved.
Patent Information
- Application Number
- CN202510539763.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-18
AI Technical Summary
In existing multimodal emotion recognition technology, simple feature splicing leads to modal bias, affecting the accuracy of emotion recognition.
By analyzing facial expression richness, audio complexity and text emotional consistency, the video, audio and text features are weighted, and trained in combination with deep learning models to obtain emotional recognition results.
It improves the accuracy of emotion recognition, overcomes the performance degradation of fixed weight models under extreme conditions, and achieves more accurate emotion recognition.
Smart Images

Figure CN120337039A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of pattern recognition, and specifically to a multimodal emotion recognition method and system based on deep learning. Background Art
[0002] Multimodal emotion recognition refers to the identification and understanding of human emotional states by combining information from different modalities, such as speech, facial expressions, text, etc. This technology combines multiple perceptual channels to more accurately and comprehensively evaluate an individual's emotional expression.
[0003] Deep learning is a branch of machine learning that mainly relies on neural network models. In emotion recognition tasks, deep learning extracts features from data automatically, avoiding the manual feature design in traditional methods.
[0004] During the training process of deep learning models for multimodal emotion recognition, usually only the features extracted from multimodal data are simply concatenated and trained and classified through a classification network. However, the amount of information in data of different modalities is different, and their contributions to the final emotion recognition are also different. Simple feature concatenation is likely to form modality bias, resulting in a low accuracy of the final emotion recognition. Summary of the Invention
[0005] To solve the above technical problems, the purpose of this application is to provide a multimodal emotion recognition method and system based on deep learning. The specific technical solutions adopted are as follows: In a first aspect, an embodiment of this application provides a multimodal emotion recognition method based on deep learning. The method includes the following steps: Obtain multimodal data of each emotional scenario with known emotional labels, where each modality includes text, video, and audio; extract text features, video features, and audio features under each emotional scenario; Analyze the position differences of facial feature points in adjacent frame video images under each emotional scenario to determine the facial expression richness of each emotional scenario; Extract the audio envelope under each emotional scenario, analyze the change rate of audio data and the fluctuation of amplitude in the audio data through the audio envelope to determine the audio complexity of each emotional scenario; Identify various emotional words in the text data under each emotional scenario, and determine the text emotion consistency of each emotional scenario based on the statistical quantity of various emotional words in the text data under each emotional scenario; Based on the facial expression richness, the audio complexity, and the text emotion consistency, respectively weight the video features, audio features, and text features under each emotional scenario, and combine with a deep learning model for training to obtain an emotion recognition result.
[0006] In one embodiment, the text feature is the word vector of all text data in each emotional scenario.
[0007] In one embodiment, the video feature includes: Perform key point detection on the faces in the video data for each emotional scenario, and form the video feature by combining the key point coordinates in all frames of video in each emotional scenario.
[0008] In one embodiment, the audio feature includes: For each emotional scenario, calculate the zero-crossing rate of all frames of audio data, obtain the prediction coefficients of all frames of audio data using linear predictive coding, and average the zero-crossing rate and the prediction coefficients respectively in the time dimension to form the audio feature of each emotional scenario.
[0009] In one embodiment, the determination of the facial expression richness includes: For each emotional scenario, calculate the metric distance of the same feature points in adjacent frame video images, and use the mean value of the metric distances of all feature points in adjacent frame video images as the position deviation of the adjacent frame video images. The facial expression richness is the normalized result of the dispersion degree of the position deviations of all adjacent frame video images in each emotional scenario.
[0010] In one embodiment, the determination of the audio complexity includes: Calculate the mean value of the ratio of the difference between all adjacent peak points in the audio envelope to the corresponding time interval in each emotional scenario, denoted as the first mean value, calculate the range of the energy intensities corresponding to all peak points in the audio envelope in each emotional scenario, and the audio complexity is the normalized result of the sum value of the first mean value and the range.
[0011] In one embodiment, the determination of the text emotion consistency includes: Obtain the maximum value and the range difference of the number of all types of emotion words in the text data for each emotional scenario, calculate the ratio of the maximum value to the number of all emotion words in the text data for each emotional scenario, denoted as the first ratio, and the text emotion consistency is the normalized result of the sum value of the range difference and the first ratio.
[0012] In one embodiment, the weighted processing of the video feature, audio feature, and text feature for each emotional scenario respectively includes: For each emotion label, calculate the mean value of the facial expression richness, the mean value of the audio complexity, and the mean value of the text emotion consistency in all emotional scenarios respectively, denoted as the second mean value, the third mean value, and the fourth mean value; Calculate the difference between the facial expression richness in each emotion scenario and the second mean value, denoted as the first difference. Calculate the difference between the audio complexity in each emotion scenario and the third mean value, denoted as the second difference. Calculate the difference between the text emotion consistency in each emotion scenario and the fourth mean value, denoted as the third difference; Use the first difference, the second difference, and the third difference to weight the video features, audio features, and text features in each emotion scenario respectively.
[0013] In one embodiment, the weight of the video features in each emotion scenario is the normalized value of the reciprocal of the first difference, the weight of the audio features in each emotion scenario is the normalized value of the reciprocal of the second difference, and the weight of the text features in each emotion scenario is the normalized value of the reciprocal of the third difference; Use the weighted text features, video features, and audio features of all emotion scenarios of all emotion labels as the training data of the deep learning model, and perform emotion recognition using the trained deep learning model.
[0014] In a second aspect, an embodiment of the present application further provides a multi-modal emotion recognition system based on deep learning, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0015] The present application has at least the following beneficial effects: This application obtains the multi-modal data of each emotional scenario with known emotional tags, where each modality includes text, video, and audio; extracts the text features, video features, and audio features under each emotional scenario; achieves comprehensive coverage of emotional information, and the synergistic effect of the three overcomes the defect of one-sidedness of single-modal information; analyzes the position differences of facial feature points in adjacent frame video images under each emotional scenario to determine the facial expression richness of each emotional scenario; the determination of facial expression richness improves the sensitivity of dynamic expression capture, can accurately identify micro-expression and macro-expression changes, and quantifying expression richness helps to distinguish emotional intensity; extracts the audio envelope under each emotional scenario, analyzes the change rate of audio data and the amplitude fluctuation in the audio data through the audio envelope to determine the audio complexity of each emotional scenario; the audio complexity reflects the fluctuation of audio data and represents the emotional intensity contained in the audio data, improving the capture accuracy of emotional transient features; identifies various emotional words in the text data under each emotional scenario, and determines the text emotional consistency of each emotional scenario based on the statistical quantity of various emotional words in the text data under each emotional scenario; the text emotional consistency optimizes the robustness of the determination of text emotional tendency and avoids misclassification caused by text ambiguity. Based on the facial expression richness, the audio complexity, and the text emotional consistency, the video features, audio features, and text features under each emotional scenario are weighted respectively, and combined with a deep learning model for training to obtain the emotional recognition result. Through the scene-based adjustment of the multi-modal contribution degree, this dynamic balance mechanism overcomes the performance degradation of the fixed-weight model under extreme conditions and improves the accuracy of the final emotional recognition. Description of the Drawings
[0016] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a flowchart of the steps of a multi-modal emotional recognition method based on deep learning provided by an embodiment of the present application; Figure 2 It is a flowchart for obtaining the emotional recognition result. Detailed Embodiments
[0018] To further elaborate on the technical means and effects adopted by this application to achieve the intended invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of the multi-modal emotion recognition method and system based on deep learning proposed according to this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0020] The following specifically describes the specific solutions of the multi-modal emotion recognition method and system based on deep learning provided by this application in conjunction with the accompanying drawings.
[0021] Please refer to Figure 1 , which shows the flowchart of the steps of the multi-modal emotion recognition method based on deep learning provided by one embodiment of this application. The method includes the following steps: S1, obtain the multi-modal data of each emotion scene with known emotion labels, where each modality includes text, video, and audio; extract the text features, video features, and audio features under each emotion scene.
[0022] In this embodiment, the CMU-MOSEI public dataset is selected as the training data for multi-modal emotion recognition. The CMU-MOSEI public dataset contains data of three modalities, specifically text, video, and audio. The data of the three modalities under the same emotion scene correspond to the same emotion label. Among them, the emotion labels include six categories: happy, sad, angry, fearful, surprised, and disgusted. An emotion label is divided into different emotion scenes. For example, for the emotion label "happy", emotion scene 1 is "I solved a difficult problem that has puzzled me for a long time today. So happy!", and emotion scene 2 is "I got into college. So excited!". Both of these emotion scenes express the emotion of happiness. Therefore, the emotion labels of these two emotion scenes are both "happy".
[0023] It should be noted that the implementer can choose other existing multi-modal public datasets by himself / herself, and this embodiment does not limit this.
[0024] Furthermore, for the multi-modal data of each emotion scene of each emotion label, that is, text, video, and audio, extract the corresponding text features, video features, and audio features. The text features, video features, and audio features are all features crucial for the final emotion recognition. The specific extraction process in this embodiment is as follows: For text features, all text data under each emotional scenario are segmented using the BERT-WWM model to obtain the word vectors corresponding to all the segmented words under each emotional scenario. The word vectors of all the segmented words under each emotional scenario are arranged in chronological order to form the text features under each emotional scenario. Among them, the BERT-WWM model is a well-known existing technology, and implementers can choose other existing feasible methods to obtain word vectors by themselves. This embodiment does not make any restrictions here.
[0025] For video features, for each frame of video image under each emotional scenario, face detection is performed using the multi-task convolutional neural network MTCNN model, and at the same time, the OpenFace tool is used to obtain the key points of facial features, specifically including 68 key points of the face. The coordinate points of the key points of each video image are extracted frame by frame to construct the feature sequence of each video frame image. The feature sequences corresponding to all video frames under each emotional scenario are spliced in the image order to obtain the image features. Among them, the 68 key points of the face are well-known existing knowledge and will not be elaborated in detail here.
[0026] In another embodiment, the facial action units in each frame of video image under each emotional scenario can be extracted to form the image features under each emotional scenario.
[0027] In other embodiments, the fixation points in each frame of video image under each emotional scenario can be extracted to form the image features under each emotional scenario.
[0028] Both the multi-task convolutional neural network MTCNN model and the OpenFace tool are well-known existing technologies, and the specific processes will not be elaborated.
[0029] For audio features, when extracting audio features, first, the FFmpeg tool is used to separate the audio content in each frame of video. Then, the Librosa library is used to extract the zero-crossing rate of each frame of audio data at a sliding interval of 512, and the linear predictive coding is used to obtain the prediction coefficients of each frame of audio data. The zero-crossing rate and the prediction coefficients of all frame audio data under each emotional scenario are averaged respectively in the time dimension to form the audio features of each emotional scenario.
[0030] In another embodiment, the Mel cepstral coefficients and constant Q transform of each frame of audio data can also be extracted using the Librosa library at a sliding interval of 512, and they are also averaged in the time dimension to form the audio features of each emotional scenario.
[0031] Both the FFmpeg tool and the Librosa library are well-known existing technologies, and the specific processes will not be elaborated.
[0032] S2. Analyze the position differences of the facial feature points in adjacent frame video images under each emotional scenario to determine the facial expression richness of each emotional scenario.
[0033] In the traditional multi-modal feature fusion method, the multi-modal features obtained by feature extraction, namely text features, image features, and audio features, are simply concatenated through a concatenation function, and the fused features are used to classify the emotions expressed in each emotional scenario through a deep learning model. However, in the actual process, the amount of data information contained in different modal features is different, and even the emotions expressed by image, audio, and text data in the same emotional scenario are opposite. Therefore, if all modal data are treated equally, it is easy to form modal bias, resulting in a deviation between the final emotion recognition result and the actual situation. Therefore, it is necessary to further analyze according to the distribution of multi-modal data to guide multi-modal feature fusion.
[0034] For a single emotional scenario, the text data, audio data, and image data are all extracted from a video clip at the same time. Therefore, it is possible to achieve temporal consistency of video frames, audio, and text data through temporal alignment. That is, the video image, audio, and text are corresponding at a single moment.
[0035] In multi-modal data, video frames and audio features are often more capable of reflecting the personal emotional information of the speaker. Therefore, in this embodiment, the modal bias in each emotional scenario is first analyzed from the image data and audio data.
[0036] When a human's emotion changes, it is mainly reflected in the change of facial micro-expressions. Therefore, when a person's mood is unstable, the facial expression is relatively complex and the overall change is large. And the manifestation in video data is that there will be a large position deviation of facial key points in a certain emotional scenario. Therefore, the position difference of facial key points in two consecutive video frame images is analyzed.
[0037] In this embodiment, considering that the key points of the mouth between different video frames may be affected by speaking, resulting in a large position deviation of the key points of the mouth. Therefore, in order to avoid the influence of speaking on the position deviation of facial key points, in each emotional scenario, the key points of the mouth in each frame of video image are removed, and the remaining key points after removing the key points of the mouth are denoted as each feature point. The metric distance between the position coordinates of the same feature point in each frame of video image and its previous frame of video image is calculated, and the mean value of the metric distances of all feature points in each frame of video image and its previous frame of video image is used as the position deviation between each frame of video image and its previous frame of video image. The normalized result of the dispersion degree of the position deviation of all frames of video images in each emotional scenario is used as the facial expression richness of each emotional scenario.
[0038] It should be noted that for the first-frame video images in each emotional scenario, the calculation of the position deviation is not performed. In this embodiment, the Euclidean distance is used as the calculation method for the metric distance. Implementers can choose other existing feasible calculation methods for the metric distance, such as the DTW distance, Manhattan distance, etc. In this embodiment, the standard deviation is used as the calculation method for the degree of dispersion. Implementers can choose other existing feasible calculation methods for the degree of dispersion, such as variance, coefficient of variation, etc. In this embodiment, the Sigmoid function is used as the method for obtaining the normalized result. Implementers can choose other existing feasible normalization methods. This embodiment does not make any restrictions here.
[0039] S3. Extract the audio envelopes in each emotional scenario, analyze the change rate of the audio data and the fluctuation of the amplitude in the audio data through the audio envelopes, and determine the audio complexity of each emotional scenario.
[0040] Since the human voice is realized through the vibration of the vocal cords, which is manifested as a fluctuating electrical signal in the audio data, and there is a close relationship between the fluctuation of the audio signal and personal emotional changes. During the emotional expression process, the waveform and envelope of the audio signal often change significantly. For example, when angry, the waveform of the audio may be more intense and the change of the envelope is larger; while when calm or sad, the audio waveform may be smoother and the envelope change is smaller.
[0041] Based on the above analysis, in this embodiment, the envelope extraction algorithm is first used to obtain the envelopes of the audio signals in each emotional scenario, denoted as audio envelopes. Calculate the difference between each peak point in the audio envelope and its previous peak point, that is, calculate the difference in the energy intensity corresponding to each peak point in the audio envelope and its previous peak point. In addition, calculate the time interval between each peak point and its previous peak point, and calculate the ratio of the difference between each peak point and its previous peak point to the time interval. Denote the mean value of the ratios of all peak points in the audio envelopes in each emotional scenario as the first mean value. The first mean value reflects the change rate of the audio data. The larger the first mean value, the greater the change of the audio envelope, the more intense the fluctuation of the audio data, indicating that the personal emotion may be more excited and more likely to be angry or excited.
[0042] It should be noted that during the calculation of the difference, the first peak of the audio envelope is not calculated.
[0043] In addition, calculate the range of the energy intensity corresponding to all peak points in the audio envelope. The range reflects the fluctuation range of the amplitude in the audio data. Combine the first mean value to determine the audio complexity of each emotional scenario. Specifically: use the normalized result of the sum of the first mean value and the range as the audio complexity of each emotional scenario. The greater the audio complexity, the more personal emotional expression information is contained in the audio data, which is more important for subsequent emotion recognition. Among them, the normalized result is obtained by using the Sigmoid function.
[0044] S4. Identify various emotional words in the text data under each emotional scenario. Based on the statistical quantity of various emotional words in the text data under each emotional scenario, determine the text emotion consistency of each emotional scenario.
[0045] Furthermore, the expression of personal emotions by text data will be more direct. When personal emotions are clearly expressed, it often involves words that have a strong correlation with such emotions. And the more concentrated the emotional words are, the more single and direct the emotion is. Therefore, based on the emotion dictionary, through emotion word matching, count the quantity of various emotional words in all text data under each emotional scenario. Among them, in this embodiment, the emotion dictionary selects the publicly available HowNet emotion dictionary, and the implementer can choose other existing feasible emotion dictionaries by himself, such as the Dalian University of Technology Emotion Lexical Ontology, NTUSD Emotion Dictionary, AFINN Emotion Dictionary, etc.
[0046] Therefore, based on the above analysis, this embodiment calculates the text emotion consistency C of each emotional scenario. The specific expression is: , where represents the ratio of the maximum value of the quantity of all types of emotional words in the text data under each emotional scenario to the quantity of all emotional words in the text data, denoted as the first ratio. is the range value of the quantity of all types of emotional words in the text data under each emotional scenario, and norm() is the normalization function.
[0047] Through the text emotion consistency, the distribution of emotional words under each emotional scenario can be measured. If the distribution is more concentrated, it indicates that the consistency of the text data is higher and the importance of the emotional information contained is higher.
[0048] S5. Based on the facial expression richness, the audio complexity, and the text emotion consistency, weight the video features, audio features, and text features under each emotional scenario respectively, and combine with a deep learning model for training to obtain the emotion recognition result.
[0049] Finally, by comprehensively considering the facial expression richness, the audio complexity, and the text sentiment consistency in each emotion scenario, the video features, audio features, and text features in each emotion scenario are weighted respectively. The aim is to assign different contribution weights to different modality features, amplify the contribution weights of the modality features with more sufficient emotional information expression, and reduce the contribution weights of the modality features with insufficient emotional information expression or ambiguous emotional information expression, which helps to improve the accuracy of the final emotion recognition.
[0050] In this embodiment, for each emotion label in the dataset, the mean value of the facial expression richness in all emotion scenarios is calculated, denoted as the second mean value; the mean value of the audio complexity in all emotion scenarios is calculated, denoted as the third mean value; and the mean value of the text sentiment consistency in all emotion scenarios is calculated, denoted as the fourth mean value. It should be understood that the second mean value, the third mean value, and the fourth mean value represent the average levels of the facial expression richness, the audio complexity, and the text sentiment consistency under each emotion label, which helps to analyze the deviation degrees of the facial expression richness, the audio complexity, and the text sentiment consistency in each emotion scenario from the average level.
[0051] The difference between the facial expression richness in each emotion scenario and the second mean value is calculated, denoted as the first difference; the difference between the audio complexity in each emotion scenario and the third mean value is calculated, denoted as the second difference; and the difference between the text sentiment consistency in each emotion scenario and the fourth mean value is calculated, denoted as the third difference. The first difference, the second difference, and the third difference reflect the deviation degrees of the facial expression richness, the audio complexity, and the text sentiment consistency in each emotion scenario from the average level. The greater the deviation degree, the less emotional information the corresponding modality feature contains, or the lower the accuracy of the emotional expression. When performing the final emotion recognition, a smaller weight should be assigned to this modality feature.
[0052] It should be noted that in this embodiment, the first difference, the second difference, and the third difference are all calculated in the way of the absolute value of the difference. In other embodiments, the first difference, the second difference, and the third difference are all calculated in the way of the square of the difference.
[0053] Using the first difference, the second difference, and the third difference, the video features, audio features, and text features in each emotion scenario are weighted respectively. Specifically, the weight of the video features in each emotion scenario is the normalized value of the reciprocal of the first difference; the weight of the audio features in each emotion scenario is the normalized value of the reciprocal of the second difference; and the weight of the text features in each emotion scenario is the normalized value of the reciprocal of the third difference. The normalized values are all obtained by using the Sigmoid function.
[0054] Thus, in the multi-modal feature fusion stage of multi-modal sentiment recognition using deep learning, that is, when fusing image features, text features, and audio features, the weights of each modal feature are added. Specifically, the weighted text features, video features, and audio features of all sentiment scenarios of all sentiment labels in the CMU-MOSEI public dataset are used as the training data for the deep learning model. Among them, the ratio of the training set, test set, and validation set is 7:2:1. In this embodiment, the deep learning model is an LSTM neural network. Implementers can choose other existing feasible neural networks by themselves, and this embodiment does not limit this. During the training process of the LSTM neural network, the L1 loss function is used for training, the batch size of training is 64, the optimizer is Adam, the learning rate is set to 0.002, and the decay coefficient is set to 0.0001.
[0055] Finally, the trained LSTM neural network is used for multi-modal sentiment recognition to output sentiment labels in each sentiment scenario. The LSTM neural network is a publicly known technology in the art, and the specific process will not be elaborated here. The flowchart for obtaining the sentiment recognition result is as Figure 2 shown.
[0056] Based on the same inventive concept as the above method, the embodiment of the present application also provides a multi-modal sentiment recognition system based on deep learning, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above multi-modal sentiment recognition methods based on deep learning.
[0057] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. In addition, the above specific embodiments of this specification have been described. Moreover, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0058] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
[0059] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A multi-modal sentiment recognition method based on deep learning, characterized in that, The method includes the following steps: Obtain the multi-modal data of each emotional scenario with known emotional labels, where each modality includes text, video, and audio; extract the text features, video features, and audio features under each emotional scenario; Analyze the position differences of facial feature points in adjacent frame video images under each emotional scenario to determine the facial expression richness of each emotional scenario; Extract the audio envelope under each emotional scenario, analyze the change rate of audio data and the fluctuation of amplitude in the audio data through the audio envelope to determine the audio complexity of each emotional scenario; Identify various emotional words in the text data under each emotional scenario, and determine the text emotional consistency of each emotional scenario based on the statistical quantity of various emotional words in the text data under each emotional scenario; Based on the facial expression richness, the audio complexity, and the text emotional consistency, weight the video features, audio features, and text features under each emotional scenario respectively, and combine with a deep learning model for training to obtain the emotion recognition result.
2. The multimodal emotion recognition method based on deep learning according to claim 1, characterized in that, The text features are the word vectors of all text data under each emotional scenario.
3. The multimodal emotion recognition method based on deep learning according to claim 1, characterized in that, The video features include: Perform key point detection on the faces in the video data under each emotional scenario, and form video features by combining the key point coordinates in all frames of video under each emotional scenario.
4. The multimodal sentiment recognition method based on deep learning according to claim 1, wherein The audio features include: For each emotional scenario, calculate the zero-crossing rate of all frame audio data, obtain the prediction coefficients of all frame audio data using linear predictive coding, and average the zero-crossing rate and the prediction coefficients respectively in the time dimension to form the audio features of each emotional scenario.
5. The multimodal sentiment recognition method based on deep learning according to claim 1, characterized in that, The determination of the facial expression richness includes: For each emotional scenario, calculate the metric distance of the same feature points in adjacent frame video images, and take the mean of the metric distances of all feature points in adjacent frame video images as the position deviation of the adjacent frame video images. The facial expression richness is the normalized result of the dispersion degree of the position deviations of all adjacent frame video images under each emotional scenario.
6. The multi-modal sentiment recognition method based on deep learning according to claim 1, characterized in that The determination of the audio complexity includes: Calculate the mean of the ratio of the difference between all adjacent peak points in the audio envelope under each emotional scenario to the corresponding time interval, denoted as the first mean, and calculate the range of the energy intensity corresponding to all peak points in the audio envelope under each emotional scenario. The audio complexity is the normalized result of the sum value of the first mean and the range.
7. The multimodal sentiment recognition method based on deep learning according to claim 1, wherein The determination of the text emotional consistency includes: Obtain the maximum value and the extreme difference of the number of all types of emotional words in the text data under each emotional scenario, calculate the ratio of the maximum value to the number of all emotional words in the text data under each emotional scenario, denoted as the first ratio. The text emotional consistency is the normalized result of the sum value of the extreme difference and the first ratio.
8. The multimodal emotion recognition method based on deep learning according to claim 1, characterized in that, The weighting of the video features, audio features, and text features under each emotional scenario respectively includes: For each emotional label, calculate the mean of the facial expression richness, the mean of the audio complexity, and the mean of the text emotional consistency under all emotional scenarios respectively, denoted as the second mean, the third mean, and the fourth mean; Calculate the difference between the facial expression richness in each emotional scenario and the second mean value, denoted as the first difference. Calculate the difference between the audio complexity in each emotional scenario and the third mean value, denoted as the second difference. Calculate the difference between the text emotional consistency in each emotional scenario and the fourth mean value, denoted as the third difference. Use the first difference, the second difference, and the third difference to weight the video features, audio features, and text features in each emotional scenario respectively.
9. The multimodal sentiment recognition method based on deep learning according to claim 8, wherein The weight of the video features in each emotional scenario is the normalized value of the reciprocal of the first difference. The weight of the audio features in each emotional scenario is the normalized value of the reciprocal of the second difference. The weight of the text features in each emotional scenario is the normalized value of the reciprocal of the third difference. Use the weighted text features, video features, and audio features of all emotional scenarios of all emotional labels as the training data of the deep learning model, and perform emotion recognition using the trained deep learning model.
10. A multi-modal emotion recognition system based on deep learning, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.