An Emotion Recognition Method, Device, Storage Medium and Equipment
By extracting and decoding non-target modal signals in emotional recognition to generate target modal signals, and combining multiple features for emotional recognition, the problem of low accuracy caused by the loss of modal signals is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202310651876.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-06-01
AI Technical Summary
The existing emotion recognition methods cannot guarantee the accuracy of the generated features when partial modal signals are missing, resulting in a low accuracy of the emotion recognition results.
By extracting the signal characteristics of the non-target modal signal in the target video for decoding, the missing target modal signal is generated, and emotional recognition is performed by combining the signal characteristics of the target modal signal and the non-target modal signal and text characteristics for emotional recognition, the facial action unit recognition model and decoder are used for feature generation and constraints.
The accuracy of emotional recognition results is improved, and a more accurate identification basis is provided through the fusion of multiple signal characteristics.
Smart Images

Figure CN116612543B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an emotion recognition method, device, storage medium, and equipment. Background Art
[0002] With the rapid development of artificial intelligence technology, human-computer interaction appears more and more frequently in people's daily work and life, which can bring great convenience to people. As an important part of human-computer interaction, emotion recognition has been widely applied to scenarios such as medical treatment, vehicle-mounted, and customer service.
[0003] Existing emotion recognition methods usually adopt multi-modal emotion recognition methods. However, since multi-modal emotion recognition needs to receive information of multiple modalities simultaneously, in actual use, the condition of receiving information of multiple modalities simultaneously is not always available. In this regard, the currently commonly used emotion recognition method that supports missing modalities is to directly generate the features of the missing modality through the features of the remaining modalities. On the one hand, this end-to-end feature generation method lacks the generation of the corresponding intermediate modality signals, and it is impossible to determine whether the generated features can accurately represent the corresponding modality signals; on the other hand, since the generated features are not directly constrained, the accuracy of the generated features cannot be guaranteed, which may lead to a low accuracy rate of the emotion recognition result obtained based on the generated features subsequently. Therefore, how to improve the accuracy rate of the emotion recognition result in the case of partial modality signal loss is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to provide an emotion recognition method, device, storage medium, and equipment, which can effectively improve the accuracy rate of the emotion recognition result in the case of partial modality signal loss.
[0005] The embodiments of this application provide an emotion recognition method, including:
[0006] Obtain a target video to be recognized, where the target video is a video lacking a target modality signal;
[0007] Extract the signal features of the non-target modality signals in the target video, and decode the signal features of the non-target modality signals to generate the target modality signal;
[0008] Extract the signal features of the target modality signal, and generate the text features corresponding to the target video by using the target modality signal or the non-target modality signal;
[0009] Perform emotion recognition on the target user in the target video according to the signal features of the target modality signal, the signal features of the non-target modality signals, and the text features, and obtain the emotion recognition result corresponding to the target user.
[0010] In a possible implementation, the target modality signal is an audio signal, and the non-target modality signal is an image signal; extracting the signal features of the non-target modality signal in the target video and decoding the signal features of the non-target modality signal to generate the target modality signal includes:
[0011] Extracting the image features of the image signal in the target video and using a first decoder to decode the image features to generate the spectrogram of the audio signal, so as to use the spectrogram of the audio signal to determine the audio signal.
[0012] In a possible implementation, using the target modality signal or the non-target modality signal to generate the text features corresponding to the target video includes:
[0013] Inputting each frame of the image signal in the target video into a pre-constructed facial action unit recognition model to predict the facial action unit corresponding to each frame of the image signal;
[0014] Determining the text information corresponding to the target video according to the facial action unit, and extracting the text features of the text information as the text features corresponding to the target video.
[0015] In a possible implementation, the facial action unit recognition model is constructed as follows:
[0016] Obtaining sample facial action unit data, where the sample facial action unit data includes five facial action units: frowning, cheek raising, mouth corner pulling, lower lip raising, and lip parting;
[0017] Using the sample facial action unit data and a target loss function to train an initial facial action unit recognition model to obtain the facial action unit recognition model.
[0018] In a possible implementation, the first decoder is constructed as follows:
[0019] Obtaining a first sample video, and separating a first sample audio signal and a first sample image signal from the first sample video;
[0020] Extracting the image features of each frame of the first sample image signal in the first sample video, and inputting the image features into an initial first decoder to generate a predicted spectrogram of the predicted audio signal of the first sample video;
[0021] Using a first loss function, perform consistency constraints on the predicted spectrogram and the spectrogram of the first sample audio signal, and using a second loss function, perform consistency constraints on the acoustic features of the predicted audio signal and the acoustic features of the first sample audio signal, and train to obtain the first decoder.
[0022] In a possible implementation, the target modal signal is an image signal, and the non-target modal signal is an audio signal; extracting the signal features of the non-target modal signal in the target video and decoding the signal features of the non-target modal signal to generate the target modal signal includes:
[0023] Extract the acoustic features of the audio signal in the target video, and use a second decoder to decode the acoustic features to generate the image signal.
[0024] In a possible implementation, generating the text features corresponding to the target video using the target modal signal or the non-target modal signal includes:
[0025] Input each frame of the image signal in the generated target video into a pre-constructed facial action unit recognition model, and predict the facial action unit corresponding to each frame of the image signal;
[0026] According to the facial action unit, determine the text information corresponding to the target video, and extract the text features of the text information as the text features corresponding to the target video.
[0027] In a possible implementation, the second decoder is constructed as follows:
[0028] Obtain a second sample video; and separate a second sample audio signal and a second sample image signal from the second sample video;
[0029] Extract the acoustic features of each frame of the second sample audio signal in the second sample video, and input the acoustic features into an initial second decoder to generate a predicted image signal of the second sample video;
[0030] Using a third loss function, perform consistency constraints on the predicted image signal and the second sample image signal, and using a fourth loss function, perform consistency constraints on the image features of the predicted image signal and the image features of the second sample image signal, and train to obtain the second decoder.
[0031] In a possible implementation, performing emotion recognition on the target user in the target video according to the signal features of the target modal signal, the signal features of the non-target modal signal, and the text features to obtain the emotion recognition result corresponding to the target user includes:
[0032] Fuse the image features of the image signal in the target video, the acoustic features of the audio signal in the target video, and the text features, and input the obtained fusion result into a fully connected layer for emotion recognition to obtain the emotion recognition result corresponding to the target user.
[0033] The embodiment of the present application also provides an emotion recognition device, including:
[0034] A first acquisition unit, configured to acquire a target video to be recognized, where the target video is a video lacking a target modality signal;
[0035] A first generation unit, configured to extract the signal features of the non-target modality signal in the target video, and decode the signal features of the non-target modality signal to generate the target modality signal;
[0036] A second generation unit, configured to extract the signal features of the target modality signal, and use the target modality signal or the non-target modality signal to generate the text features corresponding to the target video;
[0037] An identification unit, configured to perform emotion recognition on the target user in the target video according to the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features, to obtain the emotion recognition result corresponding to the target user. <l
[0038] In a possible implementation manner, the target modality signal is an audio signal, and the non-target modality signal is an image signal; specifically, the first generation unit is configured to:
[0039] Extract the image features of the image signal in the target video, and use a first decoder to decode the image features to generate a spectrogram of the audio signal, so as to determine the audio signal by using the spectrogram of the audio signal.
[0040] In a possible implementation manner, the second generation unit includes:
[0041] A first prediction sub-unit, configured to input each frame of image signal in the target video into a pre-constructed facial action unit recognition model, and predict the facial action unit corresponding to each frame of image signal;
[0042] A first extraction sub-unit, configured to determine the text information corresponding to the target video according to the facial action unit, and extract the text features of the text information as the text features corresponding to the target video.
[0043] In a possible implementation manner, the device further includes:
[0044] A second acquisition unit, configured to acquire sample facial action unit data, where the sample facial action unit data includes five facial action units: frowning, cheek raising, mouth corner pulling, lower lip raising, and lip parting.
[0045] A first training unit, configured to train an initial facial action unit recognition model by using the sample facial action unit data and a target loss function, to obtain the facial action unit recognition model.
[0046] In a possible implementation manner, the apparatus further includes:
[0047] A third acquisition unit, configured to acquire a first sample video, and separate a first sample audio signal and a first sample image signal from the first sample video;
[0048] A third generation unit, configured to extract image features of each frame of the first sample image signal in the first sample video, and input the image features into an initial first decoder to generate a predicted spectrogram of a predicted audio signal of the first sample video;
[0049] A second training unit, configured to perform consistency constraint on the predicted spectrogram and the spectrogram of the first sample audio signal by using a first loss function, and perform consistency constraint on the acoustic features of the predicted audio signal and the acoustic features of the first sample audio signal by using a second loss function, to train and obtain the first decoder.
[0050] In a possible implementation manner, the target modality signal is an image signal, and the non-target modality signal is an audio signal; the first generation unit is specifically configured to:
[0051] Extract the acoustic features of the audio signal in the target video, and decode the acoustic features by using a second decoder to generate the image signal.
[0052] In a possible implementation manner, the second generation unit includes:
[0053] A second prediction subunit, configured to input each frame of the image signal in the generated target video into a pre-constructed facial action unit recognition model, and predict the facial action unit corresponding to each frame of the image signal;
[0054] A second extraction subunit, configured to determine text information corresponding to the target video according to the facial action unit, and extract text features of the text information as the text features corresponding to the target video.
[0055] In a possible implementation manner, the apparatus further includes:
[0056] A fourth acquisition unit, configured to acquire a second sample video; and separate a second sample audio signal and a second sample image signal from the second sample video;
[0057] A fourth generation unit, configured to extract acoustic features of each frame of the second sample audio signal in the second sample video, input the acoustic features into an initial second decoder, and generate a predicted image signal of the second sample video;
[0058] A third training unit, configured to use a third loss function to perform consistency constraint on the predicted image signal and the second sample image signal, and use a fourth loss function to perform consistency constraint on the image features of the predicted image signal and the image features of the second sample image signal, and train to obtain the second decoder.
[0059] In a possible implementation manner, the recognition unit is specifically configured to:
[0060] Perform feature fusion on the image features of the image signal in the target video, the acoustic features of the audio signal in the target video, and the text features, and input the obtained fusion result into a fully connected layer for emotion recognition to obtain an emotion recognition result corresponding to the target user.
[0061] An embodiment of the present application further provides an emotion recognition device, including: a processor, a memory, and a system bus;
[0062] The processor and the memory are connected through the system bus;
[0063] The memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor is caused to execute any one of the implementation manners of the above emotion recognition method.
[0064] An embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is caused to execute any one of the implementation manners of the above emotion recognition method.
[0065] An embodiment of the present application further provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is caused to execute any one of the implementation manners of the above emotion recognition method.
[0066] An emotion recognition method, device, storage medium and equipment provided by an embodiment of the present application first obtain a target video to be recognized, where the target video is a video lacking a target modality signal, then extract the signal features of the non-target modality signal in the target video, and decode the signal features of the non-target modality signal to generate a target modality signal; then, extract the signal features of the target modality signal, and use the target modality signal or the non-target modality signal to generate text features corresponding to the target video; furthermore, the target user in the target video can be subjected to emotion recognition according to the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features, and an emotion recognition result corresponding to the target user can be obtained. It can be seen that since the present application first uses the signal features of the non-target modality signal that is not lacking in the target video to generate the lacking target modality signal, and then uses the signal features of the generated target modality signal, the signal features of the non-target modality signal, and the text features to jointly perform emotion recognition on the target user in the target video, the recognition basis is more accurate, thereby further improving the accuracy of the final emotion recognition result. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0068] Figure 1 It is a schematic flowchart of an emotion recognition method provided by an embodiment of the present application;
[0069] Figure 2 It is a schematic diagram of the process of constructing a first decoder provided by an embodiment of the present application;
[0070] Figure 3 It is a schematic diagram of the process of constructing a second decoder provided by an embodiment of the present application;
[0071] Figure 4 It is a schematic diagram of the training process in the case where the modality signal is not missing provided by an embodiment of the present application;
[0072] Figure 5 It is a schematic diagram of the composition of an emotion recognition device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] Emotion recognition is a crucial component of human-computer interaction, with a wide range of applications, including healthcare, automotive, and customer service. Single-modal emotion recognition has seen some progress, including facial, speech, and text-based emotion recognition. However, due to the limited information it receives, single-modal emotion recognition has not achieved satisfactory results.
[0074] To further improve the effectiveness of emotion recognition, multimodal emotion recognition has become a current research hotspot. This is because multimodal emotion recognition can obtain richer information, and the information between multiple modalities can usually complement each other, which is conducive to improving recognition accuracy. However, since multimodal emotion recognition requires the simultaneous reception of information from multiple modalities, the conditions for simultaneous reception of multiple modal information are not always available in actual use. For example, in some scenarios, when only cameras are deployed but no microphones, audio signals cannot be collected. It is also possible that the signal quality of some modalities is low and cannot be used. For example, in some scenes with long distances, it is more difficult to collect high-quality audio signals. In some cases, the face may be occluded, and the image signal of the visual modality cannot be used.
[0075] In this regard, the currently commonly used emotion recognition methods that support modality loss usually directly generate features of the missing modality from the features of the remaining modalities to supplement the missing modality. On the one hand, this end-to-end feature generation method lacks the generation of intermediate corresponding modal signals, and it is impossible to determine whether the generated features can accurately represent the corresponding model signals. On the other hand, since there are no direct constraints on the generated features, the accuracy of the generated features cannot be guaranteed, which may lead to low accuracy of the emotion recognition results obtained based on the generated features. Therefore, how to improve the accuracy of emotion recognition results when some modal signals are missing is a technical problem that needs to be solved urgently.
[0076] To address the above-mentioned defects, the present application provides an emotion recognition method, which first obtains a target video to be identified, wherein the target video is a video lacking a target modal signal, then extracts the signal features of the non-target modal signal in the target video, and decodes the signal features of the non-target modal signal to generate a target modal signal; then, extracts the signal features of the target modal signal, and uses the target modal signal or the non-target modal signal to generate text features corresponding to the target video; and then, based on the signal features of the target modal signal, the signal features of the non-target modal signal, and the text features, the emotion of the target user in the target video can be recognized, and the emotion recognition result corresponding to the target user can be obtained.
[0077] It can be seen that since this application first uses the signal features of non-target modality signals that are not missing in the target video to generate the missing target modality signals, and then uses the signal features of the generated target modality signals, the signal features of non-target modality signals, and text features to jointly perform emotion recognition on the target user in the target video, the recognition basis is more accurate, thereby being able to further improve the accuracy of the final emotion recognition result.
[0078] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0079] First Embodiment
[0080] See Figure 1 , which is a schematic flowchart of an emotion recognition method provided in this embodiment. The method includes the following steps:
[0081] S101: Obtain a target video to be recognized, where the target video is a video lacking a target modality signal.
[0082] In this embodiment, any video that needs to perform user emotion recognition is defined as the target video to be recognized. And the user who needs to perform emotion recognition included in the target video is defined as the target user. It should be noted that this embodiment does not limit the color type of the target video. For example, the target video can be a color video or a grayscale video, etc. And this embodiment does not limit the format type of the target video. For example, the target video can be in video formats such as mp4 or wmv (Windows Media Video). And this application also does not limit the scene type of the target video. For example, the target video can be a video clip of a movie or a short video in the teaching field, etc. In addition, this application also does not limit the length of the target video. For example, the target video can be a short video of 5 seconds or 8 seconds, etc.
[0083] It can be understood that the target video can be obtained by means such as shooting with a camera according to actual needs. For example, a 5-second video containing at least one user intercepted from a video stream, etc. And in order to be able to effectively improve the accuracy of the emotion recognition result in the case of missing some modality signals through the method provided in this embodiment, the target video obtained in this embodiment is a video lacking a target modality signal. Among them, the target modality signal can be an audio signal or an image signal missing in the target video.
[0084] Further, after obtaining the target video, existing or future video stream separation methods can be used to perform modal signal separation processing on the target video. For example, the open-source computer program FFmpeg (Fast Forward Mpeg) can be used to perform modal signal separation processing on the target video to extract the non-target modal signals contained in the target video, that is, to extract the model signals that are not missing in the target video as non-target modal signals, so as to execute the subsequent steps S102 - S104 to realize the emotion recognition of the target user in the target video and obtain a more accurate recognition result.
[0085] S102: Extract the signal features of the non-target modal signals in the target video, and decode the signal features of the non-target modal signals to generate target modal signals.
[0086] In this embodiment, after obtaining the target video to be recognized through step S101 and separating the non-target modal signals therefrom, in order to accurately identify the emotion category of the target user in the target video, existing or future feature extraction methods can be further used to extract the signal features of the non-target modal signals in the target video, and existing or future feature decoding methods can be used to perform decoding processing on the signal features of the non-target modal signals, so as to generate target modal signals according to the decoding results for executing the subsequent step S103.
[0087] S103: Extract the signal features of the target modal signals, and use the target modal signals or non-target modal signals to generate text features corresponding to the target video.
[0088] In this embodiment, after generating the target modal signals missing in the target video through step S102, in order to accurately identify the emotion category of the target user in the target video, existing or future feature extraction methods can be further used to extract the signal features of the target modal signals, and existing or future feature generation methods can be used to process the target modal signals or non-target modal signals, so as to generate text features corresponding to the target video according to the processing results for executing the subsequent step S104.
[0089] In a possible implementation of the embodiment of the present application, when the missing signal in the target video is an audio signal, that is, when the target modal signal is an audio signal and the non-target modal signal is an image signal, the specific implementation process of the above step S102 may be as follows: After separating the modal signals of the target video by using methods such as the open-source computer program FFmpeg to extract the non-target modal signals (i.e., image signals) contained in the target video, image feature extraction methods existing or emerging in the future can be used to extract the image features (such as feature maps or feature vectors, etc.) of the image signals in the target video, and the pre-trained first decoder is used to decode the image features to generate the spectrogram of the audio signal, so as to use the spectrogram of the audio signal to determine the missing audio signal in the target video.
[0090] On this basis, the specific implementation process of the above step S103 may be as follows: First, each frame of image signal in the target video is input into the pre-constructed facial action unit recognition model, so as to predict the facial action unit corresponding to each frame of image signal. Then, the text information corresponding to the target video can be determined according to the predicted facial action units, and the text features of the text information are extracted as the text features corresponding to the target video. The text features refer to the word vector sequence composed of the word vectors corresponding to all the words contained in the text information.
[0091] Next, the construction process of the facial action unit recognition model will be introduced in this embodiment. Among them, an optional implementation method is that the construction process of the facial action unit recognition model may specifically include: First, sample facial action unit data is obtained. The sample facial action unit data may include, but is not limited to, five facial action units: frowning, cheek raising, mouth corner pulling, lower lip raising, and lip parting. Then, the initial facial action unit recognition model is trained by using the sample facial action unit data and the target loss function to obtain the facial action unit recognition model.
[0092] Specifically, in this implementation method, in order to construct the facial action unit recognition model, a large number of preparatory works need to be carried out in advance. First, a large number of single-frame face image data sets containing facial action units (including but not limited to frowning, cheek raising, mouth corner pulling, lower lip raising, and lip parting, etc.) are collected as sample facial action unit data to form the model training data. And the facial action unit recognition results corresponding to these sample facial action unit data are manually labeled. Then, the initial facial action unit recognition model can be trained according to these sample facial action unit data, the facial action unit recognition results corresponding to the sample facial action unit data, and the target loss function, and then the facial action unit recognition model is generated.
[0093] Among them, an optional implementation is that the initial facial action unit recognition model can be (but not limited to) a Convolutional Neural Networks (CNN) model. Moreover, the value of the target loss function in this embodiment is not limited either, and it can be set according to the actual situation and empirical values. For example, it can be set as the Cross Entropy (CE) loss function, etc.
[0094] Specifically, during model training, a sample facial action unit data can be sequentially extracted from the training data as the model input, and the corresponding facial action unit recognition result as the output. Multiple rounds of model training are carried out, and the facial action unit recognition result obtained in each round of training is compared with the corresponding manually marked result, and the model parameters are updated according to the difference between the two until the preset conditions are met. For example, the value of the target loss function is very small and basically unchanged, then the update of the model parameters is stopped, and the training of the facial action unit recognition model is completed, generating a trained facial action unit recognition model.
[0095] It should be noted that although the video only contains two modalities, i.e., images and audio, in order to improve the accuracy of emotion recognition for users in the video, this application also introduces the features of the text modality as an auxiliary recognition basis to further improve the accuracy of the emotion recognition basis, so as to improve the accuracy of the subsequent recognition results. The current solutions for using text modality features for recognition usually obtain the text based on the result of audio recognition, that is, recognize based on the content spoken by the obtained speaker. However, the main drawback of this method is that: in most cases, the content of audio recognition is not strongly correlated with the emotional state. Using this text with weak emotional correlation to assist in emotion recognition helps little in improving the accuracy of emotion recognition. Therefore, this application proposes to generate text information more relevant to the emotional state by using the changes in the user's facial action units obtained by the facial action unit recognition model. Then, the corresponding text features are extracted from it as an auxiliary basis for emotion recognition.
[0096] Specifically, in the process of using the facial action unit recognition model to determine the text features corresponding to the target video, the first step is to input the image signal of each frame of the target video into the facial action unit recognition model for facial action unit classification. Taking the target video as a 5-second or 8-second short video as an example, since the changes in emotional state contained in the short video clip usually conform to the five situations of from weak to strong, from strong to weak, always strong, always weak, or sometimes strong and sometimes weak, the changes in the character's eyebrow movement, cheek movement, and mouth movement can be described in text based on the facial action unit classification results. Taking frowning as an example, if the predicted face in the entire target video always has a frowning action, the text information corresponding to the generated target video can be a text description such as "continuous frowning"; if the predicted frowning goes from presence to absence, the generated text information can be "brows from frowning to relaxing"; if the predicted frowning goes from non-existence to presence, the generated text information can be "brows from relaxing to frowning"; if the predicted frowning appears intermittently, the generated text information can be "brows sometimes relax and sometimes frown"; if there is no frowning action in the entire target video, the generated text information can be "brows always relaxed".
[0097] Similarly, other facial action units can be described in this way. For example, if the target user's face remains smiling throughout the entire target video, the resulting text information corresponding to the target video may be "brows are always relaxed, cheeks are always raised, corners of the mouth are always pulled upward, lower lip is not raised, and lips are always parted." After extracting the text features of this text information, they can be used as the text features corresponding to the target video to perform the subsequent step S104.
[0098] Next, this embodiment will introduce the construction process of the first decoder, wherein an optional implementation method is as follows: Figure 2 As shown, the construction process of the first decoder may specifically include: first obtaining a first sample video; and separating a first sample audio signal and a first sample image signal from the first sample video, and then using methods such as a facial action unit recognition model to extract image features of the first sample image signal of each frame in the first sample video, and inputting the obtained image features into the initial first decoder to generate a predicted spectrum graph of the predicted audio signal of the first sample video, and then using a first loss function to perform consistency constraints on the predicted spectrum graph and the spectrum graph of the first sample audio signal, and using a second loss function to perform consistency constraints on the acoustic features of the predicted audio signal and the acoustic features of the first sample audio signal, and thus training to obtain the first decoder.
[0099] Specifically, in this implementation manner, in order to construct the first decoder, a large number of preparatory works need to be carried out in advance. First of all, it is necessary to collect a large number of videos containing the speech and images emitted by users when speaking. For example, it can be achieved through microphone arrays for sound pickup and cameras for shooting. The sound pickup device can be a tablet computer, or intelligent hardware devices such as smart speakers, TVs, and air conditioners. And the collected video data of each item can be used as the first sample video respectively. Moreover, after obtaining the first sample video, it cannot be directly used to train and generate the first decoder. Instead, methods such as the open-source computer program FFmpeg need to be adopted to separate the first sample audio signal and the first sample image signal contained in each first sample video. And the emotion recognition results corresponding to the sample users in these first sample videos are manually labeled.
[0100] Then, as Figure 2 shown, the first sample image signal is input into the image feature extractor to obtain image features, and the first sample image signal is input into the facial action unit recognition model to obtain text information, and the text information is input into the text feature extractor to obtain text features.
[0101] Next, the image features are input into the initial first decoder to generate the predicted spectrogram of the predicted audio signal of the first sample video, and by adjusting the first loss function, consistency constraints are imposed on this spectrogram and the spectrogram of the first sample audio signal. The purpose is to make the generated predicted audio signal as close as possible to the real first sample audio signal. Among them, the function content of the first loss function is not limited in this application and can be set according to the actual situation and empirical values. A preferred implementation manner is that the first loss function can be set as the mean square error (MSE) loss function. The specific calculation formula is as follows:
[0102]
[0103] Among them, A represents the spectrogram of the first sample audio signal; represents the predicted spectrogram of the predicted audio signal of the first sample video generated.
[0104] Moreover, after generating the predicted spectrogram of the predicted audio signal of the first sample video, the predicted audio signal can be determined using the predicted spectrogram. After extracting the acoustic features of the predicted audio signal, the second loss function can be adjusted to impose a consistency constraint on the acoustic features of the predicted audio signal and the acoustic features of the first sample audio signal, with the aim of making the predicted acoustic features corresponding to the predicted audio signal as close as possible to the sample acoustic features of the first sample audio signal. Herein, the content of the second loss function is not limited in this application and can be set according to actual situations and empirical values. A preferred implementation is that the second loss function can also be set as the mean square error (MSE) loss function, and the specific calculation formula is as follows:
[0105]
[0106] Wherein, FA represents the sample acoustic features of the first sample audio signal; represents the predicted acoustic features corresponding to the generated predicted audio signal.
[0107] Furthermore, the generated image features, predicted acoustic features, and text features can be fused, and the processing result can be input into the fully connected layer to obtain the emotion recognition result of the first sample user in the first sample video.
[0108] Specifically, during model training, a first sample video and its corresponding emotion recognition result can be sequentially extracted from the training data for multiple rounds of training of the first decoder. The emotion prediction result obtained in each round of training is compared with the corresponding manually annotated result, and the parameters of the first decoder are updated according to the difference between the two until a preset condition is met, such as the values of the first loss function and the second loss function are very small and basically unchanged. Then, the update of the parameters of the first decoder is stopped, and the training of the first decoder is completed to generate a trained first decoder.
[0109] In another possible implementation manner of the embodiment of the present application, when the missing signal in the target video is an image signal, that is, when the target modal signal is an image signal and the non-target modal signal is an audio signal, the specific implementation process of the above step S102 may be as follows: After separating the modal signals of the target video by using methods such as the open-source computer program FFmpeg to extract the non-target modal signal (i.e., the audio signal) contained in the target video, the acoustic features of the audio signal in the target video (such as Mel-scale Frequency Cepstral Coefficients (MFCC) features or filterbank features, etc.) can be extracted by using existing or future acoustic feature extraction methods, and the second decoder pre-trained is used to decode the acoustic features to generate the missing image signal in the target video.
[0110] On this basis, the specific implementation process of the above step S103 may be as follows: First, each frame of the image signal in the generated target video is input into the pre-constructed facial action unit recognition model, so as to predict the facial action unit corresponding to each frame of the image signal. Then, according to the predicted facial action units, the text information corresponding to the target video can be determined, and the text features of the text information are extracted as the text features corresponding to the target video. Among them, the text features refer to the word vector sequence composed of the word vectors corresponding to all the words contained in the text information.
[0111] Next, this embodiment will introduce the construction process of the second decoder. Among them, an optional implementation manner is, as Figure 3 shown, the construction process of the second decoder may specifically include: First, obtain the second sample video; and separate the second sample audio signal and the second sample image signal from the second sample video. Then, the acoustic features of each frame of the second sample audio signal in the second sample video can be extracted, and the acoustic features are input into the initial second decoder to generate the predicted image signal of the second sample video. Next, the third loss function can be used to perform consistency constraints on the predicted image signal and the second sample image signal, and the fourth loss function can be used to perform consistency constraints on the image features of the predicted image signal and the image features of the second sample image signal to train and obtain the second decoder.
[0112] Specifically, in this implementation, to construct the second decoder, a large amount of preparatory work needs to be done in advance. First, a large number of videos containing the speech and images emitted by users during speech need to be collected. For example, sound can be picked up by a microphone array and images can be captured by a camera. The sound pickup device can be a tablet computer or intelligent hardware devices such as smart speakers, TVs, and air conditioners. And each piece of video data collected can be used as a second sample video respectively. Moreover, after obtaining the second sample videos, they cannot be directly used to train and generate the second decoder. Instead, methods such as the open-source computer program FFmpeg are needed to separate the second sample audio signals and second sample image signals contained in each second sample video. And the emotion recognition results corresponding to the sample users in these second sample videos are manually labeled.
[0113] Then, as Figure 3 shown, the second sample audio signal is input into the acoustic feature extractor to obtain acoustic features, and the acoustic features are input into the initial second decoder to generate the predicted image signal of the second sample video.
[0114] Next, by adjusting the third loss function, consistency constraints are imposed on the predicted image signal and the second sample image signal. The purpose is to make the generated predicted image signal as close as possible to the real second sample image signal. Among them, the function content of the third loss function is not limited in this application and can be set according to the actual situation and empirical values. A preferred implementation is that the third loss function can be set as the mean square error (MSE) loss function, and the specific calculation formula is as follows:
[0115]
[0116] Among them, V represents the second sample image signal; represents the predicted image signal of the generated second sample video.
[0117] Moreover, after generating the predicted image signal of the second sample video, the predicted image features of the predicted image signal can also be extracted. Thus, by adjusting the fourth loss function, consistency constraints are imposed on the predicted image features of the predicted image signal and the sample image features of the second sample image signal. The purpose is to make the predicted image features of the predicted image signal as close as possible to the sample image features of the second sample image signal. Among them, the function content of the fourth loss function is not limited in this application and can be set according to the actual situation and empirical values. A preferred implementation is that the fourth loss function can also be set as the mean square error (MSE) loss function, and the specific calculation formula is as follows:
[0118]
[0119] Among them, FVrepresenting the sample image features of the second sample image signal; representing the predicted image features corresponding to the generated predicted image signal.
[0120] Furthermore, the generated predicted image signal can be input into a facial action unit recognition model to generate the text information corresponding to the second sample video, and the text information can be input into a text feature extractor to obtain text features. Further, the generated predicted image features, acoustic features, and text features can be fused and processed, and the processing result can be input into a fully connected layer to obtain the emotion recognition result of the second sample user in the second sample video.
[0121] Specifically, when performing model training, a second sample video and its corresponding emotion recognition result can be sequentially extracted from the training data for multiple rounds of training of the second decoder, and the emotion prediction result obtained in each round of training can be compared with the corresponding manually annotated result, and the parameters of the second decoder can be updated according to the difference between the two until the preset conditions are met, such as the values of the third loss function and the fourth loss function are very small and basically unchanged, then the update of the parameters of the second decoder is stopped, the training of the second decoder is completed, and a trained second decoder is generated.
[0122] It should be noted that this application also uses sample video data that simultaneously includes audio signals and image signals as training data, and uses a facial action unit recognition model to perform emotion recognition training on the fully connected layer for scenarios where both modalities exist (i.e., when the modality signals are not missing). The training process is as Figure 4 shown, and the specific implementation process will not be elaborated here.
[0123] In this way, through the above training process, the facial action unit recognition model, the first decoder, the second decoder, and the fully connected layer obtained can support both multi-modal emotion recognition and multi-modal emotion recognition, and effectively improve the accuracy of the recognition result.
[0124] S104: Perform emotion recognition on the target user in the target video according to the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features, to obtain the emotion recognition result corresponding to the target user.
[0125] In this embodiment, after obtaining the signal features of the non-target modality signal in the target video through step S102, and obtaining the signal features of the target modality signal and the text features corresponding to the target video through step S103, further, the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features can be subjected to feature fusion processing, and according to the processing result, emotion recognition is performed on the target user in the target video to obtain the emotion recognition result corresponding to the target user.
[0126] Specifically, an optional implementation manner is that the image features of the image signal in the target video, the acoustic features of the audio signal in the target video, and the text features can be subjected to feature fusion, and the obtained fusion result is input into a fully connected layer for emotion recognition, so as to predict the emotion recognition result corresponding to the target user. Among them, the recognition result output by the fully connected layer can be a multi-dimensional vector. The emotion category corresponding to the highest dimension, or the emotion category corresponding to the dimension exceeding a preset threshold (the specific value is not limited), is used as the final emotion recognition result corresponding to the target user, that is, as the emotion recognition to which the target user belongs.
[0127] For example: Suppose the preset emotion categories are: "Joy", "Trust", "Fear", "Surprise", "Sadness", "Disgust", "Anger". The recognition result output by the fully connected layer is a 7-dimensional emotion category prediction vector [0.05, 0.07, 0.03, 0.41, 0.04, 0.08, 0.32]. Among them, each vector value represents the probability value corresponding to each preset emotion category. This probability value represents the degree to which the target user belongs to the corresponding emotion category. The larger the probability value, the higher the degree that the target user is of this emotion category. On the contrary, it indicates that the degree of the target user being of this emotion category is lower. It can be seen that the emotion type "Surprise" corresponding to the highest probability value (0.41) in the foregoing example is the emotion category to which the target user belongs.
[0128] In summary, an emotion recognition method provided in this embodiment first obtains a target video to be recognized, where the target video is a video lacking a target modality signal. Then, it extracts the signal features of the non-target modality signal in the target video, and decodes the signal features of the non-target modality signal to generate a target modality signal. Next, it extracts the signal features of the target modality signal, and uses the target modality signal or the non-target modality signal to generate the text features corresponding to the target video. Furthermore, it can perform emotion recognition on the target user in the target video according to the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features, so as to obtain the emotion recognition result corresponding to the target user. It can be seen that since this application first uses the signal features of the non-target modality signal that is not lacking in the target video to generate the lacking target modality signal, and then uses the signal features of the generated target modality signal, the signal features of the non-target modality signal, and the text features to jointly perform emotion recognition on the target user in the target video, the recognition basis is more accurate, thereby being able to further improve the accuracy of the final emotion recognition result.
[0129] <> Second Embodiment
[0130] This embodiment will introduce an emotion recognition device. For related content, please refer to the foregoing method embodiment.
[0131] See Figure 5 , which is a schematic diagram of the composition of an emotion recognition device provided in this embodiment. The device 500 includes:
[0132] A first acquisition unit 501, configured to acquire a target video to be recognized, where the target video is a video lacking a target modality signal;
[0133] A first generation unit 502, configured to extract signal features of non-target modality signals in the target video, and decode the signal features of the non-target modality signals to generate the target modality signal;
[0134] A second generation unit 503, configured to extract signal features of the target modality signal, and generate text features corresponding to the target video by using the target modality signal or the non-target modality signal;
[0135] An identification unit 504, configured to perform emotion recognition on a target user in the target video according to the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features, to obtain an emotion recognition result corresponding to the target user.
[0136] In an implementation manner of this embodiment, the target modality signal is an audio signal, and the non-target modality signal is an image signal; specifically, the first generation unit 502 is configured to:
[0137] Extract image features of the image signal in the target video, and decode the image features by using a first decoder to generate a spectrogram of the audio signal, so as to determine the audio signal by using the spectrogram of the audio signal.
[0138] In an implementation manner of this embodiment, the second generation unit 503 includes:
[0139] A first prediction sub-unit, configured to input each frame of image signal in the target video into a pre-constructed facial action unit recognition model, and predict the facial action unit corresponding to each frame of image signal;
[0140] A first extraction sub-unit, configured to determine text information corresponding to the target video according to the facial action unit, and extract text features of the text information as text features corresponding to the target video.
[0141] In an implementation manner of this embodiment, the device further includes:
[0142] A second acquisition unit, configured to acquire sample facial action unit data, where the sample facial action unit data includes five facial action units: frowning, cheek raising, mouth pulling, lower lip raising, and lip parting;
[0143] A first training unit, configured to train an initial facial action unit recognition model by using the sample facial action unit data and a target loss function, so as to obtain the facial action unit recognition model.
[0144] In an implementation manner of this embodiment, the apparatus further includes:
[0145] A third acquisition unit, configured to acquire a first sample video, and separate a first sample audio signal and a first sample image signal from the first sample video;
[0146] A third generation unit, configured to extract image features of each frame of the first sample image signal in the first sample video, input the image features into an initial first decoder, and generate a predicted spectrogram of a predicted audio signal of the first sample video;
[0147] A second training unit, configured to perform consistency constraint on the predicted spectrogram and the spectrogram of the first sample audio signal by using a first loss function, and perform consistency constraint on the acoustic features of the predicted audio signal and the acoustic features of the first sample audio signal by using a second loss function, so as to train and obtain the first decoder.
[0148] In an implementation manner of this embodiment, the target modality signal is an image signal, and the non-target modality signal is an audio signal; specifically, the first generation unit 502 is configured to:
[0149] Extract the acoustic features of the audio signal in the target video, and decode the acoustic features by using a second decoder to generate the image signal.
[0150] In an implementation manner of this embodiment, the second generation unit 503 includes:
[0151] A second prediction subunit, configured to input each frame of the generated image signal in the target video into a pre-constructed facial action unit recognition model, and predict the facial action unit corresponding to each frame of the image signal;
[0152] A second extraction subunit, configured to determine the text information corresponding to the target video according to the facial action unit, and extract the text features of the text information as the text features corresponding to the target video.
[0153] In an implementation manner of this embodiment, the apparatus further includes:
[0154] A fourth acquisition unit, configured to acquire a second sample video, and separate a second sample audio signal and a second sample image signal from the second sample video;
[0155] A fourth generation unit, configured to extract acoustic features of each frame of the second sample audio signal in the second sample video, input the acoustic features into an initial second decoder, and generate a predicted image signal of the second sample video;
[0156] A third training unit, configured to use a third loss function to perform consistency constraint on the predicted image signal and the second sample image signal, and use a fourth loss function to perform consistency constraint on the image features of the predicted image signal and the image features of the second sample image signal, and train to obtain the second decoder.
[0157] In an implementation manner of this embodiment, the recognition unit is specifically configured to:
[0158] Perform feature fusion on the image features of the image signal in the target video, the acoustic features of the audio signal in the target video, and the text features, and input the obtained fusion result into a fully connected layer for emotion recognition to obtain an emotion recognition result corresponding to the target user.
[0159] Furthermore, an embodiment of the present application further provides an emotion recognition device, including: a processor, a memory, and a system bus;
[0160] The processor and the memory are connected through the system bus;
[0161] The memory is configured to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor is caused to execute any implementation method of the above emotion recognition method.
[0162] Furthermore, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is caused to execute any implementation method of the above emotion recognition method.
[0163] Furthermore, an embodiment of the present application further provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is caused to execute any implementation method of the above emotion recognition method.
[0164] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0165] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to describe the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0166] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0167] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An emotion recognition method, characterized in that Including: Obtain a target video to be recognized, where the target video is a video lacking a target modality signal; Extract the signal features of the non-target modality signals in the target video, and decode the signal features of the non-target modality signals to generate the target modality signal; Extract the signal features of the target modality signal, and use the target modality signal or the non-target modality signal to generate the text features corresponding to the target video; According to the signal features of the target modality signal, the signal features of the non-target modality signal, and the text features, perform emotion recognition on the target user in the target video to obtain the emotion recognition result corresponding to the target user; When the target modality signal is an audio signal and the non-target modality signal is an image signal, the generating the text features corresponding to the target video by using the target modality signal or the non-target modality signal includes: Input each frame of image signal in the target video into a pre-constructed facial action unit recognition model, predict the facial action unit corresponding to each frame of image signal, determine the text information corresponding to the target video according to the predicted facial action unit, and extract the text features of the text information as the text features corresponding to the target video; Or, when the target modality signal is an image signal and the non-target modality signal is an audio signal, the generating the text features corresponding to the target video by using the target modality signal or the non-target modality signal includes: Input each frame of image signal in the generated target video into a pre-constructed facial action unit recognition model, predict the facial action unit corresponding to each frame of image signal, determine the text information corresponding to the target video according to the predicted facial action unit, and extract the text features of the text information as the text features corresponding to the target video.
2. The method according to claim 1, characterized in that The target modality signal is an audio signal and the non-target modality signal is an image signal; the extracting the signal features of the non-target modality signals in the target video and decoding the signal features of the non-target modality signals to generate the target modality signal includes: Extract the image features of the image signals in the target video, and use a first decoder to decode the image features to generate the spectrogram of the audio signal, so as to determine the audio signal by using the spectrogram of the audio signal.
3. The method according to claim 1, wherein The construction method of the facial action unit recognition model is as follows: Obtain sample facial action unit data, where the sample facial action unit data includes five facial action units: frowning, cheek raising, mouth corner pulling, lower lip raising, and lip parting; Use the sample facial action unit data and a target loss function to train an initial facial action unit recognition model to obtain the facial action unit recognition model.
4. The method according to claim 2, wherein The construction method of the first decoder is as follows: Obtain a first sample video, and separate a first sample audio signal and a first sample image signal from the first sample video; Extract the image features of each frame of the first sample image signal in the first sample video, input the image features into an initial first decoder to generate the predicted spectrogram of the predicted audio signal of the first sample video; Using a first loss function, perform consistency constraints on the predicted spectrogram and the spectrogram of the first sample audio signal, and using a second loss function, perform consistency constraints on the acoustic features of the predicted audio signal and the acoustic features of the first sample audio signal, and train to obtain the first decoder.
5. The method according to claim 1, wherein The target modal signal is an image signal, and the non-target modal signal is an audio signal; extracting the signal features of the non-target modal signal in the target video and decoding the signal features of the non-target modal signal to generate the target modal signal includes: Extracting the acoustic features of the audio signal in the target video and using a second decoder to decode the acoustic features to generate the image signal.
6. The method according to claim 5, characterized in that, The second decoder is constructed as follows: Obtain a second sample video, and separate a second sample audio signal and a second sample image signal from the second sample video; Extract the acoustic features of each frame of the second sample audio signal in the second sample video, input the acoustic features into an initial second decoder, and generate a predicted image signal of the second sample video; Using a third loss function, perform consistency constraints on the predicted image signal and the second sample image signal, and using a fourth loss function, perform consistency constraints on the image features of the predicted image signal and the image features of the second sample image signal, and train to obtain the second decoder.
7. The method according to any one of claims 1-6, characterized in that, The method for performing emotion recognition on the target user in the target video according to the signal features of the target modal signal, the signal features of the non-target modal signal, and the text features to obtain the emotion recognition result corresponding to the target user includes: Performing feature fusion on the image features of the image signal in the target video, the acoustic features of the audio signal in the target video, and the text features, and inputting the obtained fusion result into a fully connected layer for emotion recognition to obtain the emotion recognition result corresponding to the target user.
8. An emotion recognition device, characterized in that, Including: A first acquisition unit for acquiring a target video to be recognized, where the target video is a video lacking a target modal signal; A first generation unit for extracting the signal features of the non-target modal signal in the target video and decoding the signal features of the non-target modal signal to generate the target modal signal; A second generation unit for extracting the signal features of the target modal signal and using the target modal signal or the non-target modal signal to generate the text features corresponding to the target video; An identification unit for performing emotion recognition on the target user in the target video according to the signal features of the target modal signal, the signal features of the non-target modal signal, and the text features to obtain the emotion recognition result corresponding to the target user; When the target modal signal is an audio signal and the non-target modal signal is an image signal, the second generation unit includes: A first prediction subunit for inputting each frame of the image signal in the target video into a pre-constructed facial action unit recognition model to predict the facial action unit corresponding to each frame of the image signal; A first extraction subunit, configured to determine text information corresponding to a target video according to predicted facial action units, and extract text features of the text information as text features corresponding to the target video; Alternatively, when the target modality signal is an image signal and the non-target modality signal is an audio signal, the second generation unit includes: A second prediction subunit, configured to input each frame of image signal in the generated target video into a pre-constructed facial action unit recognition model to predict facial action units corresponding to each frame of image signal; A second extraction subunit, configured to determine text information corresponding to a target video according to predicted facial action units, and extract text features of the text information as text features corresponding to the target video.
9. An emotion recognition device, characterized in that, Comprising: A processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is configured to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is caused to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Audio-visual information mutual-generation device and training system based on cyclic generative adversarial network
CN108256627A
Emotion analysis system and method based on probability emotion dictionary
CN111859925A