Emotion recognition model training method and device, electronic equipment and storage medium
By performing feature processing and alignment of multimodal emotional features, eliminating intermodal heterogeneity and optimizing network parameters of the emotion recognition model, the problem of low multimodal emotional recognition accuracy in the prior art is solved, and a more efficient emotion recognition effect is achieved.
Patent Information
- Application Number
- CN202311651098.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-06
AI Technical Summary
The existing emotion recognition method based on multimodal features fails to effectively consider the differences between different modes, resulting in poor feature extraction accuracy and affecting the accuracy of emotion recognition.
By acquiring multiple sets of sample data, feature processing and feature spatial alignment are performed, heterogeneity between different modes is eliminated, and network parameters of the emotion recognition model are optimized based on the alignment loss results.
It improves the interaction efficiency when multimodal feature fusion, enhances the representation ability of the emotion recognition model, and improves the accuracy of the emotion recognition results.
Smart Images

Figure CN120105136A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of neural network technology, and in particular to a training method, device, electronic device and storage medium for an emotion recognition model. Background Art
[0002] Good emotion recognition capabilities can be widely used in interactive games, interactive movies, virtual tour guides, virtual assistants, and customer service agents, helping humans and computers to interact and understand each other, and improving the efficiency and quality of human-computer interaction. Emotion recognition based on multimodal features can effectively improve the accuracy of recognition results.
[0003] At present, emotion recognition methods based on multimodal features often directly fuse the features, decision results or models of different modalities without considering the differences between different modalities, which leads to poor feature extraction accuracy and affects the accuracy of emotion recognition. Summary of the invention
[0004] The purpose of this application is to provide a training method, device, electronic device and storage medium for an emotion recognition model, so as to improve the emotion recognition results under multimodal emotion data.
[0005] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows:
[0006] In a first aspect, an embodiment of the present application provides a method for training an emotion recognition model, comprising:
[0007] Acquire multiple sets of sample data, each set of sample data includes: emotion features of multiple modalities and emotion labels;
[0008] Inputting multiple groups of sample data into the initial emotion recognition model, performing feature processing on the emotion features of each mode in each group of sample data, and obtaining the target emotion features of each mode in each group of sample data;
[0009] Performing feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identifying the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data;
[0010] According to the target emotion features of each modality in each group of sample data, the alignment loss results between the two modalities are determined, and based on the alignment loss results between the two modalities, the network parameters of the initial emotion recognition model are iteratively corrected to obtain the emotion recognition model.
[0011] In a second aspect, the embodiment of the present application also provides a training device for an emotion recognition model, including: a collection module, a processing module, and a correction module;
[0012] The acquisition module is used to obtain multiple groups of sample data, each group of sample data includes: emotional features of multiple modalities and emotional labels;
[0013] The processing module is used to input multiple groups of sample data into the initial emotion recognition model, perform feature processing on the emotion features of each modality in each group of sample data, and obtain the target emotion features of each modality in each group of sample data;
[0014] The processing module is used to perform feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identify the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data;
[0015] The correction module is used to determine the alignment loss results between two modalities according to the target emotion characteristics of each modality in each group of sample data, and iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results between two modalities to obtain the emotion recognition model.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the machine-readable instructions to execute the training method of the emotion recognition model provided in the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the training method of the emotion recognition model provided in the first aspect is executed.
[0018] The beneficial effects of this application are:
[0019] The present application provides a training method, device, electronic device and storage medium for an emotion recognition model, which can be used to train a model for emotion recognition based on multimodal emotion features by collecting multiple groups of emotion features containing multiple modalities as sample data. Among them, by performing feature processing on the emotion features of each modality, the expression ability of the emotion features can be improved. Then, by aligning the multimodal emotion features, the heterogeneity between different modalities can be eliminated and the interaction efficiency during the fusion of multimodal features can be improved. The network parameters of the emotion recognition model are optimized based on the alignment loss result constructed based on the multimodal alignment processing, so that the trained emotion recognition model can better fuse the multimodal emotion features, make full use of the multimodal emotion features for comprehensive analysis and recognition, and improve the accuracy of the emotion recognition results of the trained emotion recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A flowchart of a method for training an emotion recognition model provided in an embodiment of the present application;
[0022] Figure 2 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application;
[0023] Figure 3 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application;
[0024] Figure 4 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application;
[0025] Figure 5 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application;
[0026] Figure 6 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application;
[0027] Figure 7 A schematic diagram of a flow chart of a recognition method of an emotion recognition model provided in an embodiment of the present application;
[0028] Figure 8A schematic diagram of the architecture of an emotion recognition model provided in an embodiment of the present application;
[0029] Fig. 9 A schematic diagram of a training device for an emotion recognition model provided in an embodiment of the present application;
[0030] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.
[0032] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0033] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0034] First, the relevant background of this plan is explained:
[0035] Humans can interpret the emotions of their interlocutors through a variety of signals such as spoken language, synchronized speech, and facial expressions, and giving machines the ability to understand emotions has always been the goal of engineers and researchers. Multimodal emotion recognition refers to the analysis and recognition of human emotional states through a variety of different information sources (such as facial expressions, speech, text, etc.). This ability can be widely used in scenarios such as interactive games, interactive movies, virtual tour guides, virtual assistants, and customer service agents to help humans interact and understand each other with computers, and improve the efficiency and quality of human-computer interaction. At present, the understanding and fusion of multimodal information remains an important challenge for multimodal emotion recognition. Multimodal emotion recognition requires the fusion of multimodal features from different information sources to extract more accurate and robust emotional features. However, the signals between different information sources are closely related, but also different and complex. Therefore, how to effectively fuse them is a difficult problem.
[0036] Existing technical solutions usually identify emotions from a single modality or directly fuse multiple different modalities to obtain the final expression features, while ignoring the differences between multimodal features. These methods cannot handle the nonlinear relationship and complex interaction effects between modalities, and the inability to fully extract effective emotion features may lead to a decrease in emotion recognition accuracy.
[0037] The following are some of the commonly used emotion recognition methods:
[0038] First, manual features were widely used in early sentiment analysis. This method describes faces by manually defining some visual features, such as local binary patterns (LBP), histograms of oriented gradients (HOG), etc., and then uses machine learning classifiers, such as support vector machines, to perform sentiment classification.
[0039] The advantages of this type of method are fast processing speed, easy implementation and explanation, and suitable for smaller data sets and simple sentiment classification tasks. However, since manual features are greatly affected by human subjective experience and the environment, their robustness and generalization ability are poor. In complex natural environments, sentiment analysis based on manual features often cannot achieve accurate recognition results.
[0040] Second: Sentiment analysis based on deep learning uses the powerful feature extraction capabilities of convolutional neural networks. This method is suitable for dealing with complex occlusion, lighting and other issues. This type of method is widely used in sentiment analysis in natural environments. This method uses deep neural networks to fuse information from different modalities, such as using multi-layer perceptron (MLP), convolutional neural network (CNN), recurrent neural network (RNN) and other models for feature extraction and fusion, thereby achieving emotion recognition. Emotion recognition in natural environments can be divided into the following two categories:
[0041] (1) Emotion recognition based on single modality: Single modality sentiment analysis is based on input information of a single modality (such as images, voice or text, etc.), and uses convolutional neural networks to extract it into single modality emotion representation. Single modality emotion recognition cannot fully utilize information from other modalities, which may lead to inaccuracy in emotion recognition and insufficient generalization ability. In addition, the reliability of single modality emotion recognition may not be high, because it only relies on information from one modality and may be interfered by factors such as environmental noise and language habits. Therefore, this method has high requirements for input information of a specific modality, and its robustness and anti-interference ability are poor, and it cannot obtain more effective information. In addition, since this method can only recognize emotions of a specific modality, such as only recognizing voice emotions, image emotions or text emotions, it cannot fully recognize human emotions.
[0042] (2) Multimodal emotion recognition: Multimodal sentiment analysis requires learning a multimodal fusion representation from the features of different modalities (such as images, speech or text, etc.) for the final emotion recognition. According to the stages and strategies of different modal fusion, the existing multimodal emotion recognition methods can be roughly divided into feature-level fusion, decision-level fusion, model-level fusion and other methods.
[0043] Feature-level fusion is to fuse features from different modalities, for example, concatenate or weighted sum the features of speech, image, and text. The disadvantage is that it cannot handle nonlinear relationships and complex interaction effects between modalities.
[0044] Decision-level fusion is to fuse the decisions from different modalities. For example, the decision results of speech, image and text are weighted averaged or voted. The disadvantage is that the information and characteristics of different modalities cannot be fully utilized.
[0045] Model-level fusion is to use multiple models to process information of different modalities separately, and then fuse their outputs. For example, use a speech recognition model, an image recognition model, and a text classification model to process speech, images, and text respectively, and then fuse their outputs. The disadvantage is that multiple models need to be trained, and the computational complexity is high.
[0046] In summary, the existing emotion recognition methods generally have the following problems: Traditional manual feature extraction methods are greatly affected by human subjective experience and environmental factors, resulting in poor robustness and generalization. In complex natural environments, sentiment analysis based on manual features often finds it difficult to achieve accurate recognition results. In addition, manual features have poor portability. When the data source changes, manual features cannot adapt to the new data and need to be redesigned. In addition, the time cost of extracting manual features is high.
[0047] Single-modal emotion recognition based on convolutional neural networks cannot fully utilize information from other modalities, which may lead to inaccurate emotion recognition and insufficient generalization ability. The reliability of single-modal emotion recognition may not be high because it relies only on information from one modality and may be affected by factors such as environmental noise and language habits. In addition, since this method can only recognize emotions in a specific modality, such as voice emotions, image emotions, or text emotions, it cannot fully recognize human emotions.
[0048] At present, some multimodal emotion methods based on convolutional neural networks often directly fuse the features, decision results or models of different modalities without considering the differences between different modalities. The heterogeneity between different modalities leads to complex interaction effects and nonlinear relationships between different modalities. If the heterogeneous features are not processed before feature fusion to eliminate the gap between different modal features, it will lead to the inability to fully extract effective emotion features, which may lead to a decrease in emotion recognition accuracy.
[0049] Based on this, this scheme proposes a training method for the emotion recognition model. By performing dimensionality enhancement on the input multimodal sample data, the sample features of each modality can be fully captured; by performing nonlinear transformation on the multimodal sample data, the nonlinear ability of the features can be enhanced; by aligning the multimodal features, the heterogeneity between different modalities can be eliminated, and the interaction efficiency during multimodal feature fusion can be improved; by performing feature fusion on the aligned multimodal features, the representation ability of the model can be enhanced, and more robust and effective multimodal fusion features can be extracted, thereby training a high-performance emotion recognition model. When applied to the multimodal emotion recognition scenario, it can obtain accurate emotion recognition results based on multimodal emotion data.
[0050] Figure 1 A flowchart of a method for training an emotion recognition model provided in an embodiment of the present application; the execution subject of the method may be a computer device or a server or other device with data processing capabilities, such as Figure 1 As shown, the method may include:
[0051] S101. Acquire multiple groups of sample data, each group of sample data including: emotion features of multiple modalities and emotion labels.
[0052] Optionally, multiple sets of emotion data may be collected from a video or a database, and a set of emotion data may include emotion data of at least two modes, such as image segments, text segments, or voice segments, etc. The emotion data of multiple modes included in a set of emotion data all correspond to the same emotion, that is, the image segments, text segments, and voice segments are all used to represent the same emotion.
[0053] The multimodal emotion data in the same set of emotion data can be captured from the video, for example, capturing an image segment in the video, and taking the character voice corresponding to the image segment as a voice segment, and taking the subtitles corresponding to the image segment as a text segment.
[0054] The trained convolutional neural network can be used as a unimodal feature extractor to extract the emotion features of each modality from the emotion data of each modality in each group of emotion data, and obtain the multimodal emotion features in each group of sample data.
[0055] Since the emotion data of each mode in each set of emotion data corresponds to the same emotion, the emotion corresponding to each set of emotion data can be used as the emotion label in the sample data corresponding to each set of emotion data.
[0056] In one feasible manner, pre-trained convolutional neural networks can be set as single-modal feature extractors to extract the emotional features of each modality from each set of emotional data.
[0057] Among them, the emotional data of each modality can be input into the pre-trained model of the corresponding modality to extract the emotional features of each modality to obtain multiple groups of sample data.
[0058] Taking multi-modality including image, text, and speech as an example, in this embodiment, Facet can be used to extract facial expression features in image data under image modality, COVAREP (Collaborative Speech Analysis Library) model can be used to extract speech features in speech data under speech modality, and GloVe (GloVe is Global Vectors for Word Representation, an algorithm for word vector representation) can be used to extract text features in text data under text modality. The current actual modalities are not limited to the listed ones, and the pre-trained models used are not limited to the listed ones.
[0059] S102, inputting multiple groups of sample data into an initial emotion recognition model, performing feature processing on the emotion features of each mode in each group of sample data, and obtaining target emotion features of each mode in each group of sample data.
[0060] In this embodiment, for the emotional features of each modality in each group of sample data, at least one feature processing method can be used to process the emotional features of each modality separately. The purpose of feature processing is to enhance the expressive power of the emotional features of each modality so that the emotional features of each modality can be better used for model training to obtain a more accurate model.
[0061] By performing feature processing on the emotional features of each mode in each group of sample data, the target emotional features of each mode in each group of sample data can be obtained.
[0062] S103, performing feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and based on the aligned emotion features of each modality corresponding to each group of sample data, identifying and obtaining the emotion prediction results corresponding to each group of sample data.
[0063] Since there are differences between different modal information, if the differences between different modal emotional features can be eliminated or reduced, the fusion between different modal emotional features can be improved.
[0064] In this embodiment, based on the target emotion features of each modality in each group of sample data, contrastive learning can be introduced to align the multimodal emotion features. Optionally, for each group of sample data, the target emotion features of each modality in each group of sample data can be aligned to the same feature space to eliminate the heterogeneity between different modal features, and obtain the aligned emotion features of each modality corresponding to each group of sample data.
[0065] Next, the initial emotion recognition model is used to perform emotion prediction on the aligned emotion features of each modality corresponding to each group of sample data to obtain the emotion prediction results corresponding to each group of sample data.
[0066] S104. Determine the alignment loss results between the two modalities according to the target emotion features of each modality in each group of sample data, and iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities to obtain the emotion recognition model.
[0067] Optionally, this solution can construct a calculation framework of alignment loss based on alignment processing to calculate the alignment loss result, so that the network parameters of the initial emotion recognition model can be iteratively modified based on the alignment loss result. Among them, the goal of optimizing the initial emotion recognition model based on alignment loss is to maximize the matching degree between the target emotion features of each modality in the same group of sample data, and minimize the matching degree between the target emotion features of each modality in different groups of sample data, so as to improve the accuracy of the target emotion features of each modality in each group of sample data, and then improve the recognition accuracy of the trained emotion recognition model.
[0068] In summary, the training method of the emotion recognition model provided in this embodiment can be used to train a model for emotion recognition based on multimodal emotion features by collecting multiple groups of emotion features containing multiple modalities as sample data. Among them, by performing feature processing on the emotion features of each modality, the expressiveness of the emotion features can be improved. Then, by aligning the multimodal emotion features, the heterogeneity between different modalities can be eliminated and the interaction efficiency during the fusion of multimodal features can be improved. The network parameters of the emotion recognition model are optimized based on the alignment loss result constructed based on the multimodal alignment processing, so that the trained emotion recognition model can better fuse the multimodal emotion features, make full use of the multimodal emotion features for comprehensive analysis and recognition, and improve the accuracy of the emotion recognition results of the trained emotion recognition model.
[0069] Optionally, in step S102, feature processing is performed on the emotional features of each modality in each group of sample data to obtain target emotional features of each modality in each group of sample data, which may include: performing time series aggregation processing, dimensionality enhancement processing and nonlinear transformation processing on the emotional features of each modality in each group of sample data to obtain target emotional features of each modality in each group of sample data.
[0070] First, the emotional features of each modality in each group of sample data are processed in time series aggregation. Since the emotional features of each modality in each group of input sample data are composed of multiple frames of data, which are features composed of a time series segment, they can include multiple feature vectors. In order to facilitate the subsequent fusion of multimodal features, a global pooling layer along the time dimension can be used to perform time series aggregation on the multimodal emotional features in each group of sample data, thereby aggregating multiple feature vectors of each modality into a complete feature vector.
[0071] Secondly, the multi-modal emotional features after time series aggregation are processed by dimensionality enhancement. Multi-layer perceptrons can be used to deepen the emotional features of each modality and enhance the dimension of the emotional features of each modality to enhance the feature expression ability. Then, the emotional features of each modality are processed by nonlinear transformation to obtain the emotional features of each modality with nonlinear fitting ability, thereby improving the fitting ability of the model.
[0072] Optionally, the initial emotion recognition model may include: a multimodal alignment module. In step S103, feature space alignment processing is performed on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, which may include: inputting the target emotion features of each modality in the current group of sample data into the multimodal alignment module, and converting the target emotion features of each modality from the feature space corresponding to each modality to a preset feature space through the multimodal alignment module, so as to obtain the aligned emotion features of each modality corresponding to the current group of sample data.
[0073] Features of different modalities have their own feature spaces, and the feature spaces of different modalities are different. This can be understood as the way features of different modalities are expressed differently, and it is impossible to accurately identify feature information of different modalities through one model.
[0074] Therefore, a preset feature space can be set to convert the target emotional features of each modality in a set of sample data from their respective feature spaces to a uniformly set preset feature space to eliminate the barriers between features of different modalities.
[0075] After being processed by the multimodal alignment module, the aligned emotional features of each modality corresponding to each group of sample data can be obtained.
[0076] Figure 2 A flow chart of another method for training an emotion recognition model provided in an embodiment of the present application; Optionally, in step S103, based on the aligned emotion features of each modality corresponding to each group of sample data, an emotion prediction result corresponding to each group of sample data is identified, including:
[0077] S201, performing feature concatenation on the aligned emotion features of each modality corresponding to the current group of sample data to obtain a multimodal fusion feature corresponding to the current group of sample data.
[0078] In order to fully integrate the emotional features in each modality, the feature splicing method can be used to fuse the aligned emotional features of each modality in a set of sample data and integrate them into a comprehensive multimodal fusion feature. That is, the aligned emotional features of each modality contained in a set of sample data are fused, and a set of sample data corresponds to a multimodal fusion feature.
[0079] Since the emotional features of each modality are aligned, the barriers between different modalities are alleviated, multimodal interaction and fusion can be completed more effectively, and more robust multimodal fusion features can be extracted.
[0080] S202: Identify and obtain the emotion prediction result corresponding to the current group of sample data according to the multimodal fusion features corresponding to the current group of sample data.
[0081] Optionally, based on the obtained multimodal fusion features corresponding to each group of sample data, the emotion classification capability of the initial emotion recognition model can be trained to obtain the emotion prediction results corresponding to each group of sample data.
[0082] Figure 3A flow chart of another method for training an emotion recognition model provided in an embodiment of the present application; Optionally, the target emotion features of each modality include: image target emotion features, text target emotion features, and speech target emotion features; In step S104, according to the target emotion features of each modality in each group of sample data, determining the alignment loss results between the two modalities may include:
[0083] S301, determining the alignment loss result between the image emotion feature and the text emotion feature according to the image target emotion feature and the text target emotion feature in each group of sample data.
[0084] The target emotion features of each modality in each group of sample data are extracted, and contrastive learning is introduced to perform multi-modal feature alignment processing. Specifically, the target emotion features of each modality in multiple modalities in each group of sample data can be aligned.
[0085] During training, given N groups of target emotion features containing multiple modalities, the multimodal feature alignment process is used to predict the matching between emotion features of different modalities. Here we take images and texts as examples. The input of the multimodal feature alignment process is n target image emotion feature vectors and n target text emotion feature vectors. Therefore, the result of the alignment process is to predict n×n matching situations. Here we use cosine similarity to measure the matching degree between the target emotion features of the two modalities. Our training goal is to maximize the cosine similarity of the n text-image pairs that are truly matched in each group of sample data, while minimizing the cosine similarity of the remaining (n^2-n) unmatched image-text pairs. During training, InfoNCE loss can be used as the loss function for the alignment process. Then, according to the target emotion features of each modality in each group of sample data, the alignment loss results between the two modalities can be calculated. The first loss result is also the network loss caused by feature alignment.
[0086] In this embodiment, the calculation of the alignment loss result is explained using three modalities of speech, image and text as examples.
[0087] An alignment loss result can be calculated correspondingly between each modality, and the alignment loss result between the image emotion feature and the text emotion feature can be calculated using the first calculation formula based on the image target emotion feature and the text target emotion feature in each group of sample data input into the multimodal alignment module.
[0088] The first formula can be shown as follows:
[0089]
[0090] Among them, I k , T kThey represent the target emotion features of images and texts in the kth group of sample data, respectively. τ is the temperature coefficient, which is used to adjust the model's discrimination of negative samples.
[0091] S302: Determine the alignment loss result between the image emotion feature and the speech emotion feature according to the image target emotion feature and the speech target emotion feature in each group of sample data.
[0092] The alignment loss result between the image emotion features and the speech emotion features can be calculated using the corresponding second calculation formula based on the image target emotion features and the speech target emotion features in each group of sample data input into the multimodal alignment module.
[0093] The second formula can be shown as follows:
[0094]
[0095] Among them, I k , A k They represent the target emotion features of images and speech in the kth group of sample data, respectively. τ is the temperature coefficient, which is used to adjust the model's discrimination of negative samples.
[0096] S303: Determine the alignment loss result between the text emotion feature and the speech emotion feature according to the text target emotion feature and the speech target emotion feature in each group of sample data.
[0097] The alignment loss result between the text emotion features and the speech emotion features can be calculated using the corresponding third calculation formula based on the text target emotion features and the speech target emotion features in each group of sample data input into the multimodal alignment module.
[0098] The third formula can be shown as follows:
[0099]
[0100] Among them, T k , A k They represent the text target emotion features and speech target emotion features in the kth group of sample data respectively. τ is the temperature coefficient, which is used to adjust the model's discrimination of negative samples.
[0101] In this way, the alignment loss results between each two modes can be calculated separately.
[0102] Figure 4 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application; optionally, the method may further include:
[0103] S401, determining a classification loss result according to the emotion prediction results corresponding to each group of sample data and the emotion labels in each group of sample data.
[0104] In some embodiments, the loss result when modifying the network parameters of the initial emotion recognition model also includes the classification loss result, that is, the loss result of the initial emotion recognition model is constructed based on the classification loss result and the alignment loss result, so that the trained emotion recognition model can not only better eliminate the heterogeneity between different modal features, but also have a better classification effect.
[0105] Optionally, a classification loss result may be calculated using a cross entropy loss calculation method according to the emotion prediction results corresponding to each group of sample data and the emotion labels corresponding to each group of sample data.
[0106] The calculation formula for the classification loss result is as follows:
[0107]
[0108] Among them, y i represents the emotion label in the i-th group of sample data, Represents the emotion prediction result corresponding to the i-th group of sample data.
[0109] In step S104, based on the alignment loss results between the two modalities, iteratively correcting the network parameters of the initial emotion recognition model may include:
[0110] S402: based on the alignment loss results between the two modalities and the classification loss results, iteratively correct the network parameters of the initial emotion recognition model.
[0111] Then, the alignment loss results between the two modalities and the classification loss results can be combined as the total loss results of the model to iteratively correct the network parameters of the initial emotion recognition model.
[0112] Figure 5 A flowchart of another method for training an emotion recognition model provided in an embodiment of the present application; Optionally, in step S402, based on the alignment loss result and the classification loss result between the two modalities, iteratively correcting the network parameters of the initial emotion recognition model may include:
[0113] S501. Determine the total loss result of the current initial emotion recognition model according to the alignment loss result between the image emotion features and the text emotion features, the alignment loss result between the image emotion features and the speech emotion features, the alignment loss result between the text emotion features and the speech emotion features, and the classification loss result in the current iteration.
[0114] In any round of iterative training, the total loss result of the current initial emotion recognition model in the current iteration can be calculated based on the alignment loss result between the two modes and the classification loss result calculated in the current iteration.
[0115] The total loss result can be calculated using the following formula:
[0116] L total =(1-λ)L CE +λ(L IT +L IA +L TA )
[0117] Among them, λ is the weight adjustment coefficient of the loss function, L total Represents the total loss result, L CE Represents the classification loss result, L IT represents the alignment loss result between image emotion features and text emotion features, L IA represents the alignment loss result between image emotion features and speech emotion features, L TA Represents the alignment loss result between text emotion features and speech emotion features.
[0118] S502: According to the total loss result of the current initial emotion recognition model, the network parameters of the current initial emotion recognition model are modified.
[0119] The total loss result can be used to determine whether the total loss result meets the preset accuracy. If not, the network parameters of the current initial emotion recognition model are modified. When the total loss result meets the preset accuracy, the model training is stopped to obtain the initial emotion recognition model.
[0120] Figure 6 A flow chart of another method for training an emotion recognition model provided in an embodiment of the present application; the initial emotion recognition model also includes: a time series aggregation unit and a feature deepening unit.
[0121] Optionally, in the above steps, performing time series aggregation processing, dimension enhancement processing and nonlinear transformation processing on the emotional features of each modality in each group of sample data to obtain the target emotional features of each modality in each group of sample data may include:
[0122] S601. Input the emotion features of each modality in the current group of sample data into a time series aggregation unit. Through the time series aggregation unit, perform time series aggregation on the emotion features of each modality in the sample data to obtain aggregated emotion features of each modality, wherein the current group of sample data is any group of sample data among multiple groups of sample data.
[0123] In this embodiment, the time series aggregation unit can use a global pooling layer along the time dimension. Time series aggregation is used to aggregate or integrate data arranged in chronological order. It is usually used to process time series data. The main purpose of time series aggregation is to aggregate or integrate time series data in time for further data analysis or visualization. It can be used to calculate statistical indicators such as the average, sum, maximum, minimum, etc. of time series data, or to perform data smoothing, anomaly detection, etc.
[0124] There are many methods for time series aggregation, including simple average, moving average, weighted average, cumulative sum, cumulative percentage, etc. The aggregation method you choose depends on the characteristics of the data and the analysis requirements.
[0125] In this embodiment, a time series aggregation unit can be used to perform time series aggregation on the initial emotion features of each input single modality. For example, the initial emotion feature vectors of the single modality at each moment are summed up, or weighted averaged, etc., to obtain a complete feature vector of the single modality. For example: the initial emotion features of the single modality include 3*3 feature vectors at 10 moments, then after time series aggregation processing, the 10 3*3 feature vectors can be aggregated into 1 3*3 feature vector.
[0126] S602: Input the aggregated emotional features of each modality into a feature deepening unit. Through the feature deepening unit, convolution processing is performed on the aggregated emotional features of each modality to obtain the dimensionality-enhanced emotional features of each modality.
[0127] On the basis of the time series aggregation unit, a feature deepening unit may be further provided to perform feature deepening. In this embodiment, an encoder may be used to perform feature deepening to enhance the dimension of the emotion feature after aggregation.
[0128] The aggregated emotional features of each modality can be input into the encoder of the corresponding modality for dimensionality enhancement. Three fully connected layers can be used as encoders to enhance the dimension of features. Through feature deepening, low-dimensional aggregated emotional feature vectors can be converted into higher-dimensional emotional feature vectors. That is, the dimension of the feature is refined, and the feature is represented from more dimensions to enhance the expressiveness of the feature.
[0129] S603: Use a nonlinear activation function to perform data conversion on the emotional features of each modality after dimension enhancement to obtain the target emotional features of each modality.
[0130] In this embodiment, a nonlinear activation function may be introduced as a mapping head to map the emotional features of each modality after dimensionality enhancement to the same space, so as to improve the nonlinear fitting ability of the model.
[0131] Nonlinear activation functions play a very important role in neural networks. They are used to increase the nonlinearity of neural network models, enabling the network to better learn and represent complex input-output relationships.
[0132] Some common non-linear activation functions include:
[0133] Sigmoid function: maps input to (0, 1) and is often used to predict output probability. Its advantage is that it is differentiable and has a range between 0 and 1, which standardizes the output of neurons. However, its output is not centered on zero, which means that the weights can only be updated in one direction, thus affecting the convergence speed.
[0134] Tanh function: maps the input to between (-1,1). Similar to the Sigmoid function, the Tanh function is also differentiable and has a limited range, but its output is centered on zero.
[0135] ReLU function: The input is divided into two segments for mapping. When the input value is less than 0, the original value is mapped to 0. If the input value is greater than 0, it is passed as the original value. The advantage of the ReLU function is that it has a fast calculation speed and does not lose a large amount of features during forward calculation. However, during reverse calculation, if the input is a negative number, the gradient is 0, which may cause the problem of "dead neurons" during training.
[0136] Optionally, the initial emotion recognition model also includes: an emotion classification unit; in step S202, based on the multimodal fusion features corresponding to the current group of sample data, an emotion prediction result corresponding to the current group of sample data is identified, including: inputting the multimodal fusion features corresponding to the current group of sample data into the emotion classification unit, and classifying the multimodal fusion features through the emotion classification unit to obtain the emotion prediction result corresponding to the current group of sample data.
[0137] In some embodiments, the initial emotion recognition model may perform feature analysis on the multimodal fusion features corresponding to the current group of sample data to predict emotion prediction results corresponding to the multimodal fusion features based on the features, wherein one emotion prediction result is obtained for each group of sample data.
[0138] In the above embodiment, the classification loss result can be calculated based on the emotion prediction result corresponding to each sample data and the emotion label corresponding to each sample data.
[0139] Figure 7 A schematic diagram of a recognition method flow of an emotion recognition model provided in an embodiment of the present application; optionally, Figure 8 A schematic diagram of the architecture of an emotion recognition model provided in an embodiment of the present application; Figure 8As shown, the emotion recognition model may include a time series aggregation unit, a feature deepening unit and an emotion classification unit, and a pre-training unit may also be provided at the front end of the time series aggregation unit.
[0140] like Figure 7 As shown, the method may also include:
[0141] S701, obtaining emotion features to be recognized in multiple modalities.
[0142] The acquired emotion data to be recognized of each modality can be input into the pre-training unit corresponding to the modality, so as to extract the emotion features to be recognized of each modality respectively.
[0143] S702: Input the emotion features to be recognized in multiple modalities into the emotion recognition model.
[0144] The emotion features to be recognized of each modality are input into the emotion recognition model as input data, and the emotion recognition model performs emotion recognition.
[0145] S703 . Perform feature processing on the emotion features to be recognized of each modality through the emotion recognition model to obtain the target emotion features to be recognized of each modality.
[0146] Optionally, the emotion features to be identified of each modality are first input into a time series aggregation unit for time series aggregation processing to obtain aggregated emotion features of each modality, and further, the aggregated emotion features of each modality are input into a feature deepening unit to obtain dimensionally enhanced emotion features of each modality. The aggregated emotion features of each modality are input into a feature deepening unit corresponding to the modality for dimension enhancement processing.
[0147] In addition, the output end of the feature deepening unit can also be connected to a nonlinear activation function to perform nonlinear transformation processing on the emotional features of each modality after dimensionality enhancement, so as to obtain the target emotional features to be identified of each modality.
[0148] S704: performing feature concatenation on target emotion features to be identified in each modality to obtain multi-modal fusion features to be identified.
[0149] Optionally, the emotion recognition model may also include a feature fusion unit, which is connected to the back end of the nonlinear conversion unit to perform feature splicing on the target emotion features to be recognized obtained after the nonlinear conversion processing of each modality to obtain multimodal fusion features to be recognized.
[0150] S705 , classify and identify the multimodal fusion features to be identified, and obtain emotion recognition results corresponding to the emotion features to be identified in multiple modalities.
[0151] By taking the multimodal fusion features to be identified as the input of the emotion classification unit, classification processing can be performed. By performing feature analysis on the multimodal fusion features to be identified, emotion recognition results corresponding to the emotion features to be identified in multiple modalities can be obtained.
[0152] In summary, the training method of the emotion recognition model provided in this embodiment can be used to train a model for emotion recognition based on multimodal emotion features by collecting multiple groups of emotion features containing multiple modalities as sample data. Among them, by performing feature processing on the emotion features of each modality, the expressiveness of the emotion features can be improved. Then, by aligning the multimodal emotion features, the heterogeneity between different modalities can be eliminated and the interaction efficiency during the fusion of multimodal features can be improved. The network parameters of the emotion recognition model are optimized based on the alignment loss result constructed based on the multimodal alignment processing, so that the trained emotion recognition model can better fuse the multimodal emotion features, make full use of the multimodal emotion features for comprehensive analysis and recognition, and improve the accuracy of the emotion recognition results of the trained emotion recognition model.
[0153] The following describes the apparatus, equipment, storage medium, etc. used to execute the training method of the emotion recognition model provided in this application. The specific implementation process and technical effects are described above and will not be repeated below.
[0154] Fig. 9 A schematic diagram of a training device for an emotion recognition model provided in an embodiment of the present application, wherein the functions implemented by the training device for an emotion recognition model correspond to the steps performed by the above method. The device can be understood as the above server, or a processor of the server, or can be understood as a component independent of the above server or processor that implements the functions of the present application under the control of the server, such as Fig. 9 As shown, the device may include: a collection module 910, a processing module 920, and a correction module 930;
[0155] The acquisition module 910 is used to obtain multiple groups of sample data, each group of sample data includes: emotion features and emotion labels of multiple modalities;
[0156] The processing module 920 is used to input the multiple groups of sample data into the initial emotion recognition model, perform feature processing on the emotion features of each mode in each group of sample data, and obtain the target emotion features of each mode in each group of sample data;
[0157] The processing module 920 is used to perform feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identify the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data;
[0158] The correction module 930 is used to determine the alignment loss results between the two modalities according to the target emotion characteristics of each modality in each group of sample data, and iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities to obtain the emotion recognition model.
[0159] Optionally, the processing module 920 is specifically used to perform time series aggregation processing, dimension enhancement processing and nonlinear transformation processing on the emotional features of each modality in each group of sample data to obtain the target emotional features of each modality in each group of sample data.
[0160] Optionally, the initial emotion recognition model includes: a multimodal alignment module; a processing module 920, which is specifically used to input the target emotion features of each modality in the current group of sample data into the multimodal alignment module, and through the multimodal alignment module, convert the target emotion features of each modality from the feature space corresponding to each modality to a preset feature space to obtain the aligned emotion features of each modality corresponding to the current group of sample data.
[0161] The processing module 920 is specifically used to perform feature splicing on the aligned emotion features of each modality corresponding to the current group of sample data to obtain a multimodal fusion feature corresponding to the current group of sample data;
[0162] According to the multimodal fusion features corresponding to the current group of sample data, the emotion prediction results corresponding to the current group of sample data are identified.
[0163] Optionally, the target emotion features of each modality include: image target emotion features, text target emotion features and speech target emotion features; the correction module 930 is specifically used to determine the alignment loss result between the image emotion features and the text emotion features according to the image target emotion features and the text target emotion features in each group of sample data;
[0164] According to the image target emotion features and the speech target emotion features in each group of sample data, determining the alignment loss result between the image emotion features and the speech emotion features;
[0165] According to the text target emotion features and the speech target emotion features in each group of sample data, the alignment loss results between the text emotion features and the speech emotion features are determined.
[0166] Optionally, it also includes: a determination module;
[0167] A determination module, used to determine the classification loss result according to the emotion prediction result corresponding to each group of sample data and the emotion label in each group of sample data;
[0168] The correction module 930 is specifically used to iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results and classification loss results between the two modalities.
[0169] Optionally, a correction module 930 is specifically used to determine the total loss result of the current initial emotion recognition model according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the voice emotion feature, the alignment loss result between the text emotion feature and the voice emotion feature, and the classification loss result in the current iteration;
[0170] According to the total loss result of the current initial emotion recognition model, the network parameters of the current initial emotion recognition model are corrected.
[0171] Optionally, the correction module 930 is specifically used to determine the total alignment loss result in the current iteration according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, and the alignment loss result between the text emotion feature and the speech emotion feature in the current iteration;
[0172] The total loss result of the current initial emotion recognition model is determined according to the total alignment loss result, the classification loss result, the weight coefficient corresponding to the total alignment loss result, and the weight coefficient corresponding to the classification loss result in the current iteration.
[0173] Optionally, the initial emotion recognition model further includes: a time series aggregation unit and a feature deepening unit;
[0174] The processing module 920 is specifically used to input the emotion features of each modality in the current group of sample data into the time series aggregation unit, and perform time series aggregation on the emotion features of each modality in the sample data through the time series aggregation unit to obtain aggregated emotion features of each modality, wherein the current group of sample data is any group of sample data among the multiple groups of sample data;
[0175] The aggregated emotional features of each modality are input into the feature deepening unit, and the aggregated emotional features of each modality are convolved through the feature deepening unit to obtain the dimension-enhanced emotional features of each modality;
[0176] The nonlinear activation function is used to perform data conversion on the emotional features of each modality after dimension enhancement to obtain the target emotional features of each modality.
[0177] Optionally, the initial emotion recognition model further includes: an emotion classification unit;
[0178] The processing module 920 is specifically used to input the multimodal fusion features corresponding to the current group of sample data into the emotion classification unit, and classify the multimodal fusion features through the emotion classification unit to obtain the emotion prediction results corresponding to the current group of sample data.
[0179] Optionally, it also includes: an identification module;
[0180] Recognition module, used to obtain the emotion features to be recognized in multiple modes;
[0181] Input the emotion features to be recognized in multiple modalities into the emotion recognition model;
[0182] Through the emotion recognition model, feature processing is performed on the emotion features to be recognized in each modality to obtain the target emotion features to be recognized in each modality;
[0183] Perform feature splicing on the target emotion features to be identified in each modality to obtain the multi-modal fusion features to be identified;
[0184] The multimodal fusion features to be identified are classified and identified to obtain the emotion recognition results corresponding to the emotion features to be identified in multiple modalities.
[0185] The above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more digital singnal processors (DSP), or one or more field programmable gate arrays (FPGA). For another example, when a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0186] The above modules can be connected or communicated with each other via a wired connection or a wireless connection. The wired connection may include a metal cable, an optical cable, a hybrid cable, etc., or any combination thereof. The wireless connection may include a connection in the form of a LAN, a WAN, Bluetooth, a ZigBee, or NFC, or any combination thereof. Two or more modules may be combined into a single module, and any one module may be divided into two or more units. Those skilled in the art will clearly understand that for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application.
[0187] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application includes: a processor 801, a storage medium 802 and a bus 803, wherein the storage medium 802 stores machine-readable instructions executable by the processor 801. When the electronic device runs a training method for an emotion recognition model in an embodiment, the processor 801 communicates with the storage medium 802 via the bus 803, and the processor 801 executes the machine-readable instructions to perform the following steps:
[0188] Acquire multiple sets of sample data, each set of sample data includes: emotion features of multiple modalities and emotion labels;
[0189] Inputting multiple groups of sample data into the initial emotion recognition model, performing feature processing on the emotion features of each mode in each group of sample data, and obtaining the target emotion features of each mode in each group of sample data;
[0190] Performing feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identifying the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data;
[0191] According to the target emotion features of each modality in each group of sample data, the alignment loss results between the two modalities are determined, and based on the alignment loss results between the two modalities, the network parameters of the initial emotion recognition model are iteratively corrected to obtain the emotion recognition model.
[0192] In a feasible implementation scheme, when the processor 801 performs feature processing on the emotional features of each modality in each group of sample data to obtain the target emotional features of each modality in each group of sample data, it is specifically used to: perform time series aggregation processing, dimensionality enhancement processing and nonlinear transformation processing on the emotional features of each modality in each group of sample data to obtain the target emotional features of each modality in each group of sample data.
[0193] In a feasible implementation scheme, the initial emotion recognition model includes: a multimodal alignment module; when the processor 801 performs feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, it is specifically used to: input the target emotion features of each modality in the current group of sample data into the multimodal alignment module, and through the multimodal alignment module, convert the target emotion features of each modality from the feature space corresponding to each modality to a preset feature space to obtain the aligned emotion features of each modality corresponding to the current group of sample data.
[0194] In a feasible implementation scheme, when the processor 801 executes the emotion features of each modality after alignment corresponding to each group of sample data to identify the emotion prediction results corresponding to each group of sample data, it is specifically used to: perform feature splicing on the emotion features of each modality after alignment corresponding to the current group of sample data to obtain the multimodal fusion features corresponding to the current group of sample data; and identify the emotion prediction results corresponding to the current group of sample data based on the multimodal fusion features corresponding to the current group of sample data.
[0195] In a feasible implementation scheme, the target emotion features of each modality include: image target emotion features, text target emotion features and speech target emotion features; when the processor 801 determines the alignment loss results between the modalities according to the target emotion features of each modality in each group of sample data, it is specifically used to: determine the alignment loss results between the image emotion features and the text emotion features according to the image target emotion features and the text target emotion features in each group of sample data; determine the alignment loss results between the image emotion features and the speech emotion features according to the image target emotion features and the speech target emotion features in each group of sample data; determine the alignment loss results between the text emotion features and the speech emotion features according to the text target emotion features and the speech target emotion features in each group of sample data.
[0196] In a feasible implementation scheme, the processor 801 is also used to: determine the classification loss result based on the emotion prediction results corresponding to each group of sample data and the emotion labels in each group of sample data; when the processor 801 performs iterative correction on the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities, it is specifically used to iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities and the classification loss results.
[0197] In a feasible implementation scheme, when the processor 801 performs iterative correction of the network parameters of the initial emotion recognition model based on the alignment loss results and classification loss results between the two modalities, it is specifically used to: determine the total loss result of the current initial emotion recognition model according to the alignment loss results between the image emotion features and the text emotion features, the alignment loss results between the image emotion features and the speech emotion features, the alignment loss results between the text emotion features and the speech emotion features, and the classification loss results in the current iteration; and correct the network parameters of the current initial emotion recognition model according to the total loss result of the current initial emotion recognition model.
[0198] In a feasible implementation scheme, when the processor 801 determines the total loss result of the current initial emotion recognition model according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, the alignment loss result between the text emotion feature and the speech emotion feature, and the classification loss result in the current iteration, the processor 801 is specifically used to: determine the total alignment loss result in the current iteration according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, and the alignment loss result between the text emotion feature and the speech emotion feature in the current iteration;
[0199] The total loss result of the current initial emotion recognition model is determined according to the total alignment loss result, the classification loss result, the weight coefficient corresponding to the total alignment loss result, and the weight coefficient corresponding to the classification loss result in the current iteration.
[0200] In a feasible implementation scheme, the initial emotion recognition model also includes: a time series aggregation unit and a feature deepening unit; when the processor 801 performs time series aggregation processing, dimensionality enhancement processing and nonlinear conversion processing on the emotion features of each modality in each group of sample data to obtain the target emotion features of each modality in each group of sample data, it is specifically used to: input the emotion features of each modality in the current group of sample data into the time series aggregation unit, and through the time series aggregation unit, perform time series aggregation on the emotion features of each modality in the sample data to obtain the aggregated emotion features of each modality, wherein the current group of sample data is any group of sample data among multiple groups of sample data; input the aggregated emotion features of each modality into the feature deepening unit, and through the feature deepening unit, perform convolution processing on the aggregated emotion features of each modality to obtain the dimensionality enhanced emotion features of each modality; use a nonlinear activation function to perform data conversion on the dimensionality enhanced emotion features of each modality to obtain the target emotion features of each modality.
[0201] In a feasible implementation scheme, the initial emotion recognition model also includes: an emotion classification unit; when the processor 801 executes the multimodal fusion features corresponding to the current group of sample data to identify the emotion prediction results corresponding to the current group of sample data, it is specifically used to input the multimodal fusion features corresponding to the current group of sample data into the emotion classification unit, and the multimodal fusion features are classified and processed by the emotion classification unit to obtain the emotion prediction results corresponding to the current group of sample data.
[0202] In a feasible implementation scheme, processor 801 is also used to obtain emotion features to be identified in multiple modalities; input the emotion features to be identified in multiple modalities into an emotion recognition model; perform feature processing on the emotion features to be identified in each modality through the emotion recognition model to obtain target emotion features to be identified in each modality; perform feature splicing on the target emotion features to be identified in each modality to obtain multimodal fusion features to be identified; classify and identify the multimodal fusion features to be identified to obtain emotion recognition results corresponding to the emotion features to be identified in multiple modalities.
[0203] In the above manner, the electronic device can be used to train a model for emotion recognition based on multimodal emotion features by collecting multiple groups of emotion features containing multiple modalities as sample data. Among them, by performing feature processing on the emotion features of each modality, the expression ability of the emotion features can be improved. Then, by aligning the multimodal emotion features, the heterogeneity between different modalities can be eliminated and the interaction efficiency during the fusion of multimodal features can be improved. The network parameters of the emotion recognition model are optimized based on the alignment loss result constructed based on the multimodal alignment processing, so that the trained emotion recognition model can better fuse the multimodal emotion features, make full use of the multimodal emotion features for comprehensive analysis and recognition, and improve the accuracy of the emotion recognition results of the trained emotion recognition model.
[0204] Among them, the storage medium 802 stores program code, and when the program code is executed by the processor 801, the processor 801 executes various steps in the training method of the emotion recognition model according to various exemplary embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0205] Processor 801 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor (DigitalSignal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware processor to be executed, or a combination of hardware and software modules in the processor can be executed.
[0206] Storage medium 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (Random Access Memory, RAM), static random access memory (Static Random Access Memory, SRAM), programmable read-only memory (Programmable Read Only Memory, PROM), read-only memory (Read Only Memory, ROM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), magnetic memory, disk, optical disk, etc. The memory is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The storage medium 802 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0207] Optionally, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor performs the following steps:
[0208] Acquire multiple sets of sample data, each set of sample data includes: emotion features of multiple modalities and emotion labels;
[0209] Inputting multiple groups of sample data into the initial emotion recognition model, performing feature processing on the emotion features of each mode in each group of sample data, and obtaining the target emotion features of each mode in each group of sample data;
[0210] Performing feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identifying the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data;
[0211] According to the target emotion features of each modality in each group of sample data, the alignment loss results between the two modalities are determined, and based on the alignment loss results between the two modalities, the network parameters of the initial emotion recognition model are iteratively corrected to obtain the emotion recognition model.
[0212] In a feasible implementation scheme, when the processor 801 performs feature processing on the emotional features of each modality in each group of sample data to obtain the target emotional features of each modality in each group of sample data, it is specifically used to: perform time series aggregation processing, dimensionality enhancement processing and nonlinear transformation processing on the emotional features of each modality in each group of sample data to obtain the target emotional features of each modality in each group of sample data.
[0213] In a feasible implementation scheme, the initial emotion recognition model includes: a multimodal alignment module; when the processor 801 performs feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, it is specifically used to: input the target emotion features of each modality in the current group of sample data into the multimodal alignment module, and through the multimodal alignment module, convert the target emotion features of each modality from the feature space corresponding to each modality to a preset feature space to obtain the aligned emotion features of each modality corresponding to the current group of sample data.
[0214] In a feasible implementation scheme, when the processor 801 executes the emotion features of each modality after alignment corresponding to each group of sample data to identify the emotion prediction results corresponding to each group of sample data, it is specifically used to: perform feature splicing on the emotion features of each modality after alignment corresponding to the current group of sample data to obtain the multimodal fusion features corresponding to the current group of sample data; and identify the emotion prediction results corresponding to the current group of sample data based on the multimodal fusion features corresponding to the current group of sample data.
[0215] In a feasible implementation scheme, the target emotion features of each modality include: image target emotion features, text target emotion features and speech target emotion features; when the processor 801 determines the alignment loss results between the modalities according to the target emotion features of each modality in each group of sample data, it is specifically used to: determine the alignment loss results between the image emotion features and the text emotion features according to the image target emotion features and the text target emotion features in each group of sample data; determine the alignment loss results between the image emotion features and the speech emotion features according to the image target emotion features and the speech target emotion features in each group of sample data; determine the alignment loss results between the text emotion features and the speech emotion features according to the text target emotion features and the speech target emotion features in each group of sample data.
[0216] In a feasible implementation scheme, the processor 801 is also used to: determine the classification loss result based on the emotion prediction results corresponding to each group of sample data and the emotion labels in each group of sample data; when the processor 801 performs iterative correction on the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities, it is specifically used to iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities and the classification loss results.
[0217] In a feasible implementation scheme, when the processor 801 performs iterative correction of the network parameters of the initial emotion recognition model based on the alignment loss results and classification loss results between the two modalities, it is specifically used to: determine the total loss result of the current initial emotion recognition model according to the alignment loss results between the image emotion features and the text emotion features, the alignment loss results between the image emotion features and the speech emotion features, the alignment loss results between the text emotion features and the speech emotion features, and the classification loss results in the current iteration; and correct the network parameters of the current initial emotion recognition model according to the total loss result of the current initial emotion recognition model.
[0218] In a feasible implementation scheme, when the processor 801 determines the total loss result of the current initial emotion recognition model according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, the alignment loss result between the text emotion feature and the speech emotion feature, and the classification loss result in the current iteration, the processor 801 is specifically used to: determine the total alignment loss result in the current iteration according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, and the alignment loss result between the text emotion feature and the speech emotion feature in the current iteration;
[0219] The total loss result of the current initial emotion recognition model is determined according to the total alignment loss result, the classification loss result, the weight coefficient corresponding to the total alignment loss result, and the weight coefficient corresponding to the classification loss result in the current iteration.
[0220] In a feasible implementation scheme, the initial emotion recognition model also includes: a time series aggregation unit and a feature deepening unit; when the processor 801 performs time series aggregation processing, dimensionality enhancement processing and nonlinear conversion processing on the emotion features of each modality in each group of sample data to obtain the target emotion features of each modality in each group of sample data, it is specifically used to: input the emotion features of each modality in the current group of sample data into the time series aggregation unit, and through the time series aggregation unit, perform time series aggregation on the emotion features of each modality in the sample data to obtain the aggregated emotion features of each modality, wherein the current group of sample data is any group of sample data among multiple groups of sample data; input the aggregated emotion features of each modality into the feature deepening unit, and through the feature deepening unit, perform convolution processing on the aggregated emotion features of each modality to obtain the dimensionality enhanced emotion features of each modality; use a nonlinear activation function to perform data conversion on the dimensionality enhanced emotion features of each modality to obtain the target emotion features of each modality.
[0221] In a feasible implementation scheme, the initial emotion recognition model also includes: an emotion classification unit; when the processor 801 executes the multimodal fusion features corresponding to the current group of sample data to identify the emotion prediction results corresponding to the current group of sample data, it is specifically used to input the multimodal fusion features corresponding to the current group of sample data into the emotion classification unit, and the multimodal fusion features are classified and processed by the emotion classification unit to obtain the emotion prediction results corresponding to the current group of sample data.
[0222] In a feasible implementation scheme, processor 801 is also used to obtain emotion features to be identified in multiple modalities; input the emotion features to be identified in multiple modalities into an emotion recognition model; perform feature processing on the emotion features to be identified in each modality through the emotion recognition model to obtain target emotion features to be identified in each modality; perform feature splicing on the target emotion features to be identified in each modality to obtain multimodal fusion features to be identified; classify and identify the multimodal fusion features to be identified to obtain emotion recognition results corresponding to the emotion features to be identified in multiple modalities.
[0223] In the above manner, the electronic device can be used to train a model for emotion recognition based on multimodal emotion features by collecting multiple groups of emotion features containing multiple modalities as sample data. Among them, by performing feature processing on the emotion features of each modality, the expression ability of the emotion features can be improved. Then, by aligning the multimodal emotion features, the heterogeneity between different modalities can be eliminated and the interaction efficiency during the fusion of multimodal features can be improved. The network parameters of the emotion recognition model are optimized based on the alignment loss result constructed based on the multimodal alignment processing, so that the trained emotion recognition model can better fuse the multimodal emotion features, make full use of the multimodal emotion features for comprehensive analysis and recognition, and improve the accuracy of the emotion recognition results of the trained emotion recognition model.
[0224] In the embodiment of the present application, the computer program can also execute other machine-readable instructions when run by the processor to execute other methods described in the embodiment. For the specific execution method steps and principles, please refer to the description of the embodiment, which will not be repeated here.
[0225] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0226] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0227] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0228] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (English: Read-Only Memory, abbreviated: ROM), random access memory (English: Random Access Memory, abbreviated: RAM), disk or optical disk and other media that can store program codes.
Claims
1. A training method for an emotion recognition model, It is characterized in that include: Acquire multiple sets of sample data, each set of sample data includes: emotion features of multiple modalities and emotion labels; Inputting multiple groups of sample data into the initial emotion recognition model, performing feature processing on the emotion features of each mode in each group of sample data, and obtaining the target emotion features of each mode in each group of sample data; Performing feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identifying the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data; According to the target emotion features of each modality in each group of sample data, the alignment loss results between the two modalities are determined, and based on the alignment loss results between the two modalities, the network parameters of the initial emotion recognition model are iteratively corrected to obtain the emotion recognition model.
2. The method according to claim 1, It is characterized in that The feature processing of the emotion features of each modality in each group of sample data to obtain the target emotion features of each modality in each group of sample data includes: The emotional features of each mode in each group of sample data are subjected to time series aggregation processing, dimension enhancement processing and nonlinear transformation processing to obtain the target emotional features of each mode in each group of sample data.
3. The method according to claim 1, It is characterized in that The initial emotion recognition model includes: a multimodal alignment module; the feature space alignment processing is performed on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, including: The target emotion features of each modality in the current group of sample data are input into the multimodal alignment module. Through the multimodal alignment module, the target emotion features of each modality are converted from the feature space corresponding to each modality to a preset feature space to obtain the aligned emotion features of each modality corresponding to the current group of sample data.
4. The method according to claim 1, It is characterized in that The step of identifying the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data includes: Perform feature splicing on the aligned emotion features of each modality corresponding to the current group of sample data to obtain a multimodal fusion feature corresponding to the current group of sample data; According to the multimodal fusion features corresponding to the current group of sample data, the emotion prediction results corresponding to the current group of sample data are identified.
5. The method according to claim 1, It is characterized in that The target emotion features of each modality include: image target emotion features, text target emotion features, and speech target emotion features; the alignment loss results between the two modalities are determined according to the target emotion features of each modality in each group of sample data, including: According to the image target emotion features and the text target emotion features in each group of sample data, the alignment loss result between the image emotion features and the text emotion features is determined; According to the image target emotion features and the speech target emotion features in each group of sample data, determining the alignment loss result between the image emotion features and the speech emotion features; According to the text target emotion features and the speech target emotion features in each group of sample data, the alignment loss results between the text emotion features and the speech emotion features are determined.
6. The method according to claim 1, It is characterized in that Also includes: Determine the classification loss result based on the emotion prediction results corresponding to each group of sample data and the emotion labels in each group of sample data; The iterative correction of the network parameters of the initial emotion recognition model based on the alignment loss results between the two modalities includes: Based on the alignment loss results and classification loss results between the two modalities, the network parameters of the initial emotion recognition model are iteratively corrected.
7. The method according to claim 6, It is characterized in that The iterative correction of the network parameters of the initial emotion recognition model based on the alignment loss results and the classification loss results between the two modalities includes: Determine the total loss result of the current initial emotion recognition model according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, the alignment loss result between the text emotion feature and the speech emotion feature, and the classification loss result in the current iteration; According to the total loss result of the current initial emotion recognition model, the network parameters of the current initial emotion recognition model are corrected.
8. The method according to claim 7, It is characterized in that The method determines the total loss result of the current initial emotion recognition model according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the voice emotion feature, the alignment loss result between the text emotion feature and the voice emotion feature, and the classification loss result in the current iteration, including: Determine the total alignment loss result in the current iteration according to the alignment loss result between the image emotion feature and the text emotion feature, the alignment loss result between the image emotion feature and the speech emotion feature, and the alignment loss result between the text emotion feature and the speech emotion feature in the current iteration; The total loss result of the current initial emotion recognition model is determined according to the total alignment loss result, the classification loss result, the weight coefficient corresponding to the total alignment loss result, and the weight coefficient corresponding to the classification loss result in the current iteration.
9. The method according to claim 2, It is characterized in that The initial emotion recognition model further includes: a time series aggregation unit and a feature deepening unit; The emotional features of each modality in each group of sample data are subjected to time series aggregation processing, dimension enhancement processing, and nonlinear transformation processing to obtain the target emotional features of each modality in each group of sample data, including: Inputting the emotion features of each modality in the current group of sample data into the time series aggregation unit, and performing time series aggregation on the emotion features of each modality in the sample data through the time series aggregation unit to obtain aggregated emotion features of each modality, wherein the current group of sample data is any one group of sample data among the multiple groups of sample data; Inputting the aggregated emotional features of each modality into the feature deepening unit, and performing convolution processing on the aggregated emotional features of each modality through the feature deepening unit to obtain the dimension-enhanced emotional features of each modality; The nonlinear activation function is used to perform data conversion on the emotional features of each modality after dimension enhancement to obtain the target emotional features of each modality.
10. The method according to claim 4, It is characterized in that The initial emotion recognition model further includes: an emotion classification unit; The step of identifying the emotion prediction result corresponding to the current group of sample data according to the multimodal fusion features corresponding to the current group of sample data includes: The multimodal fusion features corresponding to the current group of sample data are input into the emotion classification unit, and the multimodal fusion features are classified by the emotion classification unit to obtain the emotion prediction results corresponding to the current group of sample data.
11. The method according to claim 1, It is characterized in that Also includes: Obtain the emotion features to be identified in multiple modalities; Inputting the emotion features to be recognized in multiple modalities into the emotion recognition model; Through the emotion recognition model, feature processing is performed on the emotion features to be recognized in each modality to obtain the target emotion features to be recognized in each modality; Perform feature splicing on the target emotion features to be identified in each modality to obtain the multi-modal fusion features to be identified; The multimodal fusion features to be identified are classified and identified to obtain emotion recognition results corresponding to the emotion features to be identified in the multiple modalities.
12. A training device for an emotion recognition model, It is characterized in that include: Acquisition module, processing module, correction module; The acquisition module is used to obtain multiple groups of sample data, each group of sample data includes: emotional features of multiple modalities and emotional labels; The processing module is used to input multiple groups of sample data into the initial emotion recognition model, perform feature processing on the emotion features of each modality in each group of sample data, and obtain the target emotion features of each modality in each group of sample data; The processing module is used to perform feature space alignment processing on the target emotion features of each modality in each group of sample data to obtain the aligned emotion features of each modality corresponding to each group of sample data, and identify the emotion prediction results corresponding to each group of sample data based on the aligned emotion features of each modality corresponding to each group of sample data; The correction module is used to determine the alignment loss results between two modalities according to the target emotion characteristics of each modality in each group of sample data, and iteratively correct the network parameters of the initial emotion recognition model based on the alignment loss results between two modalities to obtain the emotion recognition model.
13. An electronic device, It is characterized in that include: A processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor, and when the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the program instructions to execute the training method of the emotion recognition model as described in any one of claims 1 to 11.
14. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the training method of the emotion recognition model as described in any one of claims 1 to 11 is executed.