Multimodal-based emotion recognition methods and related equipment
Through the multimodal network model, the problem of low emotional recognition accuracy is solved and higher accuracy emotion recognition is achieved.
Patent Information
- Application Number
- CN202211500520.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-11-28
AI Technical Summary
In the prior art, the accuracy of emotion recognition is poor and it is difficult to effectively improve.
A multimodal network model is adopted, including speech recognition model, electrocardiogram recognition model and natural language recognition model. By obtaining speech data, electrocardiogram data and natural language data, feature extraction and modal fusion are performed to identify emotions.
Through multimodal fusion, the accuracy of emotion recognition is improved and the emotional type can be identified more accurately.
Smart Images

Figure CN115775565B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing technology, and in particular to a multimodal emotion recognition method and related equipment. Background Art
[0002] With the rapid development of science and technology, more and more IT scholars are applying increasingly advanced computer and network communication technologies to personnel information analysis. In practical applications, emotion recognition has also become a research hotspot. However, at present, the accuracy of emotion analysis results is relatively poor. Therefore, the problem of how to improve the accuracy of emotion recognition needs to be solved urgently. Summary of the Invention
[0003] The embodiments of the present application provide a multimodal emotion recognition method and related equipment, which can improve the accuracy of emotion recognition.
[0004] In a first aspect, an embodiment of the present application provides a multimodal emotion recognition method, which is applied to an electronic device. The electronic device is configured with a multimodal network model, and the multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model, and a modal fusion model. The method includes:
[0005] Acquire multimodal data of a target object, wherein the multimodal data includes: speech data, electrocardiogram data, and natural language data;
[0006] Inputting the speech data into the speech recognition model to obtain speech features;
[0007] Inputting the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features;
[0008] Inputting the natural language data into the natural language recognition model to obtain text features;
[0009] The speech features, the electrocardiogram features and the text features are input into the modality fusion model to obtain an emotion recognition result.
[0010] In a second aspect, an embodiment of the present application provides a multimodal emotion recognition device, which is applied to an electronic device. The electronic device is configured with a multimodal network model, and the multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model. The device includes: an acquisition unit, an extraction unit and a recognition unit, wherein:
[0011] The acquisition unit is configured to acquire multimodal data of the target object, wherein the multimodal data includes: voice data, electrocardiogram data, and natural language data;
[0012] The extraction unit is configured to input the speech data into the speech recognition model to obtain speech features; input the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features; and input the natural language data into the natural language recognition model to obtain text features;
[0013] The recognition unit is used to input the speech features, the electrocardiogram features and the text features into the modal fusion model to obtain an emotion recognition result.
[0014] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program comprises instructions for executing the steps in the first aspect of the embodiment of the present application.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the above-mentioned computer-readable storage medium stores a computer program for electronic data exchange, wherein the above-mentioned computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.
[0016] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0017] The implementation of the embodiments of this application has the following beneficial effects:
[0018] It can be seen that the multimodal emotion recognition method and related equipment described in the embodiments of the present application are applied to electronic devices, and the electronic devices are configured with a multimodal network model. The multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model to obtain multimodal data of the target object. The multimodal data includes: speech data, electrocardiogram data, and natural language data. The speech data is input into the speech recognition model to obtain speech features. The electrocardiogram data is input into the electrocardiogram recognition model to obtain electrocardiogram features. The natural language data is input into the natural language recognition model to obtain text features. The speech features, electrocardiogram features and text features are input into the modal fusion model to obtain emotion recognition results. The three-dimensional features of speech, electrocardiogram and text of the same object are modally fused, and the corresponding emotions are identified, which can improve the accuracy of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1A This is a flowchart of a multimodal emotion recognition method provided in an embodiment of the present application;
[0021] Figure 1B This is the structural intention of a multimodal network model provided in an embodiment of the present application;
[0022] Figure 1C This is a flowchart of another multimodal emotion recognition method provided in an embodiment of the present application;
[0023] Figure 2 This is a flowchart of another multimodal emotion recognition method provided in an embodiment of the present application;
[0024] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0025] Figure 4 This is a block diagram of the functional units of a multimodal emotion recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0028] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0029] The electronic devices described in the embodiments of the present application may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, PDAs, driving recorders, servers, laptops, mobile Internet devices (MIDs) or wearable devices (such as smart watches, Bluetooth headsets), etc. The above are only examples and not exhaustive, including but not limited to the above electronic devices.
[0030] In the embodiments of the present application, an anchor example can be understood as a reference example used to construct positive and negative samples. A positive example can be understood as a sample belonging to the same class as the anchor example. A negative example can be understood as a sample that does not belong to the same class as the anchor example. A positive sample pair (anchor-positive) can be understood as a sample pair consisting of an anchor sample and a positive sample. A negative sample pair (anchor-negative) can be understood as a sample pair consisting of an anchor sample and a negative sample.
[0031] The following is a detailed introduction to the embodiments of the present application.
[0032] See also Figure 1A , Figure 1A This is a flow chart of a multimodal emotion recognition method provided by an embodiment of the present application. As shown in the figure, it is applied to an electronic device. The electronic device is configured with a multimodal network model. The multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model. The multimodal emotion recognition method includes:
[0033] 101. Acquire multimodal data of a target object, where the multimodal data includes: voice data, electrocardiogram data, and natural language data.
[0034] In an embodiment of the present application, the target object may be a person, and a multimodal signal of the target object may be obtained. The multimodal data may include: voice data, electrocardiogram data, and natural language data, wherein the voice data, electrocardiogram data, and natural language data may be data of the same time period, and the natural language data may be text data corresponding to the voice data.
[0035] In specific implementation, such as Figure 1B As shown, the multimodal network model may include: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modality fusion model.
[0036] In a specific implementation, voice data can be collected through a microphone, and electrocardiogram data can be collected through a wearable device. The voice data can be converted into natural language data.
[0037] 102. Input the speech data into the speech recognition model to obtain speech features.
[0038] In the embodiment of the present application, the speech recognition model can be pre-set or system default. The speech recognition model can include at least one of the following: a convolutional neural network model (CNN), a fully connected neural network model, a recurrent neural network model, a Wav2vec2.0 network model, etc., which are not limited here.
[0039] In the specific implementation, speech features can be extracted through Wav2vec2.0 and CNN network models, and speech data can be input into Wav2vec2.0 and CNN network models to obtain speech features.
[0040] 103. Input the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features.
[0041] In an embodiment of the present application, the electrocardiogram recognition model can be pre-set or set by the system default. The electrocardiogram recognition model can be used to extract features from the electrocardiogram data to obtain electrocardiogram features. The electrocardiogram recognition model can include at least one of the following: a convolutional neural network model, a fully connected neural network model, a recurrent neural network model, a 2DRCN module, etc., which are not limited here. In a specific implementation, the sift feature can be used to obtain a spectrogram for the electrocardiogram data, and then the visual representation of the spatial domain can be learned using 2DCNN, and the LSTM can be applied to model the spectrogram of the electrocardiogram to obtain the electrocardiogram features.
[0042] 104. Input the natural language data into the natural language recognition model to obtain text features.
[0043] In an embodiment of the present application, a natural language recognition model can be pre-set or system default. The natural language recognition model can be used to extract features from natural language data to obtain text features. The natural language recognition model may include at least one of the following: a convolutional neural network model, a fully connected neural network model, a recurrent neural network model, a natural language processing model (Natural Language Processing, NLP), etc., which are not limited here. In a specific implementation, natural language data can be input into a natural language recognition model to obtain text features. For example, NLP can be used to embed text to obtain a 128-dimensional vector. Of course, the feature dimension after each of the above single-modal embeddings can be 128.
[0044] 105. Input the speech features, the electrocardiogram features, and the text features into the modality fusion model to obtain an emotion recognition result.
[0045] In the embodiment of the present application, the emotion recognition result is used to describe the emotion type, and the emotion recognition result may include at least one of the following: happiness, sadness, depression, anger, worry, fear, shock, calmness, etc., which are not limited here.
[0046] In specific implementation, such as Figure 1C As shown, the corresponding speech features, electrocardiogram features and text features can be obtained by inputting speech data, electrocardiogram data and natural language data into the corresponding model. The speech features, electrocardiogram features and text features can be input into the modal fusion model to obtain emotion recognition results.
[0047] In the embodiment of the present application, it can be used to identify the emotional activities of a person.
[0048] Optionally, the above step 105, inputting the speech features, the electrocardiogram features, and the text features into the modality fusion model to obtain the emotion recognition result, may include the following steps:
[0049] 51. Divide the speech feature into a equal parts to obtain a speech feature parts, where a is an integer greater than 1;
[0050] 52. Divide the electrocardiogram feature into a equal parts to obtain a parts of electrocardiogram features;
[0051] 53. Divide the natural language data into a equal parts, and obtain a parts of text features;
[0052] 54. Mix the a portions of speech features, the a portions of electrocardiogram features, and the a portions of text features to obtain a portions of mixed features, each portion of the mixed features including a portion of speech features, a portion of electrocardiogram features, and a portion of text features;
[0053] 55. Determine the proportion of various modal uncertainties in each of the a mixed features as a weight, then determine corresponding weighted features based on the weight, and use the weighted features of various modes as candidate features to obtain a candidate features;
[0054] 56. Concatenate the a selected features to obtain a fused feature;
[0055] 57. Determine the emotion recognition result based on the fusion feature.
[0056] In an embodiment of the present application, the speech features are divided into a equal parts to obtain a speech features, where a is an integer greater than 1; the electrocardiogram features are divided into a equal parts to obtain a electrocardiogram features; the natural language data are divided into a equal parts to obtain a text features; the a speech features, a electrocardiogram features and a text features are mixed to obtain a mixed features, each mixed feature includes a speech feature, an electrocardiogram feature and a text feature; the proportion of various modal uncertainties in each of the a mixed features is determined as a weight, and then the corresponding weighted features are determined based on the weight; the weighted features of various modalities are used as candidate features to obtain a candidate features; the a candidate features are spliced to obtain a fused feature, so that the features of different modalities can be fused; and then the emotion recognition result is determined based on the fused feature. Specifically, the fused feature can be input into a neural network model to obtain the corresponding emotion recognition result.
[0057] In the specific implementation, the voice, text and heartbeat data can be mixed up to update the fused features and improve the fusion feature effect. For example, the x of the three modes can be mixed up. i The features are divided into four equal parts x i,1 , x i,2 , x i,3 , x i,4 Then, calculate the fusion features of each equal part and calculate the uncertainty d of each modal feature according to the uncertainty formula i,1 , according to the uncertainty, the proportion of the feature uncertainty of different modes is calculated as the weight, the weight and feature of each mode are weighted and summed to obtain the corresponding weighted feature, which is used as the candidate feature s1. Furthermore, each feature is spliced together to obtain the total feature s = concat(s1,s2,s3,s4).
[0058] Optionally, the above step 57, determining the emotion recognition result according to the fusion feature, can be implemented as follows:
[0059] The fusion features are sequentially input into a translation encoder, a translation decoder and a fully connected module to obtain the emotion recognition result.
[0060] In the embodiment of the present application, the fusion features can be sequentially input into the translation encoder, the translation decoder and the fully connected module to obtain the emotion recognition result, so that emotion recognition can be recognized.
[0061] Optionally, before step 101, the following steps may also be included:
[0062] Training the multimodal network model using a preset loss function to obtain a trained multimodal network model;
[0063] The preset loss function consists of an intra-modal contrastive learning loss function and a cross-modal contrastive learning loss function.
[0064] In the embodiment of the present application, the preset loss function can be pre-set or system default.
[0065] The widespread use of smart devices has made multimodal data fusion possible. Data selection strategies for multimodal models can improve model accuracy and stability. Furthermore, pairing positive and negative samples can enhance generalization performance and mitigate the negative impact of datasets. Therefore, designing effective sample mining strategies is crucial.
[0066] In the embodiment of the present application, during the training phase, data collection can be performed. For example, the management department can obtain continuous signal information such as signals, electrocardiogram signals, and recordings from the electronic bracelets of multiple community correction personnel, and then preprocess these data. Specifically, the voice signal in the collected data can be translated into text using a voice recognition model. After manual inspection and correction of the text information, this segment of voice, text, and corresponding electrocardiogram signal are combined into a piece of data, and then qualified data is screened out as the initial training set and test set. It is also possible to select positive and negative sample pairs from the training set as the training set based on the designed positive and negative sample pair selection mechanism, input the generated data into the multimodal model network, and use the model to learn the data features of each modality, fully explore cross-modal interactions, learn the relationship between samples and classes, and reduce the modal gap. The multimodal features are then fused together according to a certain strategy, and the model is made to learn emotion classification (happy, angry) through a supervised method.
[0067] In the embodiment of the present application, an n-pair joint loss of speech and heartbeat can be used to optimize model training. Specifically, the generated data can be input into a multimodal network model, and the model can be used to learn the data features of each modality, fully explore cross-modal interactions, learn the relationship between samples and classes, and reduce the modality gap. The multimodal features are then fused together according to a certain strategy, and through supervision methods, the model can learn emotion classification (happy, angry), etc.
[0068] In this embodiment, two contrastive losses operating on the encoded unimodal representations are provided for hybrid contrastive learning to perform intra-modal and inter-modal learning during the training phase. Through the designed losses, the model can fully understand the dynamics within and between modalities, explore inter-class relationships, and minimize modality gaps.
[0069] The design approach for contrastive learning can be divided into two parts: intra-modal contrastive learning and cross-modal contrastive learning. Intra-modal contrastive learning involves performing intra-modal learning in a supervised manner to learn the intra-modal dynamics between different samples, considering multiple positive and negative pairs in a mini-batch. Cross-modal contrastive learning also involves performing this approach in a supervised manner to learn cross-modal dynamics. Both approaches explore inter-class relationships.
[0070] In its implementation, IAMCL uses a supervised intra-modal contrastive learning method to learn intra-modal dynamics and inter-class relationships. Positive pairs are two unimodal representations of two different samples of the same modality and class; negative pairs are two unimodal representations of two samples of the same modality but different classes.
[0071] Specifically, we can use anchor sample a m Generate a batch of size K, the collection is as follows:
[0072] S={p1 m ,p2 m ,...,p N m ,n1 m ,n2 m ,...,n M m}
[0073] The above set can generate N positive examples and M negative examples from one anchor point (the anchor point is not included in N), so K is fixed, but N and M are random, that is, the number of positive and negative pairs is not fixed. The specific loss in the IAMCL mode is as follows:
[0074]
[0075]
[0076]
[0077] Among them, a m is the anchor point sample, p m is a positive sample of the same type as the anchor sample, n i is a negative sample of a different class from the anchor point. l represents text, a represents speech, and b represents ECG waveform. R To refine the loss, let am and p m The vectors of E are as similar as possible. s is the expectation of a small batch of data s. α is the modal margin between different modes. The final loss is: L IAMCL and
[0078] In the specific implementation, N_pair uses N-1 negative samples and one positive sample each time. In the embodiment of the present application, the inner product operation can be used to represent the distance between two vectors. Specifically, the cosine distance algorithm can be used. The larger the distance, the closer the two vectors are, and the smaller the distance, the farther they are.
[0079] In its implementation, when N = 2, it approximates triplet loss. The drawback of triplet loss is that it only considers the distance of one negative class at a time, without considering all other negative classes. Consequently, in randomly generated data pairs, each pair cannot effectively guarantee that the current optimization direction will increase the distance of all negative samples. This often leads to unstable convergence or stuck in local optima during training.
[0080] In the embodiments of this application, IEMCL (Intermodal Contrastive Learning) is based on the principle that intermodal contrastive learning is supervised, with cross-modal dynamics and interactions between different samples and modalities. Positive pairs are used to represent two different modalities of samples from the same category; negative pairs are used to represent two different modalities of samples from different categories.
[0081] In this embodiment of the application, due to the presence of three modalities, the number of positive and negative pairs for a small batch of anchors of size K is twice that of IAMCL. After softmax normalization of all unimodal representations, the IEMCL loss can be formulated as:
[0082]
[0083] Add refinement loss on this basis It is defined as follows:
[0084]
[0085]
[0086] where a m is the anchor point sample, p m is a positive sample of the same type as the anchor sample, n i is a negative sample of a different class from the anchor point. l represents text, a represents speech, and b represents ECG waveform. R To refine the loss, let a m and pm The vectors of E are as similar as possible. s is the expectation of a small batch s. α is the modal margin between different modes, and the loss between modes is: L IEMCL and sum.
[0087] In the embodiment of the present application, for modal fusion, after extracting sample features from multiple modalities, fusion is required to train the model. This method designs an effective cross-modal Mixup method. The mixup probability of each sample is determined according to the weight of the entropy of each modal feature. Assuming that after modal extraction, the feature of a single modality is n*1 dimensional, the n-dimensional vector x m Each dimension is considered as an evaluation index. The three modal features are considered as the ratings of three people. n defaults to 128 dimensions, and the vector x m Divide into 4 parts and set them as x m,1 , x m,2 , x m,3 , x m,4 Then calculate the fusion features of each equal part, taking the first feature as an example: calculate the uncertainty d of each mode according to the uncertainty formula m,l , take the feature corresponding to the largest uncertainty as the candidate feature s m,1 The characteristics of the second, third, and fourth parts are: m,1 , s m,2 , s m,3 , s m,4 The uncertainty of each characteristic is calculated as follows:
[0088]
[0089]
[0090] Where L is each vector x i The characteristic dimension is 32 again. is the eigenvector x of the mth mode m,i The average value of the L dimensions, d m,i is the mth mode, at x m,i Uncertainty on the vector. m is one of text, speech, and ECG waveform. The value range of i is 1, 2, 3, and 4, which means that the 128 feature vectors of a certain modality are evenly divided into the i-th part. After obtaining the uncertainty of a certain modality, calculate the features of multimodal fusion:
[0091] λ m,i =d m,i / (d a,i +d b,i +d l,i )
[0092] S i =λ a,i *x a,i +λ b,i *x b,i +λ l,i *x l,i i∈(1,2,3,4)
[0093] Among them, λ m,i represents the weight of the i-th part of the m-th modal feature, S i is the weighted sum of the i parts of the three modal features. The final multimodal fusion feature is:
[0094] S = concat(s1,s2,s3,s4)
[0095] The final loss value of the entire model is: L hybird =λ1*L IAMCL +λ2*LI AMCL Among them, λ1 and λ2 are hyperparameters that control the ratio of the two losses. The default values of λ1 and λ2 are 0.5.
[0096] Optionally, the above step of training the multimodal network model using a preset loss function to obtain the trained multimodal network model may include the following steps:
[0097] S1. Constructing positive and negative sample pairs within each modality to obtain multiple first positive and negative sample pair sets, each first positive and negative sample pair set including multiple positive and negative sample pairs;
[0098] S2. constructing positive and negative sample pairs between each modality to obtain multiple second positive and negative sample pair sets, each second positive and negative sample pair set including multiple positive and negative sample pairs;
[0099] S3. Based on the preset loss function, use the multiple first positive-negative sample pair sets and the multiple second positive-negative sample pair sets to train the multimodal network model to obtain the trained multimodal network model.
[0100] In an embodiment of the present application, positive and negative sample pairs can be constructed within each modality to obtain multiple first positive and negative sample pair sets, each first positive and negative sample pair set includes multiple positive and negative sample pairs, and positive and negative sample pairs between each modality are constructed to obtain multiple second positive and negative sample pair sets, each second positive and negative sample pair set includes multiple positive and negative sample pairs. Based on a preset loss function, multiple first positive and negative sample pair sets and multiple second positive and negative sample pair sets are used to train a multimodal network model to obtain a trained multimodal network model. According to the designed positive and negative sample pair selection mechanism, positive and negative sample pairs are selected as training sets in the training set, which can improve the feature extraction efficiency of each modality and help improve the accuracy of emotion recognition.
[0101] Optionally, the above step S1, constructing positive and negative sample pairs within each modality, may include the following steps:
[0102] S11. Obtain n positive samples and m negative samples of sample b of a first modality, where the first modality is any modality in the multimodality, sample b is any sample of the first modality, and n and m are both positive integers;
[0103] S12, determine n corresponding to the n positive samples c The m positive sample cluster centers and the m negative samples corresponding to the m c negative sample cluster centers, where n c Less than n / 2, m c Less than m / 2;
[0104] S13, based on the n c The positive sample cluster centers determine n hard positive pairs;
[0105] S14, based on the m c The negative sample cluster centers determine m hard negative pairs;
[0106] S15. Determine positive and negative sample pairs of the first modality according to the n hard positive pairs and the m hard negative pairs.
[0107] In the embodiment of the present application, n positive samples and m negative samples of sample b of the first modality are obtained, the first modality is any modality in the multimodality, sample b is any sample of the first modality, n and m are both positive integers, and n positive samples corresponding to n negative samples are determined. c There are m positive sample cluster centers and m negative samples corresponding to m c negative sample cluster centers, where n c Less than n / 2, m c Less than m / 2; based on n c The cluster centers of positive samples determine n hard positive pairs, based on m c The negative sample cluster centers determine m hard negative pairs, and the positive and negative sample pairs of the first modality are determined based on n hard positive pairs and m hard negative pairs. This can ensure uniform sampling in the data set, ensure the richness of the samples, and try to select more difficult samples to accelerate model training.
[0108] Among them, pairing generation: in the modality comparison generation stage, in an embodiment of the present application, a pairing method is used to generate corresponding data. Specifically, it is assumed that k samples constitute each batch, wherein each sample can include audio, electrocardiogram, voice and other data. Then, in the training stage, in order to speed up the convergence of the model, positive and negative sample pairs are constructed. When the model comparison is summarized, N positive examples and M negative examples are used. The k samples can be classified according to the same modality, and the positive samples of the same modality are clustered into N centers Nc. The samples closer to the center Nc are dot-producted with the anchor feature in turn to calculate the similarity between each sample and the anchor sample. A smaller dot product means less similarity, and the smallest sample is selected as the hard positive sample p i , each cluster center Nc selects the sample with the smallest dot product with the anchor point as the positive sample; the negative samples of the same modality are clustered into M centers Mc, and the sample n with the largest dot product with the anchor point feature is found i This method can ensure uniform sampling in the dataset and the richness of samples, while also selecting more difficult samples as much as possible to accelerate model training.
[0109] In this embodiment, when constructing positive and negative sample pairs from the same modality, one sample corresponds to n positive samples and m negative samples, where n and m are random. From this batch of samples, data clustering is performed to design n positive sample cluster centers and m negative sample centers. Hard positive pairs are formed by selecting the samples with the smallest similarity from the positive sample clusters, and hard negative pairs are formed by selecting the samples with the largest similarity from the negative sample cluster centers. This method achieves both average sampling and hard sample mining.
[0110] It can be seen that the multimodal emotion recognition method described in the embodiment of the present application is applied to electronic equipment, and the electronic equipment is configured with a multimodal network model. The multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model to obtain multimodal data of the target object. The multimodal data includes: speech data, electrocardiogram data, and natural language data. The speech data is input into the speech recognition model to obtain speech features. The electrocardiogram data is input into the electrocardiogram recognition model to obtain electrocardiogram features. The natural language data is input into the natural language recognition model to obtain text features. The speech features, electrocardiogram features and text features are input into the modal fusion model to obtain emotion recognition results. The three-dimensional features of speech, electrocardiogram and text of the same object are modally fused, and the corresponding emotions are identified, which can improve the accuracy of emotion recognition.
[0111] With the above Figure 1A For details on the embodiments shown, please refer to Figure 2 , Figure 2This is a flow chart of another multimodal emotion recognition method provided in an embodiment of the present application, which is applied to an electronic device. The electronic device is configured with a multimodal network model, which includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model, and a modal fusion model. As shown in the figure, this multimodal emotion recognition method includes:
[0112] 201. Use a preset loss function to train the multimodal network model to obtain the trained multimodal network model; the preset loss function is composed of an intra-modal contrastive learning loss function and a cross-modal contrastive learning loss function.
[0113] 202. Acquire multimodal data of a target object, where the multimodal data includes: voice data, electrocardiogram data, and natural language data.
[0114] 203. Input the speech data into the speech recognition model to obtain speech features.
[0115] 204. Input the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features.
[0116] 205. Input the natural language data into the natural language recognition model to obtain text features.
[0117] 206. Input the speech features, the electrocardiogram features, and the text features into the modality fusion model to obtain an emotion recognition result.
[0118] The detailed description of steps 201 to 206 can refer to the above Figure 1A The corresponding steps of the described multimodal emotion recognition method will not be repeated here.
[0119] It can be seen that the multimodal emotion recognition method described in the embodiments of the present application can, on the one hand, use the model to learn the data features of each modality, fully explore cross-modal interactions, learn the relationships between samples and classes, and reduce the modal gap. On the other hand, it can perform modal fusion of the three-dimensional features of the speech, electrocardiogram and text of the same object, and then identify the corresponding emotions, which can improve the accuracy of emotion recognition.
[0120] In accordance with the above embodiment, please refer to Figure 3 , Figure 3This is a structural diagram of an electronic device provided in an embodiment of the present application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs, which are applied to the electronic device. The one or more programs are stored in the memory and are configured to be executed by the processor. The electronic device is configured with a multimodal network model, which includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model, and a modal fusion model. In the embodiment of the present application, the program includes instructions for performing the following steps:
[0121] Acquire multimodal data of a target object, wherein the multimodal data includes: speech data, electrocardiogram data, and natural language data;
[0122] Inputting the speech data into the speech recognition model to obtain speech features;
[0123] Inputting the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features;
[0124] Inputting the natural language data into the natural language recognition model to obtain text features;
[0125] The speech features, the electrocardiogram features and the text features are input into the modality fusion model to obtain an emotion recognition result.
[0126] Optionally, in the aspect of inputting the speech features, the electrocardiogram features, and the text features into the modal fusion model to obtain the emotion recognition result, the program includes instructions for executing the following steps:
[0127] Dividing the speech feature into a equal parts to obtain a parts of speech features, where a is an integer greater than 1;
[0128] Dividing the electrocardiogram feature into a equal parts to obtain a parts of electrocardiogram features;
[0129] Divide the natural language data into a equal parts, and obtain a parts of text features;
[0130] Mixing the a portions of speech features, the a portions of electrocardiogram features, and the a portions of text features to obtain a portion of mixed features, each portion of the mixed features including a portion of speech features, a portion of electrocardiogram features, and a portion of text features;
[0131] Determine the proportion of various modal uncertainties in each of the a mixed features as a weight, then determine corresponding weighted features according to the weight, and use the weighted features of various modes as candidate features to obtain a candidate features;
[0132] Splicing the a selected features to obtain a fusion feature;
[0133] The emotion recognition result is determined according to the fusion feature.
[0134] Optionally, in determining the emotion recognition result according to the fusion feature, the program includes instructions for executing the following steps:
[0135] The fusion features are sequentially input into a translation encoder, a translation decoder and a fully connected module to obtain the emotion recognition result.
[0136] Optionally, the program further includes instructions for executing the following steps:
[0137] Training the multimodal network model using a preset loss function to obtain a trained multimodal network model;
[0138] The preset loss function consists of an intra-modal contrastive learning loss function and a cross-modal contrastive learning loss function.
[0139] Optionally, in the aspect of training the multimodal network model using a preset loss function to obtain the trained multimodal network model, the program includes instructions for executing the following steps:
[0140] Constructing positive and negative sample pairs within each modality to obtain multiple first positive and negative sample pair sets, each first positive and negative sample pair set including multiple positive and negative sample pairs;
[0141] Constructing positive and negative sample pairs between each modality to obtain multiple second positive and negative sample pair sets, each second positive and negative sample pair set including multiple positive and negative sample pairs;
[0142] Based on the preset loss function, the multimodal network model is trained using the multiple first positive-negative sample pair sets and the multiple second positive-negative sample pair sets to obtain the trained multimodal network model.
[0143] Optionally, in terms of constructing the positive and negative sample pairs within each modality, the program includes instructions for performing the following steps:
[0144] Obtain n positive samples and m negative samples of sample b of a first modality, where the first modality is any modality in the multimodality, sample b is any sample of the first modality, and n and m are both positive integers;
[0145] Determine the n corresponding to the n positive samples c The m positive sample cluster centers and the m negative samples corresponding to the m c negative sample cluster centers, where n c Less than n / 2, m c Less than m / 2;
[0146] Based on the n c The positive sample cluster centers determine n hard positive pairs;
[0147] Based on the m c The negative sample cluster centers determine m hard negative pairs;
[0148] Positive and negative sample pairs of the first modality are determined according to the n hard positive pairs and the m hard negative pairs.
[0149] It can be seen that the electronic device described in the embodiment of the present application is configured with a multimodal network model, which includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model, to obtain multimodal data of the target object, and the multimodal data includes: speech data, electrocardiogram data, and natural language data. The speech data is input into the speech recognition model to obtain speech features, the electrocardiogram data is input into the electrocardiogram recognition model to obtain electrocardiogram features, the natural language data is input into the natural language recognition model to obtain text features, the speech features, electrocardiogram features and text features are input into the modal fusion model to obtain emotion recognition results, and the three-dimensional features of speech, electrocardiogram and text of the same object are modally fused, and the corresponding emotions are identified, which can improve the accuracy of emotion recognition.
[0150] Figure 4 This is a functional unit block diagram of a multimodal emotion recognition device 400 involved in an embodiment of the present application. The multimodal emotion recognition device 400 is applied to an electronic device, and the electronic device is configured with a multimodal network model, and the multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model. The multimodal emotion recognition device 400 may include: an acquisition unit 401, an extraction unit 402 and a recognition unit 403, wherein,
[0151] The acquisition unit 401 is used to acquire multimodal data of the target object, wherein the multimodal data includes: voice data, electrocardiogram data, and natural language data;
[0152] The extraction unit 402 is configured to input the speech data into the speech recognition model to obtain speech features; input the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features; and input the natural language data into the natural language recognition model to obtain text features;
[0153] The recognition unit 403 is used to input the speech features, the electrocardiogram features and the text features into the modality fusion model to obtain an emotion recognition result.
[0154] Optionally, in inputting the speech feature, the electrocardiogram feature, and the text feature into the modal fusion model to obtain the emotion recognition result, the recognition unit 403 is specifically configured to:
[0155] Dividing the speech feature into a equal parts to obtain a parts of speech features, where a is an integer greater than 1;
[0156] Dividing the electrocardiogram feature into a equal parts to obtain a parts of electrocardiogram features;
[0157] Divide the natural language data into a equal parts, and obtain a parts of text features;
[0158] Mixing the a portions of speech features, the a portions of electrocardiogram features, and the a portions of text features to obtain a portion of mixed features, each portion of the mixed features including a portion of speech features, a portion of electrocardiogram features, and a portion of text features;
[0159] Determine the proportion of various modal uncertainties in each of the a mixed features as a weight, then determine corresponding weighted features according to the weight, and use the weighted features of various modes as candidate features to obtain a candidate features;
[0160] Splicing the a selected features to obtain a fusion feature;
[0161] The emotion recognition result is determined according to the fusion feature.
[0162] Optionally, in determining the emotion recognition result according to the fusion feature, the recognition unit 403 is specifically configured to:
[0163] The fusion features are sequentially input into a translation encoder, a translation decoder and a fully connected module to obtain the emotion recognition result.
[0164] Optionally, the device 400 is further specifically configured to:
[0165] Training the multimodal network model using a preset loss function to obtain a trained multimodal network model;
[0166] The preset loss function consists of an intra-modal contrastive learning loss function and a cross-modal contrastive learning loss function.
[0167] Optionally, in the aspect of training the multimodal network model using a preset loss function to obtain the trained multimodal network model, the apparatus 400 is specifically configured to:
[0168] Constructing positive and negative sample pairs within each modality to obtain multiple first positive and negative sample pair sets, each first positive and negative sample pair set including multiple positive and negative sample pairs;
[0169] Constructing positive and negative sample pairs between each modality to obtain multiple second positive and negative sample pair sets, each second positive and negative sample pair set including multiple positive and negative sample pairs;
[0170] Based on the preset loss function, the multimodal network model is trained using the multiple first positive-negative sample pair sets and the multiple second positive-negative sample pair sets to obtain the trained multimodal network model.
[0171] Optionally, in terms of constructing the positive and negative sample pairs within each modality, the apparatus 400 is specifically configured to:
[0172] Obtain n positive samples and m negative samples of sample b of a first modality, where the first modality is any modality in the multimodality, sample b is any sample of the first modality, and n and m are both positive integers;
[0173] Determine the n corresponding to the n positive samples c The m positive sample cluster centers and the m negative samples corresponding to the m c negative sample cluster centers, where n c Less than n / 2, m c Less than m / 2;
[0174] Based on the n c The positive sample cluster centers determine n hard positive pairs;
[0175] Based on the m c The negative sample cluster centers determine m hard negative pairs;
[0176] Positive and negative sample pairs of the first modality are determined according to the n hard positive pairs and the m hard negative pairs.
[0177] It can be seen that the multimodal emotion recognition device described in the embodiment of the present application is applied to an electronic device, and the electronic device is configured with a multimodal network model. The multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model to obtain multimodal data of the target object. The multimodal data includes: speech data, electrocardiogram data, and natural language data. The speech data is input into the speech recognition model to obtain speech features. The electrocardiogram data is input into the electrocardiogram recognition model to obtain electrocardiogram features. The natural language data is input into the natural language recognition model to obtain text features. The speech features, electrocardiogram features and text features are input into the modal fusion model to obtain emotion recognition results. The three-dimensional features of the speech, electrocardiogram and text of the same object are modally fused, and the corresponding emotions are identified, which can improve the accuracy of emotion recognition.
[0178] It can be understood that the functions of each program module of the multimodal emotion recognition device of this embodiment can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description of the above method embodiment and will not be repeated here.
[0179] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.
[0180] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may comprise an electronic device.
[0181] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0182] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0183] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0184] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0185] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0186] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0187] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0188] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A multimodal emotion recognition method, characterized in that: Applied to an electronic device, the electronic device is configured with a multimodal network model, the multimodal network model including: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model, the method comprising: Acquire multimodal data of a target object, the multimodal data including: voice data, electrocardiogram data, and natural language data; the voice data, the electrocardiogram data, and the natural language data are data from the same time period; Inputting the speech data into the speech recognition model to obtain speech features; Inputting the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features; Inputting the natural language data into the natural language recognition model to obtain text features; Inputting the speech features, the electrocardiogram features, and the text features into the modality fusion model to obtain an emotion recognition result; The step of inputting the speech features, the electrocardiogram features, and the text features into the modality fusion model to obtain an emotion recognition result includes: Dividing the speech feature into a equal parts to obtain a parts of speech features, where a is an integer greater than 1; Dividing the electrocardiogram feature into a equal parts to obtain a parts of electrocardiogram features; Divide the natural language data into a equal parts, and obtain a parts of text features; Mixing the a portions of speech features, the a portions of electrocardiogram features, and the a portions of text features to obtain a portion of mixed features, each portion of the mixed features including a portion of speech features, a portion of electrocardiogram features, and a portion of text features; Determine the proportion of various modal uncertainties in each of the a mixed features as a weight, then determine corresponding weighted features according to the weight, and use the weighted features of various modes as candidate features to obtain a candidate features; Splicing the a selected features to obtain a fusion feature; The emotion recognition result is determined according to the fusion feature.
2. The method according to claim 1, characterized in that The determining the emotion recognition result according to the fusion feature includes: The fusion features are sequentially input into a translation encoder, a translation decoder and a fully connected module to obtain the emotion recognition result.
3. The method according to claim 1 or 2, characterized in that The method further comprises: Training the multimodal network model using a preset loss function to obtain a trained multimodal network model; The preset loss function consists of an intra-modal contrastive learning loss function and a cross-modal contrastive learning loss function.
4. The method according to claim 3, characterized in that The multimodal network model is trained using a preset loss function to obtain the trained multimodal network model, including: Constructing positive and negative sample pairs within each modality to obtain multiple first positive and negative sample pair sets, each first positive and negative sample pair set including multiple positive and negative sample pairs; Constructing positive and negative sample pairs between each modality to obtain multiple second positive and negative sample pair sets, each second positive and negative sample pair set including multiple positive and negative sample pairs; Based on the preset loss function, the multimodal network model is trained using the multiple first positive-negative sample pair sets and the multiple second positive-negative sample pair sets to obtain the trained multimodal network model.
5. The method according to claim 4, characterized in that The construction of positive and negative sample pairs within each modality includes: Obtain n positive samples and m negative samples of sample b of a first modality, where the first modality is any modality in the multimodality, sample b is any sample of the first modality, and n and m are both positive integers; Determine the n corresponding to the n positive samples c The m positive sample cluster centers and the m negative samples corresponding to the m c negative sample cluster centers, where n c Less than n / 2, m c Less than m / 2; Based on the n c The cluster centers of positive samples determine n hard positive pairs; Based on the m c The negative sample cluster centers determine m hard negative pairs; Positive and negative sample pairs of the first modality are determined according to the n hard positive pairs and the m hard negative pairs.
6. A multimodal emotion recognition device, characterized in that: Applied to electronic equipment, the electronic equipment is configured with a multimodal network model, the multimodal network model includes: a speech recognition model, an electrocardiogram recognition model, a natural language recognition model and a modal fusion model, the device includes: an acquisition unit, an extraction unit and a recognition unit, wherein, The acquisition unit is configured to acquire multimodal data of a target object, wherein the multimodal data includes: voice data, electrocardiogram data, and natural language data; the voice data, the electrocardiogram data, and the natural language data are data of the same time period; The extraction unit is configured to input the speech data into the speech recognition model to obtain speech features; input the electrocardiogram data into the electrocardiogram recognition model to obtain electrocardiogram features; and input the natural language data into the natural language recognition model to obtain text features; The recognition unit is used to input the speech feature, the electrocardiogram feature and the text feature into the modal fusion model to obtain an emotion recognition result; Wherein, in inputting the speech features, the electrocardiogram features, and the text features into the modal fusion model to obtain the emotion recognition result, the recognition unit is specifically used to: Dividing the speech feature into a equal parts to obtain a parts of speech features, where a is an integer greater than 1; Dividing the electrocardiogram feature into a equal parts to obtain a parts of electrocardiogram features; Divide the natural language data into a equal parts, and obtain a parts of text features; Mixing the a portions of speech features, the a portions of electrocardiogram features, and the a portions of text features to obtain a portion of mixed features, each portion of the mixed features including a portion of speech features, a portion of electrocardiogram features, and a portion of text features; Determine the proportion of various modal uncertainties in each of the a mixed features as a weight, then determine corresponding weighted features according to the weight, and use the weighted features of various modes as candidate features to obtain a candidate features; Splicing the a selected features to obtain a fusion feature; The emotion recognition result is determined according to the fusion feature.
7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store one or more programs and is configured to be executed by the processor, wherein the programs include instructions for executing the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that A computer program for electronic data exchange is stored, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-mode based emotion recognition method
CN108805089A
Multi-modal emotion recognition method based on attention feature fusion
CN109614895A
Text multi-labeling method, device and apparatus and storage medium
CN112560463A
Emotion recognition method, system and equipment and medium
CN113380271A