A method, apparatus, device, and medium for emotion recognition

By constructing multi-dimensional tensors and performing deep fusion, the attention representation in multi-modal emotion recognition is solved, and a higher recognition accuracy is achieved.

CN114168823BActive Publication Date: 2025-06-17HANGZHOU DIANZI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111463362.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-06-17
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

The recognition results of multimodal emotion recognition in the prior art are inaccurate, mainly due to insufficient performance caused by the one-way cross-modal attention mechanism and sequential superposition feature fusion method.

Method used

By obtaining multiple modal features of the target object, constructing multi-dimensional tensors and performing deep fusion, the attention representation of each modal is calculated, and the emotional classification results are obtained using these representations. The specific steps include constructing the first tensor and the second tensor, as query values ​​and key values ​​of the attention mechanism neural network, calculating the kernel tensor and pooling matrix, and performing fusion of modal features and calculating interaction information.

Benefits of technology

The accuracy of multimodal emotion recognition is improved, and through deep fusion and attention representation, the capture of emotional common characteristics and the inhibition of non-common characteristics are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114168823B_ABST
    Figure CN114168823B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and medium for emotion recognition. The method includes: obtaining modal features corresponding to multiple modalities of a target object; respectively constructing a first tensor and a second tensor based on the multiple modal features, wherein the dimensions of the first tensor and the second tensor are determined based on the number of modalities; using the first tensor and the second tensor as the query value and the key value of a neural network based on an attention mechanism respectively; obtaining a kernel tensor corresponding to each modality and a pooling matrix obtained based on the kernel tensor according to the query value and the key value; for a target modality among the multiple modalities, fusing the kernel tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality to obtain an attention representation corresponding to the target modality; and obtaining an emotion classification result of the target object based on the attention representation of each modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and more particularly, to a method, device, equipment, and medium for emotion recognition. Background Art

[0002] With the progress of technology, the network environment has become more and more efficient, and artificial intelligence based on the network environment has also developed faster. Among them, emotion recognition has always been a hot topic in the field of artificial intelligence, which is beneficial to people's social communication and activities. In the process of people's life and communication, there are various ways to express emotions.

[0003] The ways to express emotions can include any one of the following modal data: text data, video data, voice data, etc. At present, emotion analysis focuses on the research of emotion analysis of single-modal data. However, existing research has proved that there are feature information related to emotion state discrimination between different modal data, and there are also complementary and consistent emotion information between modal features. At present, when performing emotion analysis on multiple modal data, the learning model is basically a one-way cross-modal attention mechanism, and multiple one-way cross-modal features are stacked in sequence between multiple modalities to obtain the emotion analysis results of multiple modal data. However, this analysis method has weak processing performance. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide a method, device, equipment, and medium for emotion recognition, which is used to solve the problem of inaccurate recognition results in multi-modal emotion recognition in the prior art.

[0005] In a first aspect, an embodiment of the present application provides a method for emotion recognition, including:

[0006] Obtaining modal features corresponding to multiple modalities of a target object;

[0007] Constructing a first tensor and a second tensor based on the multiple modal features respectively, where the dimensions of the first tensor and the second tensor are determined based on the number of modalities;

[0008] Using the first tensor and the second tensor as the query value and the key value of a neural network based on the attention mechanism respectively;

[0009] Obtaining a kernel tensor corresponding to each modality and a pooling matrix obtained based on the kernel tensor according to the query value and the key value;

[0010] For a target modality among the multiple modalities, fusing the kernel tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality to obtain an attention representation corresponding to the target modality;

[0011] Obtain the emotion classification result of the target object based on the attention representation of each modality.

[0012] Optionally, obtaining the emotion classification result of the target object based on the attention representation includes:

[0013] Take the modality features corresponding to each modality as the value of the neural network;

[0014] Based on the value and the attention representation, obtain the inter-modal emotion interaction information dominated by the modality features of each modality through the neural network;

[0015] Obtain the emotion classification result of the target object based on each piece of the emotion interaction information.

[0016] Optionally, the method further includes:

[0017] Perform average pooling operation on the kernel tensor of each modality along the time dimension corresponding to the modality to obtain the pooling matrix of the kernel tensor.

[0018] Optionally, obtaining the emotion classification result of the target object based on each piece of the emotion interaction information includes:

[0019] Input the superposition result of each piece of the emotion interaction information into the linear classification layer of the neural network to obtain the emotion classification result of the target object.

[0020] Optionally, before obtaining the emotion classification result of the target object based on each piece of the emotion interaction information, the method further includes:

[0021] Repeat the following steps a preset number of times:

[0022] For each modality, take the emotion interaction information of the modality as the new modality feature;

[0023] Construct a first tensor and a second tensor respectively based on the multiple modality features, and the dimensions of the first tensor and the second tensor are determined based on the number of modalities;

[0024] Take the first tensor and the second tensor as the query value and the key value of the neural network based on the attention mechanism respectively;

[0025] Obtain the kernel tensor corresponding to each modality respectively according to the query value and the key value, and the pooling matrix obtained based on the kernel tensor;

[0026] For a target modality among the multiple modalities, fuse the core tensor corresponding to the target modality and the pooling matrices respectively corresponding to the other modalities except the target modality to obtain the attention representation corresponding to the target modality;

[0027] Use the modality feature corresponding to each modality as the value of the neural network;

[0028] Based on the value and the attention representation, obtain the cross-modal sentiment interaction information dominated by the modality feature of each modality through the neural network.

[0029] Optionally, the first tensor includes a plurality of first low-rank tensors respectively corresponding to the multiple modalities, and the second tensor includes a plurality of second low-rank tensors respectively corresponding to the multiple modalities;

[0030] The obtaining the core tensor corresponding to each modality according to the query value and the key value includes:

[0031] For a target modality among the multiple modalities, obtain the core tensor corresponding to the target modality according to the first low-rank tensor and the second low-rank tensor corresponding to the target modality.

[0032] Optionally, the fusing the core tensor corresponding to the target modality and the pooling matrices respectively corresponding to the other modalities except the target modality to obtain the attention representation corresponding to the target modality includes:

[0033] For a target modality among the multiple modalities, perform a tensor contraction operation on the core tensor corresponding to the target modality and the pooling matrices respectively corresponding to the other modalities except the target modality to obtain the attention representation corresponding to the target modality.

[0034] In a second aspect, an embodiment of the present application provides an emotion recognition device, including:

[0035] An acquisition module, configured to acquire modality features respectively corresponding to multiple modalities of a target object;

[0036] A construction module, configured to respectively construct a first tensor and a second tensor based on the multiple modality features, where the dimensions of the first tensor and the second tensor are determined based on the number of modalities;

[0037] A first determination module, configured to use the first tensor and the second tensor as the query value and the key value of a neural network based on an attention mechanism respectively;

[0038] A second determination module, configured to obtain the core tensor corresponding to each modality according to the query value and the key value, and a pooling matrix obtained based on the core tensor;

[0039] A third determination module, configured to fuse, for a target modality among the multiple modalities, a kernel tensor corresponding to the target modality and pooling matrices corresponding to other modalities except the target modality, to obtain an attention representation corresponding to the target modality;

[0040] A classification module, configured to obtain an emotion classification result of the target object based on the attention representation of each modality.

[0041] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above method are executed.

[0043] In some embodiments, multi-dimensional first and second tensors are constructed using modality features of multiple modalities of a target object, and the number of dimensions is the same as the number of multiple modalities. A kernel tensor and pooling matrices calculated using the first and second tensors are used for deep fusion of modality features among multiple modalities, and an attention representation of each modality is calculated. An emotion classification result of the target object can be calculated using the attention representation of each modality. By the above tensor calculation and feature fusion, the accuracy of determining the emotion classification result is improved.

[0044] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides detailed descriptions as follows. Description of the Drawings

[0045] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 A flowchart of a method for emotion recognition provided by an embodiment of the present application;

[0047] Figure 2 A schematic diagram of a neural network provided by an embodiment of the present application;

[0048] Figure 3A schematic structural diagram of an emotion recognition device provided by an embodiment of the present application;

[0049] Figure 4 A schematic structural diagram of a computer device provided by an embodiment of the present application;

[0050] Figure 5 A schematic flowchart of another emotion recognition method provided by an embodiment of the present application. Detailed implementation manners

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present application. The components of the embodiments of the present application described and illustrated herein generally can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0052] Different modalities refer to different forms of data representation. The representation forms of different modality data include: time, controls, and content, etc. The content includes any one of the following content information: text information, voice information, and video information. For example, different description information in terms of content refers to the difference in the display form of the description information when evaluating the same target object. For example, under the same time and the same space, the text information, voice information, and video information generated when the user evaluates the target object. Different description information in terms of time refers to the difference in the time of the description information when evaluating the same target object. For example, under the same space, the text information generated when the user evaluates the same target object at two different times. Different description information in terms of space refers to the difference in the space of the description information when evaluating the same target object. For example, under the same time (8 am on two adjacent days), the video information generated when the user evaluates the same target object in two different spaces. There are feature information related to emotion state discrimination between different modality data, and there are also complementary and consistent emotion information between modality features. At present, when performing emotion analysis on multiple modality data, the learning model is basically a one-way cross-modal attention mechanism, and multiple one-way cross-modal features are stacked in sequence between multiple modalities to obtain the emotion analysis results of multiple modality data. This analysis method has weak processing performance.

[0053] Based on the above defects, an embodiment of the present application provides a method for emotion recognition, as Figure 1 shown, including the following steps:

[0054] S101, obtaining modal features corresponding to multiple modalities of a target object;

[0055] S102, respectively constructing a first tensor and a second tensor based on the multiple modal features, where the dimensions of the first tensor and the second tensor are determined based on the number of modalities;

[0056] S103, respectively using the first tensor and the second tensor as the query value and the key value of a neural network based on an attention mechanism;

[0057] S104, obtaining a kernel tensor corresponding to each modality and a pooling matrix obtained based on the kernel tensor according to the query value and the key value;

[0058] S105, for a target modality among the multiple modalities, fusing the kernel tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality to obtain an attention representation corresponding to the target modality;

[0059] S106, obtaining an emotion classification result of the target object based on the attention representation of each modality.

[0060] In the above step S101, the target object is an object that represents multiple modal data, which can be a user. The modal features corresponding to the multiple modalities are the modal features corresponding to the multiple modal data represented by the target object. The modal features include the temporal feature, spatial feature, and content feature of the modal data. Since the content includes any one of the following information: text information, voice information, and video information, the content feature includes any one of the following features: text feature, voice feature, and video feature. Among them, when the content is text information, the text feature includes any one or two of the following features: grammar feature and semantic feature. The text feature can be extracted by a text feature extraction model, where the text feature extraction model can be a TF-IDF (term frequency–inverse document frequency) model, an N-gram model, a Long Short-Term Memory (LSTM) network, etc. When the content is voice information, the voice feature includes any one or more of the following features: sound intensity feature, loudness feature, pitch feature, fundamental period feature, and fundamental frequency feature. The voice feature can be extracted by a voice feature extraction model, where the voice feature extraction model can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), etc. When the content is video information, the video feature includes any one or two of the following features: facial feature and limb movement feature. The video feature can be extracted by an image feature extraction model, where the image feature extraction model can be a neural network such as CNN, RNN, LSTM, etc.

[0061] In the above steps S102 and S103, the present application uses a neural network with an attention mechanism to determine the emotional classification result of the target object. In the neural network with an attention mechanism, there are background variables (hidden states). In the prior art, the above background variables are abstracted into query values Query (hereinafter simply referred to as Q), key values Key (hereinafter simply referred to as K), and value values Value (hereinafter simply referred to as V). In the neural network with an attention mechanism, a large number of parameters require high resource support. Among them, the query value Q and the key value K are parameters with relatively large amounts of data in the neural network with an attention mechanism. Therefore, in order to reduce the amount of data processing, the present application uses a tensor network to decompose a large tensor into a tensor chain containing multiple small tensors, that is, decomposes the modal features into tensor chains corresponding to the query value Q and the key value K respectively. Among them, the first tensor is the tensor chain corresponding to the query value Q, and the tensor chain corresponding to the first tensor includes multiple first low-rank tensors respectively corresponding to the multiple modalities. The second tensor is the tensor chain corresponding to the key value K, and the tensor chain corresponding to the second tensor includes multiple second low-rank tensors respectively corresponding to the multiple modalities.

[0062] The modal features of different modalities exist independently of each other, and the interactivity between modal features is not strong. In order to increase the interactivity between different modal features, different modal features are mapped to the same spatial coordinate system, and each dimension of this spatial coordinate system corresponds to the time dimension of each modality respectively. In the present application, the first tensor and the second tensor are a prelude to increasing the interactivity between different modal features. Therefore, the first tensor and the second tensor are also constructed in the above spatial coordinate system, that is, the dimensions of the first tensor and the second tensor are determined based on the number of modalities.

[0063] The present application takes obtaining the modal features of three modalities of the target object (the three modalities are a, v, and t respectively, and the modal features corresponding to the three modalities are X a 、X v 、X t ) as an example to establish a three-dimensional spatial coordinate system, and continue to use this example to explain the formula of each step in the following text. And the present application uses a neural network with an attention mechanism to determine the emotional classification result of the target object. In order to understand the processing mechanism in the neural network with an attention mechanism in more detail, the present application combines Figure 2 for further explanation.

[0064] For each modality, the following formula is used to construct the first tensor corresponding to the query value according to the modal feature of this modality:

[0065]

[0066] Among them, i is the i-th modality among the multiple modalities, The weight matrices of the tensor network of the first tensor of the i-th modality, and are all two-dimensional matrices, d i is the eigenvector dimension at each moment, R w and R s are the ranks corresponding to the tensor network (w, s ∈ {1,..., M}), M is the number of multiple modalities, Q represents the query, is the key-value tensor of the i-th modality, X i is the modality feature of the i-th modality, ⊙1 is the Khatri-Rao operation, reshape is the matrix transformation function; represents is a three-dimensional tensor, which is the user description feature of the i-th modality in the time domain dimension, R w dimension and R s dimension, T i is the dimension of the i-th modality in the time domain dimension.

[0067] That is, as Figure 2 shown, X a 、X v 、X t After passing through formula 1 corresponding to the first stage, the first tensor in the second stage can be obtained, that is, Figure 2 the Q in, the first tensor (Q) includes the first low-rank tensor corresponding to the a modality the first low-rank tensor corresponding to the v modality the first low-rank tensor corresponding to the t modality and the connection ranks between each tensor. Among them, and these two tensors are connected by R1, and these two tensors are connected by R2, and these two tensors are connected by R3.

[0068] For each modality, use the following formula to construct the second tensor corresponding to the key value according to the modality feature of this modality:

[0069]

[0070] where i is the i-th modality among multiple modalities, The weight matrices of the tensor network of the second tensor of the i-th modality, and are all two-dimensional matrices, d i is the eigenvector dimension at each moment, R w and R s are the ranks corresponding to the tensor network (w, s ∈ {1,..., M}), M is the number of multiple modalities, i is the i-th modality among multiple modalities, K represents the key value, is the second tensor of the i-th modality, X i is the modal feature of the i-th modality, ⊙1 is the Khatri-Rao operation, and reshape is the matrix transformation function; characterizes is a three-dimensional tensor, which is the user description feature of the i-th modality in the time domain dimension, R w dimension and R s dimension, T i is the size of the i-th modality in the time domain dimension.

[0071] That is, as Figure 2 shown, X a , X v , X t After passing through the corresponding formula 2 in the first stage, the

[0072] That is, as Figure 2 shown, X a , X v , X t After passing through the corresponding formula 1 in the first stage, the second tensor in the second stage can be obtained, that is, Figure 2 Q in, and the second tensor (Q) includes the second low-rank tensor corresponding to the v modality of the second low-rank tensor corresponding to the a modality second low-rank tensor corresponding to the t modality and the connection ranks between the tensors. Among them, and these two tensors are connected by R1, and these two tensors are connected by R2, and these two tensors are connected by R3.

[0073] In the above step S104, in the neural network of the attention mechanism, the kernel tensor can be calculated using the query value and the key value.

[0074] Specifically, for the target modality among the multiple modalities, the kernel tensor corresponding to the target modality is obtained according to the first low-rank tensor and the second low-rank tensor corresponding to the target modality.

[0075] For the target modality among the multiple modalities, the following formula is used to obtain the kernel tensor corresponding to the target modality according to the first low-rank tensor and the second low-rank tensor corresponding to the target modality:

[0076]

[0077] Among them, i is the i-th modality among multiple modalities, is the hadamard operation, is the second low-rank tensor of the i-th modality, is the first low-rank tensor of the i-th modality, is the core tensor of the i-th modality in the three-dimensional space; represents is a three-dimensional tensor, which is the user description feature of the i-th modality in the time domain dimension, R w dimension and R s dimension, and T i is the size of the i-th modality in the time domain dimension.

[0078] That is, as Figure 2 shown, in the second stage, and can output Specifically, for modality a, in the second stage represents formula three, and it can be seen from the figure that and through formula 3, the first cube outside the dashed box in the second stage can be calculated For modality v, in the second stage represents formula three, and it can be seen from the figure that and through formula 3, the second cube outside the dashed box in the second stage can be calculated For modality t, in the second stage represents formula three, and it can be seen from the figure that and through formula 3, the third cube outside the dashed box in the second stage can be calculated

[0079] In the neural network of the attention mechanism, the pooling matrix can be calculated using the core tensor.

[0080] Specifically, calculating the pooling matrix using the core tensor includes the following steps:

[0081] Perform average pooling operation on the core tensor of each modality along the time dimension corresponding to the modality to obtain the pooling matrix of the core tensor.

[0082] For each modality, use the following formula to perform average pooling operation on the core tensor of this modality along the time dimension to obtain the pooling matrix of this modality:

[0083]

[0084] where 1 ≤ j ≤ R w , 1 ≤ p ≤ R s ; R w and R s are the ranks corresponding to the tensor network (w, s ∈ {1,..., M}), M is the number of multiple modalities, i is the i-th modality among the multiple modalities, is the core tensor in the three-dimensional space of the i-th modality, average is the average pooling operation, is the pooling matrix of the i-th modality in the three-dimensional space after pooling; is a two-dimensional matrix, which is the pooling matrix of the i-th modality in the R w dimension and the R s dimension.

[0085] That is, as Figure 2 shown, in the third stage, respectively calculated through formula 4 to obtain the corresponding

[0086] In the above step S105, the target modality is any one of the multiple modalities. Although there are common sentiment features in each modality feature, there are more non-common features. Therefore, when extracting the common sentiment features among multiple modalities, it is necessary to minimize the influence of non-common features. For each modality, the pooling matrix reduces the dimension of the time dimension corresponding to this modality in the core tensor, and the dimensions corresponding to other modalities remain unchanged, so the features of this modality in other dimensions will be retained. For the target modality, fusing the core tensor of the target modality with the pooling matrices corresponding to other modalities except the target modality can reduce the influence of non-common features of other modalities, increase the complexity of the common features among multiple modalities, and calculate the attention representation corresponding to the target modality using the fused features.

[0087] Specifically, for the target modality among the multiple modalities, perform a tensor contraction operation on the core tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality to obtain the attention representation corresponding to the target modality.

[0088] For the target modality among the multiple modalities, use the following formula to perform a tensor contraction operation on the core tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality, and calculate the attention representation corresponding to the target modality:

[0089]

[0090] where j m ∈ {1,..., M}, and jm ≠i, where M is the number of multiple modalities, i is the i-th modality among the multiple modalities, linear is a linear function, is the modality feature in the three-dimensional space of the i-th modality, is the pooling matrix of the j-th modality after pooling, atten M is the pooling matrix of the j-th modality after pooling, atten i is the attention representation of the i-th modality.

[0091] Specifically, as Figure 2 shown, in the fourth stage, feature fusion is performed using the kernel tensor and pooling matrix of each modality, and the fused features are input into the linear function in the fifth stage to obtain the attention representation corresponding to each modality respectively.

[0092] In the above step S106, the sentiment classification result is used to characterize the sentiment of the target object, and the sentiment classification includes any one of the following categories: happy, sad, angry, fearful, disgusted, surprised, etc.

[0093] What this application needs to identify is the common sentiment classification result among multiple modalities of the target object, so the attention representation corresponding to each modality needs to be used for calculation. The specific calculation process includes the following steps:

[0094] Step 1061, use the modality feature corresponding to each modality as the value of the neural network;

[0095] Step 1062, based on the value and the attention representation, obtain the inter-modal sentiment interaction information dominated by the modality feature of each modality through the neural network;

[0096] Step 1063, obtain the sentiment classification result of the target object based on each of the sentiment interaction information.

[0097] In the above step 1061, it was mentioned above that there is also a background variable in the neural network of the attention mechanism, that is, the value V. In this application, the modality feature corresponding to the modality is used as the value v in the neural network.

[0098] In the above step 1062, in the neural network, using the value v and attention representation corresponding to each modality mentioned above, the inter-modal sentiment interaction information corresponding to each modality can be calculated. The inter-modal sentiment interaction information dominated by the modality feature of each modality is a feature that fuses the common features with other modalities on the basis of the original modality feature.

[0099] For each modality, use the following formula to calculate the inter-modal sentiment interaction information dominated by the modality feature of this modality according to the value and attention representation of this modality:

[0100] Y i = atten i * X i + a * X i Equation (6)

[0101] where a is a coefficient, and atten i is the attention representation of the i-th modality, and X i is the modality feature of the i-th modality.

[0102] In the above step 1063, obtaining the sentiment classification result of the target object by using the sentiment interaction information of each modality can include the following two methods. Method 1: Determine the sentiment classification result corresponding to each sentiment interaction information respectively, and determine the sentiment classification result of the target object according to the proportion of the same sentiment classification results in the sentiment classification results, and determine the sentiment classification result with the largest proportion as the sentiment classification result of the target object. Method 2: Stack the sentiment interaction information of multiple modalities, and use the stacked features to determine the sentiment classification result of the target object.

[0103] Specifically, in the above Method 1, each sentiment interaction information can be input into the linear classification layer of the neural network respectively, and the sentiment classification results corresponding to different modalities can be obtained.

[0104] Specifically, in the above Method 2, the stacked result of each of the sentiment interaction information is input into the linear classification layer of the neural network to obtain the sentiment classification result of the target object.

[0105] In the above Method 1, although the sentiment interaction information of each modality is obtained after the interaction of multiple modality features, there will still be a part of the non-common features corresponding to each modality. Therefore, if the sentiment classification results corresponding to the sentiment interaction information of each modality are determined respectively, the output sentiment classification results are likely to deviate towards the non-common features. If there are situations of deviation towards the non-common features in multiple modalities, the determined sentiment classification results are not accurate. In Method 2, the sentiment interaction information of multiple modalities is stacked, so that the same common features will become more complex. In the process of analyzing the sentiment classification results, the proportion of the common features is larger, and the proportion of the non-common features corresponding to each modality is smaller. Then the analyzed sentiment classification results will be more biased towards the common features. Therefore, the sentiment classification result of the target object determined by Method 2 is more accurate.

[0106] Since the modal features extracted from the modalities are relatively complex data, if the modal features of one modality are only fused with the modal features of other modalities once through step S104, the interaction between different modalities is relatively simple, and the common emotional features that can be extracted from the modalities will also be relatively few. Therefore, it is necessary to increase the interaction depth and interaction complexity of features between different modalities to increase the common emotional features extracted from the modalities. That is, before step S1063, the method provided by the present application further includes:

[0107] Repeat the following steps a preset number of times:

[0108] Step 201, for each modality, use the emotional interaction information of the modality as the new modal feature;

[0109] Step 202, respectively construct a first tensor and a second tensor based on the multiple modal features, and the dimensions of the first tensor and the second tensor are determined based on the number of modalities;

[0110] Step 203, use the first tensor and the second tensor as the query value and the key value of the neural network based on the attention mechanism respectively;

[0111] Step 204, obtain the core tensor corresponding to each modality and the pooling matrix obtained based on the core tensor according to the query value and the key value;

[0112] Step 205, for the target modality among the multiple modalities, fuse the core tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality to obtain the attention representation corresponding to the target modality;

[0113] Step 206, use the modal feature corresponding to each modality as the value of the neural network;

[0114] Step 207, based on the value and the attention representation, obtain the emotional interaction information between modalities dominated by the modal features of each modality through the neural network.

[0115] Step 208, the preset number of times here is set manually, for example, 500 times.

[0116] In the above steps 201 - 208, for each modality, the cross-modal emotional interaction information dominated by the modality features of this modality is regarded as a new modality feature of this modality, and steps 202 - 208 are re-executed. In this way, through multiple iterations of fusion, the common features between this modality and other modalities become more prominent, and the non-common emotional features become weaker. Equivalently, during the iteration process, other modalities impose constraints on this modality, making the new modality features corresponding to this modality and other modalities more and more similar. Furthermore, in the cross-modal emotional interaction information dominated by each modality obtained after the iteration ends, there are more and more common features. Thus, when step 1063 is executed, a more accurate emotional classification result of the target object can be obtained.

[0117] The present application provides a flowchart of another method for emotion recognition. As Figure 5 shown, the neural network for executing the above method for emotion recognition includes multiple MMTs. The MMTs represent the iterative process corresponding to the above steps 201 - 208. This method for emotion recognition can recognize two emotion classification results, which are respectively Figure 5 the smiling face and the crying face in []. Among them, the smiling face represents the emotion classification result of happiness, and the crying face represents the emotion classification result of sadness. The modality features corresponding to the modality data (Audio data, Video data, Text data) of the three modalities expressed by the user are input into the neural network. The initial modality features corresponding to each modality are input into the first MMT, and the input parameter of any MMT after the second one is the output result of the previous MMT. The output result of the last MMT (that is, the cross-modal emotional interaction information corresponding to each modality output for the last time) is input into the Classification in the figure (that is, the linear classification layer of the neural network corresponding to the above step 1063), and then the probabilities that the modality data of the three modalities are respectively each emotion classification result are output. The emotion classification result with the highest probability is determined as the emotion of the user. If the probability of the smiling face is relatively high in the figure, then the emotion classification result corresponding to the user is happiness.

[0118] When extracting the corresponding modality features for each modality, different modality feature extraction models need to be used for different modalities. This may lead to a large data distribution and relatively discrete data among the modality features extracted from different modalities. Furthermore, in the subsequent data processing process, the complexity of data processing is increased and the data processing efficiency is reduced. Therefore, it is necessary to preprocess the features extracted from the modality to obtain modality features that can optimize the calculation. Step S101 includes:

[0119] Step 1021, for each modality, extract the modality features that have not been preprocessed from the modality, and perform normalization processing on the modality features that have not been preprocessed through a preprocessing linear network to obtain the modality features of this modality.

[0120] In the above step 1021, the modality features without preprocessing are data information extracted by using the corresponding feature extraction model for each modality. The data distribution differences among the modality features without preprocessing are relatively large. For example, the data extracted from text information is between 0 and 1, while the data extracted from video information may be between -100 and +100. The data corresponding to the modality features after preprocessing may all be between 0 and 1.

[0121] In specific implementation, the present application uses the following formula to perform normalization processing on the modality features without preprocessing to obtain the modality features of the modality:

[0122] X i = f(X i ) = W i *P i + b i Formula (1)

[0123] where, X i is the modality feature of the i-th modality, P i is the modality feature without preprocessing of the i-th modality, W i is the corresponding linear network weight matrix of the i-th modality, b i is the bias vector of the linear network, then represents that X i is a two-dimensional feature, which is the modality feature of the i-th modality in the time dimension and the feature dimension, T i is the size of the i-th modality in the time domain dimension, and d i is the size of the feature vector at each moment.

[0124] During the execution of the above multi-modal emotion recognition method, it is implemented by a neural network based on the attention mechanism. The reason why the neural network can achieve the recognition of emotion classification is that it has learned a large amount of experience from a large number of training data. The present application provides a training method for an emotion category recognition model, and the training method includes:

[0125] Step 107, obtaining an emotion training sample set, where the training sample set includes at least one training sample, and each training sample includes the modality data of at least two modalities of a target object and the common true emotion classification result corresponding to at least two modalities;

[0126] Step 108: For each training sample, the neural network to be trained extracts the modal features of each modality respectively; based on the multiple modal features, a first tensor and a second tensor are constructed respectively, and the dimensions of the first tensor and the second tensor are determined based on the number of modalities; the first tensor and the second tensor are used as the query value and the key value of the neural network based on the attention mechanism respectively; according to the query value and the key value, the core tensor corresponding to each modality and the pooling matrix obtained based on the core tensor are obtained; for the target modality among the multiple modalities, the core tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality are fused to obtain the attention representation corresponding to the target modality; based on the attention representation of each modality, the predicted sentiment classification result of the target object is predicted; based on the comparison result between the predicted sentiment classification result and the true sentiment classification result, the sentiment category recognition model to be trained is trained.

[0127] In the above step 108, during the neural network training process, the values of W in formula (7) are continuously adjusted i and b i , and the values of and in formula (1) are adjusted, the values of and in formula (2), and the value of a in formula (6).

[0128] The multi-modal sentiment recognition method provided by this application is mainly used to recognize the common sentiment among multiple modalities expressed by users. The above multi-modal sentiment recognition method can be applied to many usage scenarios. For example, applying the above multi-modal sentiment recognition method to the commodity recommendation scenario, on a certain shopping platform, after a user purchases a certain commodity, the user will evaluate the commodity, and the evaluation includes various forms of descriptions, such as text information, video information, and text information in a follow-up review after a period of time. Taking the user's evaluation as the modal data corresponding to the modality, using this multi-modal sentiment recognition method to perform sentiment analysis on the user's evaluation on the shopping platform, and determining the sentiment classification result of the user for the commodity. According to the category of the commodity and the sentiment classification result of the user for the commodity, it is determined whether it is still necessary to recommend commodities of this category to the user. For example, if the commodity is a whitening product and the user's sentiment category for the commodity is happy, the shopping platform will increase the proportion of recommending whitening products to the user; if the commodity is a whitening product and the user's sentiment category for the commodity is angry, the shopping platform will reduce the proportion of recommending whitening products to the user. That is, in the commodity recommendation scenario, through the above sentiment recognition method, according to the user's evaluation of the commodity, the sentiment of the user for the commodity can be accurately determined, and then according to this sentiment, the accuracy of the shopping platform in recommending commodities to the user can be improved.

[0129] For another example, when the above multi-modal emotion recognition method is applied to the online customer service scenario, during the process of communicating with the online customer service, video data containing the player's image, voice data expressed by the player, and text data input by the player in the dialog box will be obtained in real time. According to the obtained multi-modal data, the current emotion classification result of the user can be analyzed. Furthermore, according to the emotion classification result of the user, the online customer service will automatically screen out the reply content that matches the emotion classification result of the user, improving the accuracy of the reply from the online customer service to the user.

[0130] This application provides an emotion recognition device, as Figure 3 shown, the device includes:

[0131] An acquisition module 301, configured to acquire modal features corresponding to multiple modalities of a target object;

[0132] A construction module 302, configured to respectively construct a first tensor and a second tensor based on the multiple modal features, where the dimensions of the first tensor and the second tensor are determined based on the number of modalities;

[0133] A first determination module 303, configured to respectively use the first tensor and the second tensor as the query value and the key value of a neural network based on the attention mechanism;

[0134] A second determination module 304, configured to obtain a kernel tensor corresponding to each modality and a pooling matrix obtained based on the kernel tensor according to the query value and the key value;

[0135] A third determination module 305, configured to fuse the kernel tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality among the multiple modalities to obtain an attention representation corresponding to the target modality;

[0136] A classification module 306, configured to obtain an emotion classification result of the target object based on the attention representation of each modality.

[0137] Optionally, the classification module includes:

[0138] A first determination unit, configured to use the modal feature corresponding to each modality as the value of the neural network;

[0139] A second determination unit, configured to obtain emotion interaction information between modalities dominated by the modal features of each modality through the neural network based on the value and the attention representation;

[0140] A classification unit, configured to obtain an emotion classification result of the target object based on each emotion interaction information.

[0141] Optionally, the device further includes:

[0142] A pooling unit, configured to perform an average pooling operation on the kernel tensor of each modality along the time dimension corresponding to the modality, to obtain a pooling matrix of the kernel tensor.

[0143] Optionally, the classification unit includes:

[0144] A classification subunit, configured to input the superposition result of each piece of the emotion interaction information into a linear classification layer of the neural network, to obtain an emotion classification result of the target object.

[0145] Optionally, the device further includes:

[0146] An iterative module, configured to repeatedly execute the following steps a preset number of times: for each modality, use the emotion interaction information of the modality as a new modality feature; respectively construct a first tensor and a second tensor based on the multiple modality features, where the dimensions of the first tensor and the second tensor are determined based on the number of modalities; use the first tensor and the second tensor as a query value and a key value of a neural network based on an attention mechanism respectively; obtain a kernel tensor corresponding to each modality and a pooling matrix obtained based on the kernel tensor according to the query value and the key value; for a target modality among the multiple modalities, fuse the kernel tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality, to obtain an attention representation corresponding to the target modality; use the modality feature corresponding to each modality as a value of the neural network; based on the value and the attention representation, obtain emotion interaction information between modalities mainly dominated by the modality feature of each modality through the neural network.

[0147] Optionally, the first tensor includes a plurality of first low-rank tensors respectively corresponding to the multiple modalities, and the second tensor includes a plurality of second low-rank tensors respectively corresponding to the multiple modalities;

[0148] The second determination module includes:

[0149] A third determination unit, configured to, for a target modality among the multiple modalities, obtain a kernel tensor corresponding to the target modality according to the first low-rank tensor and the second low-rank tensor corresponding to the target modality.

[0150] Optionally, the third determination module includes:

[0151] A fourth determination unit, configured to, for a target modality among the multiple modalities, perform a tensor contraction operation on the kernel tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality, to obtain an attention representation corresponding to the target modality.

[0152] Corresponding to Figure 1 the method for emotion recognition in, embodiments of the present application further provide a computer device 400, as Figure 4 shown, the device includes a memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402. Among them, when the above-mentioned processor 402 executes the above-mentioned computer program, the above-mentioned method for emotion recognition is implemented.

[0153] Specifically, the above-mentioned memory 401 and processor 402 can be general memories and processors, which are not specifically limited here. When the processor 402 runs the computer program stored in the memory 401, it can execute the above-mentioned method for emotion recognition, solving the problem of inaccurate recognition results in multi-modal emotion recognition in the prior art.

[0154] Corresponding to Figure 1 the method for emotion recognition in, embodiments of the present application further provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the above-mentioned method for emotion recognition are executed.

[0155] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned method for emotion recognition, solving the problem of inaccurate recognition results in multi-modal emotion recognition in the prior art. The present application constructs a multi-dimensional first tensor and second tensor using the modal features of multiple modalities of a target object. The number of dimensions is the same as the number of multiple modalities. The kernel tensor and pooling matrix calculated using the first tensor and second tensor are used for deep fusion of modal features between multiple modalities, and the attention representation of each modality is calculated. The emotion classification result of the target object can be calculated using the attention representation of each modality. The accuracy of determining the emotion classification result is improved through the above-mentioned tensor calculation and feature fusion.

[0156] In the embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0157] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, each functional unit in the embodiments provided in this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0159] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0160] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0161] Finally, it should be noted that: the above-described embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for emotion recognition, characterized in that, Including: Obtaining modal features corresponding to multiple modalities of a target object; Decomposing the multiple modal features into a first tensor and a second tensor through a tensor network; wherein, the first tensor represents a tensor chain including multiple first low-rank tensors respectively corresponding to the multiple modalities, the second tensor represents a tensor chain including multiple second low-rank tensors respectively corresponding to the multiple modalities, and the dimensions of the first tensor and the second tensor are determined based on the number of modalities; Taking the first tensor and the second tensor as the query value and the key value of a neural network based on an attention mechanism respectively; Obtaining a kernel tensor corresponding to each modality and a pooling matrix obtained based on the kernel tensor according to the query value and the key value; wherein, for each modality, obtaining the kernel tensor corresponding to this modality according to the first low-rank tensor and the second low-rank tensor corresponding to this modality; performing an average pooling operation on the kernel tensor of this modality along the time dimension corresponding to the modality to obtain the pooling matrix corresponding to this modality; For a target modality among the multiple modalities, performing feature fusion on the kernel tensor corresponding to the target modality and the pooling matrices corresponding to other modalities except the target modality, and inputting the fused features into a linear function to obtain an attention representation corresponding to the target modality; Obtaining an emotion classification result of the target object based on the attention representation of each modality.

2. The method according to claim 1, characterized in that, Obtaining the emotion classification result of the target object based on the attention representation includes: Taking the modal feature corresponding to each modality as the value of the neural network; Based on the value and the attention representation, obtaining emotion interaction information between modalities mainly dominated by the modal features of each modality through the neural network; Obtaining the emotion classification result of the target object based on each emotion interaction information.

3. The method according to claim 2, characterized in that, The obtaining the emotion classification result of the target object based on each emotion interaction information includes: Inputting the superposition result of each emotion interaction information into a linear classification layer of the neural network to obtain the emotion classification result of the target object.

4. The method according to claim 2, characterized in that, Before obtaining the emotion classification result of the target object based on each emotion interaction information, the method further includes: Repeating the following steps a preset number of times: For each modality, taking the emotion interaction information of this modality as a new modal feature; Decomposing the multiple modal features into a first tensor and a second tensor through a tensor network; wherein, the first tensor represents a tensor chain including multiple first low-rank tensors respectively corresponding to the multiple modalities, the second tensor represents a tensor chain including multiple second low-rank tensors respectively corresponding to the multiple modalities, and the dimensions of the first tensor and the second tensor are determined based on the number of modalities; Taking the first tensor and the second tensor as the query value and the key value of a neural network based on an attention mechanism respectively; Obtain the kernel tensor corresponding to each modality respectively according to the query value and the key value, and obtain the pooling matrix based on the kernel tensor; wherein, for each modality, obtain the kernel tensor corresponding to this modality according to the first low-rank tensor and the second low-rank tensor corresponding to this modality; perform average pooling operation on the kernel tensor of this modality along the time dimension corresponding to this modality to obtain the pooling matrix corresponding to this modality; For the target modality among the multiple modalities, perform feature fusion on the kernel tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality, and input the fused features into a linear function to obtain the attention representation corresponding to the target modality; Take the modality feature corresponding to each modality as the value of the neural network; Based on the value and the attention representation, obtain the cross-modal emotional interaction information dominated by the modality feature of each modality through the neural network.

5. The method according to claim 1, characterized in that, The step of, for the target modality among the multiple modalities, fusing the kernel tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality, includes: For the target modality among the multiple modalities, perform tensor contraction operation on the kernel tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality to obtain the attention representation corresponding to the target modality.

6. A multi-modal emotion recognition device, characterized in that, Comprising: An acquisition module, configured to acquire the modality features corresponding to multiple modalities of a target object; A construction module, configured to decompose the multiple modality features into a first tensor and a second tensor through a tensor network; wherein, the first tensor represents a tensor chain including multiple first low-rank tensors respectively corresponding to the multiple modalities, the second tensor represents a tensor chain including multiple second low-rank tensors respectively corresponding to the multiple modalities, and the dimensions of the first tensor and the second tensor are determined based on the number of modalities; A first determination module, configured to use the first tensor and the second tensor as the query value and the key value of a neural network based on an attention mechanism respectively; A second determination module, configured to obtain the kernel tensor corresponding to each modality respectively according to the query value and the key value, and obtain the pooling matrix based on the kernel tensor; wherein, for each modality, obtain the kernel tensor corresponding to this modality according to the first low-rank tensor and the second low-rank tensor corresponding to this modality; perform average pooling operation on the kernel tensor of this modality along the time dimension corresponding to this modality to obtain the pooling matrix corresponding to this modality; A third determination module, configured to, for the target modality among the multiple modalities, perform feature fusion on the kernel tensor corresponding to the target modality and the pooling matrices corresponding to the other modalities except the target modality, and input the fused features into a linear function to obtain the attention representation corresponding to the target modality; A classification module, configured to obtain the emotional classification result of the target object based on the attention representation of each modality.

7. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1-5 above.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is run by a processor, it executes the steps of the method according to any one of the above claims 1-5.

Citation Information

Patent Citations

  • Text sentiment classification algorithm based on convolutional neural network and attention mechanism

    CN108664632A

  • Social media sentiment analysis method and system based on tensor fusion network

    CN113064968A