Sentiment Recognition Method, System and Device Based on Multimodal Cross-Attention Network
Through a multimodal cross attention network, combined with visual and text attention models, the problem of underutilizing the influence of low-level features and emotional words in image-text multimodal emotion recognition is solved, and the correction and accuracy of emotional recognition results are achieved.
Patent Information
- Application Number
- CN202310377299.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-04-07
AI Technical Summary
The prior art fails to fully consider the impact of low-level visual features and affective words on emotions in image-text multimodal emotion recognition, and lacks an effective decision-making correction mechanism when multimodal data expresses emotional inconsistency.
A multimodal cross attention network is adopted to achieve the fusion and correction of image and text features through visual attention model, text attention model, image-guided text attention model, and text-guided image attention model, combined with interpretable explicit features and unexplainable implicit features.
The accuracy of multimodal emotion recognition is improved, especially when multimodal data expresses emotional inconsistency, it can effectively correct the emotional recognition results and improve the recognition performance.
Smart Images

Figure CN116434023B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal emotion recognition, and particularly to an emotion recognition method, system and device based on a multimodal cross-attention network. Background Art
[0002] Multimodal emotion recognition is a way to break the data barriers between modalities and fuse various data features for emotion recognition. Multimodal data is becoming increasingly popular on social networking sites (such as Facebook, Twitter, and Flickr). Emotion recognition for large-scale multimodal data helps to better understand people's feelings towards specific events. For example, product managers can understand customers' views on products based on emotion recognition results, thereby conducting product evaluation and improvement. The advantages of multimodal emotion recognition are reflected in the correlation and complementarity of emotion features in various data types, and usually achieve more accurate results than unimodal emotion recognition.
[0003] The correlation between multimodal data can enrich the feature representation of the same emotion unit, the complementarity between multimodal data provides emotional decision-making guidance from different angles, and multimodal data fusion can make up for the defect of one-sided emotional expression of unimodal data.
[0004] Currently, methods for fusing multimodal data based on the correlation and complementarity between multimodal data have attracted wide attention. Yang et al. proposed a multimodal emotion analysis model based on a multi-view (object and scene view) attention network (MVAN), which interactively models the cross-view dependence between images and texts, and uses an ever-updating memory network to obtain the deep semantic features of image texts. Huang et al. proposed an attention-based modality gating network (AMGN), which utilizes the correlation between image and text modalities to extract discriminative features for multimodal emotion analysis. Li et al. compared the shallow features of texts and images through cosine similarity, input the obtained results into the decision layer, and participated in the final emotion decision together with the results of texts and images respectively.
[0005] Multimodal emotion recognition has achieved excellent recognition results by utilizing the correlation and complementarity of modal data, and deep networks and attention mechanisms have also been widely studied in emotion recognition tasks. However, the existing technologies have certain limitations for image-text multimodal emotion recognition. First, most current methods do not fully consider the influence of low-level visual features of images and special emotion words in texts on emotions. Second, in actual scenarios, the emotional expressions of texts and pictures posted by people are often inconsistent or even opposite, but few people notice how to correct emotional decisions when the emotions expressed by multimodal data are inconsistent (the text has a sad emotion while the corresponding picture shows a happy emotion). Summary of the Invention
[0006] The object of the present invention is to provide an emotion recognition method, system and device based on a multi-modal cross-attention network, so as to correct the emotion recognition result and improve the accuracy of the emotion recognition result when the emotions expressed by multi-modal data are inconsistent.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] In a first aspect, the present invention provides an emotion recognition method based on a multi-modal cross-attention network, including:
[0009] Obtain a target image-text pair; the target image-text pair is an image-text pair to be recognized; the target image-text pair includes target image data and target text data corresponding to the target image data;
[0010] Process the target image-text pair to obtain the image feature, text feature, image emotion label vector and text emotion label vector of the target image-text pair; the image feature is a feature obtained by combining the explicit image interpretability feature and the implicit image non-interpretability feature, and the text feature is a feature obtained by combining the explicit text interpretability feature and the implicit text non-interpretability feature; the image emotion label vector is obtained by inputting the image data into a pre-trained ResNet model, and the text emotion label vector is obtained by inputting the text data into a trained Bert model;
[0011] Input the image feature, text feature, image emotion label vector and text emotion label vector of the target image-text pair into a trained multi-modal cross-attention network model to obtain the emotion label of the target image-text pair;
[0012] The multi-modal cross-attention network model includes a visual attention model, a text attention model, an image-guided text attention model, a text-guided image attention model, a feature fusion model and a classifier;
[0013] The visual attention model is used to determine an image weighted feature according to the image feature and the image emotion label vector;
[0014] The text attention model is used to determine a text weighted feature according to the text feature and the text emotion label vector;
[0015] The image-guided text attention model is used to determine an image-guided text weighted feature according to the text feature and the image emotion label vector;
[0016] The attention model for the text-guided image is used to determine the image weighted features guided by the text according to the image features and the text sentiment label vector;
[0017] The feature fusion model is used to fuse the image weighted features output by the visual attention model, the text weighted features output by the text attention model, the image-guided text weighted features output by the attention model for image-guided text, and the text-guided image weighted features output by the attention model for text-guided image to obtain the fused features;
[0018] The classifier is used to classify the corresponding sentiment labels according to the fused features.
[0019] In a second aspect, the present invention provides an emotion recognition system based on a multi-modal cross-attention network, including:
[0020] A target image-text pair acquisition module, configured to acquire a target image-text pair; the target image-text pair is an image-text pair to be recognized; the target image-text pair includes target image data and target text data corresponding to the target image data;
[0021] A target image-text pair processing module, configured to process the target image-text pair to obtain the image features, text features, image sentiment label vector, and text sentiment label vector of the target image-text pair; the image features are features obtained by combining the explicit features of image interpretability and the implicit features of image non-interpretability, and the text features are features obtained by combining the explicit features of text interpretability and the implicit features of text non-interpretability; the image sentiment label vector is obtained by inputting the image data into a pre-trained ResNet model, and the text sentiment label vector is obtained by inputting the text data into a trained Bert model;
[0022] An emotion label determination module, configured to input the image features, text features, image sentiment label vector, and text sentiment label vector of the target image-text pair into a trained multi-modal cross-attention network model to obtain the emotion label of the target image-text pair;
[0023] The multi-modal cross-attention network model includes a visual attention model, a text attention model, an attention model for image-guided text, an attention model for text-guided image, a feature fusion model, and a classifier;
[0024] The visual attention model is used to determine image weighted features according to the image features and the image sentiment label vector;
[0025] The text attention model is used to determine text weighted features according to the text features and the text sentiment label vector;
[0026] The attention model for image-guided text is used to determine text weighted features guided by the image according to text features and image sentiment label vectors;
[0027] The attention model for text-guided image is used to determine image weighted features guided by the text according to image features and text sentiment label vectors;
[0028] The feature fusion model is used to fuse the image weighted features output by the visual attention model, the text weighted features output by the text attention model, the text weighted features guided by the image output by the attention model for image-guided text, and the image weighted features guided by the text output by the attention model for text-guided image to obtain fused features;
[0029] The classifier is used to classify the corresponding sentiment labels according to the fused features.
[0030] In a third aspect, the present invention provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the sentiment recognition method based on a multi-modal cross-attention network according to the first aspect.
[0031] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0032] The purpose of the present invention is to overcome the shortcomings of the existing technical solutions, starting from two aspects: the extraction of modal features and the correction of inconsistent modal sentiment expressions, to improve the performance of multi-modal sentiment recognition. The specific approach is as follows: adopt an intermediate fusion method to fuse multi-modal features, and at the same time use a single-modal attention mechanism and a cross-modal attention mechanism to achieve the purpose of independently and interactively training a sentiment recognizer. Specifically, fuse interpretable explicit features and non-interpretable implicit features as the feature representations of images and texts to enhance relevant sentiment features. Then, two single-modal attention models (i.e., the visual attention model and the text attention model) and two cross-modal attention models (i.e., the attention model for image-guided text and the attention model for text-guided image) are proposed. Among them, the two single-modal attention models use the complementarity of modal data to respectively learn the sentiment features of different modalities, and the two cross-modal attention mechanisms use the correlation of modal data for interactive learning to achieve joint feature extraction. Finally, the four attention models are integrated through intermediate layer fusion to achieve sentiment recognition. Therefore, the present invention can correct the sentiment recognition result and improve the accuracy of the sentiment recognition result when the sentiments expressed by multi-modal data are inconsistent. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0034] Figure 1 It is a schematic flowchart of an emotion recognition method based on a multi-modal cross-attention network provided by an embodiment of the present invention;
[0035] Figure 2 It is a training flowchart of a multi-modal cross-attention network model provided by an embodiment of the present invention;
[0036] Figure 3 is a graph showing the accuracy trend of the NVTD and MVSA-Multiple training sets provided by an embodiment of the present invention; Figure 3(a) is a graph showing the accuracy trend of the NVTD training set provided by an embodiment of the present invention; Figure 3(b) is a graph showing the accuracy trend of the MVSA-Multiple training set provided by an embodiment of the present invention;
[0037] Figure 4 It is a schematic structural diagram of an emotion recognition system based on a multi-modal cross-attention network provided by an embodiment of the present invention. Detailed implementation manners
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0039] The diversity of multi-modal data emotion expression is the main reason for the failure of model emotion recognition, and few people pay attention to the decision correction of multi-modal emotion recognition. The present invention provides an emotion recognition method, system and device based on a multi-modal cross-attention network, which can correct the mistakes of multi-modal emotion recognition by adjusting the modal data fusion method. The present invention not only explores the feature enhancement representation of text and images, but also uses the attention mechanism to guide the model emotion decision. Especially when the multi-modal data emotion expressions are inconsistent, the emotion labels are used to guide the model to make decision correction.
[0040] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0041] MCAN: Multimodal Cross-Attentional Network, a multi-modal cross-attention network, which is the abbreviation of the method proposed in this invention.
[0042] MVAN: Multi-view Attentional Network, a multi-view attention network, which is the prior art related to this invention.
[0043] ResNet: Deep residual network
[0044] Bert: Bidirectional Encoder Representation from Transformers, an encoder of bidirectional Transformers, which is a pre-trained language representation model.
[0045] NVTD: News Visual-Text Dataset
[0046] MVSA-Multiple: A commonly used dataset in the field of image-text sentiment analysis, and its samples are image-text comments collected from Twitter.
[0047] SOTA: State Of The Art, the most advanced technology, which refers to the best / most advanced model in this research task currently.
[0048] Example 1
[0049] A sentiment recognition method based on a multi-modal cross-attention network provided by an embodiment of this invention. First, for image data and text data, different interpretable explicit features are defined and mined, and are used to combine with un-interpretable implicit features; Second, two independent single-modal attention models and two cross-modal attention models in interactive modes are proposed. The two single-modal attention models respectively learn the multi-feature fusion weights in image data and text data. Among them, the visual attention mechanism is used to automatically focus on the emotional region, while the text attention mechanism is used to improve the attention to emotionally colored words. The two cross-modal attention models realize the interactive learning of emotional features in different modalities. Finally, the four attention models are integrated through intermediate layer fusion, and then the final decision of sentiment classification is obtained.
[0050] As Figure 1 shown, a sentiment recognition method based on a multi-modal cross-attention network provided by an embodiment of this invention includes the following steps.
[0051] Step 100: Obtain a target image-text pair; the target image-text pair is an image-text pair to be recognized; the target image-text pair includes target image data and target text data corresponding to the target image data.
[0052] Step 200: Process the target image-text pair to obtain the image feature, text feature, image sentiment label vector, and text sentiment label vector of the target image-text pair; the image feature is a feature obtained by combining the explicit image interpretability feature and the implicit image non-interpretability feature, and the text feature is a feature obtained by combining the explicit text interpretability feature and the implicit text non-interpretability feature; the image sentiment label vector is obtained by inputting the image data into a pre-trained ResNet model, and the text sentiment label vector is obtained by inputting the text data into a pre-trained Bert model.
[0053] Step 300: Input the image feature, text feature, image sentiment label vector, and text sentiment label vector of the target image-text pair into a trained multi-modal cross-attention network model to obtain the sentiment label of the target image-text pair. Preferably, the sentiment label is one of moved, joyful, angry, sad, fearful, and surprised.
[0054] The multi-modal cross-attention network model includes a visual attention model, a text attention model, an image-guided text attention model, a text-guided image attention model, a feature fusion model, and a classifier; the visual attention model is used to determine the image weighted feature according to the image feature and the image sentiment label vector; the text attention model is used to determine the text weighted feature according to the text feature and the text sentiment label vector; the image-guided text attention model is used to determine the image-guided text weighted feature according to the text feature and the image sentiment label vector; the text-guided image attention model is used to determine the text-guided image weighted feature according to the image feature and the text sentiment label vector; the feature fusion model is used to fuse the image weighted feature output by the visual attention model, the text weighted feature output by the text attention model, the image-guided text weighted feature output by the image-guided text attention model, and the text-guided image weighted feature output by the text-guided image attention model to obtain a fused feature; the classifier is used to classify the corresponding sentiment label according to the fused feature.
[0055] As a preferred implementation, as Figure 2 shown, the training process of the multi-modal cross-attention network model is:
[0056] (1) Construct a sample data set; the sample data set includes multiple sample arrays; the sample array includes image features, text features, image sentiment label vectors, text sentiment label vectors, and sample sentiment labels determined by the same sample image-text pair; the image sentiment label vector determined by the sample image-text pair is obtained by inputting the image data into a pre-trained ResNet model, and the text sentiment label vector is obtained by inputting the text data into a pre-trained Bert model.
[0057] (2) Train a multi-modal cross-attention network model according to the sample data set to obtain a trained multi-modal cross-attention network model.
[0058] Further, the construction of the sample data set specifically includes:
[0059] 1) Obtain multiple sample image-text pairs {M, T}. The sample image-text pair includes a sample image and the sample text corresponding to the sample image; where the sample image H, W, and C respectively refer to the length, height, and number of channels of the sample image; the sample text T = {w1, w2,..., w i ,..., w n}, w i refers to the i-th word in the sample text, and n is the length of the sample text.
[0060] 2) Perform sentiment label prediction on all sample image-text pairs to obtain the image sentiment label vector, text sentiment label vector, and sample sentiment label corresponding to each sample image-text pair. Among them, it is represented in the form of a sentiment label vector y ∈ {love, joy, anger, sad, fear, surprise}, where love means touched, joy means happy, anger means angry, sad means sad, fear means frightened, and surprise means surprised.
[0061] 3) Perform feature extraction on all sample image-text pairs to obtain the image feature and text feature corresponding to each sample image-text pair.
[0062] As a preferred implementation, for image data, the feature after combining the interpretable explicit feature and the non-interpretable implicit feature is used as its feature representation.
[0063] The image feature is represented as: m = {m color , m texture , m shape , m cnn}.
[0064] Among them, m color represents the color feature, mtexture represents texture features, m shape Represents shape characteristics, m cnn Represents the convolution feature. m color , m texture , m shape is the explicit feature of image interpretability, m cnn It is an implicit feature of text unexplainability that is automatically mined.
[0065] See below for details:
[0066] Color features are visual features of the surface of an object. They are insensitive to size and direction and are one of the main perceptual features for humans to identify image content. Images are usually represented in RGB color space, but studies have found that HSV color space is more in line with human visual perception and is not affected by lighting and viewing angles. By using color features to label emotional semantics, we mainly consider specific colors and distributions. This paper extracts HS color features that play a significant role in the image to obtain the color features of the image. Among them, h represents hue and s represents saturation.
[0067] Texture features are information containing periodic or structural distribution in a certain area of an image. Its influence on emotional color is not as obvious as color, but texture features contain spatial frequency factors. The difference between image textures has an impact on visual effects and thus on human emotions. The present invention uses three attributes of Tamura texture with intuitive visual meaning to represent texture features and obtains the texture features of the image. Among them, coar represents coarseness, cont represents contrast, and dire represents directionality.
[0068] Shape features are features that characterize the essence of an object, and their representation must be invariant to translation, rotation, and scale. The shape features of an image can also have an impact on the generation of human emotions. The invariant moment method remains unchanged after translation, rotation, and scaling transformations. The present invention selects Hu invariant moments to extract shape features and obtains Among them, φ1-φ7 represent the seven characteristic quantities of Hu invariant moments.
[0069] Select the features extracted by the last convolutional layer before the fully connected layer in the pre-trained ResNet as the convolutional features Specifically: m cnn =conv(M;θ c ). Where conv represents the convolution operation, M represents the input image, and θ cRepresents the parameters in the convolution operation, p represents the number of regions of the image features in this layer, each region has a size of N×N, and q4 represents the number of channels in this layer.
[0070] For the input image M, color features, texture features, shape features, and convolution features in the image are obtained through various extraction techniques. Since the value ranges and dimensions of each type of feature are different, before performing multi-feature fusion, each feature needs to be processed. First, perform different convolution operations on m color , m texture , m shape to obtain the transformed features Specifically as follows:
[0071]
[0072]
[0073]
[0074] Among them, conv represents the convolution operation, and respectively represent the parameters in three convolution operations, and q1, q2, and q3 respectively represent the number of channels of m color ', m texture ', and m shape '.
[0075] Then, splice three explicitly interpretable features of the image and one implicitly interpretable feature of the image to obtain the feature of this image as Specifically as follows:
[0076] R = concat(m color ', m texture ', m shape ', m cnn ).
[0077] Among them, q represents the number of channels of the spliced feature, q = q1 + q2 + q3 + q cnn .
[0078] Furthermore, the specific operations of image feature extraction are as follows:
[0079] Extract color features, texture features, and shape features from the input image as the interpretable explicit features of the image. The present invention uses OpenCV to extract the three types of features. For color features, first convert the input image from the RGB format to the HSV format, and then use the hue-saturation two-dimensional histogram (H-S) method to extract the H-S color features that play a significant role in the HSV-format image. For texture features, first convert the input image from the RGB format to the grayscale format, and then obtain the texture features of the grayscale-format image according to the calculation formula of Tamura texture. For shape features, first convert the input image from the RGB format to the grayscale format, and then obtain the shape features of the grayscale-format image according to the calculation formula of Hu invariant moments. For convolutional features, input the input image into the pre-trained ResNet50, and use the features extracted from the convolutional layer of the pre-trained ResNet50 as the convolutional features of the image. Among them, the input image is the target image or the sample image;
[0080] Since the picture size in the dataset is small, to retain the features of the picture to the greatest extent, select the convolutional features at the end of the first convolutional stage of ResNet50 as the unexplainable implicit features. Since the above features have different physical meanings and sizes, some operations are required for multi-feature fusion. First, input the color features, texture features, and shape features into a neural network containing a convolutional layer, a max pooling layer, and an activation layer respectively. Among them, a BN layer is connected after the convolutional layer to prevent the model from overfitting, and features with the same size as the unexplainable implicit features are obtained for all three types of features. Secondly, splice these features in the channel direction to obtain the joint feature representation of the image.
[0081] As a preferred implementation, for text data, use the features after combining the interpretable explicit features and the unexplainable implicit features as its feature representation.
[0082] The text feature representation is: t = (t adj , t adv , t verb , t noun , t neg , t trans ).
[0083] Among them, t adj represents the adjective feature, t adv represents the adverb feature, t verb represents the verb feature, t noun represents the noun feature, t neg represents the negation word feature, t trans represents the text pre-training feature. t adj , t adv , t verb , tnoun , t neg is an explicit feature of text interpretability, t trans is an implicit feature of text non-interpretability automatically mined. See the following for details:
[0084] Select adjectives, adverbs, verbs, nouns, and negation words in the text as explicit features of text interpretability. These types of feature words can all reflect the sentiment of the text. The word embedding technology can map words to a low-dimensional feature space to become dense word vectors. The present invention uses the word embedding technology to embed words into a vector space to obtain the word vector t of the input text = Specifically as follows:
[0085] t adj = W em w adj .
[0086] Among them, dim represents the word embedding dimension, W em represents the embedding matrix, and w adj represents the adjective in the text. The embedding methods of the remaining words are similar to this.
[0087] Use the encoder of the pre-trained bidirectional Transformer to extract the high-order feature vector of the input text as the implicit feature of text non-interpretability, and obtain Specifically as follows:
[0088] t trans = Transformer(T).
[0089] Among them, T represents the input text, and trans represents the length of the text.
[0090] For the given text T, concatenate the explicit feature of text interpretability and the implicit feature of text non-interpretability to obtain the feature of this text as Specifically as follows:
[0091] S = concat(t adj , t adv , t verb , t noun , t neg , t trans ).
[0092] Among them, len = x + trans represents the length of the concatenated feature, and x represents the number of adjectives, adverbs, verbs, nouns, and negation words in a piece of text.
[0093] The specific operation of text feature extraction is as follows:
[0094] The input text is fed into the Bert model as a whole, and the features of the last layer before the classifier are taken as the text's unexplainable implicit features. Adjectives, adverbs, verbs, nouns, and negation words are extracted from the input text as the text's explainable explicit feature words. The extracted text's explainable explicit feature words are input into the pre-trained Bert model to obtain their feature representations. The 768-dimensional tensor obtained after average pooling is regarded as the explainable explicit feature of the input text. For text features, both the text's explainable explicit feature and the text's unexplainable implicit feature are 768-dimensional tensors. After concatenation, the joint representation of the text features is obtained.
[0095] As a preferred implementation, the key technology of the present invention is the multi-modal cross-attention network model, which is composed of two independent single-modal attention models (a visual attention model and a text attention model) and two cross-modal attention models (an image-guided text attention model and a text-guided image attention model). Among them, the visual attention model is used to automatically focus on the emotional region, while the text attention model is used to improve the attention to emotionally colored words. The cross-modal attention module automatically identifies the contribution degree of image visual features and text semantic features to emotions. Especially when the emotions expressed in the image and text are inconsistent, a larger weight is given to the image features or text features that conform to the true emotion to achieve the correction of emotional decision-making.
[0096] 1) Visual attention model
[0097] The importance of the four features of the above-mentioned extracted image is different. The visual attention model uses the image emotion label vector as a guide to dynamically focus attention on specific image regions to extract the visual features most relevant to emotions.
[0098] For the input image M, use the ResNet model to obtain its emotion label After looking up the emotion embedding table for the emotion label, the image emotion label vector is obtained where h represents the number of emotion categories.
[0099] For the input image M, input the image feature R and the image emotion label vector e img into the visual attention model, where R = {r1, r2,..., r i ,... r p} is used as the query vector and value vector, and the image emotion label vector e img is used as the key vector. According to the correlation between each image feature region r i and emotion, attention is allocated to each image feature region. First, use the attention function to obtain the attention score Secondly, use the softmax function to calculate the attention weight The finally obtained output of attention, namely the image weighted feature The whole process is as follows:
[0100]
[0101]
[0102]
[0103] where is a learnable parameter vector, are all learnable parameter matrices.
[0104] 2) Text attention model
[0105] Similar to image regions, some words in the text are usually more important for sentiment expression than others. Recently, text attention mechanisms have been proven to be beneficial for many natural language processing related tasks. The text attention model uses the text sentiment label vector as a guide to dynamically focus attention on specific words to extract the text features most relevant to sentiment.
[0106] For the input text T, use the Bert model to obtain its sentiment label After looking up the sentiment embedding table for the sentiment label, obtain the text sentiment label vector
[0107] Input the text features S = {s1, s2, …, s j , …, s len} and the text sentiment label vector e text into the text attention model, where S serves as the query vector and value vector, and e text serves as the key vector. According to the relevance of each word to the sentiment, allocate attention to each word s j . First, use the attention function to obtain the attention scores Secondly, use the softmax function to calculate the attention weights Finally, obtain the output of attention, namely the text weighted feature The whole process is as follows:
[0108]
[0109]
[0110]
[0111] where is a learnable parameter vector, are all learnable parameter matrices.
[0112] 3) Attention Model for Image-Guided Text
[0113] The attention model for image-guided text mainly learns text features that contribute more to emotions through image sentiment labels.
[0114] Taking the text feature S as the query vector and value vector, and the image sentiment label vector e img as the key vector and inputting them into the model, attention is assigned to each word s according to the importance of each word to the emotion j . First, the attention function is used to obtain the attention scores Secondly, the softmax function is used to calculate the attention weights Finally, the output of the attention, that is, the text weighted feature guided by the image, is obtained The whole process is as follows:
[0115]
[0116]
[0117]
[0118] where is a learnable parameter vector, are all learnable parameter matrices.
[0119] 4) Attention Model for Text-Guided Image
[0120] The attention model for text-guided image mainly learns image features that contribute more to emotions through text sentiment labels.
[0121] Taking the fused image feature R as the query vector and value vector, and the text sentiment label vector e text as the key vector and inputting them into the model, attention is assigned to each region r according to the importance of each image region to the emotion classification i . First, the attention function is used to obtain the attention scores Secondly, the softmax function is used to calculate the attention weights Finally, the output of the attention, that is, the image weighted feature guided by the text, is obtained The whole process is as follows:
[0122]
[0123]
[0124]
[0125] where is a learnable parameter vector, are all learnable parameter matrices.
[0126] 5) Feature fusion model
[0127] The present invention uses an intermediate fusion method to fuse the image weighted feature U, text weighted feature V, text weighted feature K guided by the image, and image weighted feature Z guided by the text obtained above. First, the features obtained by the above 4 attention mechanisms are transformed to the same dimension, and then they are fused in a per-pixel bitwise addition manner to obtain the fused feature The fusion process is as follows:
[0128]
[0129] In the formula represents bitwise addition, are all learnable parameter matrices.
[0130] 6) Classifier
[0131] The fused feature H obtained above is input into a softmax classifier after convolution and fully connected operations to obtain the final sentiment label of the image-text pair
[0132]
[0133]
[0134] In the formula, fc and conv respectively represent fully connected and convolution calculations, θ c , and θ s respectively represent the parameters of the convolution operation, fully connected operation, and softmax operation.
[0135] As a preferred embodiment, the present invention uses a cross-entropy loss function to train the classifier, and the model loss function can be expressed as:
[0136]
[0137] In the formula represents the predicted sentiment label, and y represents the true sample sentiment label.
[0138] The present invention fuses various single-modal features and multi-modal features obtained by the attention mechanism to achieve emotional semantic interaction between modalities. During the process of minimizing the loss, the model is continuously optimized.
[0139] The following uses a specific example to illustrate the technical solution protected by the present invention.
[0140] The method proposed by the present invention performs multi-modal sentiment classification on the NVTD dataset and multi-modal polarity classification on the MVSA-Multiple dataset.
[0141] 1) NVTD News Dataset
[0142] A new dataset called News Visual-Text Dataset (NVTD) is created by collecting news images and texts from websites. The facial expression regions in the images are selected and paired with the corresponding news texts to form text-image pairs. The sentiment labels of the text-image pairs are obtained based on manual scoring and are divided into six categories: "touched", "joyful", "angry", "sad", "fearful", and "surprised". The sample statistics of the dataset are shown in Table 1.
[0143] Table 1 NVTD Data Distribution Table
[0144]
[0145] 2) MVSA-multiple Dataset
[0146] MVSA-Multiple is a commonly used dataset in the field of image-text sentiment analysis. Its samples are image-text comments collected from Twitter, and the sentiment polarity labels can be divided into three categories: "positive", "neutral", and "negative". The sample statistics of the dataset are shown in Table 2.
[0147] Table 2 MVSA-Multiple Data Distribution Table
[0148]
[0149] The parameter selection and detailed description of the model training process are as follows:
[0150] Since the news data is scarce, the NVTD news dataset is divided into a training set and a test set at a ratio of 8:2, and the MVSA-multiple dataset is divided into a training set, a validation set, and a test set at a ratio of 8:1:1. Due to the problem of unbalanced data distribution, a weighted loss function and an early stopping method are used for training. Adam is used as the optimizer method with a learning rate of 0.001. The batch size of the NVTD news dataset is 16, and the batch size of the MVSA-multiple dataset is 32. The evaluation metrics used in the experiment are Accuracy and F1-score.
[0151] The experimental results of the method proposed by the present invention and the baseline method are shown in Table 3, where MCAN represents the method proposed by the present invention.
[0152] Table 3 Results of Each Model on Different Datasets
[0153]
[0154] By observation, the following information can be obtained:
[0155] On the NVTD dataset, the performance of the model MCAN of the present invention is superior to that of the baseline model. For the text sentiment recognition model, the performance of Bert exceeds that of BiLSTM by 4.15% in terms of the accuracy metric. For the image sentiment recognition model, the performance of ResNet50 is superior to that of SimpleNet. And the accuracy of image sentiment recognition is significantly lower than that of text sentiment recognition, mainly for two reasons: one is that the emotional features of images are more ambiguous compared to text; the other is that in the news dataset, some pictures do not have people, so the emotions cannot be distinguished. MCAN only improves by 1.08% compared to Bert in terms of the F1 value, because the image quality of this dataset is not high enough, the number of samples is small, and the role of the text modality is greater.
[0156] On the MVSA-Multiple dataset, the performance of the proposed model is also superior to that of the baseline model. For the text sentiment recognition model, the performance of Bert exceeds that of BiLSTM by 1.6% in terms of the accuracy metric. For the image sentiment recognition model, the performance of ResNet50 is comparable to that of SimpleNet. For the multi-modal sentiment recognition model, MCAN has better performance, exceeding HSAN by nearly 3% and DSN by 1.3% in terms of the F1 value. MCAN improves by 5.67% compared to Bert in terms of the accuracy, which also illustrates the effectiveness of multi-modal data fusion.
[0157] By comparing the performance of each model on the NVTD and MVSA-Multiple datasets, it is found that the performance of the text sentiment recognition models BiLSTM and Bert on the MVSA-Multiple dataset is not as good as that on the news dataset. This is mainly because the text of the news dataset is more formal than Twitter comments, and the emotional expression is more obvious. The performance of the image sentiment recognition models SimpleNet and ResNet50 on the news dataset is not as good as that on the MVSA-Multiple news dataset. The reason is that people mainly use Twitter comments to express emotions, while the pictures in news data are often just objective related event pictures or people pictures, and the emotional purpose is not strong.
[0158] Overall, first of all, on both datasets, the performance of the multimodal sentiment recognition model is better than that of the unimodal sentiment recognition model. This shows that the fusion of multimodal data helps with the sentiment recognition task. Secondly, the model MCAN of the present invention is superior to other multimodal models in terms of accuracy and F1 value. HSAN only considers a single image or text feature, ignoring the influence of text part-of-speech features or image texture features on sentiment. MVAN only considers image information to assist in learning text features, while ignoring the auxiliary role of text in learning image features. Obviously, the model we proposed can make full use of the features of images and text, achieve sentiment decision correction, and improve sentiment recognition performance.
[0159] The accuracy trend of the method proposed by the present invention in the training sets of the two datasets is shown in Figure 3. The early stopping method makes the number of training times of each model inconsistent, so the accuracy of 10 iterations is selected to calculate the training trend after the alignment operation. On the NVTD dataset, the training accuracy of the two image models reaches about 44%, mainly because the sentiment features of some images are implicit and ambiguous. Due to the formality of news data, the training accuracy of the two text models reaches over 99%. Relying on the dominant role of the text content, MCAN achieves a fusion of over 99%. On the MVSA-Multiple dataset, the two visual models reach about 60%. The pictures on Twitter express emotions more purposefully than news pictures. However, due to the abstractness of images, image sentiment recognition is still more challenging than text. The accuracy of Bert is better than that of BiLSTM because Bert has greater learning ability on large datasets. Through cross-modal interaction learning, the complementarity of image and text features enables MCAN to reach over 97%. There are differences between the two datasets, so the precision curves of the five models do not converge at the same point. Overall, the precision of all models becomes stable after multiple iterations.
[0160] The method proposed by the present invention can make reasonable decisions when the image and text express inconsistent sentiment information. Table 4 shows examples of inconsistent image-text sentiment expressions in the MVSA-Multiple dataset. In the first example, the sentiment is dominated by the image. Although the sentiment word "crowded" in the text is likely to predict the wrong emotion, MCAN can correctly recognize the sentiment of the image-text pair based on the vivid atmosphere and colorful lights in the image. In contrast, in the second and third examples, the sentiment is dominated by the text. The flowing road and black background in the visual information seem to convey negative emotions, but the combination of visual content and text content enables the classifier to predict the correct sentiment. The MCAN method proposed by the present invention can flexibly assign weights to visual and text information and predict more accurate sentiment than the baseline.
[0161] Table 4 Examples of inconsistent image-text sentiment expressions in the MVSA-Multiple dataset
[0162]
[0163] Embodiment 2
[0164] In order to execute the method corresponding to the above Embodiment 1 to achieve the corresponding functions and technical effects, the following provides an emotion recognition system based on a multi-modal cross-attention network.
[0165] As Figure 4 shown, an emotion recognition system based on a multi-modal cross-attention network provided by an embodiment of the present invention includes:
[0166] A target image-text pair acquisition module 1, configured to acquire a target image-text pair; the target image-text pair is an image-text pair to be recognized; the target image-text pair includes target image data and target text data corresponding to the target image data.
[0167] A target image-text pair processing module 2, configured to process the target image-text pair to obtain an image feature, a text feature, an image emotion label vector, and a text emotion label vector of the target image-text pair; the image feature is a feature obtained by combining an explicitly interpretable image feature and an implicitly interpretable image feature, and the text feature is a feature obtained by combining an explicitly interpretable text feature and an implicitly interpretable text feature; the image emotion label vector is obtained by inputting the image data into a pre-trained ResNet model, and the text emotion label vector is obtained by inputting the text data into a trained Bert model.
[0168] An emotion label determination module 3, configured to input the image feature, the text feature, the image emotion label vector, and the text emotion label vector of the target image-text pair into a trained multi-modal cross-attention network model to obtain an emotion label of the target image-text pair.
[0169] The multi-modal cross-attention network model includes a visual attention model, a text attention model, an image-guided text attention model, a text-guided image attention model, a feature fusion model, and a classifier. The visual attention model is used to determine an image weighted feature based on an image feature and an image sentiment label vector. The text attention model is used to determine a text weighted feature based on a text feature and a text sentiment label vector; the image-guided text attention model is used to determine an image-guided text weighted feature based on a text feature and an image sentiment label vector; the text-guided image attention model is used to determine a text-guided image weighted feature based on an image feature and a text sentiment label vector; the feature fusion model is used to fuse the image weighted feature output by the visual attention model, the text weighted feature output by the text attention model, the image-guided text weighted feature output by the image-guided text attention model, and the text-guided image weighted feature output by the text-guided image attention model to obtain a fused feature; the classifier is used to classify the corresponding sentiment label according to the fused feature.
[0170] Embodiment III
[0171] An embodiment of the present invention provides an electronic device including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the emotion recognition method based on the multi-modal cross-attention network in Embodiment I.
[0172] Optionally, the above electronic device may be a server.
[0173] In addition, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the emotion recognition method based on the multi-modal cross-attention network in Embodiment I is implemented.
[0174] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for emotion recognition based on a multi-modal cross-attention network, characterized in that, Including: Obtain a target image-text pair; the target image-text pair is an image-text pair to be recognized; the target image-text pair includes target image data and target text data corresponding to the target image data; Process the target image-text pair to obtain the image feature, text feature, image sentiment label vector, and text sentiment label vector of the target image-text pair; the image feature is a feature obtained by combining the explicitly interpretable image feature and the implicitly interpretable image feature; the explicitly interpretable image feature includes color feature, texture feature, and shape feature, and the implicitly interpretable image feature includes convolutional feature; the text feature is a feature obtained by combining the explicitly interpretable text feature and the implicitly interpretable text feature; the explicitly interpretable text feature includes adjective feature, adverb feature, verb feature, noun feature, and negation word feature; The implicitly interpretable text feature includes text pre-training feature; the image sentiment label vector is obtained by inputting the image data into a pre-trained ResNet model, and the text sentiment label vector is obtained by inputting the text data into a trained Bert model; Input the image feature, text feature, image sentiment label vector, and text sentiment label vector of the target image-text pair into a trained multi-modal cross-attention network model to obtain the sentiment label of the target image-text pair; The multi-modal cross-attention network model includes a visual attention model, a text attention model, an image-guided text attention model, a text-guided image attention model, a feature fusion model, and a classifier; The visual attention model is used to determine the image weighted feature according to the image feature and the image sentiment label vector; The text attention model is used to determine the text weighted feature according to the text feature and the text sentiment label vector; The image-guided text attention model is used to determine the image-guided text weighted feature according to the text feature and the image sentiment label vector; The text-guided image attention model is used to determine the text-guided image weighted feature according to the image feature and the text sentiment label vector; The feature fusion model is used to fuse the image weighted feature output by the visual attention model, the text weighted feature output by the text attention model, the image-guided text weighted feature output by the image-guided text attention model, and the text-guided image weighted feature output by the text-guided image attention model to obtain a fused feature; The classifier is used to classify the corresponding sentiment label according to the fused feature.
2. The method for emotion recognition based on a multi-modal cross-attention network according to claim 1, characterized in that, The training process of the multi-modal cross-attention network model is as follows: Construct a sample data set; the sample data set includes multiple sample arrays; the sample array includes the image feature, text feature, image sentiment label vector, text sentiment label vector, and sample sentiment label determined by the same sample image-text pair; Train the multi-modal cross-attention network model according to the sample data set to obtain a trained multi-modal cross-attention network model.
3. The method for emotion recognition based on a multi-modal cross-attention network according to claim 1 or 2, characterized in that, The image features are features obtained by splicing color features, texture features, shape features, and convolutional features; the convolutional features are implicitly uninterpreted features of the image automatically mined.
4. The method for emotion recognition based on a multi-modal cross-attention network according to claim 1 or 2, characterized in that, The text features are features obtained by splicing adjective features, adverb features, verb features, noun features, negation word features, and text pre-training features; the text pre-training features are implicitly uninterpreted features of the text automatically mined.
5. The method for emotion recognition based on a multi-modal cross-attention network according to claim 3, characterized in that, The process of determining the image features is as follows: Convert the input image from RGB format to HSV format, and use the hue-saturation two-dimensional histogram method to determine the H-S color features of the HSV format image. Convert the input image from RGB format to grayscale format, and obtain the texture features of the grayscale format image according to the calculation formula of Tamura texture. Use the calculation formula of Hu invariant moment to obtain the shape features of the grayscale format image. Send the input image to the pre-trained ResNet50, and use the features extracted from the convolutional layer of the pre-trained ResNet50 as the convolutional features of the image. Input the H-S color features, texture features, and shape features into a neural network containing a convolutional layer, a max-pooling layer, and an activation layer to obtain H-S color features, texture features, and shape features with the same size as the convolutional features. Splice the convolutional features, H-S color features, texture features, and shape features with the same size as the convolutional features in the channel direction to obtain the image features.
6. The method for emotion recognition based on a multi-modal cross-attention network according to claim 4, characterized in that, The process of determining the text features is as follows: Use word embedding technology to determine the adjective features, adverb features, verb features, noun features, and negation word features of the input text. Use the pre-trained Bert encoder to extract the text pre-training features of the input text. Splice the adjective features, adverb features, verb features, noun features, negation word features, and text pre-training features to obtain the text features.
7. The method for emotion recognition based on a multi-modal cross-attention network according to claim 1, characterized in that,In terms of fusing the image weighted features output by the visual attention model, the text weighted features output by the text attention model, the text weighted features guided by the image output by the image-guided text attention model, and the image weighted features guided by the text output by the text-guided image attention model to obtain the fusion features, the feature fusion model is further used for: Transform the image weighted features output by the visual attention model, the text weighted features output by the text attention model, the text weighted features guided by the image output by the image-guided text attention model, and the image weighted features guided by the text output by the text-guided image attention model to the same dimension, and then fuse the features of the same dimension in a pixel-by-pixel bitwise addition manner to obtain the fusion features.
8. The emotional recognition method based on a multi-modal cross-attention network according to claim 1, characterized in that, During the training process of the classifier, the loss function of the classifier is the cross-entropy loss function.
9. An emotional recognition system based on a multi-modal cross-attention network, characterized in that, Including: A target image-text pair acquisition module for acquiring a target image-text pair; the target image-text pair is an image-text pair to be recognized; the target image-text pair includes target image data and target text data corresponding to the target image data. A target image-text pair processing module for processing the target image-text pair to obtain the image features, text features, image sentiment label vector, and text sentiment label vector of the target image-text pair; the image features are features obtained by combining image interpretable explicit features and image non-interpretable implicit features; the image interpretable display features include color features, texture features, and shape features, and the image non-interpretable implicit features include convolutional features; the text features are features obtained by combining text interpretable explicit features and text non-interpretable implicit features; the text interpretable explicit features include adjective features, adverb features, verb features, noun features, and negation word features; The text non-interpretable implicit features include text pre-training features; the image sentiment label vector is obtained by inputting image data into a pre-trained ResNet model, and the text sentiment label vector is obtained by inputting text data into a trained Bert model; A sentiment label determination module for inputting the image features, text features, image sentiment label vector, and text sentiment label vector of the target image-text pair into a trained multi-modal cross-attention network model to obtain the sentiment label of the target image-text pair; The multi-modal cross-attention network model includes a visual attention model, a text attention model, an image-guided text attention model, a text-guided image attention model, a feature fusion model, and a classifier; The visual attention model is used to determine image weighted features according to image features and image sentiment label vectors; The text attention model is used to determine text weighted features according to text features and text sentiment label vectors; The image-guided text attention model is used to determine image-guided text weighted features according to text features and image sentiment label vectors; The text-guided image attention model is used to determine text-guided image weighted features according to image features and text sentiment label vectors; The feature fusion model is used to fuse the image weighted features output by the visual attention model, the text weighted features output by the text attention model, the image-guided text weighted features output by the image-guided text attention model, and the text-guided image weighted features output by the text-guided image attention model to obtain fused features; The classifier is used to classify the corresponding sentiment label according to the fused features.
10. An electronic device, characterized in that, It includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the sentiment recognition method based on a multi-modal cross-attention network according to any one of claims 1 to 8.
Citation Information
Patent Citations
Target-oriented multi-modal sentiment classification method
CN113065577A
Cross-modal image-text mutual indexing method based on self-attention reasoning
CN114461821A