Emotion prediction method and device, equipment, storage medium and computer program product

By generating and fusion of emotional vector data of pictures and text in graphic data, the problem of ignoring emotional alignment between images and text in the prior art is solved, and a higher accuracy of emotional recognition of graphic data is achieved.

CN120047729APending Publication Date: 2025-05-27CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510096528.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When identifying the emotional tendency of graphic and text data, the prior art ignores the emotional alignment between the image and the text, resulting in inaccurate emotional recognition.

Method used

By determining the emotional probability distribution data of pictures and text in the graphic data, the emotional vector data of pictures and text are generated and fused, the emotional feature vector data is obtained to improve the accuracy of emotion recognition.

Benefits of technology

Through the emotional alignment of images and text, the emotions of graphic and text data can be more accurately identified and the accuracy of emotional recognition can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047729A_ABST
    Figure CN120047729A_ABST
Patent Text Reader

Abstract

The invention provides an emotion prediction method and device, equipment, a storage medium and a computer program product, and the method comprises the steps: determining picture emotion probability distribution data corresponding to picture data in first image-text data, and determining text emotion probability distribution data corresponding to text data in the first image-text data; determining first text emotion vector data according to the picture emotion probability distribution data; determining first picture emotion vector data according to the text emotion probability distribution data; fusing the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data; predicting the emotion in the first image-text data according to the emotion feature vector data; the accuracy of emotion recognition of image-text data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly to an emotion prediction method, apparatus, device, storage medium, and computer program product. Background Art

[0002] With the booming rise of self-media and social networks, people are increasingly inclined to express personal opinions and views on social media platforms, and these expressions show diverse characteristics, including forms with both pictures and texts. In the process of identifying the emotional tendency of picture-text data, there may be problems that both the text and the emotion contain their own independent emotional tendencies. Furthermore, when obtaining image and text features separately, the emotional alignment between the image and the text is ignored, resulting in inaccurate emotional recognition of picture-text data. Summary of the Invention

[0003] Embodiments of the present application provide an emotion prediction method, apparatus, device, storage medium, and computer program product, which can improve the accuracy of emotional recognition of picture-text data.

[0004] To achieve the above object, the technical solution of the embodiments of the present application is implemented as follows:

[0005] In a first aspect, the present application provides an emotion prediction method, and the method includes:

[0006] Determine the picture emotion probability distribution data corresponding to the picture data in the first picture-text data, and determine the text emotion probability distribution data corresponding to the text data in the first picture-text data;

[0007] Determine the first text emotion vector data according to the picture emotion probability distribution data; determine the first picture emotion vector data according to the text emotion probability distribution data;

[0008] Fuse the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data;

[0009] Predict the emotion in the first picture-text data according to the emotion feature vector data.

[0010] In a second aspect, the present application proposes an emotion prediction apparatus, and the apparatus includes:

[0011] A determination unit, configured to determine the picture emotion probability distribution data corresponding to the picture data in the first picture-text data, and determine the text emotion probability distribution data corresponding to the text data in the first picture-text data;

[0012] The determining unit is further configured to determine first text sentiment vector data according to the picture sentiment probability distribution data; and determine first picture sentiment vector data according to the text sentiment probability distribution data.

[0013] The fusion unit is configured to fuse the first text sentiment vector data and the first picture sentiment vector data to obtain sentiment feature vector data.

[0014] The prediction unit is configured to predict the sentiment in the first text-image data according to the sentiment feature vector data.

[0015] In a third aspect, the present application provides an emotion prediction device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in any one of the above are implemented.

[0016] In a fourth aspect, the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0017] In a fifth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0018] The present application provides an emotion prediction method, apparatus, device, storage medium, and computer program product. The method includes: determining picture emotion probability distribution data corresponding to picture data in first graphic and text data, and determining text emotion probability distribution data corresponding to text data in the first graphic and text data; determining first text emotion vector data according to the picture emotion probability distribution data; determining first picture emotion vector data according to the text emotion probability distribution data; fusing the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data; and predicting the emotion in the first graphic and text data according to the emotion feature vector data. By adopting the above implementation solution, by determining the picture probability distribution data corresponding to the picture data in the first graphic and text data, and determining the first text emotion vector data according to the picture probability distribution data; determining the text probability distribution data corresponding to the text data in the first graphic and text data, and determining the first picture emotion vector data according to the text probability distribution data; fusing the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data, and predicting the emotion in the first graphic and text data according to the emotion feature vector data, using the picture probability distribution data to determine the first text emotion vector data, and using the text probability distribution data to determine the first picture emotion vector data, the emotion information corresponding to the picture and the text respectively can be better obtained. By fusing the first text emotion vector and the first picture emotion vector, the emotion alignment relationship between the text emotion and the picture emotion can be obtained, thereby improving the accuracy of emotion recognition for graphic and text data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 FIG. is a schematic diagram of an exemplary graphic and text data provided by an embodiment of the present application;

[0020] Figure 2 FIG. is a schematic flowchart of an exemplary multi-modal emotion analysis provided by an embodiment of the present application;

[0021] Figure 3 FIG. is a flowchart example diagram of an emotion prediction method provided by an embodiment of the present application;

[0022] Figure 4 FIG. is a schematic diagram of an exemplary model architecture provided by an embodiment of the present application;

[0023] Figure 5 FIG. is a schematic diagram of an exemplary visual and text feature extraction module provided by an embodiment of the present application;

[0024] Figure 6 FIG. is a schematic diagram of an exemplary user emotion feature pre-coding module provided by an embodiment of the present application;

[0025] Figure 7A schematic structural diagram of an exemplary graphic and text disambiguation module provided by an embodiment of the present application;

[0026] Figure 8 A schematic structural diagram of an exemplary disambiguation unit provided by an embodiment of the present application;

[0027] Figure 9 A schematic structural diagram of an emotion prediction device provided by an embodiment of the present application;

[0028] Figure 10 A schematic structural diagram of an emotion prediction device provided by an embodiment of the present application. Detailed implementation manners

[0029] In order to be able to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and illustration, and are not used to limit the embodiments of the present application.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0031] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. It should also be noted that the terms "first / second" etc. involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" etc. can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0032] Nowadays, with the booming rise of self-media and social networks, people are increasingly inclined to express their personal opinions and views on social media platforms, and these expressions show diverse characteristics, including forms such as pictures and texts. A huge amount of data floods into the network platform every day. These data come from different social media platforms, and many of them are presented in the form of a combination of text and images, forming a huge multi-modal data set. This multi-modal data set contains various emotional tendencies, such as positive, negative, positive and negative, etc. Analyzing the emotional tendency of multi-modal data helps to deeply understand people's views and opinions on objective things and the development of events, and plays a key role in public opinion analysis. Therefore, one of the important challenges currently faced in the field of emotion analysis is to efficiently identify the emotional attributes of multi-modal data on social media platforms.

[0033] Identifying the emotional tendency of graphic and text data on social platforms mainly faces two problems. One is the problem of users' emotional perspectives. Just as "one thousand readers have one thousand Hamlets", for the same thing, different people's emotional reactions to it may be completely different. That is, from different people's perspectives, the same content may have different emotional tendencies. For example, for the matter of "only half a glass of water in the desert", an optimistic person might think: "Great, there's still half a glass of water to drink. It should be enough for me to walk quite a long way, and maybe even get out of the desert." While a pessimistic person might say: "Terrible, why is there only half a glass of water left? How can I possibly get out of the desert?" Therefore, various objective factors such as users' gender, age, occupation, and the environment they are in will affect their understanding of things, thus generating different emotional tendencies. Therefore, for the content on social platforms, it is necessary to consider the situation of "each person has a different face" and comprehensively judge the emotional tendency based on user characteristics.

[0034] The second problem faced by the identification of the emotional tendency of graphic and text data on social platforms is that both the text and the image contain their own independent emotional tendencies, and although the two emotions are different from each other, they complement each other. Figure 1 A schematic diagram of an exemplary graphic and text data provided for the embodiments of this application; as Figure 1 shown, Figure 1 In (a), the text is "In the Cairo leg of the 2023 UCI Track Cycling Nations Cup held in Cairo, the Chinese team won the gold medal in the women's team sprint". The emotion conveyed by the text is positive, and the image is of the athletes raising their arms in celebration. The emotion conveyed by the image is also positive. Therefore, both the image and the text clearly show that the emotion conveyed by this post is positive, and the emotional attributes expressed by both are consistent; Figure 1 In (b), the image is of flowers blooming by the street, and the text is "Artificial Intelligence (AI) translation, please perk up quickly. Pure manual translation is really unbearable". The graphic and text are irrelevant, and there is actually no relationship between what the picture and the text express. Therefore, there is no association of emotional attributes between the two; Figure 1The text in (c) is "The weather is really nice today. It's suitable to go out and buy an ice cream." If only looking at the emotional attribute of the text, it is positive. However, the image is a thunderstorm. When it is paired with a picture of a thunderstorm, this may indicate that the poster has an ironic tone, and the emotional tendency of the text and image is in an opposing conflict. In multi-modal data of text and images, the information between the two modalities of text and pictures can often complement each other, which helps to better understand the emotion. Compared with the single-modal data of text or images, multi-modal data contains more comprehensive and rich information, and can better express and reveal the true emotional tendency of the poster. Therefore, how to fuse the emotional attributes of multi-modal data so that multi-modal data can help improve the emotion recognition effect rather than have a negative impact requires making full use of the alignment information between different modal data to effectively model the association between different modal data.

[0035] Generally, multi-modal emotion analysis includes two basic steps: processing the single-modal data separately, and fusing the processed multi-modal data. Neither of these two steps can be omitted. If the single-modal data cannot be effectively processed, it will have a negative impact on the emotion analysis results of multiple modalities; while poor performance of multi-modal fusion will disrupt the interaction of multi-modal data and affect the recognition of the overall emotion. Specifically, Figure 2 is a schematic flow diagram of an exemplary multi-modal emotion analysis provided by an embodiment of this application; as Figure 2 shown, given an image-text pair, the system first extracts visual features and text features through a visual encoder and a text encoder respectively. Then the multi-modal fusion module receives the visual and text features to generate an interactive representation of the modalities, and finally inputs it into the emotion classifier to generate the final emotion label.

[0036] However, in most of the related technical solutions, single-modal pre-trained models (such as Vision Transformer (ViT) for images and Bidirectional Encoder Representations from Transformers (BERT) for text) are first used as visual encoders and text encoders to obtain visual and text features respectively. There is no modality interaction between images and text in the independent pre-training stage, so the emotional alignment between images and text is ignored when expressing model features. This will have a negative impact on the final multi-modal emotion recognition; moreover, the encoder structure adopted by the model leads to a relatively simple multi-modal fusion method, which is mainly divided into two types: one is to combine visual and text encoding features together and then input them into a single Transformer block to fuse multi-modal inputs by combining attention; the other is not to combine them together, but to independently input them into two different Transformer blocks to achieve cross-modal interaction through cross-attention. Then only a single emotion classifier produces the final output. This single target task cannot fully learn the emotional features of multi-modal data.

[0037] Based on this, an embodiment of the present application provides an emotion prediction method. Figure 3 It is a flowchart example of an emotion prediction method provided by an embodiment of the present application; as Figure 3 shown, the method includes:

[0038] S301. Determine the picture emotion probability distribution data corresponding to the picture data in the first picture-text data, and determine the text emotion probability distribution data corresponding to the text data in the first picture-text data.

[0039] It should be noted that the first picture-text data can be understood as including both picture data and text data. Its source can be determined according to the actual situation and is not limited here. The specific content of the picture data and the text data can be determined according to the actual situation and is not limited here; among them, the emotion corresponding to the picture data and the emotion corresponding to the text data can match or not match.

[0040] In the embodiment of the present application, the process of determining the picture emotion probability distribution data corresponding to the picture data in the first picture-text data and determining the text emotion probability distribution data corresponding to the text data in the first picture-text data specifically includes: obtaining the picture emotion feature encoding data corresponding to the picture data, and obtaining the text emotion feature encoding data corresponding to the text data; inputting the picture emotion feature encoding data into the decoder to obtain the picture emotion probability distribution data; inputting the text emotion feature encoding data into the decoder to obtain the text emotion probability distribution data.

[0041] In the embodiments of the present application, the picture emotion feature encoding data can be understood as the data obtained after pre-encoding the emotion features in the picture data. The process of obtaining the picture emotion feature encoding data corresponding to the picture data specifically includes: using a convolutional neural network to extract features from the picture data to obtain picture feature data; inputting the picture feature data into an encoder to obtain the picture emotion feature encoding data.

[0042] It should be noted that the convolutional neural network can be any convolutional neural network. As an example, the convolutional neural network can be a Region-based Fast Convolutional Neural Networks (Faster R-CNN). Using a convolutional neural network to extract features from the picture data to obtain picture feature data can be exemplified as using Faster R-CNN to extract visual features. Specifically, Faster R-CNN is used to extract the 48 regions with the highest confidence values from the input image, and at the same time, the semantic category distribution of each region is retained and applied to the subsequent "mask region modeling" task. For the retained regions, their visual features use the mean pooling convolutional features processed by Faster R-CNN. Let R = {r 1 ,..., r 48} represent the visual features, where r i ∈ R 2048 represents the visual feature of the i-th region. To be consistent with the representation of the text features, the obtained visual features are fed into a linear transformation layer to project the original 2048-dimensional vector into a d-dimensional vector, denoted as V ∈ R d×48 , where d is the hidden layer size of the text representation, and in the present application, it can be set to 768.

[0043] It should be noted that the encoder is a multi-layer bidirectional labeled Transformer structure. Each Transformer block consists of a multi-head attention layer and a fully connected layer. The input and output vectors of the multi-head attention layer are connected through a residual connection and input to the fully connected layer after layer normalization. Residual connection and layer normalization operations are also performed on the output of the fully connected layer. The encoder can be understood as being obtained after training the initial encoder according to the sample emotion feature vectors. Among them, the process of training the initial encoder can be exemplified as follows: obtaining the emotion feature vectors of all users through the initial encoder of the model and saving them offline; randomly selecting user training samples and the corresponding offline user emotion features for regular supervised emotion prediction training; after the model completes one round of training iteration, updating and saving the emotion feature vectors of the users through the iterated encoder, and then continuously performing supervised training until the model converges to obtain the trained encoder.

[0044] In the embodiments of the present application, the text sentiment feature encoding data can be understood as the data obtained by pre-encoding the sentiment features in the text data. The process of obtaining the text sentiment feature encoding data corresponding to the text data specifically includes: performing tokenization on the text data to obtain the first tokens corresponding to the text data; using a first token to replace any token in the first tokens to obtain the second tokens after replacement; inputting the second tokens into an embedding matrix to obtain the text feature data corresponding to the second tokens; and inputting the text feature data into an encoder to obtain the text sentiment feature encoding data.

[0045] It should be noted that the process of performing tokenization on the text data can be determined according to the actual situation and is not limited herein. The first tokens can be understood as the token list obtained by performing tokenization on the text data. The first token can be any replacement token, which can be specifically determined according to the actual situation and is not limited herein. As an example, the first token can be denoted as <mask>Among them, the first token can be used to replace one or more tokens in the first token element. The specific number of tokens to be replaced and which token to replace are random and can be determined according to the actual situation, which is not limited here. The second token element can be understood as using <mask>The replaced token list is obtained by replacing any token in the token list with a token pair. The embedding matrix can be selected according to the actual situation and is not limited here. The text feature data can be represented by the embedding vectors of text tokens. This encoder has the same meaning as the above-mentioned encoder, and is obtained by training the initial encoder according to the sample sentiment feature vector; the training process has been described above and will not be elaborated here.

[0046] It should be noted that the text data is tokenized to obtain the first token corresponding to the text data; any token in the first token is replaced with a first token pair to obtain the replaced second token; the second token is input into the embedding matrix to obtain the text feature data corresponding to the second token; it can be illustrated by an example that the text data is tokenized, and then the obtained token list is randomly damaged, with <mask>Tokens randomly replace the word tokens in the text. The tokens obtained after random replacement are then input into the embedding matrix, and the text features are the embedding vectors of the text tokens.

[0047] The solution of the embodiment of the present application uses <mask>Tokens are randomly replaced in the text. During subsequent "masked language modeling", the damaged tokens will be reconstructed in the decoder, and this denoising strategy is used to enhance the robustness and generalization of the model.

[0048] In the embodiment of the present application, the decoder is a multi-layer Transformer. The picture emotion feature data is input into the decoder to obtain the picture emotion probability distribution data; the text emotion feature encoding data is input into the decoder to obtain the text emotion probability distribution data. It can be illustrated as taking the obtained picture emotion feature encoding data and text emotion feature encoding data as the inputs of the decoder to obtain the picture emotion probability distribution data and the text emotion probability distribution data.

[0049] S302. Determine the first text emotion vector data according to the picture emotion probability distribution data; determine the first picture emotion vector data according to the text emotion probability distribution data.

[0050] In the embodiment of the present application, the first text emotion vector data can be understood as the text emotion vector data generated according to the picture emotion probability distribution data; the first picture emotion vector data can be understood as the picture emotion vector data generated according to the text emotion probability distribution data. The process of determining the first text emotion vector data according to the picture emotion probability distribution data and determining the first picture emotion vector data according to the text emotion probability distribution data specifically includes: respectively determining the second picture emotion vector data corresponding to the picture emotion probability distribution data and the second text emotion vector data corresponding to the text emotion probability distribution data; inputting the second picture emotion vector data into the multi-modal derivation module to obtain the generated first text emotion vector data; inputting the second text emotion vector data into the multi-modal derivation module to obtain the generated first picture emotion vector data.

[0051] It should be noted that the second picture emotion vector data can be understood as the high-dimensional vector data obtained by projecting the picture emotion probability distribution data into a linear transformation layer. The second text emotion vector data can be understood as the high-dimensional vector data obtained by projecting the text emotion probability distribution data into a linear transformation layer. Respectively determining the second picture emotion vector data corresponding to the picture emotion probability distribution data and the second text emotion vector data corresponding to the text emotion probability distribution data can be illustrated as inputting the picture emotion probability distribution data and the text emotion probability distribution data into a linear transformation layer respectively to project them into high-dimensional vector data, obtaining the second picture emotion vector data and the second text emotion vector data.

[0052] It should be noted that the multi-modal derivative module can be understood as a multi-modal derivative generator. In practical applications, the multi-modal derivative generator can adopt a text and image generation model, which is a basic model that can generate images from text and generate text from images. Input the second picture sentiment vector data into the multi-modal derivative module to obtain the generated first text sentiment vector data; input the second text sentiment vector data into the multi-modal derivative module to obtain the generated first picture sentiment vector data; it can be exemplified that input the second picture sentiment vector data into the text and image generation model to obtain the generated first text sentiment vector data; input the second text sentiment vector data into the text and image generation model to obtain the generated first picture sentiment vector data.

[0053] S303. Fuse the first text sentiment vector data and the first picture sentiment vector data to obtain sentiment feature vector data.

[0054] In the embodiment of the present application, the sentiment feature vector data can be understood as the sentiment feature vector data after fusing text sentiment and picture sentiment. The process of fusing the first text sentiment vector data and the first picture sentiment vector data to obtain the sentiment feature vector data specifically includes: input the first picture sentiment vector data and the first text sentiment vector data into the sentiment matching module to obtain a first matching result; in the case where the first matching result indicates that the first picture sentiment vector data matches the second picture sentiment vector data successfully and the first text sentiment vector data matches the second text sentiment vector data successfully, fuse the first picture sentiment vector data and the first text sentiment vector data to obtain the sentiment feature vector data.

[0055] It should be noted that the sentiment matching module can be understood as a multi-modal sentiment matching module; the combination of this sentiment matching module and the above-mentioned multi-modal derivative module can be called a disambiguation module, or also called a text-picture sentiment disambiguation module. In the case where the first matching result indicates that the first picture sentiment vector data matches the second picture sentiment vector data successfully and the first text sentiment vector data matches the second text sentiment vector data successfully, fuse the first picture sentiment vector data and the first text sentiment vector data to obtain the sentiment feature vector data, which can be exemplified as, in the case where the real text sentiment vector data and the generated text sentiment vector data, the real picture sentiment vector data and the generated picture sentiment vector data all match successfully, fuse the generated text sentiment vector data and the generated picture sentiment vector data into a unified multi-modal representation, which is the sentiment feature vector data. Among them, the method of fusing the generated text sentiment vector data and the generated picture sentiment vector data can be splicing or weighted average, etc., and the specific fusion method can be determined according to the actual situation and is not limited here.

[0056] It should be noted that the above-mentioned multi-modal generator and multi-modal sentiment matching module have been trained according to the sample picture data and sample text data. The training process of the multi-modal generator and multi-modal sentiment matching module can be illustrated as follows: Prepare text and image data including sentiment matching. For training, ensure that each text and image sample has a corresponding match. In addition, create some unmatched samples, which can be randomly selected combinations of text and images to represent the unmatched situation; Input the real text vector and picture vector into the multi-modal generator to generate the corresponding picture vector and text vector, and use the real text vector and the generated text vector, as well as the real picture vector and the generated picture vector as sample pairs to input into the multi-modal sentiment matching module. The multi-modal sentiment matching module outputs a probability indicating whether the input is a match, and trains the multi-modal generator and sentiment matching module according to the matching result until the loss function of the training converges.

[0057] In the solution of the embodiment of the present application, the disambiguation module in the present application fuses the high-dimensional features of picture sentiment and text sentiment, and is used to help capture the sentiment alignment relationship between the two types of modal data, so as to improve the accuracy of sentiment recognition in picture-text data.

[0058] S304. Predict the sentiment in the first picture-text data according to the sentiment feature vector data.

[0059] In the embodiment of the present application, the process of predicting the sentiment in the first picture-text data according to the sentiment feature vector data specifically includes: performing splicing processing on the picture sentiment feature encoding data corresponding to the picture data and the text sentiment feature encoding data corresponding to the text data to obtain the spliced sentiment feature encoding data; Inputting the spliced sentiment feature encoding data and the sentiment feature vector data into the sentiment prediction model to obtain the sentiment prediction result corresponding to the first picture-text data.

[0060] It should be noted that the acquisition process and meaning of the picture sentiment feature encoding data and the text sentiment feature encoding data have been described above and will not be repeated here. The splicing of the picture sentiment feature encoding data and the text sentiment feature encoding data can be determined according to the actual situation and is not limited here. The sentiment prediction module can be understood as a multi-modal sentiment prediction module, and specifically can be a multi-layer perceptron (MLP) classifier. Inputting the spliced sentiment feature encoding data and the sentiment feature vector data into the sentiment prediction model to obtain the sentiment prediction result corresponding to the first picture-text data can be illustrated as: Inputting the spliced sentiment feature encoding data and the sentiment feature vector data into the MLP classifier to obtain the sentiment prediction result corresponding to the first picture-text data.

[0061] The solution of the embodiment of the present application determines the first text emotion vector data using the picture probability distribution data, and determines the first picture emotion vector data using the text probability distribution data, which can better obtain the emotion information corresponding to the picture and the text respectively. By fusing the first text emotion vector and the first picture emotion vector, the emotion alignment relationship between the text emotion and the picture emotion can be obtained, thereby improving the accuracy of emotion recognition for picture-text data.

[0062] For the convenience of understanding, an example is given here. Figure 4 It is a schematic diagram of an exemplary model architecture provided by the embodiment of the present application; as Figure 4 shown, it includes a feature extraction module, an encoder-decoder, a picture-text emotion disambiguation module, and a multi-modal emotion prediction module. The feature extraction module is used to obtain the feature data corresponding to the picture data and the text data, and input the picture feature data and the text feature data into the encoder to obtain the user emotion feature. The user emotion feature is used as the input of the decoder. The decoder outputs the picture probability distribution data and the text probability distribution data, and inputs the picture probability distribution data and the text probability distribution data into the picture-text emotion disambiguation module. The picture-text emotion disambiguation module inputs the emotion feature vector data into the multi-modal emotion prediction module to perform emotion prediction on the picture data and the text data. <cls>Used to generate user emotional features.

[0063] Feature extraction module.

[0064] 1. Image representation. Use Faster R-CNN to extract visual features. Specifically, use Faster R-CNN to extract the 48 regions with the highest confidence values from the input image, and at the same time retain the semantic category distribution of each region for use in the subsequent "masked region modeling" task. For the retained regions, their visual features use the mean-pooled convolutional features processed by Faster R-CNN. Let R = {r 1 ,..., r 48} represent the visual features, where r i ∈ R 2048 represents the visual feature of the i-th region. To be consistent with the representation of text features, the obtained visual features are fed into a linear transformation layer to project the original 2048-dimensional vector into a d-dimensional vector, denoted as V ∈ R d×48 , where d is the hidden layer size of the text representation, which is set to 768 in this application.

[0065] 2. Text representation. For text input, follow the general text processing method. First, tokenize the input text, and then randomly disrupt the obtained token list, using <mask>Tokens are randomly substituted for the word tokens in the text. During subsequent "masked language modeling", the corrupted word tokens will be reconstructed in the decoder. This denoising strategy is used to enhance the robustness and generalization of the model. The tokens obtained after random substitution are then input into the embedding matrix, and the text features are the embedding vectors of the text tokens. Let E = {e 1 ,.., e T} denote the token indices of the text input, where T represents the length of the input text. W = {w 1 ,.., w T} denotes the embedding vectors after mapping through the embedding matrix. It is formulated as W = ME, where M is the embedding matrix.

[0066] 3. Multimodal representation. For the overall representation of multimodal inputs, Figure 5 is a schematic diagram of an exemplary visual and text feature extraction module provided by an embodiment of the present application. As Figure 5 shown, at the beginning of the input, <cls>Markers are used to generate user emotional features; to distinguish different modal inputs, <pic>and< / pic> indicating the start and end of visual features, <sos>and <eos>Indicates text input. The image data is subjected to feature extraction using Faster R-CNN, and the text data is subjected to <mask>Tokens are randomly substituted for the lemmas in the text. In practical applications, X can be used to represent the concatenated multi-modal input.

[0067] Encoder-decoder module. The encoder of the model is a multi-layer bidirectional labeled Transformer structure. Each Transformer block consists of a multi-head attention layer and a fully connected layer. The input and output vectors of the multi-head attention layer are connected through a residual connection and are input to the fully connected layer after layer normalization. Residual connection and layer normalization operations are also performed on the output of the fully connected layer. The decoder of the model is also a multi-layer Transformer. However, different from the encoder, the decoder is unidirectional when generating the output and generates the result in an autoregressive manner, while the encoder represents bidirectionally according to the context.

[0068] User emotion feature pre-encoding module. Figure 6 Schematic diagram of an exemplary user emotion feature pre-encoding module provided by an embodiment of the present application; as Figure 6 shown, the encoder obtains the feature data of User 1, User 2, and User 3, and then outputs the user emotion feature h1 cls 、user emotion feature h2 cls and user emotion feature h3 cls , all of which are pre-encoded data and are used as the input of the decoder.

[0069] The user's personalized emotion feature is in the form of a latent variable representing the user's emotion appeal, taking values from the encoder input <cls>Output of the labeled hidden layer. At the decoder end of the model, it is spliced into the decoder input as a kind of information gain of the decoder. The construction and application of user sentiment features are mainly divided into two stages: model training and inference. During the model training process, the construction steps are as follows:

[0070] (1) Obtain the sentiment feature vectors of all users through the encoder of the model and save them offline;

[0071] (2) Randomly select user training samples and the corresponding offline user sentiment features for conventional supervised sentiment prediction training;

[0072] (3) After the model completes one round of training iteration, update the user sentiment feature vectors through the iterated encoder and save them, and then continue with supervised training until the model converges;

[0073] (4) After the model training is completed, obtain the final user sentiment feature vectors as the pre-coded features of the user.

[0074] During the model inference process, perform the conventional forward calculation process, and select different pre-coded features for different users to predict the personalized sentiment labels of the users.

[0075] Graphical and textual sentiment disambiguation module. Figure 7 It is a schematic structural diagram of an exemplary graphical and textual disambiguation module provided by an embodiment of the present application; as Figure 7 shown, a disambiguation unit is used to fuse the image sentiment and the text sentiment. The decoder obtains the image sentiment feature encoded data h cls and the text sentiment feature encoded data h cls , and respectively obtains an image sentiment probability distribution p v and a text sentiment probability distribution p t ; then p v and p t are respectively input into a linear transformation layer to be projected into high-dimensional vectors; then, a disambiguation unit is used to fuse the high-dimensional vectors of the image sentiment distribution and the text sentiment distribution to obtain a multi-modal sentiment feature vector p m , and sentiment prediction is performed on p m through the MLP.

[0076] Specifically, the graphical and textual sentiment disambiguation module uses a disambiguation unit to fuse the image sentiment and the text sentiment. First, an image sentiment probability distribution p v and a text sentiment probability distribution p t are respectively obtained through the decoder. For the decoder input, the input for image sentiment prediction is <sos> <vsp>, the input for text sentiment prediction is <sos> <tsp>; Then input p v and p t into a linear transformation layer respectively, project them into high-dimensional vectors; then fuse the high-dimensional vectors of the picture emotion distribution and the text emotion distribution through a disambiguation unit to obtain a multi-modal emotion feature vector p m .

[0077] Figure 8 is a schematic structural diagram of an exemplary disambiguation unit provided by an embodiment of the present application; as Figure 8 shown, it includes a multi-modal generator G and a multi-modal emotion matching module D. Input the real text x into the multi-modal generator G to obtain the generated picture y', input the real picture y into the multi-modal generator G to obtain the generated text x', use the real picture y and the generated picture y' as a sample pair (y, y') and input them into the multi-modal emotion matching module D to obtain a matching or non-matching result, and use the real text x and the generated text x' as a sample pair (x, x') and input them into the multi-modal emotion matching module D to obtain a matching or non-matching result.

[0078] Specifically, the training and inference steps of the disambiguation unit are as follows:

[0079] 1. Data preparation. Prepare text and image data including emotion matching. For training, ensure that each text and image sample has a corresponding match. In addition, create some non-matching samples, which can be randomly selected text and image combinations to represent non-matching situations.

[0080] 2. Multi-modal generator. The multi-modal generator G uses the currently relatively advanced text and image generation model CM3Leon, which is a basic model that can generate pictures from text and generate text from pictures. It takes the real text vector x and the picture vector y as inputs and generates the corresponding image vector y' and text vector x'. The goal of the multi-modal generator is to generate realistic matching samples: y' = G(x) and x' = G(y).

[0081] 3. Multi-modal emotion matching module. Design the multi-modal emotion matching module D. For each sample pair (x, y), it accepts two pairs of real text vectors x and image vectors y, and the generated text vectors x' and y' as inputs and outputs a probability indicating whether the inputs match. The multi-modal emotion matching module is a binary classifier: D(x, x') = Probability that (x, x') is a matched pair, D(y, y') = Probability that (y, y') is a matched pair.

[0082] 4. Model training. In this step, the multi-modal generator and the emotion matching module will be trained. The loss function for training includes:

[0083] Sample pair (x, x') loss: \(L_{D_x}=-\log D(x,x')-\log(1 - D(x,x'))\).

[0084] Image pair (y, y') loss: \(L_{D_y}=-\log D(y,y')-\log(1 - D(y,y'))\).

[0085] 5. Evaluation and inference. Once the training is completed, use the multimodal generator \(G\) in the network to generate text and image vectors, and then fuse them into a unified multimodal representation: \(z_{multimodal}=f(x',y')\).

[0086] Here, \(f\) represents a function for fusing text and generating images, which can be simple concatenation, weighted average, or other fusion strategies.

[0087] Multimodal Sentiment Prediction (MSP) module. Finally, pass through an MLP classifier to predict the multimodal sentiment distribution \(P(s)\): \(p\) v =\(Decoder(H\) e ; \(e\) vsp ), \(h\) v =\(W\) v \(p\) v ; \(p\) t =\(Decoder(H\) e ; \(e\) tsp ), \(h\) t =\(W\) t \(p\) t ; \(z\) multimodel =\(f(G(p\) v ), \(G(p\) v )); \(P(s)=Softmax(MLP(z\) multimodel \(h)), \(h = [h\) v , \(h\) t . Here, \(H_e\) is the output of the encoder for the multimodal concatenated input, and \(E_{vsp}\), \(E_{tsp}\) are the word embeddings corresponding to the two special tokens of the image sentiment input and the text sentiment input respectively. The multimodal sentiment prediction task MSP uses the sentiment label as the supervision signal. Formally, model the MSP task as a classification task and use the cross-entropy loss for the MSP task: \(L\) MSP =-\(E\) X~D \(\log P(y|X)\). Here, \(y\) is the true sentiment label annotated in the dataset.

[0088] Model training and inference. The model uses the encoder-decoder structure based on Transformer as the framework. Specifically, both the encoder and the decoder have 6 layers of transformers. When training the model, hyperparameter tuning is performed through a fixed development set. The batch size is set to 64. The learning rate is set to 5e-5. The hidden layer size of the model is set to 768. The trade-off hyperparameters λ 1 , λ 2 and λ 3 are all set to 1. When performing model inference, the output of the MSP task is used as the emotion prediction result of the multi-modal sentiment analysis.

[0089] In the solution of the embodiment of the present application, the main framework uses an encoder-decoder based architecture to generate the output emotion recognition result in an autoregressive manner; wherein, the user's emotion requirements are pre-encoded by the encoder and spliced into the decoder input as an information gain of the decoder, so that the model has the ability to output customized emotion tendencies for different users.

[0090] In the solution of the embodiment of the present application, through the graphic-text emotion disambiguation module, the high-dimensional features of the picture emotion and the text emotion are fused to help capture the emotion alignment relationship between the two modal data.

[0091] The embodiment of the present application provides an emotion prediction device, Figure 9 which is a structural schematic diagram of an emotion prediction device provided by the embodiment of the present application; as Figure 9 shown, the emotion prediction device 900 includes:

[0092] A determination unit 901, configured to determine the picture emotion probability distribution data corresponding to the picture data in the first graphic-text data, and determine the text emotion probability distribution data corresponding to the text data in the first graphic-text data;

[0093] The determination unit 901 is further configured to determine the first text emotion vector data according to the picture emotion probability distribution data; and determine the first picture emotion vector data according to the text emotion probability distribution data;

[0094] A fusion unit 902, configured to fuse the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data;

[0095] A prediction unit 903, configured to predict the emotion in the first graphic-text data according to the emotion feature vector data.

[0096] Optionally, the determining unit 901 is further configured to obtain the picture emotion feature encoding data corresponding to the picture data, and obtain the text emotion feature encoding data corresponding to the text data; input the picture emotion feature encoding data into a decoder to obtain the picture emotion probability distribution data; input the text emotion feature encoding data into the decoder to obtain the text emotion probability distribution data.

[0097] Optionally, the determining unit 901 is further configured to perform feature extraction on the picture data by using a convolutional neural network to obtain picture feature data; input the picture feature data into an encoder to obtain the picture emotion feature encoding data.

[0098] Optionally, the determining unit 901 is further configured to tokenize the text data to obtain a first token corresponding to the text data; replace any token in the first token with a first tag to obtain a second token after replacement; input the second token into an embedding matrix to obtain text feature data corresponding to the second token; input the text feature data into an encoder to obtain the text emotion feature encoding data.

[0099] Optionally, the determining unit 901 is further configured to respectively determine second picture emotion vector data corresponding to the picture emotion probability distribution data and second text emotion vector data corresponding to the text emotion probability distribution data; input the second picture emotion vector data into a multimodal derivation module to obtain the generated first text emotion vector data; input the second text emotion vector data into the multimodal derivation module to obtain the generated first picture emotion vector data.

[0100] Optionally, the fusion unit 902 is further configured to input the first picture emotion vector data and the first text emotion vector data into an emotion matching module to obtain a first matching result; in a case where the first matching result indicates that the first picture emotion vector data matches the second picture emotion vector data and the first text emotion vector data matches the second text emotion vector data, fuse the first picture emotion vector data and the first text emotion vector data to obtain the emotion feature vector data.

[0101] Optionally, the prediction unit 903 is further configured to perform splicing processing on the picture emotion feature encoding data corresponding to the picture data and the text emotion feature encoding data corresponding to the text data to obtain spliced emotion feature encoding data; input the spliced emotion feature encoding data and the emotion feature vector data into an emotion prediction model to obtain an emotion prediction result corresponding to the first picture-text data.

[0102] An embodiment of this application further provides an emotion prediction device Figure 10 The structural schematic diagram of an emotion prediction device provided by an embodiment of the present application; as Figure 10 shown, the emotion prediction device 1000 includes: a processor 1001 and a memory 1003. Optionally, the emotion prediction device 1000 may further include a communication bus 1002.

[0103] In the process of a specific embodiment, the above-mentioned processor 1001 may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing image processing device (DSPD), a programmable logic image processing device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above-mentioned processor functions may also be others, and this embodiment does not make specific limitations.

[0104] In the embodiment of the present application, the above-mentioned communication bus 1002 is used to implement the connection and communication between the processor 1001 and the memory 1003; when the above-mentioned processor 1001 executes the running program stored in the memory 1003, the following emotion prediction method is implemented:

[0105] Determine the picture emotion probability distribution data corresponding to the picture data in the first picture-text data, and determine the text emotion probability distribution data corresponding to the text data in the first picture-text data; determine the first text emotion vector data according to the picture emotion probability distribution data; determine the first picture emotion vector data according to the text emotion probability distribution data; fuse the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data; predict the emotion in the first picture-text data according to the emotion feature vector data.

[0106] Further, the above-mentioned processor 501 is further configured to obtain the picture emotion feature encoding data corresponding to the picture data, and obtain the text emotion feature encoding data corresponding to the text data; input the picture emotion feature encoding data into a decoder to obtain the picture emotion probability distribution data; input the text emotion feature encoding data into the decoder to obtain the text emotion probability distribution data.

[0107] Further, the above-mentioned processor 501 is further configured to extract features from the picture data by using a convolutional neural network to obtain picture feature data; and input the picture feature data into an encoder to obtain the picture emotion feature encoding data.

[0108] Further, the above-mentioned processor 501 is further configured to tokenize the text data to obtain a first token corresponding to the text data; replace any token in the first token with a first tag to obtain a second token after replacement; input the second token into an embedding matrix to obtain text feature data corresponding to the second token; and input the text feature data into an encoder to obtain the text emotion feature encoding data.

[0109] Further, the above-mentioned processor 501 is further configured to respectively determine a second picture emotion vector data corresponding to the picture emotion probability distribution data and a second text emotion vector data corresponding to the text emotion probability distribution data; input the second picture emotion vector data into a multi-modal derivation module to obtain the generated first text emotion vector data; and input the second text emotion vector data into the multi-modal derivation module to obtain the generated first picture emotion vector data.

[0110] Further, the above-mentioned processor 501 is further configured to input the first picture emotion vector data and the first text emotion vector data into an emotion matching module to obtain a first matching result; and in the case where the first matching result indicates that the first picture emotion vector data matches successfully with the second picture emotion vector data and the first text emotion vector data matches successfully with the second text emotion vector data, fuse the first picture emotion vector data and the first text emotion vector data to obtain the emotion feature vector data.

[0111] Further, the above-mentioned processor 501 is further configured to perform splicing processing on the picture emotion feature encoding data corresponding to the picture data and the text emotion feature encoding data corresponding to the text data to obtain spliced emotion feature encoding data; and input the spliced emotion feature encoding data and the emotion feature vector data into an emotion prediction model to obtain an emotion prediction result corresponding to the first picture and text data.

[0112] An embodiment of the present application provides a storage medium, on which a computer program is stored. The above-mentioned computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors, and the computer program implements the emotion prediction method as described above.

[0113] Based on the above embodiments, an embodiment of the present application provides a computer program product, including a computer program, which can be executed by one or more processors, and the computer program implements the emotion prediction method as described above.

[0114] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing an image display device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present disclosure.

[0116] The above is only a preferred embodiment of the present application and is not used to limit the protection scope of the present application.< / tsp> < / sos> < / vsp> < / sos> < / cls> < / mask> < / eos> < / sos> < / cls> < / mask> < / cls> < / mask> < / mask> < / mask> < / mask>

Claims

1. A sentiment prediction method, characterized in that: The method comprises: Determine the picture emotion probability distribution data corresponding to the picture data in the first picture and text data, and determine the text emotion probability distribution data corresponding to the text data in the first picture and text data; Determine first text emotion vector data according to the picture emotion probability distribution data; Determine first picture emotion vector data according to the text emotion probability distribution data; Fusion of the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data; The emotion in the first graphic data is predicted according to the emotion feature vector data.

2. The method according to claim 1, characterized in that The step of determining the image emotion probability distribution data corresponding to the image data in the first image-text data, and determining the text emotion probability distribution data corresponding to the text data in the first image-text data, includes: Obtaining picture emotion feature encoding data corresponding to the picture data, and obtaining text emotion feature encoding data corresponding to the text data; The image emotion feature encoding data is input into a decoder to obtain the image emotion probability distribution data; the text emotion feature encoding data is input into the decoder to obtain the text emotion probability distribution data.

3. The method according to claim 2, characterized in that The step of obtaining the picture emotion feature encoding data corresponding to the picture data comprises: Using a convolutional neural network to extract features from the image data to obtain image feature data; The image feature data is input into an encoder to obtain the image emotion feature encoding data.

4. The method according to claim 2, characterized in that: The step of obtaining text emotion feature encoding data corresponding to the text data includes: Tokenizing the text data to obtain a first token corresponding to the text data; Replacing any word in the first word with a first tag to obtain a replaced second word; Inputting the second word-unit into an embedding matrix to obtain text feature data corresponding to the second word-unit; The text feature data is input into an encoder to obtain the text emotion feature encoding data.

5. The method according to claim 1, characterized in that The step of determining the first text emotion vector data according to the picture emotion probability distribution data; and determining the first picture emotion vector data according to the text emotion probability distribution data comprises: Respectively determining the second picture emotion vector data corresponding to the picture emotion probability distribution data, and the second text emotion vector data corresponding to the text emotion probability distribution data; Input the second picture emotion vector data into a multimodal derivation module to obtain the generated first text emotion vector data; input the second text emotion vector data into the multimodal derivation module to obtain the generated first picture emotion vector data.

6. The method according to claim 1, characterized in that The step of fusing the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data includes: Inputting the first picture emotion vector data and the first text emotion vector data into an emotion matching module to obtain a first matching result; When the first matching result indicates that the first picture emotion vector data successfully matches the second picture emotion vector data, and the first text emotion vector data successfully matches the second text emotion vector data, the first picture emotion vector data and the first text emotion vector data are fused to obtain the emotion feature vector data.

7. The method according to claim 1, characterized in that The predicting the emotion in the first graphic data according to the emotion feature vector data includes: Performing splicing processing on the picture emotion feature coding data corresponding to the picture data and the text emotion feature coding data corresponding to the text data to obtain spliced ​​emotion feature coding data; The spliced ​​emotion feature coding data and the emotion feature vector data are input into an emotion prediction model to obtain an emotion prediction result corresponding to the first graphic data.

8. An emotion prediction device, characterized in that: The device comprises: A determination unit, used to determine the picture emotion probability distribution data corresponding to the picture data in the first picture and text data, and determine the text emotion probability distribution data corresponding to the text data in the first picture and text data; The determining unit is further used to determine the first text emotion vector data according to the picture emotion probability distribution data; determine the first picture emotion vector data according to the text emotion probability distribution data; A fusion unit, used for fusing the first text emotion vector data and the first picture emotion vector data to obtain emotion feature vector data; A prediction unit is used to predict the emotion in the first graphic data according to the emotion feature vector data.

9. An emotion prediction device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the steps of the method according to any one of claims 1 to 7 are implemented when the processor executes the program.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.