A method for picture emotion recognition that integrates multi-scale information
By integrating the multi-scale features extracted by ViT and ResNet networks and combining KL loss function and cross entropy for learning, multi-task recognition that dominates emotion recognition and emotion distribution prediction in image sentiment analysis is achieved, solving the problem that emotion distribution cannot be fully displayed in the existing technology, and improving the accuracy and robustness of the recognition.
Patent Information
- Application Number
- CN202111481080.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-06
AI Technical Summary
Existing image sentiment analysis methods rarely consider the relative importance of different emotions expressed in pictures, and it is difficult to identify dominant emotions and emotional distribution at the same time, and emotions are highly subjective.
By integrating the local features extracted by ViT network and the global features extracted by ResNet network, and combining KL loss function and cross entropy for learning, multi-task emotion recognition, including dominant emotion recognition and emotion distribution prediction.
It effectively solves the problem of blurred labels and inadequate expression of emotions, improves the accuracy and robustness of image emotion recognition, and can better represent the emotional area and emotional distribution of the image.
Smart Images

Figure CN114170411B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the problem of picture sentiment analysis in the field of deep learning, and particularly to a picture sentiment recognition method that fuses multi-scale information. Background Art
[0002] With the rapid development of Internet technology, people use pictures more to express their emotions. Therefore, the sentiment analysis of pictures is an urgent and valuable research problem. Most of the existing research completes picture annotation through single-label or multi-label learning and uses convolutional neural networks for feature extraction, and all have achieved good results. In recent years, with the great success of the ViT network in the field of natural language processing and the gradual popularization and application of label distribution learning, the field of picture sentiment analysis has also begun to draw on related ideas to better predict the emotion distribution in pictures and fully represent it. At present, picture sentiment analysis has extensive applications and deeper research needs in fields such as aesthetic analysis, intelligent advertising, and social media public opinion detection.
[0003] Existing methods rarely consider the relative importance of different emotions expressed by pictures and can only recognize the dominant emotion. In fact, emotions are highly subjective, and the same picture may arouse different emotions in different individuals. Therefore, it is very important to learn the emotion distribution of pictures. In view of this, this patent fuses local features and global features and uses multi-scale information to simultaneously complete the recognition of the dominant emotion and the prediction of the emotion distribution. First, the ViT network is used to extract local features, which learn the local information and the correlation information between local parts, facilitating the representation of the emotional regions of pictures and obtaining small-scale emotional features. Second, for the overall features, the ResNet network is used for extraction to make the results more robust. At the same time, the KL loss function and cross-entropy are used for learning, which is beneficial to measuring the information loss caused by the inconsistency between the predicted distribution and the labeled distribution. Summary of the Invention
[0004] The purpose of the present invention is to provide a picture sentiment recognition method that fuses multi-scale information. It uses the method of label distribution learning to annotate pictures and fuses the local features and global features of pictures for multi-task sentiment recognition, solving the problems of label ambiguity and insufficient display of emotion distribution in picture sentiment analysis.
[0005] For the convenience of description, first introduce the following concepts:
[0006] Vision Transformer (ViT): A neural network based on the multi-head self-attention mechanism.
[0007] Residual Networks (ResNets): Solve the network degradation problem by introducing the identity mapping. Common network types include ResNet-18, ResNet-34, ResNet-50, ResNet-101, and ResNet-152.
[0008] Label Distribution Learning (LDL): A label assignment method that characterizes images with emotion distributions.
[0009] Kullback-Leibler (KL) loss function: A loss function for distribution learning that measures the information loss due to the inconsistency between the predicted distribution and the labeled distribution.
[0010] The present invention specifically adopts the following technical solutions:
[0011] An image emotion recognition method that fuses multi-scale information, characterized in that:
[0012] a. Extract small-scale emotion features of the correlation between local parts of the image through the ViT network;
[0013] b. Extract deep global emotion features of the image through the ResNet network;
[0014] c. Use the KL loss function and cross-entropy for image emotion recognition learning;
[0015] d. Fuse local features and global features for multi-task emotion recognition, including the dominant emotion recognition task and the label distribution learning prediction task;
[0016] The method mainly includes the following steps:
[0017] (1) Image preprocessing: Unify the sizes of the images in the dataset, and then use data augmentation methods such as random cropping and horizontal flipping for data augmentation;
[0018] (2) Label preprocessing: Use label distributions to characterize the image data, perform normalization and other processing on the original multi-person voting scores in the dataset as the true values for distribution learning; the dominant emotion labels are used as the true values for classification learning;
[0019] (3) Local feature extraction: Use the ViT network pre-trained on ImageNet to extract small-scale emotion features;
[0020] (4) Global feature extraction: Use the ResNet convolutional architecture based on the residual structure to extract global large-scale emotion features;
[0021] (5) Feature fusion: Fuse the 1024-dimensional features extracted in step (3) and the 1024-dimensional features extracted in step (4) at the feature level, and splice them into 2048-dimensional features;
[0022] (6) Picture emotion recognition: Input the fused features in (5) into the fully connected layer to obtain the recognition result of the dominant emotion of the picture and the prediction result of the label distribution;
[0023] (7) Model training: Train in an end-to-end manner, and use the KL loss function and cross-entropy for learning;
[0024] (8) Result verification: Verify on a large public dataset, obtain the experimental results by comparing with various indicators, and conduct ablation experiments to prove the rationality of the method.
[0025] The beneficial effects of the present invention are:
[0026] (1) Using the ViT network to extract local features is beneficial to characterizing the emotional regions of pictures and obtaining small-scale emotional features.
[0027] (2) Adopting the ResNet network to extract global features to avoid gradient disappearance or gradient explosion caused by too deep a network.
[0028] (3) Using the KL loss function and cross-entropy for learning is beneficial to measuring the information loss caused by the inconsistency between the predicted distribution and the labeled distribution.
[0029] (4) Conducting multi-task emotion recognition by fusing multi-scale information, including the dominant emotion recognition task and the label distribution learning prediction task. Brief Description of the Drawings
[0030] Figure 1 is the model structure.
[0031] Figure 2 is the result of the present invention on the Flickr_LDL dataset.
[0032] Figure 3 is the result of the present invention on the Twitter_LDL dataset.
[0033] Figure 4 is the result of the ablation experiment. Detailed Embodiment
[0034] The present invention will be further described in detail below in conjunction with the drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the protection scope of the present invention. Those skilled in the art make some non-essential improvements and adjustments to the present invention according to the above-mentioned invention content and conduct specific implementations, which should still fall within the protection scope of the present invention.
[0035] A method for image emotion recognition that integrates multi-scale information, specifically including the following steps:
[0036] (1) Image preprocessing
[0037] Randomly split the Flickr_LDL and Twitter_LDL datasets into a training set (80%) and a test set (20%). Uniformly adjust the image size to 500*500, and randomly crop it to a size of 224*224. At the same time, perform horizontal flipping with a probability of 0.5, and change the image color with 10% automatic contrast to enhance the data and improve the training effect.
[0038] (2) Label preprocessing
[0039] Normalize the original data of the multiple people's votes on eight types of emotions in the dataset to obtain the image emotion distribution labels, and perform label distribution learning; take the category with the most votes among the eight types of emotions as the dominant emotion for emotion classification.
[0040] (3) Local feature extraction
[0041] Use the ViT network pre-trained on ImageNet as the backbone network for feature extraction. During the feature extraction process, the ViT network first divides the original image into blocks, then unfolds it into a one-dimensional sequence and inputs it into the encoder part of the original Transformer model, where it is processed by methods such as multi-head attention. Finally, change the output features to 1024 dimensions. Learn the correlation between local information and local regions, characterize the emotional regions of the image, and obtain small-scale emotional features.
[0042] (4) Global feature extraction
[0043] Adopt the ResNet convolutional architecture based on the residual structure to extract global deep emotional features, and cancel the fully connected layer for output classification in the last layer. The ResNet network structure can deepen the network depth by stacking basic residual units, while avoiding the vanishing gradient or exploding gradient caused by the overly deep network, learn the overall visual features of the image, increase the depth of representation, and obtain large-scale global features.
[0044] (5) Feature fusion
[0045] The feature fusion method is as Figure 1 shown. Combine the 1024-dimensional features extracted in step (2) and the 1024
[0046] The 1D features are concatenated into 2048D features. Before inputting into the fully connected layer, the local features and global features are fused at the feature level and concatenated into a sentiment feature vector containing multi-scale information, increasing the accuracy of image sentiment recognition.
[0047] (6)Image sentiment recognition
[0048] The 2048D features concatenated in step (5) are input into the fully connected layer to complete multi-task sentiment recognition, obtaining the final dominant emotion recognition result and the predicted result of label distribution learning. As shown by DominantLabel and DistributionLabel in Figure 1 .
[0049] (7)Model training
[0050] Training is conducted in an end-to-end manner with an initial learning rate of 0.001, divided by 10 every 10 epochs, and a total of 50 epochs. The input is the dataset images, and the result of image sentiment recognition is directly output, reducing the complexity of manual operations. The relative entropy KL loss function and cross-entropy Cross Entropy loss are used for learning. The KL loss function is a loss function for distribution learning, which can measure the information loss caused by the inconsistency between the predicted distribution and the labeled distribution. Its formula is shown in formula (1):
[0051] (1)
[0052] where y represents the emotion distribution labeled from the dataset, represents the predicted emotion distribution, N represents the number of images in a specific dataset, and C represents the emotion categories involved. By optimizing the KL loss, the distribution of visual emotions is learned, and by optimizing the Cross Entropy loss, the dominant emotion classification is learned, achieving simultaneous optimization and improvement of multi-tasks.
[0053] (7)Result verification
[0054] Six commonly used distribution learning measurement methods are used for verification on the large publicly available Flickr_LDL and Twitter_LDL datasets. Among them, the distance metrics include Chebyshev distance, Clark distance, Canberra metric, and KL divergence. The similarity metrics include Cosine coefficient and Intersection similarity. In addition, the maximum values of Clark distance and Canberra metric are determined by the number of emotion categories. For standardized comparison, the same operations as previous work are adopted: divide the Clark distance by the square root of the number of emotion categories, and divide the Canberra metric by the number of emotion categories. In addition, top-1 accuracy is further introduced as an evaluation indicator to compare the prediction of the dominant emotion. Figure 2 andFigure 3 The validation results on the Flickr_LDL and Twitter_LDL datasets are shown respectively, where the downward arrow indicates the lower the better, and the upward arrow indicates the higher the better. It can be seen that after comprehensively considering the global deep features and the correlation between local features and their predecessors, the present invention has obtained better classification and distribution results than the baseline method on two widely used datasets, proving the superiority of the present invention.
[0055] The results of the ablation experiment are as Figure 4 shown. When only using the ResNet network for feature extraction and learning, the results are the worst, and the six indicators of distribution learning and the accuracy indicator of dominant emotion classification are all very poor. After integrating the ViT module for feature extraction and using the KL loss module, the accuracy of dominant emotion classification and distribution learning indicators has been improved, indicating the compensation of ViT for insufficient global features and the effectiveness of the KL loss module for distribution learning. The finally proposed model has achieved the best distribution learning and classification results, proving the effectiveness of the model and the necessity of each part of the model.
Claims
1. A method for image emotion recognition that fuses multi-scale information, characterized in that: a. Extract small-scale emotion features of the correlation between local parts of the image through the ViT network; b. Extract deep global emotion features of the image through the ResNet network; c. Use the KL loss function and cross-entropy for image emotion recognition learning; d. Fuse local features and global features for multi-task emotion recognition, including dominant emotion recognition tasks and label distribution learning prediction tasks; This method mainly includes the following steps: (1) Image preprocessing: Unify the size of the images in the dataset, and use two data augmentation methods, random cropping and horizontal flipping, to augment the data; (2) Label preprocessing: Normalize the multi-person voting values for 8 types of emotions to generate emotion distribution labels, and at the same time use the highest-voted category as the dominant emotion label; (3) Local feature extraction: Extract local correlation features through the ViT network to obtain small-scale emotion features; (4) Global feature extraction: Extract global depth features through the ResNet network to obtain large-scale emotion features; (5) Feature fusion: Concatenate the features in steps (3) and (4) into fused features; (6) Image emotion recognition: Input the fused features into the fully connected layer, and output the dominant emotion classification result and the label distribution prediction result at the same time; (7) Model training: Optimize label distribution learning with the KL loss function, optimize dominant emotion classification with cross-entropy loss, and perform end-to-end training; (8) Result verification: Verify the effectiveness of the model through distance metrics and ablation experiments on the public dataset.
2. The method for image emotion recognition that fuses multi-scale information according to claim 1, characterized in that In step (1), the size of the dataset images is unified to 500*500, randomly cropped to 224*244, and horizontally flipped with a probability of 0.5, and the image color is changed with 10% automatic contrast to enhance the data and improve the training effect.
3. The method for image emotion recognition that fuses multi-scale information according to claim 1, characterized in that In step (2), the original data of multi-person voting on 8 types of emotions in the dataset is normalized to obtain image emotion distribution labels for label distribution learning, and at the same time the category with the most votes among the 8 types of emotions is used as the dominant emotion for emotion classification.
4. The method for image emotion recognition that fuses multi-scale information according to claim 1, characterized in that In step (3), the local channels for feature extraction are extracted using the ViT network pre-trained on ImageNet, so as to learn the correlation between local information and local parts, represent the emotional regions of the image, and obtain 1024-dimensional small-scale emotion features.
5. The method for image emotion recognition that fuses multi-scale information according to claim 1, characterized in that In step (4), the global channels for feature extraction use the ResNet convolutional architecture based on the residual structure to learn the overall visual features of the image, increase the depth of representation, and obtain 1024-dimensional large-scale global features.
6. The method for image emotion recognition that fuses multi-scale information according to claim 1, characterized in that In step (5), the local features and global features are fused at the feature level before inputting into the fully connected layer, and concatenated into a 2048-dimensional sentiment feature vector containing multi-scale information, increasing the accuracy of image sentiment recognition.
7. The method for image sentiment recognition by fusing multi-scale information according to claim 1, wherein in step (6), the fused 2048-dimensional features pass through the fully connected layer to simultaneously obtain the multi-task results of sentiment recognition, including the dominant emotion recognition result and the label distribution prediction result.
8. The method for image sentiment recognition by fusing multi-scale information according to claim 1, wherein in step (7), an end-to-end training method is used, with an initial learning rate of 0.001, divided by 10 every 10 epochs, a total of 50 epochs, the input being the dataset images, and directly outputting the results of image sentiment recognition, reducing the complexity of manual operations.
9. The method for image sentiment recognition by fusing multi-scale information according to claim 1, wherein in step (7), the relative entropy KL loss and cross-entropy CrossEntropy loss, which measure the information loss caused by the inconsistency between the predicted distribution and the labeled distribution, are used for learning. The distribution of visual sentiment is learned by optimizing the KL loss, and the dominant emotion classification is learned by optimizing the Cross Entropy loss, achieving the simultaneous optimization and improvement of multiple tasks.
10. The method for image sentiment recognition by fusing multi-scale information according to claim 1, wherein in step (8), it is tested on two large public datasets, and the distance metric and similarity metric are used for verification respectively, including Chebyshev distance, Clark distance, Canberra metric, KL divergence, Cosine coefficient, and Intersection similarity; and ablation experiments are carried out to verify the effectiveness.
Citation Information
Patent Citations
Image sentiment classification method based on class activation mapping and visual saliency
CN111832573A