Multi-modal sentiment analysis method and device based on image description and medium

Through the multimodal sentiment analysis method based on image description, and using technical means such as semantic extraction and feature fusion modules, the problems of large semantic gaps between modals and high demand for computing resources in multimodal sentiment analysis are solved, and more efficient and accurate sentiment analysis is achieved.

CN120011990APending Publication Date: 2025-05-16SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510028299.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the multimodal sentiment analysis, the existing technology has problems such as large semantic gaps between modals and high demand for computing resources, which leads to low accuracy of the analysis results.

Method used

A multimodal sentiment analysis method based on image description is designed, and image description is generated through semantic extraction modules, combined with feature extraction modules, semantic reconstruction modules, feature fusion modules and classification modules, to realize weighted fusion of text and image descriptions and output sentiment analysis results.

Benefits of technology

It effectively reduces the semantic gap between modes, improves the fusion effect of multimodal information and the accuracy of analysis results, and reduces the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011990A_ABST
    Figure CN120011990A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal sentiment analysis method and device based on image description and a medium, and the method comprises the following steps: obtaining image data and text data, inputting a multi-modal sentiment analysis model based on image description, and outputting a sentiment analysis result; the model comprises a semantic extraction module, a feature extraction module, a semantic reconstruction module, a feature fusion module and a classification module, wherein the semantic extraction module is used for generating image description according to image data; the feature extraction module is used for performing feature extraction on the text data and the image description; the semantic reconstruction module is used for reconstructing the text features and the image description features respectively; the feature fusion module dynamically adjusts the weights of the text reconstruction features and the image description reconstruction features through a gating mechanism, and performs weighted fusion on the text reconstruction features and the image description reconstruction features; the classification module is used for outputting sentiment analysis results according to the fusion features. Compared with the prior art, the accuracy of an analysis result can be improved, and meanwhile the requirement for computing resources can be lowered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual language multimodal information fusion, and in particular relates to a multimodal sentiment analysis method, device and medium based on image description. Background Art

[0002] Visual Language Models (VLMs) organically combine image and text data, and can simultaneously obtain information from both visual and language dimensions, thereby enhancing the model's understanding and cognitive capabilities, and can effectively handle practical tasks such as sentiment analysis. However, there is a natural semantic gap between images and texts, which makes it difficult for multimodal models to align semantics, which in turn limits their ability to understand semantics, reducing the overall performance and application effects of the model. Therefore, bridging the gap between the two has become a key challenge.

[0003] Image description is an interesting multimodal task proposed by Oriol Vinyals et al. It can be understood as asking a computer to generate complete descriptive text based on the content of an image. Early image descriptions relied on manually designed rules and templates for generation. Usually, a list of possible objects, actions, and scenes needs to be predefined, and these elements are identified and filled into the description template using classification algorithms. This method is too dependent on manually defined rules and templates, which makes it inflexible when facing complex scenes and difficult to adapt to the changing image content. Subsequent studies began to use machine learning methods, using feature engineering to extract key image features (such as SIFT, HOG, etc.), and perform more complex image processing and text generation based on these features. Although such methods are more flexible, they usually need to process complex visual information, which requires high computing resources and may lead to slow processing speed and inefficiency. In addition, how to fuse image description and text data is also a major problem in multimodal deep learning. Existing methods mainly learn to fuse the extracted heterogeneous features and project them into a common representation space, which places extremely high demands on the modal interaction module. Therefore, although existing vision-language pre-trained models (VLPs) perform well on tasks such as image description generation, they do not perform well when directly applied to multimodal tasks due to the lack of cross-modal correlation learning.

[0004] In summary, it is necessary to design a multimodal sentiment analysis method that can effectively reduce the semantic gap between modalities, improve the fusion effect of multimodal information, and thus improve the accuracy of the analysis results, while also reducing the demand for computing resources. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a multimodal sentiment analysis method, device and medium based on image description, so as to effectively reduce the semantic gap between modalities, improve the fusion effect of multimodal information, thereby improving the accuracy of the analysis results, and at the same time reduce the demand for computing resources.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] The present invention provides a multimodal sentiment analysis method based on image description, comprising the following steps:

[0008] Obtain image data and text data, input a multimodal sentiment analysis model based on image description, and output sentiment analysis results;

[0009] Among them, the multimodal sentiment analysis model based on image description includes a semantic extraction module, a feature extraction module, a semantic reconstruction module, a feature fusion module and a classification module. The semantic extraction module is used to generate a corresponding image description according to the image data through a pre-trained visual language model; the feature extraction module is used to perform feature extraction on the text data and the image description respectively to obtain text features and image description features; the semantic reconstruction module includes multiple linear layers, which are used to reconstruct the text features and the image description features respectively to obtain text reconstruction features and image description reconstruction features; the feature fusion module dynamically adjusts the weights of the text reconstruction features and the image description reconstruction features through a gating mechanism, and weightedly fuses the text reconstruction features and the image description reconstruction features to obtain fusion features; the classification module is used to output the sentiment analysis results according to the fusion features through a classifier.

[0010] Furthermore, the pre-trained visual language model is BLIP-2-OPT-2.7B.

[0011] Furthermore, feature extraction is performed on the text data and the image description respectively through the pre-trained GPT-2.

[0012] Furthermore, the semantic reconstruction module includes two single-layer linear layers.

[0013] Furthermore, the text reconstruction features and the image description reconstruction features are weightedly fused through early fusion and late fusion.

[0014] Furthermore, the early fusion is performed based on bilinear pooling, and the specific process is as follows:

[0015] Mapping the text reconstruction features and the image description reconstruction features through a multi-layer perceptron respectively to obtain a text reconstruction feature vector and an image description reconstruction feature vector;

[0016] The outer product of the text reconstruction feature vector and the image description reconstruction feature vector is calculated and a pooling operation is performed.

[0017] Furthermore, the late fusion is performed based on an average method, and the specific process is as follows:

[0018] Inputting the text reconstruction feature and the image description reconstruction feature into a linear classifier respectively to obtain corresponding classifier decisions;

[0019] The classifier decisions of the text reconstruction features and the image description reconstruction features are weighted averaged.

[0020] Furthermore, the classifier is a linear classifier.

[0021] The present invention also provides an electronic device, comprising a memory, a processor, and a program stored in the memory, wherein the processor implements the above method when executing the program.

[0022] The present invention also provides a computer-readable storage medium on which a computer program is stored, and the program implements the above method when executed by a processor.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. The present invention proposes a multimodal sentiment analysis method based on image description, designs a multimodal sentiment analysis model based on image description, and can generate sentiment analysis results according to image data and text data; the model includes a semantic extraction module, a feature extraction module, a semantic reconstruction module, a feature fusion module and a classification module, wherein the semantic extraction module is used to generate corresponding image descriptions according to image data as inputs of subsequent modules, which can reduce the semantic gap between modalities on the one hand, and simplify the processing process of visual information on the other hand, greatly reducing the demand for computing resources while retaining the richness and accuracy of image content, and generating image descriptions through pre-trained visual language models can reduce training costs, significantly enhance multimodal effects, and extract richer and more objective information; the feature extraction module is used to perform feature extraction on text data and image descriptions respectively. Feature extraction, obtaining text features and image description features, can accurately reflect the image content and text semantics, and lay a solid foundation for subsequent processing; the semantic reconstruction module includes multiple linear layers, which are used to reconstruct the text features and the image description features respectively, and can ensure that the features of different modalities are effectively compared and combined in the same space; the feature fusion module dynamically adjusts the weights of the text reconstruction features and the image description reconstruction features through a gating mechanism, and weightedly fuses the text reconstruction features and the image description reconstruction features to obtain fusion features, so that the model can better learn the relationship between different modalities and greatly improve the final recognition and generation capabilities; therefore, the above method can effectively reduce the semantic gap between modalities, improve the fusion effect of multimodal information, and thereby improve the accuracy of the analysis results, while also reducing the demand for computing resources.

[0025] 2. The present invention specifically performs weighted fusion of text reconstruction features and image description reconstruction features through early fusion and late fusion, wherein the early fusion is performed based on bilinear pooling, and the specific process is as follows: the text reconstruction features and the image description reconstruction features are respectively mapped through a multi-layer perceptron to obtain a text reconstruction feature vector and an image description reconstruction feature vector, and then the outer product of the text reconstruction feature vector and the image description reconstruction feature vector is calculated and a pooling operation is performed; the late fusion is performed based on an averaging method, and the specific process is as follows: the text reconstruction features and the image description reconstruction features are respectively input into a linear classifier to obtain corresponding classifier decisions, and then the classifier decisions of the text reconstruction features and the image description reconstruction features are weighted averaged; the early fusion is a feature-level fusion, which can retain the complete information of each modality, and the late fusion is a decision-level fusion, which can make full use of the complementarity of data of different modalities and improve the robustness and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a structural diagram of the multimodal sentiment analysis model based on image description;

[0027] Figure 2 Schematic diagram of early fusion;

[0028] Figure 3 Schematic diagram of late fusion. DETAILED DESCRIPTION

[0029] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0030] Example:

[0031] This embodiment proposes a framework called CaLB (Caption and Language-Based), which is a simple fusion network based on description and language. Based on this, a multimodal sentiment analysis model based on image description is constructed, which can output sentiment analysis results based on image data and text data. Figure 1 As shown in the figure, the multimodal sentiment analysis model based on image description specifically includes a semantic extraction module, a feature extraction module, a semantic reconstruction module, a feature fusion module and a classification module. The specific description of each module is as follows:

[0032] (1) Semantic Extraction Module

[0033] The semantic extraction module is used to generate corresponding image descriptions according to image data through a pre-trained visual language model. The image descriptions generated in this way are not affected by human subjective judgment, and are therefore superior to image descriptions obtained by manual annotation methods. At the same time, this method can better understand the complex interactions and relationships between images and languages, accurately capture semantic information in images, and reduce semantic loss and noise. In this embodiment, BLIP-2-OPT-2.7B is selected as the pre-trained visual language model.

[0034] (2) Feature extraction module

[0035] The feature extraction module is used to extract features from the text data and the image description generated by the semantic extraction module, respectively, to obtain text features and image description features. In this embodiment, the pre-trained GPT-2 is selected as the text extractor to accurately extract corresponding features from the image description and text data.

[0036] Using image descriptions instead of traditional images effectively reduces the semantic gap between modalities, significantly improves processing efficiency, simplifies the processing of visual information, and significantly reduces the demand for computing resources while retaining the richness and accuracy of image content.

[0037] (3) Semantic reconstruction module

[0038] In this embodiment, the semantic reconstruction module includes two single-layer linear layers, which are used to reconstruct text features and image description features respectively to obtain text reconstruction features and image description reconstruction features. Through such a reconstruction method, it is possible to ensure that features of different modalities are effectively compared and combined in the same space.

[0039] (4) Feature Fusion Module

[0040] The feature fusion module uses the GMF gate structure to dynamically adjust the weights of text reconstruction features and image description reconstruction features through the gate control mechanism, and then weightedly fuse the text reconstruction features and image description reconstruction features to obtain fusion features. This dynamic fusion strategy enables the model to better learn the relationship between different modalities and greatly improve the final recognition and generation capabilities.

[0041] Specifically, let the text reconstruction feature be T and the image description reconstruction feature be I. First, concatenate the two to obtain the joint feature representation H:

[0042] H = [T; I]

[0043] Then, through the learnable linear transformation W g and bias b g , combined with the Sigmoid activation function to calculate the gate vector g:

[0044] g=σ(W g H+b g )

[0045] Among them, each element of g ranges from [0,1], indicating the weight distribution of the two modal features. Based on the gate vector g, the text reconstruction feature and the image description reconstruction feature are weighted and summed to obtain the fusion feature F:

[0046] F=g⊙T+(1-g)⊙I

[0047] During the training process, the gate vector g is dynamically adjusted according to the differences in the data. When the text feature is more important, the g value is larger, thereby increasing the weight of g⊙T; conversely, when the image feature is more important, the g value is smaller, making the weight of (1-g)⊙I higher. Through this dynamic adjustment strategy, the model can adaptively learn the information contribution of different modalities, thereby effectively improving recognition and generation capabilities.

[0048] At the same time, this module also adopts two representative methods of early fusion (feature fusion) and late fusion (decision fusion), namely bilinear pooling and averaging, to maximize the interaction and expression between different modal information.

[0049] Early fusion is feature-level fusion, focusing on higher-order interactions at the feature level, such as Figure 2 As shown in the figure, the specific process is as follows: first, the text reconstruction features and image description reconstruction features are mapped to the same or appropriate dimensions through a multi-layer perceptron to obtain the text reconstruction feature vector and the image description reconstruction feature vector; then, the outer product of the text reconstruction feature vector and the image description reconstruction feature vector is calculated and pooled to obtain a higher-order cross-modal interaction representation. Subsequently, the interaction representation is input into the GMF gate structure together with the original reconstruction features of the text and image, and adaptively weighted through the gate vector, so as to fully integrate multimodal information in the middle layer and realize the flexible allocation of text and image feature weights.

[0050] Late fusion is decision-level fusion, focusing on the integration of the decision-making level, such as Figure 3 As shown in the figure, the specific process is: first, input the text reconstruction features and image description reconstruction features into independent linear classifiers (or generators) respectively to obtain the corresponding classifier decisions; then perform weighted merging, voting or averaging on the classifier decisions of the text reconstruction features and image description reconstruction features. Since the GMF gate structure is mainly responsible for generating the fusion features of the intermediate layer, the late fusion does not conflict with it: you can first use GMF (combined with optional early fusion operations such as bilinear pooling) to obtain fully fused intermediate features, and then re-integrate the final decisions of different branches / paths. Through the multi-level fusion architecture of "early fusion + GMF + late fusion", not only can more detailed cross-modal associations be mined in the intermediate feature stage, but also the differentiated advantages of different modal features can be retained in the decision-making stage, thereby significantly improving the recognition and generation performance in multimodal scenarios.

[0051] (5) Classification module

[0052] The classification module is used to generate sentiment analysis results based on the fusion features output by the feature fusion module through a linear classifier.

[0053] The above method provides a new perspective for future multimodal learning research, that is, in scenarios where resources are limited or rapid processing is required, it provides new ideas and methods for how to effectively achieve cross-modal learning and information fusion. To verify the effectiveness of this method, this embodiment is verified on two data sets, namely the sarcasm detection data set and the PHEME data set (a data set based on rumors and non-rumors). The sarcasm data set contains 24,635 text-image pairs, of which 19,816 samples are used as training sets, 2,410 samples are used as validation sets, and 2,409 samples are used as test sets. The data set includes two categories: Positive (sarcastic expressions) and Negative (non-sarcastic expressions). The PHEME data set consists of 9 real-time events from 2012 to 2015. The PHEME data set contains 2,948 text-image pairs, which are divided into training sets and test sets in a ratio of 8:2.

[0054] The experiments were conducted on the two datasets and compared the following settings: (1) BLIP and BLIP-2-OPT-2.7B were used in the semantic extraction module to generate image descriptions, where BLIP is a visual language pre-training (VLP) framework suitable for a wide range of downstream tasks, and BLIP-2-OPT-2.7B is achieved by making full use of pre-trained vision and language models; (2) the input combinations of the model are {image, text} and {image description, text}.

[0055] The experimental results are shown in Table 1, and the following conclusions can be drawn: (1) Compared with using a general model, the use of a pre-trained model to generate image descriptions has significantly improved accuracy, F1 value and other indicators. For example, in the sarcasm detection dataset, when the input combination is {image description, text}, the accuracy of the BLIP-2-OPT-2.7B model is improved from 96.06% to 96.63% compared to using BLIP; (2) Compared with directly using the original image, using image description as visual input has significantly improved accuracy, F1 value and other indicators. For example, in the sarcasm detection dataset, after replacing the image with the image description as input, the accuracy of the BLIP-2-OPT-2.7B model is improved from 95.63% to 96.63%.

[0056] Table 1 Experimental results on two datasets

[0057]

[0058] If the above method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0059] The above description of the embodiments is to facilitate the understanding and use of the invention by those skilled in the art. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the present invention should be within the scope of protection of the present invention.

Claims

1. A multimodal sentiment analysis method based on image description, characterized in that: The following steps are involved: Obtain image data and text data, input a multimodal sentiment analysis model based on image description, and output sentiment analysis results; The multimodal sentiment analysis model based on image description includes a semantic extraction module, a feature extraction module, a semantic reconstruction module, a feature fusion module and a classification module. The semantic extraction module is used to generate a corresponding image description according to the image data through a pre-trained visual language model; the feature extraction module is used to extract features from the text data and the image description respectively to obtain text features and image description features; the semantic reconstruction module includes a plurality of linear layers, which are used to reconstruct the text features and the image description features respectively to obtain text reconstruction features and image description reconstruction features; the feature fusion module dynamically adjusts the weights of the text reconstruction features and the image description reconstruction features through a gating mechanism, and weightedly fuses the text reconstruction features and the image description reconstruction features to obtain fusion features; The classification module is used to output a sentiment analysis result according to the fusion feature through a classifier.

2. The multimodal sentiment analysis method based on image description according to claim 1, characterized in that: The pre-trained visual language model is BLIP-2-OPT-2.7B.

3. The multimodal sentiment analysis method based on image description according to claim 1, characterized in that: Feature extraction is performed on the text data and the image description respectively through the pre-trained GPT-2.

4. The multimodal sentiment analysis method based on image description according to claim 1, characterized in that: The semantic reconstruction module includes two single-layer linear layers.

5. The multimodal sentiment analysis method based on image description according to claim 1, characterized in that: The text reconstruction features and the image description reconstruction features are weightedly fused through early fusion and late fusion.

6. The multimodal sentiment analysis method based on image description according to claim 5, characterized in that: The early fusion is based on bilinear pooling, and the specific process is as follows: Mapping the text reconstruction features and the image description reconstruction features through a multi-layer perceptron respectively to obtain a text reconstruction feature vector and an image description reconstruction feature vector; The outer product of the text reconstruction feature vector and the image description reconstruction feature vector is calculated and a pooling operation is performed.

7. The multimodal sentiment analysis method based on image description according to claim 5, characterized in that: The late fusion is performed based on the average method, and the specific process is as follows: Inputting the text reconstruction feature and the image description reconstruction feature into a linear classifier respectively to obtain corresponding classifier decisions; The classifier decisions of the text reconstruction features and the image description reconstruction features are weighted averaged.

8. The multimodal sentiment analysis method based on image description according to claim 1, characterized in that: The classifier is a linear classifier.

9. An electronic device comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Flow casting material effect prediction method and system based on multi-modal deep learning

    CN120374200A

  • A method and system for predicting the effect of streaming material based on multimodal deep learning

    CN120374200B