Multi-modal sentiment analysis and understanding method for social media comments

By combining BERT, EMOJIALL and Visual Transformer to extract the multimodal emotional characteristics in social media comments, the problem of integrating multiple information forms in social media comments is solved, and efficient and accurate sentiment analysis is achieved to adapt to the needs of large-scale data analysis.

CN120579012APending Publication Date: 2025-09-02SHENZHEN KIM DAI INTELLIGENCE INNOVATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510504411.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately integrate and analyze multiple forms of information in social media comments, especially emotional information in text, emojis and pictures, resulting in poor emotional analysis and inability to adapt to the rapidly changing large-scale data analysis needs.

Method used

The BERT model is used to extract text emotional features, the EMOJIALL website extracts emoji emotional features, the convolutional neural network or visual Transformer extracts image emotional features, and multimodal data emotional features are formed through cross-modal alignment and fusion technology, and emotional analysis is used for multimodal large language model to generate emotional tags.

Benefits of technology

It improves the accuracy and comprehensiveness of sentiment analysis, can efficiently process multimodal data in social media comments, provide fast and accurate emotional classification results, and adapt to large-scale data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579012A_ABST
    Figure CN120579012A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal sentiment analysis and understanding method for social media comments, which comprises the following steps: respectively extracting text sentiment features, expression sentiment features and picture sentiment features from an input social media comment screenshot, and fusing the sentiment features to form multi-modal data sentiment features; the classification model carries out sentiment analysis on multi-modal data sentiment features to obtain sentiment labels of the comments, multi-modal information such as texts, emoticons and pictures in the social media comments is effectively fused, joint training is carried out by utilizing a large model, the relevance among various types of data can be fully mined during sentiment analysis, and the sentiment labels of the comments are obtained. And the accuracy and comprehensiveness of sentiment analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a multimodal sentiment analysis and understanding method for social media comments. [Background Technology]

[0002] The ubiquity of social media and mobile devices has led to the emergence of a vast amount of user-generated content, particularly in the comment sections of social platforms. This content includes not only rich text but also a variety of non-text elements such as emoticons and images. These elements carry important information such as emotion, attitude, and context, making them crucial for a comprehensive understanding of user sentiment and public opinion trends.

[0003] However, in the field of multimodal sentiment analysis, how to efficiently and accurately process these complex non-text objects still faces many challenges. Although natural language processing and computer vision technologies have made significant progress in text sentiment analysis and image recognition, there are still many technical bottlenecks in the fusion of multimodal data and sentiment analysis.

[0004] In this context, the industry urgently needs a unified sentiment analysis framework that can effectively integrate multiple forms of information. This will not only improve processing efficiency, but also comprehensively enhance the accuracy and real-time performance of sentiment analysis. At the same time, existing models are usually limited to single-task analysis when processing rich and diverse social media content, and are unable to fully utilize the sentiment information of non-text elements such as emoticons and pictures, resulting in poor sentiment analysis results. With the diversification of content on social platforms and the rapid changes in information flows, existing methods are difficult to adapt to the needs of large-scale data analysis. There is an urgent need for an efficient and flexible multimodal sentiment analysis method to cope with the complex challenges of social media comments and public opinion monitoring. At present, how to develop a unified multimodal sentiment analysis framework to efficiently integrate and analyze multiple forms of information in social media comments has become one of the research hotspots in the industry. [Summary of the invention]

[0005] The present invention overcomes the shortcomings of the prior art and provides a multimodal sentiment analysis and understanding method for social media comments.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A multimodal sentiment analysis and understanding method for social media comments, characterized by:

[0008] S1. Extract text emotion features, expression emotion features, and image emotion features from the input social media comment screenshots, and fuse these emotion features to form multimodal data emotion features.

[0009] S2. The classification model performs sentiment analysis on the sentiment features of the multimodal data to obtain the sentiment label of the comment.

[0010] The multimodal sentiment analysis and understanding method for social media comments as described above is characterized in that: in S1, text sentiment features are extracted from the input social media comments through the BERT model.

[0011] The multimodal sentiment analysis and understanding method for social media comments described above is characterized in that the BERT model is pre-trained using a dataset containing text and sentiment labels.

[0012] The multimodal sentiment analysis and understanding method for social media comments as described above is characterized in that: in S1, expression and emotion features are extracted from input social media comments through the EMOJIALL website.

[0013] The multimodal sentiment analysis and understanding method for social media comments described above is characterized by: the EMOJIALL website provides a sentiment analysis graph for each emoticon, including the proportion of negative, neutral, and positive emotions, and generates a mapping from the emoticon to a vector embedding based on the proportion to obtain the emoticon emotion feature.

[0014] The multimodal sentiment analysis and understanding method for social media comments as described above is characterized in that: S1 extracts image sentiment features from input social media comments through a convolutional neural network or a visual Transformer.

[0015] The multimodal sentiment analysis and understanding method for social media comments described above is characterized in that a visual Transformer is pre-trained using a dataset containing comment images and labels.

[0016] The multimodal sentiment analysis and understanding method for social media comments as described above is characterized in that: S1 adopts cross-modal alignment and fusion technology to fuse text sentiment features, expression sentiment features and image sentiment features into the representation space to form multimodal data sentiment features.

[0017] The multimodal sentiment analysis and understanding method for social media comments as described above is characterized in that the multimodal large language model in S2 is trained by a data set containing multimodal data sentiment features and sentiment labels to obtain a classification module.

[0018] The multimodal sentiment analysis and understanding method for social media comments as described above is characterized in that the sentiment labels of the comments in S2 include positive sentiment labels, negative sentiment labels, or neutral sentiment labels.

[0019] The beneficial effects of the present invention are:

[0020] This invention effectively integrates multimodal information such as text, emoticons, and images in social media comments, and uses a large model for joint training. It can fully explore the correlation between various types of data when performing sentiment analysis, thereby improving the accuracy and comprehensiveness of sentiment analysis. [Brief Description of the Drawings]

[0021] Figure 1 It is a feature extraction schematic diagram of the present invention;

[0022] Figure 2 It is a schematic diagram of the model architecture of the present invention;

[0023] Figure 3 It is a prototype system diagram designed by the present invention. [Specific implementation method]

[0024] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings.

[0025] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back...) are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly. In addition, the descriptions of "preferred", "sub-preferred", etc. in the present invention are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "preferred" or "sub-preferred" may explicitly or implicitly include at least one such feature.

[0026] The present invention provides a multimodal sentiment analysis and understanding method for social media comments, and proposes an MMESA-SMC model, which includes a feature extraction module and a sentiment analysis module.

[0027] like Figure 1 As shown, the feature extraction module.

[0028] S1. The feature extraction module extracts text emotion features, expression emotion features, and image emotion features from the input social media comment screenshots, and fuses the above emotion features to form multimodal data emotion features.

[0029] First, for text, the BERT model is trained on a dataset of text and sentiment labels, allowing the trained BERT model to extract text sentiment features from the text in the comments. For emoticons, analysis is performed using the EMOJIALL website, which provides a sentiment analysis graph for each emoticon, including the proportion of negative, neutral, and positive sentiment. Based on the sentiment proportions, a mapping from emoticons to vector embeddings is generated to obtain image sentiment features. For images, a convolutional neural network (CNN) or a visual transformer (ViT) is used to extract image sentiment features from the images in the comments. The visual transformer (ViT) is trained on a dataset containing comment images and sentiment labels. The above text sentiment features, expression sentiment features, and image sentiment features are processed and mapped into a unified feature space in the feature extraction module, where they are fused to form multimodal data sentiment features to support subsequent sentiment analysis.

[0030] like Figure 2 As shown, sentiment analysis module.

[0031] S2. The classification model in the sentiment analysis module performs sentiment analysis on the sentiment features of the multimodal data to obtain the sentiment label of the comment. The multimodal large language model is trained using a dataset containing the sentiment features and sentiment labels of the multimodal data to obtain the classification module.

[0032] During the training phase of the multimodal large language model, the input social media comment screenshots are first processed through the feature extraction module to extract text sentiment features, emoticon sentiment features, and image sentiment features. These multimodal features are then fused and fed into the large language model along with sentiment labels for training. During the prediction phase, the trained multimodal large language model is converted into a classification model, which performs sentiment analysis on the comments in the input comment screenshots and outputs sentiment labels such as positive, negative, or neutral. Through joint learning of multimodal data, the sentiment analysis module is able to perform collaborative analysis based on data from different modalities, providing more accurate and comprehensive sentiment analysis results.

[0033] like Figure 3 As shown in the figure, this case designed a prototype system that allows users to upload screenshots of social media comments through a simple interface. The system automatically performs sentiment analysis on the uploaded comment screenshots and provides sentiment classification results. Testing of the prototype system will verify the performance of the MMESA-SMC model in real-world applications, ensuring its efficiency and accuracy.

[0034] The MMESA-SMC model proposed in this case aims to efficiently analyze and understand multimodal sentiment data in social media comments, including text, emoticons, and images. The model consists of a feature extraction module and a sentiment analysis module, which work closely together to significantly improve the accuracy and processing efficiency of sentiment analysis. First, by fine-tuning all parameters on multiple labeled social media comment datasets, the large language model's capabilities in multimodal sentiment analysis are effectively enhanced. The datasets used cover tasks such as text analysis, emoticon parsing, and image recognition. The model effectively transfers knowledge between modalities while maintaining high performance, significantly improving its generalization in multimodal sentiment analysis tasks. Second, based on the fine-tuned large language model, a sentiment analysis module is designed to accurately analyze multimodal information such as text, emoticons, and images in social media comments. Whether performing sentiment classification on text or analyzing the emotional information in emoticons and images, MMESA-SMC efficiently accomplishes this task, demonstrating strong adaptability and scalability. This model enables users to quickly obtain accurate sentiment classification results and supports real-time processing of large-scale data. Furthermore, this case also built a corresponding prototype system that can conduct multimodal sentiment analysis tests on user-uploaded social media comments, further validating the model's practicality and effectiveness. In summary, this case provides a new technical solution for sentiment analysis of social media comments, advancing the development of multimodal sentiment analysis technology. This technology demonstrates strong practical value and technical potential in areas such as public opinion monitoring, social media analysis, and automated sentiment recognition, and has broad application prospects.

[0035] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. All equivalent structural transformations made based on the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are included in the patent protection scope of the present invention.

Claims

1. A multimodal sentiment analysis and understanding method for social media comments, characterized by: Including S1. Extract text emotion features, expression emotion features, and image emotion features from the input social media comment screenshots, and fuse these emotion features to form multimodal data emotion features. S2. The classification model performs sentiment analysis on the sentiment features of the multimodal data to obtain the sentiment label of the comment.

2. The multimodal sentiment analysis and understanding method for social media comments according to claim 1, characterized in that: In S1, the BERT model is used to extract text sentiment features from input social media comments.

3. The multimodal sentiment analysis and understanding method for social media comments according to claim 2, characterized in that: The BERT model is pre-trained on a dataset containing text and sentiment labels.

4. The multimodal sentiment analysis and understanding method for social media comments according to claim 1, characterized in that: In S1, the expression emotion features are extracted from the input social media comments through the EMOJIALL website.

5. The multimodal sentiment analysis and understanding method for social media comments according to claim 4, characterized in that: The EMOJIALL website provides a sentiment analysis graph for each emoji, including the proportion of negative, neutral, and positive emotions. Based on the proportion, a mapping from emoji to vector embedding is generated to obtain the emotional features of the emoji.

6. The multimodal sentiment analysis and understanding method for social media comments according to claim 1, characterized in that: In S1, image sentiment features are extracted from input social media comments through convolutional neural networks or visual Transformers.

7. The multimodal sentiment analysis and understanding method for social media comments according to claim 6, characterized in that: The visual transformer is pre-trained on a dataset containing reviewed images and labels.

8. The multimodal sentiment analysis and understanding method for social media comments according to claim 1, characterized in that: S1 uses cross-modal alignment and fusion technology to integrate text emotion features, expression emotion features, and image emotion features into the representation space to form multimodal data emotion features.

9. The multimodal sentiment analysis and understanding method for social media comments according to claim 1, characterized in that: The multimodal large language model in S2 is trained with a dataset containing multimodal data sentiment features and sentiment labels to obtain a classification module.

10. The multimodal sentiment analysis and understanding method for social media comments according to claim 1, characterized in that: The sentiment labels of comments in S2 include positive sentiment labels, negative sentiment labels, or neutral sentiment labels.