Social media multi-mode sentiment classification method based on prompt driving and comparative learning

By introducing a prompt-driven mechanism and a comparison learning strategy in multimodal emotion classification, the problem of insufficient processing of modal noise and redundant information on social media text is solved, and higher accuracy and robustness of emotional classification are achieved.

CN120011565AActive Publication Date: 2025-05-16YUNNAN UNIV

Patent Information

Application Number
CN202510083399.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing multimodal emotion classification method ignores noise and redundant information when dealing with social media text modalities, and the design of comparative learning tasks is too simple to fully tap emotional clues in multimodal information.

Method used

Using a method based on prompt-driven and contrast learning, we use the text backbone structure prompt information to extract text and visual modal feature sequences, and fuse features in the attention layer to design contrast loss and consistency loss functions to optimize model performance.

Benefits of technology

It effectively reduces the impact of noise and redundant information in text mode on emotional classification, improves the accuracy and robustness of emotional classification, and enhances the model's ability to process complex emotional information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011565A_ABST
    Figure CN120011565A_ABST
Patent Text Reader

Abstract

The invention provides a social media multi-mode sentiment classification method based on prompt driving and comparative learning, and aims to improve the accuracy of sentiment classification in social media content. The text trunk structure is optimized by introducing prompt information, noise and redundant information in the text are removed, and the understanding ability of the model for complex emotion information is enhanced in combination with a visual data enhancement technology. Text and visual features are extracted and fused by using ResNet-50, RoBERTa, Transform and other models, the relation between modals is deepened through a contrast learning strategy, and finally the model performance is optimized through consistency loss and classification loss. The method is outstanding in performance when processing MVSA-Single, MVSA-Multi and HFM data sets, and a more accurate sentiment analysis tool is provided for the fields of social media monitoring, customer service, market strategies and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a social media multimodal sentiment classification method based on prompt-driven and contrastive learning. Background Art

[0002] With the rapid development of information technology, social media has become an important platform for people to express their feelings, attitudes and opinions. People are no longer limited to single text communication, but use multiple modal forms to express themselves in an all-round and multi-angle manner. Sentiment classification aims to identify and classify the emotional tendencies contained in each modal information and classify them into positive, negative or neutral. Most of the early sentiment classification models used single modal data to mine sentiment information, which easily led to the lack of context and deviation of classification results, and it was difficult to meet the needs of accurate and comprehensive sentiment recognition and classification. Multimodal sentiment classification can integrate multiple modal information sources, capture sentiment clues more comprehensively, and improve the accuracy of sentiment classification.

[0003] Although the existing multimodal sentiment classification methods have improved the accuracy of sentiment recognition to a certain extent, they still face some challenges. Among them, the text modality contains important sentiment clues and can provide more detailed and accurate sentiment descriptions. However, the text data in social media often contains incorrect syntactic structures and redundant information, which may introduce additional noise information. This information affects the accurate expression of the text and may even mislead the sentiment classification. It not only increases the complexity of model processing, but also may reduce the accuracy of sentiment recognition. The existing methods often ignore this when processing text modality information. Secondly, the existing multimodal sentiment classification methods based on contrastive learning often only focus on simple feature matching when designing contrastive learning tasks. Although it can improve the generalization ability of the model to a certain extent, the overly simple contrastive learning task may not be able to fully explore and utilize the sentiment clues in multimodal information, thereby limiting the performance of the model in sentiment classification. In addition, ignoring the contrastive learning between sentiment polarities may not enable the model to fully learn the subtle feature differences between samples of different sentiment polarities, thereby affecting the accuracy and robustness of the model in sentiment classification.

[0004] To this end, the present invention aims to solve the problem of insufficient processing of errors and redundant information in text modality, and optimize the design of contrastive learning tasks to apply multimodal sentiment classification technology to the fields of social media monitoring, customer service, and market strategy specification. In terms of social media monitoring, this technology can analyze the public's emotional tendencies towards specific brands, products, or events in real time, providing companies with valuable market insights; in the field of customer service, by analyzing the information posted by customers, companies can accurately identify the emotional state of customers, thereby adjusting service strategies and improving customer satisfaction; in terms of market strategy formulation, it helps companies gain a deep understanding of the audience's emotional response to advertising content, providing a scientific basis for optimizing marketing strategies. Summary of the invention

[0005] The purpose of the present invention is to provide a social media multimodal sentiment classification method based on prompt-driven and contrastive learning to solve the problem of insufficient processing of text modal noise and redundant information faced by users in social media when posting multimedia content such as informal language and emoticons.

[0006] The technical solution adopted by the present invention is a social media multimodal sentiment classification method based on prompt-driven and contrastive learning, comprising the following steps:

[0007] Step S1, obtaining text and visual modality data samples of social media, and preprocessing the data samples;

[0008] Step S2, constructing text trunk structure prompt information;

[0009] Step S3, extracting the original text and the text modality feature sequence after data enhancement, dividing the text data into sentences, and adding special identifiers [CLS] and [SEP] to mark the beginning and end of the sentence respectively;

[0010] Step S4, extracting a visual feature sequence for the visual modality of each branch;

[0011] Step S5, construct the network and input the fused features into an attention layer to deeply explore the potential connections between modalities;

[0012] Step S6, comparing and learning the original text-visual sample pairs after feature fusion and the text-visual sample pairs with added prompts with the text-visual sample pairs with data enhancement respectively;

[0013] Step S7, calculating the consistency loss of the original sample pair and the sample pair with the hint text added and the final classification loss, and further optimizing the model;

[0014] Step S8, obtaining a sentiment classification result through a multimodal sentiment classifier.

[0015] Furthermore, step S1 specifically includes:

[0016] S11, divide the training set, validation set and test set;

[0017] S12, removes visual-text pairs with inconsistent labels;

[0018] S13, uniformly process data of different modalities, use data enhancement technology to expand the data set, use synonym replacement, random insertion, and reverse translation data enhancement technology to expand the text modality data set, and use cropping, color transformation, rotation, and chromaticity separation data enhancement technology to expand the visual modality data set.

[0019] Furthermore, in step S2, the specific formula for constructing the text trunk structure prompt information is as follows:

[0020] P=Llama(t1,t2,…,t s )

[0021] Among them, P represents the text trunk structure prompt information constructed by the Llama3 model, t1, t2, …, t s represents the original text slice, s represents the length of the original text, and Llama(·) represents the process of generating prompt text information by the Llama3 model.

[0022] Furthermore, in step S3, the segmented text is converted into an ID list that can be recognized by the BERT model, which can be specifically expressed as:

[0023] T origin =token_to_id([[CLS],t1,t2,…,t s ,[SEP]])

[0024] T aug =token_to_id([[CLS],a1,a2,…,a k ,[SEP]])

[0025] Among them, T origin represents the original text after preprocessing, T aug Represents the text of data enhancement after preprocessing, t1, t2, …, t s represents the original text slice, s represents the length of the original text, a, a2,…, a k represents the text slice after data enhancement, k represents the length of the text after data enhancement, and token_to_id(·) represents the operation of converting the segmented text into an ID list recognizable by the BERT model;

[0026] Then, the text modal feature sequence of the prompt information added in step S2 is extracted, and the BERT word segmenter is used to divide the sentences. The original text and the prompt information are connected with the special mark [SEP], and the beginning and the end are marked by [CLS] and [SEP] respectively. The text with the prompt information after word segmentation is converted into an ID list that can be recognized by the BERT model, which can be specifically expressed as:

[0027] T prompt =token_to_id([[CLS],t1,t2,…,t s ,[SEP],p1,p2,…,p m ,[SEP]])

[0028] Among them, T prompt Represents the preprocessed text with added prompt information, t1, t2, …, t s represents the original text slice, s represents the length of the original text, p1, p2, ..., p m represents the prompt text slice generated by the Llama3 model, m represents the length of the prompt text, and token_to_id(·) represents the operation of converting the tokenized text into an ID list recognizable by the BERT model;

[0029] Then T origin , T aug and T prompt Input to the BERT model for encoding to extract the original text and the text feature sequence after data enhancement, which is specifically expressed as follows:

[0030]

[0031] in, Represents the original text feature sequence extracted by BERT, represents the text feature sequence after data enhancement extracted by BERT, represents the sequence of text features extracted by BERT with added prompts. BERT(·) represents the process of feature extraction using the BERT model. origin represents the original text after preprocessing, T aug represents the text after preprocessing data enhancement, T prompt Represents the text of the preprocessed prompt message.

[0032] Furthermore, in step S4, the specific visual modality feature sequence is expressed as:

[0033]

[0034]

[0035] Among them, I origin represents the original visual input, I aug represents enhanced visual input; ResNet(·) represents the use of ResNet-50 model for visual feature extraction, represents the original visual feature sequence extracted by the ResNet model, represents the enhanced visual feature sequence extracted by the ResNet model, RoBERTa(·) represents the use of the RoBERTa model to further extract visual modality features, represents the final extracted original visual feature sequence, Represents the final extracted visual feature sequence after data enhancement.

[0036] Furthermore, in step S5, the specific formula for achieving the fusion and alignment of cross-modal features and inputting the fused features into an attention layer to deeply explore the potential connection between the modalities is:

[0037]

[0038] Among them, Cat(·) represents the cascade operation, H represents the feature of simple connection of text-visual modality, and H o Represents the simple connection feature of the original sample pair, H a Represents the simple connection features of the sample pairs after data enhancement, H p represents the simple connection feature of the sample pair with the prompt, f represents the fusion feature obtained by Transformer, and f o represents the original text-visual fusion feature obtained by Transformer, f a represents the text-visual fusion feature after data enhancement obtained by Transformer, f p Q represents the text-visual fusion feature with added prompt information obtained by Transformer. f represents the query vector obtained by f, represents the key vector K obtained by f f The transposed vector, V f represents the value vector obtained by f, F represents the final fusion feature further obtained by the attention layer, and F o represents the fusion features of the original data obtained by the attention layer, F a represents the fusion features of the data after data enhancement obtained by the attention layer, F p represents the fused features of the prompted data obtained by the attention layer, represents the final extracted original visual feature sequence, represents the visual feature sequence after the final extraction of data enhancement, Represents the original text feature sequence extracted by BERT, represents the text feature sequence after data enhancement extracted by BERT, H t p represents the sequence of text features extracted by BERT with hints added, TF(·) represents the process of cross-modal feature fusion using the Transformer fusion network, and softmax(·) represents the normalization process.

[0039] Furthermore, in step S6, the specific representation of contrastive learning is as follows:

[0040]

[0041] Among them, F oa represents the matrix product between the original data sample and the data-enhanced data sample, F o represents the fusion features of the final extracted original text-visual pair, represents the transpose of the final extracted fusion features of the original text-visual pair, F pa represents the matrix product between the data sample with added prompt information and the data sample with data enhancement, F p represents the fusion feature of the text-visual pair with the added prompt information, τ represents the temperature parameter, and its value is 0.07, L oacl represents the contrast loss between the original data sample and the data-enhanced sample, L pacl represents the contrast loss between the data sample with added prompt information and the data augmented sample, y i =y j Represents samples with the same sentiment polarity, y i ≠y j Represents samples with different sentiment polarities, y i represents the i-th data sample, y j represents the jth data sample, Dot(·) represents the dot product similarity calculation process, n represents the number of batch samples, represents the matrix product of the i-th original data sample and the data-enhanced data sample, represents the matrix product of the jth original data sample and the data-enhanced data sample, represents the matrix product of the i-th data sample with added prompt information and the data enhanced data sample, Represents the matrix product of the jth data sample with added prompt information and the data augmented data sample.

[0042] Furthermore, in step S7, the specific steps of calculating the loss are as follows:

[0043] S71, the consistency loss between the original sample pair and the sample with the hint text added is specifically expressed as:

[0044]

[0045] Among them, L opcl Represents the consistency loss between the original sample pair and the sample pair with the hint text added, Represents the fusion features finally extracted from the i-th original sample. represents the fusion feature finally extracted by the sample with the i-th added prompt information, cos(·) represents the calculation of cosine similarity, n represents the number of batch samples, i represents the i-th sample, τ represents the temperature parameter, and its value is 0.07, γ1 represents the hyperparameter coefficient of the mean square error calculation process, and its value is 0.8, and γ2 represents the hyperparameter coefficient of the cosine similarity consistency loss calculation process, and its value is 0.2;

[0046] S72, calculate the final classification loss, which is specifically expressed as:

[0047]

[0048] Among them, L CE represents classification loss, CrossEntropyLoss(·) represents cross entropy loss, Represents the model prediction output, label represents the true label of the sample;

[0049] S73, calculate the total loss, which is specifically expressed as:

[0050] L total =L CE +λ1L opcl +λ2L oacl +λ3L pacl

[0051] Among them, L total represents the total loss, L CE represents the classification loss, L opcl represents the consistency loss between the original sample pair and the sample with the hint text added, L oacl represents the contrast loss between the original data sample and the data-enhanced sample, L pacl It represents the contrast loss between the data samples with the added prompt information and the data enhanced samples, λ1 represents the hyper-parameter coefficient of the contrast loss between the original data samples and the data enhanced samples, and its value is 0.4, λ2 represents the hyper-parameter coefficient of the contrast loss between the data samples with the added prompt information and the data enhanced samples, and its value is 0.3, λ3 represents the hyper-parameter coefficient of the consistency loss between the original samples and the samples with the added prompt text, and its value is 0.3.

[0052] Furthermore, in step S8, the sentiment classification result is specifically expressed as:

[0053]

[0054] Among them, FC(·) represents the fully connected layer, GeLU(·) represents the activation function, represents the model prediction output, F o Represents the fused features of the original data obtained by the attention layer.

[0055] The beneficial effects of the present invention are:

[0056] 1. The present invention effectively reduces the impact of noise and redundant information in social media texts on the accuracy of sentiment classification by introducing a prompt-driven mechanism and contrastive learning strategy, enabling the model to more accurately capture the sentiment clues in text and visual modal information, thereby greatly improving the accuracy of sentiment classification.

[0057] 2. The present invention fully considers the complementarity and correlation between multimodal data. By constructing contrast loss and consistency loss functions, it promotes the effective fusion of different modal features, which not only enhances the model's ability to process complex emotional information, but also enables the model to more comprehensively understand the emotional expression in social media content. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0059] Figure 1 Schematic diagram of the process of the multimodal sentiment classification method in an embodiment of the present invention.

[0060] Figure 2 Schematic diagram of the structure of the multimodal sentiment classification model in an embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] Example 1

[0063] like Figures 1-2 As shown, an embodiment of the present invention provides a social media multimodal sentiment classification method based on prompt-driven and contrastive learning, and the specific steps include:

[0064] Step S1: obtaining text and visual modality data samples of social media, denoted as T and I respectively, and preprocessing the data samples;

[0065] S11, divide the training set, validation set and test set into a ratio of 8:1:1;

[0066] S12, remove visual-text pairs with inconsistent labels. Specifically, for the MVSA-Single dataset marked by one annotator, if the labels of the text and visual modalities indicate positive and negative respectively, the sample pair is considered invalid and needs to be removed; for the MVSA-Multiple dataset marked by three annotators, in the single-modal label, the label is considered to be a valid label of the modality when at least two of the three annotators have the same label. Then, samples with inconsistent text and visual modality labels are removed to reduce the impact on the model.

[0067] S13, uniformly process data of different modalities, use data enhancement technology to expand the data set, use synonym replacement, random insertion, and reverse translation data enhancement technology to expand the text modality data set; use cropping, color transformation, rotation, and chromaticity separation data enhancement technology to expand the visual modality data set.

[0068] In this embodiment, the text and visual modality datasets about social media include MVSA-Single, MVSA-Multiple and HFM. Among them, the MVSA dataset is a three-category task (positive, neutral and negative), MVSA-Single contains 5129 visual-text pairs, each pair of samples is labeled by one annotator, and MVSA-Multiple contains 19600 visual-text pairs, each pair of samples is labeled by three annotators. HFM is a two-category task dataset (positive and negative), containing 24635 visual-text pairs. In order to reduce the adverse effects of erroneous samples, samples with inconsistent visual-text labels in the MVSA dataset are removed, and then the training set, validation set and test set are divided. The final dataset statistical results are shown in Table 1.

[0069] Table 1 Statistical results of different data sets

[0070] Dataset category Positive Neutral Negative total MVSA-Single 3 2683 470 1358 4511 MVSA-Multiple 3 11320 4408 1088 16816 HFM 2 10560 -- 14075 24635

[0071] Step S2: Use the Llama3 model to construct text backbone structure prompt information, which is specifically expressed as follows:

[0072] P=Llama(t1,t2,…,t s )

[0073] Among them, P represents the text trunk structure prompt information constructed by the Llama3 model, t1, t2, …, t s represents the original text slice, s represents the length of the original text, and Llama(·) represents the process of generating prompt text information by the Llama3 model.

[0074] P is connected to the original text as a prompt information to form a data sample pair branch, thereby enhancing the model's understanding of the text and alleviating the adverse effects caused by text grammatical errors, interference and redundant information. For example, the text "MT@explincolnshire@LostVillages Holloway,once abusyroad,atGainsthorpe#deserted#medieval village,Lincs" published on social media contains redundant information such as "MT@explincolnshire" and "#deserted", which is not conducive to the model's understanding of the correct meaning of the text and may be an incorrect grammatical structure introduced by the text. The text trunk structure information extracted by the Llama3 model is: "Hollowaywas once abusyroad". The generated trunk structure information removes the interference and redundant information in the original text and extracts the main content, which helps the model understand the text.

[0075] Step S3: Extract the original text and the text modal feature sequence after data enhancement, use the BERT tokenizer to segment the text data into sentences, add special identifiers [CLS] and [SEP] to mark the beginning and end of the sentence respectively, and convert the segmented text into an ID list that can be recognized by the BERT model, which can be specifically expressed as:

[0076] T origin =token_to_id([[CLS],t1,t2,…,t s ,[SEP]])

[0077] T aug =token_to_id([[CLS],a1,a2,…,a k ,[SEP]])

[0078] Among them, T origin represents the original text after preprocessing, T aug Represents the text of data enhancement after preprocessing, t1, t2, …, t s represents the original text slice, s represents the length of the original text, a, a2,…, ak represents the text slice after data augmentation, k represents the length of the text after data augmentation, and token_to_id(·) represents the operation of converting the tokenized text into an ID list recognizable by the BERT model.

[0079] The text modal feature sequence of the prompt information added in step S2 is extracted, and the BERT word segmenter is used to divide the sentences. The original text and the prompt information are connected with a special mark [SEP], and the beginning and end are marked by [CLS] and [SEP] respectively. The text with the prompt information after word segmentation is converted into an ID list that can be recognized by the BERT model, which can be specifically expressed as:

[0080] T prompt =token_to_id([[CLS],t1,t2,…,t s ,[SEP],p1,p2,…,p m ,[SEP]])

[0081] Among them, T prompt Represents the preprocessed text with added prompt information, t1, t2, …, t s represents the original text slice, s represents the length of the original text, p1, p2, ..., p m represents a prompt text slice generated by the Llama3 model, m represents the length of the prompt text, and token_to_id(·) represents the operation of converting the tokenized text into a list of IDs recognizable by the BERT model.

[0082] Then T origin , T aug and T prompt Input to the BERT model for encoding to extract the original text and the text feature sequence after data enhancement, which is specifically expressed as follows:

[0083]

[0084] in, Represents the original text feature sequence extracted by BERT, represents the text feature sequence after data enhancement extracted by BERT, represents the sequence of text features extracted by BERT with added prompts. BERT(·) represents the process of feature extraction using the BERT model. origin represents the original text after preprocessing, T aug represents the text after preprocessing data enhancement, T prompt Represents the text of the preprocessed prompt message.

[0085] Step S4: ResNet-50 and RoBERTa models are used to extract visual feature sequences for the visual modality of each branch. The specific visual modality feature sequence is expressed as:

[0086]

[0087] Among them, I origin represents the original visual input, I aug represents enhanced visual input; ResNet(·) represents the use of ResNet-50 model for visual feature extraction, represents the original visual feature sequence extracted by the ResNet model, represents the enhanced visual feature sequence extracted by the ResNet model; RoBERTa(·) represents the use of the RoBERTa model to further extract visual modality features; represents the final extracted original visual feature sequence, Represents the final extracted visual feature sequence after data enhancement.

[0088] Step S5: Build a Transformer-based network to achieve the fusion and alignment of cross-modal features, and input the fused features into an attention layer to deeply explore the potential connections between modalities. Specifically, it is expressed as:

[0089]

[0090] Among them, Cat(·) represents the cascade operation, H represents the feature of simple connection of text-visual modality, and H o Represents the simple connection feature of the original sample pair, H a Represents the simple connection features of the sample pairs after data enhancement, H p represents the simple connection feature of the sample pair with the prompt, f represents the fusion feature obtained by Transformer, and f o represents the original text-visual fusion feature obtained by Transformer, f a represents the text-visual fusion feature after data enhancement obtained by Transformer, f p Q represents the text-visual fusion feature with added prompt information obtained by Transformer. f represents the query vector obtained by f, represents the key vector K obtained by f f The transposed vector, V f represents the value vector obtained by f, F represents the final fusion feature further obtained by the attention layer, and F o represents the fusion features of the original data obtained by the attention layer, Fa represents the fusion features of the data after data enhancement obtained by the attention layer, F p represents the fused features of the prompted data obtained by the attention layer, represents the final extracted original visual feature sequence, represents the visual feature sequence after the final extraction of data enhancement, Represents the original text feature sequence extracted by BERT, represents the text feature sequence after data enhancement extracted by BERT, represents the sequence of text features extracted by BERT with hints added, TF(·) represents the process of cross-modal feature fusion using the Transformer fusion network, and softmax(·) represents the normalization process.

[0091] Step S6: The original text-visual sample pairs after feature fusion and the text-visual sample pairs with added prompts are respectively compared with the text-visual sample pairs with data enhancement. By comparing positive samples and negative samples, the effective representation of data is learned, thereby improving the performance of the model in the sentiment classification task, so that the distance between samples with the same sentiment polarity in the feature space is getting closer and closer, while the distance between samples with different sentiment polarities is getting farther and farther, which is specifically expressed as:

[0092]

[0093]

[0094] Among them, F oa represents the matrix product between the original data sample and the data-enhanced data sample, F o represents the fusion features of the final extracted original text-visual pair, represents the transpose of the final extracted fusion features of the original text-visual pair, F pa represents the matrix product between the data sample with added prompt information and the data sample with data enhancement, F p represents the fusion feature of the text-visual pair with the added prompt information, τ represents the temperature parameter, and its value is 0.07, L oacl represents the contrast loss between the original data sample and the data-enhanced sample, L pacl represents the contrast loss between the data sample with added prompt information and the data augmented sample, y i =y j Represents samples with the same sentiment polarity, y i ≠y j Represents samples with different sentiment polarities, y i represents the i-th data sample, yj represents the jth data sample, Dot(·) represents the dot product similarity calculation process, n represents the number of batch samples, represents the matrix product of the i-th original data sample and the data-enhanced data sample, represents the matrix product of the jth original data sample and the data-enhanced data sample, represents the matrix product of the i-th data sample with added prompt information and the data enhanced data sample, Represents the matrix product of the jth data sample with added prompt information and the data augmented data sample.

[0095] Step S7: Calculate the consistency loss of the original sample pair and the sample pair with the hint text added and the final classification loss to further optimize the model;

[0096] S71, the consistency loss between the original sample pair and the sample with the hint text added is specifically expressed as:

[0097]

[0098] Among them, L opcl Represents the consistency loss between the original sample pair and the sample pair with the hint text added, Represents the fusion features finally extracted from the i-th original sample. represents the fusion feature finally extracted from the i-th sample with the hint information added, cos(·) represents the calculation of cosine similarity, n represents the number of batch samples, i represents the i-th sample, τ represents the temperature parameter, and its value is 0.07, γ1 represents the hyperparameter coefficient of the mean square error calculation process, and its value is 0.8, and γ2 represents the hyperparameter coefficient of the cosine similarity consistency loss calculation process, and its value is 0.2.

[0099] S72, calculate the final classification loss, which is specifically expressed as:

[0100]

[0101] Among them, L CE represents classification loss, CrossEntropyLoss(·) represents cross entropy loss, Represents the model prediction output, and label represents the true label of the sample.

[0102] S73, calculate the total loss, which is specifically expressed as:

[0103] L total =L CE +λ1L opcl +λ2L oacl +λ3L pacl

[0104] Among them, L total represents the total loss, L CE represents the classification loss, L opcl represents the consistency loss between the original sample pair and the sample with the hint text added, L oacl represents the contrast loss between the original data sample and the data-enhanced sample, L pacl It represents the contrast loss between the data samples with the added prompt information and the data enhanced samples, λ1 represents the hyper-parameter coefficient of the contrast loss between the original data samples and the data enhanced samples, and its value is 0.4, λ2 represents the hyper-parameter coefficient of the contrast loss between the data samples with the added prompt information and the data enhanced samples, and its value is 0.3, λ3 represents the hyper-parameter coefficient of the consistency loss between the original samples and the samples with the added prompt text, and its value is 0.3.

[0105] Step S8: Obtain sentiment classification results through a multimodal sentiment classifier, which is specifically expressed as:

[0106]

[0107] Among them, FC(·) represents the fully connected layer, GeLU(·) represents the activation function, represents the model prediction output, F o Represents the fused features of the original data obtained by the attention layer.

[0108] Sentiment classification is used to analyze the sentiment of multimodal data in the fields of social media user comments, customer service, and market strategy formulation. For binary classification tasks, the classification results are positive or negative; for ternary classification tasks, the classification results are positive, negative, or neutral. Through sentiment classification, we can further understand the emotional reactions of user groups to improve service satisfaction.

[0109] Experimental verification

[0110] In order to verify the effectiveness of the model proposed in this paper, this paper conducted comparative experiments on three multimodal sentiment classification datasets: MVSA-Single, MVSA-Multiple and HFM. Tables 2 and 3 summarize the performance comparison of the baseline methods on these three datasets. The performance of all baselines is derived from the open source code and parameters given in the paper.

[0111] Table 2: Experimental results of each model on the MVSA dataset

[0112]

[0113]

[0114] Table 3: Experimental results of each model on the HFM dataset

[0115]

[0116] It can be seen from the results that most multimodal sentiment classification models outperform unimodal models on the three datasets, indicating that multimodal data provides more effective information for sentiment classification tasks. At the same time, the present invention has good competitiveness among a series of baseline models on the MVSA and HFM datasets. Specifically, as shown in Table 2, in the three-category sentiment classification task, the present invention outperforms other methods in ACC and F1 on the MVSA-Single dataset, where the accuracy is 8.38% higher than MultiSentiNet and 2.89% higher than CLMLF. For the MVSA-Multiple dataset, the accuracy of this method is 2.27% higher than the ITIN method and 6.93% higher than the MultiSentiNet method. As shown in Table 3, in the two-category multimodal sentiment classification task, the present invention outperforms other baseline models, and the accuracy ACC and F1 are increased by 1.29% and 1.81% respectively compared with CLMLF.

[0117] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0118] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A social media multimodal sentiment classification method based on cue-driven and contrastive learning, characterized in that: The following steps are involved: Step S1, obtaining text and visual modality data samples of social media, and preprocessing the data samples; Step S2, constructing text trunk structure prompt information; Step S3, extracting the original text and the text modality feature sequence after data enhancement, dividing the text data into sentences, and adding special identifiers [CLS] and [SEP] to mark the beginning and end of the sentence respectively; Step S4, extracting a visual feature sequence for the visual modality of each branch; Step S5, construct the network and input the fused features into an attention layer to deeply explore the potential connections between modalities; Step S6, comparing and learning the original text-visual sample pairs after feature fusion and the text-visual sample pairs with added prompts with the text-visual sample pairs with data enhancement respectively; Step S7, calculating the consistency loss of the original sample pair and the sample pair with the hint text added and the final classification loss, and further optimizing the model; Step S8, obtaining a sentiment classification result through a multimodal sentiment classifier.

2. According to claim 1, a method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning is characterized in that: The step S1 specifically includes: S11, divide the training set, validation set and test set; S12, removes visual-text pairs with inconsistent labels; S13, uniformly process data of different modalities, use data enhancement technology to expand the data set, use synonym replacement, random insertion, and reverse translation data enhancement technology to expand the text modality data set, and use cropping, color transformation, rotation, and chromaticity separation data enhancement technology to expand the visual modality data set.

3. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S2, the specific formula for constructing the text trunk structure prompt information is as follows: P=Call(t1,t2,…,t s ) Among them, P represents the text trunk structure prompt information constructed by the Llama3 model, t1, t2, …, t s represents the original text slice, s represents the length of the original text, and Llama(·) represents the process of generating prompt text information by the Llama3 model.

4. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S3, the segmented text is converted into an ID list that can be recognized by the BERT model, which can be specifically expressed as: T origin =token_to_id([[CLS],t1,t2,…,t s ,[SEP]]) T aug =token_to_id([[CLS],a1,a2,…,a k ,[SEP]]) Among them, T origin represents the original text after preprocessing, T aug Represents the text of data enhancement after preprocessing, t1, t2, …, t s represents the original text slice, s represents the length of the original text, a, a2,…, a k represents the text slice after data enhancement, k represents the length of the text after data enhancement, and token_to_id(·) represents the operation of converting the segmented text into an ID list recognizable by the BERT model; Then, the text modal feature sequence of the prompt information added in step S2 is extracted, and the BERT word segmenter is used to divide the sentences. The original text and the prompt information are connected with the special mark [SEP], and the beginning and the end are marked by [CLS] and [SEP] respectively. The text with the prompt information after word segmentation is converted into an ID list that can be recognized by the BERT model, which can be specifically expressed as: T prompt =token_to_id([[CLS],t1,t2,…,t s ,[SEP],p1,p2,…,p m ,[SEP]]) Among them, T prompt Represents the preprocessed text with added prompt information, t1, t2, …, t s represents the original text slice, s represents the length of the original text, p1, p2, ..., p m represents the prompt text slice generated by the Llama3 model, m represents the length of the prompt text, and token_to_id(·) represents the operation of converting the tokenized text into an ID list recognizable by the BERT model; Then T origin , T aug and T prompt Input to the BERT model for encoding to extract the original text and the text feature sequence after data enhancement, which is specifically expressed as follows: in, Represents the original text feature sequence extracted by BERT, represents the text feature sequence after data enhancement extracted by BERT, represents the sequence of text features extracted by BERT with added prompts. BERT(·) represents the process of feature extraction using the BERT model. origin represents the original text after preprocessing, T aug represents the text after preprocessing data enhancement, T prompt Represents the text of the preprocessed added prompt information.

5. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S4, the specific visual modality feature sequence is expressed as: Among them, I origin represents the original visual input, I aug represents enhanced visual input; ResNet(·) represents the use of ResNet-50 model for visual feature extraction, represents the original visual feature sequence extracted by the ResNet model, represents the enhanced visual feature sequence extracted by the ResNet model, RoBERTa(·) represents the use of the RoBERTa model to further extract visual modality features, represents the final extracted original visual feature sequence, Represents the final extracted visual feature sequence after data enhancement.

6. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S5, the specific formula for achieving the fusion and alignment of cross-modal features and inputting the fused features into an attention layer to deeply explore the potential connection between the modalities is: f=TF(H),f∈{f o ,f a ,f p },H∈{H o ,H a ,H p } Among them, Cat(·) represents the cascade operation, H represents the feature of simple connection of text-visual modality, and H o Represents the simple connection feature of the original sample pair, H a Represents the simple connection features of the sample pairs after data enhancement, H p represents the simple connection feature of the sample pair with the prompt, f represents the fusion feature obtained by Transformer, and f o represents the original text-visual fusion feature obtained by Transformer, f a represents the text-visual fusion feature after data enhancement obtained by Transformer, f p Q represents the text-visual fusion feature with added prompt information obtained by Transformer. f represents the query vector obtained by f, represents the key vector K obtained by f f The transposed vector, V f represents the value vector obtained by f, F represents the final fusion feature further obtained by the attention layer, and F o represents the fusion features of the original data obtained by the attention layer, F a represents the fusion features of the data after data enhancement obtained by the attention layer, F p represents the fused features of the prompted data obtained by the attention layer, represents the final extracted original visual feature sequence, represents the visual feature sequence after the final extraction of data enhancement, Represents the original text feature sequence extracted by BERT, represents the text feature sequence after data enhancement extracted by BERT, represents the sequence of text features extracted by BERT with hints added, TF(·) represents the process of cross-modal feature fusion using the Transformer fusion network, and softmax(·) represents the normalization process.

7. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S6, the specific representation of contrastive learning is as follows: Among them, F oa represents the matrix product between the original data sample and the data-enhanced data sample, F o represents the fusion features of the final extracted original text-visual pair, represents the transpose of the final extracted fusion features of the original text-visual pair, F pa represents the matrix product between the data sample with added prompt information and the data sample with data enhancement, F p represents the fusion feature of the text-visual pair with the added prompt information, τ represents the temperature parameter, and its value is 0.07, L oacl represents the contrast loss between the original data sample and the data-enhanced sample, L pacl represents the contrast loss between the data sample with added prompt information and the data augmented sample, y i =y j Represents samples with the same sentiment polarity, y i ≠y j Represents samples with different sentiment polarities, y i represents the i-th data sample, y j represents the jth data sample, Dot(·) represents the dot product similarity calculation process, n represents the number of batch samples, represents the matrix product of the i-th original data sample and the data-enhanced data sample, represents the matrix product of the jth original data sample and the data-enhanced data sample, represents the matrix product of the i-th data sample with added prompt information and the data enhanced data sample, Represents the matrix product of the jth data sample with added prompt information and the data augmented data sample.

8. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S7, the specific steps of calculating the loss are as follows: S71, the consistency loss between the original sample pair and the sample with the hint text added is specifically expressed as: Among them, L opcl Represents the consistency loss between the original sample pair and the sample pair with the hint text added, Represents the fusion features finally extracted from the i-th original sample. represents the fusion feature finally extracted by the sample with the i-th added prompt information, cos(·) represents the calculation of cosine similarity, n represents the number of batch samples, i represents the i-th sample, τ represents the temperature parameter, and its value is 0.07, γ1 represents the hyperparameter coefficient of the mean square error calculation process, and its value is 0.8, and γ2 represents the hyperparameter coefficient of the cosine similarity consistency loss calculation process, and its value is 0.2; S72, calculate the final classification loss, which is specifically expressed as: Among them, L CE represents classification loss, CrossEntropyLoss(·) represents cross entropy loss, Represents the model prediction output, label represents the true label of the sample; S73, calculate the total loss, which is specifically expressed as: L total =L CE +λ1L opcl +λ2L oacl +λ3L pacl Among them, L total represents the total loss, L CE represents the classification loss, L opcl represents the consistency loss between the original sample pair and the sample with the hint text added, L oacl represents the contrast loss between the original data sample and the data-enhanced sample, L pacl It represents the contrast loss between the data samples with the added prompt information and the data enhanced samples, λ1 represents the hyper-parameter coefficient of the contrast loss between the original data samples and the data enhanced samples, and its value is 0.4, λ2 represents the hyper-parameter coefficient of the contrast loss between the data samples with the added prompt information and the data enhanced samples, and its value is 0.3, λ3 represents the hyper-parameter coefficient of the consistency loss between the original samples and the samples with the added prompt text, and its value is 0.

3.

9. The method for multimodal sentiment classification of social media based on prompt-driven and contrastive learning according to claim 1, characterized in that: In step S8, the sentiment classification result is specifically expressed as: Among them, FC(·) represents the fully connected layer, GeLU(·) represents the activation function, represents the model prediction output, F o Represents the fused features of the original data obtained by the attention layer.

Citation Information

Patent Citations

  • Emotion enhancement continuous training method combining knowledge distillation and comparative learning

    CN117115505A

  • Multi-modal sentiment classification method based on comparative learning and aspect enhancement

    CN117407525A

  • Multi-modal sentiment analysis method based on multi-granularity feature comparison and fusion framework

    CN117893948A

  • Method for embedding data and system thereof

    US20240020578A1

Cited By

  • Few-sample multi-mode sentiment classification method based on prompt tuning and contrast decoding

    CN122045427A

  • Few-shot multi-modal sentiment classification method based on prompt tuning and contrastive decoding

    CN122045427B