A media stance identification method based on modal specific representation learning

By combining adversarial learning and self-supervised multi-task learning, modality-specific and consistency information is extracted from multimodal data, solving the problem of insufficient modality feature capture in multimodal sentiment stance recognition, improving recognition accuracy, reducing manual annotation costs, and promoting the efficiency of online public opinion supervision.

CN116992261BActive Publication Date: 2026-03-31TIANJIN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing multimodal sentiment and stance recognition technologies struggle to effectively capture modality-specific information and intermodal complementary information, resulting in insufficient recognition accuracy. Furthermore, the reliance on manual annotation is costly, making it difficult to achieve efficient public opinion guidance management in online public opinion supervision.

Method used

An adversarial learning mechanism is used to extract intermodal consistency and intramodal difference information. Combined with a self-supervised multi-task learning strategy, a cross-modal attention mechanism and modality-specific feature fusion are used to jointly train the multimodal representation using network-generated single-modal labels.

Benefits of technology

It improves the accuracy of media data sentiment identification, reduces the cost of manual annotation, enables efficient management of public opinion guidance for media content, and promotes the development of cybersecurity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992261B_ABST
    Figure CN116992261B_ABST
Patent Text Reader

Abstract

The application discloses a media stance recognition method based on modal specific feature learning, comprising the following steps: extracting modal consistent features and modal specific features from multi-modal data by using adversarial learning; inputting the modal specific features into label generation based on a self-supervised learning strategy to obtain independent single-modal supervision; using a cross-modal attention mechanism on the modal specific features to obtain inter-modal complementary information with the text modality as the main modality, and updating the modal specific features; splicing the updated modal specific features and the modal consistent features as multi-modal features of the media data; using each modal label generated by the label generation part based on the self-supervised learning strategy and the real media data emotional stance label to jointly train the multi-modal and single-modal tasks, and performing constraint optimization on the multi-modal features of the media data; and performing recognition on the optimized multi-modal features of the media data to obtain the stance category to which each media data belongs. The application improves the media emotional stance recognition precision based on a self-supervised multi-task learning architecture, and effectively promotes the development of the network security industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of media stance recognition, and more particularly to a media stance recognition method based on modality-specific representation learning. Background Technology

[0002] In recent years, with the rapid development of internet technology and the widespread adoption of smart devices, the internet has gradually replaced traditional channels such as newspapers and television as a crucial way for people to obtain information in their daily lives. The development of online social media platforms has played a vital role in this transformation. Social media platforms such as Sina Weibo, Twitter, and Facebook have experienced exponential growth in popularity, and social media is gradually becoming the primary carrier of mass information consumption. [1] Effective online public opinion control is a significant need. Identifying sentiment and stance in multimedia data generated on social media is a crucial technology essential for controlling public opinion and achieving effective online public opinion management.

[0003] Media stance identification aims to determine whether the sentiment information contained in user-posted media data is supportive, opposing, or neutral. With the rapid development of short video platforms, social media users are no longer limited to expressing their opinions solely through text; they are increasingly using multiple methods, such as text, images, or videos, to express their views. To effectively identify media stance, it is necessary to represent and fuse the multimodal information, including text, audio, and images, contained in user-posted media data. Against this backdrop, multimodal sentiment identification has received increasing attention. [2-4] Compared to text-based sentiment identification, multimodal sentiment identification models are more stable when processing social media data, can be applied to more scenarios, and have achieved significant improvements in final media sentiment identification results. Multimodal sentiment identification work mainly focuses on two aspects: multimodal information representation learning and multimodal fusion. Regarding multimodal information representation learning methods, Wang et al. [5] A recurrent differential embedding network was constructed to generate multimodal transitions. (Hazarika et al.) [6] Modality-invariant and modality-specific representations for multimodal representation learning are proposed. For multimodal fusion, previous work can be divided into two categories based on the fusion stage: early fusion and late fusion. Early fusion methods typically use sophisticated attention mechanisms for cross-modal fusion. Zadeh et al. [7] A memory fusion network was designed for cross-view interaction. (Tsai et al.) [3] A cross-modal transformer is proposed, which learns cross-modal attention to enhance the target modality. The late-fusion method first learns intra-modal representations and then performs inter-modal fusion. (Zadeh et al.) [2]Tensor fusion networks are used to obtain tensor representations by computing the outer product between unimodal representations. Liu et al. [8] A low-rank multimodal fusion method is proposed to reduce the computational complexity of tensor-based methods. Despite previous work achieving remarkable improvements on benchmark datasets, multimodal sentiment stance recognition remains a challenging task. Baltrusaitis et al. [9] Five core challenges in multimodal learning were identified, including multimodal alignment, translation, representation, fusion, and co-learning. Multimodal representation learning is fundamental to multimodal learning. In recent work, Hazarika et al. [6] This paper points out that multimodal representations should include both consistency and difference information. Based on the different guidance methods used in representation learning, existing methods are divided into two categories: forward guidance and backward guidance. Forward guidance methods focus on designing modules to capture cross-modal information interactions. [3][7-9] However, due to the uniformity of multimodal label annotations, they struggle to capture modality-specific information. Backward-guided methods propose using an additional loss function as a prior constraint to ensure that the learned modality representations contain both consistent and dissimilar information. [6]

[10] Yu et al.

[10] Independent unimodal manual annotations are introduced. By jointly learning unimodal and multimodal tasks, the proposed multi-task multimodal framework simultaneously learns modality-specific and modality-invariant representations. (Hazarika et al.) [6] Two different encoders were designed, projecting each modality into modality-invariant and modality-specific spaces. Two regularization components were used to aid in the learning of the modality-invariant and modality-specific representations. However, the former requires additional manual annotation for single modality, while the latter struggles to represent differences in modality-specific information through spatial discrepancies. Furthermore, they require manual weight balancing among the constraint components in the global loss function, which is highly dependent on human experience. Summary of the Invention

[0004] This invention provides a media stance identification method based on modality-specific representation learning. The invention employs an adversarial learning mechanism to mine the consistency information of sentiment stance among multimodal information and the specific sentiment stance information within each modality of media data, while simultaneously acquiring complementary information between modalities. Based on a self-supervised multi-task learning architecture, it improves the accuracy of media sentiment stance identification, effectively promoting the development of cybersecurity. Details are described below:

[0005] A media stance identification method based on modality-specific representation learning, the method comprising:

[0006] An adversarial learning structure is used to extract modality consistency features between modes and intramodal specific features of intramodal differences from the obtained text, speech, and video single-modal features, respectively.

[0007] By inputting specific features within a modality into label generation based on a self-supervised learning strategy, independent single-modal supervision is obtained;

[0008] A cross-modal attention mechanism is used to obtain intermodal complementary information with text modality as the main modality, and the specific features within the modality are updated.

[0009] The spliced ​​and updated intramodal specific features and modal consistency features serve as multimodal representations of media data;

[0010] By utilizing the modal labels generated by the label generation part based on the self-supervised learning strategy and the sentiment and stance labels of real media data, we jointly train multimodal and unimodal tasks to constrain and optimize the multimodal representation of media data. We then identify the stance category of each media data by identifying the optimized multimodal representation of media data.

[0011] The beneficial effects of the technical solution provided by this invention are:

[0012] 1. This invention addresses the problem of media sentiment identification by proposing a modality-specific representation learning network based on adversarial learning. By employing adversarial learning and introducing a modality discriminator, the obtained modality-specific representations contain modality-specific differential information. Simultaneously, a cross-modal interactive attention module is introduced to obtain complementary information between modalities, thereby improving the sentiment discriminativeness of the obtained media data representations.

[0013] 2. This invention utilizes a self-supervised multi-task learning strategy, which does not require manually annotated unimodal labels in media data. Instead, it uses network-generated unimodal labels and sentiment labels from media data to jointly train multimodal and unimodal tasks, learning intermodal consistency and differences respectively, thereby improving the accuracy of media sentiment recognition.

[0014] 3. Based on the effective identification of the emotional stance of media content, this invention can, on the one hand, combine the emotional stance of mainstream media content to measure the online public opinion orientation of a certain hot event; on the other hand, it can be used to consider the emotional stance of the same media user's historical content to determine the overall public opinion orientation of the media user's published content.

[0015] 4. This invention improves the accuracy of identifying media sentiment and stance, promotes the effective management of large-scale media data, enhances the efficiency of online public opinion supervision, and effectively promotes the development of cybersecurity. Attached Figure Description

[0016] Figure 1 This is a flowchart of a media stance recognition method based on modality-specific representation learning. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.

[0018] To address the aforementioned issues, inspired by independent unimodal labeling and advanced modality-specific representation learning, and leveraging the advantages of both forward and backward guidance, a novel media stance identification method based on modality-specific representation learning is proposed. This differs from the approach proposed by Hazarika et al. [6] The proposed method employs adversarial learning, introducing a modality discriminator to ensure that the resulting modality-specific representations contain modality-specific differential information. Simultaneously, a cross-modal interaction attention module is introduced to obtain complementary information about inter-modal interactions. This differs from the approach taken by Yu et al.

[10] The proposed method employs a self-supervised multi-task learning strategy, eliminating the need for manually annotated unimodal labels and instead utilizing network-generated unimodal labels. The design of this unimodal label generation module is based on two considerations: first, label discrepancies are positively correlated with the distance difference between modality representations and class centers; second, unimodal and multimodal labels are highly correlated. Considering the initial instability of automatically generated unimodal labels, a dynamic update method is introduced, applying greater weight to later-generated unimodal labels. Furthermore, a self-adjusting strategy is introduced when integrating the final multi-task loss function to adjust the weights of each subtask. Subtasks with small label differences between automatically generated unimodal labels and manually annotated multimodal labels struggle to learn modality-specific representations. Therefore, the weights of subtasks are positively correlated with label discrepancies. The proposed method is validated on the publicly available Chinese multimodal dataset CH-SIMS, and its effectiveness in media stance identification is verified through multiple experiments.

[0019] The work of this invention aims at representation learning based on a late-stage fusion structure. Unlike previous studies, this invention jointly learns unimodal and multimodal tasks with a self-supervised strategy, learning intermodal semantic consistency information from multimodal tasks and intramodal dissimilarity information from unimodal tasks. Based on the learned features containing intramodal dissimilarity information, cross-modal attention is used to capture complementary intermodal interaction information.

[0020] Example 1

[0021] A media stance identification method based on modality-specific representation learning, see [link to relevant documentation]. Figure 1 The method includes the following steps:

[0022] 101: Obtain multimodal information from each media data and extract the corresponding features of the three modalities: text, speech, and image;

[0023] 102: An adversarial learning structure is used to extract semantic consistency information between modalities and intramodal difference information from the obtained text, speech, and video single-modal features, respectively;

[0024] 103: Based on the obtained modality-specific features containing intramodal differential information, input them into the label generation based on the self-supervised learning strategy to obtain independent single-modality supervision;

[0025] 104: Employ cross-modal attention mechanisms for specific features within a modality to obtain complementary information between modalities, with the text modality as the primary modality, and update specific features within the modality;

[0026] 105: The spliced ​​and updated intramodal specific features and the obtained modal consistency features are used as the final representation of the discourse video segment;

[0027] 106: By utilizing the modal labels generated by the label generation part based on the self-supervised learning strategy and the sentiment labels of real media data, multimodal and unimodal tasks are jointly trained to identify the stance category of each media data in the final multimodal representation of media data, thereby improving the efficiency of online public opinion supervision and effectively promoting the development of cybersecurity.

[0028] In summary, the embodiments of the present invention, through steps 101-106, can measure the online public opinion orientation of a certain hot topic by combining the emotional stance of the content published by mainstream media; it can be used to consider the emotional stance of the historical content published by the same media user to determine the overall public opinion orientation of the content published by that media user; it can improve the accuracy of media emotional stance identification, promote the effective management of large-scale media data, improve the efficiency of online public opinion supervision, and effectively promote the development of cybersecurity.

[0029] Example 2

[0030] The following section provides a further explanation of the scheme in Example 1, using specific calculation formulas and examples. See the description below for details:

[0031] Step 201: Obtain the multimodal information of each media data in the database and extract the corresponding features of the three modalities: text, speech, and video;

[0032] Specifically, text feature extraction: All videos in the dataset have manually generated subtitles, including both Chinese and English versions. Only Chinese subtitles are used. Two unique markers, [CLS] and [SEP], are added to indicate the start and end of each subtitle line. Then, word vectors are extracted from the subtitle text using a pre-trained Chinese BERT base model. Ultimately, each word is represented as a 768-dimensional word vector, and the text features are represented as F... TAudio feature extraction: The commonly used speech feature extraction model OpenSMILE is employed.

[13] Extracting audio features. OpenSMILE

[11] OpenSMILE is an open-source software developed by Eyben et al., used to extract audio features that make up speech. It generates high-dimensional vectors summarizing multiple audio statistical descriptors, including loudness, Mel spectrum, Mel frequency cepstral coefficients, pitch, etc. For example, previous work has shown...

[12] Audio features play a crucial role in multimodal sentiment analysis. Specifically, the eGeMAPS profile is used to extract 88-dimensional features for a utterance. A unidirectional long short-term memory (sLSTM) network is then employed to capture the temporal information of the feature sequence, outputting the audio features represented as F. A Video feature extraction: Utilizing the commonly used image feature extraction model OpenFace.

[13] Visual features were extracted, with a feature dimension of 465. A unidirectional long short-term memory network (sLSTM) was then used to capture the temporal information of the feature sequence, and the output image feature representation is F. V .

[0033] 202: An adversarial learning structure is used to extract semantic consistency information between modalities and intramodal difference information from the obtained text, speech, and video single-modal features, respectively;

[0034] In this embodiment, to effectively and comprehensively utilize multimodal information, this invention proposes extracting semantically consistent information from different modalities and enriching the specific diversity information of individual modalities. Cross-modal interaction and complementary information also need to be considered during the multimodal learning stage. Inspired by adversarial learning networks, this invention designs an adversarial multimodal information decomposition module to achieve the above objectives. Like other adversarial learning networks, this module includes a generator and a discriminator, which are trained with opposing objectives to compete against each other and find potential feature spaces. In this module, features from multiple modalities are decomposed and projected into two disjoint feature spaces, extracting intermodal consistent semantic information and intramodal specific information, respectively.

[0035] Specifically, in this embodiment of the invention, the single-modal features are first mapped to feature vectors of dimension 100 through a linear layer. Then, a generator G projects the single-modal features into a common latent subspace to extract multimodal consistency semantic information, resulting in semantic consistency features between modalities.

[0036] C M =G(F M ;θ M ),

[0037] Here, M∈{T, A, V}, where T, A, and V represent the modalities of text, audio, or image, respectively, F is the unimodal feature, and θ is the generator parameter.

[0038] In addition to commonalities, each mode contains specific positional information that can complement other modes. Single-modal features are projected onto fully connected (FC) layers with corresponding parameters. The specific representation within a mode can be expressed as:

[0039] S M =FC(F M ;θ M ),

[0040] To ensure that different aspects of multimodal data are encoded, orthogonal loss is used to achieve C. M and S M The non-redundancy between modes and the guarantee of the difference in information contained in specific features within each mode are as follows:

[0041] D(I;θ D ) = softmax(I T W+b);

[0042] Where I represents the input feature matrix, θ D The parameters of the discriminator are represented by W. W is the weight matrix, b represents the bias, and T represents the transpose of the feature matrix. The true modality label of I can be represented as... The label is for the Nth sample belonging to the M mode.

[0043] The generator projects unimodal features into a shared latent space that tends to have the same distribution, with the aim of learning semantic consistency information across different modalities. Therefore, the discriminator D is encouraged not to distinguish between generated consistent representations C. M The source mode, corresponding to the loss function L C For modality-specific features, to ensure they contain corresponding modality information and that there are differences between different modes, the discriminator D is encouraged to distinguish modality-specific representations S. M The corresponding loss function is L D The loss function constrained by this process is defined as:

[0044]

[0045] Where, N b Represents batch size. L is trained using a gradient inversion layer. CDuring forward propagation, the input remains unchanged, while the gradient is multiplied by -1 during backward propagation. Then, to ensure that modality-consistent representations exhibit the same semantics and are relevant to sentiment stance, a semantic loss for sentiment classification based on modality-consistent representations was designed and implemented.

[0046]

[0047] in, The category prediction value is a representation of modal consistency. This is a real label.

[0048] Based on the above steps, adversarial learning was used to decompose multimodal information, resulting in modality consistency features containing common emotional stance information and modality-specific features containing differential emotional stance information.

[0049] 203: Based on the obtained intramodal specific features containing intramodal difference information, input them into the label generation module based on the self-supervised learning strategy to obtain independent single-modal supervision;

[0050] In this embodiment, a multi-task learning module is proposed to improve the accuracy of multimodal video position information recognition by utilizing modality-specific representations. This module is designed to generate single-modal supervision values ​​based on multimodal labels and modality data representations. To avoid unnecessary interference with the network parameter update process, the multi-task learning module is designed as a non-parametric module. Typically, single-modal supervision values ​​are highly correlated with multimodal labels, so the offset is calculated based on the relative distance from the modality data representation to the category center.

[0051] Since different modal representations exist in different feature spaces, directly using absolute distance values ​​for measurement is not accurate enough. Therefore, a relative distance value that is independent of spatial differences is proposed.

[0052] First, during model training, the positive sample centers of the data representations for each modality are maintained. and negative sample center

[0053]

[0054]

[0055] Where i∈{m,t,a,v}, and N represents the number of training samples. It is the global feature representation of sample j in mode i.

[0056] For the feature representation of each modality, the L2 normalization method is used for measurement. Distance from the category center:

[0057]

[0058]

[0059] Then, a relative distance value is defined, which evaluates the relative distance of the modal representation to the positive and negative sample centers:

[0060]

[0061] Where i∈{m,t,a,v}, m,t,a,v represent multimodal, text, audio, and image modalities, respectively, and ∈ is introduced to avoid the case of zero values, which is a local minimum.

[0062] It can be observed that α i Positively correlated with the final result.

[0063] To obtain the relationship between supervised information and predicted values, consider the following two relationships:

[0064]

[0065]

[0066] Where s∈{t,a,v}, ∝ represents a direct proportional relationship.

[0067] Clearly, by comprehensively considering the above relationships, the label for the single modality can be obtained through equal-weighted summation:

[0068]

[0069] Where s∈{t,a,v}. This represents the offset value of single-modal supervision over multimodal supervision.

[0070] Due to the dynamic changes in modal representation, the modal labels generated by the above method are not stable enough. To mitigate the adverse effects, a momentum-based update strategy is designed, which combines newly generated values ​​with historical values:

[0071]

[0072] Where s∈{t,a,v}. These are the single-modal labels obtained for each epoch.

[0073] Step 204: Apply a cross-modal attention mechanism to specific features within a modality to obtain complementary information between modalities with the text modality as the primary modality, and update the specific features within the modality;

[0074] Simply concatenating these representations without considering the complementary information between modalities will negatively impact the final identification of video stance information. Therefore, a modal information interaction module was designed to capture the interactions and complementary information between modalities. This module is based on the original Transformer...

[14] Encoder layer design.

[0075] Since textual information is the modality with the richest semantic information, it can be used to enhance the positional information of audio and video modalities. The greater the differences between modalities, the richer the complementary information that can be obtained. Therefore, modality-specific representations S are employed. T S A , and S V Learning complementary information from intermodal interactions. Utilizing a cross-modal attention layer, S a and S v respectively with S t The process of exchanging information can be represented as follows:

[0076]

[0077]

[0078] Step 205: Concatenate the updated intramodal specific features and the obtained modality consistency features as the final representation of the discourse video segment;

[0079] To maximize the acquisition of multimodal position information, C and S T , and Connecting these elements, we obtain discourse features that discriminate emotional stances, where C = C t +C a +C v This feature is used for the emotional stance classification of the final spoken video clips.

[0080] Step 206: Using the modal labels generated by the label generation module based on the self-supervised learning strategy and the sentiment stance labels of real videos, jointly train multimodal and unimodal tasks to classify the final discourse video segment representations to obtain the sentiment stance category of each video.

[0081] For unimodal tasks, the L1 loss function is used as the basic optimization objective, and the difference between u-labels and m-labels is used as the weight of the loss function:

[0082]

[0083] N is the number of training samples. It is the weight of the i-th sample in the auxiliary task s.

[0084] Finally, considering all the loss terms mentioned above, we can obtain the overall loss function of the network:

[0085] L all =γ1L O +γ2L A +γ3L sem +γ4L M ,

[0086] Wherein, γ1, γ2, γ3, and γ4 are the weight values ​​that control the proportion of each loss term.

[0087] Step 207: Monitor online public opinion based on the stance category of each media data point to improve cybersecurity.

[0088] In summary, the embodiments of the present invention, through steps 201-207 above, determine the overall public opinion orientation of the content published by the media user; improve the accuracy of media sentiment identification, promote the effective management of large-scale media data, improve the efficiency of online public opinion supervision, and effectively promote the development of cybersecurity.

[0089] Example 3

[0090] The feasibility of the methods in Examples 1 and 2 is verified below with reference to Tables 1-5. In the experiments, CH-SIMS was applied to the embodiments of the present invention.

[10] Dataset, see the description below:

[0091] The CH-SIMS dataset is a Chinese unimodal and multimodal sentiment analysis dataset. It contains 2,200 video clips. Each video clip has both multimodal and independent unimodal annotations, allowing for the exploration of interactions between multimodal and independent unimodal expressions, or the use of independent unimodal annotations for unimodal sentiment analysis. The sample distribution of the dataset is shown in Table 1.

[0092] Table 1. Sample distribution of the CH-SIMS dataset

[0093] training set Validation set test set 1368 456 376

[0094] Evaluation metrics: We employ multi-class classification metrics within the sentiment analysis domain: weighted F1 score and multi-class accuracy Acc-k, where k = {2, 3, 5}. Higher values ​​indicate better performance.

[0095] Table 2 Method Parameter Settings

[0096]

[0097]

[0098] This method is trained end-to-end. During training, SGD (Stochastic Gradient Descent) is used for optimization. Specific parameter settings are shown in Table 2. Dropout is adjusted to 0.5. The method is implemented using Python 3, PyTorch 1.7.1, and CUDA 10.1. All experiments were conducted on a server equipped with an NVIDIA 3090TI GPU.

[0099] Table 3 shows the comparison results of our proposed method with existing multimodal sentiment position recognition methods on the CH-SIMS dataset. For a fair comparison, our method reproduced all the comparison methods using the same input features. In the experiments, firstly, compared with the traditional multimodal fusion methods TFN and LMF, our method achieved significant progress on all evaluation metrics. Even compared with existing Transformer-based models BERT-MAG and MULT, our method achieved competitive results on all metrics. Furthermore, we reproduced the two best baselines, “MISA” and “MAG-BERT”, under the same conditions. We found that our model outperformed them in most evaluations. It is worth noting that the results achieved by existing Transformer-based models do not demonstrate their advantage over traditional multimodal fusion methods. The superior results achieved by our proposed method can be attributed to the following three points:

[0100] 1. For multimodal information, the proposed method uses an adversarial learning mechanism to effectively model the consistency information of sentiment stance between modalities and the feature information within modalities, and makes full and effective use of the complementarity between multimodal information;

[0101] 2. In order to ensure that the obtained intramodal feature information is non-redundant and maximizes the difference, and considering the inter-class differences of multimodal data, multimodal marginal loss is used to constrain the interaction before and after the interaction of modality-specific representations.

[0102] 3. After obtaining modality-specific representations that maximize differentiated information, the model uses a parameter-free multi-task learning module to acquire labels for each modality and performs joint training to achieve multi-task learning, effectively improving the model's sentiment stance recognition performance.

[0103] To further verify the effectiveness of the proposed method, based on the CH-SIMS dataset which includes unimodal labels, the traditional multimodal fusion method was extended to a multi-task learning framework, and its performance was compared with that of the proposed method. The experimental results are shown in Table 4. Table 4 shows that the proposed method outperforms the comparison methods on all metrics, achieving the best results. Compared with the comparison methods, the proposed method models multimodal information from two aspects: common position information between modalities and specific position information within modalities. This more comprehensive approach uncovers the correlations between modalities, thus achieving superior model performance. The confusion matrix on the test set was also visualized to fully analyze the model performance. The visualization results show that the proposed method significantly outperforms the positive category in the negative category. The main reason for this is that the number of negative category samples in the CH-SIMS dataset is much higher than that of positive category samples. Furthermore, as the categories become more refined, the model performance gradually declines, demonstrating that the model struggles to capture fine-grained sentiment position information for distinguishing sentiment position categories.

[0104] Table 3 shows the experimental results compared with existing multimodal sentiment identification algorithms.

[0105] Number of lines method Acc-2 F1-2 Acc-3 F1-3 Acc-5 F1-5 1 <![CDATA[LMF [8] ]]> 79.26 78.86 64.42 64.49 37.77 38.03 2 <![CDATA[TFN [2] ]]> 88.94 89.00 74.63 75.21 47.87 49.36 3 <![CDATA[MULT [3] ]]> 87.39 86.13 74.52 72.57 45.74 45.36 4 <![CDATA[BERT-MAG

[15] ]]> 79.15 72.46 64.78 55.17 32.02 23.63 5 <![CDATA[MMIM

[16] ]]> 80.53 75.91 64.41 58.06 30.91 23.10 6 <![CDATA[MFN [7] ]]> 87.56 87.24 70.59 70.29 41.86 42.32 7 OURS 90.68 90.75 76.69 76.61 51.54 52.22

[0106] Table 4 shows the experimental results comparing the algorithm with the multimodal sentiment identification algorithm based on the multi-task learning framework.

[0107] Number of lines method Acc-2 F1-2 Acc-3 F1-3 Acc-5 F1-5 1 MLMF 89.42 89.65 76.23 76.51 44.79 45.04 2 MTFN 74.47 74.30 57.45 59.79 35.00 35.46 3 MLF_DNN 88.25 88.36 75.43 75.67 51.17 51.28 4 OURS 90.68 90.75 76.69 76.61 51.54 52.22

[0108] Ablation experiments were conducted on the main components of the proposed method, and their impact on algorithm performance was evaluated. Table 5 presents the experimental results of the ablation. From Table 5, the following observations can be obtained.

[0109] Table 5 Ablation Experiment Results

[0110]

[0111] The comparison between lines 7 and 6 shows that removing the adversarial multimodal information decomposition process leads to a decrease in algorithm performance. Experimental results indicate that fully utilizing multimodal information from the perspectives of intermodal semantic consistency and intramodal differences can effectively improve the accuracy of multimodal sentiment stance classification. The comparison between lines 7 and 5 shows that removing the modal interaction module leads to a decrease in algorithm performance. Experimental results demonstrate the effectiveness of mining complementary information from intermodal interactions for sentiment stance recognition. The comparison between lines 7 and 4 shows that removing the multi-task learning stage leads to a decrease in algorithm performance. Experimental results verify that using generated single-modal labels for multimodal specific representation optimization learning can improve the algorithm accuracy of multimodal sentiment stance recognition tasks. The overall ablation experiment results verify the rationality of the model design and the effectiveness of each module in this invention embodiment. Through the above experimental verification, the media stance recognition method based on modal specific representation learning proposed in this invention embodiment can effectively process common media data, including text, audio, and image data, extract features of each modality, and mine sentiment stance information from media data. The proposed method innovatively decomposes the various modalities of media data using an adversarial learning structure to facilitate subsequent multimodal information fusion. It fully utilizes the sentiment and stance information within each modality, thereby improving the accuracy of media sentiment and stance identification. On the one hand, this method can effectively identify the sentiment and stance of content before it is published by media users, thus filtering the content and achieving effective control of online public opinion, preventing the spread of negative influences, and effectively reducing the cost and error of manual judgment. On the other hand, by analyzing the sentiment and stance identification results of past content published by media users, it can effectively filter media types, assisting relevant departments in determining whether a media is suitable for dissemination or high-risk, thereby contributing to the overall control of online public opinion dissemination and the cybersecurity environment.

[0112] References

[0113] [1] Hermida, Alfred. "Twittering the news: The emergence of ambientjournalism." Journalism practice 4, no.3 (2010): 297-308.

[0114] [2]Amir Zadeh,Minghai Chen,Soujanya Poria,Erik Cambria,Louis-PhilippeMorency:Tensor Fusion Network for Multimodal Sentiment Analysis.EMNLP 2017:1103-1114.https: / / doi.org / 10.24963 / ijcai.2017 / 427

[0115] [3]Yao-Hung Hubert Tsai,Shaojie Bai,Paul Pu Liang,J.Zico Kolter,Louis-Philippe Morency,Ruslan Salakhutdinov:Multimodal Transformer forUnaligned Multimodal Language Sequences.ACL(1)2019:6558-6569.

[0116] [4]Soujanya Poria,Devamanyu Hazarika,Navonil Majumder,Rada Mihalcea:Beneath the Tip of the Iceberg:Current Challenges and New Directions inSentiment Analysis Research.CoRR abs / 2005.00357(2020).

[0117] [5]Yansen Wang,Ying Shen,Zhun Liu,Paul Pu Liang,Amir Zadeh,Louis-Philippe Morency:Words Can Shift:Dynamically Adjusting Word RepresentationsUsing Nonverbal Behaviors.AAAI 2019:7216-7223.

[0118] [6]Devamanyu Hazarika,Roger Zimmermann,Soujanya Poria:MISA:Modality-Invariant and-Specific Representations for Multimodal Sentiment Analysis.ACMMultimedia 2020:1122-1131,https: / / dl.acm.org / doi / abs / 10.1145 / 2433396.2433443.

[0119] [7]Amir Zadeh,Paul Pu Liang,Navonil Mazumder,Soujanya Poria,ErikCambria,Louis-Philippe Morency:Memory Fusion Network for Multi-viewSequential Learning.AAAI 2018:5634-5641.

[0120] [8]Liu,Zhun,Ying Shen,Varun Bharadhwaj Lakshminarasimhan,Paul PuLiang,AmirAli Bagher Zadeh,and Louis-Philippe Morency."Efficient Low-rankMultimodal Fusion With Modality-Specific Factors."In Proceedings of the 56thAnnual Meeting of the Association for Computational Linguistics(Volume 1:LongPapers),pp.2247-2256.2018.

[0121] [9]Tadas Baltrusaitis,Chaitanya Ahuja,Louis-Philippe Morency:Multimodal Machine Learning:A Survey and Taxonomy.IEEE Trans.PatternAnal.Mach.Intell.41(2):423-443(2019).

[0122]

[10] Wenmeng Yu,Hua Xu,Fanyang Meng,Yilin Zhu,Yixiao Ma,Jiele Wu,JiyunZou,Kaicheng Yang:CH-SIMS:A Chinese Multimodal Sentiment Analysis Datasetwith Fine-grained Annotation of Modality.ACL 2020:3718-3727,https: / / dl.acm.org / doi / abs / 10.1145 / 2964284.2964335.

[0123]

[11] Eyben,Florian,Martin and Schuller."Opensmile:themunich versatile and fast open-source audio feature extractor."In Proceedingsof the 18th ACM international conference on Multimedia,pp.1459-1462.2010.

[0124]

[12] Mingli Song,Jiajun Bu,Chun Chen,and Nan Li,“Audio-visual basedemotion recognition-a new approach,”in Proceedings of the 2004IEEE ComputerSociety Conference on Computer Vision and Pattern Recognition,2004.CVPR2004.,vol.2,June 2004,pp.II–II.

[0125]

[13] Tadas, Peter Robinson, and Louis-Philippe Morency. "Openface: an open source facial behavior analysis toolkit." In 2016 IEEE WinterConference on Applications of Computer Vision (WACV), pp. 1-10. IEEE, 2016.

[0126]

[14] Ashish Vaswani,Noam Shazeer,Niki Parmar,Jakob Uszkoreit,LlionJones,Aidan N.Gomez,Lukasz Kaiser,Illia Polosukhin:Attention is All youNeed.NIPS 2017:5998-6008.

[0127]

[15] Wasifur Rahman,Md.Kamrul Hasan,Sangwu Lee,AmirAli Bagher Zadeh,Chengfeng Mao,Louis-Philippe Morency,Mohammed E.Hoque: Integrating MultimodalInformation in Large Pretrained Transformers.ACL 2020:2359-2369.

[0128]

[16] Wei Han, Hui Chen, Soujanya Poria: Improving Multimodal Fusion withHierarchical Mutual Information Maximization for Multimodal SentimentAnalysis.EMNLP(1)2021:9180-9192.

[0129] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0130] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A media stance identification method based on modal-specific representation learning, characterized in that, The method comprises: Adopting an adversarial learning structure to extract inter-modal consistency features and intra-modal specific features of inter-modal difference information of the obtained text, speech and video single modal features respectively; Inputting the intra-modal specific features into a label generation based on a self-supervised learning strategy to obtain independent single modal supervision; Adopting a cross-modal attention mechanism to the intra-modal specific features to obtain inter-modal complementary information with the text modal as the main modal, and updating the intra-modal specific features; Splicing the updated intra-modal specific features and the inter-modal consistency features as the multi-modal representation of the media data; Using the labels of each modal generated by the label generation part based on the self-supervised learning strategy and the real media data emotion stance labels to jointly train the multi-modal and single modal tasks, and performing constraint optimization on the multi-modal representation of the media data, and performing recognition on the optimized multi-modal representation of the media data to obtain the stance category to which each media data belongs; The cross-modal attention mechanism to the intra-modal specific features to obtain inter-modal complementary information with the text modal as the main modal, and updating the intra-modal specific features are as follows: Adopting modal specific representation , , and Inter-modal complementary information learning is performed by using a cross-modal attention layer, and respectively interact with information, which is represented as: ; ; wherein, is a text modality specific representation; is an image modality specific representation after interaction with the text information; is a text modality specific representation; is an audio modality specific representation; is an image modality specific representation.

2. The media stance identification method based on modal-specific representation learning according to claim 1, characterized in that, The inter-modal consistency features are as follows: ; wherein, , respectively represent the modalities of the text, audio or image, F is a single-modal feature, is a parameter of the generator, G is a generator.

3. The media stance identification method based on modal-specific representation learning according to claim 2, characterized in that, The intra-modal specific features are as follows: ; Wherein, FC is a full connection layer.

4. The media stance identification method based on modal-specific representation learning according to claim 3, characterized in that, The loss function of the inter-modal consistency features and the intra-modal specific features is defined as: wherein, representing a batch-size. Training using a gradient reversal layer , A semantic loss of the modal consistency representation for emotion classification is designed and implemented: wherein, is the class prediction value for modality consistency representation, is the true label.

5. The media stance identification method based on modal-specific representation learning according to claim 1, characterized in that, The single modal supervision is as follows: ; wherein, , represent text, audio, and image modalities, respectively; is the learned single-modal label; is the true multi-modal label; is the relative distance value of the modality-specific representation to the positive and negative centers; is the relative distance value of the multi-modal representation to its positive and negative centers; denotes the shift value of the single-modal supervision with respect to the multi-modal supervision; A momentum-based update strategy is designed to combine the newly generated value with the historical value: ; wherein, is the single modality label obtained for the epoch; is the single modality label obtained after epochs.

6. The media stance identification method based on modal-specific representation learning according to claim 1, characterized in that, For the single modal task, an L1 loss function is used as the basic optimization target, and the difference between the u-labels and the m-labels is used as the weight of the loss function: ; wherein, is the number of training samples, is the weight of the th sample of the auxiliary task, is the predicted multi-modal label of the th sample, is the true multi-modal label of the th sample, is the predicted single-modal label of the th epoch, is the single-modal label obtained after epochs. ​

Citation Information

Patent Citations

  • Multi-modal fusion emotion recognition system and method based on multi-task learning and attention mechanism and experimental evaluation method

    CN113420807A

  • News event search method and system based on multi-level image-text semantic alignment model

    WO2023093574A1