Aspect-level multi-modal Mongolian sentiment analysis method based on cross-modal attention mechanism

Through the cross-modal attention mechanism and hierarchical disentanglement technology, the problems of modal feature alignment and complex language phenomena in Mongolian sentiment analysis are solved, efficient sentiment recognition and accurate sentiment classification are achieved, and the accuracy and generalization ability of Mongolian sentiment analysis are improved.

CN120706419APending Publication Date: 2025-09-26INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510802130.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing technology in Mongolian sentiment analysis has a small amount of sentiment corpus, which leads to poor performance of deep learning models, cross-modal feature alignment problems in multimodal sentiment analysis, and sentiment analysis challenges brought about by complex language phenomena. In particular, in aspect-level sentiment analysis, it is difficult for the model to accurately judge the emotional tendency of the text.

Method used

An aspect-level multimodal Mongolian sentiment analysis method based on a cross-modal attention mechanism is adopted. By constructing target-opinion-sentiment triple data, a multi-head attention mechanism is used for modal translation and hierarchical disentanglement. Combined with an adaptive sentiment polarity classification network, the sentiment classification strategy is dynamically adjusted to handle complex language phenomena.

Benefits of technology

The accuracy and generalization ability of Mongolian sentiment analysis have been improved. It can maintain efficient sentiment recognition performance in the absence of modality, and can effectively correct spelling or grammatical errors to provide more accurate sentiment classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706419A_ABST
    Figure CN120706419A_ABST
Patent Text Reader

Abstract

An aspect-level multi-modal Mongolian sentiment analysis method based on a cross-modal attention mechanism comprises the steps that firstly, multi-modal data are preprocessed, and visual and audio modals are converted into text modals through modal translation; and meanwhile, the modal features are separated into public, private and noise representations by utilizing layered de-entanglement. Secondly, a staged aspect-level emotion extraction module is constructed, in the first stage, boundary labeling of target words and viewpoint words is achieved based on a BiLSTM network and a BIO label system, and a target guidance module is introduced to fuse syntactic dependency features to optimize viewpoint word prediction; in the second stage, the effectiveness of the target word-viewpoint word pair is judged through position embedding and a logistic regression model, and emotion two-tuple extraction is completed. And finally, designing an adaptive emotional polarity classification network, generating each modal dynamic weight by using a feedforward neural network to realize feature adaptive fusion, balancing classification accuracy and weight distribution through a multi-loss joint optimization mechanism, suppressing noise interference and outputting a triple result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to an aspect-level multimodal Mongolian sentiment analysis method based on a cross-modal attention mechanism. Background Art

[0002] Sentiment analysis has become an active research topic in the fields of data mining and natural language processing (NLP). It analyzes people's opinions, feelings, and emotions towards entities such as products, services, organizations, and events. This information and ideas are published in various data forms such as text and video, and more or less carry the user's personal emotional tendencies and contain a large amount of emotional information.

[0003] Existing sentiment analysis methods have achieved good results in some scenarios, but there are still problems with sentiment analysis of minority languages. Taking Mongolian as an example, these problems are mainly manifested in the following aspects:

[0004] First, the limited amount of Mongolian sentiment data leads to poor deep learning model performance. As a low-resource language, Mongolian has a scarce sentiment data base and complex grammar, making it difficult for traditional deep learning models to effectively train. In multimodal sentiment analysis, the diversity and quality of annotated data are crucial. However, the limited amount of Mongolian sentiment text data and the insufficient annotation granularity hinder model training effectiveness.

[0005] Second, there's the issue of cross-modal feature alignment in multimodal sentiment analysis. In sentiment analysis tasks, data from a single modality often cannot fully reflect sentiment. This is especially true in multimodal (text, audio, video) sentiment analysis, where information from different modalities is often misaligned. In particular, aspect-level sentiment analysis requires a richer and more diverse set of sentimental data to meticulously identify the correlations between sentiment elements.

[0006] Third, complex language phenomena pose challenges to sentiment analysis: In sentiment analysis, complex language phenomena such as irony, metaphor, and polysemy often make it difficult for models to accurately determine the emotional orientation of a text. This is particularly true in languages ​​like Mongolian. Summary of the Invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide an aspect-level multimodal Mongolian sentiment analysis method based on a cross-modal attention mechanism. The aspect-level sentiment analysis strategy is used to explore the depth and breadth of the sentiment information in the data. The data characteristics and processing methods of different modalities have an important influence on the analysis effect. By dynamically adjusting the sentiment classification strategy, multimodal information can be used to automatically adjust when facing complex language phenomena, providing more accurate sentiment classification results.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is:

[0009] A method for aspect-level multimodal Mongolian sentiment analysis based on a cross-modal attention mechanism includes the following steps:

[0010] Step 1: obtain Mongolian text, image and Mongolian speech, and construct target-viewpoint-emotion triple data using the Mongolian text and Mongolian sentiment dictionary;

[0011] Step 2: Use a modality translation method based on a multi-head attention mechanism to convert the visual modality and audio modality into textual modality, and decompose the features of each modality into public representation, private representation, and noise representation through hierarchical disentanglement;

[0012] Step 3: Use the aspect-level sentiment extraction module to learn the target-opinion bigrams in the corpus based on the triple data of the text modality as the supervision signal;

[0013] Step 4, using an adaptive emotion polarity classification network to obtain emotion polarity based on the public representation, the noise representation, and the private representations of the visual modality and the audio modality;

[0014] Step 5: Match the target-viewpoint bigram with the sentiment polarity to obtain a valid target-viewpoint-sentiment triplet, thereby realizing Mongolian sentiment analysis.

[0015] In one embodiment, the step 1 generates triple data using a Mongolian sentiment dictionary and heuristic rules. The Mongolian sentiment dictionary is obtained by translating and expanding a Chinese sentiment dictionary, and the expansion is achieved based on a point mutual information method.

[0016] In one embodiment, the step 2 is to construct a unified modality coding model to perform the modality translation and feature decomposition; the unified modality coding model includes:

[0017] The modal translation module uses the Transformer's multi-head attention mechanism to achieve cross-modal feature alignment. The formula is:

[0018] D vt =MultiHead(E t ,E v ,E v )

[0019] D at =MultiHead(E t ,E a ,E a )

[0020] Among them, E t is the text feature, E a is the audio feature, E vis the visual feature, D vt For vision-to-text translation features, D at It is the audio to text translation feature, MultiHead() represents the multi-head attention mechanism;

[0021] Hierarchical disentanglement module, using E t , D at , D vt Construct public representation C and private representation P n and noise characterization N n , where C = PublicEncoder(E t ,D at ,D vt θ c )=MultiHead(E t ,D at ,D vt ), n∈{t,a,v};

[0022] PublicEncoder() represents public sentiment learning, θ c Refers to the parameter weights learned by the PublicEncode network after learning representations containing data of different modalities, and the private representation p n The calculation method is as follows:

[0023]

[0024] After initializing the module parameters, the noise has the following loss function as the objective:

[0025]

[0026] in, It is a non-public representation, composed of the original features F of different modalities n Subtract the common representation C to get .

[0027] In one embodiment, the text features are obtained by a BiLSTM-based text sentiment analysis model;

[0028] The audio features are obtained by first obtaining a logarithmic Mel-spectrogram and prosody features from the audio data, and then inputting the logarithmic Mel-spectrogram into a BiGRU-based audio sentiment analysis model to capture the emotional spatiotemporal features in the audio;

[0029] The image features are obtained through the ViT image encoder.

[0030] In one embodiment, the text feature is represented as:

[0031] E t =BiLSTM(Xt )

[0032] where X t For text data;

[0033] The audio feature is expressed as:

[0034]

[0035] Among them, S mel (t) represents the mth Mel filter energy value of the tth frame of the logarithmic Mel spectrum graph, and MFCC(t) is the tth Mel frequency cepstral coefficient. The formulas are as follows:

[0036]

[0037] Where t is the frame index of the audio data, t = 1, 2, ..., T, T is the maximum index length of the data, m is the subscript order of the Mel filter, M is the number of Mel filters, which determines the dimension of the Mel spectrum graph, |X(f)| 2 is the power spectrum obtained after short-time Fourier transform processing of audio, H m (f) is the transfer function of the mth Mel filter, f min ,f max is the frequency range,

[0038]

[0039] Among them, k is the output index of MFCC coefficient, K is the retained MFCC feature dimension, K <M;

[0040] The image features are expressed as:

[0041] E v =ViT(Preprocess(U v ))

[0042] Among them, U v The initial image features are extracted using the facial behavior analysis toolkit, Preprocess represents the preprocessing operation on the extracted features, and ViT is the feature mapping function of the long short-term memory network.

[0043] In one embodiment, the aspect-level sentiment extraction module is a target-opinion binary extraction model that sequentially performs the following two stages:

[0044] In the first stage, target words and opinion words are extracted from Mongolian text input, and prediction is achieved through word-level boundary annotation;

[0045] In the second stage, based on the sentence context, the matching relationship between the target and the viewpoint is judged to complete the pairing of the tuple.

[0046] In one embodiment, in the first stage, the target word and opinion word extraction module based on the BiLSTM network and the GCN network and the BIO labeling system are used to implement boundary annotation of the target word and opinion word. The formula is:

[0047]

[0048] in, is the context encoding of BiLSTM output, W T and b T are model parameters, is the target word boundary prediction result at time step t.

[0049] In the second stage, the relative distance d between the target word and the opinion word is encoded by position embedding ij =|pos(i)-pos(j)|, combined with the logistic regression model to judge the effectiveness of the target word-opinion word pair, the classification function is:

[0050]

[0051] Among them, W and b are classifier parameters, p ij is the pairing probability, is the target word embedding, Embedding for opinion words.

[0052] In one embodiment, in the first stage, a target guidance module is introduced to represent the target boundary output by BiLSTM h T The syntactic dependency feature h generated by GCN GCN Splicing, the formula is:

[0053]

[0054] in, This is the opinion word boundary prediction result.

[0055] In one embodiment, the adaptive sentiment polarity classification network includes:

[0056] Dynamic weight generation module, learning weighted feature attention matrix through feedforward neural network And generate the dynamic weight w of each modality through Softmax normalization i , the formula is:

[0057]

[0058] in, Z is the multimodal feature concatenation vector of the public representation, noise representation, and private representations of the visual modality and audio modality, and is the weight matrix of the fully connected layer, h is the hidden layer dimension, and is the corresponding bias vector; tanh(·) is the hyperbolic tangent activation function, σ(·) is the Sigmoid activation function;

[0059] Multi-loss joint optimization mechanism, the total loss function is:

[0060] L=L classification +λ1L entropy +λ2L noisy

[0061] Among them, L classification is the cross entropy loss, L entropy is the entropy-based weight regularization loss, L noisy is the noise representation suppression loss, and λ1 and λ2 are hyperparameters.

[0062] In one embodiment, the dynamic weight w is used i , perform weighted combination of each feature to form the final joint representation Z joint :

[0063]

[0064] Among them, F i is the i-th feature, w i is the corresponding dynamic weight;

[0065] The joint representation Z joint Input to the fully connected layer to produce the final sentiment prediction result

[0066]

[0067] The optimization goal is formulated as a multi-objective problem:

[0068] min w,θ L(w,θ)

[0069] Where: w is the parameter of the dynamic weight generation module, θ is the parameter of the feature extraction and classifier.

[0070] Compared with the prior art, the present invention has the following beneficial effects:

[0071] (1) A modality translation method is used to convert visual and audio modalities into textual modalities, improving the representation quality between modalities and thus maintaining efficient sentiment recognition performance even when modality is missing. In addition, a hierarchical disentanglement technique is used to separate public, private, and noise representations, effectively improving the accuracy and generalization ability of multimodal sentiment analysis.

[0072] (2) A sentiment word extraction model is proposed to capture sentence-level sentiment details. The model processes target words and opinion words in two stages, respectively. A phased strategy is used to select sentiment-related words and predict their sentiment contribution in a sentence. This approach enhances the interpretability of sentiment analysis and can effectively correct spelling or grammatical errors in social media or user comments, improving the accuracy of sentiment analysis.

[0073] (3) To address complex language phenomena such as irony and metaphor, an adaptive generation strategy network is designed. Reinforcement learning is used to dynamically adjust the sentiment analysis strategy, enabling the model to flexibly adapt to different modal inputs, especially providing more accurate sentiment analysis results in complex contexts. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a schematic diagram of the main process of the present invention.

[0075] Figure 2 This is a data annotation diagram of an embodiment of the present invention.

[0076] Figure 3 2 is a schematic diagram of a GRU unit processing an audio block according to an embodiment of the present invention.

[0077] Figure 4 Schematic diagram of the structure of a ViT image encoder according to an embodiment of the present invention.

[0078] Figure 5 Schematic diagram of a modal translation module according to an embodiment of the present invention.

[0079] Figure 6 It is the public, noise, and private representation of the embodiment of the present invention.

[0080] Figure 7 Schematic diagram of the sentiment binary extraction model according to an embodiment of the present invention.

[0081] Figure 8 Schematic diagram of a binary optimal matching module according to an embodiment of the present invention.

[0082] Figure 9 Schematic diagram of the bidirectional LSTM network structure of an embodiment of the present invention.

[0083] Figure 10 2 is a schematic diagram of a sentiment polarity analysis module according to an embodiment of the present invention. DETAILED DESCRIPTION

[0084] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.

[0085] For the multimodal sentiment analysis of Mongolian, this paper proposes an aspect-level multimodal Mongolian sentiment analysis method based on a cross-modal attention mechanism, aiming to improve the ability to capture complex emotional information and the analysis accuracy. Figure 1 , the present invention mainly comprises the following steps:

[0086] Step 1: Obtain Mongolian text, images, and Mongolian speech, and use Mongolian text and Mongolian sentiment dictionary to construct target-viewpoint-sentiment triplet data.

[0087] The dataset used in the present invention is a Mongolian multimodal emotion dataset containing seven discrete emotions: happiness, anger, sadness, surprise, fear, disgust, and neutrality (i.e., a relatively stable voice without emotion). Each emotion uses 200 sentences. The dataset is divided into a supervised training set and an unlabeled training set for the three modalities. For the text-based dataset, the preprocessing performed by the present invention is to remove special characters and some punctuation marks. Then, comments of a certain length are randomly extracted.

[0088] At the same time, we collected 50,000 labeled Mongolian sentiment analysis text corpora, over 50,000 labeled GIF short video sentiment corpora, and 1.26 million Mongolian unlabeled corpora. In addition, public multimodal sentiment datasets such as CH-SIMS, CMU-MOSI, CMU-MOSEI, and IEMOCAP are available for reference.

[0089] At the same time, due to the lack of Mongolian language resources, this paper introduces data augmentation. When the amount of data is limited, it generates synthetic data by slightly modifying existing data. The proper use of data augmentation can increase the diversity and quantity of training samples, thereby improving the quality of training datasets and helping to build more effective deep learning models. As an important means to address the problem of data scarcity, data augmentation has received widespread attention in recent years and has been applied in many fields such as image processing and natural language processing, effectively enhancing labeled data and improving the performance of deep networks.

[0090] The data augmentation and expansion in this step are achieved using the synonym replacement method, and are applied to the fine-grained ABSA (Aspect-Based Sentiment Analysis) task.

[0091] To construct the target-opinion-sentiment triplet data for the triplet dataset, a sentiment lexicon and heuristic rules were used to automatically generate high-quality triplet data. To accurately describe the positional information of words, BIO (Begin, Inside, Outside) annotation was employed. This method is a common sequence annotation method, often used to annotate target words, opinion words, and other sequences with boundary information. This method, in the absence of large amounts of manually annotated data, expanded the training dataset and improved model performance.

[0092] like Figure 2 As shown, given a set of Mongolian sentences Each sentence S i Consists of a sequence of words The goal is to get t BIO tags are used to identify words in sentences that describe entities or attributes and to mark their boundaries in sentences.

[0093] Opinion Words Tagging using sentiment lexicon Annotate each word with opinions.

[0094]

[0095]

[0096] The annotation attributes include positive dictionaries and negative dictionaries B-OP is the starting position of the opinion word, and I-OP is the middle position of the opinion word. Aspect Terms Tagging uses part-of-speech tagging (POS Tagging) and dependency parsing to mark possible target words.

[0097]

[0098] Among them, B-ASP is the starting position of the target word, I-ASP is the middle position of the target word (excluding the start and end), and O is a non-target word.

[0099] Sentiment Polarity Tagging determines the sentiment polarity based on the category of opinion words in the sentiment dictionary:

[0100]

[0101] Among them, B-POS represents the starting position of positive opinion words. I-POS represents the middle position of positive opinion words. B-NEG represents the starting position of negative opinion words. I-NEG represents the middle position of negative opinion words. O represents non-sentiment related words.

[0102] In step 2, a modality translation method based on a multi-head attention mechanism is used to convert the visual modality and audio modality into text modality, and the features of each modality are decomposed into public representation, private representation, and noise representation through hierarchical disentanglement.

[0103] This step constructs a unified modality encoding model for processing Mongolian sentiment data across different modalities. First, a modality translation method is used to translate visual and audio modalities into textual modalities to enhance the quality of intermodal representations and ensure the stability of sentiment recognition when either modality is missing. Next, feature decomposition is performed, breaking down modal features into common, private, and noisy representations. This eliminates noise interference while preserving the diversity of sentiment information, improving the accuracy and effectiveness of multimodal sentiment analysis.

[0104] Specifically, the present invention defines that the multimodal data for sentiment analysis includes three modes: P = [X v , X a , X t ], where X v , X a and X t Representing visual, audio, and text modalities, respectively. The aspect-level sentiment analysis task mines deep representations of multimodal data to obtain private, public, and noise representations of text, audio, and images. Within these rich representations, the aspect-level sentiment extraction module extracts aspect terms and opinion terms from the text representation. The audio, image, public, and noise representations are then comprehensively analyzed to derive sentiment polarity, ultimately outputting aspect-level sentiment triples.

[0105] Representation of text features:

[0106] First, a Mongolian sentiment lexicon was constructed. Using a combination of machine and human translation, the Chinese sentiment lexicon was translated into a Mongolian one. To simplify the lexicon, identical words were merged. Furthermore, the sentiment lexicon was expanded using the pointwise mutual information (PMI) method. Words with strong sentiment polarity and high word frequency were selected as seed words. The semantic similarity between unknown words and the seed words was calculated, and sentiment words with high similarity were included in the lexicon. Finally, the translated sentiment words were merged with the expanded sentiment words to generate the Mongolian sentiment lexicon.

[0107] Pointwise Mutual Information (PMI) is an indicator that measures the correlation between two objects. It is widely used in fields such as natural language processing. The formula is as follows:

[0108]

[0109] Among them, P(w i ,w j ) represents the word w i and w j The probability of co-occurrence, P(w i ) and P(w j ) represent the vocabulary w i and w j The probability of appearing alone, the former is an unknown word, and the latter is a seed word.

[0110] Secondly, for text data The WordPiece algorithm is used for training to generate word indexes and word vectors, and then an index dictionary and a vector dictionary are established and converted into array form as the input of the BiLSTM-based text sentiment analysis model. Considering that the lack of labeled data may lead to limited model accuracy of neural network training, pre-training technology is introduced to solve this problem. Pre-training can not only significantly improve model performance, but also effectively model the phenomenon of polysemy. The present invention represents text features as follows:

[0111] E t =BiLSTM(X t )

[0112] Use the pre-trained BiLSTM model to extract the embedded text representation. Where t represents the text modality, and the output of the last layer in the BiLSTM model hierarchy is taken as E t .

[0113] Representation of audio features:

[0114] This paper proposes to build an audio sentiment analysis model based on the Bidirectional Gate Recurrent Unit (BiGRU). First, the OPENSMILE tool is used to extract low-level audio features, and then the BiGRU model is used to further extract suitable audio spatiotemporal sentiment features to support subsequent multimodal fusion.

[0115] Specifically, the present invention first converts the original audio file into an audio spectrogram and extracts key features therefrom. The extraction of audio features mainly uses logarithmic mel-band energy and mel-frequency cepstral coefficients (MFCCs). First, a logarithmic mel spectrogram and prosodic features are obtained from the original audio data, and the logarithmic mel spectrogram is input into an audio emotion analysis model based on BiGRU to capture the emotional spatio-temporal features in the audio:

[0116]

[0117] Where: S mel (m, t) is the energy value of the m-th mel filter in the t-th frame of the logarithmic mel spectrogram, t is the frame index of the audio data, and its value range is determined by the total number of frames of each data. The duration of each audio segment is on average 10 ms. It is agreed that the maximum index length of a data is T, and t = 1, 2,..., T. |X(f)| 2 is the power spectrum obtained by processing the audio with the short-time Fourier transform, and H m (f) is the transfer function of the m-th mel filter. Usually, the maximum value of m is 64 (this value simulates the non-linear frequency characteristics of the human ear's auditory characteristics and is related to the human auditory discrimination ability), and it specifically depends on the task requirements and data characteristics. f min , f max is the minimum and maximum values of the frequency range. The t-th mel-frequency cepstral coefficient MFCC(t) is calculated from the logarithmic mel spectrogram through the discrete cosine transform (DCT):

[0118] [[ID=2I]]

[0119] Where k is the output index of the MFCC coefficient (i.e., the k-th MFCC feature), K is the retained MFCC feature dimension, M is the number of mel filters, which determines the dimension of the mel spectrogram. Generally, K < M to balance the feature complexity and information retention ability, and N is the retained MFCC representation dimension. The mk in the cosine term is the product of the mel filter index m and the MFCC coefficient index k, and its function is to capture the frequency-domain envelope features of the mel spectrogram through the DCT basis function. For example: when k is small, the frequency of the cosine term is low, corresponding to the global spectrum trend represented by the MFCC coefficient; when k is large, it corresponds to local detail features.

[0120] In summary, the present invention obtains the primary representation and gets the audio features through BiGRU:

[0121]

[0122] Through the above method, the BiGRU model can effectively integrate multi-level audio features and provide accurate input support for sentiment analysis, such as Figure 3 shown.

[0123] Representation of image features:

[0124] This step uses the ViT image encoder as the image feature extractor. After pre-training, ViT has performed well in multiple upstream and downstream tasks, and its encoded image features contain rich label information. Image feature extraction first extracts key frames from the video and selects representative and informative key frames as the objects for image feature extraction. For these key frames, OpenFace 2.0 Toolkit is used to extract multi-dimensional facial features, including 68 facial key points, 20 facial shape parameters, HoG (Histogram of Oriented Gradients) features, head posture (pitch, yaw, roll) and line of sight direction, etc., to form the initial image feature U v .

[0125] Then, in order to extract deeper image features, these features are input into Figure 4 The ViT (Vision Transformer) image encoder shown in Figure 1 obtains deep feature representations of key frames by dividing the image into fixed-size patches and modeling global visual information through a self-attention mechanism. Then, to model the temporal dependencies between key frames, the features of the key frames are input into the ViT image encoder to capture the temporal dynamic patterns of the video, and finally a comprehensive image feature X is obtained. v . Its representation process can be formalized as:

[0126] E v =ViT(Preprocess(U v ))

[0127] Among them, U v The initial image features are extracted using the facial behavior analysis toolkit, Preprocess represents the preprocessing operation on the extracted features, and ViT is the feature mapping function of the long short-term memory network.

[0128] After the above steps, we get E t , E a , E v However, their feature dimensions are different. In this case, if they are directly input into the classification layer, the calculation will be complicated. Therefore, a fully connected layer is used for dimensionality reduction. On the one hand, it can reduce the feature dimension and reduce the training difficulty. On the other hand, it explicitly specifies the dimension size to facilitate subsequent classification. The present invention transforms different initial features into a specified size through the fully connected layer FC. The characterization result of the corresponding mode is expressed as:

[0129] F t =FC(E t )

[0130] F a =FC(E a )

[0131] F v =FC(E v )

[0132] in d represents the unified feature dimension. t Indicates the length of the text sequence, that is, the number of basic units (such as words, subwords or characters) contained in the input text. a Indicates the number of frames after audio signal processing, which is determined by the frame length and the average value of the data set. v Indicates the number of frames or key frames in a video, and is usually determined by downsampling or key frame extraction. t 、T a 、T v This parameter reflects the temporal resolution differences between modal data and is a fundamental hyperparameter for multimodal fusion. d represents the unified feature dimension. Given the limited resource requirements of Mongolian, the value of d is set to 256 based on the dataset size, text length, audio length, video length, and modal translation to avoid potential overfitting with 512 dimensions.

[0133] After obtaining the preliminary representation of each modality, in order to achieve the interaction and alignment between cross-modal features, a modal translation module based on Transformer is introduced, such as Figure 5 As shown in Figure 2, the core of this module is a cross-modal multi-head attention mechanism, which effectively captures the semantic connections and potential relationships between different modalities. The attention mechanism plays a crucial role in this process, modeling the relevance of information across modalities by dynamically assigning weights. Specifically, a multi-head self-attention module is used to enable cross-modal information interaction.

[0134] The multi-head attention mechanism projects the embedding vector of each modality into multiple hidden subspaces by introducing multiple independent attention heads, enabling the model to capture information from different feature subspaces in parallel.

[0135] Scaled Dot-Product Attention is then used to calculate the similarity between the embedding vectors and generate a fused contextual representation. This design effectively improves the alignment between representations of different modalities while preserving the uniqueness of each modality.

[0136]

[0137] where q i is the query vector (query); k j is the key vector (key); v j The value vector (value) takes the features of each modality. The multi-head attention mechanism is used to extract information from the different semantic spaces of each modality. The formula of the multi-head attention mechanism is as follows:

[0138] E M =MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O

[0139] Text features, audio features, and image features are used as key and value vectors respectively, and dot product attention is used to dynamically calculate weights. is the weight matrix, h is the number of branches of multi-head attention. The calculation method of the i-th attention head is as follows

[0140]

[0141] in, as well as are the weight matrices of the i-th query, key, and value respectively.

[0142] formula It is used to represent the process of extracting features from the context of a single modality (visual modality) through a multi-head self-attention mechanism. The following is the translation representation of audio and image modality features to text modality:

[0143] D vt =MultiHead(E t ,E v ,E v )

[0144] D at =MultiHead(E t ,E a ,E a )

[0145] Among them, D vt For vision-to-text translation features, D at It is the audio to text translation feature, and MultiHead() represents the multi-head attention mechanism.

[0146] The principle and purpose of modal translation is to address the consistency issue of multimodal data. Modal translation maps fundamentally different modalities into the textual semantic space, achieving a unified representation and providing a foundation for the subsequent disentanglement step. In this disentanglement step, public, private, and noise representations are hierarchically separated, enabling the model to capture emotional consistency while preserving the uniqueness of each modality.

[0147] To construct public representation C and private representation P n and noise characterization N n , the present invention introduces a hierarchical disentanglement module.

[0148] Representation of public representations:

[0149] The public sentiment learning layer, the public encoder generates a public representation for the original features of each modality:

[0150] C=PublicEncoder(E t ,D at ,D vt θ c )=MultiHead(E t ,D at ,D vt )

[0151] Where n∈{t,a,v}. PublicEncoder() represents public sentiment learning, θ c Refers to the parameter weights learned by the PublicEncode network after learning the representation of data with different modalities. The PublicEncode module simultaneously captures the expression characteristics of the same emotion in different modalities in the same sample. t ,D at ,D vt ) expression, E t Represents a semantic space with high reliability (this is because the Mongolian multimodal dataset has the most texts in each modality and the best quality). vt As an example, the multi-head attention mechanism is based on E t As a query, E v As the key and value, the purpose is to let the model learn "how to map visual features to the distribution of text features", that is, to calculate the correlation between text and visual features through the attention mechanism, guide the visual features to align with the text space, and generate new features D that are compatible with the text semantics. vt (i.e., the translated visual features), so that D vt The distribution is close to E t .

[0152] The above formula focuses on preserving the key information and representative features of each modality in the fusion result, ensuring the integrity of multimodal features while improving the recognition ability of the final fused representation.

[0153] Representation of noise characterization:

[0154] The separation of non-public representations is achieved by combining the original features E of different modalities n Subtract the public representation C to get the non-public representation The noise obtained by the present invention is expressed as Private Representation The noise loss γ nl Use the square of the L2 norm to evaluate the size of the noise representation After initializing the module parameters, the noise has the following loss function as the objective:

[0155] Finally, we get the public representation C and private representation P n and noise characterization N n , where private representation refers to the textual and visual modal translation features, such as Figure 6 shown.

[0156] In step 3, the aspect-level sentiment extraction module is used to learn the target-opinion bigrams in the corpus based on the triple data of the text modality as the supervision signal.

[0157] Target words are the components in a sentence that cause emotional change, and are typically the subject and object. Opinion words modify the target words. In this step, aspect words and target words are extracted through a phased aspect-level sentiment extraction module, improving the accuracy and interpretability of sentiment analysis.

[0158] Specifically, to address the challenge of capturing sentence-level sentiment details in sentiment classification, the aspect-level sentiment extraction module employs a phased sub-problem framework to analyze the sentiment expressed in a sentence and then identify the target aspects and opinion terms for that sentiment. This approach improves exploration and sample efficiency while also considering interactions between triple components. Furthermore, the model possesses a certain level of error correction capability, effectively handling text with spelling or grammatical errors.

[0159] like Figure 7 As shown in Figure 2, the phased aspect-level sentiment extraction module is based on two sequence labeling tasks and significantly improves model performance by sharing information between targets and opinions. In the first phase, the model focuses on extracting target words and opinion words, and achieves efficient prediction through word-level boundary labeling. In the second phase, the model determines the matching relationship between targets and opinions based on the sentence context, thereby completing the pairing of bigrams, such as Figure 8As shown in the figure, a joint annotation system integrates target boundary and opinion word information, simplifying the traditional step-by-step approach and improving prediction accuracy. The combination of GCN and dependency relations effectively utilizes syntactic information to identify opinion words and share signals with target information. The two-stage design separates prediction and pairing, making the model more modular and achieving higher performance.

[0160] Furthermore, in the first stage of target word and opinion word extraction, the model first performs the target extraction (AspectExtraction) task. The core of this task is the bidirectional LSTM model based on BIO (Begin-Inside-Outside) sequence annotation, such as Figure 9 As shown, the model aims to predict the boundary of the target word. Through the unified modal encoding model in the previous step, the present invention obtains the embedding and features of the text, and inputs the sentence X t ={x1,x2,…,x T Each word of} is converted into word embedding e(x i ), and encoded through a BiLSTM network.

[0161] h T =BiLSTM(e(X))

[0162] The boundaries of the target word (start, end, etc.) are marked by the space Y T = {B, I, E, S, O} is predicted, and the label contains the target boundary (such as B, I, E, S), based on the training h T , predict the target word boundary label through the Softmax layer. The target word boundary label prediction is to use Softmax to obtain the maximum possible labeling result, which is calculated as follows:

[0163]

[0164] in, It is the result of BiLSTM encoding the input sentence context. W T and b T are the weight matrix and bias terms in the model parameters; is the target boundary prediction result at time step t. In this way, the model assigns a boundary label to each word, such as B (beginning), I (middle), E (end), etc. To further optimize the prediction performance, the target boundary supervised training uses the cross entropy loss function to optimize the model parameters as follows:

[0165]

[0166] in, is the true label at time step t, and T is the length of the sentence. This target extraction method based on sequence labeling accurately captures target word boundaries and provides a foundation for subsequent sentiment labeling. After completing target boundary labeling, to ensure accurate positioning of target words, the next task is word-by-word sentiment labeling (Sentiment Labeling). A challenge of this task is maintaining consistent sentiment labeling for the same target phrase, that is, avoiding inconsistent sentiment labeling for the same target phrase. To address this issue, the model designs a consistency module (SC).

[0167] The hidden state at the current time step and the hidden state at the previous time step The hidden states of the current time step and the previous time step are fused through the gating mechanism. The gating value is calculated as follows Among them, W g and b g is the parameter of the gating mechanism, σ is the sigmoid activation function, g t It is the gate value that controls the degree of fusion between the current and previous time step states.

[0168] The hidden state is updated as follows: Among them, σ is the sigmoid function, g t Is the gate value. Using sigmoid output can effectively smooth the sentiment labeling results and avoid drastic changes in sentiment polarity due to syntactic or semantic noise. In order to further combine the information of target boundary prediction, the model introduces the transformation matrix W tr , mark the target boundary z T Transformed into label distribution z with sentiment polarity S :

[0169]

[0170] Among them, α t is a weighted coefficient based on the confidence of target boundary prediction, defined as:

[0171]

[0172] ∈ is a hyperparameter used to balance the weight of confidence. t It is a weighted coefficient based on the confidence of the target boundary prediction, reflecting the degree of trust the model has in the boundary prediction.

[0173] The first stage ends with the Opinion Extraction task. The model uses a graph convolutional network (GCN) to process dependencies and captures the syntactic associations between the target and opinion words by modeling a dependency graph. The logical structure of the dependency graph is an adjacency matrix A that represents the dependency relationships between words in a sentence. If the i-th word has a dependency relationship with the j-th word, then A ij =1, otherwise A ij = 0. This representation can effectively capture the syntactic structure. The initial representation of each word is composed of its word embedding H. GCN updates the word representation through the adjacency matrix and weight matrix:

[0174] h GCN =ReLU(AHW+b)

[0175] Among them, A is the dependency graph adjacency matrix, H is the input word embedding matrix, W and b are the weight matrix and bias term of GCN. In order to use the information of target words to assist opinion extraction, a target guidance module (TG) is designed to represent the target boundary h T The syntactic dependency feature h generated by GCN GCN Splicing is used to guide the prediction of opinion word boundaries and predict through the softmax layer:

[0176]

[0177] in, This is the opinion word boundary prediction result.

[0178] The first stage includes two subtasks: target extraction and opinion extraction. The final loss function of the model is the weighted sum of the losses of these two subtasks:

[0179]

[0180] The second stage is the pairing of target words and opinion words. In this stage, the targets and opinions extracted in the first stage are paired with each other to determine whether they form a valid triple. In this stage, the model needs to determine the strength of the correlation between the target word and the opinion word to finally output a valid triple. First, the relative distance d between the target word and the opinion word is encoded by position embedding. ij =|pos(i)-pos(j)| Combined with the logistic regression model, the validity of the target word-opinion word pair is judged. The classification function is: Here W and b are the weight matrix and bias term in the classifier parameters, p ij is the pairing probability, is the target word embedding, is the opinion word embedding, in It is the hidden state of the jth word after being processed by the graph convolutional network (GCN), which encodes the syntactic and semantic information of the word, and is used to model the dependency relationship between words and capture the syntactic association between the opinion word and the target word. A is the dependency graph adjacency matrix, H is the input word embedding matrix, W and b are the weight matrix and bias term of GCN. It represents the concatenation of the target embedding and the opinion embedding. The Softmax function generates a probability distribution for the pairing, which is used to judge the effectiveness of the pairing. The output of the classifier is a binary prediction value, indicating whether the target word-opinion word pair is valid.

[0181] In order to optimize the performance of the classifier, the model uses the cross entropy loss function as the training target. The cross entropy loss measures the predicted probability p ij and the true label y ij The classifier training goal is to maximize the probability of correct pairing, that is, to improve the accuracy of pairing by minimizing the deviation:

[0182]

[0183] where y ij represents the true label, p ij is the predicted pairing probability, y ij =1 means the goal-viewpoint pairing is valid, y ij =0 means the pairing is invalid.

[0184] Step 4: Using an adaptive emotion polarity classification network, the emotion polarity is obtained based on the public representation, the noise representation, and the private representations of the visual modality and the audio modality.

[0185] In order to solve the complex language phenomena (such as irony, metaphor, ambiguity, etc.) in sentiment analysis, the present invention proposes an adaptive generation strategy. This strategy trains a policy network through dynamic parameters that can automatically configure to adapt to different input forms. Through this policy network, the model can dynamically adapt to inputs of different modalities, especially when faced with complex language phenomena, to provide more accurate sentiment analysis results. A multimodal sentiment feature fusion model will be constructed, and the extracted text, short video and audio sentiment features will be fused in a cross-modal layered manner at the fusion layer to capture the most effective sentiment information in the context. The final fusion model will provide a more accurate sentiment semantic vector representation, thereby improving the effect of Mongolian multimodal sentiment analysis.

[0186] Specifically, in the multimodal sentiment analysis task, the features of different modalities have different importance for the judgment of sentiment polarity. Dynamic weights allow the model to automatically adjust the loss weights of different tasks according to the difficulty or importance of each task. In order to effectively fuse the features of each modality while suppressing the interference of noise representation, the present invention introduces a dynamic weight allocation module (Dynamic Weight Allocation Module). This module aims to assign dynamic weights to each feature, thereby highlighting the role of important features in the feature fusion process. The basic idea of ​​dynamic weights is that the importance of each modality may be unbalanced, so it is necessary to introduce a feature weight, which is a learnable parameter. During training, the features of each modality will be multiplied by these dynamic weights to control the influence of each modality on model prediction. The overall model is such as Figure 10 shown.

[0187] First, define the following weight w common ,w private_image ,w private_audio ,w noisy Then build a shared feature pool and concatenate all the features.

[0188]

[0189] Among them, d is the representation dimension.

[0190] Calculate the weighted feature attention matrix and learn the weighted feature attention matrix through the feedforward neural network

[0191]

[0192] Where Z is the multimodal feature concatenation vector of public representation, noise representation, and private representation of visual modality and audio modality, and is the weight matrix of the fully connected layer, h is the hidden layer dimension, and is the corresponding bias vector; tanh(·) is the hyperbolic tangent activation function, and σ(·) is the Sigmoid activation function, ensuring that the output is in the range of (0,1).

[0193] Weight normalization, in order to ensure the comparability and stability of weights, this paper averages the attention matrix and applies the Softmax function for normalization:

[0194]

[0195] here, The representation dimensions are averaged to obtain the global weight of each feature. The Softmax function ensures that the sum of all weights is 1, making the contribution of each feature comparable when fused. Feature fusion and final representation, using the generated dynamic weights Perform weighted combination of each feature to form the final joint representation Z joint :

[0196]

[0197] Among them, F i is the i-th feature, w i is the corresponding dynamic weight. This process realizes the adaptive weighted fusion of each modal feature, enabling the model to dynamically adjust the contribution of each modality according to the current input features, thereby improving the accuracy of sentiment prediction.

[0198] Then, the fused joint representation Z joint Input to the fully connected layer to produce the final sentiment prediction result

[0199]

[0200] To optimize the performance of the model, we first minimize the cross entropy loss to improve classification accuracy:

[0201]

[0202] Among them, y i is the actual emotion label, is the probability predicted by the model, and N is the total number of samples.

[0203] The role of noise representation is usually interference, and ideally its weight w noisy Should be close to zero. Noise characterization suppression loss can be introduced:

[0204]

[0205] Where α is the noise suppression coefficient. By minimizing L noisy , the model is encouraged to noisy The value of is close to zero, reducing the interference of noise representation.

[0206] In order to avoid excessive dominance of a single modal feature, this paper introduces an entropy-based weight regularization loss:

[0207]

[0208] When the weight distribution is too biased towards a certain feature, the entropy value decreases, L entropyIncrease. Minimizing this loss term helps promote the diversity and balance of weights. Through adaptive weight adjustment, the contribution of different modalities to the task is adapted; the labels of each modality are made consistent with the multimodal fusion representation, thereby improving the model's ability to understand emotions. When the uncertainty of a certain modality is large, the weight of its loss will automatically decrease, thereby reducing its impact on the final representation. This method can effectively balance the contribution of each modality to multimodal fusion and avoid a "dominant" influence of one modality on other modalities. The final total loss function combines the above items:

[0209]

[0210] Among them, L classification is the cross entropy loss, λ1 and λ2 are hyperparameters used to adjust the importance of each loss term. These hyperparameters are adjusted through experiments to find the best combination.

[0211] In the framework of dynamic weight generation, the optimization objective can be formalized as a multi-objective problem:

[0212] min w,θ L(w,θ)

[0213] Where: w is the parameter of the dynamic weight generation module, θ is the parameter of the feature extraction and classifier.

[0214] By introducing the weight dynamic allocation module, the model can dynamically adjust the weight of each modality according to the characteristics of the input data, thus achieving effective fusion of multimodal features. In particular, using the weighted feature attention matrix It can capture the dynamic contribution of different modalities to emotional expression. Combining noise suppression with weight regularization loss, the model mitigates noise interference while fully utilizing information from each modality, improving the accuracy and robustness of emotion prediction. Ultimately, the optimized total loss function comprehensively considers classification accuracy and the rationality of weight distribution, providing an efficient and robust solution for multimodal sentiment analysis.

[0215] In step 5, the target-opinion bigram is matched with the sentiment polarity to obtain a valid target-opinion-sentiment triplet, thus realizing Mongolian sentiment analysis.

[0216] Through the above steps, the present invention effectively solves the problems of data scarcity, feature alignment and complex semantic processing in Mongolian multimodal sentiment analysis, and improves the accuracy and robustness of sentiment analysis.

Claims

1. A method for aspect-level multimodal Mongolian sentiment analysis based on a cross-modal attention mechanism, characterized by: The steps include: Step 1: obtain Mongolian text, image and Mongolian speech, and construct target-viewpoint-emotion triple data using the Mongolian text and Mongolian sentiment dictionary; Step 2: Use a modality translation method based on a multi-head attention mechanism to convert the visual modality and audio modality into textual modality, and decompose the features of each modality into public representation, private representation, and noise representation through hierarchical disentanglement; Step 3: Use the aspect-level sentiment extraction module to learn the target-opinion bigrams in the corpus based on the triple data of the text modality as the supervision signal; Step 4, using an adaptive emotion polarity classification network to obtain emotion polarity based on the public representation, the noise representation, and the private representations of the visual modality and the audio modality; Step 5: Match the target-viewpoint bigram with the sentiment polarity to obtain a valid target-viewpoint-sentiment triplet, thereby realizing Mongolian sentiment analysis.

2. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 1 is characterized in that: In the step 1, triple data is generated using a Mongolian sentiment dictionary and heuristic rules, wherein the Mongolian sentiment dictionary is obtained by translating and expanding a Chinese sentiment dictionary, and the expansion is achieved based on a point mutual information method.

3. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 1 is characterized in that: The step 2 is to construct a unified modal coding model to perform the modal translation and feature decomposition; the unified modal coding model includes: The modal translation module uses the Transformer's multi-head attention mechanism to achieve cross-modal feature alignment. The formula is: D vt =MultiHead(E t ,HAVE BEEN v ,HAVE BEEN v ) D at =MultiHead(E t ,HAVE BEEN a ,HAVE BEEN a ) Among them, E t is the text feature, E a is the audio feature, E v is the visual feature, D vt For vision-to-text translation features, D at It is the audio to text translation feature, MultiHead() represents the multi-head attention mechanism; Hierarchical disentanglement module, using E t , D at , D vt Construct public representation C and private representation P n and noise characterization N n , where C = PublicEncoder(E t ,D at ,D vt ; γ c )=MultiHead(E t ,D at ,D vt ), n∈{t,a,v}; PublicEncoder() represents public sentiment learning, θ c Refers to the parameter weights learned by the PublicEncode network after learning representations containing data of different modalities, and the private representation P n The calculation method is as follows: After initializing the module parameters, the noise has the following loss function as the objective: in, It is a non-public representation, composed of the original features F of different modalities n Subtract the common representation C to get .

4. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 3 is characterized in that: The text features are obtained through a BiLSTM-based text sentiment analysis model; The audio features are obtained by first obtaining a logarithmic Mel-spectrogram and prosody features from the audio data, and then inputting the logarithmic Mel-spectrogram into a BiGRU-based audio sentiment analysis model to capture the emotional spatiotemporal features in the audio; The image features are obtained through the ViT image encoder.

5. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 4 is characterized in that: The text feature is expressed as: E t =BiLSTM(X t ) where X t For text data; The audio feature is expressed as: Among them, S mel (t) represents the mth Mel filter energy value of the tth frame of the logarithmic Mel spectrum graph, and MFCC(t) is the tth Mel frequency cepstral coefficient. The formulas are as follows: Where t is the frame index of the audio data, t = 1, 2, ..., T, T is the maximum index length of the data, m is the subscript order of the Mel filter, M is the number of Mel filters, which determines the dimension of the Mel spectrum graph, |X(f)| 2 is the power spectrum obtained after short-time Fourier transform processing of audio, H m (f) is the transfer function of the mth Mel filter, f min ,f max is the frequency range, Among them, k is the output index of MFCC coefficient, K is the retained MFCC feature dimension, K <M; The image features are expressed as: E v =ViT(Preprocess(U v )) Among them, U v The initial image features are extracted using the facial behavior analysis toolkit, Preprocess represents the preprocessing operation on the extracted features, and ViT is the feature mapping function of the long short-term memory network.

6. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 1, characterized in that: The aspect-level sentiment extraction module is a target-opinion binary extraction model that performs the following two stages in sequence: In the first stage, target words and opinion words are extracted from Mongolian text input, and prediction is achieved through word-level boundary annotation; In the second stage, based on the sentence context, the matching relationship between the target and the viewpoint is judged to complete the pairing of the tuple.

7. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 6 is characterized in that: In the first stage, the target word and opinion word extraction module and BIO labeling system based on the BiLSTM network and GCN network are used to implement boundary annotation of target words and opinion words. The formula is: in, is the context encoding of BiLSTM output, W T and b T are model parameters, is the target word boundary prediction result at time step t. In the second stage, the relative distance d between the target word and the opinion word is encoded by position embedding ij =|pos(i)-pos(j)|, combined with the logistic regression model to judge the effectiveness of the target word-opinion word pair, the classification function is: Among them, W and b are classifier parameters, p ij is the pairing probability, is the target word embedding, Embedding for opinion words.

8. The aspect-level multimodal Mongolian sentiment analysis method based on the cross-modal attention mechanism according to claim 6 is characterized in that: In the first stage, the target guidance module is introduced to represent the target boundary output by BiLSTM h T The syntactic dependency feature h generated by GCN GCN Splicing, the formula is: in, This is the opinion word boundary prediction result.

9. The aspect-level multimodal Mongolian sentiment analysis method based on cross-modal attention mechanism according to claim 1, characterized in that: The adaptive sentiment polarity classification network includes: Dynamic weight generation module, learning weighted feature attention matrix through feedforward neural network And generate the dynamic weight w of each modality through Softmax normalization i , the formula is: in, Z is the multimodal feature concatenation vector of the public representation, noise representation, and private representations of the visual modality and audio modality, and is the weight matrix of the fully connected layer, h is the hidden layer dimension, and is the corresponding bias vector; tanh(·) is the hyperbolic tangent activation function, and λ(·) is the Sigmoid activation function; Multi-loss joint optimization mechanism, the total loss function is: L=L classification +λ1L entropy +λ2L noisy Among them, L classification is the cross entropy loss, L entropy is the entropy-based weight regularization loss, L noisy is the noise representation suppression loss, and λ1 and λ2 are hyperparameters.

10. The aspect-level multimodal Mongolian sentiment analysis method based on cross-modal attention mechanism according to claim 9, characterized in that: Using the dynamic weight w i , perform weighted combination of each feature to form the final joint representation Z joint : Among them, F i is the i-th feature, w i is the corresponding dynamic weight; The joint representation Z joint Input to the fully connected layer to produce the final sentiment prediction result The optimization goal is formulated as a multi-objective problem: minutes w,θ L(w,θ) Where: w is the parameter of the dynamic weight generation module, θ is the parameter of the feature extraction and classifier.