Aspect-level sentiment analysis optimization method based on multivariate external knowledge fusion
By fusing diverse external knowledge from text, video, and audio, a cross-modal sentiment factor is constructed and counterfactual gain verification is performed. This solves the problems of insufficient alignment of candidate evidence and low credibility of sentiment factors in multimodal sentiment analysis, thereby improving the accuracy and stability of sentiment analysis.
Patent Information
- Application Number
- CN202511400378.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-01-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multimodal aspect-level sentiment analysis methods suffer from problems such as insufficient alignment of candidate evidence, low credibility of sentiment factors, and lack of uncertainty control in prediction results, leading to unstable sentiment analysis results and insufficient robustness.
By using a multi-source external knowledge fusion method, the input text, video and audio evidence is structurally analyzed to construct cross-modal sentiment factors. Counterfactual gain and gating weights are used for weighted fusion to generate aspect-level representations, which are then decoded to output sentiment polarity and intensity.
It improves the accuracy, stability, and interpretability of multimodal sentiment analysis, ensures the authenticity and integrity of candidate evidence, reduces sentiment mismatch rate, and enables the management of prediction confidence and uncertainty.
Smart Images

Figure CN121302248A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-source data analysis technology, specifically to an aspect-level sentiment analysis optimization method based on the fusion of multiple external knowledge sources. Background Technology
[0002] With the development of the internet and artificial intelligence, users generate a large amount of multimodal information, including text, images, audio, and video, on social media, e-commerce platforms, and online services. Enterprises and research institutions urgently need to accurately extract users' subjective sentiments from this data for product improvement, service optimization, and market analysis. Traditional sentiment analysis methods are mainly based on text input, predicting sentiment tendencies through keywords, dependency syntax, or deep learning models. However, this approach relies excessively on a single modality and struggles to identify the true emotional signals implicit in speech tone, image expressions, or video scenes. On the other hand, while existing aspect-level sentiment analysis can pinpoint fine-grained aspects such as "price," "service," and "environment," it generally suffers from inaccurate candidate evidence, insufficient modality alignment, and a lack of credibility control in the output results. For example, when transitions or negations appear in the text, the model is prone to polarity errors; when there are conflicts between multimodal information, the system cannot reasonably explain the relationships between the modalities, leading to unstable sentiment analysis results. Furthermore, most existing methods output a single sentiment prediction, failing to consider prediction uncertainty and ensemble coverage issues, which can easily lead to insufficient robustness and overconfidence risks in practical applications. Therefore, there is an urgent need for a novel aspect-level sentiment analysis method that can integrate diverse external knowledge, construct cross-modal sentiment factors, use counterfactual gain and consistency measures for evidence verification, and improve the reliability of results through gating fusion and ensemble prediction, so as to enhance interpretability and stability in multimodal scenarios. Summary of the Invention
[0003] In view of the above-mentioned problems, the present invention is proposed.
[0004] Therefore, this invention aims to solve the problems of insufficient alignment of candidate evidence, low credibility of sentiment factors, and lack of uncertainty control in prediction results in multimodal aspect-level sentiment analysis.
[0005] To address the aforementioned technical problems, this invention provides the following technical solution: an aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion, comprising:
[0006] The system performs structural analysis on multimodal evidence from input text, video, and audio to obtain aspect-level knowledge atoms that are combined with context, and generates processed content vectors. At the same time, it constructs cross-modal sentiment factors from candidates of different modalities.
[0007] Construct masked or isometric replacement controls on the same sample, and use the sentiment factor to generate counterfactual gains in sentiment at different levels.
[0008] A granularity consistency measure is obtained between aspect terms and candidate local units;
[0009] By using gating weights, the candidate and context are weighted and fused to form an aspect-level representation;
[0010] On the aspect-level representation, decoding is performed to output aspect polarity and intensity; and a controlled set prediction is generated.
[0011] As a preferred embodiment of the aspect-level sentiment analysis optimization method based on multi-dimensional external knowledge fusion described in this invention, the structural analysis includes: decomposing the input multimodal evidence into the smallest unit content vector through the recognition models corresponding to each multimodal mode; and obtaining the association and strong coupling relationship between each content vector and other content vectors in the same modality based on the recognition of the contextual relationship of the evidence in each modality according to the corresponding recognition model.
[0012] Each content vector corresponds to a candidate, and the two have the same index; the association relationship includes the direction and intensity of the influence of other content vectors on content vector i.
[0013] If n content vectors must exist simultaneously in order to make a correct semantic expression, then there is a strong coupling relationship between these n content vectors.
[0014] By clustering content vectors at the aspect level, a cluster family is obtained for each aspect level j. Each candidate in the family is a knowledge atom selected at aspect level j.
[0015] As a preferred embodiment of the aspect-level sentiment analysis optimization method based on multi-external knowledge fusion described in this invention, the cross-modal sentiment factors include candidate i and content features under different modalities, which are mapped to the dimensions of the other through a pre-trained neural network, and the recognition model of the other dimension is called to identify the correlation.
[0016] We obtain the relationship between candidate i and different content vectors in each modality under each modality dimension; and the relationship between the content vector in each modality and the content vector of candidate i under the candidate i dimension.
[0017] Finally, for every two cross-modal candidates, a correlation is formed in two directions.
[0018] As a preferred embodiment of the aspect-level sentiment analysis optimization method based on multivariate external knowledge fusion described in this invention, the counterfactual gain includes determining, within the same sample, a minimum sufficient set composed of text trigger fragments, corresponding salient regions of images, and speech keyframes based on the strong coupling relationship and cross-modal association between the candidate and the candidate.
[0019] Semantic-preserving masking interventions are implemented on the minimum sufficient set, wherein text fragments are replaced with neutral placeholders while keeping part-of-speech and dependency roles unchanged, salient regions of images are locally weakened or prototype-repaired while keeping categories unchanged, and speech keyframes are replaced with energy or frequency band equivalents while keeping prosodic rhythms consistent.
[0020] An isometric substitution intervention is performed on the minimum sufficient set, wherein text trigger words are replaced by neutral neighbor words within the same part of speech and synonym range, image regions are replaced by similar regions within the same category prototype, and speech segments are replaced by adjacent segments within the same speaker and context.
[0021] Through rule-based screening and small-sample exploration, the most stable replacement option is selected under the conditions of semantic consistency, opinion holder consistency, and scope consistency, ensuring that the control sample does not change the non-target aspects of the judgment.
[0022] By comparing the outcome assessments before and after the intervention, the emotional differences of the target candidate are obtained. The direction of the difference is corrected and truncated by combining the intensity of enhancement and suppression, thereby generating the counterfactual gain value and credibility of the candidate.
[0023] The granular consistency measure includes distinguishing data of different modalities in the aspect-level selected knowledge atoms corresponding to each aspect word; obtaining the content vector under each modality; and using the content vector under each modality to input the pre-trained encoder to obtain the consistency feature vector under each modality.
[0024] As a preferred embodiment of the aspect-level sentiment analysis optimization method based on multi-dimensional external knowledge fusion described in this invention, the gating weights include weighted fusion of candidates on each aspect word through multi-level weight rules to obtain the aspect-level representation.
[0025] The specific process is as follows:
[0026] By analyzing the consistency feature vector under each modality k and the distribution of all candidate encoded words under the corresponding aspect word in modality k, the first weight of each candidate in the discrete state is output by a pre-trained mapping function through discreteness judgment.
[0027] By encoding the consistent feature vector under each modality, the total feature vector under each aspect is obtained; the consistent feature vector under each modality is then mapped onto the total feature vector after dimensionality transformation, and the change ratio of the vector length is used as the second weight of each candidate under the current modality.
[0028] For each modality k, the discrete distribution of all candidates is similar. After feature dimensionality reduction through principal component analysis, the similarity between the discrete state of each candidate i and the discrete state of other modalities is analyzed. The goal is to capture the largest range with similarity greater than a threshold. The ratio of the largest range to the discrete range of modality k is obtained as the third weight of candidate i. Let the two modes for comparison be k and g, and obtain the set K under modality k and the set G under modality g. Based on the discrete position of each candidate after dimensionality reduction, the elements between sets K and G are paired to make the correspondence between all paired elements and the discrete relationship the strongest. Using the pairing results, according to the corresponding association relationship, a preset transformation function is used to map the fourth weight of each candidate.
[0029] For candidates that did not participate in the pairing, the fourth weight was set to 1.
[0030] Meanwhile, the candidates for other modalities are analyzed in the same way as the fourth weight to obtain the fifth weight; after multiplying all weights and candidates by the counterfactual gain and summing them by weight, the aspect-level representations in each dimension are obtained.
[0031] As a preferred embodiment of the aspect-level sentiment analysis optimization method based on multi-external knowledge fusion described in this invention, the aspect-level representation is decoded by a decoder so that the aspect-level representation is mapped onto each of the directional words, including sentiment description and sentiment intensity.
[0032] As a preferred embodiment of the aspect-level sentiment analysis optimization method based on multi-external knowledge fusion described in this invention, the generation process of the set prediction includes: searching for similar descriptions based on the sentiment descriptions to obtain all sentiment descriptions, and then superimposing their intensities as first-level set elements.
[0033] Based on the emotional description and emotional intensity, similar descriptive content is retrieved as a strong coupling relationship to obtain secondary set elements;
[0034] The set prediction is obtained by integrating the elements of the first-level set and the elements of the second-level set.
[0035] An aspect-level sentiment analysis optimization system based on multi-external knowledge fusion using the method described in this invention, wherein: a collection unit performs structural analysis on multimodal evidence of input text, video and audio to obtain aspect-level selected knowledge atoms combined with context, and generates a processed content vector; simultaneously, cross-modal sentiment factors are constructed from candidates of different modalities;
[0036] The analysis unit constructs masked or equidistant replacement controls on the same sample and uses the sentiment factor to generate counterfactual gains in sentiment at different levels.
[0037] The processing unit obtains a granularity consistency measure between aspect terms and candidate local units;
[0038] The fusion unit uses gating weights to perform weighted fusion of candidates and context to form aspect-level representations;
[0039] The output unit performs structured decoding on the aspect-level representation, outputting aspect polarity and intensity; and generates a controlled set prediction.
[0040] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, it implements the steps of the method described in any one of the present invention.
[0041] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the present invention.
[0042] The beneficial effects of this invention are as follows: By using multimodal structural analysis, the input text, video, and audio are decomposed into the smallest unit content vector. Combined with the strong coupling relationships and cross-modal associations between candidates, a cross-modal sentiment factor is constructed, thus ensuring the authenticity and completeness of candidate evidence. Based on this, the importance of candidates in different modalities is verified using the counterfactual gain method, and the interpretability and robustness of the model are enhanced through masking and equidistant substitution experiments. Furthermore, by using optimal transmission matching and granular consistency measurement, accurate alignment between candidates and aspect words in different modalities is ensured, effectively reducing the sentiment mismatch rate. This invention also designs a multi-level gating weight mechanism to dynamically fuse candidates and context, generating a reliable aspect-level representation. In the output stage, through structured decoding and ensemble prediction, not only aspect polarity and intensity are output, but also a controlled ensemble result is provided, achieving prediction confidence and uncertainty management. Compared with existing technologies, this invention can significantly improve the accuracy, stability, and interpretability of sentiment analysis in multimodal scenarios. Attached Figure Description
[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 The first embodiment of the present invention provides an overall flowchart of an aspect-level sentiment analysis optimization method based on the fusion of multiple external knowledge.
[0045] Figure 2 The following is a model framework diagram of an aspect-level sentiment analysis optimization method based on multivariate external knowledge fusion, provided for the second embodiment of the present invention. Detailed Implementation
[0046] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0047] Example 1, referring to Figure 1 As an embodiment of the present invention, an aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion is provided, comprising:
[0048] S1: Perform structural analysis on the multimodal evidence of input text, video and audio to obtain aspect-level selected knowledge atoms combined with context, and generate the processed content vector; at the same time, construct cross-modal sentiment factors from candidates of different modalities.
[0049] The structural analysis includes decomposing the input multimodal evidence into content vectors, the smallest unit, using the corresponding recognition models for each modality; and, based on the recognition models' identification of the contextual relationships of the evidence in each modality, obtaining the association and strong coupling relationships between each content vector and other content vectors within the same modality. Assume each content vector corresponds to a candidate, with both having the same index; the association relationship includes the direction and intensity of the influence of other content vectors on content vector i. If n content vectors must exist simultaneously to produce a correct semantic expression, then the strong coupling relationship exists between these n content vectors.
[0050] It's important to note that counterfactual experiments require "masking" or "replacing" local evidence to see if it affects the results. Without the constraint of strong coupling, it's possible to only mask "degree words" while retaining "emotion words," causing the model to mistakenly believe that the emotion still holds. With strong coupling, counterfactual gains will intervene in these necessary elements in groups, ensuring that the experiment truly reflects the necessity of the candidate for emotion determination. By clustering content vectors at the aspect level, a cluster family is obtained for each aspect level j. Each candidate in the family is a knowledge atom selected at aspect level j.
[0051] The cross-modal sentiment factor includes candidate i and content features from different modalities, which are mapped to each other's dimensions using a pre-trained neural network. The recognition model for each other's dimensions is then invoked to identify the correlation. This yields the correlation between candidate i and different content vectors in each modality's dimension; and the correlation between content vectors in each modality and candidate i's content vectors in each modality's dimension. Finally, for every two cross-modal candidates, two-way correlations are formed.
[0052] Specifically, S1 includes two parallel paths: one performs structural analysis on text, video, and audio evidence, producing the smallest unit content vector and its contextual relationship within the same modality; the other constructs cross-modal sentiment factors at the candidate level, providing comparable and referable intermediate quantities for subsequent counterfactual comparison, optimal transmission matching, and gating fusion. The input is first normalized to the timeline and segment index according to modality, establishing a three-level index of sample → segment → smallest unit to ensure that results from different modalities can be referenced and replayed in subsequent steps.
[0053] The text segmentation process involves dividing sentences into clauses, words, and sub-words, labeling them with parts of speech, dependency relationships, and semantic roles, identifying negation, contrast, comparison, degree, and emotion trigger words, and supplementing with referential resolution and opinion holder annotations to obtain several "text content vectors." Each vector includes: the start and end positions of character / word fragments, part-of-speech and role labels, candidate polarity clues, the clause and scope boundary, and the relative distance to aspect words. A co-modal relationship graph is constructed based on syntactic edges and rhetorical markers to calculate the direction (modification, negation, contrast, causation) and intensity (combined by dependency distance, trigger word weight, and co-occurrence frequency) of influence between units. When multiple units must co-occur to express complete semantics (e.g., degree words and emotion words, negation words and target predicates), they are labeled as strongly coupled relationships and their grouped occurrences are recorded as constraints.
[0054] On the video / image side, the video is first segmented into shots and keyframes are extracted. Object detection / segmentation and scene recognition are then performed on the keyframes and neighboring frames to obtain "region-level content vectors," which include spatial location, category and attributes, salience weights, and alignment information with the subtitle / dialogue timeline. Within the same frame, regions are established into comodal relationships based on spatial proximity, interaction, and shared visual emotional cues, and possible rhetorical cues are labeled (e.g., the subjectivity of on-screen subtitles, facial expressions, and gestures). When multiple regions collectively constitute a certain semantic meaning (e.g., "service counter + queuing crowd" both pointing to "service congestion"), they are categorized as strongly coupled groups.
[0055] On the audio side, speech endpoint detection and speaker separation segmentation are used to extract prosody, energy, fundamental frequency, and emotion-related acoustic features. When necessary, these features are combined with automatic speech recognition to form a "speech content vector." For consecutive segments from the same speaker or keyframes within the same semantic unit, homomodal relationships are established based on temporal proximity and prosodic consistency. If multiple acoustic features must appear simultaneously to express a stable subjective attitude (e.g., high energy + rapid speech rate + rising intonation), they are labeled as strongly coupled relationships, and the necessary combination is recorded.
[0056] The calculation of association relationships is performed separately in three modalities, and the output is a bidirectional record of the influence exerted and the influence applied to each content vector, including direction labels and intensity scales, for use in clustering and subsequent selection. (In this embodiment, the "association relationship" is implemented collaboratively by an explanatory model and a structured analyzer: the explanatory model is used for semantic and causal tendency discrimination based on prompts, and the structured analyzer is used to provide syntactic, rhetorical, and temporal boundaries. The outputs of both are fused into direction and intensity through consistency rules; the explanatory model can be a language model for fine-tuning instructions, and the structured analyzer can be a combination of dependency parsing, semantic role labeling, scene / emotion recognition, and speaker separation models.)
[0057] When performing aspect-level clustering, a candidate set is first selected based on the similarity between the aspect vocabulary and the candidate context, dependency path matching, and holder consistency. Then, intra-aspect clustering is performed on a heterogeneous graph. The nodes of the graph are trimodal content vectors and candidates, and the edges are intramodal relationships and cross-modal weak links (from time alignment and caption alignment). The clustering output is a cluster family for each aspect j, and the elements within the family are the aspect-level selected knowledge atoms. All elements and their relationships are preserved so that subsequent counterfactual and transport matching operate only within the necessary scope, avoiding interference with irrelevant evidence.
[0058] Cross-modal sentiment factors are constructed at the granularity of candidate i. First, the content vector of candidate i in its original modality is input into a pre-trained encoder to obtain a standardized representation. Then, it is projected onto the semantic spaces of other modalities through a cross-modal mapping network to obtain comparable representations in the other modalities. After projection, the recognition model of the other modality is invoked to retrieve the content vector most relevant to the projected representation within the other modality. The similarity, spatiotemporal / temporal proximity, consistency of subjective cues, and source credibility between the two are recorded, generating a "candidate → other modality" association set. Simultaneously, the content vector in the other modality is mapped back to the candidate's modality and compared with the candidate's original representation to generate a "other modality → candidate" association set. The bidirectional results jointly form the association relationship between candidate i and different content vectors in each modal dimension, and the credibility level and direction label (same-direction support or opposite-direction cancellation) are labeled according to consistency, alignment quality, and conflict degree.
[0059] For any two candidates belonging to different modalities, cross-modal relationships are constructed in two directions based on "the matching situation of candidate A after mapping to candidate B modality and its neighborhood" and "the matching situation of candidate B after mapping to candidate A modality and its neighborhood". These relationships describe the mutual support or suppression of cross-modal evidence in this aspect. When both directions are strongly supportive and the source is credible, it is denoted as a cross-modal enhancement factor. When at least one direction is strongly contradictory and the consistency is insufficient, it is denoted as a cross-modal suppression factor. All factors are written into an attribute table with the candidate as the key, including enhancement / suppression flags, strength levels, involved local units of the other party, and time / space indices. A backtracking path is also provided so that the most affected local range can be prioritized in subsequent counterfactual comparisons, used as a source of matching cost in optimal transmission matching, and used as upper / lower bound signals in gated fusion for candidate selection.
[0060] In this embodiment, two advanced techniques, GloVe and BERT, are employed to obtain word vector representations, thereby fully mining the semantic information in the text. GloVe is an unsupervised learning algorithm specifically designed to generate high-quality word embeddings. This model accurately calculates the co-occurrence probabilities between word pairs in a massive text corpus, constructs a co-occurrence matrix, and learns word vectors based on this matrix. Specifically, the GloVe model first constructs a co-occurrence matrix covering global information based on the provided corpus. In this matrix, the element Xij represents the frequency with which word i, as an aspect word, and word j, as a context word, co-occur within a pre-defined context window. In this way, GloVe can capture the co-occurrence patterns of words in the global corpus, thereby learning vector representations that reflect the semantic similarity and association of words. For example, in a corpus of movie reviews, the words "exciting" and "shocking" may frequently appear in similar contexts. The GloVe model will learn that their vector representations are close in distance in the vector space, reflecting their semantic similarity.
[0061] BERT improves upon the Transformer architecture by employing a multi-layered bidirectional Transformer autoencoder. BERT treats sentences as complete scenarios, extracting semantic and lexical information from massive amounts of unlabeled text data through unsupervised learning, and dynamically capturing the polysemy of words in different contexts, thus effectively addressing the challenge of polysemy.
[0062] The core of the BERT framework lies in two stages: pre-training and fine-tuning. During pre-training, the model learns word meaning and semantic information from massive amounts of unlabeled corpora. In the fine-tuning stage, the model parameters acquired during pre-training are used to fine-tune the model for specific downstream tasks, effectively improving the training efficiency of downstream models. The pre-training stage of the BERT model mainly includes two tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP).
[0063] In MLM tasks, the model processes the input corpus, randomly selecting 15% of the words. For these selected words, the model performs a masking operation, replacing them with the [Mask] symbol. The model then uses contextual information to predict the masked words. For example, in processing an unlabeled sentence "the food is delicious," "delicious" might be masked with [Mask], making the sentence "the food is [Mask]." The model is then trained to predict the word at the [Mask] position, with the training goal of maximizing the probability of predicting "delicious." This masking mechanism allows the model to simultaneously focus on local grammatical features and global semantic relationships during prediction, thereby learning accurate representations of words in different contexts.
[0064] NSP, on the other hand, is a binary classification task that focuses on modeling discourse coherence. It aims to train the model to determine whether a second sentence immediately follows the first, thereby learning the coherence between sentences and capturing the structure and contextual relationships of the text. The collaborative training of these two tasks enables BERT to capture both the fine-grained semantics of words and the high-level pragmatic structure of text, forming a general language representation with transferability.
[0065] By combining GloVe and BERT to obtain word vector representations, their respective advantages can be fully utilized. GloVe provides word vectors based on global co-occurrence information, reflecting the semantic relationships of words in the overall corpus; while BERT can capture the dynamic semantics of words in different contexts and handle the problem of polysemy. The fusion of these two word vector representations will provide richer and more accurate input for subsequent semantic information extraction and fusion, helping to improve the performance of natural language processing tasks such as sentiment analysis.
[0066] However, simply obtaining word vector representations of individual words is insufficient for a comprehensive understanding of text semantics. This is because text semantics is often expressed by combinations of multiple words. Therefore, this module further employs different semantic information acquisition methods to mine deeper semantic information from text. These methods include, but are not limited to, neural network-based language models, such as recurrent neural networks (RNNs) and their variants (LSTM, GRU), which can capture long-range dependencies in text; and convolutional neural networks (CNNs), which can extract local features from text. Through these methods, semantic information at multiple different levels is extracted from text, such as lexical semantics, sentence semantics, and discourse semantics.
[0067] To fully utilize these different levels and types of semantic information, this module fuses them. Fusion can be achieved through a simple concatenation operation, directly joining the vectors corresponding to different semantic information; or through a more complex weighted fusion method, assigning different weights to different semantic information based on their importance, and then performing a weighted sum. Through this fusion operation, a multi-semantic feature representation containing rich semantic information is ultimately obtained, which can more comprehensively and accurately reflect the semantic content of the text.
[0068] In semantic analysis, relying solely on the information within the text itself often falls short of achieving ideal sentiment analysis results. Therefore, introducing external knowledge enriches the text representation and enhances the model's ability to understand text sentiment. First, external sentiment knowledge (SenticNet6) is integrated into the grammatical dependency tree. A grammatical dependency tree is a tree-like structure that describes the grammatical relationships between words in a sentence, clearly showing the dependencies between various components. SenticNet6 is a resource containing rich sentiment knowledge, assigning sentiment information such as sentiment polarity and intensity to a large number of words and phrases. By combining the sentiment knowledge from SenticNet6 with the grammatical dependency tree, corresponding sentiment information can be added to each node (i.e., word) in the grammatical dependency tree. For example, in a sentence describing a movie, the grammatical dependency tree might show that the adjective "wonderful" modifies the noun "movie," but by integrating SenticNet6, it can be determined that "wonderful" has a strong positive sentiment polarity. In this way, the grammatical dependency tree not only contains the grammatical relationships between words but also rich sentiment information, thus enriching the grammatical dependency tree information of the text and enabling the model to better understand the sentiment tendency of words in a sentence and the sentimental connections between them. Simultaneously, this module also integrates concept knowledge (Microsoft ConceptGraph) into aspect words. Aspect words are a key focus in sentiment analysis; for example, in the sentence "This phone's screen is very clear," "screen" is an aspect word. The Microsoft ConceptGraph is a large-scale concept knowledge graph containing a large number of concepts and the relationships between them. By matching and associating aspect words with the Microsoft ConceptGraph, more conceptual information related to the aspect words can be obtained. For example, for the aspect word "screen," related concepts such as "resolution," "color display," and "touch sensitivity" can be obtained from the Microsoft ConceptGraph. Then, this information from related concepts is integrated into the aspect word representation, resulting in an expanded aspect word representation. This expanded aspect word representation can more comprehensively describe the features and attributes of aspect words, enabling the model to consider more factors related to aspect words when analyzing sentiment, thereby improving the accuracy and detail of sentiment analysis.
[0069] Furthermore, in text processing, a multi-dimensional strategy can be adopted to obtain comprehensive and accurate information to improve classification performance. First, Graph Convolutional Networks (GCNs) are used to deeply analyze the text and accurately capture its hidden grammatical structures, laying the foundation for text understanding. Simultaneously, mutual information (PMI) and cosine similarity (COS) methods are used to effectively extract semantic structural information from the text and clarify the semantic relationships between words.
[0070] By organically integrating semantic and grammatical information, a comprehensive understanding of the text can be obtained, leading to a more thorough comprehension. Further integration of conceptual knowledge can expand the word count within the text, helping to resolve ambiguities and polysemy, thus clarifying the text's meaning. Moreover, introducing sentiment knowledge enhances the ability to identify emotional words in the text, accurately grasping the emotional tone implied by the text. Through this comprehensive approach, by deeply mining textual information from multiple levels—grammatical, semantic, conceptual, and sentiment—the effectiveness of text classification is significantly improved.
[0071] S2: Construct masking or isometric substitution controls on the same sample, and use the sentiment factor to generate counterfactual gains in sentiment at different levels.
[0072] The counterfactual gain includes determining, within the same sample, a minimum sufficient set consisting of text trigger segments, corresponding image salient regions, and speech keyframes, based on the candidate's strong coupling relationship and cross-modal association.
[0073] Semantic-preserving masking interventions are implemented on the minimum sufficient set, wherein text fragments are replaced with neutral placeholders while preserving part-of-speech and dependency roles, salient regions of images are locally weakened or prototype-repaired while retaining categories, and speech keyframes are replaced with energy or frequency band equivalents while maintaining prosodic rhythm.
[0074] An isometric substitution intervention is performed on the minimum sufficient set, wherein text trigger words are replaced by neutral nearest neighbor words within the same part of speech and synonym range, image regions are replaced by similar regions within the same category prototype, and speech segments are replaced by adjacent segments within the same speaker and context.
[0075] Through rule-based screening and small-sample exploration, the most stable replacement option is selected under the conditions of semantic consistency, opinion holder consistency, and scope consistency, ensuring that the control sample does not change the judgment of non-target aspects.
[0076] It's important to understand that when constructing counterfactual comparison samples, all possible replacements must first be screened according to rules to ensure that the replaced samples remain structurally and semantically comparable. Rule screening involves checking candidate replacements item by item according to preset conditions; only replacements that meet these conditions can proceed to the subsequent small-sample exploration phase. These conditions mainly include: semantic consistency, meaning the replacement must not disrupt the original semantic framework; words replaced in text should maintain the same part of speech and similar semantics; and replaced areas or segments in videos and audio should vary within the same category or rhythmic range. Opinion holder consistency, meaning the replacement must not change the subject expressing the sentiment, for example, maintaining the customer's evaluation rather than the merchant's. Scope consistency, meaning the replacement must not disrupt the boundaries of semantic scopes such as negation, contrast, or comparison; for example, words within a negative scope must remain within that scope after replacement. Non-target aspect stability, meaning the replacement must not cause changes in the sentiment judgment of irrelevant aspects, avoiding polarity errors in other aspects due to the operation of a single candidate. By filtering through the aforementioned rules layer by layer, it can be ensured that the differences between the control sample and the original sample are concentrated only on the target candidate itself. This allows the intervention results to truly reflect the necessity and strength of the candidate in sentiment assessment, while avoiding interference caused by invalid replacements or excessive modifications. This process not only improves the credibility of the counterfactual gain value but also provides a reliable foundation for subsequent gain calculations, granularity consistency measurements, and gating weight allocation.
[0077] By comparing the outcome assessments before and after the intervention, the emotional differences of the target candidate are obtained. The direction of the difference is then corrected and truncated by combining the intensity of enhancement and suppression, thereby generating the counterfactual gain value and credibility of the candidate.
[0078] Specifically, S2 conducts intervention and control within the same sample for each candidate. First, constrained by strong coupling and cross-modal association, it locates the minimum sufficient set of candidates: on the text side, dependency syntax and semantic role labeling locate emotion triggers, degree words, negation / contrast markers, and their scope; on the video side, shot segmentation, object detection / segmentation, and saliency analysis extract salient regions aligned with the text trigger segments on the timeline; on the audio side, speech endpoint detection and speaker separation are used, combined with automatic speech recognition timestamps and emotional acoustic features, to extract keyframes synchronized with the text segments. A cross-modal mapping network projects the three modal units onto a unified representation space to establish comparability, while temporal and spatial indices are preserved for replayable localization; this yields a local evidence set covering only the necessary semantic components, serving as the boundary for intervention.
[0079] Within the aforementioned boundaries, semantically preserving masking interventions are performed. Text fragments are replaced with neutral placeholders while preserving part-of-speech, syntactic function, and scope. The quality of the replacement is jointly determined by a context-sensitive language model and a constraint checker. Video regions are locally weakened, blurred, or prototype-repaired. The repair mask is provided by a segmentation model and a saliency heatmap, and category consistency is verified by a re-identification network. Audio keyframes undergo equivalent replacements for energy, frequency band, formants, or fundamental frequency. Time-domain or frequency-domain signal processing algorithms are used to maintain prosodic rhythm and speaker embedding stability. This type of intervention masks candidate evidence without disrupting syntactic structure, visual categories, and phonological prosody, thus creating a comparable "de-evidence" contrast.
[0080] To reduce distribution drift, equidistant replacement interventions are constructed within the same boundary. On the text side, neutral nearest neighbors are retrieved from the same part of speech, synonyms, or synonym groups; semantic similarity and style consistency are jointly constrained by a pre-trained text encoder and style discriminator. On the video side, semantically and visually similar local regions are retrieved from a prototype library of the same category and seamlessly replaced, maintaining continuity in scale, perspective, and background. On the audio side, adjacent or similar prosodic segments are selected from a segment library of segments from the same speaker and context for splicing, using duration / pitch-independent time-frequency transformations to ensure a natural transition. A small-scale candidate pool is established using candidates as keys; replacements are further filtered after being scored for cross-modal consistency and time alignment quality.
[0081] All masking and replacement results are subjected to rule-based filtering and small-sample exploration. Rule-based filtering checks four types of constraints: semantic consistency (semantic framework unchanged, same part of speech or category, no audio prosody distortion), opinion holder consistency (expression subject unchanged), scope consistency (negation, contrast, comparison, and citation boundaries not disrupted), and non-target aspect stability (initial judgments in other aspects are unaffected). Replacements that fail any of these constraints are directly eliminated. Small-sample exploration performs multiple rounds of evaluation on the remaining candidate pool: each replacement is repeatedly inferred under a fixed random seed and the same inference configuration, using variance and consistency thresholds to filter out items that have unstable impacts on the target aspect or fluctuate on non-target aspects; if necessary, an early stopping strategy and a small amount of adaptive probing are used, prioritizing replacements that perform stably in multiple rounds of evaluation.
[0082] After intervention and screening, aspect-level assessments were performed on both the original and control samples, recording the polarity and intensity changes of each aspect, and separately measuring the magnitude and direction of changes in text, video, and audio. Cross-modal alignment quality, source credibility, and inter-candidate conflict were used as weighting factors to weight and summarize the changes in the three channels to obtain a single sentiment difference. If the direction of the difference contradicted the enhancement / suppression state in the previous cross-modal sentiment factors, the direction was corrected according to the degree of contradiction and alignment quality; if the magnitude of the difference was abnormal and accompanied by low consistency or high conflict, the magnitude was truncated to suppress spurious effects. The same candidate was repeatedly and independently intervened and evaluated, and the results from multiple trials were robustly aggregated to obtain a counterfactual gain value. Credibility was jointly measured by repeatability consistency, alignment quality, source credibility, and rule triggering records. Counterfactual gain and credibility serve as necessity and risk signals for subsequent gating solutions, while intervention logs (boundary indexes, mask and replacement pointers, screening and exploration records) are retained to support auditing and replayability.
[0083] S3: Obtain a granularity consistency measure between aspect terms and candidate local units.
[0084] In this context, a candidate local unit is the smallest fragment of content in multimodal evidence (text, video, audio) that has a potential connection to a specific aspect term. It is not the full information of the entire modality, but rather a "candidate" that, after fine-grained segmentation, may influence the emotional expression of that aspect.
[0085] The granular consistency measure includes distinguishing data of different modalities in the aspect-level selected knowledge atoms corresponding to each aspect word; obtaining the content vector under each modality; and using the content vector under each modality to input the pre-trained encoder to obtain the consistency feature vector under each modality.
[0086] Specifically, in this embodiment, the text encoder is a context-sensitive transformer-type language model that outputs a fixed-dimensional semantic vector after fine-grained word pooling; the video encoder is a region-level visual encoder that extracts RoI features and encodes the location of candidate regions obtained from detection / segmentation before outputting region vectors; and the audio encoder is a self-supervised acoustic model or a pre-trained emotion recognition model that generates vectors containing energy, prosody, and emotional cues within a short time window. In other embodiments, the text encoder can be replaced with a sentence vector model or a lightweight convolutional / gated RNN, the video encoder can be replaced with a visual Transformer or a lightweight backbone network, and the audio encoder can be replaced with an MFCC+bidirectional RNN or a Conformer. For small-scale deployment scenarios, a distilled version of the model can be used to reduce computational overhead.
[0087] It's important to note that the raw data formats of text, video, and audio differ significantly. Direct comparison can lead to inconsistencies in information hierarchy and granularity, thereby compromising the accuracy of aspect-level sentiment analysis. Therefore, it's necessary to use pre-trained encoders for each modality to transform the raw segments into structured vector representations while maintaining the same scale and semantic granularity.
[0088] On the text side, a context-sensitive transformer model is used to capture context-dependent and fine-grained subjective expressions, preventing emotional coloring from being diluted in long texts or complex semantics. On the video side, a region-level visual encoder is used to highlight salient areas corresponding to aspect words, so that emotion judgment is not limited to the overall picture, but falls on the local visual details related to the aspect. On the audio side, acoustic or emotion recognition models are used to extract implicit cues reflecting subjective attitudes in rhythm, energy, and tone, thereby making up for the emotional information that cannot be reflected in the text.
[0089] By employing this method of independent encoding within a modality and unified dimensionality across modalities, it is ensured that candidate local units corresponding to each aspect word can output directly aligned "consistent feature vectors" under multimodal conditions. This allows subsequent optimal transmission matching to be measured at the same granularity, guaranteeing fairness and reliability during cross-modal fusion, while preventing information from one modality from dominating the results or being ignored, fundamentally improving the robustness and interpretability of sentiment analysis.
[0090] S4: Using gating weights, candidates and context are weighted and fused to form aspect-level representations.
[0091] The gating weights include weighted fusion of candidates for each aspect term through multi-level weighting rules to obtain the aspect-level representation.
[0092] The specific process is as follows:
[0093] By analyzing the consistency feature vector under each modality k and the distribution of all candidate encodings (dimension-assimilated encoders, which are different for each modality) under modality k of the corresponding aspect word, the first weight of each candidate in the discrete state is output by a pre-trained mapping function through discreteness judgment.
[0094] By encoding the consistent feature vector under each modality (this encoding is different from the above, it is a synthetic encoder of multiple vectors), the total feature vector under each aspect is obtained; the consistent feature vector under each modality is encoded (the dimension transformation encoder is converted to the same dimension as the output of the "synthetic encoder") and then mapped to the total feature vector. The change ratio of the vector length is used as the second weight of each candidate under the current modality.
[0095] For each modality k, the discrete distribution of all candidates is similar. After feature dimensionality reduction through principal component analysis, the similarity between the discrete state of each candidate i and the discrete state of other modalities is analyzed. The goal is to capture the largest range with similarity greater than a threshold. The ratio of the largest range to the discrete range of modality k is used as the third weight of candidate i. Let the two modes being compared be k and g, and obtain the set K under modality k and the set G under modality g. Based on the discrete position of each candidate after dimensionality reduction, the elements between sets K and G are paired to make the correspondence between all paired elements and the discrete relationship the strongest (actually the matching degree or the highest similarity). Using the pairing results, according to the corresponding association relationship, a preset transformation function is used to map the fourth weight of each candidate.
[0096] Among them, for candidates who did not participate in the pairing, the fourth weight is set to 1, indicating that they have strong autonomy or that the relationship is too hidden.
[0097] Meanwhile, the candidates for other modalities are analyzed in the same way as the fourth weight to obtain the fifth weight; after multiplying all weights and candidates by the counterfactual gain and summing them by weight, the aspect-level representations in each dimension are obtained.
[0098] It's worth noting that introducing gating weights in the weighted fusion of candidates and context aims to avoid a disproportionate impact of a single modality or a single piece of evidence on the final sentiment judgment. Through multi-level weighting rules, candidates from different modalities can be dynamically adjusted based on their stability, necessity, and cross-modal consistency, thereby ensuring a more comprehensive and reliable aspect-level representation.
[0099] The first layer of weights measures the distribution of candidates within the current modality, aiming to highlight candidates that are semantically or structurally stable while weakening components with ambiguous boundaries or outliers. The second layer of weights reflects the fit between candidates and the main aspect through mapping to the aspect's total feature vector, intending to strengthen local units that are more representative in the overall representation. The third to fifth layers of weights are geared towards cross-modal consistency, aiming to improve the coordination of candidates across different modalities through structural matching, similarity calculation, and cross-modal association, preventing overfitting or "dogmatism" in any particular modality.
[0100] By combining the aforementioned multidimensional weights with counterfactual gains and introducing a gating mechanism, the final candidates entering the fusion must not only satisfy intramodal and extramodal consistency, but also play a genuine role in sentiment determination in a causal sense. The purpose of this design is to enable aspect-level representations to simultaneously reflect semantic stability, cross-modal consistency, and sentiment causality, thereby providing a reliable, sparse, and interpretable foundational representation for subsequent decoding and ensemble prediction.
[0101] S5: On the aspect-level representation, decode to output aspect polarity and intensity; and generate a controlled set prediction.
[0102] The aspect-level representation is decoded by a decoder, so that the aspect-level representation is mapped onto each of the directional words, including sentiment description and sentiment intensity.
[0103] In this embodiment, the decoder employs a neural network structure based on a bidirectional gated recurrent unit (Bi-GRU) combined with an attention mechanism. The input is a vector sequence of aspect-level representations. First, the sequence is context-modeled using a Bi-GRU to capture the temporal dependencies of the aspect-level representations on semantic, sentiment, and subjective cues. Then, an attention layer assigns weights to features at different time steps or different modalities, resulting in an aggregated sentiment representation. Finally, this sentiment representation is input to a fully connected layer, which outputs the corresponding sentiment description category (e.g., positive, negative, neutral) and sentiment intensity score (e.g., sentiment confidence or polarity magnitude). In other optional embodiments, the decoder can also employ different neural network structures: a convolutional neural network (CNN) decoder, which uses multi-scale convolutional kernels to extract local features from the aspect-level representations, suitable for capturing phrase-level or local modality combinations of sentiment patterns; and a Transformer-based decoder, which directly models the aspect-level representations through a multi-head self-attention mechanism, capable of capturing long-range dependencies across modalities and aspects, and modeling complex subjective sentiment shifts.
[0104] The generation process of the set prediction includes: searching for similar descriptions based on the sentiment descriptions to obtain all sentiment descriptions, and then superimposing their intensities to form a first-level set element.
[0105] Based on the emotional description and emotional intensity, similar descriptive content is retrieved as a strong coupling relationship to obtain secondary set elements.
[0106] The set prediction is obtained by integrating the elements of the first-level set and the elements of the second-level set.
[0107] Firstly, decoding at the aspect-level representation aims to transform abstract multimodal fusion vectors into directly understandable and usable sentiment information. The aspect-level representation itself is merely a high-dimensional feature representation after fusion; without decoding, it can only represent the structure and relationships between data, but cannot output specific subjective sentiment conclusions. Therefore, a decoder is needed to project these high-dimensional vectors onto explicit sentiment label and sentiment intensity spaces, achieving the transformation "from structured representation to deducible results." A decoder employing bidirectional gated recurrent units combined with an attention mechanism ensures that the decoding process captures both contextual dependencies and focuses on the features that most contribute to sentiment polarity and intensity, thereby enhancing the accuracy and interpretability of the output. In other embodiments, introducing CNN or Transformer decoders provides more suitable modeling capabilities for different application scenarios, ensuring the system's flexibility and scalability.
[0108] Secondly, the generation of ensemble predictions aims to improve the robustness and coverage of sentiment output. Individual predictions may be unstable due to noise or modal bias, while ensemble predictions, by integrating similar sentiment descriptions and accumulating and weighting them according to intensity, can cover more possibilities and suppress random errors. The generation of first-level ensemble elements ensures the integrity of expression at the individual description level, while the introduction of second-level ensemble elements ensures consistent sentiment across modalities and contexts through strong coupling. Finally, through integration, a controlled-coverage ensemble prediction is obtained, ensuring that the output maintains diversity while being controlled by constraints, avoiding excessive divergence or ambiguous judgments.
[0109] On the other hand, this embodiment also provides an aspect-level sentiment analysis optimization system based on the fusion of multiple external knowledge sources, which includes:
[0110] The data acquisition unit performs structural analysis on the multimodal evidence from the input text, video, and audio, obtains aspect-level knowledge atoms combined with context, and generates a processed content vector; at the same time, it constructs cross-modal sentiment factors from candidates in different modalities.
[0111] The analysis unit constructs masked or equidistant replacement controls on the same sample and uses the sentiment factor to generate counterfactual gains in sentiment at different levels.
[0112] The processing unit obtains a granularity consistency measure between aspect terms and candidate local units.
[0113] The fusion unit uses gating weights to perform weighted fusion of candidates and context to form aspect-level representations.
[0114] The output unit performs structured decoding on the aspect-level representation, outputting aspect polarity and intensity; and generates a controlled set prediction.
[0115] Example 2, refer to Figure 2 As an embodiment of the present invention, an aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion is provided, comprising: Figure 2 This demonstrates the overall framework of a text sentiment analysis model that integrates multi-source knowledge. Its core idea is to not only rely on word vector representations of sentences but also incorporate three types of information: grammatical structure, sentiment knowledge, and semantic relationships. These are then fused through graph convolutional networks and interactive attention mechanisms to improve the model's classification performance and interpretability.
[0116] At the input layer, sentences are first mapped to a word vector space, where GloVe or BERT can be chosen as the embedding model. Simultaneously, aspect words are extracted from the sentence and combined with conceptual knowledge from an external knowledge base, all transformed into vector representations to provide a foundation for subsequent feature fusion. Syntactic analysis tools then parse the sentence's dependency relations, forming a grammatical dependency tree and constructing an adjacency matrix to capture the structural relationships between words. Furthermore, the model introduces a sentiment knowledge graph to characterize the sentimental associations between words; simultaneously, based on cosine similarity and mutual information (PMI), semantic relevance between words is calculated, constructing a semantic graph, thereby modeling different types of linguistic knowledge at the graph structure level.
[0117] In the feature extraction stage, the model uses graph convolution operations to process the syntax graph and semantic graph respectively, extracting structured high-level feature representations. Syntax graph convolution can highlight dependency paths and syntactic structure features in sentences, while semantic graph convolution strengthens the similarity and co-occurrence relationships of words in the semantic space. At the same time, the combination of sentiment knowledge and semantic features can help the model better identify key nodes closely related to sentiment polarity, thereby enhancing its discriminative ability.
[0118] In the fusion and attention mechanism stage, the model introduces an interactive attention module. Through this interactive attention mechanism, the model can establish connections between syntactic and semantic features, thereby highlighting the parts that contribute most to the classification task. Specifically, syntactic interactive attention strengthens components in dependency structures related to aspect words or sentiment words, while semantic interactive attention captures the most discriminative regions at the semantic level. These multi-channel attention outputs are then integrated to ensure that the model can capture key information from different knowledge perspectives.
[0119] Finally, in the classification stage, the fused features are combined through weighted or concatenated methods and input into the classifier for final prediction. The classifier typically uses a softmax layer to determine the polarity of text sentiment, such as positive, negative, or neutral, and can also be extended to multi-classification tasks.
[0120] Overall, this model overcomes the problem of insufficient utilization of external knowledge by single deep learning models by organically combining syntactic, semantic, and sentiment knowledge. Graph convolutional networks effectively handle non-Euclidean structured information, while interactive attention mechanisms enhance the synergistic effect between different features. This design not only improves the model's accuracy in sentiment analysis tasks but also enhances the interpretability of the results, providing a feasible solution for knowledge fusion methods in natural language processing.
[0121] If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0122] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0123] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0124] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0125] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for optimizing aspect-level sentiment analysis based on the fusion of multiple external knowledge sources, characterized in that, include: The system performs structural analysis on multimodal evidence from input text, video, and audio to obtain aspect-level knowledge atoms that are combined with context, and generates processed content vectors. At the same time, it constructs cross-modal sentiment factors from candidates of different modalities. Construct masked or isometric replacement controls on the same sample, and use the sentiment factor to generate counterfactual gains in sentiment at different levels. A granularity consistency measure is obtained between aspect terms and candidate local units; By using gating weights, the candidate and context are weighted and fused to form an aspect-level representation; On the aspect-level representation, decoding is performed to output aspect polarity and intensity; and a controlled set prediction is generated.
2. The aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion as described in claim 1, characterized in that: The structural analysis includes decomposing the input multimodal evidence into the smallest unit of content vectors using the recognition models corresponding to each multimodal mode; Based on the identification of the contextual relationship of evidence in each modality using the corresponding identification model, the association and strong coupling relationship between each content vector and other content vectors in the same modality are obtained. Each content vector corresponds to a candidate, and the two have the same index; the association relationship includes the direction and intensity of the influence of other content vectors on content vector i. If n content vectors must exist simultaneously in order to make a correct semantic expression, then there is a strong coupling relationship between these n content vectors. By clustering content vectors at the aspect level, a cluster family is obtained for each aspect level j. Each candidate in the family is a knowledge atom selected at aspect level j.
3. The aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion as described in claim 2, characterized in that: The cross-modal sentiment factor includes candidate i and content features under different modalities, which are mapped to the dimensions of the other through a pre-trained neural network, and the recognition model of the other dimension is called to identify the correlation. We obtain the relationship between candidate i and different content vectors in each modality under each modality dimension; and the relationship between the content vector in each modality and the content vector of candidate i under the candidate i dimension. Finally, for every two cross-modal candidates, a correlation is formed in two directions.
4. The aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion as described in claim 3, characterized in that: The counterfactual gain includes determining, within the same sample, a minimum sufficient set consisting of text trigger segments, corresponding image salient regions, and speech keyframes based on the candidate and its strong coupling relationship and cross-modal association; Semantic-preserving masking interventions are implemented on the minimum sufficient set, wherein text fragments are replaced with neutral placeholders while keeping part-of-speech and dependency roles unchanged, salient regions of images are locally weakened or prototype-repaired while keeping categories unchanged, and speech keyframes are replaced with energy or frequency band equivalents while keeping prosodic rhythms consistent. An isometric substitution intervention is performed on the minimum sufficient set, wherein text trigger words are replaced by neutral neighbor words within the same part of speech and synonym range, image regions are replaced by similar regions within the same category prototype, and speech segments are replaced by adjacent segments within the same speaker and context. Through rule-based screening and small-sample exploration, the most stable replacement option is selected under the conditions of semantic consistency, opinion holder consistency, and scope consistency, ensuring that the control sample does not change the non-target aspects of the judgment. By comparing the outcome assessments before and after the intervention, the emotional differences of the target candidate are obtained. The direction of the difference is corrected and truncated by combining the intensity of enhancement and suppression, thereby generating the counterfactual gain value and credibility of the candidate. The granularity consistency measure includes distinguishing data of different modalities within the aspect-level selected knowledge atom corresponding to each aspect term; The content vector for each modality is obtained, and then the content vector for each modality is input into the pre-trained encoder to obtain the consistency feature vector for each modality.
5. The aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion as described in claim 4, characterized in that: The gating weights include weighted fusion of candidates on each aspect term through multi-level weighting rules to obtain the aspect-level representation; The specific process is as follows: By analyzing the consistency feature vector under each modality k and the distribution of all candidate encoded words under the corresponding aspect word in modality k, the first weight of each candidate in the discrete state is output by a pre-trained mapping function through discreteness judgment. By encoding the consistent feature vector under each modality, the total feature vector under each aspect is obtained; the consistent feature vector under each modality is then mapped onto the total feature vector after dimensionality transformation, and the change ratio of the vector length is used as the second weight of each candidate under the current modality. For each modality k, the discrete distribution of all candidates is judged for similarity: after feature dimensionality reduction by principal component analysis, the similarity between the discrete state of each candidate i and the discrete state of other modalities is analyzed. The target is to capture the largest range with similarity greater than the threshold. The ratio of the largest range to the discrete range of modality k is obtained as the third weight of candidate i. Let the two contrasting modes be k and g, and obtain the set K under mode k and the set G under mode g; and according to the discrete position of each candidate after dimensionality reduction, pair the elements between sets K and G so that the correspondence between all paired elements and discrete relations is the strongest. Using the pairing results, and based on the corresponding association, a fourth weight is mapped to each candidate using a preset transformation function; For candidates that did not participate in the pairing, the fourth weight was set to 1. Simultaneously, the candidates for other modalities are analyzed in the same manner as the fourth weight to obtain the fifth weight; By multiplying all weights and candidates by counterfactual gains and then summing the weights, we obtain the aspect-level representations in each dimension.
6. The aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion as described in claim 5, characterized in that: The aspect-level representation is decoded by a decoder, so that the aspect-level representation is mapped onto each of the directional words, including sentiment description and sentiment intensity.
7. The aspect-level sentiment analysis optimization method based on multi-source external knowledge fusion as described in claim 6, characterized in that: The generation process of the set prediction includes: searching for similar descriptions based on the sentiment descriptions to obtain all sentiment descriptions, and then superimposing their intensities as first-level set elements; Based on the emotional description and emotional intensity, similar descriptive content is retrieved as a strong coupling relationship to obtain secondary set elements; The set prediction is obtained by integrating the elements of the first-level set and the elements of the second-level set.
8. An aspect-level sentiment analysis optimization system based on multivariate external knowledge fusion, employing the method described in any one of claims 1-7, characterized in that: The data acquisition unit performs structural analysis on the multimodal evidence from the input text, video, and audio, obtains aspect-level knowledge atoms combined with context, and generates a processed content vector; at the same time, it constructs cross-modal sentiment factors from candidates of different modalities. The analysis unit constructs masked or equidistant replacement controls on the same sample and uses the sentiment factor to generate counterfactual gains in sentiment at different levels. The processing unit obtains a granularity consistency measure between aspect terms and candidate local units; The fusion unit uses gating weights to perform weighted fusion of candidates and context to form aspect-level representations; The output unit performs structured decoding on the aspect-level representation, outputting aspect polarity and intensity; and generates a controlled set prediction.
9. A computer device, comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of an aspect-level sentiment analysis optimization method based on the fusion of multiple external knowledge.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of an aspect-level sentiment analysis optimization method based on the fusion of multiple external knowledge.