A multi-modal bias detection method and system based on style features
Patent Information
- Application Number
- CN202610752521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]有鉴于此,本发明要解决的问题是提供一种基于风格特征的多模态偏见检测方法及系统,以解决现有方法在特征建模层面缺乏对风格隐性偏见信号的系统刻画、在解释机制层面解释与模型决策之间缺乏结构一致性约束、以及在评价体系层面解释质量缺乏统一标准的问题
[0015]本发明具有的优点和积极效果是:
Smart Images

Figure CN122594780A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural language processing and multimodal information analysis technology, specifically relating to a multimodal bias detection method and system based on style features. Background Technology
[0002] As a crucial intermediary in information dissemination, news media play a key role in shaping public perception and public opinion structures through their reporting methods and content selection. However, influenced by multiple factors such as institutional stance, ideological bias, and commercial interests, some media outlets often exhibit a degree of systemic bias in their news reporting, particularly in topic selection, expression, and presentation strategies. In the digital and social media environment, the barriers to information production and dissemination have significantly decreased, and content dissemination paths have become more diverse. For example, the forms of expression of political bias have become more complex and implicit, profoundly impacting public cognitive structures and the formation of social consensus.
[0003] However, existing news bias detection technologies have significant shortcomings in feature modeling. Current multimodal methods primarily focus on explicit content features, such as keywords and named entities in text, or object and scene recognition in images, lacking systematic modeling of implicit elements such as textual rhetorical structure, visual composition strategies, and cross-modal expressive styles. In fact, stylistic elements such as the syntactic complexity, emotional distribution, and argumentative structure of text, as well as the color distribution, composition, and visual complexity of images, often influence audience emotions and cognitive judgments without directly expressing a stance. Compared to content features dependent on specific events, stylistic features are more likely to reflect the media's long-term stable expressive habits and value orientations, exhibiting consistency across cross-event contexts. However, stylistic factors are usually introduced only as auxiliary variables or local features, and a unified multimodal style representation mechanism has not yet been formed, making it difficult to reveal the stable role of style in bias construction.
[0004] At the level of explanation mechanisms, existing deep learning methods, while improving detection accuracy, exacerbate the black-box nature of the model's decision-making process. Traditional interpretable AI methods primarily rely on feature weight visualization, which, while revealing the model's focus areas to some extent, often struggles to distinguish between statistically relevant clues and the key factors truly influencing decisions. They also fail to provide causal narratives consistent with cognitive logic and to clearly define the causal roles of different modalities in decision-making. While generative explanation methods offer expressive advantages, the lack of clear constraints between the explained content and the model's internal decisions means that explanations often remain at the level of post-hoc supplementary explanations, with uncertainties regarding fidelity and consistency. Furthermore, the lack of tight coupling with the model's internal representation space makes it difficult to establish stable causal constraints.
[0005] At the evaluation system level, the lack of a unified standard for explanatory quality and the significant differences in evaluation indicators used by different studies limit the comparability and verifiability of the explanatory results. At the same time, models may learn from pre-existing biased data and amplify potential social biases, raising concerns about algorithmic fairness. Summary of the Invention
[0006] In view of this, the problem to be solved by the present invention is to provide a multimodal bias detection method and system based on style features, so as to solve the problems of existing methods lacking a systematic characterization of style implicit bias signals at the feature modeling level, lacking structural consistency constraints between interpretation and model decision at the interpretation mechanism level, and lacking a unified standard for interpretation quality at the evaluation system level.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A multimodal bias detection method based on style features includes the following steps; S1: Construct a multimodal dataset containing news text, accompanying images, and political bias labels, and extract and statistically analyze the multi-level style features in the dataset; S2: Extract deep semantic information from the news text and deep semantic information from the accompanying images, and integrate them to obtain a multimodal semantic feature representation; S3: Construct text style features, image style features, and cross-modal consistency features to form a multi-level style feature representation; S4: Using multimodal semantic features as the backbone representation and multi-level style features as the moderating signal, the style guides the fusion of semantics through a structured fusion mechanism, and outputs the prediction results of political bias category; S5: Based on a multimodal large language model, using news text, accompanying images, model prediction results, and preset explanation prompt templates as input, generate natural language explanation text, and encode the explanation text into an explanation representation; S6: Using the interpretation as a mediating variable, variational causal inference is used to construct the causal relationship between multimodal latent variables, interpretations, and prediction results, and consistency constraint loss is introduced during the training phase.
[0008] In S1, multi-level gradient political bias category labels are predefined, and ordered numerical codes are completed according to the category arrangement order. Samples are classified and grouped based on the numerical codes. Non-parametric test methods are used to analyze the distribution differences of style features under different groups. A preset significance judgment threshold is used as the evaluation criterion. If a corresponding distribution difference is determined, the subsequent feature construction process is entered. If no distribution difference is determined, the style feature extraction operation is carried out again.
[0009] In S2, the deep semantic information of the text is extracted to obtain the text semantic representation vector through a lightweight pre-trained language model; the news text is split into sub-words and word embeddings are generated, and after fusion with position encoding, bidirectional context modeling is completed through a multi-layer encoding structure to obtain the deep semantic representation of the text. The image deep semantic information is extracted to obtain the image semantic representation vector through a multimodal pre-trained visual model; visual feature extraction processing is performed on the news images, and visual semantic modeling is completed through a multi-layer visual coding structure to obtain the deep semantic representation of the image.
[0010] In S3, the text style features include text sentiment features, text writing style features, and text theme features. The text sentiment features are obtained by using sentiment analysis to mine the text sentiment distribution information. The text writing style features are obtained by using syntax, vocabulary, and writing structure indicators to characterize the text expression habits. The text theme features are obtained by generating a text theme distribution representation through theme mining.
[0011] In S3, the cross-modal consistency feature obtains the similarity between text style features and image style features through a measurement calculation method, which is used to characterize the collaborative relationship between text and image in the emotion dimension or the theme dimension.
[0012] In S4, multi-level style feature representations are retrieved as style adjustment information; relying on a learnable fusion weight mechanism, this style information is transformed into adaptive constraint weights, and the dynamic regulation of style representation on semantic representation is achieved through an attention structure, thus completing the style-guided fusion of semantics.
[0013] In S5, the preset explanation prompt template includes basic explanation prompts and style-enhanced explanation prompts; the basic explanation prompts rely solely on the original news content to complete the bias attribution inference explanation, while the style-enhanced explanation prompts introduce a structured style feature set to assist in explanation generation based on the original input conditions.
[0014] A detection system for a multimodal bias detection method based on style features includes a data construction module, a semantic feature extraction module, a style representation construction module, a fusion prediction module, an interpretation generation module, and a causal constraint module connected by data streams. The data construction module is used to construct a multimodal dataset containing news text, images, and political bias labels, and to extract and statistically analyze the multi-level style features in the dataset. The semantic feature extraction module is used to perform deep semantic encoding on text and image respectively, to obtain text semantic representation vector and image semantic representation vector, and to obtain multimodal semantic feature representation; The style representation construction module is used to construct text style features, image style features, and cross-modal consistency features to form a multi-level style feature representation. The fusion prediction module is used to use multimodal semantic feature representation as the backbone representation and multi-level style feature representation as the adjustment signal. It achieves style-guided fusion of semantics through a structured fusion mechanism and outputs the prediction result of political bias category. The explanation generation module is used to generate natural language explanation text and encode it into an explanation representation based on a multimodal large language model, taking news text, accompanying images, model prediction results and preset explanation prompt templates as input; The causal constraint module is used to use the interpretation representation as a mediating variable to deduce the intrinsic causal relationship between multimodal latent variables, interpretation representation and prediction results, and introduces consistency constraint loss during the training phase.
[0015] The advantages and positive effects of this invention are: (1) By constructing a multi-level style feature system that includes text style, image style and cross-modal consistency, this invention breaks through the limitation of traditional detection methods that rely solely on explicit content features. It can effectively mine implicit and stable bias expression signals in news reports and significantly improve the accuracy and reliability of bias detection.
[0016] (2) The present invention adopts a dynamic fusion mechanism of style features to guide semantic features, and uses style information as an adaptive adjustment signal to optimize multimodal semantic representation, avoiding redundancy and interference caused by simple feature splicing, making the model more stable in ordered classification tasks and significantly reducing detection error.
[0017] (3) This invention combines generative interpretation with causal constraint mechanism, realizes understandable natural language interpretation output through dual template prompts, and establishes strong correlation between interpretation and model decision by using variational causal reasoning and consistency constraint loss, effectively solving the black box problem of deep learning model and greatly improving the interpretability and credibility of detection results.
[0018] (4) This invention constructs a high-quality multimodal dataset containing text, images and seven-class gradient labels, and establishes a multi-dimensional evaluation system that includes classification accuracy, prediction error and interpretation quality, so as to realize the synchronous quantitative evaluation of detection performance and interpretation effect, and improve the comprehensiveness, comparability and verifiability of model evaluation. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0020] In the attached diagram: Figure 1 This is an overall flowchart of a multimodal bias detection method based on style features according to the present invention.
[0021] Figure 2 This is a flowchart of the multi-level style feature construction and cross-modal consistency feature extraction in the method of this invention.
[0022] Figure 3 This is a diagram illustrating the consistency constraint modeling framework for interpreting generation and variational causal reasoning in the method of this invention. Figure 4 This is a graph showing the performance comparison between the method of this invention and existing baseline models on a self-constructed multimodal news bias dataset.
[0023] Figure 5 is a schematic diagram of the experimental results of multi-level style feature ablation in the method of the present invention.
[0024] Figure 6 This is an example diagram comparing the interpretation generation results under basic interpretation prompts and style-enhanced interpretation prompts in the method of this invention. Detailed Implementation
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0027] To enable those skilled in the art to more clearly understand the technical solution of the present invention, the specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0028] The overall technical framework of this invention is as follows: Figure 1 As shown, this demonstrates the complete technical roadmap from multimodal news data input to bias prediction and interpretation output, comprising four main parts: data input, feature extraction, feature fusion, and bias prediction. In the data input stage, the system receives multimodal data containing news text, accompanying images, and political bias labels. In the feature extraction stage, the system extracts textual semantic features, image semantic features, and multi-level style features, respectively. In the feature fusion stage, multimodal feature fusion is achieved through a unified dimensionality mapping and style-guided mechanism. In the bias prediction stage, the classifier module outputs a seven-class ordered prediction result.
[0029] In the implementation, the first step was to construct a multimodal news political bias dataset and analyze its style features. Political bias labels for news media were obtained from a third-party media bias assessment agency. This labeling system adopted a seven-category structure, specifically including extreme left-leaning, left-leaning, left-leaning, neutral, right-leaning, right-leaning, and extreme right-leaning, each corresponding to an ordered numerical code from 1 to 7. During the data collection phase, the headlines, body text, and corresponding images of each news media report were collected simultaneously. Data cleaning and quality control were then performed, removing samples without images, samples with excessively short text, and samples with damaged images. After completing the dataset construction, non-parametric tests were used to analyze the inter-group distribution differences of the seven political bias categories, verifying whether there were statistically significant distribution differences across style dimensions for different political bias categories, providing a statistical basis for subsequent style feature modeling.
[0030] In the multimodal deep semantic feature extraction stage, the system performs deep semantic encoding on both text and images. For text semantic feature modeling, a lightweight pre-trained language model is used as the text encoder, with DistilBERT being the preferred choice. This model is obtained by compressing the BERT model through knowledge distillation, significantly reducing the model parameter size and training cost while maintaining high classification performance. Given a news text... First, word-level segmentation is performed using the WordPiece word segmenter to obtain word embedding representations. These are then combined with positional encoding and input into a multi-layer Transformer encoder for bidirectional contextual modeling, ultimately yielding a deep semantic representation vector of the text. This is used for subsequent feature fusion operations. For image semantic feature modeling, a multimodal pre-trained model is adopted, with the CLIP model being preferred. This model is pre-trained on a large number of image-text pairs through a contrastive learning mechanism, enabling it to obtain visual semantic representations with strong cross-modal transfer capabilities. Given a news image... Deep visual semantic representation vectors of images are extracted using the CLIP model's visual encoder. .
[0031] In the multi-level style feature modeling stage, the system constructs multi-level style feature representations based on semantic features to characterize the expression style of the news rather than the content itself. The specific process of this stage is as follows: Figure 2 As shown, the process of extracting text style features, the process of extracting image style features, and the process of calculating cross-modal consistency features are illustrated.
[0032] Text style feature modeling involves three extraction operations. The first is sentiment feature extraction: using sentiment analysis methods to extract the sentiment distribution of the news text, specifically including the proportion and intensity of positive, neutral, and negative sentiment. This feature is typically represented as a three-dimensional vector. The second is writing style feature extraction: by analyzing syntactic complexity, lexical diversity, and argumentation structure indicators, the rhetorical strategies and expression habits of the news text are characterized. Syntactic complexity is measured by indicators such as average sentence length and nesting depth; lexical diversity is quantified by the ratio of classifiers to formifiers; and argumentation structure is measured by the frequency of argumentation markers. The third is topic feature extraction: using a topic model to extract the topic distribution vector of the text, representing the probability distribution of the text across different topics.
[0033] Image style feature modeling involves four extraction operations. The first is image composition feature extraction: analyzing the image's composition layout, subject position, and visual balance. Composition layout is measured using metrics such as the golden ratio and symmetry; subject position is measured by the subject's coordinates within the image; and visual balance is measured by the evenness of visual weight distribution. The second is image color feature extraction: extracting statistical information about the image's color distribution, including hue distribution, saturation distribution, and brightness distribution. This feature is typically represented as a color histogram or color moments. The third is image emotion feature extraction: identifying the emotional tendency conveyed by the image through a visual emotion analysis model, including the probability distribution of positive, negative, and neutral emotions. The fourth is image theme feature extraction: obtaining the image's theme distribution through visual theme analysis methods, calculated based on the categories of objects and scenes appearing in the image.
[0034] Cross-modal consistency feature modeling is used to reflect the synergistic relationship between text and image at the emotion or topic level. Specifically, firstly, emotion vectors are extracted from text style features and image style features, respectively. Then, the cosine similarity between these two emotion vectors is calculated to obtain the cross-modal consistency score at the emotion level. Next, text topic feature vectors and image topic feature vectors are extracted, and the cosine similarity between these two topic vectors is calculated to obtain the cross-modal consistency score at the topic level. Finally, the cross-modal consistency scores at the emotion and topic levels are weighted and fused to obtain the cross-modal consistency feature. When the text emotion and image emotion are consistent, the cross-modal consistency is high, indicating that the text and image reinforce each other in terms of expression; conversely, it indicates that there are differences or contradictions in the expression of the text and image.
[0035] In the style-guided multimodal feature fusion and bias prediction stage, the system uses semantic features as the backbone representation and style features as moderating signals, achieving style-guided semantic fusion through a structured fusion mechanism. First, the text semantic representation... Image semantic representation Text style features Image style features and cross-modal consistency features A unified dimensional mapping is performed, enabling the fusion of features from different sources within a unified space. Then, a learnable fusion weight mechanism transforms style information into learnable constraints or weights. This mechanism employs an attention mechanism to dynamically adjust style to semantics, rather than using static concatenation to avoid redundant layering. In practice, the attention mechanism dynamically calculates the attention weights between style and semantic features based on the characteristics of the input samples, allowing the model to dynamically select a discrimination path that is more dependent on semantic content or more dependent on expressive style in different samples. The fused multimodal representation is input into a classifier module, which uses a multilayer perceptron and outputs a seven-class ordered prediction result, covering the complete political spectrum from extreme left to extreme right.
[0036] In the natural language interpretation and generation stage, the system generates natural language interpretation text based on a multimodal large language model. The framework for this stage is as follows: Figure 3 As shown, the process involves parallel implementation of a natural language interpretation generation path and a causal consistency constraint path. The interpretation generation process for each news sample uses three types of information as input conditions: the first type is the news text and accompanying images, including the news title, body text, and corresponding images; the second type is the model prediction result, i.e., the political stance category output by the bias prediction model; and the third type is the interpretation generation prompt template, i.e., a preset interpretation prompt template used to guide the model to make explanatory statements based on the criteria for judging political bias.
[0037] In the design of the prompt template, an experimental design approach with controlled variables was adopted to construct two types of explanation generation conditions. The basic explanation prompt relies solely on the news content itself for bias attribution explanation. Its prompt template structure is: Based on the following news content, please determine the political bias tendency of this news and explain the basis for your judgment. News Title: {Title}, News Body: {Body}, News Image: {Image}, Prediction Result: {Prediction Category}. The style-enhanced explanation prompt, based on the basic explanation prompt, additionally introduces a set of structured style features, including text sentiment distribution indicators, image visual statistical features, and cross-modal consistency indicators. Its prompt template structure is: Based on the following news content and its style features, please determine the political bias tendency of this news and explain the basis for your judgment. News Title: {Title}, News Body: {Body}, News Image: {Image}, Text Sentiment Distribution: {Text Sentiment Features}, Image Visual Features: {Image Style Features}, Cross-modal Consistency: {Cross-modal Consistency Features}, Prediction Result: {Prediction Category}.
[0038] The generation process is modeled as a conditional probability problem, assuming the multimodal input of the news sample is... The explanatory text consists of a sequence of words. The generation process can then be expressed as a product of conditional probabilities:
[0039] in To predict the stance category, Indicates the first In practice, the explanatory text content generated at each position is generated word by word using an autoregressive generation method until an end marker is generated or the maximum length limit is reached.
[0040] In the modeling stage based on variational causal inference and the constraint of explanatory consistency, the system uses the explanatory representation as a causal mediator variable and models the causal dependency between multimodal latent variable representations, explanatory representations, and prediction results through a variational causal inference mechanism. First, the generated explanatory text is mapped to an explanatory embedding representation E through a BERT encoder. Then, a causal link is constructed: multimodal latent variable representation H → explanatory embedding representation E → prediction result Y. The explanatory representation E, as a mediator variable, needs to satisfy two conditions: first, given H, E should be able to effectively predict Y; second, E should maintain semantic consistency with H.
[0041] In variational modeling, variational distributions are introduced. Approximate posterior distribution And design a joint optimization objective function:
[0042] in For the classification loss of the bias prediction task, For the reconstruction loss of the variational autoencoder, Loss due to causal consistency constraints and To balance the weights, the causal consistency constraint loss is used to ensure structural consistency between the explanatory representation, the latent variable representation, and the prediction results.
[0043] The objective function consists of three parts: the first part is the classification loss for the bias prediction task, calculated using the cross-entropy loss function; the second part is the reconstruction loss of the variational autoencoder, used to ensure that the interpretive representation can reconstruct the original interpretive text; and the third part is the causal consistency constraint loss, used to ensure structural consistency between the interpretive representation, the latent variable representation, and the prediction result. In practice, the causal consistency constraint loss is calculated as follows: first, latent variables are sampled from the multimodal latent variable representation H; then, the conditional probability distribution between the interpretive representation E and the prediction result Y under the given latent variables is calculated; finally, the divergence is used to measure the difference between this distribution and the true distribution.
[0044] During the training phase, a causal consistency loss function is introduced, enabling the model to optimize prediction performance while constraining the latent variable representations to effectively explain the prediction results through explanatory variables. Through this mechanism, the explanation no longer merely answers why the model makes this judgment, but becomes a structural variable that constrains what information the model should base its judgment on, thus establishing a structural closed loop between the prediction path and the explanation path.
[0045] In the explanation effectiveness evaluation and system output phase, the system output includes the following: the predicted political bias category of the news sample, the corresponding natural language explanation text, and explanation quality evaluation metrics. The explanation quality evaluation metrics include three items: label alignment, content relevance, and cross-modal consistency. Label alignment measures the directional consistency between the generated explanation and the model's predicted labels, obtained by calculating the matching degree between sentiment words in the explanation text and the predicted category; content relevance assesses the semantic association between the explanation text and the original news content, obtained by calculating the cosine similarity between the explanation text embedding and the news content embedding; cross-modal consistency determines whether the explanation simultaneously addresses information from both textual and image modalities, obtained by statistically analyzing the proportion of explanation texts that mention both textual and image cues.
[0046] In its implementation, the system employs a multi-indicator joint evaluation strategy to assess bias detection performance. Accuracy is used to evaluate the proportion of correctly classified samples among the total samples; the macro-average F1 score is used to evaluate classification performance and class balance; mean absolute error (MAE) measures the degree of error in the ordered political spectrum prediction task, obtained by calculating the average of the absolute differences between predicted and true values; and root mean square error (RMSE) measures the overall magnitude of the prediction error, obtained by taking the square root of the mean of the squares of the errors between predicted and true values. The combined use of multiple indicators helps to comprehensively evaluate the effectiveness and robustness of the multimodal political bias detection model from different perspectives.
[0047] In specific application scenarios, the method of this invention can be applied to automated political bias detection and analysis systems for news media. For example, in a public opinion monitoring system, the system can collect news reports from major media outlets in real time, automatically identify their political bias tendencies, and generate corresponding explanatory texts to help users understand the potential stance behind the news reports. In specific implementation, the system first collects news data from target media through a web crawler, including news headlines, body text, and accompanying images. Then, the collected data is input into a trained multimodal bias detection model to obtain political bias prediction results and corresponding natural language explanations. Finally, the results are displayed to the user in the form of a visual interface.
[0048] This invention constructs a style-guided multimodal fusion mechanism by introducing text style features, image style features, and cross-modal consistency features, effectively alleviating the problem that simple semantic modeling is insufficient to characterize implicit political tendencies. Simultaneously, through conditional interpretation generation modeling and structured prompt template design, the model can provide textual explanations for prediction results, and the internal representation is adjusted through consistency constraint loss, transforming the explanation from a result description into a modelable structural variable. Experimental results show that the method reduces the mean absolute error by approximately 6.3% and the root mean square error by approximately 5.1%, achieving stable improvements in accuracy and F1 score, verifying the effectiveness of the style modeling mechanism.
[0049] Furthermore, this application also provides a detection system for a multimodal bias detection method based on style features, including a data construction module, a semantic feature extraction module, a style representation construction module, a fusion prediction module, an interpretation generation module, and a causal constraint module connected by data streams; The data construction module is used to construct a multimodal dataset containing news text, images, and political bias labels, and to extract and statistically analyze the multi-level style features in the dataset. The semantic feature extraction module is used to perform deep semantic encoding on text and image respectively, to obtain text semantic representation vector and image semantic representation vector, and to obtain multimodal semantic feature representation; The style representation construction module is used to construct text style features, image style features, and cross-modal consistency features to form a multi-level style feature representation. The fusion prediction module is used to use multimodal semantic feature representation as the backbone representation and multi-level style feature representation as the adjustment signal. It achieves style-guided fusion of semantics through a structured fusion mechanism and outputs the prediction result of political bias category. The explanation generation module is used to generate natural language explanation text and encode it into an explanation representation based on a multimodal large language model, taking news text, accompanying images, model prediction results and preset explanation prompt templates as input; The causal constraint module is used to use the interpretation representation as a mediating variable to deduce the intrinsic causal relationship between multimodal latent variables, interpretation representation and prediction results, and introduces consistency constraint loss during the training phase.
[0050] The working principle and process of this invention are as follows: We constructed a multimodal dataset containing news text, images, and political bias labels, and statistically verified the distribution differences of multi-level style features under different political bias categories, providing data support and theoretical basis for style feature modeling. Subsequently, deep semantic encoding was performed on the news text and accompanying images respectively to extract text semantic representation vectors and image semantic representation vectors, which were then integrated to form a unified multimodal semantic feature representation. Based on this, text style features, image style features, and cross-modal consistency features are constructed respectively to form a multi-level style feature representation that can comprehensively depict expression habits and collaborative relationships. Next, multimodal semantic feature representation is used as the backbone representation, and multi-level style feature representation is used as the moderating signal. Through learnable fusion weights and attention structures, style is dynamically guided to fuse semantics, achieving efficient aggregation of multimodal features and outputting the prediction result of political bias category. At the same time, based on the multimodal large language model, news text, accompanying images, model prediction results and preset explanation prompt templates are used as inputs to generate corresponding natural language explanation texts according to two modes: basic explanation and style-enhanced explanation. The explanation texts are then encoded into explanation representations. Finally, the explanatory representation is used as a mediating variable. Variational causal reasoning is used to construct the causal relationship between multimodal latent variables, explanatory representation, and prediction results. Consistency constraint loss is introduced during the training process to form a structural closed loop between explanation generation and model decision-making. This improves the accuracy of bias detection while ensuring the fidelity and consistency of the explanation results.
[0051] The embodiments of the present invention have been described in detail above, but the content described is only a preferred embodiment of the present invention and should not be considered as limiting the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of this patent.
Claims
1. A multimodal bias detection method based on style features, characterized in that, Includes the following steps; S1: Construct a multimodal dataset containing news text, accompanying images, and political bias labels, and extract and statistically analyze the multi-level style features in the dataset; S2: Extract deep semantic information from the news text and deep semantic information from the accompanying images, and integrate them to obtain a multimodal semantic feature representation; S3: Construct text style features, image style features, and cross-modal consistency features to form a multi-level style feature representation; S4: Using multimodal semantic features as the backbone representation and multi-level style features as the moderating signal, the style guides the fusion of semantics through a structured fusion mechanism, and outputs the prediction results of political bias category; S5: Based on a multimodal large language model, using news text, accompanying images, model prediction results, and preset explanation prompt templates as input, generate natural language explanation text, and encode the explanation text into an explanation representation; S6: Using the interpretation as a mediating variable, variational causal inference is used to construct the causal relationship between multimodal latent variables, interpretations, and prediction results, and consistency constraint loss is introduced during the training phase.
2. The multimodal bias detection method based on style features according to claim 1, characterized in that, In S1, multi-level gradient political bias category labels are predefined, and ordered numerical codes are completed according to the category arrangement order. Samples are classified and grouped based on the numerical codes. Non-parametric test methods are used to analyze the distribution differences of style features under different groups. A preset significance judgment threshold is used as the evaluation criterion. If a corresponding distribution difference is determined, the subsequent feature construction process is entered. If no distribution difference is determined, the style feature extraction operation is carried out again.
3. The multimodal bias detection method based on style features as described in claim 1, characterized in that, In S2, the deep semantic information of the text is extracted to obtain the text semantic representation vector through a lightweight pre-trained language model; the news text is split into sub-words and word embeddings are generated, and after fusion with position encoding, bidirectional context modeling is completed through a multi-layer encoding structure to obtain the deep semantic representation of the text. The image depth semantic information is extracted into an image semantic representation vector through a multimodal pre-trained visual model; Visual features are extracted from news images, and visual semantic modeling is performed through a multi-layer visual coding structure to obtain deep semantic representations of the images.
4. The multimodal bias detection method based on style features as described in claim 1, characterized in that, In S3, the text style features include text sentiment features, text writing style features, and text theme features. The text sentiment features are obtained by using sentiment analysis to mine the text sentiment distribution information. The text writing style features are obtained by using syntax, vocabulary, and writing structure indicators to characterize the text expression habits. The text theme features are obtained by generating a text theme distribution representation through theme mining.
5. The multimodal bias detection method based on style features as described in claim 1, characterized in that, In S3, the cross-modal consistency feature obtains the similarity between text style features and image style features through a measurement calculation method, which is used to characterize the collaborative relationship between text and image in the emotion dimension or the theme dimension.
6. The multimodal bias detection method based on style features as described in claim 1, characterized in that, In S4, multi-level style feature representations are retrieved as style adjustment information; relying on a learnable fusion weight mechanism, this style information is transformed into adaptive constraint weights, and the dynamic regulation of style representation on semantic representation is achieved through an attention structure, thus completing the style-guided fusion of semantics.
7. The multimodal bias detection method based on style features as described in claim 1, characterized in that, In S5, the preset explanation prompt template includes basic explanation prompts and style-enhanced explanation prompts; The basic explanation prompt relies solely on the original news content to complete the bias attribution inference, while the style-enhanced explanation prompt introduces a structured style feature set to assist in the generation of explanations based on the original input conditions.
8. A detection system for a multimodal bias detection method based on style features, characterized in that, It includes a data construction module, a semantic feature extraction module, a style representation construction module, a fusion prediction module, an interpretation generation module, and a causal constraint module, all connected via data streams. The data construction module is used to construct a multimodal dataset containing news text, images, and political bias labels, and to extract and statistically analyze the multi-level style features in the dataset. The semantic feature extraction module is used to perform deep semantic encoding on text and image respectively, to obtain text semantic representation vector and image semantic representation vector, and to obtain multimodal semantic feature representation; The style representation construction module is used to construct text style features, image style features, and cross-modal consistency features to form a multi-level style feature representation. The fusion prediction module is used to use multimodal semantic feature representation as the backbone representation and multi-level style feature representation as the adjustment signal. It achieves style-guided fusion of semantics through a structured fusion mechanism and outputs the prediction result of political bias category. The explanation generation module is used to generate natural language explanation text and encode it into an explanation representation based on a multimodal large language model, taking news text, accompanying images, model prediction results and preset explanation prompt templates as input; The causal constraint module is used to use the interpretation representation as a mediating variable to deduce the intrinsic causal relationship between multimodal latent variables, interpretation representation and prediction results, and introduces consistency constraint loss during the training phase.