Dual-mode forged information detection method fusing emotion features
The emotional characteristics of text and images are extracted through large language models, combined with domain features, and used the AdaIN network to adjust the modal distribution, solving the problem of unfusion of emotional information in traditional methods and improving the accuracy of fake news detection.
Patent Information
- Application Number
- CN202510515988.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional multimodal detection methods fail to effectively integrate emotional information, resulting in insufficient accuracy of fake news detection and failure to optimize feature representations based on different fields.
The emotional characteristics of text and images are extracted using a large language model, the domain characteristics are processed through a mixed expert network, and the modal distribution is adjusted using the AdaIN network, and the final input classifier is input to classify.
The accuracy of information detection is improved, and by integrating emotional characteristics and domain information, the ability to identify fake news is enhanced.
Smart Images

Figure CN120387138A_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a dual-modal forged information detection method integrating emotional features. [Background Technology]
[0002] Traditional multimodal detection methods typically rely on basic fusion techniques, fusing the image and text modalities of news to analyze modal consistency, but different fields place varying emphasis on features. Furthermore, this approach is limited by its failure to consider sentiment and to optimize steps based on different fields through any additional feature representation, which can easily lead to the loss of information crucial for classification. On the other hand, when it comes to identifying sentiment in news content, unlike relatively objective real news, fake news often uses more inflammatory language to garner more attention and views. However, most methods only use sentiment as an independent feature for fake news detection, without integrating sentiment features with other modalities (especially images). [Summary of the invention]
[0003] The present invention overcomes the shortcomings of the prior art and provides a dual-modal forged information detection method integrating emotional features.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] A dual-modal forged information detection method integrating emotional features, characterized by:
[0006] S1. Emotional feature extraction: Use a large language model to extract the emotional categories of text and images, and give the emotional intensity of each category and the emotional score; analyze the emotional consistency of mixed modal features of text and image;
[0007] S2. The hybrid expert network processes features based on domain classes and calculates the domain probabilities of the text modality, image modality, and hybrid modality through the text domain encoder, image domain encoder, and image-text hybrid feature domain encoder. The hybrid expert network learns the feature information of the text modality, image modality, and hybrid modality based on the domain information.
[0008] S3. The AdaIN network adjusts the modal distribution. The AdaIN network adjusts the feature distribution according to the contribution of each modality, fuses the adjusted features, and inputs them into the classifier to obtain the final classification result.
[0009] The above-mentioned dual-modal forged information detection method integrating emotional features is characterized in that: S1 includes
[0010] S11. Extract the emotional features of text and images, set the text prompt learning template and the image prompt learning template, and the large language model extracts the emotional category, emotional intensity, and emotional score of text and images respectively according to the text prompt learning template and the image prompt learning template;
[0011] S12. Analyze the emotions of the text and image hybrid modality features, set the hybrid modality prompt learning template, and the large language model analyzes the emotional consistency of the text and image hybrid modality features according to the hybrid modality prompt learning template.
[0012] A dual-modal forged information detection method integrating emotional features as described above, characterized in that: the large language model is the GPT 4V large language model.
[0013] A dual-modal forged information detection method integrating emotional features as described above, characterized in that: S2 includes
[0014] S21. Train the domain encoder, and train the domain encoders of the text modality, image modality, and image-text hybrid modality respectively through the training set data with domain labels;
[0015] S22. Calculate the domain probability, use the trained domain encoder to calculate the domain probability of each modality respectively, and obtain the domain probability of the text modality as The domain probability of the image modality is The domain probability of the hybrid modality is
[0016] S23. The hybrid expert network learns feature information, and the output of the hybrid expert network structure of the text modality is denoted as The output of the hybrid expert network structure of the image modality is denoted as The output of the hybrid expert network structure of the hybrid modality is denoted as:
[0017] A dual-modal forged information detection method integrating emotional features as described above, characterized in that: the training set data is the weibo training set data.
[0018] A dual-modal forged information detection method integrating emotional features as described above, characterized in that: 64 samples are set for each batch training, and the iteration is 50 times.
[0019] A dual-modal forged information detection method integrating emotional features as described above, characterized in that: the domain categories included in the weibo dataset are politics, economy, culture, international, military, education, and entertainment.
[0020] A dual-modal forged information detection method integrating emotional features as described above, characterized in that: the domain encoder is a domain classification multi-layer perceptron.
[0021] A method for detecting dual - modal forged information integrating emotional features as described above, characterized in that: The AdaIN network structure includes multiple experts, each expert is a multi - layer perceptron, denoted as f1, f2, …, fn. The inputs of each expert are the same, and each expert does not share weights.
[0022] A method for detecting dual - modal forged information integrating emotional features as described above, characterized in that: S3 includes
[0023] S31. Feature readjustment. By using a multi - layer perceptron to train the mean and variance of the original features, the trained mean and variance guide the adjustment of the normalized feature representation. The AdaIN network adjusts the feature distribution according to the contribution of each modality, and there is
[0024]
[0025] where μ r and σ r respectively represent the mean and standard deviation of the input features, e represents the adjusted features. After the text, image, and hybrid modalities are adjusted by the AdaIN network, the adjusted text features e t , image features e i and hybrid - modality features e m are obtained respectively;
[0026] S32. Feature aggregation and classification. Concatenate the text features e t , image features e i and hybrid - modality features e m to obtain the feature f mix . Input f mix into the classifier to obtain the classification result
[0027] The beneficial effects of the present invention are:
[0028] The present invention uses the GPT 4V large - language model to extract the emotional features of text and images respectively, and analyze the emotional consistency between text and images. It uses a mixture - of - experts network to process image features, text features, and hybrid features based on domain categories, and uses the AdaIN network to adjust the feature distribution according to the contribution of each modality, and fuses the adjusted features and inputs them into the classifier to obtain the final classification result. Integrating and analyzing emotional features with text modality, image modality, and hybrid modality greatly improves the accuracy of information detection. [Description of the Drawings]
[0029] Figure 1 It is the schematic diagram of the present invention. **Detailed Implementation Modes**
[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0031] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the descriptions involving "preferred", "sub-preferred", etc. in the present invention are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "preferred" or "sub-preferred" may explicitly or implicitly include at least one of such features.
[0032] As Figure 1 shown, a dual-modal forged information detection method integrating emotional features includes S1, emotional feature extraction, which extracts the emotional categories of text and images respectively through a large language model, and gives the emotional intensity of the respective categories, as well as scores the emotions; analyzes the emotional consistency of the mixed-modal features of text and images; among them, the emotional category represents the possibility of a specific emotion, the emotional intensity quantifies the emotional intensity of different words in the same category of emotions, the emotional score is the coarse-grained emotional score for the entire text, and the auxiliary feature is the vectorized non-word element.
[0033] Specifically, S1 includes
[0034] S11, extraction of emotional features of text and images.
[0035] Set the text prompt learning template and the image prompt learning template for the GPT 4V large language model, so that the GPT 4V large language model extracts the emotional category, emotional intensity, and emotional score of the image and text according to the prompt learning template. Among them, the prompt template is: "You are an expert in sentiment analysis.Please extract the sentimentcategory,sentiment intensity,and sentiment score from the images and texts.";
[0036] S12, emotional analysis of the mixed-modal features of text and images.
[0037] Set a hybrid-modal prompt learning template for the GPT 4V large language model, so that the GPT 4V large language model analyzes the sentiment consistency of the hybrid modality according to the prompt learning template. Among them, the prompt learning template is: "You are an expert in sentiment analysis. Please analyze whether the sentiments of the images and text are consistent."
[0038] S2. The mixture-of-experts network processes features based on domain classes, and calculates the domain probabilities of the text modality, image modality, and hybrid modality through a text domain encoder, an image domain encoder, and an image-text hybrid feature domain encoder respectively; the mixture-of-experts network learns the feature information of the text modality, image modality, and hybrid modality according to the domain information;
[0039] Specifically, S2 includes
[0040] S21. Train the domain encoder.
[0041] Use the weibo dataset with domain labels as the training set data, and train the domain classification multi-layer perceptron (MLP) for each modality of the text modality, image modality, and hybrid modality respectively. The trained domain classification MLP can calculate the probability of each domain in its respective modality. Among them, 64 samples are set for each batch training, and it is iterated 50 times; the domain categories included in the weibo dataset are politics, economy, culture, international, military, education, and entertainment;
[0042] S22. Calculate the domain probability.
[0043] Use the trained domain classification MLP to calculate the domain probabilities of each modality respectively, and the domain probability of the text modality is The domain probability of the image modality is The domain probability of the hybrid modality is The purpose of introducing the domain category features is to model and express the features from the perspective of the domain, and to separately express the domain features through the mixture-of-experts structure, aiming to comprehensively analyze and express the features from the perspectives of multiple domains.
[0044] S23. The mixture-of-experts network learns the feature information.
[0045] Existing methods generally participate in subsequent fusion processing after directly extracting features, usually ignoring or failing to consider the more targeted domain information and the characteristics that the dependence of news in different domains on sentiment is different. Therefore, by using the mixture-of-experts network, modeling the domain categories and combining them with the features of each modality, features more conducive to identifying fake news can be obtained.
[0046] The structure of the Mixture of Experts (MoE) network includes multiple experts, where each expert is a Multi-Layer Perceptron (MLP), denoted as f1, f2, …, fn. The structure of each expert is the same as that of a fully connected layer with an activation function. The input of each expert is the same, namely the explicit image features. Since each expert does not share weights, different outputs are obtained for each expert.
[0047] The common MoE network structure usually combines the outputs of each expert by taking the average in order to better fuse the outputs of each expert. However, the average output means that the relevance of each expert cannot be highlighted, that is, the best targeted results cannot be obtained. Based on the domain category vector in the text modality By combining the outputs of each expert with the weights for each expert. In this way, the importance of each expert can be quantified according to the domain category vector. The output of the MoE structure in the text modality is denoted as The output of the MoE structure in the image modality is denoted as The output of the MoE structure in the mixed modality is denoted as:
[0048] S3. The AdaIN network adjusts the modality distribution. The AdaIN network adjusts the feature distribution according to the contribution of each modality, fuses the adjusted features, and inputs them into the classifier to obtain the final classification result.
[0049] Specifically, S3 includes
[0050] S31. Feature readjustment.
[0051] In the process of knowledge fusion, domain-specific knowledge containing text, image, and mixed-modal sentiment information is obtained. In order to effectively utilize this knowledge, domain-specific knowledge and cross-domain knowledge are adaptively reweighted and aggregated at this stage.
[0052] The adaptive instance normalization (AdaIN) network method is used for reweighting. By using a Multi-Layer Perceptron (MLP) to train the mean and variance of the original features, and then using the trained mean and variance to guide the adjustment of the normalized feature representation. The AdaIN network adjusts the feature distribution according to the contribution of each modality, and there is
[0053]
[0054] where μ r and σ r respectively represent the mean and standard deviation of the input features, e represents the adjusted features. After the text, image, and mixed modality are adjusted by the AdaIN network, the adjusted text features e t and image features ei and the hybrid modal feature e m ;
[0055] S32. Feature aggregation and classification. Aggregate the text feature e t , the image feature e i and the hybrid modal feature e m to splice and obtain the feature f mix . Input f mix into the classifier to obtain the classification result
[0056] The above is only the preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the content of the specification and drawings of the present invention under the inventive concept of the present invention, or directly or indirectly applied to other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A dual-modal forged information detection method integrating emotional features, characterized in that: including S1. Emotional feature extraction: The large language model extracts the emotional categories of text and images respectively, gives the emotional intensity of the respective categories, and scores the emotions; analyze the emotional consistency of the text and image mixed-modal features; S2. Mixture of Experts Network processes features based on domains. Calculate the domain probabilities of the text modality, image modality, and image-text mixed modality respectively through the text domain encoder, image domain encoder, and image-text mixed feature domain encoder; The Mixture of Experts Network learns the feature information of the text modality, image modality, and mixed modality according to the domain information; S3. AdaIN network adjusts the modality distribution. The AdaIN network adjusts the feature distribution according to the contribution of each modality, fuses the adjusted features, and inputs them into the classifier to obtain the final classification result.
2. The dual-modal forged information detection method integrating emotional features according to claim 1, characterized in that: In S1, it includes S11. Emotional feature extraction of text and images: Set the text prompt learning template and image prompt learning template. The large language model extracts the emotional categories, emotional intensity, and emotional scores of text and images respectively according to the text prompt learning template and image prompt learning template; S12. Emotional analysis of text and image mixed-modal features: Set the mixed-modal prompt learning template. The large language model analyzes the emotional consistency of the text and image mixed-modal features according to the mixed-modal prompt learning template.
3. The dual-modal forged information detection method integrating emotional features according to claim 2, characterized in that: The large language model is the GPT 4V large language model.
4. A dual-modal forged information detection method integrating emotional features according to claim 1, characterized in that: In S2, it includes S21. Train the domain encoder: Use the training set data with domain labels to train the domain encoders of the text modality, image modality, and image-text mixed modality respectively; S22. Calculate the domain probabilities. Use the trained domain encoder to calculate the domain probabilities of each modality respectively, and obtain that the domain probability of the text modality is The domain probability of the image modality is The domain probability of the mixed modality is S23. The mixture of experts network learns the feature information, and the output of the mixture of experts network structure in the text modality is denoted as The output of the mixture of experts network structure in the image modality is denoted as The output of the mixture of experts network structure in the hybrid modality is denoted as:
5. The dual-modal forged information detection method integrating emotional features according to claim 4, characterized in that: The training set data is the weibo training set data.
6. The dual-modal forged information detection method integrating emotional features according to claim 5, characterized in that: Set 64 samples for each batch training and iterate 50 times.
7. A dual-modal forged information detection method integrating emotional features according to claim 5, characterized in that: The domain categories included in the weibo dataset are politics, economy, culture, international, military, education, and entertainment.
8. A dual-modal forged information detection method integrating emotional features according to claim 4, characterized in that: The domain encoder is a domain classification multi-layer perceptron.
9. A dual-modal forged information detection method integrating emotional features according to claim 4, characterized in that: The AdaIN network structure includes multiple experts. Each expert is a multi-layer perceptron, denoted as f1, f2, …, fn. The input of each expert is the same, and each expert does not share weights.
10. A method for detecting bimodal forged information integrating emotional features according to claim 1, characterized in that: In S3, it includes S31. Feature readjustment: Train the mean and variance of the original features through a multi-layer perceptron. The trained mean and variance guide the adjustment of the normalized feature representation. The AdaIN network adjusts the feature distribution according to the contribution of each modality, and there is Among them, μ r and σ r respectively represent the mean and standard deviation of the input features, e represents the adjusted features. After the text, image, and mixed modality are adjusted by the AdaIN network, the adjusted text feature e t , image feature e i and mixed modality feature e m ; S32. Feature aggregation and classification. Aggregate the text feature e t , the image feature e i and the hybrid modality feature e m by concatenation to obtain the feature f mix . Input f mix into the classifier to obtain the classification result
Citation Information
Cited By
Image forgery detection method based on multi-modal large language model
CN121236571A
An image forgery detection method based on a multi-modal large language model
CN121236571B