Multi-mode forged news detection method and system based on local information enhancement
By introducing a local forgery information encoder and a fine-grained contrast alignment module, combined with a multimodal feature fusion module and a large language model, the problems of cross-modal semantic contradictions and insufficient interpretability in multimodal fake news detection are solved, achieving highly accurate and interpretable fake news detection.
Patent Information
- Application Number
- CN202511087238.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multimodal fake news detection methods struggle to effectively identify cross-modal semantic contradictions, ignore localized forged information, and lack interpretability of detection results, limiting their effectiveness in key scenarios such as judicial evidence collection and public opinion guidance.
By introducing a local forgery information encoder to extract local forgery features from news images, and combining a fine-grained contrast alignment module and a multimodal feature fusion module, a forged news detection model is constructed, and a large language model is used to improve the interpretability of the detection results.
It enhances the sensitivity to subtle forgery traces, improves the accuracy and robustness of fake news detection, meets the needs of multimodal forgery detection tasks for fine-grained analysis, and improves the interpretability of detection results.
Smart Images

Figure CN120974413A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal media forgery detection technology, specifically to a multimodal forged news detection method and system based on local information enhancement. Background Technology
[0002] Breakthrough advancements in generative artificial intelligence technology are reshaping the paradigm of digital content production and dissemination. Multimodal content generation models have enabled cross-modal collaborative creation of text, images, and videos. While this technological expansion is driving innovation in the creative industries, it also raises new information security risks. Currently, multimodal media forgery based on deepfake technology and generative models exhibits an industrialized production model, with its content complexity and dissemination speed far exceeding the capabilities of traditional manual review mechanisms. This type of forgery often misleads audiences through carefully designed cross-modal semantic associations; for example, visual elements in forged images combined with misleading textual descriptions create a false narrative with apparent plausibility.
[0003] Traditional multimodal fake news detection tasks typically focus only on the semantic consistency of image-text pairs, neglecting the possibility of forgery in the content itself. Therefore, this paper extends the paradigm of this task, systematically defining the multimodal media forgery detection and localization problem for the first time. It requires simultaneously detecting the forgery types of both images and text, and locating the forged regions. Researchers constructed a dataset containing 230,000 labeled samples, covering four types of forgery: face replacement, face attribute editing, text replacement, and text attribute editing. The annotation information includes binary labels, fine-grained types, forged regions, and lexical units.
[0004] Existing detection systems face a dual dilemma: single-modal analysis methods are limited by independent verification of text or images, making it difficult to identify cross-modal semantic contradictions; while mainstream multimodal models establish intermodal correlations, they generally neglect the micro-feature analysis of local forgery information. More seriously, the insufficient interpretability of the detection system's decision-making process makes its judgments difficult to meet the transparency requirements of content governance, severely restricting the effectiveness of technological tools in key scenarios such as judicial evidence collection and public opinion guidance. Therefore, a new media forgery detection method is urgently needed, which can reasonably utilize local forgery information while endowing the detection method with stronger interpretability. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes a multimodal fake news detection method and system based on local information enhancement. By introducing a local fake information encoder to extract local fake features from news images, and combining a fine-grained contrast alignment module and a multimodal feature fusion module to construct a fake news detection model, the detection capability for multimodal fake news is enhanced.
[0006] The specific plan is as follows:
[0007] On the one hand, multimodal fake news detection methods based on local information enhancement include:
[0008] S1, the news image features and news text features of the news to be detected are extracted by the image encoder and the text encoder respectively, and the local forgery information encoder is introduced to extract the local forgery features of the news image;
[0009] S2, Construct a fake news detection model, which includes a fine-grained comparison and alignment module, a multimodal feature fusion module, and a fake news detector;
[0010] S3, through the fine-grained comparison and alignment module, the extracted news image features, news text features and local forgery features are aligned to obtain the aligned multimodal features. The aligned multimodal features are then fused through the multimodal feature fusion module to obtain the fused single feature.
[0011] S4. Based on the fused single feature, calculate the news forgery task loss through the news forgery detector. By combining the news forgery detection model with the large language model, calculate the large language model loss. Calculate the gradient based on the news forgery task loss and the large language model loss. Update the parameters of the news forgery detection model and the large language model through the gradient to obtain the updated news forgery detection model and the large language model.
[0012] S5 outputs quantitative detection results of the news to be detected through the fake news detector of the updated fake news detection model, and generates qualitative detection results of the news to be detected through the updated large language model.
[0013] Furthermore, in S1, news image features and news text features are extracted using an image encoder and a text encoder, respectively, specifically including:
[0014] For the input news images And the news text T, using a pre-trained image encoder E v With text encoder E t Feature extraction is performed on it, and the calculation formula is as follows:
[0015] e v =E v (I);
[0016] e t =E t (T);
[0017] Among them, e v Indicates the characteristics of news images; e t Indicates the characteristics of news text.
[0018] Furthermore, in S1, a local forgery information encoder is introduced to extract local forgery features of the news image. The calculation formula is as follows:
[0019] e d =E d (I);
[0020] Among them, e d Indicates local forgery features; E d () indicates a partial forgery information encoder.
[0021] Furthermore, in S3, the extracted news image features, news text features, and local forgery features are aligned through a fine-grained comparison alignment module, specifically including:
[0022] The fine-grained contrast alignment loss is calculated as follows:
[0023]
[0024] Where k represents the k-th sample in a sample set with a total sample size of K; I + T represents the positive image sample in the image-text pair; + I represents a positive sample of text in an image-text pair; d + Indicate I + Local image; I - Indicate I + Fake samples; T - T represents + Fake samples; and Represents the forged image corresponding to the k-th sample and its partial forged image; h v and h t This is a linear mapping function used to map features to the same dimension; and Represents the contrastive learning loss between images and text; and E represents the contrastive learning loss between local and global features; p() S represents the expected value; S() represents the feature similarity; τ is the temperature coefficient; L ITC This indicates the fine-grained contrast alignment loss;
[0025] News image features, news text features, and local forgery features are extracted based on fine-grained contrastive alignment loss.
[0026] Furthermore, in S3, aligned multimodal features are fused through a multimodal feature fusion module, and the calculation formula is as follows:
[0027] f d =Q(ed ,e t );
[0028] f v =Q(e v ,e t );
[0029] f = CrossAttention(f d ,f v );
[0030] Where Q() represents the feature fusion module, CrossAttentio() represents the cross-attention module; f represents the fused single feature; f d Indicates local features; f v Represents global features; e v Indicates the characteristics of news images; e t Indicates the characteristics of news text; e d This indicates a local forgery feature.
[0031] Furthermore, in S4, the fake news detector includes a binary detector for detecting the authenticity of news, a fake news type detector for detecting the type of fake news, and a text fake news location detector for detecting the location of text fake news.
[0032] The fake news detector uses the obtained fusion features f = {f cls ,f tok The binary classification loss, multi-class classification loss, and localization loss are calculated as follows:
[0033] The binary classification loss is calculated using a binary detector.
[0034]
[0035] The multi-class loss is calculated using a forgery type detector.
[0036]
[0037] The text localization loss is calculated using a text forgery localization detector.
[0038]
[0039] Among them, f cls To ensure that the fused single feature contains globally relevant category information; f tok The fused single feature contains semantic information about the image and text; H is the cross-entropy loss; C m and C bTwo multilayer perceptrons are used to convert f cls Mapped to the corresponding dimensions; L and L mul For genuine and counterfeit types, use accurate labels; D t Represents a multilayer perceptron, used to convert f tok Mapped to the corresponding dimension; L tok Labels indicating areas of text forgery; E (I,T)~P This indicates the expectation for the input news image text;
[0040] The fake news detector is updated based on binary classification loss, multi-class classification loss, and text localization loss to determine the authenticity of news, the type of fake news, and the location of fake text.
[0041] Furthermore, in S4, the fake news detection model is combined with a large language model, specifically including:
[0042] The fused single feature f is passed through a simple fully connected layer h. LLM Projecting onto the same semantic space dimension as the text input of the large language model, we obtain the fake feature input H. v ;
[0043] The projection features and text instructions are jointly encoded to obtain joint features H. The joint features H and the news image I are then input into the large language model to obtain a qualitative output result O.
[0044] The fully connected layer is optimized using cross-entropy loss, which is calculated using the following formula:
[0045]
[0046] Where L LLM Represents the loss of a large language model; O t This represents the t-th word in the final output of length T; O <t Let represent all words before the t-th word; P represents the probability.
[0047] Furthermore, in S4, gradients are calculated based on the loss from the news forgery task and the loss from the large language model. The parameters of the news forgery detection model and the large language model are updated using these gradients. The formula for calculating the total loss L is as follows:
[0048] L = L ITC +L TMG +L BLC +L MLC +L LLM ;
[0049] Where L ITC Indicates fine-grained contrast alignment loss; L TMG L represents the text localization loss; BLCL represents binary classification loss; MLC L represents multi-class classification loss; llM Indicates the loss of a large language model;
[0050] The gradients of the parameters of the fake news detection model and the large language model are calculated using the total loss L.
[0051] On the other hand, multimodal fake news detection systems based on local information enhancement include:
[0052] The feature extraction module extracts the news image features and news text features of the news to be detected through the image encoder and text encoder respectively, and introduces the local forgery information encoder to extract the local forgery features of the news image;
[0053] The model building module constructs a fake news detection model, which includes a fine-grained comparison and alignment module, a multimodal feature fusion module, and a fake news detector.
[0054] The feature fusion module aligns the extracted news image features, news text features, and local forgery features through the fine-grained comparison and alignment module to obtain aligned multimodal features. The multimodal feature fusion module then fuses the aligned multimodal features to obtain a fused single feature.
[0055] The update module calculates the news forgery task loss based on the fused single feature using the news forgery detector, calculates the large language model loss by combining the news forgery detection model with the large language model, calculates the gradient based on the news forgery task loss and the large language model loss, and updates the parameters of the news forgery detection model and the large language model using the gradient to obtain the updated news forgery detection model and the large language model.
[0056] The detection module outputs quantitative detection results of the news to be detected through the fake news detector of the updated fake news detection model, and generates qualitative detection results of the news to be detected through the updated large language model.
[0057] The present invention adopts the above technical solution and has the following beneficial effects:
[0058] (1) This invention, by introducing a local forgery information encoder, can effectively extract local forgery features from news images, thereby enhancing the sensitivity to subtle forgery traces, improving the accuracy and robustness of forged news detection, and solving the problem of limited detection performance caused by traditional methods neglecting local forgery information.
[0059] (2) The present invention uses a fine-grained comparison alignment module and a multi-modal feature fusion module to achieve cross-modal alignment and deep fusion between images, text and local forgery features, which improves the accuracy of forgery type identification and forgery region localization and meets the requirements of multi-modal forgery detection tasks for fine-grained analysis.
[0060] (3) This invention combines a fake news detection model with a large language model to generate qualitative explanations based on quantitative detection, thereby improving the interpretability and expressive power of the detection results and supporting the judgment of the authenticity of news, the identification of fake types, and the location of text forgery. Attached Figure Description
[0061] Figure 1 This is a flowchart of a multimodal fake news detection method based on local information enhancement according to an embodiment of the present invention;
[0062] Figure 2 This is a schematic diagram illustrating the main steps of an embodiment of the present invention;
[0063] Figure 3 This is a schematic diagram illustrating the task in an embodiment of the present invention;
[0064] Figure 4 This is a diagram of a multimodal fake news detection system based on local information enhancement, according to an embodiment of the present invention. Detailed Implementation
[0065] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0066] like Figure 1 As shown, the present invention provides a multimodal fake news detection method based on local information enhancement, comprising:
[0067] S1: The image encoder and text encoder extract the news image features and news text features of the news to be detected, respectively. The local forgery information encoder is introduced to extract the local forgery features of the news image.
[0068] Specifically, a local forgery information encoder is introduced to extract local forgery features from news images. The calculation formula is as follows:
[0069] e d =E d (I);
[0070] Among them, e d Indicates local forgery features; E d () indicates a partial forgery information encoder.
[0071] Specifically, a local forgery information encoder is introduced to extract local forgery features from news images. The calculation formula is as follows:
[0072] e d =E d (I);
[0073] Among them, e d Indicates local forgery features; E d () indicates a partial forgery information encoder.
[0074] Specifically, such as Figure 2 and Figure 3 As shown, this invention mainly targets the detection of forgery in news image text. Given a news image and its corresponding text description, the goal of this invention is to determine the authenticity of the image-text pair. For forged news image-text pairs, it is also necessary to provide the type of forgery and the precise location of the forged area.
[0075] S2, Construct a fake news detection model, which includes a fine-grained comparison and alignment module, a multimodal feature fusion module, and a fake news detector.
[0076] S3 aligns the extracted news image features, news text features, and local forgery features through the fine-grained comparison and alignment module to obtain aligned multimodal features. The aligned multimodal features are then fused through the multimodal feature fusion module to obtain a fused single feature.
[0077] Specifically, the extracted news image features, news text features, and local forgery features are aligned using the fine-grained comparison and alignment module, and the calculation formula is as follows:
[0078]
[0079]
[0080] Where k represents the k-th sample in a sample set with a total sample size of K; I + T represents the positive image sample in the image-text pair; + I represents a positive sample of text in an image-text pair; d + Indicate I + Local image; I - Indicate I + Fake samples; T - T represents + Fake samples; and Represents the forged image corresponding to the k-th sample and its partial forged image; h v and h t This is a linear mapping function used to map features to the same dimension; and Represents the contrastive learning loss between images and text; and E represents the contrastive learning loss between local and global features; p() S represents the expected value; S() represents the feature similarity; τ is the temperature coefficient; L ITC This indicates the fine-grained contrast alignment loss;
[0081] News image features, news text features, and local forgery features are extracted based on fine-grained contrastive alignment loss.
[0082] Specifically, the aligned multimodal features are fused through the multimodal feature fusion module, and the calculation formula is as follows:
[0083] f d =Q(e d ,e t );
[0084] f v =Q(e v ,e t );
[0085] f = CrossAttention(f d ,f v );
[0086] Where Q() represents the feature fusion module, CrossAttentio() represents the cross-attention module; f represents the fused single feature; f d Indicates local features; f v Represents global features.
[0087] S4. Based on the fused single feature, calculate the news forgery task loss using the news forgery detector. By combining the news forgery detection model with the large language model, calculate the large language model loss. Calculate the gradient based on the news forgery task loss and the large language model loss. Update the parameters of the news forgery detection model and the large language model using the gradient to obtain the updated news forgery detection model and the large language model.
[0088] Specifically, the fake news detector includes a binary detector for detecting the authenticity of news, a fake news type detector for detecting the type of fake news, and a text fake news location detector for detecting the location of text fake news.
[0089] The fused single feature f is divided into f that contains information related to the global category. cls and information containing graphic and textual semantics f tok That is, f = {f cls ,f tok};
[0090] via f clsComplete the tasks of authenticity detection and counterfeit type classification, through f tok Complete the location of the counterfeit area;
[0091] The fake news detector uses the obtained fusion features f = {f cls ,f tok The binary classification loss, multi-class classification loss, and localization loss are calculated as follows:
[0092] For a binary detector, a binary classification loss is obtained.
[0093]
[0094] For the forgery type detector, a multi-class loss is obtained.
[0095]
[0096] For the text forgery localization model detector, the text localization loss is obtained.
[0097]
[0098] Where H is the cross-entropy loss; C m and C b Two multilayer perceptrons are used to convert f cls Mapped to the corresponding dimensions; L and L mul For genuine and counterfeit types, use accurate labels; D t Represents a multilayer perceptron, used to convert f tok Mapped to the corresponding dimension; L tok Labels indicating areas of text forgery; E (I,T)~P This indicates the expectation for the input news image text;
[0099] The fake news detector is updated based on binary classification loss, multi-class classification loss, and text localization loss to determine the authenticity of news, the type of fake news, and the location of fake text.
[0100] Specifically, to more accurately identify multimodal forged content, the detector is divided into three functional modules: a binary classification module, a forgery type classification module, and a text forgery localization module. The obtained feature f can be divided into two parts: one part contains information related to the global category, and the other part represents the semantic information of the image and text, i.e., f = {f...} cls ,f tok}, using f cls Simultaneously complete the tasks of authenticity detection and counterfeit type classification, f tok Complete the location of the fake area.
[0101] Specifically, the fake news detection model is combined with a large language model, including:
[0102] The fused single feature f is passed through a simple fully connected layer h. LLM Projecting onto the same semantic space dimension as the text input of the large language model, we obtain the fake feature input H. v ;
[0103] The projection features and text instructions are jointly encoded to obtain joint features H. The joint features H and the news image I are then input into the large language model to obtain a qualitative output result O.
[0104] The fully connected layer is optimized using cross-entropy loss, which is calculated using the following formula:
[0105]
[0106] Where L LLM Represents the loss of a large language model; O t This represents the t-th word in the final output of length T; O <t Let represent all words before the t-th word; P represents the probability.
[0107] Specifically, to enhance model interpretability and expand its application boundaries in multimodal tasks, this study innovatively introduces an integrated architecture of Large Language Models (LLM). This integration strategy not only endows the model with text generation capabilities but also enables it to express complex relationships between cross-modal data in an intuitive way. This invention integrates the features extracted by the model through a simple fully connected layer h. LLM The projected features are then projected to the same semantic space dimension as the LLM text input. The projected features and text instructions are then jointly encoded and input into the LLM. The fully connected layer is then optimized using cross-entropy loss.
[0108] Specifically, the gradient is calculated based on the loss from the news forgery task and the loss from the large language model. The parameters of the news forgery detection model and the large language model are updated using the gradient. The formula for calculating the total loss L is as follows:
[0109] L = L ITC +L TMG +L BLC +L MLC +L LLM ;
[0110] Where L ITC Indicates fine-grained contrast alignment loss; L TMG L represents the text localization loss; BLC L represents binary classification loss; MLC L represents multi-class classification loss; LLM Indicates the loss of a large language model;
[0111] The gradients of the parameters of the fake news detection model and the large language model are calculated using the total loss L.
[0112] S5, the updated fake news detection model outputs quantitative detection results, and the updated large language model generates qualitative detection results. In summary, this embodiment first extracts visual and textual features of the news using a visual encoder and a text encoder, and additionally introduces a local forgery information encoder to extract local forgery features of the image; then, a contrastive learning method is used to align the extracted features from different modalities, and a feature fusion network is used to fuse the extracted features; finally, the loss for media forgery detection is calculated using the fused features, and the gradient of this loss is used to update the parameters of the weighted perceptual network. The model is then combined with a large language model to improve its interpretability. It has the following outstanding advantages: It proposes a multimodal media forgery detection framework based on local information enhancement and fully integrates it with a multimodal large language model. By introducing a pre-trained multimodal model and a forgery detection model, it fully integrates global semantic information and local forgery information, effectively improving the model's performance on multimodal media forgery detection tasks. This invention introduces pre-trained multimodal models and local prior knowledge, applying a series of multimodal models adept at extracting local features to multimodal forgery detection for the first time. A dedicated face extractor is used to enhance local information, providing a new paradigm for the task. Furthermore, by designing a fine-grained contrast alignment module and a multimodal local-global fusion module, local and global features are effectively integrated, improving the comprehensiveness and accuracy of forgery detection. Experiments demonstrate that this invention achieves better results than previous methods. On a multimodal media forgery detection dataset, this invention outperforms existing methods in binary classification, multi-label classification, and text forgery localization tasks, especially in text localization where the F1-score improvement is significantly better than current baseline models.
[0113] Furthermore, to more intuitively demonstrate the effectiveness of the present invention, the attention distribution of the model proposed in the present invention in the image is visualized. In this embodiment, a comparative analysis of image forgery and detection is performed. The upper part is the original image, including several sets of photos of people; the middle part is the corresponding forged image, showing the modified version; the lower part is the attention map, which uses a heat map to identify the key areas of the forged part, with red areas indicating the places that have received special attention or been modified.
[0114] Specifically, to demonstrate the effectiveness of this invention, tests were conducted on a multimodal media forgery dataset, and the following methods were compared: 1) Classical multimodal methods CLIP and VILT, representing different multimodal feature fusion paradigms; 2) Multimodal pre-trained models ALBEF and BLIP-2, which achieve good results in downstream vision-language tasks through feature alignment and fusion mechanisms; 3) The baseline model HAMMER, which achieves multimodal forgery detection through hierarchical inference of the multimodal pre-trained model. The comparison results are shown in Table 1. In summary, the M method of this invention... 4 -BLIP testing yields the best results.
[0115] Table 1. Test table for the multimodal media forgery dataset;
[0116]
[0117] like Figure 4 As shown, this embodiment also discloses a multimodal fake news detection system based on local information enhancement, including:
[0118] The feature extraction module 61 extracts news image features and news text features through an image encoder and a text encoder, respectively, and introduces a local forgery information encoder to extract local forgery features of the news image;
[0119] Model building module 62 constructs a fake news detection model, which includes a fine-grained comparison and alignment module, a multimodal feature fusion module, and a fake news detector.
[0120] The feature fusion module 63 aligns the extracted news image features, news text features, and local forgery features through the fine-grained comparison and alignment module to obtain aligned multimodal features. The aligned multimodal features are then fused through the multimodal feature fusion module to obtain a fused single feature.
[0121] The update module 64 calculates the news forgery task loss based on the fused single feature through the news forgery detector, calculates the large language model loss by combining the news forgery detection model with the large language model, calculates the gradient based on the news forgery task loss and the large language model loss, and updates the parameters of the news forgery detection model and the large language model through the gradient to obtain the updated news forgery detection model and the large language model.
[0122] Detection module 65 outputs quantitative detection results from the fake news detector using the updated fake news detection model, and generates qualitative detection results using the updated large language model.
[0123] The specific implementation of the multimodal fake news detection system based on local information enhancement is the same as that of the multimodal fake news detection method based on local information enhancement, and will not be described again in this embodiment.
[0124] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.
Claims
1. A multimodal news forgery detection method based on local information enhancement, characterized in that, include: S1, the news image features and news text features of the news to be detected are extracted by the image encoder and the text encoder respectively, and the local forgery information encoder is introduced to extract the local forgery features of the news image; S2, Construct a fake news detection model, which includes a fine-grained comparison and alignment module, a multimodal feature fusion module, and a fake news detector; S3, through the fine-grained comparison and alignment module, the extracted news image features, news text features and local forgery features are aligned to obtain the aligned multimodal features. The aligned multimodal features are then fused through the multimodal feature fusion module to obtain the fused single feature. S4. Based on the fused single feature, calculate the news forgery task loss through the news forgery detector. By combining the news forgery detection model with the large language model, calculate the large language model loss. Calculate the gradient based on the news forgery task loss and the large language model loss. Update the parameters of the news forgery detection model and the large language model through the gradient to obtain the updated news forgery detection model and the large language model. S5 outputs quantitative detection results of the news to be detected through the fake news detector of the updated fake news detection model, and generates qualitative detection results of the news to be detected through the updated large language model.
2. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S1, news image features and news text features are extracted using an image encoder and a text encoder, respectively, specifically including: For the input news images And the news text T, using a pre-trained image encoder E v With text encoder E t Feature extraction is performed on it, and the calculation formula is as follows: yes v =E v (I); And t =And t (T); Among them, e v Indicates the characteristics of news images; e t Indicates the characteristics of news text.
3. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S1, a local forgery information encoder is introduced to extract local forgery features of the news image. The calculation formula is as follows: yes d =E d (I); Among them, e d Indicates local forgery features; E d () indicates a partial forgery information encoder.
4. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S3, the extracted news image features, news text features, and local forgery features are aligned using a fine-grained comparison and alignment module, specifically including: The fine-grained contrast alignment loss is calculated as follows: Where k represents the k-th sample in a sample set with a total sample size of K; I + T represents the positive image sample in the image-text pair; + I represents a positive sample of text in an image-text pair; d + Indicate I + Local image; I - Indicate I + Fake samples; T - T represents + Fake samples; and Represents the forged image corresponding to the k-th sample and its partial forged image; h v and h t This is a linear mapping function used to map features to the same dimension; and Represents the contrastive learning loss between images and text; and E represents the contrastive learning loss between local and global features; p() S represents the expected value; S() represents the feature similarity; τ is the temperature coefficient; L ITC This indicates the fine-grained contrast alignment loss; News image features, news text features, and local forgery features are extracted based on fine-grained contrastive alignment loss.
5. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S3, aligned multimodal features are fused through a multimodal feature fusion module, and the calculation formula is as follows: f d =Q(e d ,e t ); f v =Q(e v ,e t ); f=CrossAttention(f d ,f v ); Where Q() represents the feature fusion module, CrossAttentio() represents the cross-attention module; f represents the fused single feature; f d Indicates local features; f v Represents global features; e v Indicates the characteristics of news images; e t Indicates the characteristics of news text; e d This indicates a local forgery feature.
6. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S4, the fake news detector includes a binary detector for detecting the authenticity of news, a fake news type detector for detecting the type of fake news, and a text fake news location detector for detecting the location of text fake news. The fake news detector uses the obtained fusion features f = {f cls ,f tok The binary classification loss, multi-class classification loss, and localization loss are calculated as follows: The binary classification loss is calculated using a binary detector. The multi-class loss is calculated using a forgery type detector. The text localization loss is calculated using a text forgery localization detector. Among them, f cls To ensure that the fused single feature contains globally relevant category information; f tok The fused single feature contains semantic information about the image and text; H is the cross-entropy loss; C m and C b Two multilayer perceptrons are used to convert f cls Mapped to the corresponding dimensions; L and L mul Real labels for genuine and counterfeit types; D t Represents a multilayer perceptron, used to convert f tok Mapped to the corresponding dimension; L tok Labels indicating areas of text forgery; E (I,T)~P This indicates the expectation for the input news image text; The fake news detector is updated based on binary classification loss, multi-class classification loss, and text localization loss to determine the authenticity of news, the type of fake news, and the location of fake text.
7. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S4, the fake news detection model is combined with a large language model, specifically including: The fused single feature f is passed through a simple fully connected layer h. LLM Projecting the pseudo-feature input h onto the same semantic space dimension as the text input of the large language model yields the pseudo-feature input h. v ; The projection features and text instructions are jointly encoded to obtain joint features H. The joint features H and the news image I are then input into the large language model to obtain a qualitative output result O. The fully connected layer is optimized using cross-entropy loss, which is calculated using the following formula: Where L LLM Represents the loss of a large language model; O t This represents the t-th word in the final output of length T; O <t Let represent all words before the t-th word; P represents the probability.
8. The multimodal fake news detection method based on local information enhancement according to claim 1, characterized in that, In S4, the gradient is calculated based on the loss from the news forgery task and the loss from the large language model. The parameters of the news forgery detection model and the large language model are updated using the gradient. The formula for calculating the total loss L is as follows: L=L ITC +L TMG +L BLC +L MLC +L LLM ; Where L ITC Indicates fine-grained contrast alignment loss; L TMG L represents the text localization loss; BLC L represents binary classification loss; MLC L represents multi-class classification loss; LLM Indicates the loss of a large language model; The gradients of the parameters of the fake news detection model and the large language model are calculated using the total loss L.
9. A multimodal fake news detection system based on local information enhancement, characterized in that, include: The feature extraction module extracts the news image features and news text features of the news to be detected through the image encoder and text encoder respectively, and introduces the local forgery information encoder to extract the local forgery features of the news image; The model building module constructs a fake news detection model, which includes a fine-grained comparison and alignment module, a multimodal feature fusion module, and a fake news detector. The feature fusion module aligns the extracted news image features, news text features, and local forgery features through the fine-grained comparison and alignment module to obtain aligned multimodal features. The multimodal feature fusion module then fuses the aligned multimodal features to obtain a fused single feature. The update module calculates the news forgery task loss based on the fused single feature using the news forgery detector, calculates the large language model loss by combining the news forgery detection model with the large language model, calculates the gradient based on the news forgery task loss and the large language model loss, and updates the parameters of the news forgery detection model and the large language model using the gradient to obtain the updated news forgery detection model and the large language model. The detection module outputs quantitative detection results of the news to be detected through the fake news detector of the updated fake news detection model, and generates qualitative detection results of the news to be detected through the updated large language model.