A Multimodal Fake News Detection Method Based on Multi-Level Semantic Enhancement
By constructing a multi-modal fake news detection method with multi-level semantic enhancement, RLprompt generates image description and adaptive hard attention mechanism, combined with knowledge graph, the problem of failure to fully mine image semantics and remove noise in the existing methods is solved, and a higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202311298800.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-10-09
AI Technical Summary
The existing multimodal fake news detection methods fail to fully tap the potential semantic information of the image and fail to effectively remove noise in the news entity, resulting in low detection accuracy.
By constructing a multi-modal fake news detection method with multi-level semantic enhancement, RLprompt unit is used to generate image descriptions related to text entities, combining adaptive hard attention mechanisms and knowledge graphs, accurate news knowledge semantic information is extracted, irrelevant entities are removed, and detection accuracy is improved.
Fully explore the potential semantic information of the image, remove the noise of the actual news, and improve the accuracy of multimodal fake news detection.
Smart Images

Figure CN117315695B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimodal fake news detection, and more specifically, it is a multimodal fake news detection method based on multi-level semantic enhancement. Background Art
[0002] As social media platforms have become increasingly integrated into people's lives, they have become a major source of information for the public. Regrettably, this has been accompanied by an explosive growth in fake news content. Due to the deceptive nature of fake news, people are often misled by it, which in turn affects their judgment and decision-making. In addition, fake news can also be used to distort and fabricate facts, guide public opinion, and have an adverse impact on social trust and stability. Therefore, in order to prevent the proliferation of fake news, there is an urgent need for automated detection methods to identify fake news and improve the credibility of the social media ecosystem. Fake news detection is a binary classification problem, the goal of which is to analyze news content to determine its authenticity. Traditional fake news detection focuses on text content and relies on extracting semantic features from text, the social media dissemination process, and user interactions to detect fake news. However, with the continuous development of multimedia technology, rumor-mongers are increasingly using multimodal content (such as attractive pictures) to attract public attention and facilitate faster dissemination. Therefore, the field of multimodal fake news detection is receiving increasing attention.
[0003] Some progress has been made in the field of multimodal fake news detection, but the existing methods do not make full use of news image information. Some methods simply extract image features through pre-trained VGG-19 or ResNet-50 networks, or combine frequency domain information as a supplement to image features. However, these methods do not fully exploit the semantic information of images, especially without combining text information to perform semantic extraction on specific image content. In addition, simply extracting image features cannot effectively narrow the modality gap between image and text features, which is not conducive to subsequent multimodal fusion. Besides the basic features of news content, the knowledge-level features of news entities are also crucial for predicting the authenticity of samples. A knowledge graph consists of entities as graph nodes and different types of relationships as edges, which contains rich background knowledge information. Therefore, in order to improve the performance of fake news detection, some methods use the high-order knowledge semantic information of news entities as an objective evidence source and incorporate it into the model. For the acquisition of visual entities, these methods only use YOLOv3 or FasterR-CNN to detect objects in images. However, this is not sufficient for the detection of visual entities, because the visual entities recognized by detectors such as YOLOv3 are limited to pre-trained entity categories, while the entity types contained in image descriptions often imply some entities in the open world. Then, these methods perform fake news detection by aligning and fusing visual entities and text entities, or discovering inconsistent semantic features at the knowledge level. But these methods often ignore the additional noise effects brought by adding external knowledge. When extracting visual or text entities, some irrelevant news entities are often recognized, which not only do not contribute to detecting the authenticity of news content, but also introduce varying degrees of noise information into the model. Summary of the Invention
[0004] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes a multimodal fake news detection method based on multi-level semantic enhancement, in order to solve the problem that traditional multimodal fake news detection methods fail to fully exploit the potential semantic information of images, and remove the noise contained in news entities to obtain accurate news knowledge semantic information, so as to more accurately detect multimodal fake news.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A multimodal fake news detection method based on multi-level semantic enhancement of the present invention is characterized in that it is carried out according to the following steps:
[0007] Step 1: Collection and preprocessing of multimodal news data;
[0008] Extract the text content of each multimodal news on the social media platform and its corresponding image to obtain a news text set and a news image set Among them, T i represents the i-th news text; I i represents the i-th i news image corresponding to T;
[0009] Set the authenticity label of the i-th news image I i and its i-th news text T i as y i , and y i ∈ {0, 1}; thus constructing the training dataset Among them, N represents the number of news in the training dataset;
[0010] Step 2: Construct a multi-modal fake news detection network, including: an image semantic enhancement module, a multi-modal fusion module, and a knowledge semantic enhancement module;
[0011] Step 3: Construct an image semantic enhancement module, including: an RLprompt unit, a prompt statement construction unit, and a BLIP model;
[0012] Step 3.1: The RLprompt unit generates n learnable prompt words {Z}1{Z}2...{Z} k ...{Z} n ;
[0013] Processing of the prompt statement construction unit:
[0014] Step 3.2.1: Use the entity linking tool TAGME to perform entity recognition on the i-th news text T i to obtain the text entity set Among them, represents the j-th text entity in the i-th news text T i , and M represents the number of text entities in each news text;
[0015] Step 3.2.2: Construct an interactive prompt statement with two text entities in the i-th news text T i as the prompt main vocabulary represents the j'-th entity in the i-th news text T i ; and represents the connector;
[0016] Construct a local prompt statement with a single text entity in the i-th news text T i as the prompt main vocabulary
[0017] Step 3.3: The BLIP model generates the image description set of the i-th multi-modal news
[0018] Feed the i-th news image I i into the BLIP model to obtain the global image description
[0019] Feed the i-th news image I i and the interactive prompt statement P con,i into the BLIP model to guide the generation of the interactive image description of the news
[0020] Feed the i-th news image I i and the local prompt statement P loc,i into the BLIP model to guide the generation of the local image description of the news
[0021] Step 4: Construct a multimodal fusion module, including: a feature extraction unit, a cross-modal feature enhancement unit;
[0022] Step 4.1: The feature extraction unit is used to extract the initial features of different modalities of the multimodal news;
[0023] Step 4.2: The cross-modal feature enhancement unit is used to process the initial features of different modalities and output the cross-modal feature F N,i ;
[0024] Step 5: Construct a knowledge semantic enhancement module, including an entity linking unit, an adaptive hard attention mechanism unit, a cross-modal knowledge interaction unit, a knowledge fusion unit;
[0025] Step 5.1: The entity linking unit is used to extract news entities and link them to the knowledge graph;
[0026] Step 5.2: The adaptive hard attention mechanism unit is used to process the entity embedding features to obtain the filtered entity features;
[0027] Step 5.3: The cross-modal knowledge interaction unit is used to enhance the features of the entity embedding to obtain the entity knowledge interaction features;
[0028] Step 5.4: The knowledge fusion unit concatenates the filtered visual entity feature F VE,i 、the filtered text entity feature F TE,i 、the text entity knowledge interaction feature and the image entity knowledge interaction feature as a group of entity embeddings, then applies the self-attention mechanism to further model the entity embeddings, and uses the fully connected layer and the average pooling layer to process the modeled entity embeddings, and finally outputs the news background knowledge feature F E,i ;
[0029] Step 6. Optimization of the multi-modal fake news detection network:
[0030] Step 6.1. Predict the probability that the i-th multi-modal news is fake news using Equation (8)
[0031]
[0032] In Equation (8), σ represents the sigmoid activation function, W c represents the weight matrix of the classifier, and b c represents the bias vector; F′ M,i represents the global multi-modal news feature after dimensionality change;
[0033] Step 6.2. Construct a cross-entropy loss function using Equation (9)
[0034]
[0035] Step 6.3. Based on the training dataset X, use the Adam optimization strategy to train the multi-modal fake news detection network until the total loss function of the network converges, so as to obtain an optimal multi-modal fake news detection model for predicting any multi-modal news.
[0036] The feature of the multi-modal fake news detection method based on multi-level semantic enhancement according to the present invention also lies in that the step 3.1 includes:
[0037] Step 3.1.1. The RLprompt unit uses the DistilGPT-2 language model to learn the to-be-learned prompt words {Z}1{Z}2...{Z} k ...{Z} n ; where {Z} k represents the k-th to-be-learned prompt word; n represents the number of prompt words;
[0038] Step 3.1.2. The start symbol <start>In the frozen DistilGPT-2 model, obtain the start token <start>Context embeddings are obtained, and a two-layer MLP layer is used to re-encode the context embeddings. After obtaining the encoded features, they are passed to the classification head of the frozen DistilGPT-2 model to output the first prompt word {Z}1;
[0039] Step 3.1.3: Based on Step 2.1.2, input the first prompt word {Z}1 into the frozen DistilGPT-2 model in an autoregressive generation manner, and gradually generate the to-be-learned prompt words {Z}1{Z}2...{Z} k ...{Z} n 。
[0040] The said Step 4.1 includes:
[0041] Step 4.1.1: Use the pre-trained BERT model to extract features from the i-th news text T i to obtain the feature sequence F i of the i-th news text T T,i =[f 1.i , f 2.i ,..., f l.i ,..., f L.i , where f l.i represents the text feature at the l-th word level in the i-th news text T i ;
[0042] Use the long short-term memory network LSTM to further extract features from the feature sequence F T,i and take the hidden state feature output at the last step of the long short-term memory network LSTM as the global feature F i of the i-th news text T G,i ;
[0043] Step 4.1.2: Use the visual feature encoder in the said frozen BLIP model to extract features from the i-th news image I i to obtain the image feature F V,i ;
[0044] Step 4.1.3: Use the text feature encoder in the said frozen BLIP model to extract features from the image description set C i to obtain the image description feature F c,i ;
[0045] Step 4.1.4: Use a shared-weight MLP layer to change the dimensions of the feature sequence F T,i and the global feature F G,i to make them the same as the image feature F V,i and the image description feature F C,i The dimensions are the same to obtain a feature sequence F' with changing dimensions T,i and the global feature F' G,i .
[0046] Step 4.2 includes:
[0047] Step 4.2.1: Construct an MLP layer consisting of a one-layer linear layer of m×n and a ReLU activation function layer, and input F' T,i , F' G,i , F V,i , F C,i into the MLP layer with shared weight W shared for processing, and obtain the text intermediate feature sequence F i of the i-th news text T Ts,i , the global text intermediate feature F i of the i-th news text T Gs,i , the image intermediate feature F i of the i-th news image I Vs,i , and the image description intermediate feature F 的 of the news image description Ci Cs,i ;
[0048] Step 4.2.2: According to Equation (1), apply the news image intermediate feature F Vs,i to perform an attention operation on the text intermediate feature sequence F Ts,i to obtain the image-enhanced text feature F VT,i ; apply the text intermediate feature sequence F Ts,i to perform an attention enhancement operation on the image intermediate feature F Vs,i and the image description intermediate feature F Cs,i respectively to obtain the text-enhanced image feature F TV,i and the text-enhanced image description feature F TC,i :
[0049]
[0050] In Equation (1), Q V,i represents the query vector based on the image intermediate feature F Vs,i , K T,i represents the key vector based on the text intermediate feature sequence F Ts,i , V T,i represents the value vector based on the text intermediate feature sequence F Ts,i , d1 represents the dimension size of the query vector, key vector, and value vector, softmax represents the mathematical function that maps a real vector to a probability distribution, T represents the transpose, and W VT represents the weight matrix of the image-text attention operation, and Q T,i Denote the query vector based on the text intermediate feature sequence F Ts,i , K V,i Denote the key vector based on the image intermediate feature F Vs,i , V V,i Denote the value vector based on the image intermediate feature F Vs,i , W TV Denote the weight matrix for the text-image attention operation, K C,i Denote the key vector based on the intermediate feature F of the image description Cs,i ; V C,i Denote the value vector based on the intermediate feature F of the image description Cs,i , W TC Denote the weight matrix for the text-image description attention operation;
[0051] Step 4.2.3, Concatenate F VT,i F TV,i and F TC,i as a set of intermediate features, and apply the self-attention mechanism to further model the intermediate features. Then, use the fully connected layer and the average pooling layer to process the modeled features, and finally output the cross-modal feature F N,i .
[0052] The said step 5.1 includes:
[0053] Step 5.1.1, Use the API of the Baidu OpenAI platform to identify the objects and celebrities in the i-th news image I i , and use the entity linking tool TAGME to extract additional visual entities from the global image description to form the news visual entity set wherein, denotes the t-th entity in the i-th news picture I i , and T represents the number of entities in each news picture; Step 5.1.2, Use the pre-trained entity representation model TransE to link the text entity set and the visual entity set to the Freebase knowledge graph, so as to obtain the text entity embedding feature and the visual entity embedding feature
[0054] The said step 5.2 includes:
[0055] Step 5.2.1, Concatenate the global text intermediate feature F Gs,i and the image intermediate feature F Vs,i to form the global multi-modal news feature F M,i , and use another MLP layer to change the global multi-modal news feature F M,i The dimension is made consistent with the dimension of the entity embedding feature to obtain the globally changed feature F′ M,i ;
[0056] Step 5.2.2: Use F′ M,i to perform an adaptive hard attention operation on the visual entity embedding feature so as to calculate and obtain the globally changed feature F′ M,i for the corresponding attention score α and similarity matrix β 1,i of the visual entity embedding 1,i ;
[0057]
[0058]
[0059] In equations (2) and (3), Q M,i represents the query vector based on the globally multimodal news feature F M,i , represents the key vector based on the visual entity embedding feature and d2 represents the dimension of the query vector, key vector, and value vector;
[0060] Step 5.2.3: Calculate the threshold δ 1,i of the attention score α 1,i :
[0061]
[0062] Step 5.2.4: When the attention score corresponding to the t-th visual entity is less than the threshold δ 1,i , it is considered that the t-th visual entity is irrelevant to the news, and its corresponding similarity is set to --oo; otherwise, the original similarity remains unchanged, thereby obtaining the updated similarity matrix;
[0063] Step 5.2.5: Re-perform the softmax operation on the updated similarity matrix to obtain the updated attention score α 1,i , and perform subsequent attention operations using equation (5) to obtain the filtered visual entity feature F VE,i ;
[0064]
[0065] In equation (5), represents based on the visual entity embedding feature The value vector represents the weight matrix of the visual entity adaptive hard attention mechanism;
[0066] Step 5.2.6, according to the process of Step 4.2.1 - Step 4.2.5, utilize the global feature F′ M,i to perform the same adaptive hard attention operation on the text entity embedding so as to obtain the updated text entity similarity matrix β 2,i and the filtered text entity feature F TE,i .
[0067] The said Step 5.3 includes:
[0068] Step 5.3.1, according to the visual entity similarity matrix β 1,i and the text entity similarity matrix β 2,i , set the values corresponding to the entities with similarity of -∞ to 0, and set the values corresponding to the entities with similarity not -∞ to 1, so as to obtain the corresponding visual entity selection sequence η 1,i and the text entity selection sequence η 2,i ;
[0069] Step 5.3.2, perform dot product on the visual entity selection sequence η 1,i and the text entity selection sequence η 2,i to obtain the attention mask mask c,i ;
[0070] Step 5.3.3, apply the visual entity embedding and the attention mask mask c,i to perform attention operation on the text entity embedding so as to obtain the text entity knowledge interaction feature by using Equation (6)
[0071]
[0072] In Equation (6), represents the query vector based on the visual entity embedding , represents the key vector based on the text entity embedding , represents the value vector based on the text entity embedding , d3 represents the dimensions of the query vector, key vector and value vector; represents the weight matrix of the visual - text entity attention operation;
[0073] Apply the text entity embedding and the transpose of the attention mask Visual entity embedding Perform an attention operation to obtain the image entity knowledge interaction feature using Equation (7)
[0074]
[0075] In Equation (7), represents the query vector based on the text entity embedding , represents the key vector based on the visual entity embedding , represents the value vector based on the visual entity embedding , represents the weight matrix of the text-visual entity attention operation.
[0076] The two-layer MLP layer in Step 3.1.2 is optimized as follows to obtain the best prompt words;
[0077] Step a: Feed the image description set C i as additional information of the i-th news T i into the optimal multi-modal fake news detection model, and calculate the probability that the i-th news T i is fake;
[0078] Step b: Use Equation (1) to obtain the probability P i of the correct label y z predicted by the optimal multi-modal fake news detection model i :
[0079]
[0080] In Equation (1), P z (y i |z i , x i ) represents the fake news probability calculated by the optimal multi-modal fake news detection model, z i represents the prompt words to be learned {Z}1{Z}2...{Z} k ...{Z} n , x i represents all the news content of the i-th multi-modal news, including the news text T i , the news image I i and the news image description C i ;
[0081] Step c: Denote the probability gap between the correct label and the wrong label as Gap z (y i ) = P z (y i ) - (1 - P z (y i ));
[0082] When the optimal multi - modal fake news detection model makes a correct prediction, Gap z (y i ) is positive, otherwise it is negative;
[0083] For a correct prediction, multiply Gap z (y i ) by a relatively large value λ to represent the desirability of the prediction, where when y i = 1, the value of λ is λ2, and when y i = 0, the value of λ is λ1;
[0084] Step d: Use the DistilGPT - 2 model to generate multiple groups of prompt words z(x i ), and then use Equation (2) to calculate the reward R(x i ), y i ), z i ), z i ), z i ) for any group of prompt words z ∈ z(x
[0085]
[0086] Step e: Normalize it by calculating the mean and standard deviation of the rewards, and then use Equation (3) to calculate the reward z - score(x i ), y i ), z(x i )): i )):
[0087]
[0088] Step f: Based on the reward z - score(x i ), y i ), z(x i ), apply the soft Q - learning algorithm to update the parameters of the two - layer MLP layer to obtain an optimized two - layer MLP layer.
[0089] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the multi - modal fake news detection method, and the processor is configured to execute the program stored in the memory.
[0090] A computer-readable storage medium of the present invention stores a computer program thereon. The feature is that when the computer program is run by a processor, it executes the steps of the multi-modal fake news detection method.
[0091] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0092] 1. The present invention discovers the best prompt format through reinforcement learning and guides the BLIP model to generate image descriptions related to specific text entities, thereby fully mining the potential semantic information of images. At the same time, the present invention also uses the generated image descriptions to expand visual entities and extracts accurate news high-order knowledge semantic information through an adaptive hard attention mechanism, thereby improving the accuracy of multi-modal fake news detection.
[0093] 2. The present invention uses news text entities as the main body vocabulary of prompts and applies reinforcement learning to discover the best prompt format to generate text-guided image descriptions, thereby solving the problem of fully mining the potential semantic information of news images and further improving the accuracy of multi-modal fake news detection.
[0094] 3. The present invention links the entities contained in the news in the knowledge graph to obtain news knowledge representations, and inputs them into the knowledge semantic enhancement module to automatically select strongly relevant news entities and remove news-irrelevant noise entities, thereby obtaining accurate news high-order knowledge semantic features and improving the accuracy of multi-modal fake news detection. Description of the Drawings
[0095] Figure 1 is the overall flowchart of the method of the present invention;
[0096] Figure 2 is the structural diagram of the Rlprompt unit of the present invention. Detailed Embodiments
[0097] In this embodiment, as Figure 1 shown, a multi-modal fake news detection method based on multi-level semantic enhancement is carried out according to the following steps:
[0098] Step 1: Collection and preprocessing of multi-modal news data;
[0099] Extract the text content of each multi-modal news on the social media platform and its corresponding one image to obtain a news text set and a news image set wherein, T i represents the i-th news text; I i represents the i-th news image corresponding to T i ;
[0100] Set the i-th news image I i and its i-th news text T i The authenticity label of is denoted as y i and y i ∈ {0, 1}; thus constructing the training dataset Among them, N represents the number of news in the training dataset.
[0101] Step 2: Construct a multimodal fake news detection network, including: an image semantic enhancement module, a multimodal fusion module, and a knowledge semantic enhancement module.
[0102] Step 3: Construct an image semantic enhancement module, including: an RLprompt unit, a prompt statement construction unit, and a BLIP model;
[0103] Step 3.1: The RLprompt unit generates n learnable prompt words {Z}1{Z}2...{Z} k ...{Z} n ; In this embodiment, n = 2;
[0104] Step 3.1.1: The RLprompt unit uses the DistilGPT-2 language model to learn the learnable prompt words {Z}1{Z}2...{Z} k ...{Z} n ; Among them, {Z} k represents the k-th learnable prompt word; n represents the number of prompt words;
[0105] Step 3.1.2: The start symbol <start>In the input frozen DistilGPT-2 model, obtain the start symbol <start>Context embedding, add two fully connected layers to the frozen DistilGPT-2 model, and use the optimized two-layer MLP layer to re-encode the context embedding. After obtaining the encoded features, pass them to the classification head of the frozen DistilGPT-2 model to output the first prompt word {Z}1.
[0106] In the specific implementation, the two-layer MLP layer in step 3.1.2 is optimized according to the following steps to obtain the best prompt word, as Figure 2 shown;
[0107] Step a: Send the image description set C i as additional information of the i-th news T i into the optimal multi-modal fake news detection model, and calculate the probability that the i-th news T i is fake;
[0108] Step b: Use Equation (1) to obtain the probability P i of the correct label y z (y i ) predicted by the optimal multi-modal fake news detection model:
[0109]
[0110] In Equation (1), P z (y i |z i , x i ) represents the fake news probability calculated by the optimal multi-modal fake news detection model, z i represents the prompt words {Z}1{Z}2...{Z} k ...{Z} n , x i represents all the news content of the i-th multi-modal news, including the news text T i , the news image I i and the news image description C i .
[0111] Step c: Denote the probability gap between the correct label and the wrong label as Gap z (y i ) = P z (y i ) - (1 - P z (y i ));
[0112] When the optimal multi-modal fake news detection model predicts correctly, Gap z (y i ) is positive, otherwise it is negative;
[0113] For correct prediction, multiply Gap z (y i ) by a relatively large value λ to represent the desirability of the prediction. Among them, when y i = 1, take the value of λ2, and when y i = 0, take the value of λ1; in this embodiment, the value of λ1 is set to 180, and the value of λ2 is set to 200;
[0114] Step d: Use the DistilGPT-2 model to generate multiple groups of prompt words z(x i ), so as to calculate the reward R(x i , y i , z i , y i , z i ) of any group of prompt words z ∈ z(x i ) using Equation (2); in this embodiment, for x i , a total of 4 groups of prompt words are generated;
[0115]
[0116] Step e: Normalize it by calculating the mean and standard deviation of the rewards, so as to calculate the reward z-score(x i ), y i , z(x i )) of multiple groups of prompt words z(x i ) using Equation (3):
[0117]
[0118] Step f: Based on the reward z-score(x i , y i , z(x i ), apply the soft Q-learning algorithm to update the parameters of the two-layer MLP layer to obtain the optimized two-layer MLP layer.
[0119] Step 3.1.3: Based on Step 2.1.2, input the first prompt word {Z}1 into the frozen DistilGPT-2 model in an autoregressive generation manner, so as to gradually generate the prompt words to be learned {Z}1{Z}2...{Z} k ...{Z} n .
[0120] Step 3.2: Processing of the prompt statement construction unit:
[0121] Step 3.2.1: Use the entity linking tool TAGME to perform entity recognition on the i-th news text T i to obtain the text entity set Among them, Denote the j-th text entity in the i-th news text T i where M represents the number of text entities in each news text; in this embodiment, M = 6
[0122] Step 3.2.2, construct an interactive prompt statement based on two text entities in the i-th news text T i as the prompt main vocabulary Denote the j-th entity in the i-th news text T i and represents a connector
[0123] Construct a local prompt statement based on a single text entity in the i-th news text T i as the prompt main vocabulary
[0124] Step 3.3, the BLIP model generates an image description set for the i-th multimodal news
[0125] Send the i-th news image I i into the BLIP model to obtain a global image description representing the global semantic information of the image
[0126] Send the i-th news image I i and the interactive prompt statement P con,i into the BLIP model to guide the generation of an interactive image description of the news representing the semantic interaction information between two text entities in the image
[0127] Send the i-th news image I i and the local prompt statement P loc,i into the BLIP model to guide the generation of a local image description of the news representing the local semantic information of a single text entity in the image
[0128] Step Four, construct a multimodal fusion module, including: a feature extraction unit, a cross-modal feature enhancement unit
[0129] Step 4.1, the feature extraction unit is used to extract initial features of different modalities of the multimodal news
[0130] Step 4.1.1, use the pre-trained BERT model to perform feature extraction on the i-th news text T i to obtain the feature sequence F i of the i-th news text T T,i =[f 1.i , f 2.i ,..., f l.i , ..., f L.i ], where fl i Represents the i-th news text T i The text feature of the lth word level in ; in this embodiment, the size of L is set to 96, and sentences with a length less than 16 will be padded with zero vectors to a length of 96.
[0131] Use the long short-term memory network LSTM to train the feature sequence F T,i Further feature extraction is performed, and the hidden state feature of the last step output of the long short-term memory network LSTM is taken as the global feature F of the i-th news text Ti G,i ;
[0132] Step 4.1.2: Use the visual feature encoder in the frozen BLIP model to decode the i-th news image I i Perform feature extraction to obtain image features F V,i ;
[0133] Step 4.1.3: Use the text feature encoder in the frozen BLIP model to describe the image set C i Perform feature extraction to obtain image description feature F C,i ;
[0134] Step 4.1.4: Use a shared weight MLP layer to transform the feature sequence F T,i and the global feature F a,i dimension, so that it is consistent with the image feature F V,i and image description feature F C,i The dimension is consistent, and the characteristic sequence F′ with dimension change is obtained T,i and the global feature F′ G,i .
[0135] Step 4.2: The cross-modal feature enhancement unit is used to process the initial features of different modalities and output the cross-modal feature F N,i ;
[0136] Step 4.2.1. Construct an MLP layer consisting of an m×n linear layer and a ReLU activation function layer. T,i , F′ G,i 、F V,i 、F C,i Enter the shared weights as W shared After processing in the MLP layer, the i-th news text T is obtained. i The text intermediate feature sequence F Ts,i , the i-th news text T i The global text intermediate feature F Gs,i , the i-th news image I i The intermediate feature F of the image Vs,i , the news image description C i The intermediate feature F of the image description Cs,i .
[0137] Step 4.2.2: According to Equation (1), apply the news image intermediate feature F Vs,i to the text intermediate feature sequence F Ts,i to perform an attention operation, thereby obtaining the text feature F with enhanced image VT,i ; Apply the text intermediate feature sequence F Ts,i to the image intermediate feature F Vs,i and the intermediate feature F of the image description Cs,i respectively to perform an attention enhancement operation, thereby obtaining the image feature F with enhanced text TV,i and the image description feature F with enhanced text TC,i :
[0138]
[0139] In Equation (1), Q V,i represents the query vector based on the image intermediate feature F Vs,i , K T,i represents the key vector based on the text intermediate feature sequence F Ts,i , V T,i represents the value vector based on the text intermediate feature sequence F Ts,i , d1 represents the dimension size of the query vector, key vector, and value vector. In this embodiment, the size of d1 is set to 256, softmax represents the mathematical function that maps a real number vector to a probability distribution, T represents the transpose, W VT represents the weight matrix of the image-text attention operation, Q T,i represents the query vector based on the text intermediate feature sequence F Ts,i , K V,i represents the key vector based on the image intermediate feature F Vs,i , V V,i represents the value vector based on the image intermediate feature F Vs,i , W TV represents the weight matrix of the text-image attention operation, K C,i represents the key vector based on the intermediate feature F of the image description Cs,i ; V C,i represents the value vector based on the intermediate feature F of the image description Cs,i , W TC represents the weight matrix of the text-image description attention operation.
[0140] Step 4.2.3: Combine F VT,i F TV,i and F TC,i After being concatenated as a group of intermediate features, the self-attention mechanism is applied to further model the intermediate features. Then, the fully connected layer and the average pooling layer are used to process the modeled features, and finally, the cross-modal feature F is output. N,i 。
[0141] Step 5: Construct a knowledge semantic enhancement module, including an entity linking unit, an adaptive hard attention mechanism unit, a cross-modal knowledge interaction unit, and a knowledge fusion unit;
[0142] Step 5.1: The entity linking unit is used to extract news entities and link them to the knowledge graph;
[0143] Step 5.1.1: Use the API of the Baidu OpenAI platform to identify the objects and celebrities in the i-th news image I i and use the entity linking tool TAGME to extract additional visual entities from the global image description to form a news visual entity set where, represents the t-th entity in the i-th news picture I i and T represents the number of entities in each news picture; in this embodiment, T = 6.
[0144] Step 5.1.2: Use the pre-trained entity representation model TransE to link the text entity set and the visual entity set to the Freebase knowledge graph, so as to obtain the text entity embedding features and the visual entity embedding features
[0145] Step 5.2: The adaptive hard attention mechanism unit is used to process the entity embedding features to obtain filtered entity features;
[0146] Step 5.2.1: Concatenate the global text intermediate feature F Gs,i and the image intermediate feature F Vs,i to form the global multi-modal news feature F M,i , and use another MLP layer to change the dimension of the global multi-modal news feature F M,i to make it consistent with the dimension of the entity embedding features, and obtain the dimension-changed global multi-modal news feature F' M,i 。
[0147] Step 5.2.2: Use F' M,i to perform an adaptive hard attention operation on the visual entity embedding features , so as to calculate and obtain the global feature F' M,i for the visual entity embedding The corresponding attention score α 1,i and the similarity matrix β 1,i ;
[0148]
[0149]
[0150] In formulas (2) and (3), Q M,i represents the query vector based on the global multi-modal news feature F M,i . represents the key vector based on the visual entity embedding feature . d2 represents the dimension of the query vector, the key vector, and the value vector. In this embodiment, the size of d2 is set to 50.
[0151] Step 5.2.3: Calculate the attention score α using formula (4) 1,i for the threshold δ 1,i :
[0152]
[0153] Step 5.2.4: When the attention score corresponding to the t-th visual entity is less than the threshold δ 1,i , it is considered that the t-th visual entity is irrelevant to the news, and its corresponding similarity is set to -∞; otherwise, the original similarity remains unchanged, thereby obtaining the updated similarity matrix.
[0154] Step 5.2.5: Re-perform the softmax operation on the updated similarity matrix to obtain the updated attention score α 1,i , and perform subsequent attention operations using formula (5) to obtain the filtered visual entity feature F VE,i ;
[0155]
[0156] In formula (5), represents the value vector based on the visual entity embedding feature , and represents the weight matrix of the visual entity adaptive hard attention mechanism.
[0157] Step 5.2.6: According to the process of steps 4.2.1 - 4.2.5, use the global feature F′ M,i to perform text entity embedding Perform the same adaptive hard attention operation to obtain the updated text entity similarity matrix β 2,i and the filtered text entity feature F TE,i .
[0158] Step 5.3. The cross-modal knowledge interaction unit is used to enhance the features of the entity embedding to obtain the entity knowledge interaction feature;
[0159] Step 5.3.1. According to the visual entity similarity matrix β 1,i and the text entity similarity matrix β 2,i , set the values corresponding to the entities with a similarity of -∞ to 0, and set the values corresponding to the entities with a similarity not equal to -∞ to 1, so as to obtain the corresponding visual entity selection sequence η 1,i and the text entity selection sequence η 2,i .
[0160] Step 5.3.2. Perform a dot product on the visual entity selection sequence η 1,i and the text entity selection sequence η 2,i to obtain the attention mask mask c,i ;
[0161] Step 5.3.3. Apply the visual entity embedding and the attention mask mask c,i to perform an attention operation on the text entity embedding so as to obtain the text entity knowledge interaction feature using Equation (6)
[0162]
[0163] In Equation (6), represents the query vector based on the visual entity embedding , represents the key vector based on the text entity embedding , represents the value vector based on the text entity embedding , d3 represents the dimensions of the query vector, key vector, and value vector. In this embodiment, the size of d3 is set to 50, represents the weight matrix of the visual-text entity attention operation.
[0164] Apply the text entity embedding and the transpose of the attention mask to perform an attention operation on the visual entity embedding so as to obtain the image entity knowledge interaction feature using Equation (7)
[0165]
[0166] In formula (f7), represents the query vector based on the text entity embedding . represents the key vector based on the visual entity embedding . represents the value vector based on the visual entity embedding . represents the weight matrix of the text-visual entity attention operation.
[0167] Step 5.4. The knowledge fusion unit is used to fuse the above entity knowledge features to obtain the news background knowledge feature F E , i ;
[0168] The four enhanced entity knowledge features F VE,i , F TE,i , and are concatenated as a group of entity embeddings, and then the self-attention mechanism is applied to further model the entity embeddings, and the fully connected layer and the average pooling layer are used to process the modeled entity embeddings, and finally the news background knowledge feature F E,i is output.
[0169] Step Six. Optimization of the multi-modal fake news detection network:
[0170] Step 6.1. Use formula (8) to predict the probability that the i-th multi-modal news is fake news
[0171]
[0172] In formula (8), σ represents the sigmoid activation function, W c represents the weight matrix of the classifier, and b c represents the bias vector;
[0173] Step 6.2. Use formula (9) to construct the cross-entropy loss function
[0174]
[0175] Step 6.3. Based on the training dataset X, use the Adam optimization strategy to train the multi-modal fake news detection network until the total loss function of the network converges, so as to obtain the optimal multi-modal fake news detection model for predicting any multi-modal news.
[0176] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0177] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.< / start> < / start> < / start> < / start>
Claims
1. A multi-modal fake news detection method based on multi-level semantic enhancement, characterized in that, It is carried out according to the following steps: Step 1: Collection and preprocessing of multimodal news data; Extract the text content of each multimodal news on the social media platform and one corresponding image to obtain a news text set and a news image set where T i represents the i-th news text; I i represents the i-th news image corresponding to T i ; Set the authenticity label of the i-th news image I i and its i-th news text T i , denoted as y i , and y i ∈{0,1}; thus constructing a training dataset where N represents the number of news in the training dataset; Step 2: Construct a multimodal fake news detection network, including: an image semantic enhancement module, a multimodal fusion module, and a knowledge semantic enhancement module; Step 3: Construct an image semantic enhancement module, including: an RLprompt unit, a prompt statement construction unit, and a BLIP model; Step 3.1, the RLprompt unit generates n to-be-learned prompt words \{Z\}_1 \{Z\}_2 … \{Z\} k … \{Z\} n ; Step 3.2: Processing of the prompt statement construction unit: Step 3.2.1: Use the entity linking tool TAGME to perform entity recognition on the i-th news text T i to obtain a text entity set where represents the j-th text entity in the i-th news text T i and M represents the number of text entities in each news text; Step 3.2.2, construct an interactive prompt statement based on the two text entities in the i-th news text T i as the prompt body vocabulary denote the j'-th entity in the i-th news text T; and denote the connector i Construct a local prompt statement based on a single text entity in the i-th news text T i as the prompt body vocabulary Step 3.3, the BLIP model generates the image description set of the i-th multimodal news Send the i-th news image I i into the BLIP model to obtain the global image description Send the i-th news image I i and the interactive prompt statement P con,i into the BLIP model to guide the generation of an interactive image description of the news Feed the i-th news image I i and the local prompt statement P loc,i into the BLIP model to guide the generation of a local image description of the news Step 4: Construct a multimodal fusion module, including: a feature extraction unit and a cross-modal feature enhancement unit; Step 4.1: The feature extraction unit is used to extract the initial features of different modalities of multimodal news; Step 4.
2. The cross-modal feature enhancement unit is used to process the initial features of different modalities and output the cross-modal feature F N,i ; Step 5: Construct a knowledge semantic enhancement module, including an entity linking unit, an adaptive hard attention mechanism unit, a cross-modal knowledge interaction unit, and a knowledge fusion unit; Step 5.1: The entity linking unit is used to extract news entities and link them to the knowledge graph; Step 5.2: The adaptive hard attention mechanism unit is used to process the entity embedding features to obtain filtered entity features; Step 5.3: The cross-modal knowledge interaction unit is used to enhance the features of the entity embedding features to obtain entity knowledge interaction features; Step 5.
4. The knowledge fusion unit concatenates the filtered visual entity feature F VE,i , the filtered text entity feature F TE,i , the text entity knowledge interaction feature and the image entity knowledge interaction feature as a group of entity embeddings, then applies the self-attention mechanism to further model the entity embeddings, and uses the fully connected layer and the average pooling layer to process the modeled entity embeddings, and finally outputs the news background knowledge feature F E,i ; Step 6: Optimization of the multimodal fake news detection network: Step 6.
1. Predict the probability that the i-th multimodal news is fake news using Equation (8). In Equation (8), σ represents the sigmoid activation function, W c represents the weight matrix of the classifier, and b c represents the bias vector; F' M,i represents the global multi-modal news feature after dimensionality change; Step 6.
2. Construct the cross-entropy loss function using Equation (9). Step 6.3: Based on the training dataset X, use the Adam optimization strategy to train the multimodal fake news detection network until the total network loss function converges, thereby obtaining an optimal multimodal fake news detection model for predicting any multimodal news.
2. The multi-modal fake news detection method based on multi-level semantic enhancement according to claim 1, characterized in that, The said Step 3.1 includes: Step 3.1.1, the RLprompt unit uses the DistilGPT-2 language model to learn the to-be-learned prompt words {Z}1{Z}2…{Z} k …{Z} n ; where {Z} k represents the k-th to-be-learned prompt word; n represents the number of prompt words; Step 3.1.2, place the start symbol <start>In the frozen DistilGPT-2 model, obtain the start symbol <start>The context embedding, and the context embedding is re-encoded by two layers of MLP layers, and after obtaining the encoded features, they are passed to the classification head of the frozen DistilGPT-2 model, so as to output the first prompt word {Z}1;< / start> < / start> Step 3.1.3: Based on Step 2.1.2, input the first prompt word {Z}1 into the frozen DistilGPT-2 model in an autoregressive generation manner, so as to gradually generate the prompt words to be learned {Z}1{Z}2…{Z} k …{Z} n 。 3. The multimodal fake news detection method based on multi-level semantic enhancement according to claim 2, wherein The said Step 4.1 includes: Step 4.1.1: Use the pre-trained BERT model to extract features from the i-th news text T i to obtain the feature sequence F i of the i-th news text T T,i = [f 1.i , f 2.i ,..., f l.i ,..., f L.i , where f l.i represents the text feature at the l-th word level in the i-th news text T i ; Using the long short-term memory network LSTM to perform further feature extraction on the feature sequence F T,i and taking the hidden state feature output at the last step of the long short-term memory network LSTM as the global feature F i of the i-th news text T G,i ; Step 4.1.2: Use the visual feature encoder in the frozen BLIP model to extract features from the i-th news image I i to obtain image features F V,i ; Step 4.1.3: Use the text feature encoder in the frozen BLIP model to perform feature extraction on the image description set C i to obtain the image description feature F C,i ; Step 4.1.4: Use an MLP layer with shared weights to change the feature sequence F T,i and the global feature F G,i in dimension to make it consistent with the image feature F V,i and the image description feature F C,i in dimension, obtaining the feature sequence F' with changed dimension T,i and the global feature F' G,i .
4. The multi-modal fake news detection method based on multi-level semantic enhancement according to claim 3, characterized in that The said Step 4.2 includes: Step 4.2.1: Construct an MLP layer consisting of a single-layer linear layer of size m×n and a ReLU activation function layer, and input F' T,i and F' G,i , F V,i , and F C,i into the MLP layer with shared weight W shared for processing. After processing, obtain the intermediate text feature sequence F i of the i-th news text T Ts,i , the global intermediate text feature F i of the i-th news text T Gs,i , the intermediate image feature F i of the i-th news image I Vs,i , and the intermediate image description feature F i of the news image description C Cs,i ; Step 4.2.
2. According to formula (1), apply the intermediate feature F of the news image Vs,i to the intermediate feature sequence F of the text Ts,i to perform an attention operation, so as to obtain the text feature F with enhanced image VT,i ; Apply the intermediate feature sequence F of the text Ts,i to perform an attention enhancement operation on the intermediate feature F of the image Vs,i and the intermediate feature F of the image description Cs,i respectively, so as to obtain the image feature F with enhanced text TV,i and the image description feature F with enhanced text TC,i : In formula (1), Q V,i represents the query vector based on the intermediate image feature F Vs,i ; K T,i represents the key vector based on the intermediate text feature sequence F Ts,i ; V T,i represents the value vector based on the intermediate text feature sequence F Ts,i ; d1 represents the dimensionality size of the query vector, key vector, and value vector; softmax represents the mathematical function that maps a real-valued vector to a probability distribution; T represents the transpose; W VT represents the weight matrix of the image-text attention operation; Q T,i represents the query vector based on the intermediate text feature sequence F Ts,i ; K V,i represents the key vector based on the intermediate image feature F Vs,i ; V V,i represents the value vector based on the intermediate image feature F Vs,i ; W TV represents the weight matrix of the text-image attention operation; K C,i represents the key vector based on the intermediate image caption feature F Cs,i ; V C,i represents the value vector based on the intermediate image caption feature F Cs,i ; W TC represents the weight matrix of the text-image caption attention operation; Step 4.2.
3. Combine F VT,i F TV,i and F TC,i in series as a group of intermediate features, and apply the self-attention mechanism to further model the intermediate features. Then, use the fully connected layer and the average pooling layer to process the modeled features, and finally output the cross-modal feature F N,i .
5. The multi-modal fake news detection method based on multi-level semantic enhancement according to claim 4, characterized in that The said Step 5.1 includes: Step 5.1.1: Use the API of Baidu OpenAI platform to identify the objects and celebrities in the i-th news image I i and use the entity linking tool TAGME to extract additional visual entities from the global image description to form a news visual entity set wherein represents the t-th entity in the i-th news picture I i and T represents the number of entities in each news picture; Step 5.1.2: Use the pre-trained entity representation model TransE to link the text entity set and the visual entity set to the Freebase knowledge graph, so as to obtain the text entity embedding features and the visual entity embedding features 6. The multi-modal fake news detection method based on multi-level semantic enhancement according to claim 5, characterized in that, The said Step 5.2 includes: Step 5.2.1: Concatenate the global text intermediate feature F Gs,i and the image intermediate feature F Vs,i to form the global multimodal news feature F M,i , and use another MLP layer to change the dimension of the global multimodal news feature F M,i to make it consistent with the dimension of the entity embedding feature, obtaining the globally feature F' M,i with changed dimension; Step 5.2.
2. Use F' M,i to perform an adaptive hard attention operation on the visual entity embedding features so as to calculate the global feature F' using equations (2) and (3) M,i for the corresponding attention scores α and similarity matrix β 1,i of the visual entity embedding 1,i ; In formulas (2) and (3), Q M,i represents the query vector based on the global multi-modal news feature F M,i , represents the key vector based on the visual entity embedding feature , and d2 represents the dimensions of the query vector, key vector, and value vector; Step 5.2.
3. Calculate the attention score α using Equation (4) 1,i threshold δ 1,i : Step 5.2.4: When the attention score corresponding to the t-th visual entity is less than the threshold δ 1,i , it is considered that the t-th visual entity is irrelevant to the news, and its corresponding similarity is set to -∞; otherwise, the original similarity remains unchanged, so as to obtain the updated similarity matrix. Step 5.2.5: Perform the softmax operation on the updated similarity matrix again to obtain the updated attention score α 1,i , and perform subsequent attention operations using Equation (5) to obtain the filtered visual entity feature F VE,i ; In formula (5), represents the value vector based on the visual entity embedding features , and represents the weight matrix of the visual entity adaptive hard attention mechanism; Step 5.2.
6. According to the process of Step 4.2.1 - Step 4.2.5, use the global feature F' M,i to perform the embedding of text entities to perform the same adaptive hard attention operation, so as to obtain the updated text entity similarity matrix β 2,i and the filtered text entity feature F TE,i .
7. The multi-modal fake news detection method based on multi-level semantic enhancement according to claim 6, characterized in that The said Step 5.3 includes: Step 5.3.1: According to the visual entity similarity matrix β 1,i and the text entity similarity matrix β 2,i , set the values corresponding to the entities with a similarity of -∞ to 0, and the values corresponding to the entities with a similarity not equal to -∞ will be set to 1, so as to obtain the corresponding visual entity selection sequence η 1,i and the text entity selection sequence η 2,i ; Step 5.3.2, perform a dot product on the visual entity selection sequence η 1,i and the text entity selection sequence η 2,i to obtain an attention mask mask c,i ; Step 5.3.3, Apply visual entity embedding and the attention mask mask c,i to perform an attention operation on the text entity embedding so as to obtain the text entity knowledge interaction feature by using Equation (6) In formula (6), represents the query vector based on the visual entity embedding , represents the key vector based on the text entity embedding , represents the value vector based on the text entity embedding , and d3 represents the dimensions of the query vector, key vector, and value vector; represents the weight matrix of the visual-text entity attention operation; Apply text entity embedding and the transpose of the attention mask to perform an attention operation on the visual entity embedding so as to obtain the image entity knowledge interaction feature by using Equation (7) In formula (7), represents the query vector based on the text entity embedding ; represents the key vector based on the visual entity embedding ; represents the value vector based on the visual entity embedding ; represents the weight matrix of the text-visual entity attention operation.
8. The multi-modal fake news detection method based on multi-level semantic enhancement according to claim 2, wherein The two layers of MLP layers in the said Step 3.1.2 are optimized according to the following steps to obtain the best prompt word; Step a: Feed the image description set C i as additional information of the i-th news T i into the optimal multi-modal fake news detection model, and calculate the probability that the i-th news T i is fake; Step b. Obtain the probability P of the correct label y predicted by the optimal multi-modal fake news detection model using Equation (1). i of z (y i ): In formula (1), P z (y i |z i ,x i ) represents the probability of fake news calculated by the optimal multi-modal fake news detection model. z i represents the prompt words to be learned {Z}1{Z}2…{Z} k …{Z} n , x i represents all the news content of the i-th multi-modal news, including the news text T i , the news image I i and the news image description C i ; Step c, denote the probability gap between the correct label and the wrong label as Gap z (y i )=P z (y i )-(1-P z (y i )); When the optimal multimodal fake news detection model makes a correct prediction, Gap z (y i ) is positive, otherwise it is negative; For a correct prediction, multiply Gap z (y i ) by a large value λ to represent the desirability of the prediction, where when y i = 1, take the value λ2, and when y i = 0, take the value λ1; Step d: Generate multiple sets of prompts z(x i ) using the DistilGPT-2 model, and then calculate the reward R(x i , y i , z i , y i , z i ) for any set of prompts z ∈ z(x i ) using equation (2); Step e: Normalize it by calculating the mean and standard deviation of the rewards, so as to calculate the reward z-score of multiple groups of prompt words z(x i ) using Equation (3) as z-score(x i ,y i ,z(x i )): Step f, base reward z-score(x i ,y i ,z(x i )) and apply the soft Q-learning algorithm to update the parameters of the two-layer MLP layer to obtain an optimized two-layer MLP layer.
9. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program for supporting the processor to execute any one of the multimodal fake news detection methods in Claims 1-8, and the processor is configured to execute the program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of any one of the multimodal fake news detection methods in Claims 1-8.
Citation Information
Patent Citations
Knowledge graph guided false news detection method
CN111061843A
Multi-modal false news detection method and system
CN116340887A