Multi-modal false news detection method based on attention mechanism
Through the multimodal fake news detection method based on attention mechanism, combined with deep learning and news classification model, the problem of multimodal information correlation and singlemodal uniqueness in the existing technology is solved, and efficient, accurate and credible fake news detection is achieved.
Patent Information
- Application Number
- CN202510001850.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-30
AI Technical Summary
The existing fake news detection methods fail to effectively consider the correlation between multimodal information and the uniqueness of singlemodal information, resulting in insufficient accuracy and credibility of the detection results, and low cost and efficiency.
The multimodal fake news detection method based on attention mechanism is adopted, and the news content is represented as feature vectors through deep learning technology. Combined with the news classification model, the interaction control vector and attention weight vector are used to capture the correlation between text and pictures and the uniqueness of single-modal data.
It improves the accuracy and credibility of fake news detection, reduces the detection cost and time, and achieves efficient classification of news authenticity and falsehood.
Smart Images

Figure CN120067967A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal learning, and particularly to a multimodal fake news detection method based on an attention mechanism. Background Art
[0002] Fake news refers to false or misleading information, usually deliberately fabricated, spread and disseminated by individuals, organizations or institutions, aiming to deceive the public, influence public opinion or obtain certain benefits, which has brought great harm to society. With hundreds of millions of pieces of information being posted on social media platforms every day, fake news can spread rapidly in a short period of time. Therefore, efficiency and cost are important criteria for fake news detection methods. In order to effectively curb the spread of fake news, reduce the harm of fake news to society, and improve the public's trust in social media, credibility and accuracy are also important criteria for measuring fake news detection methods. Traditional fake news detection methods mainly rely on manual fact-checking, which evaluates the authenticity of news by extracting the content in the news and comparing it with known facts. Manual fact-checking mainly relies on domain experts or the general public to verify news facts. The cost of domain expert verification is high, the efficiency is low, and the timeliness is poor. Moreover, the general public verification personnel are easily affected by the surrounding environment, and the results obtained often have low credibility and accuracy. Therefore, the traditional manual fact-checking method has great limitations and is difficult to meet the increasingly complex fake news detection requirements.
[0003] With the development of artificial intelligence technology, the methods and forms of fake news detection have been rejuvenated. By calculating the probability of the truth and accuracy of news texts, an intelligent fake news detection system can assess the risks of fake news, automatically detect and filter fake news, and reduce the time and labor costs of fact-checking. At present, scholars have achieved certain results in the research of fake news detection methods, but the relevance of multi-modal information fusion and the uniqueness of single-modal information have not been considered. Literature 1 (Event-Radar: Event-driven Multi-View Learning for Multimodal Fake News Detection, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5809–5821) proposed an event-based fake news detection method based on multi-view learning. This method uses the ability of a graphical structure to capture the interaction between events and parameters, constructs an event graph containing the subject-predicate logic of multi-modal entities to capture event-level multi-modal inconsistencies, and calculates the credibility of each view as a clue to learn a comprehensive and reliable news content representation. Literature 2 (QMFND: A quantum multimodal fusion-based fake news detection model for social media, Information Fusion Volume 104, April 2024, 102172) proposed a fake news detection method based on quantum multi-modal fusion. This method combines quantum coding and quantum convolutional neural networks. Multi-modal features are encoded into a variational quantum circuit through amplitude coding, which has excellent expressive and entanglement capabilities, good robustness to quantum noise, and can also alleviate the "barren plateau" phenomenon. However, these methods do not simultaneously consider the relevance between data during multi-modal information fusion and the uniqueness of single-modal data itself.
[0004] Existing fake news detection methods generally include steps such as text and image preprocessing, text feature extraction, image feature extraction, model selection, training a classification model, model evaluation, and application. However, there is a hidden correlation between text and images, and they also have certain uniquenesses of their own. The correlation between multimodal content and the uniqueness of unimodal content itself have a crucial impact on the news detection results. Traditional deep learning methods, while maintaining the uniqueness of unimodal content itself, simply concatenate multimodal content, or ignore the unique attributes of unimodal data itself while mining the correlation between multimodal content, without considering both the uniqueness of unimodal data itself and the correlation existing between multimodal data, and are unable to effectively improve the accuracy and credibility of the detection results while meeting high efficiency and low cost. Summary of the Invention
[0005] The object of the present invention is to provide a fake news detection method with high accuracy and credibility of detection results, high detection efficiency, and low cost.
[0006] The technical solution for achieving the object of the present invention is: a multimodal fake news detection method based on an attention mechanism, comprising the following steps:
[0007] Step 1, read the relevant content of the news, including the text content S′ and the image content V′;
[0008] Step 2, use deep learning technology to represent the news content as corresponding feature vectors;
[0009] Step 3, according to the feature vectors of the news content, complete the true / false classification of the news in combination with a news classification model.
[0010] Further, the use of deep learning technology in Step 2 to represent the news content as corresponding feature vectors is specifically as follows:
[0011] Step 2.1, convert each sentence in the news text into a corresponding sentence representation vector, initialize the correlation matrix between sentences, calculate the attention weight matrix between sentences according to the correlation matrix, and calculate the text representation vector in combination with the attention weight matrix between sentences and the value matrix of the text;
[0012] Step 2.2, obtain the features of the images, and merge all the image features to obtain an image representation vector, specifically as follows:
[0013] Use the VGG-19 neural network model to extract the image features of each image, and use a recurrent neural network model to merge all the image features to obtain an image representation vector P;
[0014] Step 2.3, map the text and image representation vectors to the same space to obtain corresponding text and image mapping vectors;
[0015] Step 2.4: Generate text and image interaction control vectors and alternative text and image feature vectors based on the text and image representation vectors;
[0016] Step 2.5: Use the text and image mapping vectors and the text and image interaction control vectors to obtain the preliminary text and image feature vectors;
[0017] Step 2.6: Calculate the text and image attention weight vectors based on the preliminary text and image feature vectors;
[0018] Step 2.7: Combine the text and image mapping vectors and the preliminary text and image feature vectors to calculate the similarity between the preliminary text and image feature vectors;
[0019] Step 2.8: Combine the text and image attention weight vectors, the similarity between the preliminary text and image feature vectors, and the alternative text and image feature vectors to obtain the text feature vector and the image feature vector;
[0020] Step 2.9: Eliminate the redundant information between the text feature vector and the image feature vector;
[0021] Step 2.10: Concatenate the text feature vector and the image feature vector as the feature vector of the news content, and output the feature vector of the news content.
[0022] Further, in step 2.1, converting each sentence in the news text into a corresponding sentence representation vector, initializing the association matrix between sentences, calculating the attention weight matrix between sentences according to the association matrix, and combining the attention weight matrix between sentences and the value matrix of the text to calculate the text representation vector are specifically as follows:
[0023] Step 2.1.1: Input each sentence in the news text into the open-source pre-trained BERT language model to obtain the set of representation vectors of all sentences in the news text S = {S 1 , S 2 , …, S n};
[0024] Step 2.1.2: For the representation vector of each sentence in the news text content, generate the corresponding query Query, key Key, and value Value vectors respectively, and combine the query Query, key Key, and value Value vectors generated by all sentences into the query matrix Q, the key matrix K, and the value matrix V respectively:
[0025] Q = S * W Q (1)
[0026] K = S * W K (2)
[0027] V = S * W V (3)
[0028] Among them, W Q , W K , W V is the weight matrix obtained through learning. W Q represents the weight of the query matrix Q generation function, and W K represents the weight of the key matrix K generation function, and W V represents the weight of the value matrix V generation function;
[0029] Step 2.1.3: Multiply the query matrix Q by the key matrix K to obtain the inter-sentence attention weight matrix Att. The calculation formula is:
[0030] Att = QK T (4)
[0031] Step 2.1.4: Normalize the obtained inter-sentence attention weight matrix Att and multiply it by the value matrix V of the text to obtain the representation vector T of the news text. The calculation formula is:
[0032]
[0033] Among them, d k is the scaling factor used to normalize the calculation result and ensure the stability of the gradient.
[0034] Furthermore, mapping the text and image representation vectors to the same space to obtain the corresponding text and image mapping vectors in step 2.3 is specifically as follows:
[0035] Since the calculation methods of the text representation vector and the image representation vector are different, resulting in differences in the spatial dimensions of the text and image representation vectors, it is necessary to map them to the same spatial dimension to obtain the text mapping vector e T and the image mapping vector e P which is:
[0036]
[0037] Among them is the weight that can be trained, represents the weight of the image mapping function, represents the weight of the text mapping function; is the bias that can be trained, represents the bias of the image mapping function, represents the bias of the text mapping function.
[0038] Further, the generation of the text and image interaction control vector and the alternative text and image feature vectors based on the text and image representation vectors in step 2.4 is specifically as follows:
[0039] The original text representation vector T and the image representation vector P are linearly transformed through a fully connected neural network to generate the text interaction control vector i T and the image interaction control vector i P and the alternative text feature vector and the alternative image feature vector The formula is:
[0040]
[0041] where are trainable weights, represents the weight of the text interaction control vector generation function, represents the weight of the image interaction control vector generation function, represents the weight of the alternative text feature vector generation function, represents the weight of the alternative image feature vector generation function; are trainable biases, represents the bias of the text interaction control vector generation function, represents the bias of the image interaction control vector generation function, represents the bias of the alternative text feature vector generation function, represents the bias of the alternative image feature vector generation function.
[0042] Further, the obtaining of the text and image preliminary feature vectors by using the text and image mapping vectors and the text and image interaction control vectors in step 2.5 is specifically as follows:
[0043] Multiply the text mapping vector e T , the image mapping vector e P , the text interaction control vector i T and the image interaction control vector i P to obtain the text preliminary feature vector e′ T and the image preliminary feature vector e′ P . The formula is:
[0044] e′ T =i T *e T (12)
[0045] e′ P =i P *e P (13)
[0047] Further, calculating the text and picture attention weight vectors according to the preliminary text and picture feature vectors described in step 2.6 is specifically as follows:
[0048] According to the preliminary text and picture feature vectors, use the softmax function to calculate the text attention weight vector A T and the picture attention weight vector A P , and the calculation formula is:
[0049] A T = softmax(e′ T ) (14)
[0050] A P = softmax(e′ P ) (15)
[0052] Further, combining the text and picture mapping vectors and the preliminary text and picture feature vectors described in step 2.7 to calculate the similarity between the preliminary text and picture feature vectors is specifically as follows:
[0053] The similarity α between the preliminary text feature vectors T2P and the similarity α between the preliminary picture feature vectors P2T are the products of the text and picture mapping vectors and the preliminary text and picture feature vectors, and the calculation formula is:
[0054]
[0055] Combining the text and picture attention weight vectors, the similarity between the preliminary text and picture feature vectors, and the alternative text and picture feature vectors to obtain the text feature vector and the picture feature vector is specifically as follows:
[0056] Combining the text attention weight vector A T , the picture attention weight vector A P , the similarity α between the preliminary text feature vectors T2P , the similarity α between the preliminary picture feature vectors P2T , and the alternative text feature vector the alternative picture feature vector calculate the text feature vector C T and the picture feature vector C P , and the calculation formula is:
[0057]
[0058] Further, eliminating the redundant information between the text and picture feature vectors described in step 2.9 is specifically as follows:
[0059] Eliminate the text feature vector C T and the image feature vector C P of the redundant information, and the calculation formula is:
[0060] T′ = sigmoid(PW OT +b OT ) * tanh(C T ) (20)
[0061] P′ = sigmoid(PW OP +b OP ) * tanh(C P ) (21)
[0062] where W OT , W OP are trainable weights, W OT represents the weight of the text redundancy information elimination function, and W OP represents the weight of the image redundancy information elimination function; b OT , b OP are trainable biases, b OT represents the bias of the text redundancy information elimination function, and b OP represents the bias of the image redundancy information elimination function.
[0063] Furthermore, the news classification model described in step 3 is specifically as follows:
[0064] The news classification model uses a fully connected neural network layer as the classification function, uses the softmax function as the activation function, and the loss function uses cross-entropy L, and the formula is:
[0065] Y = softmax(HW Y +b Y ) (22)
[0066]
[0067] where W Y represents the weight of the classification function and is a trainable weight; b Y represents the bias of the classification function and is a trainable bias; M is the number of news categories, with a value of 2, that is, true news or false news, and y i and respectively represent the true value and the predicted value of the news being true or false.
[0068] Compared with the prior art, the significant advantages of the present invention are as follows: (1) When detecting news content with multimodal data, the uniqueness of unimodal data and the relevance of multimodal data are comprehensively considered. By creating an interactive control vector and calculating the attention weight vector, the unique attributes of the unimodal data itself are maintained, the efficiency of fake news detection is improved, and the cost of fake news detection is reduced; (2) The similarity between different modal data is calculated using the preliminary feature vector to capture the hidden relevance between them, improving the accuracy and credibility of the detection results. Description of the Drawings
[0069] Figure 1 It is a schematic flow chart of a multimodal fake news detection method based on the attention mechanism of the present invention. Detailed Embodiments
[0070] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0071] As Figure 1 shown, a multimodal fake news detection method based on the attention mechanism of the present invention includes the following steps:
[0072] Step 1: Read the relevant content of the news, including the text content S′ and the picture content V′;
[0073] Step 2: Use deep learning technology to represent the news content as corresponding feature vectors;
[0074] Step 3: According to the feature vectors of the news content, combine the news classification model to complete the true / false classification of the news.
[0075] As a specific example, in step 2, deep learning technology is used to represent the news content as corresponding feature vectors, specifically as follows:
[0076] Step 2.1: Convert each sentence in the news text into a corresponding sentence representation vector, initialize the association matrix between sentences, calculate the attention weight matrix between sentences according to the association matrix, and calculate the text representation vector by combining the attention weight matrix between sentences and the sentence value association matrix.
[0077] As a specific example, step 2.1 is specifically as follows:
[0078] Step 2.1.1: Input each sentence in the news text into the open-source pre-trained BERT language model to obtain the set of representation vectors of all sentences in the news text S = {S 1 , S 2 , …, S n};
[0079] Step 2.1.2: For the representation vectors of each sentence in the news text content, generate the corresponding query Query, key Key, and value Value vectors respectively. Combine the query Query, key Key, and value Value vectors generated from all sentences into a query matrix Q, a key matrix K, and a value matrix V respectively:
[0080] Q = S * W Q (1)
[0081] K = S * W K (2)
[0082] V = S * W V (3)
[0083] where W Q , W K , W V is the weight matrix obtained through learning; W Q represents the weight of the query matrix Q generation function, W K represents the weight of the key matrix K generation function, W V represents the weight of the value matrix V generation function;
[0084] Step 2.1.3: The attention weight matrix Att between sentences is the similarity between different sentences. The higher the similarity between sentences, the greater their corresponding weight coefficients. Multiply the query matrix Q by the key matrix K to obtain the attention weight matrix Att between sentences. The calculation formula is:
[0085] Att = QK T (4)
[0086] Step 2.1.4: Multiply the obtained attention weight matrix Att between sentences after normalization by the value matrix V of the text to obtain the representation vector T of the news text. The calculation formula is:
[0087]
[0088] where d k is a scaling factor used to normalize the calculation result and ensure the stability of the gradient.
[0089] Step 2.2: Obtain the features of the pictures and combine all the picture features to obtain the picture representation vector, specifically as follows:
[0090] Use the VGG-19 neural network model to extract the picture features of each picture, and use the recurrent neural network model to combine all the picture features to obtain the picture representation vector P.
[0091] Step 2.3: Since the calculation methods of the text representation vector and the image representation vector are different, there are differences in the spatial dimensions of their representation vectors. It is necessary to map them to the same spatial dimension to obtain the corresponding text and image mapping vectors.
[0092] As a specific example, Step 2.3 is as follows:
[0093] Since the calculation methods of the text representation vector and the image representation vector are different, there are differences in the spatial dimensions of their representation vectors. It is necessary to map them to the same spatial dimension to obtain the text mapping vector e T and the image mapping vector e P which are:
[0094]
[0095]
[0096] where is the trainable weight, represents the weight of the image mapping function, represents the weight of the text mapping function; is the trainable bias, represents the bias of the image mapping function, represents the bias of the text mapping function.
[0097] Step 2.4: Generate text and image interaction control vectors and alternative text and image feature vectors based on the text and image representation vectors.
[0098] As a specific example, Step 2.4 is as follows:
[0099] Perform a linear transformation on the original text representation vector T and image representation vector P through a fully connected neural network to generate the text interaction control vector i T , the image interaction control vector i P and the alternative text feature vector the alternative image feature vector The formula is:
[0100]
[0101] where is the trainable weight, represents the weight of the text interaction control vector generation function, represents the weight of the image interaction control vector generation function, represents the weight of the alternative text feature vector generation function, represents the weight of the alternative image feature vector generation function; is a trainable bias, representing the bias of the text interaction control vector generation function, representing the bias of the image interaction control vector generation function, representing the bias of the alternative text feature vector generation function, representing the bias of the alternative image feature vector generation function.
[0102] Step 2.5: Using the text, image mapping vectors and text, image interaction control vectors, obtain the text, image preliminary feature vectors.
[0103] As a specific example, Step 2.5 is specifically as follows:
[0104] Multiply the text mapping vector e T , the image mapping vector e P , the text interaction control vector i T , and the image interaction control vector i P to obtain the text preliminary feature vector e′ T and the image preliminary feature vector e′ P , and the formula is:
[0105] e′ T =i T *e T (12)
[0106] e′ P =i P *e P (13)
[0107] Step 2.6: The text, image attention weight vectors reflect the importance of the text, image representation vectors. Calculate the text, image attention weight vectors according to the text, image preliminary feature vectors.
[0108] As a specific example, Step 2.6 is specifically as follows:
[0109] According to the text, image preliminary feature vectors, use the softmax function to calculate the text attention weight vector A T , the image attention weight vector A P , and the calculation formula is:
[0110] A T =softmax(e′ T ) (14)
[0111] A P =softmax(e′ P ) (15)
[0112] Step 2.7: Calculate the similarity between the text and image preliminary feature vectors by combining the text, image mapping vectors, and text and image preliminary feature vectors.
[0113] As a specific example, Step 2.7 is as follows:
[0114] The similarity α between the text preliminary feature vectors T2P and the similarity α between the image preliminary feature vectors P2T are the products of the text, image mapping vectors and text, image preliminary feature vectors, and the calculation formula is:
[0115]
[0116] The similarity between the text and image preliminary feature vectors reflects the correlation between different text feature vectors and image feature vectors. The higher the correlation of the feature vectors, the more attention will be obtained during the learning process and the greater the impact on the result;
[0117] Step 2.8: Obtain the text feature vector and image feature vector by combining the text, image attention weight vectors, the similarity between the text and image preliminary feature vectors, and the alternative text and image feature vectors.
[0118] As a specific example, Step 2.8 is as follows:
[0119] Combine the text attention weight vector A T , the image attention weight vector A P and the similarity α between the text preliminary feature vectors T2P , the similarity α between the image preliminary feature vectors P2T , as well as the alternative text feature vector the alternative image feature vector Calculate the text feature vector C T and the image feature vector C P , and the calculation formula is:
[0120]
[0121] The attention weight vector reflects the importance between different feature vectors within the feature, and the similarity reflects the correlation between different model features. Using the two in combination can not only retain the independence of the same modality features but also capture the correlation between different modality features;
[0122] Step 2.9: Eliminate the redundant information in the text feature vector and image feature vector.
[0123] As a specific example, Step 2.9 is as follows:
[0124] Eliminate the text feature vector C T and the image feature vector C P The redundant information between them, and the calculation formula is:
[0125] T′ = sigmoid(PW OT +b OT ) * tanh(C T ) (20)
[0126] P′ = sigmoid(PW OP +b OP ) * tanh(C P ) (21)
[0127] Where W OT , W OP are trainable weights, W OT represents the weight of the text redundant information elimination function, and W OP represents the weight of the image redundant information elimination function; b OT , b OP are trainable biases, b OT represents the bias of the text redundant information elimination function, and b OP represents the bias of the image redundant information elimination function.
[0128] Step 2.10: Concatenate the text feature vector and the image feature vector as the feature vector of the news content, and output the feature vector of the news content.
[0129] Step 3: According to the feature vector of the news content, combine the news classification model to complete the true or false classification of the news;
[0130] The news classification model uses a fully connected neural network layer as the classification function, uses the softmax function as the activation function, and the loss function uses cross-entropy L, and the formula is:
[0131] Y = softmax(HW Y +b Y ) (22)
[0132]
[0133] Where W Y represents the weight of the classification function, which is a trainable weight; b Y represents the bias of the classification function, which is a trainable bias; M is the number of news categories, with a value of 2, that is, true news or false news, y i and respectively represent the true value and the predicted value of the news being true or false.
[0134] The present invention provides a multi-modal fake news detection method based on an attention mechanism. There are many methods and ways to specifically implement this technical solution. The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A multimodal fake news detection method based on attention mechanism, characterized in that: The following steps are involved: Step 1: Read the relevant content of the news, including text content S′ and picture content V′; Step 2: Use deep learning technology to represent news content as corresponding feature vectors; Step 3: Based on the feature vector of the news content, the news classification model is combined to complete the classification of the truth or falsehood of the news.
2. The multimodal fake news detection method based on attention mechanism according to claim 1 is characterized in that: In step 2, the deep learning technology is used to represent the news content as a corresponding feature vector, as follows: Step 2.1, convert each sentence in the news text into a corresponding sentence representation vector, initialize the association matrix between sentences, calculate the attention weight matrix between sentences based on the association matrix, and calculate the text representation vector by combining the attention weight matrix between sentences and the value matrix of the text; Step 2.2: Get the features of the image and combine all the features to get the image representation vector, as follows: Use the VGG-19 neural network model to extract the image features of each image, and use the recurrent neural network model to merge all image features to obtain the image representation vector P; Step 2.3, map the text and image representation vectors to the same space to obtain the corresponding text and image mapping vectors; Step 2.4, generating text and picture interaction control vectors and alternative text and picture feature vectors according to the text and picture representation vectors; Step 2.5, using the text and picture mapping vectors and the text and picture interaction control vectors, obtain the text and picture preliminary feature vectors; Step 2.6: Calculate the text and picture attention weight vectors based on the preliminary feature vectors of the text and picture; Step 2.7, combining the text and picture mapping vectors with the text and picture preliminary feature vectors, and calculating the similarity between the text and picture preliminary feature vectors; Step 2.8, combining the text and picture attention weight vectors and the similarity between the text and picture preliminary feature vectors, as well as the candidate text and picture feature vectors, to obtain a text feature vector and a picture feature vector; Step 2.9, eliminating redundant information between the text feature vector and the image feature vector; Step 2.10: Concatenate the text feature vector and the image feature vector as the feature vector of the news content, and output the feature vector of the news content.
3. The multimodal fake news detection method based on attention mechanism according to claim 2 is characterized in that: As described in step 2.1, each sentence in the news text is converted into a corresponding sentence representation vector, and the association matrix between sentences is initialized. The attention weight matrix between sentences is calculated according to the association matrix, and the text representation vector is calculated by combining the attention weight matrix between sentences and the value matrix of the text, as follows: Step 2.1.1: Input each sentence in the news text into the open source pre-trained BERT language model to obtain the representation vector set S = {S1, S2, …, S n }; Step 2.1.2, for each sentence representation vector in the news text content, generate the corresponding query Query, key Key and value Value vectors respectively, and combine the query Query, key Key and value Value vectors generated by all sentences into query matrix Q, key matrix K and value matrix V respectively: Q=S*W Q (1) K=S*W K (2) V=S*W V (3) Where W Q ,W K ,W V is the weight matrix obtained through learning, W Q represents the weight of the query matrix Q generating function, W K represents the weight of the key matrix K generating function, W V represents the weights of the function generating the value matrix V; Step 2.1.3, multiply the query matrix Q by the key matrix K to obtain the attention weight matrix Att between sentences, calculated as: At=QK T (4) Step 2.1.4: Normalize the obtained attention weight matrix Att between sentences and multiply it with the value matrix V of the text to obtain the representation vector T of the news text. The calculation formula is: where d k is a scaling factor used to normalize the calculation results and ensure the stability of the gradient.
4. The multimodal fake news detection method based on attention mechanism according to claim 3 is characterized in that: Step 2.3 maps the text and image representation vectors to the same space to obtain the corresponding text and image mapping vectors, as follows: Since the calculation methods of text representation vector and image representation vector are different, the spatial dimensions of text and image representation vectors are different and need to be mapped to the same spatial dimension to obtain the text mapping vector e T and the image map vector e P for: in are the trainable weights, represents the weight of the image mapping function, Represents the weight of the text mapping function; is a trainable bias, represents the bias of the image mapping function, Indicates the bias of the text mapping function.
5. The multimodal fake news detection method based on attention mechanism according to claim 4 is characterized in that: The text and picture interaction control vectors and the alternative text and picture feature vectors are generated according to the text and picture representation vectors in step 2.4, as follows: The original text representation vector T and the image representation vector P are linearly transformed through a fully connected neural network to generate a text interaction control vector i T , picture interaction control vector i P and the candidate text feature vector Alternative image feature vector The formula is: in are the trainable weights, represents the weight of the text interaction control vector generation function, represents the weight of the image interaction control vector generation function, represents the weight of the candidate text feature vector generation function, Represents the weight of the candidate image feature vector generation function; is a trainable bias, represents the bias of the text interaction control vector generation function, represents the bias of the image interaction control vector generation function, represents the bias of the candidate text feature vector generating function, Represents the bias of the candidate image feature vector generation function.
6. The multimodal fake news detection method based on attention mechanism according to claim 5 is characterized in that: Step 2.5 uses the text and picture mapping vectors and the text and picture interaction control vectors to obtain the text and picture preliminary feature vectors, as follows: Map the text to vector e T , picture mapping vector e P and text interaction control vector i T , picture interaction control vector i P Multiply them together to get the initial feature vector e′ of the text T And the initial feature vector e′ of the image P , the formula is: e′ T =i T *e T (12) e′ P =i P *e P (13)。 7. The multimodal fake news detection method based on attention mechanism according to claim 6 is characterized in that: According to the preliminary feature vectors of text and picture described in step 2.6, the attention weight vectors of text and picture are calculated as follows: According to the preliminary feature vectors of text and pictures, the softmax function is used to calculate the text attention weight vector A T , picture attention weight vector A P , the calculation formula is: THE T =softmax(e′ T ) (14) THE P =softmax(e′ P ) (15)。 8. The multimodal fake news detection method based on attention mechanism according to claim 7 is characterized in that: Step 2.7 combines the text and picture mapping vectors with the text and picture preliminary feature vectors to calculate the similarity between the text and picture preliminary feature vectors, as follows: Similarity α between the preliminary feature vectors of the text T2P The similarity α between the initial feature vector of the image P2T It is the product of the text and picture mapping vector and the text and picture preliminary feature vector. The calculation formula is: Step 2.8 combines the text and picture attention weight vectors and the similarity between the text and picture preliminary feature vectors, as well as the candidate text and picture feature vectors to obtain the text feature vector and picture feature vector, as follows: Combined with the text attention weight vector A T , picture attention weight vector A P The similarity α between the initial feature vector of the text T2P , the similarity between the initial feature vectors of the pictures α P2T , and the candidate text feature vector Alternative image feature vector Calculate the text feature vector C T and the image feature vector C P , the calculation formula is:
9. The multimodal fake news detection method based on attention mechanism according to claim 8, characterized in that: Eliminating redundant information between text and image feature vectors as described in step 2.9 is as follows: Eliminate text feature vector C T and the image feature vector C P The redundant information in is calculated as follows: T′=sigmoid(PW OT +b OT )*tanh(C T ) (20) P′=sigmoid(PW OP +b OP )*tanh(C P ) (21) Where W OT ,W OP is the trainable weight, W OT Represents the weight of the text redundant information elimination function, W OP represents the weight of the image redundant information elimination function; b OT ,b OP is the bias that can be trained, b OT represents the bias of the text redundant information elimination function, b OP Represents the bias of the image redundant information elimination function.
10. The multimodal fake news detection method based on attention mechanism according to claim 9, characterized in that: The news classification model described in step 3 is as follows: The news classification model uses a fully connected neural network layer as the classification function, uses the softmax function as the activation function, and uses the cross entropy L as the loss function. The formula is: <h2 style=";text-align:left;direction:ltr">Y = softmax(HW<h2 style=";text-align:left;direction:ltr"> Y <h2 style=";text-align:left;direction:ltr"> +b<h2 style=";text-align:left;direction:ltr"> Y <h2 style=";text-align:left;direction:ltr"> ) (22) Where W Y represents the weight of the classification function, which is a trainable weight; b Y represents the bias of the classification function, which is a trainable bias; M is the number of news categories, and its value is 2, i.e. true news or false news, y i and Represent the true value and predicted value of whether the news is true or false respectively.