A Multimodal False News Detection Method Based on User Cognitive Consistency Reasoning
By introducing user cognitive consistency inference methods in multimodal fake news detection, cross-modal alignment, context interaction and collaborative reasoning layer structures, the problem of difficulty in capturing modal inconsistent information in the existing technology is solved, and the detection performance is significantly improved.
Patent Information
- Application Number
- CN202210574816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-24
AI Technical Summary
Existing multimodal fake news detection methods are difficult to capture inconsistent information between different modes, resulting in insufficient detection performance.
A multimodal fake news detection method (UCCIN) based on user cognitive consistency reasoning is proposed. Through cross-modal alignment layer, context interaction layer and collaborative reasoning layer, inconsistent information is mined from multiple levels between news visual information and text content, comments and comments, and news and comments.
It improves the performance of multimodal fake news detection, can effectively capture inconsistent information between news and comments, and enhances the accuracy and credibility of the detection.
Smart Images

Figure CN115964482B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic information technology, and relates to a multi-modal fake news detection method based on user cognitive consistency reasoning. Background Art
[0002] With the rapidity of information release and the equality of consumption of social media, how to efficiently and accurately identify fake news spread on social media has currently become one of the major key issues in the fields of social media analysis and information content security.
[0003] Recently, the news spread on social media has gradually changed from a text-based unimodal form to a multi-modal form mainly consisting of content such as images and videos. Multi-modal news is more attractive to readers for browsing, absorbing, sharing, and spreading, and can provide readers with an immersive reading experience. However, fake news has also taken full advantage of this to attract and mislead readers, resulting in the continuous emergence of multi-modal fake news. In order to mitigate the negative impacts brought by fake news, automatically detecting multi-modal fake news on social media has become an important issue that needs to be solved urgently by the current academic community, industry, and even government agencies.
[0004] Compared with the single-modal fake news detection task, the challenge of multi-modal fake news detection lies in how to learn valuable features from multi-model information with modal imbalance, large dimensional differences, and inconsistent semantic spaces to improve detection performance. Existing research mainly focuses on extracting and learning the correlation features between different multi-modal information for detection, which can be mainly divided into two categories: The first is the multi-modal interaction method, which aims to measure the similarity relationship between them through different cross-interaction mechanisms, so as to match similar semantic features. For example, Zhou et al. first extract text and visual features from news, and then further construct an aligned attention interaction mechanism to extract the matching relationship between cross-modal features. The second method is to introduce auxiliary tasks related to multi-modal information and learn the associated shared features between multi-tasks. Among them, Jaiswal et al. combine deep multi-modal representation learning with outlier detection methods to capture the consistency relationship between multi-modal features. At the same time, research based on pre-trained models is becoming more and more common, which uses pre-trained models such as BERT and VGG-19 to learn the deep associated semantics of text and images in news, thus significantly improving the detection performance. However, although these methods have achieved certain performance, they still have relatively serious defects, that is, it is difficult to capture the inconsistent information between different modalities. Traditional detection methods usually construct a series of aligned interaction models to capture the common similar semantics between different modalities to improve detection performance, but it is difficult to capture the inconsistent information between multi-modalities and apply this high-confidence indication feature to detection. To solve these problems, the present invention observes that a series of cognitive behaviors of users when questioning rumors can easily discover the inconsistent information in multi-modal fake news. Based on this, how to model the process of users' cognition of rumors to extract differential features in multi-modal information is the key to multi-modal fake news detection research. Summary of the Invention
[0005] Technical Problems to be Solved
[0006] Aiming at the defects existing in current multi-modal fake information detection methods and inspired by the cognitive process of users in identifying rumors, the present invention proposes a multi-modal fake news detection method based on user cognitive consistency inference (UCCIN), which comprehensively mines inconsistent information of fake news from the following three perspectives: between news visual information and text content, between comments and comments, and between news and comments, etc., to improve the multi-modal fake news detection ability. Specifically, in UCCIN, the present invention first designs a cross-modal alignment layer to semantically align the text information and visual information in the news to check the consistency of news semantics. In order to obtain valuable information widely discussed by readers from comments, the present invention develops a context interaction layer to interact each comment information with the global comment semantics, and filters out comment semantics irrelevant to the news through the designed dual-channel gating block, thereby strengthening the most concerned comment semantics. Finally, the present invention designs a collaborative inference layer to drive the interaction inference between the most concerned comment semantics and the consistent semantics of the news, and aggregates and fuses the inferred inconsistent features to enhance the detection ability between news and comments.
[0007] Technical solution
[0008] A multi-modal fake news detection method based on user cognitive consistency inference, characterized by the following steps:
[0009] S1: Embedding encoding module
[0010] Use the BERT encoder and ResNet-152 respectively to perform embedding encoding representations on the text information and image information contained in the multi-modal news;
[0011] S2: Cross-modal alignment module
[0012] Use the self-attention network to extract features from the embedded encoded text information and image information, input the features into the designed cross-attention block for semantic interaction, and integrate the consistent features captured from the two perspectives of news text features and news image features;
[0013] S3: Context interaction module
[0014] Design a context interaction layer to comprehensively interact all comment information with the global comment semantics, thereby mining and strengthening the semantic features most concerned by users in the comments;
[0015] S4: Collaborative inference module
[0016] Design a collaborative inference layer composed of a collaborative guidance block, a cross-attention block, and an aggregation and fusion block to perform collaborative inference on the multi-modal news semantics and the most concerned comment features, thereby discovering inconsistent information between news and comments and improving the detection performance of the model.
[0017] A further technical solution of the present invention: S2 includes the following steps:
[0018] S21: Self-attention network: The multi-head self-attention mechanism is adopted to learn the global correlation of all positions in the text sequence and the image respectively. Given the query Q, the key K, and the value V, the scaled dot-product attention is expressed as:
[0019]
[0020] Among them, in the text content, set In the image information, set d is the embedding dimension of the word, and N is the length of the text sequence;
[0021] S22: Project the query, key, and value h times through different linear projections, and then perform the scaled dot-product attention on these results in parallel; formally, the multi-head attention network can be expressed as:
[0022] head i = Attention(QW i q ,KW i k ,VW i v ) (4)
[0023]
[0024] Among them, W i q , W i k , W i v And are all trainable parameters, and H is d / h; O = O 3 and O = O G are the text features of the embedded encoding and the image features of the embedded encoding;
[0025] S23: Cross-attention block: Set the encoded text feature as O T as the query, and the image feature O G as the key and value; Set the encoded image feature as O G to represent the query, and the text function O T as the key and value; This process is described as:
[0026]
[0027]
[0028]
[0029]
[0030] Among them, all Ws are trainable parameters;
[0031] S24: Integrate the consistent features captured from the perspectives of news text features and news image features:
[0032]
[0033] Among them, ';' is the splicing operation and is the overall consistency feature.
[0034] A further technical solution of the present invention: S3 includes the following steps:
[0035] S31: Process all comments through the average pooling strategy to obtain the global feature C avgp ; Process all comments through the max pooling strategy to obtain the global feature C maxp ;
[0036] S32: Adopt the cross-attention block of S23 to receive C maxp as the query, C i as the key, C avgp as the value, so as to obtain the potentially useful features of the i-th comment, that is
[0037] S33: Dual-channel gating block: Use one channel to directly transmit all information downstream, and use the other channel to screen potentially useful features through the gating mechanism, that is, first transform the potentially useful features in the spatial dimension, and then filter these features using linear gating, expressed as:
[0038]
[0039]
[0040] Among them, both W and b are trainable parameters, and ⊙ is the element-wise product operation;
[0041] S34: Integrate the significantly useful information mined from k comments, so as to obtain the overall valuable voice O of all comments CC :
[0042]
[0043] A further technical solution of the present invention: S4 includes the following steps:
[0044] S41: Collaborative Guidance Block: Develop two collaborative guidance blocks to guide the learning of original valuable comment features from the perspectives of text content and image content respectively, so as to obtain valuable comment features closely related to news texts and news images respectively; among them, the process guided by text content is expressed as:
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051] Among them, all W and b are trainable parameters;
[0052] S42: In the same way as S41, the valuable comment features guided by news pictures are learned as
[0053] S43: Cross-Attention Block: Use two cross-attention blocks to interact the semantics of news texts and news images with the semantics of comments respectively. In the interaction between news texts and comments, this module uses as the query, as the key, as the value; in the interaction between news images and comments, this module uses as the query as the key, as the value, and the inconsistent semantics learned for text and image are respectively marked as O TC and O GC ;
[0054] S44: Aggregation and Fusion Block: First, respectively focus on the significant feature parts of the inconsistent semantics of the information from two perspectives through element-wise product operations, that is, and Secondly, use the residual network to fuse the significant features with text and image features, that is, T and g; finally, fully fuse these features by means of absolute value difference:
[0055]
[0056]
[0057]
[0058]
[0059] TG = [T; |T - G|; T⊙G; G] (24)
[0060] S45: Use the Softmax function to predict the learned probability distribution and perform cross - entropy training through the global loss function:
[0061] p = softmax(W p TG + b p ) (25)
[0062] Loss = -∑ylogp (26)
[0063] where W p , b p are trainable parameters and y is the true label.
[0064] A computer system, characterized in that it includes: one or more processors, a computer - readable storage medium for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the above - mentioned method.
[0065] A computer - readable storage medium, characterized in that it stores computer - executable instructions, and the instructions are used to implement the above - mentioned method when executed.
[0066] A computer program, characterized in that it includes computer - executable instructions, and the instructions are used to implement the above - mentioned method when executed.
[0067] Beneficial effects
[0068] A multi - modal fake news detection method (UCCIN) based on user - cognitive consistency reasoning proposed by the present invention comprehensively mines inconsistent information of fake news from the following three perspectives: between news visual information and text content, between comments and comments, and between news and comments at multiple levels, improving the multi - modal fake news detection ability. The present invention provides a new idea for multi - modal information fusion and improves the detection performance of multi - modal fake news. Compared with the prior art, the present invention has the following innovations:
[0069] 1: The present invention proposes a consistency reasoning network based on the user - cognitive rumor process for multi - modal fake news detection, which can discover inconsistent information between news and comments and improve the multi - modal fake news detection ability. As far as we know, this is the first time to apply the user - cognitive mechanism to the multi - modal fake news detection task;
[0070] 2: Inspired by the cognitive process of users observing multimodal news, the present invention designs a cross-modal alignment layer, which can focus on the consistency information of multimodal news from the perspectives of text and image respectively.
[0071] 3: The context interaction layer developed by the present invention can discover valuable semantics between comments through the combination of cross-interaction and dual-channel gating mechanisms.
[0072] Extensive experiments of the present invention on three competitive fake news detection datasets have confirmed the superiority of the present invention and the collaborative effectiveness of each module. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The drawings are only for the purpose of illustrating specific embodiments, and are not considered as a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0074] Figure 1 is the architecture diagram of the present invention;
[0075] Figure 2 is the experimental performance diagram of the present invention under three datasets of Weibo, Twitter and PHEME;
[0076] Figure 3 is the separation performance comparison diagram of different modules of the present invention under three datasets of Weibo, Twitter and PHEME. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0078] In view of the defects existing in current multi-modal false information detection methods and inspired by the cognitive process of users in identifying rumors, the present invention proposes a multi-modal false news detection method based on user cognitive consistency inference (UCCIN), which comprehensively mines inconsistent information of false news from the following three perspectives: between news visual information and text content, between comments and comments, and between news and comments, etc., to improve the multi-modal false news detection ability. Specifically, in UCCIN, the present invention first designs a cross-modal alignment layer to semantically align the text information and visual information in the news to check the consistency of news semantics. In order to obtain valuable information widely discussed by readers from comments, the present invention develops a context interaction layer to enable each comment information to interact with the global comment semantics, and filters out comment semantics irrelevant to the news through the designed dual-channel gating block, thereby strengthening the most concerned comment semantics. Finally, the present invention designs a collaborative inference layer to drive the interaction and inference between the most concerned comment semantics and the consistent semantics of the news, and aggregates and fuses the inferred inconsistent features to enhance the detection ability between news and comments.
[0079] The architecture diagram of the present invention is as Figure 1 shown, and it includes the following four modules:
[0080] Module 1: Embedding Encoding Module.
[0081] For the text information and visual information contained in multi-modal news, the present invention respectively uses a BERT encoder and ResNet-152 to perform embedding encoding representations on the text and vision;
[0082] Module 2: Cross-modal Alignment Module.
[0083] This module simulates the attention process of people when observing a multi-modal news, that is, usually first observing the visual information, and then looking at the text content information in detail with the content of the picture information, constructs a self-attention network to encode visual and text features, and designs a cross-attention network to enable deep semantic interaction between visual features and text features, thereby checking the consistency of news semantics;
[0084] Module 3: Context Interaction Module.
[0085] For the semantics with a large number of different views usually existing in the comments under the news, this module enables all comment information to comprehensively interact with the global comment semantics, thereby mining and strengthening the semantic features most concerned by users in the comments.
[0086] Module 4: Collaborative Inference Module.
[0087] This module constructs a collaborative guidance layer and a cross-attention layer to perform collaborative inference on multi-modal news semantics and the most concerned comment features, so as to discover inconsistent information between news and comments and improve the detection performance of the model.
[0088] The specific process of the method of the present invention is as follows:
[0089] Module 1: Embedding Encoding Module
[0090] Step 1: Multi-modal news usually includes text content and image information. For the encoding of text content, the present invention uses a pre-trained BERT model to encode the text sequence, and the encoded text sequence can be represented as T = {t1, t2, …, t N}, where is the i-th word after encoding.
[0091] Step 2: For the encoding of news images, the present invention transforms the image into 224×224 pixels, and then uses ResNet-152 to learn the representation of image information. In order to make the visual features obtain the same dimension as the text features, we project the encoded image representation ResNet(I) through a linear transformation:
[0092]
[0093] G = W r ResNet(I) (2)
[0094] where, r i is the i-th image region characterized by a 2048-dimensional vector, and W r is a trainable parameter. G is the encoded vector of the image content.
[0095] Step 3: For the encoding of the i-th comment under the news, the present invention uses a pre-trained BERT model to obtain the comment sequence representation C i = {c1, c2, …, c N}, where M is the length of the comment sequence. The embedding encoding of k comments can be represented as C1, C2, …, C k .
[0096] Module 2: Cross-modal Alignment Module.
[0097] Step 4: Intuitively, when people read a piece of news, they usually first look at vivid images and then, with this image information, view the text content in detail. This process can be abstracted as multimodal interaction, mainly learning the consistent information between the text content and the image content. To this end, the present invention designs a cross-modal alignment layer to simulate this process. First, it encodes the text and image content using a self-attention network, and then inputs the encoded features into the designed cross-attention block for semantic interaction. Specifically as follows:
[0098] Step 5: Self-attention network. The present invention uses a multi-head self-attention mechanism to learn the global correlations of all positions in the text sequence and the image respectively. Given a query Q, a key K, and a value V, the scaled dot-product attention is expressed as:
[0099]
[0100] Among them, in the text content, this module sets In the image information, this module sets d is the embedding dimension of the word, and N is the length of the text sequence;
[0101] Step 6: To obtain more valuable information from the news in parallel, this module projects the query, key, and value h times through different linear projections, and then these results perform scaled dot-product attention in parallel. Formally, the multi-head attention network can be expressed as:
[0102] head i =Attention(QW i q ,KW i k ,VW i v ) (4)
[0103]
[0104] Among them, W i q , W i k , W i v And are all trainable parameters, and H is d / h. O=O T and O=O G are the text sequence with embedding encoding and the image features with embedding encoding.
[0105] Step 7: Cross-attention block. The cross-attention block is a variant of the standard multi-head attention block that can capture global dependencies between text and images. The present invention develops two cross-attention blocks from the text-image and image-text perspectives. Among them, for the text-image cross-attention block (Equations (6) and (7)), this module sets the encoded text features as O T as the query, and the image features O G as the key and value. In this way, the text features can guide the model to focus on consistent image regions. Conversely, for the image-text cross-attention block (Equations (8) and (9)), this module sets the image such that the image features are O G represent the query, and the text function O T as the key and value. In this way, this module can capture consistent semantic features in the text content through the focus of the image features. This process can be described as:
[0106]
[0107]
[0108]
[0109]
[0110] where all W are trainable parameters.
[0111] Step 8: To obtain more extensive consistent features from the news content itself, the present invention integrates the consistent features captured from the two perspectives of news text semantics and news visual features:
[0112]
[0113] where ';' is the concatenation operation and is the overall consistency feature.
[0114] Module 3: Context Interaction Module.
[0115] Step 9: For specific news, there are often different voices and opinions in the comments, especially for fake news, where there are even more different voices and opinions. To learn valuable information widely discussed by readers (audiences) from the comments, the present invention designs a context interaction layer to enable all comments under the news to interact with each other. This module consists of two blocks, namely 1) Cross-attention block: enabling different comments to interact deeply to capture potential useful features in the comments; 2) Dual-channel gating block: strengthening more significant valuable semantics from the potential useful features.
[0116] Step 10: Cross-attention block. To capture key valuable features from all comments, it is necessary to enable all comments to interact with each other. If pairwise interactions are carried out between comments, for a news item containing k comments, k×(k - 1) interaction networks need to be designed, which will greatly inflate the model and seriously affect the model's efficiency. To overcome this problem, the present invention utilizes the global characteristics of average pooling and max pooling operations to obtain the global features of all comments, which are C avgp and C maxp .
[0117] Step 11: Then, the present invention drives each comment to interact with these two global features to obtain potentially useful segments between the two comments. Specifically, our cross-attention block (with the same structure as the cross-attention in Step 7) receives C maxp as the query, C i as the key, and C avgp as the value, so as to obtain the potentially useful features of the i-th comment, which are
[0118] Step 12: Dual-channel gating block. To filter out features unrelated to the news among these potentially useful features and extract important valuable voices, the present invention proposes a dual-channel gating block to purify these features, so as to capture important valuable features in each comment. Specifically, the present invention uses one channel to directly transmit all information downstream, and uses the other channel to screen potentially useful features through a gating mechanism, that is, first transform the potentially useful features in the spatial dimension, and then use a linear gate to filter these features, which can be expressed as:
[0119]
[0120]
[0121] where W and b are both trainable parameters. ⊙ represents the element-wise product operation.
[0122] Step 13: Then, the present invention integrates the significantly useful information mined from these k comments to obtain the overall valuable voice O CC .
[0123]
[0124] Module 4: Collaborative inference module.
[0125] Step 14: To fully discover the inconsistent features between news and comments, this module does not simply splice the learned features from multiple perspectives such as text content, image content, and comments. The present invention interacts the valuable features in the comments with the image features and text features of these news respectively, so as to mine the inconsistent features of the semantics of news and comments. Specifically, this module designs a collaborative reasoning layer composed of a collaborative guidance block, a cross-attention block, and an aggregation and fusion block. Specifically as follows:
[0126] Step 15: Collaborative guidance block. Considering the diversity of valuable information in the comments, there may be many features irrelevant to the news. This module develops two collaborative guidance blocks to guide the learning of the original valuable comment features from the perspectives of text content and image content respectively, so as to obtain valuable comment features closely related to news text and news images respectively. Taking the guidance of comment by text content as an example, this process can be expressed as:
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133] Among them, all W and b are trainable parameters.
[0134] Step 16: In the same way as in Step 15, this module can learn the valuable comment features guided by news pictures as Replace with
[0135] Step 17: Cross-attention block. To learn valuable semantics from comment features that reveal doubts about the credibility of news content, the present invention uses two cross-attention blocks to interact the semantics of news text and news images with the semantics of comments respectively. In the interaction between news text and comments, this module uses as the query, as the key, as the value. In the interaction between news images and comments, this module uses as the query as the key, as the value. In this way, the inconsistent semantics learned for text and image are respectively marked as OTC with O GC 。
[0136] Step 18: Aggregation and fusion block. The present invention constructs an aggregation and fusion block to fully fuse the inconsistent semantics obtained from the text and image perspectives. Specifically, this module first focuses on the significant feature parts of the inconsistent semantics of the information from both perspectives through element-wise product operations, that is and Secondly, the residual network is used to fuse the significant features with the text and image features, namely T and g. Finally, these features are fully fused by means of absolute value difference.
[0137]
[0138]
[0139]
[0140]
[0141] TG = [T; |T - G|; T⊙G; G] (24)
[0142] Step 19: Finally, the Softmax function predicts the learning probability distribution of this task, and the global loss is used to force the model to minimize the cross-entropy error for the training samples with the true label y:
[0143] p = softmax(W p TG + b p ) (25)
[0144] Loss = -∑ylogp (26)
[0145] The above method is applicable to the social network environment and can be used in the social media network environment that can provide a large amount of multimodal information.
[0146] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A multimodal fake news detection method based on user cognitive consistency reasoning, characterized in that The steps are as follows: S1: Embedding Encoding Module Use the BERT encoder and ResNet-152 respectively to perform embedding encoding representation on the text information and image information contained in the multi-modal news; S2: Cross-modal Alignment Module Use the self-attention network to extract features from the embedded encoded text information and image information, input the features into the designed cross-attention block for semantic interaction, and integrate the consistent features captured from the perspectives of news text features and news image features; S3: Context Interaction Module Design a context interaction layer to comprehensively interact all comment information with the global comment semantics, so as to mine and strengthen the semantic features that users are most concerned about in the comments; S4: Collaborative Inference Module Design a collaborative inference layer composed of a collaborative guidance block, a cross-attention block, and an aggregation fusion block to perform collaborative inference on the multi-modal news semantics and the most concerned comment features, so as to discover inconsistent information between the news and the comments and improve the detection performance of the model.
2. The multimodal fake news detection method based on user cognitive consistency reasoning according to claim 1, wherein S2 It includes the following steps: S21: Self-attention Network: Use the multi-head self-attention mechanism to learn the global correlation of all positions in the text sequence and the image respectively. Given the query Q, key K, and value V, the scaled dot-product attention is expressed as: Among them, in the text content, set In the image information, set d is the embedding dimension of the word, and N is the length of the text sequence; S22: Project the query, key, and value h times through different linear projections, and then perform the scaled dot-product attention on these results in parallel; formally, the multi-head attention network can be expressed as: head i = Attention(QWi i q ,KW i k ,VW i v ) (4) Among them, W i q , W i k , W i v and are all trainable parameters, and H is d / h; O = O T and O = O G are the text features of the embedded encoding and the image features of the embedded encoding; S23: Cross-attention block: Set the encoded text features to O T as the query, and the image features O G as the key and value; Set the encoded image features to O G represent the query, and the text features O T as the key and value; This process is described as: where all W are trainable parameters; S24: Integrate the consistent features captured from the perspectives of news text features and news image features: Among them, ';' is an operation for splicing and is an overall consistency feature.
3. The multi-modal fake news detection method based on user cognitive consistency reasoning according to claim 2, wherein S3 It includes the following steps: S31: Process all comments through the average pooling strategy to obtain the global feature C avgp ; Process all comments through the max pooling strategy to obtain the global feature C maxp ; S32: Receive C using the cross-attention block of S23 maxp as the query, C i as the key, C avgp as the value, thereby obtaining the potentially useful features of the i-th comment, which are S33: Dual-channel Gating Block: Use one channel to directly transmit all information downstream, and use the other channel to screen potential useful features through a gating mechanism, that is, first transform the potential useful features in the spatial dimension, and then filter these features using a linear gate, expressed as: where W and b are both trainable parameters, and ⊙ is the element-wise product operation; S34: Integrate the significantly useful information mined from k comments to obtain the overall valuable voice O of all comments CC :
4. The multimodal fake news detection method based on user cognitive consistency reasoning according to claim 3, wherein S4 includes the following steps: S41: Collaborative Guidance Block: Develop two collaborative guidance blocks to guide the learning of the original valuable comment features from the perspectives of text content and image content respectively, so as to obtain valuable comment features closely related to news text and news image respectively; The process guided by text content is expressed as: where all W and b are trainable parameters; S42: In the same way as S41, the valuable comment features guided by news pictures learned are S43: Cross-attention block: Two cross-attention blocks are used to interact the semantics of news text and news images with the semantics of comments respectively. In the interaction between news text and comments, this module uses as the query, as the key, and as the value; in the interaction between news images and comments, this module uses as the query as the key, and as the value. The inconsistent semantics learned for text and images are respectively labeled as O TC and O GC ; S44: Convergence and Fusion Block: First, the significant feature parts of the inconsistent semantics of the two perspective information are respectively focused through element-wise product operations, that is, and Secondly, the residual network is used to fuse the significant features with the text and image features, namely T and g; finally, these features are fully fused by means of absolute value difference: TG = [T; |T - G|; T⊙G; G] (24) S45: Use the Softmax function to predict the learning probability distribution and perform cross-entropy training through the global loss function: p = softmax(W p TG + b p ) (25) Loss = -∑ylogp (26) Among them, W p , b p are trainable parameters, and y is the true label.
5. A computer system, characterized in that It includes: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in claim 1.
6. A computer-readable storage medium, characterized in that Stored with computer-executable instructions, the instructions are used to implement the method described in claim 1 when executed.
7. A computer program, characterized in that Including computer-executable instructions, the instructions are used to implement the method described in claim 1 when executed.
Citation Information
Patent Citations
False news detection system and method based on evidence perception hierarchical interaction attention network
CN111581979A
Deep learning-based rumor detection method and device, equipment and storage medium
CN113722484A