False news detection method and device based on comparative learning and multi-modal fusion
Through the common attention network to enhance and integrate the text, image and social graph structural characteristics of fake news, and use contrast learning to perform fake news detection, the problem of low accuracy of fake news detection in the existing technology is solved, and more efficient false news recognition is achieved.
Patent Information
- Application Number
- CN202510202338.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-22
AI Technical Summary
The existing fake news detection methods are relatively weak in integrating multimodal features, making it difficult to guarantee the accuracy of fake news detection.
The initial text, image and graph structure features of the news are extracted through the pre-trained feature extraction model, and modal enhancement and fusion are used to use the common attention network to perform modal enhancement and fusion, and finally the predicted probability of false news is obtained through comparative learning.
Through the common attention network fusion of three modal features, the semantic consistency between different modalities is ensured and the accuracy of false news detection is improved. Comparative learning effectively distinguishes real and fake news, improving detection performance.
Smart Images

Figure CN120046111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and more specifically, to a false news detection method and device based on contrast learning and multimodal fusion. Background Art
[0002] In the current information age, the Internet has achieved leapfrog development. Social media has become a platform for most people to obtain and exchange information due to its rich information, convenient information sharing, fast and wide dissemination, etc.
[0003] The rapid development of social networks has made it an ideal place for news dissemination, which has also led to the rapid spread of false news and distorted views. The development of multimedia technology has provided opportunities for self-media news, transforming it from a simple text post to a multimedia post with images or videos. Compared with news with only text, these news have richer information and attract more audiences. Similarly, false news also takes full advantage of this to attract and mislead readers.
[0004] False news detection methods mainly use the co-attention mechanism to perform pairwise fusion between modalities based on text, image, and social Figure 3 three types of features, and then concatenate the obtained enhanced features into the final multimodal features. However, existing false news detection methods ignore the interaction between the three modalities. In addition, due to the large differences in the number and quality of modalities contained in different news samples, there is an imbalance between modalities. Although the attention fusion mechanism can adaptively allocate weights to different modalities, it cannot directly constrain and enhance the semantic consistency between different modalities. Therefore, it cannot fully ensure the high coordination of different modalities at the semantic level. Therefore, existing false news detection methods are relatively weak in fusing multimodal features, making it difficult to guarantee the accuracy of false news detection. Summary of the Invention
[0005] In view of this, the present invention provides a false news detection method and device based on contrast learning and multimodal fusion, which are used to solve the problem that existing false news detection methods based on multimodal fusion are relatively weak in fusing multimodal features and it is difficult to guarantee the accuracy of false news detection.
[0006] To achieve the above object, the following solutions are proposed:
[0007] A false news detection method based on contrast learning and multimodal fusion, the method includes:
[0008] Respectively extract the initial text feature, initial image feature, and initial graph structure feature of the news to be detected through a pre-trained feature extraction model, where the initial graph structure feature is the structure feature of the social graph;
[0009] The features in the initial text feature, the initial image feature, and the initial graph structure feature are respectively subjected to modality enhancement with the other two modality features through a co-attention network to obtain an enhanced text feature, an enhanced image feature, and an enhanced graph structure feature;
[0010] The enhanced text feature, the enhanced image feature, and the enhanced graph structure feature are concatenated through a co-attention network to obtain a target multi-modal output feature;
[0011] The target multi-modal output feature is subjected to contrastive learning and input into a fully connected layer to obtain the prediction probability that the news to be detected is false news.
[0012] Preferably, the process of respectively extracting the initial text feature, the initial image feature, and the initial graph structure feature of the news to be detected through a pre-trained feature extraction model includes:
[0013] The initial text feature of the news to be detected is extracted through a pre-trained convolutional neural network;
[0014] The initial image feature of the news to be detected is extracted through a pre-trained ResNet50;
[0015] The initial graph structure feature of the news to be detected is extracted through a pre-trained graph attention network.
[0016] Preferably, the process of enhancing the initial text feature with the other two modality features includes:
[0017] The initial text feature is enhanced based on the initial image feature to obtain a cross-modal text feature;
[0018] The cross-modal text feature is enhanced based on the initial graph structure feature to obtain an enhanced text feature.
[0019] Preferably, the enhancing the initial text feature based on the initial image feature includes:
[0020] The initial text feature is enhanced through the following formula:
[0021]
[0022] where is the cross-modal text feature, is the initial image feature query vector, is the initial image feature, is the transpose of the initial text feature chain vector, is the initial text feature key vector, is the initial text feature value vector, R is the number of heads, is a linear transformation, h represents the head index in the multi-head attention mechanism, i is an index variable, d represents the feature dimension, H is the number of heads in the multi-head attention mechanism, and Q is related to the query;
[0023] Enhance the cross-modal text features through the following formula:
[0024]
[0025] where is to enhance the text features, is the initial graph structure feature query vector, is the initial graph structure feature, is a linear transformation.
[0026] A fake news detection device based on contrast learning and multi-modal fusion, the device includes:
[0027] A feature extraction unit that extracts the initial text features, initial image features, and initial graph structure features of the news to be detected through a pre-trained feature extraction model, where the initial graph structure feature is the structure feature of the social graph;
[0028] A feature enhancement unit that enhances each feature in the initial text features, initial image features, and initial graph structure features with the other two modal features respectively through a co-attention network to obtain enhanced text features, enhanced image features, and enhanced graph structure features;
[0029] A feature fusion unit that concatenates the enhanced text features, enhanced image features, and enhanced graph structure features through a co-attention network to obtain the target multi-modal output features;
[0030] A news detection unit that performs contrast learning on the target multi-modal output features and inputs them into a fully connected layer to obtain the prediction probability that the news to be detected is fake news.
[0031] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0032] The false news detection method and device based on contrast learning and multimodal fusion provided by the present invention enhance each of the initial text features, initial image features, and initial graph structure features with the features of the other two modalities through a co-attention network; then, the co-attention network concatenates the enhanced text features, enhanced image features, and enhanced graph structure features; finally, contrast learning is performed on the target multimodal output features and input into a fully connected layer to obtain the prediction probability that the news to be detected is false news. The present invention fuses the three modality features through a co-attention network to obtain enhanced features with three modality information, ensuring semantic consistency between different modalities and improving the accuracy of false news detection.
[0033] The present invention maps the differences between real news and false news to the feature space through contrast learning, making the representations of positive samples as close as possible. The present invention can better distinguish real and false news and improve the accuracy of false news detection by bringing the probability representations of samples with the same label closer through contrast learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.
[0035] Figure 1 It is a flowchart of a false news detection method based on contrast learning and multimodal fusion provided by an embodiment of the present invention;
[0036] Figure 2 It is an architecture diagram of a false news detection network model based on contrast learning and multimodal fusion provided by an embodiment of the present invention;
[0037] Figure 3 It is an architecture diagram of a multimodal feature fusion network provided by an embodiment of the present invention;
[0038] Figure 4 It is a schematic diagram of the test results of the model performance provided by an embodiment of the present invention;
[0039] Figure 5 It is a schematic diagram of the structure of a false news detection device based on contrast learning and multimodal fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] First, in combination with Figure 1 An introduction is given to a false news detection method based on contrastive learning and multi-modal fusion provided by an embodiment of the present invention. As Figure 1 shown, the detection method includes:
[0042] Step S01, respectively extract the initial text feature, initial image feature, and initial graph structure feature of the news to be detected through a pre-trained feature extraction model.
[0043] Specifically, the embodiments of the present invention utilize text, visual, and social Figure 3 features for multi-modal feature fusion. The embodiments of the present invention are implemented through the network model architecture as Figure 2 shown. As Figure 2 shown, the network model is divided into three parts. The first part is feature extraction, and three different sub-models are used to extract features from text, images, and social graphs respectively. Among them, the nodes in the social graph represent posts, comments, or users, etc., and the edges in the social graph represent the follow-up relationships between users, the publishing relationships between users and posts / comments, and the subordinate relationships between posts and comments, etc.
[0044] In terms of text feature extraction, the embodiments of the present invention use a CNN fusion strategy to extract the semantic characteristics of sentences, and extract the initial text feature of the news to be detected through a pre-trained convolutional neural network (CNN). By using convolutional kernels of different sizes, the CNN can obtain feature representations of different granularities. Among them, smaller convolutional kernels can capture word-level features, while larger convolutional kernels can capture sentence-level features over a larger range. This multi-scale feature representation helps to understand the text more comprehensively. The present invention obtains feature maps through convolutional layers with different receptive field sizes. Then, a max pooling layer is used to extract the maximum value of each feature map, and finally, the maximum values of all feature maps are connected together to form the overall initial text feature.
[0045] In terms of image feature extraction, in the embodiments of the present invention, the initial image features of the news to be detected are extracted by the pre-trained ResNet50. The output of the second layer of ResNet50 is extracted and transformed through a fully connected layer to obtain the initial image features with the same dimension as the initial text features. Among them, ResNet50 can be pre-trained based on the UmageNet database. The advanced features learned by ResNet50 on ImageNet are fully utilized, further improving the performance of fake news detection. The pre-trained ResNet50 ensures the quality and reliability of the image features.
[0046] In terms of the extraction of the graph structure features of the social graph, in the embodiments of the present invention, the initial graph structure features of the news to be detected are extracted by the pre-trained graph attention network (GAT). A social graph is constructed based on the social network, where the nodes represent users, comments, and posts; the edges represent the relationships such as interactions and attentions between them. For user nodes, the average value of the embedding vectors of their posts and comments is used as the initial embedding to comprehensively reflect the characteristics of the users. For post and comment nodes, the text features of the posts and comments are used as the initial embeddings. The graph attention network can extract rich features from the social graph, providing strong support for subsequent social network analysis.
[0047] Step S02, through the co-attention network, each feature in the initial text features, initial image features, and initial graph structure features is respectively enhanced with the other two modal features.
[0048] Specifically, as Figure 2 shown, the second part of the network model architecture is feature fusion, in order to effectively integrate the text features, image features, and graph structure features of the post. The embodiments of the present invention provide a multi-modal feature fusion network as Figure 3 shown for constructing effective context information between different modules and extracting high-order complementary information therefrom. As Figure 2 shown, the multi-modal feature fusion network consists of multiple co-attention modules.
[0049] The features of the three modalities extracted are fused through the co-attention network to obtain enhanced features with three-modal information. The embodiments of the present invention utilize text, visual, and social Figure 3 features, and through the multi-modal co-attention network for multi-modal feature fusion, enhanced text features, enhanced image features, and enhanced graph structure features are obtained.
[0050] Taking the fusion and enhancement of the initial text features as an example for introduction, the process is as follows:
[0051] Enhance the initial text features based on the initial image features to obtain cross-modal text features. For the initial text features Use and as query vectors, chain vectors, and value vectors for text features. Among them, is a linear transformation, and R is the number of heads. When visual features are to be used to enhance text features, is replaced with Then the cross-modal features are obtained as:
[0052]
[0053] Among them, is the cross-modal text feature, is the initial image feature query vector, is the initial image feature, is the transpose of the initial text feature chain vector, is the initial text feature chain vector, is the initial text feature value vector, R is the number of heads, is a linear transformation, h represents the head index in the multi-head attention mechanism, i is the index variable, d represents the feature dimension, H is the number of heads in the multi-head attention mechanism, and Q represents related to the query.
[0054] To continue to use the initial graph structure features to enhance text features, the cross-modal text features obtained above and the initial graph structure features are enhanced through the co-attention module, and finally the enhanced text features enhanced by the initial image features and the initial graph result features are obtained
[0055]
[0056] Among them, is the enhanced text feature, is the initial graph structure feature query vector, is the initial graph structure feature, is a linear transformation.
[0057] Step S03, concatenate the enhanced text features, enhanced image features, and enhanced graph structure features through the co-attention network.
[0058] Specifically, after the co-attention network enhances each feature, it fuses the enhanced features. After enhancing the initial text features, initial image features, and initial graph structure based on the multi-modal co-attention mechanism in the aforementioned step S02, three types of enhanced features are finally obtained. Finally, they are concatenated into the final target multi-modal output feature oi :
[0059]
[0060] in, To enhance text features, To enhance image features, To enhance the graph structure features.
[0061] Step S04, performing comparative learning on the target multimodal output features and inputting them into the fully connected layer.
[0062] Specifically, fake news detection requires the model to be able to capture subtle features and differences in text data, which are usually difficult to learn through traditional supervised learning methods. Contrastive learning can help the model learn more discriminative feature representations because it encourages the model to map the differences between real news and fake news into feature space. In this way, the model can better distinguish between real and fake news. The present invention adds contrastive learning at the end of the model, such as Figure 2 As shown in the figure, the third part of the network model architecture is contrastive learning, which is used to bring the probability representations of samples with the same label closer. This makes the representations of positive samples as close as possible while being far away from negative samples. Contrastive learning is introduced through feature learning between different modal data to improve the cross-modal performance of the model.
[0063] The target multimodal feature o i Input the fully connected layer, where the fully connected layer is a two-layer structure: the first layer has a dimension of 512 and uses the ReLU activation function; the second layer has a dimension of 2 and outputs the predicted probability that the news to be detected is false news
[0064]
[0065] Among them, W c is the weight matrix of the fully connected layer, and b is the bias term of the fully connected layer.
[0066] The embodiment of the present invention provides a method and device for detecting fake news based on contrastive learning and multimodal fusion. The method uses a common attention network to modally enhance each feature of the initial text feature, the initial image feature, and the initial graph structure feature with the remaining two modal features; then, the common attention network connects the enhanced text feature, the enhanced image feature, and the enhanced graph structure feature in series; finally, the target multimodal output feature is subjected to contrastive learning and input into the fully connected layer to obtain the predicted probability that the news to be detected is fake news. The embodiment of the present invention fuses the three modal features through a common attention network to obtain enhanced features with three modal information, thereby ensuring the semantic consistency between different modalities and improving the accuracy of fake news detection.
[0067] In the embodiments of the present invention, the differences between real news and fake news are mapped to the feature space through contrastive learning, making the representations of positive samples as close as possible. In the embodiments of the present invention, the probability representations between samples with the same label are brought closer through contrastive learning, which can better distinguish real and fake news and improve the accuracy of fake news detection.
[0068] The embodiments of the present invention introduce the model training process as follows:
[0069] 1. Obtain the dataset and select the training and test datasets
[0070] Collect two public datasets and divide them into training sets and test datasets. For example: the two public datasets can be WEIBO and PHEME. The WEIBO dataset is based on the Weibo platform and covers Weibo information from various fields. Each post contains three elements (i.e., tweet id, text, and image). The PHEME dataset is collected based on 5 breaking news, and each news contains a set of posts. Each dataset contains a large number of labeled texts and images.
[0071] 2. Model training and optimization
[0072] In the model training stage, the embodiments of the present invention first need to initialize the model, including the initialization of weights and biases. To accelerate the training process and avoid overfitting, the pre-trained model weights, especially the weights of ResNet50 and the graph attention network GAT, are adopted. These pre-trained models perform well in their respective tasks and can provide rich feature representations, thus helping to improve the performance of fake news detection.
[0073] Next, use the stochastic gradient descent optimization algorithm to train the model. During the training process, update the weights and biases of the model according to the gradient of the loss function. Among them, the loss function includes classification loss and contrastive learning loss, and the importance of the two can be balanced by adjusting the hyperparameter β.
[0074] The present invention uses the cross-entropy function as the loss function and introduces modality alignment:
[0075]
[0076] where is the probability predicted by the model, y is the true label, λ c and λ a are weight coefficients, is the alignment loss, which can be obtained through cosine similarity.
[0077] Since the cross - entropy loss uses an inter - class competition mechanism to learn information between classes, it only cares about the accuracy of the predicted probability for the correct label and ignores the differences of other incorrect labels, resulting in sub - optimal generalization and instability problems. Good generalization requires capturing the similarity between samples in a class and contrasting samples within a class with samples in other classes.
[0078] Given that the output features of news in a Batch are where N is the number of inputs in the Batch. When setting the output feature o i of the news as the anchor point, the news features O j , O j with the same label are used as positive sample pairs, that is, y i = y j . Where y i and y j are the labels of o i and O j respectively, then the contrastive loss function is:
[0079]
[0080] The final loss function is
[0081]
[0082] where is the loss of the framework, L sup is the contrastive learning loss, and β is a hyper - parameter.
[0083] 3. Model Evaluation and Result Analysis
[0084] After training is completed, the performance of the model is evaluated using metrics such as accuracy, recall, and F1 - score. These metrics can comprehensively reflect the model's ability to identify fake news. By comparing the performance metrics of the model on the test set, the generalization ability and robustness of the model can be evaluated. In addition, the model of the present invention is compared with existing fake news detection models to verify the advantages of the model.
[0085] The experimental results are as Figure 4 shown. The model of the present invention is superior to existing multi - modal fake news detection models in terms of metrics such as accuracy, recall, and F1 - score. This is mainly due to the improvements in feature extraction, multi - modal feature fusion, and contrastive learning of the model. Specifically, by using the features of text, image, and social Figure 3 modes and combining the multi - modal co - attention network and contrastive learning mechanism, the model of the present invention can more comprehensively understand the news content and capture the relevance and complementarity between different modes.
[0086] The following describes the false news detection device based on contrastive learning and multi-modal fusion provided by the embodiments of the present invention. The false news detection device based on contrastive learning and multi-modal fusion described below can be correspondingly referred to the false news detection method based on contrastive learning and multi-modal fusion described above.
[0087] First, in combination with Figure 5 , an introduction to the false news detection device based on contrastive learning and multi-modal fusion is given. As Figure 5 shown, the false news detection device based on contrastive learning and multi-modal fusion may include:
[0088] A feature extraction unit 100 extracts the initial text feature, initial image feature, and initial graph structure feature of the news to be detected through a pre-trained feature extraction model respectively, where the initial graph structure feature is the structure feature of the social graph;
[0089] A feature enhancement unit 200 performs modality enhancement on each of the initial text feature, initial image feature, and initial graph structure feature with the other two modality features through a co-attention network to obtain an enhanced text feature, an enhanced image feature, and an enhanced graph structure feature;
[0090] A feature fusion unit 300 concatenates the enhanced text feature, enhanced image feature, and enhanced graph structure feature through a co-attention network to obtain a target multi-modal output feature;
[0091] A news detection unit 400 performs contrastive learning on the target multi-modal output feature and inputs it into a fully connected layer to obtain the prediction probability that the news to be detected is false news.
[0092] The embodiments of the present invention also provide a storage medium, which can store a program suitable for being executed by a processor, and the program is used to implement each processing flow in the foregoing false news detection solution based on contrastive learning and multi-modal fusion.
[0093] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0094] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other.
[0095] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A fake news detection method based on contrastive learning and multimodal fusion, characterized in that the method include: The initial text features, initial image features and initial graph structure features of the news to be detected are respectively extracted through the pre-trained feature extraction model, wherein the initial graph structure features are the structural features of the social graph; Through the co-attention network, each feature of the initial text feature, the initial image feature and the initial graph structure feature is modally enhanced with the other two modal features to obtain enhanced text features, enhanced image features and enhanced graph structure features; The enhanced text features, enhanced image features, and enhanced graph structure features are connected in series through the co-attention network to obtain the target multimodal output features; The target multimodal output features are contrastively learned and input into the fully connected layer to obtain the predicted probability that the news to be detected is false news.
2. The fake news detection method based on contrastive learning and multimodal fusion according to claim 1 is characterized in that: The process of respectively extracting the initial text features, initial image features and initial graph structure features of the news to be detected by the pre-trained feature extraction model includes: Extract the initial text features of the news to be detected through the pre-trained convolutional neural network; Extract the initial image features of the news to be detected through the pre-trained ResNet50; The initial graph structure features of the news to be detected are extracted through the pre-trained graph attention network.
3. The fake news detection method based on contrastive learning and multimodal fusion according to claim 1 is characterized in that: The process of modality enhancement of the initial text features and the remaining two modality features includes: The initial text features are enhanced based on the initial image features to obtain cross-modal text features; The cross-modal text features are enhanced based on the initial graph structure features to obtain enhanced text features.
4. The fake news detection method based on contrastive learning and multimodal fusion according to claim 3 is characterized in that: The step of enhancing the initial text features based on the initial image features includes: The initial text features are enhanced by the following formula: in, is the cross-modal text feature, is the initial image feature query vector, is the initial image feature, is the transpose of the initial text feature key vector, is the initial text feature key vector, is the value vector of the initial text features, R is the number of heads, is a linear transformation, h represents the head index in the multi-head attention mechanism, i is the index variable, d represents the feature dimension, H is the number of heads in the multi-head attention mechanism, and Q represents the correlation with the query; The cross-modal text features are enhanced by the following formula: in, To enhance text features, is the initial graph structure feature query vector, is the initial graph structure feature, is a linear transformation.
5. A fake news detection device based on contrastive learning and multimodal fusion, characterized in that: The device includes: A feature extraction unit, which extracts initial text features, initial image features and initial graph structure features of the news to be detected respectively through a pre-trained feature extraction model, wherein the initial graph structure features are structural features of the social graph; The feature enhancement unit modally enhances each feature of the initial text feature, the initial image feature, and the initial graph structure feature with the other two modal features through a co-attention network to obtain enhanced text features, enhanced image features, and enhanced graph structure features; The feature fusion unit connects the enhanced text features, enhanced image features, and enhanced graph structure features in series through the co-attention network to obtain the target multimodal output features; The news detection unit conducts comparative learning on the target multimodal output features and inputs them into the fully connected layer to obtain the predicted probability that the news to be detected is false news.
Citation Information
Patent Citations
Collaborative attention network multi-mode rumor detection method fusing image features
CN118211122A
Multi-modal false news detection method based on dynamic propagation social graph
CN118568261A
Cited By
A multi-modal content detection method and device based on feature decoupling and a three-branch classifier
CN122712381A