A method, device, equipment and storage medium for detecting false news

Through a multi-level semantic enhancement of news text and image features, the problem of insufficient accuracy and semantic information acquisition of existing models is solved, and high-accuracy fake news detection is achieved.

CN116756310BActive Publication Date: 2025-07-29QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310548939.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-12
Publication Date
2025-07-29
Estimated Expiration
2043-05-12

AI Technical Summary

Technical Problem

The existing fake news detection models have shortcomings in accuracy, ease of use and interpretability, and have failed to fully obtain the semantic information characteristics of news, which affects the judgment of authenticity of fake news.

Method used

By extracting the text basic features, image CNN features, image entity features and text theme features of the news text, the common attention converter is used for semantic enhancement, and the text representation of visual entities and themes is obtained, and stitched and fused with image CNN features, and finally input the classifier for true and false classification.

Benefits of technology

It significantly improves the accuracy of fake news detection, which is at least 3.3 percentage points better than the existing model, and fully obtains the core semantic information of news.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116756310B_ABST
    Figure CN116756310B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and storage medium for false news detection. The method includes: extracting text-based features FT, image CNN (Convolutional Neural Network) features FV, image entity features FVE, and text topic features FP from the news to be detected; obtaining a text representation FT←VE enhanced by visual entities; obtaining a text representation FT←(VE,p) enhanced by visual entities and topic features; performing an average operation on FT←(VE,p) to obtain the final representation χt of the text, and performing an average operation on the image CNN features FV to obtain the final representation χv of the image; splicing and fusing the obtained semantically enhanced text features χt and the image CNN features χv to obtain a multimodal representation χf; and inputting the multimodal representation χf into a classifier for true / false classification. The present invention can fully obtain a multimodal feature representation with enhanced core semantic information of news content from the news text T and the attached image I for false news detection, significantly improving the accuracy of false news detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fake news detection, and in particular, to a method, device, equipment and storage medium for fake news detection. Background Art

[0002] Early fake news detection methods utilized different multi-modal feature extraction tools to first extract text features and image features, and then spliced and fused them for fake news detection. However, in the process of directly splicing text features and visual features, the same weights were assigned to the text features and visual features, resulting in the failure to obtain important features, and the accumulation of the same information in the text and image formed redundant information, reducing the accuracy of fake news detection.

[0003] Some studies have found that using image information different from text information for the same news event is a common form of fake news. Therefore, detecting the inconsistency of modal features has become the key to fake news detection. Specifically, some scholars converted visual information into text information, and then compared the similarity between visual information and text information for fake news detection, or used a weight-sharing encoder to project visual features and text features into a mutual feature space, and then calculated the similarity of the transformed multi-modal features. Due to the heterogeneity of text and visual information itself, there is a large semantic gap between visual and text features, and it is a great challenge to capture the inconsistency between multi-modalities, making the specific implementation of fake news detection very difficult.

[0004] In summary, although existing various multi-modal fake news detection models have made varying degrees of progress, there are still deficiencies in terms of the accuracy, usability, interpretability, etc. of fake news detection. More importantly, most existing detection models focus on the feature information at the news presentation level, pay insufficient attention to the semantic information of news content, and cannot fully obtain the semantic information features in the news. As a text classification task, semantic information plays a decisive role in judging the authenticity of news. If the semantic information of news is insufficiently obtained, it will greatly affect the judgment of the authenticity of fake news. Summary of the Invention

[0005] The object of the present invention is to provide a method, device, equipment and storage medium for fake news detection, which can fully obtain the core semantic information-enhanced multi-modal feature representation of news content from news text T and additional image I, and perform fake news detection, significantly improving the accuracy of fake news detection.

[0006] On the one hand, the present invention provides a method for fake news detection, including:

[0007] Extracting the text basic feature F in the news to be detected T, Image CNN (Convolutional Neural Network) Feature F V , Image Entity Feature F VE , Text Topic Feature F P ;

[0008] Text Basic Feature F T and Visual Entity Feature F VE are input into the first co-attention transformer to obtain a text representation F enhanced by visual entities T←VE ;

[0009] F T←VE and Topic Feature F P are input into the second co-attention transformer to obtain a text representation F enhanced by visual entities and topic features T←(VE,p) ;

[0010] An average operation is performed on F T←(VE,p) to obtain the final representation χ of the text t, An average operation is performed on the image CNN feature F V to obtain the final representation χ of the image v ;

[0011] The obtained semantically enhanced text feature χ t is concatenated and fused with the image CNN feature χ v to obtain the final multi-modal representation χ f ;

[0012] The obtained multi-modal representation χ of the news f is input into a classifier for true / false classification to generate the probability that the news to be detected is true or false.

[0013] In some embodiments, the step of extracting text basic features includes: inputting the text T in the news to be detected into a pre-trained language representation model to obtain a feature vector representation of the text basic: F T =[W1,…,W n , where W i represents the feature of the i-th word in the text, and n is the length of the text.

[0014] In some embodiments, the step of image CNN feature extraction includes: inputting the image of the news to be detected into a pre-trained network with the VGG19 architecture to obtain a global feature representation F of the image with the same size as the text V

[0015] F V =σ(W v R vgg )(1)

[0016] where, R vggis the visual feature representation obtained from VGG19, and W v is the weight matrix of the fully connected layer, and σ is the activation function used.

[0017] In some embodiments, the step of extracting image entity features includes: obtaining initial visual entities from the image of the news to be detected through public APIs, and inputting the initial visual entities into a pre-trained language representation model to obtain image entity features F VE .

[0018] In some embodiments, the step of extracting text topic features includes: obtaining the topic text information of the news article to be detected through a document topic generation model, and inputting it into the language representation model for encoding to obtain the topic features of the original news content. Then, the topic features are adjusted to a d×1-dimensional representation through a fully connected layer to obtain the feature vector representation of the article topic: F P = [P1,…,P n , where P i represents the feature of the i-th topic of the news article sentence, and n is the number of topics corresponding to the news sentence.

[0019] In some embodiments, the language representation model is a BERT (Bidirectional Encoder Representation from Transformers) model.

[0020] In some embodiments, the document topic generation model is an LDA (Latent Dirichlet Allocation) model.

[0021] In some embodiments, the initial visual entities include person entity E P , location entity E L , and visual concept entity E as general image context C .

[0022] On the other hand, the present invention also provides a fake news detection device, including:

[0023] A basic information feature extraction module, configured to extract text basic features F T , image CNN (Convolutional Neural Network) features F V , image entity features F VE , and text topic features F P ;

[0024] A first semantic enhancement module, configured to input the text basic features F T and the visual entity features F VE into the first co-attention transformer, so as to obtain the text representation F enhanced by visual entitiesT←VE ;

[0025] The second semantic enhancement module is used to input F T←VE and the topic feature F P into the second co-attention transformer to obtain a text representation F enhanced by visual entities and topic features T←(VE,p) ;

[0026] The average operation module is used to perform an average operation on F T←(VE,p) to obtain the final representation χ of the text, perform an average operation on the image CNN feature F t to obtain the final representation χ of the image V ; v ;

[0027] The global feature fusion module splices and fuses the obtained semantically enhanced text feature χ t and the image CNN feature χ v to obtain the final multi-modal representation χ f ;

[0028] The classification module inputs the obtained multi-modal representation χ of the news f into a classifier for true / false classification to generate the probability that the news to be detected is true or false news.

[0029] On the other hand, the present invention also provides an electronic device, including: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory for executing the above-mentioned false news detection method.

[0030] On the other hand, the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the above-mentioned false news detection method.

[0031] The present invention inputs the visual entities obtained from the news and the text basic information features together into the co-attention mechanism to enhance the entity object information described in the news text, obtains the text features with the first semantic enhancement, and then inputs the text features enhanced by the entity and the extracted text topic features together into the co-attention mechanism to obtain the text features with the second semantic enhancement. Through the above-mentioned two-step progressive enhancement of different aspects of the semantic information in the text information, a text feature representation containing the core semantics is obtained. The text feature representation is spliced and fused with the visual CNN features to obtain the final multi-modal news representation, and the obtained multi-modal representation of the news is input into a classifier for true / false classification, significantly improving the accuracy of false news detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0033] Figure 1 It is a flowchart of a false news detection method provided by an embodiment of the present invention;

[0034] Figure 2 It is a schematic diagram of a false news detection device provided by an embodiment of the present invention;

[0035] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0036] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] Please refer to Figure 1 , an embodiment of the present invention provides a false news detection method, including the following steps:

[0038] S1: Extract the text basic feature F of the news to be detected T , the image CNN (Convolutional Neural Network) feature F V , the image entity feature F VE , the text topic feature F P ;

[0039] Specifically, step S1 includes:

[0040] S1.1: Extract the text basic feature, specifically:

[0041] Input the text T in the news to be detected into the pre-trained language representation model to obtain the feature vector representation of the text basic: F T = [W1,..., W n , where W i represents the feature of the i-th word in the text, and n is the length of the text.

[0042] S1.2: Extract the image CNN feature, specifically:

[0043] Input the image of the news to be detected into the pre-trained network of the VGG19 architecture to obtain the global feature representation F of the image with the same size as the text V

[0044] F V = σ(W vf ·R vgg ) (1)

[0045] where R vgg is the visual feature representation obtained from VGG19, and W vf is the weight matrix of the fully connected layer, and σ is the activation function used

[0046] The present invention uses a pre-trained network of the VGG19 architecture trained on the ImageNet database to extract the visual feature R of the input image from the output of the last layer of VGG19 vgg , and then obtains the global feature representation F of the image with the same size as the text through multiple fully connected layers V .

[0047] S1.3: Extract image entity features, specifically:

[0048] Obtain the initial visual entities from the image of the news to be detected through public APIs, and input the initial visual entities into the pre-trained language representation model to obtain the image entity feature F VE .

[0049] The visual entities contained in the news image usually represent important information of the news content, and play an important role in the understanding of the news semantic content and the detection of fake news. The present invention extracts highly representative visual entities, including entity person entity E P , location entity E L and visual concept entity E as general image context C .

[0050] S1.4: Extract text topic features, specifically:

[0051] Obtain the topic text information of the news article to be detected through the document topic generation model, and input it into the language representation model for encoding to obtain the topic features of the original news content. Then, adjust the topic features to a d×1-dimensional representation through the fully connected layer to obtain the feature vector representation of the article topic: F P = [P1,…,P n , where P i represents the feature of the i-th topic of the news article sentence, and n is the number of topics corresponding to the news sentence

[0052] In the embodiment of the present invention, the language representation model is a BERT (Bidirectional Encoder Representation from Transformers) model, and the document topic generation model is an LDA (Latent Dirichlet Allocation) model.

[0053] S2: The text basic feature F T and the visual entity feature F VE are input into the first co-attention transformer to obtain the text representation F enhanced by visual entities T←VE ;

[0054] By this step, the text features with the first semantic enhancement are obtained. Merely obtaining the text representation enhanced by visual entities cannot fully achieve the purpose of semantic feature enhancement. It is also necessary to enhance the event topic features of the news.

[0055] S3: Input F T←VE and the topic feature F P into the second co-attention transformer to obtain the text representation F enhanced by visual entities and topic features T←(VE,p) ;

[0056] By this step, the text features with the second semantic enhancement are obtained. The text feature representation containing the core semantics is obtained through the above enhancement of different aspects of the semantic information in the text information.

[0057] S4: Perform an average operation on F T←(VE,p) to obtain the final representation χ of the text t . To ensure the integrity of the image features, an average operation is performed on the image CNN feature F V to obtain the final representation χ of the image v ; The average operation here refers to taking the average value of the obtained feature vectors.

[0058] S5: Concatenate and fuse the obtained semantically enhanced text feature χ t and the image CNN feature χ v to obtain the final multi-modal representation χ f ;

[0059] χ f = concat(χ t , χ v ), (2)

[0060] S6: Input the obtained multi-modal representation χ of the news f into the classifier for true / false classification to generate the probability that the news to be detected is true or false.

[0061] The above-obtained multimodal representation x f , which models the input multimodal news from two different aspects: local semantic feature enhancement and global feature concatenation. The present invention uses a fully connected layer with softmax activation to project the multimodal feature vector x f into the target space of two categories: real news and fake news, and obtains the probability distribution:

[0062] ρ = softmax(W χf + b), (3)

[0063] where ρ = [ρ0, ρ1] is the prediction vector, and ρ0 and ρ1 respectively indicate the prediction probabilities of the labels 0 (real news) and 1 (fake news). W is the weight matrix, and b is the bias term.

[0064] The accuracy of the predicted label is continuously improved by minimizing the cross-entropy loss function, and the formula is as follows:

[0065] L ρ = -[ylogρ0 + (1 - y)logρ1], (4)

[0066] where y ∈ {0, 1} represents the true value label.

[0067] The fake news detection method provided by the embodiments of the present invention extracts the text basic feature F T , the image CNN (Convolutional Neural Network) feature F V , the image entity feature F VE , and the text topic feature F P from the news to obtain the overall information features at the news presentation level. Then, the image entity and the text basic information feature are input into the co-attention mechanism together to enhance the entity object information described in the news text, and the text feature with the first semantic enhancement is obtained. Then, the text feature enhanced by the entity and the extracted text topic feature are input into the co-attention mechanism together to obtain the text feature with the second semantic enhancement. Through the above two-step progressive enhancement of different aspects of the semantic information in the text information, the text feature representation containing the core semantics is obtained. The text feature representation is spliced and fused with the visual CNN feature to obtain the final multimodal news representation, and the obtained multimodal representation of the news is input into the classifier for true and false classification, which significantly improves the accuracy of fake news detection.

[0068] The following is a comparison experiment of the effectiveness and high accuracy of the fake news detection method of the present invention, that is, the multi-level semantic enhanced fake news detection method (abbreviated as MLSED), with nine baseline models on two datasets of Twitter and Weibo.

[0069] 1) Datasets

[0070] Twitter Dataset: It includes 7,014 fake news articles and 5,963 real news articles;

[0071] Weibo Dataset: It includes 4,665 fake news posts and 4,689 real news posts with corresponding images.

[0072] 2) Baseline Models

[0073] Several representative methods including single-modal and multi-modal methods were selected, and also included variants of the MLSED model. As follows:

[0074] Single-modal methods:

[0075] BERT: Use pre-trained BERT to obtain the representation of news text, and use a fully connected layer for classification.

[0076] VGG19: Fine-tune VGG19 to model news image information and then classify.

[0077] Multi-modal methods:

[0078] EANN: Designed an auxiliary task, event discrimination. The event discriminator takes the concatenated multi-modal news information as input and outputs the category of the event to better understand the multi-modal information. For fair comparison, the event discriminator was removed.

[0079] SpotFake: Use pre-trained BERT to obtain text features, use VGG19 visual features, and concatenate the obtained text features and visual features for classification.

[0080] MVAE: Proposed a multi-modal variational autoencoder to learn the shared representation of text and images. Through the reconstruction task, better integrate the multi-modal information of news. Use the same model as the original text.

[0081] SAFE: Convert visual information into text information, and map the text information and visual information to the same vector space through a fully connected layer. Then compare the similarity between the visual information and the text information as an auxiliary loss for fake news classification.

[0082] MCAN: Use VGG to model the spatial information of pictures, use CNN to model the frequency domain information of pictures, and use co-attention to fuse the frequency domain information and spatial information to obtain a better image representation.

[0083] MLSED-ve: The MLSED model without entity enhancement.

[0084] MLSED-p: The MLSED model without theme feature enhancement.

[0085] 3) Experimental results

[0086] As can be seen from the experimental results shown in Table 1, the proposed MLSED model of the present invention outperforms other methods on both datasets, confirming that MLSED can fully capture the semantic information of news to detect fake news. Specifically, MLSED exceeds the corresponding state-of-the-art method by at least 3.3 percentage points in terms of accuracy.

[0087] The performance of the MLSED-p model is better than that of the MLSED-ve model, proving that the entity object information in the semantic information plays a greater role than the event theme information and is crucial for fake news detection.

[0088]

[0089] Table 1

[0090] Correspondingly, please refer to Figure 2 , the embodiment of the present invention also provides a fake news detection device, including:

[0091] The basic information feature extraction module 10 is used to extract the text basic feature F T , the image CNN (Convolutional Neural Network) feature F V , the image entity feature F VE , and the text theme feature F P ;

[0092] The first semantic enhancement module 20 is used to input the text basic feature F T and the visual entity feature F VE into the first co-attention transformer to obtain the text representation F T←VE enhanced by the visual entity;

[0093] The second semantic enhancement module 30 is used to input F T←VE and the theme feature F P into the second co-attention transformer to obtain the text representation F T←(VE,p) enhanced by the visual entity and the theme feature;

[0094] The average operation module 40 is used to perform an average operation on F T←(VE,p) to obtain the final representation χ t of the text, and perform an average operation on the image CNN feature F V to obtain the final representation χ v of the image;

[0095] The global feature fusion module 50 fuses the obtained semantically enhanced text feature χ t with the image CNN feature χ v by concatenation to obtain the final multi-modal representation χ f ;

[0096] The classification module 60 inputs the obtained multi-modal representation χ of the news f into a classifier for true / false classification to generate the probability that the news to be detected is true or false news.

[0097] The false news detection device in the embodiment of the present invention corresponds to the above false news detection method one by one, and will not be elaborated here one by one.

[0098] Correspondingly, please refer to Figure 3 , the embodiment of the present invention also provides an electronic device, including: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory for executing the above false news detection method.

[0099] The embodiment of the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the above false news detection method.

[0100] The above has introduced in detail a false news detection method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

[0101] Each embodiment in this application is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

Claims

1. A false news detection method, characterized in that, Including: Extract the basic text features F of the news to be detected T and the image CNN (Convolutional Neural Network) features F V and the image entity features F VE and the text topic features F P ; Text-based feature F T and visual entity feature F VE are input into the first co-attention transformer to obtain a text representation F enhanced by visual entities T←VE ; Input F T←VE and the topic feature F P into the second co-attention transformer to obtain the text representation F enhanced by visual entities and topic features T←(VE,p) ; Average F T←(VE,p) to obtain the final representation χ of the text t, Average the CNN features F of the image V to obtain the final representation χ of the image v ; The obtained semantically enhanced text feature χ t is concatenated and fused with the image CNN feature χ v to obtain the final multimodal representation χ f ; Input the multi-modal representation χ of the obtained news f into the classifier for true / false classification to generate the probability that the news to be detected is true or false news.

2. The false news detection method according to claim 1, wherein The specific steps for extracting the basic features of the text are as follows: Input the text T in the news to be detected into a pre-trained language representation model to obtain the feature vector representation of the basic text: F T = [W1,…,W n , where W i represents the feature of the i-th word in the text, and n is the length of the text.

3. The false news detection method according to claim 1, characterized in that, The specific steps for image CNN feature extraction include: inputting the image of the news to be detected into the pre-trained network with the VGG19 architecture to obtain the global feature representation F of the image with the same size as the text V F V = σ (W vƒ • R vgg ) (1) Among them, R vgg is the visual feature representation obtained from VGG19, and W vƒ is the weight matrix of the fully connected layer, and σ is the activation function used.

4. The false news detection method according to claim 1, characterized in that The specific steps for extracting image entity features include: obtaining initial visual entities from the images of the news to be detected through public APIs, and inputting the initial visual entities into a pre-trained language representation model to obtain image entity features F VE .

5. The false news detection method according to claim 1, characterized in that, The specific steps for extracting the topic features of the text include: obtaining the topic text information of the news article to be detected through the document topic generation model, and inputting it into the language representation model for encoding to obtain the topic features of the original news content. Then, the topic features are adjusted to a d×1-dimensional representation through the fully connected layer to obtain the feature vector representation of the article topic: F P = [P1,…,P n , where P i represents the feature of the i-th topic of the news article sentence, and n is the number of topics corresponding to the news sentence.

6. The false news detection method according to claim 2, characterized in that, The language representation model is a BERT (Bidirectional Encoder Representation from Transformers) model.

7. The false news detection method according to claim 4, wherein The initial visual entity includes a person entity E P , a location entity E L , and a visual concept entity E as a general image context C .

8. A device for false news detection and prediction, characterized in that, Including: The basic information feature extraction module is used to extract the text basic feature F in the news to be detected T , the image CNN (Convolutional Neural Network) feature F V , the image entity feature F VE , the text theme feature F P ; The first semantic enhancement module is used to input the text basic feature F T and the visual entity feature F VE into the first co-attention transformer to obtain the text representation F enhanced by visual entities T←VE ; The second semantic enhancement module is used to input F T←VE and the topic feature F P into the second co-attention transformer to obtain the text representation F enhanced by visual entities and topic features T←(VE,p) ; An average operation module for averaging F T←(VE,p) to obtain the final representation χ of the text t , and averaging the image CNN feature F V to obtain the final representation χ of the image v ; Global feature fusion module, which fuses the obtained semantically enhanced text features χ t with the image CNN features χ v by concatenation to obtain the final multi-modal representation χ f ; A classification module that inputs the multi-modal representation χ of the obtained news f into a classifier for true / false classification to generate the probability that the news to be detected is true or false news.

9. An electronic device, characterized in that, Including: A memory storing executable program code; A processor coupled to the memory; The processor invokes the executable program code stored in the memory to execute the fake news detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program causes a computer to execute the fake news detection method according to any one of claims 1 to 7.