False News Detection Method Based on Generated Comments and Multi-Perspective Comprehensive Analysis
By combining the analysis and generation capabilities of large language models in fake news detection, multi-view comprehensive analysis and adaptive comment aggregator are used to solve the problem of limited detection performance in the existing technology, and more efficient and robust fake news detection is achieved.
Patent Information
- Application Number
- CN202510389715.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art fails to effectively combine the analysis and generation capabilities of large language models in false news detection, and multimodal detection ignores the semantic correlation between multimodals at different angles, resulting in the impact of detection performance.
Using a false news detection method based on generative comments and multi-view comprehensive analysis, we generate and analyze comments through the No. 1 and No. 2 large language models, combining multi-view feature extraction and adaptive comment aggregator, multi-modal feature fusion and emotional difference analysis are carried out to improve detection accuracy and robustness.
Through the comprehensive multi-angle detection and the comprehensive utilization of comment information, the performance and efficiency of false news detection are significantly improved, and the generalization ability and robustness of the model are enhanced.
Smart Images

Figure CN119917748B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a false news detection method based on generating comments and multi-perspective comprehensive analysis. Background Art
[0002] Current social media posts usually have multiple interconnected modes. Therefore, it is no longer possible to distinguish true from false news by single-modal detection alone. For example, real images can be combined with false news, or real descriptive language can be combined with tampered images. Therefore, it is necessary to comprehensively analyze false news from multiple modal perspectives. At the same time, with the emergence of various large language models, the use of large language models for multimodal false news detection has also become an important direction. Large language models can effectively capture the parts of news content that are inconsistent with common sense and logic with their rich knowledge and keen observation, thereby improving the accuracy and robustness of false news. In addition, the powerful text generation ability of large language models can provide some information that cannot be obtained by conventional detection, such as comments from silent users. Since comment information can play a very important role in detecting fake news, and currently most of the comment data relies on crawling from traditional social platforms, a major problem faced is that the comment information is not comprehensive enough. Due to the influence of social psychology, confidentiality and other factors, most of the collected comments come from active users, and some special users are missing. Therefore, the model trained based on this comment information lacks generalization and robustness. The professional knowledge and psychological knowledge of the large language model can well simulate the comments made by these special users. Secondly, through the adaptive comment aggregator and comment sentiment analyzer designed by us, the important information of the generated comments can be used to make up for the missing information of the real comments, and the emotional differences between the two comments can be compared, which can greatly enhance the generalization ability of the model and improve the detection efficiency.
[0003] Current research methods still face two major challenges. First, although large language models have been used for fake news detection, they only use the ability of large language models to analyze news or generate text, and do not combine the two to improve the ability to detect fake news. The ability of large language models to process natural language is not only about understanding text, but also about analyzing and applying it. The potential of large language models needs to be further developed. Secondly, although many existing methods extract multimodal features, they ignore the semantic relevance between multimodal features from different angles. They will misjudge news as fake news simply because the features between the modalities are different, affecting the performance of the detection model. Summary of the invention
[0004] In view of the deficiencies of the prior art, the present invention provides a false news detection method based on generated comments and multi-perspective comprehensive analysis to solve the problems mentioned in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions: A false news detection method based on generated comments and multi-perspective comprehensive analysis. Further, it includes the following steps:
[0006] Step S1: Construct news samples and perform data preprocessing on the news samples;
[0007] After data preprocessing, text feature data, image feature data, and real comment data are obtained and labeled with emotion tags;
[0008] Step S3: Based on the prompt template, use the first large language model to generate comments of several simulated users for the news samples through prompts and comprehensively analyze the text content of the news samples and the generated comments to obtain the text data of the news content analysis;
[0009] Step S4: Based on the prompt template, use the second large language model to comprehensively analyze the text data obtained after the content analysis of the first large language model, the text content of the news samples, and the comments generated by the large language model to obtain the news principle feature data;
[0010] Step S5: The multi-perspective feature extraction model extracts features from the news principle feature data, text feature data, real comment feature data, generated comments, and image feature data to obtain the corresponding encoded feature vectors;
[0011] Step S6: The adaptive comment aggregator fuses the generated comment encoded feature vectors and the real comment encoded feature vectors to obtain the comprehensive comment encoded feature vectors, and then fuses the comprehensive comment encoded feature vectors and the text encoded feature vectors to obtain the text and comment feature fusion vectors;
[0012] Step S7: Input the generated comment encoded feature vectors and the real comment encoded feature vectors into the comment sentiment analyzer to obtain the sentiment difference coefficient by calculating the discrete Gaussian distribution and the KL divergence;
[0013] Step S8: The multi-perspective feature fusion model fuses the text encoded feature vectors, news principle encoded feature vectors, comprehensive comment encoded feature vectors, image encoded feature vectors, and text-image modality encoded feature vectors to obtain the fused feature vectors;
[0014] Step S9: Aggregate the encoded feature vectors, fused feature vectors, text and comment feature fusion vectors, and sentiment difference coefficient to obtain the final aggregated encoded feature vectors, and the classifier module predicts the true or false prediction labels of the news samples corresponding to the final aggregated encoded feature vectors.
[0015] Further, the process of constructing news samples and preprocessing the news samples in step S1 is specifically as follows:
[0016] Step S11: Crawl and preprocess the publicly available social media dataset through web crawler technology to obtain the dataset required for model training. A single piece of data in the news dataset is a news sample.
[0017] Step S12: Preprocess the dataset, delete news samples lacking text content, image descriptions, or with unclear text descriptions. The obtained single news data samples respectively include text feature data, image feature data, and sentiment labels. Randomly divide the news dataset into a training set and a test set according to preset rules.
[0018] Further, the process of using the first large language model to generate comments and analyze the news text content in step S3 is specifically as follows:
[0019] Step S31, use the Zhipu Qingyan open-source large language model as the first large language model to simulate social platform users based on the news samples in the dataset and generate comments; specifically:
[0020] First, give a prompt template, and the specific prompt template is:
[0021] This is [news text content]. Suppose you are a social platform user; you can randomly select user attributes from [Gender: male, female], [Age: under 18 years old, 18 to 29 years old, 30 to 49 years old, 50 to 64 years old, over 64 years old], [Educational level: not having received higher education, receiving higher education, having completed higher education], [Stance on the news content: support, oppose, neutral], and then generate 8 corresponding user comments for this news according to the user attributes you selected. The word limit for each comment is within 50 words.
[0022] Step S32: After obtaining the corresponding generated comments according to the prompt template in step S31, continue to use the Zhipu Qingyan open-source large language model as the first large language model to analyze the news samples and the generated comments.
[0023] The prompt template for analysis is:
[0024] This is [news content] and the corresponding news comment [comment content]. Now, please act as a fake news analysis expert to analyze this news based on the news content and the comment content. You can analyze according to the source, time, background, logic, or common sense of the news. You don't need to give a judgment result. Just give your analysis. The word limit is within 200 words.
[0025] Further, the process of using the second large language model to further analyze the analysis content, generated comments, and news content of the first large language model to eliminate a part of the hallucinations in step S4 is specifically as follows:
[0026] Use the open-source LLaMa large language model as the second large language model to re-analyze the news text content, the comments generated by the Zhipu Qingyan open-source large language model, and the analysis of the news text content.
[0027] The prompt template for the re-analysis is:
[0028] This is the news content [news text content], the comment content [generated comment], the analysis results of other large language models [news analysis content of the Zhipu Qingyan large language model]. Now, as a fake news analysis expert, please make full use of your knowledge to analyze this news. You should be logical, meet professional standards, and maintain your independent thinking. You can reflect on or supplement the analysis results of other large models, especially the hallucination analysis. You don't need to judge the results, just give your analysis within 300 words to obtain the final analysis of the large language model news text content, that is, the news principle feature data.
[0029] Further, the multi-perspective feature extraction model in step S5 includes a parallel pre-trained RoBERTa model, a pre-trained CLIP model, and a pre-trained SwinT model;
[0030] The specific steps are as follows:
[0031] Step S51: Train the RoBERTa model to process the text feature data to obtain text encoding feature vectors;
[0032] The text feature data is represented as ; where represents the text feature data, represents the first text feature data, represents the second text feature data, represents the nth text feature data;
[0033] The text encoding feature vector is represented as ; where represents the text encoding feature vector, represents the first text encoding feature vector, represents the second text encoding feature vector, represents the nth text encoding feature vector;
[0034] Step S52, the pre-trained RoBERTa model processes the news principle feature data to obtain news principle encoding feature vectors;
[0035] The news principle feature data is represented as ; where represents the news principle feature data, represents the first news principle feature data, represents the second news principle feature data, represents the nth news principle feature data;
[0036] The news principle coding feature vector is represented as ; where represents the news principle coding feature vector, represents the first news principle coding feature vector, represents the second news principle coding feature vector, represents the nth news principle coding feature vector;
[0037] Step S53: The pre-trained SwinT model processes the image feature data to obtain the image coding feature vector;
[0038] Among them, the image feature data is represented as Iϵ ; and are the height and width of the image feature data respectively; I = ; I represents the image feature data, represents the first image feature, represents the second image feature, represents the nth image feature;
[0039] The image coding feature vector is represented as ; represents the image coding feature vector, represents the first image coding feature vector, represents the second image coding feature vector, represents the nth image coding feature vector;
[0040] Step S54: The pre-trained RoBERTa model processes the generated comment feature data to obtain the generated comment coding feature vector;
[0041] The generated comment feature data is represented as ; represents the first comment feature data, represents the second comment feature data, represents the nth comment feature data;
[0042] The generated comment coding feature vector is represented as ; Represents the first comment encoding feature vector, Represents the second comment encoding feature vector, Represents the nth comment encoding feature vector;
[0043] Step S55: The pre-trained RoBERTa model processes the real comment feature data to obtain the real comment encoding feature vector;
[0044] The real comment feature data is represented as ; Represents the first comment feature data, Represents the second comment feature data, Represents the nth comment feature data;
[0045] The real comment encoding feature vector is represented as ; Represents the first comment encoding feature vector, Represents the second comment encoding feature vector, Represents the nth comment encoding feature vector;
[0046] Step S56: The pre-trained CLIP model processes the text feature data and the image feature data to obtain the text-image modality encoding feature vector , and the text-image modality encoding feature vector is divided into the text-image modality text encoding feature vector , and the text-image modality image encoding feature vector .
[0047] Furthermore, the specific processing flow of the adaptive comment aggregator in step S6 is as follows:
[0048] Step S61: Fuse the generated comment encoding feature vector and the real comment encoding feature vector to obtain the comprehensive comment encoding feature vector;
[0049] Step S62: After fusing the text encoding feature vector and the comprehensive comment encoding feature vector, the text-comment feature fusion vector is obtained: specifically:
[0050] Step S621: Calculate the cosine similarity score between the generated comment encoding feature vector and the real comment encoding feature vector , and then multiply it by the generated comment encoding feature vector to obtain the comprehensive comment feature vector , so as to evaluate the important information of the generated comment to make up for the missing information in the real comment;
[0051] Step S622: Calculate the text encoding feature vector and the comment encoding feature vector The association matrix between , so as to obtain important information in the text recognition comments and ensure the information transmission between the text and the comments;
[0052] Step S623: Take the association matrix as the key vector and the comment encoding feature vector as the value vector, and calculate the aggregated information of the key vector and the value vector; Then, use the non-linear activation function to normalize the aggregated information to obtain the corresponding attention weight coefficients; Use the attention weight coefficients to perform weighted summation on the value vector V to obtain the comment aggregated information;
[0053] Step S624: Suppress the negative comments in the comment aggregated information through the fusion gating mechanism to obtain the fusion gating features, so as to ensure that the generated negative comments do not interfere with the model detection results;
[0054] Step S625: Process the fusion gating features through residual connection, splice them with the text encoding feature vector to retain the original features in the text; Obtain the text and comment feature fusion vector;
[0055] The processing flow of Step S62 is expressed as:
[0056] ;
[0057] ;
[0058] ;
[0059] ;
[0060] ;
[0061] ;
[0062] In the formula, is the cosine similarity score between the generated comment feature vector and the real comment feature vector, is the modulus length of the real comment encoding feature vector, is the modulus length of the generated comment encoding feature vector, is the comprehensive comment encoding feature vector, means changing the input part into a query vector, means changing the input part into a key vector, means the weight matrix, means the ReLU activation function, means the fusion gating feature, means the non-linear activation function, d represents the preset dimension, and T represents the transpose. Represents the text and comment feature fusion vector, Denotes element-wise multiplication, Represents the product operation, Represents the comment aggregation information.
[0063] Furthermore, step S7 is specifically as follows:
[0064] Step S71: Calculate the discrete Gaussian distribution by generating the comment sentiment label and the true comment sentiment label to obtain the sentiment distribution of the generated comment coding feature vector and the sentiment distribution of the true comment coding feature vector ;
[0065] Step S72: Calculate the divergence of the sentiment distributions of the generated comment coding feature vector and the true comment coding feature vector to compare the sentiment differences between the generated comment coding feature vector and the true comment coding feature vector, obtaining the sentiment difference score ;
[0066] Step S73: Stack the sentiment difference scores and use the mechanism and the non-linear activation function to obtain the final sentiment difference coefficient between the generated comment and the true comment ;
[0067] The processing flow of step S7 is expressed as:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] In the formula, is the generated comment sentiment label, is the true comment sentiment label, is the preset parameter, and are the sentiment distances of the generated comment coding feature vector and the true comment coding feature vector, is divergence, represents the natural constant to the exponential function, represents network, represents a non-linear activation function, represents a concatenation operation, i represents the index, and n represents the total number of generated comments.
[0074] Furthermore, the multi-view feature fusion model in step S8 includes a concatenated co-attention layer and a mapping layer; the co-attention layer consists of two connected attention blocks and a fully connected layer;
[0075] Step S8 is specifically as follows:
[0076] After the text encoding feature vector and the news principle encoding feature vector are fused, a text-principle feature fusion vector is obtained. The text encoding feature vector and the comprehensive comment encoding feature vector are used to obtain a text-comment feature fusion vector. After the text encoding feature vector and the image encoding feature vector are fused, a text-image feature fusion vector is obtained. After the text-image modality encoding feature vector extracted by the pre-trained CLIP model is fused, a text-image cross-modal feature fusion vector is obtained. The text-principle feature fusion vector, the text-comment feature fusion vector, the text-image feature fusion vector, and the text-image cross-modal feature fusion vector are fused to obtain a cross-modal feature alignment vector. Specifically:
[0077] Step S81: After the text encoding feature vector and the news principle encoding feature vector are fused, a text-principle feature fusion vector is obtained, specifically:
[0078] Step S811, input the text encoding feature vector and the news principle encoding feature vector together into the first two attention blocks in the co-attention layer;
[0079] Step S812, in the first attention block, the text encoding feature vector serves as the query vector Q, and the news principle encoding feature vector serves as the key vector K and the value vector V to calculate the semantic similarity score; use the non-linear activation function to normalize the attention score to obtain the corresponding attention weight coefficient; then use the attention weight coefficient to perform weighted summation on the value vector V to obtain the text-principle weight vector ;
[0080] Step S813, in the second attention block, the news principle encoding feature vector serves as the query vector Q, and the text encoding feature vector serves as the key vector K and the value vector V. Use the query vector Q and the key vector K to calculate the semantic similarity score; again use the non-linear activation function to normalize the attention score to obtain the corresponding attention weight coefficient; use the attention weight coefficient to perform weighted summation on the value vector V to obtain the principle-text weight vector ;
[0081] Step S814, text principle weight vector and the principle text weight vector are concatenated and input into the fully connected layer to obtain the text and principle feature fusion vector:
[0082] The processing flow of steps S811 to S814 is expressed as:
[0083] Q = × , K = × , V = × ;
[0084] ;
[0085] ;
[0086] ;
[0087] In the formula, means changing the input part into a query vector, means changing the input part into a key vector, means changing the input part into a value vector, means the execution process of an attention mechanism, means a non-linear activation function, d represents the preset dimension of the co-attention layer, T represents transpose, represents the text and principle feature fusion vector, represents the co-attention layer, represents the concatenation operation;
[0088] Step S82: After fusing the text encoding feature vector and the image encoding feature vector, the text and image feature fusion vector is obtained: Specifically:
[0089] Step S821, input the text encoding feature vector and the image encoding feature vector together into the first two attention blocks in the co-attention layer;
[0090] Step S822, in the first attention block, the text encoding feature vector is used as the query vector Q, and the image feature encoding feature vector is used as the key vector K and the value vector V to calculate the semantic similarity score; the attention score is normalized using the non-linear activation function to obtain the corresponding attention weight coefficient; then the attention weight coefficient is used to perform weighted summation on the value vector V to obtain the text image weight vector ;
[0091] Step S823, in the second attention block, the image encoding feature vector serves as the query vector Q, and the text encoding feature vector serves as the key vector K and the value vector V. Calculate the semantic similarity score using the query vector Q and the key vector K; then use the non-linear activation function to normalize the attention score to obtain the corresponding attention weight coefficient; use the attention weight coefficient to perform weighted summation on the value vector V to obtain the image-text weight vector ;
[0092] Step S824, the text-image weight vector and the image-text weight vector are concatenated and then input into the fully connected layer to obtain the text and image feature fusion vector,
[0093] The processing flow of Steps S821 to S824 is expressed as:
[0094] Q = × K = × V = × ;
[0095] ;
[0096] ;
[0097] ;
[0098] In the formula, represents changing the input part into the query vector, represents changing the input part into the key vector, represents changing the input part into the value vector, represents the execution process of an attention mechanism, represents the non-linear activation function, d represents the preset dimension of the co-attention layer, T represents the transpose, represents the text and image feature fusion vector, represents the co-attention layer, represents the concatenation operation;
[0099] Furthermore, Step S8 further includes:
[0100] Step S83: After fusing the text and image modality encoding feature vectors extracted by the pre-trained CLIP model, obtain the text-image cross-modal feature fusion vector; specifically:
[0101] Step S831: At the first attention operation, use the text encoding feature vector of the text-image modality as the query vector Q, and use the image encoding feature vector of the text-image modality as the key vector K and the value vector V, calculate the semantic relationship between the text encoding feature vector and the image encoding feature vector of the text-image modality, and obtain the output representation of the first attention operation ; At the second attention operation, use the image encoding feature vector of the text-image modality as the query vector Q, and use the text encoding feature vector of the text-image modality as the key vector K and the value vector V, and obtain the output representation of the second attention operation ;
[0102] Step S832: Perform average pooling operations on the outputs of the previous and subsequent attention operations respectively to obtain the feature representation after average pooling, and then concatenate them to obtain the CLIP feature fusion representation : The processing flow of Steps S831~S832 is expressed as:
[0103] ;
[0104] ;
[0105] In the formula, represents the cross-modal feature fusion vector of text and image;
[0106] Step S84: After fusing the text and principle feature fusion vector, the text and comment feature fusion vector, the text and image feature fusion vector, and the cross-modal feature fusion vector of text and image, obtain the cross-modal feature alignment vector; specifically:
[0107] Step S841, input the text encoding feature vector and the image encoding feature vector into the fully connected layer and then into the co-attention layer;
[0108] Step S842, when performing the first attention operation, use the text encoding feature vector as the query vector Q, and use the image encoding feature vector as the key vector K and the value vector V, calculate the semantic relationship between the text encoding feature vector and the image encoding feature vector, and obtain the output representation of the first attention operation ; When performing the second attention operation, use the image encoding feature vector as the query vector, and use the text encoding feature vector as the key vector and the value vector, and obtain the output representation of the second attention block ;
[0109] Step S843: The outputs of the previous and subsequent attention operations are respectively input into the average pooling layer for average pooling operations, and then concatenated to obtain a cross-modal feature fusion representation. ; The processing flow of Steps S841 to S843 is expressed as:
[0110] ;
[0111] ;
[0112] In the formula, represents the process of average pooling, represents the operation process of the co-attention layer;
[0113] Step S844: The cross-modal feature fusion representation and the text-image cross-modal feature fusion vector are concatenated and then input into the mapping layer for fusion to obtain a cross-modal mapping representation ;
[0114] Step S845: Use the cross-modal semantic similarity score to adjust the key weights between multiple modalities, and at the same time calculate the semantic similarity relationship between the text encoding feature vector and the image encoding feature vector of the text-image modality to obtain the cross-modal semantic similarity score;
[0115] Step S846: Multiply the cross-modal mapping representation by the cross-modal semantic similarity score to obtain a cross-modal feature alignment vector ; The processing flow of Steps S844 to S846 is expressed as:
[0116] ;
[0117] ;
[0118] ;
[0119] In the formula, is the cross-modal mapping representation, MPP is the mapping layer, is the cross-modal semantic similarity score, represents the norm of the text encoding feature vector of the text-image modality, represents the norm of the image encoding feature vector of the text-image modality, represents the transposed matrix of the image encoding feature vector of the text-image modality.
[0120] Furthermore, Step S9 is specifically:
[0121] Aggregate the text encoding feature vector, image encoding feature vector, comprehensive review encoding feature vector, review sentiment difference coefficient, news principle encoding feature vector, text and principle feature fusion vector, text and image feature fusion vector, text and review feature fusion vector, text-image cross-modal feature fusion vector, and cross-modal feature alignment vector to obtain the final aggregated encoding feature vector, and input it into the classifier module to output the true / false prediction label of the news sample; specifically;
[0122] ;
[0123] ;
[0124] In the formula, represents the prediction label of the classifier module, represents the fully connected layer, represents the cross-entropy loss function, represents the true label of the sample.
[0125] Compared with the existing technologies, the present invention has the following beneficial effects:
[0126] (1) Inspired by assessment methods such as year-end summaries and final evaluations in daily life, the present invention introduces three perspectives for comprehensively analyzing the content of true and false news, namely personal detection, peer detection, and comprehensive detection. Personal detection is the analysis of news content containing self-generated comments by the first large language model. Peer detection is the re-analysis of the analysis results of the first large language model containing news content and generated comments by the second large language model to obtain the final news principle. Comprehensive detection is the joint detection of news content from multiple modalities such as text, image, comment, and news principle. The multi-angle comprehensive detection has improved detection performance compared with traditional multi-modal detection methods.
[0127] (2) The present invention generates simulated user comments by using the Zhipu Qingyan open-source large language model, which is a supplement to the information of those users who remain silent on conventional social platforms. Due to the powerful text generation ability and knowledge accumulation of the large language model, it can well understand the psychological state and news comment level of those users. Therefore, the large language model can well generate comments of silent users for fake news detection through artificial prompts. Since the comment information is more comprehensive, the model after extracting the comment features has increased generalization and robustness, and the detection efficiency has been greatly improved.
[0128] (3) Through an adaptive comment aggregator designed in the present invention, the important information contained in the generated comments is used to make up for the information missing from the real comments due to the existence of silent users. At the same time, the aggregator can eliminate single comment information and integrate different and diverse comments together, which not only improves the overall performance of the model but also better exerts the text generation ability of the large model.
[0129] (4) Through a comment sentiment analyzer designed in the present invention, the rich sentiment contained in the comments can be extracted and effectively applied to fake news detection. However, the generated comments are different from the real comments. Comparing the differences between the two at the sentiment level through the sentiment analyzer is a relatively reasonable perspective, which, together with other modalities such as images and texts, improves the ability to detect fake news from multiple perspectives.
[0130] (5) The present invention analyzes the text content of news through the open-source large language model LLaMa. Through a large amount of data training and adjustment of tens of billions of parameters, the large language model already has extremely strong knowledge reserves and preliminary abilities to distinguish true and false news. Analyzing the background, logic, common sense, and semantic context of the news through the large language model can obtain some very useful information that can contribute to subsequent detection.
[0131] (6) The present invention uses the powerful pre-trained RoBERTa model and pre-trained SwinT model to extract text encoding feature vectors and image encoding feature vectors, which helps to enhance the model's feature recognition and semantic analysis capabilities for the text content and image content of news, so as to better capture the semantic connections in news data samples. This includes using the CLIP pre-trained model to extract text features and image features in the text-image modality. When obtaining cross-modal features of text and image, contrast learning features of text and image are also obtained, significantly improving the feature comparison between the two main modalities.
[0132] (7) The present invention uses a mapping layer in the process of text-image modality fusion to extract useful information from the text features and image features of news, and at the same time eliminates some useless information, making the text features and image features in the text-image modality extracted by the CLIP pre-trained model have mutual integration, laying a foundation for subsequent feature complementarity and semantic alignment.
[0133] (8) In the initial design of the present invention, it fully considers that the large language model can be used in the method of multi-modal fake news detection. By combining the text generation ability and news analysis ability of the large language model with comprehensive detection from multiple perspectives, various modal information is effectively utilized, and a part of the hallucinations of the large language model are eliminated. This not only fills the gap in traditional multi-modal fake news detection but also gives new ideas for subsequent technological inventions. Brief Description of the Drawings
[0134] Figure 1 This is a schematic diagram of the overall framework of the present invention. Detailed implementation manners
[0135] As Figure 1 shown, the false news detection method based on generated comments and multi-perspective comprehensive analysis further includes the following steps:
[0136] Step S1: Construct news samples and perform data preprocessing on the news samples;
[0137] Step S2: After data preprocessing, obtain text feature data, image feature data, and real comment data and label them with sentiment tags;
[0138] Step S3: Based on the prompt template, use the first large language model to generate comments of several simulated users for the news samples through prompts and comprehensively analyze the text content of the news samples and the generated comments to obtain the text data of the news content analysis;
[0139] Step S4: Based on the prompt template, use the second large language model to comprehensively analyze the text data obtained after the content analysis of the first large language model, the text content of the news samples, and the comments generated by the large language model to obtain the news principle feature data;
[0140] Step S5: The multi-perspective feature extraction model extracts features from the news principle feature data, text feature data, real comment feature data, generated comments, and image feature data to obtain the corresponding encoded feature vectors;
[0141] Step S6: The adaptive comment aggregator fuses the generated comment encoded feature vectors and the real comment encoded feature vectors to obtain the comprehensive comment encoded feature vectors, and then fuses the comprehensive comment encoded feature vectors and the text encoded feature vectors to obtain the text and comment feature fusion vectors;
[0142] Step S7: Input the generated comment encoded feature vectors and the real comment encoded feature vectors into the comment sentiment analyzer to obtain the sentiment difference coefficient by calculating the discrete Gaussian distribution and the KL divergence;
[0143] Step S8: The multi-perspective feature fusion model fuses the text encoded feature vectors, news principle encoded feature vectors, comprehensive comment encoded feature vectors, image encoded feature vectors, and text-image modality encoded feature vectors to obtain the fused feature vectors;
[0144] Step S9: Aggregate the encoded feature vectors, fused feature vectors, text and comment feature fusion vectors, and sentiment difference coefficients to obtain the final aggregated encoded feature vectors, and the classifier module predicts the true or false prediction labels of the news samples corresponding to the final aggregated encoded feature vectors.
[0145] Furthermore, in step S1, the news sample is constructed and the process of data preprocessing on the news sample is specifically as follows:
[0146] Step S11: crawling and preprocessing the public social media data set through web crawler technology to obtain the data set required for model training, where a single piece of data in the news data set is a news sample;
[0147] Step S12: preprocess the data set, delete news samples that lack text content, lack image descriptions or have unclear text descriptions, and obtain a single news data sample including text feature data, image feature data and emotion labels; randomly divide the news data set into a training set and a test set according to preset rules.
[0148] Furthermore, in step S3, the process of using the No. 1 large language model to generate comments and analyze the content of the news text is specifically as follows:
[0149] Step S31, using the Zhipu Qingyan open source large language model as the No. 1 large language model to simulate social platform users and generate comments based on news samples in the data set; specifically:
[0150] First, a prompt template is given. The specific prompt template is:
[0151] This is [News text content]. If you are a user of a social platform, you can randomly select user attributes from [Gender: Male, Female], [Age: Under 18, 18 to 29, 30 to 49, 50 to 64, 64 and above], [Education: No higher education, Currently receiving higher education, Completed higher education], [Standpoint on news content: Support, Oppose, Neutral], and then generate 8 corresponding user comments for this news based on the user attributes you selected, and the word count of each comment is limited to 50 words;
[0152] Step S32: After obtaining the corresponding generated comments according to the prompt template of step S31, continue to use the Zhipu Qingyan open source large language model as the No. 1 large language model to analyze the news samples and generated comments.
[0153] The prompt template for the analysis is:
[0154] This is [news content] and the comments on the corresponding news [comment content]. Now, as an expert in true and false news analysis, please analyze this news based on the news content and the comment content. You can analyze it based on the source, time, background, logic or common sense of the news. You do not need to give a judgment result, just give your analysis. The word count is limited to 200 words.
[0155] Further, the process of using the second large language model to further analyze the analysis content, generated comments, and news content of the first large language model to eliminate some hallucinations in step S4 is specifically as follows:
[0156] Use the open-source LLaMa large language model as the second large language model to re-analyze the news text content, the comments generated by the Zhipu Qingyan open-source large language model, and the analysis of the news text content.
[0157] The prompt template for the re-analysis is:
[0158] This is the news content [news text content], the comment content [generated comment], and the analysis results of other large language models [news analysis content of the Zhipu Qingyan large language model]. Now, as a fake news analysis expert, please make full use of your knowledge to analyze this news. Your analysis should be logical, in line with professional standards, and maintain your independent thinking. You can reflect on or supplement the analysis results of other large models, especially the hallucination analysis. You don't need to judge the result, just give your analysis, with a word limit of within 300 words, to obtain the final news text content analysis of the large language model, that is, the news principle feature data.
[0159] Further, the multi-perspective feature extraction model in step S5 includes a parallel pre-trained RoBERTa model, a pre-trained CLIP model, and a pre-trained SwinT model;
[0160] The specific steps are as follows:
[0161] Step S51: Train the RoBERTa model to process the text feature data to obtain text encoding feature vectors;
[0162] The text feature data is expressed as ; where represents the text feature data, represents the first text feature data, represents the second text feature data, represents the nth text feature data;
[0163] The text encoding feature vector is expressed as ; where represents the text encoding feature vector, represents the first text encoding feature vector, represents the second text encoding feature vector, represents the nth text encoding feature vector;
[0164] Step S52: The pre-trained RoBERTa model processes the news principle feature data to obtain news principle encoding feature vectors;
[0165] The news principle feature data is represented as ; where represents the news principle feature data, represents the first news principle feature data, represents the second news principle feature data, represents the nth news principle feature data;
[0166] The news principle coding feature vector is represented as ; where represents the news principle coding feature vector, represents the first news principle coding feature vector, represents the second news principle coding feature vector, represents the nth news principle coding feature vector;
[0167] Step S53: The pre-trained SwinT model processes the image feature data to obtain the image coding feature vector;
[0168] where the image feature data is represented as Iϵ ; and are the height and width of the image feature data respectively; I = ; I represents the image feature data, represents the first image feature, represents the second image feature, represents the nth image feature;
[0169] The image coding feature vector is represented as ; represents the image coding feature vector, represents the first image coding feature vector, represents the second image coding feature vector, represents the nth image coding feature vector;
[0170] Step S54: The pre-trained RoBERTa model processes the generated comment feature data to obtain the generated comment coding feature vector;
[0171] The generated comment feature data is represented as ; represents the first comment feature data, represents the second comment feature data, represents the nth comment feature data;
[0172] The generated comment coding feature vector is represented as ; Represents the first comment encoding feature vector, Represents the second comment encoding feature vector, Represents the nth comment encoding feature vector;
[0173] Step S55: The pre-trained RoBERTa model processes the real comment feature data to obtain the real comment encoding feature vector;
[0174] The real comment feature data is represented as ; Represents the first comment feature data, Represents the second comment feature data, Represents the nth comment feature data;
[0175] The real comment encoding feature vector is represented as ; Represents the first comment encoding feature vector, Represents the second comment encoding feature vector, Represents the nth comment encoding feature vector;
[0176] Step S56: The pre-trained CLIP model processes the text feature data and the image feature data to obtain the text-image modality encoding feature vector , and the text-image modality encoding feature vector is divided into the text-image modality text encoding feature vector and the text-image modality image encoding feature vector .
[0177] Furthermore, the specific processing flow of the adaptive comment aggregator in step S6 is as follows:
[0178] Step S61: Fuse the generated comment encoding feature vector and the real comment encoding feature vector to obtain the comprehensive comment encoding feature vector;
[0179] Step S62: After fusing the text encoding feature vector and the comprehensive comment encoding feature vector, obtain the text-comment feature fusion vector: specifically:
[0180] Step S621: Calculate the cosine similarity score between the generated comment encoding feature vector and the real comment encoding feature vector , and then multiply it by the generated comment encoding feature vector to obtain the comprehensive comment feature vector , so as to evaluate the important information of the generated comment to make up for the missing information in the real comment;
[0181] Step S622: Calculate the text encoding feature vector and the comment encoding feature vector The association matrix between , so as to obtain important information in the text recognition comments and ensure the information transmission between the text and the comments;
[0182] Step S623: Take the association matrix as the key vector and the comment encoding feature vector as the value vector, and calculate the aggregated information of the key vector and the value vector; Use the non-linear activation function to normalize the aggregated information again to obtain the corresponding attention weight coefficients; Use the attention weight coefficients to perform weighted summation on the value vector V to obtain the comment aggregated information;
[0183] Step S624: Suppress the negative comments in the comment aggregated information through the fusion gating mechanism to obtain the fusion gating features, so as to ensure that the generated negative comments do not interfere with the model detection results;
[0184] Step S625: Process the fusion gating features through residual connection, splice them with the text encoding feature vector to retain the original features in the text; Obtain the text and comment feature fusion vector;
[0185] The processing flow of Step S62 is expressed as:
[0186] ;
[0187] ;
[0188] ;
[0189] ;
[0190] ;
[0191] ;
[0192] In the formula, is the cosine similarity score between the generated comment feature vector and the real comment feature vector, is the modulus length of the real comment encoding feature vector, is the modulus length of the generated comment encoding feature vector, is the comprehensive comment encoding feature vector, means changing the input part into a query vector, means changing the input part into a key vector, means the weight matrix, means the ReLU activation function, means the fusion gating feature, means the non-linear activation function, d represents the preset dimension, and T represents the transpose, represents the fusion vector of text and comment features represents element-wise multiplication represents the product operation represents the comment aggregation information
[0193] Further, step S7 is specifically as follows:
[0194] Step S71: Calculate the discrete Gaussian distribution by generating the comment sentiment label and the true comment sentiment label to obtain the sentiment distribution of the generated comment coding feature vector and the sentiment distribution of the true comment coding feature vector ;
[0195] Step S72: Calculate the divergence of the sentiment distributions of the generated comment coding feature vector and the true comment coding feature vector to compare the sentiment differences between the generated comment coding feature vector and the true comment coding feature vector, and obtain the sentiment difference score ; ;
[0196] Step S73: Superimpose the sentiment difference scores and use the mechanism and the non-linear activation function to obtain the final sentiment difference coefficient between the generated comment and the true comment ;
[0197] The processing flow of step S7 is expressed as:
[0198] ;
[0199] ;
[0200] ;
[0201] ;
[0202] ;
[0203] In the formula, is the generated comment sentiment label, is the true comment sentiment label, is the preset parameter, and are the sentiment distances of the generated comment coding feature vector and the true comment coding feature vector, is divergence, represents the natural constant to the exponential function, represents network represents a non-linear activation function, represents a splicing operation, i represents the index, and n represents the total number of generated comments.
[0204] Furthermore, the multi-view feature fusion model in step S8 includes a concatenated co-attention layer and a mapping layer; the co-attention layer consists of two connected attention blocks and a fully connected layer;
[0205] Step S8 is specifically as follows:
[0206] After the text encoding feature vector and the news principle encoding feature vector are fused, a text and principle feature fusion vector is obtained. The text encoding feature vector and the comprehensive comment encoding feature vector are used to obtain a text and comment feature fusion vector. After the text encoding feature vector and the image encoding feature vector are fused, a text and image feature fusion vector is obtained. After the text-image modality encoding feature vector extracted by the pre-trained CLIP model is fused, a text-image cross-modal feature fusion vector is obtained. After the text and principle feature fusion vector, the text and comment feature fusion vector, the text and image feature fusion vector, and the text-image cross-modal feature fusion vector are fused, a cross-modal feature alignment vector is obtained. Specifically:
[0207] Step S81: After the text encoding feature vector and the news principle encoding feature vector are fused, a text and principle feature fusion vector is obtained, specifically:
[0208] Step S811, input the text encoding feature vector and the news principle encoding feature vector together into the first two attention blocks in the co-attention layer;
[0209] Step S812, in the first attention block, the text encoding feature vector serves as the query vector Q, and the news principle encoding feature vector serves as the key vector K and the value vector V to calculate the semantic similarity score; use the non-linear activation function to normalize the attention score to obtain the corresponding attention weight coefficient; then use the attention weight coefficient to perform weighted summation on the value vector V to obtain the text principle weight vector ;
[0210] Step S813, in the second attention block, the news principle encoding feature vector serves as the query vector Q, and the text encoding feature vector serves as the key vector K and the value vector V. Use the query vector Q and the key vector K to calculate the semantic similarity score; again use the non-linear activation function to normalize the attention score to obtain the corresponding attention weight coefficient; use the attention weight coefficient to perform weighted summation on the value vector V to obtain the principle text weight vector ;
[0211] Step S814, text principle weight vector and the principle text weight vector are concatenated and input into the fully connected layer to obtain a text and principle feature fusion vector:
[0212] The processing flow of steps S811 to S814 is expressed as:
[0213] Q = × , K = × , V = × ;
[0214] ;
[0215] ;
[0216] ;
[0217] In the formula, means changing the input part into a query vector, means changing the input part into a key vector, means changing the input part into a value vector, means the execution process of an attention mechanism, means a non-linear activation function, d represents the preset dimension of the co-attention layer, T represents transpose, means the text and principle feature fusion vector, means the co-attention layer, means the concatenation operation;
[0218] Step S82: After the text encoding feature vector and the image encoding feature vector are fused, a text and image feature fusion vector is obtained: Specifically:
[0219] Step S821, the text encoding feature vector and the image encoding feature vector are input into the first two attention blocks in the co-attention layer together;
[0220] Step S822, in the first attention block, the text encoding feature vector is used as the query vector Q, and the image feature encoding feature vector is used as the key vector K and the value vector V to calculate the semantic similarity score; the attention score is normalized using the non-linear activation function to obtain the corresponding attention weight coefficient; then the attention weight coefficient is used to perform weighted summation on the value vector V to obtain the text-image weight vector ;
[0221] Step S823, in the second attention block, the image-encoded feature vector is used as the query vector Q, and the text-encoded feature vector is used as the key vector K and the value vector V. The semantic similarity score is calculated using the query vector Q and the key vector K; the attention score is normalized again using the non-linear activation function to obtain the corresponding attention weight coefficient; the value vector V is weighted and summed using the attention weight coefficient to obtain the image-text weight vector ;
[0222] Step S824, the text-image weight vector and the image-text weight vector are concatenated and input into the fully connected layer to obtain the text and image feature fusion vector,
[0223] The processing flow of Steps S821 to S824 is expressed as:
[0224] Q = × K = × V = × ;
[0225] ;
[0226] ;
[0227] ;
[0228] In the formula, represents changing the input part into the query vector, represents changing the input part into the key vector, represents changing the input part into the value vector, represents the execution process of an attention mechanism, represents the non-linear activation function, d represents the preset dimension of the co-attention layer, T represents the transpose, represents the text and image feature fusion vector, represents the co-attention layer, represents the concatenation operation;
[0229] Furthermore, Step S8 further includes:
[0230] Step S83: The text and image modality encoded feature vectors extracted by the pre-trained CLIP model are fused to obtain the text and image cross-modal feature fusion vector; specifically:
[0231] Step S831: At the first attention operation, take the text encoding feature vector of the text-image modality as the query vector Q, and the image encoding feature vector of the text-image modality as the key vector K and the value vector V, calculate the semantic connection between the text encoding feature vector and the image encoding feature vector of the text-image modality, and obtain the output representation of the first attention operation ; At the second attention operation, take the image encoding feature vector of the text-image modality as the query vector Q, and the text encoding feature vector of the text-image modality as the key vector K and the value vector V, and obtain the output representation of the second attention operation ;
[0232] Step S832: Perform average pooling operations on the outputs of the previous and subsequent attention operations respectively to obtain the feature representations after average pooling, and then concatenate them to obtain the CLIP feature fusion representation : The processing flow of Steps S831 to S832 is expressed as:
[0233] ;
[0234] ;
[0235] In the formula, represents the cross-modal feature fusion vector of text and image;
[0236] Step S84: After fusing the text and principle feature fusion vector, the text and comment feature fusion vector, the text and image feature fusion vector, and the cross-modal feature fusion vector of text and image, obtain the cross-modal feature alignment vector; specifically:
[0237] Step S841, input the text encoding feature vector and the image encoding feature vector into the fully connected layer and then into the co-attention layer;
[0238] Step S842, when performing the first attention operation, take the text encoding feature vector as the query vector Q, and the image encoding feature vector as the key vector K and the value vector V, calculate the semantic connection between the text encoding feature vector and the image encoding feature vector, and obtain the output representation of the first attention operation ; At the second attention operation, take the image encoding feature vector as the query vector, and the text encoding feature vector as the key vector and the value vector, and obtain the output representation of the second attention block ;
[0239] Step S843: The outputs of the previous and subsequent attention operations are respectively input into the average pooling layer for average pooling operations, and then concatenated to obtain a cross-modal feature fusion representation. ; The processing flow of Steps S841 to S843 is expressed as:
[0240] ;
[0241] ;
[0242] In the formula, represents the process of average pooling, represents the operation process of the co-attention layer;
[0243] Step S844: The cross-modal feature fusion representation and the text-image cross-modal feature fusion vector are concatenated and then input into the mapping layer for fusion to obtain a cross-modal mapping representation ;
[0244] Step S845: Use the cross-modal semantic similarity score to adjust the key weights between multiple modalities, and at the same time calculate the semantic similarity relationship between the text encoding feature vector and the image encoding feature vector of the text-image modality to obtain the cross-modal semantic similarity score;
[0245] Step S846: Multiply the cross-modal mapping representation by the cross-modal semantic similarity score to obtain a cross-modal feature alignment vector ; The processing flow of Steps S844 to S846 is expressed as:
[0246] ;
[0247] ;
[0248] ;
[0249] In the formula, is the cross-modal mapping representation, MPP is the mapping layer, is the cross-modal semantic similarity score, represents the norm of the text encoding feature vector of the text-image modality, represents the norm of the image encoding feature vector of the text-image modality, represents the transpose matrix of the image encoding feature vector of the text-image modality.
[0250] Furthermore, Step S9 is specifically:
[0251] Aggregate the text encoding feature vector, image encoding feature vector, comprehensive comment encoding feature vector, comment sentiment difference coefficient, news principle encoding feature vector, text and principle feature fusion vector, text and image feature fusion vector, text and comment feature fusion vector, text-image cross-modal feature fusion vector, and cross-modal feature alignment vector to obtain the final aggregated encoding feature vector, and input it into the classifier module to output the true / false prediction label of the news sample; specifically;
[0252] ;
[0253] ;
[0254] wherein, represents the prediction label of the classifier module, represents the fully connected layer, represents the cross-entropy loss function, represents the true label of the sample.
[0255] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A fake news detection method based on generated comments and multi-perspective comprehensive analysis, characterized in that: The steps include: Step S1: construct news samples and perform data preprocessing on the news samples; Step S2: After data preprocessing, text feature data, image feature data, and real comment data are obtained and annotated with emotion tags; Step S3: Based on the prompt template, a large language model is used to generate several simulated user comments on the news sample through prompts, and the text content of the news sample and the generated comments are comprehensively analyzed to obtain text data for news content analysis; Step S4: Based on the prompt template, the second large language model is used to perform a comprehensive analysis on the text data obtained after the content analysis of the first large language model, the text content of the news sample and the comments generated by the large language model through prompts to obtain news principle feature data; Step S5: The multi-view feature extraction model extracts features from news principle feature data, text feature data, real comment feature data, generated comment and image feature data to obtain corresponding encoding feature vectors; Step S6: The adaptive comment aggregator fuses the generated comment encoding feature vector and the real comment encoding feature vector to obtain a comprehensive comment encoding feature vector, and then fuses the comprehensive comment encoding feature vector and the text encoding feature vector to obtain a text and comment feature fusion vector; Step S7: input the generated comment encoding feature vector and the real comment encoding feature vector into the comment sentiment analyzer to obtain the sentiment difference coefficient by calculating the discrete Gaussian distribution and KL divergence; Step S8: The multi-view feature fusion model fuses the text encoding feature vector, the news principle encoding feature vector, the comprehensive comment encoding feature vector, the image encoding feature vector, and the graphic modality encoding feature vector to obtain a fused feature vector; Step S9: Aggregate the coded feature vector, fused feature vector, text and comment feature fusion vector and sentiment difference coefficient to obtain the final aggregated coded feature vector. The classifier module predicts the true or false prediction label of the news sample corresponding to the final aggregated coded feature vector.
2. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 1, characterized in that: In step S1, a news sample is constructed and the process of data preprocessing of the news sample is specifically as follows: Step S11: crawling and preprocessing the public social media data set through web crawler technology to obtain the data set required for model training, where a single piece of data in the news data set is a news sample; Step S12: preprocess the data set, delete news samples that lack text content, lack image descriptions or have unclear text descriptions, and obtain a single news data sample including text feature data, image feature data and emotion labels; randomly divide the news data set into a training set and a test set according to preset rules.
3. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 2, characterized in that: The specific process of using the No. 1 large language model to generate comments and analyze the content of the news text in step S3 is as follows: Step S31, using the Zhipu Qingyan open source large language model as the No. 1 large language model to simulate social platform users and generate comments based on news samples in the data set; specifically: First, a prompt template is given. The specific prompt template is: This is [News text content]. If you are a user of a social platform, you can randomly select user attributes from [Gender: Male, Female], [Age: Under 18, 18 to 29, 30 to 49, 50 to 64, 64 and above], [Education: No higher education, Currently receiving higher education, Completed higher education], [Standpoint on news content: Support, Oppose, Neutral], and then generate 8 corresponding user comments for this news based on the user attributes you selected, and the word count of each comment is limited to 50 words; Step S32: After obtaining the corresponding generated comments according to the prompt template of step S31, continue to use the Zhipu Qingyan open source large language model as the No. 1 large language model to analyze the news samples and generated comments. The prompt template for the analysis is: This is [news content] and the comments on the corresponding news [comment content]. Now, as an expert in true and false news analysis, please analyze this news based on the news content and the comment content. You can analyze it based on the source, time, background, logic or common sense of the news. You do not need to give a judgment result, just give your analysis. The word count is limited to 200 words.
4. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 3, characterized in that: In step S4, the second large language model is used to further analyze the analysis content, generated comments and news content of the first large language model to eliminate part of the hallucination process. Specifically, the process is as follows: Use LLaMa open source big language model as the second big language model to re-analyze the news text content, the comments generated by Zhipu Qingyan open source big language model, and the analysis of the news text content. The prompt template for reanalysis is: This is the news content [news text content], comment content [generated comments], and the analysis results of other large language models [Zhipu Qingyan large language model news analysis content]. Now, as an expert in true and false news analysis, please make full use of your knowledge to analyze this news. The logic should be clear and in line with professional standards, and you should maintain your independent thinking. You can reflect on or supplement the analysis results of other large models, especially some illusion analysis. You do not need to judge the results, but only give your analysis. The word count is limited to 300 words to obtain the final large language model news text content analysis, that is, the news principle feature data.
5. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 4, characterized in that: The multi-view feature extraction model in step S5 includes a parallel pre-trained RoBERTa model, a pre-trained CLIP model, and a pre-trained SwinT model; The specific steps are: Step S51: training the RoBERTa model to process the text feature data to obtain a text encoding feature vector; Text feature data is represented as ;in, Represents text feature data, Represents the first text feature data, Represents the second text feature data, Represents the nth text feature data; The text encoding feature vector is represented as ;in, represents the text encoding feature vector, represents the first text encoding feature vector, represents the second text encoding feature vector, Represents the nth text encoding feature vector; Step S52, the pre-trained RoBERTa model processes the news principle feature data to obtain a news principle encoding feature vector; The news principle feature data is expressed as ;in, Represents news principle feature data, Indicates the first news principle feature data, Indicates the second news principle feature data, Represents the nth news principle feature data; The news principle encoding feature vector is expressed as ;in, represents the news principle encoding feature vector, represents the first news principle encoding feature vector, represents the second news principle encoding feature vector, represents the nth news principle encoding feature vector; Step S53: The pre-trained SwinT model processes the image feature data to obtain an image coding feature vector; Among them, the image feature data is represented by Iϵ ; and are the height and width of the image feature data respectively; I= ]; I represents image feature data, Represents the first image feature, Represents the features of the second image, Represents the features of the nth image; The image encoding feature vector is expressed as ; represents the image encoding feature vector, represents the first image encoding feature vector, represents the second image encoding feature vector, represents the nth image encoding feature vector; Step S54: The pre-trained RoBERTa model is used to process the generated comment feature data to obtain a generated comment encoding feature vector; Generate comments Comment feature data is represented as ; Represents the first comment feature data, Represents the second comment feature data, Represents the feature data of the nth comment; Generate comments Comment encoding feature vector is represented as ; represents the first comment encoding feature vector, represents the second comment encoding feature vector, represents the nth comment encoding feature vector; Step S55: The pre-trained RoBERTa model processes the real comment feature data to obtain the real comment encoding feature vector; The real review feature data is represented as ; Represents the first comment feature data, Represents the second comment feature data, Represents the feature data of the nth comment; The real review encoding feature vector is represented as ; represents the first comment encoding feature vector, represents the second comment encoding feature vector, represents the nth comment encoding feature vector; Step S56: The pre-trained CLIP model processes the text feature data and the image feature data to obtain a text-image modality encoding feature vector , image-text modality encoding feature vector Divided into image-text modal text encoding feature vector , Image encoding feature vector of text-image modality .
6. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 5, characterized in that: The specific processing flow of the adaptive comment aggregator in step S6 is as follows: Step S61: fusing the generated comment coding feature vector and the real comment coding feature vector to obtain a comprehensive comment coding feature vector; Step S62: The text encoding feature vector and the comprehensive comment encoding feature vector are fused to obtain a text and comment feature fusion vector: specifically: Step S621: Calculate and generate comment encoding feature vector and the true review encoding feature vector The cosine similarity score between them is then multiplied to generate the comment encoding feature vector Get the comprehensive review feature vector , thereby evaluating the important information of generated reviews to make up for the missing information in real reviews; Step S622: Calculate text encoding feature vector and the comment encoding feature vector The correlation matrix between , thereby obtaining important information in text recognition comments and ensuring information transmission between text and comments; Step S623: Convert the correlation matrix as the key vector and the comment encoding feature vector As the value vector, calculate the aggregate information of the key vector and the value vector; use the nonlinear activation function again to normalize the aggregate information to obtain the corresponding attention weight coefficient; use the attention weight coefficient to perform weighted summation on the value vector V to obtain the comment aggregation information; Step S624: Suppressing comment aggregation information through fusion gating mechanism Obtain fusion gating features from negative comments in order to ensure that the generated negative comments do not interfere with the model detection results; Step S625: The gated features are fused through residual connection processing and concatenated with the text encoding feature vector to retain the original features in the text; and a text and comment feature fusion vector is obtained; The processing flow of step S62 is expressed as follows: ; ; ; ; ; ; In the formula, To generate the cosine similarity score between the comment feature vector and the real comment feature vector, is the modulus of the feature vector encoding the real review, To generate the modulus length of the comment encoding feature vector, Encode the feature vector for the comprehensive review, Indicates that the input part is transformed into a query vector, Indicates that the input part is changed into a key vector. Represents the weight matrix, represents the ReLU activation function, represents the fusion gated feature, represents a nonlinear activation function, d represents the preset dimension, T represents transposition, represents the fusion vector of text and comment features, represents element-wise multiplication, represents the product operation, Represents comment aggregation information.
7. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 6, characterized in that: Step S7 is specifically as follows: Step S71: Generate comment sentiment labels and real review sentiment labels Calculate the discrete Gaussian distribution to get the sentiment distribution of the generated comment encoding feature vector and sentiment distribution of the encoded feature vector of the real review ; Step S72: Calculate the sentiment distribution of the generated comment encoding feature vector and the real comment encoding feature vector Divergence, to compare the sentiment difference between the generated comment encoding feature vector and the real comment encoding feature vector, and get the sentiment difference score ; Step S73: Overlay the sentiment difference scores, using Mechanism and nonlinear activation function to obtain the final sentiment difference coefficient between the generated comments and the real comments ; The processing flow of step S7 is expressed as follows: ; ; ; ; ; In the formula, To generate comment sentiment labels, is the real comment sentiment label, is the preset parameter, and To generate the sentiment distance between the comment encoding feature vector and the real comment encoding feature vector, for Divergence, The natural constants are expressed as The exponential function of express network, represents a nonlinear activation function, represents the concatenation operation, i represents the index, and n represents the total number of generated comments.
8. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 7, characterized in that: The multi-view feature fusion model in step S8 includes a co-attention layer and a mapping layer connected in series; the co-attention layer consists of two connected attention blocks and a fully connected layer; Step S8 is specifically as follows: The text encoding feature vector and the news principle encoding feature vector are fused to obtain the text and principle feature fusion vector, the text encoding feature vector and the comprehensive comment encoding feature vector are fused to obtain the text and comment feature fusion vector, the text encoding feature vector and the image encoding feature vector are fused to obtain the text and image feature fusion vector, and the image and text modality encoding feature vector extracted by the pre-trained CLIP model is fused to obtain the image and text cross-modal feature fusion vector; the text and principle feature fusion vector, the text and comment feature fusion vector, the text and image feature fusion vector and the image and text cross-modal feature fusion vector are fused to obtain the cross-modal feature alignment vector; specifically: Step S81: The text coding feature vector and the news principle coding feature vector are fused to obtain a text and principle feature fusion vector, specifically: Step S811: Encode the text feature vector and news principle encoding feature vector The first two attention blocks are fed together into the co-attention layer; Step S812, in the first attention block, the text encoding feature vector As the query vector Q, the news principle encoding feature vector As the key vector K and the value vector V, the semantic similarity score is calculated; the attention score is normalized using a nonlinear activation function to obtain the corresponding attention weight coefficient; the value vector V is then weighted and summed using the attention weight coefficient to obtain the text principle weight vector ; Step S813, in the second attention block, the news principle encoding feature vector As the query vector Q, the text encoding feature vector As the key vector K and the value vector V, the query vector Q and the key vector K are used to calculate the semantic similarity score; the attention score is normalized again using the nonlinear activation function to obtain the corresponding attention weight coefficient; the value vector V is weighted and summed using the attention weight coefficient to obtain the principle text weight vector ; Step S814: text principle weight vector With the principle text weight vector After concatenation, the vector is input into the fully connected layer to obtain the fusion vector of text and principle features: The processing flow of step S811 to step S814 is expressed as follows: Q= × ,K= × ,V= × ; ; ; ; In the formula, Indicates that the input part is turned into a query vector, Indicates that the input part is turned into a key vector, It means to change the input part into a value vector. Represents the execution process of an attention mechanism, represents a nonlinear activation function, d represents the preset dimension of the common attention layer, T represents transposition, Represents the fusion vector of text and principle features, represents the shared attention layer, Represents a splicing operation; Step S82: The text coding feature vector and the image coding feature vector are fused to obtain a text and image feature fusion vector: specifically: Step S821: Encode the text feature vector and the image encoding feature vector The first two attention blocks are fed together into the co-attention layer; Step S822, in the first attention block, the text encoding feature vector As the query vector Q, the image feature encoding feature vector As the key vector K and the value vector V, the semantic similarity score is calculated; the attention score is normalized using a nonlinear activation function to obtain the corresponding attention weight coefficient; the value vector V is then weighted summed using the attention weight coefficient to obtain the text image weight vector ; Step S823, in the second attention block, the image encoding feature vector As the query vector Q, the text encoding feature vector As the key vector K and the value vector V, the query vector Q and the key vector K are used to calculate the semantic similarity score; the attention score is normalized again using the nonlinear activation function to obtain the corresponding attention weight coefficient; the value vector V is weighted and summed using the attention weight coefficient to obtain the image text weight vector ; Step S824: text image weight vector and image text weight vector After splicing, it is input into the fully connected layer to obtain the text and image feature fusion vector. The processing flow of step S821 to step S824 is expressed as follows: Q= × ,K= × ,V= × ; ; ; ; In the formula, Indicates that the input part is transformed into a query vector, Indicates that the input part is changed into a key vector. Indicates that the input part is changed into a value vector, Represents the execution process of an attention mechanism, represents a nonlinear activation function, d represents the preset dimension of the common attention layer, T represents transposition, represents the fusion vector of text and image features, represents the shared attention layer, Represents a concatenation operation.
9. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 8, characterized in that: Step S8 also includes: Step S83: fusing the image-text modality encoding feature vectors extracted by the pre-trained CLIP model to obtain an image-text cross-modality feature fusion vector; specifically: Step S831: In the first attention operation, the text encoding feature vector of the image-text modality is As the query vector Q, the image encoding feature vector of the text-image modality As the key vector K and value vector V, the semantic connection between the text encoding feature vector of the text-to-image modality and the image encoding feature vector of the text-to-image modality is calculated to obtain the output representation of the first attention operation ; In the second attention operation, the image encoding feature vector of the text-image modality As the query vector Q, the text encoding feature vector of the image-text modality As the key vector K and value vector V, we get the output representation of the second attention operation ; Step S832: Perform average pooling operations on the outputs of the previous and next attention operations to obtain the feature representation after average pooling, and then cascade to obtain the CLIP feature fusion representation The processing flow of step S831 to step S832 is expressed as follows: ; ; In the formula, Represents the cross-modal feature fusion vector of images and texts; Step S84: The text and principle feature fusion vector, the text and comment feature fusion vector, the text and image feature fusion vector and the image-text cross-modal feature fusion vector are fused to obtain a cross-modal feature alignment vector; specifically: Step S841: Encode the text feature vector and the image encoding feature vector After being input into the fully connected layer, it is then input into the co-attention layer; Step S842, when performing the first attention operation, the text encoding feature vector As the query vector Q, the image encoding feature vector As the key vector K and value vector V, the semantic connection between the text encoding feature vector and the image encoding feature vector is calculated to obtain the output representation of the first attention operation ; In the second attention operation, the image encoding feature vector As the query vector, the text encoding feature vector is used as the key vector and the value vector to get the output representation of the second attention block. ; Step S843: Input the outputs of the two attention operations into the average pooling layer for average pooling operation, and then cascade to obtain the cross-modal feature fusion representation. ; The processing flow of step S841 to step S843 is expressed as: ; ; In the formula, represents the average pooling process, Represents the operation process of the shared attention layer; Step S844: Fusing cross-modal features and image-text cross-modal feature fusion vector After cascading, the input mapping layer is fused to obtain a cross-modal mapping representation ; Step S845: using the cross-modal semantic similarity score to adjust the key weights between multiple modalities, and calculating the semantic similarity relationship between the text encoding feature vector of the graphic modality and the image encoding feature vector of the graphic modality to obtain the cross-modal semantic similarity score; Step S846: Represent the cross-modal mapping Multiply it with the cross-modal semantic similarity score to obtain the cross-modal feature alignment vector ; The processing flow of step S844 to step S846 is expressed as: ; ; ; In the formula, is the cross-modal mapping representation, MPP is the mapping layer, is the cross-modal semantic similarity score, represents the modulus length of the text encoding feature vector of the image-text modality, represents the modulus length of the image encoding feature vector of the text-to-image modality, Represents the transposed matrix of the encoded feature vector of the text-to-image modality image.
10. The method for detecting fake news based on generating comments and comprehensive analysis from multiple perspectives according to claim 9, characterized in that: Step S9 is specifically as follows: Aggregate text encoding feature vector, image encoding feature vector, comprehensive comment encoding feature vector, comment sentiment difference coefficient, news principle encoding feature vector, text and principle feature fusion vector, text and image feature fusion vector, text and comment feature fusion vector, image and text cross-modal feature fusion vector, cross-modal feature alignment vector to obtain the final aggregated encoding feature vector, input it into the classifier module, and output the true or false prediction label of the news sample; specifically; ; ; In the formula, represents the predicted label of the classifier module, represents the fully connected layer, represents the cross entropy loss function, Represents the true label of the sample.
Citation Information
Patent Citations
Multi-modal false information detection system based on co-situation theory guidance
CN117851894A
False news detection method based on multi-view and hierarchical fusion
CN118114188A