False news detection method based on multi-modal information fusion and medium

Through semantic enhanced cross-modal common attention networks (SCCNs), combined with improved graph attention networks (GAT) and self-supervised learning, the problems of insufficient fusion of cross-modal information in fake news detection, noise interference between social relationship networks and semantic gaps between modals are solved, and high-accurate fake news detection is achieved.

CN120045977APending Publication Date: 2025-05-27NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510106005.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing fake news detection methods are inadequate detection accuracy due to insufficient cross-modal information fusion, noise interference in social relationship networks and semantic gaps between modals.

Method used

A semantically enhanced cross-modal common attention network (SCCNs) is proposed to process social relationship graphs through improved graph attention networks (GATs), complement the connections of social relationship graphs and remove noise. Combining self-supervised learning and common attention mechanisms to achieve cross-modal alignment and fusion.

Benefits of technology

It significantly improves the accuracy and robustness of fake news detection, effectively solving key problems in multimodal fake news detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045977A_ABST
    Figure CN120045977A_ABST
Patent Text Reader

Abstract

The invention provides a false news detection method based on multi-modal information fusion and a medium. The method comprises the following steps: firstly, acquiring news texts, images, users and comment information, and constructing a social relation graph containing news nodes, user nodes and comment nodes; low-similarity connection is removed through node similarity calculation, potential connection is deduced, and an improved social relation graph is generated. Performing feature coding on the text and the image by using the pre-training model to generate a text feature vector, an image feature vector and a social relation feature vector; and text and image entity embedding enhancement feature vectors are extracted. And mapping the enhanced text, image and social relation feature vectors to a common semantic space, and realizing modal alignment by optimizing mean square error loss to obtain cross-modal alignment feature vectors. And finally, fusing the multi-modal information through a common attention mechanism, generating a multi-modal fusion feature vector, inputting the multi-modal fusion feature vector into a classifier, and outputting true and false news labels. According to the method, the precision and robustness of false news detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and specifically relates to a fake news detection method and medium based on multimodal information fusion. Background Art

[0002] With the rapid development of the Internet and social media, the way and form of information dissemination have undergone profound changes. Multimodal news (i.e., news that simultaneously contains text, images, and other information forms) has been widely used on social media platforms. This form of news can not only attract users' attention but also convey information in a more vivid and intuitive way. However, the popularity of multimodal news has also provided conditions for the generation and spread of fake news. Fake news usually misleads the public by distorting facts, forging pictures, or manipulating social emotions, thus having an adverse impact on the normal order and public opinion environment of society. Therefore, how to effectively identify and detect fake news has become an urgent problem to be solved.

[0003] The essence of fake news detection is a binary classification problem, aiming to judge the authenticity of news by analyzing the multimodal features of news content. Traditional methods mainly focus on unimodal content analysis, such as text-based semantic analysis or image-based feature extraction. However, with the progress of fake news generation technology, unimodal detection methods are difficult to cope with fake news with complex content. For example, some fake news has no obvious errors in text, while its pictures may be tampered with; vice versa. Therefore, multimodal fusion detection methods have gradually become a research hotspot. Multimodal-based fake news detection methods usually try to combine text and image features to capture the correlation between different modalities, thereby improving the detection accuracy.

[0004] In recent years, deep learning-based models have demonstrated powerful performance in fake news detection tasks. For example, multimodal learning methods based on variational autoencoders (VAEs), generative adversarial networks (GANs), or attention mechanisms can perform deeper expression and fusion of text and image features. However, these methods still have the following deficiencies:

[0005] 1. Ignoring the deep fusion of cross-modal information: Most existing methods focus on a single modality or simply fuse different modality information, without fully utilizing the complementarity between modalities, resulting in limited performance of the detection model.

[0006] 2. Lack of utilization of social relationship information: In the social media environment, the dissemination of news is often accompanied by social relationships such as comments and forwards. These relationships contain a lot of important structured information, which has potential auxiliary effects on fake news detection. However, existing research has rarely deeply explored these social relationship features.

[0007] 3. Incompleteness and noise interference of social relationship graphs: Connections in social relationship networks may be incomplete, and there may even be a large number of irrelevant noise connections, which may lead to a decline in model performance.

[0008] 4. Insufficient modal alignment and noise processing: During the multi-modal fusion process, the semantic gap between different modalities and inconsistent information introduced by noise will reduce the robustness and detection performance of the model.

[0009] Based on the above problems, this paper proposes a novel Semantic-enhanced Cross-modal Co-attention Network (SCCNs) aiming to achieve deep fusion of multi-modal information and accurate detection of fake news. This method processes the social relationship graph by introducing an improved Graph Attention Network (GAT) to complete the connections of the social relationship graph and remove noise. In addition, entity features are used to enhance the semantic expressions of text and images, and cross-modal alignment and fusion are achieved by combining self-supervised learning and co-attention mechanisms, thus effectively solving the key problems in multi-modal fake news detection. Experiments show that the performance of this method on real datasets is significantly better than existing methods, providing an innovative solution for fake news detection technology. Summary of the Invention

[0010] The present invention provides a method and medium for detecting fake news based on multi-modal information fusion, aiming to solve the problem of insufficient cross-modal information fusion, noise interference in social relationship networks, and insufficient detection accuracy caused by the semantic gap between modalities in existing fake news detection.

[0011] To achieve the above object, the first aspect of the present invention provides a method and medium for detecting fake news based on multi-modal information fusion, including the following steps:

[0012] Obtain the text, image, user, and comment information of the news;

[0013] Construct a social relationship graph including news nodes, user nodes, and comment nodes according to the interaction relationship between users and comments in the news content;

[0014] Based on the calculation of node similarity in the social relationship graph, remove low-similarity connections, infer potential connections, and generate an improved social relationship graph;

[0015] Use pre-trained models to perform feature encoding on the text and image respectively to generate text feature vectors and image feature vectors, and obtain social relationship feature vectors from the improved social relationship graph;

[0016] Extract text entities and image entities from the text and image;

[0017] The extracted text entities and image entities are respectively embedded to enhance the information of the text feature vector and the image feature vector, generating an enhanced text feature vector and an enhanced image feature vector;

[0018] The enhanced text feature vector, the enhanced image feature vector and the social relationship feature vector are mapped to a common semantic space, and inter-modal alignment is performed by optimizing the mean square error loss to obtain the feature vectors of the image, text and social relationship graph after cross-modal alignment;

[0019] The feature vectors of the aligned image, text and social relationship graph are fused using a common attention mechanism to generate a multi-modal fusion feature vector containing all modal information;

[0020] The multi-modal fusion feature vector is input into a classifier, and a true / false news label is output through the classifier.

[0021] Further, the method for generating an improved social relationship graph includes:

[0022] Calculate the initial embedding features of the nodes. Among them, for news nodes, extract the embedding features according to their text content; for comment nodes, extract the embedding features according to their content; for user nodes, generate the embedding features of the user according to the mean value of the embedding features of the comments they publish.

[0023] Based on the initial embedding features of the nodes, calculate the similarity between any two nodes, and use the cosine similarity to quantify the degree of association between the nodes; determine the range of the similarity threshold between nodes, and set the upper threshold and the lower threshold for judging the node connection relationship;

[0024] According to the similarity calculation result, judge the connections in the social relationship graph: if there is a connection between nodes and the similarity is lower than the lower threshold, it is marked as a noise connection; if there is no connection between nodes but the similarity is higher than the upper threshold, it is marked as a potential connection; the rest of the connections remain unchanged;

[0025] According to the identified noise connections and potential connections, adjust the social relationship graph: delete the noise connections, add potential connections to complete the hidden associations in the social relationship graph, and update the adjacency matrix to reflect the improved connection relationship;

[0026] Based on the optimized adjacency matrix and node information, generate an improved social relationship graph.

[0027] Further, the method for generating a text feature vector includes:

[0028] Extract the text of each news from the social media news dataset;

[0029] Preprocess the text, including word segmentation, removing stop words, and normalizing the text length to make the text length of each news the same;

[0030] Use the pre-trained BERT model to encode the processed text to generate a word embedding sequence;

[0031] Input the word embedding sequence into a bidirectional long short-term memory network to model the context information of the text;

[0032] Aggregate the output of the bidirectional long short-term memory network by weighted summation, and calculate the final text feature vector through a fully connected layer.

[0033] Furthermore, the method for generating an image feature vector includes:

[0034] Extract the image corresponding to each news from the social media news dataset;

[0035] Preprocess the image, including normalizing the image size, denoising, and normalization;

[0036] Use the pre-trained ResNet-50 model to extract features from the image, and obtain the output of its penultimate layer as the initial image feature;

[0037] Input the initial image feature into a fully connected layer, and perform a non-linear transformation on the initial image feature through an activation function to obtain the final image feature vector.

[0038] Furthermore, the method for obtaining a social relationship feature vector from the improved social relationship graph includes:

[0039] Input the improved social relationship graph into a graph attention network;

[0040] Based on the graph attention network, model the relationship between each node and its neighboring nodes, calculate the attention weights using the node embedding vectors, and obtain the initial attention weights between nodes through a combination of linear transformation and dot product operations of node embeddings, which are used to represent the relationship strength of neighboring nodes to the target node;

[0041] Expand the initial attention weights between nodes into positive and negative weights, where the positive weights are used to reflect the positive impact of neighboring nodes on the target node, and the negative weights are used to capture the negative impact of neighboring nodes on the target node; perform normalization processing on the positive and negative weights respectively to enable them to correctly represent the positive and negative contributions of neighboring nodes;

[0042] Based on the attention weights after positive and negative normalization, the neighborhood node features of the target node are weighted and summed, and the weighted results of the positive and negative weights are respectively aggregated to generate the feature representation of the target node, and the positive and negative aggregation results are concatenated into the initial feature vector of the target node;

[0043] Apply the multi-head attention mechanism to the feature representations of all nodes for further fusion, extract the feature relationships of the target node in the complex graph structure from different attention heads, and generate the final feature representation of each node;

[0044] Summarize and fuse the final feature representations of all nodes in the graph to generate the social relationship feature vector.

[0045] Further, the method for generating the enhanced text feature vector and the enhanced image feature vector includes:

[0046] Extract the text and image of each news from the social media news data;

[0047] Perform entity recognition on the text of the news to extract the key entities in the text, and generate the text entity embedding vector through the entity embedding model. At the same time, perform object detection on the image to extract the entity information in the image, and generate the image entity embedding vector;

[0048] Concatenate the text feature vector and the text entity embedding vector, and generate the enhanced text feature vector through the multi-layer perceptron; concatenate the image feature vector and the image entity embedding vector, and generate the enhanced image feature vector through the multi-layer perceptron.

[0049] Further, the method for obtaining the feature vectors of the cross-modal aligned image, text, and social relationship graph includes:

[0050] Map the enhanced text feature vector, the enhanced image feature vector, and the social relationship feature vector to a common semantic space to obtain the aligned feature vectors:

[0051]

[0052] Among them, represents the feature vector after the enhanced image feature vector Z I is mapped to the unified semantic space, represents the feature vector after the enhanced text feature vector Z T is mapped to the unified semantic space, represents the social relationship feature vector Z R is mapped to the unified semantic space, and are learnable matrices, Z I is the enhanced image feature vector, ZT is the enhanced text feature vector, Z R is the social relationship feature vector;

[0053] Use the mean squared error loss function to calculate the loss of alignment between the enhanced text feature vector and the enhanced image feature vector;

[0054] Use the mean squared error loss function to calculate the loss of alignment between the enhanced text feature vector and the social relationship feature vector;

[0055] Use the mean squared error loss function to calculate the loss of alignment between the enhanced image feature vector and the social relationship feature vector;

[0056] Sum up the above mean squared error losses to obtain the total alignment loss.

[0057] Furthermore, the method for generating a multi-modal fusion feature vector containing all modal information includes:

[0058] Fuse the aligned text feature vector, image feature vector, and feature vector of the social relationship graph, and use the co-attention mechanism to capture the interaction information between modalities:

[0059] The cross-modal feature of text and image fusion is calculated by the following formula:

[0060]

[0061] where, represents the query vector of the text feature, and is calculated by the linear transformation matrix ; represents the key vector of the image feature, and is calculated by the linear transformation matrix ; represents the value vector of the image feature, and is calculated by the linear transformation matrix ; represents the query vector of the image feature, and is calculated by the linear transformation matrix ; represents the key vector of the text feature, and is calculated by the linear transformation matrix ; represents the value vector of the text feature, and is calculated by the linear transformation matrix ; W TI 、W IT are linear transformation matrices for obtaining the final cross-modal feature representation; is the normalization coefficient for scaling, where d is the dimension of the feature vector; K is the number of attention heads, representing the number of heads in the multi-head attention mechanism; f TI represents the text feature enhanced by the image feature; f ITImage features representing text feature enhancement;

[0062] The cross-modal feature calculation formula for text and social relationship graph is as follows:

[0063]

[0064]

[0065] Among them, The query vector representing social relationship features is calculated through the linear transformation matrix ; The key vector representing social relationship features is calculated through the linear transformation matrix ; The value vector representing social relationship features is calculated through the linear transformation matrix ; W TR 、W RT are linear transformation matrices for obtaining the final cross-modal feature representation; f TR represents the text features enhanced by social relationship features; f RT represents the social relationship features enhanced by text features;

[0066] The cross-modal feature calculation formula for image and social relationship graph is as follows:

[0067]

[0068]

[0069] Among them, The query vector representing image features is calculated through the linear transformation matrix ; The key vector representing social relationship features is calculated through the linear transformation matrix ;

[0070] The value vector representing social relationship features is calculated through the linear transformation matrix ; W IR 、W RI are linear transformation matrices for obtaining the final cross-modal feature representation; f IR represents the image features enhanced by social relationship features; f RI represents the social relationship features enhanced by image features;

[0071] Concatenate the cross-modal features between the above modalities to obtain the final multi-modal fusion feature:

[0072] Z = concat(f TI , f IT, f TR , f RT , f IR , f RI )

[0073] Among them, concat represents the feature concatenation operation.

[0074] Furthermore, the method of inputting the multi-modal fusion feature vector into a classifier and outputting true / false news labels through the classifier includes:

[0075] Obtain and concatenate cross-modal fusion feature vectors between different modalities;

[0076] Input the multi-modal fusion feature vector into a classifier composed of a fully connected layer, and the output of the classifier passes through the softmax activation function to obtain the prediction score. The specific calculation formula is as follows:

[0077]

[0078] Among them, represents the prediction score, and MLP represents the multi-layer perceptron;

[0079] Output the true / false label of the news according to the prediction score. Among them, the false news label is 0, and the true news label is 1;

[0080] Use the cross-entropy loss function to calculate the classification loss. The specific formula is as follows:

[0081]

[0082] Among them, y represents the actual label;

[0083] Comprehensively consider the loss of the cross-modal alignment task and the classification loss, set the weight parameters to obtain the final total loss function. The specific calculation formula is as follows:

[0084]

[0085] Among them, λ a and λ b are weight parameters; is the classification loss, is the loss of the cross-modal alignment task;

[0086] Optimize the parameters of the classifier based on the total loss function until the classifier converges and outputs the final true / false news label.

[0087] To achieve the above object, the second aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the false news detection method based on multi-modal information fusion.

[0088] Advantages of the present invention:

[0089] Compared with the prior art, a fake news detection method and medium based on multi-modal information fusion provided by the present invention improve the structured information expression ability of the social relationship graph by introducing an improved graph attention network (GAT) to complete and denoise the social relationship graph; extract semantic information from text and images using entity features, enhance the intra-modal semantic expression while narrowing the semantic gap between modalities; adopt a cross-modal co-attention mechanism to deeply fuse the features of text, images, and social relationship graphs, fully exploring the correlation between modalities; and align the feature vectors of different modalities through self-supervised learning, thereby effectively reducing the inconsistency and noise interference between modalities, and achieving a significant improvement in the accuracy and robustness of fake news detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.

[0091] Figure 1 It is a cross-modal fake news detection framework diagram based on semantic enhancement disclosed in the embodiments of the present invention.

[0092] Figure 2 It is a schematic diagram of an improved GAT disclosed in the embodiments of the present invention.

[0093] Figure 3 It is a pseudo-code diagram of an operation for calculating the text feature gradient during training disclosed in the embodiments of the present invention.

[0094] Figure 4 It is a visualization result diagram of SCCNs and its variants on the PHEME dataset disclosed in the embodiments of the present invention.

[0095] Figure 5 It is a visualization result diagram of SCCNs and its variants on the Weibo dataset disclosed in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0096] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0097] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the following methods, in some cases, the steps shown or described can be executed in a different order than here.

[0098] Figure 1 A cross-modal fake news detection framework based on semantic enhancement provided by the present invention first obtains image, text, and relationship information from the news. Then, their feature vectors are obtained through feature extraction, where vision and text are content information, and social relationships are structural information. At the same time, entities in the image and text are extracted and encoded to strengthen the content information. To facilitate subsequent fusion, modal alignment is performed. In addition, enhanced features between modalities are obtained through co-attention. Finally, the enhanced features of all modalities are integrated to obtain cross-modal fusion features for fake news detection. The technical solutions of the present invention will be described in detail below:

[0099] The present invention provides a method for detecting fake news based on multi-modal information fusion, including the following steps:

[0100] Step S100, obtaining the text, image, user, and comment information of the news;

[0101] Step S200, constructing a social relationship graph including news nodes, user nodes, and comment nodes according to the interaction relationship between users and comments in the news content;

[0102] Step S300, based on the calculation of node similarity in the social relationship graph, removing low-similarity connections, inferring potential connections, and generating an improved social relationship graph;

[0103] Step S400, using a pre-trained model to perform feature encoding on the text and image respectively to generate a text feature vector and an image feature vector, and obtaining a social relationship feature vector from the improved social relationship graph;

[0104] Step S500, extracting text entities and image entities from the text and image;

[0105] Step S600, using the extracted text entities and image entities to embed and enhance the text feature vector and the image feature vector respectively to generate an enhanced text feature vector and an enhanced image feature vector;

[0106] Step S700, mapping the enhanced text feature vector, the enhanced image feature vector, and the social relationship feature vector to a common semantic space, performing modal alignment by optimizing the mean square error loss, and obtaining the feature vectors of the cross-modally aligned image, text, and social relationship graph;

[0107] Step S800: The feature vectors of the aligned images, texts, and social relationship graphs are fused using a co-attention mechanism to generate a multi-modal fusion feature vector containing all modal information.

[0108] Step S900: Input the multi-modal fusion feature vector into a classifier, and output the true / false news label through the classifier.

[0109] In this embodiment, as described in step S100 above, first obtain the multi-modal information constituting the news content from a preset dataset. This information includes:

[0110] Text information t j : The text description part of the news, including the title and the body text.

[0111] Image information v j : The pictures or video frames associated with the news.

[0112] User information u j : The relevant information of the social media user who posts the news or comments.

[0113] Comment information c j : The user comments on the news and their interaction situations.

[0114] To effectively perform subsequent processing, it is necessary to ensure that each news item collects the above four types of information and forms a quadruple record. Especially for comment information, it is further organized into a set where represents the i-th comment, related to news x j and it is clear that this comment is posted by user . The preliminary dataset D = {X, Y} formed in this way enables each news to be represented by a quadruple x j = {t j , v j , u j , c j} ∈ D, where D represents the news dataset, containing multi-modal information and its labels, X represents the set of news, that is, the multi-modal information part in {X, Y}, Y is the set of news labels (true or false), and x j represents a single news.

[0115] In this embodiment, as described in step S200 above, based on the data obtained in step S100, in order to intuitively display the complex relationships among news, users, and comments, construct a heterogeneous social relationship graph G = {V, A, E}, where G represents the heterogeneous social relationship graph, composed of a node set V, an adjacency matrix A, and an edge set E. The specific implementation process is as follows:

[0116] Node construction: Define a set of nodes V, including three types of nodes:

[0117] News node: Represents each news item x j .

[0118] User node: Represents all independent users u who participate in news publication and commenting j .

[0119] Comment node: Represents each comment

[0120] Adjacency matrix construction Initialize the adjacency matrix A, which reflects the link relationship between nodes. A value of 1 indicates the existence of a connection, and a value of 0 indicates no connection. Specifically:

[0121] If a certain user u j issues a comment then set the corresponding position in A to 1.

[0122] If a certain comment c j belongs to a certain news item x j , then set the corresponding position in A to 1.

[0123] If user u j directly publishes a certain news item x j , then set the corresponding position in A to 1.

[0124] Edge set construction (E): Collect all the connections marked as 1 in the adjacency matrix A. These are the edges in the social relationship graph, thus forming a complete edge set E.

[0125] The heterogeneous social relationship graph constructed in this way can more efficiently represent and analyze the interaction relationships among news, users, and comments, providing basic support for subsequent feature extraction and detection models.

[0126] In this embodiment, as described in step S300 above, the constructed social relationship graph is further processed to infer potential connections and remove noise. According to the homogeneity of the graph neural network, it is speculated that there is a higher probability of an association between similar nodes than between different nodes. Therefore, by calculating the similarity of features between nodes, the links between highly similar nodes can be inferred and the noise between low-similarity nodes can be removed. Specifically, for the embeddings of the nodes in G, text vectors and sentence vectors are used as the initial embeddings for news nodes and comment nodes respectively. Then, the mean of all the comment embeddings published by the user is used as its embedding vector. For the convenience of subsequent calculations, a node embedding matrix is constructed where each row represents the embedding of a node, and d is the dimension of the embedding vector.

[0127] Next, the cosine similarity is used to characterize node n i and n j The similarity coefficient a between them ij :

[0128]

[0129] where b i and b j represent the embeddings of nodes n i and n j respectively. It is considered that under the condition of no connection between nodes, if the similarity coefficient is greater than 0.5, it can be regarded as the existence of a potential connection. Similarly, if there is already a connection between nodes, but the similarity coefficient is less than 0.2, it can be considered as a noisy connection. The above statement can be expressed by the formula:

[0130]

[0131] where δ ij is a transitional variable used to measure the similarity relationship between nodes n i and n j Next, based on δ ij the original adjacency matrix A is improved, that is, the noisy edges are removed and the potential edges are added. The formula is:

[0132]

[0133] where a ij and a i ′ j are the elements in the initial and improved adjacency matrices A and A′ respectively. So far, the improved social relationship graph G′ = {V, A′, E} is obtained. After completing the above steps, the improved social relationship graph G′ is input into the second stage together with the image and text features for further processing. In this way, the filtered relationship graph can more accurately reflect the real connections in social interactions and provide a reliable data basis for the fake news detection task.

[0134] In this embodiment, as described in step S400 above, in the information extraction and feature encoding part, a synchronous processing mode is adopted for the three types of input data to obtain the required feature vectors. Then, entities with a confidence coefficient greater than 0.2 are retained as the results of the final entity extraction. In addition, a pre-trained transE model connects them to the Freebase knowledge graph to obtain the background knowledge features e I 、e T as entity embeddings.

[0135] Secondly, use the encoder to extract feature vectors from the text and images. Specifically, use the pre-trained BERT and ResNet-50 models as the text and image encoders respectively. However, for each news, the length of its text is basically different. To facilitate subsequent operations, make the text length of each news be L through padding or truncation operations:

[0136]

[0137] where d is the dimension of the embedding, T i represents the text embedding matrix of the i-th news, which is composed of word embeddings of L words, and each (where j ranges from 1 to L) is a d-dimensional word embedding vector, and this embedding sequence is fed into a bi-LSTM to obtain the final text feature a T :

[0138] a T = W T (Bi-LSTM(τ i )) + b T

[0139] where a T is the final text feature vector, which is the result of the linear transformation of the output obtained from the Bi-LSTM, W T is a learnable weight matrix, b T is a bias vector, and Bi-LSTM is a Bidirectional Long Short-Term Memory network, which is used to process sequence data and generate context-sensitive text features.

[0140] For images, extract the output of the penultimate layer of ResNet-50, and then obtain a feature vector a with the same dimension as the text feature through a fully connected layer: I :

[0141] a I = sigmoid(W I * R I ),

[0142] where R I is the output of the penultimate layer of ResNet-50, which is used to represent the original image features, W I is the weight matrix of the fully connected layer, which is used to transform the output of ResNet-50 into a feature vector with the same dimension as the text feature, a I is the final image feature vector, which is obtained by performing a non-linear transformation on the output passing through the fully connected layer using the sigmoid activation function.

[0143] Different from the above two types of representations, the social relationship graph contains rich structured information. Inspired by Velickovic et al., GAT can capture the structural information of the graph. However, traditional GAT has problems such as weak interpretability and inconsistent performance across datasets when the graph is complex. Therefore, this embodiment proposes an improved GAT to capture the correlation of neighbor nodes in order to obtain better graph feature representations, such as Figure 2 As shown, combining the parts of the two dashed boxes results in a new hybrid form, where the circles represent unnormalized attention and the diamonds are transition nodes.

[0144] For node n in graph G′ i and its set of neighbor nodes First, calculate the attention weights between node n i and all its neighbor nodes through the following formula where represents the attention weight between n i and . Combining two common attention mechanisms of traditional GAT, namely single-layer neural network (GO) and dot-product (DP), the formula is as follows

[0145]

[0146] where represents the inter-node attention coefficient calculated by combining the GO and DP attention mechanisms, is a parameter in GO, W is a learnable weight matrix, b i and b k respectively represent the embeddings of node n i and its neighbor node . || represents the vector concatenation operator, which concatenates two vectors to form a new vector, Applying the Leaky ReLU activation function to the attention coefficient to introduce non-linearity and prevent the problem of gradient vanishing.

[0147] Next, in order to facilitate the next calculation, it is necessary to normalize the obtained attention weights ε i using the softmax function. In addition, it is found that some attention weights become very small after normalization, which means that the influence of this neighbor node on node n i is very small. This situation occurs when the attention weight is negative. In fact, the attention weight objectively reflects the influence of the neighbor node on node n iPotential impacts, which include both positive and negative impacts, are reflected as positive and negative values in the attention weight numerical values. However, directly using the softmax function will ignore the negative impacts, which also need attention. For example, for a specific node n p , after calculation, its attention weight ε p = {0.7, 0.2, 0.1, -0.2, -0.8}. After normalization, its attention coefficient Φ p = {0.36, 0.22, 0.20, 0.14, 0.08}. It can be seen that the attention coefficient of the neighbor node with an attention weight value of -0.8 after normalization is 0.08, indicating that its contribution to the output result is almost negligible. However, this large negative impact implies that the embeddings of these two nodes are in opposite directions, which may be beneficial for fake news detection. This situation may be a kind of "fraud" and "disguise" behavior. For example, there may be some "bought" real comments or some malicious comments in the comments of certain fake news to slander a real news.

[0148] Therefore, a sign mechanism is introduced to correctly handle the positive and negative relationships between nodes. Specifically, for node n i , after taking the opposite of the attention weight, we get ε~ i . Then, softmax is used to normalize the two attention weights ε i and respectively to obtain their respective attention coefficients:

[0149]

[0150] where, and are both attention coefficients. To fully capture the interaction information between nodes, we use Φ i and Φ' i to perform feature weighted summation with the embeddings of the neighbor nodes of n i respectively. Then, the above vectors are concatenated and fed into a fully connected layer to obtain the vector representation of the final node n i . Note that in this process, multi-head attention is adopted to adapt to the complex graph structure, enabling it to fully consider the correlations and importance between different nodes, thereby improving the expression ability of the model, as shown in the following formula:

[0151]

[0152] where, K is the number of heads, σ is the activation function, W k is the weight matrix of the fully connected layer, and B i is the embedding matrix of the neighbor nodes.

[0153] So far, the final embeddings of all nodes can be calculated by the above formula and denoted as the node embedding matrix B'. Finally, the features of the social network graph are obtained using the multi-head attention mechanism:

[0154]

[0155] where K is the number of heads, the i-th row of G represents the graph feature of the i-th news, and B' is the node embedding matrix.

[0156] In this technical solution, an improved Graph Attention Network (GAT) method is proposed to better extract the structured information in the social relationship graph for fake news detection. First, the attention weights between nodes are calculated through a formula that combines a single-layer neural network and a dot product mechanism, and then the attention weights are processed by the LeakyReLU activation function. To solve the problem that negative weights are ignored by the traditional softmax, a sign mechanism is introduced to normalize the positive and negative weights respectively, so as to capture the positive and negative impacts of neighbor nodes on the target node. On this basis, the final representation of the node is obtained by weighted summation of the neighbor node features, and the multi-head attention mechanism is used to combine the correlations between different nodes to generate the complete social network graph features. This method not only enhances the interpretability of the model but also improves the performance consistency across datasets, providing richer and more accurate feature expressions for fake news detection.

[0157] In this embodiment, as described in the above steps S500 and S600, information enhancement and cross-modal fusion are performed using the feature vectors of the text, image, and social relationship graph obtained above. For the feature vectors of text and image, self-information enhancement is performed on themselves using their respective entity embeddings. Specifically, taking the image as an example, a I and e I are concatenated and then input into a multi-layer perceptron to obtain the information-enhanced feature vector Z I , and the calculation method of the feature vector Z I of text is similar:

[0158]

[0159] where σ is the activation function, W I ', W T ' are learnable matrices, and b I ' and b' T are bias vectors. For the social network graph, G is fed into a multi-layer perceptron to obtain a feature vector Z I with the same dimension as Z R .

[0160] In this embodiment, as described in the above step S700, it should not be overlooked that when performing cross-modal fusion operations, inevitable loss of intrinsic information between modalities will occur, and the representations of different modalities in the original news should have an inline relationship. This leads to a large semantic gap between features from different modalities. To solve this problem, a cross-modal alignment method with self-supervised loss is introduced to refine the feature representations. For example, for the feature vectors Z T 、Z I and Z R , map them to the same semantic space:

[0161]

[0162] where, represents the feature vector after the enhanced image feature vector Z I is mapped to the unified semantic space, represents the feature vector after the enhanced text feature vector Z T is mapped to the unified semantic space, represents the feature vector after the social relationship feature vector Z R is mapped to the unified semantic space, and are learnable matrices, Z I is the enhanced image feature vector, Z T is the enhanced text feature vector, Z R is the social relationship feature vector;

[0163] After that, taking images and texts as examples, the mean squared error loss is used to reduce the distance between and :

[0164]

[0165] where n is the total number of news. Similarly, Z T and Z R , Z I and Z R can be mapped to the same semantic space and their MSE losses can be calculated. Then, add the above three loss functions to obtain the total loss function:

[0166]

[0167] So far, the feature vectors of the cross-modally aligned images, texts, and social relationship graphs have been obtained.

[0168] In this embodiment, as described in the above step S800, considering that there are three types of feature vectors from different modalities here, it is necessary to integrate their embeddings before detection to improve the credibility. This embodiment adopts a cross-fusion method with a co-attention mechanism. Specifically, the aligned text feature vectors, image feature vectors, and social relationship feature vectors are fused, and the co-attention mechanism is used to capture the interaction information between modalities: The cross-modal features of text and image fusion are calculated by the following formula:

[0169]

[0170]

[0171] Where, represents the query vector of text features, which is calculated by the linear transformation matrix ; represents the key vector of image features, which is calculated by the linear transformation matrix ; represents the value vector of image features, which is calculated by the linear transformation matrix ; represents the query vector of image features, which is calculated by the linear transformation matrix ; represents the key vector of text features, which is calculated by the linear transformation matrix ; represents the value vector of text features, which is calculated by the linear transformation matrix ; W TI 、W IT are linear transformation matrices, which are used to obtain the final cross-modal feature representation; is the normalization coefficient for scaling, where d is the dimension of the feature vector; K is the number of attention heads, which represents the number of heads in the multi-head attention mechanism; f TI represents the text features enhanced by image features; f IT represents the image features enhanced by text features;

[0172] The cross-modal feature calculation formula of text and social relationship graph is as follows:

[0173]

[0174]

[0175] Where, represents the query vector of social relationship features, which is calculated by the linear transformation matrix ; represents the key vector of social relationship features, which is calculated by the linear transformation matrix ; The value vector representing the social relationship feature is obtained through a linear transformation matrix calculated; W TR and W RT are linear transformation matrices used to obtain the final cross-modal feature representation; f TR represents the text feature enhanced by the social relationship feature; f RT represents the social relationship feature enhanced by the text feature;

[0176] The cross-modal feature calculation formula for the image and the social relationship graph is as follows:

[0177]

[0178]

[0179] Among them, The query vector representing the image feature is obtained through a linear transformation matrix calculated; The key vector representing the social relationship feature is obtained through a linear transformation matrix calculated;

[0180] The value vector representing the social relationship feature is obtained through a linear transformation matrix calculated; W IR and W RI are linear transformation matrices used to obtain the final cross-modal feature representation; f IR represents the image feature enhanced by the social relationship feature; f RI represents the social relationship feature enhanced by the image feature;

[0181] Concatenate the cross-modal features between the above modalities to obtain the final multi-modal fusion feature:

[0182] Z = concat(f TI , f IT , f TR , f RT , f IR , f RI )

[0183] Among them, concat represents the feature concatenation operation.

[0184] In this embodiment, as described in step S800 above, feed the final multi-modal fusion feature Z into a fully connected layer to predict the label of the news:

[0185]

[0186] Among them, denotes the predicted score, and MLP denotes the multi-layer perceptron;

[0187] Output the true or false label of the news according to the predicted score, where the false news label is 0 and the true news label is 1;

[0188] Since the fake news detection task is a binary classification problem, the cross-entropy loss function is used to calculate the prediction loss, and the specific formula is as follows:

[0189]

[0190] where y represents the actual label;

[0191] Taking into account the loss of the cross-modal alignment task and the classification loss, set the weight parameters to obtain the final total loss function, and the specific calculation formula is as follows:

[0192]

[0193] where λ a and λ b are weight parameters; is the classification loss, is the loss of the cross-modal alignment task;

[0194] Optimize the parameters of the classifier based on the total loss function until the classifier converges and outputs the final true or false label of the news.

[0195] However, it is found in the actual operation process that the texts of some news do not strictly follow the grammar rules, which to a certain extent reduces the efficiency of the model operation. To solve the above problems, an adversarial perturbation mechanism is added when extracting text embeddings to enhance the robustness of the model. Specifically, when calculating the gradient of the text features during training, it is also necessary to calculate the perturbation added to the text features. Then recalculate its gradient and repeat the above process T times. Finally, all the adversarial perturbations are accumulated on the original gradient and the parameters are updated. Finally, all the adversarial perturbations are accumulated on the original gradient and the parameters are updated. Some specific operations can be seen in Figure 3 the pseudocode of.

[0196] In Figure 3 's pseudocode, sign(g) is a function that returns the sign of the gradient g, indicating the direction of perturbing the input to maximize the loss. In addition, clip and project are functions to ensure that the perturbation is within an acceptable range and project the perturbed input back to the valid data space.

[0197] The following will demonstrate the performance that the above scheme of the present invention can achieve in fake news detection in combination with the specific experimental process:

[0198] First, the performance of the model SCCNs and baselines was evaluated on two widely used datasets, PHEME and Weibo.

[0199] PHEME is a collection of Twitter posts about multiple breaking news and their related information. Similarly, Weibo is a Chinese dataset that collects a large number of posts on China's most popular social media. Both of the above datasets contain information such as text, images, users, and comments. In this paper, text, images, and social relationship networks are mainly used to achieve fake news detection. Therefore, the original datasets are processed, including data cleaning, filtering, etc. News with only text or pictures are deleted because the model of the present invention focuses on multi-modal news. The relevant data of the processed PHEME and Weibo datasets are shown in Table 1.

[0200] Table 1: Statistical information of the experimental datasets after removal

[0201]

[0202] The baselines used in this paper are listed below for comparison with the model of the present invention.

[0203] 1) EANN uses a multi-modal feature extractor and a fake post detector to support fake news detection. It can derive event-invariant features, facilitating the detection of newly arrived events.

[0204] 2) MVAE uses a bimodal variational autoencoder to model images and text and achieve classification.

[0205] 3) QSAN integrates a quantum-driven text encoding and signature mechanism in the framework, which can use conflict information to provide clues for detection, and the method is interpretable.

[0206] 4) SAFE is a similarity-aware fake news detection method that pays more attention to the similarity between text and visual information than other methods.

[0207] 5) EBGCN identifies the unreliable relationships existing in rumors and realizes fake news detection by training an edge consistency framework.

[0208] 6) GLAN is a global-local network that captures the structural information for fake news detection by jointly encoding global and local information.

[0209] 7) MPFN can identify the information levels represented in different modalities and use it to construct a powerful hybrid modality.

[0210] 8) KMAGCN is an adaptive graph convolutional network that converts posts into graphs to capture discontinuous semantic relationships.

[0211] 9) MFAN introduces comment elements in posts while considering the complementation and alignment between different modalities for better integration.

[0212] First, the PHEME and Weibo datasets are divided into training, validation and test sets according to the ratio of 7:1:2. Second, all word embeddings in the model of the present invention are initialized with vectors of dimension 300, i.e., d = 300. Those words that are not in the pre-trained word vectors are initialized from a uniform distribution. In this paper, the number of heads K involved in the multi-head attention mechanism is set to 8. And the parameter value λ in the loss function is calculated a and λ b They are 2.15 and 1.55 respectively. The above parameter values ​​were finally determined based on the performance of the model and previous experience. In the experiment, the learning rate during data training was set to 0.002. In addition, the average value of the results after running 5 times was taken as the final result of the model. Finally, the most commonly used indicators, accuracy, precision, recall rate and F1 score, were taken as indicators to measure the performance of the model of the present invention.

[0213] In Table 2, the performance of the model is compared with the baselines on the PHEME and Weibo datasets. For ease of comparison, the best baseline is underlined and the best result is bolded. It can be seen that the SCCNs of the present invention outperform the baseline model in both PHEME and Weibo, with accuracy rates of 90.4% and 89.5%, respectively. In addition, it is found that the performance of KMAGCN and MFAN in the baseline model is much higher than that of the others. Their common point is that structural information is introduced to enhance the model. This proves the importance of structural information for fake news detection. However, since the PHEME dataset is derived from Twitter, most posts are related to certain specific things and the correlation between them is very small, which is more likely to cause overfitting. In contrast, SCCNs can perform better on both datasets. This shows that the use of entity features to enhance text and image expressions can refine feature vectors, which is more conducive to subsequent fusion. At the same time, the social network relationship composed of news, users and their comments is introduced, and the potential structural information in the news is mined here. Specifically, the model of the present invention is higher than the baseline model in all indicators, especially the F1 score, which further illustrates that the performance of SCCNs is very good.

[0214] Table 2: Performance comparison of SCCNs with other methods on PHEME and Weibo datasets

[0215]

[0216] To further explore the impact of each part of the model of the present invention on performance, a series of ablation experiments were conducted. According to the idea of controlling variables, the corresponding results were obtained after removing one part of the model in each experiment, which can reflect the importance of this part to the overall model.

[0217] Specifically, the variants of the model SCCNs of the present invention can be classified into 4 categories, and they are listed below:

[0218] 1) SCCNs w / o A, which removes the processing of all original data in the data preparation part and instead directly uses the original data as input.

[0219] 2) SCCNs w / o B, which removes the social relationship part and only considers two modalities, text and image, to achieve fake news detection.

[0220] 3) SCCNs w / o C, which removes the information enhancement part and simply uses the original embedding vectors as the input for cross-modal fusion.

[0221] 4) SCCNs w / o D, which removes the cross-modal fusion module and uses the sum of the feature vectors of text, image, and social relationship as an alternative.

[0222] Table 3: Ablation study and efficiency analysis of different modules in SCCN on two datasets

[0223]

[0224] Table 3 intuitively and clearly lists the results of the ablation experiments. Here, only two indicators that can best reflect the model performance, accuracy and F1 score, are listed. Generally speaking, all variants of SCCNs perform worse than the original model in terms of accuracy and F1 score, which indicates that each module of our model makes an important contribution to performance. Specifically, the following points can be summarized: 1) Comparing the four variants of SCCNs, it is found that except for the relatively high indicator values of SCCNs w / o D, the indicator values of the other three variants are similar, which shows that data processing, the introduction of social relationships, and information enhancement all enhance the performance of the model of the present invention. 2) Although the performance of SCCNs w / o D is the best among the variants, it still shows relatively weak performance compared to SCCNs, which proves that the cross-modal fusion module implemented by the co-attention mechanism can also help the model of the present invention improve performance. 3) All variants show similar performance on the two datasets of PHEME and Weibo, and SCCNs perform better on PHEME than on Weibo, which indicates that these modules or parts play a greater role on the PHEME dataset.

[0225] In the ablation experiment, the impact of different modules on the whole can be basically understood. To more intuitively show the contributions of each part of the model of the present invention, SCCNs and its variants are respectively visualized on two datasets, PHEME and Weibo. Two different methods are adopted to realize the visualization of data. Specifically, heatmaps are used for visualization in PHEME, while T-SNE is used for visualization in Weibo. Considering that most posts in PHEME are related to certain specific events, 20 news items are randomly selected from them, including 10 true news items and 10 false news items, and then they are respectively input into SCCNs and its variants to obtain Figure 4 , Figure 4 which are the visualization results of SCCNs and its variants on the PHEME dataset. The values in the boxes represent the correlations between news items, which are represented by the correlations of news prediction labels. At the same time, T-SNE visualization is used to deeply analyze the features in the model proposed by the present invention, and these features are all learned on SCCNs and its variants.

[0226] Figure 4 The fake news detection capabilities of the model and variants of the present invention are compared. It can be noted that the boundaries between true news and fake news in SCCNs are obvious, while similarities are shown among their own categories. However, the boundaries between true news and fake news in the remaining methods seem to be blurred, which means their performance is poor and it is not easy to identify fake news. The visualization results here are similar to those of the ablation experiment, that is, the performance of SCCNs w / o B and SCCNs w / o C is inferior to that of SCCNs w / o A and SCCNs w / o D. This shows that the experimental results of the present invention have a certain degree of credibility. Social relationships are not considered in SCCNs w / o B, and only text and image features are combined, which may lead to the omission of some key information hidden in the comments. Social relationships are not considered in SCCNs w / o B, and only text and image features are combined, which may lead to the omission of some key information hidden in the comments. Compared with SCCNs, SCCNs w / o C not only do not use entity features to enhance the information of text and image modalities, but also do not process the information of the social relationship graph with the improved GAT. Compared with SCCNs, SCCNs w / o C not only do not use entity features to enhance the information of text and image modalities, but also do not process the information of the social relationship graph with the improved GAT.

[0227] Different from the above, the present invention adopts the T-SNE visualization method to realize the visualization of the test set data of Weibo, as Figure 5 shown Figure 5Visualization results of SCCNs and its variants on the Weibo dataset. Dots of the same color belong to the same label. It can be observed that the visualization results are also consistent with the ablation experiments. Generally speaking, whether from the ablation experiments or the quantitative analysis, each module in the model proposed by the present invention plays a beneficial role in fake news detection.

[0228] In summary, the present invention proposes a new semantic-enhanced cross-modal co-attention network to study fake news detection, introduces social relationships with structural information, and uses the co-attention mechanism to achieve cross-modal fusion with text and images. It introduces social relationships with structural information and uses the co-attention mechanism to achieve cross-modal fusion with text and images. In addition, hidden links in the social relationship graph are inferred and possible noises are removed. At the same time, a perturbation mechanism is used to enhance the robustness of the model. Finally, the results of all experiments show that the performance of SCCNs proposed by the present invention is superior to other existing models.

[0229] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0230] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.

[0231] In addition, each functional unit in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0232] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0233] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A fake news detection method based on multimodal information fusion, characterized in that: The steps include: Get news text, images, users and comments information; According to the interactive relationship between users and comments in news content, a social relationship graph including news nodes, user nodes and comment nodes is constructed; Based on the node similarity calculation in the social relationship graph, low-similarity connections are removed, potential connections are inferred, and an improved social relationship graph is generated; Use the pre-trained model to encode the features of text and image respectively, generate text feature vector and image feature vector, and obtain the social relationship feature vector from the improved social relationship graph; Extracting text entities and image entities from the text and image; The extracted text entities and image entities are respectively embedded to enhance the information of the text feature vector and the image feature vector, thereby generating an enhanced text feature vector and an enhanced image feature vector; The enhanced text feature vector, enhanced image feature vector and social relationship feature vector are mapped to a common semantic space, and inter-modal alignment is performed by optimizing the mean square error loss to obtain the feature vectors of the image, text and social relationship graph after cross-modal alignment. The feature vectors of the aligned image, text, and social graph are fused using a common attention mechanism to generate a multimodal fusion feature vector containing all modal information. The multimodal fusion feature vector is input into a classifier, and the classifier outputs true or false news labels.

2. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: The method for generating an improved social relationship graph includes: Calculate the initial embedding features of the nodes. For news nodes, extract embedding features based on their text content; for comment nodes, extract embedding features based on their content; for user nodes, generate user embedding features based on the mean of the embedding features of their comments; Based on the initial embedding features of the nodes, the similarity between any two nodes is calculated, and the degree of association between the nodes is quantified using cosine similarity; the similarity threshold range between the nodes is determined, and the upper and lower thresholds for judging the node connection relationship are set; According to the similarity calculation results, the connections in the social relationship graph are judged: if there is a connection between the nodes and the similarity is lower than the lower threshold, it is marked as a noise connection; if there is no connection between the nodes but the similarity is higher than the upper threshold, it is marked as a potential connection; the rest of the connections remain unchanged; According to the identified noise connections and potential connections, the social relationship graph is adjusted: the noise connections are deleted, potential connections are added to complete the hidden associations in the social relationship graph, and the adjacency matrix is ​​updated to reflect the improved connection relationship; Based on the optimized adjacency matrix and node information, an improved social relationship graph is generated.

3. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: Methods for generating text feature vectors include: Extract the text of each news item from the social media news dataset; Preprocess the text, including word segmentation, removal of stop words, and standardization of text length to make the text length of each news item consistent; Use the pre-trained BERT model to encode the processed text and generate a word embedding sequence; Input the word embedding sequence into the bidirectional long short-term memory network to model the contextual information of the text; The output of the bidirectional long short-term memory network is aggregated by weighted summation, and the final text feature vector is calculated through the fully connected layer.

4. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: Methods for generating image feature vectors include: Extract the image corresponding to each news item from the social media news dataset; Preprocess the image, including image size standardization, denoising and normalization; Use the pre-trained ResNet-50 model to extract features from the image and obtain its penultimate layer output as the initial image features; The initial image features are input into the fully connected layer, and the initial image features are nonlinearly transformed through the activation function to obtain the final image feature vector.

5. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: The method of obtaining the social relationship feature vector from the improved social relationship graph includes: Input the improved social relationship graph into the graph attention network; Based on the graph attention network, the relationship between each node and its neighboring nodes is modeled, and the attention weight is calculated using the node embedding vector. By combining the linear transformation of the node embedding and the dot product operation, the initial attention weight between the nodes is obtained to represent the relationship strength of the neighboring node to the target node. The initial attention weights between nodes are expanded into two types of weights: positive and negative. The positive weights are used to reflect the positive impact of the neighboring nodes on the target node, and the negative weights are used to capture the negative impact of the neighboring nodes on the target node. The positive and negative weights are normalized respectively to correctly characterize the positive and negative contributions of the neighboring nodes. Based on the positive and negative normalized attention weights, the neighborhood node features of the target node are weighted and summed, the weighted results of the positive weights and the negative weights are aggregated respectively to generate the feature representation of the target node, and the positive and negative aggregation results are spliced ​​into the initial feature vector of the target node; The multi-head attention mechanism is applied to further fuse the feature representations of all nodes, extracting the feature relationships of target nodes in complex graph structures from different attention heads to generate the final feature representation of each node; The final feature representations of all nodes in the graph are aggregated and fused to generate a social relationship feature vector.

6. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: The method of generating the enhanced text feature vector and the enhanced image feature vector includes: Extract text and images of each news item from social media news data; Perform entity recognition on the news text to extract key entities in the text, and generate text entity embedding vectors through the entity embedding model. At the same time, perform object detection on the image to extract entity information in the image and generate image entity embedding vectors. The text feature vector is concatenated with the text entity embedding vector, and an enhanced text feature vector is generated through a multi-layer perceptron; the image feature vector is concatenated with the image entity embedding vector, and an enhanced image feature vector is generated through a multi-layer perceptron.

7. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: Methods for obtaining feature vectors of cross-modal aligned images, texts, and social graphs include: Map the enhanced text feature vector, enhanced image feature vector and social relationship feature vector to a common semantic space to obtain the aligned feature vector: in, Represents the enhanced image feature vector Z I The feature vector after being mapped to the unified semantic space, Represents the enhanced text feature vector Z T The feature vector after being mapped to the unified semantic space, Represents the social relationship feature vector Z R The feature vector after being mapped to the unified semantic space, and is the learnable matrix, Z I is the enhanced image feature vector, Z T is the enhanced text feature vector, Z R is the social relationship feature vector; The loss of aligning the enhanced text feature vector and the enhanced image feature vector is calculated using the mean square error loss function; The mean square error loss function is used to calculate the loss of alignment between the enhanced text feature vector and the social relationship feature vector; The mean square error loss function is used to calculate the loss of alignment between the enhanced image feature vector and the social relationship feature vector; The above mean squared error losses are added together to get the total alignment loss.

8. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: The method of generating a multimodal fusion feature vector containing all modal information includes: The aligned text feature vector, image feature vector and social graph feature vector are fused, and the interaction information between modalities is captured using the joint attention mechanism: The cross-modal features of text and image fusion are calculated by the following formula: in, The query vector representing the text feature is transformed by the linear transformation matrix Calculated; The key vector representing the image features is transformed by the linear transformation matrix Calculated; The value vector representing the image features is transformed by the linear transformation matrix Calculated; The query vector representing the image features is transformed by the linear transformation matrix Calculated; The key vector representing the text feature is transformed by the linear transformation matrix Calculated; The value vector representing the text feature is transformed by the linear transformation matrix Calculated; W TI , W IT is a linear transformation matrix used to obtain the final cross-modal feature representation; The normalization coefficient used for scaling, where d is the dimension of the feature vector; K is the number of attention heads, indicating the number of heads in the multi-head attention mechanism; f TI represents the text features enhanced by image features; f IT Represents image features enhanced with text features; The formula for calculating the cross-modal features of text and social relationship graph is as follows: in, The query vector representing the social relationship features is transformed by the linear transformation matrix Calculated; The key vector representing the social relationship features is transformed by the linear transformation matrix Calculated; The value vector representing the social relationship characteristics is transformed by the linear transformation matrix Calculated; W TR , W RT is a linear transformation matrix used to obtain the final cross-modal feature representation; f TR Represents text features enhanced with social relationship features; f RT Represents social relationship features enhanced by text features; The formula for calculating the cross-modal features of images and social relationship graphs is as follows: in, The query vector representing the image features is transformed by the linear transformation matrix Calculated; The key vector representing the social relationship features is transformed by the linear transformation matrix Calculated; The value vector representing the social relationship characteristics is transformed by the linear transformation matrix Calculated; W IR , W RI is a linear transformation matrix used to obtain the final cross-modal feature representation; f IR Image features that represent social relationship feature enhancement; f RI Represents social relationship features enhanced by image features; The cross-modal features between the above modalities are spliced ​​to obtain the final multimodal fusion features: Z=concat(f TI ,f IT ,f TR ,f RT ,f IR ,f RI ) Among them, concat represents the feature concatenation operation.

9. The fake news detection method based on multimodal information fusion as claimed in claim 1, characterized in that: The method of inputting the multimodal fusion feature vector into a classifier and outputting true or false news labels through the classifier includes: Obtain and concatenate cross-modal fusion feature vectors between different modalities; The multimodal fusion feature vector is input into a classifier consisting of a fully connected layer, and the output of the classifier is used to obtain a prediction score through a softmax activation function. The specific calculation formula is as follows: in, represents the prediction score, MLP represents multi-layer perceptron; Output the true or false label of the news based on the predicted score, where the false news label is 0 and the true news label is 1; The cross entropy loss function is used to calculate the classification loss. The specific formula is as follows: Among them, y represents the actual label; Taking into account the loss of the cross-modal alignment task and the classification loss, the weight parameters are set to obtain the final total loss function. The specific calculation formula is as follows: Among them, λ a and λ b is the weight parameter; is the classification loss, is the loss for the cross-modal alignment task; The classifier parameters are optimized based on the total loss function until the classifier converges and outputs the final news true or false label.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the fake news detection method based on multimodal information fusion according to any one of claims 1 to 9 are performed.

Citation Information

Cited By

  • Social governance work order dispatching method based on emergency degree evaluation, medium and equipment

    CN120822923A

  • False information detection method and system based on multi-modal decoupling learning

    CN121032528A

  • Transform-based multi-modal graph noise reduction false news detection method and system

    CN121167465A

  • Transform-based multi-modal graph denoising fake news detection method and system

    CN121167465B