False information detection method and device based on multi-dimensional feature collaborative fusion network and computer readable storage medium thereof

Through the multi-dimensional feature collaborative fusion network, the problem of single feature fusion strategy and noise interference in multi-modal false information detection is solved, and more accurate false information detection and cross-modal fusion accuracy are achieved, which is suitable for a variety of multi-modal data processing.

CN120338812APending Publication Date: 2025-07-18ZHENGZHOU UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510328809.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing multimodal false information detection technology has the problems of single feature fusion strategy, serious noise interference, insufficient dynamic context perception ability, and weak generalization across data sets, making it difficult to effectively deal with multimodal forgery methods.

Method used

A multi-dimensional feature collaborative fusion network is adopted to achieve dynamic interaction and fusion of image and text features through frequency domain, airspace, semantics and emotional features extraction, combined with a two-layer collaborative attention mechanism and CLIP model, and optimize the semantic alignment of graphics and text.

Benefits of technology

It significantly improves the richness and robustness of feature representations, can detect false information more accurately, has stable performance in high noise environments, and is suitable for a variety of multimodal data processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338812A_ABST
    Figure CN120338812A_ABST
Patent Text Reader

Abstract

The invention discloses a false information detection method and device based on a multi-dimensional feature collaborative fusion network and a computer readable storage medium. According to the method, a multi-dimensional feature extraction network is constructed, and fine-grained features are extracted from an image frequency domain, a spatial domain, text semantics and emotion dimensions; a cross-modal collaborative attention fusion module (CCF) is designed, multi-modal feature fusion of dynamic context sensing is achieved, and multi-source data noise is effectively eliminated; in combination with a CLIP-based cross-modal consistency learning branch, the image-text semantic consistency detection capability is enhanced; and finally, multi-dimensional features are weighted and fused through an attention mechanism, and a detection result is optimized. The method is suitable for efficient identification of multi-modal false information in a social media platform, and has high robustness and generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal false information detection, and specifically relates to a false information detection method, device and computer-readable storage medium based on a multi-dimensional feature collaborative fusion network. Background Art

[0002] With the rapid development of social media platforms, the speed of information dissemination has increased significantly, but at the same time, it has become a breeding ground for false information. False information often misleads the public's perception through carefully fabricated graphic and text content, causing far-reaching negative impacts on key social fields such as political elections and public health events. Traditional single-modal detection methods (such as text-based sentiment analysis or image forgery detection) are no longer sufficient to cope with the increasingly complex multimodal forgery methods. Therefore, multimodal false information detection technology has gradually become a research hotspot, and its core challenge lies in how to effectively integrate multi-source information such as text and images, and accurately capture the semantic consistency or contradiction between modalities.

[0003] Existing methods are mainly divided into two categories: detection techniques based on feature fusion and detection techniques based on cross-modal consistency. The former realizes detection by extracting and fusing multimodal features, but the feature fusion strategy of such methods is too simple (such as direct splicing), fails to effectively handle the noise interference in multi-source data, and ignores the dynamic interaction relationship between modalities. In addition, methods based on pre-trained models (such as BERT, ViT) can efficiently extract single-modal features, but lack targeted analysis of frequency-domain forgery traces (such as image recompression, splicing artifacts), and are difficult to cope with highly concealed forgery means. The latter focuses on mining the semantic consistency between graphics and text, but such methods may misjudge real information due to over-reliance on semantic alignment. In addition, the existing research defines the inconsistency between modalities rather generally, lacking targeted modeling of specific forgery scenarios (such as the combination of text sentiment misdirection and local image tampering), resulting in limited generalization ability.

[0004] The current technology has the following core problems: the feature fusion strategy is single, most methods rely on simple splicing or shallow attention mechanisms, and it is difficult to dynamically perceive context information, resulting in noise features interfering with the detection results. The frequency domain and sentiment features are not fully utilized. Frequency domain analysis can effectively identify image forgery traces, but existing research mainly focuses on the extraction of spatial domain features, ignoring the discriminative value of frequency domain clues. In terms of text, there is a significant correlation between sentiment features and the spread of fake news, but traditional methods do not construct a sentiment knowledge graph, making it difficult to capture context-sensitive sentiment semantics. The lack of dynamic context awareness ability, existing fusion modules mostly use fixed weight allocation and cannot adaptively adjust the feature contribution according to the input content. Weak cross-dataset generalization ability, most models perform well on a single dataset, but their performance drops significantly when transferred to other fields, especially in the face of cross-language or cross-topic scenarios, lacking robust feature representation capabilities. Summary of the Invention

[0005] The present invention provides a method, apparatus and computer-readable storage medium for detecting false information based on a multi-dimensional feature collaborative fusion network, aiming to solve the limitations in multi-modal false information detection in the prior art.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] Design a method for detecting false information based on a multi-dimensional feature collaborative fusion network, including the following steps:

[0008] S1: Extract frequency-domain features from image data, generate a frequency-domain view using the Error Level Analysis (ELA) algorithm, and extract frequency-domain tampering features through ResNet50;

[0009] S2: Extract spatial domain semantic features of the image, and extract spatial domain semantic features using a Masked Autoencoder (MAE);

[0010] S3: Extract text semantic features from text data using BERT;

[0011] S4: First construct an emotion knowledge graph, and then generate text emotion knowledge embeddings through a Graph Convolutional Network (GCN);

[0012] S5: Construct a cross-modal collaborative attention fusion module through a double-layer collaborative attention mechanism to achieve dynamic interaction and fusion of image and text features;

[0013] S6: Extract semantic features of images and texts based on the CLIP model, optimize the image-text semantic alignment through cosine embedding loss, and calculate similarity scores to guide the fusion;

[0014] S7: After splicing the image frequency-domain and spatial-domain features, text semantic and emotion features, and cross-modal fusion features, dynamically adjust the contributions of each feature through an attention weight matrix to achieve multi-feature dynamic fusion. Finally, output multi-dimensional fusion features for false information detection.

[0015] Further, in step S1, first use the ELA algorithm to obtain a full-size frequency-domain view ELA(I). To obtain deeper features, use ResNet50 trained on the ImageNet dataset as a tampering feature extractor, so as to obtain the frequency-domain tampering feature vector of the false information image:

[0016] h ME = ResNet50(ELA(I))

[0017] where I represents the input original image, ELA(I) represents the image processed by ELA, and h MERepresents the tampering features extracted by ResNet50.

[0018] Further, in step S2, a masked autoencoder (MAE) is used as an image semantic analyzer to extract h I = MAE(I).

[0019] Further, in step S3, BERT is used to extract text semantic features h T = BERT(T).

[0020] Further, in step S4, the construction of a sentiment knowledge graph based on SenticNet is introduced. Building a sentiment knowledge graph first requires preprocessing the original text to retain only words containing letters and numbers. The sentiment knowledge graph is represented as G = (V, E, H), where V and E are the sets of sentiment nodes and edges respectively, and H represents the set of features for each sentiment node. There are two types of nodes in the graph, V1 is the words in the original text, V2 is the external sensitivity and sentiment knowledge of words in SenticNet, and V1 ∪ V2 = V; each node v i in the graph has an associated lexical embedding vector h i , which is extracted by the bert-base-uncased pre-trained model.

[0021] Further, the adjacency matrix A is an n×n matrix, where n is the number of nodes. Suppose there is an edge e i between the lexical node v j in the original text and the lexical node v ij in SenticNet. The weight w ij of the edge is proportional to the number of times the two words co-occur, and the positive pointwise mutual information (PPMI) matrix is used to capture the co-occurrence information of sentiment nodes PPMI(v i , v j ).

[0022]

[0023]

[0024] Among them, the first formula defines the adjacency matrix, P(v i ) represents the frequency of occurrence of node v i , and P(v i , v j ) is the frequency of co-occurrence of nodes v i and v j in the sampled random walk.

[0025] Furthermore, graph embedding calculation is performed. Since the graph-level sentiment knowledge representation depends on each node, the sentiment features of each node are first updated. For node v in the graph i , information is obtained from neighboring nodes N(v i ) through a graph convolutional network (GCN), and the node features are updated. The feature update formula for node v i is as follows:

[0026]

[0027] where represents the set of neighboring nodes of node v i , d i and d j are the degrees of node v i and v j respectively (i.e., the number of edges directly connected to the node), W (l) is the learnable weight matrix of the l-th layer, and σ is the activation function (such as ReLU).

[0028] Furthermore, based on the GCN, the graph-level sentiment knowledge embedding h SA is finally generated, which can be done through the following formula:

[0029]

[0030] where is the readout function of graph G, and here the global average operation of each node is adopted. K represents the number of layers of the GCN.

[0031] Furthermore, in step S5, a double-layer collaborative attention mechanism is adopted. First, the left visual stream uses its own query Q I and the key K T and value V T in the right language stream for attention calculation. Then, the text-image fusion feature is deeply fused with the text semantic feature again to strengthen the text information, so as to obtain the text-image semantic fusion feature of the enhanced text information processed by the cross-modal collaborative attention fusion module. Similarly, the image fusion feature of the enhanced spatial domain information and the text fusion feature

[0032] of the enhanced text sentiment knowledge are obtained.

[0033] Furthermore, a cross-modal consistency learning auxiliary task is set up to maximize the semantic similarity of positive sample pairs and suppress negative sample pairs through the cosine embedding loss with a margin of d. After that, the text features h extracted by the BERT and MAE encoders are concatenated T and the image features h I , and they are input into a linear layer with a sigmoid function. Then, a multiplicative gate is used to weight the fusion of text and image semantic features to obtain the cross-modal fusion features

[0034] Furthermore, in step S7, the image frequency domain features h ME , the spatial domain features h I , the text semantic features h T , the text sentiment features h SA and the cross-modal fusion features are concatenated into a multi-modal representation h. Then, through a learnable attention weight matrix W Att the contribution weights of each feature channel are dynamically adjusted to obtain the output joint representation h Fusion .

[0035] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0036] 1. The present invention extracts features from four dimensions: frequency domain, spatial domain, semantics, and sentiment, significantly enhancing the richness and robustness of feature representation. Specifically, the image branch combines frequency domain (ELA algorithm) and spatial domain (MAE model) features, which can simultaneously capture the physical tampering traces and semantic content of images; the text branch combines semantic (BERT) and sentiment (SenticNet) features, which can more comprehensively understand the context information and sentiment tendency of text. Through multi-dimensional feature extraction, the present invention can more accurately represent multi-modal data and provide high-quality feature input for subsequent false information detection tasks.

[0037] 2. The present invention significantly improves the ability of noise suppression and key information enhancement through multi-dimensional feature collaborative fusion and a dynamic attention mechanism. Specifically, the complementarity of frequency domain, spatial domain, semantic, and sentiment features can effectively distinguish noise from effective features; the attention weight matrix dynamically adjusts feature contributions, suppressing irrelevant noise and enhancing key information. This design makes the present invention show stronger robustness in multi-modal data processing and can maintain high performance in a high-noise environment.

[0038] 3. The present invention optimizes the semantic alignment capability of images and texts through a consistency learning module based on CLIP. Specifically, the CLIP model uses large-scale pre-training data to extract semantic features of images and texts and capture cross-modal semantic associations; the cosine embedding loss optimizes the semantic alignment of images and texts to enhance the model's adaptability to local and global consistency. This design enables the present invention to more accurately capture the alignment relationship of image and text semantics and improve the accuracy and robustness of cross-modal fusion.

[0039] 4. The present invention is not only applicable to the processing of image-text data, but can also be extended to other multimodal scenarios (such as video-audio, image-speech, etc.). Through modular design, each component (such as CCF module, consistency learning module) can be flexibly replaced or expanded to meet different task requirements. This design makes the present invention widely applicable and scalable, and can meet the needs of various multimodal data processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the structure of the multi-dimensional feature collaborative fusion network of the present invention.

[0041] Figure 2 Flowchart for text sentiment feature extraction based on SenticNet.

[0042] Figure 3 Schematic diagram of the cross-modal collaborative attention fusion module.

[0043] Figure 4 Schematic diagram of the collaborative attention layer.

[0044] Figure 5 Table 1 is the comparative experimental results of different methods in the present invention.

[0045] Figure 6 Table 2 is the ablation experiment results in the present invention. DETAILED DESCRIPTION

[0046] The specific implementation modes of the present invention are described below in conjunction with the accompanying drawings and examples. However, the following examples are only used to illustrate the present invention in detail and do not limit the scope of the present invention in any way.

[0047] Example 1: A multi-dimensional feature collaborative fusion network, see Figure 1 ,It consists of four stages, namely multi-dimensional feature extraction, cross-modal feature fusion, CLIP-based consistency learning and multi-feature fusion.,In the first two stages, they are further divided into image branch and text branch.

[0048] 1) Multi-dimensional feature extraction part: Image feature extraction is divided into two parts: frequency-domain features and spatial-domain features. First, the Error Level Analysis (ELA) algorithm is used to generate the frequency-domain view of the image. The ELA algorithm can effectively detect the tampered areas of the image by analyzing the image compression traces. Then, the generated frequency-domain view is input into the pre-trained ResNet50 model to extract the frequency-domain tampering features. After that, the original image is input into the pre-trained Masked Autoencoder (MAE) model to extract the spatial-domain semantic features. The MAE model can capture the semantic content of the image due to its powerful global information extraction ability. Text feature extraction is divided into two parts: semantic features and sentiment features. First, the pre-processed text is input into the pre-trained BERT model to extract the text semantic features. The BERT model can generate high-quality text representations due to its bidirectional encoding mechanism and deep understanding of context semantics. After that, for text sentiment feature extraction, a sentiment knowledge graph is constructed, where the nodes include the words in the original text and the sentiment words in SenticNet, and the edge weights are calculated by Positive Pointwise Mutual Information (PPMI). Then, the Graph Convolutional Network (GCN) is used to update the node features to generate the sentiment knowledge embedding.

[0049] 2) Cross-modal feature fusion part: The dynamic interaction and fusion of image and text features are realized through a two-layer collaborative attention mechanism. With the two-layer collaborative attention mechanism, first, the left visual stream uses its own query Q I and the key K T and value V T in the right language stream for attention calculation. Then, the image-text fusion features are deeply fused with the text semantic features again to strengthen the text information, so as to obtain the image-text semantic fusion features with enhanced text information processed by the cross-modal collaborative attention fusion module. Similarly, the image fusion features with enhanced spatial-domain information are obtained. The text fusion features with enhanced text sentiment knowledge.

[0050] 3) CLIP-based consistency learning part: Based on the CLIP model, the semantic features of images and texts are extracted, and the image-text semantic alignment is optimized through the cosine embedding loss. Specifically, positive text-image pairs and negative text-image pairs are randomly sampled from the original dataset to generate a synthetic dataset. The pre-trained CLIP model is used to extract text features and image features, and a shared semantic embedding is generated through a Multi-Layer Perceptron (MLP). Finally, the cosine embedding loss is calculated to optimize the image-text semantic alignment. In addition, the image-text similarity score is calculated to achieve similarity-aware image-text fusion.

[0051] 4) Multi - feature fusion part: After concatenating the frequency - domain features, spatial - domain features, semantic features, sentiment features, and similarity - aware image - text fusion features, the contributions of each feature are dynamically adjusted through an attention weight matrix. Specifically, first, the attention weight matrix is calculated, and then it is multiplied by the concatenated multi - dimensional features to generate the final fusion features. These final features are used for single - modality and multi - modality fake information detection.

[0052] The present invention uses the Weibo, Twitter, and GossipCop datasets for training and testing. The Twitter dataset was released for the MediaEval Verify Multimodal Usage task, and its posts contain text content, images / videos, and additional social context information. We focus on combining text and image information to detect fake information. Therefore, the present invention filters out tweets with attached videos and non - English tweets; the Weibo dataset is a publicly available dataset in the field of multi - modal fake news detection. In this dataset, real news posts are collected from official news sources in China (such as Xinhua News Agency), and fake news posts are verified by the Weibo official rumor - busting platform; the Gossipcop dataset is a multi - modal news dataset collected from the entertainment domain of the FakeNewsNet knowledge base.

[0053] Through Figure 5 comparative experiments, it can be seen that the present invention can achieve better recognition effects. Through Figure 6 ablation experiments, it can be seen the importance of each module in the present invention.

[0054] The above has described the present invention in detail with reference to the accompanying drawings and embodiments. However, those skilled in the art can understand that without departing from the gist of the present invention, various specific parameters in the above - mentioned embodiments can be changed to form multiple specific embodiments, which are all within the common variation range of the present invention and will not be elaborated one by one here.

Claims

1. A false information detection method based on a multi-dimensional feature collaborative fusion network, characterized in that It includes the following modules: S1: Multi-dimensional feature extraction module: used to extract the frequency-domain feature h of the image, the spatial-domain feature h ME , the semantic feature h of the text I , and the sentiment feature h T from the input data; SA ​ S2: Cross-modal Collaborative Attention Fusion Module (CCF): Through a dynamic and context-aware attention mechanism, it realizes the interactive fusion of image and text features to generate cross-modal enhanced features; S3: Cross-modal Consistency Learning Module Based on CLIP: It uses the pre-trained CLIP model to extract image-text semantic features and optimizes cross-modal semantic alignment through cosine embedding loss; S4: Multi-feature Weighted Fusion Module: Based on the attention mechanism, it dynamically adjusts the weights of each feature to generate the final joint representation.

2. The false information detection method based on the multi-dimensional feature collaborative fusion network according to claim 1, wherein The multi-dimensional feature extraction module includes: S1: Image Frequency Domain Feature Extraction Sub-module: It uses the Error Level Analysis (ELA) algorithm to extract image frequency domain tampering features and performs deep feature encoding through the pre-trained ResNet model; S2: Image Spatial Domain Feature Extraction Sub-module: It extracts the global semantic features of the image based on the Masked Autoencoder (MAE); S3: Text Semantic Feature Extraction Sub-module: It extracts the context semantic features of the text through the pre-trained BERT model; S4: Text Sentiment Feature Extraction Sub-module: It constructs a text sentiment association graph based on the Sentiment Knowledge Graph (SenticNet) and generates sentiment embeddings through the Graph Convolutional Network (GCN).

3. The false information detection method based on the multi-dimensional feature collaborative fusion network according to claim 1, characterized in that The Cross-modal Collaborative Attention Fusion Module (CCF) realizes feature interaction through the following steps: S1: Adopt a double-layer collaborative attention mechanism. First, the left visual stream uses its own query Q I and the key K in the right language stream T and the value V T to perform attention calculation. Then, the graph-text fusion feature is deeply fused with the text semantic feature again to strengthen the text information, so as to obtain the graph-text semantic fusion feature that strengthens the text information processed by the cross-modal collaborative attention fusion module S2: Similarly, obtain the image fusion features that enhance the spatial domain information Text fusion features that enhance text emotion knowledge 4. The false information detection method based on the multi-dimensional feature collaborative fusion network according to claim 1, wherein, The cross-modal consistency learning module based on CLIP includes: S1: Set up a cross-modal consistency learning auxiliary task, and maximize the semantic similarity of positive sample pairs and suppress negative sample pairs through the cosine embedding loss with a margin of d; S2: Concatenate the text features h extracted by the BERT and MAE encoders T and the image features h I , and input them into a linear layer with a sigmoid function. Then, use a multiplicative gate to weight the cross-modal semantic feature fusion to obtain cross-modal fusion features 5. The false information detection method based on a multi-dimensional feature collaborative fusion network according to claim 1, characterized in that, The multi-feature weighted fusion module generates the joint representation through the following steps: S1: Concatenate the image frequency-domain feature h ME , the spatial-domain feature h I , the text semantic feature h T , the text sentiment feature h SA and the cross-modal fusion feature into a multi-modal representation h; S2: Dynamically adjust the contribution weights of each feature channel through the learnable attention weight matrix W Att to obtain the output joint representation h Fusion .

6. A false information detection device based on a multi-dimensional feature collaborative fusion network, characterized in that The device includes: a memory that stores one or more computer-executable instructions, and the computer-executable instructions implement each step of the method for detecting false information based on the multi-dimensional feature collaborative fusion network described in any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed to implement each step of the method for detecting false information based on the multi-dimensional feature collaborative fusion network described in any one of claims 1 to 5.

Citation Information

Cited By

  • Mixed reality intelligent inspection method and system based on multi-modal fusion

    CN120579148A