Data-enhanced graph learning multi-modal false information detection method and system
Through the data-enhanced graph learning method, combined with OCR technology, positive and negative attention mechanism and coordinated attention mechanism, the problem of improper user interaction and embedded word processing in the existing false information detection technology is solved, and more accurate false information detection is achieved.
Patent Information
- Application Number
- CN202510179172.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-24
AI Technical Summary
Existing false information detection technologies are difficult to effectively utilize user interaction and target information internal knowledge during information dissemination, and improper processing of image samples embedded in text will affect feature learning, ignoring the importance of negative comments.
The data-enhanced graph learning method is adopted to extract embedded text through OCR technology to achieve text completion and image mask; improve the potential connection of graph structure nodes based on similarity calculation; use positive and negative attention mechanism to learn the views of approval and objection in comments; and use the synergistic attention mechanism to achieve the fusion of text, image and graph structure modalities.
It effectively improves the accuracy of false information detection, makes full use of user interaction and internal information knowledge, improves the processing of embedded text and negative comments, and improves the accuracy of detection results.
Smart Images

Figure CN120198922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer and network information technology, and in particular to a data-enhanced graph learning multimodal false information detection method and system. Background Art
[0002] With the rapid development of the Internet and self-media technology, uncensored information can be easily misinterpreted and turned into false information during the dissemination process. These false information have a huge impact in many fields due to their diverse formats, wide dissemination methods, and fuzzy judgment boundaries. How to make full use of user interaction during the dissemination process of information and the internal knowledge of the target information to achieve detection is the key.
[0003] In the context of the digital age, social networking platforms have profoundly affected the daily lives of the public due to their real-time and convenience. They have not only changed the way people obtain and exchange information, but also greatly affected the dissemination and reception patterns of information. However, the review mechanism of social platforms often only targets political and inappropriate content, while ignoring the review of false information, which leads to the spread of false information. Studies have shown that information is easily misinterpreted and deformed during the dissemination process, resulting in a variety of false information versions, which spread rapidly in many areas of public concern such as politics, economy, health, technology, and entertainment. Its dissemination will lead to misleading public cognition, affecting the information security and property security of the people, and causing serious impacts on the credibility of mainstream media. Therefore, false information detection technology has emerged. How to fully learn and analyze the potential knowledge within and between information to judge the authenticity of information, and to replace manual verification to a certain extent to reduce the cost of manpower and financial resources, is a key issue that needs to be solved urgently.
[0004] After a long period of development, false information detection technology initially focused on exploring the internal knowledge of information content such as text or images, and gradually evolved into exploring the potential connections between information. For example, SpotFake performs feature transformation on text and images respectively, splices the representations of the two modalities, and uses a fully connected neural network to discriminate false information; constructs a dynamic news environment, and perceives potential environmental factors based on the popularity analysis module and novelty analysis. The graph structure can intuitively display the connections and interactions between information through nodes and edges. Methods based on graph learning have shown significant advantages and great potential in the field of true and false information recognition. By modeling the complex relationships in the information dissemination process in the form of a graph structure, it can not only effectively capture the potential associations between information, but also make full use of the rich knowledge contained in interactive data such as user comments. GET represents each word in the article as a node, fixes the sliding window size to establish the connections (edges) between words, embeds the graph through GCN, and inputs it into a fully connected layer with an attention mechanism for information detection. Propagation-based methods focus on the dissemination process of posts / articles. Throughout the process, users interact by posting, forwarding, or replying to a given article. These users and their interactions form a tree (or graph) structure. By examining the dissemination structure of news within the dissemination network and the credibility of users, the potential authenticity of a given news article can be inferred. UPFD focuses on forwarding interactions and user attributes, and constructs a graph composed of the source news article and the users who forward the article, where the user node features are extracted from their historical posts. The final graph model captures information related to user credibility from the user's historical activities through joint content and graph modeling. A key difference between propagation-based methods and methods based on heterogeneous social contexts is the process of applying graph modeling. Propagation-based methods tend to use a graph to model the dissemination process of a single article, while methods based on heterogeneous social contexts tend to use a graph to model multiple articles and users involved in a larger social context interacting with different or related articles.
[0005] While these studies have made progress technically, they have also revealed several key issues. First, the detection strongly depends on data, and a carefully designed dataset is required to achieve high-accuracy classification. Statistical analysis of data on platforms such as Weibo and Xiaohongshu shows that there are embedded texts in nearly 40% of the sample images. Treating them as part of the image will instead affect the learning of image features and result in the loss of the original content, leading to difficulties in actual implementation despite advanced technology. Additionally, positive and negative comments are equally important for information authenticity detection. Only using the traditional attention mechanism to achieve feature learning of the graph structure will greatly ignore the importance of negative comments. Therefore, feature learning needs to be carefully designed to focus on actually helpful content. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0007] The present invention proposes a data augmentation graph learning multi-modal fake information detection method, which focuses on posts containing image and text modalities and user interactions such as related comments on current social platforms. For image samples with embedded text, OCR technology is used to extract text to achieve text completion and image masking. Based on similarity calculation, potential connections of graph structure nodes are improved. Positive and negative attention mechanisms are adopted to fully learn the approving and opposing views in comments. Based on the co-attention mechanism, modality fusion and information authenticity detection are realized.
[0008] Another object of the present invention is to propose a data augmentation graph learning multi-modal fake information detection system.
[0009] To achieve the above object, on the one hand, the present invention proposes a data augmentation graph learning multi-modal fake information detection method, including:
[0010] Construct graph structure data based on multi-source data and generate a multi-modal data set; wherein each sample in the multi-modal data set includes data of three modalities: text, image, and graph structure.
[0011] Splice the text embedded in the image of the extracted multi-modal data set into the text content, calculate the average color value of the surrounding area of the text to generate an image mask, and complete the potential edges between graph structure nodes based on similarity calculation to obtain an enhanced data set.
[0012] Based on the enhanced data set, use the BERT model to extract the feature representation of the text, use the ResNet50 model to extract the visual features of the image, and extract the relationship features between the nodes of the graph structure data through the GAT with positive and negative attention strategies to respectively output the feature vectors corresponding to the three modalities.
[0013] Use the cross-attention strategy to perform pairwise feature fusion on the feature vectors of the three modalities, and input the fused feature vectors into a classifier to output the information authenticity detection result.
[0014] The data augmentation graph learning multi-modal fake information detection method according to the embodiments of the present invention may also have the following additional technical features:
[0015] In an embodiment of the present invention, splicing images with a quantity exceeding a preset number in each sample includes:
[0016] Define the number of images as n;
[0017] Calculate the square root of n and round down to get
[0018] Define the number of rows as
[0019] Stitch the images. The first s*s images form a complete square. If n > s 2 , the remaining images are arranged in the form of row rows and col columns. For the last row where it is not completely filled, the remaining places are set to all 0s;
[0020] Compress the stitched images.
[0021] In an embodiment of the present invention, the text embedded in the images of the extracted multi-modal data set is stitched into the text content, and the color mean value of the peripheral area of the text is calculated to generate an image mask, and the potential edges between the nodes of the complementary graph structure are calculated based on similarity to obtain an enhanced data set, including:
[0022] Use OCR technology to scan the images of all samples, mark the samples with embedded text images, and perform denoising, binarization processing and rotation correction operations to identify the embedded text and stitch it after the text corresponding to the sample;
[0023] Perform a masking operation on the text part in the image. Define the text area as represented by the coordinates (x1, y1, x2, y2), and the extended area as represented by (x′1, y′1, x′2, y′2), where x′1 = max(0, x1 - 20), y′1 = max(0, y1 - 20), x′2 = min(x max , x2 + 20), y′2 = min(y max , y2 + 20);
[0024] Use the color mean value of the annular area between the extended area and the text area as the filling of the text area, which is represented by the following formula:
[0025]
[0026] where R represents the annular area, and I(x, y) represents the color value of the coordinate point;
[0027] For the graph feature embedding of the target information p i , define the graph network to include target information, users, and comments; define the node embedding matrix X ∈ R |V|*d , where |V| is the number of nodes and d is the embedding dimension; calculate the similarity between nodes. If the similarity between two nodes is greater than the threshold β ij , it means that there is a potential edge among them, and add the potential edge to the original adjacency matrix A:
[0028]
[0029] In one embodiment of the present invention, based on the enhanced dataset, the BERT model is used to extract the feature representation of the text, the ResNet50 model is used to extract the visual features of the image, and the Graph Attention Network (GAT) with positive and negative attention strategies is used to extract the relationship features between the nodes of the graph structure data, so as to output the feature vectors corresponding to the three modalities respectively, including:
[0030] For the text content, BERT is used as the embedding model to obtain the text feature R t , and it is mapped to 900 dimensions through the fully connected layer W t to output E t :
[0031] E t = σ(W t *R t + b)
[0032] where W t is the weight matrix of the text fully connected layer, σ(·) is the sigmoid activation function, b is the bias vector, and E t ∈ R n*d , where n is the number of samples and d is the vector dimension;
[0033] For the image content, ResNet50 is used as the embedding model to obtain the image feature R i , and it is mapped to 900 dimensions through the fully connected layer W i to output E i :
[0034] E i = σ(W i *R i + b)
[0035] where W i is the weight matrix of the image fully connected layer;
[0036] The Cross-Attention (CA) layer is used to align the features of the two modalities to achieve modality fusion; among them, the CA layer contains two parallel CA blocks;
[0037] For node i, the unnormalized attention weights ε i between its neighbors (j ∈ N ij ) and itself are calculated one by one:
[0038]
[0039] where and W are the parameters to be learned, e g represents the initial embedding of the node, [·||·] represents the concatenation of the transformed features of nodes i and j, and LeakyReLU is the activation function;
[0040] Calculate the positive attention weight ε while performing normalization pos and the negative attention weight ε nag :
[0041]
[0042] where ε i ={ε ij |j∈N i}, representing the unnormalized attention weights between node i and all its neighbor nodes j;
[0043] After performing a weighted sum on the positive and negative attention weights and concatenating with the original node, a new feature of the node fused with neighborhood information is obtained:
[0044]
[0045] where W is the weight matrix of the fully connected layer, is the initial feature vector of neighbor node j.
[0046] In an embodiment of the present invention, a cross-attention strategy is used to perform pairwise feature fusion on the feature vectors of three modalities, and the fused feature vectors are input into a classifier to output the information authenticity detection result, including:
[0047]
[0048]
[0049] E = concat(E ti , E tg , E gi )
[0050] The loss is divided into two parts, namely the modality alignment loss and the classification loss; the feature vectors of the three modalities are mapped to the same feature space, and the MSE loss is used for modality alignment:
[0051]
[0052] where W is the mapping matrix to be learned;
[0053] The feature vector E is input into the fully connected layer for prediction, and finally the prediction result is obtained and the cross-entropy loss is used as the classification loss:
[0054]
[0055] Loss = λ c L class + λa L align
[0056] where σ() is the softmax activation function; the final loss function is jointly composed of the alignment loss and the classification loss.
[0057] To achieve the above object, on the other hand, the present invention proposes a data-augmented graph learning multi-modal false information detection system, including:
[0058] A multi-modal dataset construction module, configured to construct graph-structured data based on multivariate data and generate a multi-modal dataset; wherein each sample in the multi-modal dataset includes data of three modalities: text, image, and graph structure.
[0059] A dataset augmentation module, configured to splice the text embedded in the image of the extracted multi-modal dataset into the text content, calculate the average color value of the surrounding area of the text to generate an image mask, and complete the potential edges between the graph structure nodes based on similarity calculation to obtain an augmented dataset.
[0060] A multi-modal feature learning module, configured to extract the feature representation of the text using the BERT model, extract the visual features of the image using the ResNet50 model, and extract the relationship features between the nodes of the graph-structured data through the GAT with positive and negative attention strategies, so as to output the feature vectors corresponding to the three modalities respectively.
[0061] A feature fusion and information detection module, configured to perform pairwise feature fusion on the feature vectors of the three modalities using the cross-attention strategy, and input the fused feature vectors into a classifier to output the detection result of the authenticity of the information.
[0062] The data-augmented graph learning multi-modal false information detection method and system according to the embodiments of the present invention not only perform data augmentation on different modalities of information respectively, but also fully weigh the influence of positive and negative comments based on the positive and negative attention mechanism, and improve the accuracy of the detection result based on multi-modal feature learning. By scanning the image samples with embedded text using OCR technology, grabbing the text content for text completion and image masking, improving the missing graph structure using similarity calculation, balancing the influence weights of the node neighbors in the graph based on the positive and negative attention mechanism, and combining the collaborative attention mechanism to realize the fusion of the three modalities of text, image, and graph structure, finally realizing the detection of the authenticity of the information.
[0063] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of embodiments in conjunction with the accompanying drawings, wherein:
[0065] Figure 1 is a flowchart of a data augmentation-based graph learning multi-modal misinformation detection method according to an embodiment of the present invention;
[0066] Figure 2 is an architecture diagram of a data augmentation-based graph learning multi-modal misinformation detection model according to an embodiment of the present invention;
[0067] Figure 3 is a schematic diagram of text and image augmentation according to an embodiment of the present invention;
[0068] Figure 4 is a structural diagram of a data augmentation-based graph learning multi-modal misinformation detection system according to an embodiment of the present invention. Detailed Embodiments
[0069] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0070] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0071] The data augmentation-based graph learning multi-modal misinformation detection method and system according to an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0072] Figure 1 is a flowchart of a data augmentation-based graph learning multi-modal misinformation detection method according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0073] S1, constructing graph-structured data based on multivariate data and generating a multi-modal data set; wherein, each sample in the multi-modal data set includes data of three modalities: text, image, and graph structure;
[0074] S2, splicing the text embedded in the image of the extracted multi-modal data set into the text content, calculating the average color value of the surrounding area of the text to generate an image mask, and complementing the potential edges between the graph structure nodes based on similarity calculation to obtain an enhanced data set;
[0075] S3. Based on the enhanced dataset, use the BERT model to extract the feature representation of the text, use the ResNet50 model to extract the visual features of the image, and use the GAT with positive and negative attention strategies to extract the relationship features between the nodes of the graph structure data, so as to output the feature vectors corresponding to the three modalities respectively;
[0076] S4. Use the cross-attention strategy to perform pairwise feature fusion on the feature vectors of the three modalities, and input the fused feature vectors into the classifier to output the information authenticity detection result.
[0077] The present invention proposes a graph learning multi-modal false information detection method with data augmentation, which respectively performs data augmentation on the text modality, image modality, and graph structure modality constructed by user interaction of the sample, adopts a positive and negative attention mechanism to mine two valuable pieces of information, approval and opposition, in the comments, and realizes authenticity detection by fusing the three modalities of text, image, and graph structure based on the cross-attention mechanism. As Figure 2 shown, it is the complete architecture diagram of the detection model proposed by the present invention.
[0078] In an embodiment of the present invention, information on the Weibo platform, including metadata such as text, images, comments, and user information, is collected based on the automated collection module, and the comment content is stored in a tree structure to ensure that each sample contains data of the three modalities.
[0079] Specifically, data is crawled based on the Weibo platform, and the dataset includes three modalities: text, images, and comments. The dataset is divided into JSON files and corresponding image files, where JSON is composed of the fields in Table 1, and the images are named with image_id.
[0080] Table 1
[0081]
[0082]
[0083] The specific process includes: crawling Weibo platform posts, and performing preprocessing such as deleting samples with missing data and splicing the image modality to form a dataset, ensuring that each sample in the dataset includes information such as text, images, and comments. Define each sample as consisting of a long text t, an image i, and a comment set g. For samples with more than 1 image, the images need to be spliced, and the specific logic is as follows:
[0084] Define the number of images as n;
[0085] Calculate the square root of n and round down to get
[0086] Define the number of rows as
[0087] Stitch the pictures. The first s*s pictures form a complete square. If n>s 2 , the remaining pictures are arranged in the form of row rows and col columns, where the last row may not be completely filled, and the remaining places are set to all 0s;
[0088] Compress the stitched large picture and output an image of 900*900.
[0089] In an embodiment of the present invention, scan the image samples with embedded text based on OCR technology, extract and stitch them after the text content; perform image masking based on the average color value of the surrounding area of the text to prevent the text from interfering with the image feature learning; complete the potential edges between the graph structure nodes based on similarity calculation.
[0090] It can be understood that data augmentation is performed on the three modalities of text, image, and graph structure of the samples respectively to achieve text completion, image masking, and graph structure potential edge inference, as Figure 3 shown.
[0091] Specifically, scan the images of all samples based on OCR technology, mark the samples with embedded text images, and identify and stitch the embedded text after the text corresponding to the sample through preprocessing operations such as denoising, binarization, and rotation correction. Perform a masking operation on the text part in the image. Define the text area as represented by the coordinates (x1, y1, x2, y2), and the extended area as represented by (x′1, y′1, x′2, y′2), where x′1 = max(0, x1 - 20), y′1 = max(0, y1 - 20), x′2 = min(x max , x2 + 20), y′2 = min(y max , y2 + 20) to ensure the validity of the coordinates of the extended area. To prevent the embedded text from introducing noise during the image embedding process, use the average color value of the annular area between the extended area and the text area as the filling of the text area, as shown in the following formula, where R represents the annular area and I(x, y) represents the color value of the coordinate point (represented by three values in the RGB space):
[0092]
[0093]
[0094] For the graph feature embedding of the target information p i , define that the graph network consists of target information, users, and comments. For the target information and comment nodes, their initial features are text embeddings, and for the user, the initial feature is the average of the published articles and comment embeds. Define the node embedding matrix X ∈ R|V|*d , where |V| is the number of nodes and d is the embedding dimension. Calculate the similarity between nodes. If the similarity between two nodes is greater than the threshold β ij , it means that there is a potential edge among them. Add the potential edge to the original adjacency matrix A as shown in the following formula:
[0095]
[0096] In one embodiment of the present invention, feature learning of text and images is implemented based on BERT and ResNet50. For the graph structure, the graph features are learned by GAT with positive and negative attention mechanisms to fully learn the influence of comments from different viewpoints on information.
[0097] It can be understood that feature learning is performed on the text, text and graph modalities of the samples respectively, and then feature fusion is achieved based on cross-attention.
[0098] Specifically, for the text content, BERT is used as the embedding model to obtain the text feature R t , and it is mapped to 900 dimensions through the fully connected layer W t to output E t , as shown in the formula, where W t is the weight matrix of the text fully connected layer, σ(·) is the sigmoid activation function, b is the bias vector, and E t ∈R n*d , where n is the number of samples and d is the vector dimension.
[0099] E t =σ(W t *R t +b)
[0100] For the image content, ResNet50 is used as the embedding model to obtain the image feature R i , and it is mapped to 900 dimensions through the fully connected layer W i to output E i , as shown in the following formula, where W i is the weight matrix of the image fully connected layer.
[0101] E i =σ(W i *R i +b)
[0102] Multi-modal feature fusion is achieved by the CA layer (co-attention layer), as Figure 3As shown, the CA layer contains two parallel CA blocks (co-attention blocks). The difference between co-attention and traditional self-attention is that the input Q (query) of the multi-head Attention comes from one modality, and the K (key) and V (value) come from another modality. In this way, the features of the two modalities are aligned to achieve modality fusion.
[0103] In the current graph structure, each node only has initial text features and has not effectively integrated information from neighboring nodes. Therefore, a graph attention network (GAT) is used to learn the features of graph nodes, hoping that the nodes can truly integrate the information of their neighboring nodes to achieve a more accurate feature representation. For node i, the unnormalized attention weights ε i between its neighbors (j ∈ N ij ) and itself are calculated one by one, as shown in the following formula, where and W are parameters to be learned, and e g represents the initial embed of the node, [·||·] concatenates the transformed features of nodes i and j, and LeakyReLU is the activation function.
[0104]
[0105] The normalization operation of traditional GAT is based on the softmax function, which assigns higher weights to similar nodes and lower weights to irrelevant nodes, thus ignoring the content with opposing opinions in the comments. To fully consider both positive and negative correlations, the positive attention weight ε pos and the negative attention weight ε nag are calculated simultaneously during normalization, as shown in the following formula, where ε i = {ε ij | j ∈ N i}, representing the unnormalized attention weights between node i and all its neighbor nodes j.
[0106]
[0107] After weighting and summing the positive and negative attention weights and concatenating them with the original node, a new feature of the node that fuses neighborhood information is formed, as shown in the following formula, where W is the weight matrix of the fully connected layer, is the initial feature vector of neighbor node j.
[0108]
[0109] In one embodiment of the present invention, based on the feature vectors output from the above steps, cross-attention is used to learn pairwise among the three modalities to achieve the fusion of the three modalities, and the final vector is input into a classifier to achieve the detection of the authenticity of information. Modal alignment is regarded as a regression task, and authenticity detection is regarded as a classification task. The losses of the two tasks are fully considered for optimization to realize the design of the overall model.
[0110] Specifically, based on the cross-attention mechanism, pairwise fusion of the three modalities of text, image, and graph structure is realized, and the final vector is input into a classifier for true / false detection, as shown in the following formula:
[0111]
[0112]
[0113] E = concat(E ti , E tg , E gi )
[0114] The loss is divided into two parts, namely the modal alignment loss and the classification loss. The three modal features are mapped to the same feature space, and the MSE loss is used for modal alignment, as shown in the following formula, where W is the mapping matrix to be learned.
[0115]
[0116] The feature vector E is input into a fully connected layer for prediction, and finally the prediction result is obtained The cross-entropy loss is used as the classification loss, which is defined as the following formula, where σ() is the softmax activation function, and the final loss function is composed of the alignment loss and the classification loss together.
[0117]
[0118] Loss = λ c L class + λ a L align
[0119] Furthermore, the present invention selects the SAFE and MFAN models used by research institutes in the same field for performance and effectiveness comparison. The results of the binary classification experiment show that the new model is significantly better than the existing models in predicting the authenticity of information.
[0120] According to the data augmentation graph learning multimodal false information detection method of the embodiments of the present invention, it focuses on posts containing image and text modalities and user interactions such as related comments on current social platforms. For image samples with embedded text, the OCR technology is used to extract text to achieve text completion and image masking. Based on similarity calculation, the potential connections of the graph structure nodes are improved. The positive and negative attention mechanisms are used to fully learn the approving and opposing views in the comments. Based on the co-attention mechanism, modality fusion and information authenticity detection are realized, effectively improving the accuracy of the detection results.
[0121] To implement the above embodiments, as Figure 4 shown, in this embodiment, a data augmentation graph learning multimodal false information detection system 10 is further provided, including:
[0122] A multimodal dataset construction module 100, configured to construct graph structure data based on multivariate data and generate a multimodal dataset; wherein, each sample in the multimodal dataset includes data of three modalities: text, image, and graph structure;
[0123] A dataset augmentation module 200, configured to splice the text embedded in the images of the extracted multimodal dataset into the text content, calculate the average color value of the surrounding area of the text to generate an image mask, and complete the potential edges between the graph structure nodes based on similarity calculation to obtain an augmented dataset;
[0124] A multimodal feature learning module 300, configured to, based on the augmented dataset, use the BERT model to extract the feature representation of the text, use the ResNet50 model to extract the visual features of the image, and extract the relationship features between the nodes of the graph structure data through the GAT with positive and negative attention strategies, so as to output the feature vectors corresponding to the three modalities respectively;
[0125] A feature fusion and information detection module 400, configured to use the cross-attention strategy to perform pairwise feature fusion on the feature vectors of the three modalities, and input the fused feature vectors into a classifier to output the information authenticity detection result.
[0126] Further, for each sample, splicing images with a quantity exceeding a preset number includes:
[0127] Define the number of images as n;
[0128] Calculate the square root of n and round down to get
[0129] Define the number of rows as
[0130] Splice the pictures. The first s*s pictures form a complete square. If n>s 2, the remaining images are arranged in the form of row * row and column * column. For the last row where there is an incomplete filling situation, the remaining places are set to all 0;
[0131] Compress the spliced image.
[0132] Furthermore, the dataset enhancement module 200 is also used for:
[0133] Use OCR technology to scan the images of all samples, mark the samples with embedded text images, and perform denoising, binarization processing and rotation correction operations to identify the embedded text and splice it to the corresponding text of the sample;
[0134] Perform a masking operation on the text part in the image. Define the text area as represented by the coordinates (x1, y1, x2, y2), and the extended area as represented by (x′1, y′1, x′2, y′2), where x′1 = max(0, x1 - 20), y′1 = max(0, y1 - 20), x′2 = min(x max , x2 + 20), y′2 = min(y max , y2 + 20);
[0135] Based on the color mean value of the annular area between the extended area and the text area as the filling of the text area, which is represented by the following formula:
[0136]
[0137] Where R represents the annular area, and I(x, y) represents the color value of the coordinate point;
[0138] For the graph feature embedding of the target information p i , define the graph network to include target information, users and comments; define the node embedding matrix X ∈ R |V|*d , where |V| is the number of nodes and d is the embedding dimension; calculate the similarity between nodes. If the similarity between two nodes is greater than the threshold β ij , it means that there is a potential edge among them, and add the potential edge to the original adjacency matrix A:
[0139]
[0140] Furthermore, the multimodal feature learning module 300 is also used for:
[0141] For the text content, use BERT as the embedding model to obtain the text feature R t , and map it to 900 dimensions through the fully connected layer W t , and output E t :
[0142] E t = σ(Wt *R t +b)
[0143] where W t is the weight matrix of the text fully-connected layer, σ(·) is the sigmoid activation function, b is the bias vector, and e t ∈R n*d , where n is the number of samples and d is the vector dimension;
[0144] For the image content, ResNet50 is used as the embedding model to obtain the image feature R i , and it is mapped to 900 dimensions through the fully-connected layer W i to output E i :
[0145] E i = σ(W i *R i +b)
[0146] where W i is the weight matrix of the image fully-connected layer;
[0147] The CA layer is used to align the two-modal features to achieve modal fusion; among them, the CA layer contains two parallel CA blocks;
[0148] For node i, the unnormalized attention weights ε i of its neighbors (j ∈ N ij ) and itself are calculated one by one:
[0149]
[0150] where and W are parameters to be learned, e g represents the initial embed of the node, [·||·] represents concatenating the transformed features of nodes i and j, and LeakyReLU is the activation function;
[0151] When performing normalization, the positive attention weight ε pos and the negative attention weight ε nag are calculated simultaneously:
[0152]
[0153] where ε i = {ε ij |j ∈ N i}, representing the unnormalized attention weights of node i and all its neighbor nodes j;
[0154] After performing a weighted sum on the positive and negative attention weights and concatenating them with the original nodes, a new feature of the node integrating neighborhood information is obtained:
[0155]
[0156] where W is the weight matrix of the fully connected layer, is the initial feature vector of neighbor node j.
[0157] Furthermore, the feature fusion and information detection module 400 is further configured to:
[0158]
[0159] E = concat(E ti , E tg , E gi )
[0160] The loss is divided into two parts, namely the modality alignment loss and the classification loss; the feature vectors of the three modalities are mapped to the same feature space, and the MSE loss is used for modality alignment:
[0161]
[0162] where W is the mapping matrix to be learned;
[0163] The feature vector E is input into the fully connected layer for prediction, and finally the prediction result is obtained and the cross-entropy loss is used as the classification loss:
[0164]
[0165] Loss = λ c L class + λ a L align
[0166] where σ() is the softmax activation function; the final loss function is jointly composed of the alignment loss and the classification loss.
[0167] According to the data augmentation-based graph learning multi-modal false information detection system of the embodiments of the present invention, it focuses on posts containing image and text modalities on the current social platform and user interactions such as related comments. For image samples with embedded text, the OCR technology is used to extract text to achieve text completion and image masking. The potential connections of the nodes in the graph structure are improved based on similarity calculation. The positive and negative attention mechanisms are used to fully learn the approving and opposing views in the comments. The modality fusion and information authenticity detection are realized based on the co-attention mechanism, effectively improving the accuracy of the detection results.
[0168] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0169] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
Claims
1. A data-enhanced graph learning multimodal false information detection method, characterized in that: include: Constructing graph structure data based on multivariate data and generating a multimodal data set; wherein each sample in the multimodal data set includes data in three modes: text, image and graph structure; The text embedded in the extracted image of the multimodal dataset is spliced into the text content, and the color mean of the area around the text is calculated to generate an image mask, and the potential edges between the nodes of the completion graph structure are calculated based on the similarity to obtain an enhanced dataset; Based on the enhanced dataset, the BERT model is used to extract the feature representation of text, the ResNet50 model is used to extract the visual features of the image, and the GAT with positive and negative attention strategies is used to extract the relationship features between the nodes of the graph structure data to output the feature vectors corresponding to the three modalities respectively; The cross-attention strategy is used to fuse the feature vectors of the three modalities in pairs, and the fused feature vectors are input into the classifier to output the information authenticity detection result.
2. The method according to claim 1, characterized in that The images exceeding the preset number in each sample are stitched together, including: Define the number of images as n; Calculate the square root of n and round down to get Define the number of rows as The pictures are stitched together, and the first s*s pictures form a complete square. If n>s 2 , the remaining images are arranged in row and col columns. If the last row is not completely filled, the remaining space is set to all 0s; Compress the spliced images.
3. The method according to claim 1, characterized in that The text embedded in the extracted image of the multimodal dataset is spliced into the text content, and the color mean of the area around the text is calculated to generate an image mask, and the potential edges between the nodes of the completion graph structure are calculated based on the similarity to obtain an enhanced dataset, including: Use OCR technology to scan the images of all samples, mark the samples with embedded text images, and perform denoising, binarization and rotation correction operations to identify the embedded text and splice it into the corresponding text of the sample; For the text part of the image, the mask operation is performed. The text area is defined as (x1, y1, x2, y2) coordinates, and the extended area is represented by (x′1, y′1, x′2, y′2), where x′1 = max(0, x1-20), y′1 = max(0, y1-20), x′2 = min(x max ,x2+20),y′2=min(y max ,y2+20); The color average of the annular area between the extended area and the text area is used as the filling of the text area, which is expressed by the following formula: Where R represents the annular area, and I(x,y) represents the color value of the coordinate point; For the target information p i Graph feature embedding, define the graph network including target information, users and comments; define the node embedding matrix X∈R |V|*d , where |V| is the number of nodes and d is the embedding dimension; calculate the similarity between nodes, if the similarity between two nodes is greater than the threshold β ij , it means there is a potential edge, add the potential edge to the original adjacency matrix A:
4. The method according to claim 1, characterized in that: Based on the enhanced dataset, the BERT model is used to extract the feature representation of text, the ResNet50 model is used to extract the visual features of the image, and the GAT with positive and negative attention strategies is used to extract the relationship features between the nodes of the graph structure data to output the feature vectors corresponding to the three modalities, including: For text content, BERT is used as an embedding model to obtain text features R t , and through the fully connected layer W t Map to 900 dimensions, output E t : E t σ(W t *R t +b) Where W t is the weight matrix of the text fully connected layer, σ(·) is the sigmoid activation function, b is the bias vector, and E t ∈R n*d , where n is the number of samples and d is the vector dimension; For image content, ResNet50 is used as the embedding model to obtain image features R i , and through the fully connected layer W i Map to 900 dimensions, output E i : E i σ(W i *R i +b) Where W i is the weight matrix of the fully connected layer of the image; The CA layer is used to align the features of the two modalities to achieve modal fusion. The CA layer contains two parallel CA blocks. For node i, calculate the neighbors (j∈N i ) and its own unnormalized attention weight ε ij : in and W are the parameters to be learned, e g represents the initial embedding of the node, [·||·] represents the concatenation of the transformed features of nodes i and j, and LeakyReLU is the activation function; When normalizing, the positive attention weight ε is calculated at the same time pos and negative attention weight ε nag : where ε i ={ε ij |j∈N i }, represents the unnormalized attention weights of node i and all its neighbor nodes j; The positive and negative attention weights are weighted and concatenated with the original nodes to obtain new features that incorporate neighborhood information nodes: Where W is the weight matrix of the fully connected layer, is the initial feature vector of neighbor node j.
5. The method according to claim 1, characterized in that The cross-attention strategy is used to fuse the feature vectors of the three modalities in pairs, and the fused feature vectors are input into the classifier to output the information authenticity detection results, including: E=concat(E ti ,AND tg ,AND gi ) The loss is divided into two parts, namely modality alignment loss and classification loss. The feature vectors of the three modalities are mapped to the same feature space, and the MSE loss is used for modality alignment: Where W is the mapping matrix to be learned; Input the feature vector E into the fully connected layer for prediction, and finally get the prediction result And use the cross entropy loss as the classification loss: Loss=λ c L class +λ a L align Where σ() is the softmax activation function; the final loss function is composed of alignment loss and classification loss.
6. A data-enhanced graph learning multimodal false information detection system, characterized in that: include: A multimodal data set construction module, used to construct graph structure data based on multivariate data and generate a multimodal data set; wherein each sample in the multimodal data set includes data in three modes: text, image and graph structure; A data set enhancement module, used to splice the text embedded in the extracted image of the multimodal data set into the text content, calculate the color mean of the area around the text to generate an image mask, and calculate the potential edges between the nodes of the completion graph structure based on similarity to obtain an enhanced data set; A multimodal feature learning module, for extracting feature representations of text using a BERT model based on the enhanced dataset, extracting visual features of images using a ResNet50 model, and extracting relational features between graph structure data nodes using a GAT with positive and negative attention strategies, so as to output feature vectors corresponding to the three modalities respectively; The feature fusion and information detection module is used to use the cross-attention strategy to fuse the feature vectors of the three modalities in pairs, and input the fused feature vectors into the classifier to output the information authenticity detection result.
7. The system according to claim 6, characterized in that The images exceeding the preset number in each sample are stitched together, including: Define the number of images as n; Calculate the square root of n and round down to get Define the number of rows as The pictures are stitched together, and the first s*s pictures form a complete square. If n>s 2 , the remaining images are arranged in row and col columns. If the last row is not completely filled, the remaining space is set to all 0s; Compress the spliced images.
8. The system according to claim 6, characterized in that The dataset enhancement module is also used to: Use OCR technology to scan the images of all samples, mark the samples with embedded text images, and perform denoising, binarization and rotation correction operations to identify the embedded text and splice it into the corresponding text of the sample; For the text part of the image, the mask operation is performed. The text area is defined as (x1, y1, x2, y2) coordinates, and the extended area is represented by (x′1, y′1, x′2, y′2), where x′1 = max(0, x1-20), y′1 = max(0, y1-20), x′2 = min(x max ,x2+20),y′2=min(y max ,y2+20); The color average of the annular area between the extended area and the text area is used as the filling of the text area, which is expressed by the following formula: Where R represents the annular area, and I(x,y) represents the color value of the coordinate point; For the target information p i Graph feature embedding, define the graph network including target information, users and comments; define the node embedding matrix X∈R |V|*d , where |V| is the number of nodes and d is the embedding dimension; calculate the similarity between nodes, if the similarity between two nodes is greater than the threshold β ij , it means there is a potential edge, add the potential edge to the original adjacency matrix A:
9. The system according to claim 6, characterized in that The multimodal feature learning module is also used for: For text content, BERT is used as an embedding model to obtain text features R t , and through the fully connected layer W t Map to 900 dimensions, output E t : E t σ(W t *R t +b) Where W t is the weight matrix of the text fully connected layer, σ(·) is the sigmoid activation function, b is the bias vector, and E t ∈R n*d , where n is the number of samples and d is the vector dimension; For image content, ResNet50 is used as the embedding model to obtain image features R i , and through the fully connected layer W i Map to 900 dimensions, output E i : E i σ(W i *R i +b) Where W i is the weight matrix of the fully connected layer of the image; The CA layer is used to align the features of the two modalities to achieve modal fusion; the CA layer contains two parallel CAblocks; For node i, calculate the neighbors (j∈N i ) and its own unnormalized attention weight ε ij : in and W are the parameters to be learned, e g represents the initial embedding of the node, [·||·] represents the concatenation of the transformed features of nodes i and j, and LeakyReLU is the activation function; When normalizing, the positive attention weight ε is calculated at the same time pos and negative attention weight ε nag : where ε i ={ε ij |j∈N i }, represents the unnormalized attention weights of node i and all its neighbor nodes j; The positive and negative attention weights are weighted and concatenated with the original nodes to obtain new features that incorporate neighborhood information nodes: Where W is the weight matrix of the fully connected layer, is the initial feature vector of neighbor node j.
10. The system according to claim 6, characterized in that The feature fusion and information detection module is also used for: E=concat(E ti ,AND tg ,AND gi ) The loss is divided into two parts, namely modality alignment loss and classification loss. The feature vectors of the three modalities are mapped to the same feature space, and the MSE loss is used for modality alignment: Where W is the mapping matrix to be learned; Input the feature vector E into the fully connected layer for prediction, and finally get the prediction result And use the cross entropy loss as the classification loss: Loss=λ c L class +λ a L align Where σ() is the softmax activation function; the final loss function is composed of alignment loss and classification loss.