Multi-modal Rumor Detection Method Based on Data Augmentation and Global Information Fusion
The method addresses multi-modal false news detection by using data augmentation and global information fusion to enhance feature extraction and alignment, resulting in improved detection performance by reducing noise interference and improving feature alignment across text and image modalities.
Patent Information
- Application Number
- CN202510519247.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing multimodal rumor detection methods have shortcomings in single-modal feature representation and cross-modal collaborative learning, and it is difficult to effectively deal with multimodal information, especially in frequency domain characterization and cross-modal collaborative learning, where there are limitations and efficiency bottlenecks at the feature optimization level.
Using a method based on data augmentation and global information fusion, a multi-modal rumor detection network architecture is built through EMA module and spectrum compression technology, parallel multi-scale convolution and dynamic attention fusion mechanisms are used, and spectrum compression technology is combined to achieve dynamic fusion of local details and global features and efficient collaborative integration of cross-modal information.
It improves the performance of multi-modal rumor detection, can more effectively identify and evaluate the authenticity of rumor content, enhances the model's ability to capture key information, and improves the detection effect in different data environments.
Smart Images

Figure CN120046120B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rumor detection, and specifically to a multi-modal rumor detection method based on data augmentation and global information fusion. Background Art
[0002] With the popularization of the Internet and the rapid development of artificial intelligence technology, social media has become the main channel for information dissemination globally. Currently, global social media users cover multiple platforms, which greatly facilitate information exchange and enable news, opinions, and social hotspots to spread rapidly. However, the convenience of information circulation has also given rise to the proliferation of false information. Rumors are usually spread by exaggerating, distorting, or fabricating facts and take advantage of the anonymity and decentralization characteristics of social media to spread rapidly. Especially in the current online environment, rumors not only appear in text form but may also combine multiple modalities such as pictures, videos, and audio, making the detection difficulty significantly increase.
[0003] Currently, the detection research on false news still mainly focuses on traditional manually written news content. However, with the rapid development of social media, news content automatically generated by large language models (LLMs) has begun to emerge in large numbers, posing greater challenges to traditional detection methods. Although existing research has focused on the detection of machine-generated news, most still remains limited to the analysis of pure text data and has not fully considered the dissemination characteristics of false news in a multi-modal environment. In fact, on today's social media platforms, false news usually combines manually written news and machine-generated content and may be accompanied by various modal information such as text, images, or videos, making the detection task more complex. Therefore, how to construct a detection framework that can not only identify manually written news but also effectively detect machine-generated content has become an important issue that urgently needs to be solved.
[0004] Although important breakthroughs have been made in the frequency domain characterization of current multi-modal rumor detection, two core challenges still need to be urgently addressed: (1) The inherent defects of single-modal feature representation. Existing methods rely on Fourier transform for global dependence modeling (such as FSRU enhancing feature discriminability through frequency domain sparsity), but there are still limitations in the feature optimization level. For example, the dynamic range mismatch problem of text features leads to high-frequency noise interference, and the direct transformation of uncalibrated text embeddings will amplify the amplitude differences of extreme word vectors (such as sentiment polarity words), resulting in the attenuation of semantic fundamental frequency energy; the image block embedding mechanism is difficult to effectively process multi-scale targets, and the high-frequency noise (such as crowd textures) in complex backgrounds and key features (such as tampered edge harmonics) produce aliasing effects in the frequency domain, resulting in a decrease in the signal-to-noise ratio of the low-frequency structure in the foreground text area. (2) The efficiency bottleneck of cross-modal collaborative learning. Although existing frequency domain fusion methods avoid the O(n 2) Complexity, however, due to the significant differences in the spectral distributions between the low-frequency semantic information of the text and the high-frequency edge features of the image, severely restricts the effectiveness of feature alignment. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a multi-modal rumor detection method based on data augmentation and global information fusion, aiming to solve the problems in the background art.
[0006] To achieve the above object, the present invention provides the following technical solution: A multi-modal rumor detection method based on data augmentation and global information fusion, characterized by comprising the following steps:
[0007] Step S1: Construct a data set, which contains a number of graphic news data;
[0008] Step S2: Obtain the text embedding matrix and the image embedding matrix in the graphic news data;
[0009] Step S3: Obtain the spectral features of the text embedding matrix and the spectral features of the image embedding matrix ;
[0010] Step S4: Compress the spectral features of the text embedding matrix and the spectral features of the image embedding matrix ;
[0011] Step S5: Strengthen the compressed spectral features of the text embedding matrix and the spectral features of the image embedding matrix to obtain the enhanced text spectral feature representation and the enhanced image spectral feature representation ;
[0012] Step S6: Perform inverse discrete Fourier transform on the enhanced text spectral feature representation and the enhanced image spectral feature representation to obtain the spatial domain information of the text and the spatial domain information of the image ;
[0013] Step S7: Perform weighted summation on the spatial domain information of the text and the spatial domain information of the image to obtain the final multi-modal representation , and input the final multi-modal representation into the fully connected layer for classification prediction.
[0014] Furthermore, given a set of graphic news data , denote the One graphic news data; Denote the th graphic news data as , , where represents the text content, represents the image.
[0015] Furthermore, the specific process of step S2 is as follows:
[0016] Step S2.1: Define the text content in the graphic news data , denotes a word with a length of in the text content;
[0017] Use the Word2Vec function to obtain the word embedding representation of each word in the text content , expressed as:
[0018] ;
[0019] In the formula, represents 's word embedding representation, denotes a word with a length of in the text content, ; , represents the set of real numbers, is the dimension of the word embedding representation; represents model;
[0020] Use the sine function and cosine function to perform positional encoding on each word in the text content to obtain the positional encoding representation of each word ;
[0021] Add the word embedding representation of each word in the text content and the positional encoding representation of each word element-wise to obtain the text embedding matrix , expressed as:
[0022] ;
[0023] ;
[0024] In the formula, is the embedding representation of each word in the text content, ;
[0025] Step S2.2: Define the image in the graphic news data ; Denote the image Divided into non - overlapping blocks, indicating the th block of the division, indicating rows, indicating columns, ; perform block embedding on the blocks using two - dimensional convolution and align with the text embedding matrix to obtain the block embedding of the image , expressed as:
[0026] ;
[0027] In the formula, represents the two - dimensional convolution operation; represents the flattening operation;
[0028] Use the sine function and cosine function to perform position encoding on the image to obtain the position encoding representation of the image ;
[0029] Add the block embedding of the image element - by - element with the position encoding representation of the image to obtain the image embedding matrix , expressed as:
[0030] .
[0031] Furthermore, the specific process of step S3 is as follows:
[0032] Step S3.1: Use the Sigmoid activation function to adaptively adjust the weights of the text embedding matrix , and fuse the text embedding matrix with the adjusted weights with the original text embedding matrix to obtain the text feature representation :
[0033] ;
[0034] In the formula, represents the Sigmoid activation function; represents the Hadamard product;
[0035] Use the discrete Fourier transform to extract the spectral features of the text feature representation , that is, the spectral features of the text embedding matrix , expressed as:
[0036] ;
[0037] In the formula, is the text feature representation of the th word in the text content; represents the spectrum value at the frequency of ; represents the frequency index; represents the imaginary unit; represents the base of the natural logarithm;
[0038] Step S3.2: Use the EMA module to optimize the image embedding matrix . The EMA module will first divide along the channel dimension into sub-feature maps, and each sub-feature map contains channels, which is expressed as:
[0039] ;
[0040] In the formula, represents the th sub-feature map, and represent the height and width respectively;
[0041] Perform average pooling on in the horizontal and vertical directions respectively, and obtain the feature representation after pooling along the height direction and the feature representation after pooling along the width direction ;
[0042] Concatenate and , and perform feature fusion through one-dimensional convolution to obtain the feature representation that fuses the height and width direction information:
[0043] ;
[0044] In the formula, represents concatenating the features;
[0045] After that, is split back into and , generate attention weights through the Sigmoid function, recalibrate and , and perform group normalization operations to obtain the enhanced global feature representation :
[0046] ;
[0047] In the formula, represents the group normalization operation;
[0048] For a three-dimensional convolution is used to obtain an enhanced local feature representation , and through the function and matrix multiplication operations, spatial attention weights are generated , expressed as:
[0049] ;
[0050] In the formula, represents the average pooling operation;
[0051] By matrix dot product, and are fused to obtain the image feature representation , expressed as:
[0052] ;
[0053] The discrete Fourier transform is used to extract the spectral features of the image feature representation , that is, the spectral features of the image embedding matrix , expressed as: , expressed as:
[0054] ;
[0055] In the formula, represents the spectral value at the frequency of ; represents the -th dimensional feature representation of the image; is the spatial dimension of the image.
[0056] Furthermore, the specific process of step S4 is: For the spectral features of the text embedding matrix and the image embedding matrix , a filter bank is introduced to compress the spectral components of different modalities, expressed as:
[0057] ;
[0058] In the formula, represents the spectral features of the compressed text embedding matrix and the image embedding matrix , represents the compressed , represents the compressed ; is the length; is the number of filters; represents the th filter.
[0059] Furthermore, enhance the spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix to obtain the enhanced text spectral feature representation and the enhanced image spectral feature representation :
[0060] ;
[0061] ;
[0062] In the formula, and respectively represent the mapping functions of the image and the text; represents a one-dimensional convolution operation, represents an average pooling operation.
[0063] Furthermore, perform an inverse discrete Fourier transform on the enhanced text spectral feature representation and the enhanced image spectral feature representation to losslessly convert the spectral representations of the text and the image back to the spatial domain, obtaining the spatial domain information of the text and the spatial domain information of the image :
[0064] ;
[0065] ;
[0066] In the formula, represents the spatial domain feature representation of the th word of the text content; represents the spatial domain feature representation of the th dimension of the image; represents the enhanced text spectral feature representation on the th frequency component; represents the enhanced image spectral feature representation on the th frequency component.
[0067] Furthermore, perform a weighted sum on the spatial domain information of the text and the spatial domain information of the image to obtain the final multi-modal representation , and input the final multi-modal representation into the fully connected layer For classification prediction, it is expressed as:
[0068] ;
[0069] ;
[0070] In the formula, and are both trainable parameters; is the similarity score; represents the prediction label of the rumor. When is less than 0.5, it is determined that the corresponding graphic news data is non-rumor. When is greater than 0.5, it is determined that the corresponding graphic news data is a rumor. When is equal to 0.5, the corresponding graphic news data is randomly divided into rumor or non-rumor.
[0071] Furthermore, the similarity score is expressed as:
[0072] ;
[0073] In the formula, represents the JS divergence; represents the posterior probability distribution of the text feature representation under the given text content ; represents the posterior probability distribution of the image feature representation under the given image ;
[0074] Compared with the existing technologies, the present invention has the following beneficial effects:
[0075] (1) The present invention constructs a new network architecture for multi-modal rumor detection based on data augmentation and global information fusion. At the front end of frequency domain feature extraction, an EMA module is introduced, and through parallel multi-scale convolution and dynamic attention fusion mechanism, the dynamic fusion of local details and global features is realized. In addition, the spectrum compression technology is adopted to adjust the dynamic range of text features, suppress the influence of extreme values on the spectrum, and enhance the model's ability to capture key information.
[0076] (2) The present invention combines spectrum compression and EMA visual enhancement strategies to effectively suppress the cross-modal spectrum distribution differences, enhance the mutual information content between text and image features, achieve efficient collaborative integration of cross-modal information, and improve the ability to evaluate the authenticity of rumor content; in the feature extraction stage, a frequency-domain filter bank is used for spectrum compression to generate a spectrum compression representation, revealing the potential feature patterns in each modality. At the same time, making full use of the complementarity between modalities, key spectrum components helpful for rumor identification are selected from different modality information. Compared with the current multi-modal fusion methods based on spatial representation, this framework can more effectively improve the performance of rumor detection and shows better detection effects in different data environments. Brief Description of the Drawings
[0077] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0078] As Figure 1 shown, the present invention provides a technical solution: a multi-modal rumor detection method based on data augmentation and global information fusion, including the following steps:
[0079] Step S1: Construct a data set, which contains a number of graphic news data.
[0080] The specific process of Step S1 is as follows: Due to the lack of a publicly available graphic fake news data set generated by large language models, the present invention constructs a new Chinese graphic news data set and an English graphic data set based on common news topics.
[0081] These data sets are all generated by the ChatGPT-4 model and screened manually to ensure the representativeness of the data. The diversity of the data sets helps to train a more generalizable detection model.
[0082] Based on these news topics, multiple specific events are set under each news topic in a more fine-grained manner to avoid the generation of repeated news events, which can also enrich the diversity of the news data set. Thus, a number of different instructions are obtained, and these instructions are input into the ChatGPT-4 model to generate Chinese and English graphic news data.
[0083] The written instructions will require the ChatGPT-4 model to generate graphic news data based on historical real events, which is equivalent to tampering with historical real events to make the generated news as credible as possible, rather than allowing news readers to immediately determine its truth or falsehood after reading.
[0084] To conform to the format of short news texts, for Chinese graphic news data, the text length is controlled at about 140 characters; for English graphic news data, the text length is controlled within 140 words.
[0085] Finally, two Chinese text-image news datasets and an English text-image dataset covering multiple news topics were obtained. Among them, the English text-image news dataset contains 605 English text-image news data, and the Chinese text-image news dataset contains 1420 Chinese text-image news data. Each English text-image news data and Chinese text-image news data includes text content and an image.
[0086] Define the multi-modal rumor detection task as a binary classification problem. Given a set of text-image news data , denote the th text-image news data; denote the th text-image news data as , , where represents the text content, represents the image. The purpose of the present invention is to comprehensively utilize the features of both text and image to predict the label of each text-image news data.
[0087] Step S2: Obtain the text embedding matrix and the image embedding matrix in the text-image news data.
[0088] Step S2.1: Define the text content in the text-image news data , denote the word with length in the text content; Encode the text content by combining word embedding mapping and positional encoding.
[0089] Use the Word2Vec function to obtain the word embedding representation of each word in the text content , expressed as:
[0090] ;
[0091] In the formula, denotes the word embedding representation of , denotes the word with length in the text content, ; , denotes the set of real numbers, is the dimension of the word embedding representation; denotes model for generating word embeddings.
[0092] Perform positional encoding on each word in the text content using sine and cosine functions to obtain the positional encoding representation of each word ; Specifically, for the text content Position in and the embedding dimension index , where the even and odd dimensions of the position encoding vector are generated by sine and cosine functions respectively, and are expressed as:
[0093] ;
[0094] ;
[0095] In the formula, represents the position encoding representation; represents the position index of a certain word in the text content. The exponential term is used to control the frequency difference between different dimensions, so that the sine and cosine functions of different dimensions have different periods, in order to capture absolute and relative position information and enhance the model's understanding of the sequence structure. Finally, the position encoding representation of each word can be obtained .
[0096] Add the word embedding representation of each word in the text content and the position encoding representation of each word element by element to obtain the text embedding matrix , which is expressed as:
[0097] ;
[0098] ;
[0099] In the formula, is the embedding representation of each word in the text content, .
[0100] Step S2.2: Define the image in the graphic news data ; Divide the image into non-overlapping blocks, represents the th block of the division, represents the row, represents the column, ; Use two-dimensional convolution to perform block embedding on the blocks and align them with the text embedding matrix to obtain the block embedding of the image, which is expressed as:
[0101] ;
[0102] In the formula, represents the two-dimensional convolution operation; represents the flattening operation.
[0103] Use the sine function and cosine function for the image Perform positional encoding to obtain the position encoding representation of the image of the image .
[0104] Add the patch embeddings of the image element-wise to the position encoding representation of the image to obtain the image embedding matrix , denoted as:
[0105] .
[0106] Step S3: Obtain the spectral features of the text embedding matrix and the spectral features of the image embedding matrix of the image
[0107] Step S3.1: Use the Sigmoid activation function to adaptively adjust the weights of the text embedding matrix , and fuse the text embedding matrix with adjusted weights with the original text embedding matrix to obtain the text feature representation :
[0108] ;
[0109] wherein denotes the Sigmoid activation function; denotes the Hadamard product
[0110] Extract the spectral features of the text feature representation using the discrete Fourier transform (DFT), that is, the spectral features of the text embedding matrix of the image , denoted as:
[0111] ;
[0112] wherein is the text feature representation of the th word in the text content; denotes the spectral value at the frequency of , denotes the frequency index; denotes the imaginary unit; denotes the base of the natural logarithm
[0113] Step S3.2: Use the EMA module to optimize the image embedding matrix . To reduce the computational overhead, the EMA module first divides along the channel dimension into Sub-feature maps, each sub-feature map containing channels, denoted as:
[0114] ;
[0115] In the formula, represents the th sub-feature map, and represent the height and width respectively.
[0116] To capture the global information of the feature map in the height and width directions, is respectively subjected to average pooling in the horizontal and vertical directions, and the feature representation after pooling along the height direction and the feature representation after pooling along the width direction are obtained.
[0117] Combine and through simple concatenation and perform feature fusion through one-dimensional convolution to obtain the feature representation that fuses the information in the height and width directions:
[0118] ;
[0119] In the formula, represents concatenating the features.
[0120] After that, is split back into and ; is split back into and which is equivalent to splitting the information in the height and width directions again, so that and fuse the global information in the height and width directions, enabling to retain the width-direction perception information in the height direction, to retain the height-direction perception information in the width direction, and enhancing the feature representation ability. After splitting, and are not just the original and , but an enhanced version; simply put, the purpose of this step is to restore the information in the height and width directions while enhancing the feature representation ability.
[0121] Generate attention weights through the Sigmoid function to recalibrate and , and perform group normalization operation to make the eigenvalue distribution more stable, obtaining an enhanced global feature representation :
[0122] ;
[0123] In the formula, represents the group normalization operation.
[0124] In addition, for perform 3D convolution to obtain an enhanced local feature representation , and generate spatial attention weights through function and matrix multiplication operation , expressed as:
[0125] ;
[0126] In the formula, represents the average pooling operation.
[0127] Finally, fuse and through matrix dot multiplication to obtain the image feature representation , expressed as:
[0128] .
[0129] Adopt discrete Fourier transform to extract the spectral features of the image feature representation , that is, the spectral features of the image embedding matrix , expressed as:
[0130] ;
[0131] In the formula, represents the spectral value at the frequency of ; represents the feature representation of the th dimension of the image; is the spatial dimension of the image.
[0132] In the frequency domain, the spatial features within different frequency components can be effectively integrated, enabling the effective extraction of key information in text and images.
[0133] Step S4: Compress the spectral features of the text embedding matrix and the spectral features of the image embedding matrix .
[0134] After obtaining the spectral representations of text and image, for the text embedding matrix and the image embedding matrix Spectrum characteristics Introduce a filter bank to compress the spectral components of different modalities to obtain important features related to rumors, expressed as:
[0135] ;
[0136] In the formula, represents the compressed text embedding matrix and the spectrum characteristics of the image embedding matrix ; represents the compressed , represents the compressed ; is the length of; is the number of filters; represents the th filter; is used to concentrate energy more efficiently and strengthen the information highly related to rumor features in the spectrum.
[0137] Step S5: Strengthen the spectrum characteristics of the compressed text embedding matrix and the spectrum characteristics of the image embedding matrix to obtain the enhanced text spectrum feature representation and the enhanced image spectrum feature representation
[0138] ;
[0139] ;
[0140] In the formula, and respectively represent the mapping functions of the image and the text, used to map the original features to a new feature space; represents a one-dimensional convolution operation, represents an average pooling operation.
[0141] Step S6: Perform the inverse discrete Fourier transform (IDFT) on the enhanced text spectrum feature representation and the enhanced image spectrum feature representation to losslessly transform the spectral representations of the text and the image back to the spatial domain, obtaining the spatial domain information of the text and the spatial domain information
[0142] ;
[0143] ;
[0144] In the formula, represents the spatial domain feature representation of the th word of the text content; represents the spatial domain feature representation of the th dimension of the image; represents the enhanced text spectrum feature representation on the th frequency component; represents the enhanced image spectrum feature representation on the th frequency component.
[0145] Step S7: Perform weighted summation on the spatial domain information of the text and the spatial domain information of the image to obtain the final multi-modal representation , and input the final multi-modal representation into the fully connected layer for classification prediction.
[0146] ;
[0147] ;
[0148] In the formula, and are both trainable parameters; is the similarity score. In this embodiment, is used as a hyperparameter to adaptively adjust the fusion weight of cross-modal features; represents the prediction label of the rumor. When is less than 0.5, it is determined that the corresponding text-image news data is not a rumor. When is greater than 0.5, it is determined that the corresponding text-image news data is a rumor. When is equal to 0.5, the corresponding text-image news data is randomly divided into a rumor or not a rumor.
[0149] Among them, the similarity score is expressed as:
[0150] ;
[0151] In the formula, represents the Jensen-Shannon (JS) divergence; represents the posterior probability distribution of the text feature representation given the text content ; represents the posterior probability distribution of the image feature representation given the image ;
[0152] Finally, rumor detection is regarded as a binary classification task, and cross-entropy loss is adopted to optimize the classification performance.
[0153] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-modal rumor detection method based on data augmentation and global information fusion, characterized in that, Including the following steps: Step S1: Construct a dataset, which contains a number of graphic news data; the graphic news data includes text content and images; Step S2: Obtain the text embedding matrix and the image embedding matrix in the graphic news data; Step S3: Obtain the text embedding matrix and the spectral features of the image embedding matrix ; Step S4: Compress the spectral features of the text embedding matrix and the spectral features of the image embedding matrix ; Step S5: Enhance the spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix to obtain the enhanced text spectral feature representation and the enhanced image spectral feature representation ; Step S6: Perform inverse discrete Fourier transform on the enhanced text spectral feature representation and the enhanced image spectral feature representation to obtain the spatial domain information of the text and the spatial domain information of the image ; Step S7: Perform weighted summation on the spatial domain information of the text and the spatial domain information of the image to obtain the final multi-modal representation , and input the final multi-modal representation into the fully connected layer for classification prediction; The specific process of Step S3 is as follows: Step S3.1: Use the Sigmoid activation function to adaptively adjust the weights of the text embedding matrix and fuse the text embedding matrix with adjusted weights and the original text embedding matrix to obtain the text feature representation : ; In the formula, represents the Sigmoid activation function; represents the Hadamard product; Use the discrete Fourier transform to extract the text feature representation of the spectral features , that is, the text embedding matrix of the spectral features , which is expressed as: ; In the formula, is the text feature representation of the th word in the text content; represents the spectral value at the frequency of , represents the frequency index; represents the imaginary unit; represents the base of the natural logarithm; represents the length of the word in the text content; Step S3.2: Use the EMA module to optimize the image embedding matrix The EMA module will first divide it along the channel dimension into sub-feature maps, and each sub-feature map contains channels, expressed as: ; In the formula, represents the th sub-feature map, represents the set of real numbers, and represent the height and width respectively; Pair Perform average pooling in the horizontal and vertical directions respectively to obtain Feature representation after pooling along the height direction And Feature representation after pooling along the width direction ; Concatenate and and perform feature fusion through one-dimensional convolution to obtain a feature representation that combines the height and width direction information : ; In the formula, denotes the splicing of features; denotes a one-dimensional convolution operation; After that, split back into and , generate attention weights through the Sigmoid function, and recalibrate and , and perform group normalization operations to obtain an enhanced global feature representation : ; In the formula, represents the group normalization operation; Pair Use three-dimensional convolution to obtain an enhanced local feature representation , and generate spatial attention weights through function and matrix multiplication operations, expressed as: : ; In the formula, represents the average pooling operation; Fusion by matrix dot product and to obtain the image feature representation which is expressed as: ; Using the discrete Fourier transform to extract the image feature representation of the spectral features, that is, the spectral features of the image embedding matrix are represented as: , expressed as: ; In the formula, represents the spectral value at the frequency of ; represents the -dimensional feature representation of the image; is the spatial dimension of the image; The specific process of step S4 is: Embed a matrix for the text and an image embedding matrix of the spectral characteristics Introduce a filter bank to compress the spectral components of different modalities, expressed as: ; In the formula, represents the spectral characteristics of the compressed text embedding matrix and the image embedding matrix ; represents the compressed , represents the compressed ; is the length of; is the number of filters; represents the th filter.
2. The multimodal rumor detection method based on data augmentation and global information fusion according to claim 1, characterized in that: Given a set of graphic news data , indicating the th graphic news data; Denote the th graphic news data as , , where represents the text content, and represents the image.
3. The multimodal rumor detection method based on data augmentation and global information fusion according to claim 2, characterized in that: The specific process of Step S2 is as follows: Step S2.1: Define the text content in the graphic news data , denote the word with a length of in the text content; Use the Word2Vec function to obtain the word embedding representation of each word in the text content , which is expressed as: ; In the formula, represents the word embedding representation of which represents a word of length in the text content; ; , where is the dimension of the word embedding representation; represents the model; Use the sine function and cosine function to perform positional encoding on each word in the text content to obtain the positional encoding representation of each word ; The word embedding representation of each word in the text content and the positional encoding representation of each word are added element by element to obtain the text embedding matrix , which is expressed as: ; ; In the formula, is the embedding representation of each word in the text content, ; Step S2.2: Define the images in the graphic news data ; Divide the image into non-overlapping blocks, where represents the -th block of the division, represents the row, represents the column, ; Perform block embedding on the blocks using two-dimensional convolution and align with the text embedding matrix to obtain the block embedding of the image , expressed as: ; In the formula, represents a two-dimensional convolution operation; represents a flattening operation; Use the sine function and cosine function to encode the position of the image to obtain the position encoding representation of the image ; ; Embed the patches of the image with the positional encoding representation of the image element-wise add them to obtain the image embedding matrix , denoted as: 。 4. The multi-modal rumor detection method based on data augmentation and global information fusion according to claim 3, wherein: The spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix are enhanced to obtain the enhanced text spectral feature representation and the enhanced image spectral feature representation : ; ; In the formula, and represent the mapping functions of the image and the text respectively, represents the average pooling operation.
5. The multimodal rumor detection method based on data augmentation and global information fusion according to claim 4, wherein: Perform inverse discrete Fourier transform on the enhanced text spectral feature representation and the enhanced image spectral feature representation to convert the spectral representations of the text and the image back to the spatial domain losslessly, obtaining the spatial domain information of the text and the spatial domain information of the image : ; ; In the formula, represents the spatial domain feature representation of the th word of the text content; represents the spatial domain feature representation of the th dimension of the image; represents the enhanced text spectral feature representation on the th frequency component; represents the enhanced image spectral feature representation on the th frequency component.
6. The multi-modal rumor detection method based on data augmentation and global information fusion according to claim 5, wherein: Spatial domain information of the text and spatial domain information of the image are weighted and summed to obtain the final multi-modal representation , and the final multi-modal representation is input into the fully connected layer for classification prediction, expressed as: ; ; In the formula, and are both trainable parameters; is the similarity score; represents the predicted label of the rumor. When is less than 0.5, it is judged that the corresponding picture-text news data is non-rumor. When is greater than 0.5, it is judged that the corresponding picture-text news data is a rumor. When is equal to 0.5, the corresponding picture-text news data is randomly divided into rumor or non-rumor.
7. The multi-modal rumor detection method based on data augmentation and global information fusion according to claim 6, characterized in that: Similarity score Expressed as: ; In the formula, represents the JS divergence; represents the posterior probability distribution of the text feature representation under the condition of the given text content ; represents the posterior probability distribution of the image feature representation under the condition of the given image .
Citation Information
Patent Citations
False news detection method and system based on progressive multi-modal fusion network
CN114528912A
False news detection method based on multi-view and hierarchical fusion
CN118114188A