Multi-modal rumor detection method based on data enhancement and global information fusion
By introducing EMA module and spectrum compression technology in multimodal rumor detection, the inherent defects of single-modal feature characterization and the efficiency bottleneck of cross-modal collaborative learning are solved, and more efficient multimodal information collaborative integration and rumor detection performance are achieved.
Patent Information
- Application Number
- CN202510519247.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing multimodal rumor detection methods have inherent flaws in dealing with single-modal feature representations, such as dynamic range mismatch of text features and the difficulty of image chunking embedding mechanisms to deal with multi-scale targets, resulting in poor detection results. In addition, the efficiency bottleneck of cross-modal collaborative learning also limits the effectiveness of feature alignment.
The multimodal detection method based on data augmentation and global information fusion is adopted to dynamically fusion of local details and global features through the EMA module, and the dynamic range of text features is adjusted using spectrum compression technology to suppress the impact of extreme values on the spectrum. At the same time, combining spectrum compression and EMA visual enhancement strategies, it effectively suppresses the differences in cross-modal spectrum distribution and improves the mutual information of text and image features.
This method effectively suppresses the differences in cross-modal spectrum distribution, improves the mutual information of text and image features, improves the ability to evaluate the authenticity of rumor content, and significantly improves the performance of multimodal rumor detection.
Smart Images

Figure CN120046120A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rumor detection, and specifically to a multimodal rumor detection method based on data enhancement and global information fusion. Background Art
[0002] With the popularization of the Internet and the rapid development of artificial intelligence technology, social media has become the main channel for information dissemination worldwide. At present, global social media users cover multiple platforms, which greatly facilitate information exchange and enable news, opinions and social hot spots to spread rapidly. However, the convenience of information flow has also spawned the proliferation of false information. Rumors are usually spread by exaggerating, distorting or falsifying facts, and they spread rapidly by taking advantage of the anonymity and decentralized nature of social media. Especially in the current network environment, rumors not only appear in text form, but may also be combined with multiple modalities such as pictures, videos, and audio, making detection significantly more difficult.
[0003] At present, the research on fake news detection still mainly focuses on traditional manually written news content. However, with the rapid development of social media, news content automatically generated by large language models (LLMs) has begun to emerge in large quantities, making traditional detection methods face greater challenges. Although existing studies have focused on the detection of machine-generated news, most of them are still limited to the analysis of pure text data, and have not fully considered the propagation characteristics of fake news in a multimodal environment. In fact, on today's social media platforms, fake news usually combines manually written news with machine-generated content, and may be accompanied by multiple modal information such as text, images or videos, making the detection task more complicated. Therefore, how to build a detection framework that can both identify manually written news and effectively detect machine-generated content has become an important issue that needs to be solved urgently.
[0004] Although multimodal rumor detection has made important breakthroughs in frequency domain representation, it still faces two core challenges that need to be addressed: (1) The inherent defects of single-modal feature representation. Existing methods rely on Fourier transform for global dependency modeling (such as FSRU to enhance feature discrimination through frequency domain sparsity), but there are still limitations in the feature optimization level. For example, the dynamic range mismatch problem of text features leads to high-frequency noise interference. Uncalibrated direct transformation of text embedding will amplify the amplitude difference of extreme word vectors (such as sentiment polarity words), resulting in semantic fundamental frequency energy attenuation; image block embedding mechanism is difficult to effectively handle multi-scale targets. High-frequency noise in complex backgrounds (such as crowd texture) and key features (such as tampering edge harmonics) produce aliasing effects in the frequency domain, resulting in a decrease in the signal-to-noise ratio of low-frequency structures in the foreground text area. (2) The efficiency bottleneck of cross-modal collaborative learning. Although the existing frequency domain fusion method circumvents the O(n) computational complexity of spatial domain attention, it can effectively deal with multi-scale targets. 2) complexity, but the effectiveness of feature alignment is severely restricted due to the significant difference in spectral distribution between the low-frequency semantic information of text and the high-frequency edge features of images. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a multimodal rumor detection method based on data enhancement and global information fusion, which aims to solve the problems in the background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a multimodal rumor detection method based on data enhancement and global information fusion, characterized in that it comprises the following steps: Step S1: construct a data set, which includes a number of graphic news data; Step S2: Obtaining a text embedding matrix and an image embedding matrix in the graphic news data; Step S3: Get text embedding matrix Spectral features and image embedding matrix Spectral characteristics of Step S4: Text embedding matrix Spectral features and image embedding matrix The spectral characteristics of Step S5: Spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix Enhanced text spectrum feature representation is obtained And enhance the image spectral feature representation ; Step S6: Enhance the text spectral feature representation And enhance the image spectral feature representation Perform inverse discrete Fourier transform to obtain the spatial domain information of the text And the spatial domain information of the image ; Step S7: Spatial domain information of the text And the spatial domain information of the image Perform weighted summation to obtain the final multimodal representation , input the final multimodal representation into the fully connected layer to make classification predictions.
[0007] Furthermore, given a set of graphic news data , Indicates picture and text news data; Graphic news data , ,in, Represents the text content. Represents an image.
[0008] Furthermore, the specific process of step S2 is as follows: Step S2.1: Define the text content in the graphic news data , Indicates that the length of the text content is words; Use the Word2Vec function to obtain the word embedding representation of each word in the text content , expressed as: ; In the formula, express The word embedding representation of Indicates that the length of the text content is Words, ; , represents the set of real numbers, is the dimension of word embedding representation; express Model; Use sine and cosine functions to analyze text content Each word in the position is encoded to obtain the position encoding representation of each word ; Embed each word in the text content and the positional encoding representation of each word Add element by element to get the text embedding matrix , expressed as: ; ; In the formula, is the embedding representation of each word in the text content, ; Step S2.2: Defining images in graphic news data ; The image Divide into non-overlapping blocks, Indicates the division Blocks, Indicates the line, Represents a column, ; Use two-dimensional convolution to embed the blocks and align them with the text embedding matrix to obtain the block embedding of the image , expressed as: ; In the formula, Represents a two-dimensional convolution operation; Represents a flattening operation; Use sine and cosine functions to transform the image Perform position encoding to obtain image Positional encoding representation ; Embed the image into blocks Positional encoding representation of the image Add element by element to get the image embedding matrix , expressed as: .
[0009] Furthermore, the specific process of step S3 is as follows: Step S3.1: Adaptively embed the text matrix using the Sigmoid activation function Adjust the weights and compare the weight-adjusted text embedding matrix with the original text embedding matrix Fusion to obtain text feature representation : ; In the formula, Represents the Sigmoid activation function; represents the Hadamard product; Using discrete Fourier transform to extract text feature representation The spectrum characteristics of , that is, the text embedding matrix The spectrum characteristics of , expressed as: ; In the formula, For the text content Text feature representation of words; express At a frequency of The spectrum value at represents the frequency index; represents an imaginary unit; represents the base of natural logarithms; Step S3.2: Use EMA module to embed the image into matrix To optimize, the EMA module first converts Divided along the channel dimension into sub-feature graphs, each of which contains channels, expressed as: ; In the formula, Indicates Sub-feature graphs, and Represents height and width respectively; right Perform average pooling in the horizontal and vertical directions respectively, and get Feature representation after pooling along the height direction as well as Feature representation after pooling along the width direction ; Will and The splicing is performed and the feature fusion is performed through one-dimensional convolution to obtain the feature representation of the fusion height and width direction information. : ; In the formula, Indicates the splicing of features; Afterwards Split back and , generate attention weights through the Sigmoid function and recalibrate and , and processed by group normalization operation to obtain enhanced global feature representation : ; In the formula, Represents the group normalization operation; right Enhanced local feature representation using 3D convolution and through Function and matrix multiplication operations generate spatial attention weights , expressed as: ; In the formula, represents the average pooling operation; Fusion through matrix multiplication and , get the image feature representation , expressed as: ; Using Discrete Fourier Transform to Extract Image Feature Representation The spectral features of the image are embedded in the matrix The spectrum characteristics of , expressed as: ; In the formula, express At a frequency of The spectrum value at ; Represents the image Dimensional feature representation; is the spatial dimension of the image.
[0010] Furthermore, the specific process of step S4 is: embedding matrix and the image embedding matrix The spectrum characteristics of A filter bank is introduced to compress the spectral components of different modes, which can be expressed as: ; In the formula, Represents the compressed text embedding matrix and the image embedding matrix The spectrum characteristics of Indicates the compressed , Indicates the compressed ; for Length; is the number of filters; Indicates A filter.
[0011] Furthermore, the spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix Enhanced text spectrum feature representation is obtained And enhance the image spectral feature representation : ; ; In the formula, and Represent the mapping functions of images and texts respectively; represents a one-dimensional convolution operation, Represents an average pooling operation.
[0012] Furthermore, the spectral feature representation of the enhanced text And enhance the image spectral feature representation Perform inverse discrete Fourier transform to convert the spectral representation of text and image back to the spatial domain losslessly to obtain the spatial domain information of the text And the spatial domain information of the image : ; ; In the formula, Indicates the text content The spatial domain feature representation of each word; Represents the image Dimensional spatial domain feature representation; Indicated in Enhanced text spectral feature representation on frequency components; Indicated in Enhanced image spectral feature representation on frequency components.
[0013] Furthermore, the spatial domain information of the text And the spatial domain information of the image Perform weighted summation to obtain the final multimodal representation , input the final multimodal representation into the fully connected layer To make classification predictions, it is expressed as: ; ; In the formula, and All are trainable parameters; is the similarity score; Represents the predicted label of the rumor, when When it is less than 0.5, the corresponding graphic news data is judged as non-rumor. When it is greater than 0.5, the corresponding graphic news data is judged to be a rumor. When it is equal to 0.5, the corresponding graphic news data will be randomly divided into rumors or non-rumors.
[0014] Furthermore, the similarity score It is expressed as: ; In the formula, represents JS divergence; Indicates that in the given text content In the case of The posterior probability distribution of ; Indicates that in a given image In the case of The posterior probability distribution of .
[0015] Compared with the existing technology, the present invention has the following beneficial effects:
[0016] (1) The present invention constructs a new network architecture for multimodal rumor detection based on data enhancement and global information fusion. In the frequency domain feature extraction front end, the EMA module is introduced to achieve dynamic fusion of local details and global features through parallel multi-scale convolution and dynamic attention fusion mechanism. In addition, spectrum compression technology is used to adjust the dynamic range of text features, suppress the influence of extreme values on the spectrum, and enhance the model's ability to capture key information.
[0017] (2) The present invention combines spectrum compression with EMA visual enhancement strategy to effectively suppress the difference in cross-modal spectrum distribution, improve the mutual information of text and image features, realize efficient collaborative integration of cross-modal information, and improve the ability to evaluate the authenticity of rumor content; in the feature extraction stage, the frequency domain filter group is used to perform spectrum compression to generate a spectrum compression representation, which reveals the potential feature patterns in each modality. At the same time, the complementarity between modalities is fully utilized to select key spectral components that are helpful in identifying rumors from different modal information. Compared with the current multimodal fusion method based on spatial representation, this framework can more effectively improve the performance of rumor detection and show better detection effects in different data environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION
[0019] like Figure 1 As shown, the present invention provides a technical solution: a multimodal rumor detection method based on data enhancement and global information fusion, comprising the following steps:
[0020] Step S1: construct a data set, which contains a number of graphic news data.
[0021] The specific process of step S1 is: due to the lack of a public large language model generated fake news image and text dataset, the present invention constructs a new Chinese image and text news dataset and an English image and text dataset based on 15 common news topics (such as culture and entertainment, social politics, international fields, diffusion and forwarding, official administration and anti-corruption, education and employment, food safety, life tips, emergencies, transportation, natural disasters, scientific research, health and epidemic prevention, financial news, military fields, etc.).
[0022] These data sets are all generated by the ChatGPT-4 model and are manually screened to ensure the representativeness of the data. The diversity of the data sets helps to train detection models with better generalization capabilities.
[0023] Based on these 15 major news topics, we set 30-50 specific events under each news topic in a fine-grained manner to avoid the generation of repeated news events, which can also enrich the diversity of the news data set. As a result, we obtained about 450 different instructions, which were input into the ChatGPT-4 model to generate Chinese and English graphic news data.
[0024] The instructions written will require the ChatGPT-4 model to base the generation of graphic news data on real historical events, which is equivalent to tampering with real historical events, so as to make the generated news as credible as possible, rather than allowing news readers to immediately determine its authenticity after reading it.
[0025] To comply with the format of short news text, for Chinese graphic news data, the length of the text is controlled within 140 characters; for English graphic news data, the length of the text is controlled within 140 words.
[0026] Finally, we obtained two Chinese picture news datasets and English picture news datasets covering 15 news topics. Among them, the English picture news dataset contains 605 English picture news data, and the Chinese picture news dataset contains 1420 Chinese picture news data. Each English picture news data and Chinese picture news data includes text content and an image.
[0027] The multimodal rumor detection task is defined as a binary classification problem. Given a set of graphic news data , Indicates picture and text news data; Graphic news data , ,in, Represents the text content. The purpose of the present invention is to comprehensively utilize the features of both text and image to predict the label of each graphic news data.
[0028] Step S2: Obtain the text embedding matrix and image embedding matrix in the graphic news data.
[0029] Step S2.1: Define the text content in the graphic news data , Indicates that the length of the text content is words; the text content is encoded by combining word embedding mapping and position encoding.
[0030] Use the Word2Vec function to obtain the word embedding representation of each word in the text content , expressed as: ; In the formula, express The word embedding representation of Indicates that the length of the text content is Words, ; , represents the set of real numbers, is the dimension of word embedding representation; express Model for generating word embeddings.
[0031] Use sine and cosine functions to analyze text content Each word in the position is encoded to obtain the position encoding representation of each word Specifically, for text content Position in and embedding dimension index , the even and odd dimensions of the position encoding vector are generated by sine and cosine functions respectively, expressed as: ; ; In the formula, represents positional encoding representation; Indicates the position index of a word in the text content. Index item It is used to control the frequency difference between different dimensions so that the sine and cosine functions of different dimensions have different periods to capture absolute and relative position information and enhance the model's understanding of the sequence structure. Finally, the position encoding representation of each word can be obtained. .
[0032] Embed each word in the text content and the positional encoding representation of each word Add element by element to get the text embedding matrix , expressed as: ; ; In the formula, is the embedding representation of each word in the text content, .
[0033] Step S2.2: Defining images in graphic news data ; The image Divide into non-overlapping blocks, Indicates the division Blocks, Indicates the line, Represents a column, ; Use two-dimensional convolution to embed the blocks and align them with the text embedding matrix to obtain the block embedding of the image , expressed as: ; In the formula, Represents a two-dimensional convolution operation; Represents a flatten operation.
[0034] Use sine and cosine functions to transform the image Perform position encoding to obtain image Positional encoding representation .
[0035] Embed the image into blocks Positional encoding representation of the image Add element by element to get the image embedding matrix , expressed as: .
[0036] Step S3: Get text embedding matrix Spectral features and image embedding matrix spectral characteristics.
[0037] Step S3.1: Adaptively embed the text matrix using the Sigmoid activation function Adjust the weights and compare the weight-adjusted text embedding matrix with the original text embedding matrix Fusion to obtain text feature representation : ; In the formula, Represents the Sigmoid activation function; represents the Hadamard product.
[0038] Using Discrete Fourier Transform (DFT) to extract text feature representation The spectrum characteristics of , that is, the text embedding matrix The spectrum characteristics of , expressed as: ; In the formula, For the text content Text feature representation of words; express At a frequency of The spectrum value at represents the frequency index; represents an imaginary unit; Represents the base of natural logarithms.
[0039] Step S3.2: Use EMA module to embed the image into matrix To optimize the process, in order to reduce the computational overhead, the EMA module first converts Divided along the channel dimension into sub-feature graphs, each of which contains channels, expressed as: ; In the formula, Indicates Sub-feature graphs, and Represents height and width respectively.
[0040] In order to capture the global information of the feature map in the height and width directions, Perform average pooling in the horizontal and vertical directions respectively, and get Feature representation after pooling along the height direction as well as Feature representation after pooling along the width direction .
[0041] Will and Simple splicing is performed, and feature fusion is performed through one-dimensional convolution to obtain a feature representation that fused the height and width information. : ; In the formula, Indicates that features are concatenated.
[0042] Afterwards Split back and ;Will Split back and This is equivalent to splitting the information in the height and width directions again, making and Fusion of global information in height and width allows Preserve the perceptual information in the width direction in the height direction, The perceptual information in the height direction is retained in the width direction to enhance the feature representation capability. and Not just the original and , but an enhanced version; to put it simply, the purpose of this step is to restore the information in the height and width directions while enhancing the feature representation capability.
[0043] Generate attention weights through Sigmoid function and recalibrate and , and the group normalization operation is used to make the eigenvalue distribution more stable, and the enhanced global feature representation is obtained : ; In the formula, Represents a group normalization operation.
[0044] In addition, Enhanced local feature representation using 3D convolution and through Function and matrix multiplication operations generate spatial attention weights , expressed as: ; In the formula, Represents an average pooling operation.
[0045] Finally, through matrix dot multiplication fusion and , and get the image feature representation , expressed as: .
[0046] Using Discrete Fourier Transform to Extract Image Feature Representation The spectral features of the image are embedded in the matrix The spectrum characteristics of , expressed as: ; In the formula, express At a frequency of The spectrum value at ; Represents the image Dimensional feature representation; is the spatial dimension of the image.
[0047] In the frequency domain, the spatial features within different frequency components can be effectively integrated, so that the key information in text and images can be effectively extracted.
[0048] Step S4: Text embedding matrix Spectral features and image embedding matrix The spectral features are compressed.
[0049] After obtaining the spectral representation of text and image, we embed the text matrix and the image embedding matrix The spectrum characteristics of A filter bank is introduced to compress the spectral components of different modes to obtain important features related to rumors, which can be expressed as: ; In the formula, Represents the compressed text embedding matrix and the image embedding matrix The spectrum characteristics of Indicates the compressed , Indicates the compressed ; for Length; is the number of filters; Indicates filters; Used to focus energy more efficiently and strengthen information in the spectrum that is highly relevant to rumor characteristics.
[0050] Step S5: Spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix Enhanced text spectrum feature representation is obtained And enhance the image spectral feature representation : ; ; In the formula, and Respectively represent the mapping functions of images and texts, used to map the original features to the new feature space; represents a one-dimensional convolution operation, Represents an average pooling operation.
[0051] Step S6: Enhance the text spectral feature representation And enhance the image spectral feature representation Perform an inverse discrete Fourier transform (IDFT) to losslessly convert the spectral representation of text and image back to the spatial domain to obtain the spatial domain information of the text And the spatial domain information of the image : ; ; In the formula, Indicates the text content The spatial domain feature representation of each word; Represents the image Dimensional spatial domain feature representation; Indicated in Enhanced text spectral feature representation on frequency components; Indicated in Enhanced image spectral feature representation on frequency components.
[0052] Step S7: Spatial domain information of the text And the spatial domain information of the image Perform weighted summation to obtain the final multimodal representation , input the final multimodal representation into the fully connected layer to make classification predictions. ; ; In the formula, and All are trainable parameters; is the similarity score. As a hyperparameter, it adaptively adjusts the fusion weight of cross-modal features; Represents the predicted label of the rumor, when When it is less than 0.5, the corresponding graphic news data is judged as non-rumor. When it is greater than 0.5, the corresponding graphic news data is judged to be a rumor. When it is equal to 0.5, the corresponding graphic news data will be randomly divided into rumors or non-rumors.
[0053] Among them, the similarity score It is expressed as: ; In the formula, represents the Jensen-Shannon (JS) divergence; Indicates that in the given text content In the case of The posterior probability distribution of ; Indicates that in a given image In the case of The posterior probability distribution of .
[0054] Finally, rumor detection is regarded as a binary classification task, and cross entropy loss is used to optimize the classification performance.
[0055] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multimodal rumor detection method based on data enhancement and global information fusion, characterized in that: The steps include: Step S1: construct a data set, which includes a number of graphic news data; Step S2: Obtaining a text embedding matrix and an image embedding matrix in the graphic news data; Step S3: Get text embedding matrix Spectral features and image embedding matrix Spectral characteristics of Step S4: Text embedding matrix Spectral features and image embedding matrix The spectral characteristics of Step S5: Spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix Enhanced text spectrum feature representation is obtained And enhance the image spectral feature representation ; Step S6: Enhance the text spectral feature representation And enhance the image spectral feature representation Perform inverse discrete Fourier transform to obtain the spatial domain information of the text And the spatial domain information of the image ; Step S7: Spatial domain information of the text And the spatial domain information of the image Perform weighted summation to obtain the final multimodal representation , input the final multimodal representation into the fully connected layer to make classification predictions.
2. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 1 is characterized in that: Given a set of graphic news data , Indicates picture and text news data; Graphic news data , ,in, Represents the text content. Represents an image.
3. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 2 is characterized in that: The specific process of step S2 is: Step S2.1: Define the text content in the graphic news data , Indicates that the length of the text content is words; Use the Word2Vec function to obtain the word embedding representation of each word in the text content , expressed as: ; In the formula, express The word embedding representation of Indicates that the length of the text content is Words, ; , represents the set of real numbers, is the dimension of word embedding representation; express Model; Use sine and cosine functions to analyze text content Each word in the position is encoded to obtain the position encoding representation of each word ; Embed each word in the text content and the positional encoding representation of each word Add element by element to get the text embedding matrix , expressed as: ; ; In the formula, is the embedding representation of each word in the text content, ; Step S2.2: Defining images in graphic news data ; The image Divide into non-overlapping blocks, Indicates the division Blocks, Indicates the line, Represents a column, ; Use two-dimensional convolution to embed the blocks and align them with the text embedding matrix to obtain the block embedding of the image , expressed as: ; In the formula, Represents a two-dimensional convolution operation; Represents a flattening operation; Use sine and cosine functions to transform the image Perform position encoding to obtain image Positional encoding representation ; Embed the image into blocks Positional encoding representation of the image Add element by element to get the image embedding matrix , expressed as: 。 4. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 3 is characterized in that: The specific process of step S3 is: Step S3.1: Adaptively embed the text matrix using the Sigmoid activation function Adjust the weights and compare the weight-adjusted text embedding matrix with the original text embedding matrix Fusion to obtain text feature representation : ; In the formula, Represents the Sigmoid activation function; represents the Hadamard product; Using discrete Fourier transform to extract text feature representation The spectrum characteristics of , that is, the text embedding matrix The spectrum characteristics of , expressed as: ; In the formula, For the text content Text feature representation of words; express At a frequency of The spectrum value at represents the frequency index; represents an imaginary unit; represents the base of natural logarithms; Step S3.2: Use EMA module to embed the image into matrix To optimize, the EMA module first converts Divided along the channel dimension into sub-feature graphs, each of which contains channels, expressed as: ; In the formula, Indicates Sub-feature graphs, and Represents height and width respectively; right Perform average pooling in the horizontal and vertical directions respectively, and get Feature representation after pooling along the height direction as well as Feature representation after pooling along the width direction ; Will and The splicing is performed and the feature fusion is performed through one-dimensional convolution to obtain the feature representation of the fusion height and width direction information. : ; In the formula, Indicates the splicing of features; Afterwards Split back and , generate attention weights through the Sigmoid function and recalibrate and , and processed by group normalization operation to obtain enhanced global feature representation : ; In the formula, represents the group normalization operation; right Enhanced local feature representation using 3D convolution and through Function and matrix multiplication operations generate spatial attention weights , expressed as: ; In the formula, represents the average pooling operation; Fusion through matrix multiplication and , get the image feature representation , expressed as: ; Using Discrete Fourier Transform to Extract Image Feature Representation The spectral features of the image are embedded in the matrix The spectrum characteristics of , expressed as: ; In the formula, express At a frequency of The spectrum value at ; Represents the image Dimensional feature representation; is the spatial dimension of the image.
5. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 4 is characterized in that: The specific process of step S4 is: embedding matrix for text and the image embedding matrix The spectrum characteristics of A filter bank is introduced to compress the spectral components of different modes, which can be expressed as: ; In the formula, Represents the compressed text embedding matrix and the image embedding matrix The spectrum characteristics of Indicates the compressed , Indicates the compressed ; for Length; is the number of filters; Indicates A filter.
6. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 5 is characterized in that: Spectral features of the compressed text embedding matrix and the spectral features of the image embedding matrix Enhanced text spectrum feature representation is obtained And enhance the image spectral feature representation : ; ; In the formula, and Represent the mapping functions of images and texts respectively; represents a one-dimensional convolution operation, Represents an average pooling operation.
7. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 6 is characterized in that: Enhanced Text Spectral Feature Representation And enhance the image spectral feature representation Perform inverse discrete Fourier transform to convert the spectral representation of text and image back to the spatial domain losslessly to obtain the spatial domain information of the text And the spatial domain information of the image : ; ; In the formula, Indicates the text content The spatial domain feature representation of each word; Represents the image Dimensional spatial domain feature representation; Indicated in Enhanced text spectral feature representation on frequency components; Indicated in Enhanced image spectral feature representation on frequency components.
8. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 7 is characterized in that: Spatial domain information of text And the spatial domain information of the image Perform weighted summation to obtain the final multimodal representation , input the final multimodal representation into the fully connected layer To make classification predictions, it is expressed as: ; ; In the formula, and All are trainable parameters; is the similarity score; Represents the predicted label of the rumor, when When it is less than 0.5, the corresponding graphic news data is judged as non-rumor. When it is greater than 0.5, the corresponding graphic news data is judged to be a rumor. When it is equal to 0.5, the corresponding graphic news data will be randomly divided into rumors or non-rumors.
9. The multimodal rumor detection method based on data enhancement and global information fusion according to claim 8, characterized in that: Similarity score It is expressed as: ; In the formula, represents JS divergence; Indicates that in the given text content In the case of The posterior probability distribution of ; Indicates that in a given image In the case of The posterior probability distribution of .
Citation Information
Patent Citations
False news detection method and system based on progressive multi-modal fusion network
CN114528912A
False information detection method and device, equipment and medium
CN114579876A
Multimodal fusion Mongolian rumor detection method based on knowledge perception attention network
CN114925682A
False news early detection method, system, equipment and medium
CN117874607A
False news detection method based on multi-view and hierarchical fusion
CN118114188A