A fake news image detection method based on multimodal data fusion neural network
Through the detection method of multimodal data fusion neural network, combined with visual and physical features, XGBoost is used to identify fake news images, solving the problem of inefficient relying on manual detection in the existing technology, and achieving automated, fast and accurate detection effects.
Patent Information
- Application Number
- CN202210173800.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-02-24
AI Technical Summary
The existing methods of detecting fake news images mainly rely on manual detection, which is inefficient and time-consuming.
The detection method of multimodal data fusion neural network is adopted, including visual modal module, visual feature fusion module, physical feature module and integration module. Through error level analysis, pre-trained ResNet50 feature extraction, discrete cosine transformation and Fourier transform, combined with the physical features of the image and visual modal prediction results, XGBoost is used to identify fake news images.
It effectively avoids the inefficiency and time-consuming and labor-intensive problems of manual detection, and realizes automated, fast and accurate detection of fake news images.
Smart Images

Figure CN114612679B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technology, and in particular to a false news image detection method using a multimodal data fusion neural network. Background Art
[0002] The detection of fake news images plays a vital role in timely curbing the spread of rumors and maintaining national order. Current image detection methods are still unable to comprehensively detect whether news images have been digitally modified and whether they contain false semantic information.
[0003] Existing methods for detecting fake news images are generally done manually, which is not only inefficient but also time-consuming and labor-intensive. Summary of the invention
[0004] In view of the shortcomings of the existing technology, the present invention provides a fake news image detection method based on a multimodal data fusion neural network, which overcomes the shortcomings of the existing technology and aims to solve the problem that the existing false news image detection method is generally performed manually, which is not only inefficient but also time-consuming and labor-intensive.
[0005] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a fake news image detection method of a multimodal data fusion neural network, comprising four modules: a visual modality module, a visual feature fusion module, a physical feature module and an integration module. The visual modality module applies an error level analysis (ELA) algorithm to an input image to obtain the location of the tampered area; uses a pre-trained ResNet50 to extract features from the input image; then divides it into three RGB channels, and then performs discrete cosine transform (DCT) on the three channels respectively, and then combines them to perform Fourier transform on the image to obtain a feature representation domain of the frequency;
[0006] The visual feature fusion module first connects the features in series, performs dimensionality reduction on the extracted features, performs matrix calculation, and decomposes the eigenvalues, and then obtains the projection matrix, and then uses the SoftMax function to obtain the probability distribution of the target space of the true and false images;
[0007] The physical feature module is used to extract the physical features of the image, and the integration module is used to combine the physical features of the image with the visual modality prediction results of the image, and then use XGBoost to identify the final fake news image.
[0008] As a preferred technical solution of the present invention, the visual modality module extracts image features from the pixel domain, frequency domain and gradient histogram of the image, and realizes more comprehensive image detection from the three branches of tampering detection, semantic detection and frequency domain detection.
[0009] As a preferred technical solution of the present invention, the visual feature fusion module includes feature fusion, principal component analysis and SoftMax function. The feature fusion is used to concatenate the features of the image, the principal component analysis is used to perform matrix calculation and eigenvalue decomposition on the image features, and the SoftMax function is used to recognize the image.
[0010] As a preferred technical solution of the present invention, the physical feature module is used to extract physical features of the image, such as the size, clarity, and quantized values of the length and width of the image as physical features for detecting fake news images.
[0011] As a preferred technical solution of the present invention, the integration module is used to collect the physical features of the image extracted by the physical feature module, and combine the physical features of the image with the visual modality prediction result of the image.
[0012] As a preferred technical solution of the present invention, the visual feature fusion module is used to reduce the error rate of positive samples being judged as negative samples during category classification, so that the convolutional neural network can obtain more effective information during training and learning.
[0013] As a preferred technical solution of the present invention, the error level analysis (ELA) algorithm uses a pre-trained residual network (ResNet50) to extract features and obtain the location of the tampered area accordingly.
[0014] As a preferred technical solution of the present invention, the discrete cosine transform (DCT) + Fourier transform is first divided into three RGB channels, and discrete cosine transform (DCT) is performed on the three channels respectively. Finally, the image is combined and Fourier transformed to obtain the frequency feature representation domain, which is then used as the input of ResNet50.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] This application first uses a visual modality module to input an image for feature extraction, then connects the features in series through a visual feature fusion module, reduces the dimension of the extracted features, centralizes matrix calculations and eigenvalue decomposition to obtain a projection matrix, and then uses a SoftMax function to obtain the probability distribution of the target space of true and false images. Next, the physical features of the image are extracted through a physical feature module, and finally input into an integration module. The integration module combines the physical features of the image with the visual modality prediction results of the image and then uses XGBoost to identify the final false news image, thereby avoiding the existing false news image detection method, which is generally performed manually. Manual detection is not only inefficient, but also time-consuming and labor-intensive. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the detection framework of the present invention;
[0018] In the figure: 1. Visual modality module; 2. Visual feature fusion module; 3. Physical feature module; 4. Integration module. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] See also Figure 1 ,A fake news image detection method based on a multimodal data fusion neural network, comprising four modules: a visual modality module 1, a visual feature fusion module 2, a physical feature module 3 and an integration module 4.,In the visual modality module 1, an error level analysis (ELA) algorithm is applied to the input image to obtain the location of the tampered area;,the input image is feature extracted using a pre-trained ResNet50;,it is then divided into three RGB channels, and discrete cosine transform (DCT) is performed on the three channels respectively, and then they are combined to perform Fourier transform on the image to obtain the feature representation domain of the frequency;
[0021] The visual feature fusion module first concatenates the features, performs dimensionality reduction on the extracted features, centralizes matrix calculation, and decomposes the eigenvalues, then obtains the projection matrix, and then uses the SoftMax function to obtain the probability distribution of the target space of the true and false images;
[0022] The physical feature module is used to extract the physical features of the image, and the integration module is used to combine the physical features of the image with the visual modality prediction results of the image, and then use XGBoost to identify the final fake news image.
[0023] A fake news image detection method based on a multimodal data fusion neural network includes four modules: a visual modality module 1, a visual feature fusion module 2, a physical feature module 3 and an integration module 4. The visual modality module 1 applies an error level analysis (ELA) algorithm to an input image to obtain the location of a tampered area; uses a pre-trained ResNet50 to extract features from the input image; then divides it into three RGB channels, and then performs discrete cosine transform (DCT) on the three channels respectively, and then combines them to perform Fourier transform on the image to obtain a feature representation domain of the frequency;
[0024] The visual feature fusion module first concatenates the features, performs dimensionality reduction on the extracted features, centralizes matrix calculation, and decomposes the eigenvalues, then obtains the projection matrix, and then uses the SoftMax function to obtain the probability distribution of the target space of the true and false images;
[0025] The physical feature module is used to extract the physical features of the image, and the integration module is used to combine the physical features of the image with the visual modality prediction results of the image, and then use XGBoost to identify the final fake news image.
[0026] Optionally, the visual modality module 1 extracts image features from the pixel domain, frequency domain and gradient histogram of the image, and implements more comprehensive image detection from the three branches of tampering detection, semantic detection and frequency domain detection.
[0027] Optionally, the visual feature fusion module 2 includes feature fusion, principal component analysis and SoftMax function, wherein feature fusion is used to concatenate features of the image, principal component analysis is used to perform matrix calculation and eigenvalue decomposition on image features, and SoftMax function is used to recognize images.
[0028] Optionally, the physical feature module 3 is used to extract physical features of the image, such as the size, clarity, and quantized values of the length and width of the image as physical features for detecting fake news images.
[0029] Optionally, the integration module 4 is used to collect the physical features of the image extracted by the physical feature module 3, and combine the physical features of the image with the visual modality prediction result of the image.
[0030] Optionally, the visual feature fusion module 2 is used to reduce the error rate of positive samples being judged as negative samples during category classification, so that the convolutional neural network can obtain more effective information during training and learning.
[0031] Optionally, the Error Level Analysis (ELA) algorithm uses a pre-trained residual network (ResNet50) to extract features and obtain the location of the tampered area accordingly.
[0032] Optionally, discrete cosine transform (DCT) + Fourier transform is first divided into three RGB channels, and discrete cosine transform (DCT) is performed on each of the three channels. Finally, the image is combined and Fourier transformed to obtain the frequency feature representation domain, which is then used as the input of ResNet50.
[0033] The detection of fake news images by the visual modality module in the method for detecting fake news images includes the following steps:
[0034] In the visual modality module, in order to comprehensively detect fake news images, we designed three branch networks to respectively realize image tampering detection, semantic detection, and frequency domain detection.
[0035] In the tampering detection part, we first applied the Error Level Analysis (ELA) algorithm to the input image. The main idea of ELA is that the tampered area of the image is significantly different from the tampered area of the original image after compression with a fixed quality, and the location of the tampered area is obtained accordingly. For the image processed by ELA, we use the pre-trained ResNet50 to extract features. In addition, we add a fully connected layer of 2048 neurons to ResNet50 to obtain the representation as:
[0036] F t =[v1 t ,v2 t ,v3 t ,...,v 2048 t ] T (1)
[0037] The eigenvector of
[0038] F t =ReLU(Wv f Ft p ) (2)
[0039] in are visual features obtained from a pre-trained ResNet50 network, represents the weight of the fully connected layer and ReLU represents the ReLU activation function.
[0040] (2) In the semantic detection part, we directly use the pre-trained ResNet50 to extract features from the input image. The network structure of the semantic detection part is the same as that of the tampering detection part. The extracted features can be expressed as:
[0041] F s =[v1 s ,v2 s ,v3 s ,...,v 2048 s ] T (3)
[0042] In the frequency domain detection part, considering that the heavy compression of the image can be well reflected in the frequency domain, we first divide it into three RGB channels, and then perform discrete cosine transform (DCT) in the three channels respectively, and then combine them to perform Fourier transform on the image to obtain the feature representation of the frequency domain. This feature is then used as the input of ResNet50. The structure of ResNet50 in the frequency domain detection part is the same as that of the tampering detection and semantic detection parts. The extracted features can be expressed as:
[0043] F f =[v1 f ,v2 f ,v3 f ,...,v 2048 f ] T (4)
[0044] The visual feature fusion module in the method for detecting fake news images detects fake news images, including the following steps:
[0045] In the visual feature fusion module, for
[0046] F t =[v1 t ,v2 t ,v3 t ,...,v 2048 t ] T , (5)
[0047] F s =[v1 s ,v2 s ,v3 s ,...,v 2048 s ] T , (6)
[0048] F f =[v1 f ,v2 f ,v3 f ,...,v 2048 f ] T , (7)
[0049] We first concatenate the features to get a new feature vector:
[0050] F c =[F t ,F s ,F f ] T (8)
[0051] Considering the high feature dimension after extraction, we use principal component analysis (PCA) to reduce the dimension of the extracted features and map the 6144-dimensional feature vector to a 1024-dimensional feature vector to obtain a more compact feature. We first calculate its centralization matrix as:
[0052]
[0053] Then the covariance matrix of V is calculated, and then the eigenvalue decomposition is performed to obtain the eigenvalue λ and eigenvector U as:
[0054] C==VV T (10)
[0055] C=UλU T (11)
[0056] Sort λ as λ1>λ2>λi>…>λn, and take the first d′ eigenvector corresponding to the sorted eigenvalue to obtain the projection matrix W=[u1,u2,…,U0], where d′ is the dimension after dimensionality reduction, which is determined by the cumulative contribution rate of each component:
[0057]
[0058] Using the obtained matrix W, the projection is calculated as:
[0059] F p =V*W (13)
[0060] In order to obtain a high-level representation of the Fp input image, a fully connected layer with SoftMax activation is used to project the vector into the target space of fake news images and real news images, and its probability distribution function is obtained:
[0061] p = Softmax(W c F p +b c ) (14)
[0062] Where p is the probability that the image is identified as a real news image, v1 t 、v1 s 、v1 f The neurons representing tampering detection, semantic detection, and frequency domain detection respectively.
[0063] The physical feature module in the method for detecting fake news images detects fake news images, including the following steps:
[0064] In the physical feature module, we noticed that fake news images usually have a longer propagation time than real news images, that is, more recompression time. Therefore, we use the image file size, image length, image width and image clarity quantization value as physical features for detecting fake news images, that is, Fp = [p1, p2, p3, p4], where p1 represents the length of the image, p2 represents the width of the image, p3 represents the file size of the image, and p4 represents the quantization value of the image. In order to calculate the quantization value of the image, we first convolve the original image with a 3×3 Laplace operator, and then use the variance of the convolution operation result as the quantization value image of the image.
[0065] The integrated learning module in the method for detecting fake news images includes the following steps for detecting fake news images:
[0066] In the ensemble learning module, we combine the physical features of the image and the visual modality prediction results Fa = [p, p1, p2, p3, p4] of the image, and then use XGBoost to identify the final fake news images.
[0067] Working principle: First, the image is input through the visual modality module for feature extraction, and then the features are connected in series through the visual feature fusion module, the extracted features are reduced in dimension, and the projection matrix is obtained by centralized matrix calculation and eigenvalue decomposition. The SoftMax function is then used to obtain the probability distribution of the target space of true and false images. Next, the physical features of the image are extracted through the physical feature module and finally input into the integration module. The integration module combines the physical features of the image with the visual modality prediction results of the image and then uses XGBoost to identify the final false news image, thus avoiding the existing false news image detection method, which is generally performed manually. Manual detection is not only inefficient, but also time-consuming and labor-intensive.
[0068] Finally, it should be noted that in the description of the present invention, it should be noted that the terms "vertical", "up", "down", "horizontal", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention.
[0069] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0070] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting fake news images using a multimodal data fusion neural network, comprising four modules: a visual modality module (1), a visual feature fusion module (2), a physical feature module (3) and an integration module (4), characterized in that: The visual modality module (1) applies an error level analysis (ELA) algorithm to the input image to obtain the location of the tampered area; uses a pre-trained ResNet50 to extract features from the input image; then divides it into three RGB channels, and then performs discrete cosine transform (DCT) on the three channels respectively, and then combines them to perform Fourier transform on the image to obtain a frequency feature representation domain; The visual feature fusion module first connects the features in series, performs dimensionality reduction on the extracted features, performs matrix calculation, and decomposes the eigenvalues, and then obtains the projection matrix, and then uses the SoftMax function to obtain the probability distribution of the target space of the true and false images; The physical feature module is used to extract the physical features of the image, and use the size of the image file, the length of the image, the width of the image, and the quantitative value of the image clarity as the physical features for detecting fake news images. The integration module is used to combine the physical features of the image and the visual modality prediction results of the image, and then use XGBoost to identify the final fake news image.
2. The method for detecting fake news images using a multimodal data fusion neural network according to claim 1 is characterized in that: The visual modality module (1) extracts image features from the pixel domain, frequency domain and gradient histogram of the image, and implements image detection from three branches: tampering detection, semantic detection and frequency domain detection.
3. The method for detecting fake news images using a multimodal data fusion neural network according to claim 1, characterized in that: The visual feature fusion module (2) includes feature fusion, principal component analysis and SoftMax function, wherein the feature fusion is used to connect the features of the picture in series, the principal component analysis is used to perform matrix calculation and eigenvalue decomposition on the picture features, and the SoftMax function is used to recognize the image.
4. The method for detecting fake news images using a multimodal data fusion neural network according to claim 1, characterized in that: The integration module (4) is used to collect the physical features of the image extracted by the physical feature module (3), and combine the physical features of the image with the visual modality prediction result of the image.
5. The method for detecting fake news images using a multimodal data fusion neural network according to claim 1, characterized in that: The visual feature fusion module (2) is used to reduce the error rate of positive samples being judged as negative samples during category classification.
6. The method for detecting fake news images using a multimodal data fusion neural network according to claim 1, characterized in that: The Error Level Analysis (ELA) algorithm uses a pre-trained residual network (ResNet50) to extract features and obtain the location of the tampered area accordingly.
7. The method for detecting fake news images using a multimodal data fusion neural network according to claim 1, characterized in that: The discrete cosine transform (DCT) + Fourier transform is first divided into three RGB channels, and discrete cosine transform (DCT) is performed on the three channels respectively. Finally, the image is combined and Fourier transformed to obtain the frequency feature representation domain, which is then used as the input of ResNet50.
Citation Information
Patent Citations
News image detection method, system and device based on multi-domain visual features
CN110889430A
False news detection method and system fusing multi-scale visual information
CN111797326A