False image detection method based on contrast learning
By adopting a method based on contrast learning in false image detection, a false image discrimination model including frequency domain analysis and noise extraction modules is constructed, which solves the problem of insufficient generalization ability of false image recognition models in the prior art, and realizes efficient recognition of existing and new false images.
Patent Information
- Application Number
- CN202510182195.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively identify existing false images and brand new false images that have never been seen before, and the generalization ability of false image recognition models is insufficient.
Using a false image detection method based on contrast learning, a false image discrimination model is constructed, which includes a frequency domain analysis module, a noise extraction module, a fusion module, a feature extraction module and a classifier. The model is trained using contrast learning methods to improve its generalization ability.
It achieves a high recognition rate of existing false images and achieves a good recognition effect on new false images, which significantly improves the generalization ability of false image recognition models.
Smart Images

Figure CN120107679A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image detection, and in particular relates to a false image detection method based on contrast learning. Background Art
[0002] In recent years, image generation models have made significant progress and can be divided into two categories: methods based on generative adversarial networks (GANs) and methods based on diffusion models. Since Goodfellow et al. proposed GAN in 2014, this technology of generating realistic images by pitting two neural networks against each other has developed rapidly and has been widely used in art creation, data enhancement and other fields. However, with the evolution of technology, especially the rise of methods based on diffusion models (such as DDPMs, Denoising Diffusion Probabilistic Models) in recent years, the quality and diversity of image generation have reached new heights. These models generate images by gradually denoising, which can produce more delicate and natural visual effects.
[0003] Although image generation technology has brought revolutionary changes to many fields, its abuse has also caused serious social problems. For example, deepfake technology can easily transplant a person's facial features onto another subject and be used to create illegal activities such as false news, defamation or fraud, which seriously undermines the authenticity and trust of information. In addition, false content generated using these technologies may also affect public opinion.
[0004] Faced with this challenge, researchers have proposed a variety of approaches from both active and passive perspectives.
[0005] From the perspective of active prevention, the idea is to add disturbance noise or seamless watermark to the original image so that it cannot be used as valid data for the generative model without changing the image content and quality. Once it is passed into the generative model, it will only produce erroneous images, thereby avoiding the generation of false images.
[0006] From the perspective of passive detection, early researchers proposed some detection techniques based on traditional methods, such as analyzing inconsistencies or abnormal patterns in images, including lighting conditions, geometric distortion, etc.; in recent years, machine learning, especially deep learning methods, have been used to identify features generated by specific generative models. For example, classifiers are trained to distinguish between real images and images generated by specific generative models. Some work has also explored combining blockchain technology to verify the authenticity of the source of images, ensuring that the entire process from image creation to release is transparent and traceable.
[0007] At the same time, with the continuous advancement of generative model technology and the continuous iterative output of new models, how to ensure that the fake image recognition model has strong generalization capabilities has become a key challenge. Summary of the invention
[0008] Purpose of the invention: The technical problem to be solved by the present invention is to provide a false image detection method based on contrastive learning in view of the shortcomings of the prior art. The method can achieve a high recognition rate for existing false images and also achieve a good recognition effect for completely new and never-before-seen false images.
[0009] In order to solve the above technical problems, the present invention discloses a false image detection method based on contrastive learning, comprising the following steps:
[0010] Step 1: Collect real images and virtual images to build an image dataset;
[0011] Step 2: construct a false image discrimination model, wherein the false image discrimination model is used to discriminate whether an image is a real image or a false image;
[0012] Step 3, using the image data set to train the false image discrimination model to obtain a trained false image discrimination model;
[0013] Step 4: Input the image to be detected into the trained false image discrimination model to obtain the detection result.
[0014] Further, step 1 comprises:
[0015] Step 1-1, using the collected real image and the false image as the input of the generative model to obtain a real reconstructed image and a false reconstructed image, wherein the image category of the collected real image is true, and the image categories of the collected false image, the real reconstructed image and the false reconstructed image are false;
[0016] Step 1-2, data enhancement is performed on the obtained four types of images, each image after data enhancement is segmented and reorganized in a shuffled order to obtain an enhanced image with destroyed semantic information, and the enhanced images with destroyed semantic information constitute an image data set.
[0017] By performing data augmentation on the four types of images, the fake image discrimination model trained subsequently can be made more generalized.
[0018] Furthermore, the false image discrimination model in step 2 includes a frequency domain analysis module, a noise extraction module, a fusion module, a feature extraction module and a classifier, wherein the frequency domain analysis module is used to perform frequency domain analysis on the input image to obtain frequency domain analysis features;
[0019] The noise extraction module is used to extract the noise of the input image and obtain the noise characteristics;
[0020] The fusion module is used to perform feature splicing and fusion on the input image, frequency domain analysis features and noise features to obtain frequency domain noise features;
[0021] The feature extraction module is used to extract the frequency domain noise features to obtain image features;
[0022] The classifier is used to calculate the probability value of the image category according to the image features and determine the category of the input image.
[0023] By introducing the frequency domain analysis module and the noise extraction module, the false image discrimination model can obtain lower-level image features and more accurately judge the authenticity of the image.
[0024] Further, step 3 includes:
[0025] Step 3-1, inputting the enhanced image with destroyed semantic information into the frequency domain analysis module and the noise extraction module to obtain processed frequency domain analysis features and noise features, inputting the enhanced image with destroyed semantic information, the frequency domain analysis features and the noise features into the fusion module for feature splicing and fusion operation to obtain frequency domain noise features;
[0026] Step 3-2, input the frequency domain noise features into the feature extraction module to obtain the final image features;
[0027] Step 3-3, perform comparative learning based on image features and their image categories, narrow the distance between similar features, and make features of different categories farther away, and obtain feature loss values;
[0028] Step 3-4: Input the image features into the classifier to obtain the probability value of the category to which the image belongs, and calculate the cross entropy loss value based on the category to which it belongs;
[0029] Step 3-5, use the feature loss value and the cross entropy loss value to perform back propagation to calculate the gradient of the model parameters, and use the gradient descent algorithm to update the model parameters;
[0030] Step 3-6, repeat step 1-2, step 3-1 to step 3-5, until the number of training times reaches the predetermined number of iterations, and the final trained false image discrimination model is obtained.
[0031] Furthermore, step 2 includes:
[0032] Steps 1-2 include:
[0033] Step 1-2-1, each of the four types of images is denoted as I, and image enhancement technology is used to enhance the representativeness of image I to obtain a preliminary preprocessed image. The enhancement technology used is applied to image I under the condition of probability p, including image compression, horizontal flipping, adding Gaussian noise or blurring, cropping to a specified image size after enhancement, and standardizing the pixel value, where 0 <p≤1;
[0034] Step 1-2-2: split each pre-processed image obtained in step 1-2-1 into small images and reorganize them in a disordered order to obtain an enhanced image I that destroys the semantic information of the image. en ; H is the height of the original image, W is the width of the original image, h is the height of the divided small image, and w is the width of the divided small image.
[0035] Furthermore, the frequency domain analysis features obtained in step 3-1 include:
[0036] Step 3-1-1: Enhance the image I en Input into the frequency domain analysis module, first perform Fourier transform on the image to obtain the enhanced image I en The frequency domain characteristics x freq ;
[0037] Step 3-1-2, from the frequency domain feature x freq The amplitude characteristics x are obtained in abs and phase characteristics x angle ;
[0038] Step 3-1-3, for the amplitude feature x abs and phase characteristics x angle Update to extract the unique high- and low-frequency information and represent it through amplitude and phase features:
[0039] x abs =process abs (x abs )=ReLU(BN(RRG(x abs )))
[0040] x angle =process angle (x angle )=ReLU(BN(RRG(x angle )))
[0041] Among them, RRG stands for Recursive Residual Group, which is composed of multiple dual attention blocks connected in series and combined with skip connections; BN stands for Batch Normalization; ReLU stands for ReLU activation function;
[0042] Step 3-1-4, for the amplitude feature x obtained in step 3-1-3 abs and phase characteristics x angle Perform feature splicing and fuse the spliced features to obtain the updated amplitude feature x abs and phase characteristics x angle :
[0043] x temp =concat(x abs ,x angle , dim=1)
[0044] x abs =merge abs (x temp )
[0045] x angle =merge angle (x temp )
[0046] concat means concatenation operation, dim=1 means concatenation operation on 1 dimension of feature, merge abs and merge angle It is a fusion operation, both of which are 1×1 convolution modules; this involves the amplitude feature x abs and phase characteristics x abs The filtering and updating are used to obtain two new features as the input of the subsequent modules. The purpose is to better integrate the amplitude and phase features, promote the information interaction between the two, enable the model to learn more complex patterns, better cope with noise interference or data distribution changes, and improve the generalization ability of the model. Secondly, it allows the model to readjust the information weights of amplitude and phase according to the fused features during training, so that the updated features are more suitable for subsequent tasks.
[0047] Step 3-1-5, repeat steps 3-1-3 to 3-1-4 n times, and finally obtain the updated amplitude feature x abs and phase characteristics x angle; Where n ≥ 1; By repeatedly stacking multiple submodules, the model can gradually extract deeper feature representations, thereby better capturing complex patterns in the data. The ReLU nonlinear activation function introduced in step 3-1-3 can significantly increase the expressive power of the model after stacking. Each layer of submodules will perform feature updates based on the output of the previous layer, thereby gradually optimizing the final feature representation;
[0048] Step 3-1-6, the amplitude characteristic x obtained in step 3-1-5 abs and phase characteristics x angle Splice and perform inverse Fourier transform to obtain the extracted frequency domain analysis feature x freq-anal :
[0049]
[0050] Among them, real is to extract the real part of the complex number, ifft is the inverse Fourier transform operation, M conv is a 3×3 convolution operation, M merge1 Represents a fusion operation, including a convolution module of size 3×3 and stride 1, a 2D batch normalization module, and a ReLU activation module.
[0051] This step uses the inverse Fourier transform to switch back to the image space after the previous steps are completed in order to keep the feature distribution consistent with the original image. Switching back to the image space does not lose the extracted frequency domain information, but better retains these features in the image space to facilitate subsequent feature extraction.
[0052] Furthermore, the noise characteristics and frequency domain noise characteristics obtained in step 3-1 include:
[0053] Step 3-1-7, enhance the image I en Input to the noise extraction module M denoise And obtain the noise characteristic I noise :
[0054] I noise =I en -M denoise (I en )
[0055] Step 3-1-8, by enhancing the image I en , frequency domain analysis characteristics x freq-anal and noise characteristics I noise Perform feature splicing and fusion to obtain the frequency noise feature x fn :
[0056] x fn =M merge2 (concat(I en ,xfreq-anal ,I noise ,dim=1))
[0057] Among them, M merge2 It consists of four groups of 3×3 convolution modules, normalization modules and ReLU modules connected in series.
[0058] Here we will enhance the image I en , frequency domain analysis characteristics x freq-anal and noise characteristics I noise Splicing in the channel dimension can integrate information from different sources and improve feature representation capabilities. It can simultaneously utilize multiple information related to the spatial domain, frequency domain, and noise to avoid feature loss caused by independent processing. After integrating these features, the model is more robust to noise interference or changes in data distribution, and the generalization performance of the model is also improved;
[0059] Further, step 3-3 includes: forming a pair of any two image features from the multiple image features obtained in step 3-2, and calculating the feature loss value L by using a contrastive learning method. contrastive :
[0060]
[0061] Among them, D W (i) represents the Euclidean distance between a pair of image features, N represents the number of image feature pairs, i represents the image feature pair index, Y represents whether the pair of image features belong to the same category, and m is the specified minimum distance between features of different categories.
[0062] Further, step 3-4 includes: the image feature obtained in step 3-2 is denoted as x feature , the image feature x feature Input to classifier M classifier In the output, the predicted probability of the corresponding category is calculated and the cross entropy loss L is calculated. cross-entropy :
[0063] logits = M classifier (x feature )
[0064] probs=softmax(logits,dim=-1)
[0065]
[0066] Among them, logits is the score output by the classifier, and softmax calculates the probability probs of each category on the last dimension based on logits to obtain the predicted category; p jis the probability value that the jth sample image in the image dataset belongs to the predicted category, and M represents the total number of samples in the image dataset.
[0067] Further, steps 3-5 include: combining the cross entropy loss L cross-entropy And feature loss L contrastive Perform back propagation calculations and use the gradient descent algorithm to optimize the false image discrimination model parameters:
[0068] L total =λL contrastive +(1-λ)L cross-entropy
[0069]
[0070] Among them, L total represents the loss function of the fake image discrimination model, λ is a hyperparameter that controls the ratio between the two loss values, θ model are the parameters of the false image discrimination model, and lr is the learning rate that controls the update speed of the false image discrimination model parameters.
[0071] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: (1) high recognition accuracy. The frequency domain analysis module, noise extraction module, etc. designed and introduced by the present invention can enable the model to obtain lower-level image features and more accurately judge the authenticity of the image; (2) the model method is simple. By adding a frequency domain analysis module and a noise extraction module to any existing backbone network and using a comparative learning method, the recognition ability of the model can be greatly improved; (3) the method has strong universality. Since the present invention mines and analyzes the underlying features of the image, it can have a good recognition effect on false images that have been seen or not seen during the model training process, and the application scenarios are more extensive and the universality is stronger. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0073] Figure 1 This is a flow chart of a false image detection method based on contrast learning provided in an embodiment of the present application.
[0074] Figure 2 It is a flowchart of training a false image discrimination model in a false image detection method based on contrastive learning provided in an embodiment of the present application.
[0075] Figure 3 It is a schematic diagram of a contrastive learning method used in a false image detection method based on contrastive learning provided in an embodiment of the present application.
[0076] Figure 4 This is a false picture sample identified by a false image detection method based on contrastive learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0077] The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0078] With the continuous advancement of generative model technology and the continuous iteration of new models, how to ensure that fake image recognition models have strong generalization capabilities has become a key challenge. This means that the recognition model must not only be able to accurately detect known types of generated images, but also be able to recognize images that have never been seen, including existing but unseen images and those that may appear in the future. From this perspective, there are two major solutions. One is to continuously update the model through continuous learning or incremental learning frameworks to keep up with the generation algorithm, but there are problems such as the possibility of forgetting past knowledge and excessive continuous training overhead; the other is to dig deep into the underlying features of a type of generation algorithm to find the optimal recognition model that is suitable for this type of algorithm. Regardless of how the generation model is iteratively optimized, the generation model based on this underlying generation algorithm can be effectively identified. The problem with training this type of model is how to mine the most generalized feature patterns of these images.
[0079] As an important machine learning paradigm, contrastive learning has made significant progress in the field of unsupervised and self-supervised learning in recent years. Its core idea is to learn useful feature representations by comparing the similarities and differences between data samples. With the rise of deep learning in recent years, especially the successful application of convolutional neural networks (CNNs) in image processing, contrastive learning has begun to receive more attention. This type of method can train the model without manual annotation, greatly reducing the dependence on large-scale annotated data sets, making it very suitable for scenarios where it is difficult to obtain high-quality annotated data. And by utilizing a large amount of unlabeled data, contrastive learning can help the model learn more representative feature expressions, thereby improving the generalization ability and robustness of the model. Now, contrastive learning methods are not only applicable to image recognition tasks, but can also be extended to multiple fields such as natural language processing and speech recognition, and even combined with multiple fields. For example, the CLIP (Contrastive Language-Image Pre-Training) model uses contrastive learning to enable the model to learn accurate features of both text and image modalities at the same time, showing strong cross-domain adaptability. Some advanced contrastive learning frameworks, such as MoCo, have improved the training efficiency and convergence speed by improving the selection of negative samples and the encoder update mechanism, further promoting its popularity in practical applications. With the deepening of research and the development of technology, it is expected that more innovative algorithms and applications based on contrastive learning will appear in the future.
[0080] like Figure 1 As shown, the embodiment of the present application discloses a false image detection method based on contrastive learning, comprising the following steps:
[0081] Step 1: Collect real images and virtual images to build an image dataset, including:
[0082] Step 1-1 includes: collecting real images and false images, real and false image set I fake On the other hand, we use a generative model Get the reconstructed image set I of the corresponding category real_rec and I fake_rec The image category of the collected real image is true, and the image category of the collected false image, true reconstructed image and false reconstructed image is false.
[0083] Step 1-2, data enhancement is performed on the obtained four types of images, each image after data enhancement is segmented and reorganized in a shuffled order to obtain an enhanced image with destroyed semantic information, wherein the enhanced images with destroyed semantic information constitute an image data set, including:
[0084] Step 1-2-1: Denote each image in the four types of images obtained as I. Use image enhancement techniques to enhance the representativeness of image I to obtain a preliminarily preprocessed image. For the enhancement techniques used, they are all applied to image I under the condition of probability p, including image compression, horizontal flipping, adding Gaussian noise or blurring. After enhancement, it is cropped to the specified image size, and the pixel values are normalized, where 0 < p ≤ 1. In the specific implementation process, the default value is set to 0.5;
[0085] Step 1-2-2: Cut each preprocessed image obtained in Step 1-2-1 into small images, shuffle the order and reorganize them to obtain an enhanced image I that destroys the semantic information of the image en ; H is the height of the original image, W is the width of the original image, h is the height of the divided small image, and w is the width of the divided small image.
[0086] Step 2: Construct a fake image discrimination model, which is used to discriminate whether an image is a real image or a fake image;
[0087] The fake image discrimination model includes a frequency domain analysis module, a noise extraction module, a fusion module, a feature extraction module, and a classifier. The frequency domain analysis module is used to perform frequency domain analysis on the input image to obtain frequency domain analysis features;
[0088] The noise extraction module is used to extract the noise of the input image to obtain noise features;
[0089] The fusion module is used to perform feature splicing and fusion on the input image, frequency domain analysis features, and noise features to obtain frequency domain noise features;
[0090] The feature extraction module is used to perform feature extraction on the frequency domain noise features to obtain image features;
[0091] The classifier is used to calculate the probability value of the image category according to the image features and discriminate the input image category.
[0092] Step 3: Use the image dataset to train the fake image discrimination model to obtain a trained fake image discrimination model, as Figure 2 shown, including:
[0093] Step 3-1: Input the enhanced image that destroys the semantic information into the frequency domain analysis module and the noise extraction module to obtain the processed frequency domain analysis features and noise features. Input the enhanced image that destroys the semantic information, frequency domain analysis features, and noise features into the fusion module for feature splicing and fusion operations to obtain frequency domain noise features; Specifically, it includes the following steps:
[0094] Step 3-1-1: Enhance the image I en Input into the frequency domain analysis module, first perform Fourier transform on the image to obtain the enhanced image I en The frequency domain characteristics x freq :
[0095]
[0096] Among them, u, v represent the horizontal and vertical coordinates of the pixel points in the frequency domain feature image, fft represents Fourier transform, and x and y represent the enhanced image I en The horizontal and vertical coordinates of the pixel point in the middle, j represents the imaginary unit, and the obtained frequency domain feature x freq Its value is plural;
[0097] Step 3-1-2, from the frequency domain feature x freq The amplitude characteristics x are obtained in abs and phase characteristics x angle :
[0098]
[0099]
[0100] Where a and b represent the frequency domain features x freq The real and imaginary parts of the ,σ is a very small positive value, in the specific implementation process, it can be taken as 10 -9 ;
[0101] Step 3-1-3, for the amplitude feature x abs and phase characteristics x angle To update:
[0102] x abs =process abs (x abs )=ReLU(BN(RRG(x abs )))
[0103] x angle =process angle (x angle )=ReLU(BN(RRG(x angle )))
[0104] Among them, RRG represents the recursive residual group network, which is composed of multiple dual attention modules in series and combined with skip connections; BN represents batch normalization; ReLU represents the ReLU activation function;
[0105] Step 3-1-4, for the amplitude feature x obtained in step 3-1-3 abs and phase characteristics x anglePerform feature splicing and fuse the spliced features to obtain the updated amplitude feature x abs and phase characteristics x angle :
[0106] x temp =concat(x abs ,x angle , dim=1)
[0107] x abs =merge abs (x temp )
[0108] x angle =merge angle (x temp )
[0109] concat means concatenation operation, dim=1 means concatenation operation on 1 dimension of feature, merge abs and merge angle It is a fusion operation, both of which are 1×1 convolution modules.
[0110] Step 3-1-5, repeat the above steps 3-1-3 to 3-1-4 n times, and finally obtain the updated amplitude feature x abs and phase characteristics x angle Where n≥1, and n can be set to 5 in a specific implementation.
[0111] Step 3-1-6, the amplitude characteristic x obtained in step 3-1-5 abs and phase characteristics x angle Splice and perform inverse Fourier transform to obtain the extracted frequency domain analysis feature x freq-anal :
[0112] x freq-anal =M merge1 (concat(M conv (I en ),real(ifft(x abs ,x angle )),dim=1))
[0113] Among them, real is to extract the real part of the complex number, ifft is the inverse Fourier transform operation, M conv is a 3×3 convolution operation, M merge1 Represents a fusion operation, including a convolution module of size 3×3 and stride 1, a 2D batch normalization module, and a ReLU activation module.
[0114] Step 3-1-7, enhance the image I enInput to the noise extraction module M denoise And obtain the noise characteristic I noise :
[0115] I noise =I en -M denoise (I en )
[0116] Step 3-1-8, by enhancing the image I en , frequency domain analysis characteristics x fneq-anal and noise characteristics I noise Perform feature splicing and fusion to obtain the frequency noise feature x fn :
[0117] x fn =M merge2 (concat(I en ,x req-anal ,I noise ,dim=1))
[0118] Among them, M merge2 It consists of four groups of 3×3 convolution modules, normalization modules and ReLU modules connected in series.
[0119] Step 3-2: transform the frequency domain noise feature x fn Input to the feature extraction module to obtain the final image feature x feature :
[0120] x feature =M backbone (x fn )
[0121] Among them, M backbone It is a designated general feature extraction backbone network. In the specific implementation process, classic visual model architectures such as ResNet, ViT, etc., or Conv-B's ConvNeXt, etc. can be used, as long as the final output is a feature vector. In the specific implementation process of this application, Conv-B's ConvNeXt is used.
[0122] Step 3-3, conduct comparative learning based on image features and their image categories, narrow the distance between similar features and keep features of different categories away, such as Figure 3 As shown, obtaining the feature loss value includes: combining any two of the multiple image features obtained in step 3-2 into a pair, and calculating the feature loss value L using a contrastive learning method contrastive :
[0123]
[0124] Among them, DW (i) represents the Euclidean distance between a pair of image features, N represents the number of image feature pairs, i represents the image feature pair index, Y represents whether the pair of image features belongs to the same category, Y=1 represents the same category, Y=0 represents different categories, m is the specified minimum distance between features of different categories, and in the specific implementation process, m can take the value of 1, 2 or other values.
[0125] Step 3-4: Input the image features into the classifier to obtain the probability value of the category to which the image belongs, and calculate the cross entropy loss value based on the category to which it belongs, including: feature Input to classifier M classifier In the output, the predicted probability of the corresponding category is calculated and the cross entropy loss L is calculated. cross-entropy :
[0126] logits = M classifier (x feature )
[0127] probs=softmax(logits,dim=-1)
[0128]
[0129] Among them, logits is the score output by the classifier, and softmax calculates the probability probs of each category on the last dimension based on logits to obtain the predicted category; p j is the probability value that the jth sample image in the image dataset belongs to the predicted category, and M represents the total number of samples in the image dataset. In the specific implementation process, the classifier can use a simple linear layer for classification, or a multi-layer perceptron or other complex classifier to map image features to the probability of each category.
[0130] Step 3-5, use the feature loss value and the cross entropy loss value to perform back propagation to calculate the gradient of the model parameters, and use the gradient descent algorithm to update the model parameters, including: combining the cross entropy loss L cross-entropy And feature loss L contrastive Perform back propagation calculations and use the gradient descent algorithm to optimize the false image discrimination model parameters:
[0131] L total =λL contrastive +(1-λ)L cross-entropy
[0132]
[0133] Among them, L total represents the loss function of the fake image discrimination model, λ is a hyperparameter that controls the ratio between the two loss values, θmodel are the parameters of the false image discrimination model, and lr is the learning rate that controls the update speed of the false image discrimination model parameters.
[0134] Step 3-6, repeat step 1-2, step 3-1 to step 3-5, until the number of training times reaches the predetermined number of iterations, and the final trained false image discrimination model is obtained.
[0135] Step 4: Input the image to be detected into the trained false image discrimination model to obtain the detection result.
[0136] Example:
[0137] In order to verify the effectiveness of the embodiments of the present application, pictures in the test set of the DRCT-2M data set are selected as the experimental data set. These data contain 5,000 real pictures and image data generated by multiple different generative models. There are 5,000 false pictures generated by the generative model in each category. The real picture set is shared during the test, and a total of 70,000 pictures are involved in the test.
[0138] The compared model methods are previous discriminative models or frameworks of the same type, including CNNSpot, F3Net, CLIP / RN50, GramNet, De-fake, Conv-B, UnivFD, DIRE and DRCT / Conv-B. Among the compared models, CNNSpot and CLIP / RN50 are both based on the classic ResNet model, using it as the backbone network; F3Net is based on ResNet, adds a large number of CFM modules and performs multi-dimensional feature fusion operations, which greatly increases the model's reasoning time; GramNet is also based on ResNet and adds Gram Block as a bifurcated network to extract the texture features of the image, which involves a large number of Gram Matrix calculation processes; De-fake needs to obtain the description text for the input image before reasoning, so that the model needs to do more preparation before reasoning; DIRE not only uses reconstructed images during training, but also needs to use a generative model to generate reconstructed images during reasoning, which has great overhead in terms of time and space; DRCT and this method both use Conv-B's ConvNeXt as the backbone network. ConvNeXt is a further optimization of the ResNet network, which reduces the model parameters and other additional operations used in it, and has a smaller computational overhead. The evaluation indicator used is the classification accuracy, that is, the ratio of the number of correct predicted categories to the total number of test images.
[0139] The first row of Table 1 represents different false image recognition models, and the first column represents different generative models. From the experimental data in the table, it can be found that the embodiment of the present application performs better in the recognition of most false pictures. Specifically, all the methods listed in the table are obtained by training experiments on false pictures based on SD1.4. Each method can achieve an accurate recognition rate of more than 99%. This method also has an accuracy rate of 98%, with a 1% performance degradation. In the recognition of false pictures obtained by other SD1.4 generation methods, that is, the first five rows of data in the table, this method has an accuracy rate of more than 98% in the recognition of pictures such as SDXL. The recognition effects of other methods on these pictures remain unchanged or decline. In the recognition of SDXL pictures, this method also achieves the second most accurate recognition accuracy, with a recognition effect of about 86%. All methods still maintain a high accuracy rate in the recognition of pictures similar to the training set SD1.4.
[0140] However, when it comes to pictures generated by models with more different implementation methods, the method of the present invention shows good stability and generalization. Except for the recognition rate of LCM SDXL pictures, which only reached 82.2%, the accuracy of the remaining false pictures can be maintained at more than 90%, while other methods have shown a great degree of effect degradation. Finally, the average recognition rate of this method on images generated by 13 types of generative models reached 95%, which is much higher than other methods, and is still 4% higher than the second place. It also shows that this method makes full use of the frequency and noise information in the image, and can extract more unique and easy-to-recognize image features in combination with contrastive learning methods. It has stronger generalization performance and can still maintain strong recognition capabilities on more unseen pictures. Compared with other methods, the model obtained by this method does not require additional preprocessing operations in the actual reasoning process, such as calling other generative models to extract reconstruction features or reconstruction errors, etc., which significantly speeds up the model recognition speed and saves the video memory space required for model recognition.
[0141] Table 1
[0142]
[0143]
[0144] Figure 4Some examples of false images successfully identified using the method provided by the embodiment of the present application are shown. These images are all generated by the SDXL-Refiner model, which is generated by inputting the descriptive text of the image to be generated into the model. The generated images involve different contents, ranging from food, people to general scenery, which can be effectively identified, and are not limited to false image detection of a certain type of content. Generally speaking, these generated high-quality pictures are not easy to identify directly by the naked eye, but some unreasonable aspects can still be seen by analyzing the shallow details, such as the bread surface in the first picture is too perfect and lacks some defects in the real scene; the missing of the human hand in the second picture; the unreasonable environment of the fire hydrant in the third picture and the asymmetry of its own structure; the illusion of overlap of the sculpture legs in the fourth picture, etc. With the continuous development of generative model technology, such shallow detail errors will be more difficult to detect, and detection models are needed to explore the underlying feature patterns of the image for analysis and discrimination.
[0145] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, the invention content of a false image detection method based on contrast learning provided by the present invention and some or all of the steps in each embodiment can be executed. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0146] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of computer programs and their corresponding general hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention can be essentially or partly contributed to the prior art in the form of a computer program, i.e., a software product, which can be stored in a storage medium and includes several instructions for enabling a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, a MUU or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.
[0147] The present invention provides a false image detection method based on contrastive learning. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A false image detection method based on contrastive learning, characterized in that: The following steps are involved: Step 1: Collect real images and virtual images to build an image dataset; Step 2: construct a false image discrimination model, wherein the false image discrimination model is used to discriminate whether an image is a real image or a false image; Step 3, using the image data set to train the false image discrimination model to obtain a trained false image discrimination model; Step 4: Input the image to be detected into the trained false image discrimination model to obtain the detection result.
2. The false image detection method based on contrastive learning according to claim 1, characterized in that: Step 1 includes: step 1-1, taking the collected real image and the false image as inputs of the generative model to obtain a real reconstructed image and a false reconstructed image, wherein the image category of the collected real image is true, and the image categories of the collected false image, the real reconstructed image and the false reconstructed image are false; Step 1-2, data enhancement is performed on the obtained four types of images, each image after data enhancement is segmented and reorganized in a shuffled order to obtain an enhanced image with destroyed semantic information, and the enhanced images with destroyed semantic information constitute an image data set.
3. The false image detection method based on contrastive learning according to claim 2, characterized in that: The false image discrimination model in step 2 includes a frequency domain analysis module, a noise extraction module, a fusion module, a feature extraction module and a classifier, wherein the frequency domain analysis module is used to perform frequency domain analysis on the input image to obtain frequency domain analysis features; The noise extraction module is used to extract the noise of the input image and obtain the noise characteristics; The fusion module is used to perform feature splicing and fusion on the input image, frequency domain analysis features and noise features to obtain frequency domain noise features; The feature extraction module is used to extract the frequency domain noise features to obtain image features; The classifier is used to calculate the probability value of the image category according to the image features and determine the category of the input image.
4. The false image detection method based on contrastive learning according to claim 3 is characterized in that: Step 3 includes: Step 3-1, inputting the enhanced image with destroyed semantic information into the frequency domain analysis module and the noise extraction module to obtain processed frequency domain analysis features and noise features, inputting the enhanced image with destroyed semantic information, the frequency domain analysis features and the noise features into the fusion module for feature splicing and fusion operation to obtain frequency domain noise features; Step 3-2, input the frequency domain noise features into the feature extraction module to obtain the final image features; Step 3-3, perform comparative learning based on image features and their image categories, narrow the distance between similar features, and make features of different categories farther away, and obtain feature loss values; Step 3-4: Input the image features into the classifier to obtain the probability value of the category to which the image belongs, and calculate the cross entropy loss value based on the category to which it belongs; Step 3-5, use the feature loss value and the cross entropy loss value to perform back propagation to calculate the gradient of the model parameters, and use the gradient descent algorithm to update the model parameters; Step 3-6, repeat step 1-2, step 3-1 to step 3-5, until the number of training times reaches the predetermined number of iterations, and the final trained false image discrimination model is obtained.
5. The false image detection method based on contrastive learning according to claim 4, characterized in that: Steps 1-2 include: Step 1-2-1, each of the four types of images is denoted as I, and image enhancement technology is used to enhance the representativeness of image I to obtain a preliminary preprocessed image. The enhancement technology used is applied to image I under the condition of probability p, including image compression, horizontal flipping, adding Gaussian noise or blurring, cropping to a specified image size after enhancement, and standardizing the pixel value, where 0 <p≤1; Step 1-2-2: split each pre-processed image obtained in step 1-2-1 into small images and reorganize them in a disordered order to obtain an enhanced image I that destroys the semantic information of the image. en ; H is the height of the original image, W is the width of the original image, h is the height of the divided small image, and w is the width of the divided small image.
6. The false image detection method based on contrastive learning according to claim 5, characterized in that: The frequency domain analysis features obtained in step 3-1 include: Step 3-1-1: Enhance the image I en Input into the frequency domain analysis module, first perform Fourier transform on the image to obtain the enhanced image I en The frequency domain characteristics x freq ; Step 3-1-2, from the frequency domain feature x freq The amplitude characteristics x are obtained in abs and phase characteristics x angle ; Step 3-1-3, for the amplitude feature x abs and phase characteristics x angle To update: x abs =process abs (x abs )=ReLU(BN(RRG(x abs ))) x angle =process angle (x angle )=ReLU(BN(RRG(x angle ))) Among them, RRG represents the recursive residual group network, which is composed of multiple dual attention modules in series and combined with skip connections; BN represents batch normalization; ReLU represents the ReLU activation function; Step 3-1-4, for the amplitude feature x obtained in step 3-1-3 abs and phase characteristics x angle Perform feature splicing and fuse the spliced features to obtain the updated amplitude feature x abs and phase characteristics x angle : x temp =concat(x abs ,x angle ,dim=1) x abs =goes abs (x temp ) x angle =goes angle (x temp ) concat means concatenation operation, dim=1 means concatenation operation on 1 dimension of feature, merge abs and merge angle It is a fusion operation, both of which are 1×1 convolution modules; Step 3-1-5, repeat steps 3-1-3 to 3-1-4 n times, and finally obtain the updated amplitude feature x abs and phase characteristics x angle ; where n≥1; Step 3-1-6, the amplitude characteristic x obtained in step 3-1-5 abs and phase characteristics x angle Splice and perform inverse Fourier transform to obtain the extracted frequency domain analysis feature x freq-anal : x freq-anal =M merge1 (concat(M conv (I en ),real(ifft(x abs ,x angle )),dim=1)) Among them, real is to extract the real part of the complex number, ifft is the inverse Fourier transform operation, M conv is a 3×3 convolution operation, M merge1 Represents a fusion operation, including a convolution module of size 3×3 and stride 1, a 2D batch normalization module, and a ReLU activation module.
7. The false image detection method based on contrastive learning according to claim 6, characterized in that: The noise characteristics and frequency domain noise characteristics obtained in step 3-1 include: Step 3-1-7, enhance the image I en Input to the noise extraction module M denoise And obtain the noise characteristic I noise : I noise =I en -M denoise (I en ) Step 3-1-8, by enhancing the image I en , frequency domain analysis characteristics x freq-anal and noise characteristics I noise Perform feature splicing and fusion to obtain the frequency noise feature x fn : x fn =M merge2 (concat(I en ,x freq-anal ,I noise ,dim=1)) Among them, M merge2 It consists of four groups of 3×3 convolution modules, normalization modules and ReLU modules connected in series.
8. The false image detection method based on contrastive learning according to claim 7, characterized in that: Step 3-3 includes: combining any two of the multiple image features obtained in step 3-2 into a pair, and calculating the feature loss value L using a contrastive learning method. contrastive : Among them, D W (i) represents the Euclidean distance between a pair of image features, N represents the number of image feature pairs, i represents the image feature pair index, Y represents whether the pair of image features belong to the same category, and m is the specified minimum distance between features of different categories.
9. The false image detection method based on contrastive learning according to claim 8, characterized in that: Step 3-4 includes: let the image feature obtained in step 3-2 be x feature , the image feature x feature Input to classifier M classifier In the output, the predicted probability of the corresponding category is calculated and the cross entropy loss L is calculated. cross-entropy : logits=M classifier (x feature ) probs=softmax(logits,dim=-1) Among them, logits is the score output by the classifier, and softmax calculates the probability probs of each category on the last dimension based on logits to obtain the predicted category; p j is the probability value that the jth sample image in the image dataset belongs to the predicted category, and M represents the total number of samples in the image dataset.
10. The false image detection method based on contrastive learning according to claim 9, characterized in that: Steps 3-5 include: combining the cross entropy loss L cross-entropy And feature loss L contrastive Perform back propagation calculations and use the gradient descent algorithm to optimize the false image discrimination model parameters: THE total =λL contrastive +(1-λ)L cross-entropy Among them, L total represents the loss function of the fake image discrimination model, λ is a hyperparameter that controls the ratio between the two loss values, θ model are the parameters of the false image discrimination model, and lr is the learning rate that controls the update speed of the false image discrimination model parameters.