Artificial intelligence generated image detection method based on consistency verification

Through the detection method based on consistency verification, the pre-trained natural image basic model and image transformation function are used to detect images generated by artificial intelligence, solving the problem of relying on specific training images and high maintenance costs in the existing technology, and achieving efficient and low-cost detection effects.

CN120125554APending Publication Date: 2025-06-10UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235277.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When detecting images generated by artificial intelligence, the prior art relies on specific manual annotations to train images, resulting in a decrease in detection accuracy and an increase in data collection and maintenance costs, and it is difficult to adapt to new generation technologies.

Method used

The detection method based on consistency verification is adopted, and the image to be detected is randomly cropped or completed by setting the target size of the image, and the pre-trained natural image basic model and image transformation function are used to calculate the feature difference of the image to determine whether the image is generated by artificial intelligence.

Benefits of technology

Without relying on specific hand-noted training images, it can effectively detect various types of artificial intelligence-generated images, reduce the risk of false and forged content to sensitive areas, and reduce training costs and data collection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125554A_ABST
    Figure CN120125554A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence generated image detection method based on consistency verification. The method comprises the steps of 1, extracting features of a to-be-detected image by using a large visual basic model pre-trained on a large number of natural images; 2, performing image transformation on the to-be-detected image to generate a transformed image; 3, extracting features of the transformed image by using the same model; and 4, calculating the similarity between the original image features and the transformed image features. According to the method, the image generated by the artificial intelligence model can be reliably detected without training, so that the potential risk brought by false and forged contents to the highly sensitive field can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and computer vision, and specifically relates to a method for detecting artificial intelligence-generated images based on consistency verification. Background Art

[0002] In recent years, with the rapid development of generative models, the technology of artificial intelligence-generated images has made remarkable progress. Such technologies can generate highly realistic images and are widely used in fields such as entertainment and advertising. However, as the quality of generated images continues to improve, they may be maliciously used for purposes such as the spread of false information and malicious manipulation, posing a huge security risk to society.

[0003] Existing detection technologies usually rely on binary classification models and are trained with a large number of manually labeled natural images and generated images. This method has two significant drawbacks: one is that the generative model used in training may be inconsistent with the model encountered during detection, resulting in a decrease in detection accuracy; the other is that the model needs to be continuously updated to adapt to new generative technologies, increasing the cost of data collection and maintenance. Therefore, how to improve the detection ability and reduce the training dependence of the detection model has become an urgent problem in this field. Summary of the Invention

[0004] The present invention is proposed to solve the above-mentioned deficiencies of the existing technologies, and provides a method for detecting artificial intelligence-generated images based on consistency verification, aiming to be able to detect various different types of artificial intelligence-generated images without relying on specific manually labeled training images, thereby reducing the potential risks brought by false and forged content to highly sensitive fields.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for detecting artificial intelligence-generated images based on consistency verification according to the present invention is characterized by including:

[0007] Step 1: Set the target size of the image, and perform N times of random cropping or padding on the image to be detected according to the set target size, so as to obtain a processed image set ... .. }; where x i represents the i-th preprocessed image; ;

[0008] Step 2: Select a base model trained by self-supervised learning from natural images to process it, and obtain an original feature set F 1 = { ... ... }, where represents the original feature;

[0009] Step 3: Use the image transformation function to process to obtain the transformed feature set F 2 = { ... ... }, where represents the transformed feature; and ;

[0010] Step 4: Use Equation (1) or Equation (2) to obtain the feature difference of the image to be detected :

[0011] (1)

[0012] In Equation (1), represents the transpose of and respectively represent and the norms of , ;

[0013] Step 5: If > , then determine that the image to be detected is an AI-generated image; otherwise, determine that the image to be detected is a natural image, where represents the threshold.

[0014] Another feature of the method for detecting AI-generated images based on consistency verification according to the present invention is that the image transformation function is a composite transformation function composed of any one or more image transformation methods, and the image transformation methods include: horizontal flipping, randomly changing brightness, contrast, saturation, and hue, and Gaussian blur.

[0015] A feature of an electronic device according to the present invention, which includes a memory and a processor, is that the memory is used to store a program that supports the processor to execute the method for detecting AI-generated images, and the processor is configured to execute the program stored in the memory.

[0016] A computer-readable storage medium of the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the artificial intelligence-generated image detection method.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] 1. By virtue of the ability of the large-scale visual foundation model itself to distinguish natural images and artificial intelligence-generated images, the present invention uses the feature similarity of the test image and its transformed image on the foundation model as the basis for judging whether the image is generated by an artificial intelligence model, overcoming the dependence of traditional methods on specific types of training images and alleviating the problem of insufficient generalization.

[0019] 2. The present invention utilizes a pre-trained model without further training, making full use of the powerful visual representation ability of the pre-trained large-scale visual foundation model. Compared with traditional methods, the present invention greatly reduces the training cost and data collection cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is the overall flowchart of the disclosed embodiment of the present invention;

[0021] Figure 2 It is the experimental result diagram on the ImageNet dataset provided by the embodiment of the present invention;

[0022] Figure 3 It is the experimental result diagram on the LSUN-BEDROOM dataset provided by the embodiment of the present invention;

[0023] Figure 4 It is the experimental result diagram on the GenImage dataset provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] Please refer to Figure 1 , in this embodiment, an artificial intelligence-generated image detection method based on consistency verification includes the following steps:

[0026] Step 1: Set the target size of the image, and perform N times of random cropping or padding on the image to be detected according to the set target size. If the size of the image to be detected is smaller than the required input size, pad zeros around it to the specified size; if the size is larger than the required input size, crop it to the specified size. The operations of padding zeros or cropping can avoid distortion caused during the image scaling process. In this way, a processed image set is obtained. ... .. }; where, x i represents the i-th preprocessed image; .

[0027] Step 2: Select a base model trained by self-supervision from natural images to perform processing, and obtain the original feature set F1 = { ... ... }, where, represents the original feature of; f represents Figure 1 the visual base model in.

[0028] Step 3: Use the image transformation function to perform processing, and obtain the transformed feature set F2 = { ... ... }, where, represents the transformed feature of; and ; where, represents Figure 1 the image transformation function in, which is designed as an image transformation function along the tangent space of the natural distribution manifold. Given that the image transformation function used in the base model training stage is an approximation of the local tangent space of the manifold, in this embodiment, this function can be instantiated as a composite transformation function composed of multiple image transformations such as horizontal flipping, randomly changing brightness, contrast, saturation and hue, and Gaussian blur.

[0029] Step 4: Use Equation (1) to obtain the feature difference of the image to be detected :

[0030] (1)

[0031] In Equation (1), represents the transpose of, and respectively represent and of norm, and , .

[0032] Step 5: In the pre-training stage, the base model is trained to be insensitive to image transformations along the tangent space of the natural distribution manifold. Therefore, in this embodiment, the similarity between the test image and its transformed image in the feature space of the base model can be used as an important indicator for determining whether the image is a natural image or a generated image. Specifically, if > , it is determined that the image to be detected is an AI-generated image; otherwise, it is determined that the image to be detected is a natural image, where represents a threshold value.

[0033] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0034] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

[0035] To verify the effectiveness of the present invention, comparative experiments will be conducted on three datasets below, and the experimental results will be given. Moreover, comparisons will be made with four advanced artificial intelligence image detection methods, namely CNNspot (Wang S Y, Wang O, Zhang R, et al. CNN-generated images are surprisingly easy to spot...for now[C] / / Proceedings of the IEEE / CVF conference on computer vision andpattern recognition. 2020: 8695-8704.), DIRE (Wang Z, Bao J, Zhou W, et al.Dire for diffusion-generated image detection[C] / / Proceedings of the IEEE / CVFInternational Conference on Computer Vision. 2023: 22445-22455.), UniversalFake Detect (abbreviated as UFD, Ojha U, Li Y, Lee Y J. Towards universal fake imagedetectors that generalize across generative models[C] / / Proceedings of theIEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023: 24480-24489.) and RIGID (He Z, Chen P Y, Ho T Y. Rigid: A training-free and model-agnostic framework for robust ai-generated image detection[J]. arXiv preprintarXiv:2405.20112, 2024.). (The meaning of the abbreviation itself is unclear and requires a Chinese definition or an English full name introduction)

[0036] The three datasets are the ImageNet (Deng J, Dong W, Socher R, et al. Imagenet: A large-scale hierarchical image database[C] / / 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009: 248-255.) generated image dataset, the LSUN-BEDROOM generated image dataset (Yu F, Seff A, Zhang Y, et al. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop[J]. arXiv preprint arXiv:1506.03365, 2015.), and GenImage (Zhu M, Chen H, Yan Q, et al. Genimage: A million-scale benchmark for detecting ai-generated image[J]. Advances in Neural Information Processing Systems, 2023, 36: 77771-77782.). ImageNet includes those generated by ADM (Dhariwal P, Nichol A. Diffusion models beat ganson image synthesis[J]. Advances in neural information processing systems, 2021, 34: 8780-8794.), ADMG (Dhariwal P, Nichol A. Diffusion models beat ganson image synthesis[J]. Advances in neural information processing systems, 2021, 34: 8780-8794.), LDM (Rombach R, Blattmann A, Lorenz D, et al.High-resolution image synthesis with latent diffusion models[C] / / Proceedings ofthe IEEE / CVF conference on computer vision and pattern recognition. 2022:10684-10695.),DiT(Peebles W, Xie S. Scalable diffusion models withtransformers[C] / / Proceedings of the IEEE / CVF international conference oncomputer vision. 2023: 4195-4205.),BigGAN(Brock A, Donahue J, Simonyan K.Large scale GAN training for high fidelity natural image synthesis[J]. arXivpreprint arXiv:1809.11096, 2018.),GigaGAN(Kang M, Zhu J Y, Zhang R, et al.Scaling up gans for text-to-image synthesis[C] / / Proceedings of the IEEE / CVFconference on computer vision and pattern recognition. 2023: 10124-10134.),StyleGAN XL(Karras T, Laine S, Aila T. A style-based generator architecturefor generative adversarial networks[C] / / Proceedings of the IEEE / CVFconference on computer vision and pattern recognition. 2019: 4401-4410.),RQ-Transformer(Lee D, Kim C, Kim S, et al.Generated images synthesized by these generative models such as Autoregressive image generation using residual quantization [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022: 11523-11532.) and Mask GIT (Chang H, Zhang H, Jiang L, et al. Maskgit: Masked generative image transformer [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 11315-11325.). LSUN-BEDROOM includes those generated by ADM, DDPM (Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models [J]. Advances in neural information processing systems, 2020, 33: 6840-6851.), iDDPM (Nichol A Q, Dhariwal P. Improved denoising diffusion probabilistic models [C] / / International conference on machine learning. PMLR, 2021: 8162-8171.), Diffusion GAN (Wang Z, Zheng H, He P, et al. Diffusion-gan: Training gans with diffusion [J]. arXiv preprint arXiv:2206.02262, 2022.), Projected GAN (Wang Z, Zheng H, He P, et al. Diffusion-gan: Training gans with diffusion [J]. arXiv preprint arXiv:2206.02262, 2022.), the generated images synthesized by generative models such as StyleGAN and UnleashingTransformer (Bond-Taylor S, Hessey P, Sasaki H, et al. Unleashingtransformers: Parallel token prediction with discrete absorbing diffusion forfast high-resolution image generation from vector-quantized codes[C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland,2022: 170-188.). GenImage includes those generated by Midjourney (Midjourney. https: / / www.midjourney.com / home / . 2022.), Stable Diffusion (Rombach R, Blattmann A, Lorenz D, et al. High-resolution image synthesiswith latent diffusion models[C] / / Proceedings of the IEEE / CVF conference oncomputer vision and pattern recognition. 2022: 10684-10695.), ADM, GLIDE (NicholA, Dhariwal P, Ramesh A, et al. Glide: Towards photorealistic imagegeneration and editing with text-guided diffusion models[J]. arXiv preprintarXiv:2112.10741, 2021.), Wukong (Wukong. https: / / xihe.mindspore.cn / modelzoo / wukong. 2022.), VQDM (Gu S, Chen D, Bao J, et al.Vector quantized diffusion model for text-to-image synthesis[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 10696-10706.) and the generated images synthesized by generative models such as BigGAN. On the ImageNet and LSUN-BEDROOM datasets, in this embodiment, the area under the receiver operating characteristic (AUROC) and the average precision (AP) are used as evaluation metrics. On GenImage, in this embodiment, the classification accuracy (ACC) is used as the evaluation metric.

[0037] The experimental results are as Figures 2 - 4 shown. ConV in the figure represents the method of the present invention. The results show that the method of the present invention achieves the best detection performance on various datasets. Most existing large vision-based models are trained on natural image data, which leads to a deviation between the model in natural images and generated images. Based on this, a method for detecting generated images based on consistency verification proposed by the present invention utilizes the consistency of natural images and their transformed images in the feature space to achieve efficient detection performance. This method does not need to rely on the training of generated image data and can be seamlessly integrated into future large vision-based models, providing an efficient and low-cost solution for the detection of generated images.

[0038] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for detecting artificial intelligence generated images based on consistency verification, characterized in that: include: Step 1: Set the target size of the image, and perform random cropping or completion N times on the image to be detected according to the set target size, so as to obtain a processed image set. ... .. }; where x i represents the i-th preprocessed image; ; Step 2: Select a base model pair trained with natural images through self-supervision After processing, the original feature set F1 = { ... ... },in, express The original characteristics of Step 3: Using image transformation function right After processing, the transformed feature set F2 = { ... ... } in, express The transformation characteristics of ; Step 4: Use formula (1) or formula (2) to obtain the feature difference of the image to be detected : (1) In formula (1), express The transpose of and Respectively and of norm, and , ; Step 5: If > , then the image to be detected is determined to be an artificial intelligence generated image, otherwise, the image to be detected is determined to be a natural image, where Indicates the threshold value.

2. The method for detecting an artificial intelligence generated image based on consistency verification according to claim 1, characterized in that: Image transformation function It is a composite transformation function composed of any one or more image transformation modes, wherein the image transformation modes include: horizontal flipping, random changes in brightness, contrast, saturation and hue, and Gaussian blur.

3. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the artificial intelligence generated image detection method described in claim 1 or 2, and the processor is configured to execute the program stored in the memory.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the artificial intelligence generated image detection method according to claim 1 or 2 are executed.