Detection of generated image

By reconstructing and applying the classification model of the target image, the smoothness difference between the real image and the generated image is solved, and the problem of identifying and generating images is achieved is achieved with efficient and accurate detection effect.

WO2025146094A1PCT designated stage expired Publication Date: 2025-07-10ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/070194
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2025-01-02
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify generated images and real images, resulting in possible infringement risks and data security issues, especially when the generated images generated by the model may be used for illegal operations.

Method used

通过获取目标图像,选取部分图像进行重构处理,利用预先训练的分类模型根据重构图像的平滑度差异进行分类,识别生成图像和真实图像。

Benefits of technology

It improves the efficiency and accuracy of image generation detection, reduces interference factors, and can quickly and effectively distinguish between images and real images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070194_10072025_PF_FP_ABST
    Figure CN2025070194_10072025_PF_FP_ABST
Patent Text Reader

Abstract

One or more embodiments of the present description disclose a detection method and apparatus for a generated image. The method comprises: first, acquiring a target image; second, selecting a partial image of the target image, and reconstructing the partial image on the basis of the remaining image of the target image except the partial image, so as to acquire a reconstructed image consisting of the remaining image and the reconstructed partial image; then, inputting the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image; and finally, on the basis of the reconstruction effect category of the reconstructed image, determining whether the target image is a real image or a generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Generating image detection Technical Field

[0001] This document relates to the field of image recognition technology, and in particular to the detection of generated images. Background Art

[0002] With the development of artificial intelligence (AI), the application of various generative models is becoming increasingly widespread. For example, the text-to-image model can generate corresponding images based on a piece of text. Generative models based on diffusion models can significantly improve the quality of generated images. Therefore, various open source communities support users to create related works based on different open source generative models.

[0003] However, the widespread use of generative models also carries with it corresponding impacts. On the one hand, images generated using text-based image processing techniques pose potential infringement risks. For example, celebrity photos and paintings generated using large text-based image processing models can raise copyright issues for the artists. On the other hand, images generated using image-based image processing techniques can pose data security risks, hindering the dissemination of authentic and valid information. For example, image-based image processing techniques can be used to stylize and edit real images, potentially leading to illegal operations related to image authentication. With increasing attention to privacy concerns and compliance requirements for generated images, there is an urgent need for a method to detect generated images and identify them promptly. Summary of the Invention

[0004] On the one hand, one or more embodiments of the present specification provide a method for detecting a generated image, comprising: acquiring a target image, wherein the target image includes a real image and / or a generated image, wherein the real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions; selecting a partial image from the target image, and reconstructing the partial image based on the remaining image in the target image except the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image; inputting the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, the classification model being used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image; and determining whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

[0005] On the other hand, one or more embodiments of the present specification provide a detection device for generating an image, comprising: an image acquisition module, which acquires a target image, wherein the target image includes a real image and / or a generated image, wherein the real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions; an image filling processing module, which selects a partial image from the target image, and reconstructs the partial image based on the remaining image in the target image except the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image; a classification module, which inputs the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, and the classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image; a generated image determination module, which determines whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

[0006] On the other hand, one or more embodiments of the present specification provide an electronic device, comprising: a processor; and a memory arranged to store computer-executable instructions, which, when the executable instructions are executed, enable the processor to: acquire a target image, wherein the target image includes a real image and / or a generated image, wherein the real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions; select a partial image from the target image, and reconstruct the partial image based on the remaining image in the target image except the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image; input the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, wherein the classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image; and determine whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

[0007] On the other hand, one or more embodiments of the present specification provide a storage medium for storing a computer program, which can be executed by a processor to implement the following process: obtaining a target image, wherein the target image includes a real image and / or a generated image, wherein the real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions; selecting a partial image from the target image, and reconstructing the partial image based on the remaining image in the target image except the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image; inputting the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, the classification model being used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image; and determining whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0009] FIG1 is a schematic flow chart of a detection method for generating an image according to an embodiment of this specification;

[0010] FIG2 is a schematic diagram illustrating the implementation principle of a detection method for generating an image according to an embodiment of this specification;

[0011] FIG3 is a schematic block diagram of a detection device for generating an image according to an embodiment of this specification;

[0012] FIG4 is a schematic block diagram of an electronic device according to an embodiment of this specification. DETAILED DESCRIPTION

[0013] One or more embodiments of this specification provide a detection method and apparatus for generating an image.

[0014] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0015] As shown in Figure 1, an embodiment of this specification provides a detection method for generating an image. The execution subject of this method can be a terminal device or a server, wherein the terminal device can be a certain terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car-mounted device, etc.). The server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server for a financial business or an online shopping business, or a background server for an application, etc. In this embodiment, a detailed description is given using a server as an example. For the execution process of the terminal device, please refer to the following relevant content, which will not be repeated here. The method can specifically include the following steps:

[0016] In step S102 , a target image is acquired, where the target image includes a real image and / or a generated image.

[0017] In the embodiments of this specification, the target image is the current image to be detected. This target image can be a single image, specifically a real image or a generated image. The target image can also be a collection of multiple images to be detected, which can be a collection of real images, a collection of generated images, or a collection composed of a mixture of real images and generated images. The detection method of the embodiments of this specification can identify all generated images from this collection.

[0018] It should be noted that the real image and the generated image can be images of the same content or images of different contents. Taking the target image in the form of a set of multiple images to be detected as an example, the target image may include: a real image of a duckling, a generated image of a running puppy, a generated image of a kitten, and a real image of a document. Through the detection method in the embodiments of this specification, the generated image of a running puppy and the generated image of a kitten can be identified.

[0019] In addition, the real images in the embodiments of this specification are images captured by an image acquisition device and are unprocessed images, such as images captured by a camera. Generated images are images generated or synthesized based on preset conditions, typically images processed by image processing algorithms or image processing tools, such as images generated based on AIGG (AI Generated Content), images generated by the Wenshengtu model, and images generated by generative models.

[0020] In step S104, a partial image is selected from the target image, and the partial image is reconstructed based on the remaining image in the target image except the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image.

[0021] The reconstructed image is composed of the remaining image and the reconstructed partial image, that is, the reconstructed image corresponding to the target image. In order to obtain the reconstructed image corresponding to the target image, the embodiment of this specification first selects a partial image from the target image, and then reconstructs the partial image based on the remaining image in the target image except the partial image.

[0022] In implementation, the method of reconstructing part of the image can be to cover part of the image with other images, for example: using a mask to perform mask processing and then cover part of the image, and then perform image filling processing on the covered area; it can also be to remove the selected part of the image from the target image and then perform image filling processing on the removed area; it can also be to identify the image content and select content with specified semantics, and then use other images to cover the selected content with specified semantics or remove the selected content with specified semantics, and then perform image filling processing on the covered or removed area.

[0023] In step S106, the reconstructed image is input into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image. The classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image based on a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image.

[0024] In implementation, the classification model can be obtained by model training based on multiple real image samples and generated image samples and a preset loss function. The classification model can adopt a binary classification model constructed based on a deep neural network. The input data of the classification model is the reconstructed image corresponding to the target object, and the output result is the reconstruction effect category of the reconstructed image. The reconstruction effect category includes two categories: good and poor. The larger the reconstruction error, the worse the reconstruction effect, and the smaller the reconstruction error, the better the reconstruction effect. According to the preset reconstruction error threshold, if the reconstruction error exceeds the threshold, it is judged as poor reconstruction effect, and if it is lower than the preset reconstruction error threshold, it is judged as good reconstruction effect. Since the reconstruction effect of the generated image and the real image is quite different, the reconstruction effect category of the generated image is good, and the reconstruction effect category of the real image is poor, so that the generated image and the real image can be distinguished.

[0025] In step S108 , it is determined whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

[0026] Since the smoothness difference between the real image and the generated image is large, and the large smoothness difference will lead to a large difference in the reconstruction effects of the reconstructed images corresponding to the real image and the generated image. Specifically, the smoothness of the pixels inside the generated image is high. When the image filling process is performed on the area where the partial image is located, a smoother reconstructed image will be obtained by performing the filling process based on the image information around the partial image (i.e., the image information of the remaining image in the target image except the partial image), so that the reconstruction error between the reconstructed image corresponding to the generated image and the generated image is small, and the reconstruction effect is good; while the smoothness of the pixels around the real image is low. When the image filling process is performed on the area where the partial image is located, a relatively smooth reconstructed image will be obtained by performing the filling process based on the image information around the partial image, so that the reconstruction error between the reconstructed image corresponding to the real image and the real image is large, and the reconstruction effect is poor. Therefore, the present embodiment can effectively identify the generated image by comparing the reconstruction effects of the reconstructed images of the real image and the generated image, thereby improving the accuracy and detection efficiency of the generated image detection results.

[0027] In implementation, if the reconstruction effect category of the reconstructed image corresponding to the target image is good, the current target image is determined to be a generated image; if the reconstruction effect category of the reconstructed image corresponding to the target image is poor, the current target image is determined to be a real image.

[0028] The embodiments of this specification provide a detection method for a generated image, which first obtains a target image, then selects a partial image from the target image, and reconstructs the partial image based on the remaining image in the target image except the partial image, obtaining a reconstructed image composed of the remaining image and the reconstructed partial image, and then inputs the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, and finally determines whether the target image is a real image or a generated image based on the reconstruction effect category of the reconstructed image. The embodiments of this specification make full use of the principle that the smoothness difference between the real image and the generated image is large, and the large smoothness difference will lead to a large difference in the reconstruction effect of the reconstructed images corresponding to the real image and the generated image respectively. By reconstructing a partial area of ​​the target image to obtain the reconstructed image corresponding to the target image, and by classifying the reconstruction effect of the reconstructed image, it is possible to quickly and effectively determine whether the current target image is a real image or a generated image, thereby effectively improving the detection efficiency of the generated image and the accuracy of the detection results. In addition, in the embodiment of this specification, the input data of the classification model is a reconstructed image. The classification model is based on the real image and its corresponding reconstructed image, and the reconstruction effect category is divided based on the generated image and its corresponding reconstructed image. Therefore, the images used for reconstruction effect classification have the same content. This method is conducive to reducing interference factors and enables the classification model to make category judgments based on the actual content gap of the image, thereby improving the accuracy of the classification results, and then improving the accuracy of the generated image detection results.

[0029] In the embodiment of this specification, the processing of the above step S104 can be varied. An optional processing method is provided below. For details, please refer to the processing of the following steps S1042-S1044.

[0030] In step S1042 , mask processing is performed on the selected partial image according to a preset mask to obtain a target image to be filled consisting of the remaining image and the partial image after the mask processing.

[0031] Here, the target image to be filled is composed of the remaining image and the partial image after mask processing, and the area to be filled in the target image to be filled is the area where the mask is located.

[0032] The area where the partial image is located is masked using a preset mask, which can cover the selected partial image in the target image, thereby obtaining the target image to be filled consisting of the remaining image and the partial image after masking.

[0033] In one implementation, the size ratio of the preset mask to the target image is less than or equal to 1 / 4. In practice, the size ratio of the preset mask to the target image can be selected to be 1 / 4, 1 / 8, etc., and the preset mask can be set to any shape. If the target image is a collection of multiple images to be detected, the size ratio of the preset mask to the image with the lowest area ranking in the collection of multiple images to be detected is less than or equal to 1 / 4. This preset mask size can not only improve the efficiency of the image filling process, but also meet the accuracy of the image filling process, which is conducive to improving the operability of the image filling process.

[0034] In step S1044, image filling processing is performed on the target image to be filled based on the image filling model, so as to reconstruct a portion of the image in the area where the mask is located based on the remaining image to obtain a reconstructed image corresponding to the target image.

[0035] In implementation, the image filling model may adopt an existing image filling model, for example, an image filling model constructed based on a diffusion model.

[0036] In the embodiment of this specification, the processing of the above step S104 can be varied. Another optional processing method is provided below. For details, please refer to the processing of the following steps S1046-S1048.

[0037] In step S1046 , a partial image is removed from the target image, and the target image to be filled is obtained based on the remaining image and the area where the partial image is located.

[0038] The target image to be filled here consists of the remaining image and the area where the partial image is located (i.e., the blank image). The area to be filled of the target image to be filled is the corresponding area in the target image where the partial image is removed (or the area where the blank image is located).

[0039] In implementation, the existing cutout tool can be used to directly cut out the part of the image to be selected from the target image, and retain the area where the part of the image is located in the target image, so as to obtain the target image to be filled based on the remaining image and the area where the part of the image is located. This method can quickly obtain the target image to be filled, which is conducive to improving the efficiency of image filling processing.

[0040] In step S1048, image filling processing is performed on the target image to be filled based on a preset image filling algorithm, so as to reconstruct the partial image in the area where the partial image is located based on the remaining image to obtain a reconstructed image corresponding to the target image.

[0041] In practice, the preset image filling algorithm may adopt an injection filling area algorithm, a seed filling algorithm, a scan line filling algorithm, an edge filling algorithm, etc., which is not limited in the embodiments of this specification. Alternatively, an image filling model may be used to perform image filling processing on the target image to be filled to obtain a reconstructed image corresponding to the target image.

[0042] In the embodiment of this specification, the training method of the classification model in the above step S106 can be various. An optional processing method is provided below. For details, please refer to the processing steps A1-A3 below.

[0043] In step A1, a plurality of image samples with label information are obtained, where the image samples include real image samples and generated image samples, and the label information is used to mark the image samples as real image samples or generated image samples.

[0044] The multiple image samples obtained can be of the same or different content. The image samples carry label information for supervised training, which helps improve the accuracy of the classification model's results. Multiple image samples can be obtained by directly acquiring multiple real image samples and multiple generated image samples from open source datasets. Alternatively, multiple real image samples can be obtained from open source datasets alone and then corresponding generated image samples generated based on the acquired real image samples.

[0045] Step A2: Select a partial image from each image sample, and reconstruct the partial image based on the remaining image in each image sample except the partial image, to obtain a reconstructed image composed of the remaining image corresponding to each image sample and the reconstructed partial image.

[0046] In step A3, a plurality of image samples are used as input data, and the reconstruction effect categories of the reconstructed images corresponding to the plurality of image samples are used as output results. The classification model is trained based on a preset loss function to obtain a trained classification model.

[0047] In implementation, the preset loss function is a classification model loss function, and a cross entropy loss function can be used.

[0048] In the embodiments of this specification, the processing method of the above step A1 can be various. An optional processing method is provided below. For details, please refer to the processing of the following steps A11-A12.

[0049] Step A11: Acquire multiple real image samples based on an open source dataset.

[0050] Step A12: Generate multiple generated image samples based on multiple real image samples using the generative model.

[0051] In implementation, the generative model may adopt a diffusion model, a VAE (Variational Autoencoder) generative model, a flow-based generative model, and the like.

[0052] Based on the open source dataset, only multiple real image samples are obtained, and corresponding generated image samples are generated based on the obtained real image samples. This method can realize model training with a small number of samples, and the generated images and real images in this method are consistent in content, which is conducive to the model's directional learning of the actual content differences between real pictures and generated pictures. It is more targeted and conducive to further improving the efficiency of model training and improving the accuracy of classification results of classification models.

[0053] In the embodiment of the present specification, the classification model in the above step S106 is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to the first reconstruction error and the second reconstruction error. The first reconstruction error and the second reconstruction error can be implemented in various ways. In one implementation, the first reconstruction error is determined based on a preset similarity index between the real image and its corresponding reconstructed image, and the second reconstruction error is determined based on a preset similarity index between the generated image and its corresponding reconstructed image, and both the first reconstruction error and the second reconstruction error are inversely proportional to the smoothness of the target image. The larger the first reconstruction error (the second reconstruction error), the smaller the smoothness of the target image, and the smaller the first reconstruction error (the second reconstruction error), the greater the smoothness of the target image.

[0054] In one implementation, the preset similarity indicators include one or more of a structural similarity parameter, a peak signal-to-noise ratio parameter, and a learning-perceptual image block similarity parameter. That is, only one similarity indicator may be used, or two or three of the above similarity indicators may be selected simultaneously to determine the first reconstruction error or the second reconstruction error. The structural similarity parameter is inversely proportional to the first and second reconstruction errors, and directly proportional to the reconstruction effect; the peak signal-to-noise ratio parameter is inversely proportional to the first and second reconstruction errors, and directly proportional to the reconstruction effect; and the learning-perceptual image block similarity parameter is directly proportional to the first and second reconstruction errors, and inversely proportional to the reconstruction effect. That is, the larger the structural similarity parameter, the smaller the first reconstruction error (second reconstruction error) and the better the reconstruction effect; the larger the peak signal-to-noise ratio parameter, the smaller the first reconstruction error (second reconstruction error) and the better the reconstruction effect; and the larger the learning-perceptual image block similarity parameter, the larger the first reconstruction error (second reconstruction error) and the worse the reconstruction effect.

[0055] The implementation principle of the detection method for generating an image in the embodiment of this specification can be seen in Figure 2. Taking the generated image output by the generation model based on the diffusion model as an example, since the diffusion model performs image generation processing by a multi-step iterative denoising method, in the hierarchical denoising process, the noise irrelevant to the set conditional text is gradually removed by the guidance of the input conditions, thereby generating a generated image that meets the set conditional text. Therefore, the pixels of the generated image obtained by the above method are too smooth globally (the smoothness is high). However, the real image taken by the image acquisition device will cause the pixels of the real image to be sharper (the smoothness is lower) due to the different natural scenes and camera parameters when shooting. Therefore, the smoothness of the generated image and the real image is quite different. The embodiment of this specification distinguishes the smoothness of the generated image and the real image by comparing the reconstruction effect of the reconstructed image corresponding to the real image with the reconstructed image corresponding to the generated image, so that the generated image and the real image can be distinguished quickly and effectively.

[0056] Taking 2500 generated images based on Stable Diffusion (an intelligent AI drawing generation tool, a generation model based on a diffusion model) and 2500 real images taken by a camera as examples, by randomly generating a mask with a ratio smaller than 1 / 8 of the area of ​​the real image and a mask smaller than 1 / 4 of the area of ​​the real image, the generated images and the real images are respectively masked to obtain the area to be filled (i.e., the mask area), and then the image filling model based on Stable Diffusion is used to perform image filling processing on the area to be filled, and the reconstructed images corresponding to the generated images and the reconstructed images corresponding to the real images are obtained. The similarity indicators between the generated images and the real images after the image filling processing are respectively calculated, and the contents shown in Tables 1 and 2 below are obtained. Table 1 Similarity indicators when the mask area is 1 / 8 of the original image area Table 2 Similarity indicators when the mask area is 1 / 4 of the original image area

[0057] As shown in Tables 1 and 2, the SSIM (structural similarity index measurement) and PSNR (peak signal-to-noise ratio) of the real image are lower than those of the generated image, and the LPIPS (learned perceptual image patch similarity) of the real image is higher than that of the generated image. Specifically, due to the high pixel smoothness within the generated image, when filling the area where the partial image is located, a smoother reconstructed image is obtained by filling the area based on the image information around the partial image. Therefore, the reconstruction error between the generated image and its corresponding reconstructed image is small, and the reconstruction effect is good. However, due to the low pixel smoothness around the real image, when filling the area where the partial image is located, a relatively smooth reconstructed image is also obtained by filling the area where the partial image is located based on the image information around the partial image. Therefore, the reconstruction error between the real image and its corresponding reconstructed image is large, and the reconstruction effect is poor. Combining Tables 1 and 2, it can be seen that the reconstruction effect of the real image is far inferior to that of the generated image. Based on this, after training the classification model with multiple image samples in Figure 2, the classification model divides the reconstructed images corresponding to the target image into two categories: images with good reconstruction effect and images with poor reconstruction effect. Finally, it is determined whether the current target image is a generated image or a real image based on the reconstruction effect.

[0058] The embodiments of this specification provide a detection method for a generated image, which first obtains a target image, then selects a partial image from the target image, and reconstructs the partial image based on the remaining image in the target image except the partial image, obtaining a reconstructed image composed of the remaining image and the reconstructed partial image, and then inputs the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, and finally determines whether the target image is a real image or a generated image based on the reconstruction effect category of the reconstructed image. The embodiments of this specification make full use of the principle that the smoothness difference between the real image and the generated image is large, and the large smoothness difference will lead to a large difference in the reconstruction effect of the reconstructed images corresponding to the real image and the generated image respectively. By reconstructing a partial area of ​​the target image to obtain the reconstructed image corresponding to the target image, and by classifying the reconstruction effect of the reconstructed image, it is possible to quickly and effectively determine whether the current target image is a real image or a generated image, thereby effectively improving the detection efficiency of the generated image and the accuracy of the detection results. In addition, in the embodiment of this specification, the input data of the classification model is a reconstructed image. The classification model is based on the real image and its corresponding reconstructed image, and the reconstruction effect category is divided based on the generated image and its corresponding reconstructed image. Therefore, the images used for reconstruction effect classification have the same content. This method is conducive to reducing interference factors and enables the classification model to make category judgments based on the actual content gap of the image, thereby improving the accuracy of the classification results, and then improving the accuracy of the generated image detection results.

[0059] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.

[0060] The above is a detection method for generating an image provided in one or more embodiments of this specification. Based on the same idea, one or more embodiments of this specification also provide a detection device for generating an image, as shown in FIG3 .

[0061] The detection device for generating an image includes: an image acquisition module 210 , an image filling processing module 220 , a classification module 230 and a generated image determination module 240 .

[0062] An image acquisition module 210 acquires a target image, where the target image includes a real image and / or a generated image. The real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions.

[0063] An image filling processing module 220 selects a partial image from a target image and reconstructs the partial image based on a remaining image in the target image except the partial image, thereby obtaining a reconstructed image composed of the remaining image and the reconstructed partial image.

[0064] a classification module 230 that inputs the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, wherein the classification model is configured to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image based on a first reconstruction error between the real image and its corresponding reconstructed image, and a second reconstruction error between the generated image and its corresponding reconstructed image;

[0065] The generated image determination module 240 determines whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

[0066] In an embodiment of the present specification, the image filling processing module 220 may include: a mask processing unit, which performs mask processing on the selected portion of the image according to a preset mask, and obtains a target image to be filled composed of the remaining image and the portion of the image after mask processing; a first image filling processing unit, which performs image filling processing on the target image to be filled based on an image filling model, so as to reconstruct the portion of the image in the area where the mask is located based on the remaining image, and obtain a reconstructed image corresponding to the target image.

[0067] In one embodiment, the size ratio of the mask preset in the mask processing unit to the target image is less than or equal to 1 / 4.

[0068] In an embodiment of the present specification, the image filling processing module 220 may further include: a subtraction unit, which subtracts a portion of the image from the target image, and obtains the target image to be filled based on the remaining image and the area where the partial image is located; a second image filling processing unit, which performs image filling processing on the target image to be filled based on a preset image filling algorithm, so as to reconstruct the partial image in the area where the partial image is located based on the remaining image, and obtain a reconstructed image corresponding to the target image.

[0069] In an embodiment of the present specification, in the classification model of the classification module 230, the first reconstruction error is determined based on a preset similarity index between a real image and its corresponding reconstructed image, and the second reconstruction error is determined based on a preset similarity index between a generated image and its corresponding reconstructed image, and both the first reconstruction error and the second reconstruction error are inversely proportional to the smoothness of the target image.

[0070] In one embodiment, the similarity indicators preset in the classification model of the classification module 230 include: one or more of a structural similarity parameter, a peak signal-to-noise ratio parameter, and a learning-perceived image block similarity parameter, and the structural similarity parameter is inversely proportional to the first reconstruction error and the second reconstruction error, the peak signal-to-noise ratio parameter is inversely proportional to the first reconstruction error and the second reconstruction error, and the learning-perceived image block similarity parameter is directly proportional to the first reconstruction error and the second reconstruction error.

[0071] In an embodiment of the present specification, the detection device for generating an image also includes a classification model training module, which trains the classification model based on multiple image samples with label information and a preset loss function to obtain a trained classification model. The classification model training module includes: an image sample acquisition unit, which acquires multiple image samples with label information, where the image samples include real image samples and generated image samples, and the label information is used to mark the image samples as real image samples or generated image samples; a reconstructed image acquisition unit, which selects a partial image from each image sample, and reconstructs the partial image based on the remaining image other than the partial image in each image sample to obtain a reconstructed image consisting of the remaining image corresponding to each image sample and the reconstructed partial image; a model training unit, which uses multiple image samples as input data and the reconstruction effect category of the reconstructed images corresponding to the multiple image samples as output results, and trains the classification model based on a preset loss function to obtain a trained classification model.

[0072] In an embodiment of this specification, the image sample acquisition unit includes: a real image sample acquisition subunit, which acquires multiple real image samples based on an open source dataset; and a generated image sample generation subunit, which uses a generation model to generate multiple generated image samples based on multiple real image samples.

[0073] Those skilled in the art should understand that the above-mentioned detection device for generating an image can be used to implement the detection method for generating an image described above, and the detailed description should be similar to the description of the method part above. To avoid tediousness, it will not be repeated here.

[0074] The embodiment of the present specification provides a detection device for generating an image. First, a target image is acquired through an image acquisition module. Second, a partial image is selected from the target image using an image filling processing module. The partial image is reconstructed based on the remaining image in the target image except the partial image to obtain a reconstructed image composed of the remaining image and the reconstructed partial image. Then, the reconstructed image is input into a pre-trained classification model based on a classification module to obtain a reconstruction effect category of the reconstructed image. Finally, a generated image determination module is used to determine whether the target image is a real image or a generated image based on the reconstruction effect category of the reconstructed image. The embodiment of the present specification fully utilizes the principle that the smoothness difference between the real image and the generated image is large, and the large smoothness difference will lead to a large difference in the reconstruction effect of the reconstructed images corresponding to the real image and the generated image. By reconstructing a partial area of ​​the target image to obtain the reconstructed image corresponding to the target image, and by classifying the reconstruction effect of the reconstructed image, it is possible to quickly and effectively determine whether the current target image is a real image or a generated image, thereby effectively improving the detection efficiency of the generated image and the accuracy of the detection results. In addition, in the embodiment of this specification, the input data of the classification model is a reconstructed image. The classification model is based on the real image and its corresponding reconstructed image, and the reconstruction effect category is divided based on the generated image and its corresponding reconstructed image. Therefore, the images used for reconstruction effect classification have the same content. This method is conducive to reducing interference factors and enables the classification model to make category judgments based on the actual content gap of the image, thereby improving the accuracy of the classification results, and then improving the accuracy of the generated image detection results.

[0075] Based on the same idea, one or more embodiments of this specification also provide an electronic device, as shown in Figure 4. Electronic devices may have relatively large differences due to different configurations or performances, and may include one or more processors 301 and memory 302, and the memory 302 may store one or more storage applications or data. Among them, the memory 302 can be a temporary storage or a persistent storage. The application stored in the memory 302 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the electronic device. Furthermore, the processor 301 can be configured to communicate with the memory 302 to execute a series of computer-executable instructions in the memory 302 on the electronic device. The electronic device may also include one or more power supplies 303, one or more wired or wireless network interfaces 304, one or more input and output interfaces 305, and one or more keyboards 306.

[0076] Specifically, in this embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the electronic device, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following:

[0077] Acquire a target image, where the target image includes a real image and / or a generated image. The real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions.

[0078] Selecting a partial image from a target image, and reconstructing the partial image based on a remaining image in the target image excluding the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image;

[0079] Inputting the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, wherein the classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image;

[0080] According to the reconstruction effect category of the reconstructed image, it is determined whether the target image is a real image or a generated image.

[0081] One or more embodiments of this specification provide a storage medium for storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the following process is implemented:

[0082] Acquire a target image, where the target image includes a real image and / or a generated image. The real image is an image captured by an image acquisition device, and the generated image is an image generated based on preset conditions.

[0083] Selecting a partial image from a target image, and reconstructing the partial image based on a remaining image in the target image excluding the partial image, to obtain a reconstructed image composed of the remaining image and the reconstructed partial image;

[0084] Inputting the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, wherein the classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image;

[0085] According to the reconstruction effect category of the reconstructed image, it is determined whether the target image is a real image or a generated image.

[0086] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0087] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0088] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0089] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0090] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0091] It will be understood by those skilled in the art that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] One or more embodiments of this specification are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0093] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0095] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0096] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0097] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0098] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0099] One or more embodiments of this specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In distributed computing environments, program modules may be located in local and remote computer storage media, including storage devices.

[0100] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0101] The foregoing is merely one or more embodiments of this specification and is not intended to limit this application. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of the claims of one or more embodiments of this specification.

Claims

1. A detection method for generating an image, comprising: Obtaining a target image, where the target image includes a real image and / or a generated image, the real image is an image captured based on an image acquisition device, and the generated image is an image generated based on preset conditions; Selecting a partial image from the target image, and performing reconstruction processing on the partial image based on the remaining image in the target image except the partial image to obtain a reconstructed image composed of the remaining image and the reconstructed partial image; Inputting the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, where the classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image; Determining whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

2. The method according to claim 1, where the performing reconstruction processing on the partial image based on the remaining image in the target image except the partial image to obtain a reconstructed image composed of the remaining image and the reconstructed partial image includes: Performing masking processing on the selected partial image according to a preset mask to obtain a target image to be filled composed of the remaining image and the masked partial image; Performing image filling processing on the target image to be filled based on an image filling model to perform reconstruction processing on the partial image in the area where the mask is located based on the remaining image, and obtaining a reconstructed image corresponding to the target image.

3. The method according to claim 2, where the ratio of the size of the preset mask to the size of the target image is less than or equal to 1 / 4.

4. The method according to claim 1, where the performing reconstruction processing on the partial image based on the remaining image in the target image except the partial image to obtain a reconstructed image composed of the remaining image and the reconstructed partial image includes: Cropping out the partial image from the target image, and obtaining a target image to be filled based on the remaining image and the area where the partial image is located; Performing image filling processing on the target image to be filled based on a preset image filling algorithm to perform reconstruction processing on the partial image in the area where the partial image is located based on the remaining image, and obtaining a reconstructed image corresponding to the target image.

5. The method according to claim 1, where the first reconstruction error is determined according to a preset similarity index between the real image and its corresponding reconstructed image, the second reconstruction error is determined according to a preset similarity index between the generated image and its corresponding reconstructed image, and both the first reconstruction error and the second reconstruction error are inversely proportional to the smoothness of the target image.

6. The method according to claim 5, wherein the similarity metric includes: One or more of a structural similarity parameter, a peak signal-to-noise ratio parameter, and a learned perceptual image patch similarity parameter, wherein the structural similarity parameter is inversely proportional to the first reconstruction error and the second reconstruction error, the peak signal-to-noise ratio parameter is inversely proportional to the first reconstruction error and the second reconstruction error, and the learned perceptual image patch similarity parameter is directly proportional to the first reconstruction error and the second reconstruction error.

7. The method according to claim 1, wherein the training method of the classification model comprises: Obtaining a plurality of image samples with label information, wherein the image samples include real image samples and generated image samples, and the label information is used to label the image samples as real image samples or generated image samples; Selecting a part of the images from each image sample, and performing a reconstruction process on the part of the images based on the remaining images in each image sample except the part of the images, to obtain a reconstructed image composed of the remaining images and the reconstructed part of the images corresponding to each image sample; Using the plurality of image samples as input data and the reconstruction effect categories of the reconstructed images corresponding to the plurality of image samples as output results, training the classification model based on a preset loss function to obtain a trained classification model.

8. The method according to claim 7, wherein the obtaining of the plurality of image samples comprises: Obtaining a plurality of real image samples based on an open-source data set; Using a generation model to generate a plurality of generated image samples based on the plurality of real image samples.

9. A detection device for a generated image, comprising: An image acquisition module, which acquires a target image, wherein the target image includes a real image and / or a generated image, the real image is an image captured based on an image acquisition device, and the generated image is an image generated based on preset conditions; An image filling and processing module, which selects a part of the images from the target image, and performs a reconstruction process on the part of the images based on the remaining images in the target image except the part of the images, to obtain a reconstructed image composed of the remaining images and the reconstructed part of the images; A classification module, which inputs the reconstructed image into a pre-trained classification model to obtain a reconstruction effect category of the reconstructed image, and the classification model is used to classify the reconstructed image corresponding to the real image and the reconstructed image corresponding to the generated image according to a first reconstruction error between the real image and its corresponding reconstructed image and a second reconstruction error between the generated image and its corresponding reconstructed image; A generated image determination module, which determines whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

10. An electronic device, comprising: A processor; And A memory arranged to store computer-executable instructions, which when executed, can cause the processor to: Acquire a target image, wherein the target image includes a real image and / or a generated image, the real image is an image captured based on an image acquisition device, and the generated image is an image generated based on preset conditions; Select a partial image from the target image, and perform reconstruction processing on the partial image based on the remaining image in the target image except the partial image to obtain a reconstructed image composed of the remaining image and the reconstructed partial image; Input the reconstructed image into a pre-trained classification model to obtain the reconstruction effect category of the reconstructed image. The classification model is used to classify the reconstructed images corresponding to the real image and the generated image according to the first reconstruction error between the real image and its corresponding reconstructed image and the second reconstruction error between the generated image and its corresponding reconstructed image; Determine whether the target image is a real image or a generated image according to the reconstruction effect category of the reconstructed image.

Citation Information

Patent Citations

  • Image processing method and image processing device

    CN113538273A

  • Training method of image detection model, and image detection method and device

    CN116486199A

  • Generated image detection method and device

    CN117523323A

  • Image reconfiguration device, image reconfiguration method, and image reconfiguration program

    JP2017138829A