Artificial intelligence generated image identification method and device, terminal and storage medium

By using preset images to generate large models to extract features and fuse features, and construct a training data set of fused feature images carrying labels, the problem of insufficient generalization ability of existing AI-generated image identification networks is solved, and stronger identification accuracy and generalization ability are achieved.

CN120147792APending Publication Date: 2025-06-13SHENZHEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064533.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-13

Smart Images

  • Figure CN120147792A_ABST
    Figure CN120147792A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence generated image identification method and device, a terminal and a storage medium, and the method comprises the steps: carrying out the feature extraction of an obtained real image and an artificial intelligence generated image through a preset image generation large model, and carrying out the fusion of a plurality of feature images corresponding to each extracted image, labeling the fused feature images to construct a training data set of the fused feature images which respectively correspond to the real image and the artificial intelligence generated image and carry labels, and training an initial image identification model constructed based on a deep learning network by using the training data set, and inputting the target fused feature image corresponding to the to-be-detected image into the trained target image identification model to identify whether the to-be-detected image belongs to the artificial intelligence generated image. According to the method, the feature image extracted by the large model is generated based on the preset image to train the initial image identification model, so that the model has higher generalization ability, and the applicability of the model in different scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimedia information security, and particularly to a method, device, terminal and storage medium for identifying AI-generated images. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, especially the continuous progress of generative models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and diffusion models, the quality of AI-generated images has reached an unprecedented level. These generative models can generate realistic images based on given conditions or random noise and are widely used in fields such as art creation, entertainment, and advertising design. However, this progress in image generation technology also brings new challenges, especially in aspects such as image authentication, content review, copyright protection, and security monitoring. Therefore, the identification of AI-generated images has important practical significance.

[0003] In recent years, deep learning technology has made significant progress in the field of identifying AI-generated images, especially in identifying the subtle differences between generated images and original images. However, deep learning network models have the problem of insufficient generalization ability. Existing AI-generated image identification networks mainly rely on comparing the differences in the spectrograms of AI-generated images and real images for identification. Although the images generated by AI image generation tools or algorithms have differences in spectrograms from real images, there are still significant differences between the spectrograms of images generated by different AI image generation tools or algorithms. This makes the identification network need to cover as many AI generation tools or algorithms as possible in the training dataset to improve the identification effect.

[0004] However, existing identification networks are usually trained only on specific generation tools, algorithms, or datasets. Therefore, when faced with new AI image generation tools or unseen AI image generation algorithms, their identification performance is poor, that is, the existing identification networks for AI-generated image identification have the problem of weak generalization ability. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, device, terminal and storage medium for identifying AI-generated images to solve the problem of weak generalization ability of existing networks for identifying AI-generated images in view of the above-mentioned defects of the prior art.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] An artificial intelligence-generated image identification method, wherein the method includes:

[0008] Obtain a real image dataset and an artificial intelligence-generated image dataset, and use a preset image generation large model to perform multi-level feature extraction on the images in the real image dataset and the artificial intelligence-generated image dataset respectively to obtain multiple feature images corresponding to each image;

[0009] Fuse the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image, and label the fused feature image corresponding to each image to construct a training dataset containing the fused feature images with labels corresponding to real images and artificial intelligence-generated images respectively;

[0010] Use the training dataset to train an initial image identification model constructed based on a deep learning network to obtain a trained target image identification model;

[0011] Input the target fused feature image corresponding to the image to be tested into the target image identification model to identify whether the image to be tested belongs to the artificial intelligence-generated image to obtain a corresponding identification result.

[0012] In one implementation, the fusing the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image includes:

[0013] Fuse the multiple feature images corresponding to each image to obtain a two-dimensional fused feature image corresponding to each image;

[0014] Convert the two-dimensional fused feature image corresponding to each image into a three-dimensional fused feature image.

[0015] In one implementation, the fusing the multiple feature images corresponding to each image to obtain a two-dimensional fused feature image corresponding to each image includes:

[0016] Perform an averaging operation on each of the multiple feature images corresponding to each image to obtain an averaged feature image, and perform a re-averaging operation on the averaged feature image to obtain a two-dimensional fused feature image.

[0017] In one implementation, the converting the two-dimensional fused feature image corresponding to each image into a three-dimensional fused feature image includes:

[0018] Perform normalization processing on the two-dimensional fused feature image corresponding to each image to obtain a two-dimensional normalized feature image;

[0019] Convert the two-dimensional normalized feature image into a three-dimensional fused feature image using a color mapping table.

[0020] In one implementation, the step of performing label annotation on the fused feature image corresponding to each image to construct a training data set including the fused feature images with labels corresponding to real images and AI-generated images respectively includes:

[0021] Perform label annotation on the three-dimensional fused feature image corresponding to each image to obtain a training data set including the three-dimensional fused feature images with labels corresponding to real images and AI-generated images respectively.

[0022] In one implementation, the difference in the number of images between the real image data set and the AI-generated image data set is not greater than a preset number threshold.

[0023] In one implementation, after performing label annotation on the fused feature image corresponding to each image to construct a training data set including the fused feature images with labels corresponding to real images and AI-generated images respectively, it further includes:

[0024] Use data augmentation techniques to perform corresponding image variant operations on the fused feature images with labels in the training data set to obtain variant images with labels;

[0025] Construct a target training data set including the variant images with labels and the fused feature images with labels;

[0026] Among them, training an initial image discrimination model constructed based on a deep learning network using the training data set to obtain a trained target image discrimination model includes:

[0027] Train an initial image discrimination model constructed based on a deep learning network using the target training data set to obtain a trained target image discrimination model.

[0028] The present invention also discloses an AI-generated image discrimination device, where the device includes:

[0029] An image acquisition module for acquiring a real image data set and an AI-generated image data set;

[0030] A feature extraction module for performing multi-level feature extraction on the images in the real image data set and the AI-generated image data set respectively using a preset image generation large model to obtain multiple feature images corresponding to each image;

[0031] A feature fusion module for fusing multiple said feature images corresponding to each image to obtain a fused feature image corresponding to each image;

[0032] A training dataset construction module for label - annotating the fused feature images corresponding to each image to construct a training dataset containing the fused feature images with labels corresponding to real images and AI - generated images respectively;

[0033] A model training module for training an initial image discrimination model constructed based on a deep - learning network using the training dataset to obtain a trained target image discrimination model;

[0034] An image discrimination module for inputting the target fused feature image corresponding to a to - be - tested image into the target image discrimination model to discriminate whether the to - be - tested image belongs to the AI - generated image and obtain a corresponding discrimination result.

[0035] The present invention also discloses a terminal, which includes: a memory, a processor, and an AI - generated image discrimination program stored on the memory and executable on the processor. When the AI - generated image discrimination program is executed by the processor, the steps of the above - mentioned AI - generated image discrimination method are implemented.

[0036] The present invention also discloses a computer - readable storage medium, where the computer - readable storage medium stores a computer program that can be executed to implement the steps of the above - mentioned AI - generated image discrimination method.

[0037] An artificial intelligence generated image identification method, device, terminal and storage medium provided by the present invention. The artificial intelligence generated image identification method includes: obtaining a real image data set and an artificial intelligence generated image data set, and respectively performing multi-level feature extraction on the images in the real image data set and the artificial intelligence generated image data set by using a preset image generation large model to obtain multiple feature images corresponding to each image; fusing the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image, and performing label annotation on the fused feature image corresponding to each image to construct a training data set including the fused feature images with labels corresponding to real images and artificial intelligence generated images respectively; using the training data set to train an initial image identification model constructed based on a deep learning network to obtain a trained target image identification model; inputting the target fused feature image corresponding to the image to be tested into the target image identification model to identify whether the image to be tested belongs to the artificial intelligence generated image to obtain a corresponding identification result. It can be seen from this that the present invention respectively performs multi-level feature extraction on the obtained real images and artificial intelligence generated images by using an image generation large model, and then fuses the feature images of different levels and assigns labels to them to generate fused feature images with labels corresponding to real images and artificial intelligence generated images respectively. The fused feature image can more comprehensively represent the generation characteristics of the image. Then, by using these fused feature images with labels to train an initial image identification model constructed based on a deep learning network, not only the accuracy of model identification is significantly improved, but also the generalization ability of the model is greatly enhanced, which can effectively adapt to different types of artificial intelligence generated images and ensure good recognition performance and stability in various application scenarios. That is, the technical solution of the present application trains the initial image identification model based on the feature images extracted by the preset image generation large model, which can make the model have stronger generalization ability, can adapt to a variety of AI generation tools and generation algorithms, and significantly improves the applicability of the model in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flowchart of a preferred embodiment of the artificial intelligence generated image identification method in the present invention;

[0039] Figure 2 is a specific feature extraction schematic diagram disclosed in the present invention.

[0040] Figure 3 is a specific deep learning network schematic diagram disclosed in the present invention;

[0041] Figure 4 is a flowchart of a specific artificial intelligence generated image identification method disclosed in the present invention;

[0042] Figure 5 It is a schematic diagram of a specific feature image fusion disclosed by the present invention;

[0043] Figure 6 It is a flowchart of a specific method for identifying AI-generated images disclosed by the present invention;

[0044] Figure 7 It is a functional principle block diagram of a preferred embodiment of the AI-generated image identification device in the present invention;

[0045] Figure 8 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. Specific Embodiments

[0046] To make the objectives, technical solutions and advantages of the present invention clearer and more explicit, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] Please refer to Figure 1 , Figure 1 It is a flowchart of the method for identifying AI-generated images in the present invention. As Figure 1 shown, the method for identifying AI-generated images described in the embodiments of the present invention includes:

[0048] Step S11: Obtain a real image dataset and an AI-generated image dataset, and use a preset image generation large model to perform multi-level feature extraction on the images in the real image dataset and the AI-generated image dataset respectively to obtain multiple feature images corresponding to each image.

[0049] In this embodiment, a real image dataset and an AI-generated image dataset are obtained. For example, corresponding real images are selected from the known LSUN (Large-Scale Scene Understanding) dataset to obtain the real image dataset, that is, a subset of the LSUN dataset, and the known Crovi training set is obtained to obtain the AI-generated image dataset. Among them, the Crovi training set represents the AI-generated image set used in the AI-generated image identification method proposed by Crovi.

[0050] In this embodiment, after obtaining the real image dataset and the AI-generated image dataset, the images in the obtained real image dataset and AI-generated image dataset are used as image prompts and input into a preset image generation large model respectively for multi-level feature extraction, so as to obtain multiple feature images corresponding to each image output by the large model. That is, the preset image generation large model is used to perform feature extraction on the images to generate at least one feature image. That is to say, the feature extraction process is to use the large model to perform multi-scale feature extraction on the input images to capture feature images at different levels.

[0051] It should be noted that the preset image generation large model can extract representative feature information from the input images after pre-training, and the preset image generation large model has undergone multi-stage training during the pre-training process, which can improve the accuracy and generalization ability of feature extraction. Therefore, when using the image generation large model to perform multi-level feature extraction on the input images, different levels of feature information output by the large model can be obtained. Moreover, the preset image generation large model can be a deep learning network, specifically including but not limited to convolutional neural networks (CNNs), Transformer models or their variants, etc. These network architectures can effectively capture local and global features in the images and have strong expressive and generalization abilities.

[0052] To achieve feature extraction, an automatically written feature extraction program is called, the Stable Diffusion large model (i.e., the preset image generation large model) is downloaded, and the pre-trained image-variation-stable-diffusion weights are loaded. The image is input as an image prompt of the large model to perform feature extraction on the image.

[0053] For example, input images are selected from the real image dataset and the AI-generated image dataset, and the input images are input into the preset image generation large model as image prompts (image prompt) of the preset image generation large model, so as to obtain the outputs of at least one internal sub-module during the process of the preset image generation large model generating new images, that is, to obtain at least one multi-dimensional array. The multi-dimensional array contains the intermediate representation of the input images by the internal sub-modules in the preset image generation large model. The multi-dimensional array is converted into an image form to obtain at least one feature image. Specifically, as shown in Figure 2 For each real image or AI-generated image, the API of the Stable Diffusion large model is called for automatic feature extraction to obtain multiple feature images corresponding to each image. Figure 2On the left side in the middle is the network structure of the Stable Diffusion large model. The outputs of the two sub-modules marked with 64×64 inside above the structure are used as the extracted features and saved as jpg images to obtain the Figure 2 feature image shown on the right side in the middle.

[0054] It should also be noted that the difference in the number of images between the real image dataset and the artificial intelligence-generated image dataset is not greater than a preset number threshold, so that the number of real images and artificial intelligence-generated images in the subsequently constructed training dataset remains relatively balanced, avoiding the model from being biased towards a certain type of prediction and ensuring the fairness and accuracy of the discrimination results.

[0055] Step S12: Fuse the multiple feature images corresponding to each image to obtain the fused feature image corresponding to each image, and label the fused feature image corresponding to each image to construct a training dataset containing the fused feature images with labels corresponding to real images and artificial intelligence-generated images respectively.

[0056] In this embodiment, the feature images corresponding to each image are fused to generate a fused feature image. The fused feature image can be used for subsequent image classification, and the fusion process can include but is not limited to weighted average, splicing, convolution operation, etc. Then, the fused feature image corresponding to each image is labeled to obtain the fused feature image with a label, thereby constructing a training dataset including the fused feature images and corresponding labels corresponding to real images and artificial intelligence-generated images respectively. That is to say, the training dataset includes the fused feature images with labels corresponding to real images and artificial intelligence-generated images respectively. For example, the fused feature image corresponding to the real image is given the label 0_real, and the fused feature image corresponding to the AI-generated image is given the label 1_fake.

[0057] Step S13: Use the training dataset to train an initial image discrimination model constructed based on a deep learning network to obtain a trained target image discrimination model.

[0058] In this embodiment, the initial image discrimination model is trained using a training dataset that includes the fused feature images corresponding to the real images and the AI-generated images respectively, and the corresponding labels, to obtain the trained target image discrimination model. It can be understood that the preset image generation large model is used to perform multi-level feature extraction on the input image, and the multi-level feature images extracted are fused to generate the fused feature image. Subsequently, the fused feature image and its label are used to train the initial image discrimination model constructed based on the deep learning network, and finally, a target image discrimination model capable of automatically discriminating AI-generated images is obtained. That is, by introducing the image generation large model to extract features, more comprehensive image feature expressions can be achieved, and at the same time, the generalization ability of the features is significantly improved, so that the trained target image discrimination model can still maintain high-efficiency and stable discrimination performance when processing images generated by unseen generation tools or new generation algorithms, and is applicable to diverse actual application scenarios.

[0059] For example, as shown in Figure 3 The deep learning network can be Resnet34. The pre-trained weights of each module in the initial image discrimination model in the image classification task can be pre-loaded, and then the fused feature image and its corresponding label in the training dataset are used as training samples and input into the initial image discrimination model constructed based on the deep learning network for training to update the weights, so as to obtain the trained target image discrimination model.

[0060] It should be noted that loading the pre-trained weights of the initial image discrimination model in the image classification task can accelerate the model convergence process. At the same time, using the features already learned in the pre-trained weights can improve the model's ability to express image features, thereby further enhancing the accuracy and generalization of the discrimination task.

[0061] Step S14: Input the target fused feature image corresponding to the image to be tested into the target image discrimination model to discriminate whether the image to be tested belongs to the AI-generated image and obtain the corresponding discrimination result.

[0062] In this embodiment, the image to be tested is input as an image prompt into the preset image generation large model for feature extraction to obtain multiple target feature images corresponding to the image to be tested. Then, the multiple target feature images are subjected to the aforementioned feature fusion process to obtain the corresponding target fused feature image. Next, the target fused feature image corresponding to the image to be tested is input into the trained target image discrimination model to determine whether the image to be tested is an AI-generated image and obtain the corresponding discrimination result.

[0063] It can be seen that in the embodiments of the present invention, the image generation large model is used to perform multi-level feature extraction on the obtained real images and AI-generated images respectively. Then, by fusing the feature images of different levels and assigning labels to them, the fused feature images with labels corresponding to the real images and AI-generated images are generated. These fused feature images can more comprehensively characterize the generation characteristics of the images. Then, by using these fused feature images with labels to train the initial image discrimination model constructed based on the deep learning network, not only the accuracy of model discrimination is significantly improved, but also the generalization ability of the model is greatly enhanced, which can effectively adapt to different types of AI-generated images and ensure good recognition performance and stability in various application scenarios. That is, the technical solution of this application trains the initial image discrimination model based on the feature images extracted by the preset image generation large model, enabling the model to have stronger generalization ability, adapt to a variety of AI generation tools and generation algorithms, significantly improving the applicability of the model in different scenarios, and having a wide range of application prospects in the field of information security, such as having important application value in image forensics.

[0064] See Figure 4 As shown, the embodiments of the present invention disclose a specific method for discriminating AI-generated images. Compared with the previous embodiment, this embodiment further elaborates and optimizes the technical solution.

[0065] Step S21: Obtain a real image dataset and an AI-generated image dataset, and use the preset image generation large model to perform multi-level feature extraction on the images in the real image dataset and the AI-generated image dataset respectively, to obtain multiple feature images corresponding to each image.

[0066] Step S22: Fuse the multiple feature images corresponding to each image to obtain a two-dimensional fused feature image corresponding to each image, and convert the two-dimensional fused feature image corresponding to each image into a three-dimensional fused feature image.

[0067] In this embodiment, a large image generation model is used to perform multi-level feature extraction on an image to obtain feature images at different levels. For each image, at least one corresponding feature image is fused to obtain a two-dimensional fused feature image for each image. Specifically, an averaging operation is performed on each feature image among the multiple feature images corresponding to each image to obtain an averaged feature image, and then a re-averaging operation is performed on the averaged feature image to obtain a two-dimensional fused feature image. Then, the two-dimensional fused feature image corresponding to each image is converted into a three-dimensional fused feature image. Specifically, the two-dimensional fused feature image corresponding to each image is normalized to obtain a two-dimensional normalized feature image, and then the two-dimensional normalized feature image is converted into a three-dimensional fused feature image using a color mapping table. It can be understood that a preset large image generation model is used to perform feature extraction on an image, obtaining at least one corresponding feature image for each image, fusing the feature images, and converting them into three-dimensional fused feature images.

[0068] For example, referring to Figure 5 As shown, to achieve feature image fusion, a pre-written automatic feature image fusion program is called to average each feature image output by two sub-modules with 64×64 markings in the Stable Diffusion large model network structure, obtaining an averaged feature image. All the averaged feature images are stored in an array, and an averaging operation is performed on all the averaged feature images in the array to obtain a two-dimensional fused feature image. Then, the two-dimensional fused feature image is normalized, and the two-dimensional normalized feature image is converted into a three-dimensional fused feature image using a color mapping table (Viridis).

[0069] Step S23: Perform label annotation on the three-dimensional fused feature image corresponding to each image to obtain a training data set containing the three-dimensional fused feature images with labels corresponding to real images and AI-generated images respectively.

[0070] In this embodiment, the feature images corresponding to each image are fused to generate a three-dimensional fused feature image, and then label annotation is performed on the three-dimensional fused feature image corresponding to each image to obtain a three-dimensional fused feature image with a label, thereby constructing a training data set containing the three-dimensional fused feature images and corresponding labels corresponding to real images and AI-generated images respectively.

[0071] For example, the LSUN dataset contains millions of real images. 180,000 images are randomly selected from the LSUN dataset. Among them, the Crovi training set contains 180,000 images generated by the Stable Diffusion algorithm. Then, the subset of the LSUN dataset and each image in the Crovi training set are subjected to the aforementioned feature extraction and feature fusion processing, and are all converted into three-dimensional fused feature images. A 0_real label is assigned to the three-dimensional fused feature image corresponding to each real image in the LSUN dataset subset, and a 1_fake label is assigned to the three-dimensional fused feature image corresponding to each AI-generated image in the Crovi training set.

[0072] Step S24: Train the initial image discrimination model constructed based on the deep learning network using the training dataset to obtain a trained target image discrimination model.

[0073] Step S25: Input the target fused feature image corresponding to the image to be tested into the target image discrimination model to discriminate whether the image to be tested belongs to the AI-generated image to obtain a corresponding discrimination result.

[0074] For the specific content of the above Step S21 and Steps S24 to S25, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0075] It can be seen that in the embodiment of the present invention, the image generation large model is used to perform multi-level feature extraction on the obtained real images and AI-generated images respectively. Then, by fusing feature images of different levels and assigning labels to them, three-dimensional fused feature images with labels corresponding to real images and AI-generated images are generated respectively. The three-dimensional fused feature image can more comprehensively characterize the generation characteristics of the image. Furthermore, by using the three-dimensional fused feature images corresponding to the real images and AI-generated images in the constructed training dataset and their corresponding labels as training samples, the initial image discrimination model constructed based on the deep learning network is trained, so as to obtain a network model that can accurately identify AI-generated images. By utilizing the generalization of the features extracted by the image generation large model, the generalization of the image discrimination model is greatly improved. The trained target image discrimination model shows excellent performance when processing images generated by different AI generation tools or algorithms, and plays an important role in overcoming the limitation that the current research in the field of AI-generated image discrimination is limited to specific generation tools.

[0076] See Figure 6 As shown, the embodiment of the present invention discloses a specific method for discriminating AI-generated images. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution.

[0077] Step S31: Obtain a real image dataset and an AI-generated image dataset, and use a preset image generation large model to perform multi-level feature extraction on the images in the real image dataset and the AI-generated image dataset respectively, to obtain multiple feature images corresponding to each image.

[0078] Step S32: Fuse the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image, and perform label annotation on the fused feature image corresponding to each image, so as to construct a training dataset including the fused feature images with labels corresponding to real images and AI-generated images respectively.

[0079] Step S33: Use data augmentation technology to perform corresponding image variant operations on the fused feature images with labels in the training dataset to obtain variant images with labels.

[0080] In this embodiment, after constructing a training dataset including the fused feature images with labels corresponding to real images and AI-generated images respectively based on the fused feature images with labels, data augmentation technology can be used to perform corresponding image variant operations on the fused feature images with labels in the training dataset. Among them, the image variant operations can include but are not limited to operations such as rotation, flipping, cropping, color jittering, noise adding, etc., so as to obtain variant images corresponding to the fused feature images.

[0081] Step S34: Construct a target training dataset including the variant images with labels and the fused feature images with labels.

[0082] In this embodiment, after using data augmentation technology to perform corresponding image variant operations on the fused feature images with labels in the training dataset to obtain corresponding variant images, a target training dataset including these variant images and fused feature images is constructed, and these variant images also carry labels indicating that they are real images or AI-generated images. That is to say, in addition to the original fused feature images, the training dataset also includes variant images generated by data augmentation technology, so as to increase the quantity and diversity of training samples.

[0083] Step S35: Use the target training dataset to train an initial image discrimination model constructed based on a deep learning network to obtain a trained target image discrimination model.

[0084] In this embodiment, the initial image discrimination model is trained using the target training data including the original fused feature images and the variant images corresponding to the original fused feature images, so as to obtain a trained target image discrimination model.

[0085] Step S36: Input the target fused feature image corresponding to the image to be tested into the target image discrimination model to discriminate whether the image to be tested belongs to the AI-generated image and obtain the corresponding discrimination result.

[0086] For the specific contents of the above steps S31 to S23 and step S36, reference can be made to the corresponding contents disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0087] It can be seen that in the embodiments of the present invention, the image generation large model is used to perform multi-level feature extraction on the obtained real image and AI-generated image respectively, and then by fusing the feature images of different levels and assigning labels to them, the fused feature images with labels corresponding to the real image and AI-generated image are generated. The fused feature image can more comprehensively characterize the generation characteristics of the image. Then, by using these fused feature images with labels and their variant images to train the initial image discrimination model constructed based on the deep learning network, not only the accuracy of model discrimination is significantly improved, but also the generalization ability of the model is greatly enhanced, which can effectively adapt to different types of AI-generated images and ensure good recognition performance and stability in various application scenarios. That is, the technical solution of this application trains the initial image discrimination model based on the feature images extracted by the preset image generation large model, enabling the model to have stronger generalization ability, adapt to a variety of AI generation tools and generation algorithms, significantly improving the applicability of the model in different scenarios, and having a wide application prospect in the field of information security, such as having important application value in image forensics.

[0088] Among them, in order to illustrate that the AI-generated image discrimination technical solution of this application can play an important role in the image forensics method, the effectiveness of the AI-generated image discrimination method proposed in this application can be verified by comparing some currently disclosed AI-generated image discrimination methods, such as the AI-generated image discrimination method based on the CLIP (Contrastive Language-Image Pre-training) algorithm, the AI-generated image discrimination method based on the CNN algorithm, the AI-generated image discrimination method based on the DIRE (DIffusion Reconstruction Error) algorithm, and the AI-generated image discrimination method proposed by Crovi. That is, using the Crovi training set and a subset of the LSUN dataset as training samples for training, and using the Crovi test set to test the performance of each algorithm on the test set generated by multiple AI image generation tools or algorithms. The performance evaluation index used is the detection accuracy (%), and the relevant test results are shown in Table 1:

[0089] Table 1

[0090]

[0091]

[0092] Among them, ProGAN (Progressive Growing of Generative Adversarial Networks) represents an image generation tool or method based on the ProGAN algorithm; StyleGAN2 (Second-generation style-based GAN architecture) represents an image generation tool or method based on StyleGAN2; StyleGAN3 (Third-generation style-based GAN architecture) represents an image generation tool or method based on StyleGAN3; BigGAN (Big Generative Adversarial Networks) represents an image generation tool or method based on BigGAN; EG3D (Efficient geometry-aware 3d generative adversarial networks) represents an image generation tool or method based on EG3D; TamingTran (TamingTransformers) represents an image generation tool or method based on TamingTran; DALL E2 (the second generation of DALL E, which was launched by the American artificial intelligence non-profit organization OpenAI in January 2021) represents an artificial intelligence system that generates images according to written text; DALLE Mini (miniature DALL E, which was launched by the American artificial intelligence non-profit organization OpenAI in January 2021) represents another artificial intelligence system that generates images according to written text; GLIDE (Guided Language to Image Diffusion for Generation and Editing); ADM (Ablated Diffusion Model) represents an image generation tool or method based on ADM; Latent Diffusion large model represents an image generation tool or method based on the Latent Diffusion large model; Stable Diffusion large model represents an image generation tool or method based on the Stable Diffusion large model.

[0093] Moreover, as can be seen from Table 1 above, the CLIP method tends to misclassify most real images as AI-generated images, indicating that its ability to distinguish real images is weak; the CNN method shows high discrimination accuracy on images generated by Latent Diffusion and Stable Diffusion, but has low accuracy on other generation methods and limited generalization ability; the performance of the DIRE method is similar to that of CNN and also lacks generalization ability; the Crovi method performs mediocrely on most generation tools, without obvious shortcomings but also without showing significant advantages; while the method of the present invention performs excellently on multiple generation tools, and its average detection accuracy is the highest value, fully demonstrating its excellent comprehensive performance and strong generalization ability. That is, the artificial intelligence-generated image discrimination method provided in this application has the following significant advantages compared with existing AI-generated image discrimination methods, namely:

[0094] First, it has stronger generalization ability: Through multi-level feature extraction and fusion, the adaptability to various generation tools and algorithms is significantly improved.

[0095] Second, it has higher discrimination accuracy: It performs excellently on various AI image generation tools, and the average accuracy is significantly ahead of other methods.

[0096] Third, it has wide applicability: It is not only applicable to common generation algorithms, but can also effectively handle new generation tools, breaking through the limitations of existing methods.

[0097] In one embodiment, as Figure 7 shown, based on the above artificial intelligence-generated image discrimination method, the present invention also correspondingly provides an artificial intelligence-generated image discrimination device, including:

[0098] An image acquisition module 11, configured to acquire a real image data set and an artificial intelligence-generated image data set;

[0099] A feature extraction module 12, configured to use a preset image generation large model to perform multi-level feature extraction on the images in the real image data set and the artificial intelligence-generated image data set respectively, to obtain multiple feature images corresponding to each image;

[0100] A feature fusion module 13, configured to fuse the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image;

[0101] A training data set construction module 14, configured to perform label annotation on the fused feature image corresponding to each image, so as to construct a training data set including the fused feature images with labels corresponding to real images and artificial intelligence-generated images respectively;

[0102] A model training module 15, configured to train an initial image discrimination model constructed based on a deep learning network by using the training data set to obtain a trained target image discrimination model;

[0103] An image discrimination module 16, configured to input the target fused feature image corresponding to the image to be tested into the target image discrimination model to discriminate whether the image to be tested belongs to the artificial intelligence generated image to obtain a corresponding discrimination result.

[0104] Figure 8 The following is a schematic structural diagram of the terminal provided by the embodiment of the present application. The terminal may include:

[0105] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0106] When the processor 502 executes the program, it implements the artificial intelligence generated image discrimination method provided in the above embodiment.

[0107] Further, the terminal further includes:

[0108] A communication interface 503, configured to communicate between the memory 501 and the processor 502.

[0109] The memory 501 is used to store a computer program executable on the processor 502.

[0110] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0111] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0112] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.

[0113] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0114] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the artificial intelligence generated image identification method as described above is implemented.

[0115] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0116] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0117] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device.

[0118] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0119] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. An artificial intelligence generated image identification method, characterized in that: The method comprises: Acquire a real image dataset and an AI-generated image dataset, and use a preset image generation model to perform multi-level feature extraction on the images in the real image dataset and the AI-generated image dataset, respectively, to obtain multiple feature images corresponding to each image; Fusing the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image, and labeling the fused feature image corresponding to each image to construct a training data set containing fused feature images with labels corresponding to real images and artificial intelligence generated images, respectively; Using the training data set to train an initial image identification model built based on a deep learning network to obtain a trained target image identification model; The target fused feature image corresponding to the image to be tested is input into the target image identification model to identify whether the image to be tested belongs to the artificial intelligence generated image to obtain a corresponding identification result.

2. The artificial intelligence generated image identification method according to claim 1, characterized in that: The fusing the plurality of feature images corresponding to each image to obtain a fused feature image corresponding to each image includes: Fusing the plurality of feature images corresponding to each image to obtain a two-dimensional fused feature image corresponding to each image; The two-dimensional fused feature image corresponding to each image is converted into a three-dimensional fused feature image.

3. The artificial intelligence generated image identification method according to claim 2, characterized in that: The fusing of the plurality of feature images corresponding to each image to obtain a two-dimensional fused feature image corresponding to each image includes: An average operation is performed on each of the plurality of feature images corresponding to each image to obtain an averaged feature image, and a re-average operation is performed on the averaged feature image to obtain a two-dimensional fused feature image.

4. The artificial intelligence generated image identification method according to claim 2, characterized in that: The converting the two-dimensional fused feature image corresponding to each image into a three-dimensional fused feature image comprises: Normalizing the two-dimensional fused feature image corresponding to each image to obtain a two-dimensional normalized feature image; The two-dimensional normalized feature image is converted into a three-dimensional fused feature image using a color mapping table.

5. The artificial intelligence generated image identification method according to any one of claims 2 to 4, characterized in that: The labeling of the fused feature image corresponding to each image to construct a training data set containing fused feature images with labels corresponding to real images and artificial intelligence generated images, respectively, includes: The three-dimensional fused feature image corresponding to each image is labeled to obtain a training data set containing three-dimensional fused feature images with labels corresponding to real images and artificial intelligence generated images, respectively.

6. The artificial intelligence generated image identification method according to claim 5, characterized in that: The difference in the number of images between the real image dataset and the artificial intelligence generated image dataset is no greater than a preset number threshold.

7. The artificial intelligence generated image identification method according to claim 1, characterized in that: After labeling the fused feature image corresponding to each image to construct a training data set containing fused feature images with labels corresponding to real images and artificial intelligence generated images, the method further includes: Using data augmentation technology, performing corresponding image variant operations on the fused feature images carrying labels in the training data set to obtain variant images carrying labels; Constructing a target training data set including the variant images carrying the labels and the fused feature images carrying the labels; The training data set is used to train the initial image identification model constructed based on the deep learning network to obtain a trained target image identification model, including: The target training data set is used to train an initial image identification model constructed based on a deep learning network to obtain a trained target image identification model.

8. An artificial intelligence generated image identification device, characterized in that: The device comprises: Image acquisition module, used to acquire real image datasets and AI-generated image datasets; A feature extraction module, used to perform multi-level feature extraction on the images in the real image dataset and the artificial intelligence generated image dataset respectively using a preset image generation large model to obtain multiple feature images corresponding to each image; A feature fusion module, used for fusing the multiple feature images corresponding to each image to obtain a fused feature image corresponding to each image; A training data set construction module, used to label the fused feature image corresponding to each image, so as to construct a training data set containing fused feature images with labels corresponding to real images and artificial intelligence generated images respectively; A model training module, used to train an initial image identification model constructed based on a deep learning network using the training data set to obtain a trained target image identification model; The image identification module is used to input the target fused feature image corresponding to the image to be tested into the target image identification model to identify whether the image to be tested belongs to the artificial intelligence generated image to obtain a corresponding identification result.

9. A terminal, characterized in that: include: A memory, a processor, and an artificial intelligence generated image identification program stored in the memory and executable on the processor, wherein the artificial intelligence generated image identification program, when executed by the processor, implements the steps of the artificial intelligence generated image identification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed to implement the steps of the artificial intelligence generated image identification method as described in any one of claims 1 to 7.