Method, system and related device for identifying artificially generated images
By pre-training a deep learning neural network model on a public dataset and combining a feature pyramid network with a convolutional attention module, the problems of insufficient accuracy and speed in AI-generated image identification are solved, and efficient image identification under conditions of few samples is achieved.
Patent Information
- Application Number
- CN202510750066.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Existing technologies suffer from insufficient accuracy in identifying AI-generated images, slow processing speed, and difficulty in effectively training deep learning neural network models with few samples.
A two-stage training strategy is adopted. First, pre-training is performed using a large-scale public image dataset. Then, fine-tuning is performed using a small dataset from the target platform. A deep learning neural network model is constructed by combining a feature pyramid network and a convolutional attention module for image recognition.
It achieves efficient recognition of AI-generated images under limited sample conditions, with high accuracy and low computational overhead, making it suitable for practical application scenarios and improving the efficiency and accuracy of image identification.
Smart Images

Figure CN120298810B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image identification technology, and in particular to an identification method, system and related equipment for artificial intelligence-generated images. Background Art
[0002] With the rapid development of artificial intelligence (AI), AI technologies for generating realistic images have become widely used. The emergence of models such as generative adversarial networks (GANs) and diffusion models (DMs) has brought image generation capabilities to unprecedented heights. These models, capable of producing highly realistic images, have found widespread application not only in entertainment, design, and advertising, but are also gradually penetrating into fields such as film and television production and virtual reality. However, this technological advancement has also brought with it significant challenges in distinguishing true from false images. With the prevalence of fake images, effectively and rapidly verifying their authenticity has become a pressing issue.
[0003] While existing image authentication technologies, including deep learning-based detection models and traditional forensic methods, can detect traces of forged images to a certain extent, they still face numerous limitations when faced with increasingly sophisticated forgeries. Deep learning neural network models typically rely on massive amounts of data for adequate training. However, the number of AI-generated images available on target platforms for certain authentication tasks is relatively limited. Therefore, effectively training deep learning neural network models with such a small sample size presents a pressing challenge.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The present invention provides a method, system and related equipment for identifying images generated by artificial intelligence. The main purpose of the present invention is to solve the technical problems mentioned in the background technology of the existing technology.
[0006] A first aspect of the present invention provides a method for identifying an artificial intelligence-generated image, comprising:
[0007] Obtaining a first order of magnitude of AI-generated images and real images from a public image dataset as a pre-training dataset, where the first order of magnitude is greater than or equal to 10,000;
[0008] Obtain a second order of magnitude of AI-generated images from the target platform of the identification task as a fine-tuning dataset, where the second order of magnitude is greater than or equal to ten and less than or equal to one hundred;
[0009] Preprocessing the pre-training dataset and the fine-tuning dataset, wherein the preprocessing includes random cropping and color space conversion;
[0010] The deep learning neural network model for AI-generated image discrimination is constructed by combining a feature pyramid network and a convolutional attention module.
[0011] The deep learning neural network model is pre-trained by the pre-processed pre-training data set to obtain a pre-training model.
[0012] The pre-training model is fine-tuned by the pre-processed fine-tuning data set to obtain an AI-generated image discrimination model.
[0013] The AI-generated image discrimination model is used to discriminate the to-be-detected image on the target platform, and a discrimination result is output.
[0014] In an optional implementation of the first aspect of the present application, the AI-generated image discrimination model is used to discriminate the to-be-detected image on the target platform, and a discrimination result is output.
[0015] The to-be-detected image on the target platform is input into the AI-generated image discrimination model.
[0016] The feature pyramid network in the AI-generated image discrimination model is used to extract multi-level features from the to-be-detected image, and shallow high-resolution features and deep semantic features in the extracted multi-level features are fused to enhance the perception ability of different scale artifacts in the image.
[0017] The channel attention mechanism in the convolutional attention module is used to learn the channel weights of each level of features, and the response to high-frequency noise and color distortion features in the image is strengthened.
[0018] The spatial attention mechanism in the convolutional attention module is used to generate a spatial mask for each level of features to locate abnormal areas in the image.
[0019] The inherent noise pattern of the image is extracted by the SRM filter and frequency domain analysis, and the deviation of the image content from the priori of the real world is detected by the visual unit, and finally the discrimination result is output by the joint discriminator.
[0020] In an optional implementation of the first aspect of the present application, the pre-processing of the pre-training data set and the fine-tuning data set includes:
[0021] For each training image in the pre-training data set and the fine-tuning data set, an image block of a preset specification is randomly cropped from the training image, and if the original size of the training image is smaller than the preset specification, the training image is duplicated and spliced into the image block of the preset specification.
[0022] convert the image block from an RGB space to a YCbCr space to obtain an input image for model training.
[0023] In an optional implementation of the first aspect of the present application, the deep learning neural network model for AI-generated image discrimination is constructed by combining the feature pyramid network and the convolution attention module, and comprises:
[0024] The feature pyramid network extracts features from different convolution layers and fuses them to generate multi-scale feature maps to capture information at different levels in the input image.
[0025] The convolution attention module further enhances the features extracted by the feature pyramid network and outputs classification labels for discrimination.
[0026] In an optional implementation of the first aspect of the present application, the pre-trained model is obtained by pre-training the deep learning neural network model using the pre-processed pre-training data set, and comprises:
[0027] The multi-scale fused features of the pre-processed pre-training data set are input into the feature pyramid network of the deep learning neural network model to obtain multi-scale fused features of the pre-training image.
[0028] The multi-scale fused features are input into the classification network of the convolution attention module integrated with the channel attention mechanism and the spatial attention mechanism to obtain pre-training classification results.
[0029] The pre-training classification results and image labels are used to calculate a loss function, and the model parameters are updated through backpropagation until the loss function converges to a preset target, obtaining the pre-training model.
[0030] In an optional implementation of the first aspect of the present application, the AI-generated image discrimination model is obtained by fine-tuning the pre-training model using the fine-tuning data set after pre-processing, and comprises:
[0031] The fine-tuning data set after pre-processing is input into the pre-training model.
[0032] Each image of the fine-tuning data set is discriminated by the pre-training model.
[0033] The discrimination results of each image of the fine-tuning data set are obtained and the discrimination accuracy is calculated.
[0034] The weight parameters of the network in the pre-training model are adjusted based on the discrimination accuracy, so that the pre-training model transitions from general features for discriminating AI-generated images to specific task features for discriminating AI-generated images on the target platform, obtaining an AI-generated image discrimination model.
[0035] In an optional implementation of the first aspect of the present application, the disclosed image dataset includes a Genimage image dataset, a DiFF image dataset, and a Fake2M image dataset, and the target platform includes an Adobe Firefly platform, a Bing platform, a Canva platform, a Dreamstudio platform, a JasperArt platform, a Midjourney platform, a Nightcafe platform, a Playground platform, and a Prodia platform.
[0036] The second aspect of the present application provides an artificial intelligence generated image identification system, which comprises:
[0037] A pre-training dataset acquisition module is configured to acquire a first order of magnitude of AI generated images and real images from a disclosed image dataset as a pre-training dataset, wherein the first order of magnitude is greater than or equal to ten thousand;
[0038] A fine-tuning dataset acquisition module is configured to acquire a second order of magnitude of AI generated images from a target platform of an identification task as a fine-tuning dataset, wherein the second order of magnitude is greater than or equal to ten and less than or equal to one hundred;
[0039] A dataset preprocessing module is configured to preprocess the pre-training dataset and the fine-tuning dataset, wherein the preprocessing includes random cropping and color space conversion;
[0040] An identification model construction module is configured to combine a feature pyramid network and a convolutional attention module to construct a deep learning neural network model for AI generated image identification;
[0041] A model pre-training module is configured to pre-train the deep learning neural network model through the preprocessed pre-training dataset to obtain a pre-training model;
[0042] A model fine-tuning module is configured to fine-tune the pre-training model through the preprocessed fine-tuning dataset to obtain an AI generated image identification model;
[0043] A platform image identification module is configured to identify a to-be-detected image on the target platform based on the AI generated image identification model and output an identification result.
[0044] The third aspect of the present application provides an artificial intelligence generated image identification device, which comprises a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected through a circuit;
[0045] The at least one processor calls the instructions in the memory to enable the artificial intelligence-generated image identification device to perform the artificial intelligence-generated image identification method as described in any one of the first aspects of the present invention.
[0046] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for identifying an artificial intelligence-generated image as described in any one of the first aspects of the present invention is implemented.
[0047] Beneficial effects: The present invention provides a method, system and related equipment for identifying images generated by artificial intelligence. The method adopts a two-stage training strategy: first, the image identification model is pre-trained based on a public image dataset, and then fine-tuned using a small amount of target platform image data to achieve model adaptation and accurate detection. In the image preprocessing stage, the model's adaptability to image size is improved through random cropping, and color space conversion is performed to improve feature recognition. In terms of network architecture, the present invention combines feature pyramids to achieve multi-scale feature fusion, and embeds a convolution module with an attention mechanism to focus on key areas, thereby reducing the number of parameters while ensuring detection accuracy. The technical solution of the present invention does not require the target platform to have a large amount of training data, and can efficiently identify false images generated by the target platform. It is suitable for actual application scenarios, has high accuracy and low computational overhead, and provides technical support for the identification of AI-generated images. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of an embodiment of a method for identifying an artificial intelligence-generated image according to the present invention;
[0049] Figure 2 A schematic diagram of an embodiment of a convolutional block attention module of the present invention;
[0050] Figure 3 A schematic diagram of an embodiment of a channel attention mechanism of the present invention;
[0051] Figure 4 A schematic diagram of an embodiment of a spatial attention mechanism of the present invention;
[0052] Figure 5 A schematic diagram of an embodiment of a feature pyramid network according to the present invention;
[0053] Figure 6 A schematic diagram of an embodiment of a classification network based on a convolutional block attention module of the present invention;
[0054] Figure 7 A schematic diagram of an embodiment of an artificial intelligence-generated image identification system of the present invention;
[0055] Figure 8 The figure is a schematic diagram of an embodiment of an artificial intelligence-generated image identification device of the present invention. DETAILED DESCRIPTION
[0056] The terms "first," "second," "third," "fourth," and so on (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that shown or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or apparatus.
[0057] Existing image authentication technologies, including deep learning-based detection models and traditional forensic methods, can detect traces of forged images to a certain extent, but they still face numerous limitations when faced with increasingly sophisticated forgeries. First, processing speeds are slow, especially when performing real-time authentication on large-scale image data, often failing to meet the demands of practical applications. Second, authentication accuracy is also insufficient, prone to misjudgment when faced with high-precision forged images. Furthermore, the diversity of different generative models and the high fidelity of the generated images complicate the development of universal authentication technologies.
[0058] Deep learning neural network models, especially large-scale ones, typically rely on massive amounts of data for thorough training to achieve excellent generalization and accuracy. However, for many companies and platforms, obtaining sufficient large-scale annotated data is both time-consuming and labor-intensive. In some areas, obtaining high-quality data resources is even extremely difficult. Furthermore, the number of images available from target platforms is relatively limited, making effective training of deep learning neural network models with such a small sample size a pressing challenge.
[0059] To address the above technical issues, this paper proposes a novel deep learning neural network model. This model, with the limited number of training samples, can be pre-trained using a large-scale pre-training dataset, enabling it to learn common features relevant to AI-generated image identification. The pre-trained model is then further fine-tuned using a fine-tuning dataset, ultimately resulting in an AI-generated image identification model. This model, with its lightweight architecture and strong learning capabilities, demonstrates superior performance in distinguishing AI-generated images from real ones.
[0060] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 The first aspect of the present application is a method for identifying AI-generated images, comprising:
[0061] S100, obtaining a first order of magnitude of AI-generated images and real images from a public image dataset as a pre-training dataset, the first order of magnitude being greater than or equal to ten thousand; in the present application, the public image dataset includes Genimage image dataset, DiFF image dataset and Fake2M image dataset.
[0062] In the present application, in order to make the model learn large-scale and diversified data, and make it learn general features and patterns in the pre-training process, the pre-training dataset needs to contain feature information of several AI-generated images. Therefore, a large number of AI-generated images and real images in the public dataset are used as the pre-training dataset.
[0063] In an exemplary embodiment of the present application, the AI-generated images can specifically select SDv14, Glide, Midjourney, ADM and VQDM in the public dataset Genimage; SDXL, SD_Refiner and HPS in the public dataset DiFF; IF_v1.0 in the public dataset Fake2M, and the real images select ImageNet in the public dataset Genimage. Each image category in each of the above datasets collects one hundred thousand images. The results are shown in Table 1.
[0064] Table 1 Composition of pre-training dataset
[0065]
[0066] S200, obtaining a second order of magnitude of AI-generated images from the target platform of the identification task as a fine-tuning dataset, the second order of magnitude being greater than or equal to ten and less than or equal to one hundred; the target platform includes Adobe Firefly platform, Bing platform, Canva platform, Dreamstudio platform, JasperArt platform, Midjourney platform, Nightcafe platform, Playground platform and Prodia platform.
[0067] In the present application, in order to further train the pre-trained model for specific tasks, so that the model can transition from the pre-trained general features and patterns to task-specific features and models, the fine-tuning dataset usually contains highly relevant data for the target task. Therefore, in order to identify AI-generated images for the target platform, a small amount of AI-generated images from the target platform of the identification task are needed as the fine-tuning dataset.
[0068] In an exemplary embodiment of the present application, fifty AI-generated images are obtained from each of the nine target platforms of Adobe Firefly, Bing, Canva, Dreamstudio, JasperArt, Midjourney, Nightcafe, Playground, and Prodia. The results are shown in Table 2.
[0069] Table 2 Composition of fine-tuning dataset
[0070]
[0071] S300, pre-processing the pre-training dataset and the fine-tuning dataset, the pre-processing including random cropping and color space conversion; in an optional embodiment of step S300 of the present application, the pre-processing of the pre-training dataset and the fine-tuning dataset includes: for each training image in the pre-training dataset and the fine-tuning dataset, randomly cropping an image block of a preset specification from the training image, if the original size of the training image is smaller than the preset specification, then replicating the training image to splice into the image block of the preset specification; converting the image block from RGB space to YCbCr space to obtain an input image for model training.
[0072] Specifically, an exemplary embodiment of step S300 of the present application is as follows: random cropping operation: the random cropping operation refers to randomly cropping an image into a 256*256 image block, and if the image is smaller than 256*256, then replicating the image to splice into a 256*256 image. The random cropping operation not only makes the model adaptive to image sizes, thereby enhancing the model's processing capability for input images of different sizes, but also increases the diversity of images, enabling the model to have stronger learning ability for the features of AI-generated images. Space conversion operation: the space conversion operation refers to converting an image from RGB space to YCbCr space. Since image compression is performed in YCbCr space, learning features in YCbCr space helps to improve the robustness of the model.
[0073] S400, combining a feature pyramid network and a convolutional attention module to construct a deep learning neural network model for AI-generated image discrimination; in an optional embodiment of step S400 of the present application, the combination of the feature pyramid network and the convolutional attention module to construct a deep learning neural network model for AI-generated image discrimination includes: extracting features from different convolutional layers and fusing them through the feature pyramid network to generate multi-scale feature maps to capture information at different levels in the input image; further enhancing the features extracted by the feature pyramid network through the convolutional attention module to output classification labels for discrimination.
[0074] Specifically, the model used in the present invention is a deep learning neural network model, which has a lightweight architecture and strong learning ability. The feature pyramid extracts features from different convolutional layers and fuses them to generate multi-scale feature maps, capturing information at different levels of the image. By combining low-level features (containing high-resolution information) and high-level features (containing high semantic information), the feature pyramid improves the accuracy of target detection. Low-level features are particularly important for small target detection, while high-level features are helpful for large target detection. The comprehensive improvement of the model's feature extraction capabilities for images of different sizes. The convolutional attention module is responsible for further enhancing and classifying the features extracted by the feature pyramid network. By applying the convolutional block attention module (CBAM, such as Figure 2 As shown), that is, the channel attention mechanism (CA, as Figure 3 as shown) and spatial attention mechanism (SA, as Figure 4 As shown in Figure 2, we can focus on more important features and improve the performance of convolutional neural networks and the accuracy of the model.
[0075] S500, pre-training the deep learning neural network model using the pre-processed pre-training data set to obtain a pre-training model; in an optional implementation of step S500 of the present invention, pre-training the deep learning neural network model using the pre-processed pre-training data set to obtain a pre-training model includes: inputting the pre-processed pre-training data set into the feature pyramid network of the deep learning neural network model to obtain multi-scale fusion features of the pre-training image; inputting the multi-scale fusion features into a classification network of a convolutional attention module that integrates a channel attention mechanism and a spatial attention mechanism to obtain a pre-training classification result; calculating a loss function using the pre-training classification result and the image label, and updating the model parameters through back propagation until the loss function converges to a preset target to obtain the pre-training model.
[0076] Specifically, the present invention pre-trains the deep learning neural network model using the pre-training dataset to obtain the pre-trained deep learning neural network model. The main process is as follows: the pre-processed pre-training dataset image is input into the feature pyramid network (FPN, such as Figure 5 As shown in ), the feature information of the image is obtained; the feature information of the image is input into the convolutional attention module (CNN with CBAM, as shown in Figure 6The classification result is obtained by using the classification result and the image label to calculate a loss function, and the model parameters are updated through back propagation to finally converge to obtain a pre-training model. Through pre-training of the model on a large-scale pre-training data set in advance, the application learns general feature representation, and then migrates these features to the target task, thereby improving the performance of the model. The main functions of pre-training include: improving performance on a small amount of data through feature migration and fine-tuning; speeding up the training process, significantly reducing the computing resources and time cost; solving the problem of insufficient data, enabling the model to perform well under limited data conditions; accelerating the convergence speed of the model and improving the overall performance.
[0077] S600, fine-tuning the pre-training model through the pre-processed fine-tuning data set to obtain an AI-generated image discrimination model; in an optional embodiment of step S600 of the application, fine-tuning the pre-training model through the pre-processed fine-tuning data set to obtain an AI-generated image discrimination model comprises: inputting the pre-processed fine-tuning data set into the pre-training model; discriminating each image of the fine-tuning data set through the pre-training model; obtaining the discrimination result of each image of the fine-tuning data set and counting the discrimination accuracy; adjusting the weight parameters of the network in the pre-training model based on the discrimination accuracy, so that the pre-training model transitions from general features for discriminating AI-generated images to specific task features for discriminating AI-generated images of the target platform, and obtains an AI-generated image discrimination model.
[0078] Specifically, the purpose of step S600 of the application is to further fine-tune the pre-training model on the fine-tuning data set, so that it transitions from learned general features to features adapted to specific tasks, i.e. features of AI-generated images of the target platform. Fine-tuning can effectively solve the problem of incomplete matching between the pre-training model and the specific task. Especially in the case of limited data, through careful adjustment of the model weight, the model can be optimized for the target task, improving performance and accuracy, thereby achieving better results.
[0079] S700, based on the AI-generated image discrimination model, the target platform on the image to be detected is discriminated, and the discrimination result is output. In an optional embodiment of step S700 of the present application, the AI-generated image discrimination model is used to discriminate the image to be detected on the target platform, and the discrimination result is output, which includes: inputting the image to be detected on the target platform into the AI-generated image discrimination model; performing multi-level feature extraction on the image to be detected through the feature pyramid network in the AI-generated image discrimination model, and fusing the shallow high-resolution features and deep semantic features extracted from the multi-level features to enhance the perception ability of different scale artifacts in the image; learning the channel weight of each level feature through the channel attention mechanism in the convolution attention module, and strengthening the response to high-frequency noise and color distortion features in the image; generating a spatial mask for each level feature through the spatial attention mechanism in the convolution attention module, and locating the abnormal area in the image; extracting the inherent noise pattern of the image through the SRM filter and frequency domain analysis, and using the visual unit to detect the deviation of the image content from the priori in the real world, and finally outputting the discrimination result through the joint discriminator.
[0080] Specifically, in the feature pyramid network of the AI-generated image discrimination model of the present application, the encoder gradually extracts multi-level features (such as low-level edge / texture, medium-level semantic, and high-level global information) through down-sampling, and the decoder restores the resolution through up-sampling and connects with the encoder features horizontally to supplement the detail information. At each level, the multi-level features obtained by up-sampling and down-sampling are connected horizontally, the shallow high-resolution features are fused with the deep semantic features, and the perception ability of different scale artifacts (such as local texture abnormalities and global light and shadow inconsistencies) is enhanced. In the convolution attention enhancement module, the channel attention mechanism is embedded in each level of the feature pyramid, dynamically learns the channel weight, and strengthens the response of key features such as high-frequency noise and color distortion in the generated image; the spatial attention mechanism generates a spatial mask through a convolution layer to locate the abnormal area (such as unnatural facial features and blurred object edges), and suppresses irrelevant background interference. When classifying images, high-frequency components are extracted from shallow features (such as DCT transform or SRM filter) to capture statistical noise patterns in generated images, analyze low-level features, extract global semantic embeddings using visual units (such as CLIP) to detect object co-occurrence logic errors (such as hand deformities and light and shadow direction contradictions) to implement high-level semantic analysis, and finally concatenate the high-level semantic features and low-level resolution features in the channel dimension, and output the classification probability through the multi-layer perception (MLP).
[0081] In order to better illustrate the effect of the technical scheme of the present application, the present application is compared with the existing model in the following aspects. The image discrimination model of the present application using the feature pyramid network and the convolution attention module has the advantage of lightweight compared with the mainstream Convnext_tiny model. As shown in Table 3, under the same hardware conditions, the self-developed model of the present application has less parameters.
[0082] Table 3 Comparison of parameter quantity between the model of the present application and the mainstream classification model Convnext_tiny
[0083]
[0084] As shown in Table 4, the processing time of the image discrimination model of the present application on a single image on a CPU-only device is less than 3 seconds, and the memory occupation is 0.8 GB. On a device with GPU, the processing time of a single image is less than one second, the memory occupation is 1.4 GB, and the video memory occupation is 0.4 GB, which can meet the performance requirements of most users.
[0085] Table 4 Processing time and resource occupation of the model of the present application on CPU and GPU for processing images of different resolutions
[0086]
[0087] In the present embodiment, the discrimination model of the present application has strong learning ability and stronger AI-generated image discrimination ability. Compared with the mainstream classification model Convnext_tiny, the accuracy is improved by at least 15%, which is obviously advantageous. Compared with another AI-generated image discrimination platform AI or Not, whether it is the accuracy of AI-generated images or the accuracy of discriminating real images, the discrimination accuracy of the discrimination model of the present application is about 5 percentage points higher. The results are shown in Table 5.
[0088] Table 5 Comparison of discrimination accuracy between the self-developed model of the present application and the classification model Convnext_tiny and the discrimination platform AI or Not
[0089]
[0090] From the above, the model based on the deep learning neural network model and adopting the few-shot transfer learning strategy is trained. The model has stronger learning ability for the features of the AI generated image through random cropping and spatial conversion operations on the image. And through the application of the feature pyramid module (FPN) and the convolution block attention module (CBAM), the model has stronger feature extraction ability and classification ability. Compared with the mainstream classification model Convnext, the model provided by the application has a lighter architecture and stronger learning ability. Compared with the mainstream AI generated image identification platform AI or Not, the model has superior performance in the identification task of AI generated images and real images.
[0091] Referring to Figure 7 The second aspect of the application provides an artificial intelligence generated image identification system, which comprises:
[0092] A pre-training data set acquisition module 10 is configured to acquire a first order of magnitude of AI generated images and real images from a public image data set as a pre-training data set, wherein the first order of magnitude is greater than or equal to ten thousand;
[0093] A fine-tuning data set acquisition module 20 is configured to acquire a second order of magnitude of AI generated images from a target platform of an identification task as a fine-tuning data set, wherein the second order of magnitude is greater than or equal to ten and less than or equal to one hundred;
[0094] A data set preprocessing module 30 is configured to preprocess the pre-training data set and the fine-tuning data set, wherein the preprocessing includes random cropping and color space conversion;
[0095] An identification model construction module 40 is configured to combine a feature pyramid network and a convolution attention module to construct a deep learning neural network model for AI generated image identification;
[0096] A model pre-training module 50 is configured to pre-train the deep learning neural network model through the preprocessed pre-training data set to obtain a pre-training model;
[0097] A model fine-tuning module 60 is configured to fine-tune the pre-training model through the preprocessed fine-tuning data set to obtain an AI generated image identification model;
[0098] A platform image identification module 70 is configured to identify the to-be-detected images on the target platform based on the AI generated image identification model and output an identification result.
[0099] In an optional embodiment of the second aspect of the application, the platform image identification module comprises:
[0100] An image to be detected input unit is configured to input an image to be detected on a target platform to the AI-generated image discrimination model;
[0101] A multi-level feature extraction unit is configured to perform multi-level feature extraction on the image to be detected by the feature pyramid network in the AI-generated image discrimination model, and fuse shallow high-resolution features and deep semantic features in the extracted multi-level features to enhance the perception ability of different scale artifacts in the image;
[0102] A channel attention processing unit is configured to learn channel weights of each level feature by a channel attention mechanism in the convolution attention module, and strengthen the response to high-frequency noise and color distortion features in the image;
[0103] A spatial attention processing unit is configured to generate a spatial mask for each level feature by a spatial attention mechanism in the convolution attention module, and locate abnormal areas in the image;
[0104] A joint discrimination unit is configured to extract the inherent noise pattern of the image by an SRM filter and frequency domain analysis, and detect the deviation of the image content from the priori of the real world by a visual unit, and finally output a discrimination result by a joint discriminator.
[0105] In an optional implementation of the second aspect of the present application, the data set preprocessing module comprises:
[0106] A random cropping unit is configured to randomly crop an image block of a preset specification from each training image in the pre-training data set and the fine-tuning data set, and if the original size of the training image is smaller than the preset specification, the training image is replicated and spliced into the image block of the preset specification.
[0107] A color space conversion unit is configured to convert the image block from RGB space to YCbCr space to obtain an input image for model training.
[0108] In an optional implementation of the second aspect of the present application, the discrimination model construction module comprises:
[0109] A feature pyramid network unit is configured to extract features from different convolution layers and fuse them by the feature pyramid network to generate multi-scale feature maps to capture information at different levels in the input image.
[0110] A convolution attention unit is configured to further enhance the features extracted by the feature pyramid network by the convolution attention module, and output classification labels for discrimination.
[0111] In an optional implementation of the second aspect of the present application, the model pre-training module comprises:
[0112] a feature pyramid network processing unit configured to input the preprocessed pre-training data set into the feature pyramid network of the deep learning neural network model to obtain multi-scale fusion features of pre-training images;
[0113] a convolutional attention processing unit configured to input the multi-scale fusion features into a classification network of a convolutional attention module integrated with a channel attention mechanism and a spatial attention mechanism to obtain pre-training classification results;
[0114] a loss calculation set back propagation unit configured to calculate a loss function by using the pre-training classification results and image labels, and update model parameters through back propagation until the loss function converges to a preset target to obtain the pre-training model.
[0115] In an optional implementation of the second aspect of the present application, the model fine-tuning module comprises:
[0116] a data set input unit configured to input the preprocessed fine-tuning data set into the pre-training model;
[0117] an image discrimination unit configured to discriminate each image of the fine-tuning data set by using the pre-training model;
[0118] a structure and statistics unit configured to obtain discrimination results of each image of the fine-tuning data set and to statistically analyze discrimination accuracy;
[0119] a weight parameter adjustment unit configured to adjust the weight parameters of the network in the pre-training model based on the discrimination accuracy, so that the pre-training model is transitioned from discriminating AI-generated images based on general features to discriminating AI-generated images of the target platform based on specific task features, thereby obtaining an AI-generated image discrimination model.
[0120] In an optional implementation of the second aspect of the present application, the public image data set comprises a Genimage image data set, a DiFF image data set, and a Fake2M image data set, and the target platform comprises an Adobe Firefly platform, a Bing platform, a Canva platform, a Dreamstudio platform, a JasperArt platform, a Midjourney platform, a Nightcafe platform, a Playground platform, and a Prodia platform.
[0121] Figure 8Fig. 1 is a schematic diagram of an apparatus for identifying AI-generated images according to an embodiment of the present application. The apparatus for identifying AI-generated images can have a large difference due to different configurations or performances, and can include one or more processors 80 (central processing units, CPUs) (e.g., one or more processors) and a memory 90, and one or more storage media 100 (e.g., one or more mass storage devices) for storing applications or data. The memory and the storage media can be temporary storage or persistent storage. The programs stored in the storage media can include one or more modules (not shown in the figure), and each module can include a series of instruction operations for the apparatus for identifying AI-generated images. Further, the processor can be configured to communicate with the storage media and execute the series of instruction operations in the storage media on the apparatus for identifying AI-generated images.
[0122] The apparatus for identifying AI-generated images can further include one or more power supplies 110, one or more wired or wireless network interfaces 120, one or more input / output interfaces 130, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that the apparatus for identifying AI-generated images can include more or fewer components than those shown, or combine some components, or arrange different components. Figure 8 The structure of the apparatus for identifying AI-generated images shown does not constitute a limitation on the apparatus for identifying AI-generated images, and can include more or fewer components than those shown, or combine some components, or arrange different components.
[0123] The present application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium has instructions stored therein, and when the instructions are executed on a computer, the computer performs the steps of the apparatus for identifying AI-generated images.
[0124] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system or system, unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0125] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, an artificial intelligence generated image identification device, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0126] The above-described embodiments are merely used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying images generated by artificial intelligence, characterized in that: include: Obtaining a first order of magnitude of AI-generated images and real images from a public image dataset as a pre-training dataset, where the first order of magnitude is greater than or equal to 10,000; Obtain a second order of magnitude of AI-generated images from the target platform of the identification task as a fine-tuning dataset, where the second order of magnitude is greater than or equal to ten and less than or equal to one hundred; Preprocessing the pre-training dataset and the fine-tuning dataset, wherein the preprocessing includes random cropping and color space conversion; Combining the feature pyramid network with the convolutional attention module, a deep learning neural network model for AI-generated image identification was constructed; Pre-training the deep learning neural network model using the pre-processed pre-training data set to obtain a pre-training model; Fine-tune the pre-trained model using the pre-processed fine-tuning dataset to obtain an AI-generated image identification model; Identify the image to be detected on the target platform based on the AI-generated image identification model and output an identification result; The image to be detected on the target platform is identified based on the AI-generated image identification model, and the identification result is outputted, including: Inputting the image to be detected on the target platform into the AI-generated image identification model; Performing multi-level feature extraction on the image to be detected through the feature pyramid network in the AI-generated image identification model, and fusing the shallow high-resolution features and deep semantic features extracted from the multi-level features to enhance the perception of artifacts of different scales in the image; The channel attention mechanism in the convolutional attention module learns the channel weights of features at each level to enhance the response to high-frequency noise and color distortion features in the image; Generate spatial masks for features at each level through the spatial attention mechanism in the convolutional attention module to locate abnormal areas in the image; The inherent noise pattern of the image is extracted through SRM filter and frequency domain analysis, and the deviation of the image content from the real world prior is detected by the visual unit, and finally the identification result is output through the joint discriminator; Fine-tuning the pre-trained model using the pre-processed fine-tuning dataset to obtain an AI-generated image identification model includes: Inputting the pre-processed fine-tuning dataset into the pre-training model; Identify each image of the fine-tuning dataset using the pre-trained model; Obtaining identification results for each image in the fine-tuning dataset and calculating the identification accuracy; Based on the identification accuracy, the weight parameters of the network in the pre-trained model are adjusted so that the pre-trained model transitions from identifying the general features of AI-generated images to adapting to the specific task features of identifying the AI-generated images of the target platform, thereby obtaining an AI-generated image identification model.
2. The method for identifying artificial intelligence-generated images according to claim 1, characterized in that: The preprocessing of the pre-training dataset and the fine-tuning dataset includes: For each training image in the pre-training dataset and the fine-tuning dataset, randomly cropping an image block of a preset size from the training image; if the original size of the training image is smaller than the preset size, copying multiple copies of the training image and splicing them into the image block of the preset size; The image block is converted from RGB space to YCbCr space to obtain an input image for model training.
3. The method for identifying artificial intelligence-generated images according to claim 1, wherein: The deep learning neural network model for AI-generated image identification constructed by combining the feature pyramid network and the convolutional attention module includes: Extracting and fusing features from different convolutional layers through the feature pyramid network to generate a multi-scale feature map to capture information at different levels in the input image; The features extracted by the feature pyramid network are further enhanced by the convolutional attention module, and classification labels for identification are output.
4. The method for identifying artificial intelligence-generated images according to claim 1, wherein: Pre-training the deep learning neural network model using the pre-processed pre-training data set to obtain a pre-training model includes: Inputting the pre-processed pre-training data set into the feature pyramid network of the deep learning neural network model to obtain multi-scale fusion features of the pre-training image; Inputting the multi-scale fusion features into a classification network of a convolutional attention module that integrates a channel attention mechanism and a spatial attention mechanism to obtain a pre-trained classification result; The loss function is calculated using the pre-trained classification results and image labels, and the model parameters are updated through back propagation until the loss function converges to a preset target to obtain the pre-trained model.
5. The method for identifying artificial intelligence-generated images according to claim 1, wherein: The public image datasets include the Genimage image dataset, the DiFF image dataset and the Fake2M image dataset, and the target platforms include the Adobe Firefly platform, the Bing platform, the Canva platform, the Dreamstudio platform, the JasperArt platform, the Midjourney platform, the Nightcafe platform, the Playground platform and the Prodia platform.
6. An artificial intelligence generated image identification system, characterized in that: The artificial intelligence generated image identification system includes: A pre-training dataset acquisition module, configured to acquire a first order of magnitude of AI-generated images and real images from a public image dataset as a pre-training dataset, wherein the first order of magnitude is greater than or equal to 10,000; a fine-tuning dataset acquisition module, configured to acquire AI-generated images of a second order of magnitude from a target platform for the identification task as a fine-tuning dataset, where the second order of magnitude is greater than or equal to ten and less than or equal to one hundred; A dataset preprocessing module, configured to preprocess the pre-training dataset and the fine-tuning dataset, wherein the preprocessing includes random cropping and color space conversion; The identification model building module is used to combine the feature pyramid network and the convolutional attention module to build a deep learning neural network model for AI-generated image identification; A model pre-training module is used to pre-train the deep learning neural network model using the pre-processed pre-training data set to obtain a pre-trained model; A model fine-tuning module, configured to fine-tune the pre-trained model using the pre-processed fine-tuning dataset to obtain an AI-generated image identification model; A platform image identification module is used to identify the image to be detected on the target platform based on the AI-generated image identification model and output an identification result; The platform image identification module includes: An image input unit for detecting, configured to input the image to be detected on the target platform into the AI-generated image identification model; A multi-level feature extraction unit, configured to perform multi-level feature extraction on the image to be detected using the feature pyramid network in the AI-generated image identification model, and to fuse shallow high-resolution features and deep semantic features extracted from the multi-level features to enhance the perception of artifacts of different scales in the image; A channel attention processing unit, configured to learn the channel weights of features at each level through the channel attention mechanism in the convolutional attention module, and enhance the response to high-frequency noise and color distortion features in the image; A spatial attention processing unit, configured to generate spatial masks for features at each level through the spatial attention mechanism in the convolutional attention module to locate abnormal areas in the image; The joint identification unit is used to extract the inherent noise pattern of the image through the SRM filter and frequency domain analysis, and uses the visual unit to detect the deviation of the image content from the real-world prior, and finally outputs the identification result through the joint discriminator; The model fine-tuning module includes: A data set input unit, configured to input the pre-processed fine-tuning data set into the pre-trained model; An image identification unit, configured to identify each image in the fine-tuning dataset using the pre-trained model; a structured statistics unit, configured to obtain identification results of each image in the fine-tuning dataset and to calculate the identification accuracy; A weight parameter adjustment unit is used to adjust the weight parameters of the network in the pre-trained model based on the identification accuracy, so that the pre-trained model transitions from identifying the general features of AI-generated images to adapting to the specific task features of identifying the AI-generated images of the target platform, thereby obtaining an AI-generated image identification model.
7. An artificial intelligence generated image identification device, characterized in that The artificial intelligence generated image identification device includes: a memory and at least one processor, the memory storing instructions, the memory and the at least one processor being interconnected via a circuit; The at least one processor calls the instructions in the memory to enable the artificial intelligence-generated image identification device to perform the artificial intelligence-generated image identification method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying an artificial intelligence-generated image according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Small target detection method based on attention mechanism
CN114202672A
Image tampering detection method based on cross-window self-attention correlation network
CN118711008A