Method and system for detecting GAN generated image based on local extreme value

Through the detection method based on local most value, the features of the GAN generated image are extracted and combined with a classifier for detection, which solves the problems of low accuracy and limited generalization performance caused by artifacts in the existing model, and achieves higher detection accuracy and generalization capabilities.

CN119942312APending Publication Date: 2025-05-06GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411787261.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing model for detecting GAN images depends on artifacts. With the improvement of the GAN structure, the artifacts are effectively hidden, resulting in low detection accuracy and limited generalization performance.

Method used

A detection method based on local most value is proposed. The filtered image is extracted through the feature extraction network and detected with a classifier to avoid dependence on artifacts.

Benefits of technology

It improves the accuracy of the model to detect different generated images, improves generalization performance, and can effectively detect images generated by unknown GANs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942312A_ABST
    Figure CN119942312A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for detecting a GAN generated image based on local extreme values, and the method comprises the steps: obtaining the GAN generated image and a real image used during the training of a GAN, and constructing a training data set, an evaluation data set, and a test data set; performing data enhancement processing on the images of the training data set; performing filtering processing on the images in the training data set; based on the local extreme value, performing feature extraction on the filtering graph through a feature extraction network; and detecting the generated image through the classifier to obtain the detection probability belonging to the real image and the generated image, determining the label attribute of the to-be-detected image according to the size of the probability value, obtaining feedback according to the actual label, and further updating the parameter. After multi-round iteration, training of the detection model is completed; and testing and evaluating the detection model to obtain a final target model. The embodiment of the invention can improve the accuracy of detecting different generated images by the model, and can be widely applied to the technical field of computers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and system for generating images based on local minimum detection GAN. Background Art

[0002] GAN (Generative Adversarial Networks) is a generative model based on deep learning technology. GAN consists of a generator and a discriminator, where the generator is responsible for generating data according to given conditions or using random noise, and its goal is to generate sufficiently realistic samples to deceive the discriminator. The discriminator is responsible for judging whether the input real samples are true or false. After adversarial training, the generator can generate samples similar to real data. GAN has a wide range of applications, including generating realistic images. However, some people maliciously use GAN to forge images and abuse them in the political and pornographic fields, posing a serious threat to personal privacy and society. Therefore, it is of great practical significance to detect whether an image is a generated image through an effective method. It can provide effective technical means for protecting the authenticity of information, social security and individual privacy, and it is also helpful to protect the public interest and enhance social credibility. With the improvement of the quality of generated images, it is obviously unrealistic to rely on human observation or experience. Therefore, researchers have proposed a large number of detection models to detect images generated by GAN based on computer vision technology.

[0003] At present, the detection models for GAN-generated images are mainly divided into detection models based on traditional digital image forensics and detection models based on deep learning technology. The former mainly uses manually designed features such as statistical features, watermarks, and illumination analysis to detect GAN-generated images. The detection accuracy of the model depends on the quality of the manual features, and such features can often only be used to detect images generated by specific GANs. The accuracy of detecting unknown GAN-generated images is low, that is, the cross-model generalization performance of the detection model is weak. Models based on deep learning technology generally need to use the artifacts in the generated images caused by the imperfect GAN design to detect the generated images. Since deep learning technology is based on a large amount of data for training, and neural networks have strong learning capabilities, this type of model has strong generalization performance and has attracted many scholars to study.

[0004] Mainstream models generally need to use artifacts to detect generated images. However, with the improvement of GAN structure, the obvious artifacts in generated images have been effectively hidden. In addition, these models are prone to overfitting the artifacts unique to the generated images in the training set. However, since the artifacts brought by different GANs are also different, the generalization performance of the detection model that relies on artifacts is limited, and the accuracy of detecting images generated by unknown GANs is low, and the effect is not satisfactory. Summary of the invention

[0005] The main purpose of the embodiments of the present invention is to propose a method and system for detecting GAN generated images based on local minimum values, which can improve the accuracy of the model in detecting different generated images.

[0006] To achieve the above objective, an embodiment of the present invention provides a method for generating an image based on a local minimum detection GAN, comprising the following steps:

[0007] Obtain GAN-generated images and real images used to train GAN, and construct training data sets, evaluation data sets, and test data sets;

[0008] Perform data enhancement processing on the images of the training data set to obtain a training data set with increased data volume;

[0009] Performing filtering processing on the images in the training data set to obtain a filtering image;

[0010] Based on the local maximum, extracting features from the filter image through a feature extraction network to obtain features for detecting the generated image;

[0011] According to the features, the generated image is detected by the classifier to obtain the detection probability of judging whether the image belongs to the real image or the generated image, the label attribute of the image to be detected is determined according to the probability value, and feedback is obtained according to the actual label, and then the parameters are updated to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed;

[0012] The detection model is tested and evaluated based on the evaluation data set and the test data set to obtain a final target model.

[0013] In some embodiments, performing data enhancement processing on the images of the training data set to obtain the training data set with increased data volume includes the following steps:

[0014] Randomly flip the images in the training dataset horizontally;

[0015] Resize and randomly cut the images to limit all training images to the same size range;

[0016] Use the built-in transform module of pytorch to convert the training images into tensor data and load the data to obtain the training data set with increased data volume.

[0017] In some embodiments, filtering the images in the training data set to obtain a filter image comprises the following steps:

[0018] Through MaxSel filtering, this filtering method splits the three-channel RGB image channel by channel, and then uses 4 convolution kernels to perform convolution operations channel by channel to obtain the convolution results of 4 directions for each point in the three channels;

[0019] The local features are selected based on the maximum value strategy. Specifically, the convolution results of the four directions of the corresponding position are compared one by one within the group, and the maximum value is taken as the convolution value at the corresponding point to obtain the filtering value at each point of each channel.

[0020] After obtaining the filter value of each point in each channel, the filter graphs of the three channels are reassembled to form a complete filter graph; before the convolution operation, by 3×H×W Fill the RGB image to make it X∈R 3 ×(H+2)×(W+2) To ensure that the filtered image is the same size as the RGB image.

[0021] In some embodiments, the method of extracting features from the filter image based on the local maximum through a feature extraction network to obtain features for detecting the generated image includes the following steps:

[0022] Use the Pytorch deep learning framework to build a feature extraction network MResNet, and input the filter graph into MResNet;

[0023] Feature extraction is performed through the MResNet module, where MResNet consists of multiple convolutional layers for preprocessing, 8 ResNet basic blocks, and 9 designed MA blocks; the MA block consists of a local maximum pooling layer and a mean pooling layer;

[0024] Among them, the formula of the process of MA block emphasizing the maximum value is: F out =MP(F in )+λ×abs(F in -AP(F in )), where λ represents an updateable weight parameter used to balance the weights of the two filtering features; F in Represents the input features; MP represents maximum pooling; AP represents mean filtering;

[0025] The MResNet changes a 7×7 average pooling layer before the ResNet output feature to a maximum pooling layer to retain the local maximum value feature as the feature F of the generated image. d .

[0026] In some embodiments, the MA block emphasizes the maximum value in the local range through the maximum pooling operation on the one hand, and smoothes the feature through mean filtering on the other hand, and then subtracts it from the original feature and takes the absolute value to emphasize the outliers in the local area, including points with too large or too small values, and finally merges the two features.

[0027] In some embodiments, the generated image is detected by a classifier according to the features to obtain the detection probability of judging whether the image belongs to a real image or a generated image, the label attribute of the image to be detected is determined according to the probability value, and feedback is obtained according to the actual label, and then the parameters are updated to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed, including the following steps:

[0028] Use MaxPix's classifier to detect the image to be detected and obtain the probability of the real image and the generated image;

[0029] Among them, MaxPix's classifier consists of two layers of fully connected neural networks. MaxPix uses the 8192 unit features F output by MResNet d Flatten, and then use a fully connected neural network to convert it into the predicted probabilities of true and false categories, and output 2 unit feature values; among them, if the probability value corresponding to the output being a real image exceeds 0.5, the input image is confirmed to be a real image;

[0030] Among them, the expression of the loss function of MResNet and the classifier is:

[0031]

[0032] Among them, Loss represents the loss function; N represents the total number of images in a batch; y represents the true label of the image; y′ represents the predicted probability value of the classifier.

[0033] In some embodiments, completing the testing and evaluation of the detection model according to the evaluation data set and the test data set to obtain the final target model includes the following steps:

[0034] When the detection accuracy of the detection model in the evaluation set reaches 100% or starts to decrease, the detection model is subjected to a final effect test using images of the test data set, and the detection effect of the detection model is evaluated according to the evaluation index;

[0035] The evaluation indicators that need to be calculated are: accuracy and average precision, where the average precision is determined by precision and recall;

[0036] The calculation formulas for each evaluation index are:

[0037]

[0038]

[0039] Among them, Acc stands for accuracy; TP stands for true positive, which is used to determine that the actual target is a positive sample; TN stands for true negative, which is used to determine that the actual sample is a negative sample; FP stands for false positive, which is used to determine negative samples; FN stands for false negative, which is used to determine positive samples; Precision stands for precision; and Recall stands for recall.

[0040] Another aspect of the embodiments of the present invention further provides a system for generating images based on a local maximum detection GAN, including:

[0041] The first module is used to obtain GAN-generated images and real images used in training GAN, and to construct training data sets, evaluation data sets, and test data sets;

[0042] The second module is used to perform data enhancement processing on the images of the training data set to obtain a training data set with an increased data volume;

[0043] The third module is used to filter the images in the training data set to obtain a filter image;

[0044] A fourth module is used to extract features of the filter image through a feature extraction network based on the local maximum value to obtain features for detecting the generated image;

[0045] The fifth module is used to detect the generated image through the classifier according to the features, obtain the detection probability of judging whether the image belongs to the real image and the generated image, determine the label attribute of the image to be detected according to the probability value, and obtain feedback according to the actual label, and then update the parameters to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed;

[0046] The sixth module is used to complete the testing and evaluation of the detection model according to the evaluation data set and the test data set to obtain the final target model.

[0047] To achieve the above objective, another aspect of an embodiment of the present invention provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.

[0048] To achieve the above objective, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0049] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0050] The embodiments of the present invention include at least the following beneficial effects: the present invention provides a method and system for detecting GAN generated images based on local extrema, the scheme constructs a training data set, an evaluation data set and a test data set by acquiring GAN generated images and real images used in training GAN; performs data enhancement processing on the images in the training data set to obtain a training data set with an increased data volume; performs filtering processing on the images in the training data set to obtain a filter graph; based on the local extrema, extracts features from the filter graph through a feature extraction network to obtain features for detecting generated images; based on the features, detects the generated images through a classifier to obtain the detection probabilities of real images and generated images, thereby completing the training of the detection model; completes the testing and evaluation of the detection model based on the evaluation data set and the test data set to obtain the final target model. The embodiments of the present invention can improve the accuracy of the model in detecting different generated images. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0052] Figure 2 is a flow chart of the overall steps provided by an embodiment of the present invention;

[0053] Figure 3 It is an overall implementation flow chart provided by an embodiment of the present invention;

[0054] Figure 4 It is a schematic diagram of the ResNet and MResNet structures provided by an embodiment of the present invention;

[0055] Figure 5 is a MaxPix model structure diagram provided by an embodiment of the present invention;

[0056] Figure 6 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.

[0058] It is understood that the terms "first", "second", etc. used in the present invention may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0059] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0061] The method and system for generating images based on the detection GAN of local extrema provided in the embodiment of the present invention relate to the field of computer technology. The method for generating images based on the detection GAN of local extrema provided in the embodiment of the present invention can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the method for generating images based on the detection GAN of local extrema, etc., but is not limited to the above forms.

[0062] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0063] like Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected to a network wirelessly or wired to complete data transmission and exchange.

[0064] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0065] In addition, the server 101 can also be a node server in the blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.

[0066] The terminal 102 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. The terminal 102 may also be a vehicle-mounted terminal of various device types described above, but is not limited thereto. The terminal 102 and the server 101 may be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the present invention.

[0067] Based on the example Figure 1 In the implementation environment shown, an embodiment of the present invention provides a method for generating an image based on a local minimum value detection GAN. The following is an example of applying the method for generating an image based on a local minimum value detection GAN to a server 101. It can be understood that the method can also be applied to the terminal 102.

[0068] Reference Figure 2 , Figure 2 A flowchart of a method for generating an image based on a local maximum detection GAN applied to a server according to an embodiment of the present invention, wherein the execution subject of the method may be any of the aforementioned computer devices (including a server or a terminal). Figure 2 , the method may include the following steps:

[0069] Obtain GAN-generated images and real images used to train GAN, and construct training data sets, evaluation data sets, and test data sets;

[0070] Perform data enhancement processing on the images of the training data set to obtain a training data set with increased data volume;

[0071] Performing filtering processing on the images in the training data set to obtain a filtering image;

[0072] Based on the local maximum, extracting features from the filter image through a feature extraction network to obtain features for detecting the generated image;

[0073] According to the features, the generated image is detected by the classifier to obtain the detection probability of judging whether the image belongs to the real image or the generated image, the label attribute of the image to be detected is determined according to the probability value, and feedback is obtained according to the actual label, and then the parameters are updated to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed;

[0074] The detection model is tested and evaluated based on the evaluation data set and the test data set to obtain a final target model.

[0075] In some embodiments, performing data enhancement processing on the images of the training data set to obtain the training data set with increased data volume includes the following steps:

[0076] Randomly flip the images in the training dataset horizontally;

[0077] Resize and randomly cut the images to limit all training images to the same size range;

[0078] Use the built-in transform module of pytorch to convert the training images into tensor data and load the data to obtain the training data set with increased data volume.

[0079] In some embodiments, filtering the images in the training data set to obtain a filter image comprises the following steps:

[0080] Through MaxSel filtering, this filtering method splits the three-channel RGB image channel by channel, and then uses 4 convolution kernels to perform convolution operations channel by channel to obtain the convolution results of 4 directions for each point in the three channels;

[0081] The local features are selected based on the maximum value strategy. Specifically, the convolution results of the four directions of the corresponding position are compared one by one within the group, and the maximum value is taken as the convolution value at the corresponding point to obtain the filtering value at each point of each channel.

[0082] After obtaining the filter value of each point in each channel, the filter graphs of the three channels are reassembled to form a complete filter graph; before the convolution operation, by 3×H×W Fill the RGB image to make it X∈R 3 ×(H+2)×(W+2) To ensure that the filtered image is the same size as the RGB image.

[0083] In some embodiments, the method of extracting features from the filter image based on the local maximum through a feature extraction network to obtain features for detecting the generated image includes the following steps:

[0084] Use the Pytorch deep learning framework to build a feature extraction network MResNet, and input the filter graph into MResNet;

[0085] Feature extraction is performed through the MResNet module, where MResNet consists of multiple convolutional layers for preprocessing, 8 ResNet basic blocks, and 9 designed MA blocks; the MA block consists of a local maximum pooling layer and a mean pooling layer;

[0086] Among them, the formula of the process of MA block emphasizing the maximum value is: F out =MP(F in )+λ×abs(F in -AP(F in )), where λ represents an updateable weight parameter used to balance the weights of the two filtering features; F in Represents the input features; MP represents maximum pooling; AP represents mean filtering;

[0087] The MResNet changes a 7×7 average pooling layer before the ResNet output feature to a maximum pooling layer to retain the local maximum value feature as the feature F of the generated image. d .

[0088] In some embodiments, the MA block emphasizes the maximum value in the local range through the maximum pooling operation on the one hand, and smoothes the feature through mean filtering on the other hand, and then subtracts it from the original feature and takes the absolute value to emphasize the outliers in the local area, including points with too large or too small values, and finally merges the two features.

[0089] In some embodiments, the generated image is detected by a classifier according to the features to obtain the detection probability of judging whether the image belongs to a real image or a generated image, the label attribute of the image to be detected is determined according to the probability value, and feedback is obtained according to the actual label, and then the parameters are updated to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed, including the following steps:

[0090] Use MaxPix's classifier to detect the image to be detected and obtain the probability of the real image and the generated image;

[0091] Among them, MaxPix's classifier consists of two layers of fully connected neural networks. MaxPix uses the 8192 unit features F output by MResNet d Flatten, and then use a fully connected neural network to convert it into the predicted probabilities of true and false categories, and output 2 unit feature values; among them, if the probability value corresponding to the output being a real image exceeds 0.5, the input image is confirmed to be a real image;

[0092] Among them, the expression of the loss function of MResNet and the classifier is:

[0093]

[0094] Among them, Loss represents the loss function; N represents the total number of images in a batch; y represents the true label of the image; y′ represents the predicted probability value of the classifier.

[0095] In some embodiments, completing the testing and evaluation of the detection model according to the evaluation data set and the test data set to obtain the final target model includes the following steps:

[0096] When the detection accuracy of the detection model in the evaluation set reaches 100% or starts to decrease, the detection model is subjected to a final effect test using images of the test data set, and the detection effect of the detection model is evaluated according to the evaluation index;

[0097] The evaluation indicators that need to be calculated are: accuracy and average precision, where the average precision is determined by precision and recall;

[0098] The calculation formulas for each evaluation index are:

[0099]

[0100] Among them, Acc stands for accuracy; TP stands for true positive, which is used to determine that the actual target is a positive sample; TN stands for true negative, which is used to determine that the actual sample is a negative sample; FP stands for false positive, which is used to determine negative samples; FN stands for false negative, which is used to determine positive samples; Precision stands for precision; and Recall stands for recall.

[0101] The following describes the specific implementation process of the present invention in detail by taking a specific application scenario as an example:

[0102] In view of the problems existing in the prior art, the present invention relies on deep learning technology to study a detection model that does not need to use artifacts to detect generated images, and proposes a MaxPix detection model through local maximum features. The purpose of the present invention is to improve the accuracy of the model in detecting different generated images, and to provide a detection model for detecting GAN generated images through local maximums.

[0103] refer to Figure 3 The present invention provides a detection model for detecting GAN generated images through local maximum. The overall process of the model includes the following steps:

[0104] Step 1: Get the training data set, evaluation data set and test data set respectively.

[0105] Step 2: First, perform data enhancement processing on the training data set images to increase the data volume of the data set and enhance the complexity of the image content, so that the trained network structure has stronger expression ability. The present invention uses four data enhancement processing methods: (1) random horizontal flip; (2) resize; (3) random cropping.

[0106] The present invention requires that the input image size is X∈R 3×299×299 In real life, the size of images is inconsistent. In order to allow images of various sizes to be input into the model, it is necessary to resize and randomly cut the images at this stage. After resizing and random cutting, all training images are limited to the same size range, that is, X∈R 3 ×299×299 The present invention applies Resize, on the one hand, to adjust the size, and on the other hand, to reduce the difference in details between the generated image and the real image. After preprocessing, the built-in transform module of pytorch is used to convert the data into tensor data and load the data. The resulting data shape is (N, 3, 299, 299), where N is the number of images loaded at one time.

[0107] Step 3: Use the designed MaxSel method to filter the image. As shown in Formula 1, the MaxSel method designed by the present invention uses the following four matrices (or convolution kernels) as filter kernels to filter the image. This function can use pytorch's Conv2d and use Formula 1 as the convolution kernel to implement filtering.

[0108]

[0109] MaxSel first splits the three-channel RGB image channel by channel, and then uses four convolution kernels to perform convolution operations on each channel to obtain the convolution results X in four directions for each point in the three channels. (c,i,j) (α1, α2, α3, α4). Then, local features are selected based on the maximum value strategy. MaxSel compares the convolution results of the four directions of the corresponding position one by one within the group, and takes the maximum value as the convolution value at the corresponding point, which is also the filter value. (c,i,j) (α1, α2, α3, α4), select the maximum value among α1, α2, α3, and α4. As shown in formula (2), where X (c,i,j) Indicates the filter value at the position (i, j) of the image c channel, Max indicates selecting the maximum value among the four convolution results. And so on, the filter value at each point of each channel is obtained.

[0110] X (c,i,j) =Max(α1,α2,α3,α4) (2)

[0111] Finally, after obtaining the filter value of each point in each channel, MaxSel rejoins the filter graphs of the three channels to form a complete filter graph F in Before the convolution operation, MaxPix converts X∈R 3×H×W The original image is filled with X∈R 3×(H+2)×(W+2) , so that the filter graph F obtained by the convolution operation in ∈R 3×H×W , to keep the image size unchanged. Step 3 does not need to change any parameters during training. In fact, the MaxSel method proposed in the present invention uses 4 convolution kernels for filtering mainly to replace the operation of formula (3).

[0112] Each convolution kernel actually performs an operation in one direction.

[0113]

[0114] Step 4: Use the Pytorch deep learning framework to build a feature extraction network (MResNet) and input the filter graph obtained in Step 3 into MResNet. The present invention designs the MResNet module to perform feature extraction. The design of MResNet is modified from ResNet, and is composed of multiple convolutional layers for preprocessing, 8 ResNet basic blocks, and 9 designed MA blocks, where the MA block is composed of a local maximum pooling layer and a mean pooling layer. The main difference between MResNet and ResNet is that MResNet has 9 more MA blocks than ResNet, such as Figure 4 shown.

[0115] The process of MA block emphasizing the maximum value is shown in formula (4), where λ is an updateable weight parameter used to balance the weights of the two filtering features. in Represents the input feature. MP stands for maximum pooling, and AP stands for mean filtering. MA block emphasizes the maximum value in the local range through the maximum pooling operation, and smoothes the feature through mean filtering, then subtracts it from the original feature and takes the absolute value to emphasize the outliers in the local area, including points with too large or too small values, and finally fuses the two features.

[0116] F out =MP(F in )+λ×abs(F in -AP(F in )) (4)

[0117] In addition, MResNet changes a 7×7 average pooling layer before the ResNet output feature to a maximum pooling layer to retain the local maximum value feature as the feature F for detecting the generated image. d The parameters of MResNet are as follows: k represents kernel_size, s represents stride, and p represents padding, which are all parameters in convolution and pooling in pytorch. Layer1, layer2, layer3, and layer4 each contain two basicBlocks.

[0118] MResNet:

[0119] layer0:

[0120] Conv2d(3,64,k=7,s=2,p=3),BatchNorm2d(64);

[0121] (MA block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0122] ReLU(inplace=True),MaxPool2d(k=3,s=2,p=1);

[0123] layer1:

[0124] BasicBlock0(

[0125] Conv2d(64,64,k=3,s=1,p=1),BatchNorm2d(64);

[0126] (MA block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0127] ReLU(inplace=True),Conv2d(64,64,k=3,s=1,p=1),BatchNorm2d(64); );

[0129] BasicBlock1(

[0130] Conv2d(64,64,k=3,s=1,p=1),BatchNorm2d(64);

[0131] (MA block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0132] ReLU(inplace=True),Conv2d(64,64,k=3,s=1,p=1),BatchNorm2d(64); );

[0134] layer2:

[0135] BasicBlock0(

[0136] Conv2d(64,128,k=3,s=2,p=1),BatchNorm2d(128);

[0137] (MA Block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0138] ReLU(inplace=True),Conv2d(128,128,k=3,s=1,p=1),BatchNorm2d(128);

[0139] (downsample):Conv2d(64,128,k=1,s=2),BatchNorm2d(128); );

[0141] BasicBlock1(

[0142] Conv2d(128,128,k=3,s=1,p=1),BatchNorm2d(128);

[0143] (MA Block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0144] ReLU(inplace=True),Conv2d(128,128,k=3,s=1,p=1),BatchNorm2d(128); );

[0146] layer3:

[0147] BasicBlock0(

[0148] Conv2d(128,256,k=3,s=2,p=1),BatchNorm2d(256);

[0149] (MA Block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);ReLU(inplace=True),Conv2d(256,256,k=3,s=1,p=1),BatchNorm2d(256);

[0150] (downsample):Conv2d(128,256,k=1,s=2),BatchNorm2d(256); );

[0152] BasicBlock1(

[0153] Conv2d(256,256,k=3,s=2,p=1),BatchNorm2d(256);

[0154] (MA Block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0155] ReLU(inplace=True),Conv2d(256,256,k=3,s=1,p=1),BatchNorm2d(256); );

[0157] (layer4):

[0158] BasicBlock0(

[0159] Conv2d(256,512,k=3,s=2,p=1),BatchNorm2d(512);

[0160] (MA Block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0161] ReLU(inplace=True),Conv2d(512,512,k=3,s=1,p=1),BatchNorm2d(512);

[0162] (downsample):Conv2d(256,512,k=1,s=2),BatchNorm2d(512); );

[0164] BasicBlock1(

[0165] Conv2d(512,512,k=3,s=2,p=1),BatchNorm2d(512);

[0166] (MA Block):MaxPool2d(k=3,s=1,p=1),AvgPool2d(k=3,s=1,p=1);

[0167] ReLU(inplace=True),Conv2d(512,512,k=3,s=1,p=1),BatchNorm2d(512); );

[0169] MaxPool2d(kernel_size=7, stride=1, padding=0);

[0170] Step 5: Use the classifier to detect the generated image. The classifier designed by the present invention can output two probability results, corresponding to the probability that the image is a real image and a generated image. The classifier C of MaxPix consists of two layers of fully connected neural networks. MaxPix uses the 8192 unit features F output by MResNet d Flatten, and then use a fully connected neural network to convert it into the predicted probability of true and false categories, that is, output 2 unit feature values. According to the setting of the present invention, if the probability value corresponding to the output of the real image exceeds 0.5, the input image will be regarded as a real image. The loss function of the present invention for training MResNet and the classifier is shown in formula (5), where C represents the prediction using the classifier, y represents the real label of the image, and y′ represents the predicted probability value of the classifier.

[0171]

[0172] After completing each round of model training, the model detection capability can be verified using the validation set images, and the need for the next round of training is determined based on the detection effect of the current model on the validation set. If the next round of training is required, the model hyperparameters are tuned using the back propagation strategy based on the cross entropy loss. Through multiple adjustments to the hyperparameters, the model performance is optimized, and the model training is completed.

[0173] The network model structure of the MaxPix detection model proposed in the present invention is as follows: Figure 5 As shown. It consists of a filtering module, an MResNet feature extraction network and a classifier C. The filtering module uses the MaxSel filtering strategy proposed in the present invention to filter the image to select appropriate features, so that MResNet can easily extract deep distinguishable features from it to detect the image generated by GAN.

[0174] Step 7: When the detection model training is completed, that is, when the detection accuracy of the model in the evaluation set reaches 100% or begins to decline, the detection model is tested with the test set images for the final effect test. The detection effect of the detection model is evaluated according to the evaluation indicators. The evaluation indicators that need to be calculated are: Accuracy (Acc) and Average Precision (AP), where the average precision is determined by Precision (Precision) and Recall (Recall). The calculation formulas of each evaluation indicator can be seen in Equations 6 to 8.

[0175]

[0176] Among them, TP represents true positive, that is, the actual target is a positive sample, and the model also judges the target as a positive sample. TN represents true negative, which means that the actual sample is a negative sample, and the model also judges the negative sample as a negative sample. FP represents false positive, that is, it is actually a negative sample, but the model mistakenly judges it as a positive sample, and FN represents false negative, that is, it is actually a positive sample, but the model mistakenly judges it as a negative sample. Precision rate represents the proportion of samples predicted by the model as positive samples that are truly positive samples. Recall rate represents the proportion of real positive samples predicted by the model as positive samples. With all recall rates as the horizontal coordinates and all precision rates as the vertical coordinates, all points on the coordinate axis are connected into a curve to form a PR curve. The average precision AP can be obtained by calculating the area under the PR curve.

[0177] In summary, previous detection models use artifacts of generated images to detect generated images. When the artifacts of the generated images are significantly different from the artifacts of the model fitting, the detection accuracy of the model decreases significantly. The present invention proposes the MaxSel algorithm. MaxSel is used to filter the image and obtain a filter graph by emphasizing the maximum value in the local range. Then, the present invention designs MABlock to improve ResNet to obtain MResNet. Finally, MaxPix uses MResNet to extract features in the filter graph to detect the generated image. MaxPix removes irrelevant information in the image through filtering, so that the detection model can easily learn general features from the image for detecting the generated image. A large number of experiments have shown that the MaxPix model can effectively detect the generated image. The advantages of the present invention are:

[0178] (1) The MaxPix model is proposed to detect generated images. It is tested on image datasets generated by multiple GANs, including StyleGAN, BigGAN, and StarGAN, and an average accuracy of 85.9% and an average precision of 94.6% are obtained, which shows strong cross-model generalization performance.

[0179] (2) The MaxSel algorithm is proposed for image filtering, and the MA Block is designed and embedded in ResNet to obtain MResNet, which is used to extract the features of the filtered image to detect the images generated by GAN.

[0180] Another aspect of the embodiments of the present invention further provides a system for generating images based on a local maximum detection GAN, including:

[0181] The first module is used to obtain GAN-generated images and real images used in training GAN, and to construct training data sets, evaluation data sets, and test data sets;

[0182] The second module is used to perform data enhancement processing on the images of the training data set to obtain a training data set with an increased data volume;

[0183] The third module is used to filter the images in the training data set to obtain a filter image;

[0184] A fourth module is used to extract features of the filter image through a feature extraction network based on the local maximum value to obtain features for detecting the generated image;

[0185] The fifth module is used to detect the generated image through the classifier according to the features, obtain the detection probability of judging whether the image belongs to the real image and the generated image, determine the label attribute of the image to be detected according to the probability value, and obtain feedback according to the actual label, and then update the parameters to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed;

[0186] The sixth module is used to complete the testing and evaluation of the detection model according to the evaluation data set and the test data set to obtain the final target model.

[0187] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0188] An embodiment of the present invention further provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method of generating an image based on the detection of local maximum GAN when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0189] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0190] See also Figure 6 , Figure 6 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0191] The processor 601 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention;

[0192] The memory 602 may be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 may store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 602, and the processor 601 calls and executes the method for generating images based on the detection of local maximum values ​​by GAN in the embodiment of the present invention;

[0193] Input / output interface 603, used to implement information input and output;

[0194] Communication interface 604, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0195] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );

[0196] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .

[0197] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method for generating an image based on the local minimum detection GAN is implemented.

[0198] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0199] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0200] It should be noted that in various specific embodiments of the present invention, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present invention needs to obtain the user's sensitive personal information, it will obtain the user's separate permission or consent through a pop-up window or jump to a confirmation page, and after clearly obtaining the user's separate permission or consent, it will obtain the necessary user-related data for the normal operation of the embodiment of the present invention.

[0201] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0202] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0203] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0205] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0206] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0207] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0208] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0209] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including multiple instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store programs.

[0211] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present invention is not limited thereby. Any modification, equivalent substitution and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.

Claims

1. A method for generating images based on local minimum detection GAN, characterized in that: The following steps are involved: Obtain GAN-generated images and real images used to train GAN, and construct training data sets, evaluation data sets, and test data sets; Perform data enhancement processing on the images of the training data set to obtain a training data set with increased data volume; Performing filtering on the images in the training data set to obtain a filtering image; Based on the local maximum, extracting features from the filter image through a feature extraction network to obtain features for detecting the generated image; According to the features, the generated image is detected by the classifier to obtain the detection probability of judging whether the image belongs to the real image or the generated image, the label attribute of the image to be detected is determined according to the probability value, and feedback is obtained according to the actual label, and then the parameters are updated to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed; The detection model is tested and evaluated based on the evaluation data set and the test data set to obtain a final target model.

2. The method for generating images based on local minimum detection GAN according to claim 1, characterized in that: The step of performing data enhancement processing on the images of the training data set to obtain the training data set with increased data volume comprises the following steps: Randomly flip the images in the training dataset horizontally; Resize and randomly cut the images to limit all training images to the same size range; Use the built-in transform module of pytorch to convert the training images into tensor data and load the data to obtain the training data set with increased data volume.

3. The method for generating images based on local minimum detection GAN according to claim 1, characterized in that: The filtering process of the images in the training data set to obtain a filter image comprises the following steps: Through MaxSel filtering, this filtering method splits the three-channel RGB image channel by channel, and then uses 4 convolution kernels to perform convolution operations channel by channel to obtain the convolution results of 4 directions for each point in the three channels; The local features are selected based on the maximum value strategy. Specifically, the convolution results of the four directions of the corresponding position are compared one by one within the group, and the maximum value is taken as the convolution value at the corresponding point to obtain the filtering value at each point of each channel. After obtaining the filter value of each point in each channel, the filter graphs of the three channels are reassembled to form a complete filter graph; before the convolution operation, by 3×H×W The RGB image is filled to make it X∈R 3 ×(H+2)×(W+2) To ensure that the filtered image is the same size as the RGB image.

4. The method for generating images based on local minimum detection GAN according to claim 1, characterized in that: The method of extracting features from the filter image based on the local maximum through a feature extraction network to obtain features for detecting and generating an image comprises the following steps: Use the Pytorch deep learning framework to build a feature extraction network MResNet, and input the filter graph into MResNet; Feature extraction is performed through the MResNet module, where MResNet consists of multiple convolutional layers for preprocessing, 8 ResNet basic blocks, and 9 designed MA blocks; the MA block consists of a local maximum pooling layer and a mean pooling layer; Among them, the formula of the process of MA block emphasizing the maximum value is: F out =MP(F in )+λ×abs(F in -AP(F in )), where λ represents an updateable weight parameter used to balance the weights of the two filtering features; F in Represents the input features; MP represents maximum pooling; AP represents mean filtering; The MResNet changes a 7×7 average pooling layer before the ResNet output feature to a maximum pooling layer to retain the local maximum value feature as the feature F of the generated image. d .

5. The method for generating images based on local minimum detection GAN according to claim 4, characterized in that: The MA block emphasizes the maximum value in the local range through the maximum pooling operation on the one hand, and smoothes the feature through the mean filter on the other hand, and then subtracts it from the original feature and takes the absolute value to emphasize the outliers in the local area, including points with too large and too small values, and finally merges the two features.

6. The method for generating images based on local minimum detection GAN according to claim 1, characterized in that: According to the features, the generated image is detected by a classifier to obtain the detection probability of judging whether the image belongs to a real image or a generated image, the label attribute of the image to be detected is determined according to the probability value, and feedback is obtained according to the actual label, and then the parameters are updated to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed, including the following steps: The image to be detected is detected by MaxPix's classifier to obtain the probability of the real image and the generated image; Among them, MaxPix's classifier consists of two layers of fully connected neural networks. MaxPix uses the 8192 unit features F output by MResNet d Flatten, and then use a fully connected neural network to convert it into the predicted probabilities of true and false categories, and output 2 unit feature values; among them, if the probability value corresponding to the output being a real image exceeds 0.5, the input image is confirmed to be a real image; Among them, the expression of the loss function of MResNet and the classifier is: Among them, Loss represents the loss function; N represents the total number of images in a batch; y represents the true label of the image; y ′ Represents the predicted probability value of the classifier.

7. The method for generating images based on local minimum detection GAN according to claim 1, characterized in that: The step of completing the testing and evaluation of the detection model according to the evaluation data set and the test data set to obtain a final target model comprises the following steps: When the detection accuracy of the detection model in the evaluation set reaches 100% or starts to decrease, the detection model is subjected to a final effect test using images of the test data set, and the detection effect of the detection model is evaluated according to the evaluation index; The evaluation indicators that need to be calculated are: accuracy and average precision, where the average precision is determined by precision and recall; The calculation formulas for each evaluation index are: Among them, Acc stands for accuracy; TP stands for true positive, which is used to determine that the actual target is a positive sample; TN stands for true negative, which is used to determine that the actual sample is a negative sample; FP stands for false positive, which is used to determine negative samples; FN stands for false negative, which is used to determine positive samples; Precision stands for precision; and Recall stands for recall.

8. A system for generating images based on local minimum detection GAN, characterized in that: include: The first module is used to obtain GAN-generated images and real images used in training GAN, and to construct training data sets, evaluation data sets, and test data sets; The second module is used to perform data enhancement processing on the images of the training data set to obtain a training data set with an increased data volume; The third module is used to filter the images in the training data set to obtain a filter image; A fourth module is used to extract features of the filter image through a feature extraction network based on the local maximum value to obtain features for detecting the generated image; The fifth module is used to detect the generated image through the classifier according to the features, obtain the detection probability of judging whether the image belongs to the real image and the generated image, determine the label attribute of the image to be detected according to the probability value, and obtain feedback according to the actual label, and then update the parameters to complete a round of training. After multiple rounds of iterations, the training of the detection model is completed; The sixth module is used to complete the testing and evaluation of the detection model according to the evaluation data set and the test data set to obtain the final target model.

9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.