Artificial Intelligence Generated Image Detection Method, Device, Storage Medium and Electronic Device
By using the image encoder and fully connected network of the CLIP model, combined with single-center loss metric learning and metric learning, the problems of detecting multi-class generated images and robustness of post-processing operations in the prior art are solved, and higher detection accuracy and generalization are achieved.
Patent Information
- Application Number
- CN202510238554.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art is difficult to effectively detect multiple types of artificial intelligence-generated images simultaneously, especially when facing new methods, unseen architectures or images under different conditions, the generalization ability is poor and the robustness of post-processing operations such as compression and blur is insufficient.
The object detection model is constructed using the image encoder of the CLIP model and the fully connected network. The real image features are aggregated through single-center loss metric learning, the boundary between the real image and the generated image is expanded, and the similar representation of the image features before and after post-processing is achieved through metric learning.
The generalization and detection accuracy of the detection method for unknown generation methods are improved, and the robustness of post-processing problems such as compression and fuzziness is enhanced.
Smart Images

Figure CN119741396B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence generated image detection. Specifically, it relates to a method, device, storage medium, and electronic device for detecting artificial intelligence generated images. Background Art
[0002] Artificial intelligence generated images refer to image content automatically generated or assisted in creation using artificial intelligence technology. Currently, the quality of generated images is already sufficient to be indistinguishable from real ones, and there are various mature tools available on the Internet for the public to use flexibly. If these tools or technologies are maliciously applied to individuals, it will not only infringe on personal portrait rights and privacy rights, but may also lead to privacy leakage, reputation damage, or be used to release false information to mislead public opinion, thereby damaging social trust. If maliciously applied to national leaders and false images related to national defense and diplomacy are released, it may lead to serious political crises. If exploited by fraudsters, it may cause serious economic losses to individuals, enterprises, etc. Therefore, researching the detection technology for artificial intelligence generated images has far-reaching significance in maintaining information authenticity, combating the abuse of generation technology, and the spread of generated content.
[0003] Currently, there are many studies on the detection of artificial intelligence generated images, mainly targeting two major mainstream artificial intelligence generated image methods, namely the generative model based on generative adversarial networks and the generative model based on diffusion. Currently, the methods for detecting artificial intelligence generated images are mainly divided into three categories: (1) detection based on visual artifact features; (2) detection based on frequency domain features; (3) detection based on data-driven.
[0004] Current artificial intelligence generated image detection models can well solve the in-distribution scenarios with perfectly aligned training sets and test sets within a single type of image generation method, and have good detection performance. However, they still cannot simultaneously well detect images generated by multiple types of generation methods, have poor generalization ability for images generated by new and unseen generation methods by the model, and are even less able to handle problems such as compression and blurring faced in real scenarios.
[0005] The main problems existing in current detection research mainly include: First, even if there are multiple types of generated images in the training set, a single model cannot simultaneously have good detection effects on the same multiple types of generated images, and has poor generalization ability for images generated by new and unseen generation methods by the model. And the test images in real scenarios are usually generated by new methods, unseen architectures, or even known architectures retrained under different conditions. Second, the artificial intelligence generated images spread on social networks face problems such as compression and blurring. Current detection methods are not robust enough to these post-processing operations, resulting in poor practicality. Summary of the Invention
[0006] Embodiments of the present application provide an artificial intelligence-generated image detection method, apparatus, storage medium, and electronic device to solve the technical problems existing in the prior art.
[0007] Other features and advantages of the present application will become apparent from the following detailed description, or will be learned in part through the practice of the present application.
[0008] According to the first aspect of the embodiments of the present application, an artificial intelligence-generated image detection method is provided, including:
[0009] Collect original images from a typical generated image dataset as a basic dataset, where the original images include generated images and real images;
[0010] Post-process the original images in the basic dataset to generate image pairs;
[0011] Construct an object detection model based on the image encoder of the CLIP model and a fully connected network;
[0012] Train the object detection model based on the image pairs;
[0013] Detect the image to be detected based on the trained object detection model.
[0014] In some embodiments of the present application, based on the foregoing solution, the post-processing of the original images in the basic dataset to generate image pairs includes:
[0015] Perform JPEG compression operation and Gaussian blur on the original images to obtain the image pairs.
[0016] In some embodiments of the present application, based on the foregoing solution, the construction of the object detection model based on the image encoder of the CLIP model and a fully connected network includes:
[0017] Select the image encoder of the CLIP model as the basic feature extraction model;
[0018] Construct a four-layer fully connected network as a classification network, where the last layer of the classification network uses the Sigmoid function, and the other layers use the RELU activation function;
[0019] Splice the basic feature extraction model and the classification network to obtain the object detection model.
[0020] In some embodiments of the present application, based on the foregoing solution, the training of the object detection model based on the image pairs includes:
[0021] Based on the image pairs, use the basic feature extraction model to extract image feature encodings;
[0022] Based on the image feature encoding, calculate the metric loss function and the binary cross-entropy loss function of the image by using the classification network;
[0023] Based on the metric loss function and the binary cross-entropy loss function, confirm the total loss of the target detection model;
[0024] Use the mini-batch gradient descent algorithm to train the parameters of the classification network until the total loss drops to the minimum value.
[0025] In some embodiments of the present application, based on the foregoing solution, the calculation formula of the binary cross-entropy loss function is as follows:
[0026] (1);
[0027] Wherein, represents the binary cross-entropy loss function, N represents the number of samples, represents the true label of the i-th sample, represents the predicted probability of the i-th sample.
[0028] In some embodiments of the present application, based on the foregoing solution, calculating the metric loss function of the image includes:
[0029] Calculate the single-center loss function of the image based on formula (2) ;
[0030] (2);
[0031] Wherein, represents the average Euclidean distance between the real image and the center point C in a batch, represents the average Euclidean distance between the generated image and the center point C in a batch, is the boundary distance between the generated image feature representation and the real image feature representation;
[0032] Calculate the loss function representing the difference in image features before and after post-processing based on formula (3) ;
[0033] (3);
[0034] Where i and j respectively represent the images before and after the post-processing operation in an image pair, M is the number of image pairs in a batch, and respectively represent that the neural network composed of CLIP ViT L / 14 plus the first two fully-connected layers embeds the image and the image into a 2048-dimensional vector.
[0035] In some embodiments of the present application, based on the foregoing solution, the calculation formula of the total loss is as follows:
[0036] (4);
[0037] Wherein, represents the total loss, represents the single-center loss function of the image, represents the loss function of the difference in image features before and after post-processing, and are both hyperparameters.
[0038] According to the second aspect of the embodiments of the present application, an artificial intelligence generated image detection device is provided, including:
[0039] A collection unit, configured to collect original images from a typical generated image dataset as a basic dataset, wherein the original images include generated images and real images;
[0040] A processing unit, configured to perform post-processing on the original images in the basic dataset to generate image pairs;
[0041] A construction unit, configured to construct an object detection model based on the image encoder and the fully connected network of the CLIP model;
[0042] A training unit, configured to train the object detection model based on the image pairs;
[0043] A detection unit, configured to detect the image to be detected based on the trained object detection model.
[0044] According to the third aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which computer instructions are stored, and when the computer instructions run on a computer, the computer is enabled to execute the method described in the first aspect.
[0045] According to the fourth aspect of the embodiments of the present application, an electronic device is provided, including: a memory and a processor;
[0046] The memory is configured to store computer instructions;
[0047] The processor is configured to call the computer instructions stored in the memory, so that the electronic device executes the method described in the first aspect.
[0048] The technical solution of this application uses the image encoder of the CLIP model with better feature expression ability as the basic feature extraction model, introduces single-center loss metric learning to aggregate real image features and expand the boundary between real images and generated images. At the same time, through metric learning, similar representations of image features before and after post-processing are achieved, improving the detection accuracy and generalization ability for unknown generation methods, and enhancing the robustness of the detection method when facing post-processing problems such as compression and blurring in real scenarios.
[0049] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and should not limit this application. Brief Description of the Drawings
[0050] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0051] Figure 1 It shows a schematic flowchart of an artificial intelligence-generated image detection method according to an embodiment of this application;
[0052] Figure 2 It shows a schematic diagram of the working principle of a classification network according to an embodiment of this application;
[0053] Figure 3 It shows a block diagram of an artificial intelligence-generated image detection device according to an embodiment of this application;
[0054] Figure 4 It shows a block diagram of an electronic device according to an embodiment of this application;
[0055] Figure 5 It shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of this application. Detailed Description of the Embodiments
[0056] Now, the exemplary embodiments will be described more comprehensively with reference to the drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art.
[0057] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0058] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0059] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0060] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0061] The following will describe in detail some embodiments of the present application in conjunction with the accompanying drawings. Without conflict, the embodiments and features in the following embodiments may be combined with each other.
[0062] See Figure 1 , which shows a schematic flowchart of an artificial intelligence-generated image detection method according to an embodiment of the present application.
[0063] As Figure 1 shown, an artificial intelligence-generated image detection method is presented, specifically including steps S100 to S500.
[0064] Refer to Figure 1 , step S100, collect original images from a typical generated image dataset as a basic dataset, where the original images include generated images and real images.
[0065] It should be noted that in this embodiment, the generated images and the real images are generated based on two major types of methods: generative adversarial networks and diffusion models.
[0066] Exemplarily, the process of obtaining the generated images and the real images is as follows:
[0067] Obtain the generated images based on generative adversarial networks such as ProGAN, StyleGAN, and CycleGAN from the dataset published in the paper "Wang, Sheng-Yu, et al. “CNN-Generated Images Are Surprisingly Easy to Spot… for Now.” CVPR, 2020."; according to the image and text description pairs provided in the LAION dataset, input the text description into models such as LDM, Glide, and DALL-E to generate the generated images based on diffusion models; the real images are also obtained from the real images in the above two datasets.
[0068] Continue to refer to Figure 1 , step S200, perform post-processing on the original images in the basic dataset to generate image pairs.
[0069] In some feasible embodiments, based on the foregoing solution, the performing post-processing on the original images in the basic dataset to generate image pairs includes:
[0070] Perform JPEG compression operation and Gaussian blur on the original images to obtain the image pairs.
[0071] It can be understood that the original images include the generated images and the real images. Performing post-processing on the original images means performing post-processing on both the generated images and the real images. Specifically, it means performing post-processing such as compression and blurring on all the images (including real images and generated images) in the basic dataset to form image pairs with the original images.
[0072] Continue to refer to Figure 1 , step S300, construct an object detection model based on the image encoder of the CLIP model and a fully connected network.
[0073] In some feasible embodiments, based on the foregoing solution, the constructing an object detection model based on the image encoder of the CLIP model and a fully connected network includes:
[0074] Select the image encoder of the CLIP model as the basic feature extraction model;
[0075] Construct a four-layer fully connected network as the classification network, where the last layer of the classification network uses the Sigmoid function, and the remaining layers use the RELU activation function;
[0076] The base feature extraction model and the classification network are spliced to obtain the target detection model.
[0077] It should be noted that the CLIP model refers to a vision-language model.
[0078] It should be noted that the output of the last layer of the base feature extraction model is the image feature encoding, which is 768-dimensional. During the later training process, the parameters of the image encoder CLIP ViT L / 14 remain unchanged.
[0079] Exemplarily, the network structure of the target detection model is as Figure 2 shown. The image encoder CLIP ViT L / 14 of the pre-trained CLIP model is selected as the base feature extraction model and spliced at the front end of the classification network; a four-layer fully connected network is constructed as the classification network, which includes 2048, 2048, 500, and 2 nodes. The sigmoid function is used in the last layer, and the ReLU activation function is used in the remaining layers.
[0080] It should be noted that in this embodiment, the main function of this classification network is:
[0081] The image feature encoding extracted by the pre-trained model is mapped to a 2048-dimensional high-dimensional space through the first two fully connected networks, and it is expected that the representation in the high-dimensional space has the following characteristics: one is to aggregate normal image features and expand the boundary between real images and generated images, so that the detection method can improve the detection accuracy while maintaining the generalization of the detection method to unknown generation methods; the other is to minimize the difference in the feature representation of the image before and after the post-processing operation, and hope to improve its robustness when facing post-processing problems such as compression and blurring in the real scene. In order to make the image features mapped to the high-dimensional space have these characteristics, it is mainly achieved through metric learning, that is, a metric loss is introduced in the second fully connected layer to reduce or increase the distance between image features.
[0082] Continue to refer to Figure 1 , step S400, training the target detection model based on the image pair.
[0083] In some feasible embodiments, based on the foregoing solution, training the target detection model based on the image pair includes:
[0084] Based on the image pair, use the base feature extraction model to extract image feature encoding;
[0085] Based on the image feature encoding, use the classification network to calculate the metric loss function and binary cross-entropy loss function of the image;
[0086] Confirm the total loss of the object detection model based on the metric loss function and the binary cross-entropy loss function;
[0087] Use the mini-batch gradient descent algorithm to train the parameters of the classification network until the total loss drops to the minimum value.
[0088] Exemplarily, the specific process of extracting the image feature encoding is as follows:
[0089] Take the image pair as input data and input it into the basic feature extraction model to extract the image feature encoding.
[0090] In some feasible embodiments, based on the foregoing solution, the calculation formula of the binary cross-entropy loss function is as follows:
[0091] ;
[0092] Where represents the binary cross-entropy loss function, N represents the number of samples, represents the true label of the i-th sample, represents the predicted probability of the i-th sample.
[0093] It should be noted that in order to make the feature distribution of the training data easier to classify, the metric loss function consists of two parts:
[0094] One is the metric loss function used to aggregate the normal image features and expand the boundary between the real image and the generated image. Here, the single-center loss function is adopted; the other is the loss function used to minimize the difference in feature representation between the images before and after post-processing.
[0095] Exemplarily, the calculation process of the single-center loss function is as follows:
[0096] Assume a given training set consisting of N samples and their labels which form a feature dimension of 2048 in the second fully connected layer, and its feature vector is represented by Assume that the center point of the real image feature is C. For a batch of data, the value of the center point C is the center point of the real image features in a batch. The single-center loss function is:
[0097] ;
[0098] Where is the average Euclidean distance between the real images and the center point C in a batch, is the average Euclidean distance between the generated images and the center point C in a batch, is the boundary distance between the generated image feature representation and the real image feature representation.
[0099] and The calculation formulas are as follows:
[0100] ;
[0101] ;
[0102] Where and represent the real image and the generated image sets respectively.
[0103] The calculation formula of the loss function used to minimize the difference in feature representations of the images before and after post - processing is as follows:
[0104] ;
[0105] Where i and j respectively represent the images before and after the post - processing operation in an image pair, M is the number of image pairs in a batch, and respectively represent the neural network composed of CLIP ViT L / 14 plus the first two fully - connected layers embedding the image and the image into a 2048 - dimensional vector.
[0106] In some feasible embodiments, based on the foregoing solution, the calculation formula of the total loss is as follows:
[0107] ;
[0108] Where, represents the total loss, represents the single - center loss function of the image, represents the loss function of the difference in image features before and after post - processing, and are both hyperparameters.
[0109] Continuing to refer to Figure 1 , in step S500, the image to be detected is detected based on the trained object detection model.
[0110] Exemplarily, referring to Figure 2 , the image to be detected is input into the trained object detection model, and according to the output result of the softmax function of the last fully - connected layer in the object detection model, it is determined whether the image is a real image or a generated image, thereby realizing image detection.
[0111] In summary, the present method utilizes the powerful feature representation ability of the image encoder of the CLIP model to extract image features, which is beneficial to improving the generalization of the detection method to unknown generation methods. When mapping the image features to a high-dimensional space, metric learning is used to aggregate the real image features and expand the boundary between the real image and the generated image, but without overly constraining the feature representation of the generated image, which can not only improve the detection accuracy but also avoid overfitting to the generation methods existing in the training dataset, maintaining the generalization of the detection method to unknown generation methods. At the same time, metric learning is used to achieve a similar representation of the image features before and after post-processing, improving its robustness when facing post-processing problems such as compression and blurring in the real scenario.
[0112] The following introduces the device embodiments of the present application, which can be used to execute an artificial intelligence-generated image detection method in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application above.
[0113] Refer to Figure 3 As shown, an artificial intelligence-generated image detection device 300 according to an embodiment of the present application includes:
[0114] A collection unit 301, configured to collect original images from a typical generated image dataset as a basic dataset, where the original images include generated images and real images;
[0115] A processing unit 302, configured to perform post-processing on the original images in the basic dataset to generate image pairs;
[0116] A construction unit 303, configured to construct an object detection model based on the image encoder of the CLIP model and a fully connected network;
[0117] A training unit 304, configured to train the object detection model based on the image pairs;
[0118] A detection unit 305, configured to detect the image to be detected based on the trained object detection model.
[0119] As Figure 4 shown, an embodiment of the present application further provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored on the memory 410 and executable on the processor. When the processor 420 executes the computer program 411, the steps of the above artificial intelligence-generated image detection method are implemented.
[0120] Since the electronic device introduced in this embodiment is the device used to implement an artificial intelligence generated image detection device in the embodiments of the present application, based on the method introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of the present application belongs to the scope protected by the present application.
[0121] In the specific implementation process, when the computer program 411 is executed by the processor, it can implement any implementation manner in the corresponding embodiments of the first aspect.
[0122] Figure 5 The structural schematic diagram of the computer system of the electronic device suitable for implementing the embodiments of the present application is shown.
[0123] It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present application.
[0124] As Figure 5 shown, the computer system 500 includes a central processing unit 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory 502 or the program loaded from the storage section 508 into the random access memory 503, such as executing the method described in the above embodiments. In the random access memory 503, various programs and data required for system operation are also stored. The central processing unit 501, the read-only memory 502, and the random access memory 503 are connected to each other through a bus 504. The input / output interface 505 is also connected to the bus 504.
[0125] The following components are connected to the input / output interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that the computer program read from it can be installed into the storage portion 508 as needed.
[0126] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, various functions defined in the system of the present application are executed.
[0127] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0129] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0130] As another aspect, the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes an artificial intelligence-generated image detection method described in the above embodiments.
[0131] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements an artificial intelligence-generated image detection method described in the above embodiments.
[0132] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0133] Those skilled in the art can easily understand from the description of the above embodiments that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0134] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An artificial intelligence generated image detection method, characterized in that: include: Collecting original images from a typical generated image dataset as a basic dataset, wherein the original images include generated images and real images; Post-process the original image in the basic data set to form an image pair with the original image; Build an object detection model based on the image encoder and fully connected network of the CLIP model; Training the object detection model based on the image pairs; Detect the image to be detected based on the trained target detection model; The CLIP model-based image encoder and fully connected network construct a target detection model, including: Select the image encoder of the CLIP model as the basic feature extraction model; Construct a four-layer fully connected network as a classification network, wherein the last layer of the classification network adopts a Sigmoid function and the remaining layers adopt a RELU activation function; The basic feature extraction model is spliced with the classification network to obtain the target detection model; The training of the target detection model based on the image pair includes: Based on the image pair, extracting image feature coding using the basic feature extraction model; Based on the image feature encoding, using the classification network to calculate the metric loss function and the binary cross entropy loss function of the image; Determining a total loss of the object detection model based on the metric loss function and the binary cross entropy loss function; Using a mini-batch gradient descent algorithm to train the parameters of the classification network until the total loss is reduced to a minimum value; The calculation formula of the binary cross entropy loss function is as follows: (1); in, represents the binary cross entropy loss function, N represents the number of samples, represents the true label of the i-th sample, represents the predicted probability of the i-th sample; Calculate the metric loss function of the image, including: The single center loss function of the image is calculated based on formula (2) ; (2); in, represents the average Euclidean distance between the real image and the center point C in a batch, represents the average Euclidean distance between the generated image and the center point C in a batch, The boundary distance between the generated image feature representation and the real image feature representation; Based on formula (3), the loss function representing the difference in image features before and after post-processing is calculated ; (3); Where i and j represent the images before and after the post-processing operation in an image pair, respectively, and M is the number of image pairs in a batch. and They represent the neural network composed of CLIP ViT L / 14 plus the first two fully connected layers to transform the image and images Embedded into a 2048-dimensional vector.
2. The method according to claim 1, characterized in that The post-processing of the original images in the basic data set to generate image pairs includes: The original image is subjected to JPEG compression and Gaussian blur to obtain the image pair.
3. The method according to claim 1, characterized in that The total loss is calculated as follows: (4); in, represents the total loss, represents the single center loss function of the image, The loss function representing the difference in image features before and after post-processing, and are all hyperparameters.
4. An artificial intelligence generated image detection device, applied to the method according to any one of claims 1 to 3, characterized in that: include: A collecting unit, used to collect original images from a typical generated image dataset as a basic dataset, wherein the original images include generated images and real images; A processing unit, used for post-processing the original image in the basic data set to form an image pair with the original image; A construction unit for building an object detection model based on the image encoder and fully connected network of the CLIP model; A training unit, configured to train the object detection model based on the image pairs; The detection unit is used to detect the image to be detected based on the trained target detection model.
5. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer executes the method according to any one of claims 1 to 3.
6. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer instructions; The processor is used to call the computer instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-3.
Citation Information
Patent Citations
Image semantic segmentation model, detection method and device, equipment and storage medium
CN110555437A
Target detection method and device, computer readable medium and electronic equipment
CN115115906A