Image detection method, device and equipment and computer readable storage medium

By generating differential images through diffusion model and combining image detection model and frequency domain feature analysis, the problems of poor generalization and low accuracy of existing false image detection methods are solved, and efficient identification and accurate detection of false images generated by generative large models are realized.

CN120259857APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410027304.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing false image detection methods, especially those based on binary classification networks and frequency domain feature analysis, have problems such as poor generalization and low accuracy in detection of low-quality images, making it difficult to effectively identify false images generated through generative large models.

Method used

The trained diffusion model is used to predict the image to be detected, the first predicted image is generated and the differential image is constructed, and the trained image detection model is used to predict the differential image, combining information entropy and frequency domain feature analysis to improve detection accuracy.

Benefits of technology

By constructing differential images and frequency domain feature analysis, the accuracy and confidence of false image detection are significantly improved, the risk of false image misidentification is reduced, and the authenticity and security of information are guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259857A_ABST
    Figure CN120259857A_ABST
Patent Text Reader

Abstract

The invention provides an image detection method, device and equipment and a computer readable storage medium. The method comprises the following steps: acquiring a to-be-detected image, a trained diffusion model and a trained image detection model; performing prediction processing on the to-be-detected image by using the trained diffusion model to obtain a first prediction image; determining a first difference image based on the first prediction image and the to-be-detected image; performing prediction processing on the first difference image by using the trained image detection model to obtain a first prediction result; based on the first prediction result, a detection result of the to-be-detected image is determined, the detection result of the to-be-detected image is output, and the detection result is used for representing whether the to-be-detected image is a false image or not. According to the invention, the confidence of the image detection result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to image processing technology, and in particular to an image detection method, device, equipment and computer-readable storage medium. Background Art

[0002] With the continuous introduction of generative large models, a large number of artificial intelligence generated content (AIGC) has begun to appear on various social platforms. The false images generated by models mainly based on the Stable Diffusion model bring certain information identification difficulties to users, increasing the risks of public opinion and deception. In order to identify whether an image is a false image, false detection is often required. In the digital age, false detection is mainly applied to verify the authenticity and integrity of digital information such as digital images, videos, and documents to ensure the credibility and security of information. There are mainly two methods for false detection in related technologies: false image detection based on a binary classification network and false image detection based on frequency domain feature analysis. False image detection based on a binary classification network is very sensitive to training data and has poor generalization, while false image detection based on frequency domain feature analysis has a low detection accuracy for images with poor quality. Summary of the Invention

[0003] Embodiments of the present application provide an image detection method, device, equipment and computer-readable storage medium, which can improve the accuracy of detecting false images.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] Embodiments of the present application provide an image detection method, the method comprising:

[0006] Obtain an image to be detected, a trained diffusion model and a trained image detection model;

[0007] Use the trained diffusion model to perform prediction processing on the image to be detected to obtain a first predicted image;

[0008] Based on the first predicted image and the image to be detected, determine a first difference image;

[0009] Use the trained image detection model to perform prediction processing on the first difference image to obtain a first prediction result;

[0010] Based on the first prediction result, determine the detection result of the image to be detected, and output the detection result of the image to be detected, where the detection result is used to characterize whether the image to be detected is a false image

[0011] Embodiments of the present application provide an image detection device, comprising:

[0012] A first acquisition module, configured to acquire an image to be detected, a trained diffusion model, and a trained image detection model;

[0013] A first prediction module, configured to perform prediction processing on the image to be detected by using the trained diffusion model to obtain a first predicted image;

[0014] A difference image reconstruction module, configured to determine a first difference image based on the first predicted image and the image to be detected;

[0015] A second prediction module, configured to perform prediction processing on the first difference image by using the trained image detection model to obtain a first prediction result;

[0016] A result output module, configured to determine a detection result of the image to be detected based on the first prediction result, and output the detection result of the image to be detected, where the detection result is used to characterize whether the image to be detected is a false image.

[0017] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0018] A memory, configured to store computer-executable instructions;

[0019] A processor, configured to implement the image detection method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, which are used to implement the image detection method provided by the embodiment of the present application when being executed by a processor.

[0021] An embodiment of the present application provides a computer program product, including computer-executable instructions, which are used to implement the image detection method provided by the embodiment of the present application when being executed by a processor.

[0022] The embodiment of the present application has the following beneficial effects:

[0023] In the embodiments of the present application, when performing authenticity image detection on the image to be detected, first, the trained diffusion model is used to perform prediction processing on the image to be detected. Through the processes of forward noise addition and backward denoising, a first predicted image is generated. Then, a first difference image is constructed from the image to be detected and the first predicted image. If the image to be detected is fake through the diffusion model, then the content of the first predicted image generated by the diffusion model is relatively close to that of the image to be detected. Therefore, the information content of the constructed first difference image is less. If the image to be detected is a real image, then the difference between the first predicted image generated by the diffusion model and the image to be detected will be relatively large. Therefore, the information content of the constructed first difference image is more. Therefore, the trained image detection model can be used to perform prediction processing on the first difference image to obtain a more accurate prediction result. Finally, the detection result is determined based on the prediction result, so as to improve the accuracy and confidence of the image detection result and enhance the authenticity of the network environment. Description of the Drawings

[0024] Figure 1A is a schematic implementation flowchart of a fake image detection method based on a binary classification network in the related art;

[0025] Figure 1B is a schematic implementation flowchart of a fake image detection method based on frequency domain feature analysis in the related art;

[0026] Figure 2 is a schematic network architecture diagram of the image detection system 100 provided by the embodiments of the present application;

[0027] Figure 3 is a schematic structural diagram of the server 400 provided by the embodiments of the present application;

[0028] Figure 4A is a schematic implementation flowchart of an image detection method provided by the embodiments of the present application;

[0029] Figure 4B is a schematic implementation flowchart of determining the first difference image provided by the embodiments of the present application;

[0030] Figure 5A is a schematic implementation flowchart of determining the detection result of the image to be detected based on the first prediction result provided by the embodiments of the present application;

[0031] Figure 5B is a schematic implementation flowchart of determining the information entropy of the first difference image provided by the embodiments of the present application;

[0032] Figure 6A is a schematic implementation flowchart of training the diffusion model and the image detection model provided by the embodiments of the present application;

[0033] Figure 6B It is a schematic diagram of the implementation process for constructing the third training dataset and the fourth training dataset provided by an embodiment of the present application;

[0034] Figure 6C It is a schematic diagram of the implementation process for training an image detection model provided by an embodiment of the present application;

[0035] Figure 7 It is another schematic diagram of the implementation process for the image detection method provided by an embodiment of the present application. Detailed implementation manners

[0036] To make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0037] In the following descriptions, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and they can be combined with each other without conflict.

[0038] In the following descriptions, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0040] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0041] 1) False detection: The process of detecting and identifying possible false, tampered, or fraudulent behaviors through technical means or manual judgment. In the digital age, false detection is mainly applied to verify the authenticity and integrity of digital information such as digital images, videos, and documents to ensure the credibility and security of the information.

[0042] 2) Artificial Intelligence Generated Content (AIGC), also known as generative AI, means... For example, AI text continuation, AI images generated from text-to-image, AI hosts, etc. all belong to the applications of AIGC.

[0043] 3) Diffusion models, which are deep generative models, work by adding noise (Gaussian noise) to available training data (also known as the forward diffusion process), and then reversing this process (called denoising or reverse diffusion process) to recover the data. The model gradually learns to remove the noise. This learned denoising process generates new high-quality images from a random seed (a random noise image).

[0044] 4) Stable Diffusion (SD), a deep learning text-to-image generation model. It is mainly used to generate detailed images based on text descriptions, although it can also be applied to other tasks such as inpainting, outpainting, and generating image-to-image translations under the guidance of prompts.

[0045] To better understand the image detection method provided by the embodiments of the present application, first, the image detection methods and existing drawbacks in the related art will be described.

[0046] In the related art, there are usually two methods for detecting the authenticity of images: the false image detection method based on a binary classification network and the false image detection method based on frequency domain feature analysis.

[0047] The implementation process of the false image detection method based on a binary classification network is as Figure 1A shown. The data to be detected is input into the binary classification network to directly obtain the detection result of whether the image is a real image or a false image. This type of solution regards the false detection task as a conventional binary classification task, that is, real images are one class and false images are another class. By training a classifier composed of ResNet, the false detection function of the image is realized.

[0048] Usually, when using a generative model to upsample to obtain the final generated image, the frequency domain abnormal points formed by interpolation are used to detect false images. The implementation process of the false image detection method based on frequency domain feature analysis is as Figure 1B shown. First, perform a two-dimensional Fourier transform on the image to be detected, and filter and denoise the transformation result to highlight the frequency domain abnormal points caused by upsampling. Then, perform matching through a pre-set template, and judge whether the image is falsely generated according to the matching result.

[0049] The problem with the false image detection method based on the binary classification network is that the binary classification method is usually very sensitive to training data and has poor generalization ability. As a result, the method using the binary classification task for false detection usually cannot handle data types that have not been seen before. The Stable Diffusion method can generate false data with different styles based on various datasets. It is usually impossible to include all styles of data when constructing the dataset for the binary classification model. Therefore, the widespread use of Stable Diffusion significantly increases the difficulty of the binary classification model in detecting false images.

[0050] The problem with the false image detection method based on frequency domain feature analysis is that the abnormal points in the frequency domain are usually difficult to distinguish. Although image enhancement means such as filtering can be used to improve the visibility of the abnormal points in the frequency domain. However, the false images collected from various channels on the Internet usually go through various image processing means such as compression and cropping, resulting in a decrease in image quality and making the abnormal points in their frequency domain even more difficult to distinguish. Therefore, this method usually has many limitations when facing actual usage scenarios.

[0051] Based on this, the embodiments of the present application provide an image detection method, device, equipment and computer-readable storage medium, which can improve the accuracy of image detection results. The following describes the exemplary applications of the electronic equipment provided by the embodiments of the present application. The electronic equipment provided by the embodiments of the present application can be implemented as a terminal or a server.

[0052] See Figure 2 , Figure 2 is the schematic architecture diagram of the image detection system 100 provided by the embodiments of the present application. As Figure 2 shown, the image detection system 100 includes a terminal 200, a network 300 and a server 400. The terminal 200 is connected to the server 400 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0053] The terminal 200 is used to obtain an image to be published to a social platform. The image to be published to the social platform can be a real image captured by the terminal 200 through an image acquisition device, or a fake image synthesized by other means. The terminal 200 sends an image publishing request to the server 400 via the network 300. The image publishing request carries the image to be published. After receiving the image publishing request, the server 400 obtains the image to be published and determines the image to be published as the image to be detected. The server 400 obtains a trained diffusion model and a trained image detection model, and uses the trained diffusion model to perform prediction processing on the image to be detected to obtain a first predicted image. Then, based on the first predicted image and the image to be detected, a first difference image is determined. Using the trained image detection model, prediction processing is performed on the first difference image to obtain a first prediction result. Finally, based on the first prediction result, the detection result of the image to be detected is determined, and the detection result of the image to be detected is output. The detection result is used to indicate whether the image to be detected is a fake image. When the detection result indicates that the image to be detected is a real image, the server 400 publishes the image to be detected and returns a notification message indicating successful publication to the terminal 200. When the detection result indicates that the image to be detected is a fake image, the server 400 does not publish the image to be detected and returns a notification message indicating failed publication to the terminal 200.

[0054] In some embodiments, the server 400 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 200 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart TV, a vehicle terminal, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0055] See Figure 3 , Figure 3 is a schematic structural diagram of the server 400 provided by the embodiments of the present application. Figure 3 The server 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the server 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3Various buses are labeled as bus system 440.

[0056] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0057] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.

[0058] Memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Memory 450 optionally includes one or more storage devices that are physically remote from processor 410.

[0059] Memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), and volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0060] In some embodiments, memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0061] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0062] Network communication module 452, for reaching other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, wireless compatibility certification (WiFi), and universal serial bus (USB), etc.;

[0063] A presentation module 453 for enabling the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.);

[0064] An input processing module 454 for detecting and translating one or more user inputs or interactions from one of one or more input devices 432.

[0065] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 3 Shown is an image detection device 455 stored in the memory 450, which may be software in the form of a program, a plug-in, etc., including the following software modules: a first acquisition module 4551, a first prediction module 4552, a differential image reconstruction module 4553, a second prediction module 4554, and a result output module 4555. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.

[0066] In other embodiments, the device provided by the embodiments of the present application may be implemented in hardware. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the image detection method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0067] The image detection method provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server provided by the embodiments of the present application.

[0068] Next, the image detection method provided by the embodiments of the present application will be described. As mentioned above, the electronic device for implementing the image detection method of the embodiments of the present application may be a terminal, a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.

[0069] See Figure 4A ,Figure 4A It is a schematic diagram of an implementation process of the image detection method provided by an embodiment of the present application, and will be described in combination with Figure 4A the steps shown below. Figure 4A The execution subject of the steps is described by taking the server as an example.

[0070] In step 101, obtain the image to be detected, the trained diffusion model, and the trained image detection model.

[0071] In some embodiments, the image to be detected can be a grayscale image or a color image. For example, it can be an RGB color image, or it can also be a CMYK image. The image to be detected can be an image collected by the terminal through its own image acquisition device (such as a camera), or an image sent by other terminals, or an image downloaded from the network, or a synthetic fake image. For example, it can be an image generated through the text-to-image conversion function, or another generated fake image generated according to the image.

[0072] The diffusion model is a generative model, that is, a model capable of generating synthetic images. The principle of the diffusion model comes from physics. In physics, gas molecules diffuse from a high-concentration region to a low-concentration region, which is similar to the information loss caused by the interference of noise. Therefore, by introducing noise and then trying to generate an image by denoising. Through multiple iterations over a period of time, the diffusion model learns to generate new images each time given some noise inputs. The working principle of the diffusion model is to learn the information attenuation caused by noise and then use the learned pattern to generate an image. In some embodiments, the diffusion model can be a stable diffusion model, a denoising diffusion probabilistic model, etc. The specific form of the diffusion model is not limited in the embodiments of the present application.

[0073] In some embodiments, the trained image detection model is used to detect the authenticity of the image to be detected to determine whether the image to be detected is a real image or a fake image. The trained image detection model can be a binary classification model.

[0074] In step 102, use the trained diffusion model to perform prediction processing on the image to be detected to obtain a first predicted image.

[0075] In some embodiments, when the trained diffusion model performs prediction processing on the image to be detected, it first performs a forward processing process on the image to be detected, that is, gradually adding noise to the image to be detected until a white noise image is obtained, and then through a reverse processing process, predicting the noise added in each step of the forward processing process and performing denoising processing, so as to gradually restore and obtain a noiseless first predicted image. Among them, the image to be detected and the first predicted image have the same size.

[0076] In step 103, based on the first predicted image and the image to be detected, a first difference image is determined.

[0077] In some embodiments, as Figure 4B shown, step 103 can be implemented through the following steps 1031 to 1033, which are specifically described below.

[0078] In step 1031, the first pixel values of each first pixel point in the first predicted image are obtained, and the second pixel values of each second pixel point in the image to be detected are obtained.

[0079] In some embodiments, when the image to be detected is a grayscale image, the first predicted image generated by the diffusion model is also a grayscale image. At this time, the first pixel values of each pixel point obtained from the first predicted image are grayscale values. Similarly, the second pixel values of each second pixel point obtained from the image to be detected are also grayscale values. When the image to be detected is a color image, the first predicted image generated by the diffusion model is also a color image. Assuming that both the image to be detected and the first predicted image are RGB color images, then the first pixel values of each first pixel point obtained from the first predicted image include the color pixel values corresponding to each of the three color components. Similarly, the second pixel values of each second pixel point obtained from the image to be detected also include the color pixel values corresponding to each of the three color components.

[0080] In step 1032, based on the first pixel values of each first pixel point and the second pixel values of the second pixel points at the same positions as the first pixel points, each difference pixel value is determined.

[0081] In some embodiments, since the size of the image to be detected is the same as that of the first predicted image, when implementing step 1032, the absolute value of the difference obtained by subtracting the second pixel value of the second pixel point from the first pixel value of the first pixel point at the same position can be determined as the differential pixel value. Exemplarily, the first pixel value of the first pixel point located in the first row and first column of the first predicted image is subtracted from the second pixel value of the second pixel point located in the first row and first column of the image to be detected, and the absolute value of the obtained pixel difference is determined as the differential pixel value. Among them, if the first predicted image and the image to be detected are grayscale images, the first pixel value can be directly subtracted from the second pixel value to obtain a pixel difference. If the first predicted image and the image to be detected are color images, the pixel value corresponding to the red component in the first pixel value can be subtracted from the pixel value corresponding to the red component in the second pixel value to obtain a red pixel difference, the pixel value corresponding to the green component in the first pixel value can be subtracted from the pixel value corresponding to the green component in the second pixel value to obtain a green pixel difference, the pixel value corresponding to the blue component in the first pixel value can be subtracted from the pixel value corresponding to the blue component in the second pixel value to obtain a blue pixel difference, and the absolute values of the red pixel difference, the green pixel difference, and the blue pixel difference are determined as the pixel difference. In some embodiments, the square root of the sum of the squares of the red pixel difference, the green pixel difference, and the blue pixel difference can also be determined as the pixel difference.

[0082] In step 1033, a first difference image is determined based on each differential pixel value.

[0083] In some embodiments, after obtaining the differential pixel values corresponding to each pixel position, the first difference image can be determined based on each differential pixel value.

[0084] Through the above steps 1031 to 1033, after using the trained diffusion model to perform prediction processing on the image to be detected and generating the first predicted image, the reconstruction difference between the image to be detected and the first predicted image is determined, thereby providing a necessary data basis for subsequent determination of the authenticity of the image to be detected. This is because if the image to be detected itself is a false image generated by the diffusion model, then the first predicted image generated based on the image to be detected by using the diffusion model again is closer in content to the image to be detected due to the same generation method, so the amount of information contained in the obtained first difference image is less. However, if the image to be detected is a real image, then the first predicted image generated based on the image to be detected by using the trained diffusion model has a relatively large distribution difference, so the amount of information contained in the first difference image is more. In this way, using the first difference image to classify the image to be detected in subsequent steps can improve the accuracy of the classification result.

[0085] Continue to refer to Figure 4A, continue to describe according to step 103.

[0086] In step 104, using the trained image detection model, perform prediction processing on the first difference image to obtain a first prediction result.

[0087] In some embodiments, the first difference image is used as the input data of the trained image detection model. The trained image detection model extracts features from the first difference image to obtain a feature vector of the first difference image. Then, based on the feature vector of the first difference image, a first prediction result is determined. The first prediction result represents the probability that the image to be detected is a real image, which is a real number between 0 and 1. The closer the first prediction result is to 1, the higher the probability that the image to be detected is a real image. The closer the first prediction result is to 0, the higher the probability that the image to be detected is a false image.

[0088] In step 105, based on the first prediction result, determine the detection result of the image to be detected, and output the detection result of the image to be detected.

[0089] Among them, the detection result is used to represent whether the image to be detected is a false image. In some embodiments, when determining the detection result of the image to be detected based on the first prediction result, it can be achieved by judging the size relationship between the first prediction result and a preset first result threshold. Among them, when the first prediction result is greater than the preset first result threshold, it is determined that the detection result of the image to be detected is a real image; when the first prediction result is less than or equal to the first result threshold, it is determined that the detection result of the image to be detected is a false image. The first result threshold can be a pre-set real number between 0 and 1. For example, the first result threshold can be 0.8, or it can also be 0.7, etc.

[0090] In some embodiments, refer to Figure 5A ,"Determining the detection result of the image to be detected based on the first prediction result" in step 105 can also be implemented through the following steps 1051 to 1055, which are specifically described below.

[0091] In step 1051, based on the pixel values of each third pixel point in the first difference image, determine the information entropy of the first difference image.

[0092] In some embodiments, information entropy is a statistical expression of the amount of information in information theory. The information entropy of the first difference image can represent the amount of information included in the first difference image. Refer to Figure 5B ,"Step 1051 can be implemented through the following steps 511 to 514, which are specifically described below.

[0093] In step 511, obtain a plurality of preset gray levels and the gray ranges corresponding to each of the gray levels.

[0094] In some embodiments, the total number of gray levels and the gray range corresponding to each gray level are preset. Exemplarily, 8 gray levels can be preset, namely gray level 0 to gray level 7, where the gray range corresponding to gray level 0 is [0, 31], the gray range corresponding to gray level 1 is [32, 63], the gray range corresponding to gray level 2 is [64, 95], and so on. The gray range corresponding to gray level 7 is [224, 255].

[0095] In step 512, based on the pixel values of each third pixel point in the first difference image, the number of each pixel point of the third pixel points whose pixel values are within the gray ranges corresponding to each gray level is determined.

[0096] In some embodiments, the initial value of the number of pixel points corresponding to each gray level is 0, then the pixel values of each third pixel point in the first difference image are obtained in sequence, then it is determined which gray range the pixel value of the third pixel point in the first difference image is located in, and then the number of pixel points corresponding to this gray level is incremented by 1 until the gray levels corresponding to all pixel points in the first difference image are determined, and the number of pixel points corresponding to each gray level is obtained.

[0097] Exemplarily, assuming that the pixel value of the third pixel point in the first difference image is 30, then through the example in step 511, it can be obtained that this third pixel point is located within the gray range corresponding to gray level 0, then the number of pixel points corresponding to gray level 0 is incremented by 1, and then the gray level corresponding to the next third pixel point is judged and the number of pixel points corresponding to this gray level is updated.

[0098] In step 513, the ratio of the number of each pixel point corresponding to each gray level to the total number of pixel points of the first difference image is determined.

[0099] Here, the total number of pixel points of the first difference image is obtained, and then the number of pixel points corresponding to each gray level is divided by the total number of pixel points to obtain the ratio corresponding to each gray level. Obviously, this ratio is a real number between 0 and 1, and the sum of the ratios corresponding to all gray levels is 1.

[0100] In step 514, the information entropy of the first difference image is determined based on each ratio.

[0101] In some embodiments, the information entropy of the first difference image can be determined by formula (1-1):

[0102]

[0103] where p iRepresents the ratio corresponding to the gray level i, and L is the total number of gray levels. Exemplarily, L can be 8, 16, 4, etc.

[0104] In step 1052, based on the information entropy and the first prediction result, determine the target prediction result.

[0105] In some embodiments, the first prediction result can be corrected by the following formula (1-2) using the information entropy to obtain the target prediction result:

[0106]

[0107] where S pred is the first prediction result, and S pred represents the prediction probability that the image to be detected is a real image, which is a real number between 0 and 1.

[0108] In step 1053, determine whether the target prediction result is greater than the second result threshold.

[0109] where, when the target prediction result is greater than the second result threshold, go to step 1054, and when the target prediction result is less than or equal to the second result threshold, go to step 1055.

[0110] In step 1054, determine that the detection result of the image to be detected is a real image.

[0111] In step 1055, determine that the detection result of the image to be detected is a false image.

[0112] Since the amount of information carried in the first difference image is large when the image to be detected is a real image, the information entropy is high. When the image to be detected is a false image, the amount of information carried in the first difference image is small, so the information entropy is low. In the above steps 1051 to 1055, the information entropy of the first difference image is used to correct the first prediction result, so as to further improve the accuracy of the detection result.

[0113] In the image detection method provided in the embodiments of the present application, when performing authenticity detection on the image to be detected, first, the trained diffusion model is used to perform prediction processing on the image to be detected. Through the processes of forward noise addition and backward denoising, a first predicted image is generated. Then, a first difference image is constructed by using the image to be detected and the first predicted image. If the image to be detected is false through the diffusion model, the content of the first predicted image generated by using the diffusion model is relatively close to that of the image to be detected. Therefore, the information content of the constructed first difference image is less. If the image to be detected is a real image, the difference between the first predicted image generated by using the diffusion model and the image to be detected will be relatively large. Therefore, the information content of the constructed first difference image is more. Therefore, the trained image detection model can be used to perform prediction processing on the first difference image to obtain a more accurate first prediction result. Finally, when determining the detection result based on the first prediction result, the information entropy of the first difference image can also be used to correct the first prediction result to obtain the target detection result, and then the detection result is determined based on the target detection result. In this way, the accuracy and confidence of the image detection result can be further improved, and the authenticity of the network environment can be enhanced.

[0114] In some embodiments, before step 101, the initial diffusion model can be trained through Figure 6A the steps 001 to 009 shown to obtain the trained diffusion model. Then, based on the trained diffusion model, the training data for the preset image detection model is constructed, and thus the preset image detection model is trained to obtain the trained image detection model. The following is described in conjunction with Figure 6A this.

[0115] In step 001, a preset diffusion model and a first training dataset are obtained.

[0116] In some embodiments, the preset diffusion model can be a stable diffusion model, a denoising diffusion probability model, or other types of diffusion models. The first training dataset includes multiple first training images. The first training images can be real images or synthetic images, that is, false images.

[0117] In step 002, the preset diffusion model is used to perform forward noise addition processing and backward denoising processing on each first training image, and the corresponding second predicted images are obtained.

[0118] In some embodiments, when the diffusion model performs forward noise addition processing on the first training image, noise is gradually added to the first training image until a noise image is obtained. Adding noise to the first training image means randomly adding pixels with random pixel values to the first training image. The number of noise addition steps for the forward noise addition processing is preset, for example, it can be 1000 steps.

[0119] After forward noise addition processing to obtain a noisy image, backward denoising processing is then performed on the noisy image. The backward denoising processing is also carried out step by step. In each step, the noise added in each step of the forward noise addition processing is predicted, and then the predicted noise is removed until all the noise is removed, obtaining a first predicted image. Exemplarily, assume that the forward noise addition processing includes 1000 steps of noise addition. The first step of the backward denoising processing predicts the noise added in the 1000th step of the forward noise addition processing and removes the predicted noise to obtain the first denoised image. Then the second step predicts the noise added in the 999th step of the forward noise addition processing and removes the predicted noise from the first denoised image to obtain the second denoised image, and so on, until after 1000 steps of denoising processing, a second predicted image is obtained.

[0120] In step 003, based on each first training image and its corresponding second predicted image, a first loss value is determined.

[0121] In some embodiments, based on a preset loss function (which can also be called an optimization function), using each first training image and its corresponding second predicted image, the first loss value can be determined.

[0122] In step 004, the preset diffusion model is trained by backpropagation using the first loss value until the training end condition is reached, obtaining a trained diffusion model.

[0123] In some embodiments, the first loss value is backpropagated to the diffusion model, and the parameters of the preset diffusion model are adjusted using the gradient descent algorithm until the training end condition is reached, obtaining a trained diffusion model. Among them, the training end condition can be that the first loss value is less than a preset first loss threshold, or it can be that a preset number of training times is reached, or it can also be that the difference between the first loss values of two consecutive trainings is less than a preset first difference threshold.

[0124] In step 005, a preset image detection model, a trained diffusion model, and a second training data set are obtained.

[0125] Among them, the second training data set includes multiple second training images and the label information of each second training image. The label information of the second training image can be true or false to represent whether the second training image is a real image. The multiple second training images included in the second training data set can be the same as the multiple first training images included in the first training data set, or different, or a part of the first training images in the first training data set can be selected as the second training images, and then a part of the images can be obtained through other means as the second training images. For example, real images can be collected using the image acquisition device of the terminal, or virtual images can be generated using a trained generation model.

[0126] In step 006, the trained diffusion model is used to perform prediction processing on each second training image to obtain each third prediction image corresponding to each second training image.

[0127] In some embodiments, the trained diffusion model is used to perform forward noise addition processing on each second training image to obtain a noise image corresponding to each second training image, and then backward denoising processing is performed on each noise image to obtain each third prediction image corresponding to each second training image.

[0128] It should be noted that the implementation process of step 006 is similar to that of step 002, and the implementation process of step 002 can be referred to.

[0129] In step 007, based on each second training image and each corresponding third prediction image, each second difference image is determined.

[0130] In some embodiments, for a second training image and the third prediction image corresponding to the second training image, the implementation process of determining the second difference image corresponding to the second training image is similar to that of step 103. The implementation process of step 103 can be referred to during implementation.

[0131] In step 008, a third training data set and a fourth training data set are constructed based on each second difference image.

[0132] Among them, the third training data set includes each second difference image and the label information of each second difference image, and the fourth training data set includes the amplitude images corresponding to each second difference image and the label information of each amplitude image.

[0133] In some embodiments, referring to Figure 6B , the above step 008 can be implemented through the following steps 0081 to 0085, which will be specifically described below.

[0134] In step 0081, the label information of each second training image corresponding to each second difference image is determined as the label information of each second difference image.

[0135] In some embodiments, it is assumed that the label information of the second training image corresponding to the second difference image is true, that is to say, the second difference image is determined by the true second training image and the third predicted image corresponding to the true second training image. Then the label information of the second difference image is true. It is assumed that the label information of the second training image corresponding to the second difference image is false, that is to say, the second difference image is determined by the false second training image and the third predicted image corresponding to the false second training image. Then the label information of the second difference image is false.

[0136] In step 0082, a third training data set is constructed using each second difference image and the label information of each second difference image.

[0137] In step 0083, a Fourier transform is performed on each second difference image to obtain each amplitude image corresponding to each second difference image.

[0138] Performing a Fourier transform on an image converts the image from the spatial domain to the frequency domain. The pixel values in the original image are under the x,y coordinate axes (i.e., the spatial domain), while the pixel values after the Fourier transform are under the u,v coordinate axes (i.e., the frequency domain). The two-dimensional Fourier transform can decompose a two-dimensional signal (image) into the sum of triangular plane waves. In image processing, performing a Fourier transform on an image, the obtained frequency domain can reflect the degree of intensity change of the image in the spatial domain, that is, the change speed of the image gray level, that is, the gradient magnitude of the image. For an image, the edge part of the image is the mutated part with a fast change, so it is a high-frequency component in the frequency domain; most of the image noise is the high-frequency part; the gently changing part of the image is the low-frequency component. That is to say, the Fourier transform provides another angle to observe the image, and the image can be observed from the gray level distribution to the frequency distribution to observe the characteristics of the image.

[0139] In some embodiments, the discrete Fourier transform function can be used to perform a Fourier transform on the second difference image to obtain the Fourier transform result of the second difference image, and then based on the Fourier transform result of the second difference image, the frequency domain amplitude corresponding to each pixel point is determined, so as to obtain the amplitude image corresponding to the second difference image using the frequency domain amplitude corresponding to each pixel point.

[0140] In step 0084, the label information of each second difference image corresponding to each amplitude image is determined as the label information of each amplitude image.

[0141] Exemplarily, if the label information of the second difference image corresponding to the amplitude image is true, then the label information of the amplitude image is also true. Similarly, if the label information of the second difference image corresponding to the amplitude image is false, then the label information of the amplitude image is also false.

[0142] In step 0085, a fourth training dataset is constructed using each amplitude image and the label information of each amplitude image.

[0143] Through the above steps 0081 to 0085, a third training dataset in the spatial domain is constructed using the second difference image and the label information of the second difference image, and a fourth training dataset in the frequency domain is constructed using the amplitude image obtained by performing a Fourier transform on the second difference image and the label information of the amplitude image. This can not only increase the data volume and diversity of the training data, but also improve the prediction accuracy of the image detection model trained using the training data.

[0144] Continue to refer to Figure 6A , and the description continues with step 008.

[0145] In step 009, the preset image detection model is trained using the third training dataset and the fourth training dataset to obtain a trained image detection model.

[0146] In some embodiments, referring to Figure 6C , step 009 can be implemented through the following steps 0091 to 0096, which are specifically described below.

[0147] In step 0091, each second difference image is subjected to a prediction process using the preset image detection model to obtain a second prediction result for each second difference image.

[0148] In some embodiments, the image detection model can be a binary classification model. When using the preset image detection model to perform a prediction process on the second difference image, first, the image features of the second difference image are extracted, and then, based on the image features, the second prediction result of the second difference image is determined. The second prediction result is used to represent the probability that the second training image corresponding to the second difference image is a real image.

[0149] In step 0092, a second loss value is determined based on the second prediction results of each second difference image and the label information of each second difference image.

[0150] In some embodiments, a preset loss function can be used to determine the second loss value based on the second prediction results of each second difference image and the label information of each second difference image. Exemplarily, the preset loss function can be a cross-entropy loss function, a mean squared error loss function, etc.

[0151] In step 0093, each amplitude image is subjected to a prediction process using the preset image detection model to obtain a third prediction result for each amplitude image.

[0152] In step 0094, based on the third prediction results of each amplitude image and the label information of each amplitude image, a third loss value is determined.

[0153] It should be noted that the implementation processes of step 0093 and step 0094 are similar to those of step 0091 and step 0092. When implementing, the implementation processes of step 0091 and step 0092 can be referred to.

[0154] In step 0095, based on the second loss value and the third loss value, a comprehensive loss value is determined.

[0155] In some embodiments, the sum of the second loss value and the third loss value can be determined as the comprehensive loss value. It is also possible to obtain a first weight for the second loss value and a second weight for the third loss value, and then perform a weighted sum of the second loss value and the third loss value using the first weight and the second weight to obtain the comprehensive loss value.

[0156] In step 0096, the preset image detection model is trained by backpropagation using the comprehensive loss value until the training end condition is reached, and a trained image detection model is obtained.

[0157] In some embodiments, the comprehensive loss value is backpropagated to the image detection model, and the parameters of the preset image detection model are adjusted using the gradient descent algorithm until the training end condition is reached, and a trained image detection model is obtained. Among them, the training end condition can be that the comprehensive loss value is less than a preset second loss threshold, or it can be that a preset number of training times is reached, or it can be that the difference between the comprehensive loss values of two consecutive trainings is less than a preset second difference threshold.

[0158] In the above steps 001 to 009, first, the initial diffusion model is trained using the first training images in the first training dataset to obtain a trained diffusion model. Then, the diffusion model is used to perform prediction processing on the second training images to obtain the third prediction images of each second training image, and the reconstruction difference images between the third prediction images and the second training images, that is, the second difference images, are determined. Then, the second difference images and the second difference images are used to construct a third training dataset in the spatial domain and a fourth training dataset in the frequency domain. Then, the third training dataset and the fourth training dataset are used to train the preset image detection model to obtain a trained image detection model. That is to say, the training data used to train the preset image detection model are the difference images in the spatial domain and the amplitude images of the difference images in the frequency domain. In this way, the image detection model can learn the characteristics of real images and fake images in the gray-scale distribution and frequency distribution, so as to obtain more accurate prediction results.

[0159] Based on the foregoing embodiments, an embodiment of the present application further provides an image detection method, which is applied to Figure 2 the image detection system shown in Figure 7 Another schematic flowchart of the image detection method provided by the embodiment of the present application is shown below, and will be specifically described.

[0160] In step 201, the terminal acquires the image to be detected.

[0161] In some embodiments, the image to be detected may be an image whose authenticity needs to be detected. The image to be detected may be a portrait image, a product image, or a building image, etc.

[0162] In step 202, the terminal sends an image detection request to the server.

[0163] In some embodiments, the image detection request carries the image to be detected.

[0164] In step 203, the server parses the received image detection request to acquire the image to be detected.

[0165] In step 204, the server acquires the trained diffusion model and the trained image detection model.

[0166] In step 205, the server performs prediction processing on the image to be detected by using the trained diffusion model to obtain a first predicted image.

[0167] In step 206, the server determines a first difference image based on the first predicted image and the image to be detected.

[0168] In step 207, the server performs prediction processing on the first difference image by using the trained image detection model to obtain a first prediction result.

[0169] In step 208, the server determines the detection result of the image to be detected based on the first prediction result.

[0170] It should be noted that the implementation processes of steps 204 to 208 are similar to those of steps 101 to 105. In practical applications, the implementation processes of steps 101 to 105 can be referred to.

[0171] In step 209, the server sends the detection result of the image to be detected to the terminal.

[0172] In step 210, the terminal displays the detection result of the image to be detected.

[0173] In the image detection method provided in the embodiments of the present application, after the terminal obtains the image to be detected, it sends an image detection request to the server. The server first uses the trained diffusion model to generate a fake first predicted image from the image to be detected, and determines a difference reconstruction image between the first predicted image and the image to be detected, that is, the first difference image. Then, the server uses the trained image detection model to perform prediction processing on the first difference image to obtain a first prediction result, and determines the detection result of the image to be detected based on the first prediction result. Finally, the server sends the detection result to the terminal. Since the trained image detection model is used to perform authenticity detection on the image to be detected based on the difference image between the image to be detected and the generated fake image, the accuracy of the detection result can be guaranteed.

[0174] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0175] With the continuous introduction of generative large models, a large number of artificial intelligence-generated contents have begun to appear on various social platforms. False pictures have brought certain information identification difficulties to users, increasing the risks of public opinion and deception. Based on this situation, when a user posts a picture on a social platform, the social platform can detect the picture to be posted. If a false picture is detected, a prompt can be given to the user, or the false picture can be prevented from being posted, so as to provide a safer and more reliable social experience for users and reduce the risks of users being deceived and misled. In addition, the image detection method provided in the embodiments of the present application can also be used to perform authenticity detection on the pictures already posted on the Internet, so as to quickly and accurately identify and filter out false pictures, ensuring the authenticity and credibility of information. In addition, the image detection method provided in the embodiments of the present application can also help social platforms strengthen information security management, improve user satisfaction and loyalty, and thus enhance the brand image and market competitiveness of products. Therefore, image detection technology has very important significance and value in product design and development, and it will bring a safer and more reliable user experience and commercial value to products.

[0176] In the image detection method provided in the embodiments of the present application, the trained diffusion model is used to perform prediction processing on the image to be detected to obtain a fake first predicted image. Then, the reconstruction difference image between the first predicted image and the image to be detected is determined. Furthermore, the trained binary classification network (corresponding to the image detection model in other embodiments) is used to perform prediction processing on the reconstruction difference image to obtain a first prediction result, and the final detection result is determined based on the first prediction result.

[0177] The trained diffusion model is used to perform prediction processing on the image to be detected, including a forward processing process and a backward processing process. Among them, in the forward processing process, Gaussian noise is gradually added to the image to be detected until a pure noise image is obtained, and then the noise is gradually removed from the pure noise until an image without noise is obtained. In the embodiments of the present application, for the convenience of explaining the principle, the denoising diffusion probabilistic model (Denoising Diffusion Probabilistic Models, DDPM) is used as an example for description below.

[0178] The forward process of DDPM can be described as follows:

[0179]

[0180] Where x t and x t-1 are the latent variables at the t-th step and the (t - 1)-th step in the forward process respectively. Formula (2-1) describes that at the t-th step of the forward process, the generated latent variable x t follows a normal distribution with a mean of and a variance of β t I, and β t is the value of the variance schedule at the current moment.

[0181] Similarly, the reverse process of DDPM can be described as follows:

[0182]

[0183] Where μ θ (x t , t) and ∑ θ (x t , t) are the mean and variance related to the current state x t and can be optimized during training.

[0184] In the diffusion model, the forward process aims to add noise to a given image x0 T times to obtain a pure noise image x T . The reverse process aims to gradually denoise the pure noise image until the real image is obtained. The pure noise image x T finally obtained by the forward process follows the standard normal distribution That is:

[0185]

[0186] It can be seen that the image generated by the diffusion model must follow a certain normal distribution p′ θ(x), the parameters of this distribution are obtained from the reverse process of training the diffusion model. Since real images are not generated from white noise, the normal distribution they follow must be different from p′ θ (x), denoted as p r (x).

[0187] According to the principle of the diffusion model, the optimization objective of the diffusion model is ultimately to minimize the KL divergence between the two distributions of q(x t-1 |x t , x0) and p θ (x t-1 |x t ). Among them, q(x t-1 |x t , x0) can be directly calculated by formula (2-4):

[0188]

[0189] Then the optimization function of the diffusion model can be expressed as:

[0190]

[0191] That is, to make the Gaussian distributions of the forward process and the reverse process as close as possible.

[0192] For a fake image x f generated by DDPM, if it is sent into the DDPM model again, a pure noise image is obtained after the forward process, and then a new fake image y f is generated after the reverse process. Since the generation methods of x f and y f are the same, there should be no obvious difference in their contents. However, if a real image x r is sent into the DDPM model, a pure noise image is obtained after the forward process, and then a new fake image y r is generated after the reverse process. Since the distribution difference between x r and y r is relatively large, there should be a more obvious difference in their contents.

[0193] For any diffusion model, given an image x i (whether it is a real image or a fake image), after the forward process of the diffusion model, a pure noise image n i is obtained. After the reverse process of n i , a reconstructed image y i is obtained. The image reconstruction difference can be described by the following formula:

[0194] d(x i , y i ) = ||xi -y i ||2 (2 - 6);

[0195] where n i = F(x i ), y i = B(n i ), F() represents the forward processing function, and B() represents the backward processing function.

[0196] In the embodiments of the present application, based on the image reconstruction difference d(x i , y i ) between the real image and the fake image, the real image and the fake image are classified.

[0197] For the reconstructed difference image d f (x i , y i ) of the fake image, the information it contains should be less. For the image reconstruction difference d r (x i , y i ) of the real image, it should contain more information. Therefore, a binary classification network can be trained with the reconstructed difference image as the input and whether the input image is real or fake as the label. The loss function is shown in formula (2 - 7):

[0198]

[0199] d i represents the true label, and d' i represents the label predicted by the model. N is the batch size fed into the network each time during training.

[0200] To further improve the discrimination ability of the network for real images and fake images, in addition to the image reconstruction difference d(x i , y i ), in the embodiments of the present application, the frequency domain characteristics of d(x i , y i ) are also used simultaneously for auxiliary judgment. The Fourier transform of d(x i , y i ) is performed to obtain its amplitude image f(x i , y i ). After being superimposed with d(x i , y i ), they are fed into the network simultaneously. The loss function for f(x i , y i ) can be expressed as:

[0201]

[0202] The total loss function is as shown in formula (2-9):

[0203] L = L d + L f (2-9);

[0204] Using the loss function shown in (2-9) to train the binary classification network, after obtaining the trained binary classification network and using the trained binary classification network to predict the reconstructed difference image to obtain the first prediction result, considering the difference in the amount of information contained in d r (x i , y i ) and d f (x i , y i ), the amount of information can be used to correct the first prediction result to improve the accuracy of fake image detection. The information entropy of an image can be calculated by formula (2-10):

[0205]

[0206] p i represents the proportion of pixel points with gray values in the gray range corresponding to the i-th gray level in the image to the total number of pixel points, which is a real number between 0 and 1. Among them, i is an integer from 0 to L-1, and L is the total number of gray levels of the image. Assume that the prediction score of the binary classification network is a value S pred from 0 to 1. The closer it is to 1, the greater the probability of predicting a real image. Calculate the information entropy H(d) corresponding to d(x i , y i ). Then use formula (2-11) to determine the final prediction score:

[0207]

[0208] where log2L is the normalization of the information entropy, and L represents the total number of gray levels.

[0209] Suppose the structure information of the binary classification network provided in the embodiments of the present application is shown in Table 1. Exemplarily, the structure of ResNet 18 can be used as the basis, the input is adjusted to 6 channels, and the output is adjusted to two outputs, and L d and L f are calculated respectively.

[0210] Table 1. Structure information table of the binary classification network provided in the embodiments of the present application

[0211]

[0212]

[0213] Since the binary classification task based on deep learning simply extracts image features and classifies based on prior knowledge. This process is a black box process and cannot utilize the known information of generative models such as DDPM. Therefore, in the image detection method provided in the embodiments of the present application, the false detection task of images is no longer regarded as a simple binary classification task, but a first predicted image is generated based on a diffusion model, and a binary classification is performed on the reconstructed difference image between the first predicted image and the image to be detected. Since the image detection method provided in the embodiments of the present application is designed based on the principle of the diffusion model, there is sufficient theoretical support. And the frequency domain characteristics and information entropy of the image reconstruction difference are additionally introduced. Further improving the confidence level during actual prediction.

[0214] It can be understood that in the embodiments of the present application, when it comes to data related to images to be detected, etc., when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0215] Next, continue to describe the exemplary structure of the implementation of the image detection device 455 provided in the embodiments of the present application as software modules. In some embodiments, as Figure 3 shown, the software modules stored in the image detection device 455 in the memory 450 may include:

[0216] The first acquisition module 4551 is used to acquire the image to be detected, the trained diffusion model, and the trained image detection model;

[0217] The first prediction module 4552 is used to perform prediction processing on the image to be detected by using the trained diffusion model to obtain a first predicted image;

[0218] The difference image reconstruction module 4553 is used to determine a first difference image based on the first predicted image and the image to be detected;

[0219] The second prediction module 4554 is used to perform prediction processing on the first difference image by using the trained image detection model to obtain a first prediction result;

[0220] The result output module 4555 is used to determine the detection result of the image to be detected based on the first prediction result and output the detection result of the image to be detected, and the detection result is used to characterize whether the image to be detected is a false image.

[0221] In some embodiments, the device further includes:

[0222] The second acquisition module is used to acquire a preset diffusion model and a first training data set, and the first training data set includes a plurality of first training images;

[0223] A third prediction module, configured to perform forward noise addition processing and backward denoising processing on each first training image by using the preset diffusion model, and correspondingly obtain each second prediction image;

[0224] A second determination module, configured to determine a first loss value based on each of the first training images and the corresponding second prediction images;

[0225] A first training module, configured to perform backpropagation training on the preset diffusion model by using the first loss value until a training end condition is reached, and obtain a trained diffusion model.

[0226] In some embodiments, the apparatus further includes:

[0227] A third acquisition module, configured to acquire a preset image detection model, a trained diffusion model, and a second training data set, where the second training data set includes a plurality of second training images and label information of each second training image;

[0228] A fourth prediction module, configured to perform prediction processing on each of the second training images by using the trained diffusion model, and obtain each third prediction image corresponding to each of the second training images;

[0229] A third determination module, configured to determine each second difference image based on each of the second training images and the corresponding third prediction images;

[0230] A data construction module, configured to construct a third training data set and a fourth training data set based on each of the second difference images, where the third training data set includes each of the second difference images and label information of each of the second difference images, and the fourth training data set includes magnitude images corresponding to each of the second difference images and label information of each of the magnitude images;

[0231] A second training module, configured to train the preset image detection model by using the third training data set and the fourth training data set, and obtain a trained image detection model.

[0232] In some embodiments, the data construction module is further configured to:

[0233] Determine the label information of each of the second training images corresponding to each of the second difference images as the label information of each of the second difference images;

[0234] Construct a third training data set by using each of the second difference images and the label information of each of the second difference images;

[0235] Perform Fourier transform on each of the second difference images to obtain respective amplitude images corresponding to each of the second difference images;

[0236] Determine the label information of each of the second difference images corresponding to each of the amplitude images as the label information of each of the amplitude images;

[0237] Construct a fourth training dataset using each of the amplitude images and the label information of each of the amplitude images.

[0238] In some embodiments, the second training module is further configured to:

[0239] Perform prediction processing on each of the second difference images using the preset image detection model to obtain second prediction results of each of the second difference images;

[0240] Determine a second loss value based on the second prediction results of each of the second difference images and the label information of each of the second difference images;

[0241] Perform prediction processing on each of the amplitude images using the preset image detection model to obtain third prediction results of each of the amplitude images;

[0242] Determine a third loss value based on the third prediction results of each of the amplitude images and the label information of each of the amplitude images;

[0243] Determine a comprehensive loss value based on the second loss value and the third loss value;

[0244] Use the comprehensive loss value to perform backpropagation training on the preset image detection model until the training end condition is reached to obtain a trained image detection model.

[0245] In some embodiments, the first determination module is further configured to:

[0246] Obtain the first pixel values of each first pixel point in the first prediction image, and obtain the second pixel values of each second pixel point in the image to be detected;

[0247] Determine respective difference pixel values based on the first pixel values of each first pixel point and the second pixel values of the second pixel points at the same positions as each of the first pixel points;

[0248] Determine a first difference image based on each of the difference pixel values.

[0249] In some embodiments, the result output module is further configured to:

[0250] When the first prediction result is greater than a preset first result threshold, determine that the detection result of the image to be detected is a real image;

[0251] When the first prediction result is less than or equal to the first result threshold, determine that the detection result of the image to be detected is a false image.

[0252] In some embodiments, the result output module is further configured to:

[0253] Determine the information entropy of the first difference image based on the pixel values of each third pixel point in the first difference image;

[0254] Determine a target prediction result based on the information entropy and the first prediction result;

[0255] When the target prediction result is greater than the second result threshold, determine that the detection result of the image to be detected is a real image;

[0256] When the target prediction result is less than or equal to the second result threshold, determine that the detection result of the image to be detected is a false image.

[0257] In some embodiments, the result output module is further configured to:

[0258] Obtain a plurality of preset gray levels and the gray ranges corresponding to each of the gray levels;

[0259] Determine the number of each pixel point of the third pixel points whose pixel values are within the gray ranges corresponding to each of the gray levels based on the pixel values of each third pixel point in the first difference image;

[0260] Based on the ratio of the number of each pixel point corresponding to each of the gray levels to the total number of pixel points of the first difference image;

[0261] Determine the information entropy of the first difference image based on each of the ratios.

[0262] An embodiment of the present application provides a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the image detection method described above in the embodiments of the present application.

[0263] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the processor will be caused to execute the image detection method provided by the embodiments of the present application, for example, Figure 4A the image detection method shown.

[0264] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0265] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0266] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).

[0267] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0268] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. An image detection method, characterized in that, The method includes: Obtaining an image to be detected, a trained diffusion model, and a trained image detection model; Performing prediction processing on the image to be detected by using the trained diffusion model to obtain a first predicted image; Determining a first difference image based on the first predicted image and the image to be detected; Performing prediction processing on the first difference image by using the trained image detection model to obtain a first prediction result; Determining a detection result of the image to be detected based on the first prediction result, and outputting the detection result of the image to be detected, where the detection result is used to characterize whether the image to be detected is a fake image.

2. The method according to claim 1, wherein The method further includes: Obtaining a preset diffusion model and a first training data set, where the first training data set includes a plurality of first training images; Performing forward noise addition processing and backward denoising processing on each first training image by using the preset diffusion model to respectively obtain corresponding second predicted images; Determining a first loss value based on each first training image and its corresponding second predicted image; Performing backpropagation training on the preset diffusion model by using the first loss value until a training end condition is reached to obtain a trained diffusion model.

3. The method according to claim 1, wherein The method further includes: Obtaining a preset image detection model, a trained diffusion model, and a second training data set, where the second training data set includes a plurality of second training images and label information of each second training image; Performing prediction processing on each second training image by using the trained diffusion model to obtain corresponding third predicted images of each second training image; Determining respective second difference images based on each second training image and its corresponding third predicted image; Constructing a third training data set and a fourth training data set based on each second difference image, where the third training data set includes each second difference image and label information of each second difference image, and the fourth training data set includes magnitude images corresponding to each second difference image and label information of each magnitude image; Training the preset image detection model by using the third training data set and the fourth training data set to obtain a trained image detection model.

4. The method according to claim 3, wherein The constructing the third training data set and the fourth training data set based on each second difference image includes: Determining label information of each second training image corresponding to each second difference image as label information of each second difference image; Constructing the third training data set by using each second difference image and label information of each second difference image; Performing Fourier transform on each second difference image to obtain corresponding magnitude images of each second difference image; Determining label information of each second training image corresponding to each magnitude image as label information of each magnitude image; Constructing the fourth training data set by using each magnitude image and label information of each magnitude image.

5. The method according to claim 3, characterized in that, Training the preset image detection model using the third training dataset and the fourth training dataset to obtain a trained image detection model includes: Performing prediction processing on each of the second difference images using the preset image detection model to obtain second prediction results for each of the second difference images; Determining a second loss value based on the second prediction results of each of the second difference images and the label information of each of the second difference images; Performing prediction processing on each of the amplitude images using the preset image detection model to obtain third prediction results for each of the amplitude images; Determining a third loss value based on the third prediction results of each of the amplitude images and the label information of each of the amplitude images; Determining a comprehensive loss value based on the second loss value and the third loss value; Performing backpropagation training on the preset image detection model using the comprehensive loss value until a training end condition is reached to obtain a trained image detection model.

6. The method according to any one of claims 1 to 5, characterized in that, The determining the first difference image based on the first predicted image and the image to be detected includes: Obtaining the first pixel values of each first pixel point in the first predicted image and obtaining the second pixel values of each second pixel point in the image to be detected; Determining respective difference pixel values based on the first pixel values of each first pixel point and the second pixel values of the second pixel points at the same positions as each of the first pixel points; Determining the first difference image based on each of the difference pixel values.

7. The method according to any one of claims 1 to 5, characterized in that The first prediction result is the probability that the image to be detected is a real image. The determining the detection result of the image to be detected based on the first prediction result includes: When the first prediction result is greater than a preset first result threshold, determining that the detection result of the image to be detected is a real image; When the first prediction result is less than or equal to the first result threshold, determining that the detection result of the image to be detected is a fake image.

8. The method according to any one of claims 1 to 5, characterized in that The determining the detection result of the image to be detected based on the first prediction result includes: Determining the information entropy of the first difference image based on the pixel values of each third pixel point in the first difference image; Determining a target prediction result based on the information entropy and the first prediction result; When the target prediction result is greater than a second result threshold, determining that the detection result of the image to be detected is a real image; When the target prediction result is less than or equal to the second result threshold, determining that the detection result of the image to be detected is a fake image.

9. The method according to claim 8, wherein The determining the information entropy of the first difference image based on the pixel values of each third pixel point in the first difference image includes: Obtaining a plurality of preset gray levels and the gray ranges corresponding to each of the gray levels; Determining the respective numbers of each third pixel point whose pixel values fall within the gray ranges corresponding to each of the gray levels based on the pixel values of each third pixel point in the first difference image; Based on the ratio of each of the numbers of pixel points corresponding to each of the gray levels to the total number of pixel points of the first difference image; Determining the information entropy of the first difference image based on each of the ratios.

10. An image detection device, characterized in that, The device includes: A first acquisition module, configured to acquire an image to be detected, a trained diffusion model, and a trained image detection model; A first prediction module, configured to perform prediction processing on the image to be detected by using the trained diffusion model to obtain a first predicted image; A difference image reconstruction module, configured to determine a first difference image based on the first predicted image and the image to be detected; A second prediction module, configured to perform prediction processing on the first difference image by using the trained image detection model to obtain a first prediction result; A result output module, configured to determine a detection result of the image to be detected based on the first prediction result and output the detection result of the image to be detected, where the detection result is used to characterize whether the image to be detected is a false image.

11. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the method according to any one of claims 1 to 9 when executing the computer-executable instructions stored in the memory.

12. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 9.

13. A computer program product comprising computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 9.