Method and device for realizing synthetic face detection based on multi-generation model reconstruction difference analysis, processor and readable storage medium thereof
Through multi-generating model reconstruction difference analysis, combined with image reconstruction difference analysis of StyleGAN and ADM models, the problems of low detection accuracy and insufficient ethnic diversity in the prior art are solved, and high-precision synthetic face detection is achieved, especially in Asian populations.
Patent Information
- Application Number
- CN202510426530.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has problems in synthetic face detection with low detection accuracy, poor cross-model adaptability and insufficient racial diversity. Especially in the face of insufficient mixed generation models and Asian face data, it is difficult to effectively distinguish real images from synthetic images of multiple generation models.
The multi-generating model reconstruction difference analysis method is used to construct image acquisition and data sets, and image reconstruction is carried out using pre-trained StyleGAN and ADM models. The LPIPS value and ternary cross entropy loss function are combined for classification detection. The Asian synthetic face data set ASFD is constructed to analyze the reconstruction differences of images under different generative models.
It realizes high-precision and robust synthetic face detection, which can accurately distinguish real images from GAN/DM generated images, has cross-ethnic adaptability, and improves the generalization ability of the detection model in Asian populations.
Smart Images

Figure CN120340093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and image processing, and particularly to the field of synthetic image detection. Specifically, it refers to a method, device, processor, and computer-readable storage medium for synthetic face detection based on the reconstruction difference analysis of multiple generative models. Background Art
[0002] With the rapid development of image generation technologies such as generative adversarial networks (GANs) and diffusion models (DMs), the quality of synthetic face images has approached that of real faces and is difficult to distinguish by the naked eye or traditional methods. The abuse of such technologies may lead to security risks such as identity theft and the spread of false information. Therefore, the development of high-precision synthetic face detection technologies has become an urgent problem to be solved.
[0003] The existing technologies have the following deficiencies in the field of synthetic image detection: Most existing methods only regard synthetic image detection as a binary classification problem (real vs synthetic), for example, directly classifying through a convolutional neural network, but do not consider the differences between different generative models (such as generative adversarial networks and diffusion models). As a result, when there are images of multiple generative models at the same time, it is impossible to distinguish the generative model type, and the detection accuracy drops significantly. In addition, some methods rely on the prior knowledge of specific generative models to design detection algorithms, such as frequency domain analysis or reconstruction residuals of diffusion models, but it is difficult to adapt to new models or hybrid scenarios, and the generalization performance is limited. At the same time, existing datasets (such as FFHQ) have biases in racial diversity, and the proportion of Asian face samples is extremely low, resulting in limited application effects of the model in the Asian population. Further, existing reconstruction analysis methods only use the reconstruction residuals of a single generative model as the classification basis, and do not fully exploit the reconstruction differences of different generative models for the same image. For example, GAN-generated images have higher reconstruction quality in the GAN model, while DM-generated images have better reconstruction quality in the DM model. This difference has not been systematically utilized, resulting in insufficient dimensions of detection information. In summary, the existing technologies have significant defects in detection accuracy, cross-model adaptability, and support for racial diversity, and there is an urgent need for a detection method that can simultaneously distinguish real images from synthetic images of multiple generative models (GAN / DM) and has cross-racial adaptability. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned shortcomings of the existing technologies and provide a method, device, processor, and computer-readable storage medium for synthetic face detection based on the reconstruction difference analysis of multiple generative models, which has good generalization ability, good robustness, and a relatively wide range of applications.
[0005] In order to achieve the above purpose, the method, device, processor, and computer-readable storage medium for synthetic face detection based on the reconstruction difference analysis of multiple generative models of the present invention are as follows:
[0006] The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models is mainly characterized in that the method includes the following steps:
[0007] (1) Construct an image acquisition and dataset;
[0008] (2) Perform image reconstruction;
[0009] (3) Perform reconstruction difference analysis;
[0010] (4) Perform classification detection.
[0011] Preferably, step (1) specifically includes the following steps:
[0012] (1.1) Use the FFHQ dataset as the source of real images, and select multiple images containing features from it as the real image part;
[0013] (1.2) Use the classical GAN model and DM model to generate synthetic images;
[0014] (1.3) Divide the real images and synthetic images into training sets, validation sets, and test sets;
[0015] (1.4) Perform category annotation on each image.
[0016] Preferably, step (2) specifically includes the following steps:
[0017] (2.1) Use a pre-trained StyleGAN encoder to extract the features of the input image, generate hierarchical latent codes and feature codes; input the extracted latent codes into the pre-trained StyleGAN for GAN reconstruction to generate a reconstructed image;
[0018] (2.2) Use DDIM inversion to map the input image into the latent space of DM, and input the inverted latent codes into the pre-trained ADM for DM reconstruction to generate a reconstructed image.
[0019] Preferably, for the GAN reconstruction in step (2.1), specifically:
[0020] Perform GAN reconstruction according to the following formula:
[0021]
[0022] where z * ∈Z represents the optimal latent vector in the latent space Z, G(·) is the pre-trained generator function of the GAN, which maps the latent space Z to the image space, and the function is a loss function used to measure the difference between the generated image G(z) and the target image X.
[0023] Preferably, in the step (2.2), DM reconstruction is performed specifically as follows:
[0024] DM reconstruction is performed according to the following formula:
[0025]
[0026] where x t is the latent code at time step t, x0 is the original image, and α t is the noise scheduling parameter at time step t, and ∈ is Gaussian noise.
[0027] Preferably, the step (3) specifically includes the following steps:
[0028] (3.1) Use a pre-trained convolutional neural network to extract the feature maps of the original image, the GAN-reconstructed image, and the DM-reconstructed image;
[0029] (3.2) Analyze the performance differences of different generation models during the reconstruction process by calculating the LPIPS values between the original image and the GAN-reconstructed image, and the DM-reconstructed image.
[0030] Preferably, the step (4) specifically includes the following steps:
[0031] (4.1) Take the original image, the GAN-reconstructed image, and the DM-reconstructed image as inputs and input them into a convolutional neural network for classification;
[0032] (4.2) Use a triplet cross-entropy loss function for optimization.
[0033] Preferably, in the step (4.2), using a triplet cross-entropy loss function for optimization is specifically as follows:
[0034] Use a triplet cross-entropy loss function for optimization according to the following formula:
[0035]
[0036] where y is the one-hot encoded vector of the true label, is the predicted probability distribution, C is the number of classification results, and C = 3.
[0037] The device for realizing synthetic face detection based on the reconstruction difference analysis of multiple generation models is mainly characterized in that the device includes:
[0038] A processor configured to execute computer-executable instructions;
[0039] A memory stores one or more computer-executable instructions, which, when executed by the processor, implement each step of the above method for synthetic face detection based on multi-generation model reconstruction difference analysis.
[0040] The processor for synthetic face detection based on multi-generation model reconstruction difference analysis is characterized in that the processor is configured to execute computer-executable instructions, which, when executed by the processor, implement each step of the above method for synthetic face detection based on multi-generation model reconstruction difference analysis.
[0041] The computer-readable storage medium is characterized in that a computer program is stored thereon, and the computer program can be executed by a processor to implement each step of the above method for synthetic face detection based on multi-generation model reconstruction difference analysis.
[0042] By adopting the method, device, processor and computer-readable storage medium for synthetic face detection based on multi-generation model reconstruction difference analysis of the present invention, it is possible to quickly, accurately and reasonably and effectively distinguish real images, GAN-generated images and DM-generated images, with high detection accuracy, strong generalization ability and robustness, and make up for the deficiency of Asian face data in the existing dataset. Description of the Drawings
[0043] Figure 1 It is a flowchart of the method for synthetic face detection based on multi-generation model reconstruction difference analysis of the present invention.
[0044] Figure 2 It is a schematic diagram of a synthetic face detector for multi-generation model reconstruction difference analysis of the method for synthetic face detection based on multi-generation model reconstruction difference analysis of the present invention.
[0045] Figure 3 It is a schematic diagram of the reconstruction effects of images of each category under different reconstruction methods of the method for synthetic face detection based on multi-generation model reconstruction difference analysis of the present invention.
[0046] Figure 4 It is a schematic diagram of the reconstruction difference distribution of the method for synthetic face detection based on multi-generation model reconstruction difference analysis of the present invention.
[0047] Figure 5 It is a schematic diagram of the constructed Asian synthetic face dataset ASFD of the method for synthetic face detection based on multi-generation model reconstruction difference analysis of the present invention, where all faces are synthetic images. Detailed Embodiments
[0048] To more clearly describe the technical content of the present invention, the following will be further described in conjunction with specific embodiments.
[0049] The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generative models of the present invention includes the following steps:
[0050] (1) Construct an image acquisition and dataset;
[0051] (2) Perform image reconstruction;
[0052] (3) Perform reconstruction difference analysis;
[0053] (4) Perform classification detection.
[0054] As a preferred embodiment of the present invention, the step (1) specifically includes the following steps:
[0055] (1.1) Use the FFHQ dataset as the source of real images, and select multiple images containing features from it as the real image part;
[0056] (1.2) Use the classical GAN model and DM model to generate synthetic images;
[0057] (1.3) Divide the real images and synthetic images into training sets, validation sets, and test sets;
[0058] (1.4) Perform category annotation on each image.
[0059] As a preferred embodiment of the present invention, the step (2) specifically includes the following steps:
[0060] (2.1) Use a pre-trained StyleGAN encoder to extract the features of the input image, generate hierarchical latent codes and feature codes; input the extracted latent codes into the pre-trained StyleGAN for GAN reconstruction to generate a reconstructed image;
[0061] (2.2) Use DDIM inversion to map the input image into the latent space of DM, and input the inverted latent codes into the pre-trained ADM for DM reconstruction to generate a reconstructed image.
[0062] As a preferred embodiment of the present invention, in the step (2.1), the GAN reconstruction is specifically:
[0063] Perform GAN reconstruction according to the following formula:
[0064]
[0065] where z *∈Z represents the optimal latent vector in the latent space Z, and G(·) is the pre-trained generator function of the GAN, which maps the latent space Z to the image space. The function is a loss function used to measure the difference between the generated image G(z) and the target image X.
[0066] As a preferred embodiment of the present invention, in the step (2.2), DM reconstruction is performed specifically as follows:
[0067] DM reconstruction is performed according to the following formula:
[0068]
[0069] where x t is the latent code at time step t, x0 is the original image, α t is the noise scheduling parameter at time step t, and ∈ is Gaussian noise.
[0070] As a preferred embodiment of the present invention, the step (3) specifically includes the following steps:
[0071] (3.1) Use a pre-trained convolutional neural network to extract the feature maps of the original image, the GAN reconstructed image, and the DM reconstructed image;
[0072] (3.2) By calculating the LPIPS values between the original image and the GAN reconstructed image, and the DM reconstructed image, analyze the performance differences of different generation models during the reconstruction process.
[0073] As a preferred embodiment of the present invention, the step (4) specifically includes the following steps:
[0074] (4.1) Take the original image, the GAN reconstructed image, and the DM reconstructed image as inputs and input them into a convolutional neural network for classification;
[0075] (4.2) Use a triplet cross-entropy loss function for optimization.
[0076] As a preferred embodiment of the present invention, in the step (4.2), using a triplet cross-entropy loss function for optimization is specifically as follows:
[0077] Optimize using a triplet cross-entropy loss function according to the following formula:
[0078]
[0079] where y is the one-hot encoded vector of the true label, is the predicted probability distribution, C is the number of classification results, and C = 3.
[0080] The device for realizing synthetic face detection based on multi - generation model reconstruction difference analysis according to the present invention, wherein the device includes:
[0081] A processor configured to execute computer - executable instructions;
[0082] A memory storing one or more computer - executable instructions, which, when executed by the processor, implement each step of the method for realizing synthetic face detection based on multi - generation model reconstruction difference analysis as described above.
[0083] The processor for realizing synthetic face detection based on multi - generation model reconstruction difference analysis according to the present invention, wherein the processor is configured to execute computer - executable instructions, which, when executed by the processor, implement each step of the method for realizing synthetic face detection based on multi - generation model reconstruction difference analysis as described above.
[0084] The computer - readable storage medium according to the present invention, on which a computer program is stored, and the computer program can be executed by a processor to implement each step of the method for realizing synthetic face detection based on multi - generation model reconstruction difference analysis as described above.
[0085] In the specific implementation manner of the present invention, a synthetic face detection algorithm based on multi - generation model reconstruction difference analysis is provided to solve the problems of poor generalization ability and insufficient robustness of the synthetic face detection method in the prior art when facing a mixed scenario (i.e., images generated by both GAN and DM exist simultaneously), and an Asian synthetic face dataset is constructed to make up for the deficiency of Asian face data in the existing dataset.
[0086] The algorithm of the present invention analyzes the reconstruction differences of images under different generation models by combining the multi - reconstruction strategies of the generative adversarial network (GAN) and the diffusion model (DM) to achieve accurate detection of real faces, GAN - generated faces, and DM - generated faces. The specific scheme includes: (1) constructing an Asian synthetic face dataset (ASFD), which contains 11,000 Asian real faces and synthetic images generated by various GAN / DM models; (2) using a pre - trained model to perform GAN inversion reconstruction and DM inversion reconstruction on the input image to generate corresponding reconstructed images; (3) calculating the perceptual differences between the original image and the reconstructed image to capture the characteristics of different generation models; (4) cascading the original image and the multi - reconstructed images and inputting them into a convolutional neural network to achieve detection based on ternary classification.
[0087] Figures 1 to 5 The face images therein are all synthetic faces, not real faces. Figure 2 It is a schematic diagram of a synthetic face detector for multi - generation model reconstruction difference analysis, and the faces appearing therein are also all synthetic face images. Figure 3The figure shows the schematic diagrams of the reconstruction effects of images of each category under different reconstruction methods, divided into three rows. The first row is the original image, where the first one in the first row is a real face, and both it and the reconstructed image have been blurred. The second row is the synthetic image reconstructed by the GAN model, and the third row is the synthetic image reconstructed by the DM model. Figure 5 It is a schematic diagram of the constructed Asian synthetic face dataset ASFD, where all faces are synthetic images.
[0088] Such as Figures 1 to 4 , a synthetic face detection algorithm based on the analysis of reconstruction differences of multiple generative models, at least including the following steps:
[0089] (1) Image acquisition and dataset construction: Use the FFHQ dataset as the source of real images, and select 11,000 images with Asian face features from it as the real image part; Use four classic GAN models (StyleGAN1, StyleGAN2, ProGAN, VQGAN) and four DM models (ADM, IDDPM, LDM, SDE) to generate 10,000 synthetic images with a resolution of 256x256 pixels; Divide the real images and synthetic images into a training set, a validation set, and a test set according to a ratio of 7:2:1. The training set contains images generated by ADM and StyleGAN and their corresponding real images, and the test set contains images generated by StyleGAN2, VQGAN, IDDPM, and LDM and their corresponding real images; Perform category annotation on each image, and the annotation categories include real images, GAN-generated images, and DM-generated images.
[0090] (2) Image reconstruction: Use a pre-trained StyleGAN encoder to extract the features of the input image, generate hierarchical latent codes and feature codes, and map them into the latent space of the GAN, and then use a pre-trained StyleGAN to generate the reconstructed image; Use DDIM (Denoising Diffusion Implicit Models) inversion to map the input image into the latent space of the DM, and then use a pre-trained ADM (Advanced Diffusion Model) to generate the reconstructed image.
[0091] (3) Reconstruction difference analysis: Analyze the performance differences of different generative models in the reconstruction process by calculating the differences between the original image and the GAN-reconstructed image and the DM-reconstructed image.
[0092] (4) Classification detection: Use the original image, the GAN-reconstructed image, and the DM-reconstructed image as inputs, input them into a convolutional neural network for classification, and use a triple cross-entropy loss function for optimization.
[0093] Preferably, step (1) specifically includes the following steps:
[0094] (1.1) Select 11,000 images with Asian face features from the FFHQ dataset to ensure the representativeness of Asian faces in the dataset; the screening criteria include face features, skin color, and facial structure to ensure that the selected images have typical Asian face features.
[0095] (1.2) Use four classic GAN models (StyleGAN1, StyleGAN2, ProGAN, VQGAN) and four DM models (ADM, IDDPM, LDM, SDE) to generate 10,000 synthetic images with a resolution of 256x256 pixels; during the generation process, ensure that the number of images generated by each model is balanced to avoid having too many or too few images generated by a certain type of generation model in the dataset. Name the constructed Asian synthetic face dataset ASFD.
[0096] (1.3) Divide the real images and synthetic images into a training set, a validation set, and a test set according to a ratio of 7:2:1; the training set contains the images generated by ADM and StyleGAN and their corresponding real images, and the test set contains the images generated by StyleGAN2, VQGAN, IDDPM, and LDM and their corresponding real images; ensure the balanced distribution of various types of images in the training set and the test set to avoid data skew.
[0097] (1.4) Perform class annotation on each image, and the annotation classes include real images, GAN-generated images, and DM-generated images; during the annotation process, ensure the accuracy of the class of each image to avoid the impact of misannotation on the model training and test results.
[0098] Preferably, step (2) specifically includes the following steps:
[0099] (2.1) Use a pre-trained StyleGAN encoder to extract the features of the input image, generating hierarchical latent codes and feature codes; input the extracted latent codes into the pre-trained StyleGAN to generate a reconstructed image; during the GAN reconstruction process, ensure that the reconstructed image is as consistent as possible with the original image in terms of structure and details to capture the features of the GAN-generated image. The formula for GAN reconstruction is as follows:
[0100]
[0101] where z * ∈Z represents the optimal latent vector in the latent space Z, and G(·) is the pre-trained generator function of the GAN that maps the latent space Z to the image space. The function represents a loss function used to measure the difference between the generated image G(z) and the target image X.
[0102] (2.2) Use DDIM inversion to map the input image into the latent space of the DM; input the inverted latent code into the pre-trained ADM to generate a reconstructed image; during the DM reconstruction process, ensure that the reconstructed image is as consistent as possible with the original image in terms of color and texture to capture the characteristics of the images generated by the DM. The formula for DM reconstruction is as follows:
[0103]
[0104] where x t represents the latent code at time step t, x0 represents the original image, α t represents the noise scheduling parameter at time step t, and ∈ represents Gaussian noise.
[0105] Preferably, step (3) specifically includes the following steps:
[0106] (3.1) Use a pre-trained convolutional neural network (such as VGG or AlexNet) to extract the feature maps of the original image, the GAN reconstructed image, and the DM reconstructed image; during the feature extraction process, ensure that the extracted feature maps can fully reflect the details and structural information of the images.
[0107] (3.2) By calculating the LPIPS values between the original image and the GAN reconstructed image, and the DM reconstructed image, analyze the performance differences of different generation models during the reconstruction process; the smaller the LPIPS value, the more similar the reconstructed image is to the original image in terms of perception.
[0108] (3.3) For GAN-generated images, the LPIPS value between the GAN reconstructed image and the original image is smaller, while the LPIPS value between the DM reconstructed image and the original image is larger; for DM-generated images, the LPIPS value between the DM reconstructed image and the original image is smaller, while the LPIPS value between the GAN reconstructed image and the original image is larger; for real images, the LPIPS values of both the GAN reconstructed image and the DM reconstructed image with the original image are larger.
[0109] Preferably, step (4) specifically includes the following steps:
[0110] (4.1) Use ResNet-50 as the backbone network and a triplet classifier as the output layer; the input is the original image, the GAN reconstructed image, and the DM reconstructed image, and the output is the classification results of three categories (real image, GAN-generated image, DM-generated image).
[0111] (4.2) Use the triplet cross-entropy loss function for optimization, and the formula is as follows. The training dataset is ASFD; perform validation on the validation set and adjust the model parameters to ensure the generalization ability of the model on the test set.
[0112]
[0113] Among them, y is the one-hot encoded vector of the true label, is the predicted probability distribution, and C = 3 represents three-class classification.
[0114] The technical solution of the present invention can achieve the analysis of the reconstruction differences of multiple generation models. The technical solution of the present invention combines the generative adversarial network (GAN) and the diffusion model (DM), and uses the different generative characteristics of these two models to analyze the reconstruction differences of images. It deeply explores the internal relationship between the synthetic image and its generation technology, and finds that there are significant reconstruction differences when different generation methods are used for images. This is innovative compared with the prior art because most of the prior art fails to consider the reconstruction differences of multiple generation models.
[0115] The technical solution of the present invention has the Asian Synthetic Face Dataset (ASFD). To make up for the bias towards Asian faces in the existing datasets, a representative Asian Synthetic Face Dataset (ASFD) is constructed, which is not explicitly mentioned in the prior art. This dataset plays an important role in improving the generalization ability of the detection model in the Asian population.
[0116] The technical solution of the present invention can achieve ternary classification and reconstruction difference analysis. Taking the original data of the image, the GAN reconstructed image, and the DM reconstructed image as inputs, ternary classification is used to distinguish real images, GAN-generated images, and DM-generated images.
[0117] For the specific implementation solution of this embodiment, reference can be made to the relevant descriptions in the above embodiments, and details will not be elaborated here.
[0118] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be referred to the same or similar content in other embodiments.
[0119] It should be noted that in the description of the present invention, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality" means at least two.
[0120] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0121] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0122] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried out in implementing the above embodiment methods can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0123] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0124] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0125] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0126] The method, device, processor and computer-readable storage medium for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models of the present invention can quickly, accurately and effectively distinguish real images, GAN-generated images and DM-generated images based on evidence, have high detection accuracy, strong generalization ability and robustness, and make up for the deficiency of Asian face data in the existing dataset.
[0127] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it is obvious that various modifications and transformations can still be made without departing from the spirit and scope of the present invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method for synthetic face detection based on reconstruction difference analysis of multiple generation models, characterized in that, The method described above includes the following steps: (1) Construct image acquisition and dataset; (2) Perform image reconstruction; (3) Conduct reconstruction difference analysis; (4) Conduct classification detection.
2. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 1, wherein The specific steps of step (1) include the following steps: (1.1) Use the FFHQ dataset as the source of real images, and select multiple images containing features from it as the real image part; (1.2) Use classical GAN models and DM models to generate synthetic images; (1.3) Divide the real images and synthetic images into training sets, validation sets, and test sets; (1.4) Perform class annotation on each image.
3. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 1, wherein, The specific steps of step (2) include the following steps: (2.1) Use a pre-trained StyleGAN encoder to extract the features of the input image, generate hierarchical latent codes and feature codes; input the extracted latent codes into the pre-trained StyleGAN for GAN reconstruction to generate reconstructed images; (2.2) Use DDIM inversion to map the input image into the latent space of DM, and input the inverted latent codes into the pre-trained ADM for DM reconstruction to generate reconstructed images.
4. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 3, wherein, In step (2.1), the GAN reconstruction is specifically as follows: Perform GAN reconstruction according to the following formula: where z * ∈Z represents the optimal latent vector in the latent space Z, G(·) is the pre-trained generator function of the GAN that maps the latent space Z to the image space, and the function is a loss function used to measure the difference between the generated image G(z) and the target image X.
5. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 3, wherein In step (2.2), the DM reconstruction is specifically as follows: Perform DM reconstruction according to the following formula: where x t is the latent code at time step t, x0 is the original image, and α t is the noise schedule parameter at time step t, and ∈ is Gaussian noise.
6. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 1, wherein The specific steps of step (3) include the following steps: (3.1) Use a pre-trained convolutional neural network to extract the feature maps of the original image, GAN reconstructed image, and DM reconstructed image; (3.2) Analyze the performance differences of different generation models during the reconstruction process by calculating the LPIPS values between the original image and the GAN reconstructed image, and the DM reconstructed image.
7. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 1, wherein The specific steps of step (4) include the following steps: (4.1) Use the original image, GAN reconstructed image, and DM reconstructed image as inputs and input them into a convolutional neural network for classification; (4.2) Use the triplet cross-entropy loss function for optimization.
8. The method for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models according to claim 7, characterized in that, In step (4.2), using the triplet cross-entropy loss function for optimization is specifically as follows: Perform optimization using the triplet cross-entropy loss function according to the following formula: Among them, y is the one-hot encoded vector of the true label, is the predicted probability distribution, C is the number of classification results, and C = 3.
9. An apparatus for synthetic face detection based on reconstruction difference analysis of multiple generative models, characterized in that, The device described above includes: A processor configured to execute computer-executable instructions; A memory storing one or more computer-executable instructions, and when the computer-executable instructions are executed by the processor, the various steps of the method for synthetic face detection based on reconstruction difference analysis of multiple generation models described in any one of claims 1 to 6 are implemented.
10. A processor for realizing synthetic face detection based on reconstruction difference analysis of multiple generation models, characterized in that The processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, the various steps of the method for synthetic face detection based on reconstruction difference analysis of multiple generation models described in any one of claims 1 to 6 are implemented.
11. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by the processor to implement the various steps of the method for synthetic face detection based on reconstruction difference analysis of multiple generation models described in any one of claims 1 to 6.