Deep fake face image traceability evidence obtaining method and system based on diffusion model
Through the DiffMark framework, the diffusion model is used to generate and extract robust watermark images, which solves the problem of insufficient robustness in existing technologies, realizes effective traceability and evidence collection of new counterfeiting technologies, and improves the visual quality and robustness of the watermark.
Patent Information
- Application Number
- CN202510822507.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-30
AI Technical Summary
Existing deep fake facial image forensics technology is not robust enough when facing different forgery techniques, especially the diffusion model generation method has not yet been applied, resulting in limited generalization ability of watermarking methods when facing new forgery techniques.
A DiffMark framework was designed, which adopts a robust watermarking framework of the diffusion model. Through the diffusion model encoder, watermark decoder and autoencoder with frozen parameters, combined with the cross-information fusion module and DDIM sampling process, it generates and extracts more robust watermark images to adapt to specific deep fake models.
It improves the visual quality and robustness of the watermark image, can effectively resist various deep fake tampering, has stronger adaptability and generalization ability, and achieves more reliable deep fake traceability and evidence collection.
Smart Images

Figure CN120725848A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep fake facial image forensics, and in particular to a deep fake facial image source tracing and forensics method and system based on a diffusion model. Background Art
[0002] In recent years, with the development of artificial intelligence and generative models, deepfake technology has made significant progress, capable of generating highly realistic images and videos by manipulating facial identities or attributes. This technology has been widely used in fields such as film, television, and advertising, but it has also been abused to create illegal fake videos, seriously damaging personal reputations and social stability. To address the security risks posed by deepfake technology, various forensic methods have emerged. Existing deepfake forensics techniques are mainly divided into two main areas: passive forensics and active forensics. Forensic research on deepfake facial images primarily focuses on passive forensics, verifying the authenticity of facial images by analyzing traces of tampering. In contrast, active forensics, as an emerging area, mostly relies on deep watermarking technology to verify and trace the authenticity of content. However, existing deepfake facial image forensics techniques all have shortcomings.
[0003] Passive forensics techniques analyze traces of tampering left in images or videos (such as facial artifacts, inconsistent lighting conditions, and unusual biometric features) to verify authenticity. While this method doesn't require pre-processing of the original image, it often struggles to maintain stable detection performance in the face of increasingly sophisticated deepfake technology. In contrast, provenance forensics techniques achieve more reliable forensics by pre-embedding protective watermarks in the original image. These techniques typically offer greater robustness and traceability.
[0004] In the field of active deepfake forensics, technical solutions based on deep watermarking dominate. The core of this type of technology is to embed invisible watermarks in facial images or videos. These watermarks can withstand common image processing operations (such as compression, scaling, etc.) and deepfake tampering, and can be reliably extracted for tracing the source during the forensics stage. In addition, watermarks can also be combined with additional functions, such as forgery detection. The embedded watermark may not be limited to tracing the source in terms of functionality, but the function of tracing the source is important and almost indispensable. Although existing robust watermarking schemes perform well under common image processing operations (such as compression, scaling, etc.), due to the large amount of semantic disturbances brought about by deepfake tampering, the existing watermarking methods are not robust enough in the face of deepfake tampering and their generalization ability is limited. Effective deepfake tracing forensics cannot be achieved.
[0005] Furthermore, there are currently no methods based on diffusion models in the field of active deepfake forensics. As an emerging generative model, diffusion models offer advantages in image generation due to their progressive generation mechanism and conditional guidance. Therefore, exploring provenance forensics solutions based on diffusion models and developing deep watermarking methods with enhanced robustness and generalization capabilities are key issues that are both valuable and urgently needed to be addressed in the current field of active deepfake forensics.
[0006] The present invention aims to fill the gap in the field of deep fake active evidence collection, which is still lacking in methods based on diffusion models, and at the same time solve the problem of insufficient robustness of existing deep watermarking methods when dealing with different counterfeiting techniques, and further enhance the generalization ability of watermarks. With the diversification of counterfeiting techniques, existing robust watermarking methods face greater challenges, resulting in a decrease in their effectiveness when facing new counterfeiting techniques. The diffusion model has shown excellent performance in image generation. The present invention designs DiffMark, a new robust watermarking framework based on the diffusion model, for deep fake traceability and evidence collection, aiming to effectively deal with the challenges brought by deep fake technology. The present invention innovatively constructs facial images and watermarks as conditions, guiding the diffusion model to gradually denoise and generate corresponding watermark images. In order to enhance the robustness of the watermark against deep fake operations, the present invention adds a specific deep fake model to the sampling process for guidance, guides the image generation process, and generates a more robust watermark image.
[0007] Through the above-mentioned innovative design, the present invention can effectively deal with different counterfeiting techniques, provide a more robust and reliable deep counterfeit traceability and evidence collection solution, and has strong adaptability and broad application prospects. Summary of the Invention
[0008] The purpose of the present invention is to provide a deep fake facial image tracing and evidence collection method and system based on a diffusion model to solve the problem that the existing deep watermarking method proposed in the above background technology is insufficiently robust when dealing with different forgery techniques.
[0009] To achieve the above objectives, the present invention adopts the following technical solutions:
[0010] In a first aspect, the present invention proposes a method for tracing the source of deep fake facial images based on a diffusion model, comprising the following steps:
[0011] S1. Construct a robust watermarking framework based on a diffusion model; the framework includes a diffusion model encoder, a watermark decoder, and an autoencoder. The diffusion model encoder is used to generate a watermark image, the autoencoder is used to simulate the image reconstruction process in a conventional deep fake model to reconstruct the watermark image, and the watermark decoder is used to extract a traceable watermark from the reconstructed image.
[0012] S2. Train the robust watermarking framework based on the diffusion model. During the training phase, the face image and watermark are constructed as conditions to guide the diffusion model encoder to denoise and predict the corresponding watermark image at time t.
[0013] S3. Use the trained robust watermarking framework based on the diffusion model for testing; add a specific deep fake model to the sampling process in the test phase for guidance, iteratively denoise to generate a more robust watermark image; obtain a robust watermarking framework based on the diffusion model that is adapted to the specific deep fake model;
[0014] S4. Using the adapted watermark decoder of the robust watermark framework based on the diffusion model, the traceable watermark in the forged image to be forensically verified is extracted.
[0015] Preferably, the S1 is as follows:
[0016] The diffusion model encoder is designed using the backbone network U-Net of the diffusion model, and a cross-information fusion module is designed during the upsampling process of the diffusion model encoder;
[0017] The autoencoder adopts a VQGAN model with pre-trained frozen parameters;
[0018] The watermark decoder is designed using the downsampling part of the diffusion model encoder.
[0019] Furthermore, the cross-information fusion module is specifically as follows:
[0020] The cross-information fusion module adopts a learnable embedding table and a cross-attention mechanism;
[0021] Feature extraction is performed through a learnable embedding table. Continuous embedding is established for the bit values at each position in the binary watermark sequence. Watermark features are adaptively extracted through an embedding table lookup mechanism. The embedding table adjusts the 1-dimensional binary watermark to a 2-dimensional feature through indexing. Image features are also reduced in dimensionality, from 3-dimensional to 2-dimensional.
[0022] Feature fusion is performed through a cross-attention mechanism; image features and watermark features are projected through separate linear layers, and the fusion process uses cross-attention. The image features are used as the query tensor Q, and the transformed watermark feature tensors are used as the key-value tensors K and V. Feature fusion is performed through the attention formula, and finally a residual connection is made with the initial image features to obtain the fused features.
[0023] Preferably, the S2 is as follows:
[0024] The diffusion model encoder, watermark decoder and autoencoder together constitute an end-to-end training framework of a robust watermark framework based on the diffusion model;
[0025] Add t-step noise to the original image x0 to obtain the noisy image x t ; The diffusion model encoder performs the watermarking task and predicts the watermarked image An autoencoder with pre-trained frozen parameters is used to simulate the image reconstruction process in a conventional deep fake model to reconstruct the watermarked image. Improve the robustness of the watermark; use a watermark decoder to extract a traceable watermark from the image reconstructed by the autoencoder to trace the image.
[0026] Furthermore, when the diffusion model encoder performs the watermarking task, the noise image x t and time step t as a conditional input to the diffusion model encoder, and uses the coefficients of the noise term ∈ Dynamically scale the original image x0 and As another conditional input to the diffusion model encoder, the watermark w is also input.
[0027] Preferably, the S3 is as follows:
[0028] The trained diffusion model encoder and watermark decoder are used in the testing phase of the robust watermarking framework based on the diffusion model;
[0029] The diffusion model encoder is applied to DDIM sampling. Starting from the standard Gaussian distribution, the corresponding watermark image is generated by guiding the face and watermark conditions. A specific deep fake model is introduced into the DDIM sampling process to improve the robustness of the watermark image. The watermark image is tampered with by a specific deep fake model to obtain a forged image, and a watermark decoder is used to extract a traceable watermark from the forged image.
[0030] Furthermore, the generation of the corresponding watermark image is as follows:
[0031] From the standard Gaussian distribution x T Start sampling and denoise x by T-step iterative denoising. T Gradually passing through x t to x t-1 The conversion process is converted to x0;
[0032] In each iteration step, the coefficient in front of the noise term of the previous step is used For the original image x co Zoom in and get the face condition x c ; and in each iteration step, the t-step noise image x t , face condition x c , time step t and watermark w are fed into the diffusion model encoder to predict the watermark image Watermark image predicted by diffusion model encoder DDIM sampling is performed. In this process, a specific deep fake model is used to tamper with the fake image, and then the watermark is extracted using a watermark decoder. The bit error between the extracted watermark and the embedded watermark is used to compare x t Find the gradient, affecting x t to x t-1 The conversion process does not require targeted retraining, which saves time and effectively improves the robustness of the watermark against new forgery models.
[0033] Finally, after T steps of iteration, the watermark image corresponding to the original image is generated.
[0034] Preferably, the S4 is specifically as follows:
[0035] By performing a single-step decoding operation using a watermark decoder, the watermark can be extracted from the forged image to be forensically verified, thus achieving traceability. This does not rely on the multi-step iteration of the diffusion model sampling process, saving time.
[0036] In a second aspect, the present invention proposes a deep fake facial image source tracing and evidence collection system based on a diffusion model, comprising:
[0037] Diffusion model encoder, which uses the backbone network U-Net of the diffusion model to generate watermark images;
[0038] The autoencoder uses the VQGAN model to simulate deep fake tampering and reconstruct the watermarked image;
[0039] The watermark decoder uses the downsampling part of the backbone network U-Net of the diffusion model to extract the traceable watermark from the reconstructed watermark image.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] (1) The robust watermarking framework based on the diffusion model in the present invention has better visual quality. The traditional pixel space watermarking method directly embeds the watermark into the face image in the pixel space through a deep neural network; the latent space watermarking method converts the face image into a feature representation in the latent space, then embeds the watermark in the latent space, and finally maps it back to the pixel space to obtain the watermark image. In comparison, the method in the present invention relies on the powerful image generation ability of the diffusion model, constructs the face image and the watermark as conditions, guides the diffusion model to start from the standard Gaussian distribution, gradually denoises and generates the required watermark image, and improves the overall visual quality of the watermark image without affecting the watermark extraction effect.
[0042] (2) The robust watermarking framework based on the diffusion model in the present invention has better robustness. Existing watermarking methods often lack robustness against deep forgery and tampering, and often show poor generalization ability when facing new forgery technologies. The present invention introduces a frozen parameter autoencoder in the training stage to simulate deep forgery and tampering, thereby improving the robustness of the watermark; in addition, it utilizes the multi-step iterative characteristics of the diffusion model sampling process, adds a specific deep forgery model to guide the sampling process, and generates a more robust watermark image that can resist deep forgery and tampering. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Schematic diagram of the training phase and the testing phase of the robust watermarking framework based on the diffusion model in the present invention;
[0044] Figure 2 This is a structural block diagram of the cross-information fusion module in the present invention. DETAILED DESCRIPTION
[0045] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0046] Example 1:
[0047] The key technical point of the deep fake facial image traceability and evidence collection method based on the diffusion model is the proposed DiffMark framework, a new robust watermarking framework based on the diffusion model, which is used for deep fake facial image traceability and evidence collection. The present invention innovatively constructs facial images and watermarks as conditions, guiding the diffusion model to gradually denoise and generate corresponding watermark images. In addition, in order to address the problem that existing watermarking methods are not robust enough in the face of various counterfeiting techniques, the present invention adds a specific deep fake model to the sampling process of the diffusion model for guidance, generating a more robust watermark image that can resist deep fake tampering.
[0048] Step 1: Construct a robust watermarking framework based on the diffusion model.
[0049] The robust watermarking framework based on the diffusion model includes a diffusion model encoder, a watermark decoder and an autoencoder with frozen parameters. The encoder performs the image generation task with the watermark task. The autoencoder can simulate the image reconstruction process in most deep fake models. The watermark decoder is used to extract the watermark and assist the encoder in predicting the watermark image.
[0050] First, a conditional diffusion model that fuses face and watermark is constructed as the diffusion model encoder;
[0051] The diffusion process of DiffMark is performed in pixel space rather than latent space, because the mapping from pixel space to latent space often discards image details and only retains semantic consistency, which is detrimental to the imperceptibility of watermarks. The backbone network of the diffusion model selected by this invention is U-Net, that is, Figure 1 For the watermarking task, it is obvious that the watermark needs to be the input of the encoder. In order to achieve the fusion of watermark conditions, the present invention designs a cross information fusion (CIF) module in the upsampling process of the encoder.
[0052] In this embodiment, the cross information fusion (CIF) module is specifically as follows:
[0053] In order to achieve the fusion of watermark conditions, a cross-information fusion (CIF) module is designed in the encoder (U-Net), which is based on a learnable embedding table and a cross-attention mechanism to achieve deep fusion of watermark and facial features.
[0054] like Figure 2 As shown, the cross-information fusion module includes a learnable embedding table that establishes a continuous embedding for the bit value at each position in the binary watermark sequence. The embedding table lookup mechanism enables adaptive extraction of watermark features. The embedding table converts the one-dimensional binary watermark into a two-dimensional feature through indexing. To facilitate feature fusion, the image features undergo dimensionality reduction, adjusting from three dimensions to two. Feature fusion is then performed using cross-attention. The facial image and watermark features are first projected through separate linear layers. The fusion process uses cross-attention, with the image features serving as the query tensor Q and the transformed watermark feature tensors serving as the key-value tensors K and V. Feature fusion is performed using an attention formula and finally a residual connection is made with the initial image features to obtain the fused features. This effectively integrates the semantic features of the watermark while preserving the original facial image features.
[0055] Second, choose an autoencoder with frozen parameters;
[0056] This paper introduces a pre-trained frozen parameter autoencoder (VQGAN). By using the autoencoder, the image reconstruction process in most deepfake models is simulated, which improves the robustness of the watermark against deepfake tampering to a certain extent.
[0057] Finally, design the watermark decoder;
[0058] Unlike the traditional diffusion model prediction noise, this invention uses the coefficient in front of the noise term to dynamically scale the original image and provides it as a conditional input to the diffusion model. Therefore, it is easier to make the model converge by directly predicting the original image x0 instead of the noise ∈. However, since this is a watermarking task, the model expects the prediction to be the watermark image In order to ensure the existence and extractability of the watermark, the present invention extracts the downsampling part of the encoder (U-Net) network of the diffusion model as a watermark decoder to extract the watermark and assist the encoder in predicting the watermark image.
[0059] In this way, the encoder, decoder, and frozen autoencoder together constitute the end-to-end training framework of DiffMark.
[0060] Step 2: Train the robust watermarking framework based on the diffusion model.
[0061] like Figure 1 As shown in the upper part of the figure, in the training phase, the present invention modifies the training process of the diffusion model to adapt to the watermark task. The traditional diffusion model requires the noise image x t (obtained by adding t-step noise to the original image x0) and time step t as input, the encoder (U-Net) is trained to predict the noise added at time step t, thereby achieving iterative denoising generation, which is the image generation task.
[0062] However, the encoder trained in this way has great randomness when denoising a pure Gaussian distribution image. This is very beneficial for image generation tasks and can generate images with diverse styles, but it is not suitable for watermarking tasks because the present invention needs to watermark a specific image. Therefore, in order to reduce the randomness, the present invention uses the coefficient of the noise term ∈ To dynamically scale the original image x0 and use it as another conditional input to the encoder. This makes it so that as the time step t increases during the diffusion process, the noise intensity increases, that is, the uncertainty becomes stronger, and the guidance condition becomes stronger, increasing the certainty of denoising while achieving progressive conditional guidance to adapt to the deterministic watermarking task. The encoder performs the watermarking task and predicts the watermarked image It represents the original image x0 containing the invisible watermark w.
[0063] Since the purpose of DiffMark is to be used for deep fake facial image tracing and evidence collection, the present invention uses a pre-trained frozen parameter autoencoder (VQGAN) during the training phase. By using the autoencoder, the present invention simulates the image reconstruction process in most deep fake models, and to a certain extent improves the robustness of the watermark against deep fake operations. The downsampling part of the encoder (U-Net) network of the diffusion model is then used as a watermark decoder to perform a single-step decoding operation to extract the watermark from the watermarked image that has been tampered with by deep fakes for tracing and evidence collection.
[0064] Step 3: Test the robust watermarking framework based on the diffusion model.
[0065] DiffMark mainly uses DDIM sampling to gradually remove noise and generate watermark images. T The sampling process from x0 to x0 includes many steps. The present invention applies the encoder (U-Net) trained by the end-to-end training framework to DDIM sampling to help x t Convert to x t-1 .like Figure 1 shown in the lower half of the .
[0066] A significant advantage of the diffusion model is that once it is trained, the image generation process can be controlled by introducing guidance during sampling. In this paper, deep fake guidance is introduced during the DDIM sampling process to obtain more robust watermarked images.
[0067] First, from the standard Gaussian distribution x T Start sampling, and gradually reduce x by T steps of iterative denoising. T Convert to x0. In each iteration step, first use the coefficient in front of the noise term of the previous step For the original image x co Zoom in and get the face condition x c Then, the t-step noise image x t , face condition x c , time step t and watermark w are fed into the encoder to predict the watermark image. In addition, a specific deep fake model can be introduced during the sampling process to guide the denoising process. Specifically, the watermark image predicted by the encoder is then tampered with by the deep fake model to obtain the forged image. The decoder is then used to extract the watermark, and the cross entropy loss between the extracted watermark and the embedded watermark and the mean square error between the original image and the watermarked image are used to influence x. t to x t-1 Finally, after T steps of iteration, the watermark image corresponding to the original image will be generated, which has better robustness.
[0068] While watermark embedding is achieved through a multi-step denoising process within a diffusion model, which utilizes the face and watermark conditions to generate the watermarked image, watermark extraction is independent of this embedding process. Tracing the watermark requires only a single-step decoding operation via the decoder. This eliminates the need for targeted retraining, saving time and effectively improving the watermark's robustness against new forgery techniques.
[0069] Experimental verification:
[0070] In the watermark embedding invisibility experiment, we evaluated the average PSNR, SSIM, and LPIPS between the original and watermarked face images. These three metrics are used to assess the differences in image quality, structural similarity, and perceptual similarity between the watermarked and original images. The results are shown in Table 1.
[0071] Table 1 Image quality evaluation after watermark embedding
[0072] Methods PSNR↑ SSIM↑ LPIPS↓ MBRS 35.3970 0.8935 0.1027 CIN 39.7044 0.9308 0.0248 ARWGAN 38.5746 0.9733 0.0183 SepMark 38.3129 0.9599 0.0196 EditGuard 37.1664 0.9516 0.0746 Ours 41.3005 0.9776 0.0090
[0073] As shown in Table 1, the proposed method maintains the best visual quality at a resolution of 128 × 128, outperforming previous watermarking methods. This indicates that the watermarked image generated by the proposed DiffMark is very similar to the original image, with almost no perceptible visual difference.
[0074] In the watermark robustness experiment, the present invention uses the bit error rate as an evaluation indicator. A lower bit error rate indicates a higher accuracy of watermark extraction, and a lower average bit error rate under various interferences indicates better robustness of the watermark. In Table 2, the present invention evaluates the bit error rate under attacks by representative deep fake models, including SimSwap, UniFace, CSCS, StarGAN, FSRT, and VQGAN. It should be noted that VQGAN is only used as an automatic encoder for image reconstruction. Because it is included in the training framework of the present invention, it is used as a reference for comparison.
[0075] Table 2 Bit error rate analysis of watermark extraction after deep fake tampering
[0076] Distortion MBRS CIN ARWGAN SepMark EditGuard Ours SimSwap 24.52% 40.08% 46.60% 20.02% 45.32% 5.58% UniFace 0.43% 11.80% 26.28% 0.34% 9.17% 0.01% CSCS 9.88% 0.49% 6.29% 0.68% 0.99% 0.13% StarGAN 11.68% 56.93% 36.78% 0.11% 7.62% 4.63% FSRT 49.84% 3.20% 4.34% 50.22% 5.77% 0.12% VQGAN 0.05% 39.60% 35.06% 1.28% 6.38% 0.02% Average 16.07% 25.35% 25.89% 12.11% 12.54% 1.75%
[0077] As shown in Table 2, the experimental results show that our proposed method outperforms all compared methods in average bit error rate at 128×128 resolution. Notably, SepMark exhibits an exceptionally low bit error rate on StarGAN, likely due to targeted optimizations of StarGAN within SepMark's training framework. However, our proposed DiffMark maintains the best average bit error rate, demonstrating its robustness against various deepfake manipulations.
[0078] The above description is only used to help understand the method and core essence of the present invention, but the scope of protection of the present invention is not limited thereto. For those skilled in the art, equivalent replacements or modifications based on the technical solutions and inventive concepts of the present invention within the technical scope disclosed by the present invention should be included in the scope of protection of the present invention. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A deep fake face image tracing and evidence collection method based on a diffusion model, characterized by: The following steps are involved: S1. Construct a robust watermarking framework based on a diffusion model; the framework includes a diffusion model encoder, a watermark decoder, and an autoencoder. The diffusion model encoder is used to generate a watermark image, the autoencoder is used to simulate the image reconstruction process in a conventional deep fake model to reconstruct the watermark image, and the watermark decoder is used to extract a traceable watermark from the reconstructed image. S2, training the robust watermarking framework based on the diffusion model; In the training phase, the face image and watermark are constructed as conditions to guide the diffusion model encoder to denoise and predict the corresponding watermark image; S3,testing using the trained robust watermarking framework based on diffusion model; During the sampling process of the test phase, a specific deep fake model is added for guidance, and watermark images are generated through iterative denoising. A robust watermarking framework based on a diffusion model is obtained that adapts to specific deep fake models; S4. Using the adapted watermark decoder of the robust watermark framework based on the diffusion model, the traceable watermark in the forged image to be forensically verified is extracted.
2. The deep fake face image tracing and evidence collection method based on the diffusion model according to claim 1 is characterized in that: The S1 is specifically as follows: The diffusion model encoder is designed using the backbone network U-Net of the diffusion model, and a cross-information fusion module is designed during the upsampling process of the diffusion model encoder; The autoencoder adopts a VQGAN model with pre-trained frozen parameters; The watermark decoder is designed using the downsampling part of the diffusion model encoder.
3. The deep fake face image tracing and evidence collection method based on the diffusion model according to claim 2 is characterized in that: The cross information fusion module is specifically as follows: The cross-information fusion module adopts a learnable embedding table and a cross-attention mechanism; Feature extraction is performed through a learnable embedding table. Continuous embedding is established for the bit values at each position in the binary watermark sequence. Watermark features are adaptively extracted through an embedding table lookup mechanism. The embedding table adjusts the 1-dimensional binary watermark to a 2-dimensional feature through indexing. Image features are also reduced in dimensionality, from 3-dimensional to 2-dimensional. Feature fusion is performed through a cross-attention mechanism; image features and watermark features are projected through separate linear layers, and the fusion process uses cross-attention. The image features are used as the query tensor Q, and the transformed watermark feature tensors are used as the key-value tensors K and V. Feature fusion is performed through the attention formula, and finally a residual connection is made with the initial image features to obtain the fused features.
4. The deep fake face image tracing and evidence collection method based on the diffusion model according to claim 1 is characterized in that: The S2 is specifically as follows: The diffusion model encoder, watermark decoder and autoencoder together constitute an end-to-end training framework of a robust watermark framework based on the diffusion model; Add t-step noise to the original image x0 to obtain the noisy image x t ; The diffusion model encoder performs the watermarking task and predicts the watermarked image An autoencoder with pre-trained frozen parameters is used to simulate the image reconstruction process in a conventional deep fake model to reconstruct the watermarked image. A watermark decoder is used to extract a traceable watermark from the image reconstructed by the autoencoder to trace the image.
5. The deep fake face image tracing and evidence collection method based on the diffusion model according to claim 4 is characterized in that: When the diffusion model encoder performs the watermarking task, the noise image x t and time step t as a conditional input to the diffusion model encoder, and uses the coefficients of the noise term ∈ Dynamically scale the original image x0 and As another conditional input to the diffusion model encoder, the watermark w is also input.
6. The deep fake face image source tracing and evidence collection method based on the diffusion model according to claim 1 is characterized in that: The S3 is specifically as follows: The trained diffusion model encoder and watermark decoder are used in the testing phase of the robust watermarking framework based on the diffusion model; The diffusion model encoder is applied to DDIM sampling. Starting from the standard Gaussian distribution, the corresponding watermark image is generated by guiding the face and watermark conditions. A specific deep fake model is introduced into the DDIM sampling process to improve the robustness of the watermark image against deep fake tampering. The watermark image is tampered with by a specific deep fake model to obtain a forged image, and a watermark decoder is used to extract a traceable watermark from the forged image.
7. The deep fake face image tracing and evidence collection method based on the diffusion model according to claim 6 is characterized in that: The generation of the corresponding watermark image is as follows: From the standard Gaussian distribution x T Start sampling and denoise x by T-step iterative denoising. T Gradually passing through x t to x t-1 The conversion process is converted to x0; In each iteration step, the coefficient in front of the noise term of the previous step is used For the original image x co Zoom in and get the face condition x c ; and in each iteration step, the t-step noise image x t , face condition x c , time step t and watermark w are fed into the diffusion model encoder to predict the watermark image Watermark image predicted by diffusion model encoder DDIM sampling is performed. In this process, a specific deep fake model is used to tamper with the fake image, and then the watermark is extracted using a watermark decoder. The bit error between the extracted watermark and the embedded watermark is used to compare x t Find the gradient, affecting x t to x t-1 The conversion process; Finally, after T steps of iteration, the watermark image corresponding to the original image is generated.
8. A deep fake facial image tracing and evidence collection system based on a diffusion model applied to the method according to any one of claims 1 to 7, characterized in that: include: Diffusion model encoder, which uses the backbone network U-Net of the diffusion model to generate watermark images; The autoencoder uses the VQGAN model to simulate deep fake tampering and reconstruct the watermarked image; The watermark decoder uses the downsampling part of the backbone network U-Net of the diffusion model to extract the traceable watermark from the reconstructed watermark image.
Citation Information
Cited By
Deep forgery detection method based on quantum control robust feature watermark
CN120931466A
Watermark anti-counterfeiting evaluation method and system based on diffusion model, electronic equipment and storage medium
CN120931468A
A watermark anti-counterfeiting evaluation method and system based on a diffusion model, an electronic device, and a storage medium
CN120931468B
Robust adversarial watermark generation method, device and equipment for face deep counterfeiting
CN121837009A