Active defense method for image forgery based on multiple watermark fusion and cross-domain learning
By embedding multiple watermarks in the image and combining cross-domain learning technology, the problems of insufficient generalization and low detection efficiency of existing deep forgery defense technologies are solved, and the image identification and traceability functions without computing resource occupation are realized, which improves the robustness and efficiency of defense.
Patent Information
- Application Number
- CN202411918332.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing deep forgery defense technology has insufficient generalization and low detection efficiency of new models. The active interference method is fragile when spreading on social media, affecting the benign use of deep forgery.
An active defense method for image forgery based on multiple watermark fusion and cross-domain learning is adopted. Invisible and visible watermarks are embedded in the image through a watermark encoder, and false warning signs are generated through the visible watermark after the noise layer is processed, and the image traceability and detection are carried out in combination with the watermark decoder.
Deep fake image identification without occupancy of computing resources is realized. The naked eye can determine whether the image has been forged, and verify the authenticity of the image source through invisible watermark traceability and detection, which improves the robustness and efficiency of defense.
Smart Images

Figure CN119379524B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to an active defense method for image forgery based on multiple watermark fusion and cross-domain learning. Background Art
[0002] Deep fake technology uses deep learning, especially generative adversarial network (GAN) technology, to generate or tamper with realistic audio, video and images. Deep fakes initially demonstrated the amazing effects of deep learning on image generation, such as image restoration, cross-age generation, expression transfer, etc. However, after the widespread application of deep fake technology, it has also brought serious social problems, such as privacy leakage, dissemination of false information and identity fraud, especially in the fields of film and television, media, social networks, etc., posing potential threats. Therefore, deep fake defense has become an urgent research direction.
[0003] At present, deep fake defense technology is mainly divided into active defense and passive defense. Passive defense mainly involves passive detection, which is the traditional binary classification problem, to identify true and false. Many studies use deep convolutional neural networks to classify deep fake images. Typical methods include building decision models through neural networks such as CNN, ResNet, and Xception, and achieving high accuracy in face deep fake detection. However, this defense has the problem of insufficient generalization and low efficiency in detecting new models. Active defense is mainly divided into active interference, active tracing, and active detection. Active interference is to add carefully designed interference information to the image, so that the forgery algorithm cannot work properly, and the generated forged content is distorted or unnatural. The current mainstream addition of carefully designed interference information is adversarial perturbation, which is to add tiny imperceptible perturbations to the image, so that the deep fake model misjudges or outputs abnormal results. This perturbation usually finds small changes in the model input space that are confusing to the model, and realizes active interference of the model. However, this imperceptible perturbation is very fragile when it is spread on social media, and active interference is easy to fail. At the same time, this defense method will also affect the benign use of deep fakes to a certain extent. Active tracing focuses on tracking the generation process of deep fake images or identifying the source of the generation tool in order to locate the source or tool of the forged content. The current mainstream method is to embed a secret watermark into the image, which is often used for copyright protection. For the tracing of deep fake images, many studies add secret watermarks to the image, and extract the watermark to ensure the true source information of the image. Active detection also adds a secret watermark to the image, but the watermark is robust to common image processing methods and sensitive to deep fake models, so that the watermark information of the image embedded with the watermark will be destroyed after deep fake, while the watermark information can be correctly extracted after normal image processing. This method can be used to determine whether the image has been deeply faked. The robustness of the method based on active tracing and active detection is better than that of active interference, but it consumes a lot of computing resources when tracing or detecting, and it is not possible to directly identify whether the image has been deeply faked by the naked eye. Summary of the invention
[0004] In order to solve the above problems, the present invention provides an active defense method for image forgery based on multiple watermark fusion and cross-domain learning.
[0005] In order to achieve the above object, the present invention is implemented through the following technical solutions:
[0006] The present invention provides an active defense method for image forgery based on multiple watermark fusion and cross-domain learning, comprising the following steps:
[0007] S1. Obtain the image to be processed;
[0008] S2. The image to be processed is embedded with an invisible watermark and a visible watermark through a watermark encoder to obtain an image embedded with an invisible watermark and an image embedded with a visible watermark respectively;
[0009] S3. The image embedded with the invisible watermark is processed through the noise layer to obtain a noise image;
[0010] S4. The image embedded with the visible watermark is processed through the noise layer, and an obvious false warning mark is generated at the image position where the random noise is embedded through joint optimization of the visible watermark, which is used to determine whether the image has been deeply forged;
[0011] S5. The noise image is traced and detected by the watermark decoder to determine the authenticity of the image;
[0012] S6. Perform supervised training of the loss function.
[0013] Furthermore, in step S1, the image to be processed is obtained from the image dataset and the video dataset. .
[0014] Furthermore, the watermark encoder in step S2 includes an invisible watermark embedding module and a visible watermark injection module, specifically:
[0015] S21. Invisible watermark embedding module: image to be processed After the color space conversion from RGB color space to Lab color space, the converted image is obtained , the formula is as follows:
[0016] ,
[0017] in, Represents the operation of converting from RGB color space to Lab color space; the converted image Characteristics of the L channel The feature map of the L channel is obtained by extracting features through the encoder part of the U-net network ; To process the image Convert it to get a 128-bit binary bit stream as watermark information , watermark information After expansion and shape change, it is transformed into perturbation information matching the size of the L channel to obtain the watermark feature ; The watermark feature Feature map injected into L channel through diffusion model In the fusion feature , the formula is as follows:
[0018] ,
[0019] ,
[0020] in, represents the disturbance of the watermark feature, represents the adjustment factor, represents element-wise multiplication, Represents a multiplication operation; the fusion feature Concatenate the converted images The a channel and b channel of the image are obtained by ;
[0021] The complete Lab color space image Convert from Lab color space back to RGB color space to get an image with invisible watermark embedded , the formula is as follows:
[0022] ,
[0023] in, Represents the inverse operation of converting from RGB color space to Lab color space;
[0024] S22. Visible watermark embedding module: embedding invisible watermark into the image After the visible watermark injection module adds random noise, the image with embedded visible watermark is obtained. .
[0025] Furthermore, step S3 is specifically as follows:
[0026] The noise layer includes a general image processing module, a deep fake general modeling model and a deep fake model; an image with an invisible watermark embedded After the noise layer randomly selects one of the three parts: the general image processing module, the deep fake general modeling model and the deep fake model for processing to obtain the noise image ;
[0027] The general image processing module includes JPEG compression, cropping, Gaussian filtering, median filtering, brightness adjustment, contrast adjustment, saturation adjustment, Gaussian noise, and salt and pepper noise;
[0028] Images with invisible watermarks embedded in a general deepfake model Enter the noise layer and get the noise image through the deep fake general modeling model , the formula is as follows:
[0029] ,
[0030] in, Indicates that the convex hull of the facial core area and the convex hull of the facial non-core area are used as masks;
[0031] The deep fake models include the face-changing model SimSwap and the attribute editing model StarGAN.
[0032] Further, in step S4, the image with the visible watermark is embedded After passing through the general image processing module and deep fake model in the noise layer, and then through the joint optimization of the visible watermark, an obvious false warning mark is generated at the image location where the random noise is embedded.
[0033] Furthermore, step S5 is specifically as follows:
[0034] The watermark decoder includes a traceability decoder and a detection decoder;
[0035] Noisy Image The binary bit stream watermark obtained by tracing back the source decoder ; Noise image After the detection decoder, the binary bit stream watermark is obtained. ;calculate and When the traceability bit error rate is less than 2%, the source of the image is considered to be credible. and When the detection bit error rate is greater than 5%, the image is considered to be forged.
[0036] Furthermore, step S6 is specifically as follows:
[0037] In the process of invisible watermark embedding, cross-domain robust steganography constraints are adopted. The loss function formula of cross-domain robust steganography constraints is expressed as follows:
[0038] ,
[0039] in, Indicates the selection of the center of the 2D Fourier spectrum The area low pass filter, represents the Fourier transform function, Indicates an invisible watermark, represented by get, represents the cross-domain robustness steganographic constraint loss;
[0040] In order to make the image embedded with visible watermark produce warning signs, the following generative adversarial network regularization loss function is designed, and the formula is as follows:
[0041] ,
[0042] in, represents the identification symbol, SSIM represents the structural similarity function; the generative adversarial network regularization loss function is directly added to the training process of the face-changing model SimSwap and the attribute editing model StarGAN. The loss function after the loss regularization of the generative adversarial network model GAN is ,in, represents the original loss function of the generative adversarial network model, represents the regularization parameter;
[0043] The watermark encoder loss function formula is as follows:
[0044] ,
[0045] in, represents the watermark encoder loss, represents the Euclidean function, , Respectively represent calculation and as well as and The Euclidean distance between
[0046] In the watermark decoder, the source decoder loss function formula is expressed as follows:
[0047] ,
[0048] in, represents the Euclidean function, Indicates watermark information. Represents the binary bit stream watermark of source decoding; the detection decoder loss function formula is as follows:
[0049] ,
[0050] in, Represents a binary bit stream watermark for detection decoding;
[0051] In the deep fake model of the noise layer, the loss function of the detection decoder's processing operation on the deep fake model is designed, and the formula is expressed as follows:
[0052] ,
[0053] in, Represents a binary bit stream watermark for detection decoding;
[0054] The discriminator receives the image to be processed and images with visible watermarks , the loss function is used to distinguish the image to be processed and the watermarked image. The discriminator loss function is as follows:
[0055] ,
[0056] in, represents the discriminator loss, represents the discriminator; represents the expected logarithm of the probability that the discriminator identifies the original image as true, It represents the expectation of the logarithm of the probability that the discriminator judges the watermarked image to be false.
[0057] Furthermore, the process of obtaining the convex hull of the facial core area and the convex hull of the facial non-core area is as follows:
[0058] Use the existing face detection model to process the image Mark the position information of facial key points to obtain the corresponding facial key point information;
[0059] Based on the extracted facial key points, the convex hull of the core facial area is extracted, including the left eye, right eye, left eyebrow, right eyebrow, nose, and mouth; then the convex hull of the non-core facial area is extracted, including the left cheek, right cheek, and forehead.
[0060] The present invention also provides a system for applying an active defense method for image forgery based on multiple watermark fusion and cross-domain learning, comprising:
[0061] Image acquisition module: used to acquire the image to be processed;
[0062] Watermark embedding module: used to embed invisible watermark and visible watermark into the image to be processed through the watermark encoder, and obtain the image embedded with invisible watermark and the image embedded with visible watermark respectively;
[0063] Noise image generation module: used to process the image embedded with invisible watermark through the noise layer to obtain a noise image;
[0064] Warning sign generation module: used to process the image embedded with visible watermark through the noise layer, and generate an obvious false warning sign at the image position embedded with random noise through joint optimization of visible watermark;
[0065] Source tracing detection module: used to trace and detect the noise image through the watermark decoder to determine the authenticity of the image;
[0066] Training module: used for supervised training of loss function.
[0067] The advantages of the present invention are:
[0068] The present invention does not need to use any detectors that occupy computing resources to identify deep fake images. The naked eye can directly determine whether the image has been forged by identifying whether the warning mark exists. Even if this method is not applicable, we can still use invisible watermark tracing and detection methods to determine whether the image has been forged and verify the authenticity of the image source. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0070] Figure 1 is a flow chart of the steps of the method of the present invention;
[0071] Figure 2 The actual application effect of the visible watermark of the present invention;
[0072] Figure 3 This is the invisible watermark experimental effect of the present invention. DETAILED DESCRIPTION
[0073] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0074] Example 1
[0075] In this embodiment, Figure 1 As shown, the present invention provides an active defense method for image forgery based on multiple watermark fusion and cross-domain learning, and the specific steps include:
[0076] S1. Obtain the image to be processed;
[0077] Specifically, the image to be processed is obtained from the image dataset and the video dataset. .
[0078] S2. The image to be processed is embedded with an invisible watermark and a visible watermark through a watermark encoder to obtain an image embedded with an invisible watermark and an image embedded with a visible watermark respectively;
[0079] Specifically, S21. Invisible watermark embedding module: image to be processed After the color space conversion from RGB color space to Lab color space, the converted image is obtained , the formula is as follows:
[0080] ,
[0081] in, Represents the operation of converting from RGB color space to Lab color space; the converted image Characteristics of the L channel The feature map of the L channel is obtained by extracting features through the encoder part of the U-net network ; To process the image Convert it to get a 128-bit binary bit stream as watermark information , watermark information After expansion and shape change, it is transformed into perturbation information matching the size of the L channel to obtain the watermark feature ; The watermark feature Feature map injected into L channel through diffusion model In the fusion feature , the formula is as follows:
[0082] ,
[0083] ,
[0084] in, represents the disturbance of the watermark feature, represents the adjustment factor, represents element-wise multiplication, Represents a multiplication operation; the fusion feature Concatenate the converted images The a channel and b channel of the image are obtained by ;
[0085] The complete Lab color space image Convert from Lab color space back to RGB color space to get an image with invisible watermark embedded , the formula is as follows:
[0086] ,
[0087] in, Represents the inverse operation of converting from RGB color space to Lab color space;
[0088] S22. Visible watermark embedding module: embedding invisible watermark into the image After the visible watermark injection module adds random noise, the image with embedded visible watermark is obtained. .
[0089] S3. The image embedded with the invisible watermark is processed through the noise layer to obtain a noise image;
[0090] Specifically, the noise layer includes a general image processing module, a deep fake general modeling model and a deep fake model; an image with an invisible watermark embedded After the noise layer randomly selects one of the three parts: the general image processing module, the deep fake general modeling model and the deep fake model for processing to obtain the noise image ;
[0091] The general image processing module includes JPEG compression, cropping, Gaussian filtering, median filtering, brightness adjustment, contrast adjustment, saturation adjustment, Gaussian noise, and salt and pepper noise;
[0092] Images with invisible watermarks embedded in a general deepfake model Enter the noise layer and get the noise image through the deep fake general modeling model , the formula is as follows:
[0093] ,
[0094] in, Indicates that the convex hull of the facial core area and the convex hull of the facial non-core area are used as masks;
[0095] The deep fake models include the face-changing model SimSwap and the attribute editing model StarGAN.
[0096] Specifically, the process of obtaining the convex hull of the facial core area and the convex hull of the facial non-core area is as follows: the image to be processed is processed by the existing face detection model The facial key point position information is labeled to obtain the corresponding facial key point information; based on the extracted facial key points, the convex hull of the facial core area is extracted, including the left eye, right eye, left eyebrow, right eyebrow, nose, and mouth; then the convex hull of the facial non-core area is extracted, including the left cheek, right cheek, and forehead.
[0097] S4. Images with visible watermarks After passing through the general image processing module and deep fake model in the noise layer, and then through the joint optimization of the visible watermark, an obvious false warning mark is generated at the image position where the random noise is embedded, which is used to determine whether the image has been deeply faked.
[0098] S5. The noise image is traced and detected by the watermark decoder to determine the authenticity of the image; on the one hand, we check whether the image has been deeply forged by detecting the bit error rate. On the other hand, even if the image has been forged, we cannot determine the true source of the image. By decoding it with the traceability decoder, we can trace and authenticate the image.
[0099] Specifically, the watermark decoder includes a source tracing decoder and a detection decoder; the source tracing decoder performs source tracing to track the true source of the image, which can be used for copyright authentication, etc. The detection decoder is mainly used for deep fake detection.
[0100] Noisy Image The binary bit stream watermark obtained by tracing back the source decoder ; Noise image After the detection decoder, the binary bit stream watermark is obtained. ;calculate and When the traceability bit error rate is less than 2%, it can be judged that the source of the image is credible, that is, traceability authentication can be performed; calculation and The detection error rate is greater than 5%, and the image is considered to be forged. The closer the detection error rate is to 50%, the stronger the forgery is.
[0101] S6. Perform supervised training of the loss function.
[0102] Specifically, in the process of embedding invisible watermarks, cross-domain robust steganography constraints are adopted, and the loss function formula of cross-domain robust steganography constraints is expressed as follows:
[0103] ,
[0104] in, Indicates the selection of the center of the 2D Fourier spectrum The area low pass filter, represents the Fourier transform function, Indicates an invisible watermark, represented by get, represents the cross-domain robustness steganographic constraint loss;
[0105] In order to make the image embedded with visible watermark produce warning signs, the following generative adversarial network regularization loss function is designed, and the formula is as follows:
[0106] ,
[0107] in, represents the identification symbol, SSIM represents the structural similarity function; the generative adversarial network regularization loss function is directly added to the training process of the face-changing model SimSwap and the attribute editing model StarGAN. The loss function after the loss regularization of the generative adversarial network model GAN is ,in, represents the original loss function of the generative adversarial network model, represents the regularization parameter;
[0108] The watermark encoder loss function formula is as follows:
[0109] ,
[0110] in, represents the watermark encoder loss, represents the Euclidean function, , Respectively represent calculation and as well as and The Euclidean distance between
[0111] In the watermark decoder, the source decoder loss function formula is expressed as follows:
[0112] ,
[0113] in, represents the Euclidean function, Indicates watermark information. Represents the binary bit stream watermark of source decoding; the detection decoder loss function formula is as follows:
[0114] ,
[0115] in, Represents a binary bit stream watermark for detection decoding;
[0116] In the deep fake model of the noise layer, the loss function of the detection decoder's processing operation on the deep fake model is designed, and the formula is expressed as follows:
[0117] ,
[0118] in, Represents a binary bit stream watermark for detection decoding;
[0119] The discriminator receives the image to be processed and images with visible watermarks , the loss function is used to distinguish the image to be processed and the watermarked image. The discriminator loss function is as follows:
[0120] ,
[0121] in, represents the discriminator loss, represents the discriminator; represents the expected logarithm of the probability that the discriminator identifies the original image as true, It represents the expectation of the logarithm of the probability that the discriminator judges the watermarked image to be false.
[0122] Example 2
[0123] In this embodiment, taking face images as an example, we conduct an experimental comparative analysis between the existing method and the method of the present invention. The average results evaluated on the CelebA-HQ (256*256 pixels) dataset are shown in Table 1:
[0124] Table 1. The traceability detection effect and visual quality comparison of invisible watermarks in the face of deep fake models
[0125]
[0126] When the invisible watermark passes through the deep fake model, we hope that the traceability bit error rate is as low as possible, and the detection bit error rate is as close to 50% as possible (because this is in line with randomness, because our binary watermark bit stream is composed of numbers such as 0 or 1, so the closer it is to 50%, it means that this part of the watermark is destroyed, which means that the image is deeply faked.) The PSNR, SSIM and LPIPS in the last three columns are indicators for measuring visual quality. Among them, PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index) and LPIPS (Learned Perceptual Image Patch Similarity) are indicators for measuring visual quality. We hope that the visual quality will not change much after adding the invisible watermark. It can be found from Table 1 that the various indicators of the present invention are better than the current existing methods (this method is a paper "SepMark: Deep Separable Watermarking for Unified Source Tracing and Deepfake Detection" published in the CCF Class A conference ACM MM2023). As for deepfake detection, we determine that the image is a deepfake if the detection error rate is closer to 50%.
[0127] Figure 2 The actual application effect of the visible watermark of the method of the present invention. The left side is a real image, and the right side is a deep fake image. We embed a visible watermark in the upper left corner of the facial image to ensure that the watermark appears as a warning mark after deep fake to indicate the fake content. It can be seen from the figure that the method of the present invention can generate an effective warning mark for deep fake images;
[0128] Figure 3The first line is the original image, the second line is the image after adding the invisible watermark, the third line is the image processed by the noise layer, and the fourth line is the residual image of the image after embedding the invisible watermark minus the original image, that is, the visualization effect of the embedded invisible watermark. Figure 3 It can be seen that the method of the present invention can correctly encode the watermark into the facial features, hair and part of the background area of the facial image. The facial features are the areas that are extremely easy to destroy in the deep fake model, and the watermark embedded in this area can perform the detection function. The background area and the hair area are relatively difficult to forge, and the watermark embedded in these areas can perform the traceability function.
[0129] Example 3
[0130] This embodiment provides a system for active image forgery defense based on multiple watermark fusion and cross-domain learning, including:
[0131] Image acquisition module: used to acquire the image to be processed;
[0132] Watermark embedding module: used to embed invisible watermark and visible watermark into the image to be processed through the watermark encoder, and obtain the image embedded with invisible watermark and the image embedded with visible watermark respectively;
[0133] Noise image generation module: used to process the image embedded with invisible watermark through the noise layer to obtain a noise image;
[0134] Warning sign generation module: used to process the image embedded with visible watermark through the noise layer, and generate an obvious false warning sign at the image position embedded with random noise through joint optimization of visible watermark;
[0135] Source tracing detection module: used to trace and detect the noise image through the watermark decoder to determine the authenticity of the image;
[0136] Training module: used for supervised training of loss function.
[0137] In the system, for images whose deepforge is unknown, we compare the watermark decoded by the watermark detector with the original watermark m. When the matching error rate is close to 50%, we judge that the image is a deepforge; otherwise, we judge that the image is not a deepforge. At the same time, we can also compare the watermark decoded by the source decoder with the original watermark to determine that the original watermark with the lowest error rate is the copyright watermark (the watermark embedded in the image by the source user). At the same time, when the original watermark is unknown, our method can also determine whether the image has been forged by the degree of matching between the watermark decoded by the source decoder and the watermark decoded by the detection decoder, that is, when the watermarks decoded by the two decoders match to a certain extent, the image is not forged; otherwise, it is forged; by calculating and The Euclidean distance is used to calculate the match. When the match is greater than 10%, it can be judged that the image has been forged. Otherwise, it has not been deeply forged. In general, the method used by the system realizes three functions. One is that the visible watermark can generate a warning mark for the deeply forged image so that the human eye can judge whether the image has been deeply forged. At the same time, even if the visible watermark is destroyed, our invisible watermark can also realize traceability to perform copyright authentication and detect whether it has been deeply forged.
[0138] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An active defense method for image forgery based on multiple watermark fusion and cross-domain learning, characterized in that: The following steps are involved: S1. Get the image to be processed ; S2. The image to be processed is embedded with an invisible watermark and a visible watermark through a watermark encoder to obtain an image embedded with an invisible watermark and an image embedded with a visible watermark respectively; The watermark encoder includes an invisible watermark embedding module and a visible watermark injection module, specifically: S21. Invisible watermark embedding module: image to be processed After the color space conversion from RGB color space to Lab color space, the converted image is obtained , the formula is as follows: , in, Represents the operation of converting from RGB color space to Lab color space; the converted image Characteristics of the L channel The feature map of the L channel is obtained by extracting features through the encoder part of the U-net network ; To process the image Convert it to get a 128-bit binary bit stream as watermark information , watermark information After expansion and shape change, it is transformed into perturbation information matching the size of the L channel to obtain the watermark feature ; The watermark feature Feature map injected into L channel through diffusion model In the fusion feature , the formula is as follows: , , in, represents the disturbance of the watermark feature, represents the adjustment factor, represents element-wise multiplication, Represents a multiplication operation; the fusion feature Concatenate the converted images The a channel and b channel of the image are obtained by ; The complete Lab color space image Convert from Lab color space back to RGB color space to get an image with invisible watermark embedded , the formula is as follows: , in, Represents the inverse operation of converting from RGB color space to Lab color space; S22. Visible watermark embedding module: embedding invisible watermark into the image After the visible watermark injection module adds random noise, the image with embedded visible watermark is obtained. ; S3. The image embedded with the invisible watermark is processed through the noise layer to obtain a noise image; The noise layer includes a general image processing module, a deep fake general modeling model and a deep fake model; an image with an invisible watermark embedded After the noise layer randomly selects one of the three parts: the general image processing module, the deep fake general modeling model and the deep fake model for processing to obtain the noise image ; The general image processing module includes JPEG compression, cropping, Gaussian filtering, median filtering, brightness adjustment, contrast adjustment, saturation adjustment, Gaussian noise, and salt and pepper noise; Images with invisible watermarks embedded in a general deepfake model Enter the noise layer and get the noise image through the deep fake general modeling model , the formula is as follows: , in, Indicates that the convex hull of the facial core area and the convex hull of the facial non-core area are used as masks; The deep fake models include the face-changing model SimSwap and the attribute editing model StarGAN; S4. The image embedded with the visible watermark is processed through the noise layer, and an obvious false warning mark is generated at the image position where the random noise is embedded through joint optimization of the visible watermark, which is used to determine whether the image has been deeply forged; Specifically, an image with a visible watermark is embedded After passing through the general image processing module and deep fake model in the noise layer, the visible watermark is jointly optimized to produce an obvious false warning mark at the image location where the random noise is embedded; S5. The noise image is traced and detected by the watermark decoder to determine the authenticity of the image; S6. Perform supervised training of the loss function.
2. The active defense method for image forgery based on multiple watermark fusion and cross-domain learning according to claim 1 is characterized in that: Step S1, obtaining the image to be processed from the image dataset and the video dataset .
3. The active defense method for image forgery based on multiple watermark fusion and cross-domain learning according to claim 2 is characterized in that: Step S5 is specifically as follows: The watermark decoder includes a traceability decoder and a detection decoder; Noisy Image The binary bit stream watermark obtained by tracing back the source decoder ; Noise image After the detection decoder, the binary bit stream watermark is obtained. ;calculate and When the traceability bit error rate is less than 2%, the source of the image is considered to be credible. calculate and When the detection bit error rate is greater than 5%, the image is considered to be forged.
4. The active defense method for image forgery based on multiple watermark fusion and cross-domain learning according to claim 3 is characterized in that: Step S6 is specifically as follows: In the process of invisible watermark embedding, cross-domain robust steganography constraints are adopted. The loss function formula of cross-domain robust steganography constraints is expressed as follows: , in, Indicates the selection of the center of the 2D Fourier spectrum The area low pass filter, represents the Fourier transform function, Indicates an invisible watermark, represented by get, represents the cross-domain robustness steganographic constraint loss; In order to make the image embedded with visible watermark produce warning signs, the following generative adversarial network regularization loss function is designed, and the formula is as follows: , in, represents the identification symbol, SSIM represents the structural similarity function; the generative adversarial network regularization loss function is directly added to the training process of the face-changing model SimSwap and the attribute editing model StarGAN. The loss function after the loss regularization of the generative adversarial network model GAN is ,in, represents the original loss function of the generative adversarial network model, represents the regularization parameter; The watermark encoder loss function formula is as follows: , in, represents the watermark encoder loss, represents the Euclidean function, , Respectively represent calculation and as well as and The Euclidean distance between In the watermark decoder, the source decoder loss function formula is expressed as follows: , in, represents the Euclidean function, Indicates watermark information. Represents the binary bit stream watermark of source decoding; the detection decoder loss function formula is as follows: , in, Represents a binary bit stream watermark for detection decoding; In the deep fake model of the noise layer, the loss function of the detection decoder's processing operation on the deep fake model is designed, and the formula is expressed as follows: , in, Represents a binary bit stream watermark for detection decoding; The discriminator receives the image to be processed and images with visible watermarks , the loss function is used to distinguish the image to be processed and the watermarked image. The discriminator loss function is as follows: , in, represents the discriminator loss, represents the discriminator; represents the expected logarithm of the probability that the discriminator identifies the original image as true, It represents the expectation of the logarithm of the probability that the discriminator judges the watermarked image to be false.
5. The active defense method for image forgery based on multiple watermark fusion and cross-domain learning according to claim 4 is characterized in that: The process of obtaining the convex hull of the facial core area and the convex hull of the facial non-core area is as follows: Use the existing face detection model to process the image Mark the position information of facial key points to obtain the corresponding facial key point information; Based on the extracted facial key points, extract the convex hull of the facial core area, including the left eye, right eye, left eyebrow, right eyebrow, nose, and mouth; Then extract the convex hull of the non-core area of the face, including the left cheek, right cheek and forehead.
6. A system using the active image forgery defense method based on multiple watermark fusion and cross-domain learning as claimed in claim 1, characterized in that: include: Image acquisition module: used to acquire the image to be processed; Watermark embedding module: used to embed invisible watermark and visible watermark into the image to be processed through the watermark encoder, and obtain the image embedded with invisible watermark and the image embedded with visible watermark respectively; Noise image generation module: used to process the image embedded with invisible watermark through the noise layer to obtain a noise image; Warning sign generation module: used to process the image embedded with visible watermark through the noise layer, and generate an obvious false warning sign at the image position embedded with random noise through joint optimization of visible watermark; Source tracing detection module: used to trace and detect the noise image through the watermark decoder to determine the authenticity of the image; Training module: used for supervised training of loss function.
Citation Information
Patent Citations
Face deep counterfeiting evidence obtaining method, system and device and storage medium
CN118279995A
Semi-fragile diffusion watermarking method and system based on elliptic curve
CN119151766A