An active defense image generation method, system and medium for face forgery

Through adaptive anti-noise generator and low-frequency enhancement attention module, the momentum coefficient and local and global variance are dynamically adjusted, and combined with low-frequency enhancement and consistency loss functions, a highly adaptable face forged active defense images is generated, which solves the problem of poor defense effects in the existing technology and realizes efficient hidden attacks and visual nature.

CN120198531BActive Publication Date: 2025-07-22NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510670148.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-22
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In the prior art, the active defense method of face forgery cannot adapt to different input samples or attack stages, resulting in poor defense effects and inability to conduct effective attacks while maintaining the visualization effect of the face.

Method used

Adaptive anti-noise generator and low-frequency enhanced attention module are used to dynamically adjust the momentum coefficient and local and global variance to generate frequency adversarial components, and the low-frequency enhanced loss and consistency loss function training model, combined with one-time anti-noise attack, generate face forged active defense images.

Benefits of technology

It realizes the addition of effective covert attacks on face images under different postures, adapting to different input samples or attack stages, maintaining the naturalness of the image visually and improving the defense effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198531B_ABST
    Figure CN120198531B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and medium for actively defending against face forgery image generation, belonging to the technical field of image processing. The method includes: obtaining a face image to be protected; projecting the face image to be protected onto a parameterized two-dimensional coordinate system to obtain a parameterized facial texture image; extracting the frequency components of the parameterized facial texture image, inputting the frequency components into a pre-constructed active defense image generation model, and outputting frequency adversarial components; performing feature fusion on the frequency adversarial components to obtain adversarial features; and projecting the adversarial features onto a normalized device coordinate system to obtain an actively defended face forgery image. The present invention can ensure the visual effect of the face in the spatial domain while adding effective stealth attacks to face images in different poses, adapting to different input samples or attack stages, and has a good defense effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method, system and medium for generating face forgery active defense images. Background Art

[0002] Currently, face forgery defense methods are mainly divided into active defense and passive defense. Passive defense is mainly used for forgery detection and authenticity verification in post-event defense, but it cannot prevent the spread of forged face images from the source. Active defense is mainly based on adversarial machine learning, which applies tiny adversarial perturbations to the original image. Without destroying the visual effect of the face image, it can effectively attack the deep forgery model and achieve the purpose of defense from the data source.

[0003] In traditional face forgery active defense image generation methods, the momentum coefficient of the adversarial sample attack algorithm based on momentum is usually a fixed hyperparameter, which needs to be determined by manual tuning through experiments. The static momentum coefficient may not be able to adapt to different input samples or attack stages, and cannot achieve good active defense effects. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method, system and medium for generating face forgery active defense images, which can ensure the visual effect of the face in the spatial domain while adding effective stealth attacks to face images in different poses, adapting to different input samples or attack stages, and having good defense effects.

[0005] To achieve the above purpose, the present invention is implemented by the following technical solutions:

[0006] On the one hand, the present invention provides a method for generating face forgery active defense images, including:

[0007] Obtain a face image to be protected;

[0008] Project the face image to be protected onto a parameterized two-dimensional coordinate system to obtain a parameterized facial texture image;

[0009] Extract the frequency components of the parameterized facial texture image, input the frequency components into a pre-constructed active defense image generation model, and output frequency adversarial components;

[0010] Fuse the frequency adversarial components to obtain adversarial features;

[0011] Project the adversarial features onto the normalized device coordinate system to obtain a face forgery active defense image;

[0012] Among them, the active defense image generation model includes an adaptive adversarial noise generator and a low-frequency enhanced attention module connected in sequence; the low-frequency enhanced attention module is used to highlight the importance of image regions; in the adaptive adversarial noise generator, the static momentum coefficient is replaced by a dynamic adaptive momentum coefficient, and the dynamic adaptive momentum coefficient is dynamically adjusted through local variance and global variance, and the global variance and local variance are dynamically adjusted through image gradients.

[0013] Optionally, projecting the face image to be protected onto a parameterized two-dimensional coordinate system to obtain a parameterized facial texture image includes:

[0014] Inputting the face image to be protected into a 3D face reconstruction network to generate a texture map, and projecting the texture map onto the parameterized two-dimensional coordinate system according to 3D position parameters to obtain a parameterized facial texture image.

[0015] Optionally, using discrete wavelet transform to extract the frequency components of the parameterized facial texture image includes:

[0016] ;

[0017] Among them, , , , respectively represent the low-frequency component, the first high-frequency component, the second high-frequency component, and the third high-frequency component; x and y represent the spatial coordinates corresponding to the pixels of the parameterized facial texture image; represents the parameterized facial texture image of size M×N; represents the two-dimensional scaling function for the first-level decomposition of the low-pass filter used to capture ; , , respectively represent the wavelet functions for the low-pass row and high-pass column combination used to capture , the wavelet functions for the high-pass row and low-pass column combination used to capture , and the wavelet functions for the high-pass row and low-pass column combination used to capture ; m and n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components.

[0018] Optionally, the low-frequency enhanced attention module includes a channel attention unit and a spatial attention unit connected in sequence;

[0019] In the channel attention unit, channel weighting is performed on the adversarial noise components to obtain channel adversarial components; the adversarial noise components are the outputs of the adaptive adversarial noise generator.

[0020] In the spatial attention unit, the channel adversarial component is spatially weighted to obtain the frequency adversarial component.

[0021] Optionally, the active defense image generation model is expressed as:

[0022] ;

[0023] ;

[0024] where and represent the adversarial noise components at the (t + 1)-th iteration and the t-th iteration respectively, and the adversarial noise component is the output of the adaptive adversarial noise generator; represents the dynamic adaptive momentum coefficient; represents the frequency adversarial component generated in the current iteration; represents the important region feature map; represents the gradient of the frequency component at the t-th iteration; represents the element-wise multiplication; represents the L1 norm; represents the basic momentum coefficient; and represent the local variance adaptive weight and the global variance adaptive weight respectively; and represent the local variance and the global variance at the pixel position respectively; and represent the local variance and the global variance respectively; represents the frequency component at the t-th iteration; represents the maximum perturbation amplitude in the current iteration; represents the sign function; represents the truncation function; represents the size of the sliding window; represents the gradient of the frequency component at the pixel position ; represents the total number of elements of the gradient tensor; represents the gradient of the frequency component at the pixel position (i, j, l); represents the average value of the gradient of the frequency component; m and n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components; C, H, and W represent the number of channels, the height, and the width of the input image.

[0025] Optionally, the training steps of the active defense image generation model include:

[0026] Train the active defense image generation model using the total loss function composed of the low-frequency enhancement loss function and the consistency loss function to obtain a trained active defense image generation model.

[0027] Optionally, the total loss function is expressed as:

[0028] ;

[0029] ;

[0030] ;

[0031] Wherein, , respectively represent the frequency adversarial component and the frequency component generated in the current iteration; represents the total loss function; represents the consistency loss function; , respectively represent the low-frequency component loss function and the high-frequency component loss function; , respectively represent the loss function balance hyperparameters; G represents the face forgery operation; B represents the number of samples; represents the frequency adversarial component generated in the current iteration of the b-th input sample; represents the frequency component generated in the current iteration of the b-th input sample; , respectively represent the discrete wavelet transform low-frequency subband components of the frequency adversarial component and the frequency component; , respectively represent the discrete wavelet transform high-frequency subband components of the frequency adversarial component and the frequency component; s represents the high-frequency subband type; , respectively represent the L1 and L2 norms; respectively represent the types of the first high-frequency subband component, the second high-frequency subband component, and the third high-frequency subband component.

[0032] Optionally, project the adversarial feature onto the normalized device coordinate system to obtain a face active defense image, including:

[0033] Input the adversarial feature into a pre-constructed noise projection alignment unit to output a projected image;

[0034] Apply a one-time adversarial noise attack to the area outside the projected image to obtain a face active defense image;

[0035] Wherein, the processing steps of the noise projection alignment unit include:

[0036] Project the adversarial feature onto the normalized device coordinate system to obtain the adversarial feature in the device coordinate system;

[0037] Perform point-plane merging and differentiable rendering on the adversarial feature in the device coordinate system in sequence to generate a depth map and a mask;

[0038] Multiply the adversarial feature by the depth map and the mask to obtain a projected image.

[0039] In a second aspect, the present invention provides a face forgery active defense system, including:

[0040] An image acquisition module, configured to: acquire a face image to be protected;

[0041] A first projection module, configured to: project the face image to be protected onto a two-dimensional coordinate system to obtain a parameterized facial texture image;

[0042] An adversarial application module, configured to: extract the frequency components of the parameterized facial texture image, input the frequency components into a pre-constructed active defense image generation model, and output frequency adversarial components;

[0043] A feature fusion module, configured to: perform feature fusion on the frequency adversarial components to obtain an adversarial feature;

[0044] A second projection module, configured to: project the adversarial feature onto the 3D domain to obtain a face forgery active defense image;

[0045] Wherein, the active defense image generation model includes an adaptive adversarial noise generator and a low-frequency enhancement attention module connected in sequence; the low-frequency enhancement attention module is used to highlight the importance of image regions; in the adaptive adversarial noise generator, the static momentum coefficient is replaced with a dynamic adaptive momentum coefficient, and the dynamic adaptive momentum coefficient is dynamically adjusted through local variance and global variance, and the global variance and local variance are dynamically adjusted through image gradients.

[0046] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the face forgery active defense image generation method as described in the first aspect.

[0047] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: 1. The present invention extracts the frequency components of the parametric facial texture image, adjusts the gradient update direction of the frequency adversarial component by using a dynamic adaptive momentum coefficient, and highlights the regional importance by using a low-frequency enhancement attention module, which can add effective stealth attacks to face images in different poses, resist the generation of high-quality tampered images by face forgery systems, adapt to different input samples or attack stages, and has a good active defense effect; 2. The present invention uses the low-frequency enhancement loss to limit the local perturbation range of the high-frequency components in the model, and uses the consistency loss to limit the global perturbation range of all components in the model, reducing the high-frequency noise energy while retaining the attack effect, avoiding local mutations of adversarial perturbations in the high-frequency region, and maintaining the visual naturalness of the image, thereby enhancing the stealthiness of the adversarial noise; 3. The present invention applies a one-time adversarial noise attack to the area outside the projected image to obtain an active face defense image, which can ensure the visual effect of the face in the airspace while taking into account the effectiveness of the adversarial attack. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 FIG. shows a schematic flow chart of the method for generating an active face forgery defense image according to the present invention in one embodiment;

[0049] Figure 2 FIG. shows a schematic flow chart of the method for generating an active face forgery defense image according to the present invention in another embodiment;

[0050] Figure 3 FIG. shows a schematic flow chart of the active defense image generation model according to the present invention in one embodiment;

[0051] Figure 4 FIG. shows a schematic flow chart of the low-frequency enhancement attention module according to the present invention in one embodiment;

[0052] Figure 5 FIG. shows a schematic flow chart of the noise projection alignment according to the present invention in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0054] The term "and / or" merely describes an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.

[0055] Example 1

[0056] As Figure 1 shown, this example introduces a method for generating actively defensive images against face forgery, including the following steps:

[0057] Step 1: Obtain a face image to be protected;

[0058] Step 2: As Figure 2 shown, input the face image to be protected into a three-dimensional (3D) face reconstruction network to obtain a parameterized facial texture image, that is, migrate and render from the two-dimensional (2D) domain to the 3D domain. Input the face image to be protected into the 3D face reconstruction network to generate a texture map, and by mapping the texture map according to the 3D position parameters in the 3D face reconstruction network to the parameterized two-dimensional (UV) coordinate system, obtain a parameterized facial texture image.

[0059] Step 3: As Figure 2 shown, use the discrete wavelet transform to extract the frequency components of the parameterized facial texture image. Among them, by performing DWT (Discrete Wavelet Transform) frequency feature extraction on the parameterized facial texture image, obtain frequency components, and the frequency components include a low-frequency component, a first high-frequency component, a second high-frequency component, and a third high-frequency component, expressed as:

[0060] ;

[0061] Among them, , , , respectively represent the low-frequency component, the first high-frequency component, the second high-frequency component, and the third high-frequency component; x, y represent the spatial coordinates corresponding to the pixels of the parameterized facial texture image; represents the parameterized facial texture image of size M×N; represents the two-dimensional scaling function for the first-level decomposition of the low-pass filter used to capture , ; , , respectively represent the wavelet functions of the low-pass row and high-pass column combination used to capture (vertical edge transformation), the wavelet functions of the high-pass row and low-pass column combination used to capture (horizontal edge transformation), and the wavelet functions of the high-pass row and low-pass column combination used to capture (diagonal feature transformation); m, n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components.

[0062] The two-dimensional scaling function uses the orthogonal Haar wavelet function and is obtained from the one-dimensional scaling function and the wavelet basis function through the tensor product, that is: ;

[0063] The one-dimensional scaling function and the wavelet basis function are respectively:

[0064] ;

[0065] ;

[0066] Step 4: Input the low-frequency component, the first high-frequency component, the second high-frequency component, and the third high-frequency component into the pre-constructed active defense image generation model to output the frequency adversarial component, specifically:

[0067] As Figure 2 shown, the active defense image generation model includes an adaptive adversarial noise generator and a low-frequency enhanced attention module connected in sequence.

[0068] As Figure 3 shown, the active defense image generation model first calculates the local variance and gradient normalization for the input frequency components through the adaptive adversarial noise generator. Among them, the local variance and global variance are calculated based on the gradient of the frequency component, and the dynamic adaptive momentum coefficient is calculated according to the local variance and global variance to output the adversarial noise component; then the adversarial noise component is sequentially subjected to feature splicing, channel weighting, and spatial weighting through the low-frequency enhanced attention module to obtain the final weight map frequency adversarial component.

[0069] In the adaptive adversarial noise generator, the static momentum coefficient in the momentum-based adversarial sample attack algorithm is replaced with a dynamic adaptive momentum coefficient;

[0070] As Figure 4 shown, the low-frequency enhanced attention module includes a channel attention unit and a spatial attention unit connected in sequence;

[0071] The channel attention unit sequentially includes a global average pooling layer, two convolutional layers connected by the ReLU activation function, and a Sigmoid activation function for outputting the channel attention weight between 0 and 1. In the channel attention unit, the adversarial noise component is weighted by channels to obtain the channel adversarial component;

[0072] The spatial attention unit sequentially includes global max pooling and global average pooling, and connects a 1×1 convolutional layer to a Sigmoid activation function. In the spatial attention unit, the channel adversarial component is spatially weighted to output the spatial attention weight, obtaining the spatial attention feature map, that is, the frequency adversarial component.

[0073] That is, the processing steps of the active defense image generation model include:

[0074] Based on the momentum-based adversarial sample attack, an adaptive momentum adjustment mechanism is introduced, and an adaptive momentum adversarial sample attack is performed on the low-frequency component and the high-frequency component. According to the local variance and the global variance to measure the effectiveness and stability of the adversarial attack, and dynamically adjust the dynamic adaptive momentum coefficient of the adversarial sample , the global variance and the local variance are dynamically adjusted through the image gradient. When the gradient changes greatly, the momentum is increased to accelerate the model convergence. When the gradient changes little, the momentum is decreased to avoid the gradient update direction falling into the local optimum. The active defense image generation model is expressed as:

[0075] ;

[0076] Among them,

[0077] ;

[0078] In the formula, , respectively represent the momentum accumulation amounts of the (t + 1)-th iteration and the t-th iteration, that is, the adversarial noise components, and the adversarial noise components are the outputs of the adaptive adversarial noise generator; represents the dynamic adaptive momentum coefficient; C, H, and W represent the number of channels, height, and width of the input image; represents the frequency adversarial component generated in the current iteration; represents the important region feature map, which is obtained through the low-frequency enhancement attention module; represents the gradient of the frequency component at the t-th iteration; represents the element-wise multiplication; represents the L1 norm; represents the basic momentum coefficient; , respectively represent the local variance adaptive weight and the global variance adaptive weight; , respectively represent the local variance and the global variance at the pixel position ; represents the frequency component at the t-th iteration; represents the maximum perturbation amplitude in the current iteration; denotes the sign function; denotes the truncation function, which is used to limit the range of noise perturbation; denotes the size of the sliding window; denotes at the pixel position the gradient of the frequency component; denotes the total number of elements of the gradient tensor; denotes the gradient of the frequency component at the pixel position (i, j, l); denotes the average value of the gradient of the frequency component; m and n denote the discrete indices of the pixel positions corresponding to the frequency components of the wavelet transform.

[0079] To ensure the spatial concealment of the low-frequency perturbation while enhancing its attack effectiveness, according to the above local variance and the amplitude of the gradient tensor, the adaptive adversarial noise generator generates the adversarial noise component, and the importance of the adversarial noise component is weighted by the low-frequency enhancement attention module to highlight the feature map of the important region, and it is dot-multiplied with the normalized gradient to more precisely update the adaptive momentum.

[0080] The training steps of the active defense image generation model include:

[0081] As Figure 2 shown, the active defense image generation model is trained using the total loss function composed of the low-frequency enhancement loss function and the consistency loss function to obtain the trained active defense image generation model;

[0082] The low-frequency enhancement loss function is used to limit the perturbation range of the high-frequency component noise of the model. The expression of the low-frequency enhancement loss function is as follows:

[0083] ;

[0084] The consistency loss function is used to limit the global perturbation range of the adversarial noise of the model. The expression of the consistency loss function is as follows:

[0085] ;

[0086] The total loss function is expressed as:

[0087] ;

[0088] In the formula denotes the frequency component generated in the current iteration; denotes the total loss function; denotes the consistency loss function; , respectively denote the low-frequency component loss function and the high-frequency component loss function; The offset distance of the feature components before and after the attack is calculated through the discrete wavelet transform (DWT), and the low-frequency semantic-level perturbation is strengthened during the iterative attack process; The L2 regularization is added to the loss function as a penalty term to reduce the high-frequency noise energy while retaining the attack effect, avoid local mutations of the adversarial perturbation in the high-frequency region, maintain the visual naturalness of the image, and thus enhance the concealment of the adversarial noise; 、 respectively represent the balance hyperparameter weights of the loss functions for the low-frequency components and the balance hyperparameter weights of the loss functions for the high-frequency components; G represents the face forgery operation; B represents the number of samples; represents the frequency adversarial component generated in the current iteration of the b-th input sample; represents the frequency component generated in the current iteration of the b-th input sample; 、 respectively represent the discrete wavelet transform low-frequency subband components of the frequency adversarial component and the frequency component; 、 respectively represent the discrete wavelet transform high-frequency subband components of the frequency adversarial component and the frequency component; s represents the high-frequency subband type; 、 respectively represent the L1 and L2 norms; respectively represent the types of the first high-frequency subband component, the second high-frequency subband component, and the third high-frequency subband component.

[0089] The model hyperparameter weights are continuously updated using gradient descent 、 until the convergence condition is reached.

[0090] By using the low-frequency enhancement loss to limit the local perturbation range of the high-frequency components in the model and the consistency loss to limit the global perturbation range of all components in the model, the high-frequency noise energy is reduced while retaining the attack effect, local mutations of the adversarial perturbation in the high-frequency region are avoided, and the visual naturalness of the image is maintained, thereby enhancing the concealment of the adversarial noise.

[0091] Step Five: Feature fusion is performed on the low-frequency adversarial component, the first high-frequency adversarial component, the second high-frequency adversarial component, and the third high-frequency adversarial component to obtain the adversarial feature, specifically:

[0092] As Figure 2 shown, first, high-frequency feature fusion is performed on the first high-frequency adversarial component, the second high-frequency adversarial component, and the third high-frequency adversarial component to obtain the high-frequency fusion feature, and then feature fusion is performed on the low-frequency adversarial component and the high-frequency fusion feature to obtain the adversarial feature.

[0093] Step Six: The adversarial feature is projected onto the normalized device coordinate system to obtain the face forgery active defense image, specifically:

[0094] Construct a noise projection alignment unit. The noise projection alignment unit includes homogeneous coordinate transformation.

[0095] As Figure 5 shown, input the adversarial feature into the noise projection alignment unit, project it to the normalized device coordinate system, and output the projected image; the processing flow of the noise projection alignment unit includes:

[0096] Homogeneous coordinate transformation: Project the adversarial feature to the normalized device coordinate system to obtain the adversarial feature in the device coordinate system;

[0097] Process the mesh input: Merge the vertices and faces of the adversarial feature in the device coordinate system to obtain the merged feature;

[0098] Differentiable rasterization rendering: Perform differentiable rasterization rendering on the merged feature to obtain the rendered feature;

[0099] Generate the mask and depth map of the rendered feature;

[0100] Generate the projected image: Projected image = Texture map × Depth map × Mask.

[0101] Apply a one-time adversarial noise attack to the area outside the projected image to obtain the face active defense image, which can ensure the visual effect of the face spatial domain while taking into account the effectiveness of the adversarial attack. In this embodiment, the one-time adversarial noise attack is an adversarial perturbation that only applies one gradient iteration to the low-frequency features.

[0102] The traditional projection alignment unit is a part of the 3D face reconstruction network, which is used to project the UV texture map reconstructed from the 3D domain onto the 2D plane for texture alignment. Projected image = Texture map × Depth map × Illumination map. The projected image in this embodiment = Adversarial noise (noise containing the semantics of human face features) × Depth map × Facial area mask.

[0103] In a specific embodiment, select the pre-trained generative adversarial network (Star Generative Adversarial Network, StarGAN) forgery model as an example as the defense object, train the active defense image generation model, input the original face image into the trained active defense image generation model to obtain the noisy active defense image; input the active defense image into the above StarGAN forgery model to obtain the face active defense visualization effect diagram under different forgery attributes.

[0104] In this embodiment, a parameterized facial texture image is obtained by rendering based on a 3D face reconstruction network. The facial features are projected into the UV coordinate system to obtain the parameterized facial texture image. The frequency-domain features of the parameterized facial texture image are extracted through discrete wavelet transform to obtain the frequency-domain components of the UV facial texture image. An active defense image generation model is constructed. The adaptive adversarial noise generator adjusts the gradient update direction of the adversarial samples based on the dynamic adaptive momentum coefficient of the gradient change, performs adversarial sample attacks on the frequency-domain components, and uses the channel attention and spatial attention mechanisms to highlight the importance of regions. The low-frequency enhancement loss is used to limit the local perturbation range of the high-frequency components in the model. The consistency loss is used to limit the global perturbation range of all components in the model. The hyperparameter weights of the model are continuously updated using gradient descent until the convergence condition is reached. The adversarial noise in the UV facial texture image is projected into the 2D domain and cascaded with the original face image to obtain the final defense sample, which can add effective stealth attacks to face images in different poses, resist the generation of high-quality tampered images by face forgery systems, adapt to different input samples or attack stages, and has a good active defense effect.

[0105] Embodiment 2

[0106] Based on Embodiment 1, this embodiment introduces an active defense system for face forgery, including:

[0107] An image acquisition module, configured to: acquire a face image to be protected;

[0108] A first projection module, configured to: project the face image to be protected into a two-dimensional coordinate system to obtain a parameterized facial texture image;

[0109] An adversarial application module, configured to: extract the frequency components of the parameterized facial texture image, input the frequency components into a pre-constructed active defense image generation model, and output frequency adversarial components;

[0110] A feature fusion module, configured to: perform feature fusion on the frequency adversarial components to obtain adversarial features;

[0111] A second projection module, configured to: project the adversarial features into the normalized device coordinate system to obtain an active defense image for face forgery;

[0112] Wherein, the construction of the active defense image generation model includes an adaptive adversarial noise generator and a low-frequency enhancement attention module connected in sequence; the low-frequency enhancement attention module is used to highlight the importance of image regions; in the adaptive adversarial noise generator, the static momentum coefficient is replaced with a dynamic adaptive momentum coefficient, and the dynamic adaptive momentum coefficient is dynamically adjusted through local variance and global variance, and the global variance and local variance are dynamically adjusted through image gradients.

[0113] For the specific function implementation of each of the above modules, refer to the relevant content in the method of Embodiment 1, which will not be elaborated here.

[0114] Embodiment 3

[0115] This embodiment introduces a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the face forgery active defense image generation method as described in Embodiment 1.

[0116] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0118] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0120] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. An active defense image generation method for face forgery, characterized in that, Including: Obtain a face image to be protected; Project the face image to be protected onto a parametric two-dimensional coordinate system to obtain a parametric facial texture image; Extract the frequency components of the parametric facial texture image, input the frequency components into a pre-constructed active defense image generation model, and output frequency adversarial components; Perform feature fusion on the frequency adversarial components to obtain adversarial features; Project the adversarial features onto a normalized device coordinate system to obtain an active face forgery defense image; Wherein, the active defense image generation model includes an adaptive adversarial noise generator and a low-frequency enhanced attention module connected in sequence; the low-frequency enhanced attention module is used to highlight the importance of image regions; in the adaptive adversarial noise generator, the static momentum coefficient is replaced with a dynamic adaptive momentum coefficient, and the dynamic adaptive momentum coefficient is dynamically adjusted through local variance and global variance, and the global variance and local variance are dynamically adjusted through image gradients; Using discrete wavelet transform to extract the frequency components of the parametric facial texture image, including: ; Among them, , , , respectively represent the low-frequency component, the first high-frequency component, the second high-frequency component, and the third high-frequency component; x and y represent the spatial coordinates corresponding to the pixels of the parameterized facial texture image; represents the parameterized facial texture image of size M×N; represents the two-dimensional scaling function for the first-level decomposition of the low-pass filter used to capture ; , , respectively represent the wavelet function for the low-pass row and high-pass column combination used to capture , the wavelet function for the high-pass row and low-pass column combination used to capture , and the wavelet function for the high-pass row and low-pass column combination used to capture ; m and n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components; The active defense image generation model is expressed as: ; ; Among them, and respectively represent the adversarial noise components of the (t + 1)-th iteration and the t-th iteration, and the adversarial noise component is the output of the adaptive adversarial noise generator; represents the dynamic adaptive momentum coefficient; represents the frequency adversarial component generated in the current iteration; represents the important region feature map; represents the gradient of the frequency component at the t-th iteration; represents the element-wise multiplication; represents the L1 norm; represents the basic momentum coefficient; and respectively represent the local variance adaptive weight and the global variance adaptive weight; and respectively represent the local variance and the global variance at the pixel position ; and respectively represent the local variance and the global variance; represents the frequency component at the t-th iteration; represents the maximum perturbation amplitude in the current iteration; represents the sign function; represents the truncation function; represents the size of the sliding window; represents the gradient of the frequency component at the pixel position ; represents the total number of elements in the gradient tensor; represents the gradient of the frequency component at the pixel position (i, j, l); represents the average value of the gradient of the frequency component; m and n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components; C, H, and W represent the number of channels, height, and width of the input image.

2. The method for generating an actively defensive image against face forgery according to claim 1, wherein Projecting the face image to be protected onto a parametric two-dimensional coordinate system to obtain a parametric facial texture image, including: Input the face image to be protected into a 3D face reconstruction network to generate a texture map, and project the texture map onto a parametric two-dimensional coordinate system according to 3D position parameters to obtain a parametric facial texture image.

3. The face forgery active defense image generation method according to claim 1, characterized in that The low-frequency enhanced attention module includes a channel attention unit and a spatial attention unit connected in sequence; In the channel attention unit, channel weighting is performed on the adversarial noise components to obtain channel adversarial components; the adversarial noise components are the output of the adaptive adversarial noise generator; In the spatial attention unit, spatial weighting is performed on the channel adversarial components to obtain frequency adversarial components.

4. The method for generating an actively defensive image against face forgery according to claim 1, wherein, The training steps of the active defense image generation model include: Train the active defense image generation model using a total loss function composed of a low-frequency enhanced loss function and a consistency loss function to obtain a trained active defense image generation model.

5. The method for generating an actively defensive image against face forgery according to claim 4, wherein The total loss function is expressed as: ; ; ; Among them, and respectively represent the frequency adversarial component and the frequency component generated in the current iteration; represents the total loss function; represents the consistency loss function; and respectively represent the low-frequency component loss function and the high-frequency component loss function; and respectively represent the loss function balance hyperparameters; G represents the face forgery operation; B represents the number of samples; represents the frequency adversarial component generated in the current iteration of the b-th input sample; represents the frequency component generated in the current iteration of the b-th input sample; and respectively represent the low-frequency subband components of the discrete wavelet transform of the frequency adversarial component and the frequency component; and respectively represent the high-frequency subband components of the discrete wavelet transform of the frequency adversarial component and the frequency component; s represents the high-frequency subband type; and respectively represent the L1 and L2 norms; respectively represent the types of the first high-frequency subband component, the second high-frequency subband component, and the third high-frequency subband component.

6. The method for generating an actively defensive image against face forgery according to claim 1, wherein Projecting the adversarial features onto a normalized device coordinate system to obtain an active face defense image, including: Input the adversarial features into a pre-constructed noise projection alignment unit and output a projection image; Apply a one-time adversarial noise attack to the area outside the projection image to obtain an active face defense image; Wherein, the processing steps of the noise projection alignment unit include: Project the adversarial features onto a normalized device coordinate system to obtain the adversarial features in the device coordinate system; Perform point-plane merging and differentiable rendering on the adversarial features in the device coordinate system in sequence to generate a depth map and a mask; Multiply the adversarial features by the depth map and the mask to obtain a projection image.

7. An active defense image generation system for face forgery, characterized in that, Including: An image acquisition module for: obtaining a face image to be protected; A first projection module for: projecting the face image to be protected onto a parametric two-dimensional coordinate system to obtain a parametric facial texture image; An adversarial application module, configured to: extract the frequency components of the parameterized facial texture image, input the frequency components into a pre-constructed active defense image generation model, and output frequency adversarial components; A feature fusion module, configured to: perform feature fusion on the frequency adversarial components to obtain adversarial features; A second projection module, configured to: project the adversarial features onto a normalized device coordinate system to obtain a face forgery active defense image; Wherein, the active defense image generation model includes an adaptive adversarial noise generator and a low-frequency enhancement attention module connected in sequence; the low-frequency enhancement attention module is used to highlight the importance of image regions; in the adaptive adversarial noise generator, the static momentum coefficient is replaced with a dynamic adaptive momentum coefficient, and the dynamic adaptive momentum coefficient is dynamically adjusted through local variance and global variance, and the global variance and local variance are dynamically adjusted through image gradients; Using discrete wavelet transform to extract the frequency components of the parameterized facial texture image, including: ; Among them, , , , respectively represent the low-frequency component, the first high-frequency component, the second high-frequency component, and the third high-frequency component; x and y represent the spatial coordinates corresponding to the pixels of the parametric facial texture image; represents the parametric facial texture image of size M×N; represents the two-dimensional scaling function for the first-level decomposition of the low-pass filter used to capture ; , , respectively represent the wavelet function for capturing the low-pass row and high-pass column combination of , the wavelet function for capturing the high-pass row and low-pass column combination of , and the wavelet function for capturing the high-pass row and low-pass column combination of ; m and n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components. The active defense image generation model is expressed as: ; ; Among them, and respectively represent the adversarial noise components of the (t + 1)-th iteration and the t-th iteration, and the adversarial noise components are the outputs of the adaptive adversarial noise generator; represents the dynamic adaptive momentum coefficient; represents the frequency adversarial component generated in the current iteration; represents the important region feature map; represents the gradient of the frequency component at the t-th iteration; represents the element-wise multiplication; represents the L1 norm; represents the basic momentum coefficient; and respectively represent the local variance adaptive weight and the global variance adaptive weight; and respectively represent the local variance and the global variance at the pixel position ; and respectively represent the local variance and the global variance; represents the frequency component at the t-th iteration; represents the maximum perturbation amplitude in the current iteration; represents the sign function; represents the truncation function; represents the size of the sliding window; represents the gradient of the frequency component at the pixel position ; represents the total number of elements of the gradient tensor; represents the gradient of the frequency component at the pixel position (i, j, l); represents the average value of the gradient of the frequency component; m and n represent the discrete indices of the pixel positions corresponding to the wavelet transform frequency components; C, H, and W represent the number of channels, height, and width of the input image.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the face forgery active defense image generation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Face image processing method, related device and storage medium

    CN117831089A

  • Face depth counterfeiting active defense method and device based on discrete wavelet transform

    CN119399612A