Active identity information hiding method for face exchange counterfeiting
By building an identity analysis module, an image feature analysis module and an identity information hiding module, generating and optimizing perturbing images, the visual quality degradation and inefficiency of face exchange forgery processing in the prior art is solved, and the effective hiding of identity information and the defense of multiple face-changing attacks is realized.
Patent Information
- Application Number
- CN202510637495.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing active perturbation methods have problems such as visual quality degradation, difficulty in defending against facial exchange, and relying on white box and gray box settings during training to cause high energy consumption, low efficiency, and low generalization.
By building an identity analysis module, image feature analysis module and identity information hiding module, identity features and image features are extracted and processed, perturbed images are generated, and the visual quality of perturbed images is improved through adversarial training.
Effectively hide identity information, resist multiple face-changing attacks, while ensuring the visual consistency and nature of the perturbed image, improving the adaptability and robustness of the model.
Smart Images

Figure CN120163904A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image-level deep forgery, and in particular relates to an active identity information hiding method for face swapping forgery. Background Art
[0002] Deepfakes are technologies that use deep learning to generate synthetic images, audio, or video. Since the public nature of videos, audio, and images of public figures provides a large amount of material for the training of AI models, celebrities around the world often become victims of AI fakes. In addition, the pictures that ordinary people post on social media are also often used for deepfakes, so anyone can become a victim of deepfakes.
[0003] The implementation principle of deep fakes mainly relies on deep neural networks, especially generative adversarial networks. By using a large amount of facial image data for training, the deep fake model can learn the potential relationship between facial features and generate fake facial content that is highly similar to real facial features. At present, deep fake technology is widely used in film production, video editing, virtual reality and other fields, and the rapid development of this technology also brings challenges to maintaining network security.
[0004] Due to the progress of generative models, the current field has a performance bottleneck in passively detecting high-quality deep fake images. Therefore, in recent years, researchers have begun to proactively protect the images of potential victims in advance. Active measures mainly include perturbations and watermarks. Among them, perturbations provide a way to disrupt deep fake operations by inserting signals into benign images. However, existing active perturbation methods still have problems in the following aspects: 1) The visual quality of images is degraded due to the direct addition of perturbations to pixels; 2) Most of them are defense methods against attribute tampering. Due to the high difficulty of face swapping, existing perturbation methods are almost ineffective against it; 3) Inevitably rely on white-box and gray-box settings to involve generative models during training, resulting in high energy consumption, low efficiency, low generalization and other problems. Summary of the invention
[0005] In order to solve the above problems, the present invention provides an active identity information hiding method against face swapping and forgery.
[0006] In order to achieve the above object, the present invention is implemented through the following technical solutions: The present invention provides an active identity information hiding method for face swapping and forgery, comprising the following steps: S1. Obtain a face image, preprocess the image, and obtain a standardized image set; S2. Construct an identity analysis module. The images in the standardized image set are subjected to identity feature analysis through the identity analysis module to obtain identity features. The identity analysis module consists of a ConvBlock block, a max pooling layer, and four SEResBlock blocks in sequence. S3. Construct an image feature analysis module. The images in the standardized image set are subjected to image feature analysis through the image feature analysis module to obtain attribute features. The image feature analysis module consists of three ConvBlock blocks and five SEResBlock blocks in sequence. S4. Construct an identity information hiding module. The identity features are processed through the identity information hiding module to obtain perturbed images. The identity information hiding module consists of a ConvBlock block, three SEResBlock blocks, a ConvBlock block, a SEResBlock block, a DeConvBlock block, a ConvBlock block, a DeConvBlock block, two ConvBlock blocks, and a convolutional layer in sequence. S5. Construct a discriminator Discriminator for adversarial training. The images in the standardized image set and the corresponding perturbed images are input into the discriminator Discriminator for adversarial training, thereby improving the visual quality of the perturbed images. The discriminator Discriminator consists of four NormConvBlock blocks, an adaptive average pooling layer, a flattening layer, a linear layer, a LeakyReLU activation function layer, and a linear layer in sequence. S6. Input the perturbed images and the corresponding images in the standardized image set into the face swapping model for face swapping operation to obtain face swapped images. S7. Calculate the similarity of identity features between the face swapped images and the corresponding images in the standardized image set to indicate the authenticity of the face swapped images and prove whether the identity information is hidden.
[0007] Further, in step S1, the face images are all the face images in the CelebA-HQ dataset , and the images in the CelebA-HQ dataset are standardized to obtain a standardized image set of tensor type with a size of 256 ×256 , is the th face image, , represents the total number of images.
[0008] Further, in step S2, the ConvBlock block of the identity analysis module includes a convolutional layer, a batch normalization layer, and a LeakyReLU activation function layer; the SEResBlock block includes a first convolutional layer, a first batch normalization layer, a first LeakyReLU activation function layer, a second convolutional layer, a second batch normalization layer, a second LeakyReLU activation function layer, a third convolutional layer, a third batch normalization layer, a Sigmoid activation function layer, an adaptive average pooling layer, a fourth convolutional layer, a third LeakyReLU activation function layer, and a fifth convolutional layer; all convolutional layers in the identity analysis module are two-dimensional; The images in the standardized image set are subjected to convolutional processing through the ConvBlock block to obtain a first identity feature ; The first identity feature is passed through a max pooling layer to obtain a second identity feature ; The second identity feature enters the first SEResBlock block and is successively processed through a first convolutional layer, a first batch normalization layer, a first LeakyReLU activation function layer, a second convolutional layer, a second batch normalization layer, a second LeakyReLU activation function layer, a third convolutional layer, a third batch normalization layer, and a Sigmoid activation function layer to obtain a third identity feature ; The third identity feature is successively passed through an adaptive average pooling layer, a fourth convolutional layer, a third LeakyReLU activation function layer, and a fifth convolutional layer to obtain a fourth identity feature ; The second identity feature and the fifth identity feature are added element-wise to obtain a fifth identity feature ; The fifth identity feature is input into the second SEResBlock block to perform the same operations as the first SEResBlock block to obtain a sixth identity feature ; The sixth identity feature is input into the third SEResBlock block to perform the same operations as the second SEResBlock block to obtain a seventh identity feature ; The seventh identity feature is input into the fourth SEResBlock block to perform the same operations as the third SEResBlock block to obtain an eighth identity feature .
[0009] Further, the ConvBlock and SEResBlock in the image feature analysis module in step S3 have the same structure as the ConvBlock and SEResBlock in the identity analysis module; all convolutional layers in the image feature analysis module are two-dimensional; The images in the standardized image set are subjected to convolution processing through three ConvBlock blocks to obtain the first attribute feature ; The first attribute feature is subjected to feature extraction through five SEResBlock blocks to obtain the second attribute feature .
[0010] Further, the ConvBlock and SEResBlock in the identity information hiding module in step S4 have the same structure as the ConvBlock and SEResBlock in the identity analysis module; the DeConvBlock block includes an upsampling layer, a convolutional layer, a batch normalization layer, and a LeakyReLU activation function layer; all convolutional layers in the identity information hiding module are two-dimensional; The eighth identity feature is processed through the first ConvBlock block to obtain the ninth identity feature ; The ninth identity feature is processed through three SEResBlock blocks to obtain the tenth identity feature ; The tenth identity feature generates a random noise tensor with the same shape as through the torch.randn_like function in the torch library of python ; The random noise tensor is processed to obtain the second noise tensor , and the formula is as follows: , where represents the mean of the random noise tensor , represents the standard deviation of the random noise tensor ; The second noise tensor and the tenth identity feature are calculated to obtain the first perturbation feature , and the formula is as follows: , , , , wherein, represents the first learnable parameter, represents a 4-dimensional all-zero tensor with a shape of (1, 64, 1, 1), represents an element-wise multiplication operation, represents the second learnable parameter; , , respectively represent the third noise tensor, the fourth noise tensor, and the fifth noise tensor; the first perturbation feature is processed through the second ConvBlock to obtain the second perturbation feature ; the second perturbation feature obtains the third perturbation feature by taking the mean in the height and width dimensions through the torch.mean function of the torch library in python ; the third perturbation feature passes through a convolutional layer and a Sigmoid activation function layer to obtain a scaling factor , and the scaling factor is multiplied by the second perturbation feature to obtain the fourth perturbation feature ; the fourth perturbation feature and the second attribute feature are concatenated through the concatenation function of the torch library in python to obtain the first concatenated feature , and the formula is as follows: , wherein, represents concatenation in the channel dimension, represents the concatenation function; the first concatenated feature successively passes through the first convolutional layer, the first batch normalization layer, the first LeakyReLU activation function layer, the second convolutional layer, the second batch normalization layer, the second LeakyReLU activation function layer, the third convolutional layer, the third batch normalization layer, and the Sigmoid activation function layer in the fourth SEResBlock to obtain the second concatenated feature ; the second concatenated feature successively passes through an adaptive average pooling layer, the fourth convolutional layer, the third LeakyReLU activation function layer, and the fifth convolutional layer to obtain the third concatenated feature ; the first concatenated feature and the third concatenated feature are added element-wise to obtain the fourth concatenated feature ; the fourth concatenated feature The fifth concatenated feature is obtained through the first DeConvBlock ; The fifth concatenated feature is processed through the third ConvBlock to obtain the sixth concatenated feature ; The sixth concatenated feature The seventh concatenated feature is obtained through the second DeConvBlock ; The images in the standardized image set and the seventh concatenated feature are concatenated through the concatenation function torch.cat of the torch library in python to obtain the eighth concatenated feature ; The eighth concatenated feature is processed through the last two ConvBlocks to obtain the ninth concatenated feature ; The ninth concatenated feature is reconstructed through the last convolutional layer to obtain the perturbed image .
[0011] Furthermore, the loss function in the training process of steps S2 - S4 includes: The mean squared error loss is adopted to calculate the similarity between the original input image and the perturbed image, and the formula is expressed as: , where represents norm, and the Adam optimizer is used to minimize this loss to optimize the similarity between the perturbed image and the original image , and enhance the visual quality of the perturbed image ; The perceptual loss is adopted to calculate the perceptual similarity between the original image and the perturbed image , and the Adam optimizer is used to train the perceptual similarity between the original image and the perturbed image , further enhancing the perturbed image , and the formula is expressed as: , where represents the feature activation of the th layer in the network; Multiple face recognition tools are adopted to extract the identity embedding vector and from and the perturbed identity embedding vector , and the identity loss function is used to calculate the identity embedding vector recognized by the th recognition tool and the perturbed identity embedding vector The identity similarity in is expressed by the formula: , represents the dot product operation, represents the norm; an adaptive weighted loss mechanism is adopted, and the identity loss function is used in this mechanism to adaptively adjust the weighted loss according to the values returned by the face recognition tools during training; the formula of the weighted loss is expressed as follows: , wherein, represents the identity loss after iterations, represents the number of iterations, represents the round, represents the number of face recognition tools, , and 2 face recognition tools are used to participate in the training, namely ArcFace and FaceNet, represents the identity loss of the th identity extractor, represents the loss function weight.
[0012] Furthermore, step S5 specifically includes: The discriminator Discriminator consists of four NormConvBlock blocks, an adaptive average pooling layer, a flattening layer, a linear layer, a LeakyReLU activation function layer, and a linear layer in sequence; the first three of the four NormConvBlock blocks each include a spectral normalization convolutional layer, a LeakyReLU activation function layer, and an SEBlock block; the last NormConvBlock block of the four NormConvBlock blocks includes a spectral normalization convolutional layer and a LeakyReLU activation function layer; the SEBlock block consists of an adaptive average pooling layer, a first linear layer, a RELU activation function layer, a second linear layer, and a Sigmoid activation function layer in sequence; all convolutional layers in the discriminator are two-dimensional; Normalize the images in the image set Pass through the first NormConvBlock block, and sequentially pass through the spectral normalization convolutional layer and the LeakyReLU activation function layer to obtain the first image feature , and the first image feature Adjust the shape to (b, c) through the view function in python to obtain the second image feature , where b represents the batch size and c represents the number of channels of the image. The second image feature Processed by the SEBlock block to obtain the third image feature , the third image feature is adjusted to (b, c, 1, 1) through the view function in python to obtain the fourth image feature ; the fourth image feature passes through the second NormConvBlock block and performs the same operations as the first NormConvBlock block to obtain the fifth image feature ; similarly, the fifth image feature passes through the third NormConvBlock block to obtain the sixth image feature ; the sixth image feature passes through the last NormConvBlock block to obtain the seventh image feature ; the seventh image feature successively passes through the adaptive average pooling layer, flattening layer, linear layer, and LeakyReLU activation function layer to obtain the eighth image feature ; most of the eighth image features pass through the last linear layer to obtain the classification result; similarly, the perturbed image is classified by the discriminator Discriminator to obtain the classification result; the classification result is the original image or the perturbed image.
[0013] Furthermore, in step S5, the discriminator Discriminator in the formula uses the adversarial loss , where represents the expectation,[[]] represents the discriminator; the enhanced loss is used to enhance the quality of the perturbed image , and the formula is as follows: , and the Adam optimizer is used to train the visual quality of the perturbed image .
[0014] Furthermore, step S6 specifically includes: Input the perturbed image and the image in its corresponding standardized image set into the face swapping model together to perform face swapping operation to obtain the face swapped image ; use the perturbed image as the source face image, and the image in the corresponding standardized image set as the target face image; the face swapping model uses an existing face swapping model.
[0015] Furthermore, step S7 specifically includes: The face-swapped image The image in the corresponding standardized image set Pass through the ArcFace model for identity features to obtain an identity vector ; The face-swapped image Pass through the ArcFace model for identity features to obtain a face-swapped identity vector , and judge whether the identity hiding is successful by calculating the identity similarity. The formula is as follows: , wherein, represents the identity similarity; when , it proves that the face-swapped image is fake, thus proving that the identity hiding is successful.
[0016] The advantages of the present invention are as follows: The present invention extracts the identity features of the input image through the identity analysis module, and extracts the shallow image features through the image feature analysis module. Subsequently, the two are jointly input into the identity information hiding module. This module performs feature purification and perturbation initialization based on the identity features, generates perturbation features with a hiding function, and introduces shallow features for image reconstruction, so as to hide the identity information to the greatest extent while maintaining the naturalness of the image appearance. Experimental results show that if the perturbed image is used as the source identity input for face-swapping attacks, the original image identity cannot be restored in the generated forged results, verifying the effectiveness of the present invention in identity hiding. In addition, to enhance the defense versatility against different face-swapping algorithms, the present invention introduces an adaptive weighted loss mechanism, which can dynamically adjust the identity loss weight according to the feedback results of the face recognition tool during the training process, further improving the adaptability and robustness of the model. Generally speaking, the present invention can not only effectively hide identity information, resist various face-swapping attacks, but also ensure the visual consistency and naturalness of the perturbed image, and has good practicability and application prospects. Description of the Drawings
[0017] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.
[0018] Figure 1 is the step flow chart of the method of the present invention; Figure 2 is the visualization result of the image sample. Specific Embodiments
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] Embodiment 1 In this embodiment, as Figure 1 shown, the present invention provides an active identity information hiding method for face swapping forgery. The specific steps include: S1. Obtain a face image, preprocess the image, and obtain a standardized image set.
[0021] The face image is all the face images in the CelebA-HQ dataset , and the images in the CelebA-HQ dataset are standardized to obtain a standardized image set of tensor type with a size of 256 × 256, where is the th face image, and , represents the total number of images.
[0022] Specifically, through the transforms.Compose() function in the torchvision library in Python, the transforms.Compose() function sequentially includes three image preprocessing operation functions: the transforms.Resize() function, the transforms.ToTensor() function, and the transforms.Normalize() function; all the face images in the CelebA-HQ dataset are preprocessed to obtain a standardized image set of tensor type with a size of 256 × 256, where , is the th face image. The CelebA-HQ dataset consists of 30,000 face images, containing a total of 6,217 different face identities, and the resolution of each image is 1024 1024, 30,000 face images are divided into a training set, a validation set, and a test set according to the official partitioning scheme provided in the paper "Deep Learning Face Attributes in the Wild". The images in the LFW dataset contain 5,749 different face identities, and the resolution of each image is 250 250. The images corresponding to the 5,749 face identities in the LFW dataset are used for cross-dataset testing.
[0023] S2. Construct an identity analysis module. The images in the standardized image set are subjected to identity feature analysis through the identity analysis module to obtain identity features.
[0024] Specifically, the identity analysis module consists of a ConvBlock block, a max pooling layer, and four SEResBlock blocks in sequence; the ConvBlock block includes a convolutional layer, a batch normalization layer, and a LeakyReLU activation function layer; the number of channels of the convolutional layer in the ConvBlock block is 32, the convolutional kernel is 7, the stride is 2, and the padding is 3; the window size of the max pooling layer is 3, the stride is 2, and the padding is 1; the SEResBlock block includes a first convolutional layer, a first batch normalization layer, a first LeakyReLU activation function layer, a second convolutional layer, a second batch normalization layer, a second LeakyReLU activation function layer, a third convolutional layer, a third batch normalization layer, a Sigmoid activation function layer, an adaptive average pooling layer, a fourth convolutional layer, a third LeakyReLU activation function layer, and a fifth convolutional layer; all convolutional layers in the identity analysis module are two-dimensional; the number of channels of the first convolutional layer is 64, the convolutional kernel is 1, the stride is 1, and the padding is 0; the number of channels of the second convolutional layer is 64, the convolutional kernel is 3, and both the stride and padding are 1; the number of channels of the third convolutional layer is 64, the convolutional kernel is 1, and both the stride and padding are 1, with the bias set to 0; the output feature size of the adaptive average pooling layer is 1 1; the number of channels of the fourth convolutional layer is 4, the convolutional kernel is 1, and both the stride and padding are 1, with the bias set to 0; the number of channels of the fifth convolutional layer is 64, the convolutional kernel is 1, and both the stride and padding are 1, with the bias set to 0.
[0025] The images in the standardized image set are subjected to convolutional processing through the ConvBlock block to obtain the first identity feature ; the first identity feature is subjected to the max pooling layer to obtain the second identity feature ; the second identity feature Enter the first SEResBlock, and process it through the first convolutional layer, the first batch normalization layer, the first LeakyReLU activation function layer, the second convolutional layer, the second batch normalization layer, the second LeakyReLU activation function layer, the third convolutional layer, the third batch normalization layer, and the Sigmoid activation function layer in sequence to obtain the third identity feature ; The third identity feature Pass through the adaptive average pooling layer, the fourth convolutional layer, the third LeakyReLU activation function layer, and the fifth convolutional layer in sequence to obtain the fourth identity feature ; Add the second identity feature and the fifth identity feature element-wise to obtain the fifth identity feature ; Input the fifth identity feature into the second SEResBlock to perform the same operations as the first SEResBlock to obtain the sixth identity feature ; Input the sixth identity feature into the third SEResBlock to perform the same operations as the second SEResBlock to obtain the seventh identity feature ; Input the seventh identity feature into the fourth SEResBlock to perform the same operations as the third SEResBlock to obtain the eighth identity feature .
[0026] S3. Construct an image feature analysis module, and perform image feature analysis on the images in the standardized image set through the image feature analysis module to obtain attribute features.
[0027] Specifically, the image feature analysis module is sequentially composed of three ConvBlock blocks and five SEResBlock blocks; the ConvBlock blocks and SEResBlock blocks have the same structure as the ConvBlock blocks and SEResBlock blocks in the identity analysis module; the number of channels of the convolutional layer in the first ConvBlock block is 64, the convolutional kernel is 7, the stride is 2, and the padding is 3; the number of channels of the convolutional layer in the second ConvBlock block is 128, the convolutional kernel is 3, the stride is 2, and the padding is 1; the number of channels of the convolutional layer in the third ConvBlock block is 64, the convolutional kernel is 3, the stride is 1, and the padding is 1; all convolutional layers in the image feature analysis module are two-dimensional; The images in the standardized image set are subjected to convolutional processing through three ConvBlock blocks to obtain the first attribute feature ; The first attribute feature Feature extraction is performed through five SEResBlock blocks to obtain the second attribute feature .
[0028] S4. Construct an identity information hiding module. The identity feature is processed by the identity information hiding module to obtain a perturbed image.
[0029] Specifically, the identity information hiding module consists of the sequence of one ConvBlock block, three SEResBlock blocks, one ConvBlock block, one SEResBlock block, one DeConvBlock block, one ConvBlock block, one DeConvBlock block, two ConvBlock blocks and one convolutional layer; the ConvBlock block and SEResBlock block have the same structure as the ConvBlock block and SEResBlock block in the identity analysis module; the DeConvBlock block includes an upsampling layer, a convolutional layer, a batch normalization layer and a LeakyReLU activation function layer; all convolutional layers in the identity information hiding module are two-dimensional; The eighth identity feature is processed by the first ConvBlock block to obtain the ninth identity feature . The number of channels of the convolutional layer of the first ConvBlock block is 64, the convolutional kernel is 3, and the stride and padding are both 1; the ninth identity feature is processed through three SEResBlock blocks to obtain the tenth identity feature ; The tenth identity feature generates a random noise tensor with the same shape as through the torch.randn_like function in the torch library of python ; the random noise tensor is processed to obtain the second noise tensor , and the formula is expressed as follows: , where represents the mean of the random noise tensor , represents the standard deviation of the random noise tensor ; the second noise tensor and the tenth identity feature are calculated to obtain the first perturbation feature , and the formula is expressed as follows: , , , , Among them, represents the first learnable parameter; represents a 4-dimensional all-zero tensor with a shape of (1, 64, 1, 1), represents an element-wise multiplication operation, represents the second learnable parameter; , , respectively represent the third noise tensor, the fourth noise tensor, and the fifth noise tensor; the first perturbation feature is processed by the second ConvBlock to obtain the second perturbation feature . The number of channels of the convolutional layer of the second ConvBlock is 64, the convolutional kernel is 3, and the stride and padding are both 1; the second perturbation feature obtains the third perturbation feature by taking the mean in the height and width dimensions through the torch.mean function in the torch library of python; the third perturbation feature passes through a convolutional layer and a Sigmoid activation function layer to obtain a scaling factor . The scaling factor is multiplied by the second perturbation feature to obtain the fourth perturbation feature ; the fourth perturbation feature and the second attribute feature are concatenated through the torch.cat function in the torch library of python to obtain the first concatenated feature . The formula is as follows: , Among them, represents concatenation in the channel dimension, represents the concatenation function; The first concatenated feature successively passes through the first convolutional layer, the first batch normalization layer, the first LeakyReLU activation function layer, the second convolutional layer, the second batch normalization layer, the second LeakyReLU activation function layer, the third convolutional layer, the third batch normalization layer, and the Sigmoid activation function layer in the fourth SEResBlock to obtain the second concatenated feature ; the second concatenated feature successively passes through an adaptive average pooling layer, the fourth convolutional layer, the third LeakyReLU activation function layer, and the fifth convolutional layer to obtain the third concatenated feature ; In the fourth SEResBlock, the number of channels of the first convolutional layer is 128, the convolutional kernel is 1, the stride is 1, and the padding is 0; the number of channels of the second convolutional layer is 128, the convolutional kernel is 3, and both the stride and padding are 1; the number of channels of the third convolutional layer is 128, the convolutional kernel is 1, and both the stride and padding are 1, and the bias is set to 0; the output feature size of the adaptive average pooling layer is 1×1; the number of channels of the fourth convolutional layer is 8, the convolutional kernel is 1, and both the stride and padding are 1, and the bias is set to 0; the number of channels of the fifth convolutional layer is 128, the convolutional kernel is 1, and both the stride and padding are 1, and the bias is set to 0; the first concatenated feature and the third concatenated feature are added element-wise to obtain the fourth concatenated feature ; The fourth concatenated feature passes through the first DeConvBlock to obtain the fifth concatenated feature ; In the first DeConvBlock, the scaling factor of the upsampling layer is 2, the interpolation mode is bilinear interpolation, and the alignment corner is set to 1; the number of channels of the convolutional layer is 128, the convolutional kernel is 3, and both the stride and padding are 1; the fifth concatenated feature is processed by the third ConvBlock to obtain the sixth concatenated feature ; The number of channels of the convolutional layer in the third ConvBlock is 64, the convolutional kernel is 3, and both the stride and padding are 1; the sixth concatenated feature passes through the second DeConvBlock to obtain the seventh concatenated feature ; In the second DeConvBlock, the scaling factor of the upsampling layer is 2, the interpolation mode is bilinear interpolation, and the alignment corner is set to 1; the number of channels of the convolutional layer is 64, the convolutional kernel is 3, and both the stride and padding are 1; the standardized image in the image set and the seventh concatenated feature are concatenated through the concatenation function torch.cat in the torch library of python to obtain the eighth concatenated feature ; The eighth concatenated feature passes through the last two ConvBlocks to obtain the ninth concatenated feature ; In the last two ConvBlocks, the number of channels of the convolutional layer in the previous ConvBlock is 32, the convolutional kernel is 3, and both the stride and padding are 1; the number of channels of the convolutional layer in the latter ConvBlock is 16, the convolutional kernel is 3, and both the stride and padding are 1; the ninth concatenated feature is reconstructed through the last convolutional layer to obtain the perturbed image ; The number of channels of the last convolutional layer is 3, the convolutional kernel is 3, and both the stride and padding are 1.
[0030] S5. Construct a discriminator Discriminator for adversarial training. Input the images in the standardized image set and the corresponding perturbed images into the discriminator Discriminator for adversarial training, thereby improving the visual quality of the perturbed images.
[0031] Specifically, the discriminator Discriminator consists of four NormConvBlock blocks, an adaptive average pooling layer, a flattening layer, a linear layer, a LeakyReLU activation function layer, and a linear layer in sequence; the first three of the four NormConvBlock blocks each include a spectral normalization convolutional layer, a LeakyReLU activation function layer, and an SEBlock block; the last of the four NormConvBlock blocks includes a spectral normalization convolutional layer and a LeakyReLU activation function layer; the SEBlock block consists of an adaptive average pooling layer, a first linear layer, a RELU activation function layer, a second linear layer, and a Sigmoid activation function layer in sequence; all convolutional layers in the discriminator are two-dimensional.
[0032] Images in the standardized image set Through the first NormConvBlock block, pass through the spectral normalization convolutional layer and the LeakyReLU activation function layer in sequence to obtain the first image feature , the first image feature Adjust the shape to (b, c) through the view function in Python to obtain the second image feature , where b represents the batch size and c represents the number of channels of the image. The second image feature Is processed through the SEBlock block to obtain the third image feature , the third image feature Adjust to (b, c, 1, 1) through the view function in Python to obtain the fourth image feature ; The spectral normalization convolutional layer in the first NormConvBlock block is obtained by spectral normalizing a convolutional layer with 64 channels, a convolutional kernel of 4, a stride of 2, and a padding of 1 through the spectral_norm function in the torch.nn.utils module in Python; the window size of the adaptive average pooling layer in the SEBlock block of the first NormConvBlock block is 1, the number of output nodes of the first linear layer is 4, and the number of output nodes of the second linear layer is 64; the fourth image feature Through the second NormConvBlock, perform the same operations as the first NormConvBlock to obtain the fifth image feature ; The spectral normalization convolutional layer in the second NormConvBlock is obtained by performing spectral normalization on a convolutional layer with 128 channels, a convolutional kernel of 4, a stride of 2, and a padding of 1 through the spectral_norm function in the torch.nn.utils module in Python; the number of output nodes of the first linear layer in the SEBlock of the second NormConvBlock is 8, and the number of output nodes of the second linear layer is 128; similarly, the fifth image feature passes through the third NormConvBlock to obtain the sixth image feature ; The spectral normalization convolutional layer in the third NormConvBlock is obtained by performing spectral normalization on a convolutional layer with 256 channels, a convolutional kernel of 4, a stride of 2, and a padding of 1 through the spectral_norm function in the torch.nn.utils module in Python; the number of output nodes of the first linear layer in the SEBlock of the third NormConvBlock is 16, and the number of output nodes of the second linear layer is 256; the sixth image feature passes through the last NormConvBlock to obtain the seventh image feature ; The spectral normalization convolutional layer of the last NormConvBlock is obtained by performing spectral normalization on a convolutional layer with 512 channels, a convolutional kernel of 4, a stride of 2, and a padding of 1 through the spectral_norm function in the torch.nn.utils module in Python; the seventh image feature successively passes through an adaptive average pooling layer, a flattening layer, a linear layer, and a LeakyReLU activation function layer to obtain the eighth image feature , where the window size of the adaptive average pooling layer is 1, and the number of output nodes of the linear layer is 128; most of the eighth image features pass through the last linear layer to obtain the classification result, and the number of output nodes of the last linear layer is 1; similarly, the perturbed image is classified by the discriminator Discriminator to obtain the classification result; the classification result is the original image or the perturbed image.
[0033] S6. Input the perturbed image and the images in its corresponding standardized image set into the face swapping model for face swapping operation to obtain the face swapped image.
[0034] Specifically, the perturbed image The images in the corresponding standardized image set are input into the face swapping model together for face swapping operation to obtain the face swapped image ; The perturbed image is used as the source face image, and the image in the corresponding standardized image set is used as the target face image; The face swapping model uses an existing face swapping model.
[0035] The face swapping model refers to the face swapping models of existing works, such as SimSwap, InfoSwap, UniFace, E4S, DiffSwap. Among them, SimSwap uses the source code of the paper "SimSwap: An Efficient Framework For High-Fidelity Face Swapping" to implement face swapping, InfoSwap uses the source code of the paper "InfoSwap: Information Bottleneck Disentanglement for Identity Swapping" to implement face swapping, UniFace uses the source code of the paper "Designing One Unified Framework for High-Fidelity Face Reenactment and Swapping" to implement face swapping, E4S uses the source code of the paper "Fine-grained face swapping via regional gan inversion" to implement face swapping, and DiffSwap uses the source code of the paper "Diffswap: High-fidelity and controllable face swapping via 3d-aware masked diffusion" to implement face swapping.
[0036] S7. Calculate the identity feature similarity between the face swapped image and the image in its corresponding standardized image set to indicate the authenticity of the face swapped image and prove whether the identity information is hidden.
[0037] Specifically, the face swapped image and the image in the corresponding standardized image set undergo identity features through the ArcFace model to obtain the identity vector ; The face swapped image undergoes identity features through the ArcFace model to obtain the face swapped identity vector , and judge whether the identity hiding is successful by calculating the identity similarity. The formula is as follows: , where Indicates identity similarity; when it is, it proves that the face-swapped image is fake, thus proving that the identity hiding is successful.
[0038] Specifically, the loss function in the training process of steps S2 - S4 includes: To maintain the visual quality of image reconstruction while hoping to make the face recognition tool used by the deepfake model extract incorrect identity information, thus invalidating malicious face tampering. The mean squared error loss is used to calculate the similarity between the original input image and the perturbed image, and the formula is as follows: , where represents norm, and the Adam optimizer is used to minimize this loss to optimize the perturbed image and the original image to enhance the visual quality of the perturbed image .
[0039] Regarding the potential artifacts ignored by the mean squared error (MSE) loss function, such as blurring and loss of high-frequency details. The perceptual loss is used to calculate the perceptual similarity between the original image and the perturbed image , and the Adam optimizer is used to train the perceptual similarity between the original image and the perturbed image to further enhance the perturbed image , and the formula is as follows: , where, represents the feature activation of the th layer in the network; To ensure the generalization ability in hiding identity information, multiple face recognition tools are used to extract the identity embedding vectors and from and the perturbed identity embedding vector , and the identity similarity in the identity embedding vector recognized by the th recognition tool is calculated through the identity loss function and the perturbed identity embedding vector , and the formula is as follows: ,, represents the dot product operation, Denote the modulus length; the present invention aims to counter face recognition tools and make them ineffective in identity embedding extraction. To ensure generalization, multiple face recognition tools are introduced, and an adaptive weighted loss is designed to adaptively balance different losses. The weights are dynamically adjusted through loss variance (measuring stability and reducing the weights of high-noise losses) and relative progress (prioritizing the optimization of slower-progressing losses) to ensure training stability and task balance. The adaptive weighted loss mechanism is adopted, and the identity loss function is used in this mechanism to adaptively adjust the weighted loss according to the values returned by the participating face recognition tools during training; the formula for the weighted loss is as follows: , wherein, denotes the identity loss at the -th iteration, denotes the number of iterations, denotes the -th round, denotes the number of face recognition tools, , and 2 face recognition tools, namely ArcFace and FaceNet, are used to participate in the training, denotes the identity loss of the -th identity extractor, denotes the weight of the loss function; for each of the loss functions in the -th iteration in the -th round, the weight is formulated as follows: , wherein, denotes the weight corresponding to the -th face recognition tool, denotes the weight corresponding to the -th face recognition tool; the weight is formulated as follows: , wherein, denotes taking the maximum value, denotes the constant regularization factor for balancing loss variance and relative progress, denotes the variance, denotes the relative progress, denotes the parameter to be dynamically adjusted; denotes the lower limit of the denominator value, denotes the lower limit of , ; the parameter to be dynamically adjusteddepends on the current and is formulated as follows: , Among them, represents taking the minimum value, represents controlling the growth coefficient of the change speed, and respectively represent and the upper limit values of, represents the initial value of; The variance is calculated according to the most recent of the current iteration count consecutive loss values, and the formula is as follows: , Among them, represents traversing the loss values and calculating the squared deviation, represents traversing the loss values and calculating the squared deviation; The relative progress is calculated based on the loss values of the current and previous iterations, and the formula is as follows: , Among them, represents a constant used to prevent the denominator from being zero or approaching zero, .
[0040] Specifically, in step S5, the discriminator Discriminator adopts the adversarial loss for adversarial training, and the formula is as follows: , Among them, represents the expectation, represents the discriminator; To improve the visual quality of the perturbed image , guide the output result generated when it passes through the discriminator to be close to the discrimination result of the original image. This process makes the visual features of the perturbed image gradually tend to the features of the original image, thereby improving the naturalness and consistency of the perturbed image. The enhanced loss is used to enhance the quality of the perturbed image , and the formula is as follows: , and the Adam optimizer is used to train the visual quality of the perturbed image .
[0041] Example 2 In this embodiment, a comparative experiment was conducted. The experiment used two existing commonly used face image datasets, the CelebA-HQ dataset and the LFW dataset. The CelebA-HQ dataset was officially divided into a training set, a validation set, and a test set, while the LFW dataset only participated in the test. The experiment was carried out on images resized to 256 ×256, and the present invention was not involved in tampering with the model during the entire training process.
[0042] To verify the reliability of the inventive method, a comparative experiment was conducted and the experimental results were presented.
[0043] The present invention selected the Initiative watermark model, the AntiForgery watermark model, the CMUA model, and the DF-RAP watermark model for comparative experiments. And the watermark models respectively countered the SimSwap, InfoSwap, UniFace, E4S, and DiffSwap face swap models. For fairness, the publicly available model weights corresponding to the best performance of the above algorithms were directly adopted.
[0044] The bold data in the following table are all the best results.
[0045] Table 1 Quantitative visual quality assessment of perturbed images on the CelebA-HQ dataset As shown in Table 1, the present invention calculated the average peak signal-to-noise ratio, the structural similarity index, and the perceptual image patch similarity of the images in the standardized image set and the perturbed images. The method of the present invention achieved state-of-the-art visual performance. The average peak signal-to-noise ratio of the method of the present invention was higher than 41, the average peak signal-to-noise ratio of the others was lower than 40, and the structural similarity index and the perceptual image patch similarity of the method of the present invention were also higher than those of other models.
[0046] Table 2 TOP-5 and TOP-1 accuracies of identity matching performance after using the perturbation algorithm After obtaining the perturbed images it was tested whether the perturbed images still remained in the same identity cluster or were successfully hidden. For this purpose, a face recognition task evaluation of the identity labels provided by the CelebA-HQ dataset was performed on each perturbed image. Specifically, the identity embedding and each The top-5 and top-1 accuracies of the cosine similarity between embeddings. To demonstrate the generality of the method proposed in the present invention, in addition to ArcFace and FaceNet involved in the training stage, the present invention also employs VGGFace and SFace as tools for disturbing face recognition in the test stage of the present invention. As shown in Table 2, although the numerical experimental results vary due to the performance of different face recognition tools, their top-5 and top-1 accuracies for face identity recognition on clean test images (images without perturbations) generally maintain the expected highest performance. On the other hand, although the most advanced comparative active perturbation models have led to a decrease in accuracy in most works, except for DF-RAP, when facing SFace, the differences from the unperturbed setting are mostly less than 0.1 in absolute value. In addition, each comparative model even unexpectedly encountered higher top-5 and top-1 accuracies than in the unperturbed setting. Therefore, the method proposed in the present invention achieved average top-5 and top-1 face recognition accuracies of 0.711 and 0.565 respectively, showing the potential to outperform the existing state-of-the-art active perturbation algorithms and achieving the optimal overall effect in disturbing the performance of face recognition systems. The specific calculation methods for top5 and top 1 accuracies are as follows: First, for each picture, extract its identity vector. Then, for the pictures compared with it (pictures in the standard image set), extract the identity vectors respectively, and then calculate the cosine similarity between these two identity vectors. Then, take the top five and top one similarities and denote them as top5 and top 1 respectively. Next, determine whether the true identity of the image appears in the Top-1 or Top-5 matching results: If the true identity is in Top-1, then the sample is recorded as "correct" in the Top-1 accuracy calculation; if the true identity is in Top-5, then the sample is recorded as "correct" in the Top-5 accuracy calculation; otherwise, it is recorded as "wrong". Finally, by statistically counting the proportion of correct samples in the evaluation results of all test images, the overall Top-1 and Top-5 accuracies are calculated respectively. The arrow mark (↓) in the table indicates that the lower these indicators, the better the model performance.
[0047] Table 3 Identity Similarity Evaluation on CelebA-HQ Dataset Table 3 shows the evaluation of the identity similarity between the face-swapping results on perturbed images and those on clean images using various identity swapping forgery models and the perturbation algorithm on the CelebA-HQ dataset. As shown in Table 3, regardless of which face-swapping algorithm is used, the cosine similarities of the face-swapping results obtained by the Initiative model, the Anti-Forgery model, and the CMUA model are all around 0.9. Except for DiffSwap, the inserted perturbations show reasonable defensive capabilities for relatively simple tasks such as attribute editing, but usually cannot distort or counter identity swapping forgery. However, these perturbations lack universality for other face-swapping models, resulting in similarities around 0.850 or above. DiffSwap is a diffusion-based face-swapping method. Although the generated results are unstable due to its original performance, the method of the present invention can still counter it. When using the ArcFace and VGGFace identity extractors, the identity similarities are as low as 0.310 and 0.352 respectively. Therefore, the method proposed by the present invention can always counter different face-swapping models by hiding identity information, and the average similarities under the ArcFace and VGGFace settings reach 0.334 and 0.330 respectively, and each similarity is lower than 0.6.
[0048] Table 4 Identity Similarity Evaluation on the LFW Dataset Table 4 shows the evaluation of the identity similarity between the face-swapping results on perturbed images and those on clean images using different active perturbation algorithms on LFW and various identity swapping forgery models. The present invention conducted face-swapping experiments on the LFW dataset and evaluated the cross-dataset identity similarity performance. Generally speaking, the proposed method achieved the best performance in most face-swapping models except SimSwap, and DF-RAP has proven to have reliable defensive capabilities especially in terms of SimSwap. On the other hand, due to the inability to maintain satisfactory visual quality on the LFW dataset and generate texture checkerboard artifacts in the perturbed benign images, Initiative obtained the second-best defensive performance when facing InfoSwap and UniFace. Therefore, the method of the present invention always limits the identity similarity of all face-swapping models to below 0.650 or equal to 0.650, and obtains the best average similarities of 0.541 and 0.515 on ArcFace and VGGFace respectively.
[0049] In addition, the present invention also conducts experiments on Figure 2Some samples are visualized. The first column is the original image, which is also the original image without perturbation. The second column is the perturbed image. The third column is the target image without perturbation. The fourth column is the result of face swapping between the first column and the third column. The fifth column is the result of face swapping between the second column and the third column. It can be seen from the figure that the perturbed image of the present invention is almost visually indistinguishable from the original image, indicating its good visual quality. At the same time, there are obvious differences between the face-swapped images in the fourth column and the fifth column, which confirms the success of hiding identity information in the present invention.
[0050] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An active identity information hiding method against face swapping and forgery, characterized in that: The following steps are involved: S1. Obtain a face image, preprocess the image, and obtain a standardized image set; S2. Construct an identity analysis module, and analyze the identity features of the images in the standardized image set through the identity analysis module to obtain identity features; the identity analysis module is composed of a ConvBlock block, a maximum pooling layer and four SEResBlock blocks in sequence; S3. Construct an image feature analysis module, and perform image feature analysis on the images in the standardized image set through the image feature analysis module to obtain attribute features; the image feature analysis module is composed of three ConvBlock blocks and five SEResBlock blocks in sequence; S4. Construct an identity information hiding module, wherein the identity feature is processed by the identity information hiding module to obtain a perturbed image; the identity information hiding module is composed of a ConvBlock block, three SEResBlock blocks, a ConvBlock block, a SEResBlock block, a DeConvBlock block, a ConvBlock block, a DeConvBlock block, two ConvBlock blocks and a convolutional layer in sequence; S5. Construct a discriminator for adversarial training, input the images in the standardized image set and the corresponding perturbed images into the discriminator for adversarial training, so as to improve the visual quality of the perturbed images; the discriminator is composed of four NormConvBlock blocks, an adaptive average pooling layer, a flattening layer, a linear layer, a LeakyReLU activation function layer, and a linear layer in sequence; S6. Inputting the perturbed image and its corresponding image in the standardized image set into the face-changing model to perform a face-changing operation to obtain a face-changing image; S7. Calculate the identity feature similarity of the face-swapped image and the corresponding image in the standardized image set to indicate the authenticity of the face-swapped image and prove whether the identity information is hidden.
2. The method for actively hiding identity information against face swapping and forgery according to claim 1, characterized in that: Step S1 specifically includes: The face images are all the face images in the CelebA-HQ dataset , the images in the CelebA-HQ dataset are normalized to a size of 256 256 tensor type standardized image set , For the Personal face image, , Indicates the total number of images.
3. The method for actively hiding identity information against face swapping and forgery according to claim 2, characterized in that: Step S2 specifically includes: In the identity analysis module, the ConvBlock block includes a convolution layer, a batch normalization layer, and a LeakyReLU activation function layer; the SEResBlock block includes a first convolution layer, a first batch normalization layer, a first LeakyReLU activation function layer, a second convolution layer, a second batch normalization layer, a second LeakyReLU activation function layer, a third convolution layer, a third batch normalization layer, a Sigmoid activation function layer, an adaptive average pooling layer, a fourth convolution layer, a third LeakyReLU activation function layer, and a fifth convolution layer; all convolution layers in the identity analysis module are two-dimensional; The images in the standardized image set After convolution processing by ConvBlock block, the first identity feature is obtained The first identity feature The second identity feature is obtained through the maximum pooling layer The second identity feature Enter the first SEResBlock block, and go through the first convolution layer, the first batch of normalization layers, the first LeakyReLU activation function layer, the second convolution layer, the second batch of normalization layers, the second LeakyReLU activation function layer, the third convolution layer, the third batch of normalization layers, and the Sigmoid activation function layer to get the third identity feature. The third identity feature The fourth identity feature is obtained by sequentially passing through the adaptive average pooling layer, the fourth convolutional layer, the third LeakyReLU activation function layer and the fifth convolutional layer. ; The second identity feature and the fifth identity characteristic Add element by element to get the fifth identity feature ; The fifth identity feature Input to the second SEResBlock block to perform the same operation as the first SEResBlock block to obtain the sixth identity feature ; The sixth identity feature Input to the third SEResBlock block to perform the same operation as the second SEResBlock block to obtain the seventh identity feature ; The seventh identity feature Input to the fourth SEResBlock block to perform the same operation as the third SEResBlock block to obtain the eighth identity feature .
4. The method for actively hiding identity information against face swapping and forgery according to claim 3, characterized in that: Step S3 specifically includes: The structures of the ConvBlock and SEResBlock blocks in the image feature analysis module are the same as those of the ConvBlock and SEResBlock blocks in the identity analysis module; all convolutional layers in the image feature analysis module are two-dimensional; The images in the standardized image set After three ConvBlock blocks are used for convolution processing, the first attribute feature is obtained The first attribute feature After five SEResBlock blocks for feature extraction, the second attribute feature is obtained .
5. The method for actively hiding identity information against face swapping and forgery according to claim 4, characterized in that: Step S4 specifically includes: The ConvBlock and SEResBlock blocks in the identity information hiding module have the same structure as the ConvBlock and SEResBlock blocks in the identity analysis module; the DeConvBlock block includes an upsampling layer, a convolution layer, a batch normalization layer, and a LeakyReLU activation function layer; all convolution layers in the identity information hiding module are two-dimensional; The eighth identity characteristic After processing by the first ConvBlock block, the ninth identity feature is obtained ; The ninth identity characteristic After processing by three SEResBlock blocks, the tenth identity feature is obtained ; The tenth identity characteristic The torch.randn_like function in the torch library in python generates a A random noise tensor of the same shape ; For the random noise tensor Processing is performed to obtain the second noise tensor , the formula is as follows: , in, Represents a random noise tensor The mean of Represents a random noise tensor The standard deviation of the second noise tensor Tenth Identity Characteristics Calculate and get the first disturbance characteristic , the formula is as follows: , , , , in, represents the first learnable parameter, Represents a 4-dimensional all-zero tensor with a shape of (1,64,1,1). represents an element-wise multiplication operation, represents the second learnable parameter; , , Respectively represent the third noise tensor, the fourth noise tensor, and the fifth noise tensor; the first disturbance feature After being processed by the second ConvBlock block, the second perturbation feature is obtained The second disturbance characteristic The third perturbation feature is obtained by taking the average value in the height and width dimensions through the torch.mean function of the torch library in Python The third disturbance characteristic After the convolution layer and the Sigmoid activation function layer, the scaling factor is obtained , scaling factor With the second perturbation feature Multiply to get the fourth perturbation characteristic ; The fourth disturbance characteristic Second attribute characteristics The first splicing feature is obtained by splicing through the splicing function of the torch library in Python , the formula is as follows: , in, represents concatenation in the channel dimension, represents the concatenation function; The first splicing feature The second concatenated feature is obtained by sequentially passing through the first convolution layer, the first normalization layer, the first LeakyReLU activation function layer, the second convolution layer, the second normalization layer, the second LeakyReLU activation function layer, the third convolution layer, the third normalization layer, and the Sigmoid activation function layer in the fourth SEResBlock block. ; Second splicing feature The third concatenated feature is obtained by sequentially passing through the adaptive average pooling layer, the fourth convolutional layer, the third LeakyReLU activation function layer and the fifth convolutional layer. ; The first stitching feature and the third splicing feature Add element by element to get the fourth concatenation feature ; The fourth splicing feature After the first DeConvBlock block, the fifth concatenated feature is obtained ; The fifth splicing feature After processing by the third ConvBlock block, the sixth concatenated feature is obtained ; Sixth splicing feature After the second DeConvBlock block, the seventh concatenated feature is obtained ; The image in the standardized image set With the seventh splicing feature The eighth splicing feature is obtained by splicing through the splicing function torch.cat of the torch library in python The eighth splicing feature After the last two ConvBlock blocks, the ninth concatenated feature is obtained The ninth splicing feature After the last convolutional layer, the perturbed image is reconstructed .
6. The method for actively hiding identity information against face swapping and forgery according to claim 5, characterized in that: The loss function in the training process of steps S2-S4 includes: Using mean square error loss To calculate the similarity between the original input image and the perturbed image, the formula is expressed as: ,in express norm, and use the Adam optimizer to minimize the loss and optimize the perturbation image With the original image The similarity between the enhanced perturbation images Visual quality; Using perceptual loss Calculate the original image and the perturbed image The perceptual similarity between the original images is trained with the Adam optimizer and the perturbed image The perceptual similarity between them further enhances the perturbation image , the formula is: ,in, Indicates the network Feature activation of the layer; Using a variety of face recognition tools and Extract identity embedding vector from and the perturbed identity embedding vector , through the identity loss function Calculate the Identity embedding vectors identified by identification tools and the perturbed identity embedding vector The identity similarity in is expressed as: , represents the dot product operation, Represents the modulus length; an adaptive weighted loss mechanism is adopted, in which the identity loss function is used to adaptively adjust the weighted loss according to the value returned by the participating face recognition tool during training; the formula of the weighted loss is expressed as follows: , in, Represents iteration Second loss of identity, represents the number of iterations, Indicates wheel, Indicates the number of face recognition tools, , using 2 face recognition tools for training, ArcFace and FaceNet, Indicates The identity loss of the identity extractor, Represents the loss function weight.
7. The method for actively hiding identity information against face swapping and forgery according to claim 6, characterized in that: Step S5 specifically includes: The discriminator is composed of four NormConvBlock blocks, an adaptive average pooling layer, a flattening layer, a linear layer, a LeakyReLU activation function layer, and a linear layer in order; the first three NormConvBlock blocks of the four NormConvBlock blocks each include a spectral normalization convolution layer, a LeakyReLU activation function layer, and a SEBlock block; the last NormConvBlock block of the four NormConvBlock blocks includes a spectral normalization convolution layer and a LeakyReLU activation function layer; the SEBlock block is composed of an adaptive average pooling layer, a first linear layer, a RELU activation function layer, a second linear layer, and a Sigmoid activation function layer in order; all convolution layers in the discriminator are two-dimensional; Standardized image set images Through the first NormConvBlock block, through the spectral normalization convolution layer, the LeakyReLU activation function layer, and the first image feature , the first image feature The shape is adjusted to (b, c) through the view function in Python to obtain the second image feature , where b represents the batch size, c represents the number of channels of the image, and the second image feature After processing by SEBlock, the third image feature is obtained , the third image feature The fourth image feature is obtained by adjusting the view function in Python to (b, c, 1, 1) The fourth image feature Through the second NormConvBlock block, the same operation as the first NormConvBlock block is performed to obtain the fifth image feature Similarly, the fifth image feature After the third NormConvBlock block, the sixth image feature is obtained The sixth image feature After the last NormConvBlock block, the seventh image feature is obtained ; The seventh image feature The eighth image feature is obtained by sequentially passing through the adaptive average pooling layer, flattening layer, linear layer, and LeakyReLU activation function layer. ; Most eighth image features After the last linear layer, the classification result is obtained; similarly, the perturbation image The discriminator performs classification to obtain a classification result; the classification result is the original image or the perturbed image.
8. The method for actively hiding identity information against face swapping and forgery according to claim 7, characterized in that: The discriminator in step S5 adopts adversarial loss For adversarial training, the formula is as follows: , in, Express expectations, represents the discriminator; using enhanced loss To enhance the perturbation image The quality of is expressed as follows: , use Adam optimizer to train the perturbation image visual quality.
9. The method for actively hiding identity information against face swapping and forgery according to claim 8, characterized in that: Step S6 specifically includes: The perturbed image The corresponding image in the standardized image set Input them into the face-changing model to perform face-changing operation to obtain the face-changing image ; The perturbed image As the source face image, the corresponding image in the standardized image set As the target face image; the face-changing model adopts the existing face-changing model.
10. The method for actively hiding identity information against face swapping and forgery according to claim 9, characterized in that: Step S7 specifically includes: Face swap image The corresponding images in the standardized image set The ArcFace model is used to perform identity features and obtain the identity vector ; Face swap image After the ArcFace model is used to perform identity features, the face-changing identity vector is obtained. , by calculating the identity similarity, we can determine whether the identity hiding is successful. The formula is as follows: , in, Indicates identity similarity; when When the face-changing image is is fake, thus proving that the identity is hidden successfully.
Citation Information
Patent Citations
Face privacy protection method for defending deep counterfeiting
CN118194342A
Deep fake face change detection method based on robust identity perception watermark
CN118691452A
Electrocardiograph (ECG) signal enhancement method based on novel generative adversarial network (GAN)
US20240324936A1
Cited By
Deep fake face change active defense method based on face identity feature level disturbance
CN121961831A