Face anonymization method, device, equipment and medium for availability preservation

By introducing attribute retention module and identity dissociation module in the face anonymization technology, combined with the local redrawing framework of the potential diffusion model, the problem of difficulty in retaining image attribute features and removing identity features in the existing technology is solved, and high-quality face anonymization is achieved.

CN119814937BActive Publication Date: 2025-05-23NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510281011.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-23
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In the process of face anonymization, it is difficult for the prior art to retain the attribute features of the image and remove the identity features at the same time, resulting in a decrease in image quality or loss of attribute features after anonymization.

Method used

Using a method based on the decoupling of attribute features and identity features, the attribute retention module and identity dissociation module are introduced, and the local redrawing framework of the potential diffusion model is combined to realize the attribute feature retention of faces and the concealment of identity features.

Benefits of technology

While ensuring the degree of face anonymization, the attribute characteristics of the image are retained to the maximum extent and ensure the anonymous image quality and usability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814937B_ABST
    Figure CN119814937B_ABST
Patent Text Reader

Abstract

The present invention proposes a method, device, equipment and medium for face anonymization with availability preservation. The method realizes the preservation and anonymization of attribute features of a face by constructing an anonymization network, including a backbone network, an attribute preservation module and an identity dissociation module; decoupling the attribute features from the identity features by using the attribute preservation module and the identity dissociation module, and training the anonymization network by optimizing the loss function of the potential diffusion model in the backbone network; retaining the attribute features of the face image to be anonymized by using the attribute preservation module, and injecting the identity features and key point images of the face to be anonymized by using the identity dissociation module to complete anonymization. The present invention completes the decoupling of the attribute features from the identity features by introducing the attribute preservation module and the identity dissociation module, thereby realizing the anonymization of the face, which not only protects the identity privacy of the face image, but also retains the availability attributes of the face image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of image generation technology, and in particular to a method, device, equipment and medium for face anonymization with usability preservation. Background Art

[0002] With the development of deep learning and artificial intelligence technology, people's lives have become more and more convenient, but at the same time, more and more problems have also emerged. Privacy information protection has been a hot scientific topic in recent years. The rapid development of multimedia technology has made personal privacy very vulnerable to infringement. The face, as an extremely recognizable personal biological feature, may cause damage to personal reputation and property security if it is not properly handled and used by others. Therefore, many governments have issued decrees to restrict the use and dissemination of facial images. Although this has obvious benefits for facial privacy protection to a certain extent, it will lead to costly side effects in some fields that rely on facial images to achieve results.

[0003] The field of face anonymization aims to change the identity information of a face in a face image, thereby achieving the purpose of face privacy protection. Traditional face anonymization methods directly eliminate the information of the face part through masking, mosaics, and direct shielding. Although this can directly and quickly achieve the effect of anonymization, it will cause the processed face image to lose most of the original attribute characteristics of the image, making the image difficult to use. How to retain the original attributes of the image while removing the facial features has become a problem to be solved.

[0004] Some works learn the data distribution of original face images based on the encoder-decoder method, hoping that the distribution of anonymous images and original face images will be consistent. However, the limitation of this method is that the backbone network has weak generation ability, cannot generate high-quality images, and is prone to produce serious artifacts between the generated face and the original background; other works guide the face identity to change while ensuring that the latent code of the control-related attributes remains unchanged through a potential and controllable generation path, thus achieving good results. However, the method based on latent code editing will cause some attributes unrelated to the identity to change (image background, lighting and other factors), and the above methods usually require a large number of loss functions to control the model for adversarial training, which makes the training process very difficult. Summary of the invention

[0005] In view of the technical problems existing in the prior art, the present invention proposes a face anonymization method, device, equipment and medium for usability retention, and specifically provides a method for face anonymization for usability retention based on the decoupling of attribute features and identity features, which can retain the available attribute features to the maximum extent while concealing the identity features to ensure the degree of face anonymization, so that the usability of the image is not affected.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] In one aspect, the present invention provides a method for face anonymization for usability preservation, comprising:

[0008] Step 110: Obtain the face image in the face image database;

[0009] Step 120: extracting a mask of a face region in the face image; extracting a face region corresponding to the mask as a reference image; extracting identity features and key point images of the face in the face image;

[0010] Step 130: construct an anonymization network for realizing face attribute feature preservation and face anonymization, wherein the anonymization network includes a backbone network based on a latent diffusion model local redrawing framework, an attribute preservation module, and an identity dissociation module;

[0011] Step 140: inputting the face image and the mask into the backbone network, inputting the reference image into the attribute retention module, and inputting the identity features and key point image of the face into the identity separation module; using the attribute retention module and the identity separation module, decoupling the attribute features and the identity features of the face image; by optimizing the loss function of the potential diffusion model in the backbone network, controlling the learning process of the parameters in the attribute retention module and the identity separation module, completing the training of the anonymization network, and obtaining the trained anonymization network;

[0012] Step 150: Obtain the face image to be anonymized, the mask, reference image and masked image of the face image to be anonymized, as well as the identity features and key point images of the face in the face image to be anonymized; input the reference image of the face image to be anonymized into the attribute retention module in the trained anonymization network to retain the attribute features of the face to be anonymized; input the face image to be anonymized and its mask into the backbone network, use the identity dissociation module in the trained anonymization network to obtain the identity features and key point images of the face to be anonymized, and then decouple the identity features, and inject the potential diffusion model through the backbone network to control the generated image to complete the face anonymization.

[0013] Specifically, the backbone network includes a potential local redrawing input module; the potential local redrawing input module is used to perform a pixel dot product operation on the mask and the face image to achieve the splicing of the background part image retained in the face image, the face image and the mask, and obtain the masked image, and then input the face image, the mask and the masked image into the backbone network together to perform an image generation operation based on the potential diffusion model local redrawing;

[0014] The potential diffusion model includes a noise adding network and a noise removing network;

[0015] The attribute preservation module adopts a controllable noise addition strategy and ReferenceNet technology to interactively control the image generation direction of the backbone network to preserve the attribute characteristics of the face; the attribute characteristics include the texture characteristics and structural characteristics of the face;

[0016] The identity dissociation module adopts a combination of IP-Adapter technology and ControlNet technology as a post-decoupler of the backbone network to achieve the decoupling of identity features so that the backbone network can hide the identity features during the image generation process and achieve anonymization.

[0017] Preferably, the backbone network uses a lightweight potential diffusion model Stable Diffusion 1.5 for local redrawing and training; in Step 150, the real face image is used as the face image to be anonymized, the key point image of the real face is used as the key point image of the face to be anonymized, and a fake face library is generated through the image generation model StyleGAN, and the kNN algorithm is used to obtain the identity features of the matching fake face as the identity features of the face to be anonymized.

[0018] On the other hand, the present invention also protects a face anonymization device for availability preservation, the device is used to implement the steps of the above method, the device comprises:

[0019] The first module: used to obtain face images in the face image database;

[0020] The second module: a mask for extracting the face area in the face image; extracting the face area corresponding to the mask as a reference image; extracting the identity features and key point images of the face in the face image;

[0021] The third module is used to build an anonymization network for realizing face attribute feature preservation and face anonymization. The anonymization network includes a backbone network based on a local redrawing framework of a potential diffusion model, an attribute preservation module, and an identity dissociation module.

[0022] The fourth module is used to input the face image and the mask into the backbone network, input the reference image into the attribute retention module, and input the identity features and key point images of the face into the identity separation module; use the attribute retention module and the identity separation module to decouple the attribute features and the identity features of the face image; and control the learning process of the parameters in the attribute retention module and the identity separation module by optimizing the loss function of the potential diffusion model in the backbone network, so as to complete the training of the anonymization network and obtain the trained anonymization network;

[0023] The fifth module is used to obtain the face image to be anonymized, the mask, reference image and masked image of the face image to be anonymized, as well as the identity features and key point images of the face in the face image to be anonymized; input the reference image of the face image to be anonymized into the attribute retention module in the trained anonymization network to retain the attribute features of the face to be anonymized; input the face image to be anonymized and its mask into the backbone network, use the identity dissociation module in the trained anonymization network to obtain the identity features and key point images of the face to be anonymized, and then decouple the identity features, and inject the potential diffusion model through the backbone network to control the generated image and complete the face anonymization.

[0024] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the aforementioned face anonymization method for availability retention are implemented.

[0025] In addition, the present invention also protects a storage medium for storing a computer program, which, when executed by a processor, implements the steps of the aforementioned face anonymization method for availability preservation.

[0026] Compared with the prior art, the present invention has the following technical effects:

[0027] (1) The attribute preservation module and identity dissociation module are introduced to decouple the attribute features and identity features of the face through interaction with the backbone network.

[0028] (2) Aiming at the design requirement of availability retention, the present invention introduces a controllable noise addition strategy to assist ReferenceNet technology to achieve the retention operation of the attribute features of the face by the attribute retention module; and introduces a combination of IP-Adapter technology and ControlNet technology to achieve the decoupling and concealment control of the identity features by the identity dissociation module.

[0029] (3) The present invention achieves face anonymization for availability preservation by locally redrawing the latent diffusion model without fine-tuning the pre-trained latent diffusion model and without adding any additional loss function. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0031] Figure 1 is a flow chart of a face anonymization method for availability preservation provided by a first embodiment of the present invention;

[0032] Figure 2 This is a flowchart of training an anonymization network in the face anonymization method for availability preservation provided by the first embodiment of the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] The problem that the present invention needs to solve is to de-identify the identity features of the face in the face image, while retaining the attribute features of the face to the greatest extent, so as to realize the face anonymization process for usability retention. Based on this, the present invention uses the local redrawing framework of the latent diffusion model as the backbone network of the anonymization network in the technical solution of the present invention, and introduces additional attribute retention modules and identity dissociation modules into the anonymization network to complete the decoupling of the attribute features and identity features of the face in the latent space.

[0035] In the first embodiment, the present invention provides a face anonymization method for availability preservation, referring to Figure 1 As shown, including:

[0036] Step 110: Obtain the face image in the face image database;

[0037] Step 120: extracting a mask of a face region in the face image; extracting a face region corresponding to the mask as a reference image; extracting identity features and key point images of the face in the face image;

[0038] Step 130: construct an anonymization network for realizing face attribute feature preservation and face anonymization, wherein the anonymization network includes a backbone network based on a latent diffusion model local redrawing framework, an attribute preservation module, and an identity dissociation module;

[0039] Step 140: inputting the face image and the mask into the backbone network, inputting the reference image into the attribute retention module, and inputting the identity features and key point image of the face into the identity separation module; using the attribute retention module and the identity separation module, decoupling the attribute features and the identity features of the face image; by optimizing the loss function of the potential diffusion model in the backbone network, controlling the learning process of the parameters in the attribute retention module and the identity separation module, completing the training of the anonymization network, and obtaining the trained anonymization network;

[0040] Step 150: Obtain the face image to be anonymized, the mask, reference image and masked image of the face image to be anonymized, as well as the identity features and key point images of the face in the face image to be anonymized; input the reference image of the face image to be anonymized into the attribute retention module in the trained anonymization network to retain the attribute features of the face to be anonymized; input the face image to be anonymized and its mask into the backbone network, use the identity dissociation module in the trained anonymization network to obtain the identity features and key point images of the face to be anonymized, and then decouple the identity features, and inject the potential diffusion model through the backbone network to control the generated image to complete the face anonymization.

[0041] Specifically, the backbone network includes a potential local redrawing input module; the potential local redrawing input module is used to perform pixel dot product operation on the mask and the face image to realize the splicing of the background part image retained in the face image, the face image and the mask, and obtain the masked image, and then the face image, the mask and the masked image are input into the backbone network together to perform the image generation operation based on the potential diffusion model local redrawing.

[0042] Among them, the Latent Diffusion Model (LDM) is an image generation model based on diffusion models. It achieves efficient and high-quality image generation by performing diffusion and denoising operations in latent space. The main mechanisms include:

[0043] The diffusion model divides the idea of ​​data generation into two stages, the forward diffusion process and the reverse denoising process. In the forward diffusion stage based on the Markov chain, starting with the original face image, Gaussian noise is gradually added to gradually "destroy" the face image data until it becomes a pure noise image that conforms to a specific distribution; in the reverse denoising stage, the model denoises the noisy image by training a noise prediction model. During the training stage, the diffusion model can use the context-free grammar (CFG) mechanism to guide the diffusion model to generate data within the conditional distribution using prior conditions such as text.

[0044] The input image of the diffusion model is transferred into the latent space through the encoder. Through the sequence autoencoder, the image can be better represented in the latent space. Then the diffusion model is trained and inferred in the latent space, which greatly improves the efficiency of the model and reduces the resource consumption of the model.

[0045] The forward diffusion process and the reverse denoising process of the diffusion model are used to obtain the denoising network and the denoising network of the potential diffusion model.

[0046] Furthermore, the attribute preservation module adopts a controllable noise addition strategy and ReferenceNet technology to interactively control the image generation direction of the backbone network to preserve the attribute characteristics of the face; the attribute characteristics include the texture characteristics and structural characteristics of the face.

[0047] The identity dissociation module adopts a combination of IP-Adapter technology and ControlNet technology as a post-decoupler of the backbone network to achieve the decoupling of identity features so that the backbone network can hide the identity features during the image generation process and achieve anonymization.

[0048] The original idea of ​​ReferenceNet technology comes from the implementation of ControlNet technology in the sd-webuiControlNet plug-in. Later, some people applied it in the field of posture generation and dressing. The basic idea of ​​this model is to splice a reference image and the input of the self-attention module of the denoising network together, and then use it as the input of the self-attention layer. In this way, the model can learn the texture features and structural features of the reference image.

[0049] ControlNet technology is used to lock the parameters of the latent diffusion model, and by copying the encoding layer of the denoising network in the model as a trainable "encoder copy" of the backbone network, the backbone network is trained to obtain the spatial information of the conditional image while retaining the image generation capability of the latent diffusion model, thereby controlling the image generation effect of the latent diffusion model. This method can achieve more refined control generation, protect the pre-trained model in the backbone network from interference from fine-tuning, and greatly reduce the training cost.

[0050] IP-Adapter technology adds additional modules to the denoising network of the latent diffusion model to learn relevant prompts, guiding the model to integrate the prompt information into the noise prediction through the cross-attention mechanism, thereby guiding the distribution of images generated by the model to be closer to the distribution described by the prompt.

[0051] In order to fully utilize the capabilities of the latent diffusion model, the present invention extracts the text prompt words of each photo through the multimodal training model CLIP as the text input of the latent diffusion model. The identity features of the face image are extracted using the pre-trained model Arcface, and the face key point extractor is used to form the face key point image, and the face identity features and key point image are used as the input of the identity dissociation module. When training the anonymization network, all gradients except the attribute retention module and the identity dissociation module are turned off to ensure the generation capability of the latent diffusion model, and the learning of each module is controlled by optimizing the original loss function of the latent diffusion model.

[0052] CLIP (Contrastive Language-Image Pre-training) is a multimodal training model proposed by OpenAI. It maps images and texts into the same feature space through contrastive learning, thereby achieving semantic alignment between images and texts. CLIP contains an image encoder and a text encoder, and features of images and texts are extracted through their respective encoders.

[0053] ArcFace is a face recognition pre-training model based on deep metric learning. By introducing angle interval loss in deep neural networks, the extracted features have better discriminability. The core is to extract the features of face images through multi-layer convolution and pooling operations, and then normalize and angularly measure the feature vectors. During training, a large number of face images and corresponding labels are used, and the network parameters are optimized through the back propagation algorithm. The face features are normalized to make the feature vectors more comparable and stable. In addition, the cosine similarity measurement method used can more accurately measure the similarity between feature vectors than the traditional Euclidean distance measurement, and can learn a more discriminative feature space, thereby improving the accuracy of face recognition.

[0054] Common facial key point extractors, including Dlib, MTCNN, PyTorch Face Landmark, etc.

[0055] Dlib is an open source machine learning library that provides face detection and key point extraction functions. It uses Python's Dlib tool to complete the recognition and reading of 68 key point positions of face images, which can be used to quickly obtain coordinate key point information during development.

[0056] MTCNN (Multi-Task Cascaded Convolutional Networks) is a deep learning model for face detection and key point location. It achieves efficient and accurate face detection and key point extraction by cascading three convolutional neural networks of different tasks.

[0057] PyTorch Face Landmark is an open source project based on deep learning, designed for fast and accurate facial key point detection. The project leverages the power of the PyTorch framework and provides a simple and easy-to-use interface that developers can easily integrate into their own applications.

[0058] The main function of the attribute preservation module proposed in the present invention is to control the image generation direction of the backbone network in an interactive manner, learn the texture features and structural features of the face in the reference image, and thus retain the attribute features of the face while ignoring its identity features. The present invention is inspired by the idea of ​​​​equalizing work and decides to use ReferenceNet to preserve the attribute features of the image. Because it is realized that ReferenceNet has an extremely strong reference effect on the image, if ReferenceNet is simply used to preserve the attributes of the face, the difference between the generated face and the original face will not be too large. How to make the module of the present invention only learn the attribute features of the face part and ignore the identity features?

[0059] Based on the previous research results on the noise adding method of the diffusion model, the present invention makes the following assumptions: taking the noise with a noise adding time step of 1000 steps as an example, the work of 1000 steps can be divided into three stages, the chaotic stage: 1000 steps-700 steps, the image in this stage can be understood as an almost complete noise form, and no information can be obtained from the visual observation point of view; 700 steps-500 steps, this stage is the stage where Gaussian noise destroys the image structure, and the approximate facial contour can be obtained by observation, and the visual or computer vision detection means cannot obtain the identity features; 500 steps-1 step, this stage is the stage where Gaussian noise destroys the image details, and you can try to obtain texture, structure and other information. Based on this conjecture, the present invention intends to use appropriate noise in the reference graph of the input attribute retention module to enable the module to have the ability to learn attribute features without ingesting too many identity features. Therefore, a controllable noise addition strategy is selected to gradually add noise to the data of the face image in the noise addition network of the latent diffusion model. The amount of noise added at each step is controlled by the scheduling algorithm, and the data is gradually converted into pure noise. By adjusting the scheduling algorithm and the sampling algorithm, fine control of the noise addition process is achieved.

[0060] InstantID is an image generation technology based on the combination of IP-Adapter and ControlNet. IP-Adapter processes image embedding information through a decoupled cross-attention mechanism, enabling it to process facial images of various postures and expressions. ControlNet is responsible for extracting the identity features of the face in the reference image and injecting them into the generated image of the face, achieving perfect injection of identity information.

[0061] The present invention guides and controls through the identity dissociation module, making the potential diffusion model in the backbone network more controllable. The idea of ​​building the identity dissociation module is inspired by InstantID. By mixing IP-Adapter technology with Controlnet technology, the extracted identity features of the face are injected into the potential diffusion model through a trainable backbone network. At the same time, the key point image of the face is used as the control image, the posture information of the face is input into the model, and then the identity information and attribute features of the face are integrated through the mean square error loss.

[0062] Preferably, in step 140, if Figure 2 As shown in Figure 1, the process of training the anonymized network includes:

[0063] Step 141: The attribute preservation module performs an attribute feature preservation operation on the face image by using a controllable noise addition strategy to assist ReferenceNet technology, so as to control the difference between the generated image and the face image to be not too large while not taking in too many identity features, including:

[0064] The controllable noise adding strategy is used to gradually add noise to the reference image input by the model in the noise adding network of the potential diffusion model. By adjusting the scheduling algorithm and the sampling algorithm, the noise adding process is finely controlled so that the facial identity features in the reference image are masked, while the features including the facial texture and contour structure including the skin color are retained.

[0065] The ReferenceNet technology is used to learn the attribute features of the face in the reference image after noise addition, but the identity features of the face are ignored, including splicing the reference image input to the ReferenceNet network and the original noisy face image input to the backbone network together through the input of the self-attention mechanism module of the denoising network, and then using them as the input of the self-attention layer in the network, so that the latent diffusion model learns the texture features and structural features of the face in the reference image;

[0066] Step 142: The identity dissociation module guides and controls the potential diffusion model by combining IP-Adapter technology with ControlNet technology, including:

[0067] Using the backbone network to inject the identity features of the face into the latent diffusion model, including combining ControlNet technology with IP-Adapter technology, so that the latent diffusion model has identity mapping capabilities, that is, by inputting identity features, it can generate a face image with corresponding identity features; and add additional control to the entire image;

[0068] A multimodal training encoder CLIP is used to encode the identity feature information extracted by the Arcface model into a semantic feature vector;

[0069] The ControlNet technology is used to increase the spatial control capability and identity control capability of the image generation results, including:

[0070] Lock the parameters of the latent diffusion model and copy these parameters to the encoding layer of the denoising network as a trainable encoder copy of the backbone network;

[0071] The key point image of the face is used as the control image to provide the geometric structure and key position information of the face, so that the face image is more in line with the specific posture and expression requirements;

[0072] Inputting posture information of the face based on the control image into the latent diffusion model, the posture information including the positions of key points of the face;

[0073] The semantic feature vector encoded by the identity feature information is input into the ControlNet network, so that the identity information is injected into the potential diffusion model;

[0074] The IP-Adapter technology is used to increase the identity control capability of image generation results, including:

[0075] By adding an additional learning prompt module to the denoising network, the latent diffusion model is guided to use the cross-attention mechanism to fuse the semantic feature vector encoded by the identity feature information into the noise prediction, so as to guide the distribution of the image generated by the latent diffusion model to be closer to the distribution described by the corresponding prompt; the prompt information includes: text prompt and image prompt;

[0076] Step 143: By integrating the above two processes, the identity features and attribute features of the face are integrated by optimizing the loss function of the potential diffusion model, and the parameters in the ReferenceNet, ControlNet technology and IP-Adapter technology are updated; the loss function is designed using mean square error, which is specifically given by the following formula:

[0077]

[0078] in, represents the time step; Represents a text prompt, which is used as a model input in the form of text to guide the image content output by the latent diffusion model; Represents the image prompt, which is used as the model input in the form of an image to guide the image content output by the latent diffusion model; Represents additional prompts, which are used to provide detailed descriptions, control the generation process, and improve generation accuracy and efficiency; represents the true noise value, represents the noise value predicted by the denoising network; when updating and training the parameters, a training strategy of closing all gradients except the attribute preservation module and the identity dissociation module is adopted to ensure the generation ability of the potential diffusion model; the gradient is the derivative of the loss function of the potential diffusion model with respect to the parameters.

[0079] Step 144: Repeat the above process of Step 141-Step 143 until the required number of execution rounds is reached to complete the training of the anonymized network.

[0080] In the second embodiment of the present invention, the controllable noise adding strategy adopts the noise adding strategy of the diffusion model DDIM, through a hyperparameter , to control the time step of the input noise. The longer the step, the stronger the added noise.

[0081] DDIM (Denoising Diffusion Implicit Models) is an improved diffusion model that can improve sample quality by reducing the number of iterations required by traditional diffusion models. The core is to redefine the diffusion process as a non-Markov process, allowing some denoising steps to be skipped, thereby accelerating the image generation process.

[0082] In the third embodiment, the potential diffusion model of the backbone network in the anonymization network uses Stable Diffusion 1.5, and according to the method steps of the first embodiment, the backbone network is partially redrawn and trained. StableDiffusion 1.5 is a lightweight potential diffusion model with low operating resource usage, strong compatibility, and is suitable for most common graphics cards. It has the characteristics of fast generation speed and can quickly respond to text descriptions and generate images.

[0083] The training process of the anonymization network completes the decoupling operation of identity features and attribute features. In order to realize the anonymization function, in Step 150 of this embodiment, the real face image is used as the face image to be anonymized, the key point image of the real face is used as the key point image of the face to be anonymized, and StyleGAN is used to generate a fake face library. Then, the identity features of the matching fake faces are obtained through the kNN algorithm as the identity features of the face to be anonymized.

[0084] The StyleGAN (Style-Based Generative Adversarial Network) is an image generation model based on the Generative Adversarial Network (GAN), proposed by NVIDIA in 2018. It significantly improves the quality and diversity of generated images by introducing the concept of "style", making the generated images more realistic and controllable.

[0085] The kNN (k-Nearest Neighbors) algorithm is a simple and effective supervised learning algorithm, mainly used for classification and regression problems. In classification problems, kNN calculates the distance between the sample to be classified and the known category samples, finds multiple known samples (i.e., "nearest neighbors") that are closest to the sample to be classified, and then determines the category of the sample to be classified based on the categories of these multiple samples. The kNN algorithm combines simplicity and flexibility, especially in classification problems with small-scale data sets or high-dimensional data.

[0086] During the specific reasoning process, the reference image of the real face, the mask, the masked image, the key point map of the real face, and the identity features of the fake face are input into the entire trained anonymization network, which retains the attribute features of the real face usability while achieving face anonymization with controllable identity features.

[0087] In the fourth embodiment of the present invention, the potential diffusion model of the backbone network in the anonymization network still uses StableDiffusion 1.5, uses the training set of the dataset FFHQ as the training data of the anonymization network, and is tested on the validation set / test set of the dataset FFHQ and the dataset CelebA-HQ.

[0088] The dataset FFHQ (Flickr-Faces-HQ) is a high-quality face image dataset created by the research team of NVIDIA in 2019. The dataset contains 70,000 high-quality PNG images with a resolution of 1024×1024, covering face images of different ages, races, and backgrounds. These images have rich diversity in age, race, image background, and face attributes.

[0089] The dataset CelebA-HQ (CelebA-High Quality) is a high-quality face image dataset proposed by NVIDIA in the 2018 ICLR paper "Progressive Growing of GANs for Improved Quality, Stability, and Variation". This dataset is an upgraded version of the CelebA dataset, containing 30,000 high-quality images with a resolution of 1024×1024. It not only provides higher-resolution images, but also retains the detailed annotation information of the original dataset, making it an ideal choice for research in fields such as face recognition, image generation, and computer vision.

[0090] The above two image data sets are used below to conduct multiple test experiments and compare the experimental results of the fourth embodiment under different test sets, which proves that the face anonymization method for availability preservation provided by the present invention can provide better image generation quality while effectively meeting the anonymization requirements of the face.

[0091] Experiment 1:

[0092] In this experiment, we select the training set of FFHQ as the training data of the anonymization network, and use the validation set of FFHQ and CelebA-HQ as the test data.

[0093] Several representative face anonymization methods in the existing technology are selected as comparative experimental objects, including:

[0094] (1) DP2: A deep learning-based face anonymization technology, the core of which is to achieve face anonymization through generative adversarial networks (GANs);

[0095] (2) Deepprivacy: A face anonymization technology based on generative adversarial networks (GANs) that can learn from large amounts of facial data and create realistic new faces. When processing the original image, it replaces facial features so that the overall visual effect of the image remains unchanged while protecting the individual’s identity.

[0096] (3) LDFA: A face anonymization technique based on a diffusion model, which achieves face anonymization through a two-stage approach. The first stage is face detection, and the second stage is to generate realistic face images using the latent diffusion model.

[0097] (4) G2Face: A face anonymization technology that uses generative priors and geometric priors to enhance identity operations and achieve high-quality reversible face anonymization. Specifically, geometric information is extracted from a 3D face model and integrated with a pre-trained GAN-based decoder to generate realistic anonymous faces with consistent geometry.

[0098] (5) RiDDLE: A face anonymization technique based on the pre-learned StyleGAN2 generator that achieves reversible and diverse de-identification by encrypting and decrypting facial identities in the latent space. Its encryption process is password-guided, allowing diverse anonymization using different passwords.

[0099] (6) FALCO: A method for anonymous face image generation and recognition based on identity preservation. By combining face anonymization and recognition, the identity information in face images is protected. At the same time, the anonymous images can still be used for face recognition tasks.

[0100] The quality indicator of the evaluation image is FID, which is an indicator used to evaluate the difference between the generated model and the real data distribution. The generation quality is measured by calculating the Fréchet distance between the generated data distribution and the real data distribution. The lower the FID, the smaller the direct distance between the generated data distribution and the real data distribution, and the better the generation quality.

[0101] Table 1 Comparison of image quality obtained from experiments using different face anonymization methods

[0102]

[0103] As shown in Table 1, the anonymization method provided by the present invention achieves the lowest FID value compared with other methods, which also means that the data distribution of the image generated by the method of the present invention is closest to the data distribution of the original face image (bold indicates optimal, and underline indicates suboptimal).

[0104] Experiment 2:

[0105] This experiment changes the test data and conducts tests on the test set of FFHQ and CelebA-HQ. This experiment illustrates that the method of the present invention can achieve effective anonymization for face images.

[0106] The evaluation indicator of this experiment is the identity re-identification rate (ReID).

[0107] The identity re-identification rate ReID is the ratio of the anonymized face image data set that can be recognized by the face detector. The anonymous face image and the real face image are mapped to the feature space through the feature extractor of the pre-trained model Arcface, and the feature closest to the feature distance in the original face image is calculated by calculating the cosine similarity or L2 distance (Euclidean distance given by the L2 norm). Then the successful matching rate is calculated. Cos_R1 represents the successful matching rate obtained by calculating the cosine similarity, and L2_R1 represents the successful matching rate obtained by calculating the L2 distance. The lower the successful matching rate, the better the anonymity effect.

[0108] Table 2 Comparison of anonymization effects of different face anonymization methods

[0109]

[0110] The test results of this experiment are shown in Table 2, which shows that the method of the present invention is second only to FALCO and RiDDLE among many methods, and is higher than the effects of other methods. FALCO and RiDDLE, two anonymization methods, are based on latent space coding. They control the face identity information through a loss function to make it as far away from the original identity as possible, so their methods have excellent effects in anonymization. However, this method of controlling anonymization through a loss function will cause exaggerated deformation of certain specific parts of the face image. RiDDLE does not maintain good facial attributes, while FALCO has a more exaggerated deformation in the eye area of ​​the image, and the method based on latent space coding modifies attributes unrelated to the face identity, such as background environment, lighting, character hairstyle, etc. (bold indicates optimal, underline indicates suboptimal).

[0111] Experiment 3:

[0112] This experiment is still tested on the test set of FFHQ and CelebA-HQ. This experiment proves that the method of the present invention can effectively retain the utility of attribute features for face images.

[0113] The indicators for evaluating image quality are face detection rate, face posture L2 distance, expression 3D feature L2 distance, and classifier classification accuracy.

[0114] The experimental results of this experiment are shown in Table 3 and Table 4, indicating that the face detection rate of the images generated by the method of the present invention on both data sets reached 100%. In terms of availability retention, the present invention specifically divides the attributes into semantically weakly related attributes and semantically strongly related attributes in the experiment. Semantically weakly related attributes refer to the geometric constraints of the face in the image, such as the posture vector of the face, three-dimensional expression features, etc.; semantically strongly related attributes refer to the intuitive visual experience of a face image, such as the age, gender, and facial makeup of the person in the image, etc. (bold indicates the best, underlined indicates the second best).

[0115] Table 3 Quantitative experimental results of face detection rate and preservation of weakly associated semantic attributes

[0116]

[0117] For semantically weakly related attributes, the present invention uses a posture feature extractor and a facial three-dimensional expression feature extractor to extract the relevant features of the original face image and the anonymized generated image, and calculates their Euclidean distance. Compared with other methods, the present invention achieves the best in most cases, and the rest are suboptimal. Although G2Face can retain the relevant features of facial geometric constraints, a large number of artifacts will be generated in its generated image. This fully proves the ability of the present invention to retain the features of the semantically weakly related attributes of the face.

[0118] Table 4 Quantitative experimental results of retaining semantically strongly related attributes

[0119]

[0120] For the strongly semantically related attributes, this experiment used a pre-trained classifier to test its attribute retention ability. The results showed that the method of the present invention achieved the best results in terms of age and gender, and achieved the second best results in terms of makeup retention ability. The comparison of the above two types of attributes fully proves the attribute feature retention ability of the method of the present invention.

[0121] Experiment 4

[0122] This experiment conducted an ablation experiment on the FFHQ test set to demonstrate the effectiveness of each module and hyperparameter in each anonymization network. The effectiveness of the experimental results is shown in Tables 5 and 6 (bold indicates the best, underline indicates the suboptimal).

[0123] This experiment removes the attribute retention module in the method of the present invention from the anonymization network, and trains according to the same training configuration as the previous experiment. In the inference stage, this experiment inputs the same identity features and facial feature points into the identity dissociation module. Without the control of the attribute features, the identity dissociation module will directly act on the original face image. Although it changes the identity features of the original face image, it will cause the internal face to lose most of its attribute features. When there is no identity dissociation module, the anonymized part of the face is completely generated by the model, and the generated face is close to the original face, and the anonymization effect cannot be achieved, which proves the necessity of the identity dissociation module.

[0124] This experiment also uses hyperparameters Control the noise level of the reference image and observe its generation effect. It can be seen that as the noise intensity increases, the degree of attribute feature retention gradually decreases, while the degree of anonymization gradually increases.

[0125] Table 5 Ablation experiment results (image quality, detection rate and anonymization degree)

[0126]

[0127] Among them, w. / o. refer means not using the attribute preservation module, and w. / o. instant means not using the identity dissociation module. This experiment compares the results with different hyperparameters. The image generation quality under the condition of taking value is balanced, and a trade-off is made between anonymization and attribute preservation. 300 was chosen as the final value for this experiment.

[0128] Table 6 Ablation experiment results (attribute preservation degree)

[0129]

[0130] The above four groups of experiments show that the face anonymization method for usability preservation proposed in this invention uses the decoupling of attribute features and identity features to achieve face identity anonymization and attribute feature retention. Compared with other methods, while effectively anonymizing, it retains relevant attribute features as much as possible, thus achieving an effective balance between utility preservation and identity anonymity.

[0131] In addition, in a fifth embodiment, the present invention further provides a face anonymization device for availability retention, the device is used to implement the steps of the aforementioned face anonymization method for availability retention, and the device includes:

[0132] The first module: used to obtain face images in the face image database;

[0133] The second module: a mask for extracting the face area in the face image; extracting the face area corresponding to the mask as a reference image; extracting the identity features and key point images of the face in the face image;

[0134] The third module is used to build an anonymization network for realizing face attribute feature preservation and face anonymization. The anonymization network includes a backbone network based on a local redrawing framework of a potential diffusion model, an attribute preservation module, and an identity dissociation module.

[0135] The fourth module is used to input the face image and the mask into the backbone network, input the reference image into the attribute retention module, and input the identity features and key point images of the face into the identity separation module; use the attribute retention module and the identity separation module to decouple the attribute features and the identity features of the face image; and control the learning process of the parameters in the attribute retention module and the identity separation module by optimizing the loss function of the potential diffusion model in the backbone network, so as to complete the training of the anonymization network and obtain the trained anonymization network;

[0136] Fifth module: used to obtain the face image to be anonymized, the mask of the face image to be anonymized, the reference image and the masked image, as well as the identity features and key point images of the face in the face image to be anonymized; input the reference image of the face image to be anonymized into the attribute retention module in the trained anonymization network to retain the attribute features of the face to be anonymized; input the face image to be anonymized and its mask into the backbone network, use the identity dissociation module in the trained anonymization network to obtain the identity features and key point images of the face to be anonymized, then decouple the identity features, and inject them into the latent diffusion model through the backbone network to control the generation of images and complete face anonymization.

[0137] In another embodiment of the present invention, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the face anonymization method for usability retention provided in any of the above embodiments. This computer device can be a server. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample data. The network interface of the computer device is used to communicate with an external terminal through a network connection.

[0138] In another embodiment, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the face anonymization method for usability retention provided in any of the above embodiments.

[0139] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0140] Matters not covered by the present invention are known technologies.

[0141] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0142] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

[0143] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A face anonymization method for availability preservation, characterized in that: include: Step 110: Obtain the face image in the face image database; Step 120: extracting a mask of a face region in the face image; extracting a face region corresponding to the mask as a reference image; extracting identity features and key point images of the face in the face image; Step 130: construct an anonymization network for realizing face attribute feature preservation and face anonymization, wherein the anonymization network includes a backbone network based on a latent diffusion model local redrawing framework, an attribute preservation module, and an identity dissociation module; Step 140: inputting the face image and the mask into the backbone network, inputting the reference image into the attribute retention module, and inputting the identity features and key point image of the face into the identity separation module; using the attribute retention module and the identity separation module, decoupling the attribute features and the identity features of the face image; by optimizing the loss function of the potential diffusion model in the backbone network, controlling the learning process of the parameters in the attribute retention module and the identity separation module, completing the training of the anonymization network, and obtaining the trained anonymization network; Step 150: Obtain the face image to be anonymized, the mask, reference image and masked image of the face image to be anonymized, and the identity features and key point images of the face in the face image to be anonymized; input the reference image of the face image to be anonymized into the attribute preservation module in the trained anonymization network to preserve the attribute features of the face to be anonymized; The face image to be anonymized and its mask are input into the backbone network. The identity features and key point images of the face to be anonymized are obtained by using the identity dissociation module in the trained anonymization network, and then the identity features are decoupled. The latent diffusion model is injected into the backbone network to control the generated image and complete the face anonymization.

2. The face anonymization method for availability preservation according to claim 1, characterized in that: In Step 130, the backbone network includes a potential local redrawing input module; the potential local redrawing input module is used to perform a pixel dot product operation on the mask and the face image to achieve the splicing of the background part image retained in the face image, the face image and the mask, and obtain the masked image, and then input the face image, the mask and the masked image into the backbone network together to perform an image generation operation based on the potential diffusion model local redrawing; The potential diffusion model includes a noise adding network and a noise removing network; The attribute preservation module adopts a controllable noise addition strategy and ReferenceNet technology to interactively control the image generation direction of the backbone network to preserve the attribute characteristics of the face; the attribute characteristics include the texture characteristics and structural characteristics of the face; The identity dissociation module adopts a combination of IP-Adapter technology and ControlNet technology as a post-decoupler of the backbone network to achieve the decoupling of identity features so that the backbone network can hide the identity features during the image generation process and achieve anonymization.

3. The face anonymization method for availability preservation according to claim 2, characterized in that: In Step 140, the process of training the anonymized network includes: Step 141: Using the attribute preservation module, the attribute features of the face image are preserved by using a controllable noise strategy to assist the ReferenceNet technology, so as to control the difference in usability attributes between the generated image and the face image; Step 142: Using the identity dissociation module, the potential diffusion model is guided and controlled by combining the IP-Adapter technology with the ControlNet technology to decouple and control the identity characteristics; Step 143: after fusing the two processes of Step 141 and Step 142, fusing the identity features and attribute features of the face, the parameters in the ReferenceNet, ControlNet technology and IP-Adapter technology are updated by optimizing the loss function of the potential diffusion model; Step 144: Repeat the above process of Step 141-Step 143 until the required number of execution rounds is reached to complete the training of the anonymized network.

4. The face anonymization method for availability preservation according to claim 3, characterized in that: In Step 141, the attribute preservation module performs the process of preserving the attribute features of the face image by using a controllable noise addition strategy to assist the ReferenceNet technology, including: The controllable noise adding strategy is used to gradually add noise to the reference image input by the attribute preservation module in the noise adding network of the potential diffusion model, and the fine control of the noise adding process is achieved by adjusting the scheduling algorithm and the sampling algorithm, so that the identity features of the face in the reference image are masked, while the attribute features of the face are preserved; The ReferenceNet technology is used to learn the attribute features of the face in the reference image after noise addition, but the identity features of the face are ignored, including splicing the reference image input to ReferenceNet and the original noisy face image input to the backbone network together through the input of the self-attention mechanism module of the denoising network, and then using them as the input of the self-attention layer in the backbone network, so that the latent diffusion model learns the texture features and structural features of the face in the reference image.

5. The face anonymization method for availability preservation according to claim 4, characterized in that: In step 142, the process of decoupling and controlling identity features using the identity dissociation module includes: Using the backbone network to inject the identity features of the face into the latent diffusion model, including combining ControlNet technology with IP-Adapter technology, so that the latent diffusion model has identity mapping capabilities, and can generate a face image with the corresponding identity features by inputting identity features, and add additional identity feature controls to the entire image; A multimodal training encoder CLIP is used to encode the identity feature information extracted by the Arcface model into a semantic feature vector; The ControlNet technology is used to increase the spatial control capability and identity control capability of the image generation results, including: Lock the parameters of the latent diffusion model and copy them to the encoding layer of the denoising network as a trainable encoder copy of the backbone network; The key point image of the face is used as the control image to provide the geometric structure and key position information of the face, so that the face image is more in line with the specific posture and expression requirements; Inputting posture information of the face based on the control image into the latent diffusion model, the posture information including the positions of key points of the face; The semantic feature vector encoded by the identity feature information is input into the ControlNet network, so that the identity feature information is injected into the potential diffusion model; The IP-Adapter technology is used to increase the identity control capability of image generation results, including: By adding an additional learning prompt module to the denoising network, the latent diffusion model is guided to use the cross-attention mechanism to fuse the semantic feature vector encoded by the identity feature information into the noise prediction, so as to guide the distribution of the image generated by the latent diffusion model to be closer to the distribution described by the corresponding prompt; the prompt information includes text prompts and image prompts.

6. The method for face anonymization for availability preservation according to claim 5, characterized in that: In Step 143, the loss function of the potential diffusion model uses the mean square error to construct the loss function, which is given by the following formula: in, represents the time step; Represents text prompts, which are used as model input in the form of text to guide the image content output by the latent diffusion model; Represents image prompts, which are used as model input in the form of images and are used to guide the image content output by the latent diffusion model; Represents additional hints used to provide descriptions, control the generation process, and improve generation accuracy and efficiency; represents the true noise value, represents the noise value predicted by the denoising network; when updating and training the parameters, a training strategy of closing all gradients except the attribute preservation module and the identity dissociation module is adopted to ensure the generation ability of the potential diffusion model; the gradient is the derivative of the loss function of the potential diffusion model with respect to the parameters.

7. The method for face anonymization for availability preservation according to claim 6, characterized in that: The backbone network uses a lightweight potential diffusion model Stable Diffusion 1.5 for local redrawing and training; in Step 150, the real face image is used as the face image to be anonymized, the key point image of the real face is used as the key point image of the face to be anonymized, and a fake face library is generated through the image generation model StyleGAN. The kNN algorithm is used to obtain the identity features of the matched fake faces as the identity features of the face to be anonymized.

8. A face anonymization device for availability preservation, characterized in that: The device is used to implement the steps of the method according to claim 1, and the device comprises: The first module: used to obtain face images in the face image database; The second module: a mask for extracting the face area in the face image; extracting the face area corresponding to the mask as a reference image; extracting the identity features and key point images of the face in the face image; The third module is used to build an anonymization network for realizing face attribute feature preservation and face anonymization. The anonymization network includes a backbone network based on a local redrawing framework of a potential diffusion model, an attribute preservation module, and an identity dissociation module. The fourth module is used to input the face image and the mask into the backbone network, input the reference image into the attribute retention module, and input the identity features and key point images of the face into the identity separation module; use the attribute retention module and the identity separation module to decouple the attribute features and the identity features of the face image; and control the learning process of the parameters in the attribute retention module and the identity separation module by optimizing the loss function of the potential diffusion model in the backbone network, so as to complete the training of the anonymization network and obtain the trained anonymization network; The fifth module is used to obtain the face image to be anonymized, the mask, reference image and masked image of the face image to be anonymized, as well as the identity features and key point images of the face in the face image to be anonymized; input the reference image of the face image to be anonymized into the attribute retention module in the trained anonymization network to retain the attribute features of the face to be anonymized; input the face image to be anonymized and its mask into the backbone network, use the identity dissociation module in the trained anonymization network to obtain the identity features and key point images of the face to be anonymized, and then decouple the identity features, and inject the potential diffusion model through the backbone network to control the generated image and complete the face anonymization.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the face anonymization method for availability preservation according to any one of claims 1 to 7 are implemented.

10. A storage medium for storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the face anonymization method for usability preservation described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Face image recognition removal and restoration system and method

    CN113935915A

  • Face anonymization method based on multi-condition diffusion model, storage medium and equipment

    CN118658188A