A method and system for anonymizing recognizable human faces

Through the twin U-Net deep neural network, the problem of low recognition rate of anonymous images is solved, and the machine recognizability and visual privacy protection of anonymous images is realized, and it is suitable for a variety of face recognition scenarios.

CN115424314BActive Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210873245.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-07-18
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

The existing facial privacy protection methods cannot effectively maintain the recognition availability of images while ensuring visual privacy. Especially in real-time intelligent analysis and video surveillance scenarios, the recognition rate of anonymous images is not high.

Method used

A twin U-Net deep neural network is used to fuse the original image and the anonymous preprocessed image. The network is trained through identity information loss and image information loss functions to generate anonymous images that can be recognized by machines but cannot be recognized by human eyes.

Benefits of technology

The generated anonymous images are visually similar to anonymous preprocessed images, with high machine recognition and unrecognizable human eyes. They are suitable for face recognition in various scenarios, and the recognition rate is much higher than that of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424314B_ABST
    Figure CN115424314B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology, and particularly relates to a recognizable face anonymization processing method and system. The method includes performing anonymization processing on the original image, and fusing the original image with the image after anonymization preprocessing to obtain an anonymized image, and using the anonymized image as the image for face recognition; fusing the original image and the image after anonymization preprocessing through a deep image fusion network, and the deep image fusion network includes two twin U-Net deep neural networks, one network is used to process the original image, and the other is used to process the image after anonymization preprocessing, and the two U-Nets perform image fusion in the decoder to obtain a fused image; the present invention ensures that the processed image is visually similar to the anonymized image, and at the same time ensures that the processed image can be used for machine recognition, which not only protects the privacy of the original image, but also ensures the usability of the image, and can be used in various scenarios that require face privacy protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and particularly relates to a method and system for recognizable face anonymization processing. Background Art

[0002] In the wave of technology with the continuous strengthening of the empowerment of artificial intelligence in terms of depth and breadth, technologies such as face recognition and video surveillance have become increasingly mature, and the commercialization process has been accelerating, being implemented in various fields. At the technical level, we still need an effective means to protect the visual information privacy of the human face in the image while ensuring the normal operation of the face recognition system.

[0003] In the research field, the existing methods for protecting the privacy of face images can be classified into three categories:

[0004] 1) Methods based on traditional image processing. It mainly includes image obfuscation processing, visual masking methods, privacy information hiding, means based on probabilistic generative models, and transformation methods in different image domains, such as spatial domain transformation, frequency domain transformation, coding domain transformation, etc. This type of method usually lacks consideration of the usability of the privacy-protected image. For example, the protected image usually cannot be directly used for real-time machine analysis, or there are obvious processing traces, deformation and distortion, or visual defects, which are likely to attract extra attention from attackers.

[0005] 2) Methods based on adversarial perturbations or adversarial samples. This type of method deliberately adds subtle and imperceptible interferences (adversarial perturbations) to the input image, resulting in the face recognition model being unable to accurately identify the image attributes (such as identity, category, etc.), ensuring that unauthorized third parties cannot easily use the machine recognition model to violate people's privacy. Recently, the team of Zhu Jun from Tsinghua University proposed a targeted identity protection iterative method TIP-IM to generate adversarial identity masks that can be overlaid on face images, hiding the original identity without sacrificing visual quality, and achieving a privacy protection rate of over 95% for various advanced face recognition algorithms and commercial models. The privacy protection method based on adversarial samples can effectively limit the accurate recognition of the privacy attributes of the image by the machine without affecting the subjective perception of the visual information of the image by the human eye. Therefore, this type of method is more suitable for image and video sharing or publishing in scenarios such as social media, and is not suitable for video surveillance and other scenarios that have certain requirements for real-time intelligent analysis and need to prevent human eye peeping.

[0006] 3) Method for generating or editing based on anonymized faces. This type of method is based on generative models such as GANs to process or edit the input face image, generating an anonymized face that is visually realistic and natural but has a different identity from the original face. For example, Maximov et al. from the Technical University of Munich proposed an anonymized generation network CIAGAN that takes face key points, face background information, and target identity index vectors as inputs, ensuring that the generated face identity is between the original image and a certain target identity, and maintaining the same pose and background as the original image.

[0007] However, none of the above methods have considered the issue of the usability of anonymized image recognition. For recognizable anonymization methods, there are only a few studies. For example, the team of Cao Xiaochun from the Chinese Academy of Sciences proposed a face anonymization algorithm that retains identity information. By adaptively modifying the facial features of the face through a network, the modified face looks different from the original image visually, but the original identity can still be recognized by the face recognition system with a certain probability, retaining a certain degree of usability of the anonymized image. However, the recognition rate of this method on anonymized images is not high. Summary of the Invention

[0008] Aiming at the problem of the lack of recognition usability in the current mainstream face image privacy protection technology, the present invention proposes a recognizable face anonymization processing method and system. The method first performs anonymization preprocessing on the original image, and fuses the original image with the anonymized preprocessed image to obtain an anonymized image that can be recognized by machines and not recognizable by human eyes. Therefore, the anonymized image can be used as the image for face recognition.

[0009] Furthermore, a network with a twin structure is selected for image fusion. The network with a twin structure includes two sub-networks with exactly the same structure. Each sub-network includes a decoder and an encoder. In the decoder, the features of the images in the two sub-networks are fused with each other, and the images output by the two sub-networks are finally fused to obtain an anonymized image.

[0010] Furthermore, in order to make the anonymized image recognizable by machines and not recognizable by human eyes, a loss function is used to parameterize the fusion network. The loss function used at least includes the identity information loss between the anonymized image and the original image, and the image information loss between the anonymized image and the anonymized preprocessed image.

[0011] Furthermore, the loss function used is expressed as:

[0012]

[0013] wherein, represents the total loss function between the fused image and the input image; is the identity information loss between the anonymized image and the original image; is the image information loss between the anonymous image and the image after anonymization preprocessing; λ1 and λ2 are respectively weights.

[0014] Furthermore, the identity information loss between the anonymous image and the original image is expressed as:

[0015]

[0016] where is a typical triplet loss function, E() represents a pre-trained face recognition feature extraction model, and its output is the feature representation of the face identity (usually a one-dimensional vector with a length of 512), A represents the anchor sample, P represents the positive sample of the anchor sample, N represents the negative sample of the anchor sample, α is the distance threshold of the triplet loss, and the identity information loss is constructed by two triplets, and its purpose is to effectively support face recognition in the anonymous domain and cross-domain; I A represents an image input into the depth image fusion network, that is, the anchor sample; represents the image I A the fused image obtained by inputting into the depth image fusion network; represents the positive sample I A with the same identity as the input image I P the fused image obtained by inputting into the depth image fusion network; represents the negative sample IN A with a different identity from the input image I

[0017] Furthermore, the image information loss between the anonymous image and the image after anonymization preprocessing is expressed as:

[0018]

[0019] where is the image visual loss function; is the image L1 loss function; λ 21 and λ 22 are respectively weights.

[0020] Furthermore, the image visual loss function Expressed as:

[0021]

[0022] It is essentially a visual triple loss, indicating the perceptual similarity of images, such as Learned Perceptual Image Patch Similarity (LPIPS). Among them, I represents the original face image, represents the image after preprocessing of face image anonymization, represents the target image generated by the image fusion network, and β represents the threshold of the loss function.

[0023] Furthermore, a network with a twin structure is constructed using two U-Net type networks with the same structure. One U-Net type network is used to process the original image, and the other is used to process the image after preprocessing of anonymization. Both U-Net type networks perform feature fusion in the decoder stage.

[0024] The present invention also proposes a recognizable face anonymization processing system, including an image preprocessing module, a deep image fusion network, and a face image recognition network. The image preprocessing module performs at least one of image blurring operation, pixelation operation, face deformation operation, face swapping operation, or a combination of two or more of these operations on the input original image to obtain the image after preprocessing of anonymization; the deep image fusion network fuses the original image and the image after preprocessing of anonymization, and the face image recognition network recognizes the fused image.

[0025] The present invention constructs a twin deep image fusion network through a U-Net deep neural network. By extracting the feature information of the original image and performing multi-level fusion and embedding with the preprocessed anonymized image, it ensures that the generated image is anonymous to the human eye and recognizable to machines, effectively solving the problems of privacy protection and usability of face privacy images. The specific beneficial effects of the present invention include:

[0026] 1) The present invention has strong versatility. That is, in terms of privacy protection, the present invention can support face anonymization effects presenting different appearances and intensities (including blurring, pixelation, and face deformation, etc.); in terms of recognition usability, it can complete the recognition task only relying on a pre-trained face recognition model, that is, it can be used as an extension of the existing face recognition model to provide a privacy enhancement function;

[0027] 2) The present invention has high efficiency. It has been experimentally proven that the anonymization model proposed by this method only requires a small-scale deep neural network model to complete the anonymization task, with high efficiency;

[0028] 3) The present invention has strong usability and can support face recognition in different scenarios, including anonymous domain recognition (recognition and matching between anonymous images) and cross-domain recognition (recognition and matching between anonymous images and original images). In addition, through experimental verification, the recognition rate obtained by this method is much higher than that of the methods proposed by relevant research. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flowchart of a method for recognizable face anonymization processing according to the present invention;

[0030] Figure 2 It is a schematic diagram of a depth image fusion network in the present invention;

[0031] Figure 3 It is a schematic diagram of the effect after being processed by the method of the present invention;

[0032] Figure 4 It is a preferred embodiment for realizing face recognition according to the present invention;

[0033] Figure 5 It is another preferred embodiment for realizing face recognition according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] The present invention proposes a method for recognizable face anonymization processing, which anonymizes the original image, fuses the original image with the image after anonymization preprocessing to obtain an anonymous image, and uses the anonymous image as the image for face recognition.

[0036] Embodiment

[0037] In this embodiment, as Figure 1 , the present invention is a method for recognizable face anonymization processing, which specifically includes the following steps:

[0038] For an input face image to be processed, first perform anonymization preprocessing on it. The preprocessing means include but are not limited to image blurring, pixelation, face deformation, and face swapping;

[0039] The anonymized preprocessed image and the original face image are fed into a deep image fusion network trained based on a specific face recognition model. After being processed by the fusion network, the final anonymized image is output. The anonymized image is visually similar to the anonymized preprocessed image but hides some key information of the original image;

[0040] The anonymized image is fed into the above-mentioned face recognition model. The recognition model can identify the identity of the original face from the anonymized image. In this embodiment, the face recognition model is not limited. Various pre-trained face recognition models in the prior art can be used to supervise the training of the anonymized fusion model, and the face recognition model corresponding to the fusion model is used to identify the image processed by the present invention.

[0041] The construction process of the deep image fusion network adopted in this embodiment includes:

[0042] Step 1) A U-Net type network is constructed by connecting an encoder and a decoder composed of multiple convolutional dense blocks for constructing the image fusion network;

[0043] Step 2) A deep image fusion network with a siamese network structure is constructed by using two U-Net type networks with the same network structure but different weights. The two networks respectively receive the original image and the anonymized preprocessed image of the original image and perform feature fusion.

[0044] The deep image fusion network adopted in this embodiment includes two siamese U-Net type networks with the same structure, but the network parameters of the two networks are different. One network is used to process the original image, and the other network is used to process the anonymized preprocessed image. Feature fusion can occur at each stage of the U-Net type network. In this embodiment, an attempt is made to fuse the features of each layer by addition at the decoding stage of the decoder of the U-Net network. The sum of the outputs of each layer of the decoder is sent to the subsequent layer. This embodiment gives a Figure 2 specific implementation manner of a deep image fusion network as shown. In this embodiment, the decoder is composed of three downsampling convolutional layers, and the decoder is composed of three upsampling convolutional layers. In this embodiment, the fusion of the image occurs at the decoder stage. When fusing at the decoder, the image information from the encoder and the decoder of the other network is fused simultaneously. The feature maps output by the two networks are finally fused by addition or multiplication to obtain an anonymized image, which can be recognized by a machine for identity information but cannot be recognized by the human eye.

[0045] In this embodiment, two twin U-Net type networks with the same structure are selected to construct a depth image fusion network. In the art, other networks can also be selected to fuse two images, and the present invention does not make requirements on the specific structure of the network. In addition, any existing technology network can be used to fuse two images, and the fusion can occur in the decoder or the encoder. The present invention does not make other limitations on this.

[0046] As a preferred embodiment, in this embodiment, the depth image fusion network can use different types and intensities of anonymization preprocessing means each time it is trained to obtain a corresponding depth image fusion network.

[0047] The depth image fusion network is constrained by two loss functions, namely the identity loss function and the image loss function. The identity loss function is used to ensure that the generated image is similar to the original image in terms of identity feature representation, and the image loss function is used to ensure that the generated image is visually similar to the anonymized image. In this embodiment, two functions are selected to ensure that the generated image is visually similar to the anonymized image. The total loss function includes:

[0048]

[0049] Among them, represents the total loss function between the fused image and the input image; is the identity information loss between the anonymized image and the original image; is the image information loss between the anonymized image and the image after anonymization preprocessing; λ1 and λ2 are the weights of respectively. The identity information loss between the anonymized image and the original image is expressed as:

[0050]

[0051] Among them, is a typical triplet loss function. E() represents a pre-trained face recognition feature extraction model, and its output is the feature representation of the face identity (usually a one-dimensional vector with a length of 512). A represents the anchor sample, P represents the positive sample of the anchor sample, N represents the negative sample of the anchor sample, and α is the distance threshold of the triplet loss. The identity information loss is composed of two such triplets, and its purpose is to effectively support face recognition in the anonymized domain and cross-domain. Among them, I A represents an image input into the depth image fusion network, that is, the anchor sample; represents the fused image obtained by inputting the image I A into the depth image fusion network; IP represents the positive sample image with the same identity as the input image I A ; denotes the fused image obtained after the IP input depth image fusion network; I N denotes the negative sample image with a different identity from the input image I A denotes I N the fused image obtained by inputting the depth image into the fusion network.

[0052] In this embodiment, the image information loss between the final anonymized image and the pre - anonymized image is expressed as:

[0053]

[0054] where is the image visual loss function; is the image L1 loss function; λ 21 and λ 22 are respectively weights. The image visual loss function can be expressed by the triplet loss as:

[0055]

[0056] where I represents the original face image, denotes the image after pre - anonymization of the face image, denotes the target image generated by the image fusion network, β represents the distance threshold of the triplet loss function, denotes the function for measuring the visual similarity of two images, such as LPIPS. In addition, in addition to the perceptual loss, this embodiment also uses the image pixel - level L1 distance loss function to supervise the generation of the anonymized image.

[0057] Input the original face image and the pre - anonymized image into the depth image fusion network to generate a recognizable face anonymized image. Visually, this image is similar to the pre - anonymized image, but the existing machine vision face recognition model can still recognize the original image face identity from this image. The present invention can rely only on a pre - trained face recognition model without training the face recognition model, that is, the present invention does not need to use the pre - anonymized image and the original image to train the face recognition model and update the network parameters of the face recognition model. The existing trained model can directly recognize the pre - anonymized image of the present invention and has a good recognition accuracy rate.

[0058] ​This embodiment proposes a recognizable face anonymization processing system, including an image preprocessing module, a depth image fusion network, and a face image recognition network. The image preprocessing module performs at least one of image blurring operation, pixelation operation, face deformation operation, and face swapping operation on the input original image, or a combination of two or more of them, to obtain an image after anonymization preprocessing; the depth image fusion network fuses the original image with the image after anonymization preprocessing, and the face image recognition network recognizes the fused image; wherein the depth image fusion network includes two twin U-Net deep neural networks, one network is used to process the original image, and the other is used to process the image after anonymization preprocessing, and the outputs of the two U-Net deep neural networks are added together for fusion to obtain a fused image; each U-Net deep neural network includes an encoder and a decoder, and the input image is used to extract features through a convolutional module. The encoder includes three cascaded downsampling modules, and the extracted features are continuously downsampled three times by the encoder and then input into a convolutional module; the decoder includes three cascaded upsampling layers, and each upsampling module upsamples through the input of the previous layer and makes a skip connection with the output of the downsampling module of the corresponding size and then serves as the input of the next-level upsampling module.

[0059] The method described in this invention can be used in a privacy-friendly face recognition system. Specific implementation cases are as Figure 4 as Figure 5 shown. Among them, Figure 4 shows a case of face recognition in an anonymous domain image. Among them, first, the template image is anonymized through the method described in the solution, and the anonymized template image is used for registration operation, that is, the anonymized face image is used as the identity feature in the virtual identity feature library of the user in the APP. When face recognition is required, the current real-time face data is collected, and the collected data is anonymized, and the anonymized image is used for recognition to match the corresponding face information. In the above embodiment, the template image and the image to be recognized remain visually anonymous in the storage, display, transmission, and other links of the entire system, ensuring the visual privacy of users. Figure 5 shows another embodiment of a cross-domain face recognition system. Among them, the user may be required to use the original image for registration during the registration stage (such as the ID card photo used in the public security system), and the anonymization method proposed in this invention can also be used during the recognition stage to match the identity through the anonymized image of the image to be recognized and the original image.

[0060] This embodiment also gives the specific training process of the depth image anonymization network, which specifically includes the following steps:

[0061] 1) Dataset and preprocessing

[0062] CelebA dataset: It contains 202,599 face images from 10,177 identities and is annotated with approximately 40 face attributes such as whether wearing glasses, whether smiling, etc.; the training set of this dataset is used for training the model in this embodiment, and the test set is used for testing the model.

[0063] VGGFace2 dataset: It contains 3.31 million images from 9,131 identities, with an average of 362.6 images per person; approximately 59.7% are male; each image is also annotated with a face bounding box, 5 key points, as well as information such as age and pose; the test set of this dataset is used for testing the model in this embodiment.

[0064] LFW dataset: It contains 13,233 images from 5,749 identities, and 1,680 of them have 2 or more face images; this dataset provides a standardized face matching process for testing the model in this embodiment.

[0065] Use a pre-trained open-source face tool to detect, crop, and align the face images in the above datasets, keeping the face head in the central area of the image, and setting the image resolution to 112*112.

[0066] 2) Training of the network

[0067] Use the training set of CelebA to train the proposed deep image fusion network. Four anonymization preprocessing methods and five basic face recognition models are used in the training, and a total of 20 models are trained. Among them, the four anonymization preprocessing methods are respectively:

[0068] · Gaussian blur (Blur): The blur kernel size is fixed at 31, and blur kernel variances ranging from 2 to 8 are used in the training, and the variance is fixed at 5 in the test phase.

[0069] · Pixelation: Pixel blocks with sizes ranging from 4 to 10 are used in the training, and the pixel block size is fixed at 7 in the test.

[0070] · FaceShifter: Through the FaceShifter deep face swapping algorithm, perform face swapping operations with randomly selected other face images as the target.

[0071] · SimSwap: Through the SimSwap deep face swapping algorithm, perform face swapping operations with randomly selected other face images as the target.

[0072] Details of the five basic face recognition models are shown in Table 1.

[0073] Table 1 Details of the five basic face recognition models

[0074] Face recognition backbone model Training method Number of parameters LFW recognition accuracy MobileFaceNet ArcFace 1M 0.9863 InceptionResNet FaceNet 28M 0.9906 IResNet50 ArcFace 44M 0.9898 SEResNet50 ArcFace 44M 0.9896 IResNet100 ArcFace 65M 0.9983

[0075] The training process uses the Adam optimizer with β1 = 0.9, β2 = 0.999, and learning rate = 0.001 to optimize the training process.

[0076] This embodiment gives a schematic diagram of the effect as shown in Figure 3 Each row is a schematic picture of a face image after various processes. The first column is the original image (Original) of the image, the second column is the version of the original image preprocessed by blur, the third column is the anonymized fusion image (Blur*) corresponding to the blur. Similarly, the fourth to ninth columns respectively show the remaining anonymized preprocessed images and the final anonymized image. As shown in Figure 3 The final anonymized image is highly visually similar to the anonymized preprocessed image. Figure 3 This embodiment also conducts a quantitative verification of the proposed recognizable anonymization model through simulation experiments, aiming to test the privacy protection performance and recognition usability of the generated images. Some experiments are compared with the method proposed by Li and effectively prove the superiority of the method of the present invention.

[0077] In terms of privacy protection performance, the visual differences between the anonymized images and the original images are measured by subjective and objective methods respectively. Objectively, LPIPS and SSIM are used to measure the gap between the anonymized images and the original images. Tables 2 and 3 respectively show the objective indicators of the privacy protection performance of the method of the present invention on the CelebA and VGGFace2 datasets, while Table 4 shows the comparison of the objective privacy protection performance between the method of the present invention and the Li method. The results show that the privacy protection of the method of the present invention after pixelization and blurring operations is much higher than that of the Li method, and the privacy protection through face swapping operations is similar to that of the Li method.

[0078] Table 2 Objective indicators of the privacy protection performance of the method of the present invention on the CelebA dataset

[0079] Table 3 Objective indicators of the privacy protection performance of the method of the present invention on the VGGFace2 dataset

[0080]

[0081] Table 4 Comparison of the objective privacy protection performance between the method of the present invention and the Li method

[0082]

[0083] Table 4 Comparison of the objective privacy protection performance between the method of the present invention and the Li method

[0084]

[0085] Subjectively, in this embodiment, the commercial crowdsourcing platform Mechanical Turk provided by Amazon is used. Through an online questionnaire survey, crowdsourcing users are hired to identify anonymized images by human eye observation. The lower the recognition rate, the stronger the privacy protection performance. Table 5 shows the subjective recognition rates of different types of images. Through the anonymization process of the method of the present invention, the subjective recognition accuracy rate has been significantly reduced. In Table 5, the lower the recognition rate, the better the privacy protection effect. ★ indicates the image processed by the method described in the present invention.

[0086] Table 5 Subjective recognition rates of different types of images

[0087] Image type Accuracy rate Confidence level Original image 0.920 4.20 Blur 0.490 3.25 Blur★ 0.675 3.55 Pixelate 0.350 3.11 Pixelate★ 0.520 3.30 FaceShifter 0.510 3.32 FaceShifter★ 0.675 3.67 SimSwap 0.455 3.44 SimSwap★ 0.700 3.79

[0088] In terms of recognition usability, face matching experiments are carried out on three face image datasets, CelebA, VGGFace2 and LFW. The face recognition rate of anonymized images is used as a measure of usability. Table 6 shows the face recognition rates of the method of the present invention in two scenarios, the anonymized domain (ADR) and the cross-domain (XDR) (for the LFW dataset, the recognition rate is measured by TAR@FAR = 0.01 / 0.1). The results show that in both cases, the face images processed by the method of the present invention can still maintain a relatively high recognition rate. Table 7 compares the average recognition rates of the method of the present invention with the Li method through CelebA and VGGFace2. It can be seen that the face recognition rate of the method of the present invention is much higher than that of the Li method.

[0089] Table 6 Face recognition rates of the method of the present invention in two scenarios, the anonymized domain (ADR) and the cross-domain (XDR)

[0090]

[0091] Table 7 Comparison of the average recognition rates of the method of the present invention and the Li method through CelebA and VGGFace2

[0092]

[0093] In the above table, the MobileFaceNet method is from "VGGFace2: A Dataset for Recognising Faces across Pose and Age" published by Qiong Cao et al.; the InceptionResNet method is from "ArcFace: Additive Angular Margin Loss for Deep Face Recognition" published by Jiankang Deng et al.; IResNet50 and IResNet100 are from "PrivacyCam: A Privacy Preserving Camera Using uCLinux on the Blackfin DSP" published by Ankur Chattopadhyay et al.; the SEResNet50 method is from "SimSwap: An Efficient Framework For High Fidelity Face Swapping" published by Renwang Chen et al.

[0094] In summary, through the above simulation experiments, the feasibility of the solution in this embodiment is verified in the embodiments of the present invention. A general machine-recognizable face visual anonymization processing method provided by the embodiments of the present invention ensures that the generated images are anonymous to the human eye and recognizable by machines, effectively solving the problems of privacy protection and usability of face privacy images.

[0095] The present invention also proposes a computer device, which includes a processor and a memory. The processor is used to run a computer program stored in the memory to implement the above-mentioned recognizable face anonymization processing method.

[0096] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A recognizable face anonymization method, characterized in that, The original image is preprocessed anonymously, and the original image is fused with the anonymously preprocessed image to obtain an anonymous image. The anonymous image is visually similar to the anonymously preprocessed image and its original identity cannot be accurately recognized by human vision, while the pre-trained machine recognition model can still extract the original identity features of the face from the anonymous image and recognize the anonymous image; when obtaining the anonymous image, a network with a twin structure is constructed using two U-Net type networks with the same structure. One U-Net type network is used to process the original image, and the other is used to process the anonymously preprocessed image. At the same time, feature fusion is performed between the two U-Net type networks. Each U-Net type network includes an encoder and a decoder, and feature fusion of the images is performed between the two U-Net type networks. The images output by the two U-Net type networks are finally fused to obtain an anonymous image; in order to make the anonymous image recognizable by machines and unrecognizable by human eyes, a loss function is used to update the parameters of the network with a twin structure. The loss function used at least includes the identity information loss between the anonymous image and the original image, and the image information loss between the anonymous image and the anonymously preprocessed image. The loss function used is expressed as: Among them, represents the total loss function between the fused image and the input image; is the identity information loss between the anonymized image and the original image; is the image information loss between the anonymized image and the image after anonymization preprocessing; λ1 and λ2 are respectively weights.

2. The method for anonymizing recognizable human faces according to claim 1, characterized in that, Identity information loss between the anonymized image and the original image Expressed as: Among them, represents a typical triple loss function, A represents the anchor sample, P represents the positive sample of the anchor sample, N represents the negative sample of the anchor sample, and I A represents an image input into the depth image fusion network, that is, the anchor sample; represents the image I A input into the depth image fusion network to obtain the fused image; I P represents the positive sample image with the same identity as the input image I A ; represents the fused image obtained after I P is input into the depth image fusion network; I N represents the negative sample image with a different identity from the input image I A ; represents the fused image obtained after I N is input into the depth image fusion network.

3. The recognizable face anonymization method according to claim 1, characterized in that, Image information loss between an anonymous image and an image after anonymization preprocessing Expressed as: Among them, is the image visual perception loss function, which is used to measure the visual similarity of images; is the image L1-norm loss function, which is used to measure the pixel-level similarity of images; λ 21 and λ 22 are respectively weights.

4. The method for recognizable face anonymization processing according to claim 3, characterized in that, Image visual loss function It can be expressed by triplet loss as follows: Among them, I represents the original face image, represents the image after anonymization preprocessing of the face image, represents the target image generated by the image fusion network, and β represents the distance threshold of the triplet loss function, represents the function for measuring the visual similarity between two images.

5. A recognizable face anonymization method according to claim 3, characterized in that, Image L1 loss function Expressed as: Among them, represents the image after preprocessing of face image anonymization, represents the target image generated by the image fusion network, and ‖·,·‖1 represents the L1 distance on the pixels of two images.

6. A recognizable face anonymization processing system, characterized in that A method for recognizable face anonymization processing according to any one of claims 1 to 5, including an image preprocessing module, a depth image fusion network, and a face image recognition network. The image preprocessing module preprocesses the input original image anonymously to obtain an anonymously preprocessed image. Among them, the preprocessing methods include, but are not limited to, image blurring operations, pixelation operations, face deformation operations, or face swapping operations; the depth image fusion network fuses the original image with the anonymously preprocessed image, and the face image recognition network recognizes the fused image.