Face recognition method and device, electronic equipment and storage medium

The feature representation of synthetic face images is generated by generating the model and using feature representation differences for face recognition, which solves the problem of privacy leakage in face recognition, and realizes effective privacy protection and anti-reconstruction capabilities while ensuring the accuracy of recognition.

CN120375436APending Publication Date: 2025-07-25TENCENT TECH SHANGHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410096230.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing facial recognition technology poses a risk of privacy leakage when uploading facial images, making it difficult to effectively protect privacy while ensuring the accuracy of recognition.

Method used

The feature representation of the synthetic face image of the object to be identified is generated by the generation model, and facial recognition is performed based on the difference in the feature representation between the original face image and the synthetic face image. Most of the visual information is removed by using artificial intelligence technology, sufficient recognition information is retained, and anti-reconstruction ability is increased through random channel sequential transformation.

Benefits of technology

It realizes that while not affecting facial recognition, it reduces the risk of privacy leakage, and enhances the ability to resist reconstruction attacks, achieving effective privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375436A_ABST
    Figure CN120375436A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a face recognition method and device, electronic equipment and a storage medium, and relates to the fields of artificial intelligence, identity authentication, privacy security and the like. The method comprises the steps of obtaining an original face image of a to-be-recognized object, determining a first feature representation of the original face image, and generating a second feature representation of a synthetic face image corresponding to the to-be-recognized object by adopting a trained generation model based on the first feature representation. And obtaining a face recognition result of the to-be-recognized object based on a feature representation difference between the first feature representation and the second feature representation. Due to the fact that the visual information contained in the original face image and the visual information contained in the synthesized face image are similar, most visual information can be removed through the feature representation difference, determined based on the artificial intelligence technology, of the original face image and the synthesized face image, sufficient recognition information is reserved, and the purpose of face recognition is achieved. Based on the method, privacy protection is carried out on the face image while face recognition is not affected, and the risk of privacy disclosure is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and may involve fields such as artificial intelligence, identity authentication, privacy security, etc. Specifically, this application relates to a face recognition method, device, electronic device, and storage medium. Background Art

[0002] Face Recognition (FR) is a biometric method for identifying an individual's identity through a face image, and has been widely applied in scenarios such as payment, access control, and travel.

[0003] In order to overcome the resource limitations of local devices and improve the recognition accuracy, current face recognition technologies usually upload the face image to the server for feature extraction after the face image is collected by the terminal, and match the extracted face features with each face feature in the database for identity recognition.

[0004] Since face images are recognized as sensitive biometric information, it is very important to protect the privacy of face images during the face recognition process to reduce the risk of privacy leakage in the upload, processing, and other processes. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a face recognition method, device, electronic device, and storage medium that can effectively reduce the risk of privacy leakage. To achieve this purpose, the technical solutions provided by the embodiments of this application are as follows:

[0006] On the one hand, the embodiments of this application provide a face recognition method, which includes:

[0007] Obtain the original face image of the object to be recognized;

[0008] Determine the first feature representation of the original face image;

[0009] Based on the first feature representation, use the trained generation model to generate the second feature representation of the synthetic face image corresponding to the object to be recognized;

[0010] Determine the feature representation difference between the first feature representation and the second feature representation;

[0011] Based on the feature representation difference, obtain the face recognition result of the object to be recognized.

[0012] On the other hand, the embodiments of this application also provide a face recognition device, which includes:

[0013] An original image acquisition module, configured to obtain the original face image of the object to be recognized;

[0014] The first feature determination module is configured to determine a first feature representation of the original face image;

[0015] The second feature determination module is configured to generate a second feature representation of the synthesized face image corresponding to the object to be recognized by using the trained generation model based on the first feature representation;

[0016] The feature difference determination module is configured to determine the feature representation difference between the first feature representation and the second feature representation;

[0017] The face recognition module is configured to obtain a face recognition result of the object to be recognized based on the feature representation difference.

[0018] Optionally, both the first feature representation and the second feature representation include feature representations of a first number of channels;

[0019] The feature difference determination module can be configured to:

[0020] Determine the representation difference between the feature representations of each channel in the first feature representation and the second feature representation;

[0021] Based on the representation differences of each channel, determine the feature representation difference between the first feature representation and the second feature representation;

[0022] The face recognition module can be configured to:

[0023] Perform a channel order transformation on the representation differences of at least two channels in the feature representation difference to obtain a transformed feature representation difference;

[0024] Perform face feature recognition on the object to be recognized based on the transformed feature representation difference to obtain a face recognition result.

[0025] Optionally, the face recognition module can be configured to:

[0026] Generate a first sequence based on the number of channels of the feature representation difference; wherein, the number of elements in the first sequence is equal to the number of channels, and the position of each element in the first sequence corresponds to each channel in the feature representation difference;

[0027] Randomly transform the positions of at least two elements in the first sequence to obtain a second sequence;

[0028] Perform a sequence transformation on the representation differences of the channels in the feature representation difference according to the second sequence.

[0029] Optionally, the first feature determination module can be configured to:

[0030] Perform a transformation from the spatial domain to the frequency domain on the original face image to obtain a first feature representation of the original face image; wherein, the feature representation of one spatial domain channel of the original face image is transformed into the feature representations of multiple frequency domain channels.

[0031] Optionally, the face recognition module can be used to:

[0032] Perform a transformation from the frequency domain to the spatial domain on the feature representation difference to obtain a corresponding private face image of the feature representation difference;

[0033] Based on the private face image and the known face images in the server, obtain the face recognition result of the object to be recognized, wherein the known face images are obtained from the real face images of known objects, and the method for obtaining the known face images is the same as the method for obtaining the private face image.

[0034] Optionally, the face recognition device further includes a model training module, and the model training module can be used to:

[0035] Obtain multiple first samples with labels, each first sample includes the original face image of a sample object, and the label of the first sample is the real identity information of the sample object;

[0036] For each first sample, determine a third feature representation of the first sample;

[0037] Perform a training operation on the generative model to be trained based on each third feature representation until a first training end condition is met to obtain a trained generative model, and the training operation includes:

[0038] Input the third feature representations corresponding to each first sample into the generative model to be trained, and respectively obtain the fourth feature representations of the synthetic face images corresponding to each sample object through the generative model;

[0039] Respectively determine the predicted feature representation differences between the third feature representations and the fourth feature representations corresponding to each first sample;

[0040] Based on the predicted feature representation differences of each first sample, extract the predicted face features of each sample object, and the predicted face features characterize the identity information of the predicted sample object; and based on the fourth feature representations of each sample object, generate the synthetic face images corresponding to each sample object;

[0041] Based on the differences between the predicted face features corresponding to each first sample and the labels of each first sample, and the differences between the original face images and the synthetic face images of each sample object, determine the first training loss;

[0042] Adjust the model parameters in the generation model based on the first training loss.

[0043] Optionally, the model training module may be configured to:

[0044] Input the difference in each predicted feature representation into the auxiliary recognition model to be trained to obtain the predicted face features of each sample object;

[0045] For each sample object, perform a conversion from the frequency domain to the spatial domain on the fourth feature representation of the sample object to obtain the synthetic face image corresponding to the sample object;

[0046] The model training module may be configured to:

[0047] Adjust the model parameters in the generation model and the model parameters in the auxiliary recognition model based on the first training loss.

[0048] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the method provided in any optional embodiment of the present application.

[0049] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.

[0050] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.

[0051] The beneficial effects brought by the technical solution provided by the embodiment of the present application are as follows:

[0052] The face recognition method provided by the embodiment of the present application generates a feature representation of a synthetic face image corresponding to an object to be recognized through artificial intelligence technology, and performs face recognition based on the difference in feature representation between the original face image and the synthetic face image. Since the visual information contained in the original face image and the synthetic face image is similar, the difference in feature representation determined based on artificial intelligence technology can remove most of the visual information and retain sufficient recognition information to achieve the purpose of face recognition. Based on this method, while not affecting face recognition, privacy protection is performed on the face image, reducing the risk of privacy leakage.

[0053] Moreover, by further performing a transformation of the random channel order on the difference in feature representation, random material texture features are added to the image, which helps to better resist reconstruction attacks, enhances the anti-reconstruction ability, and realizes privacy protection. Brief Description of the Drawings

[0054] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0055] Figure 1 It is a schematic flowchart of a face recognition method provided by an embodiment of the present application;

[0056] Figure 2 It is a schematic flowchart of performing discrete cosine transform on the original face image provided by an embodiment of the present application;

[0057] Figure 3 It is a schematic flowchart of training a generation model provided by an embodiment of the present application;

[0058] Figure 4 It is a comparison schematic diagram of the visual information differences included in different images provided by an embodiment of the present application;

[0059] Figure 5 It is a schematic diagram of face recognition provided by an embodiment of the present application;

[0060] Figure 6 It is a schematic structural diagram of a payment system in a face-swiping payment scenario provided by an embodiment of the present application;

[0061] Figure 7 It is a schematic structural diagram of an identity verification system in a travel scenario provided by an embodiment of the present application;

[0062] Figure 8 It is a schematic structural diagram of a face recognition system provided by an embodiment of the present application;

[0063] Figure 9 It is a comparison schematic diagram of the privacy protection effects of the solution of the present application and other privacy-preserving face protection methods;

[0064] Figure 10 It is a schematic structural diagram of a face recognition device provided by an embodiment of the present application;

[0065] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0066] The following describes the embodiments of the present application with reference to the drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0067] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present invention. It should be understood that when we say an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items can refer to one, multiple or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, and it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, and A3.

[0068] In order to avoid the leakage of face images during the face recognition process, it is necessary to perform privacy protection on face images. The privacy protection of face images is mainly achieved from two aspects: on the one hand, the face visual information in the face image is hidden, such as represented in other forms (high-dimensional features, etc.), so that even if the hidden face image is obtained, the appearance information of the face cannot be viewed. On the other hand, further anti-attack designs (such as randomization, noise, perturbation, etc.) are adopted to make it difficult to reconstruct and restore the visual information of the image.

[0069] The embodiments of the present application provide a face recognition method, device, electronic device and storage medium. Based on the above two aspects, the method generates a feature representation of the synthetic face image corresponding to the object to be recognized through artificial intelligence technology, and performs face recognition based on the feature representation difference between the original face image and the synthetic face image. Since the visual information contained in the original face image and the synthetic face image is similar, the feature representation difference determined based on artificial intelligence technology can remove most of the visual information and retain sufficient recognition information to achieve the purpose of face recognition. Based on this method, while not affecting face recognition, privacy protection is performed on the face image, reducing the risk of privacy leakage.

[0070] Moreover, by further randomly permuting the channel order of the feature representation differences, this application adds random material texture features to the image, which helps better resist reconstruction attacks, enhances the anti-reconstruction ability, and achieves privacy protection.

[0071] Among them, the method provided in the embodiments of this application may involve Artificial Intelligence (AI) technology and can be implemented based on AI technology. For example, the second feature representation of the synthetic face image corresponding to the object to be recognized can be generated by a trained generation model. Among them, the trained generation model can be obtained by training based on a training sample set in a Machine Learning (ML) manner.

[0072] Artificial Intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0073] Artificial Intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of Artificial Intelligence generally include technologies such as sensors, dedicated Artificial Intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of Artificial Intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0074] Among them, computer vision technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to machine vision that uses cameras and computers to replace human eyes to identify, detect, and measure targets, and further performs graphic processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (Three Dimensions) technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0075] Machine learning (ML) is an interdisciplinary subject that involves multiple fields and disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge and skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental path to endowing computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0076] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, and intelligent transportation. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0077] Optionally, the solution provided in the embodiments of the present application may involve cloud technology. For example, the solution in the embodiments of the present application can be executed by a server or a user terminal. Among them, the server can be a cloud server. The data processing involved in the implementation process of this solution can be based on cloud technology, and the data storage involved in the implementation process can use cloud storage. For example, the matching of face features can be implemented using cloud technology. The training data sets used when training the generation model and the face recognition model can be data sets stored in the cloud server, and the face features of each object in the feature library can be stored in the cloud server.

[0078] Among them, cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. Cloud storage is a new concept extended and developed from the cloud computing concept. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world.

[0079] It should be noted that in the optional embodiments of the present application, for relevant data such as object information (such as the face image of a user), when the embodiments in the present application are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments in the present application involve data related to an object, it needs to be obtained under the authorization and consent of the object, the authorization and consent of the relevant department, and compliance with the relevant laws, regulations, and standards of the country and region. In the embodiments, if personal information is involved, the acquisition of all personal information needs to obtain the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the authorization and consent of the object.

[0080] Next, through the description of several embodiments, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described. It should be pointed out that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0081] Figure 1 The flowchart of a face recognition method provided by an embodiment of the present application is shown. This method can be executed by any electronic device, such as a user terminal or a server, or can be implemented by multiple electronic devices in cooperation.

[0082] Among them, the above-mentioned server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server (which can be called the cloud) providing cloud computing services. The terminal (which can also be called a user terminal or user device) can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0083] As Figure 1 shown, the face recognition method provided by the embodiments of this application may include the following steps S110 to step S150.

[0084] Step S110: Obtain the original face image of the object to be recognized.

[0085] Step S120: Determine the first feature representation of the original face image.

[0086] Among them, the object to be recognized can be a user who needs to perform face recognition in any face recognition scenario. The original face image refers to the face image of the object to be recognized that is truly collected. The first feature representation can be the feature representation of the original face image mapped in a high-dimensional space.

[0087] The face recognition method provided by the embodiments of this application can be applied to any face recognition scenarios such as payment, travel, access control recognition, etc. Since the above scenarios usually have relatively high security requirements, it is necessary to collect the current face image of the object to be recognized in real time as the original face image.

[0088] In the embodiments of this application, after the original face image is obtained, the original face image can be subjected to dimensionality increase mapping to obtain the first feature representation of the original face image. Among them, when performing dimensionality increase mapping, any mapping that satisfies reversibility and additive homomorphism can be used. As an optional implementation manner, the dimensionality increase mapping can be performed by means of time-frequency transformation, that is, the original face image is converted from the spatial domain to the frequency domain.

[0089] When performing feature extraction through time-frequency transformation, the original face image can be divided into several image blocks according to a preset segmentation unit. For each image block in the original face image, a transformation from the spatial domain (spatial domain) to the frequency domain (frequency domain) is performed on the image block to obtain the feature representation of each channel corresponding to the image block in the frequency domain space. By combining the feature representations of each channel corresponding to each sub-image, the first feature representation of the original face image is obtained. Among them, when performing the transformation from the spatial domain to the frequency domain, any transformation method such as the Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), or Discrete Wavelet Transform (DWT) can be used to transform each image from the spatial domain to the frequency domain.

[0090] Since the obtained original face image is usually an RGB image (composed of three color channels (spatial domain channels) of Red (R), Green (G), and Blue (B) superimposed), in order to facilitate subsequent image transmission or processing, before performing time-frequency transformation, the original face image can be converted into a YCbCr image (a color space represented by three color channels of luminance, red, and blue, where Y represents luminance, Cb represents blue chrominance, and Cr represents red chrominance), and then time-frequency transformation is performed on each color channel respectively.

[0091] In order to keep the size of the original face image and the first feature representation after time-frequency transformation consistent, based on the preset segmentation unit, the original face image can be upsampled first to obtain the first image, and then the first image is segmented according to the preset segmentation unit to obtain several image blocks (also known as pixel blocks, macro blocks, etc.). For each image block of the first image, a Discrete Cosine Transform (DCT) is performed on the image block to obtain the feature representation of each channel corresponding to the image block.

[0092] Among them, when performing the transformation from the spatial domain to the frequency domain (DCT transformation), for each color channel (Y, Cb, Cr) of the original face image, the image region (N*N) of each preset segmentation unit in the color channel is converted into a (N×N)-dimensional feature representation.

[0093] Exemplarily, such as Figure 2As shown, assume that the image size of the original face image is H*W*3, where H represents the image height, W represents the image width, and 3 represents the number of image channels (the three channels of Y, Cb, and Cr). If N is taken as 8, first perform 8-fold bilinear upsampling on the original face image to obtain the first image of 8H*8W*3. Then, for each color channel, divide the first image according to an 8*8 segmentation unit to obtain H×W 8*8 image blocks. For each image block, perform DCT transformation on the image block to convert each 8*8 image area into a 1*64 feature representation. Based on the feature representations of each (H×W) image block, obtain the first feature representation corresponding to this color channel with a size of H*W*64. By combining the first feature representations of the three color channels of the original face image, obtain the first feature representation of H*W(64×3 = 192), that is, the feature representation of each spatial domain channel is converted into the feature representation of multiple frequency domain channels. Among them, each color channel corresponds to 64 high-dimensional channels, and the number of channels of the converted image becomes 192.

[0094] It can be understood that the size of the original face image is H*W*3, and the size of this first feature representation is H*W*192. The number of channels is significantly increased, and the first feature representation can be regarded as the feature representation of the original face image corresponding in the high-dimensional space.

[0095] In the embodiment of the present application, the first feature representation of the original face image can also be obtained by extracting features from the original face image through a trained first model. Among them, the first model can be jointly trained with the subsequent generation model, and the specific training process can be referred to the subsequent detailed content. Optionally, the first model can be a neural network layer in the generation model for extracting the first feature representation.

[0096] Step S130: Based on the first feature representation, use the trained generation model to generate the second feature representation of the synthetic face image corresponding to the object to be recognized.

[0097] Among them, the synthetic face image is a non-real captured face image synthesized through AI technology. The synthetic face image is similar to the visual information contained in the original face image. The so-called visual information refers to the image content that can be directly observed by the naked eye.

[0098] In the embodiment of the present application, the first feature representation of the original face image is input into the trained generation model to generate the second feature representation of the synthetic face image corresponding to the object to be recognized.

[0099] Among them, the generation model can be trained in the following way:

[0100] A1: Obtain multiple first samples with labels; wherein, each first sample includes the original face image of a sample object, and the label of the first sample is the true identity information of the sample object.

[0101] A2: For each first sample, determine the third feature representation of the first sample;

[0102] For each first sample, through performing time-frequency transformation on the first sample, obtain the third feature representation of the first sample. Among them, the feature extraction method adopted in the model training stage is the same as the method for determining the first feature representation in the above step S120.

[0103] A3: Perform a training operation on the generative model to be trained based on each third feature representation until the first training end condition is satisfied, and obtain the trained generative model. Among them, the first training end condition can be that the first loss function is less than the first threshold, or the number of training times reaches the second threshold. This application does not limit this and can be set according to needs. Optionally, the generative model can adopt any neural network model, such as the U-Net convolutional neural network.

[0104] Among them, the training operation includes the following steps:

[0105] A31: Input the third feature representations corresponding to each first sample into the generative model to be trained, and respectively obtain the fourth feature representations of the synthetic face images corresponding to each sample object through the generative model;

[0106] A32: Respectively determine the predicted feature representation differences between the third feature representations and the fourth feature representations corresponding to each first sample; wherein, the predicted feature representation differences include the representation differences between the feature representations of each channel in the third feature representation and the fourth feature representation.

[0107] A33: Based on the predicted feature representation differences of each first sample, extract the predicted face features of each sample object, and based on the fourth feature representations of each sample object, generate the synthetic face images corresponding to each sample object;

[0108] Among them, the predicted face features of the sample object characterize the identity information of the predicted sample object.

[0109] Optionally, the predicted feature representation differences can be input into the auxiliary recognition model to be trained to obtain the predicted face features of each sample object, and for the sample object in each first sample, based on the fourth feature representation of the sample object, perform dimensionality reduction mapping to obtain the synthetic face image corresponding to the sample object. Among them, when the third feature representation is extracted by means of time-frequency transformation, correspondingly, the inverse transformation of time-frequency transformation can be adopted to perform the conversion from the frequency domain to the spatial domain, and the fourth feature representation is converted into a synthetic face image.

[0110] For example, assume that in step A2, the third feature representation of each first sample is obtained by performing discrete cosine transform (DCT) on each first sample. After the fourth feature representation is output based on the generative model to be trained, an inverse discrete cosine transform (IDCT) is performed on the fourth feature representation to obtain a synthesized face image corresponding to each sample object.

[0111] A34: Determine a first training loss based on the difference between the predicted face features corresponding to each first sample and the label of each first sample, and the difference between the original face image and the synthesized face image of each sample object.

[0112] Exemplarily, the first training loss function can be expressed as:

[0113] L = α|X, X'| + βl face (f(r), y)

[0114] where L represents the first training loss, X represents the original face image, X' represents the synthesized face image, |X, X'| represents the difference between the original face image and the synthesized face image, denoted as the image visual loss, α represents the first weight parameter, r represents the feature representation difference, f(r) represents the predicted face feature, y represents the true identity information, and l face (f(r), y) represents the face recognition loss, and β represents the second weight parameter of the face recognition loss.

[0115] Optionally, when calculating the face recognition loss, any face recognition loss function can be used, such as the additive angular margin loss for deep face recognition (ArcFace), the large margin cosine loss for deep face recognition (CosFace), the deep hypersphere embedding for face recognition (SphereFace), etc.

[0116] By constraining |X, X'|, the visual information contained in the third feature representation and the fourth feature representation is made similar, and the visual information contained in the feature representation difference obtained by subtracting the two is less. And by constraining l face (f(r), y) such that the feature representation difference contains sufficient recognition information to achieve the purpose of face recognition.

[0117] A35: Adjust model parameters in the generated model based on the first training loss.

[0118] Optionally, the auxiliary recognition model can be used to assist in training the generation model. In this case, the generation model and the auxiliary recognition model can be jointly trained during the model training process, and the model parameters in the generation model and the model parameters in the auxiliary recognition model can be adjusted based on the first training loss.

[0119] Figure 3 A flow chart of a training generation model provided in an embodiment of the present application, wherein a time-frequency transform is performed on the original facial image in the first sample to obtain a third feature representation of the first sample, and the third feature representation is input into the generation model to be trained to obtain a fourth feature representation. The fourth feature representation is subjected to an inverse transform of the time-frequency transform to obtain a synthesized facial image. Based on the fourth feature representation and the third feature representation, a predicted feature representation difference is obtained, and the predicted feature representation difference is input into the auxiliary recognition model to be trained to obtain a predicted facial feature. Based on the difference between the original facial image and the synthesized facial image, the image visual difference is determined, and based on the difference between the predicted facial feature and the sample label, the facial recognition difference is determined. Based on the image visual difference and the facial recognition difference, a first training loss is determined, and the model parameters in the generation model to be trained and the auxiliary recognition model are adjusted based on the first training loss.

[0120] In an embodiment of the present application, the face recognition method can be executed by a terminal, so after the generation model training is completed, the trained generation model can be deployed to the terminal so that the terminal generates a second feature representation through the trained generation model.

[0121] Optionally, when determining the label of the first sample, the real face image (i.e., the original face image) in the first sample can be input into a pre-trained first recognition model to obtain the real face features representing the real identity information of the first sample as the sample label for training the generated model.

[0122] The first recognition model is trained with the second samples as training samples, each second sample includes an original face image of a sample object, and the label of the second sample can be the object identifier of the sample object. The first recognition model includes several feature extraction layers and classification layers. After the real face image is input into the trained first recognition model, the face features are extracted through several feature extraction layers, and the face features are classified through the classification layer to determine the object identifier of the sample object to which they belong. Therefore, the face features output by the last feature extraction layer are used as the real face features.

[0123] Optionally, in the above step A2, the feature representation of the first sample can also be extracted by the first model. Each first sample is input into the first model to be trained, and the predicted third feature representation of the first sample is obtained. Based on the predicted third feature representations of each first sample, a joint training operation is performed on the generative model to be trained and the first model, and based on the first training loss, the model parameters in the generative model to be trained and the model parameters in the first model are adjusted.

[0124] Step S140: Determine the feature representation difference between the first feature representation and the second feature representation.

[0125] Step S150: Obtain the face recognition result of the object to be recognized based on the feature representation difference.

[0126] Among them, the second feature representation has the same size as the first feature representation (the same length, height, and number of channels), and both include the feature representation of the first number of channels.

[0127] In the embodiment of the present application, for each channel, the representation difference between the feature representation of the first feature representation corresponding to this channel and the feature representation of the second feature representation corresponding to this channel is determined. For example, the two feature representations are subtracted to obtain the representation difference corresponding to this channel. Based on the representation differences of each channel, the feature representation difference between the first feature representation and the second feature representation is determined.

[0128] In order to further improve the anti-reconstruction ability and the privacy protection ability of the face image, in the embodiment of the present application, the channel order in the feature representation difference can be randomly exchanged. For the representation differences of at least two channels in the feature representation difference, the channel order transformation is performed to obtain the transformed feature representation difference. Based on the transformed feature representation difference, face feature recognition is performed on the object to be recognized to obtain the face recognition result.

[0129] Optionally, when performing the channel order transformation, a first sequence can be generated based on the number of channels of the feature representation difference. Among them, the number of elements in the first sequence is equal to the number of channels, and the position of each element in the first sequence corresponds to each channel in the feature representation difference. Then, the positions of at least two elements in the first sequence are randomly transformed to obtain a second sequence, and according to the second sequence, the order of the representation differences of the channels in the feature representation difference is transformed.

[0130] Exemplarily, assume that the number of channels of both the first feature representation and the second feature representation is 5. Then, the number of channels of the feature representation difference between them is also 5. Based on the number of channels of the feature representation difference, the first sequence {1, 2, 3, 4, 5} is obtained. The channels in the feature representation difference are sequentially denoted as the first channel, the second channel... the fifth channel. Randomly transform the element positions in the first sequence to obtain the second sequence {5, 3, 1, 2, 4}. Then, according to the order of each element in the second sequence, sequentially transform the channels in the feature representation difference. The transformed channels are the fifth channel, the third channel, the first channel, the second channel, and the fourth channel in sequence.

[0131] It can be understood that when the discrete cosine transform (DCT) is used to extract the first feature representation from the original face image, the original face image of H*W*3 is converted into the first feature representation of H*W*192. The size of the second feature representation obtained by the generative model is also H*W*192. Therefore, the size of the feature representation difference obtained by subtracting the two is also H*W*192. There are 192×191×…×1 different transformation methods for performing channel order transformation on the feature representation difference. Therefore, even if the feature representation difference (or the latent face image after dimensionality reduction) of the object to be recognized with the channel order transformed is leaked during subsequent transmission or processing, it is impossible to guess the channel transformation method from 192×191×…×1 different transformation methods, and it is difficult to restore the original face image, greatly enhancing the anti-reconstruction ability.

[0132] Optionally, the above time-frequency transformation method can also use the discrete wavelet transform (DWT). However, when performing channel order transformation on the feature representation difference obtained based on DWT, there are 12×11×…×1 different transformation methods.

[0133] In the embodiments of the present application, when the terminal executes the face recognition method, in order to overcome the resource limitations of the terminal device locally and improve the face recognition accuracy, usually the target information to be recognized is sent to the server to obtain the face recognition result of the object to be recognized through the server.

[0134] After the terminal determines the feature representation difference between the first feature representation and the second feature representation, or determines the feature representation difference after performing channel order transformation, it can perform dimensionality reduction mapping based on at least one of the feature representation difference between the first feature representation and the second feature representation or the feature representation difference after performing channel order transformation, to obtain a private face image corresponding to the feature representation difference. And send the obtained private face image to the server for face recognition. Among them, the private face image contains less visual information and recognition information available for face recognition. The dimensionality reduction mapping can adopt the inverse mapping corresponding to the above-mentioned dimensionality increase mapping. When the dimensionality increase mapping is performed by means of time-frequency transformation, the inverse time-frequency transformation can be used for dimensionality reduction mapping, that is, the conversion from the frequency domain to the spatial domain is performed.

[0135] The server can obtain the face recognition result of the object to be recognized based on the received private face image and the known face images in the server. Among them, the known face images in the server are obtained from the real face images of known objects, and the method of obtaining the known face images is the same as the method of obtaining the private face images.

[0136] Optionally, after receiving the private face image, the server can input the private face image into a trained face recognition model, extract the face features to be recognized through the face recognition model, and determine the face features of the object to be recognized from the face features of multiple known objects stored in the feature library based on the object identifier of the object to be recognized, and match the face features to be recognized with the face features of the object to be recognized to obtain the face recognition result of the object to be recognized. If the face features to be recognized match the face features of the object to be recognized, it is determined that the face recognition of the object to be recognized is successful, otherwise it is determined that the face recognition fails. Optionally, when performing face matching, if the matching degree of the face features is greater than or equal to the fifth threshold, it is considered a successful match, otherwise it is considered a failed match.

[0137] Among them, the object identifier of the object to be recognized can be sent from the terminal to the server together with the private face image. The feature library stores the mapping relationship between the object identifiers of multiple known objects and their face features. The face features corresponding to any object identifier are extracted through the face recognition model based on the known face image corresponding to the known object. The method of obtaining the known face image is the same as the method of obtaining the private face image.

[0138] Optionally, the face features of multiple known objects stored in the feature library can be obtained by obtaining the second feature representation through a trained generation model based on the first feature representation of the real face image of each object, and then obtaining the face features through a trained face recognition model based on the feature representation difference between the first feature representation and the second feature. It can also be the real face features of known objects obtained through the above-mentioned trained first recognition model based on the real face images of each known object.

[0139] Optionally, the terminal can also directly send the feature representation difference, or the feature representation difference after channel transformation, to the server, so that the server performs face recognition based on the feature representation difference corresponding to the object to be recognized.

[0140] Optionally, the face recognition model can be trained in the following manner:

[0141] B1: Obtain a plurality of third samples with labels; wherein, each third sample includes a hidden face image corresponding to a sample object, and the label of the third sample is the object identifier of the sample object.

[0142] Optionally, the target information in the third sample can be determined based on a trained generation model. Specifically, obtain the original face images (real face images) of the sample objects in the plurality of third samples, determine the fifth feature representation of the original face image corresponding to each third sample, input the fifth feature representation corresponding to each third sample into the trained generation model, obtain the sixth feature representation corresponding to each third sample, and determine the hidden face image corresponding to each third sample based on the feature representation difference between the fifth feature representation and the sixth feature representation of each third sample.

[0143] B2: Perform a training operation on the face recognition model to be trained based on each third sample until the second training end condition is met, and obtain a trained face recognition model. The second training end condition can be that the second loss function is less than the third threshold, or the number of training times reaches the fourth threshold. This application does not limit this and can be set as needed.

[0144] Among them, the training operation includes the following steps:

[0145] B21: Input each third sample into the face recognition model to be trained, and respectively obtain the predicted object identifier corresponding to each sample object through the face recognition model.

[0146] B22: Determine the second training loss based on the difference between the predicted object identifier corresponding to each third sample and its label.

[0147] B23: Adjust the model parameters in the face recognition model based on the second training loss.

[0148] Based on Figure 1The face recognition method shown in the figure generates a feature representation of a synthetic face image corresponding to an object to be recognized through a generative model trained by AI technology, and performs face recognition based on the difference in feature representations between the original face image and the synthetic face image. Since the visual information contained in the original face image and the synthetic face image is similar, the difference in feature representations determined based on artificial intelligence technology can remove most of the visual information and retain sufficient recognition information to achieve the purpose of face recognition. Based on this method, while not affecting face recognition, privacy protection is provided for the face image, reducing the risk of privacy leakage.

[0149] It should be noted that the above-mentioned dimensionality reduction and elevation mapping are used to achieve the mutual conversion of an image between the spatial domain M and the high-dimensional space N, and this conversion is important for retaining the recognition information in the difference of feature representations. Any mapping that satisfies reversibility and additive homomorphism can be used for dimensionality reduction and elevation mapping.

[0150] Reversibility:

[0151] Additive homomorphism:

[0152] Among them, e represents the elevation mapping relationship, and d represents the dimensionality reduction mapping relationship.

[0153] The method provided in this application can not only be applied to privacy protection of face images in the face recognition process, but also be applied to any image processing process that requires privacy protection, such as fingerprint images containing object fingerprint information, iris images containing object iris information, payment code images, etc. The method provided in this application can be used for privacy protection of images.

[0154] It should be noted that in the embodiments of this application, based on the difference in feature representations, a private face image obtained by dimensionality reduction mapping has recognizability and privacy protection.

[0155] (1) Recognizability

[0156] Assume that X represents the original face image, X' represents the synthetic face image, X p represents the private face image, r represents the difference in feature representations, θ represents the channel transformation order, and s(r,θ) represents the channel transformation result obtained by performing a channel order transformation on the difference in feature representations.

[0157] Based on the properties of the above-mentioned additive homomorphism, we can obtain:

[0158] d(r + △r) = d(r) + d(△r) (1)

[0159] △r = r - s(r,θ) ≠ 0 (2)

[0160] Among them, △r represents the perturbation added to the feature representation difference. By randomly transforming the channel order of the feature representation difference, a non-zero perturbation △r is added, and some recognizable information is actually increased in the resulting channel transformation result.

[0161] In the embodiment of the present application, the original face image X is subtracted from the synthetic face image X' to obtain the image visual difference between the two images, denoted as R'. Through the dimensionality reduction mapping relationship, there is R' = d(r) between R' and r. Since the visual information contained in the original face image X and the synthetic face image X' is close, there is almost no visual information in R', and thus R' = d(r) = 0 can be obtained.

[0162] Since the feature representation difference r contains the recognition information in the original face image X, s(r, θ) and △r also contain the recognition information for face recognition in the original face image X. The private face image X obtained by performing dimensionality reduction mapping based on s(r, θ) p also contains the recognition information for face recognition.

[0163] (2) Privacy protection

[0164] Figure 4 Exemplarily shows the schematic diagrams of the original face image, the representation differences of each channel corresponding to the feature representation difference, and the private face image. Figure 4 (a) is the original face image. Figure 4 (b) exemplarily shows the representation differences corresponding to 4 channels in the feature representation difference. Figure 4 (c) exemplarily shows 4 private face images obtained by randomly swapping the channel order. Figure 4 The private face image shown in (c) contains almost no visual information of the face and cannot be reconstructed using the structural and texture information of the image. Therefore, it can achieve privacy protection of the face and cannot achieve the reconstruction of the face image.

[0165] Figure 5 For the schematic diagram of face recognition provided by the embodiment of the present application, after collecting the original face image of the object to be recognized, performing time-frequency transformation on the original face image to obtain the first feature representation, and inputting the first feature representation into the trained generation model to obtain the second feature representation of the synthetic face image corresponding to the object to be recognized. Subtract the feature representations of the corresponding channels of the first feature representation and the second feature representation to obtain the feature representation difference. Perform a random channel order transformation on the feature representation difference, and perform the inverse time-frequency transformation on the feature representation difference after transforming the channel order to obtain the private face image. From Figure 5It can be seen that there is a lot of visual information in the first feature representation and the second feature representation, while the visual information in the feature representation difference is significantly reduced, and no facial features can be observed from it. Moreover, the spatial domain image obtained after the inverse transformation also contains less visual information.

[0166] To facilitate a better understanding and description of the method provided in the embodiments of the present application, the following introduces the optional implementation manners of the method provided in the present application in combination with several specific scenario embodiments.

[0167] Taking the face recognition payment scenario as an example for illustration, the face recognition method provided in the embodiments of the present application can be applied to face recognition payment in the face recognition payment scenario. Figure 6 The structural schematic diagram of the payment system in the face recognition payment scenario is shown. The payment system includes a user terminal, a payment server, and a database. The user terminal can be any terminal running a payment application / software, such as a mobile phone, a wearable device, a face recognition payment device in a convenience store, etc. After the user terminal collects a face image and performs privacy protection processing, it sends it to the payment server for face recognition.

[0168] When performing face recognition payment, the mobile phone responds to the trigger operation of user A for face recognition payment in the payment client, and real-time collects the current face image of user A, that is, the original face image. By performing time-frequency transformation on the original face image, the first feature representation of the original face image is obtained.

[0169] The first feature representation of the original face image is input into the trained generation model to obtain the second feature representation of the synthetic face image of user A. Among them, the generation model is pre-trained and sent to the mobile phone in the installation package (or update package) of the payment client.

[0170] By subtracting the feature representations corresponding to each channel of the first feature representation and the second feature representation, the feature representation difference between the first feature representation and the second feature representation is obtained. Among them, the feature representation difference includes the representation differences of multiple channels. The channel order of the feature representation difference is randomly adjusted, and the inverse transformation of the time-frequency transformation is performed on the adjusted feature representation difference to obtain the private face image of user A.

[0171] The user terminal sends a payment request to the payment server based on the private face image and payment information of User A. After receiving the payment request, the payment server can input the private face image of User A in the payment request into the trained face recognition model to extract the face features to be recognized. And according to the account identifier carried in the payment request, it searches for the face features corresponding to the account identifier from the face features of multiple accounts stored in the feature library, and matches the face features to be recognized with the face features corresponding to the account identifier. Among them, the face features of users corresponding to multiple accounts are pre-stored in the feature library, which represent the identity information of the users. The face features corresponding to any account can be obtained by obtaining the private face image based on the collected face image and extracting the private face image through the face recognition model when the face payment is enabled.

[0172] If the match is successful, the face recognition of User A is successful, and the payment can be completed based on the payment information in the payment request, and the payment success information is fed back to the user terminal. If the match fails, the payment failure information is returned to the user terminal to prompt User A to perform face recognition again or pay in other ways.

[0173] The face recognition method provided in the embodiments of the present application can also be applied to travel scenarios. For example, when traveling by train, high-speed train, or plane, face recognition can be used for identity verification. Figure 7 The structural schematic diagram of the identity verification system in the travel scenario is shown. The identity verification system includes a face recognition turnstile, a verification server, and a database. After the face recognition turnstile collects the face image and performs privacy protection processing, it sends it to the verification server for identity verification.

[0174] After the face image of User B (i.e., the original face image) is collected in real time by the face recognition turnstile corresponding to the travel station, the first feature representation of the original face image is obtained by performing time-frequency transformation on the original face image. The first feature representation of the original face image is input into the trained generation model to obtain the second feature representation of the synthetic face image of User B. Among them, the generation model is pre-trained and deployed in the face recognition turnstile.

[0175] By subtracting the feature representations of each channel of the first feature representation and the second feature representation, the feature representation difference between the first feature representation and the second feature representation is obtained. Among them, the feature representation difference includes the representation differences of multiple channels. The channel order of the feature representation difference is randomly adjusted, and the inverse time-frequency transformation is performed on the adjusted feature representation difference to obtain the private face image of User B.

[0176] The face recognition turnstile sends an authentication request to the verification server based on the private face image of User B. After receiving the authentication request, the verification server can input the private face image of User B into the trained face recognition model to extract the face features to be recognized. And according to the user identification carried in the authentication request, it looks up the face features corresponding to the user identification from the face features of multiple users stored in the feature library, and matches the face features to be recognized with the face features corresponding to the user identification. Among them, the face features of multiple users are pre-stored in the feature library, and the face features of any user can be determined based on the collected face image to obtain the private face image when purchasing tickets online or offline, and are extracted through the face recognition model.

[0177] The embodiment of the present application also provides a face recognition system, as Figure 8 shown. The system includes a terminal 20 and a server 21, and the terminal 20 and the server 21 can be communicatively connected by wired or wireless means.

[0178] When performing face recognition, the terminal 20 can use its configured image acquisition device (such as a camera, etc.) to collect the current face image of the object to be recognized in real time as the original face image. By performing dimensionality elevation mapping on the original face image, the first feature representation of the original face image is obtained. Among them, when performing dimensionality elevation mapping, any mapping that satisfies reversibility and additive homomorphicity can be used. As an optional implementation manner, the dimensionality elevation mapping can be performed by means of time-frequency transformation, such as DFT transformation, DCT transformation, DWT transformation, etc., to transform each image from the spatial domain to the frequency domain.

[0179] Based on the first feature representation, use the trained generation model to generate the second feature representation of the synthetic face image corresponding to the object to be recognized, and determine the feature representation difference between the first feature representation and the second feature representation. Perform dimensionality reduction mapping on the feature representation difference to obtain the private face image corresponding to the feature representation difference. Send the private face image to the server 21. Among them, for each channel, determine the representation difference between the feature representation corresponding to the first feature representation of this channel and the feature representation corresponding to the second feature representation of this channel. For example, subtract the two feature representations to obtain the representation difference corresponding to this channel. Based on the representation differences of each channel, determine the feature representation difference between the first feature representation and the second feature representation.

[0180] The server 21 can match the received private face image with the known face images stored in itself to obtain the face recognition result of the object to be recognized. Among them, the known face images are obtained based on the real face images of known objects, and the method of obtaining the known face images is the same as the method of obtaining the private face images.

[0181] Optionally, the server 21 may input the received private face image into the trained face recognition model, extract the face features to be recognized through the face recognition model, and determine the face features of the object to be recognized from the face features of multiple known objects stored in the feature library based on the object identifier of the object to be recognized. Then, the face features to be recognized are matched with the face features of the object to be recognized to obtain the face recognition result of the object to be recognized. If the face features to be recognized match the face features of the object to be recognized, it is determined that the face recognition of the object to be recognized is successful; otherwise, it is determined that the face recognition fails. Among them, the face features of each known object stored in the feature library are extracted based on the known face images of the known objects, and the method for obtaining the known face images is the same as the method for obtaining the private face images.

[0182] Optionally, the terminal 20 may also directly send the feature representation difference to the server 21, so that the server 21 extracts the face features to be recognized based on the feature representation difference for subsequent face recognition.

[0183] The specific implementation methods of each end in the above face recognition system and related contents such as model training are elaborated in detail in the above steps S110 to S150. This application will not repeat them here. For details, please refer to the above content.

[0184] The solution of this application provides a private face protection method based on high-dimensional channel subtraction and channel order transformation, which can be abbreviated as MinusFace. In order to verify that the solution of this application can accurately perform face recognition, the solution of this application (11), the solution without privacy protection (1), and other private face protection solutions (2) - (10) are compared. The face recognition accuracy rates of the solution of this application and the above comparison solutions are tested on 7 face image datasets such as LFW, CFP-FP, AgeDB, CPLFW, CALFW, IJB-B, and IJB-C. As shown in the following table:

[0185]

[0186] It can be seen that the gap between the solution of this application and the solution without privacy protection (1) is less than 2% in terms of face recognition accuracy rate. In the comparison with other private face protection solutions (2) - (10), the solution of this application has a higher face recognition accuracy rate, and the privacy protection ability of the solution of this application is significantly better than that of other private face protection solutions.

[0187] Figure 9 The privacy protection effects of the solution of this application (11) and other private face protection methods (5) - (10) are compared, where Figure 9(a) shows the visual information of the privacy-preserving face images generated by various privacy-preserving face protection methods. It can be seen that the privacy-preserving face images generated by this application contain almost no visually visible information, and its privacy protection ability is significantly better than other privacy-preserving face protection methods. Figure 9 (b) shows the face images restored after reconstruction by each privacy-preserving face protection method. It can be seen that other methods can restore clear and recognizable faces, while the solution of this application cannot restore a face. Therefore, the solution of this application has significant anti-reconstruction ability. Figure 9 (c) compares the structural similarity (structural similarity index, SSIM), which is an index for measuring the similarity between two images, between the reconstructed face image and the original face image, and the peak signal-to-noise ratio (Peak Signal to Noise Ratio, PSNR) of the reconstructed face image, which is an objective standard for evaluating image quality. Among them, the SSIM and PSNR of the solution of this application are smaller than those of other privacy-preserving face protection methods, indicating that the anti-reconstruction ability of the solution of this application is stronger than that of other privacy-preserving face protection methods.

[0188] Currently, for the privacy protection of images, there is also a privacy protection method based on encryption, which encrypts the face image and extracts and recognizes image features in the encrypted domain. However, this privacy protection method based on encryption often requires hundreds or even thousands of times more computational workload, has a long computational time, and a low face recognition efficiency. Especially for scenarios with high real-time requirements, such as payment and travel, it is difficult to meet the actual needs. The solution provided by this application saves computational resources compared with this method, improves the face recognition efficiency while ensuring privacy security.

[0189] Based on the same principle as the face recognition method provided by the embodiments of this application, the embodiments of this application provide a face recognition device, as Figure 10 shown. The face recognition device 300 may include an original image acquisition module 310, a first feature determination module 320, a second feature determination module 330, a feature difference determination module 340, and a face recognition module 450.

[0190] The original image acquisition module 310 is used to acquire the original face image of the object to be recognized;

[0191] The first feature determination module 320 is used to determine the first feature representation of the original face image;

[0192] The second feature determination module 330 is used to generate the second feature representation of the synthetic face image corresponding to the object to be recognized by using the trained generation model based on the first feature representation;

[0193] A feature difference determination module 340, configured to determine a feature representation difference between the first feature representation and the second feature representation;

[0194] A face recognition module 350, configured to obtain a face recognition result of the object to be recognized based on the feature representation difference.

[0195] Optionally, both the first feature representation and the second feature representation include feature representations of a first number of channels;

[0196] The feature difference determination module 340 may be configured to:

[0197] Determine a representation difference between the feature representations of each channel in the first feature representation and the second feature representation;

[0198] Based on the representation differences of each channel, determine the feature representation difference between the first feature representation and the second feature representation;

[0199] The face recognition module 350 may be configured to:

[0200] Perform a channel order transformation on the representation differences of at least two channels in the feature representation difference to obtain a transformed feature representation difference;

[0201] Based on the transformed feature representation difference, perform face feature recognition on the object to be recognized to obtain a face recognition result.

[0202] Optionally, the face recognition module 350 may be configured to:

[0203] Generate a first sequence based on the number of channels of the feature representation difference; wherein, the number of elements in the first sequence is equal to the number of channels, and the position of each element in the first sequence corresponds to each channel in the feature representation difference;

[0204] Randomly transform the positions of at least two elements in the first sequence to obtain a second sequence;

[0205] According to the second sequence, perform an order transformation on the representation differences of the channels in the feature representation difference.

[0206] Optionally, the first feature determination module 320 may be configured to:

[0207] Perform a conversion from the spatial domain to the frequency domain on the original face image to obtain a first feature representation of the original face image; wherein, the feature representation of one spatial domain channel of the original face image is converted into feature representations of multiple frequency domain channels.

[0208] Optionally, the face recognition module 350 may be configured to:

[0209] Perform a frequency-domain to spatial-domain conversion on the feature representation difference to obtain a hidden private face image corresponding to the feature representation difference;

[0210] Based on the hidden private face image and known face images in the server, obtain a face recognition result of the object to be recognized, where the known face images are obtained from real face images of known objects, and the method for obtaining the known face images is the same as the method for obtaining the hidden private face image.

[0211] Optionally, the face recognition device further includes a model training module, and the model training module can be used for:

[0212] Obtain a plurality of first samples with labels, each first sample includes an original face image of a sample object, and the label of the first sample is the true identity information of the sample object;

[0213] For each first sample, determine a third feature representation of the first sample;

[0214] Perform a training operation on a generative model to be trained based on each third feature representation until a first training end condition is met, and obtain a trained generative model. The training operation includes:

[0215] Input the third feature representations corresponding to each first sample into the generative model to be trained, and respectively obtain fourth feature representations of synthetic face images corresponding to each sample object through the generative model;

[0216] Respectively determine the predicted feature representation differences between the third feature representations and the fourth feature representations corresponding to each first sample;

[0217] Based on the predicted feature representation differences of each first sample, extract predicted face features of each sample object, where the predicted face features represent the identity information of the predicted sample object; and based on the fourth feature representations of each sample object, generate synthetic face images corresponding to each sample object;

[0218] Based on the differences between the predicted face features corresponding to each first sample and the labels of each first sample, and the differences between the original face images and the synthetic face images of each sample object, determine a first training loss;

[0219] Adjust the model parameters in the generative model based on the first training loss.

[0220] Optionally, the model training module can be used for:

[0221] Input each predicted feature representation difference into an auxiliary recognition model to be trained to obtain predicted face features of each sample object;

[0222] For each sample object, perform a conversion from the frequency domain to the spatial domain on the fourth feature representation of the sample object to obtain a synthesized face image corresponding to the sample object;

[0223] The model training module can be used to:

[0224] Adjust the model parameters in the generation model and the model parameters in the auxiliary recognition model based on the first training loss.

[0225] The device according to the embodiments of the present application can execute the method provided by the embodiments of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function description of each module of the device, reference can specifically be made to the description in the corresponding method shown above, and details are not described herein again.

[0226] Based on Figure 10 The shown face recognition device generates a feature representation of a synthesized face image corresponding to an object to be recognized through artificial intelligence technology, and performs face recognition based on the feature representation difference between the original face image and the synthesized face image. Since the visual information contained in the original face image and the synthesized face image is similar, the feature representation difference determined based on artificial intelligence technology can remove most of the visual information and retain sufficient recognition information to achieve the purpose of face recognition. Based on this method, while not affecting face recognition, privacy protection is performed on the face image, reducing the risk of privacy leakage.

[0227] Moreover, by further performing a transformation of the random channel order on the feature representation difference, random material texture features are added to the image, which helps to better resist reconstruction attacks, enhances the anti-reconstruction ability, and realizes privacy protection.

[0228] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0229] An electronic device is provided in the embodiments of the present application, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program stored in the memory, the method according to any optional embodiment of the present application can be implemented.

[0230] Figure 11The following is a schematic structural diagram of an electronic device to which the embodiments of the present invention are applicable. As Figure 11 shown, the electronic device can be a server or a user terminal, and the electronic device can be used to implement the methods provided in any embodiment of the present invention.

[0231] As Figure 11 shown in, the electronic device 2000 mainly includes at least one processor 2001 ( Figure 11 one is shown in), a memory 2002, a communication module 2003, an input / output interface 2004 and other components. Optionally, the components can be connected and communicate with each other through a bus 2005. It should be noted that Figure 11 the structure of the electronic device 2000 shown in is only schematic and does not constitute a limitation on the electronic device applicable to the methods provided in the embodiments of the present application.

[0232] Among them, the memory 2002 can be used to store an operating system and application programs, etc. The application programs can include computer programs that implement the methods shown in the embodiments of the present invention when called by the processor 2001, and can also include programs for implementing other functions or services. The memory 2002 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and computer programs, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0233] The processor 2001 is connected to the memory 2002 via the bus 2005, and realizes corresponding functions by calling the application programs stored in the memory 2002. Among them, the processor 2001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 2001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0234] The electronic device 2000 can be connected to a network through the communication module 2003 (which can include, but is not limited to, components such as a network interface) to communicate with other devices (such as user terminals or servers, etc.) through the network to achieve data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 2003 can include a wired network interface and / or a wireless network interface, etc., that is, the communication module can include at least one of a wired communication module or a wireless communication module.

[0235] The electronic device 2000 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 2004. The electronic device 2000 itself can have a display device, and can also externally connect other display devices through the interface 2004. Optionally, a storage device, such as a hard disk, etc., can also be connected through the interface 2004 to store the data in the electronic device 2000 into the storage device, or read the data in the storage device, and can also store the data in the storage device into the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 2004 can be components of the electronic device 2000 or external devices connected to the electronic device 2000 when needed.

[0236] The bus 2005 for connecting each component may include a path for transmitting information between the above components. The bus 2005 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. According to different functions, the bus 2005 may be divided into an address bus, a data bus, a control bus, etc.

[0237] Optionally, for the solution provided in the embodiments of the present invention, the memory 2002 may be used to store a computer program for executing the solution of the present invention, and the processor 2001 runs the computer program to implement the actions of the method or device provided in the embodiments of the present invention.

[0238] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application provide a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the corresponding content of the foregoing method embodiments can be implemented.

[0239] The embodiments of the present application also provide a computer program product, which includes a computer program, and when the computer program is executed by a processor, the corresponding content of the foregoing method embodiments can be implemented.

[0240] It should be noted that the terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in words.

[0241] It should be understood that although the flowcharts in the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0242] The above are only optional implementation manners of some implementation scenarios of this application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of this application, adopting other similar implementation means based on the technical idea of this application also belongs to the protection scope of the embodiments of this application.

Claims

1. A face recognition method, characterized in that, Including: Obtain the original face image of the object to be recognized; Determine the first feature representation of the original face image; Based on the first feature representation, use the trained generation model to generate the second feature representation of the synthetic face image corresponding to the object to be recognized; Determine the feature representation difference between the first feature representation and the second feature representation; Based on the feature representation difference, obtain the face recognition result of the object to be recognized.

2. The method according to claim 1, wherein Both the first feature representation and the second feature representation include feature representations with a first number of channels; The determining the feature representation difference between the first feature representation and the second feature representation includes: Determine the representation difference between the feature representations of each channel in the first feature representation and the second feature representation; Based on the representation differences of each channel, determine the feature representation difference between the first feature representation and the second feature representation; The obtaining the face recognition result of the object to be recognized based on the feature representation difference includes: Perform a channel order transformation on the representation differences of at least two channels in the feature representation difference to obtain a transformed feature representation difference; Based on the transformed feature representation difference, perform face feature recognition on the object to be recognized to obtain a face recognition result.

3. The method according to claim 2, wherein The performing a channel order transformation on the representation differences of at least two channels in the feature representation difference to obtain a transformed feature representation difference includes: Generate a first sequence based on the number of channels of the feature representation difference; wherein, the number of elements in the first sequence is equal to the number of channels, and the position of each element in the first sequence corresponds to each channel in the feature representation difference; Randomly transform the positions of at least two elements in the first sequence to obtain a second sequence; According to the second sequence, perform an order transformation on the representation differences of the channels in the feature representation difference.

4. The method according to claim 1, characterized in that, The determining the first feature representation of the original face image includes: Perform a conversion from the spatial domain to the frequency domain on the original face image to obtain the first feature representation of the original face image; wherein, the feature representation of one spatial domain channel of the original face image is converted into the feature representations of multiple frequency domain channels.

5. The method according to claim 1 or 2, characterized in that, The obtaining the face recognition result of the object to be recognized based on the feature representation difference includes: Perform a conversion from the frequency domain to the spatial domain on the feature representation difference to obtain the private face image corresponding to the feature representation difference; Based on the private face image and the known face images in the server, obtain the face recognition result of the object to be recognized, wherein the known face images are obtained from the real face images of known objects, and the method for obtaining the known face images is the same as the method for obtaining the private face image.

6. The method according to any one of claims 1 to 4, characterized in that The generation model is trained in the following manner: Obtain a plurality of first samples with labels, each first sample includes the original face image of a sample object, and the label of the first sample is the true identity information of the sample object; For each first sample, determine the third feature representation of the first sample; Performing a training operation on the generative model to be trained based on each third feature representation until a first training end condition is met, obtaining a trained generative model, where the training operation includes: Inputting the third feature representations corresponding to each first sample into the generative model to be trained, and respectively obtaining the fourth feature representations of the synthetic face images corresponding to each sample object through the generative model; Respectively determining the predicted feature representation differences between the third feature representations and the fourth feature representations corresponding to each first sample; Based on the predicted feature representation differences of each first sample, extracting the predicted face features of each sample object, where the predicted face features characterize the identity information of the predicted sample object; and based on the fourth feature representations of each sample object, generating the synthetic face images corresponding to each sample object; Determining a first training loss based on the differences between the predicted face features corresponding to each first sample and the labels of each first sample, and the differences between the original face images and the synthetic face images of each sample object; Adjusting the model parameters in the generative model based on the first training loss.

7. The method according to claim 6, characterized in that, The extracting the predicted face features of each sample object based on the predicted feature representation differences of each first sample includes: Inputting each predicted feature representation difference into an auxiliary recognition model to be trained to obtain the predicted face features of each sample object; The generating the synthetic face images corresponding to each sample object based on the fourth feature representations of each sample object includes: For each sample object, performing a conversion from the frequency domain to the spatial domain on the fourth feature representation of the sample object to obtain the synthetic face image corresponding to the sample object; The adjusting the generative model based on the first training loss includes: Adjusting the model parameters in the generative model and the model parameters in the auxiliary recognition model based on the first training loss.

8. A face recognition device, characterized in that, Including: An original image acquisition module, configured to acquire the original face image of the object to be recognized; A first feature determination module, configured to determine the first feature representation of the original face image; A second feature determination module, configured to generate the second feature representation of the synthetic face image corresponding to the object to be recognized by using the trained generative model based on the first feature representation; A feature difference determination module, configured to determine the feature representation difference between the first feature representation and the second feature representation; A face recognition module, configured to obtain the face recognition result of the object to be recognized based on the feature representation difference.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, where a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that, The computer product includes a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.