A face recognition method with privacy protection using homomorphic encryption technology

By constructing a neural network with mixed plain text and cipher text, the plain text and cipher text information in face recognition is processed, and knowledge distillation technology is used to solve the problem of computing overhead caused by the difference in feature importance in face recognition methods under homomorphic encryption, and efficient face recognition is achieved.

CN115937939BActive Publication Date: 2025-06-10NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211561797.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-07
Publication Date
2025-06-10
Estimated Expiration
2042-12-07

AI Technical Summary

Technical Problem

The existing face recognition methods under homomorphic encryption fail to effectively consider the importance differences between features, resulting in a long model inference time and high computational overhead.

Method used

A neural network with mixed plain text ciphertexts is constructed, and the plain text information and ciphertext information are processed separately through two branches, and a cross-branch connection channel and channel integration layer are added between branches, so as to improve network performance using knowledge distillation technology.

Benefits of technology

While maintaining the original performance, the computing overhead is greatly reduced and the efficiency of face recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937939B_ABST
    Figure CN115937939B_ABST
Patent Text Reader

Abstract

The present invention discloses a face recognition method with privacy protection using homomorphic encryption technology. By constructing a network structure that mixes plaintext and ciphertext to process partially encrypted image information, the computational overhead is greatly reduced. At the same time, a cross-branch connection channel and a channel integration layer are proposed to transmit the information of the plaintext branch to the ciphertext branch, and the knowledge distillation technology is used to enable the hybrid encryption model to learn the feature representation of the "Teacher" network with better performance, so as to ensure that the performance loss under encryption is as low as possible. Users can upload partially encrypted images as needed and upload the content to be predicted to the provider of machine learning as a service, and the provider of the machine learning server can also use the hybrid encryption model in the present invention to efficiently and accurately return the prediction result to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a face recognition method with privacy protection using homomorphic encryption technology, belonging to the fields of image recognition and data homomorphic encryption technology. Background Art

[0002] Currently, with the development of multimedia technology and the popularization of intelligent devices, the acquisition of image information has become increasingly convenient. At the same time, the demand for image classification has also increased. Currently, Machine Learning inference as a Service (MLaSS) is widely used. Through this server, users can use artificial intelligence services deployed in the cloud with powerful computing power, such as object recognition, image retouching, voice assistants, etc. However, users need to upload relevant data to the server, which may lead to privacy leakage. Therefore, more and more researchers are concerned about how to build machine learning models with privacy protection, which can provide machine learning as a service to users but cannot obtain users' private information.

[0003] Face recognition is an important application of machine learning. Since faces contain privacy information, how to build a face recognition method with privacy protection is an important issue. Currently, the main face recognition methods with privacy protection are based on federated learning, differential privacy, and homomorphic encryption methods. However, federated learning requires multiple parties to participate, and its communication cost is very high; the differential privacy method can only protect the privacy security of training data weakly and will significantly reduce the model accuracy; the previous homomorphic encryption methods do not consider the privacy differences between features and directly encrypt the entire face, so the computational overhead on the ciphertext is very large. Summary of the Invention

[0004] Object of the Invention: The previous face recognition methods under homomorphic encryption often do not consider the importance differences between features. However, in actual situations, often only a few features play a major role. For example, in a face picture, the importance and sensitivity of the background of the picture are much lower than the face information appearing in the picture. Ignoring the importance differences between features also results in the previous models often having more homomorphic operations, which makes the inference time of the model longer. To address the above problems, the present invention constructs a neural network that mixes plaintext and ciphertext, which can efficiently predict homomorphically encrypted face images. The network has two branches, which can process plaintext information and ciphertext information respectively; there is a cross-branch connection channel between the two branches, which can integrate the features of the plaintext part onto the ciphertext branch. The present invention also uses the knowledge distillation technology, aiming to enable the two-branch network to better learn the feature representation and improve the network performance. Finally, on the premise of maintaining the original performance, the computational overhead is greatly reduced.

[0005] Technical solution: A face recognition method with privacy protection using homomorphic encryption technology, including a hybrid encryption model training step and an encrypted face prediction step.

[0006] The specific hybrid encryption model training step is as follows:

[0007] Step 100, select a homomorphic encryption scheme (such as BFV), and select a backbone network (such as VGG16) as the plaintext network branch of the hybrid encryption model. Modify the non-polynomial part of the backbone network structure to obtain the ciphertext network branch, ensuring that the selected homomorphic encryption scheme supports the prediction operation task of the ciphertext network branch. At the same time, add a cross-branch connection channel and a channel integration layer between the plaintext branch and the ciphertext branch.

[0008] Step 101, take the sensitive part of the face image as the input of the ciphertext network branch. The remaining part is used as the input of the plaintext network branch.

[0009] Step 102, use the backbone network to train the face image, and use the knowledge distillation technology to let the hybrid encryption model learn the feature representation of the trained backbone network.

[0010] Step 103, use the knowledge distillation technology to let the hybrid encryption model learn the output of the trained backbone network.

[0011] Step 104, the server implements the trained hybrid encryption model with a homomorphic encryption framework and deploys it on the cloud.

[0012] The encrypted face prediction step is as follows:

[0013] Step 200, the user prepares the face image to be predicted.

[0014] Step 201, the user divides the face image to be predicted into sensitive and non-sensitive parts, and encrypts the sensitive part.

[0015] Step 202, the user uploads the plaintext-ciphertext mixed image data to the cloud.

[0016] Step 203, the cloud hybrid encryption model makes a prediction on the plaintext-ciphertext mixed image data and returns the result to the user.

[0017] Step 204, the user decrypts to obtain the prediction result.

[0018] The backbone network in step 100, that is, the network used for the plaintext branch of the hybrid encryption model, and for the ciphertext branch, the operations in the backbone network need to be replaced with operations supported by the selected homomorphic encryption scheme. For example, if the homomorphic encryption scheme supports polynomial operations but does not support the maximum operation, the ReLU activation function can be replaced with a polynomial activation function, and the max pooling layer can be replaced with an average pooling layer.

[0019] The cross-branch connection channel in step 100: A channel used to transmit the information of the plaintext branch to the ciphertext branch. This component can be formalized as:

[0020]

[0021] where x c,i , x nc,i respectively represent the input tensors of the i-th layer of the ciphertext branch and the i-th layer of the plaintext branch . Note that before calculating the addition of x c,i and the output tensor of the i-th layer of the plaintext branch , the shapes of the two need to be unified using the Chop function first.

[0022] The channel integration layer in step 100: A network layer used to combine the channels of the input information of the plaintext branch in the cross-branch connection. Its purpose is to enable each channel of the ciphertext branch to learn using the information from all channels of the plaintext branch. This component can be formalized as a matrix whose parameters can be learned during training. Adding it to the cross-branch connection channel, the new cross-branch connection can be formalized as:

[0023]

[0024] The sensitive part in step 101: The part that contains more privacy information or the part that the user wants to encrypt.

[0025] The remaining part in step 101: The result obtained by filling 0 at the corresponding positions of the sensitive part in the original image.

[0026] In step 102, the knowledge distillation technique is used to let the hybrid encryption model learn the feature representation of the trained backbone network. The specific method is as follows: Select the output of a certain intermediate layer of the backbone network (for example, the output of the last convolutional block), denoted as Denote the output of the corresponding layer of the ciphertext branch as The output of the corresponding layer of the plaintext branch as Then the intermediate representation of the backbone network can be learned by minimizing the following loss function, where x is the input of the network, that is, filling 0 on both the width and height sides of to make its size consistent with , so that the addition operation in the following formula can be performed (refer to Figure 1 in the specific implementation).

[0027]

[0028] In step 103, the knowledge distillation technique is used to enable the hybrid encryption model to learn the output of the trained backbone network. Let the output of the backbone network (“Teacher”) be a T , and the output of the hybrid encryption model (“Student”) be a S . Let where the relaxation parameter τ > 1. The hybrid encryption model is trained by minimizing the following loss:

[0029]

[0030] where y True is the true label, is the cross-entropy loss function. W MCN are the parameters of the hybrid encryption model, and λ ∈ [0, 1].

[0031] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the privacy-protected face recognition method using homomorphic encryption technology as described above.

[0032] A computer-readable storage medium stores a computer program for executing the privacy-protected face recognition method using homomorphic encryption technology as described above. Description of the Drawings

[0033] Figure 1 This is an example of the hybrid encryption model according to an embodiment of the present invention, where the light gray part is the backbone network, the dark gray part is the ciphertext network branch, and the white part is the plaintext network branch;

[0034] Figure 2 This is the flowchart of the method according to an embodiment of the present invention;

[0035] Figure 3 This is the flowchart of the training steps of the hybrid encryption model according to an embodiment of the present invention;

[0036] Figure 4 This is the flowchart of the encrypted face prediction step according to an embodiment of the present invention. Detailed Embodiments

[0037] The present invention will be further clarified below with reference to specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention fall within the scope defined by the appended claims of this application.

[0038] The privacy-protected face recognition method using homomorphic encryption technology includes a hybrid encryption model training step and an encrypted face prediction step.

[0039] The training steps of the hybrid encryption model are as follows Figure 3 shown, and an example of the hybrid encryption model is as follows Figure 1 shown. Before constructing the hybrid encryption model, the provider of machine learning as a service (hereinafter referred to as the server) needs to collect a large number of face images as training data (step 10), and should select a backbone network with good performance (such as Resnet50 or VGG-16) (step 11). To ensure performance, the sensitive parts of these images can be adjusted to the same area through scaling operations. After the server selects the homomorphic encryption scheme, when constructing the hybrid encryption model (step 13), it needs to adjust the operations of the backbone network. Taking the BFG / BGV encryption scheme as an example, the server needs to replace the non-polynomial pooling layer and activation function in the backbone network with a polynomial pooling layer and activation function as the ciphertext network branch in the hybrid encryption model, and the original backbone network serves as the plaintext network branch. Then, cross-branch connection channels and channel integration layers need to be added to the network. A combination of cross-branch connection channels and channel integration layers is added after each convolutional unit to ensure that the low-level information of the plaintext network branch is transmitted to the ciphertext network branch and each channel of the ciphertext branch can obtain information from all channels of the plaintext branch. Since there is no channel dimension in the final fully connected layer of the network, only cross-branch connections need to be added. The server also needs to divide the training set images into sensitive parts and non-sensitive parts. For example, extract the middle quarter part of the image as the sensitive part, and this part is used as the input of the ciphertext network branch. The original face image is filled with 0 in the middle quarter part to obtain the remaining part, which is used as the input of the plaintext network branch (step 12).

[0040] At the same time, the server needs to train the backbone network on the prepared face image dataset (step 14). First, initialize the hybrid encryption model (step 15), use the trained backbone network model as the "Teacher" model in knowledge distillation, and the hybrid encryption model as the "Student" model to learn the intermediate representation of the backbone network model (step 16). Specifically, select the output result of a certain layer of the backbone network (denoted as the k-th layer) (for example, the output result of the last convolutional unit as the representation to be learned, select the output result of the same layer of the hybrid encryption model, that is, the output of the last convolutional unit in the plaintext branch and the output of the last convolutional unit in the ciphertext branch Considering that the image input to the ciphertext branch is only a small part of the original image, the output of the ciphertext branch needs to be filled (refer to Figure 1 ), so that the shape of the output tensor of the ciphertext branch is the same as that of the output tensor of the plaintext branch, and then add the filled output tensor of the k-th layer of the ciphertext branch and the output tensor of the k-th layer of the plaintext branch, that is Optimize the distance between the result of adding the above formula and the output result of the backbone network to learn the intermediate representation of the backbone network.

[0041] When the above optimization objective (i.e., ) is less than a certain small value ∈ (step 17), proceed to the next knowledge distillation. The method of the next knowledge distillation refers to step 103 of the technical solution in the invention content (step 18). After the model is trained, the server implements the trained model using the selected homomorphic encryption scheme (such as BGV, CKKS) and deploys it in the cloud.

[0042] The encrypted face prediction steps are as Figure 4 shown. First, the user prepares the face image data to be predicted (step 20), divides the face image to be predicted into a sensitive part and a non-sensitive part according to the position of the sensitive area, and encrypts the sensitive part on the user side (step 21). Upload the face image data to be predicted to the cloud (step 22), and the hybrid encryption model in the cloud makes a prediction on the face image data to be predicted (step 23). After the prediction is completed, the server returns the result to the user side, and the user decrypts the result (step 24) to obtain the predicted value.

[0043] Obviously, those skilled in the art should understand that each step of the above privacy-protected face recognition method using homomorphic encryption technology in the embodiments of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. A face recognition method with privacy protection using homomorphic encryption technology, characterized in that, it includes a hybrid encryption model training step and an encrypted face prediction step; The specific hybrid encryption model training step is as follows: Step 100, select a homomorphic encryption scheme, and select a backbone network as the plaintext network branch of the hybrid encryption model. Modify the non-polynomial part of the backbone network structure to obtain the ciphertext network branch, so that the selected homomorphic encryption scheme supports the prediction operation task of the ciphertext network branch; At the same time, add a cross-branch connection channel and a channel integration layer between the plaintext branch and the ciphertext branch; Step 101, take the sensitive part of the face image as the input of the ciphertext network branch, and the remaining part as the input of the plaintext network branch; Step 102, train the face image with the backbone network, and use the knowledge distillation technology to let the hybrid encryption model learn the feature representation of the trained backbone network; Step 103, use the knowledge distillation technology to let the hybrid encryption model learn the output of the trained backbone network; Step 104, the server implements the trained hybrid encryption model with a homomorphic encryption framework and deploys it on the cloud; The encrypted face prediction step is as follows: Step 200, the user prepares the face image to be predicted; Step 201, the user divides the face image to be predicted into sensitive and non-sensitive parts, and homomorphically encrypts the sensitive part; Step 202, the user uploads the plaintext-ciphertext mixed image data to the cloud; Step 203, the cloud hybrid encryption model makes a prediction on the plaintext-ciphertext mixed image data and returns the result to the user; Step 204, the user decrypts to obtain the prediction result.

2. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that, the backbone network in step 100, that is, the network used in the plaintext branch of the hybrid encryption model, and the operations in the backbone network need to be replaced with operations supported by the selected homomorphic encryption scheme in the ciphertext branch; if the homomorphic encryption scheme supports polynomial operations but does not support the maximum operation, the ReLU activation function can be replaced with a polynomial activation function, and the max pooling layer can be replaced with an average pooling layer.

3. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that, the cross-branch connection channel in step 100: a channel for transmitting the information of the plaintext branch to the ciphertext branch, and the channel is formalized as: where x c,i and x nc,i represent the input tensors of the i-th layer of the ciphertext branch and the i-th layer of the plaintext branch respectively; before calculating the addition of x c,i and the output tensor of the i-th layer of the plaintext branch the shapes of the two need to be unified using the Chop function first.

4. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that, the channel integration layer in step 100: a network layer for combining the channels of the input information of the plaintext branch in the cross-branch connection, and the purpose is to let each channel of the ciphertext branch learn using the information from all channels of the plaintext branch; The channel integration layer is formalized as a matrix whose parameters are learned during training; Adding the channel integration layer to the cross-branch connection channel, the new cross-branch connection is formalized as:

5. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that, The sensitive part of the face image in step 101 is: the part that contains more privacy information or the part that the user wishes to encrypt.

6. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that the remaining part of the face image in step 101 is: the result obtained by filling 0 at the corresponding positions of the sensitive part in the original image.

7. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that In step 102, the knowledge distillation technique is used to enable the hybrid encryption model to learn the feature representation of the trained backbone network. The specific method is as follows: Select the output of a certain intermediate layer of the backbone network and denote it as Denote the output of the corresponding layer of the ciphertext branch as The output of the corresponding layer of the plaintext branch is Then, the intermediate representation of the backbone network is learned by minimizing the following loss function 8. The face recognition method with privacy protection using homomorphic encryption technology according to claim 1, characterized in that In step 103, the knowledge distillation technology is used to enable the hybrid encryption model to learn the output of the trained backbone network. Let the backbone network be the Teacher model, and its output be a T , and the hybrid encryption model be the Student model, and its output be a s , let where the relaxation parameter τ > 1; the hybrid encryption model is trained by minimizing the following loss: where y True is the true label, is the cross-entropy loss function, W MCN are the parameters of the hybrid encryption model, and λ ∈ [0, 1].

9. A computer device, characterized in that: the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the face recognition method with privacy protection using homomorphic encryption technology as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: the computer-readable storage medium stores a computer program for executing the face recognition method with privacy protection using homomorphic encryption technology as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Privacy protection text named entity recognition method and device, equipment and storage medium

    CN113486665A

  • System for secure image recognition

    US20130148868A1