Privacy security face recognition model conversion method based on fully homomorphic encryption
By converting the deep learning model into a fully homomorphic encryption model, the problem of insufficient efficiency and accuracy of the face recognition model in the existing technology in the fully homomorphic encryption environment is solved, and the effect of high precision and rapid inference is achieved.
Patent Information
- Application Number
- CN202411974259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-06
AI Technical Summary
The existing privacy and security model identification methods have shortcomings in terms of efficiency and accuracy, especially in a fully homomorphic encryption environment, where the calculation overhead is large, the response time is long, and the data accuracy loss affects the performance of the face recognition system.
By converting the original single-branch model into a multi-branch model and training with unencrypted face data, then performing structural reparameterization adjustment and pruning techniques, the model is finally converted into a quantitative model and compiled into a fully homomorphic encryption model.
While ensuring high precision of the model, it significantly improves the inference speed in a fully homomorphic encryption environment, reduces the number of convolutional channels of the model, reduces the inference time, and improves the practical application performance.
Smart Images

Figure CN120108013A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and privacy protection, and relates to a privacy-safe face recognition model conversion method based on fully homomorphic encryption. Background Art
[0002] With the widespread application of deep learning models, especially in the fields of security access control, finance, and medical care, privacy and security issues have gradually become prominent. In traditional deep learning models, users' private data usually needs to be processed and transmitted, which may lead to the risk of data leakage. Users are worried that their private information may be abused or used improperly. Once leaked, it may be used for unauthorized authentication or surveillance, threatening users' privacy and security.
[0003] Fully homomorphic encryption is a secure and efficient method because it allows computations to be performed on encrypted data without decrypting the data. This means that facial data can be recognized while maintaining the privacy of the data even in an encrypted state. Fully homomorphic encryption can better balance the relationship between privacy and performance. Although fully homomorphic encryption is a promising privacy-safe method, it also has some drawbacks: fully homomorphic encryption often introduces significant computational overhead when performing calculations, which can result in longer response times, especially in large-scale data processing. Encryption and decryption operations may result in a loss of data accuracy, which may affect the performance of face recognition systems.
[0004] Existing privacy-preserving model identification methods usually include technologies such as differential privacy and multi-party secure computation, but they have some problems. Differential privacy introduces noise or ambiguity, which reduces the accuracy of the model identification system. Multi-party secure computation methods require collaboration between multiple parties and may involve complex communication and coordination, increasing the complexity of the system. Existing privacy-preserving model identification methods are usually not efficient and accurate. Summary of the invention
[0005] In view of the above-mentioned technical problems existing in the prior art, the present invention provides a privacy-safe face recognition model conversion method based on fully homomorphic encryption, which improves the reasoning speed in a fully homomorphic encryption environment while ensuring the accuracy of the model.
[0006] The present invention adopts the following technical solution:
[0007] A privacy-safe face recognition model conversion method based on fully homomorphic encryption, comprising the steps of:
[0008] Step 1: Convert the original single-branch model into a multi-branch model and train the multi-branch model using unencrypted face data;
[0009] Step 2: Perform structural reparameterization on the trained multi-branch model to convert it into a single-branch model, thus obtaining a model with fewer model parameters but still higher performance;
[0010] Step 3: Use pruning technology to prune some of the convolution channels of the converted single-branch model and save the pruned model weights;
[0011] Step 4: Convert the pruned single-branch model into the corresponding quantized model, and import the model weights into the quantized model for quantization-aware fine-tuning;
[0012] Step 5: Compile the trained quantization model into a fully homomorphic encryption model to realize the reasoning of encrypted images.
[0013] Furthermore, the present invention proposes a privacy-safe face recognition model conversion method based on fully homomorphic encryption. In step one, the original single-branch model is converted into a multi-branch model, and the multi-branch model is trained using unencrypted face data, that is, two branches are added to each convolutional layer of the single-path model, one of which contains a 1×1 convolutional layer and a BN layer, and the other branch only contains a BN layer.
[0014] For the face dataset, the extended boundary of the face image is resized to make the short side 256 pixels, and a 224×224 pixel area is randomly cropped from each sample, and the cropped area is used as input. The stochastic gradient descent method is used, the weight decay is set to 5e-4, the momentum is 0.9, the initial learning rate is 0.1, and it is divided by 10 at the 30th, 60th and 90th epochs. The training ends after the 100th epoch to obtain the recognition model.
[0015] Furthermore, in the privacy-safe face recognition model transformation method based on fully homomorphic encryption proposed in the present invention, in step 2, the structural reparameterization technology is applied to the trained multi-branch model, and the original convolution of the two added branches is padded to the form of a 3x3 convolution kernel, and the 3×3 convolution layer and the BN layer in each branch are fused into a new 3×3 convolution layer. The convolution kernel of the jth channel of the new convolution and the bias term The formula is as follows:
[0016]
[0017] K is the original convolution kernel, γ is the scaling factor of the BN layer, σ is the standard deviation of the BN layer, μ is the average value of the BN layer, and β is the deviation of the BN layer. After the convolution and BN of each branch are fused, the 3x3 convolutions of the three branches are added together to form a new 3x3 convolution.
[0018] Furthermore, in the privacy-safe face recognition model conversion method based on fully homomorphic encryption proposed in the present invention, in step 3, the pruning technology is used to prune some of the convolution channels of the converted single-branch model. In the single-branch model, a 1x1 convolution compression layer is added after each 3×3 convolution, so as to divide the pruned convolution layer into a memory part and a forgetting part. By introducing the penalty term, the loss function L total The formula is as follows:
[0019] L total (X,Y,θ)=L perf (X,Y,θ)+λP(K)
[0020] Among them, X and Y are data samples and labels respectively, θ is the model parameter, and L perf is the cross entropy loss function, P(K) is the penalty term for the convolution kernel K, and λ is the penalty term coefficient. Different pruning rates are set according to different models and different data sets, and the convolution layer is sparsely trained to make some channels approach zero. The convolution layer is fused with the pruned compression layer to obtain a pruned single-branch model.
[0021] Furthermore, the present invention proposes a method for converting a deep learning model into a fully homomorphic encryption model. In step four, the pruned single-branch model is converted into a corresponding quantized model, and the model weights are imported into the quantized model for quantization-aware fine-tuning. The optimized single-branch model is converted into a corresponding quantized model using the Brevitas library. Specifically, a QuantIdentity layer is added after the input for quantized input, and the convolutional layer and the ReLU layer are converted into their corresponding quantized layers, i.e., quantized convolutional layers and quantized ReLU layers, thereby converting the model into a corresponding quantized model; the model's input data, model weights, and activation functions are converted into integer form. Different quantization bits are set according to different models and different data sets, the model weights are imported into the quantized model, and the quantized model is fine-tuned at a small learning rate to reduce losses during the quantization process.
[0022] Furthermore, the privacy-safe face recognition model conversion method based on fully homomorphic encryption proposed in the present invention provides a representative input data set in step 5, and compiles the quantized model into a fully homomorphic encryption model through the Concrete ML library. The face data set is encrypted and input into the fully homomorphic encryption model to complete the encrypted reasoning.
[0023] The present invention discloses a privacy-safe face recognition model conversion method based on fully homomorphic encryption, which belongs to the field of computer vision and privacy protection. The implementation method is as follows: first, a network model is trained using plaintext data. In order to enhance the ability of feature extraction, if the network structure is a single-branch network, the single-branch structure can be converted into a multi-branch structure for training. If the network structure is already a multi-branch network, it can be directly trained; then, the trained multi-branch model is subjected to structural reparameterization adjustment, and it is converted into a single-branch model to obtain a model with fewer model parameters but still having higher performance; next, the converted single-branch model is pruned using pruning technology to prune some of its convolution channels, and the pruned model weights are saved; then, the pruned single-branch model is converted into a corresponding quantized model, and the model weights are imported into the quantized model for quantization-aware fine-tuning; finally, the trained quantized model is compiled into a fully homomorphic encryption model to realize the reasoning of encrypted images. Compared with the prior art, the present invention has a high model test accuracy and a lower model reasoning time, and the effect is significantly improved on a large data set of images.
[0024] Compared with the prior art, the privacy-safe face recognition model conversion method based on fully homomorphic encryption provided by the present invention has the following beneficial effects:
[0025] 1. The high accuracy of the model is guaranteed. The method of the present invention not only ensures the high accuracy of the model, but also significantly improves the inference speed of the model in a fully homomorphic encryption environment. The experimental results show that in a fully homomorphic encryption environment, the recognition accuracy of the model is not only retained, but even improved in some cases.
[0026] 2. Improved the model's reasoning speed:
[0027] Through structural reparameterization and pruning techniques, the deep learning model is optimized, which significantly reduces the number of convolution channels of the model, thereby reducing the inference time of the model. Experimental results show that the inference speed of the model in a fully homomorphic encryption environment has been significantly improved, greatly improving the practical application performance of the model.
[0028] 3. Expanded the application of fully homomorphic encryption technology in the field of deep learning:
[0029] Existing fully homomorphic encryption technology has not been widely used in the field of deep learning due to its high computational overhead. Through the method of the present invention, the existing deep learning model can be converted into a model suitable for a fully homomorphic encryption environment, thereby expanding the application scenarios of fully homomorphic encryption technology in deep learning, especially in fields where data privacy needs to be protected, such as face recognition in cloud computing.
[0030] 4. The method has good versatility and adaptability:
[0031] The method of the present invention is not only applicable to facial recognition models, but can also be extended to other deep learning models. Through the same technical solution, other models can be optimized and converted to improve their performance and application scope in a fully homomorphic encryption environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Schematic diagram of structural reparameterization in the present invention.
[0033] Figure 2 Schematic diagram of pruning in the present invention.
[0034] Figure 3 This is a schematic diagram of converting a deep learning model into a fully homomorphic encryption model in the present invention.
[0035] Figure 4 It is an execution flow chart of the present invention. DETAILED DESCRIPTION
[0036] The specific implementation of the present invention is further described below in conjunction with specific embodiments:
[0037] like Figure 3 As shown, the present invention comprises the following steps:
[0038] Step 1: Convert the original single-branch model to a multi-branch model as shown in the figure, and use the unencrypted face data to train this multi-branch model, that is, add two branches to each convolution layer of the single-path model, one of which contains a 1×1 convolution layer and a BN layer, and the other branch contains only a BN layer. This can enhance the expressiveness of the model and improve the performance of the model. Train the multi-branch model to improve the overall performance of the model. For the face dataset, adjust the size of the extended boundary of the face image so that the short side is 256 pixels, randomly crop a 224×224 pixel area from each sample, and use the cropped area as input. Use the stochastic gradient descent method, set the weight decay to 5e-4, the momentum to 0.9, the initial learning rate to 0.1, and divide it by 10 at the 30th, 60th and 90th epochs. The training ends after the 100th epoch to obtain the recognition model.
[0039] Step 2: Apply the structure reparameterization technique to the trained multi-branch model, such as Figure 1 , pad the original convolution of the two added branches to a 3x3 convolution kernel, and fuse the 3×3 convolution layer and BN layer in each branch into a new 3×3 convolution layer. The convolution kernel of the jth channel of the new convolution and the bias term The formula is as follows:
[0040]
[0041] K is the original convolution kernel, γ is the scaling factor of the BN layer, σ is the standard deviation of the BN layer, μ is the average value of the BN layer, and β is the deviation of the BN layer. After the convolution and BN of each branch are fused, the 3x3 convolutions of the three branches are added together to form a new 3x3 convolution.
[0042] Step 3: The privacy-safe face recognition model conversion method based on fully homomorphic encryption proposed in the present invention, in step 3, uses pruning technology to prune some of its convolution channels of the converted single-branch model. In the single-branch model, a 1x1 convolution compression layer is added after each 3×3 convolution, so as to divide the pruned convolution layer into a memory part and a forgetting part. The 3×3 convolution is used for memory, and the 1×1 convolution is used for forgetting. By introducing the penalty term, the loss function L total The formula is as follows:
[0043] L total (X,Y,θ)=L perf (X,Y,θ)+λP(K)
[0044] Among them, X and Y are data samples and labels respectively, θ is the model parameter, and L perf is the cross entropy loss function, P(K) is the penalty term for the convolution kernel K, and λ is the penalty term coefficient. Sparse training is performed on the convolution layer, such as Figure 2 , select some channels on the 1×1 convolution compression layer, set the corresponding masks, and make the channels with masks of 0 gradually approach zero. After training, prune the channels that tend to 0 on the 1×1 convolution compression layer. Finally, fuse the convolution layer with the pruned 1×1 convolution compression layer to obtain the pruned single-branch model. Tables 1 and 2 show the comparison of the number of convolution channels before and after model pruning on three small datasets. Table 3 shows the comparison of the number of convolution channels before and after model pruning on two large face datasets.
[0045] Layer Index Number of convolution channels per layer of the original model Number of convolution channels per layer on MNSIT 1 8 7 2 8 8 3 16 10 4 16 16
[0046] Table 1 Changes in the number of convolution channels of the 4-layer convolutional model before and after pruning on the MNSIT dataset
[0047] Layer Index Original Model Cifar10 Cifar100 1 64 13 38 2 64 41 41 3 128 120 123 4 128 109 128 5 256 246 256 6 256 226 256
[0048] Table 2 Changes in the number of convolution channels of the 6-layer convolutional model before and after pruning on the Cifar10 and CIfar100 datasets
[0049] Layer Index Original Model CASIA-WebFace MS-Celeb-1M 1 64 15 20 2 64 31 29 3 64 46 53 4 128 97 102 5 128 120 121 6 256 185 202 7 256 241 242 8 1280 1280 1280
[0050] Table 3 Changes in the number of convolution channels of the 8-layer convolutional model before and after pruning on the CASIA-WebFace and MS-Celeb-1M datasets
[0051] Step 4: The privacy-safe face recognition model conversion method based on fully homomorphic encryption proposed by the present invention, in step 4, the pruned single-branch model is converted into a corresponding quantized model, and the model weight is imported into the quantized model for quantization-aware fine-tuning. The Brevitas library is used to convert the optimized single-branch model into a corresponding quantized model, specifically, adding a QuantIdentity layer after the input for quantized input, converting the convolutional layer and the ReLU layer into their corresponding quantized layers, i.e., quantized convolutional layer and quantized ReLU layer, so as to convert the model into a corresponding quantized model, which is specifically corresponding to Table 4.
[0052] Torch Layer Quantization layer Conv2d QuantConv2d ReLU QuantReLU AvgPool2d AvgPool2d+QuantIdentity
[0053] Table 4 Torch layer corresponding to the quantization layer:
[0054] That is, the quantized convolution layer QuantConv2d and the quantized ReLU layer QuantReLU are used to convert the pruned model into the corresponding quantized model; the input data, model weights and activation functions of the model are converted into integer form, and the loss in the quantization process is reduced through fine-tuning. The quantization-aware fine-tuning adopts the stochastic gradient descent method, the weight decay is set to 5e-4, the momentum is 0.9, the initial learning rate is 0.00006, and it is divided by 10 at the 25th epoch. The training ends after the 30th epoch to obtain the quantized model.
[0055] Step 5: Convert the quantized model into a fully homomorphic encryption model to realize the reasoning of encrypted face images.
[0056] Provide a representative input dataset and compile the quantized model into a fully homomorphic encrypted model through the Concrete ML library. Concrete ML provides a function for compiling a quantized model into its corresponding fully homomorphic encrypted model. Use 1,000 face images as a representative input dataset to compile the quantized model. We compare the performance of the original method and our proposed method, as shown in Tables 5 and 6.
[0057] Dataset The accuracy of the original method Our method has a high accuracy MNIST 98.70% 98.71% CIFAR10 90.00% 90.83% CIFAR100 69.00% 67.11% CASIA-WebFace 57.12% 74.41% MS-Celeb-1M 71.12% 85.87%
[0058] Table 5 Comparison of accuracy of fully homomorphic encryption models
[0059] Dataset Original method reasoning time Our method inference time MNIST 5072 1268 CIFAR10 25000 6763 CIFAR100 25000 10135 CASIA-WebFace 58230 8617 MS-Celeb-1M 58598 8291
[0060] Table 6 Comparison of reasoning time of fully homomorphic encryption models.
Claims
1. A privacy-safe face recognition model transformation method based on fully homomorphic encryption, characterized in that: The following steps are involved: Step 1: Convert the original single-branch model into a multi-branch model and train the multi-branch model using unencrypted face data; Step 2: Re-parameterize the structure of the trained multi-branch model and convert it into a single-branch model to obtain a model with fewer model parameters but still higher performance; Step 3: Use pruning technology to prune some of the convolution channels of the converted single-branch model and save the pruned model weights; Step 4: Convert the pruned single-branch model into the corresponding quantized model, and import the model weights into the quantized model for quantization-aware fine-tuning; Step 5: Compile the trained quantization model into a fully homomorphic encryption model to realize the reasoning of encrypted images.
2. The method according to claim 1, characterized in that The specific process of step one is as follows: Two branches are added to each convolutional layer of the single-branch model, one of which contains a 1×1 convolutional layer and a BN layer; the other branch only contains a BN layer; a multi-branch model with stronger feature extraction capability is obtained, and the multi-branch model is trained for face recognition.
3. The method according to claim 2, characterized in that The specific process of step 2 is as follows: The original convolution of the two added branches is padded with a 3x3 convolution kernel. The 3×3 convolution layer and the BN layer in each branch are fused into a new 3×3 convolution layer. In this way, the entire multi-branch model is converted into a single-branch model, and a smaller learning rate is used for fine-tuning after the fusion is completed. The convolution kernel of the jth channel of the new convolution and the bias term The formula is as follows: K is the original convolution kernel, γ is the scaling factor of the BN layer, σ is the standard deviation of the BN layer, μ is the average value of the BN layer, and β is the deviation of the BN layer. After the convolution and BN of each branch are fused, the 3x3 convolutions of the three branches are added together to form a new 3x3 convolution.
4. The method according to claim 1, characterized in that: The specific process of step three is as follows: Different pruning rates are set according to different models and different data sets. In the single-branch model, the face data set is used to perform sparse training on the model, so that some compression layer channels are close to zero. The convolution layer is fused with the pruned compression layer to obtain the pruned single-branch model. And save the model weights.
5. The method according to claim 4, characterized in that In step 3, in the single-branch model, a 1x1 convolution compression layer is added after each 3×3 convolution to divide the pruned convolution layer into a memory part and a forgetting part; by introducing the penalty term, the loss function L total The formula is as follows: L total (X, Y, θ)=L perf (X, Y, θ)+λP(K) Among them, X and Y are data samples and labels respectively, ι is the model parameter, and L perf is the cross entropy loss function, P(K) is the penalty term for the convolution kernel K, and λ is the penalty term coefficient. Different pruning rates are set according to different models and different data sets. The convolution layer is sparsely trained to make some channels approach zero. The convolution layer is fused with the pruned compression layer to obtain a pruned single-branch model.
6. The method according to claim 1, characterized in that The specific process of step 4 is as follows: The pruned single-branch model is converted into a corresponding quantized model; different quantization bits are set according to different models and different data sets, the model weights are imported into the quantized model, and the quantized model is fine-tuned with a small learning rate to reduce the loss in the quantization process.
7. The method according to claim 1, characterized in that The specific process of step five is as follows: Provide a representative input dataset and compile the quantized model into a fully homomorphic encryption model through the Concrete ML library. Encrypt the face dataset and input it into the fully homomorphic encryption model to complete encrypted reasoning.