Target detection method based on homomorphic encryption
Through the combination of homomorphic encryption and convolutional neural networks, multithreading technology and polynomial approximation methods are used to perform object detection in the encryption domain, solving the problem of data privacy protection and model performance, and achieving safe and efficient object detection.
Patent Information
- Application Number
- CN202510467280.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively perform object detection of deep learning models without leaking original data, especially when combined with homomorphic encryption, and faces the challenge of nonlinear operation.
The homomorphic encryption algorithm is used to generate the encryption public and private key, encrypt the original image data, and image feature extraction and target positioning are performed under the encryption domain through the pre-trained convolutional neural network. Convolution operations are performed in parallel using multi-threading technology, Taylor series and Chebischev polynomial approximation excitation function are used, average pooling is used instead of maximum pooling, and suitable loss functions are designed to train the model.
It realizes the protection of data privacy without affecting the performance of the model, improves the security of data transmission and processing, and ensures that data privacy is not leaked.
Smart Images

Figure CN120388162A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and information security, and more specifically, to an object detection method based on homomorphic encryption. Background Art
[0002] With the development of artificial intelligence, deep learning models such as convolutional neural networks (CNNs) have been widely applied in fields such as image recognition and object detection. However, in practical applications, due to the involvement of a large amount of sensitive personal or commercial information, how to perform effective model inference without revealing the original data has become an urgent problem to be solved. Homomorphic encryption, as a technology that can directly perform calculations on ciphertexts, provides a new idea for solving this problem. However, it is not easy to combine homomorphic encryption with complex deep learning models. For example, many challenges are encountered when dealing with non-linear operations.
[0003] Therefore, it is of great significance to design an object detection method that can protect data privacy without affecting model performance. Summary of the Invention
[0004] In view of this, the present invention provides an object detection method based on homomorphic encryption. While protecting data privacy, homomorphic encryption can effectively support the inference of deep learning models, realizing object detection based on homomorphic encryption.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An object detection method based on homomorphic encryption, comprising:
[0007] Generating a pair of encrypted public and private keys through a homomorphic encryption algorithm on a local client, encrypting the original image data with the public key to obtain an encrypted image, and sending the encrypted image to a server;
[0008] The server preprocesses the encrypted image, and performs calculations on the preprocessed encrypted image data through a pre-trained convolutional neural network model, extracts image feature values and locates the target area in the encrypted domain, generating an encrypted object detection result, which includes the position information and category information of the target;
[0009] The server returns the encrypted object detection result to the local client;
[0010] The local client decrypts the encrypted object detection result returned by the server using the private key to obtain the final target position and category information.
[0011] Preferably, the convolutional neural network model includes an input layer, a convolutional layer, an activation layer, a pooling layer, a fully connected layer, and an output layer.
[0012] Preferably, the specific processing process of the convolutional layer is as follows:
[0013] The input image is divided into multiple small blocks according to the size of the convolution kernel, and the size of each small block is M1×N1, where M1 and N1 are the height and width of the convolution kernel respectively;
[0014] Each small block and the convolution kernel K are multiplied element by element, and the convolution operation is performed in parallel using multi-threading technology;
[0015] All the results of parallel computing are recombined into the final output matrix, and the shape of the output matrix is (H1 - M1 + 1)×(W1 - N1 + 1)×C, where H1 and W1 are the height and width of the input image, and C is the number of channels.
[0016] Preferably, the activation layer uses Taylor series or Chebyshev polynomials to directly approximate the polynomial representation of the activation function;
[0017] Taylor series approximation of the ReLU function:
[0018]
[0019] where a n is the coefficient of the Taylor series, O is the order of the polynomial, and x n represents the nth power of the variable x;
[0020] Approximation using Chebyshev polynomials:
[0021]
[0022] where T n (x) is the Chebyshev polynomial, and b n is the corresponding coefficient.
[0023] Preferably, the pooling layer adopts the average pooling method. During the average pooling operation, for an input feature map with a size of H×W×C, where H represents the height, W represents the width, and C represents the number of channels, and the pooling layer window size is k×k, the calculation formula for average pooling is:
[0024]
[0025] Sum(l,j,c) = S((l + 1)×k - 1,(j + 1)×k - 1,c) - S(l×k - 1,(j + 1)×k - 1,c) - S((l + 1)×k - 1,j×k - 1,c)
[0026] + S(l×k - 1,j×k - 1,c)
[0027] Among them, passing through the pooling layer does not change the number of channels of the feature map. For the position (l, j, c) of the pixel points in each output feature map, l represents the row index of the output feature map pixel, j represents the column index of the output feature map pixel, and c represents the channel index of the output feature map. Among them, 0 ≤ l < H', 0 ≤ j < W′, 0 ≤ c < C, S represents the cumulative sum array, which is used to store the cumulative sum of each channel in the input feature map, and the round function represents the rounding operation on the calculation result.
[0028] Preferably, each neuron in the fully connected layer is connected to all neurons in the pooling layer, and each connection is represented by a weight value, and each node outputs a weighted sum over the entire pooling layer.
[0029] Preferably, a pair of encrypted public and private keys are generated by the homomorphic encryption algorithm on the local client, where the homomorphic encryption is fully homomorphic encryption.
[0030] Preferably, it further includes the training stage of the convolutional neural network model, including the following steps:
[0031] Step a) Collect and label a set of training image data, and the training image data contains multiple target instances of different categories;
[0032] Step b) Encrypt the training image data using the homomorphic encryption algorithm;
[0033] Step c) Use the encrypted training image data to train the convolutional neural network model so that it can accurately detect the target in the encrypted domain;
[0034] Step d) During the training process, continuously adjust the model parameters to minimize the loss function until the model converges to a satisfactory detection accuracy;
[0035] Step e) Save the trained model for use in actual target detection tasks.
[0036] Preferably, in step c), the training process adopts a loss function optimized for homomorphic encrypted data:
[0037] L combined = ɑ·L noise-aware + β·L secure-aggregate + γ·L approx-gradient
[0038]
[0039] Among them, α, β, γ are hyperparameters used to weigh the importance of different loss terms, and L noise-aware is the loss function based on noise perception, and L secure-aggregateis the secure aggregation loss, L approx-gradient is the approximate gradient loss, N is the number of samples, y i is the true label of the i-th sample, is the predicted label of the i-th sample, n i is the estimated noise level, λ is a hyperparameter used to balance the prediction error and the noise penalty, w i is the weight of the i-th sample.
[0040] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a target detection method based on homomorphic encryption. When using a convolutional neural network for target detection by adjusting the deep learning model structure, the data is encrypted during transmission and processing, improving data security and protecting privacy information without affecting the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0042] Figure 1 is a flowchart of a target detection method based on homomorphic encryption provided by the present invention.
[0043] Figure 2 is a schematic diagram of the convolutional neural network model structure provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0045] The embodiments of the present invention disclose a target detection method based on homomorphic encryption, as Figure 1 shown, including:
[0046] Generating a pair of encrypted public and private keys through a homomorphic encryption algorithm on the local client, encrypting the original image data with the public key to obtain an encrypted image, and sending the encrypted image to the server; where the homomorphic encryption is fully homomorphic encryption.
[0047] The server preprocesses the encrypted image and calculates the preprocessed encrypted image data through a pre-trained convolutional neural network model, extracts the image feature values and locates the target area in the encrypted domain to generate an encrypted target detection result, which contains the location information and category information of the target;
[0048] The server returns the encrypted target detection result to the local client;
[0049] The local client uses the private key to decrypt the encrypted target detection result returned by the server to obtain the final target location and category information.
[0050] Among them, the preprocessing includes but is not limited to resizing, format conversion, etc., to adapt to the input requirements of the subsequent target detection model, and all preprocessing operations are completed in the encrypted domain.
[0051] Since homomorphic encryption only supports addition and multiplication operations, while deep learning models usually involve non-linear operations other than multiplication and addition, such as activation function operations, this makes the direct application of homomorphic encryption more complex. To solve the above problems, the structures of each layer of the deep learning model are adjusted in the homomorphic encryption environment. As Figure 2 shown, the convolutional neural network model includes an input layer, a convolutional layer, an activation layer, a pooling layer, a fully connected layer, and an output layer.
[0052] Among them, the specific processing process of the convolutional layer is:
[0053] The input image is divided into multiple small blocks according to the size of the convolutional kernel, and the size of each small block is M1×N1. M1 and N1 are the height and width of the convolutional kernel respectively, and the input image is the image output after the encrypted image passes through the input layer;
[0054] Perform element-wise multiplication on each small block and the convolutional kernel K, and use multi-threading technology to perform the convolution operation in parallel;
[0055] Recombine all the results of the parallel calculations into the final output matrix, and the shape of the output matrix is (H1 - M1 + 1)×(W1 - N1 + 1)×C, where H1 and W1 are the height and width of the input image, and C is the number of channels.
[0056] In the homomorphic encryption environment, the main challenge faced by the convolution operation is how to perform efficient calculations on encrypted data. The present invention uses multi-threading technology to process different encrypted small blocks in parallel, which can significantly improve the calculation efficiency in the homomorphic encryption environment.
[0057] The activation layer uses the Taylor series or Chebyshev polynomial to directly approximate the polynomial representation of the activation function;
[0058] Taylor series approximation of the ReLU function:
[0059]
[0060] where a n is the coefficient of the Taylor series and O is the order of the polynomial;
[0061] Approximation with Chebyshev polynomials:
[0062]
[0063] where T n (x) is the Chebyshev polynomial and b n is the corresponding coefficient. The coefficients a n or b n of the polynomial can be obtained through numerical integration or other optimization methods.
[0064] In the homomorphic encryption environment, the main challenge faced by the excitation layer is the lack of support for non-linear operations. The present invention converts the non-linear activation function into a series of addition and multiplication operations through two polynomial approximation methods, so that encrypted data can be directly calculated without decryption while maintaining the non-linear ability of the model. This method not only supports additive and multiplicative homomorphic operations, but also reduces noise accumulation, provides flexibility and adjustability, and ensures data privacy and security.
[0065] In the homomorphic encryption environment, special consideration needs to be given to the processing method of encrypted data when implementing the pooling layer. Since max pooling involves comparison operations and homomorphic encryption generally does not support direct comparison, there are certain challenges in implementation. In contrast, average pooling only involves addition and multiplication operations and is easier to implement in the homomorphic encryption environment. Therefore, the pooling layer of the present invention adopts the average pooling method. During the average pooling operation, an input feature map with a size of H×W×C, where H represents the height, W represents the width, and C represents the number of channels, and the pooling layer window size is k×k. The calculation formula for average pooling is:
[0066]
[0067] Sum(l,j,c) = S((l + 1)×k - 1,(j + 1)×k - 1,c) - S(l×k - 1,(j + 1)×k - 1,c) - S((l + 1)×k - 1,j×k - 1,c)
[0068] + S(l×k - 1,j×k - 1,c)
[0069] where the number of channels of the feature map remains unchanged after passing through the pooling layer. For the position (l,j,c) of each pixel point in the output feature map, l represents the row index of the output feature map pixel, j represents the column index of the output feature map pixel, and c represents the channel index of the output feature map, where 0 ≤ l < H'. 0 ≤ j < W′, 0 ≤ c < C, where S represents an array of cumulative sums used to store the cumulative sums of each channel in the input feature map.
[0070] The average pooling process of the present invention has the following advantages:
[0071] Reducing multiplication operations: The number of multiplication operations is reduced by means of accumulation.
[0072] Efficient calculation: The sum within the pooling window can be quickly calculated using the cumulative sum array, thereby improving the calculation efficiency.
[0073] Suitable for homomorphic encryption: All operations in this method are based on addition and subtraction, making it very suitable for the homomorphic encryption environment.
[0074] Each neuron in the fully connected layer is connected to all neurons in the pooling layer, and each connection is represented by a weight value. Each node outputs a weighted sum over the entire pooling layer.
[0075] By making the above adjustments to the neural network structure, the present invention can satisfy the direct application of encrypted data in the convolutional neural network, ensuring the privacy and security of the data.
[0076] In this embodiment, it further includes a convolutional neural network model training stage, which includes the following steps:
[0077] Step a) Collect and label a set of training image data, where the image data contains multiple target instances of different classes;
[0078] Step b) Encrypt the training image data using a homomorphic encryption algorithm;
[0079] Step c) Use the encrypted training image data to train the convolutional neural network model so that it can accurately detect targets within the encrypted domain;
[0080] Step d) During the training process, continuously adjust the model parameters to minimize the loss function until the model converges to a satisfactory detection accuracy;
[0081] Step e) Save the trained model for use in actual target detection tasks.
[0082] In step c), the training process adopts a loss function optimized for homomorphic encrypted data. For a convolutional neural network object detection method in a homomorphic encryption (FHE) environment, it is crucial to design a suitable loss function. Since directly performing complex mathematical operations (such as non-linear activation functions, gradient calculations, etc.) in the FHE environment is very inefficient or infeasible, a loss function that can effectively process encrypted data and work under additive and multiplicative homomorphic operations needs to be designed. The loss function designed in the present invention is as follows:
[0083] L combined = α·L noise-aware + β·L secure-aggregate + γ·L approx-gradient
[0084]
[0085] where α, β, and γ are hyperparameters used to weigh the importance of different loss terms. L noise-aware is a noise-aware loss function, L secureg-aggregate is a secure aggregation loss, L approx-gradient is an approximate gradient loss, N is the number of samples, y i is the true label of the i-th sample, is the predicted label of the i-th sample, n i is the estimated noise level. The estimation of n i can be achieved by analyzing the characteristics of the FHE algorithm. λ is a hyperparameter used to balance the prediction error and the noise penalty. w i is the weight of the i-th sample.
[0086] The loss function of the present invention takes into account the influence of noise, ensures privacy protection, and can be efficiently calculated in the FHE environment.
[0087] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0088] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A target detection method based on homomorphic encryption, characterized in that, Including: Generating a pair of encrypted public and private keys through a homomorphic encryption algorithm on the local client, encrypting the original image data with the public key to obtain an encrypted image, and sending the encrypted image to the server; The server preprocesses the encrypted image, and calculates the preprocessed encrypted image data through a pre-trained convolutional neural network model, extracts image feature values and locates the target area in the encrypted domain, and generates an encrypted target detection result, which contains the location information and category information of the target; The server returns the encrypted target detection result to the local client; The local client decrypts the encrypted target detection result returned by the server using the private key to obtain the final target location and category information.
2. The object detection method based on homomorphic encryption according to claim 1, wherein, The convolutional neural network model includes an input layer, a convolutional layer, an activation layer, a pooling layer, a fully connected layer, and an output layer.
3. The object detection method based on homomorphic encryption according to claim 2, wherein, The specific processing process of the convolutional layer is as follows: Dividing the input image into multiple small blocks according to the size of the convolutional kernel, and the size of each small block is M1×N1, where M1 and N1 are the height and width of the convolutional kernel respectively; Performing element-wise multiplication on each small block and the convolutional kernel K, and using multi-threading technology to perform the convolution operation in parallel; Recombining all the results of parallel computing into the final output matrix, and the shape of the output matrix is (H1 - M1 + 1)×(W1 - N1 + 1)×C, where H1 and W1 are the height and width of the input image, and C is the number of channels.
4. The object detection method based on homomorphic encryption according to claim 2, characterized in that, The activation layer uses Taylor series or Chebyshev polynomials to directly approximate the polynomial representation of the activation function; Taylor series approximates the ReLU function: where a n is the coefficient of the Taylor series, O is the order of the polynomial, and x n represents the n-th power of the variable x; Chebyshev polynomials are approximated: where T n (x) is a Chebyshev polynomial and b n is the corresponding coefficient.
5. The object detection method based on homomorphic encryption according to claim 2, wherein The pooling layer adopts the average pooling method. During the average pooling operation, an input feature map with a size of H×W×C, where H represents the height, W represents the width, and C represents the number of channels, and the pooling layer window size is k×k. The calculation formula for average pooling is: Sum(l,j,c)=S((l + 1)×k - 1,(j + 1)×k - 1,c)-S(l×k - 1,(j + 1)×k - 1,c)-S((l + 1)×k - 1,j×k - 1,c)+S(l×k - 1,j×k - 1,c) Among them, passing through the pooling layer will not change the number of channels of the feature map. For the position (l, j, c) of the pixel points in each output feature map, l represents the row index of the pixels in the output feature map, j represents the column index of the pixels in the output feature map, and c represents the channel index of the output feature map. Among them, 0 ≤ l < H', 0 ≤ j < W′, 0 ≤ c < C. S represents the cumulative sum array, which is used to store the cumulative sum of each channel in the input feature map. The round function represents the rounding operation on the calculation result.
6. The object detection method based on homomorphic encryption according to claim 2, characterized in that, Each neuron in the fully connected layer is connected to all neurons in the pooling layer, and each connection is represented by a weight value, and each node outputs a weighted sum over the entire pooling layer.
7. A target detection method based on homomorphic encryption according to claim 1, characterized in that Generating a pair of encrypted public and private keys through a homomorphic encryption algorithm on the local client, where the homomorphic encryption is fully homomorphic encryption.
8. The object detection method based on homomorphic encryption according to claim 1, wherein It also includes the training stage of the convolutional neural network model, including the following steps: Step a) Collecting and labeling a set of training image data, where the training image data contains multiple target instances of different categories; Step b) Encrypting the training image data using the homomorphic encryption algorithm; Step c) Training the convolutional neural network model using the encrypted training image data so that it can accurately detect targets in the encrypted domain; Step d) During the training process, continuously adjusting the model parameters to minimize the loss function until the model converges to the preset detection accuracy; Step e) Saving the trained model for use in actual target detection tasks.
9. The object detection method based on homomorphic encryption according to claim 8, wherein, In step c), the training process uses a loss function optimized for homomorphically encrypted data: L combined = α·L noise-aware + β·L secure-aggregate + γ·L approx-gradient Among them, α, β, γ are hyperparameters used to weigh the importance of different loss terms, L noise-aware is the noise-aware loss function, L secure-aggregate is the secure aggregation loss, L approx-gradient is the approximate gradient loss, N is the number of samples, y i is the true label of the i-th sample, is the predicted label of the i-th sample, n i is the estimated noise level, λ is a hyperparameter used to balance the prediction error and the noise penalty, w i is the weight of the i-th sample.