Monitoring system based on face recognition and homomorphic privacy security protection
Through the IResNet50 network and CKKS homomorphic encryption algorithm, the network degradation and information leakage of facial recognition technology are solved, and efficient and secure facial feature extraction and recognition are achieved, which are suitable for smart cities, finance, security and home monitoring and other fields.
Patent Information
- Application Number
- CN202510532000.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-18
AI Technical Summary
Existing facial recognition technology has degradation problems caused by the depth of network layers, inefficient training and information leakage risks. Especially in the application in the field of public safety, excessive collection and storage of personal facial information will lead to serious consequences.
The IResNet50 network is used for facial feature extraction, combining deterministic training, reverse residual structure and dynamic hole rate prediction to improve network training speed and feature capture capabilities; the homomorphic encryption algorithm CKKS is used for data encryption and calculation to ensure information security.
It improves the recognition accuracy and information security of facial recognition, adapts to big data application scenarios, reduces the video memory peak and calculation amount of network training, and enhances computing speed and privacy protection.
Smart Images

Figure CN120342629A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a monitoring system based on face recognition and homomorphic privacy security protection. Background Art
[0002] Face recognition is a biometric recognition technology that extracts feature information (e.g., facial features) from a person's face through artificial intelligence technology, synthesizes this information, and performs operations to identify the person's identity. Compared with other biometric recognition technologies, face recognition has its unique advantages: non-contact and not easily noticeable, and it has been widely used in banking systems, access control systems, and mobile payments. However, the face recognition networks adopted in current face recognition technologies, such as the FaceNet network and the InsighFace network, both use convolutional neural networks. As the number of network layers deepens, phenomena such as degradation will occur. In addition, there are also problems such as low network training efficiency and slow convergence speed.
[0003] In addition, current public safety issues are becoming increasingly prominent, and face recognition technology has a wide range of applications in the field of public safety, such as security monitoring in crowded places such as railway stations, airports, shopping malls, and schools. Since an individual's facial information (e.g., face images) is highly correlated with other information of the subject, many enterprises over-collect, store, and use face information. If the information is leaked, very serious consequences will occur.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The content part of the present disclosure is used to briefly introduce concepts that will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose a monitoring system based on face recognition and homomorphic privacy security protection to solve the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure provide a monitoring system based on face recognition and homomorphic privacy security protection, including: a front-end monitoring component, a cloud storage component, and a cloud computing component, where: the front-end monitoring component is configured to, after user authentication is passed, collect face images through a corresponding camera, extract facial features from the collected face images, encrypt the extracted facial features with the public key sent by the cloud storage component, and send the encrypted facial features to the cloud computing component; the cloud storage component is configured to, after authenticating the user login information sent by the front-end monitoring component, send the public key to the front-end monitoring component and the cloud computing component respectively, encrypt the face information in the pre-constructed face information database with the public key, send the encrypted face information to the cloud computing component, receive the encrypted face matching result fed back by the cloud computing component, and combine the encrypted face matching result and a pre-set threshold to determine whether the face image matches the face information in the face information database successfully; the cloud computing component is configured to determine the Euclidean distance between the encrypted facial features sent by the front-end monitoring component and the encrypted face information sent by the cloud storage component through a homomorphic encryption algorithm to obtain a face matching result, and send the encrypted face matching result to the cloud storage component, where the encrypted face matching result is the face matching result encrypted with the public key.
[0008] Optionally, the front-end monitoring component is further configured to: obtain user login information input by the user; send the user login information to the cloud storage component for user authentication.
[0009] Optionally, the front-end monitoring component is further configured to: receive the matching result sent by the cloud storage component and display the matching result.
[0010] Optionally, the cloud storage component is further configured to: decrypt the encrypted face matching result with the private key corresponding to the public key to obtain the face matching result; in response to the face matching result being greater than or equal to the pre-set threshold, generate a matching identifier indicating that the face image matches the face information in the face information database successfully; in response to the face matching result being less than the pre-set threshold, generate a matching identifier indicating that the face image matches the face information in the face information database successfully.
[0011] Optionally, the extracting facial features from the collected face images includes: extracting facial features from the face image through a facial feature extraction network, where the facial feature extraction network adopts an IResNet50 network.
[0012] Optionally, the above-mentioned facial feature extraction network adopts a deterministic training method for model training during the network training phase.
[0013] Optionally, the above-mentioned facial feature extraction network adopts an inverse residual structure and dynamic dilation rate prediction for model training during the network training phase.
[0014] Optionally, the above-mentioned facial feature extraction network adopts an optimizer with mixed precision during the network training phase.
[0015] In a second aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.
[0016] In a third aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0017] The monitoring system based on face recognition and homomorphic privacy security protection according to some embodiments of the present disclosure has the following beneficial effects:
[0018] (1) The facial feature extraction network adopts the IResNet network, which can adapt to deeper network levels compared to conventional convolutional networks and can adapt to problems such as gradient disappearance during network training, thus being more suitable for big data application scenarios (for example, meeting the development scenarios of smart cities with big data processing).
[0019] (2) During the network training phase, a deterministic training method is adopted for model training to lock all sources of randomness, so as to achieve the reproducibility of experiments. The training speed is greatly improved, and the data cleaning accuracy is improved while the false killing rate is reduced by using label verification and a deterministic framework.
[0020] (3) During the network training phase, an inverse residual structure and dynamic dilation rate prediction are adopted, and an optimizer with mixed precision is used, which improves the feature capture ability of the facial feature extraction network, compensates for the additional computational overhead of the dynamic network, greatly reduces the peak video memory during the network training phase, and in addition, while maintaining a low computational amount, the receptive field of the network is twice as large as that of traditional networks.
[0021] (4) The homomorphic encryption algorithm (CKKS (Cheon-Kim-Kim-Song) algorithm) is adopted, which supports floating-point calculations and can adapt to more application environments. By using relinearization, the ciphertext is reduced from a cubic ciphertext to a quadratic ciphertext, improving the calculation speed.
[0022] (5) The monitoring system based on face recognition and homomorphic privacy security protection of the present disclosure has a wide range of application scenarios. For example, in the financial field, it can be applied to mobile payment, online banking, and insurance business for real-name authentication to ensure the security of transactions and user information. Also, in the security field, it can be used for target detection at the scene of major events to ensure event security, and can also be used in areas with concentrated personal privacy such as home cameras to prevent illegal theft. Again, in the field of smart city construction, it can be applied to public transportation, large shopping malls, and the field of autonomous driving perception, making urban construction safer and more intelligent. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0024] Figure 1 is a system architecture diagram of the monitoring system based on face recognition and homomorphic privacy security protection in some embodiments of the present disclosure;
[0025] Figure 2 is a schematic diagram of the network structures of Start ResBlock, Middle ResBlock, and End ResBlock;
[0026] Figure 3 is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0028] It should also be noted that, for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions executed by these devices, modules, or units or their interdependent relationships.
[0030] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0031] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0032] In particular, for operations such as the collection, storage, and use of the user's personal information (such as user login information, face images) involved in this disclosure, before performing the corresponding operations, relevant organizations or individuals shall fulfill obligations including conducting a personal information security impact assessment, fulfilling the obligation of notification to the personal information subject, and obtaining the prior authorization and consent of the personal information subject.
[0033] The following will detail this disclosure with reference to the accompanying drawings and in conjunction with embodiments.
[0034] Reference Figure 1 , which shows the system architecture diagram of a monitoring system based on face recognition and homomorphic privacy security protection according to this disclosure. Among them, the monitoring system includes: a front-end monitoring component, a cloud storage component, and a cloud computing component. Among them:
[0035] The front-end monitoring component is configured to, after the user authentication is passed, collect face images through the corresponding camera, extract facial features from the collected face images, and encrypt the extracted facial features with the public key sent by the above-mentioned cloud storage component, and send the encrypted facial features to the above-mentioned cloud computing component.
[0036] Optionally, the front-end monitoring component is further configured to: obtain the user login information input by the user; send the above-mentioned user login information to the above-mentioned cloud storage component for user authentication.
[0037] Optionally, the above-mentioned front-end monitoring component is further configured to: receive the matching result sent by the above-mentioned cloud storage component and display the matching result.
[0038] Optionally, the extraction of facial features from the collected face images includes: extracting facial features from the above-mentioned face images through a facial feature extraction network, where the above-mentioned facial feature extraction network uses an IResNet50 network.
[0039] Optionally, the above-mentioned face feature extraction network adopts a deterministic training method for model training during the network training phase. Specifically, deterministic training is a way to train machine learning models, and its main feature is to ensure that the same result is obtained every time training is performed under the same initial conditions. It can eliminate randomness and make the model training process repeatable and predictable.
[0040] Optionally, the above-mentioned face feature extraction network adopts an inverted residual structure and dynamic dilation rate prediction for model training during the network training phase.
[0041] Optionally, the above-mentioned face feature extraction network adopts a mixed-precision optimizer during the network training phase.
[0042] Specifically, to better understand the functions of the front-end monitoring component, please refer to the following example for details:
[0043] In the first step, the user will input the corresponding user login information through the front-end monitoring component (that is, obtain the user login information input by the user through the front-end monitoring component).
[0044] For example, the user login information may include: account name, account password.
[0045] In the second step, the front-end monitoring component will send the above-mentioned user login information to the cloud storage component for user authentication.
[0046] In practice, when sending the user login information to the cloud storage component, the user login information can be encrypted to ensure the security of the information transmission process.
[0047] In the third step, when the user identity information is verified, receive the public key sent by the cloud storage component.
[0048] In the fourth step, turn on the camera corresponding to the front-end monitoring component to collect the face image of the user and obtain the face image.
[0049] In the fifth step, perform face feature extraction on the face image through the face feature extraction network (for example, the IResNet50 network).
[0050] In the sixth step, encrypt the extracted face features with the public key sent by the cloud storage component.
[0051] In the seventh step, send the encrypted face features to the cloud computing component.
[0052] In the eighth step, when receiving the matching result sent by the cloud storage component, display the matching result.
[0053] Among them, the matching result indicates whether the face image matches the face information in the face information database.
[0054] The above cloud storage component is configured to, after authenticating the user login information sent by the above front-end monitoring component, send a public key to the above front-end monitoring component and the above cloud computing component respectively, encrypt the face information in the pre-constructed face information database with the above public key, send the encrypted face information to the above cloud computing component, receive the encrypted face matching result fed back by the above cloud computing component, and combine the above encrypted face matching result and a pre-set threshold to determine whether the face image matches the face information in the face information database successfully.
[0055] Optionally, the above cloud storage component is further configured to: decrypt the result of the above encrypted face matching result with the private key corresponding to the above public key to obtain the above face matching result; in response to the above face matching result being greater than or equal to the above pre-set threshold, generate a matching identifier indicating that the face image matches the face information in the face information database successfully; in response to the above face matching result being less than the above pre-set threshold, generate a matching identifier indicating that the face image matches the face information in the face information database successfully.
[0056] Specifically, for a better understanding of the functions of the cloud storage component, please refer to the following example in detail:
[0057] The first step is to, when receiving the user login information sent by the front-end monitoring component, authenticate the user login information.
[0058] In practice, the front-end monitoring component can decrypt the user login information and match it with the user identity information in the pre-constructed user identity information database. When there is user identity information in the user identity information database that matches the decrypted user login information, generate a verification result indicating that the user authentication is passed. When there is no user identity information in the user identity information database that matches the decrypted user login information, generate a verification result indicating that the user authentication fails.
[0059] The second step is to, when the user authentication is passed, send a public key to the above front-end monitoring component and the above cloud computing component respectively.
[0060] The third step is to encrypt the face information in the pre-constructed face information database with the public key and send the encrypted face information to the cloud computing component.
[0061] The fourth step is to, when receiving the encrypted face matching result fed back by the cloud computing component, decrypt the result of the above encrypted face matching result with the private key corresponding to the above public key to obtain the face matching result.
[0062] In the fifth step, in response to the above face matching result being greater than or equal to the above preset threshold, a matching result indicating successful matching between the face image and the face information in the face information database is generated.
[0063] In the sixth step, in response to the above face matching result being less than the above preset threshold, a matching result indicating successful matching between the face image and the face information in the face information database is generated.
[0064] The above cloud computing component is configured to determine the Euclidean distance between the encrypted facial features sent by the above front-end monitoring component and the encrypted face information sent by the above cloud storage component through the fully homomorphic encryption algorithm, obtain the face matching result, and send the encrypted face matching result to the above cloud storage component, where the encrypted face matching result is the face matching result encrypted by the above public key.
[0065] Among them, homomorphic encryption refers to an encryption function that encrypts after performing addition and multiplication operations on the plaintext on a ring, and performs corresponding operations on the ciphertext after encryption, and the results are equivalent. This means that data can be analyzed or calculated by a third party while ensuring privacy. Homomorphic encryption is divided into two types: partial homomorphic and fully homomorphic. Among them, partial homomorphic only supports specific operations, such as addition or multiplication, while fully homomorphic can perform arbitrary calculations. With the in-depth research, the homomorphic encryption algorithm has been optimized, improving the computing efficiency and reducing the resources required for encryption and decryption, making the current homomorphic encryption technology widely used in many fields such as cloud computing, privacy protection, data analysis, Internet of Things, blockchain technology, secure multi-party computing, and identity authentication.
[0066] Specifically, for a better understanding of the functions of the cloud computing component, please refer to the following example in detail:
[0067] In the first step, when receiving the encrypted facial features sent by the front-end monitoring component and the encrypted face information sent by the cloud storage component, through the fully homomorphic encryption algorithm (CKKS algorithm), determine the Euclidean distance between the encrypted facial features sent by the above front-end monitoring component and the encrypted face information sent by the above cloud storage component as the face matching result.
[0068] In the second step, encrypt the face matching result with the public key to obtain the encrypted face matching result.
[0069] Among them, the encrypted face matching result is the face matching result encrypted by the above public key.
[0070] In the third step, send the encrypted face matching result to the cloud storage platform.
[0071] In particular, for the facial feature extraction network, the monitoring system of the present disclosure adopts the IResNet50 network. Among them, the IResNet network is an improvement on the ResNet (Residual Neural Network), which was proposed by Duta et al. in 2020. Among them, the trainable network corresponding to the IResNet network has more than 3000 layers, and among them, the depth is the same but the accuracy is higher. At the same time, the IResNet network can also achieve the effect of increasing the accuracy without increasing the computational amount, and the accuracy has been significantly improved in multiple computer vision tasks (such as image classification, COCO (Common Objects in Context) object detection, action recognition).
[0072] Among them, the ResNet network is composed of a large number of building blocks (residual blocks (ResBlocks)) stacked together. Its core idea is to enable the building blocks to have the ability of identity mapping, that is, the input of the building block is equivalent to the output. Specifically, the definition of each residual block is as follows:
[0073]
[0074] Among them, x [l] represents the input vector of the l-th residual block. x [l+1] represents the output vector of the l-th residual block. ReLU() represents the activation function. represents the learnable residual mapping function. characterizes the weights to be learned corresponding to x [l] . Among them, the residual mapping function characterizes the content to be learned in residual learning. When learning to 0, then x [l+1] = x [l] , thus realizing the identity mapping, that is, the input of the building block (residual block) is equivalent to the output. The size() function is used to determine the feature dimension.
[0075] Among them, when the dimensions of the input vector and the output vector are the same, at this time, When the dimensions of the input vector and the output vector are different, at this time, That is, after converting the dimension of x [l] by Wp [l] and then performing subsequent operations. [l]
[0076] In addition, in terms of the network structure, the IResNet50 network is set with 4 Start ResBlocks, 29 Middle ResBlocks, and 4 End ResBlocks. Among them, a batch normalization layer (BN (Batch Normalization) layer) is set at the end of the End ResBlock, which can effectively suppress the distribution shift of the output features to alleviate the problem of gradient disappearance. Structurally, compared with the traditional residual neural network structure of conv→BN→ReLu, adopting the reverse order of BN→conv→Bn→ReLU→conv→BNde (reverse residual structure) can greatly improve the convergence speed during the network training phase.
[0077] As an example, refer to Figure 2 the schematic diagram of the network structures of the Start ResBlock, Middle ResBlock, and End ResBlock shown in the figure, where: First, for the Start ResBlock, the input vector x [l] is input into conv (convolution layer, convolution kernel 1×1) → BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 3×3) → BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 1×1) → BN (batch normalization layer). Among them, the output of the last BN (batch normalization layer) is superimposed with the input vector x [l] to obtain the output vector x [l+1] . Second, for the MiddleResBlock, the input vector x [l] is input into BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 1×1) → BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 3×3) → BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 1×1). Among them, the output of the last conv (convolution layer, convolution kernel 1×1) is superimposed with the input vector x [l] to obtain the output vector x [l+1] . Then, for the End ResBlock, the input vector x [l] is input into the vector BN (batch normalization layer) → conv (convolution layer, convolution kernel 1×1) → BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 3×3) → BN (batch normalization layer) → ReLU (activation function) → conv (convolution layer, convolution kernel 1×1). Among them, the output of conv (convolution layer, convolution kernel 1×1) is superimposed with the input vector x [l]The input after superposition enters the BN (Batch Normalization layer) → ReLU (activation function) to obtain the output vector x [l+1] 。
[0078] Immediately afterwards, there is a linear bottleneck in the feature extractor of the traditional face recognition network for passing features between layers, which will cause a significant decline in the non-linear expression ability of deep features and an increase in intra-class features. Therefore, in the network training stage of the face feature extraction network of this disclosure, a lightweight multi-layer perceptron is introduced. The principle is to embed a two-layer MLP (Multilayer Perceptron) structure between the face feature extraction network and the classification head (classifier). The corresponding mathematical expression can be as follows:
[0079]
[0080] where x ∈ R 512 , which represents the original feature vector output by the backbone network (face feature extraction network). W1 is the dimensionality reduction matrix, which is used to compress the 512-dimensional features into a 4-dimensional hidden space. W2 is the dimensionality increase matrix, which is used to restore the features to the original dimension. b1 and b2 are model learnable parameters and do not need to be set manually. The model learnable parameters can be automatically optimized and updated through the backpropagation algorithm during the network training stage. In particular, b1 is the hidden layer bias, which is used to compensate for the loss of dimensionality reduction information. b2 is the output bias, which is used to enhance the feature expression ability. BatchNorm() is the parameter momentum, which is used to control the distribution stability of the features. LayerNorm() is the feature dimension (normalization operation) to maintain the consistency of the output feature beam rigidity. By introducing the multi-layer perceptron embedding, the face feature extraction network can be made more computationally efficient and the accuracy of the network can be improved. z (1) 、a (1) 、z (2) are all intermediate variables.
[0081] Furthermore, in the face of some complex environments, for example, when the face occlusion area is large, the recognition rate of the model will decrease significantly, and a fixed dilation rate will cause the receptive field to mismatch with the target size. The receptive field of each traditional residual block is determined by the network depth and the convolution kernel size and is difficult to adapt to targets of different scales. Among them, the traditional receptive field S function expression is as follows:
[0082]
[0083] where k i represents the convolution kernel size of the i-th layer. L represents the number of layers. s j represents the stride of the j-th layer.
[0084] It can be observed that the feature extraction ability of a convolution kernel with a fixed size of the pooling layer for the occluded area is limited. Therefore, the monitoring system of the present disclosure uses a method of dynamic dilated convolution prediction for model training, where the dynamic dilated convolution prediction formula is as follows:
[0085]
[0086] Where F (att) ∈R H×W×C , that is, the feature map after being processed by the spatial attention mechanism, where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map. represents the attention weight value of the c-th channel at the position (i, j) (the attention weight value has been processed by Sigmoid normalization). It is the channel attention weight vector, which is used to measure the contribution of different channels to the dilation rate. b dil ∈R It is the bias term, which is used to adjust the base dilation rate. softplus() is the activation function to ensure that d ij is greater than 0 and the gradient is smooth. Specifically, softplus(x) = ln(1 + e x ), where x is the independent variable. d ij represents the dilation rate at the position (i, j).
[0087] The dynamic convolution calculation function is as follows:
[0088]
[0089] Where K ∈ R (2k+1)(2k+1) , which is the convolution kernel. I ∈ R H×W which is the input feature map. O ij is the value of the output feature map at (i, j). m and n represent the position offsets equivalent to the center within the convolution kernel. Specifically, according to the current feature map (input feature map), the dilation rate d ij corresponding to each position is predicted, and then the sampling coordinates are calculated for each convolution kernel position (m, n), the value of I(x, y) is obtained through interpolation, and then the weighted sum is calculated to obtain the output value O ij . Through this operation, the robustness of the facial feature extraction network for the detection of occluded faces is significantly improved, making it more in line with the actual environmental requirements.
[0090] Finally, during the network training stage of the facial feature extraction network, an optimizer with mixed precision (the gradient magnitude adaptive scaling system in mixed-precision training) is adopted. Specifically, by jointly using two different precision data types (FP16 (half-precision) and FP32 (single-precision)) in 2, the training speed of the network is increased, and the memory consumption is reduced, so that the network can adapt to the operating environments of more different hardware devices.
[0091] Next, for the fully homomorphic encryption algorithm, the monitoring system of the present disclosure adopts the CKKS fully homomorphic encryption algorithm, which supports the calculation of encrypted floating-point numbers and is very suitable for the calculation of face feature matrices. Among them, the core of the CKKS fully homomorphic encryption algorithm is to regard the encrypted noise as part of the approximate calculation error, that is, the decrypted result is directly regarded as an approximation of the original message, and the important message is placed in the MSB (where MSB represents the high-order part (most significant bit) of the encrypted data) to prevent the data from being damaged after calculation, and the error is placed in the LSB (where LSB represents the low-order part (least significant bit) of the encrypted data) to ensure security. If the error noise is small enough compared to the message, then the noise will not damage the message. Before encryption, the message is multiplied by a scaling factor to expand the message, so as to reduce the error caused by adding noise. Thus, message + noise can be used to replace the original message. At the same time, CKKS also uses relinearization and rescaling techniques to reduce the expansion problem during ciphertext calculation, making the ciphertext scale grow linearly instead of exponentially to increase the number of multiplications, and can also eliminate the noise in the LSB bit, similar to a rounding operation. Its complete calculation process is as follows:
[0092] Given the parameter N, define the polynomial ring R = Z[X] / (X N +1), the scaling factor Δ, and the parameter χ that follows a uniform distribution. In particular, the larger the value of the parameter N, the higher the difficulty based on the RLWE (Ring Learning with Errors) problem, and at the same time, the lower the computational efficiency. Given the modulus chain {q i |i = 1, 2...L}, where each q i is the product of prime numbers. When selecting, the first modulus is larger and is used to store the initial noise and the final result. The middle moduli are smaller and are used for intermediate calculations after modulus switching.
[0093] The specific steps are as follows:
[0094] The first step is to generate the public key, private key, relinearization key, and Galois key.
[0095] Among them, randomly sample s belonging to R (polynomial ring), and denote the private key s = (1, s). Sample the random polynomial a = R q , and the error e = χ.
[0096] Specifically, calculate P k0 = (-(a·s + e), a) ∈ R q , where the public key P k = (P k0 , P k1 ).
[0097] Among them, the relinearization key is used to convert the ciphertext from quadratic form to linear form, and its essence is to decompose the private key.
[0098] Specifically, randomly select t (usually taking values of 60 or 30), decompose the private key into multiple components according to the base (2t), for example, obtain components S02, S12... Then, sample the random polynomial a i = R q and the error e i = χ, and construct the relinearization key component rlk i = [a i ·s + e i + 2 t ·s i 2 q . Determine the rotation step k that needs to be supported, where the value range of is 1, 2, 4,..., N / 2. Sample the random polynomial a k = R q and the error e k = χ, and calculate the Galois key component Galoiskey k = [a k ·s + e k - X k ·s]q.
[0099] Step 2, encrypt the plaintext.
[0100] Among them, let m be the plaintext polynomial, randomly sample u = R2, e1, e2 = χ. And calculate in combination with the following formula:
[0101]
[0102] In this way, the ciphertext CT = (c0, c1) ∈ R q 2 .
[0103] Step 3, ciphertext multiplication.
[0104] Among them, let two groups of ciphertexts be C1(c 10 , c 11 ), C2(C 20 , C 21 ), calculate the cubic polynomial CT = (c0, c1, c2) ∈ R q 3 , where:
[0105]
[0106] Step 4, relinearization.
[0107] Among them, convert the cubic ciphertext CT = (c0, c1, c2) ∈ Rq 3 Compress it into the quadratic form (c0', c1'), where:
[0108]
[0109] Among them, c2 is the decomposed component according to the base (2t).
[0110] Step 5, rotation operation.
[0111] Among them, given the ciphertext CT = (c0, c1), the rotated ciphertext is:
[0112] CT' = (c0·X k + c1·GaloisKey k , c1·X k ) mod q.
[0113] Specifically, when decrypting, the plaintext polynomial is rotated by k steps, so the noise increases by c1·e k . The specific function expression is as follows:
[0114] c0' + c1'·s = c0X k + c1·(a k s + e k - X k s) + c1X k s = c0X k + c1e k ≈ X k (c0 + c1s) Step 5, decrypt the ciphertext.
[0115] Use the private key s to recover the plaintext polynomial. The specific function expression is as follows:
[0116] m decrypted = [c0, c1·s] q
[0117] Among them, m decrypted is the recovered plaintext.
[0118] In addition, the correctness verification can refer to the following formula:
[0119] c0 + c1·s = [P ko ·u + e1 + m + (P k1 ·u + e2)·s] q
[0120] = [(-a·s - e)·u + e1 + m + a·u·s + e2·s] q
[0121] = [-e·u + e1 + m + e2·s]q
[0122] Among them, when the noise |-e·u + e1 + e2·s| < Δ / 2, m can be restored through a rounding operation.
[0123] Reference is made below to Figure 3 , which shows a schematic structural diagram of an electronic device (for example, a computing device corresponding to a monitoring system) suitable for use in implementing some embodiments of the present disclosure. Figure 3 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure. As Figure 3 shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any of the above methods. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium, and when the computer program is executed by the processor, the processor can execute any of the above methods. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 3 the structure shown in
[0124] is only a block diagram of a part of the structure related to the solution of the present disclosure and does not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0125] Among them, in one embodiment, the above-mentioned processor is used to run the computer program stored in the memory to implement the functional steps corresponding to each component in the front-end monitoring component, cloud storage component, and cloud computing component included in the monitoring system based on face recognition and homomorphic privacy security protection.
[0126] Embodiments of the present disclosure also provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. The computer program includes program instructions. The method implemented when the program instructions are executed may refer to the various embodiments of the method of the present disclosure above.
[0127] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.
[0128] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0129] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.
Claims
1. A monitoring system based on face recognition and homomorphic privacy security protection, characterized in that, Comprising: A front-end monitoring component, a cloud storage component, and a cloud computing component, where: The front-end monitoring component is configured to, after user authentication is passed, collect face images through a corresponding camera, extract facial features from the collected face images, encrypt the extracted facial features with the public key sent by the cloud storage component, and send the encrypted facial features to the cloud computing component; The cloud storage component is configured to, after authenticating the user login information sent by the front-end monitoring component, send a public key to the front-end monitoring component and the cloud computing component respectively, encrypt the face information in the pre-constructed face information database with the public key, send the encrypted face information to the cloud computing component, receive the encrypted face matching result fed back by the cloud computing component, and combine the encrypted face matching result and a preset threshold to determine whether the face image matches the face information in the face information database successfully; The cloud computing component is configured to determine the Euclidean distance between the encrypted facial features sent by the front-end monitoring component and the encrypted face information sent by the cloud storage component through a fully homomorphic encryption algorithm to obtain a face matching result, and send the encrypted face matching result to the cloud storage component, where the encrypted face matching result is the face matching result encrypted with the public key.
2. The monitoring system according to claim 1, characterized in that The front-end monitoring component is further configured to: Obtain user login information input by the user; Send the user login information to the cloud storage component for user authentication.
3. The monitoring system according to claim 2, characterized in that, The front-end monitoring component is further configured to: Receive the matching result sent by the cloud storage component and display the matching result.
4. The monitoring system according to claim 3, characterized in that The cloud storage component is further configured to: Decrypt the encrypted face matching result with the private key corresponding to the public key to obtain the face matching result; In response to the face matching result being greater than or equal to the preset threshold, generate a matching result indicating that the face image matches the face information in the face information database successfully; In response to the face matching result being less than the preset threshold, generate a matching result indicating that the face image matches the face information in the face information database successfully.
5. The monitoring system according to claim 4, characterized in that, The extracting facial features from the collected face images includes: Extracting facial features from the face image through a facial feature extraction network, where the facial feature extraction network adopts an IResNet50 network.
6. The monitoring system according to claim 5, wherein The facial feature extraction network adopts a deterministic training method for model training in the network training stage.
7. The monitoring system according to claim 6, wherein The facial feature extraction network adopts an inverse residual structure and a dynamic dilation rate prediction method for model training in the network training stage.
8. The monitoring system according to claim 7, wherein The facial feature extraction network adopts a mixed-precision optimizer in the network training stage.
9. An electronic device, characterized in that, Comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the monitoring system as described in any one of claims 1 to 8.
10. A computer-readable medium, characterized in that, A computer program is stored thereon, wherein when the computer program is executed by a processor, it implements the monitoring system according to any one of claims 1 to 8.