Face recognition method based on deep learning in cloud computing scene

By employing OpenCV, MTCNN, heteromorphic affine transformation, and RSA homomorphic encryption, the problems of facial recognition accuracy and privacy protection in cloud computing scenarios are solved, achieving efficient and secure facial information processing and recognition.

CN120954068APending Publication Date: 2025-11-14CHINA NAT BUILDING MATERIALS TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511080461.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In cloud computing scenarios, traditional facial recognition technology suffers from performance degradation in complex situations such as facial occlusion, changes in expression, scars, birthmarks, tattoos, and missing parts. Furthermore, it lacks effective privacy protection mechanisms, leading to the risk of leakage of sensitive personal information.

Method used

We employ OpenCV preprocessing, MTCNN keypoint localization, heterogeneous affine transformation, mutation similarity coefficient algorithm, and RSA homomorphic encryption technology, combined with deep learning methods, to perform face image preprocessing, keypoint localization, feature extraction, alignment, and encryption to ensure recognition accuracy and privacy protection.

Benefits of technology

It significantly improves the accuracy and robustness of facial recognition, enabling efficient and secure facial information comparison and storage, and protecting user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954068A_ABST
    Figure CN120954068A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep learning, in particular to a face recognition method based on deep learning in a cloud computing scene. The method comprises the following steps: S1, collecting face image data, and preprocessing the image data through OpenCV; s2, positioning key points of a human face based on an MTCNN large model, and decomposing the human face into different parts; s3, performing feature extraction on the key points; s4, based on a key point detection result, performing face alignment by adopting a variation affine transformation algorithm; s5, performing face comparison according to the features of the integrated parts and the facial features, and calculating a comprehensive score based on a variation similarity coefficient algorithm; s6, generating a public key and a private key based on an RSA algorithm; s7, encrypting the face information by adopting a homomorphic encryption algorithm based on the public key; and after encryption is completed, the data are stored to a large cloud computing model, so that flexible comparison of subsequent repeated data is facilitated. The application of the homomorphic encryption technology ensures the security of the face information in the storage and transmission process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and more specifically, to a deep learning-based face recognition method in a cloud computing scenario. Background Technology

[0002] With the rapid development of computer vision and deep learning technologies, facial recognition technology has made significant progress. These technologies have a wide range of applications, including security monitoring, identity verification, and personalized recommendations. However, with the widespread use of facial recognition technology, how to achieve efficient and accurate facial recognition while protecting personal privacy has become an important research topic.

[0003] Traditional deep learning-based face recognition solutions, while achieving high accuracy, may suffer performance degradation when handling complex scenarios involving facial occlusion, expression changes, scars, birthmarks, tattoos, and missing parts of the face. Furthermore, traditional solutions often lack effective privacy protection mechanisms, potentially leading to the leakage of sensitive personal information during data collection, storage, and processing. In summary, this paper proposes a deep learning-based face recognition method for cloud computing scenarios. Summary of the Invention

[0004] The purpose of this invention is to provide a deep learning-based face recognition method for cloud computing scenarios, addressing the issues raised in the background art where recognition performance may be affected by complex scenarios such as facial occlusion, expression changes, scars, birthmarks, tattoos, and missing parts. It also addresses the problem that traditional solutions often lack effective privacy protection mechanisms, leading to the risk of leakage of sensitive personal information during data collection, storage, and processing.

[0005] To achieve the above objectives, the present invention aims to provide a deep learning-based face recognition method in a cloud computing scenario, comprising the following steps:

[0006] S1, S1, Collect face image data and preprocess the image data using OpenCV;

[0007] S2. Based on the MTCNN large model, locate the key points of the face and decompose the face into different parts to facilitate feature extraction;

[0008] S3. Extract features from key points;

[0009] S4. Based on the key point detection results, face alignment is performed using the heterogeneous affine transformation algorithm;

[0010] S5. Based on the combined features of each part and facial features, perform face comparison and calculate a comprehensive score based on the variation similarity coefficient algorithm.

[0011] S6. Generate public and private keys based on the RSA algorithm;

[0012] S7. Facial information is encrypted using a homomorphic encryption algorithm based on a public key; after encryption, it is stored in a cloud computing big data model to facilitate flexible comparison of duplicate data in the future.

[0013] As a further improvement to this technical solution, in S1, OpenCV is used to perform cropping, scaling, rotation correction, grayscale conversion, and histogram equalization on the original image data in sequence.

[0014] As a further improvement to this technical solution, in S2, the MTCNN large model locates the key points of the face based on a key point localization algorithm, which is used to obtain the coordinate information of various parts of the face. The key point localization algorithm is specifically as follows:

[0015] k = FC(z) = Wz + b;

[0016] In the formula, k represents a 2L-dimensional vector representing the coordinates of all keypoints, L is the number of keypoints, and each keypoint has two coordinates; z represents the input feature vector; W represents a 2L×M weight matrix, where M is the dimension of the input feature vector; b represents a 2L-dimensional bias vector; and FC(z) represents a fully connected layer.

[0017] As a further improvement to this technical solution, in step S3, feature extraction of key points and the introduction of a weighted mask considering facial differences involve the following steps:

[0018] A weighted mask is introduced, and for each cropped sub-image, a convolution operation is applied to extract features, specifically:

[0019]

[0020] In the formula, I represents the cropped sub-image; K represents the convolution kernel; (x, y) represents the coordinates on the output feature map; F represents the size of the convolution kernel; m and n represent the coordinates on the convolution kernel; Mask represents the weight mask; Mask(x+m, y+n) represents the value of the weight mask at position (x+m, y+n).

[0021] The activation function is applied to the output of the convolutional layer, specifically:

[0022] ReLU(z) = max(0, z);

[0023] In the formula, ReLU represents the modified linear unit; z represents the input to the activation function; and max represents the maximum value function.

[0024] Pooling is applied to the activated feature map to reduce the feature dimensionality and improve the spatial invariance of the features, specifically:

[0025] MaxPool(I)(x, y) = max m,n I(x+q, y+p);

[0026] In the formula, I represents the cropped sub-image; (x, y) represents the coordinates on the output feature map; q and p represent the coordinates on the pooling window;

[0027] Finally, the output of the pooling layer is flattened into a vector and mapped to the final feature representation through a fully connected layer, specifically:

[0028] FC(z) = W·Mask(z) + b;

[0029] In the formula, W represents the weight matrix, b represents the bias vector, z represents the input vector, Mask represents the weight mask, and Mask(z) represents a weight mask vector with the same dimension as the input vector z.

[0030] As a further improvement to this technical solution, in step S4, the face alignment based on the coordinates of key points is performed using a heterogeneous affine transformation algorithm, which is used to calculate the weights of facial features, specifically:

[0031]

[0032] In the formula, x and y represent the coordinates of a point in the original image; x′ and y′ represent the coordinates of a point in the transformed image; a represents the scaling parameter of the image; d represents the rotation parameter of the image; b represents the shearing along the x-axis of the image; c represents the shearing along the y-axis of the image; e represents the translation along the x-axis of the image; and f represents the translation along the y-axis of the image.

[0033] As a further improvement to this technical solution, in step S5, the variation similarity coefficient algorithm is used to calculate the difference between facial features and standardized facial features. The specific steps involved are as follows:

[0034] S1.1 Extract the standardized facial feature vector v1 as the standardized parameter of the facial data;

[0035] S1.2. Convert the non-zero element feature vectors in the feature vectors into sets, and calculate a comprehensive score based on the mutation similarity coefficient algorithm to measure the similarity between the two sets.

[0036] As a further improvement to this technical solution, in S1.1, the feature extraction of the standardized face feature vector v1 is specifically as follows:

[0037]

[0038] In the formula, v1 represents the standardized face feature vector; W represents the weight matrix; b represents the bias vector; z represents the input vector; and α represents the learning rate parameter.

[0039] As a further improvement to this technical solution, the variation similarity coefficient algorithm in S1.2 is specifically as follows:

[0040]

[0041] In the formula, Jaccard represents the variation similarity coefficient; v1 represents the face normalization feature vector; v2 represents the face recognition feature vector; i represents the index in the feature vector; v 1i This represents the value of the i-th element in the feature vector v1; v 2i This represents the value of the i-th element in the feature vector v2.

[0042] As a further improvement to this technical solution, the steps involved in generating the public and private keys based on RSA in step S6 are as follows:

[0043] Choose two large prime numbers p and q, and calculate their product N, specifically:

[0044] N = p × q;

[0045] In the formula, p and q represent random prime numbers; N represents the modulus;

[0046] The Euler totient function is calculated based on the modulus N, specifically as follows:

[0047] φ(N)=(p-1)×(q-1);

[0048] In the formula, φ(N) represents the Euler totient function; p and q represent random prime numbers;

[0049] A fixed prime number is randomly selected, and the private key exponent is calculated based on Euler's totient function, specifically as follows:

[0050] g×h≡1modφ(N);

[0051] In the formula, g represents the private key exponent; h represents the public key exponent; φ(N) represents the Euler totient function;

[0052] A public key consists of two parts: the modulus N and the public key exponent h. Therefore, a public key can be represented as:

[0053] pk = (N, h);

[0054] In the formula, pk represents the public key; N represents the modulus; and h represents the public key exponent.

[0055] The private key also consists of two parts: the modulus N and the private key exponent g. Therefore, the private key can be represented as:

[0056] sk = (N, g);

[0057] In the formula, sk represents the private key; N represents the modulus; and g represents the private key exponent.

[0058] As a further improvement to this technical solution, in step S7, the facial information data extracted from features is encrypted using a homomorphic encryption algorithm, specifically as follows:

[0059] c = Enc(pk, mc);

[0060] In the formula, c represents the encrypted ciphertext; Enc represents the encryption function; pk represents the public key; and mc represents the facial information data extracted from features.

[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0062] 1. This deep learning-based face recognition method in a cloud computing scenario significantly improves the accuracy and robustness of face recognition by combining OpenCV preprocessing, MTCNN keypoint localization, feature extraction, and affine transformation. Specifically, OpenCV's preprocessing function ensures the quality of the input image, laying a solid foundation for subsequent steps. The use of the large MTCNN model achieves high-precision facial keypoint localization, providing accurate coordinate information for feature extraction. The feature extraction method introducing weighted masks effectively considers facial differences, enhancing the model's adaptability to different facial features. The application of the affine transformation algorithm achieves precise face alignment, further improving the accuracy of feature extraction.

[0063] 2. This deep learning-based face recognition method in a cloud computing scenario achieves efficient comparison and secure storage of facial features by introducing a mutation similarity coefficient algorithm and homomorphic encryption technology, which is not available in existing technologies. The use of the mutation similarity coefficient algorithm provides a novel feature comparison method that can more accurately measure the similarity between facial features, thereby improving recognition accuracy. The application of homomorphic encryption technology ensures the security of facial information during storage and transmission, protecting user privacy even in a cloud computing environment. Attached Figure Description

[0064] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation

[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] Example:

[0067] Please see Figure 1 As shown, this embodiment provides a face recognition method based on deep learning in a cloud computing scenario, including the following steps: S1, collecting face image data and preprocessing the image data using OpenCV; OpenCV is used to perform cropping, scaling, rotation correction, grayscale conversion, and histogram equalization on the original image data in sequence.

[0068] OptiCV is an open-source computer vision and machine learning software library that provides a series of programming functions designed to help developers quickly implement computer vision applications. OptiCV is a powerful library widely used in robotics, autonomous vehicles, security monitoring, medical image analysis, augmented reality, and other fields. Due to its open-source and cross-platform nature, OptiCV has become one of the most popular tools in the field of computer vision.

[0069] S2. Based on the MTCNN large-scale model, key points of the face are located, and the face is decomposed into different parts to facilitate feature extraction. Key point location is based on a key point localization algorithm to obtain the coordinate information of various parts of the face. The MTCNN large-scale model is based on a three-layer stepwise network, and the output of each step serves as the input of the next step. The specific steps involved are as follows:

[0070] The purpose of P-Net is to generate candidate face regions and initially locate key points of the face. Its mathematical expression can be summarized as follows:

[0071] f P (I) = CNN(I; θ) P );

[0072] In the formula, f P The P-Net prediction function is represented by I; the input image is represented by CNN; and θ represents the convolutional neural network. P This represents the network parameters of P-Net;

[0073] The output of P-Net includes predicted bounding boxes for the face region, confidence scores, and preliminary keypoint locations, which can be represented by the following expression:

[0074]

[0075] In the formula, bi represents the coordinates of the i-th predicted box; ci represents the confidence score of the i-th predicted box; ki represents the initial keypoint position of the i-th predicted box; N represents the number of predicted boxes; and i represents the index coefficient.

[0076] After receiving the candidate regions generated by P-Net, R-Net further refines them and corrects the key point positions. Its mathematical expression can be summarized as follows:

[0077] f R (br,k;θ R =CNN(br, k; θ) R );

[0078] In the formula, f R θ represents the prediction function of R-Net; br represents the candidate region output by P-Net; k represents the initial keypoint location output by P-Net; θ represents the prediction function of R-Net. R Represents the network parameters of R-Net;

[0079] The output of R-Net is a more accurate prediction and correction of face region keypoint locations, which can be represented by the following expression:

[0080]

[0081] In the formula, b′ i c′ represents the coordinates of the i-th refined prediction box; i k′ represents the confidence level of the i-th refined prediction box. i This represents the location of the i-th refined keypoint; M represents the number of refined prediction boxes.

[0082] O-Net is the last network in the MTCNN model. Its task is to output the final face detection results and precise keypoint locations. Its mathematical expression can be summarized as follows:

[0083] f O (b′,k′;θ O ) = CNN(br′, k′; θ O );

[0084] In the formula, f O θ represents the prediction function of O-Net; b′ represents the refined candidate region output by R-Net; k′ represents the corrected key point location output by R-Net; θ0 represents the network parameters of O-Net.

[0085] The output of O-Net is the final face detection result, including the face's location, size, and precise keypoint locations, which can be represented by the following expression:

[0086]

[0087] In the formula, b″ i c″ represents the coordinates of the i-th final predicted bounding box. i k″ represents the confidence level of the i-th final predicted box; i This represents the precise keypoint location of the i-th final predicted bounding box; K represents the number of final predicted bounding boxes.

[0088] The key point localization algorithm is specifically as follows:

[0089] k = FC(z) = Wz + b;

[0090] In the formula, k represents a 2L-dimensional vector representing the coordinates of all keypoints, L is the number of keypoints, and each keypoint has two coordinates; z represents the input feature vector; W represents a 2L×M weight matrix, where M is the dimension of the input feature vector; b represents a 2L-dimensional bias vector; and FC(z) represents a fully connected layer.

[0091] S3. Feature extraction is performed on key points, and a weighted mask is introduced to account for facial differences. The steps involved are as follows:

[0092] A weighted mask is introduced, and for each cropped sub-image, a convolution operation is applied to extract features, specifically:

[0093]

[0094] In the formula, I represents the cropped sub-image; K represents the convolution kernel; (x, y) represents the coordinates on the output feature map; F represents the size of the convolution kernel; m and n represent the coordinates on the convolution kernel; Mask represents the weight mask; Mask(x+m, y+n) represents the value of the weight mask at position (x+m, y+n).

[0095] The activation function is applied to the output of the convolutional layer, specifically:

[0096] ReLU(z) = max(0, z);

[0097] In the formula, ReLU represents the modified linear unit; z represents the input to the activation function; and max represents the maximum value function.

[0098] Pooling is applied to the activated feature map to reduce the feature dimensionality and improve the spatial invariance of the features, specifically:

[0099] MaxPool(I)(x, y) = max m,n I(x+q, y+p);

[0100] In the formula, I represents the cropped sub-image; (x, y) represents the coordinates on the output feature map; q and p represent the coordinates on the pooling window;

[0101] Finally, the output of the pooling layer is flattened into a vector and mapped to the final feature representation through a fully connected layer, specifically:

[0102] FC(z) = W·Mask(z) + b;

[0103] In the formula, W represents the weight matrix, b represents the bias vector, z represents the input vector, Mask represents the weight mask, and Mask(z) represents a weight mask vector with the same dimension as the input vector z.

[0104] Facial differentiation includes differentiations manifested by facial occlusion, expression changes, scars, birthmarks, tattoos, and missing parts. Mask is introduced to assign weights to facial differentiations. Feature extraction is performed on facial information based on facial differentiations, and the final output is flattened into a feature vector.

[0105] S4. Based on the keypoint detection results, a heterogeneous affine transformation algorithm is used for face alignment; based on the coordinates of the keypoints, a heterogeneous affine transformation algorithm is used for face alignment to calculate the weights of facial features, specifically:

[0106]

[0107] In the formula, x and y represent the coordinates of a point in the original image; x′ and y′ represent the coordinates of a point in the transformed image; a represents the scaling parameter of the image; d represents the rotation parameter of the image; b represents the shearing along the x-axis of the image; c represents the shearing along the y-axis of the image; e represents the translation along the x-axis of the image; and f represents the translation along the y-axis of the image.

[0108] The heteromorphic affine transformation algorithm performs scaling, rotation, shearing, and translation operations on parts of the original image to align the faces of each part, and then generates a new blurred image in the original image data with coordinates x′ and y′.

[0109] S5. Based on the combined features of various body parts and facial features, perform face comparison and calculate a comprehensive score using the variation similarity coefficient algorithm. The variation similarity coefficient algorithm is used to calculate the difference between facial features and standardized facial features. The specific steps involved are as follows:

[0110] A standardized facial feature vector v1 is extracted and used as a standardization parameter for the facial data. Specifically, the standardized facial feature vector v1 is extracted as follows:

[0111]

[0112] In the formula, v1 represents the standardized face feature vector; W represents the weight matrix; b represents the bias vector; z represents the input vector; and α represents the learning rate parameter.

[0113] The non-zero elements in the feature vector are converted into sets. A comprehensive score is calculated based on the mutation similarity coefficient algorithm to measure the similarity between the two sets. The mutation similarity coefficient algorithm is as follows:

[0114]

[0115] In the formula, Jaccard represents the variation similarity coefficient; v1 represents the face normalization feature vector; v2 represents the face recognition feature vector; i represents the index in the feature vector; V 1i This represents the value of the i-th element in the feature vector v1; v 2i This represents the value of the i-th element in the feature vector v2.

[0116] Before calculating the similarity coefficient of variation, v1 and v2 are converted into a set of non-zero element indices;

[0117] The result calculated using the mutation similarity coefficient algorithm is a value between 0 and 1, which represents the similarity between two sets. Specifically:

[0118] A result of 0 indicates that the two sets are completely dissimilar, meaning they have no common elements.

[0119] A result of 1 indicates that the two sets are completely identical, meaning that all their elements are the same.

[0120] Results between 0 and 1 indicate that the two sets are partially similar; the closer the value is to 1, the higher the similarity; the closer the value is to 0, the lower the similarity.

[0121] S6. Generate public and private keys based on the RSA algorithm. The steps involved are as follows:

[0122] Choose two large prime numbers p and q, and calculate their product N, specifically:

[0123] N = p × q;

[0124] In the formula, p and q represent random prime numbers; N represents the modulus;

[0125] The Euler totient function is calculated based on the modulus N, specifically as follows:

[0126] φ(N)=(p-1)×(q-1);

[0127] In the formula, φ(N) represents the Euler totient function; p and q represent random prime numbers;

[0128] A fixed prime number is randomly selected, and the private key exponent is calculated based on Euler's totient function, specifically as follows:

[0129] g×h≡1modφ(N);

[0130] In the formula, g represents the private key exponent; h represents the public key exponent; φ(N) represents the Euler totient function;

[0131] A public key consists of two parts: the modulus N and the public key exponent h. Therefore, a public key can be represented as:

[0132] pk = (N, h);

[0133] In the formula, pk represents the public key; N represents the modulus; and h represents the public key exponent.

[0134] The private key also consists of two parts: the modulus N and the private key exponent g. Therefore, the private key can be represented as:

[0135] sk = (N, g);

[0136] In the formula, sk represents the private key; N represents the modulus; and g represents the private key exponent.

[0137] S7. Facial information is encrypted using a homomorphic encryption algorithm based on a public key; after encryption, it is stored in a large cloud computing model (cloud) for flexible comparison of subsequent duplicate data;

[0138] Perform private set intersection (PSI) and union (PSU) cardinality operations in the cloud, returning |S query ∩S target |and|S query ∪S target The encrypted result is used by the client to decrypt and calculate the Jaccard coefficient:

[0139]

[0140] In the formula, S query S represents the set of faces to query; target Represents the set of target facial features; Jaccard represents the variation similarity coefficient;

[0141] Facial information data extracted based on features is encrypted using a homomorphic encryption algorithm, specifically:

[0142] c = Enc(pk, mc);

[0143] In the formula, c represents the encrypted ciphertext; Enc represents the encryption function; pk represents the public key; and mc represents the facial information data extracted from features.

[0144] The flexible comparison specifically refers to:

[0145] The client extracts the non-zero index set S of the queried facial features. query ;

[0146] Cloud-based computing using the PSI protocol | S query ∩S target |, PSU protocol calculation |S query ∪S target |(S target (A set of encryption features for the database);

[0147] The client decrypts the cardinality of the intersection and union sets and calculates the Jaccard coefficients;

[0148] Homomorphic encryption is a special form of encryption that allows computation to be performed directly on encrypted data without decryption. This means that data can be processed and analyzed while remaining encrypted. This is particularly useful for data analytics and machine learning in cloud computing, as it allows users to perform computations using cloud services without exposing the raw data.

[0149] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A deep learning-based face recognition method in a cloud computing scenario, characterized in that: Includes the following steps: S1. Collect face image data and preprocess the image data using OpenCV; S2. Based on the MTCNN large model, locate the key points of the face and decompose the face into different parts to facilitate feature extraction; S3. Extract features from key points; S4. Based on the key point detection results, face alignment is performed using the heterogeneous affine transformation algorithm; S5. Based on the combined features of each part and facial features, perform face comparison and calculate a comprehensive score based on the variation similarity coefficient algorithm. S6. Generate public and private keys based on the RSA algorithm; S7. Facial information is encrypted using a homomorphic encryption algorithm based on a public key; after encryption, it is stored in a cloud computing big data model to facilitate flexible comparison of duplicate data in the future.

2. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In S1, OpenCV is used to perform cropping, scaling, rotation correction, grayscale conversion, and histogram equalization on the original image data in sequence.

3. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In S2, the MTCNN large model locates facial key points based on a key point localization algorithm to obtain the coordinate information of various parts of the face. The key point localization algorithm is specifically as follows: k = FC(z) = Wz + b; In the formula, k represents a 2L-dimensional vector; z represents the input feature vector; W represents a 2L×M weight matrix; b represents a 2L-dimensional bias vector; and FC(z) represents a fully connected layer.

4. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In step S3, feature extraction of key points and the introduction of a weighted mask to account for facial differences involve the following steps: A weighted mask is introduced, and for each cropped sub-image, a convolution operation is applied to extract features, specifically: In the formula, I represents the cropped sub-image; K represents the convolution kernel; (x, y) represents the coordinates on the output feature map; F represents the size of the convolution kernel; m and n represent the coordinates on the convolution kernel; Mask represents the weight mask; Mask(x+m, y+n) represents the value of the weight mask at position (x+m, y+n). The activation function is applied to the output of the convolutional layer, specifically: ReLU(z) = max(0, z); In the formula, ReLU represents the modified linear unit; z represents the input to the activation function; and max represents the maximum value function. Pooling is applied to the activated feature map to reduce the feature dimensionality and improve the spatial invariance of the features, specifically: MaxPool(I)(x, y)=max m,n I(x+q, y+p); In the formula, I represents the cropped sub-image; (x, y) represents the coordinates on the output feature map; q and p represent the coordinates on the pooling window; Finally, the output of the pooling layer is flattened into a vector and mapped to the final feature representation through a fully connected layer, specifically: FC(z) = W·Mask(z) + b; In the formula, W represents the weight matrix, b represents the bias vector, z represents the input vector, and Mask represents the weight mask. Mask(z) represents a weight mask vector with the same dimension as the input vector z.

5. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In step S4, the face alignment based on the coordinates of key points is performed using a heterogeneous affine transformation algorithm, which is used to calculate the weights of facial features. Specifically: In the formula, x and y represent the coordinates of a point in the original image; x′ and y′ represent the coordinates of a point in the transformed image; a represents the scaling parameter of the image; d represents the rotation parameter of the image; b represents the shearing along the x-axis of the image; c represents the shearing along the y-axis of the image; e represents the translation along the x-axis of the image; and f represents the translation along the y-axis of the image.

6. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In step S5, the variation similarity coefficient algorithm is used to calculate the difference between facial features and standardized facial features. The specific steps involved are as follows: S1.1 Extract the standardized facial feature vector v1 as the standardized parameter of the facial data; S1.

2. Convert the non-zero element feature vectors in the feature vectors into sets, and calculate a comprehensive score based on the mutation similarity coefficient algorithm to measure the similarity between the two sets.

7. The face recognition method based on deep learning in a cloud computing scenario according to claim 6, characterized in that: In S1.1, the feature extraction of the standardized face feature vector v1 is specifically as follows: In the formula, v1 represents the standardized face feature vector; W represents the weight matrix; b represents the bias vector; z represents the input vector; and α represents the learning rate parameter.

8. The face recognition method based on deep learning in a cloud computing scenario according to claim 6, characterized in that: In S1.2, the specific algorithm for the variation similarity coefficient is as follows: In the formula, Jaccard represents the variation similarity coefficient; v1 represents the face normalization feature vector; v2 represents the face recognition feature vector; i represents the index in the feature vector; v 1i This represents the value of the i-th element in the feature vector v1; v 2i This represents the value of the i-th element in the feature vector v2.

9. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In step S6, the steps involved in generating the public and private keys based on RSA are as follows: Choose two large prime numbers p and q, and calculate their product N, specifically: N = p × q; In the formula, p and q represent random prime numbers; N represents the modulus; The Euler totient function is calculated based on the modulus N, specifically as follows: φ(N)=(p-1)×(q-1); In the formula, φ(N) represents the Euler totient function; p and q represent random prime numbers; A fixed prime number is randomly selected, and the private key exponent is calculated based on Euler's totient function, specifically as follows: g×h≡1modφ(N); In the formula, g represents the private key exponent; h represents the public key exponent; φ(N) represents the Euler totient function; A public key consists of two parts: the modulus N and the public key exponent h. Therefore, a public key can be represented as: pk = (N, h); In the formula, pk represents the public key; N represents the modulus; and h represents the public key exponent. The private key also consists of two parts: the modulus N and the private key exponent g. Therefore, the private key can be represented as: sk = (N, g); In the formula, sk represents the private key; N represents the modulus; and g represents the private key exponent.

10. The face recognition method based on deep learning in a cloud computing scenario according to claim 1, characterized in that: In step S7, the facial information data extracted from features is encrypted using a homomorphic encryption algorithm, specifically as follows: c = Enc(pk, mc); In the formula, c represents the encrypted ciphertext; Enc represents the encryption function; pk represents the public key; and mc represents the facial information data extracted from features.

Citation Information

Patent Citations

  • Facial expression recognition method based on pyramid structure convolutional neural network

    CN111563417A

  • Face recognition system based on machine learning

    CN117115881A

  • Multi-modal emotion recognition method based on zero-knowledge machine learning

    CN119106367A

  • Face recognition face changing method and system based on deep learning

    CN120047984A

  • Face recognition method, robot, and storage medium

    WO2022126464A1