Facet-PyTorch framework-based face recognition method
The Facenet-PyTorch framework addresses face recognition challenges by optimizing triplet loss and using data augmentation to enhance feature extraction, ensuring high precision and robustness across varying environments and scenarios.
Patent Information
- Application Number
- CN202510710506.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the face features of traditional CNN models are susceptible to light, posture and occlusion interference, the Softmax classification loss function is difficult to optimize the feature vector, the triple loss function of Facenet requires manual parameter adjustment, and the static calculation graph of TensorFlow limits training flexibility and insufficient data enhancement, resulting in low recognition accuracy, weak anti-interference ability, and insufficient real-time and generalization ability.
The Facenet-PyTorch framework is adopted, combining dynamic computing graphs and efficient GPU acceleration, and the triple loss function parameters are optimized through data augmentation technologies such as random cropping, rotation, flip and color jitter. The flexible programming interface and cross-validation of PyTorch are used to adjust model parameters to achieve high-precision face recognition.
It significantly improves the accuracy and robustness of facial recognition, adapts to a variety of scenarios, supports real-time processing of large-scale data, reduces misunderstandings and missed recognition, and improves the system's response speed and processing capabilities.
Smart Images

Figure CN120318888A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and deep learning, and specifically relates to a face recognition method based on the Facenet-PyTorch framework. Background Art
[0002] The face recognition system based on Facenet-PyTorch focuses on the forefront fields of computer vision and deep learning, and its background art covers a number of key innovations and developments. The existing technologies still have the following key problems:
[0003] The face features extracted by traditional CNN models (such as VGG, ResNet) are vulnerable to interference from light, pose, and occlusion in high-dimensional spaces, resulting in limited ability to distinguish similar faces (such as twins);
[0004] The ordinary Softmax classification loss function is difficult to optimize the intra-class compactness and inter-class separability of feature vectors, affecting the recognition accuracy;
[0005] The triplet loss of the original Facenet relies on the selection of hard samples. If the sampling strategy is inappropriate, it is easy to cause the model to converge slowly or fall into a local optimum;
[0006] Lacking an adaptive mechanism for dynamically adjusting the triplet margin parameter, manual parameter tuning is required, increasing the training complexity;
[0007] The static computational graph of early frameworks (such as TensorFlow) limits the flexibility of the training process and is difficult to adapt to diverse face data augmentation strategies;
[0008] Under a large-scale face database, the feature comparison time of traditional methods increases sharply, lacking real-time performance;
[0009] Conventional data augmentation (rotation, cropping) cannot fully simulate extreme changes in real scenarios (such as strong backlighting, partial occlusion), resulting in insufficient generalization ability of the model;
[0010] Traditional methods use fixed example thresholds, which are difficult to adapt to the recognition requirements in different scenarios, resulting in large fluctuations in the false recognition rate and the missed recognition rate.
[0011] As a deep learning architecture specifically designed for face embedding identification, the Facenet model can map face images into compact feature vectors in a high-dimensional space, maintaining high distinctiveness even under different lighting, angles, and expression variations. The Facenet-PyTorch system adopts an advanced Triplet Loss function, minimizing the distance between image pairs of the same individual and maximizing the distance between image pairs of different individuals through supervised learning to optimize the distribution of face feature vectors. This ensures that the model maintains good recognition performance in large face datasets and can even accurately distinguish extremely similar faces.
[0012] As an important tool in the field of deep learning, the PyTorch framework, with its flexible programming interface and powerful computing capabilities, has become an ideal platform for implementing the Facenet model. It supports dynamic computational graphs, allowing the adjustment of network structures during training. Combined with efficient GPU acceleration, it significantly shortens the model training cycle, making it possible to efficiently process large-scale face datasets.
[0013] Data augmentation techniques are also crucial in the Facenet-PyTorch system. By performing transformations such as rotation, cropping, flipping, and color adjustment on training images, data augmentation not only increases the model's robustness to complex environments but also effectively prevents overfitting, significantly enhancing the model's generalization ability in real-world scenarios (such as strong backlighting and partial occlusion).
[0014] In summary, the face recognition system based on Facenet-PyTorch integrates advanced model architectures, optimized loss functions, high-performance computing frameworks, and innovative data processing techniques, providing a powerful solution for efficient and accurate face recognition. This system not only promotes technological innovation in academia but also demonstrates great potential and commercial value in multiple practical application scenarios such as security monitoring, identity verification, smart devices, and social media. Summary of the Invention
[0015] Aiming at the above problems, the purpose of the present invention is to provide a highly accurate, robust, and multi-scenario applicable face recognition method based on the Facenet-PyTorch framework.
[0016] The specific technical solution to achieve the purpose of the present invention is as follows:
[0017] A face recognition method based on the Facenet-PyTorch framework, comprising the following steps:
[0018] Step 1: Collect an image dataset containing faces;
[0019] Step 2: Extract features from the image dataset based on a pre-trained deep learning model;
[0020] Step 3: Map the extracted feature vectors to a compact embedding space;
[0021] Step 4: Train the deep learning model and optimize the distribution of feature vectors by minimizing the triplet loss function;
[0022] Step 5: Extract features of unknown face images based on the trained deep learning model, and perform face recognition based on the distance metric between the unknown face feature vectors and the known face feature vectors.
[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0024] (1) Compared with the prior art, the solution of the present invention can extract highly discriminative feature vectors from images through the Facenet model of the deep convolutional neural network (DCNN), and can maintain high-precision face recognition even in complex environments, which is significantly better than the traditional method based on handcrafted features, solving the problems of low face recognition accuracy and weak anti-interference ability in the prior art, and still maintaining a high recognition rate especially under complex conditions such as illumination changes, facial occlusion, and pose changes;
[0025] (2) The solution of the present invention introduces data augmentation techniques, including random cropping, rotation, flipping, and color jitter, significantly improving the robustness of the model and enabling it to maintain stable recognition performance when facing challenges such as illumination changes, expression differences, and pose changes;
[0026] (3) The solution of the present invention finely adjusts the preset threshold τ and the parameters α of the triplet loss function, and combines a dynamic optimization strategy to ensure the accuracy of the decision logic and reduce the cases of misrecognition and missed recognition;
[0027] (4) The solution of the present invention utilizes the efficient GPU acceleration and dynamic computational graph features of the PyTorch framework. The present invention can quickly process large-scale data sets, support real-time face recognition requirements, and significantly improve the system response speed and processing capacity;
[0028] (5) The solution of the present invention can adaptively optimize model parameters, including the fixed dimension D, the distance metric method, and the parameters of the triplet loss function, through methods such as cross-validation and grid search, to adapt to different environments and requirements, and ensure that the system can achieve the best performance under various conditions;
[0029] (6) The face recognition architecture of the present invention is flexibly designed and can adapt to a variety of application scenarios, including but not limited to security monitoring, access control, payment verification, identity authentication, etc., demonstrating broad applicability and commercial value. It not only provides a complete system process from image acquisition to feature comparison, but also reserves sufficient expansion space to facilitate the integration of new technologies or modules, such as the latest neural network architecture, more efficient data augmentation strategies, etc., to continuously improve the system performance.
[0030] The following further describes the present invention in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic flowchart of the face recognition method based on the Facenet-PyTorch framework of the present invention.
[0032] Figure 2 It is a structure diagram of the pytorch model in the embodiment of the present invention.
[0033] Figure 3 It is a structure diagram of the general face recognition system FaceNet model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Embodiment
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0036] As shown in this application and the claims, unless the context clearly indicates otherwise, the words "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0037] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0038] Combined with Figure 1 , a face recognition method based on the Facenet-PyTorch framework, comprising the following steps:
[0039] Step 1, collect an image dataset containing human faces;
[0040] The image dataset includes at least M images of different individuals, with at least N images for each individual, to ensure the robustness and generalization ability of model training;
[0041] In addition, the data in the image dataset needs to be subjected to data augmentation processing, including random cropping, random rotation, horizontal flipping, or color jittering;
[0042] Among them, random cropping is to randomly select the cropping area of the image while keeping the human face main body, which can be center cropping, random position cropping, etc., to increase the robustness of the model to different image croppings. torchvision.transforms.RandomResizedCrop can be used, which can randomly crop the image according to the ratio and then resize it;
[0043] Random rotation is to randomly rotate the image within a certain angle range to help the model learn the representations of human faces in different directions. torchvision.transforms.RandomRotation can be used to set the angle range of rotation;
[0044] Horizontal flipping is to randomly perform horizontal flipping on the image, which helps the model learn symmetry and improve the recognition ability of mirror images. torchvision.transforms.RandomHorizontalFlip can be used to flip the image with a certain probability;
[0045] Color jittering is to randomly change the brightness, contrast, saturation, etc. of an image to enhance the generalization ability of the model under different lighting conditions. torchvision.transforms.ColorJitter can be used to adjust the brightness, contrast, saturation, and hue of the image.
[0046] Step 2: Extract features from the image dataset based on the pre-trained deep learning model:
[0047] Combined with Figure 2 , the deep learning model is based on the Inception-v1 architecture and is implemented through the PyTorch framework;
[0048] The process of feature extraction involves the cascading of convolutional layers, pooling layers, and fully connected layers, converting the input image into a high-dimensional feature vector.
[0049] Step 3: Map the extracted feature vectors to a compact embedding space. The dimension D of the feature vectors determines the size of the embedding vectors. The optimal dimension D is determined through cross-validation experiments. The model is trained with multiple different D values, and its performance on the validation set is evaluated. Select the D value that can balance the recognition accuracy and computational efficiency;
[0050] In Facenet, the dimension D of the embedding vectors is a key hyperparameter of the model. Assume there are N face samples, and the dimension of the embedding vector for each sample is D. The performance of the model under different D values is measured by the following formula:
[0051]
[0052] where, f i is the embedding vector of the i-th test sample, g i is the embedding vector of the most similar sample of the same person, and h i is the embedding vector of the most similar sample of another person. distance(·) represents the distance metric function, including Euclidean distance or cosine similarity.
[0053] In this embodiment, the general optimal dimension D is determined to be 128 dimensions according to cross-validation experiments, and it can also be adjusted according to specific application scenarios. Higher dimensions may capture more details, but will increase the computational cost and storage requirements. Lower dimensions are beneficial for reducing the amount of calculation, but may lose some feature information.
[0054] Step 4: Combined with Figure 3 , using the pre-trained model in Step 2 as the initial model, train the deep learning model, and optimize the distribution of the feature vectors by minimizing the triplet loss function;
[0055] The triplet loss function is defined as follows:
[0056] L(A, P, N) = max(0, m + d(A, P) - d(A, N))
[0057] Where A is the anchor point, P is the positive sample, N is the negative sample, and m is the margin value (i.e., parameter α);
[0058] During the training process of the deep learning model, by optimizing the parameter α in the triplet loss function, the distance between positive and negative samples is balanced to ensure the accuracy and robustness of the model;
[0059] The parameter α in the triplet loss function is optimized by dynamic adjustment or grid search;
[0060] Among them, dynamic adjustment means that during the training process, the value of α is dynamically adjusted according to the performance of the model to meet the learning needs at different stages;
[0061] Grid search means that within a certain range, a grid search is performed on α to find the optimal value that can maximize the performance of the training set and the validation set.
[0062] Step 5: Extract the features of the unknown face image based on the trained deep learning model, and perform face recognition based on the distance measurement between the unknown face feature vector and the known face feature vector;
[0063] Use distance metrics including Euclidean distance and cosine similarity to determine the distance between the unknown face feature vector (the newly input image to be recognized) and the known face feature vector (i.e., the individual face feature vector extracted in advance by the trained deep learning model and stored in the database), and according to the set threshold, judge whether the unknown face feature vector and the known face feature vector belong to the same person, so as to obtain the recognition result;
[0064] The Euclidean distance is as follows:
[0065]
[0066] Where D is the total dimension of the feature vector, f i represents the i-th element of the feature vector, f i (a) is the i-th element of the feature vector extracted from the unknown face image, f i (b) is the i-th element of the feature vector extracted from the known face image;
[0067] The Euclidean distance is one of the most commonly used measurement methods and is applicable to most cases;
[0068] The cosine similarity is as follows:
[0069]
[0070] Among them, · represents the dot product of vectors, and ∥·∥ represents the norm of vectors.
[0071] Cosine similarity is usually more suitable for comparing the directions of vectors in a high-dimensional space. Which distance metric to choose depends on the specific application scenario and data characteristics.
[0072] Specifically, the threshold is determined by cross-validation to achieve the best recognition accuracy and recall rate;
[0073] That is, using cross-validation technology, considering multiple performance metrics such as accuracy, recall rate, and F1 score, adjust the value of the threshold τ on different subsets of the training set, and determine the threshold τ that makes the model perform best on the validation set;
[0074] The calculation formulas for the recognition accuracy and recall rate are as follows:
[0075]
[0076] Among them, Precision represents precision, Recall represents recall rate, TP represents true positive examples, FP represents false positive examples, and FN represents false negative examples.
[0077] The following table shows the performance metrics of the model (taking the F1 score as an example) under different dimensions D and different thresholds τ:
[0078] D value τ = 0.3 τ = 0.4 τ = 0.5 τ = 0.6 τ = 0.7 64 0.82 0.85 0.87 0.84 0.81 128 0.85 0.88 0.90 0.89 0.86 256 0.84 0.87 0.89 0.88 0.85
[0079] It can be seen that when D = 128 and τ = 0.5, the F1 score of the model is the highest, reaching 0.90. Therefore, this may be the optimal combination.
[0080] In summary, the present invention clearly demonstrates how to build a face recognition solution that can adapt to complex environments, process massive data, and provide instant feedback from theory to practice. From data preprocessing to model training, and then to system integration and optimization, it improves the system response speed and processing ability, and at the same time improves the robustness of the solution. Each link reflects the power of innovative thinking and technology integration, and has a certain promotion prospect.
[0081] In addition, the present invention also provides a face recognition system based on the Facenet-PyTorch framework, including the following modules:
[0082] Dataset acquisition module: Collect an image dataset containing faces;
[0083] Feature extraction and mapping module: Used to extract features from the image dataset based on a pre-trained deep learning model, and map the extracted feature vectors to a compact embedding space;
[0084] Model training module: used to train a deep learning model and optimize the distribution of feature vectors by minimizing the triplet loss function;
[0085] Recognition module: used to extract features of unknown face images based on the trained deep learning model and perform face recognition based on the distance metric between the unknown face feature vector and the known face feature vector.
[0086] In some embodiments, the system of the present solution may further include:
[0087] User interface module: provides an interaction interface between the user and the system, and displays information such as recognition results and system status.
[0088] Data storage module: stores the trained model parameters, registered face feature vectors, and system configuration information.
[0089] Network communication module: realizes data exchange between the system and remote servers or devices, and supports cloud service integration and distributed deployment.
[0090] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0091] Step 1: Collect an image data set containing faces;
[0092] Step 2: Extract features from the image data set based on a pre-trained deep learning model;
[0093] Step 3: Map the extracted feature vectors to a compact embedding space;
[0094] Step 4: Train the deep learning model and optimize the distribution of feature vectors by minimizing the triplet loss function;
[0095] Step 5: Extract features of unknown face images based on the trained deep learning model and perform face recognition based on the distance metric between the unknown face feature vector and the known face feature vector.
[0096] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0097] Step 1: Collect an image data set containing faces;
[0098] Step 2: Extract features from the image data set based on a pre-trained deep learning model;
[0099] Step 3: Map the extracted feature vectors to a compact embedding space;
[0100] Step 4: Train the deep learning model and optimize the distribution of feature vectors by minimizing the triplet loss function;
[0101] Step 5: Extract the features of the unknown face image based on the trained deep learning model, and perform face recognition based on the distance metric between the unknown face feature vector and the known face feature vector.
[0102] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A face recognition method based on the Facenet-PyTorch framework, characterized in that, It includes the following steps: Step 1, collect an image dataset containing human faces; Step 2, perform feature extraction on the image dataset based on a pre-trained deep learning model; Step 3, map the extracted feature vectors to a compact embedding space; Step 4, train the deep learning model, and optimize the distribution of feature vectors by minimizing the triplet loss function; Step 5, perform feature extraction on unknown face images based on the trained deep learning model, and perform face recognition based on the distance metric between the unknown face feature vectors and the known face feature vectors.
2. The face recognition method based on the Facenet-PyTorch framework according to claim 1, characterized in that The image dataset includes at least M images of different individuals, and each individual has at least N images to ensure the robustness and generalization ability of model training; The data in the image dataset needs to be processed by data augmentation, including random cropping, random rotation, horizontal flipping or color jitter.
3. The face recognition method based on the Facenet-PyTorch framework according to claim 1, characterized in that, The feature extraction in Step 2 is specifically: The deep learning model is based on the Inception-v1 architecture and is implemented through the PyTorch framework; The process of feature extraction involves the cascade of convolutional layers, pooling layers and fully connected layers, and converts the input image into a high-dimensional feature vector.
4. The face recognition method based on the Facenet-PyTorch framework according to claim 1, characterized in that The dimension D of the extracted feature vectors in Step 3 determines the size of the embedding vectors, and the optimal dimension D is determined through cross-validation experiments; Suppose there are N face samples, and the embedding vector dimension of each sample is D. The performance of the model under different dimension D values is measured by the following formula: where, f i is the embedding vector of the i-th test sample, g i is the embedding vector of the most similar sample of the same person, and h i is the embedding vector of the most similar sample of another person, and distance(·) represents a distance metric function, including Euclidean distance or cosine similarity.
5. The face recognition method based on the Facenet-PyTorch framework according to claim 1, wherein During the training process of the deep learning model in Step 4, by optimizing the parameter α in the triplet loss function, the distance between positive and negative samples is balanced to ensure the accuracy and robustness of the model; The parameter α in the triplet loss function is optimized by means of dynamic adjustment or grid search.
6. The face recognition method based on the Facenet-PyTorch framework according to claim 1, characterized in that The face recognition based on the distance metric between the unknown face feature vectors and the known face feature vectors in Step 5 is specifically: Use distance metrics including Euclidean distance and cosine similarity to determine the distance between the unknown face feature vectors and the known face feature vectors, and judge whether the unknown face feature vectors and the known face feature vectors are of the same person according to the set threshold, so as to obtain the recognition result; The Euclidean distance is: where D is the total dimension of the feature vector, and f i represents the i-th element of the feature vector, and f i (a) is the i-th element of the feature vector extracted from the unknown face image, and f i (b) is the i-th element of the feature vector extracted from the known face image; The cosine similarity is: Among them, · represents the dot product of vectors, and ∥·∥ represents the norm of vectors.
7. The face recognition method based on the Facenet-PyTorch framework according to claim 6, wherein The threshold is determined by cross-validation to achieve the best recognition accuracy and recall rate; That is, using cross-validation technology, adjust the value of the threshold τ on different subsets of the training set to determine the threshold that makes the model perform best on the validation set; The calculation formulas for the recognition accuracy and recall rate are: Among them, Precision represents precision, Recall represents recall rate, TP represents true positive, FP represents false positive, and FN represents false negative.
8. A face recognition system based on the Facenet-PyTorch framework, characterized in that, It includes the following modules: Dataset acquisition module: collect an image dataset containing human faces; Feature extraction and mapping module: used to perform feature extraction on the image dataset based on a pre-trained deep learning model, and map the extracted feature vectors to a compact embedding space; Model training module: used to train a deep learning model and optimize the distribution of feature vectors by minimizing the triplet loss function; Recognition module: used to extract features of an unknown face image based on the trained deep learning model, and perform face recognition based on the distance metric between the unknown face feature vector and the known face feature vector.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1-7.