Face Recognition Method, Device, Electronic Device and Storage Medium

By using target alignment templates and autoencoders to perform image reconstruction and feature extraction in facial recognition technology, the problem of low recognition accuracy caused by occluding different angles of face images is solved, and higher recognition accuracy is achieved.

CN115713789BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110951048.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-18
Publication Date
2025-07-08
Estimated Expiration
2041-08-18

AI Technical Summary

Technical Problem

In the prior art, due to problems such as occlusion of different angles of face images, key feature point detection cannot cover all situations, resulting in low accuracy of face recognition.

Method used

By obtaining the face image to be recognized, identifying the face key points, and performing alignment processing based on the target alignment template, then image reconstruction and feature extraction are performed, and secondary fine alignment is used for autoencoder and trained models to improve the accuracy of the alignment image.

Benefits of technology

Through secondary refined alignment processing, the accuracy of face recognition is improved, more accurate alignment images are obtained, and the overall accuracy of face recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713789B_ABST
    Figure CN115713789B_ABST
Patent Text Reader

Abstract

The present application discloses a face recognition method, device, electronic device, and storage medium; it can obtain a face image to be recognized and recognize the face key points of the face image; perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; extract features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image. Embodiments of the present application can perform secondary fine alignment on the aligned face image through image reconstruction, so as to obtain a more accurate aligned image, that is, the reconstructed aligned face image, and improve the accuracy of face image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a face recognition method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of computer technology, image processing technology has been applied to more and more fields. For example, face recognition technology is widely used in many fields such as access control and attendance, information security, electronic certificates, and detection and security. Specifically, face recognition technology is a technology that automatically extracts face features from a face image and then performs identity verification based on these features. Specifically, the standard process of face recognition consists of the following four steps: detecting the face image, locating the key feature points of the face part, aligning the key feature points, and finally using a recognition network to extract the face image features and perform feature comparison. Among them, the process of aligning the key feature points of the face part will affect the accuracy of the final face recognition result.

[0003] In the current related technologies, generally, a face detector is used for multi-task training to obtain the key feature points of the face part and the detection frame at the same time, and then an aligned face image is obtained for face recognition. However, due to problems such as occlusion of the face at different angles, the detection of key feature points cannot cover all situations, and a well-quality aligned face image cannot be obtained, resulting in a low accuracy of face recognition. Summary of the Invention

[0004] Embodiments of this application provide a face recognition method, apparatus, electronic device, and storage medium, which can improve the accuracy of face recognition.

[0005] Embodiments of this application provide a face recognition method, including:

[0006] Obtaining a face image to be recognized, and recognizing the face key points of the face image;

[0007] Performing alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image;

[0008] Performing image reconstruction on the aligned face image to obtain a reconstructed aligned face image;

[0009] Performing feature extraction on the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image;

[0010] Performing face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image.

[0011] Correspondingly, embodiments of this application provide a face recognition apparatus, including:

[0012] An acquisition unit, configured to acquire a face image to be recognized and recognize face key points of the face image;

[0013] An alignment unit, configured to perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image;

[0014] A reconstruction unit, configured to perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image;

[0015] An extraction unit, configured to extract feature information of the face of the reconstructed aligned face image from the reconstructed aligned face image;

[0016] A recognition unit, configured to perform face recognition processing on the face image according to the face feature information to obtain a face recognition result of the face image.

[0017] Optionally, in some embodiments of the present application, the reconstruction unit may include an encoding subunit and a reconstruction subunit, as follows:

[0018] The encoding subunit is configured to perform feature encoding on the aligned face image to obtain encoding information of the aligned face image;

[0019] The reconstruction subunit is configured to perform feature reconstruction on the encoding information to obtain a reconstructed aligned face image.

[0020] Optionally, in some embodiments of the present application, the reconstruction unit is specifically configured to perform image reconstruction on the aligned face image through a trained alignment model to obtain a reconstructed aligned face image; the extraction unit is specifically configured to perform feature extraction on the reconstructed aligned face image through a trained face recognition model to obtain face feature information of the reconstructed aligned face image.

[0021] Optionally, in some embodiments of the present application, the face recognition device may further include a first training unit, and the first training unit is configured to train the alignment model. Specifically, the first training unit is configured to obtain first training data, where the first training data includes a first sample face image and true identity information of the first sample face image; perform alignment processing on the face key points of the first sample face image based on the target alignment template to obtain an aligned sample face image; perform image reconstruction on the aligned sample face image through the alignment model to obtain a reconstructed aligned sample face image; and train the alignment model according to the reconstructed aligned sample face image and the true identity information of the first sample face image.

[0022] Optionally, in some embodiments of the present application, the step of "training the alignment model according to the reconstructed and aligned sample face image and the true identity information of the first sample face image" may include:

[0023] Extracting face feature information of the reconstructed and aligned sample face image through the trained face recognition model;

[0024] Determining the predicted identity information of the first sample face image according to the face feature information;

[0025] Adjusting the parameters of the alignment model according to the true identity information and the predicted identity information of the first sample face image to obtain a trained alignment model.

[0026] Optionally, in some embodiments of the present application, the target alignment template is obtained by scaling the original alignment template; the trained face recognition model is trained based on the original alignment template.

[0027] Optionally, in some embodiments of the present application, the face recognition device may further include a second training unit for training the face recognition model. Specifically, the second training unit is configured to obtain second training data, where the second training data includes a second sample face image and the true identity information of the second sample face image; aligning the face key points in the second sample face image based on the original alignment template to obtain a target aligned sample face image; extracting face feature information of the target aligned sample face image through the face recognition model; and adjusting the parameters of the face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the second sample face image to obtain a trained face recognition model.

[0028] Optionally, in some embodiments of the present application, the face recognition device may further include a third training unit, and the third training unit is used to jointly train the alignment model and the face recognition model. Specifically, the third training unit is configured to obtain third training data, where the third training data includes third sample face images and the true identity information of the third sample face images; based on the target alignment template, perform alignment processing on the face key points of the third sample face images to obtain aligned sample face images; through the pre-trained alignment model, perform image reconstruction on the aligned sample face images to obtain reconstructed aligned sample face images; through the pre-trained face recognition model, extract the face feature information of the reconstructed aligned sample face images; and adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images.

[0029] Optionally, in some embodiments of the present application, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images" may include:

[0030] Based on the original alignment template, perform alignment processing on the face key points of the third sample face images to obtain target aligned sample face images, where the target alignment template is obtained by scaling the original alignment template;

[0031] Through a preset standard face recognition model, extract the reference face feature information of the third sample face images according to the target aligned sample face images and the reconstructed aligned sample face images;

[0032] Adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face images.

[0033] Optionally, in some embodiments of the present application, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face images" may include:

[0034] Calculate a first loss value between the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images;

[0035] Calculate a second loss value between the predicted reference identity information corresponding to the reference human face feature information and the true identity information of the third sample human face image;

[0036] According to the first loss value and the second loss value, adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model.

[0037] Optionally, in some embodiments of the present application, the step of "according to the first loss value and the second loss value, adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model" may include:

[0038] Fuse the first loss value and the second loss value to obtain a total loss value;

[0039] According to the total loss value, adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model.

[0040] Optionally, in some embodiments of the present application, the step of "through a preset standard human face recognition model, according to the target aligned sample human face image and the reconstructed aligned sample human face image, extract the reference human face feature information of the third sample human face image" may include:

[0041] Fuse the target aligned sample human face image and the reconstructed aligned sample human face image to obtain a fused sample human face image;

[0042] Through a preset standard human face recognition model, perform feature extraction on the fused sample human face image to obtain the reference human face feature information of the third sample human face image.

[0043] An electronic device provided by an embodiment of the present application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the human face recognition method provided by the embodiment of the present application.

[0044] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the human face recognition method provided by the embodiment of the present application are implemented.

[0045] The embodiments of the present application provide a face recognition method, apparatus, electronic device, and storage medium. It can obtain a face image to be recognized and identify the face key points of the face image; perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; extract features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; and perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image. The embodiments of the present application can perform secondary fine alignment on the aligned face image through image reconstruction, so as to obtain a more accurate aligned image, that is, the reconstructed aligned face image, and improve the accuracy of face image recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0047] Figure 1a is a schematic diagram of the scenario of the face recognition method provided by the embodiments of the present application;

[0048] Figure 1b is a flowchart of the face recognition method provided by the embodiments of the present application;

[0049] Figure 2a is another flowchart of the face recognition method provided by the embodiments of the present application;

[0050] Figure 2b is another flowchart of the face recognition method provided by the embodiments of the present application;

[0051] Figure 2c is another flowchart of the face recognition method provided by the embodiments of the present application;

[0052] Figure 2d is another flowchart of the face recognition method provided by the embodiments of the present application;

[0053] Figure 2e is another flowchart of the face recognition method provided by the embodiments of the present application;

[0054] Figure 2f is a model architecture diagram of the face recognition method provided by the embodiments of the present application;

[0055] Figure 2g is another model architecture diagram of the face recognition method provided by the embodiments of the present application;

[0056] Figure 3 is a schematic structural diagram of a face recognition device provided by an embodiment of the present application;

[0057] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0059] An embodiment of the present application provides a face recognition method, device, electronic device, and storage medium. The face recognition device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server.

[0060] It can be understood that the face recognition method in this embodiment can be executed on a terminal, on a server, or jointly executed by a terminal and a server. The above examples should not be construed as a limitation to the present application.

[0061] Such as Figure 1a shown, taking the face recognition method jointly executed by a terminal and a server as an example. The face recognition system provided by an embodiment of the present application includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected through a network, for example, through a wired or wireless network connection, etc., where the face recognition device can be integrated in the server.

[0062] Among them, the server 11 can be used to: obtain a face image to be recognized, and recognize the face key points of the face image; perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; extract features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image. Among them, the server 11 can be a single server, or a server cluster or cloud server composed of multiple servers.

[0063] Among them, the terminal 10 can be used to send a face image to be recognized to the server 11 and receive the face recognition result of the face image sent by the server 11. Among them, the terminal 10 can include a mobile phone, a smart TV, a tablet computer, a laptop computer, or a personal computer (PC, Personal Computer), etc. A client can also be set on the terminal 10, and the client can be an application client or a browser client, etc.

[0064] The steps for the server 11 to perform face recognition can also be executed by the terminal 10.

[0065] The face recognition method provided by the embodiments of this application involves computer vision technology and machine learning in the field of artificial intelligence. This application can perform secondary refined alignment on the aligned face image through image reconstruction, so as to obtain a more accurate aligned image, that is, the reconstructed aligned face image, to improve the accuracy of face image recognition.

[0066] Among them, artificial intelligence (AI, Artificial Intelligence) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Among them, artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0067] Among them, computer vision technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, detection, and measurement on targets, and further perform graphic processing to make the computer process into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0068] Among them, machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0069] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0070] This embodiment will be described from the perspective of a face recognition device, which can be specifically integrated in an electronic device, and the electronic device can be a device such as a server or a terminal.

[0071] The face recognition method of the embodiments of the present application can be applied to various face recognition scenarios, including but not limited to: mobile phone face unlocking, application (APP) face login, remote face verification, face recognition access control system, offline face payment, automatic face clearance, etc.

[0072] As Figure 1b shown, the specific process of the face recognition method can be as follows:

[0073] 101. Obtain a face image to be recognized, and recognize the face key points of the face image.

[0074] Among them, the face image is the image of the identity to be recognized. Facial key points, also known as facial feature points, usually include the points that make up the facial features (eyebrows, eyes, nose, mouth, and ears) and the facial contour.

[0075] Generally, in some embodiments, before recognizing the facial key points of a face image, face detection can be performed first. The purpose of face detection is to accurately locate the face in the picture, that is, to find the position of the face in the picture. Through face detection, the coordinate information of the face can be marked, or the face can be cut out. Specifically, a face detector can be used to perform face detection to obtain a detection box, and then the positions of five standard points of the face (left eye, right eye, left corner of the mouth, right corner of the mouth, tip of the nose) can be obtained by using the detection box and the face image.

[0076] In this embodiment, after recognizing the facial key points of the face image, face alignment can be performed. Facial alignment can align face images at different angles into the same standard shape; specifically, facial alignment can first locate the feature points on the face (i.e., facial key points), and then transform the face image into the face in the alignment template, such as through geometric transformations (affine, rotation, scaling), so that each feature point is aligned (such as moving parts such as eyes and mouth to the positions corresponding to the eyes and mouth in the alignment template).

[0077] 102. Align the facial key points of the face image based on the target alignment template to obtain an aligned face image.

[0078] Among them, the target alignment template can specifically be obtained by performing scaling processing on the original alignment template. The original alignment template can be a preset alignment template, which can be specifically set according to the actual situation. The alignment template can contain the standard position information of the facial key points, and the face image can be adjusted through the alignment template so that the position information of the facial key points in the face image is as close as possible to the standard position information of the facial key points in the alignment template.

[0079] Among them, using the target alignment template to perform alignment processing on the facial key points of the face image can specifically be to perform geometric transformations on the face image through methods such as rotation, translation, and scaling, so that the facial key points in the geometrically transformed face image meet the specified standards. For example, if the positions of the facial key points in the geometrically transformed face image meet the position information of the facial key points specified in the target alignment template, then the geometrically transformed face image is the aligned face image.

[0080] 103. Perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image.

[0081] Among them, an autoencoder (AE) can be used to reconstruct the aligned face images. The autoencoder can first compress the input into a latent space representation and then reconstruct the output through this representation. Specifically, it can be a neural network model. For example, it can be a Visual Geometry Group Network (VGGNet), a Residual Network (ResNet), a Dense Convolutional Network (DenseNet), and so on. It should be understood that the autoencoder in this embodiment is not limited to the several types listed above. In some embodiments, a network with a UNet (U-shaped network) structure can be used to reconstruct the aligned face images.

[0082] Among them, the autoencoder can be composed of two parts, namely an encoder and a decoder. The encoder can be used to perform feature encoding on the aligned face images, and the decoder can be used to perform feature reconstruction on the obtained encoded information.

[0083] Among them, this part of the encoder can compress the input into a latent space representation, which can be represented by the encoding function h = f(x). This part of the decoder can reconstruct the input from the latent space representation, which can be represented by the decoding function r = g(h). Therefore, the entire autoencoder can be described by the function g(f(x)) = r, where the output r is similar to the original input x.

[0084] In this embodiment, using the autoencoder to reconstruct the aligned face images is to adjust the positions of the facial key points in the aligned face images again. This process can be regarded as the second alignment process of the facial key points in the face images. Among them, the first alignment process is the alignment process through the target alignment template, and the second alignment process is the image reconstruction through the autoencoder. It should be noted that through the image reconstruction of the autoencoder, the positions of the facial key points in the output of the autoencoder (i.e., the reconstructed aligned face images) can be more accurate than the positions of the facial key points in the input of the autoencoder (i.e., the aligned face images).

[0085] Specifically, the target alignment template is obtained by magnifying the original alignment template. Therefore, compared with the original alignment template, the alignment process of the face key points in the face image by the target alignment template is coarser. Or rather, the position range of the face key points in the target alignment template is larger than that in the original alignment template. In this way, the adjustability of the face key points in the aligned face image obtained through the target alignment template is stronger. So that after the aligned face image is reconstructed, the face key points can be adjusted through the image reconstruction, and the positions of the face key points in the reconstructed aligned face image can be made more accurate.

[0086] Optionally, in this embodiment, the step of "performing image reconstruction on the aligned face image to obtain a reconstructed aligned face image" may include:

[0087] Performing feature encoding on the aligned face image to obtain the encoding information of the aligned face image;

[0088] Performing feature reconstruction on the encoding information to obtain a reconstructed aligned face image.

[0089] 104. Extracting face feature information of the reconstructed aligned face image from the reconstructed aligned face image.

[0090] Among them, the feature extraction of the reconstructed aligned face image may specifically be operations such as convolution calculation, non-linear activation function (Relu, Rectified Linear Unit, that is, the rectified linear function) calculation, and pooling calculation on the reconstructed aligned face image.

[0091] 105. Performing face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image.

[0092] In some embodiments, the face feature information may be compared with preset face feature information, and the preset face feature information may be the face feature information of a user stored in advance. Specifically, the similarity between the face feature information and each preset face feature information may be calculated, and according to the similarity, the matching face feature information may be selected from the preset face feature information, and the user corresponding to the matching face feature information may be used as the face recognition result of the face image, that is, the identity information of the face image is the user corresponding to the matching face feature information.

[0093] Optionally, in this embodiment, the step of "performing image reconstruction on the aligned face image to obtain a reconstructed aligned face image" may include:

[0094] Using the trained alignment model, perform image reconstruction on the aligned face image to obtain the reconstructed aligned face image;

[0095] The step of "performing feature extraction on the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image" may include:

[0096] Using the trained face recognition model, perform feature extraction on the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image.

[0097] Among them, the alignment model can be a network with a UNet structure. This embodiment does not limit the type of the alignment model. For example, the alignment model can be an autoencoder (AE), a variational autoencoder (VAE), etc.

[0098] Among them, the face recognition model can be a neural network model. For example, it can be a convolutional neural network (CNN), a visual geometry group network (VGGNet), a residual network (ResNet), a dense connection convolutional network (DenseNet), etc. It should be understood that the face recognition model in this embodiment is not limited to the above-listed types.

[0099] It should be noted that the alignment model and the face recognition model in this embodiment can be trained by multiple training data with labels. The training data in this embodiment includes multiple sample face images, and the label refers to the true identity information corresponding to the sample face image; the alignment model and the face recognition model can be specifically trained by other devices and then provided to the face recognition device, or they can also be trained by the face recognition device itself. This embodiment does not limit this.

[0100] Optionally, in this embodiment, the face recognition method may further include:

[0101] Obtain first training data, where the first training data includes a first sample face image and the true identity information of the first sample face image;

[0102] Based on the target alignment template, perform alignment processing on the face key points of the first sample face image to obtain an aligned sample face image;

[0103] Using an alignment model, perform image reconstruction on the aligned sample face image to obtain a reconstructed aligned sample face image;

[0104] Train the alignment model according to the true identity information of the reconstructed aligned sample face image and the first sample face image.

[0105] Among them, the true identity information of the first sample face image, that is, the label information of the first sample face image, is the expected face recognition result.

[0106] Optionally, in this embodiment, the step of "training the alignment model according to the true identity information of the reconstructed aligned sample face image and the first sample face image" may include:

[0107] Extract the face feature information of the reconstructed aligned sample face image through the trained face recognition model to obtain the face feature information of the reconstructed aligned sample face image;

[0108] Determine the predicted identity information of the first sample face image according to the face feature information;

[0109] Adjust the parameters of the alignment model according to the true identity information and the predicted identity information of the first sample face image to obtain a trained alignment model.

[0110] Among them, when training the alignment model, the face recognition model used is pre-trained.

[0111] Optionally, in some embodiments, the predicted identity information may specifically include the probabilities of the first sample face image belonging to each preset identity. The true identity information can be regarded as the expected probability that the first sample face image belongs to its true identity is 1, and the expected probability of belonging to a non-true identity is 0. In this embodiment, the loss value between the true identity information and the predicted identity information can be calculated, and the parameters of the alignment model can be adjusted according to the loss value. Among them, the loss function corresponding to the loss value can be a cross-entropy loss function, a mean square error loss function, etc., and this embodiment does not limit this.

[0112] Optionally, in other embodiments, the true identity information of the first sample face image may specifically be the standard face feature information of the first sample face image. The similarity between the face feature information of the first sample face image and the standard face feature information can be calculated to determine the loss value between the true identity information and the predicted identity information; the greater the similarity, the smaller the loss value; conversely, the smaller the similarity, the greater the loss value.

[0113] Among them, according to the facial feature information, the predicted identity information of the first sample face image is determined. Specifically, a classifier can be used to determine its predicted identity information. The classifier can specifically be a support vector machine (SVM, Support Vector Machine), or a recurrent neural network, or a fully connected deep neural network (DNN, Deep Neural Networks), etc. This embodiment does not limit this.

[0114] In this embodiment, the training process of the alignment model can be to first calculate the matching degree between the predicted identity information and the true identity information, and then use the backpropagation algorithm to adjust the parameters of the alignment model. Based on the loss value between the predicted identity information and the true identity information, the parameters of the alignment model are optimized so that the predicted identity information approaches the true identity information to obtain the trained alignment model. Specifically, the loss value calculated between the predicted identity information and the true identity information can be made less than a preset loss value, and the preset loss value can be set according to the actual situation.

[0115] Optionally, in this embodiment, the target alignment template is obtained by scaling the original alignment template; the trained face recognition model is trained based on the original alignment template.

[0116] Among them, the target alignment template can scale the facial key points in the original alignment template by a certain proportion. If the facial key points in the original alignment template are denoted as (X, Y), then the facial key points in the target alignment template can be denoted as (αX, βY), where α and β are scaling coefficients, α>1, β>1. Optionally, the size of the target alignment template is generally taken as 2 times the size of the original alignment template, that is, α = 2, β = 2.

[0117] Among them, the training process of the face recognition model is specifically as follows.

[0118] Optionally, in this embodiment, before the step of "extracting the facial feature information of the reconstructed aligned sample face image through the trained face recognition model", it may further include:

[0119] Obtain second training data, where the second training data includes a second sample face image and the true identity information of the second sample face image;

[0120] Based on the original alignment template, perform alignment processing on the facial key points in the second sample face image to obtain a target aligned sample face image;

[0121] Through the face recognition model, extract the facial feature information of the target aligned sample face image;

[0122] Adjust the parameters of the face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the second sample face image, so as to obtain the trained face recognition model.

[0123] Among them, the training process of the face recognition model can be to first calculate the matching degree between the predicted identity information corresponding to the face feature information and the true identity information of the second sample face image, and then use the backpropagation algorithm to adjust the parameters of the face recognition model. Based on the loss value between the predicted identity information and the true identity information, optimize the parameters of the face recognition model to make the predicted identity information close to the true identity information, so as to obtain the trained face recognition model. Specifically, the loss value calculated between the predicted identity information and the true identity information can be made less than a preset loss value, and the preset loss value can be set according to the actual situation, and this embodiment does not limit this.

[0124] In this embodiment, after separately training the alignment model and the face recognition model, in order to make the alignment model and the face recognition model more matching, the alignment model and the face recognition model can also be jointly trained, and the joint training process is specifically as follows.

[0125] Optionally, in this embodiment, the face recognition method may further include:

[0126] Obtain third training data, where the third training data includes a third sample face image and the true identity information of the third sample face image;

[0127] Based on the target alignment template, perform alignment processing on the face key points of the third sample face image to obtain an aligned sample face image;

[0128] Through the pre-trained alignment model, perform image reconstruction on the aligned sample face image to obtain a reconstructed aligned sample face image;

[0129] Through the pre-trained face recognition model, perform feature extraction on the reconstructed aligned sample face image to obtain the face feature information of the reconstructed aligned sample face image;

[0130] According to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face image, adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model.

[0131] Among them, the pre-trained alignment model can specifically be trained using the first training data described in the above embodiment, and the pre-trained face recognition model can specifically be trained using the second training data described in the above embodiment.

[0132] Optionally, a loss value between the predicted identity information corresponding to the face feature information and the true identity information of the third sample face image can be calculated. The loss function corresponding to this loss value can be a cross-entropy loss function or a mean squared error loss function, and this embodiment does not limit this. According to the calculated loss value, the parameters of the pre-trained alignment model and the pre-trained face recognition model are adjusted so that the predicted identity information calculated by these two models approaches the true identity information of the third sample face image.

[0133] Optionally, in this embodiment, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face image" may include:

[0134] Based on the original alignment template, align the face key points of the third sample face image to obtain a target aligned sample face image, where the target alignment template is obtained by scaling the original alignment template;

[0135] Through a preset standard face recognition model, according to the target aligned sample face image and the reconstructed aligned sample face image, extract the reference face feature information of the third sample face image;

[0136] According to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face image, adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model.

[0137] Among them, the preset standard face recognition model is a trained standard face recognition model, which can specifically be a large face recognition model with excellent performance and a neural network model with a relatively deep convolutional layer. Introducing the preset standard face recognition model into the joint training process of the alignment model and the face recognition model can guide the reconstructed aligned sample face image obtained by image reconstruction through the alignment model, enabling the alignment model to train to a more refined face recognition area, thereby improving the accuracy of face recognition.

[0138] Optionally, in this embodiment, the step of "extracting the reference face feature information of the third sample face image through a preset standard face recognition model according to the target aligned sample face image and the reconstructed aligned sample face image" may include:

[0139] Fuse the target aligned sample face image and the reconstructed aligned sample face image to obtain a fused sample face image;

[0140] Extract feature information of the reference face of the third sample face image by extracting features from the fused sample face image using a preset standard face recognition model.

[0141] Among them, there are various ways to fuse the target aligned sample face image and the reconstructed aligned sample face image, and this embodiment does not limit this. For example, the fusion method can be splicing processing. For example, the reconstructed aligned sample face image can be spliced after the target aligned sample face image, or the target aligned sample face image can be spliced after the reconstructed aligned sample face image. This embodiment does not limit the splicing order. Optionally, the fusion method can also be multiplication, etc.

[0142] Among them, extracting features from the fused sample face image using a preset standard face recognition model can specifically be performing convolution processing, pooling processing, etc. on the fused sample face image. Through feature extraction, the reference face feature information of the third sample face image can be obtained.

[0143] Optionally, in this embodiment, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face image" may include:

[0144] Calculate a first loss value between the predicted identity information corresponding to the face feature information and the true identity information of the third sample face image;

[0145] Calculate a second loss value between the predicted reference identity information corresponding to the reference face feature information and the true identity information of the third sample face image;

[0146] Adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the first loss value and the second loss value.

[0147] Among them, in some embodiments, the true identity information of the third sample face image can specifically be the standard face feature information of the third sample face image. The similarity between the face feature information and this standard face feature information can be calculated to determine the first loss value between the predicted identity information corresponding to the face feature information and the true identity information of the third sample face image; the greater the similarity, the smaller the first loss value; conversely, the smaller the similarity, the greater the first loss value.

[0148] Similarly, for the second loss value between the predicted reference identity information corresponding to the reference face feature information and the true identity information of the third sample face image, it can also be determined by the similarity between the reference face feature information and the standard face feature information. The greater the similarity between the two, the smaller the second loss value; conversely, the smaller the similarity between the two, the greater the second loss value.

[0149] Among them, there are various ways to calculate the similarity. For example, the similarity between face feature information can be calculated by cosine similarity.

[0150] Optionally, in this embodiment, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the first loss value and the second loss value" may include:

[0151] Fuse the first loss value and the second loss value to obtain a total loss value;

[0152] Adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the total loss value.

[0153] Among them, there are various ways to fuse the first loss value and the second loss value, and this embodiment does not limit this. For example, the fusion method can be weighted fusion, etc.

[0154] In this embodiment, the joint training process of the alignment model and the face recognition model can be to adjust the parameters of the alignment model and the face recognition model using the backpropagation algorithm. Based on the total loss value, optimize the parameters of the alignment model and the face recognition model so that the total loss value is less than a preset loss value, and this preset loss value can be set according to the actual situation.

[0155] This application can improve the accuracy of face recognition by using a two-level alignment method. Specifically, it uses a target alignment template to perform a first-level rough alignment on the face image to obtain an image that meets certain rules, and then uses the trained alignment model for a second-level fine alignment. Through image reconstruction, the image obtained after the first-level rough alignment is further finely matched to improve the adaptability between the face recognition model and the alignment model, thereby improving the accuracy of the face recognition system.

[0156] Among them, the method for improving the accuracy of face recognition by using the secondary alignment proposed in this application is equivalent to adding a small face reconstruction network (i.e., the alignment model described in the above embodiments) to the existing recognition process. The face reconstruction network is used to perform secondary fine alignment on the face area, and according to the matching degree between the recognized and aligned face images, a face image suitable for the face recognition model to recognize is learned and obtained. In addition, the training of the face reconstruction network does not require the use of additional training data, which ensures the convenience of training; this method adds a small face reconstruction network to the existing recognition system without making major changes to the existing system, ensuring its practicability in actual deployment.

[0157] As can be seen from the above, this embodiment can obtain a face image to be recognized and recognize the face key points of the face image; perform alignment processing on the face key points of the face image based on the target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; extract features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image. The embodiment of the present application can perform secondary fine alignment on the aligned face image through image reconstruction, so as to obtain a more accurate aligned image, that is, the reconstructed aligned face image, and improve the accuracy of face image recognition.

[0158] According to the method described in the previous embodiments, the following will take the specific integration of the face recognition device in the server as an example for further detailed description.

[0159] The embodiment of the present application provides a face recognition method, as Figure 2a shown, the specific process of this face recognition method can be as follows:

[0160] 201. The server trains the face recognition model based on the second training data and the original alignment template to obtain a trained face recognition model.

[0161] Optionally, in this embodiment, the step of "training the face recognition model based on the second training data and the original alignment template to obtain a trained face recognition model" may include:

[0162] Obtain the second training data, where the second training data includes the second sample face image and the true identity information of the second sample face image;

[0163] Based on the original alignment template, perform alignment processing on the face key points in the second sample face image to obtain a target aligned sample face image;

[0164] Extract the feature of the target aligned sample face image through the face recognition model to obtain the face feature information of the target aligned sample face image;

[0165] Adjust the parameters of the face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the second sample face image to obtain the trained face recognition model.

[0166] In one embodiment, as Figure 2b shown in (A) of, it shows the process of training the face recognition model, where the recognition network unit module is the face recognition model in the above embodiment; specifically as follows:

[0167] During the training process, the training data preparation module can be used to read the face training data and input the read data into the deep network unit for processing. Specifically, the training data preparation module can obtain the second training data, identify the face key points in the second sample face image in the second training data, and then align the face key points in the second sample face image through the original alignment template to obtain the target aligned sample face image.

[0168] The recognition network unit module can extract the features of the target aligned sample face image to obtain the face feature information of the target aligned sample face image, and this face feature information retains the spatial structure information of the face image.

[0169] The face recognition objective function calculation module can use the face feature information output by the fully connected mapping unit and the label information (i.e., the true identity information) of the second sample face image as inputs to calculate the objective function value of the two. Among them, the objective function can select a classification function, such as softmax, various softmax with margin types, or other types of objective functions, and this embodiment does not limit this. Among them, the softmax function can convert the output values of multi-classification into a probability distribution in the range of [0,1], and margin represents the difficulty degree of feature learning.

[0170] The face recognition objective function optimization module can train and optimize the entire network based on the gradient descent method (such as stochastic gradient descent, stochastic gradient descent with momentum, adaptive moment estimation (adam), adagard (an optimization method)). During the training process of the face recognition model, repeat the above steps until the training result meets the training termination condition, then the recognition network unit module optimization is completed. The condition for terminating the model training can be set that the number of iterations meets the set value, or the loss value calculated by the objective function is less than the preset loss value to complete the training of the model.

[0171] 202. The server trains the alignment model based on the first training data and the target alignment template to obtain a trained alignment model.

[0172] Optionally, in this embodiment, the step of "training the alignment model based on the first training data and the target alignment template to obtain a trained alignment model" may include:

[0173] Obtain the first training data, where the first training data includes the first sample face image and the true identity information of the first sample face image;

[0174] Based on the target alignment template, perform alignment processing on the face key points of the first sample face image to obtain an aligned sample face image, where the target alignment template is obtained by scaling the original alignment template;

[0175] Through the alignment model, perform image reconstruction on the aligned sample face image to obtain a reconstructed aligned sample face image;

[0176] Through the trained face recognition model, extract the face feature information of the reconstructed aligned sample face image, where the trained face recognition model is trained based on the original alignment template.

[0177] According to the face feature information, determine the predicted identity information of the first sample face image;

[0178] According to the true identity information and the predicted identity information of the first sample face image, adjust the parameters of the alignment model to obtain a trained alignment model.

[0179] Specifically, as Figure 2c shown, it is the process of performing alignment processing using the original alignment template and the target alignment template. Among them, the sample face image can be first detected by the detection network module, then the face key points in the sample face image are recognized by the key point detection module, and then the alignment template is used to perform alignment processing on the face key points in the sample face image. Among them, the aligned picture 1 is obtained by alignment processing using the original alignment template, which can specifically be the target aligned sample face image corresponding to the second sample face image in the above embodiment; the aligned picture 2 is obtained by alignment processing using the target alignment template, which can specifically be the aligned sample face image corresponding to the first sample face image in the above embodiment. Since the original alignment template and the target alignment template are different, the scales of the aligned picture 1 and the aligned picture 2 obtained by the two are different.

[0180] During the training process of the alignment model, it is necessary to utilize the face recognition model (i.e., the recognition network unit module) trained in step 201. During the training process of the alignment model, the parameters of the recognition network unit module are not updated. The specific training process of the alignment model can refer to Figure 2b shown in (B) therein, where the secondary alignment network unit module is the alignment model in the above embodiment, which is specifically described as follows:

[0181] Among them, for the training data preparation module, the function of this module is the same as that of the training data preparation module in the training of the recognition network unit module in step 201. Specifically, in this embodiment, the training data preparation module can obtain the first training data, identify the face key points in the first sample face image in the first training data, and then, perform alignment processing on the face key points in the first sample face image through the target alignment template to obtain the aligned sample face image.

[0182] The secondary alignment network unit module can perform image reconstruction on the coarsely aligned image obtained by using the target alignment template in the training data preparation module, that is, the aligned sample face image, and obtain a finely aligned picture that meets the input size of the recognition network unit module - the reconstructed aligned sample face image. The secondary alignment network unit module can be composed of a convolutional neural network and generally includes operations such as convolutional calculation, non-linear activation function calculation, and pooling calculation.

[0183] The functions of other modules are basically the same as those shown in step 201. During the training process, only the parameters of the secondary alignment network unit module are updated, while other modules only provide gradient calculation and do not participate in the parameter update process.

[0184] Specifically, the recognition network unit module can extract features from the reconstructed aligned sample face image to obtain the face feature information of the reconstructed aligned sample face image. And the face recognition objective function calculation module can calculate the objective function value between the face feature information output by the recognition network unit module and the label information (i.e., the true identity information) of the first sample face image. This objective function can select a classification function, such as softmax, etc. Then, it is judged whether the termination condition of model training is satisfied. If not, the parameters of the secondary alignment network unit module can be adjusted based on the gradient descent method through the face recognition objective function optimization module until the training result meets the training termination condition, then the optimization of the secondary alignment network unit module is completed. The condition for terminating model training can be set that the number of iterations meets the set value, or the loss value calculated by the objective function is less than the preset loss value to complete the training of the model.

[0185] 203. The server jointly trains the trained face recognition model and the trained alignment model based on the third training data to obtain a jointly trained face recognition model and alignment model.

[0186] Optionally, in this embodiment, the step of "jointly training the trained face recognition model and the trained alignment model based on the third training data to obtain a jointly trained face recognition model and alignment model" may include:

[0187] Obtain the third training data, where the third training data includes third sample face images and the true identity information of the third sample face images;

[0188] Based on the target alignment template, perform alignment processing on the facial key points of the third sample face image to obtain an aligned sample face image;

[0189] Through the pre-trained alignment model, perform image reconstruction on the aligned sample face image to obtain a reconstructed aligned sample face image;

[0190] Through the pre-trained face recognition model, perform feature extraction on the reconstructed aligned sample face image to obtain the facial feature information of the reconstructed aligned sample face image;

[0191] Based on the original alignment template, perform alignment processing on the facial key points of the third sample face image to obtain a target aligned sample face image, where the target alignment template is obtained by performing scaling processing on the original alignment template;

[0192] Through a preset standard face recognition model, extract the reference facial feature information of the third sample face image according to the target aligned sample face image and the reconstructed aligned sample face image;

[0193] According to the predicted identity information corresponding to the facial feature information, the predicted reference identity information corresponding to the reference facial feature information, and the true identity information of the third sample face image, adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model.

[0194] Optionally, in this embodiment, the step of "according to the predicted identity information corresponding to the facial feature information, the predicted reference identity information corresponding to the reference facial feature information, and the true identity information of the third sample face image, adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model" may include:

[0195] Calculate a first loss value between the predicted identity information corresponding to the facial feature information and the true identity information of the third sample face image;

[0196] Calculate a second loss value between the predicted reference identity information corresponding to the reference human face feature information and the true identity information of the third sample human face image;

[0197] Adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model according to the first loss value and the second loss value.

[0198] Optionally, in this embodiment, the step of "extracting the reference human face feature information of the third sample human face image according to the target aligned sample human face image and the reconstructed aligned sample human face image through a preset standard human face recognition model" may include:

[0199] Fuse the target aligned sample human face image and the reconstructed aligned sample human face image to obtain a fused sample human face image;

[0200] Extract the reference human face feature information of the third sample human face image by performing feature extraction on the fused sample human face image through a preset standard human face recognition model.

[0201] In one embodiment, as Figure 2d shown, it shows the process of jointly training an alignment model and a human face recognition model, where the recognition network unit module is the human face recognition model described in the above embodiment, the secondary alignment network unit module is the alignment model described in the above embodiment, and the large-scale recognition network unit module is the preset standard human face recognition model described in the above embodiment; specifically as follows:

[0202] Among them, the training data preparation module can obtain third training data, identify the human face key points in the third sample human face image in the third training data, and input the third sample human face image after identifying the human face key points into the corresponding alignment data acquisition module of the original alignment template. In addition, the training data preparation module can also perform alignment processing on the human face key points in the third sample human face image through the target alignment template to obtain an aligned sample human face image, and input the obtained aligned sample human face image into the secondary alignment network unit module.

[0203] Among them, the secondary alignment network unit module can perform image reconstruction on the aligned sample human face image output by the training data preparation module to obtain a reconstructed aligned sample human face image. Then, perform feature extraction on the reconstructed aligned sample human face image through the recognition network unit module to obtain the human face feature information of the reconstructed aligned sample human face image. The human face recognition objective function calculation module can calculate the objective function value between the human face feature information of the reconstructed aligned sample human face image and the label information (i.e., the true identity information) of the third sample human face image. This objective function can select a classification function, such as softmax, etc.

[0204] The original alignment template corresponding alignment data acquisition module can perform alignment processing on the facial key points in the third sample face image based on the third sample face image after identifying the facial key points obtained through the training data preparation module, using the original alignment template, to obtain the target aligned sample face image.

[0205] The input of the large recognition network unit module can include the target aligned sample face image output by the original alignment template corresponding alignment data acquisition module and the reconstructed aligned sample face image output by the secondary alignment network unit module. The large recognition network unit module can fuse the target aligned sample face image and the reconstructed aligned sample face image, and perform feature extraction on the fused sample face image to obtain the reference facial feature information of the third sample face image.

[0206] The feature similarity loss function calculation module can calculate the loss function value between the reference facial feature information and the label information (i.e., the true identity information) of the third sample face image.

[0207] Among them, the loss value corresponding to the objective function output by the face recognition objective function calculation module is denoted as the first loss value, and the loss value output by the feature similarity loss function calculation module is denoted as the second loss value. It can be determined whether the termination condition of model training is satisfied according to the first loss value and the second loss value. If not, the face recognition objective function optimization module can adjust the parameters of the secondary alignment network unit module and the parameters of the recognition network unit module according to the first loss value and the second loss value until the training result meets the training termination condition, then the joint network unit module optimization is completed, that is, the joint optimization of the secondary alignment network unit module and the recognition network unit module is completed. The condition for terminating model training can be set that the number of iterations meets the set value, or the total loss value calculated according to the first loss value and the second loss value is less than the preset loss value to complete the training of the model.

[0208] In this embodiment, the alignment model and the face recognition model are jointly trained, specifically, the parameters of the alignment model and the face recognition model are fine-tuned.

[0209] After obtaining the jointly trained face recognition model and alignment model, module deployment can be performed to obtain a face recognition system, as Figure 2e shown.

[0210] Specifically, in this embodiment, during the network module training phase, the recognition network unit module (i.e., the face recognition model in the above embodiment) can be trained first, and then the original alignment template can be scaled at key points to obtain the target alignment template, so as to obtain rough alignment data through the target alignment template (specifically, it can be the aligned sample face image in step 202). Based on the rough alignment data, the secondary alignment network unit module (i.e., the alignment model in the above embodiment) is trained, and then the recognition network unit module and the secondary alignment network unit module are optimized as a whole, and they are jointly trained to obtain the jointly trained recognition network unit module and secondary alignment network unit module.

[0211] During the network module integration and deployment phase, it is mainly to combine and deploy the relevant modules obtained in the module training phase to form a complete solution. As Figure 2f shown, it is the model architecture of the face recognition system obtained after deployment. It can include a detection network module, a primary alignment module, a secondary alignment module, a recognition network module, and a feature comparison and search module. Among them, the secondary alignment module is specifically the jointly trained secondary alignment network unit module, and the recognition network module is specifically the jointly trained recognition network unit module. Among them, the detection network module can be used to detect the position of the face in the face image and identify the key points of the face in the face image; the function of the primary alignment module can specifically be to perform alignment processing on the face image using the target alignment template; the feature comparison and search module can perform face recognition based on the face feature information output by the recognition network module to obtain the identity information corresponding to the face image.

[0212] In the actual application process of the face recognition system obtained after deployment, the face image to be recognized can be input into the detection network module. The detection network module obtains a detection frame, and the detection frame and the face image enter the primary alignment module. The primary alignment module can perform alignment processing using the target alignment template to obtain an aligned face image. The aligned face image passes through the secondary alignment module for image reconstruction to obtain a finely aligned image - the reconstructed aligned face image; then, the recognition network module extracts the face feature information of the reconstructed aligned face image; finally, the feature comparison and search module performs face recognition processing on the face image based on the face feature information to obtain the face recognition result of the face image.

[0213] Due to the reconstruction alignment through the secondary alignment module, the reconstructed aligned face image output has a higher degree of adaptation to the recognition network, and this secondary alignment module is obtained through supervised training of a pre-set standard face recognition model with better performance. The reconstructed face image contains stronger recognition information, thereby improving the recognition accuracy of the entire face recognition system and making it adaptable to various complex application scenarios.

[0214] Optionally, the present application may adopt the integration technology of the model to merge the secondary alignment module and the recognition network module. As Figure 2g shown, a new recognition network module is obtained, so that the face recognition device provided by the present application can be deployed in practice without making major changes to the existing recognition system.

[0215] Optionally, in some embodiments, the present application may introduce a generative adversarial network (GAN, Generative Adversarial Networks), thereby replacing the primary alignment network and directly using the generative adversarial network for refined adjustment, which can eliminate the matching errors existing between various steps.

[0216] 204. The server acquires the face image to be recognized and recognizes the face key points of the face image.

[0217] 205. The server performs alignment processing on the face key points of the face image based on the target alignment template to obtain an aligned face image.

[0218] In this embodiment, after recognizing the face key points of the face image, face alignment can be performed. Face alignment can align face images at different angles into the same standard shape; specifically, face alignment can first locate the feature points on the face (i.e., face key points), and then transform the face image into the face in the alignment template, such as through geometric transformation (affine, rotation, scaling), so that each feature point is aligned (such as moving parts such as eyes and mouth to the positions corresponding to the eyes and mouth in the alignment template).

[0219] 206. The server performs image reconstruction on the aligned face image through the jointly trained alignment model to obtain a reconstructed aligned face image.

[0220] Optionally, in this embodiment, the step of "performing image reconstruction on the aligned face image to obtain a reconstructed aligned face image" may include:

[0221] Performing feature encoding on the aligned face image to obtain the encoded information of the aligned face image;

[0222] Performing feature reconstruction on the encoded information to obtain a reconstructed aligned face image.

[0223] 207. The server extracts the face feature information of the reconstructed aligned face image through the jointly trained face recognition model; based on the face feature information, face recognition processing is performed on the face image to obtain the face recognition result of the face image.

[0224] This application uses a two - level alignment method to perform adaptive regional matching on the face images entering the face recognition system, obtaining a more suitable face image for recognition to improve the accuracy of the face recognition system. The overall technical solution of this application is mainly divided into two stages: the network module training stage and the network module integration and deployment stage. In the network module training stage, first, the face recognition model (i.e., the recognition network unit module) is trained. Then, the sample face images of the training data are realigned through the target alignment template, which is obtained by scaling the original alignment template to obtain a larger area outside the face region. Next, the alignment model (i.e., the two - level alignment network unit module) is trained while fixing the face recognition model to obtain a refined alignment model adapted to this face recognition model. Finally, the face recognition model and the alignment model are jointly fine - tuned and trained. In the network module integration and deployment stage, these modules can be integrated and cooperate with the feature comparison and search module to form a complete face recognition system.

[0225] As can be seen from the above, in this embodiment, the server can train the face recognition model based on the second training data and the original alignment template to obtain the trained face recognition model; then train the alignment model based on the first training data and the target alignment template to obtain the trained alignment model. Then, the server jointly trains the trained face recognition model and the trained alignment model based on the third training data to obtain the jointly trained face recognition model and alignment model for application in face recognition. For example, the server can obtain the face image to be recognized and identify the face key points of the face image; align the face key points of the face image based on the target alignment template to obtain the aligned face image; reconstruct the aligned face image through the jointly trained alignment model to obtain the reconstructed aligned face image; extract the face feature information of the reconstructed aligned face image through the jointly trained face recognition model; and perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image. The embodiment of this application can perform secondary refined alignment on the aligned face image through image reconstruction, thereby obtaining a more accurate aligned image, that is, the reconstructed aligned face image, to improve the accuracy of face image recognition.

[0226] To better implement the above - mentioned method, the embodiment of this application also provides a face recognition device, as Figure 3 shown. The face recognition device can include an acquisition unit 301, an alignment unit 302, a reconstruction unit 303, an extraction unit 304, and an identification unit 305, as follows:

[0227] (1) Acquisition unit 301;

[0228] An acquisition unit, configured to acquire a face image to be recognized and recognize face key points of the face image.

[0229] (2) An alignment unit 302;

[0230] The alignment unit is configured to perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image.

[0231] (3) A reconstruction unit 303;

[0232] The reconstruction unit is configured to perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image.

[0233] Optionally, in some embodiments of the present application, the reconstruction unit may include an encoding subunit and a reconstruction subunit, as follows:

[0234] The encoding subunit is configured to perform feature encoding on the aligned face image to obtain encoding information of the aligned face image;

[0235] The reconstruction subunit is configured to perform feature reconstruction on the encoding information to obtain a reconstructed aligned face image.

[0236] (4) An extraction unit 304;

[0237] The extraction unit is configured to perform feature extraction on the reconstructed aligned face image to obtain face feature information of the reconstructed aligned face image.

[0238] (5) An identification unit 305;

[0239] The identification unit is configured to perform face recognition processing on the face image according to the face feature information to obtain a face recognition result of the face image.

[0240] Optionally, in some embodiments of the present application, the reconstruction unit is specifically configured to perform image reconstruction on the aligned face image through a trained alignment model to obtain a reconstructed aligned face image; the extraction unit is specifically configured to perform feature extraction on the reconstructed aligned face image through a trained face recognition model to obtain face feature information of the reconstructed aligned face image.

[0241] Optionally, in some embodiments of the present application, the face recognition device may further include a first training unit, and the first training unit is used to train the alignment model. Specifically, the first training unit is configured to obtain first training data, where the first training data includes first sample face images and real identity information of the first sample face images; based on the target alignment template, perform alignment processing on the face key points of the first sample face images to obtain aligned sample face images; through the alignment model, perform image reconstruction on the aligned sample face images to obtain reconstructed aligned sample face images; and train the alignment model according to the reconstructed aligned sample face images and the real identity information of the first sample face images.

[0242] Optionally, in some embodiments of the present application, the step of "training the alignment model according to the reconstructed aligned sample face images and the real identity information of the first sample face images" may include:

[0243] Extract the face feature information of the reconstructed aligned sample face images through the trained face recognition model to obtain the face feature information of the reconstructed aligned sample face images;

[0244] Determine the predicted identity information of the first sample face images according to the face feature information;

[0245] Adjust the parameters of the alignment model according to the real identity information and the predicted identity information of the first sample face images to obtain a trained alignment model.

[0246] Optionally, in some embodiments of the present application, the target alignment template is obtained by scaling the original alignment template; the trained face recognition model is trained based on the original alignment template.

[0247] Optionally, in some embodiments of the present application, the face recognition device may further include a second training unit, and the second training unit is used to train the face recognition model. Specifically, the second training unit is configured to obtain second training data, where the second training data includes second sample face images and real identity information of the second sample face images; based on the original alignment template, perform alignment processing on the face key points in the second sample face images to obtain target aligned sample face images; through the face recognition model, extract the face feature information of the target aligned sample face images to obtain the face feature information of the target aligned sample face images; and adjust the parameters of the face recognition model according to the predicted identity information corresponding to the face feature information and the real identity information of the second sample face images to obtain a trained face recognition model.

[0248] Optionally, in some embodiments of the present application, the face recognition device may further include a third training unit, and the third training unit is used to jointly train the alignment model and the face recognition model. Specifically, the third training unit is configured to obtain third training data, where the third training data includes third sample face images and the true identity information of the third sample face images; perform alignment processing on the face key points of the third sample face images based on the target alignment template to obtain aligned sample face images; perform image reconstruction on the aligned sample face images through the pre-trained alignment model to obtain reconstructed aligned sample face images; extract the face feature information of the reconstructed aligned sample face images through the pre-trained face recognition model; and adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images.

[0249] Optionally, in some embodiments of the present application, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images" may include:

[0250] Perform alignment processing on the face key points of the third sample face images based on the original alignment template to obtain target aligned sample face images, where the target alignment template is obtained by scaling the original alignment template;

[0251] Extract the reference face feature information of the third sample face images according to the target aligned sample face images and the reconstructed aligned sample face images through a preset standard face recognition model;

[0252] Adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face images.

[0253] Optionally, in some embodiments of the present application, the step of "adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face images" may include:

[0254] Calculate a first loss value between the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images;

[0255] Calculate a second loss value between the predicted reference identity information corresponding to the reference human face feature information and the true identity information of the third sample human face image;

[0256] According to the first loss value and the second loss value, adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model.

[0257] Optionally, in some embodiments of the present application, the step of "According to the first loss value and the second loss value, adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model" may include:

[0258] Fuse the first loss value and the second loss value to obtain a total loss value;

[0259] According to the total loss value, adjust the parameters of the pre-trained alignment model and the pre-trained human face recognition model.

[0260] Optionally, in some embodiments of the present application, the step of "Through a preset standard human face recognition model, according to the target aligned sample human face image and the reconstructed aligned sample human face image, extract the reference human face feature information of the third sample human face image" may include:

[0261] Fuse the target aligned sample human face image and the reconstructed aligned sample human face image to obtain a fused sample human face image;

[0262] Through a preset standard human face recognition model, perform feature extraction on the fused sample human face image to obtain the reference human face feature information of the third sample human face image.

[0263] As can be seen from the above, in this embodiment, the acquisition unit 301 can acquire a human face image to be recognized and identify the human face key points of the human face image; the alignment unit 302 performs alignment processing on the human face key points of the human face image based on the target alignment template to obtain an aligned human face image; the reconstruction unit 303 performs image reconstruction on the aligned human face image to obtain a reconstructed aligned human face image; the extraction unit 304 performs feature extraction on the reconstructed aligned human face image to obtain the human face feature information of the reconstructed aligned human face image; the recognition unit 305 performs human face recognition processing on the human face image according to the human face feature information to obtain the human face recognition result of the human face image. In the embodiment of the present application, through image reconstruction, secondary refined alignment of the aligned human face image can be performed, so as to obtain a more accurate aligned image, that is, the reconstructed aligned human face image, and the accuracy of human face image recognition can be improved.

[0264] The embodiment of the present application also provides an electronic device, such asFigure 4 As shown, it shows a schematic structural diagram of an electronic device involved in an embodiment of the present application. The electronic device can be a terminal or a server, etc. Specifically:

[0265] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0266] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling the data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby performing an overall detection of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 401.

[0267] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0268] The electronic device also includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby realizing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0269] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0270] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0271] It is possible to obtain a face image to be recognized and identify the face key points of the face image; perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; perform feature extraction on the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image.

[0272] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0273] As can be seen from the above, in this embodiment, it is possible to obtain a face image to be recognized and identify the face key points of the face image; perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; perform feature extraction on the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image. In the embodiment of the present application, through image reconstruction, secondary refined alignment can be performed on the aligned face image, so as to obtain a more accurate aligned image, that is, the reconstructed aligned face image, thereby improving the accuracy of face image recognition.

[0274] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0275] To this end, an embodiment of the present application provides a computer-readable storage medium storing multiple instructions that can be loaded by a processor to execute the steps in any of the face recognition methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0276] It is possible to obtain a face image to be recognized and recognize the face key points of the face image; perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image; extract features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; and perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image.

[0277] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.

[0278] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0279] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the face recognition methods provided by the embodiments of the present application, the beneficial effects achievable by any of the face recognition methods provided by the embodiments of the present application can be realized. For details, reference can be made to the previous embodiments and will not be elaborated here.

[0280] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the methods provided in various alternative implementations of the above face recognition aspect.

[0281] The above has introduced in detail a face recognition method, device, electronic device, and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A face recognition method, characterized in that, Including: Obtain a face image to be recognized, and recognize the face key points of the face image; Perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; Perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image, where the image reconstruction is used to adjust the positions of the face key points in the aligned face image again; Extract features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; Perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image.

2. The method according to claim 1, characterized in that The performing image reconstruction on the aligned face image to obtain a reconstructed aligned face image includes: Perform feature encoding on the aligned face image to obtain the encoding information of the aligned face image; Perform feature reconstruction on the encoding information to obtain a reconstructed aligned face image.

3. The method according to claim 1, wherein The performing image reconstruction on the aligned face image to obtain a reconstructed aligned face image includes: Perform image reconstruction on the aligned face image through a trained alignment model to obtain a reconstructed aligned face image; The extracting features from the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image includes: Extract features from the reconstructed aligned face image through a trained face recognition model to obtain the face feature information of the reconstructed aligned face image.

4. The method according to claim 3, characterized in that, The method further includes: Obtain first training data, where the first training data includes a first sample face image and the true identity information of the first sample face image; Perform alignment processing on the face key points of the first sample face image based on the target alignment template to obtain an aligned sample face image; Perform image reconstruction on the aligned sample face image through an alignment model to obtain a reconstructed aligned sample face image; Train the alignment model according to the reconstructed aligned sample face image and the true identity information of the first sample face image.

5. The method according to claim 4, wherein The training the alignment model according to the reconstructed aligned sample face image and the true identity information of the first sample face image includes: Extract features from the reconstructed aligned sample face image through a trained face recognition model to obtain the face feature information of the reconstructed aligned sample face image; Determine the predicted identity information of the first sample face image according to the face feature information; Adjust the parameters of the alignment model according to the true identity information and the predicted identity information of the first sample face image to obtain a trained alignment model.

6. The method according to claim 5, characterized in that, The target alignment template is obtained by performing scaling processing on an original alignment template; the trained face recognition model is trained based on the original alignment template.

7. The method according to claim 6, wherein Before the extracting features from the reconstructed aligned sample face image through a trained face recognition model to obtain the face feature information of the reconstructed aligned sample face image, it further includes: Obtain second training data, where the second training data includes second sample face images and the true identity information of the second sample face images; Based on the original alignment template, perform alignment processing on the face key points in the second sample face images to obtain target aligned sample face images; Through a face recognition model, extract face feature information of the target aligned sample face images; According to the predicted identity information corresponding to the face feature information and the true identity information of the second sample face images, adjust the parameters of the face recognition model to obtain a trained face recognition model.

8. The method according to claim 3, characterized in that, The method further includes: Obtain third training data, where the third training data includes third sample face images and the true identity information of the third sample face images; Based on the target alignment template, perform alignment processing on the face key points of the third sample face images to obtain aligned sample face images; Through a pre-trained alignment model, perform image reconstruction on the aligned sample face images to obtain reconstructed aligned sample face images; Through a pre-trained face recognition model, extract face feature information of the reconstructed aligned sample face images; According to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images, adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model.

9. The method according to claim 8, wherein The adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images includes: Based on the original alignment template, perform alignment processing on the face key points of the third sample face images to obtain target aligned sample face images, where the target alignment template is obtained by performing scaling processing on the original alignment template; Through a preset standard face recognition model, extract reference face feature information of the third sample face images according to the target aligned sample face images and the reconstructed aligned sample face images; According to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face images, adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model.

10. The method according to claim 9, wherein The adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the predicted identity information corresponding to the face feature information, the predicted reference identity information corresponding to the reference face feature information, and the true identity information of the third sample face images includes: Calculate a first loss value between the predicted identity information corresponding to the face feature information and the true identity information of the third sample face images; Calculate a second loss value between the predicted reference identity information corresponding to the reference face feature information and the true identity information of the third sample face images; Adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the first loss value and the second loss value.

11. The method according to claim 10, wherein The adjusting the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the first loss value and the second loss value includes: Fuse the first loss value and the second loss value to obtain a total loss value; Adjust the parameters of the pre-trained alignment model and the pre-trained face recognition model according to the total loss value.

12. The method according to claim 9, wherein The extracting, by a preset standard face recognition model, the reference face feature information of the third sample face image according to the target aligned sample face image and the reconstructed aligned sample face image includes: Fuse the target aligned sample face image and the reconstructed aligned sample face image to obtain a fused sample face image; Extract the reference face feature information of the third sample face image by performing feature extraction on the fused sample face image through a preset standard face recognition model.

13. A face recognition device, characterized in that, Includes: An acquisition unit, configured to acquire a face image to be recognized and recognize the face key points of the face image; An alignment unit, configured to perform alignment processing on the face key points of the face image based on a target alignment template to obtain an aligned face image; A reconstruction unit, configured to perform image reconstruction on the aligned face image to obtain a reconstructed aligned face image, where the image reconstruction is used to adjust the positions of the face key points in the aligned face image again; An extraction unit, configured to perform feature extraction on the reconstructed aligned face image to obtain the face feature information of the reconstructed aligned face image; An identification unit, configured to perform face recognition processing on the face image according to the face feature information to obtain the face recognition result of the face image.

14. An electronic device, characterized in that, Includes a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the operations in the face recognition method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the face recognition method according to any one of claims 1 to 12.

16. A computer program product comprising computer instructions, characterized in that, The computer instructions are stored in a computer-readable storage medium, and the processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the face recognition method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Face reconstruction method based on supervised learning depth autoencoder

    CN108537133A

  • Face recognition method and device, storage medium and electronic equipment

    CN112613488A