A Pose-Robust Facial Feature Extraction Method Based on Symmetric Siamese Network

By adopting a symmetric twin network structure in the depth model, combining data preprocessing and side face data selection methods, the problem of poor effect of deep model in large-angle deflection face recognition is solved, and the recognition effect of large-angle deflection faces is improved and the performance of robust face recognition is improved.

CN114399825BActive Publication Date: 2025-06-17GUANGZHOU ZHIFU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210055607.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-06-17
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

The existing depth models have poor results in face recognition with large angle deflection, mainly due to self-occlusion and nonlinear distortion of faces, as well as the excessive deflection of face pictures at small angles in training data, resulting in the model's preference for faces deflection at small angles.

Method used

The pose robust face feature extraction method based on symmetric twin networks is adopted. Through data preprocessing, side face data selection and deep model structure construction and training, it is ensured that every two near-facing pictures of the neural network have two side face pictures during the training process, and a symmetric twin network is constructed to extract pose robust face features.

Benefits of technology

It has achieved improved recognition effect of large-angle deflected faces, balanced the neural network's preference for data, and improved the performance of pose robust face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399825B_ABST
    Figure CN114399825B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for extracting pose-robust face features based on a symmetric Siamese network, belonging to the technical field of image processing. Considering that the existing deep models have a better recognition effect on frontal faces than on profile faces and have a preference for frontal face data, the present invention designs a symmetric Siamese network to balance the neural network's preference for profile faces. By selecting profile face data and constructing a symmetric Siamese network, the profile face data and all face data are divided into two paths to train the network simultaneously to balance the neural network's preference for profile face data and near-frontal face pictures, overcoming the influence of the much larger number of near-frontal face pictures than profile face pictures in the training set on the deep model. At the same time, combining the idea of contrastive learning, pose-robust features are extracted. The present invention can improve the recognition effect of the deep model on face recognition with large pose deflections, achieve pose robustness in face recognition, and promote the development of related technical fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and relates to a method for extracting pose-robust face features based on a symmetric siamese network. Background Art

[0002] Some studies have shown that compared with the recognition effect of near-frontal faces, some deep models are not satisfactory for the recognition of faces with large-angle deflections. There are two main reasons. On the one hand, when a face undergoes an angular deflection, it is accompanied by self-occlusion and non-linear distortion of the face, which increases the difficulty of the model to extract side-face features. On the other hand, for the face dataset collected under unconstrained conditions, the number of face images with small-angle deflections far exceeds the number of face images with large-angle deflections, resulting in the trained deep model having a preference for face images with small-angle deflections.

[0003] Currently, there are also related methods to solve the pose problem of face recognition. These methods can be divided into face generation methods and face pose-robust feature extraction methods. The face generation method mainly converts faces with different poses into frontal face images and then performs recognition. This method often requires a large amount of paired frontal-side face data as training data, and most of the training data comes from the face dataset under constrained conditions. The generalization of the finally obtained model is worth considering. The face pose-robust feature extraction method aims to design a feature extractor so that the obtained features are not affected by the pose, and this method generally ignores the pose long-tail problem, that is, the problem of uneven distribution of near-frontal face images and side-face data in the training set. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method for extracting pose-robust face features based on a symmetric siamese network.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A method for extracting pose-robust face features based on a symmetric siamese network, the method comprising the following steps:

[0007] S1: Data preprocessing stage: Using a face detection network to crop the face part in the original image in an unconstrained pose and resize the size of the cutout to 128×128;

[0008] S2: Side-face data selection stage: Using the key points detected by the face detection network to estimate the face pose coefficient, and constructing side-face data according to the face deflection angle distribution and its corresponding face pose coefficient in the face dataset;

[0009] S3: Construction of the deep model structure and its training phase: Separate the entire face training dataset and the profile face dataset for training. Ensure that for every two nearly frontal face images, there are two profile face images when the neural network updates the gradient. Construct a symmetric Siamese network, which consists of two parts: a feature consistency learning sub-network and an identity consistency learning sub-network.

[0010] S4: Pose-robust face feature extraction: Input a face image with an arbitrary deflection angle into the network, and the output result of the second-to-last layer of the network is used as the pose-robust face feature.

[0011] Optionally, S1 includes the following specific steps:

[0012] S11: Normalize all face image data so that the mean of the entire image is 0 and the standard deviation is 1, even if the distribution of the original image in the r, g, b channels follows a normal distribution.

[0013] S12: Use a face detection network to crop the face part of the original image in an unconstrained pose, and resize the size of this cutout to 128×128 to ensure that the size of the feature map obtained after subsequent convolution operations on the image is consistent.

[0014] Optionally, S2 includes the following specific steps:

[0015] S21: Use the cosine definition to estimate the angle between the nose and two eyes, and take the minimum value as the face pose coefficient.

[0016] S22: Under the constraint dataset containing accurate pose information, estimate the face pose coefficients corresponding to different angles, and at the same time count the distribution of face images in each deflection angle interval of the face pose coefficients in the training set; select a specific face pose coefficient as the threshold, and the images smaller than this threshold are selected as profile data.

[0017] Optionally, S3 includes the following specific steps:

[0018] S31: Put all face images into the feature consistency learning sub-network. The input images are face images with two virtual poses. The structure of the feature consistency learning sub-network is a Siamese network.

[0019] S32: Put all profile data into the identity consistency learning sub-network. The input images are images of the same person with different poses. The structure of the identity consistency learning sub-network is a Siamese network.

[0020] S33: The feature consistency learning sub-network and the identity consistency learning sub-network share weights.

[0021] Optionally, S31 includes the following specific steps:

[0022] S311: The face images with virtual poses are generated by data augmentation, where data augmentation mainly includes: vertical flipping, rotation, and random cropping;

[0023] S312: Use the feature consistency loss to maximize the similarity of the two feature vectors obtained by the siamese network, ensure feature consistency, and thus enhance the discriminative features.

[0024] Optionally, S32 includes the following specific steps:

[0025] S321: Combine different poses of the same person in pairs to obtain the input with identity consistency;

[0026] S322: Use the identity consistency loss to maximize the similarity of the two feature vectors obtained by the siamese network, ensure feature consistency, and thus enhance the discriminative features.

[0027] Optionally, S33 includes the following specific steps:

[0028] S331: Weight sharing ensures that the feature consistency learning sub-network and the identity consistency learning sub-network are jointly trained, so as to ensure the participation of side face data when the network parameters are updated each time.

[0029] The beneficial effects of the present invention are as follows: The present invention can not only distinguish near-frontal face images, but also improve the recognition effect of face images with large pose deflections. By training the data of near-frontal faces and side faces in two paths, the preference of the neural network for data is balanced, and pose-robust face recognition is achieved.

[0030] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in preferred detail below in conjunction with the drawings, where:

[0032] Figure 1 is the core design example diagram of the present invention;

[0033] Figure 2 is the flow chart of selecting side face data according to the present invention; (a) is the diagram of selecting face key points and defining pose coefficients; (b) is the diagram of estimating multiple pose coefficients; (c) is the example diagram of selecting side face data;

[0034] Figure 3 This is a distribution diagram of different pose intervals of the VGGFace2 face dataset statistically by the present invention;

[0035] Figure 4 This is a schematic diagram of the feature consistency learning sub-network structure of the present invention;

[0036] Figure 5 This is a schematic diagram of the identity consistency learning sub-network structure of the present invention;

[0037] Figure 6 This is a schematic diagram of the effect of face image classification by the present invention. Specific implementation manners

[0038] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0039] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0040] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0041] Figure 1 This is a core design example diagram of the present invention, mainly including image preprocessing, side face data selection, network structure construction and two learning processes. The specific implementation steps are as follows:

[0042] Step 1: Data preprocessing stage: Use a face detection network to crop the face part in the original image in an unconstrained pose, and resize the size of this cutout to 128×128.

[0043] First, crop the face part through the face detection algorithm MTCNN, and then normalize all face image data so that the mean of the entire image is 0 and the standard deviation is 1, that is, the distribution of the original image in the r, g, b channels follows a normal distribution.

[0044] Step 2: Use the Figure 2 method flow in to select the side face data. Figure 2 is the flow chart for selecting side face data described in the present invention; (a) is the diagram for selecting face key points and defining pose coefficients; (b) is the diagram for estimating multi-pose coefficients; (c) is the example diagram for selecting side face data.

[0045] Step 201: Obtain the key points of the face through the MTCNN face detection algorithm. As shown in Figure (a), among them, the three key points of the left eye, right eye, and nose are the keys to evaluating the face pose coefficient. Specifically, the distance between the left eye and the right eye is denoted as a, the distance between the left eye and the nose is denoted as b, the distance between the right eye and the nose is denoted as c, the included angle formed by the nose, left eye, and right eye is denoted as θ1, and the included angle formed by the nose, right eye, and left eye is denoted as θ2. Among them, the face deflection coefficient is γ(θ), and the definition is as shown in Equation (1):

[0046] γ(θ) = min(cosθ1, cosθ2) (1)

[0047] Among them, cosθ can be obtained using the cosine theorem, as shown in Formulas (2) and (3):

[0048] cosθ1 = (a 2 + b 2 - c 2 ) / 2ab (2)

[0049] cosθ2 = (a 2 + c 2 - b 2 ) / 2ac (3)

[0050] Step 202: Estimate the face deflection coefficient corresponding to each angle on the MultiPIE face data. There is accurate face pose information in the MultiPIE face dataset. The face deflection coefficient corresponding to each angle is shown in Figure (b).

[0051] Step 203: Define the side faces in the dataset, and count the distribution of face images with different poses in the cross-pose face dataset VGGFace2 according to Steps 201 and 202. As Figure 3As shown, when the face deflection angle is greater than about 45 degrees, the number of face images is much smaller than that when the face deflection angle is less than about 45 degrees. Therefore, face images with a face deflection coefficient less than 0.2465 (face deflection angle less than about 45 degrees) are defined as profile faces.

[0052] Step 204: Design an algorithm to construct the profile face data in the VGGFace2 face dataset according to the profile face definition in Step 203. Some results of the profile face data are shown in Figure (c).

[0053] Step 3: Construct a symmetric siamese network to balance the network's preference for near-frontal and profile face images.

[0054] Step 301: Construction of the feature consistency learning sub-network, whose structure is a siamese network structure. As Figure 4 shown, first generate virtual pose face images, which are mainly generated by data augmentation methods, mainly including vertical flipping, rotation, and random cropping. Put two virtual pose face images into the feature consistency learning sub-network for training. The feature consistency loss function is used to constrain the training process to maximize the similarity of the two obtained feature vectors. Combining with the cross-entropy loss, discriminative features are obtained. The feature consistency loss is shown in Equation (4):

[0055]

[0056] where f(·) represents that the neural network maps the face image to a feature vector, T1(·) and T2(·) represent two different data augmentation methods, and X i represents the face image.

[0057] Step 302: Construction of the identity consistency learning sub-network, whose structure is a siamese network structure. As Figure 5 shown, first pair up different pose face images of the same person, and then input the images into the identity consistency learning sub-network for training. The identity consistency loss function is used to constrain the training process to maximize the similarity of the two obtained feature vectors. Combining with the cross-entropy loss, discriminative features are obtained. The feature consistency loss is shown in Equation (5):

[0058]

[0059] where and represent different pose images of the same person.

[0060] Step 4: Train the face data in two ways. For the feature consistency learning sub-network, its input is all the face images in the training set. For the identity consistency learning sub-network, its input is the side face data. The two parts of data are separated for training to ensure that the update gradient frequencies of the near frontal face and the side face for the network parameters are almost the same, thus balancing the model's preference for measurement data.

[0061] Step 401: To enable the two networks to work together, the two networks share weights to ensure that the near frontal face and the side face update the network parameters simultaneously.

[0062] Step 402: Select LightCNN as the backbone network of the symmetric network. Through a series of convolutional and fully connected operations of the neural network, finally map the face image to a 256-dimensional feature vector, which is used as the pose-robust feature.

[0063] To verify the effectiveness of the present invention, the following experiments were carried out:

[0064] 1. The effect of cross-pose face recognition;

[0065] 2. The experimental effect of feature visualization;

[0066] For the experiment, the present invention is carried out on the cross-pose face dataset (MultiPIE), mainly according to the setting-2 protocol standard in the MultiPIE dataset. This standard involves all 337 people in the MultiPIE, including the face images of the same person in four periods. Each person contains 13 pose variations and 20 illumination condition variations. All the images of the first 200 people are used in the training stage; during the test, one frontal image of the last 137 people under natural illumination conditions is selected as the gallery, and the other images of the last 137 people are used as the probe. In the experiment, the present invention only uses the trained network model and does not fine-tune on the MultiPIE dataset. The specific test process is to obtain the feature vectors of the face images in the gallery and the probe through the network, and match the features obtained from the probe with the features obtained from the gallery. The matching process is to calculate the cosine distance between the features. The larger the value, the more likely the face in the probe belongs to the category in the gallery. Table 1 shows the recognition results of the present invention on setting-2 of MultiPIE.

[0067] Table 1 Database test structure (%)

[0068] Face deflection angle ±90° ±75° ±60° ±45° ±30° ±15° Recognition rate (%) 88.31 96.62 99.29 99.29 99.75 99.94

[0069] In addition, as Figure 6As shown, the t-SNE algorithm is used to map all the obtained 256-dimensional features to 2D. Although LightCNN can well distinguish near frontal face images (solid circles), it cannot well distinguish large pose face images (semi-solid circles). Compared with the LightCNN network, the present invention can not only well separate the features of near frontal faces, but also separate the features of side faces.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for extracting pose-robust face features based on a symmetric Siamese network, characterized in that: The method includes the following steps: S1: Data preprocessing stage: Use a face detection network to crop the face part in the original image in an unconstrained pose, and resize the cropped image to 128×128; S2: Side face data selection stage: Use the key points detected by the face detection network to estimate the face pose coefficient, and construct side face data according to the face deflection angle distribution and its corresponding face pose coefficient in the face dataset; S3: Construction and training stage of the deep model structure: Separate the entire face training dataset and the side face dataset for training. Ensure that when the neural network updates the gradient, for every two nearly frontal face images, there are two side face images. Construct a symmetric siamese network, which consists of two parts: a feature consistency learning sub-network and an identity consistency learning sub-network; S4: Pose-robust face feature extraction: Input a face image with an arbitrary deflection angle into the network, and the output result of the penultimate layer of the network is used as the pose-robust face feature; The S1 includes the following specific steps: S11: Normalize all face image data so that the mean of the entire image is 0 and the standard deviation is 1, that is, the distribution of the original image in the r, g, b channels follows a normal distribution; S12: Use a face detection network to crop the face part in the original image in an unconstrained pose, and resize the cropped image to 128×128 to ensure that the feature map sizes are consistent after subsequent convolution operations on the image; The S2 includes the following specific steps: S21: Use the cosine definition to estimate the angles between the nose and two eyes, and take the minimum value as the face pose coefficient; S22: Under the constraint dataset containing accurate pose information, estimate the face pose coefficients corresponding to different angles, and at the same time count the distribution of face images in each deflection angle interval of the face pose coefficients in the training set; Select a specific face pose coefficient as the threshold, and the images smaller than this threshold are selected as side face data; The S3 includes the following specific steps: S31: Put all face images into the feature consistency learning sub-network. The input images are face images with two virtual poses. The structure of the feature consistency learning sub-network is a siamese network; S32: Put all side face data into the identity consistency learning sub-network. The input images are images of the same person with different poses. The structure of the identity consistency learning sub-network is a siamese network; S33: The feature consistency learning sub-network and the identity consistency learning sub-network share weights; The S31 includes the following specific steps: S311: The face images with virtual poses are generated by data augmentation, where data augmentation includes: vertical flipping, rotation, and random cropping; S312: Maximize the similarity of the two feature vectors obtained by the siamese network using the feature consistency loss to ensure feature consistency, thereby enhancing discriminative features; The S32 includes the following specific steps: S321: Combine different poses of the same person in pairs to obtain the input for identity consistency; S322: Maximize the similarity of the two feature vectors obtained by the siamese network using identity consistency loss to ensure feature consistency, thereby enhancing discriminative features; The said S33 includes the following specific steps: S331: Weight sharing ensures that the feature consistency learning sub-network and the identity consistency learning sub-network are jointly trained, so as to ensure the participation of profile face data when the network parameters are updated each time.

Citation Information

Patent Citations

  • Twin neural network training method for face verification

    CN109117744A

  • Cross-pose face recognition method based on progressive neural network and attention mechanism

    CN112818850A