A face alignment method, training method, device, and storage medium

By adjusting the penalty weight of the loss function using the angle prediction results in the face alignment network, the sample imbalance caused by the uneven distribution of the pose angle of the training image is solved, and the accuracy of face alignment and network performance are improved.

CN114220138BActive Publication Date: 2025-05-30ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111350466.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-05-30
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

The existing deep learning-based face alignment method leads to sample imbalance due to the uneven pose angle distribution of the training image, which in turn reduces the accuracy of face alignment.

Method used

By using the face alignment network in training, the face key points in the sample image are detected and angle prediction, the penalty weight of the first loss in the loss function is determined based on the angle prediction results, and the model parameters of the face alignment network are adjusted based on the penalty weight, the first loss and the second loss until the preset training end condition is met.

Benefits of technology

By adaptively adjusting the value of the loss function, the sample equalization is achieved, the sample imbalance is avoided, the training effect and accuracy of the face alignment network is improved, and the performance of face alignment is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114220138B_ABST
    Figure CN114220138B_ABST
Patent Text Reader

Abstract

The present application discloses a face alignment method, a training method, an apparatus, and a storage medium. The face alignment method includes: obtaining a face image; processing the face image based on a trained face alignment network to obtain an alignment result of the face to be detected; the trained face alignment network is obtained by the following method: using the face alignment network in training to detect key points of the face in a sample image to obtain a sample alignment result, and predicting the angle of the face to obtain an angle prediction result, so as to determine a penalty weight corresponding to a first loss, the first loss corresponding to the sample alignment result; based on the penalty weight, the first loss, and a second loss, adjusting the model parameters of the face alignment network until the face alignment network meets a preset training end condition, the second loss corresponding to the angle prediction result. By the above method, the present application can effectively solve the problem of unbalanced training samples of the face alignment network and improve the accuracy of face alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a face alignment method, a training method, a device, and a storage medium. Background Art

[0002] In existing deep learning-based face alignment methods, such as ECCN (Extensive Facial Landmark), DAN-Deep Alignment Network, Deep Convolutional Activation Feature (DeCaFA), or DU-Net (Quantized Densely Connected U-Nets), etc., during the process of training the face alignment network model, there will be a problem of sample imbalance caused by the different distributions of the pose angles of the training images, which in turn leads to a low face alignment accuracy of the model obtained using the training images. Summary of the Invention

[0003] This application provides a face alignment method, a training method, a device, and a storage medium, which can effectively solve the problem of unbalanced training samples of the face alignment network and improve the accuracy of face alignment.

[0004] To solve the above technical problems, the technical solution adopted by this application is: to provide a face alignment method, the method includes: obtaining a face image, the face image includes a face to be detected; based on the trained face alignment network, processing the face image to obtain an alignment result of the face to be detected; wherein, the trained face alignment network is obtained in the following manner: using the face alignment network during training, detecting the key points of the face included in the sample image to obtain a sample alignment result, and predicting the angle of the face to obtain an angle prediction result; based on the angle prediction result, determining a penalty weight corresponding to the first loss; the first loss is the loss corresponding to the sample alignment result in the loss function of the face alignment network; based on the penalty weight, the first loss, and the second loss, adjusting the model parameters of the face alignment network until the face alignment network meets the preset training end condition, obtaining the trained face alignment network; the second loss is the loss corresponding to the angle prediction result in the loss function.

[0005] To solve the above technical problems, another technical solution adopted by this application is: to provide a face alignment device, the face alignment device includes a memory and a processor connected to each other, wherein, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the face alignment method or the training method of the face alignment network in the above technical solution.

[0006] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium for storing a computer program, which, when executed by a processor, is used to implement the face alignment method or the training method of the face alignment network in the above technical solution.

[0007] Through the above solution, the beneficial effects of this application are as follows: during the process of training the face alignment network with the sample images in the training data, the key points of the faces in the sample images are detected by the face alignment network to obtain the sample alignment results, and the angles of the faces are detected by the face alignment network to obtain the angle prediction results. Then, according to the angle prediction results, the penalty weight of the first loss corresponding to the sample alignment result in the loss function of the face alignment network is determined. Then, based on the first loss, the second loss corresponding to the angle prediction result, and the penalty weight, the value of the loss function is calculated to adjust the model parameters of the face alignment network, and finally a trained face alignment network is obtained. Since the angle prediction results are used to determine the penalty weight of the first loss, the value of the loss function can be adaptively adjusted, thereby achieving sample balance and avoiding the problem of sample imbalance caused by certain factors. In actual use, first obtain a face image containing the face to be detected, and then use the face alignment network to process the face image to obtain the alignment result of the face to be detected, achieving the effect of face alignment. Since the penalty weight adjustment strategy is adopted during the training process to adjust the penalty weight of the sample alignment result in the loss function, the training effect of the face alignment network is improved, the accuracy of face alignment is improved, and the performance of the face alignment network is further enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0009] Figure 1 is a schematic flowchart of an embodiment of the face alignment method provided by this application;

[0010] Figure 2 is the effect diagram of face key point alignment;

[0011] Figure 3 is the effect diagram of image preprocessing;

[0012] Figure 4 is a schematic flowchart of an embodiment of the training method of the face alignment network provided by this application;

[0013] Figure 5It is a schematic flowchart of another embodiment of the training method of the face alignment network provided by this application;

[0014] Figure 6 It is a schematic diagram of the face alignment network provided by this application;

[0015] Figure 7 It is a schematic diagram of the structure perception layer provided by this application;

[0016] Figure 8 It is a comparison diagram between the deep feature fusion module and the ordinary convolution fusion module provided by this application;

[0017] Figure 9 It is a comparison diagram between the deep feature fusion module and the convolution module in the lightweight network MobileV2 provided by this application;

[0018] Figure 10 It is a flowchart for determining the penalty weight of the first loss provided by this application;

[0019] Figure 11 It is a schematic structural diagram of an embodiment of the face alignment device provided by this application;

[0020] Figure 12 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed Embodiments

[0021] The following will further describe this application in detail with reference to the accompanying drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate this application, but do not limit the scope of this application. Similarly, the following embodiments are only partial embodiments of this application rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0022] Referring to "embodiments" in this application means that the specific features, structures or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0023] It should be noted that the terms "first", "second", and "third" in this application are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0024] First, an explanation is given for the face alignment involved in this application. Face alignment, also known as face key point localization, is to find the position information of a person's eyes, eyebrows, nose, mouth, and contour in a given face image. The following provides a detailed description of the solution involved in this application.

[0025] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the face alignment method provided by this application. The method includes:

[0026] Step 11: Obtain a face image.

[0027] The face image includes the face to be detected. The face image can be obtained from the captured video data (including video or image) by a camera device (such as a surveillance camera or a camera, etc.) or from a database, and it can be an RGB image with a resolution of 112*112.

[0028] Step 12: Based on the trained face alignment network, process the face image to obtain the alignment result of the face to be detected.

[0029] The alignment result of the face to be detected includes the key points of the face to be detected. The face key points can be obtained by processing the face image to be detected through the face alignment network. Specifically, the face alignment network can detect the facial features and facial contour of the face, such as eyes, eyebrows, nose, mouth, and facial contour, etc., to obtain multiple face key points, as Figure 2 shown, Figure 2 (a) is the original face image. By processing this face image through the face alignment network, the corresponding face key points can be obtained (as shown in Figure 2 (b)). By superimposing the face key points on the face image, the result as shown in Figure 2(c) The results shown can be applied in many aspects, such as face recognition, expression analysis, face verification, face tracking, or face expression capture, etc.

[0030] In a specific embodiment, a face alignment network is trained using training data. The training data includes multiple sample images. The trained face alignment network is obtained in the following manner:

[0031] Using the face alignment network during training, detect the key points of the face included in the sample image to obtain a sample alignment result, and predict the angle of the face to obtain an angle prediction result; based on the angle prediction result, determine the penalty weight corresponding to the first loss, where the first loss is the loss corresponding to the sample alignment result in the loss function of the face alignment network; based on the penalty weight, the first loss, and the second loss, adjust the model parameters of the face alignment network until the face alignment network meets the preset training end condition to obtain the trained face alignment network, where the second loss is the loss corresponding to the angle prediction result in the loss function.

[0032] Furthermore, the face alignment network includes a main network and an auxiliary network. The main network is used to detect the key points of the face in the sample image to obtain a sample alignment result; the auxiliary network is used to predict the angle of the face to obtain an angle prediction result, so as to determine the penalty weight corresponding to the first loss in the loss function of the face alignment network based on the angle prediction result, which can achieve sample balance for training the face alignment network, achieve a better model training effect, and thus improve the performance of the face alignment network.

[0033] In a specific implementation manner, before using the main network to perform key point detection processing on the sample image to obtain a sample alignment result, face detection processing can be first performed on the sample image to obtain a sample face box, and then the image corresponding to the sample face box in the sample image is input into the main network, so that the main network performs key point detection on the image corresponding to the sample face box to obtain multiple sample key points. Specifically, common face detection algorithms may include: SingleShot MultiBox Detector (SSD) or Multitask Cascaded Convolutional Network (MTCNN), etc.

[0034] In other implementation manners, before using the main network to perform key point detection processing on the sample image to obtain a sample alignment result, face detection processing can also be first performed on the sample image to obtain a sample face box, then preprocess the image corresponding to the sample face box in the sample image to obtain a preprocessed sample image, and then input the preprocessed sample image into the main network so that the main network detects the key points of the face in the sample image.

[0035] Further, the preprocessing may include image cropping, image translation, image rotation, or image mirroring, etc. As Figure 3 shown, for example, the sample image as Figure 3 (a) can be cropped. For example, if the size of the sample image is 640*480, the bilinear interpolation can be used to reduce the image size to crop it into an image with a size of 112*112, so as to obtain the cropped image as Figure 3 (b). Image cropping can reduce the image size input to the main network, greatly reducing the number of parameters and the amount of calculation required for the main network to detect key points, and can optimize the training speed and effect. Or, the cropped face box can be translated according to a translation ratio of 1:10, and then the blank position after translation can be filled with black pixels; or, the translated image can be rotated or mirrored to obtain the Figure 3 image shown on the far right in Figure 3 (c)~ Figure 3 (e) are the mirror image, the image rotation image, and the rotated image after image mirroring in sequence. Among them, the rotation angle of the image is generally between -30° and 30°.

[0036] It can be understood that the solution provided in this embodiment is not limited to processing human faces, and other objects (such as vehicles or pedestrians) can also be aligned. This embodiment does not make any limitations in this regard.

[0037] In this embodiment, since the penalty weight of the first loss is determined by using the angle prediction result, the value of the loss function can be adaptively adjusted, thereby realizing sample balance and avoiding the problem of sample imbalance caused by certain factors; in actual use, the trained face alignment network can be used to process the face image to obtain the alignment result of the face to be detected in the face image, realizing the effect of face alignment; since the penalty weight adjustment strategy is adopted during the training process to adjust the penalty weight of the sample alignment result in the loss function, the training effect of the face alignment network is improved, the accuracy of face alignment is improved, and thus the performance of the face alignment network is enhanced.

[0038] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an embodiment of the training method of the face alignment network provided by this application. The method includes:

[0039] Step 41: Use the face alignment network during training to detect the key points of the face included in the sample image to obtain the sample alignment result.

[0040] Obtain training data, which includes multiple sample images. When training a face alignment network using the sample images in the training data, use the face alignment network to detect the key points of the face in the sample images, and obtain a sample alignment result. The sample alignment result includes multiple sample key points, and the sample key points are the key points of the face in the sample images. Specifically, the main network can detect the facial features and facial contours of the face, such as eyes, eyebrows, nose, mouth, and facial contours, etc., so as to obtain multiple sample key points.

[0041] Step 42: Use the face alignment network during training to predict the angle of the face included in the sample image to obtain an angle prediction result, and determine the penalty weight corresponding to the first loss based on the angle prediction result.

[0042] Use the face alignment network to predict the angle of the face in the sample image to generate an angle prediction result, so as to use the angle prediction result to determine the penalty weight corresponding to the first loss in the loss function, and then adjust the loss value of this training. The first loss is the loss corresponding to the sample alignment result in the loss function of the face alignment network.

[0043] Step 43: Based on the penalty weight, the first loss, and the second loss, adjust the model parameters of the face alignment network until the face alignment network meets the preset training end condition, and obtain the trained face alignment network.

[0044] After obtaining the sample alignment result, the penalty weight corresponding to the sample alignment result, and the angle prediction result, processing can be performed to generate the current training loss. Specifically, the loss corresponding to the sample alignment result (i.e., the first loss) and the loss corresponding to the angle prediction result in the loss function (i.e., the second loss) can be weighted and summed to obtain the current training loss.

[0045] After obtaining the current training loss, it can be determined whether the face alignment network meets the preset training end condition. If the face alignment network does not meet the preset training end condition, continue to train the face alignment network using the sample images and adjust the model parameters of the face alignment network until the face alignment network meets the preset training end condition. It can be understood that the operation of adjusting the model parameters of the face alignment network is similar to the operation of adjusting the model parameters in the related technology, and will not be elaborated here.

[0046] In a specific embodiment, the preset training end condition includes: determining whether the loss value converges, that is, whether the difference between the previous training loss and the current training loss is less than a set value. If the difference between the previous training loss and the current training loss is less than the set value, it is determined that the preset training end condition is reached; determining whether the current training loss is less than a preset loss value, where the preset loss value is a pre-set loss threshold. If the current loss value is less than the preset loss value, it is determined that the preset training end condition is reached; the number of training times reaches a set value (for example: training 10,000 times); or the accuracy rate obtained when testing using a test set reaches a set condition (for example: exceeding a preset accuracy rate), etc., which are not limited herein.

[0047] In the technical solution provided in this embodiment, the penalty weight corresponding to the sample alignment result is set based on the face angle prediction result, so as to adjust the loss value of this training, which can solve the problem of unbalanced sample quantity caused by different pose angles, improve the effect of model training, and thus ensure the performance of the face alignment network.

[0048] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of another embodiment of the training method of the face alignment network provided in this application. The method includes:

[0049] Step 51: Obtain sample images from the training data.

[0050] Step 52: Use the main network to detect the key points of the face included in the sample image, and obtain the sample alignment result based on the detected key points of the face in the sample image.

[0051] The main network may include a feature map extraction layer. The feature map extraction layer can be used to perform feature extraction processing on the sample image to obtain a training feature map. It can be understood that the training feature map is the feature map of the key points of the face in the sample image, and the number of training feature maps is the same as the number of key points of the face to be detected in the sample image (that is, the sample key points). For example: there are 68 sample key points in the sample image, and there are 68 corresponding training feature maps.

[0052] In a specific embodiment, as Figure 6 shown, the feature map extraction layer may include a first feature extraction layer, a structure perception layer, and a second feature extraction layer. The first feature extraction layer is used to perform feature extraction processing on the sample image to obtain second feature information; the structure perception layer is used to process the second feature information to obtain third feature information; the second feature extraction layer is used to perform feature extraction processing on the third feature information to obtain a training feature map.

[0053] The first feature extraction layer may include a convolutional layer and a pooling layer (not shown in the figure), which can perform convolution and pooling operations on the sample image, thereby performing preliminary feature extraction on the sample image to obtain second feature information. The structure perception layer can process the second feature information by means of graph convolution to obtain the structural relationship between the global information and the sample key points, thereby obtaining third feature information. Specifically, as Figure 7 shown, the structure perception layer can be a Graph-Based Global Reasoning Networks. The structure perception layer can first convert the feature space in the second feature information into a graph space, and then convert the graph space into a feature space. Among them, X represents the second feature information, and Y represents the third feature information. This embodiment can be an improvement based on the lightweight network MobileV2. By adding a structure perception layer, it is possible to obtain the structural relationship between the global information and the sample key points while ensuring the lightweight design of the neural network without increasing the complexity and computational amount of the model, thereby enhancing the depth of feature extraction.

[0054] Further, the second feature extraction layer may include at least one deep feature fusion module ( Figure 6 3 feature fusion modules are shown in the figure), and the processing process of the second feature extraction layer is as Figure 8 and Figure 9 shown. Figure 8 FIG. is a comparison diagram between a deep feature fusion module and a general convolutional fusion module. Among them, Figure 8 (a) is a deep feature fusion module, Figure 8 (b) is a general convolutional fusion module, and "Concat" is a feature fusion operation. Figure 9 FIG. is an operation comparison diagram between a deep feature fusion module (as shown in Figure 9 (a)) and a convolutional module in a general lightweight network MobileV2 (as shown in Figure 9 (b)) when the convolution strides are 1 and 2 respectively. Among them, Pointwise is point convolution, Depthwise is depth convolution, "conv(p*p)" is a convolution operation with a convolution kernel size of p*p, "ReLU6" is an activation operation, and "Add" is a feature superposition operation.

[0055] In this embodiment, in order to make the model lightweight, a feature fusion module is adopted to construct the main network. The deep feature fusion module adopts a strategy of fusing multi-scale information in the deep space, which can obtain more spatial correlation information in the channel space, can extract features of sample key points more fully, obtain the training feature map of sample key points, and does not need to improve the feature extraction performance of the model by adjusting the depth or width of the network structure. It can reduce the size of the model and greatly reduce the model complexity at the same time. Compared with the depthwise separable convolution, it not only controls the number of model parameters, but also can obtain richer feature information, which helps to improve the model performance.

[0056] In another specific embodiment, continue to refer to Figure 6 , the main network further includes a DSNT layer (Differentiable Spatial to Numerical Transform layer), which can perform feature extraction processing on the sample image at the feature map extraction layer. After obtaining the training feature map, the extracted training feature map is input into the DSNT layer, and then the DSNT layer is used to perform conversion processing on the training feature map, and the position coordinates of the corresponding key points are obtained according to the training feature map, so as to achieve the positioning effect based on the training feature map and obtain the sample alignment result.

[0057] Specifically, the DSNT layer is used to perform regularization processing on the training feature map to obtain a regularization result. Then, based on the coordinate information of the training feature map, the first matrix and the second matrix are calculated, and then the Hadamard product between the regularization result and the first matrix and the second matrix is calculated respectively to obtain the third matrix and the fourth matrix. Finally, based on the third matrix and the fourth matrix, the coordinates of the sample key points are obtained. By performing regularization processing on the training feature map and then using the regularization result to operate on the first matrix and the second matrix, overfitting can be prevented and the generalization ability of the model can be improved.

[0058] The steps of calculating the first matrix and the second matrix based on the coordinate information of the training feature map are specifically shown in the following formula:

[0059] X i,j = 2 * j - (n + 1) / n Formula (1)

[0060] Y i,j = 2 * i - (m + 1) / m Formula (2)

[0061] Wherein, in Formulas (1) to (2), X i,j represents the first matrix, Y i,jDenote the second matrix, where (i, j) is the coordinate information of the training image, and m and n represent the resolution of the training feature map. For example, if the resolution of the training feature map is 68*112, then m is 68 and n is 112. Substitute the coordinates and resolution of each training feature map into the above formula (1) and formula (2) respectively, and the corresponding first matrix and second matrix can be obtained. Then calculate the first matrix X i,j and the second matrix Y i,j respectively with the Hadamard product of the regularization results, so as to obtain the third matrix and the fourth matrix respectively. Then sum all the elements in the third matrix and sum all the elements in the fourth matrix to obtain the coordinates of the final sample key points.

[0062] Specifically, the coordinates of the sample key points include the abscissa and the ordinate. The sum of all the elements in the third matrix can be calculated to obtain the abscissa; the sum of all the elements in the fourth matrix can be calculated to obtain the ordinate. The number of regularization results is the same as the number of training feature maps. If there are 68 sample key points, when 68 training feature maps are extracted in the feature map extraction layer, the regularization results corresponding to the 68 training feature maps can be calculated respectively, that is, 68 regularization results. Then calculate the Hadamard product of the first matrix and the second matrix corresponding to each training feature map with the 68 regularization results respectively to obtain the third matrix and the fourth matrix respectively containing 68 elements. Then sum the 68 elements in the third matrix to obtain the abscissa corresponding to the sample key point, and sum the 68 elements in the fourth matrix to obtain the ordinate corresponding to the sample key point.

[0063] Step 53: Use the auxiliary network to process the sample alignment result output by the main network to obtain the angle prediction result.

[0064] The result output by the auxiliary network may include the angle prediction result and the expression prediction result. The feature map extraction layer in the main network is used to output the first feature information. The auxiliary network can be used to predict the expression of the face based on the first feature information to obtain the expression prediction result. The angle prediction result and the expression prediction result are the prediction results of the pose angle and the expression of the face respectively. It can be understood that in other embodiments, the auxiliary network can also identify / classify other attribute information of the sample, and can implement multi-task detection, which can assist the main network in key point localization, thereby improving the accuracy of key point localization of the main network.

[0065] Furthermore, the first feature information can be a two-dimensional image containing the face to be detected, that is, a two-dimensional face image. The auxiliary network can use a three-dimensional reconstruction algorithm (for example: the Pnp function in the OpenCV software) to convert the two-dimensional face image into a three-dimensional image, and then classify the three-dimensional image to obtain the angle prediction result, that is, the pose angle (i.e., Euler angle) of the three-dimensional image.

[0066] After using the main network to output the key coordinates of the sample and the auxiliary network to output the angle prediction result and the expression prediction result, the corresponding loss function can be used to calculate the current training loss generated by this training, so as to determine whether the face alignment network has completed training. The steps for calculating the current training loss are shown in steps 54 to 56:

[0067] Step 54: Calculate the loss between the sample key points and the labeled key points to obtain the first loss.

[0068] The training data includes the labeled key points corresponding to the sample image. The labeled key points are the reference key points in the sample image. By comparing the coordinates of the sample key points detected by the main network with the coordinates of the labeled key points, the loss between the sample key points and the corresponding labeled key points can be calculated, so as to obtain the positioning error generated by this training, that is, the first loss.

[0069] Step 55: Calculate the loss between the angle prediction result and the pose angle label to obtain the second loss.

[0070] The training data also includes the pose angle label and the expression label corresponding to the sample image. The error between the pose prediction result and the pose angle label and the error between the expression prediction result and the expression label can be calculated respectively; specifically, calculate the error between the angle prediction result and the pose angle label to obtain the second loss, and at the same time calculate the error between the expression prediction result and the expression label to obtain the third loss.

[0071] Step 56: Calculate the current training loss based on the first loss and the second loss.

[0072] The current training loss also includes a fourth loss, and the fourth loss is the regularization loss of the training feature map. Specifically, the regularization loss of the training feature map can be obtained when obtaining the training feature map, and then the first loss, the second loss, the third loss, and the fourth loss are weighted and summed to obtain the current training loss, as shown in the following formula:

[0073]

[0074] Among them, in formula (3), Loss is the current training loss, F(x) is the first loss, is the penalty weight of the first loss, G(y) is the second loss, H(z) is the third loss, μ 1 and μ 2 are the penalty weights corresponding to the second loss and the third loss respectively, and Z reg is the fourth loss; specifically, the penalty weights corresponding to the sample alignment result can be determined based on the auxiliary result, such as Figure 10As shown below, a method for determining the penalty weight of the first loss based on the angle prediction result and the preset mapping table will be introduced:

[0075] Step 61: Divide the pose angle range into multiple sub-ranges.

[0076] During the training process of the face alignment network, there is a problem of sample imbalance, which is specifically reflected in the fact that there are few large pose samples and many small pose samples in the training data. As a result, during network training, more emphasis is placed on learning the large pose samples with a large quantity, while ignoring the learning of small pose samples with a small quantity, thus affecting the generalization of the model. At this time, the corresponding results of the samples can be weighted according to the size of the pose angle, and a penalty weight is added to the first loss according to the pose angle of the samples, so as to solve the sample imbalance problem and ensure the effect of network training.

[0077] Generally speaking, large pose samples are samples with the absolute value of the pose angle in the range of 60° to 90°, small pose samples are samples with the absolute value of the pose angle in the range of 0° to 30°, and samples with the absolute value of the pose angle in the range of 30° to 60° are intermediate samples. Then, the pose angle range can be correspondingly divided into three sub-ranges, namely: 0° to 30°, 30° to 60°, and 60° to 90°.

[0078] Step 62: Statistically analyze the pose angles of all sample images to obtain the number of sample images corresponding to each sub-range.

[0079] Step 63: Based on the number of sample images corresponding to the sub-range, set the penalty weight corresponding to the sub-range to construct a preset mapping table.

[0080] The preset mapping table may include sub - ranges and penalty weights corresponding to the sub - ranges. The more the number of sample images corresponding to a sub - range, the smaller the penalty weight corresponding to the sub - range. Specifically, the corresponding penalty weights can be set according to the quantity ratio of the sample images in each sub - range, and the penalty weights are set inversely proportional to the quantity ratio. For example, it is statistically obtained that the number of sample images in the 0° - 30° sub - range among all sample images is 30, the number of sample images in the 30° - 60° sub - range is 20, and the number of sample images in the 60° - 90° sub - range is 10. That is, the quantity ratio of small - pose samples, intermediate samples, and large - pose samples is 3:2:1. Then, at this time, the penalty weight of small - pose samples can be set to 1, the penalty weight of intermediate samples can be set to 2, and the penalty weight of large - pose samples can be set to 3. It can be understood that the specific numerical values of the penalty weights can be set according to the actual value range of the penalty weights. The value range of the penalty weights can also be 1 - 10, 1 - 100, or 1 - 1000, etc. For example, when the value range of the penalty weights is 1 - 100, the penalty weight of small - pose samples can also be correspondingly set to 10, the penalty weight of intermediate samples can be set to 20, and the penalty weight of large - pose samples can be set to 30.

[0081] Step 64: Match the angle prediction result with the preset mapping table to obtain the penalty weight of the first loss.

[0082] The penalty weight of the first loss is the penalty weight corresponding to the angle prediction result. Matching the angle prediction result with the preset mapping table can obtain the penalty weight corresponding to the first loss. For example, if the angle prediction result of the sample image is 40°, then according to the matching of 40° in the preset mapping table, the corresponding penalty weight can be obtained as 2, that is, set to 2.

[0083] Step 57: Judge whether the face alignment network meets the preset training end condition based on the current training loss.

[0084] If the face alignment network does not meet the preset training end condition, it indicates that the performance of the face alignment network at this time does not meet the requirements. At this time, the model parameters of the face alignment network can be adjusted, and the face alignment network can continue to be trained using the sample images, that is, return to the step of obtaining sample images from the training data until the face alignment network meets the preset training end condition.

[0085] Step 58: If the face alignment network meets the preset training end condition, obtain the trained face alignment network.

[0086] If the face alignment network meets the preset training end condition, it means that the training of the face alignment network is completed. At this time, stop the training and obtain the trained face alignment network.

[0087] Based on the lightweight network MobileV2, in the main network for sample key point localization in this embodiment, a structure perception layer, a deep feature fusion module, and a DSNT layer are added. The structure perception layer is used to obtain the global information and the spatial structure relationship between the sample key points. By setting the deep feature fusion module, the sample image can be feature-extracted at multiple scales, which can not only reduce the size of the model (i.e., the model is a lightweight model), but also obtain rich feature information and spatial correlation in the channel space. The DSNT layer is used to convert the generated training feature map into the coordinates of the corresponding sample key points, making the model differentiable (i.e., the face alignment network in this embodiment is a fully differentiable model). Since the fully differentiable model directly optimizes the position coordinates of the key points of the face using the input instead of optimizing the heat map, it avoids the error caused by face inconsistency. At the same time, it maintains the spatial generalization ability of feature map regression, which helps to obtain a better localization effect, and it can be used on a small-resolution feature map, reducing memory occupancy and accelerating the inference speed, and the optimized face is consistent with the required face, without the problem of error lower bound. In addition, an auxiliary network is set to detect the attribute information such as the pose angle and expression of the sample image, which can achieve multi-task detection and assist the key point localization of the main network at the same time, making the localization of the sample key points more accurate. Furthermore, based on the pose angle prediction result, a penalty weight corresponding to the sample alignment result is set, and the problem of unbalanced sample numbers with different pose angles is solved through the sample weighting strategy, improving the effect of model training, so as to ensure the performance of the face alignment network.

[0088] Please refer to Figure 11 , Figure 11 FIG. is a schematic structural diagram of an embodiment of a face alignment device provided by the present application. The face alignment device 110 includes a memory 111 and a processor 112 connected to each other. The memory 111 is used to store a computer program, and when the computer program is executed by the processor 112, it is used to implement the face alignment method in the above embodiment or the training method of the face alignment network in the above embodiment.

[0089] In the solution of this embodiment, during model training, a sample weighting strategy and corresponding loss design are adopted, effectively alleviating the problem of sample imbalance; a deep feature fusion strategy is also adopted, making the model not only lightweight, but also capable of obtaining rich feature information and spatial correlation, and being more suitable for constructing a lightweight network model; in addition, for a lightweight network, since a lightweight network is generally relatively shallow and lacks information between face key points, this solution adds a structure perception layer at the front layer position of the lightweight network, which can obtain the connection between face key points and fuse it into the network; additionally, in the related art, due to the small model scale of the lightweight network, the positioning accuracy will be lost compared with the large model. To take into account the advantages of high generalization of the model based on feature map regression and overcome the disadvantages of the feature map, this solution adds a DSNT layer to the lightweight network, making the model differentiable and reducing the storage occupied by the model.

[0090] Please refer to Figure 12 , Figure 12 FIG. is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 120 is used to store a computer program 121, and when the computer program 121 is executed by a processor, it is used to implement the face alignment method in the above embodiment or the training method of the face alignment network in the above embodiment.

[0091] The computer-readable storage medium 120 can be various media that can store program codes, such as a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.

[0092] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0093] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0094] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically as individual units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0095] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.

Claims

1. A face alignment method, characterized in that, it includes: Obtain a face image, where the face image includes a face to be detected; Based on the trained face alignment network, process the face image to obtain an alignment result of the face to be detected; Among them, the trained face alignment network is obtained through the following method: Use the face alignment network during training to detect the key points of the face included in the sample image to obtain a sample alignment result, and predict the angle of the face to obtain an angle prediction result; Based on the angle prediction result, determine the penalty weight corresponding to the first loss, where the first loss is the loss corresponding to the sample alignment result in the loss function of the face alignment network; Based on the penalty weight, the first loss, and the second loss, adjust the model parameters of the face alignment network until the face alignment network meets the preset training end condition to obtain the trained face alignment network, where the second loss is the loss corresponding to the angle prediction result in the loss function.

2. The face alignment method according to claim 1, characterized in that, The face alignment network includes a main network and an auxiliary network. The step of detecting the key points of the face included in the sample image to obtain a sample alignment result includes: Use the main network to detect the key points of the face included in the sample image, and based on the detected key points of the face in the sample image, obtain the sample alignment result; The step of predicting the angle of the face to obtain an angle prediction result includes: Use the auxiliary network to process the sample alignment result output by the main network to obtain the angle prediction result.

3. The face alignment method according to claim 2, characterized in that, The step of adjusting the model parameters of the face alignment network based on the penalty weight, the first loss, and the second loss until the face alignment network meets the preset training end condition to obtain the trained face alignment network includes: Calculate the loss between the sample key points and the label key points to obtain the first loss; Calculate the loss between the angle prediction result and the pose angle label to obtain the second loss; Based on the first loss and the second loss, calculate the current training loss; Based on the current training loss, determine whether the face alignment network meets the preset training end condition; If not, adjust the model parameters of the face alignment network and continue to train the face alignment network using the sample image until the face alignment network meets the preset training end condition.

4. The face alignment method according to claim 3, characterized in that, The main network includes a feature map extraction layer and a DSNT layer. The step of using the main network to perform key point detection processing on the sample image to obtain the sample alignment result includes: Use the feature map extraction layer to perform feature extraction processing on the sample image to obtain a training feature map, where the training feature map is the feature map of the key points of the face in the sample image; The training feature map is processed by the DSNT layer to obtain the sample alignment result.

5. The face alignment method according to claim 4, wherein, the feature map extraction layer is further configured to output first feature information, and the method further includes: using the auxiliary network to predict the expression of the face based on the first feature information to obtain an expression prediction result; calculating an error between the expression prediction result and an expression label to obtain a third loss; performing weighted summation on the first loss, the second loss, the third loss, and a fourth loss to obtain the current training loss, where the fourth loss is a regularization loss of the training feature map; determining a penalty weight of the first loss based on the angle prediction result and a preset mapping table.

6. The face alignment method according to claim 5, wherein, the step of determining the penalty weight of the first loss based on the angle prediction result and the preset mapping table includes: dividing the pose angle range into multiple sub-ranges; counting the pose angles of all the sample images to obtain the number of sample images corresponding to each sub-range; setting a penalty weight corresponding to the sub-range based on the number of sample images corresponding to the sub-range to construct the preset mapping table; the preset mapping table includes the sub-range and the penalty weight corresponding to the sub-range, and the more the number of sample images corresponding to the sub-range, the smaller the penalty weight corresponding to the sub-range; matching the angle prediction result with the preset mapping table to obtain the penalty weight of the first loss.

7. The face alignment method according to claim 4, wherein, the coordinates of the sample key points include an abscissa and an ordinate, and the step of processing the training feature map by the DSNT layer to obtain the sample alignment result includes: performing regularization processing on the training feature map to obtain a regularization result; using the DSNT layer to calculate a first matrix and a second matrix based on the coordinate information of the training feature map; using the DSNT layer to calculate Hadamard products between the regularization result and the first matrix and the second matrix respectively to obtain a third matrix and a fourth matrix; summing all elements in the third matrix to obtain the abscissa; summing all elements in the fourth matrix to obtain the ordinate.

8. The face alignment method according to claim 4, wherein, the feature map extraction layer includes a first feature extraction layer, a structure perception layer, and a second feature extraction layer, and the step of using the feature map extraction layer to perform feature extraction processing on the sample image to obtain a training feature map includes: using the first feature extraction layer to perform feature extraction processing on the sample image to obtain second feature information; using the structure perception layer to process the second feature information to obtain third feature information; using the second feature extraction layer to perform feature extraction processing on the third feature information to obtain the training feature map.

9. A face alignment device, wherein, It includes a mutually connected memory and a processor. Among them, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the face alignment method described in any one of claims 1-8.

10. A computer-readable storage medium for storing a computer program, characterized in that when the computer program is executed by a processor, it is used to implement the face alignment method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Cross-pose face recognition method based on progressive neural network and attention mechanism

    CN112818850A

  • Face detection method

    US20210056293A1