Semi-supervised facial feature point detection result evaluation method, device and storage medium

Through the semi-supervised learning method, the face feature point detection model is used to evaluate the accuracy of the prediction results of face images and identify difficult samples, which solves the problem that the accuracy of the model prediction results and the identification of difficult samples in the prior art is unable to evaluate the accuracy of the model prediction results and improves the stability and accuracy of the model.

CN113935414BActive Publication Date: 2025-05-06DALIAN NEUSOFT EDUCATION TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111199251.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-14
Publication Date
2025-05-06
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

The existing face feature point detection algorithm cannot evaluate the accuracy of the model's prediction results in practical applications, resulting in uncertain performance of the model in different scenarios, affecting system stability, and unable to identify difficult samples to improve the model.

Method used

A semi-supervised face feature point detection result evaluation method is proposed. By inputting face images into the trained convolutional neural network model, the prediction results of face feature point coordinates and the evaluation results for evaluating the accuracy of the current prediction result are obtained. The evaluation tags were calculated using the head posture similarity of the human face, the predicted position similarity of the human face and the 3D face similarity to determine the image as a difficult sample and updated training.

Benefits of technology

The quality evaluation of the prediction results of the face feature point detection model without manual labeling reference is realized, and the model's ability to identify difficult samples is recognized and improved, and the stability and accuracy of the model in practical applications are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113935414B_ABST
    Figure CN113935414B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device and storage medium for evaluating the results of semi-supervised facial feature point detection. On the basis of the traditional deep learning-based face detection algorithm, an evaluation result for evaluating the accuracy of the current prediction result is added, so that the algorithm model can evaluate the prediction quality without manual label reference, thereby discovering in which cases the trained model performs poorly, and thus making targeted improvements. For example: In actual application scenarios, each face image input to the model is not marked with the position of facial feature points. At this time, it is impossible to judge the quality of the algorithm prediction by comparing the manual labels with the prediction results. However, the method of the present invention can provide additional evaluation results for evaluating the accuracy of the current prediction results while predicting the coordinates of the facial feature points. In this way, it is possible to know which data the model predicts well and which data it predicts poorly during use, which facilitates targeted improvements to the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face detection, and in particular to a method, device and storage medium for evaluating semi-supervised face feature point detection results. Background Art

[0002] Facial feature points refer to five feature points such as the eyes, nose, and two corners of the mouth, or 28 feature points including more contour points, or even 64 or 128 facial feature points. The quality of facial feature points represents the quality of the face to a certain extent. For example, if the feature points are blocked or blurred, the face cannot be effectively recognized. In the early stage, selecting facial images with high-quality feature points for face recognition can not only improve the accuracy of face recognition, but also improve the efficiency of face recognition.

[0003] The facial feature point detection algorithm is an algorithm for detecting the location of facial feature points and is widely used in face detection applications. Existing facial feature point detection algorithms are mostly based on deep learning. They use trained models to obtain the predicted results of facial feature point coordinates, and then compare the differences between the model's predicted results and the manual labels to evaluate the accuracy of the model's predicted results.

[0004] However, when the trained model is used in actual application scenarios, the accuracy of the model prediction results cannot be evaluated because the input face images at this time are not marked with manual labels. This makes the prediction results of the facial feature point detection model uncertain, which will affect the stability of the entire algorithm system. At the same time, due to the lack of such a prediction result accuracy evaluation method, the model does not know which application scenarios are more difficult to predict (difficult samples), and thus does not know how to improve the model. Summary of the invention

[0005] In view of this, the present invention discloses a semi-supervised facial feature point detection result evaluation method, device and storage medium to achieve accurate evaluation of facial feature point detection results and identification of difficult samples.

[0006] To this end, the present invention provides the following technical solutions:

[0007] On the one hand, the present invention provides a semi-supervised facial feature point detection result evaluation method, comprising:

[0008] Inputting a face image into a trained face feature point detection model; the face feature point detection model is constructed based on a convolutional neural network;

[0009] Obtaining the facial feature point coordinate prediction result output by the facial feature point detection model and the evaluation result of evaluating the accuracy of the current prediction result; the evaluation result is a value in the range of [0, 1] calculated from the facial head posture similarity, facial prediction position similarity and 3D facial similarity obtained during the training process, wherein the closer the value is to 1, the higher the prediction quality, and the closer the value is to 0, the worse the prediction quality;

[0010] When the evaluation result is less than a preset threshold, it is determined that the currently input face image is a difficult sample.

[0011] Furthermore, the facial feature point detection model is trained in the following manner:

[0012] Obtaining a training data set including face images and manual labels of their corresponding face feature points;

[0013] Taking the data in the training data set as input, the facial feature point detection model is used to predict the 2D facial feature point coordinates corresponding to the input face and the facial feature point prediction result quality assessment value;

[0014] Based on the predicted 2D facial feature point coordinates corresponding to the input face and the manual labels of the facial feature points corresponding to the input face, respectively calculating the facial head posture similarity, the facial predicted position similarity and the 3D facial similarity;

[0015] Calculating an evaluation label based on the face head posture similarity, the face predicted position similarity and the 3D face similarity;

[0016] Calculate a first distance between the predicted 2D facial feature point coordinates corresponding to the input face and the artificial label of the facial feature point using a first loss function;

[0017] Calculate a second distance between the face feature point prediction result quality assessment value and the assessment label using a second loss function;

[0018] The sum of the value of the first loss function and the value of the second loss function is taken as the value of the target loss function;

[0019] The value of the target loss function is reduced by the gradient descent algorithm until it stops decreasing, so that the coordinates of the 2D facial feature points predicted by the model are continuously close to the manual labels of the facial feature points. At the same time, an additional evaluation value describing the accuracy of the prediction result can be output.

[0020] Furthermore, calculating the similarity of face and head postures includes:

[0021] Use 3D modeling technology to create a standard 3D face model in 3D space;

[0022] Constructing a geometric relationship between the artificial labels of facial feature points corresponding to the input face and the positions of corresponding points of the 3D face model;

[0023] Calculate a first head posture [x, y, z] of the input face according to the geometric relationship;

[0024] Obtain the 2D facial feature point coordinates corresponding to the input face predicted by the facial feature point detection model, and use the 2D facial feature point coordinates to calculate the second head posture [x p ,y p , z p ];

[0025] Based on the first head posture [x, y, z] and the second head posture [x p ,y p , z p ] Calculate the face head posture similarity S according to the following formula headpose :

[0026]

[0027] in, and They are the similarity weights of the head posture in the three degrees of freedom. The weights will be initialized to 1 before the model training starts. As the training progresses, the three parameters will automatically adjust their sizes so that the algorithm will continue to converge to the optimal solution.

[0028] Furthermore, calculating the similarity of the predicted face position includes:

[0029] Determine the first compact rectangular box Box that frames the face according to the manual labels of the facial feature points corresponding to the input image;

[0030] Determine the second compact rectangular frame Box that frames the face according to the 2D facial feature point coordinates corresponding to the input face predicted by the facial feature point detection model p ;

[0031] Based on the first compact rectangular box Box and the second compact rectangular box Box p The face prediction position similarity S is calculated according to the following formula position :

[0032]

[0033] Wherein, Area(Box) represents the area of ​​the first compact rectangular box, Area(Box p ) represents the area of ​​the second compact rectangular frame.

[0034] Furthermore, calculating the 3D face similarity includes:

[0035] Convert the input image into an RGBD image;

[0036] Obtain first depth information (X, Y) = [x1, y1, z1, x2, y2, z2, ..., x68, y68, z68] of each facial feature point in the manual label of the facial feature point corresponding to the input image from the RGBD image;

[0037] Obtain second depth information (Xp, Yp) = [x1p, y1p, z1p, x2p, y2p, z2p, ..., x68p, y68p, z68p] corresponding to each facial feature point output by the facial feature point detection model network;

[0038] Based on the first depth information (X, Y) = [x1, y1, z1, x2, y2, z2, ..., x68, y68, z68] and the second depth information (Xp, Yp) = [x1p, y1p, z1p, x2p, y2p, z2p, ..., x68p, y68p, z68p], the 3D face similarity S is calculated according to the following formula 3D :

[0039]

[0040] Furthermore, the evaluation label S is calculated in the following manner:

[0041] S=0.3*S headpose +0.4*S position +0.3*S 3D ;

[0042] Among them, S headpose Indicates the similarity of face head posture, S position Indicates the similarity of face prediction position, S 3D Indicates 3D face similarity.

[0043] Furthermore, it also includes:

[0044] Saving the difficult sample locally;

[0045] The difficult samples are manually marked and added to a training data set, and the facial feature point detection model is updated and trained using the training data set.

[0046] Furthermore, the facial feature point detection model is a ShufflenetV2 convolutional neural network model, and the width of the initial convolutional layer of the ShufflenetV2 convolutional neural network model is 16.

[0047] In another aspect, the present invention further provides a semi-supervised facial feature point detection result evaluation device, the device comprising:

[0048] An image input unit, used to input a face image into a trained face feature point detection model; the face detection point model is constructed based on a convolutional neural network;

[0049] A detection and evaluation unit, used to obtain the facial feature point coordinate prediction result output by the facial feature point detection model and an evaluation result for evaluating the accuracy of the current prediction result; the evaluation result is a value in the range of [0, 1] calculated from the facial head posture similarity, facial prediction position similarity and 3D facial similarity obtained during the training process, wherein the closer the value is to 1, the higher the prediction quality, and the closer the value is to 0, the worse the prediction quality;

[0050] The difficult sample determination unit is used to determine that the currently input face image is a difficult sample when the evaluation result is less than a preset threshold.

[0051] On the other hand, the present invention also provides a computer-readable storage medium, which stores a computer instruction set. When the computer instruction set is executed by a processor, the semi-supervised facial feature point detection result evaluation method provided above is implemented.

[0052] Advantages and positive effects of the present invention:

[0053] In the present invention, an evaluation result for evaluating the accuracy of the current prediction result is added on the basis of the traditional face detection algorithm based on deep learning, so that the algorithm model can evaluate the quality of the prediction result without manual label reference, so as to discover under what circumstances the trained model performs poorly, and thus make targeted improvements. For example: in actual application scenarios, each face image input to the model is not marked with the position of facial feature points. At this time, it is impossible to judge the quality of the algorithm prediction by comparing the manual labels with the prediction results. However, the method in the present invention can give an additional evaluation result for evaluating the accuracy of the current prediction result while predicting the coordinates of the facial feature points. In this way, it is possible to know which data the model predicts well and which data it predicts poorly during use, which facilitates targeted improvements to the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0055] Figure 1 This is a schematic diagram of the structure of a traditional detection model based on deep learning methods;

[0056] Figure 2 A schematic diagram of the structure of a facial feature point detection model in an embodiment of the present invention;

[0057] Figure 3 Schematic diagram of the structure of another facial feature point detection model in an embodiment of the present invention;

[0058] Figure 4 A flowchart of a semi-supervised facial feature point detection result evaluation method provided in an embodiment of the present invention;

[0059] Figure 5 It is a schematic diagram of the structure of the improved ShufflenetV2 convolutional neural network model in an embodiment of the present invention;

[0060] Figure 6 The present invention provides a block diagram of a semi-supervised facial feature point detection result evaluation device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0062] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0063] The facial feature point detection method based on deep learning uses a large amount of data with artificial labels for training. During the training process, the prediction results of the model can be compared with the artificial labels, so as to evaluate the quality of the model detection results. When the model training is completed and used in an actual production environment, the model will only output the feature point coordinate information corresponding to the input face image. Due to the lack of artificial labels, it is difficult to evaluate the quality of the algorithm detection results at this time. The present invention proposes a semi-supervised facial feature point detection model detection result evaluation method, which can evaluate the accuracy of the current detection results of the facial feature point detection model without relying on artificial labels. It can be applied to any business scenarios and industries related to facial feature point detection, especially those that require the model to be able to complete targeted updates for difficult samples that appear in business scenarios.

[0064] In an embodiment of the present invention, the application of a semi-supervised facial feature point detection result evaluation method can be summarized as follows: when executing a facial feature point detection model in an actual environment, for each detection, the model not only gives the predicted facial feature point coordinates, but also gives an evaluation result for evaluating the accuracy of the current prediction result, and the value of the evaluation result is between 0 and 1, and the larger the value, the higher the accuracy, and the smaller the value, the lower the accuracy. When the value of the evaluation result is less than a certain threshold, the face image currently input to the facial feature point detection model is considered to be a difficult sample.

[0065] Taking the training of n-point facial feature point detection algorithm as an example, the traditional facial feature point detection model structure based on deep learning method is as follows Figure 1 As shown in the figure. Taking the face image as input, the convolutional neural network (CNN) is used to extract the feature vector f, and the prediction result with dimension 2n is obtained after the fully connected network transformation. The loss function is used to calculate the prediction value X with dimension 2n pre With manual labels X of dimension 2n label The distance between them D = loss (X pre ,X label ), and continuously reduce the value D of the loss function through the gradient descent algorithm until it stops decreasing. That is, through this method, the predicted value of the model is constantly close to the artificial label of the input face.

[0066] The present invention provides two improved solutions based on the above-mentioned face feature point detection algorithm training. One improved solution is as follows: Figure 2 As shown, it is basically similar to the traditional method, but the dimension of the predicted value becomes 2n+1, where the extra dimension X e Used to evaluate the accuracy of the current prediction results. e Tags of X elabelThe calculation method is: the face head posture similarity obtained during the training process, the face prediction position similarity and the 3D face similarity are comprehensively calculated to obtain a value in the range of [0, 1]. The closer the value is to 1, the higher the prediction quality is, and the closer it is to 0, the worse the prediction quality is. The final loss function is defined as D = loss (X pre ,X label )+loss(X e ,X elabel ). Another improved solution is Figure 3 As shown, X is used to evaluate the accuracy of the current result. e Given by another branch of the network, X elabel And the definition of loss function is Figure 2 The method is the same.

[0067] After training, the present invention uses the output value X e Compare with the set threshold to determine whether the current input image is a difficult sample. According to the actual situation, if these difficult samples can be saved locally, use manual marking to mark these images, and then update the model for training. If these difficult samples cannot be saved, analyze the attributes of these data online, summarize the reasons for inaccurate model recognition, and then guide the improvement of the algorithm training.

[0068] After understanding the difficult samples, the model can be updated in a targeted manner to improve the accuracy of the model in business scenarios. In actual operation, there are two solutions: one is to analyze these difficult samples online (no need to save these data locally, which can better protect user privacy), so as to understand which input images the model's detection capabilities are insufficient and make targeted improvements; the other is to save the difficult samples locally after obtaining the user's permission, and then manually mark these samples and add them to the original data set, and use the new data set to update the face feature point detection model.

[0069] See also Figure 4 , which shows a flow chart of a method for evaluating a semi-supervised facial feature point detection result in an embodiment of the present invention, the method comprising:

[0070] S1: Training part:

[0071] Build a deep convolutional neural network model, such as Figure 5 As shown, the model in the present invention is an improved ShufflenetV2 convolutional neural network model, which is improved to simplify the shuffle structure of the network block, and can significantly improve the computing efficiency while maintaining accuracy. The specific improvements include:

[0072] A. Reduce the width of the initial convolutional layer of ShuffleNetV2 from 24 to 16, which can effectively improve the execution efficiency of the model while maintaining good accuracy.

[0073] B. The shuffle strategy of ShuffleNetV2 is simplified, which significantly improves the inference speed of the model while maintaining the same accuracy.

[0074] During training, a face image and the labels of the positions of the facial features and facial contour feature points corresponding to the face image [x1, y1, x2, y2, …, x68, y68] (artificial labels of facial feature points) are used as input, and a facial feature point detection model is used to obtain a prediction result, which includes: the coordinate positions of the feature points corresponding to the input face and an evaluation result; the evaluation result is used to evaluate the accuracy of the current prediction value;

[0075] The reason why the trained model can output the coordinates of facial feature points is that there are corresponding manual labels as guidance during training. The existing data set has manually marked coordinates of facial feature points that can guide model training. However, in order for the model to be able to output an evaluation of the quality of the prediction results, the label of the evaluation result needs to be obtained during algorithm training. The present invention uses relevant calculation theories to calculate the required quality assessment label based on the existing data and existing labels, that is, the label of the evaluation result. [Data labels that are manually marked can be called supervised learning. If there are also labels obtained indirectly through calculation, it is called semi-supervised learning]; the evaluation result label comprehensively evaluates the accuracy of the prediction from three aspects: face head posture similarity, face prediction position similarity, and 3D face similarity, and finally gives a quality evaluation value of 0 to 1. The closer the value is to 1, the higher the prediction quality, and the closer it is to 0, the worse the prediction quality.

[0076] The specific training process is as follows:

[0077] S101, obtaining a training data set including facial images and manual labels of their corresponding facial feature points;

[0078] S102, using the data in the training data set as input, using the facial feature point detection model to predict the 2D facial feature point coordinates corresponding to the input face and the facial feature point prediction result quality assessment value;

[0079] S103, based on the predicted 2D facial feature point coordinates corresponding to the input face and the manual labels of the facial feature points corresponding to the input face, respectively calculating the facial head posture similarity, the facial predicted position similarity and the 3D facial similarity;

[0080] S104, calculating an evaluation label based on the face head posture similarity, the face predicted position similarity and the 3D face similarity; specifically, the evaluation label S is calculated in the following manner:

[0081] S=0.3*S headpose +0.4*S positoon +0.3*S 3D ;

[0082] Among them, S headpose Indicates the similarity of face head posture, S position Indicates the similarity of face prediction position, S 3D Indicates 3D face similarity.

[0083] S105, using a first loss function to calculate a first distance between the predicted 2D facial feature point coordinates corresponding to the input face and the artificial label of the facial feature point;

[0084] S106, using a second loss function to calculate a second distance between the quality evaluation value of the facial feature point prediction result and the evaluation label;

[0085] S107, taking the sum of the value of the first loss function and the value of the second loss function as the value of the target loss function;

[0086] S108. The value of the target loss function is reduced by a gradient descent algorithm until it stops decreasing, so that the coordinates of the 2D facial feature points predicted by the model are continuously close to the manual labels of the facial feature points, and the network learns to autonomously predict the positions of the feature points corresponding to the input face. At the same time, an additional evaluation value describing the accuracy of the prediction result can be output to complete the training of the facial feature point detection model.

[0087] Among them, in S103, calculating the similarity of face and head posture includes:

[0088] Use 3D modeling technology to create a standard 3D face model in 3D space;

[0089] Constructing a geometric relationship between the artificial labels of facial feature points corresponding to the input face and the positions of corresponding points of the 3D face model;

[0090] Calculate a first head posture [x, y, z] of the input face according to the geometric relationship;

[0091] Obtain the 2D facial feature point coordinates corresponding to the input face predicted by the facial feature point detection model, and use the 2D facial feature point coordinates to calculate the second head posture [x p ,y p , z p ];

[0092] Based on the first head posture [x, y, z] and the second head posture [x p ,y p , z p ] Calculate the face head posture similarity S according to the following formula headpose :

[0093]

[0094] in, and They are the similarity weights of the head posture in the three degrees of freedom. The weights will be initialized to 1 before the model training starts. As the training progresses, the three parameters will automatically adjust their sizes so that the algorithm will continue to converge to the optimal solution.

[0095] Among them, calculating the similarity of the predicted face position includes:

[0096] Determine the first compact rectangular box Box that frames the face according to the manual labels of the facial feature points corresponding to the input image;

[0097] Determine the second compact rectangular frame Box that frames the face according to the 2D facial feature point coordinates corresponding to the input face predicted by the facial feature point detection model p ;

[0098] Based on the first compact rectangular box Box and the second compact rectangular box Box p The face prediction position similarity S is calculated according to the following formula position :

[0099]

[0100] Wherein, Area(Box) represents the area of ​​the first compact rectangular box, Area(Box p ) represents the area of ​​the second compact rectangular frame.

[0101] Among them, calculating the 3D face similarity includes:

[0102] Use Opencv to convert the input image into an RGBD image; obtain the first depth information (X, Y) = [x1, y1, z1, x2, y2, z2, ..., x68, y68, z68] of each facial feature point in the artificial label of the facial feature point corresponding to the input image from the RGBD image;

[0103] Obtain second depth information (Xp, Yp) = [x1p, y1p, z1p, x2p, y2p, z2p, ..., x68p, y68p, z68p] corresponding to each facial feature point output by the facial feature point detection model network;

[0104] Based on the first depth information (X, Y) = [x1, y1, z1, x2, y2, z2, ..., x68, y68, z68] and the second depth information (Xp, Yp) = [x1p, y1p, z1p, x2p, y2p, z2p, ..., x68p, y68p, z68p], the 3D face similarity S is calculated according to the following formula 3D :

[0105]

[0106] S2: Usage process:

[0107] After the model is trained, during use, the model will output the coordinates of the facial feature points and the evaluation value X for evaluating the prediction accuracy. e .

[0108] Specifically, a face image is input into a trained face feature point detection model; a face feature point coordinate prediction result output by the face feature point detection model and an evaluation result for evaluating the accuracy of the current prediction result are obtained; the evaluation result is a value in the range of [0, 1] calculated from the face head posture similarity, face prediction position similarity and 3D face similarity obtained during the training process, and the closer the value is to 1, the higher the prediction quality is, and the closer the value is to 0, the worse the prediction quality is;

[0109] S3: If X e If the value is greater than the manually set reliability threshold, no processing will be performed. If the value is less than the threshold, the currently collected face data can be saved locally after obtaining the user's consent, manually annotated, and then the model can be updated and trained using the newly annotated data. If the data cannot be saved due to privacy protection, the properties of these difficult samples can be analyzed online: including face angle, image occlusion, image quality, lighting factors, etc., and then the model can be improved based on the analysis results.

[0110] S4: Update training based on the feedback from S3.

[0111] In the embodiment of the present invention, an evaluation result for evaluating the accuracy of the current prediction result is added on the basis of the traditional face detection algorithm based on deep learning, so that the algorithm model can evaluate the prediction quality without manual label reference, thereby discovering in which cases the trained model performs poorly, and thus making targeted improvements. For example: in the actual application scenario, each face image input to the model is not marked with the position of facial feature points. At this time, it is impossible to judge the quality of the algorithm prediction by comparing the manual labels with the prediction results. However, the method in the present invention can give an additional evaluation result for evaluating the accuracy of the current prediction result while predicting the coordinates of the facial feature points. In this way, it is possible to know which data the model predicts well and which data it predicts poorly during use, which facilitates targeted improvements to the algorithm.

[0112] Corresponding to the semi-supervised facial feature point detection result evaluation method in the present invention, the present invention also provides a semi-supervised facial feature point detection evaluation device. Figure 6 , which shows a schematic diagram of the structure of a semi-supervised facial feature point detection and evaluation device in an embodiment of the present invention, the device comprises:

[0113] The image input unit 100 is used to input a face image into a trained face feature point detection model; the face feature point detection model is constructed based on a convolutional neural network;

[0114] The detection and evaluation unit 200 is used to obtain the facial feature point coordinate prediction result output by the facial feature point detection model and the evaluation result of evaluating the accuracy of the current prediction result; the evaluation result is a value in the range of [0, 1] calculated from the facial head posture similarity, facial prediction position similarity and 3D facial similarity obtained during the training process, and the closer the value is to 1, the higher the prediction quality is, and the closer the value is to 0, the worse the prediction quality is;

[0115] The difficult sample determination unit 300 is used to determine that the currently input face image is a difficult sample when the evaluation result is less than a preset threshold.

[0116] The model training unit 400 is used to train the face detection point model in the following manner:

[0117] Obtaining a training data set including face images and manual labels of their corresponding face feature points;

[0118] Taking the data in the training data set as input, the facial feature point detection model is used to predict the 2D facial feature point coordinates corresponding to the input face and the facial feature point prediction result quality assessment value;

[0119] Based on the predicted 2D facial feature point coordinates corresponding to the input face and the manual labels of the facial feature points corresponding to the input face, respectively calculating the facial head posture similarity, the facial predicted position similarity and the 3D facial similarity;

[0120] Calculating an evaluation label based on the face head posture similarity, the face predicted position similarity and the 3D face similarity;

[0121] Calculate a first distance between the predicted 2D facial feature point coordinates corresponding to the input face and the artificial label of the facial feature point using a first loss function;

[0122] Calculate a second distance between the face feature point prediction result quality assessment value and the assessment label using a second loss function;

[0123] The sum of the value of the first loss function and the value of the second loss function is taken as the value of the target loss function;

[0124] The value of the target loss function is reduced by the gradient descent algorithm until it stops decreasing, so that the coordinates of the 2D facial feature points predicted by the model are continuously close to the manual labels of the facial feature points. At the same time, an additional evaluation value describing the accuracy of the prediction result can be output.

[0125] As for the semi-supervised facial feature point detection and evaluation device of the embodiment of the present invention, since it corresponds to the semi-supervised facial feature point detection and evaluation method in the above embodiment, the description is relatively simple. For relevant similarities, please refer to the partial description in the above embodiment, which will not be described in detail here.

[0126] An embodiment of the present invention further discloses a computer-readable storage medium, which stores a computer instruction set. When the computer instruction set is executed by a processor, the semi-supervised facial feature point detection result evaluation method provided in any of the above embodiments is implemented.

[0127] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0128] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0129] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A semi-supervised facial feature point detection result evaluation method, characterized in that: include: Input the face image into the trained face feature point detection model; The facial feature point detection model is constructed based on a convolutional neural network; Obtaining the facial feature point coordinate prediction result output by the facial feature point detection model and the evaluation result of evaluating the accuracy of the current prediction result; the evaluation result is a value in the range of [0, 1] calculated from the facial head posture similarity, facial prediction position similarity and 3D facial similarity obtained during the training process, wherein the closer the value is to 1, the higher the prediction quality, and the closer the value is to 0, the worse the prediction quality; When the evaluation result is less than a preset threshold, the currently input face image is determined to be a difficult sample; The facial feature point detection model is trained in the following manner: Obtaining a training data set including face images and manual labels of their corresponding face feature points; Taking the data in the training data set as input, the facial feature point detection model is used to predict the 2D facial feature point coordinates corresponding to the input face and the facial feature point prediction result quality assessment value; Based on the predicted 2D facial feature point coordinates corresponding to the input face and the manual labels of the facial feature points corresponding to the input face, respectively calculating the facial head posture similarity, the facial predicted position similarity and the 3D facial similarity; Calculating an evaluation label based on the face head posture similarity, the face predicted position similarity and the 3D face similarity; Calculate a first distance between the predicted 2D facial feature point coordinates corresponding to the input face and the artificial label of the facial feature point using a first loss function; Calculate a second distance between the face feature point prediction result quality assessment value and the assessment label using a second loss function; The sum of the value of the first loss function and the value of the second loss function is taken as the value of the target loss function; The value of the target loss function is reduced by the gradient descent algorithm until it stops decreasing, so that the coordinates of the 2D facial feature points predicted by the model are continuously close to the manual labels of the facial feature points. At the same time, an additional evaluation value describing the accuracy of the prediction result can be output.

2. A semi-supervised facial feature point detection result evaluation method according to claim 1, characterized in that: Calculating the similarity of face and head posture includes: Use 3D modeling technology to create a standard 3D face model in 3D space; Constructing a geometric relationship between the artificial labels of facial feature points corresponding to the input face and the positions of corresponding points of the 3D face model; Calculate a first head posture [x, y, z] of the input face according to the geometric relationship; Obtain the 2D facial feature point coordinates corresponding to the input face predicted by the facial feature point detection model, and use the 2D facial feature point coordinates to calculate the second head posture [x p ,y p ,z p ]; Based on the first head posture [x, y, z] and the second head posture [x p ,y p ,z p ] Calculate the face head posture similarity S according to the following formula headpose : in, and They are the similarity weights of the head posture in the three degrees of freedom. The weights will be initialized to 1 before the model training starts. As the training progresses, the three parameters will automatically adjust their sizes so that the algorithm will continue to converge to the optimal solution.

3. A semi-supervised facial feature point detection result evaluation method according to claim 1, characterized in that: Calculating the similarity of face prediction positions includes: Determine the first compact rectangular box Box that frames the face according to the manual labels of the facial feature points corresponding to the input image; Determine the second compact rectangular frame Box that frames the face according to the 2D facial feature point coordinates corresponding to the input face predicted by the facial feature point detection model p ; Based on the first compact rectangular box Box and the second compact rectangular box Box p The face prediction position similarity S is calculated according to the following formula position : Among them, Area(Box) represents the area of ​​the first compact rectangular box, Area(BOx p ) represents the area of ​​the second compact rectangular frame.

4. A semi-supervised facial feature point detection result evaluation method according to claim 1, characterized in that: Calculating 3D face similarity includes: Convert the input image into an RGBD image; Obtain first depth information (X, Y) = [x1, y1, z1, x2, y2, z2, ..., x68, y68, z68] of each facial feature point in the manual label of the facial feature point corresponding to the input image from the RGBD image; Obtain second depth information (Xp, Yp) = [x1p, y1p, z1p, x2p, y2p, z2p, ..., x68p, y68p, z68p] corresponding to each facial feature point output by the facial feature point detection model network; Based on the first depth information (X, Y) = [x1, y1, z1, x2, y2, z2, ..., x68, y68, z68] and the second depth information (Xp, Yp) = [x1p, y1p, z1p, x2p, y2p, z2p, ..., x68p, y68p, z68p], the 3D face similarity S is calculated according to the following formula 3D :

5. The semi-supervised facial feature point detection result evaluation method according to claim 1, characterized in that: The evaluation label S is calculated as follows: S=0.3*S headpose +0.4*S position +0.3*S 3D ; Among them, S headpose Indicates the similarity of face head posture, S position Indicates the similarity of face prediction position, S 3D Indicates 3D face similarity.

6. A semi-supervised facial feature point detection result evaluation method according to claim 1, characterized in that: Also includes: Saving the difficult sample locally; The difficult samples are manually marked and added to a training data set, and the facial feature point detection model is updated and trained using the training data set.

7. A semi-supervised facial feature point detection result evaluation method according to claim 1, characterized in that: The facial feature point detection model is a ShufflenetV2 convolutional neural network model, and the width of the initial convolutional layer of the ShufflenetV2 convolutional neural network model is 16.

8. A semi-supervised facial feature point detection result evaluation device, characterized in that: The device comprises: An image input unit, used to input a face image into a trained face feature point detection model; the face feature point detection model is constructed based on a convolutional neural network; A detection and evaluation unit, used to obtain the facial feature point coordinate prediction result output by the facial feature point detection model and an evaluation result for evaluating the accuracy of the current prediction result; the evaluation result is a value in the range of [0, 1] calculated from the facial head posture similarity, facial prediction position similarity and 3D facial similarity obtained during the training process, wherein the closer the value is to 1, the higher the prediction quality, and the closer the value is to 0, the worse the prediction quality; The facial feature point detection model is trained in the following manner: Obtaining a training data set including face images and manual labels of their corresponding face feature points; Taking the data in the training data set as input, the facial feature point detection model is used to predict the 2D facial feature point coordinates corresponding to the input face and the facial feature point prediction result quality assessment value; Based on the predicted 2D facial feature point coordinates corresponding to the input face and the manual labels of the facial feature points corresponding to the input face, respectively calculating the facial head posture similarity, the facial predicted position similarity and the 3D facial similarity; Calculating an evaluation label based on the face head posture similarity, the face predicted position similarity and the 3D face similarity; Calculate a first distance between the predicted 2D facial feature point coordinates corresponding to the input face and the artificial label of the facial feature point using a first loss function; Calculate a second distance between the face feature point prediction result quality assessment value and the assessment label using a second loss function; The sum of the value of the first loss function and the value of the second loss function is taken as the value of the target loss function; The value of the target loss function is reduced by a gradient descent algorithm until it stops decreasing, so that the coordinates of the 2D facial feature points predicted by the model are continuously close to the manual labels of the facial feature points, and an additional evaluation value describing the accuracy of the prediction result can be outputted; The difficult sample determination unit is used to determine that the currently input face image is a difficult sample when the evaluation result is less than a preset threshold.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer instruction set, and when the computer instruction set is executed by the processor, the semi-supervised facial feature point detection result evaluation method provided by any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A face beauty evaluation method based on a depth model of sorting-guided regression

    CN109344855A

  • Face key point quality evaluation method and device, computer equipment and storage medium

    CN110879981A