A method for generating a face key point detection model, a detection method, and an electronic device

By transforming the coordinates of the face frame and key point positions in the training sample set, a face key point detection model is generated, which solves the problem that the face detection frame error affects the detection accuracy, and improves the robustness and detection accuracy of the face key point positioning.

CN114596600BActive Publication Date: 2025-07-29WUHAN TCL CORP RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011387481.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-01
Publication Date
2025-07-29
Estimated Expiration
2040-12-01

AI Technical Summary

Technical Problem

The accuracy of positioning of the existing face key point detection model for face key points depends on the face detection results. The error of the face detection frame will affect the accuracy of the key point positioning, resulting in a decrease in detection accuracy.

Method used

By transforming the coordinates of the face frame and key point position coordinates in the training sample set, the second face frame and key point position coordinates are generated, and these coordinates are used to train the preset network model to generate a face key point detection model.

Benefits of technology

The dependence of the face key point detection model on face detection results is reduced, and the robustness and detection accuracy of face key point positioning are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596600B_ABST
    Figure CN114596600B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a face key point detection model, a detection method and an electronic device. The generation method performs coordinate transformation on the first face frame position coordinates and the first face key point position coordinates corresponding to the training images in the training sample set to obtain the second face frame position coordinates and the second face key point position coordinates; trains a preset network model according to the first face image corresponding to the second face frame position coordinates and the second face key point position coordinates to generate a face key point detection model. The present invention trains the face key point detection model with the first face image corresponding to the second face frame position coordinates and the second face key point position coordinates after coordinate transformation, reduces the dependence of the face key point detection model on the face detection result, and improves the robustness of the face key point detection model for face key point positioning without additionally increasing training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face detection, and particularly to a method for generating a face key point detection model, a detection method and an electronic device. Background Art

[0002] Face key point detection is the premise and breakthrough of face-related technologies such as face verification, face recognition, attribute calculation, expression recognition, and pose estimation. Due to the automatic learning and continuous learning capabilities of deep learning, more and more research has been conducted on face key point detection technology based on deep learning. This detection technology inputs the face image detected by a face detector into a pre-trained face key point detection model, and automatically locates the facial key feature points according to the input face image.

[0003] In the real scene, the face detection result is a crucial link affecting the detection accuracy of the face key point detection model. The incorrect estimation of the face global structure by the face detection box will directly lead to inaccurate face key point positioning, and the correctness of the face global often depends on the performance of the face detector. A too large face box will contain redundant background information, and a too small face box will lose the structural information of the face. Both of these results will directly affect the result output of the face key point detection model.

[0004] Therefore, the existing technologies still need to be improved and developed. Summary of the Invention

[0005] In view of the above-mentioned defects of the existing technologies, the present invention provides a method for generating a face key point detection model, a detection method and an electronic device, aiming to solve the problem that the positioning accuracy of the existing face key point detection model for face key points depends on the face detection result.

[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0007] A method for generating a face key point detection model, including:

[0008] Obtaining the position coordinates of the first face box and the position coordinates of the first face key points corresponding to the training images in the training sample set; wherein, the training images contain faces;

[0009] Performing coordinate transformation on the position coordinates of the first face box and the position coordinates of the first face key points to obtain the position coordinates of the second face box and the position coordinates of the second face key points;

[0010] Training a preset network model according to the first face image corresponding to the position coordinates of the second face box and the position coordinates of the second face key points to generate a face key point detection model.

[0011] The method for generating a face key point detection model, wherein the step of performing coordinate transformation on the first face frame position coordinates and the first face key point position coordinates to obtain the second face frame position coordinates and the second face key point position coordinates includes:

[0012] Performing coordinate transformation on the first face frame position coordinates according to a pre-generated scaling factor to obtain the second face frame position coordinates; performing coordinate transformation on the first face key point position coordinates according to the second face frame position coordinates to obtain the second face key point position coordinates.

[0013] The method for generating a face key point detection model, wherein the method for generating the scaling factor includes:

[0014] Randomly generating the scaling factor within a preset range.

[0015] The method for generating a face key point detection model, wherein the step of training a preset network model according to the first face image corresponding to the second face frame position coordinates and the second face key point position coordinates to generate a face key point detection model includes:

[0016] Inputting the first face image corresponding to the second face frame position coordinates into the preset network model to generate the third face key point position coordinates corresponding to the first face image;

[0017] Correcting the model parameters of the preset network model according to the second face key point position coordinates and the third face key point position coordinates, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the preset network model meets the preset conditions to generate a face key point detection model.

[0018] The method for generating a face key point detection model, wherein the step of correcting the model parameters of the preset network model according to the second face key point position coordinates and the third face key point position coordinates, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the preset network model meets the preset conditions includes:

[0019] Obtaining a loss value according to the second face key point position coordinates and the third face key point position coordinates. If the loss value is greater than or equal to a preset threshold

[0020] , then correcting the model parameters of the preset network model according to a preset parameter learning rate, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the loss value is less than the preset threshold.

[0021] The method for generating a face key point detection model, wherein, after the step of performing coordinate transformation on the first face frame position coordinates and the first face key point position coordinates to obtain the second face frame position coordinates and the second face key point position coordinates, the following steps are included:

[0022] Crop the training images in the training sample set according to the second face frame position coordinates to obtain the second face images corresponding to the second face position coordinates;

[0023] Perform preprocessing operations on the second face images to obtain the first face images corresponding to the second face position coordinates.

[0024] The method for generating a face key point detection model, wherein the step of performing preprocessing operations on the second face images to obtain the first face images corresponding to the second face position coordinates includes:

[0025] Perform scale scaling on the second face images, and perform normalization operations on the second face images after scale scaling to obtain the first face images corresponding to the second face position coordinates.

[0026] The method for generating a face key point detection model, wherein the step of performing scale scaling on the second face images includes:

[0027] Obtain the input dimension of the preset network model;

[0028] Perform scale scaling on the second face images according to the input dimension of the preset network model and the resolution of the second face images.

[0029] The method for generating a face key point detection model, wherein, before the step of cropping the training images in the training sample set according to the second face frame position coordinates to obtain the second face images corresponding to the second face position coordinates, the following steps are included:

[0030] Perform image enhancement on the training images in the training sample set, and use the training images after image enhancement as the training images in the training sample set.

[0031] A face key point detection method, which is applied to the face key point detection model generated by the method for generating a face key point detection model, and includes:

[0032] Obtain the target face image corresponding to the target image; wherein the target face is included in the target image;

[0033] Input the target face image into the face key point detection model to obtain the fourth face key point position coordinates corresponding to the target image;

[0034] Perform coordinate transformation on the position coordinates of the fourth face key points to obtain the position coordinates of the target face key points corresponding to the target image.

[0035] The face key point detection method, wherein the step of obtaining the target face image corresponding to the target image includes:

[0036] Obtain the position coordinates of the third face frame corresponding to the target image;

[0037] Crop the target image according to the position coordinates of the third face frame to obtain the third face image corresponding to the target image;

[0038] Perform scale scaling and normalization operations on the third face image to obtain the target face image corresponding to the target image.

[0039] The face key point detection method, wherein the step of performing coordinate transformation on the position coordinates of the fourth face key points to obtain the position coordinates of the target face key points corresponding to the target image includes:

[0040] Perform coordinate transformation on the position coordinates of the fourth face key points according to the position coordinates of the third face frame to obtain the position coordinates of the target face key points corresponding to the target image.

[0041] A face key point detection model generation device, which includes:

[0042] A coordinate acquisition module, configured to acquire the position coordinates of the first face frame and the position coordinates of the first face key points corresponding to the training images in the training sample set; wherein, the training images contain human faces;

[0043] A coordinate transformation module, configured to perform coordinate transformation on the position coordinates of the first face frame and the position coordinates of the first face key points to obtain the position coordinates of the second face frame and the position coordinates of the second face key points;

[0044] A model generation module, configured to train a preset network model according to the first face image corresponding to the position coordinates of the second face frame and the position coordinates of the second face key points to generate a face key point detection model.

[0045] A terminal, which includes: a processor and a storage medium communicatively connected to the processor, the storage medium is adapted to store multiple instructions; the processor is adapted to call the instructions in the storage medium to execute the steps in the above-mentioned face key point detection model generation method, or the steps in the face key point detection method.

[0046] A storage medium stores multiple instructions, where the instructions are adapted to be loaded and executed by a processor to perform the steps in the method for generating a face key point detection model described above, or the steps in the face key point detection method described above.

[0047] Advantages of the present invention: The present invention trains a face key point detection model using the first face image corresponding to the position coordinates of the second face frame after coordinate transformation and the position coordinates of the second face key points, reducing the dependence of the face key point detection model on the face detection result, and improving the robustness of the face key point detection model for face key point positioning without additionally increasing training data. Description of the Drawings

[0048] Figure 1 is a flowchart of an embodiment of a method for generating a face key point detection model provided in Embodiment 1 of the present invention;

[0049] Figure 2 is a flowchart of an embodiment of a face key point detection method provided in Embodiment 2 of the present invention;

[0050] Figure 3 is a functional schematic diagram of a face key point detection model generation device provided in Embodiment 3 of the present invention;

[0051] Figure 4 is a functional schematic diagram of a terminal provided in Embodiment 4 of the present invention. Detailed Embodiments

[0052] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The description of at least one exemplary embodiment is actually only illustrative and in no way restrictive of the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0054] The method for generating a face key point detection model and the face key point detection method provided by the present invention can be applied to a terminal. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, mobile phones, tablet computers, in-vehicle computers, and portable wearable devices. The terminal of the present invention uses a multi-core processor. Among them, the processor of the terminal can be at least one of a central processing unit (CPU), a graphics processing unit (GPU), a video processing unit (VPU), etc.

[0055] Embodiment 1

[0056] Based on the deep learning face key point detection technology, on the basis of face detection, the face image is input into a pre-trained face key point detection model, and the position information of the face key points is automatically located, such as the position information of the eyes, the tip of the nose, the corner of the mouth points, the eyebrows, and the contours of each part of the face. The input is the face image, and the output is the position information of the face key points.

[0057] In the real scene, the expression, illumination, and occlusion conditions of the face will affect the accuracy of the output results of the face key point model. The face detection result is a crucial link affecting the detection accuracy of the face key point detection model. The wrong estimation of the face global structure by the face detection box will directly lead to inaccurate face key point positioning, and the correctness of the face global often depends on the performance of the face detector. Too large a face box will contain redundant background information, and too small a face box will lose the structural information of the face. Both of these results will directly affect the result output of the face key point detection model.

[0058] To solve the above problems, in Embodiment 1 of the present invention, a method for generating a face key point detection model is provided. Please refer to Figure 1 , Figure 1 which is a flowchart of an embodiment of a method for generating a face key point detection model provided by the present invention.

[0059] In an embodiment of the present invention, the method for generating the face key point detection model has three steps:

[0060] S100. Obtain the position coordinates of the first face box and the position coordinates of the first face key points corresponding to the training images in the training sample set; wherein, the training images contain faces.

[0061] Since the incorrect estimation of the global structure of a human face by a face detection box will directly lead to inaccurate localization of human face key points, in this embodiment, before training a preset network model using the training images in a training sample set, the position coordinates of a first face box and the position coordinates of first human face key points corresponding to the training images in the training sample set are obtained. Wherein, the training sample set contains multiple training images for training the preset network model, and each training image contains a human face. The position coordinates of the first face box are the upper left vertex coordinates and the lower right vertex coordinates corresponding to the human face in each training image, that is, the upper left vertex coordinates of the first face box and the lower right vertex coordinates of the first face box. The rectangular area determined by the line connecting the upper left vertex coordinates of the first face box and the lower right vertex coordinates of the first face box as the diagonal is the first face box corresponding to the training image in the training sample set. The position coordinates of the first human face key points are the position coordinates of the human face key points corresponding to the human face in the training images in the training sample set. For example, the position coordinates corresponding to the eyes, nose tip, mouth corner points, eyebrows, etc. on the human face.

[0062] S200. Perform coordinate transformation on the position coordinates of the first face box and the position coordinates of the first human face key points to obtain the position coordinates of a second face box and the position coordinates of second human face key points.

[0063] Considering that directly training the preset network model according to the face image corresponding to the position coordinates of the first face box and the position coordinates of the first human face key points, the detection accuracy of the obtained human face key point detection model depends on the accuracy of the face detection result. In this embodiment, after obtaining the position coordinates of the first face box and the position coordinates of the first human face key points, coordinate transformation is performed on the position coordinates of the first face box and the position coordinates of the first human face key points to obtain the position coordinates of a second face box and the position coordinates of second human face key points, so as to train the preset network model based on the position coordinates of the second face box and the position coordinates of the second human face key points in subsequent steps, reduce the dependence of the output result of the network model on the face detection result, and improve the accuracy of the output result of the network model.

[0064] In a specific embodiment, the step S200 specifically includes:

[0065] S210. Perform coordinate transformation on the position coordinates of the first face box according to a scaling factor to obtain the position coordinates of a second face box; wherein, the scaling factor is randomly generated within a preset range;

[0066] S220. Perform coordinate transformation on the position coordinates of the first human face key points according to the position coordinates of the second face box to obtain the position coordinates of second human face key points.

[0067] The scaling factor refers to the ratio of the position change amount before and after scaling the coordinates to the position before scaling. For example, if the coordinate point (x, y) changes to (x′, y′) after scaling, the scaling factor in the x direction The scaling factor in the y direction To reduce the dependence of the face key point detection model on the size of the face bounding box, in this embodiment, a random jitter operation is used to perform coordinate transformation on the position coordinates of the first face bounding box. First, scaling factors are randomly generated within a preset range, and then the position coordinates of the first face bounding box are transformed based on the scaling factors to obtain the position coordinates of the second face bounding box. Considering that the boundary error of the output result of the general face detector affecting the face key point detection model is between -0.15 and 0.15, in this embodiment, the scaling factors are generated within the range of -0.15 to 0.15.

[0068] In the foregoing steps, it is mentioned that the position coordinates of the first face bounding box include the upper left vertex coordinates and the lower right vertex coordinates of the first face bounding box, that is, the position coordinates of the first face bounding box include two abscissa values and two ordinate values. Correspondingly, in this embodiment, four scaling factors are randomly generated within a preset range, and the upper left vertex coordinates and the lower right vertex coordinates of the first face bounding box are respectively transformed using the four scaling factors to obtain the position coordinates of the second face bounding box composed of the upper left vertex coordinates and the lower right vertex coordinates of the second face bounding box. Specifically, the coordinate transformation formula for the position coordinates of the first face bounding box is:

[0069]

[0070] where, (X min , Y min ) are the upper left vertex coordinates of the first face bounding box, (X max , Y max ) are the lower right vertex coordinates of the first face bounding box, γ1, γ2, γ3, and γ4 are the scaling factors, (X′ min , Y′ min ) are the upper left vertex coordinates of the second face bounding box, (X′ max , Y′ max ) are the lower right vertex coordinates of the second face bounding box.

[0071] Considering that after performing the random jitter operation on the first face bounding box, the position coordinates of the first face key points will also change. In this embodiment, after obtaining the position coordinates of the second face bounding box, the position coordinates of the first face key points are further transformed according to the position coordinates of the second face bounding box. Specifically, the coordinate transformation formula for the position coordinates of the second face bounding box is:

[0072]

[0073]

[0074] Among them, (X′ min , Y′ min ) is the upper left vertex coordinate of the second face frame, (X′ max , Y′ max ) is the lower right vertex coordinate of the second face frame, (x i , y i ) is the position coordinate of the key point of the first face, and (x′ i , y′ i ) is the position coordinate of the key point of the second face.

[0075] S300. Train a preset network model according to the first face image corresponding to the position coordinates of the second face frame and the position coordinates of the key points of the second face to generate a face key point detection model.

[0076] Specifically, as mentioned in the foregoing steps, the position coordinates of the second face frame are the coordinates obtained after transforming the upper left vertex coordinate and the lower right vertex coordinate corresponding to the face in the training image. The rectangular area determined by the line connecting the points in the position coordinates of the second face frame as the diagonal is the first face image. After obtaining the position coordinates of the second face frame and the position coordinates of the key points of the second face, train the preset network model according to the first face image and the position coordinates of the key points of the second face to generate a face key point detection model. Training the preset network model through the first face image corresponding to the position coordinates of the second face frame reduces the dependence of the face key point detection model on the position of the face frame and improves the detection accuracy of the face key point detection model. The preset network model can be constructed by using the bottleneck unit, the max pooling layer, and the fully connected layer in the existing MobileNet V2 network model.

[0077] In a specific embodiment, step S300 specifically includes:

[0078] S310. Input the first face image corresponding to the position coordinates of the second face frame into the preset network model to generate the position coordinates of the third face key points corresponding to the first face image;

[0079] S320. Correct the model parameters of the preset network model according to the position coordinates of the key points of the second face and the position coordinates of the third face key points, and continue to execute the step of generating the position coordinates of the third face key points according to the first face image until the preset network model meets the preset conditions to generate a face key point detection model.

[0080] Specifically, when training a preset network model, the first face image corresponding to the second face frame position coordinates is input into the preset network model to generate the third face key point position coordinates corresponding to the first face image. The third face key point position coordinates are the key point position coordinates corresponding to the first face image output by the preset network model, while the actual key point position coordinates corresponding to the first face image are the second face key point position coordinates. The model parameters of the preset network model are corrected according to the second face key point position coordinates and the third face key point position coordinates, and the step of generating the third face key point position coordinates corresponding to the first face image according to the first face image corresponding to the second face position coordinates is continued until the preset network model meets the preset conditions, so as to generate a face key point detection model.

[0081] When determining whether the preset network model meets the preset conditions, the loss function is used to calculate the loss value between the second face key point position coordinates and the third face key point position coordinates. Generally, the smaller the loss value, the better the performance of the network model. After obtaining the loss value, it is judged whether the loss value is less than the preset threshold; if so, it indicates that the preset network model meets the preset conditions; if not, it means that the preset network model does not meet the preset conditions, and the model parameters of the preset network model are corrected according to the preset parameter learning rate, and the step of generating the third face key point position coordinates corresponding to the first face image according to the first face image corresponding to the second face position coordinates is continued until the loss value is less than the preset threshold. The obtained network model is the face key point detection model. Among them, the loss function can be selected according to actual needs. In a specific embodiment, the existing wing loss function is used to regress the second face key point position coordinates and the third face key point position coordinates.

[0082] In a specific embodiment, before step S200, it further includes:

[0083] M110. Intercept the training images in the training sample set according to the second face frame position coordinates to obtain the second face image corresponding to the second face position coordinates;

[0084] M120. Perform a preprocessing operation on the second face image to obtain the first face image corresponding to the second face position coordinates.

[0085] Since the input of a general face key point detection model is usually a face image, in this embodiment, after obtaining the position coordinates of the second face frame, the training images in the training sample set are intercepted according to the position coordinates of the second face frame to obtain the second face image corresponding to the position coordinates of the second face. Then, the input dimension of the preset network model is obtained, that is, the resolution of the image allowed to be input by the preset network model. According to the input dimension of the preset network model and the resolution of the second face image, the second face image is scaled. For example, if the resolution of the face image cropped from the first training image is 256*200 pixels, the resolution of the face image cropped from the second training image is 224*180 pixels, and the input dimension of the preset network model is 128*128 pixels, then both the first training image and the face image cropped from the first training image are scaled to 128*128 pixels.

[0086] To reduce the computational complexity of the model, in this embodiment, further normalization operation is performed on the second face image. The gray value of each pixel point in the second face image is divided by 255, so that the gray value of each pixel point in the second face image is between 0 and 1. The image intercepted according to the position coordinates of the second face frame, scaled and normalized is used as the first face image corresponding to the position coordinates of the second face and input into the preset network model for training the network model.

[0087] In a specific embodiment, before the step M110, the following steps are further included:

[0088] M100. Image enhancement is performed on the training images in the training sample set, and the training images after image enhancement are used as the training images in the training sample set.

[0089] To further improve the detection accuracy of the face key point detection model, in this embodiment, before obtaining the first face image, image enhancement is performed on the training images in the training sample set, such as image color transformation, image rotation, image flipping, and Gaussian blur operations. Image color transformation means randomly changing the contrast, brightness, saturation, etc. of the image within a preset rule. For example, randomly increasing or decreasing the brightness, saturation, etc. of a certain picture, so as to increase the richness of the training images in terms of color. Image rotation means randomly rotating the image within a preset angle range. For example, rotating within -30° to 30°, so as to generate training images at different angles, which are used to construct samples at various angles of the face in the simulated real scene for training. Image flipping is to flip the image horizontally. These data augmentation transformations can improve the detection accuracy of the face key point detection model.

[0090] Embodiment 2

[0091] Based on the above method for generating a face key point detection model, this embodiment also provides a face key point detection method, as Figure 2 shown, the face key point detection method includes:

[0092] R100. Obtain a target face image corresponding to the target image; wherein the target image contains a target face.

[0093] Specifically, the target image is an image for which face key point detection is required, and it can be obtained through existing devices with a photographing function, such as cameras, mobile phones, tablet computers, etc. Since the target image not only contains a face image but also contains a background image irrelevant to the face, in this embodiment, after obtaining the target image, a target face image corresponding to the target image is further obtained. The target face image is the image containing the face remaining after removing the background image from the target image.

[0094] In a specific implementation manner, the step R100 specifically includes:

[0095] R110. Obtain the position coordinates of the third face frame corresponding to the target image;

[0096] R120. Crop the target image according to the position coordinates of the third face frame to obtain the third face image corresponding to the target image;

[0097] R130. Perform scale scaling and normalization operations on the third face image to obtain the target face image corresponding to the target image.

[0098] Specifically, the position coordinates of the third face frame are the upper left vertex coordinates and the lower right vertex coordinates corresponding to the face in the target image. After obtaining the position coordinates of the third face frame corresponding to the target image, the target image is cropped according to the position coordinates of the third face frame to obtain the third face image corresponding to the target image. The scale of the third face image is scaled, and the scale of the third face image is adjusted to the input image scale of the face key point detection model, and the gray levels of the pixels in the third face image are normalized to obtain the target face image corresponding to the target image.

[0099] R200. Input the target face image into the face key point detection model to obtain the position coordinates of the fourth face key points corresponding to the target image.

[0100] Specifically, after obtaining the target face image, the target face image is input into a trained face key point detection model to obtain the position coordinates of the fourth face key points corresponding to the target image. Since the face key point detection model is trained based on the face image after coordinate transformation and the position coordinates of the face key points, the position coordinates of the fourth face key points output by the face key point detection model are not affected by the detection result of the target face image, correcting the instability and inaccuracy of face key point detection in the case of a deviated face frame, and improving the accuracy of the output result of the face key point detection model.

[0101] R300. Perform a coordinate transformation on the position coordinates of the fourth face key points to obtain the position coordinates of the target face key points corresponding to the target image.

[0102] Since the face key point detection model is trained based on the position coordinates of the second face key points after the coordinate transformation of the position coordinates of the first face key points, the position coordinates of the fourth face key points output by the face key point detection model also need to be coordinate-transformed to obtain the true position coordinates of the face key points of the target image, that is, the position coordinates of the target face key points.

[0103] In a specific embodiment, the step R300 specifically includes:

[0104] R310. Perform a coordinate transformation on the position coordinates of the fourth face key points according to the position coordinates of the third face to obtain the position coordinates of the target face key points corresponding to the target image.

[0105] Specifically, after obtaining the position coordinates of the fourth face key points in this embodiment, a coordinate transformation is performed on the position coordinates of the fourth face key points according to the position coordinates of the third face frame to obtain the position coordinates of the target key points corresponding to the target image. Among them, the position coordinates of the third face frame include the upper left vertex coordinates of the third face frame and the lower right vertex coordinates of the third face frame, and the coordinate transformation formula of the position coordinates of the fourth face key points is:

[0106]

[0107]

[0108] where, (x i p , y i p ) are the position coordinates of the fourth face key points, (X min , Y min ) are the upper left vertex coordinates of the third face frame, (X max , Y max ) are the lower right vertex coordinates of the third face frame, ( y ip′ ) are the position coordinates of the target face key points.

[0109] Embodiment III

[0110] Based on the above embodiments, the present invention further provides a device for generating a face key point detection model, and its functional schematic diagram is as Figure 3 shown. The device includes a coordinate acquisition module 110, a coordinate transformation module 120, and a model generation module 130;

[0111] The coordinate acquisition module 110 is used to acquire the position coordinates of the first face frame and the position coordinates of the first face key points corresponding to the training images in the training sample set; wherein, the training images contain faces; specifically as described in the embodiments of the method for generating a face key point detection model.

[0112] The coordinate transformation module 120 is used to perform coordinate transformation on the position coordinates of the first face frame and the position coordinates of the first face key points to obtain the position coordinates of the second face frame and the position coordinates of the second face key points; specifically as described in the embodiments of the method for generating a face key point detection model.

[0113] The model generation module 130 is used to train a preset network model according to the first face image corresponding to the position coordinates of the second face frame and the position coordinates of the second face key points to generate a face key point detection model; specifically as described in the embodiments of the method for generating a face key point detection model.

[0114] Embodiment IV

[0115] Based on the above embodiments, the present invention further provides a terminal, and its principle block diagram can be as Figure 4 shown. The terminal includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for generating a face key point detection model and a method for detecting face key points. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor of the terminal is pre-set inside the device to detect the current operating temperature of the internal device.

[0116] Those skilled in the art can understand, Figure 4The block diagram of the principle shown only shows the block diagram of part of the structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0117] In one embodiment, a terminal is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps can be implemented at least:

[0118] Obtain the first face frame position coordinates and the first face key point position coordinates corresponding to the training image in the training sample set; wherein, the training image contains a human face;

[0119] Perform coordinate transformation on the first face frame position coordinates and the first face key point position coordinates to obtain the second face frame position coordinates and the second face key point position coordinates;

[0120] Train a preset network model according to the first face image corresponding to the second face frame position coordinates and the second face key point position coordinates to generate a face key point detection model.

[0121] In one of the embodiments, when the processor executes the computer program, it can also implement: perform coordinate transformation on the first face frame position coordinates according to a pre-generated scaling factor to obtain the second face frame position coordinates; perform coordinate transformation on the first face key point position coordinates according to the second face frame position coordinates to obtain the second face key point position coordinates.

[0122] In one of the embodiments, when the processor executes the computer program, it can also implement: randomly generate the scaling factor within a preset range.

[0123] In one of the embodiments, when the processor executes the computer program, it can also implement: input the first face image corresponding to the second face frame position coordinates into a preset network model to generate the third face key point position coordinates corresponding to the first face image; correct the model parameters of the preset network model according to the second face key point position coordinates and the third face key point position coordinates, and continue to execute the step of generating the third face key point position coordinates according to the first face image until the preset network model meets the preset conditions to generate a face key point detection model.

[0124] In one embodiment, when the processor executes the computer program, it can also implement: obtaining a loss value according to the second face key point position coordinates and the third face key point position coordinates; if the loss value is greater than or equal to a preset threshold, correcting the model parameters of the preset network model according to a preset parameter learning rate, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the loss value is less than the preset threshold.

[0125] In one embodiment, when the processor executes the computer program, it can also implement: intercepting the training images in the training sample set according to the second face frame position coordinates to obtain a second face image corresponding to the second face position coordinates; performing a preprocessing operation on the second face image to obtain a first face image corresponding to the second face position coordinates.

[0126] In one embodiment, when the processor executes the computer program, it can also implement: performing scale scaling on the second face image, and performing a normalization operation on the second face image after scale scaling to obtain a first face image corresponding to the second face position coordinates.

[0127] In one embodiment, when the processor executes the computer program, it can also implement: obtaining the input dimension of the preset network model; performing scale scaling on the second face image according to the input dimension of the preset network model and the resolution of the second face image.

[0128] In one embodiment, when the processor executes the computer program, it can also implement: performing image enhancement on the training images in the training sample set, and using the training images after image enhancement as the training images in the training sample set.

[0129] In one embodiment, when the processor executes the computer program, it can also implement: obtaining a target face image corresponding to the target image; where the target image contains a target face; inputting the target face image into the face key point detection model to obtain fourth face key point position coordinates corresponding to the target image; performing coordinate transformation on the fourth face key point position coordinates to obtain target face key point position coordinates corresponding to the target image.

[0130] In one embodiment, when the processor executes the computer program, it can also implement: obtaining third face frame position coordinates corresponding to the target image; intercepting the target image according to the third face frame position coordinates to obtain a third face image corresponding to the target image; performing scale scaling and normalization operations on the third face image to obtain a target face image corresponding to the target image.

[0131] In one of the embodiments, when the processor executes the computer program, it can also implement: performing coordinate transformation on the fourth face key point position coordinates according to the third face frame position coordinates to obtain the target face key point position coordinates corresponding to the target image.

[0132] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0133] In summary, the present invention discloses a method for generating a face key point detection model, a detection method, and an electronic device. The generation method performs coordinate transformation on the first face frame position coordinates and the first face key point position coordinates corresponding to the training images in the training sample set to obtain the second face frame position coordinates and the second face key point position coordinates; and trains a preset network model according to the first face image corresponding to the second face frame position coordinates and the second face key point position coordinates to generate a face key point detection model. The present invention trains the face key point detection model with the first face image corresponding to the second face frame position coordinates and the second face key point position coordinates after coordinate transformation, reduces the dependence of the face key point detection model on the face detection result, and improves the robustness of the face key point detection model for face key point localization without additionally increasing training data.

[0134] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for generating a face key point detection model, characterized in that Including: Obtaining the first face box position coordinates and the first face key point position coordinates corresponding to the training images in the training sample set; wherein, the training images contain human faces; Performing coordinate transformation on the first face box position coordinates and the first face key point position coordinates to obtain second face box position coordinates and second face key point position coordinates; Training a preset network model according to the first face image corresponding to the second face box position coordinates and the second face key point position coordinates to generate a face key point detection model; The step of performing coordinate transformation on the first face box position coordinates and the first face key point position coordinates to obtain second face box position coordinates and second face key point position coordinates includes: Performing coordinate transformation on the first face box position coordinates according to a scaling factor to obtain second face box position coordinates; Performing coordinate transformation on the first face key point position coordinates according to the second face box position coordinates to obtain second face key point position coordinates; The method for generating the scaling factor includes: Randomly generating the scaling factor within a preset range.

2. The method for generating a facial key point detection model according to claim 1, wherein The step of training a preset network model according to the first face image corresponding to the second face box position coordinates and the second face key point position coordinates to generate a face key point detection model includes: Inputting the first face image corresponding to the second face box position coordinates into the preset network model to generate third face key point position coordinates corresponding to the first face image; Correcting the model parameters of the preset network model according to the second face key point position coordinates and the third face key point position coordinates, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the preset network model meets the preset conditions to generate a face key point detection model.

3. The method for generating a face key point detection model according to claim 2, wherein, The step of correcting the model parameters of the preset network model according to the second face key point position coordinates and the third face key point position coordinates, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the preset network model meets the preset conditions includes: Obtaining a loss value according to the second face key point position coordinates and the third face key point position coordinates. If the loss value is greater than or equal to a preset threshold, correcting the model parameters of the preset network model according to a preset parameter learning rate, and continuing to execute the step of generating the third face key point position coordinates according to the first face image until the loss value is less than the preset threshold.

4. The method for generating a face key point detection model according to claim 1, wherein After the step of performing coordinate transformation on the first face box position coordinates and the first face key point position coordinates to obtain second face box position coordinates and second face key point position coordinates includes: Cropping the training images in the training sample set according to the second face box position coordinates to obtain second face images corresponding to the second face box position coordinates; Performing a preprocessing operation on the second face images to obtain first face images corresponding to the second face box position coordinates.

5. The method for generating a face key point detection model according to claim 4, wherein The step of preprocessing the second face image to obtain the first face image corresponding to the position coordinates of the second face frame includes: Performing scale scaling on the second face image, and performing a normalization operation on the second face image after scale scaling to obtain the first face image corresponding to the position coordinates of the second face frame.

6. The method for generating a face key point detection model according to claim 5, wherein The step of performing scale scaling on the second face image includes: Obtaining the input dimension of the preset network model; Performing scale scaling on the second face image according to the input dimension of the preset network model and the resolution of the second face image.

7. The method for generating a face key point detection model according to claim 4, wherein Before the step of intercepting the training image in the training sample set according to the position coordinates of the second face frame to obtain the second face image corresponding to the position coordinates of the second face frame, it includes: Performing image enhancement on the training images in the training sample set, and using the training images after image enhancement as the training images in the training sample set.

8. A method for facial key point detection, characterized in that, Applied to the face key point detection model generated by the method for generating a face key point detection model according to any one of claims 1-7, it includes: Obtaining a target face image corresponding to the target image; wherein the target image contains a target face; Inputting the target face image into the face key point detection model to obtain the position coordinates of the fourth face key points corresponding to the target image; Performing coordinate transformation on the position coordinates of the fourth face key points to obtain the position coordinates of the target face key points corresponding to the target image.

9. The face key point detection method according to claim 8, wherein The step of obtaining the target face image corresponding to the target image includes: Obtaining the position coordinates of the third face frame corresponding to the target image; Intercepting the target image according to the position coordinates of the third face frame to obtain the third face image corresponding to the target image; Performing scale scaling and normalization operations on the third face image to obtain the target face image corresponding to the target image.

10. The face key point detection method according to claim 9, wherein, The step of performing coordinate transformation on the position coordinates of the fourth face key points to obtain the position coordinates of the target face key points corresponding to the target image includes: Performing coordinate transformation on the position coordinates of the fourth face key points according to the position coordinates of the third face frame to obtain the position coordinates of the target face key points corresponding to the target image.

11. A device for generating a facial key point detection model, characterized in that, It includes: A coordinate acquisition module, configured to acquire the position coordinates of the first face frame and the position coordinates of the first face key points corresponding to the training images in the training sample set; wherein, the training images contain faces; A coordinate transformation module, configured to perform coordinate transformation on the position coordinates of the first face frame and the position coordinates of the first face key points to obtain the position coordinates of the second face frame and the position coordinates of the second face key points; A model generation module, configured to train a preset network model according to the first face image corresponding to the position coordinates of the second face frame and the position coordinates of the second face key points to generate a face key point detection model; The coordinate transformation module includes: Performing coordinate transformation on the position coordinates of the first face frame according to the scaling factor to obtain the position coordinates of the second face frame; Performing coordinate transformation on the first human face key point position coordinates according to the second human face frame position coordinates to obtain the second human face key point position coordinates; The method for generating the scaling factor includes: Randomly generating the scaling factor within a preset range.

12. A terminal, characterized in that, Including: A processor and a storage medium communicatively connected to the processor, the storage medium being adapted to store a plurality of instructions; the processor being adapted to call the instructions in the storage medium to execute the steps in the method for generating a human face key point detection model according to any one of claims 1-7 above, or the steps in the human face key point detection method according to any one of claims 8-10 above.

13. A storage medium having a plurality of instructions stored thereon, characterized in that, The instructions are adapted to be loaded and executed by the processor to execute the steps in the method for generating a human face key point detection model according to any one of claims 1-7 above, or the steps in the human face key point detection method according to any one of claims 8-10 above.

Citation Information

Patent Citations

  • Face key point tracking system and method applied to mobile device

    CN106909888A

  • Detection method of key points of human face

    CN107704847A