Key point recognition method and device based on human face 3D model and facial semantic features
By combining multi-view 2D key point recognition and 3D semantic feature calibration methods, the practicality and cost problems of face 3D model key point detection in the prior art are solved, and high-precision and low-cost 3D face key point recognition are achieved.
Patent Information
- Application Number
- CN202411764845.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-04
AI Technical Summary
The key point detection algorithm of the existing face 3D model is limited in practicality in scenarios requiring a large number of face key point scenarios, and machine learning methods require a large amount of training data, which has high development costs and long cycles.
The key point recognition method based on the face 3D model and facial semantic features is adopted, and images are collected through multiple perspectives, 2D key points are extracted using MVCNN, and converted into a 3D world coordinate system, weighted fusion and facial semantic feature calibration are performed to obtain the final 3D key points.
It realizes the relatively accurate 3D face key point recognition results at low development costs, reduces the possible random errors of the CNN model, and improves the accuracy and verifiability of the recognition.
Smart Images

Figure CN119229509B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of face intelligence technology, and more particularly to a key point recognition method and device based on a face 3D model and facial semantic features, an electronic device, and a computer-readable storage medium. Background Art
[0002] Key point detection of 3D face models is mainly used for:
[0003] Improve recognition accuracy: 3D key point detection provides x, y, and z coordinate information. Compared with 2D key points, it adds depth information, which helps to more accurately recognize faces in complex scenarios (such as large-angle postures and face occlusion);
[0004] Enhanced application effects: In terms of face posture estimation, 3D object wear, etc., 3D key point detection has obvious advantages and can provide more accurate face spatial position information, thus improving application effects;
[0005] Multi-field applications: Facial information obtained through 3D key point detection is of great value in fields such as human-computer interaction, entertainment, security monitoring and medical applications, and can provide richer and more accurate facial shape and posture information;
[0006] Key point detection of 3D facial models plays an important role in improving recognition accuracy, enhancing application effects, and expanding multi-field applications. It is one of the current research hotspots in the field of key point detection of facial models.
[0007] The existing key point detection algorithms for face 3D models are mainly the following:
[0008] 1.3D Morphable Model (3DMM): This is a classic method that uses statistical models to represent 3D facial shape and texture. 3D facial key point detection is achieved by optimizing model parameters to fit the facial feature points in the input image.
[0009] 2. Deep Learning Methods:
[0010] 3D Face Reconstruction Networks: Some deep learning networks, such as FaceMesh or Dlib's 3D face model, use convolutional neural networks (CNNs) to predict the 3D shape and key points of the face. These networks are usually trained on a large amount of annotated 3D face data to obtain high-precision 3D reconstruction.
[0011] 3. Multi-view Convolutional Neural Networks (MVCNN): This method collects images from multiple perspectives and uses CNN models for 3D facial reconstruction and key point detection.
[0012] 4.3D Face Alignment Network (3DFAN): This is a deep learning network designed specifically for 3D face alignment and key point detection. It predicts 3D facial feature points by performing layer-by-layer regression on the image.
[0013] 5. Face Alignment Network (FAN): This method was originally used for 2D face alignment, but it can also be extended to 3D facial key point detection by detecting key points in 2D facial images and then mapping them to 3D space through geometric transformation.
[0014] 6.3D Face Shape Models: Using 3D facial shape models, feature points in 2D images are mapped to 3D space. These models can be trained based on 3D scan data to obtain accurate facial shapes and key point locations.
[0015] However, among these traditional methods, most key point detection algorithms / methods can only locate a few key points with obvious features, so they cannot be applied to scenarios that require a large number of facial key points, and their practicality is limited; in addition, machine learning methods often require a large amount of training data, with high development costs and long cycles. Summary of the invention
[0016] In order to solve the technical problems existing in the prior art, the present invention provides the following technical solutions:
[0017] On the one hand, a key point recognition method based on a 3D face model and facial semantic features is provided, the method is implemented by an electronic device, and the method includes:
[0018] S1. Collect multi-view facial images and use the MVCNN method to extract 2D key points in the facial images at each view.
[0019] S2, converting the 2D key points at each viewing angle into a 3D world coordinate system, obtaining the 3D key points at each viewing angle and mapping them onto a preset 3D face model;
[0020] S3, according to the key point sequence number, the same 3D key points under different viewing angles are grouped together, and the group of 3D key points is weightedly fused to obtain preliminary face key points under each viewing angle;
[0021] S4. Calculate the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, use the facial semantic features to calibrate the positions of the preliminary facial key points, and obtain the final 3D key points Pw at each viewing angle.
[0022] Furthermore, in step S2, the 2D key points at each viewing angle are converted to a 3D world coordinate system according to the following formula:
[0023] ,
[0024] in:
[0025] Pw represents the 3D world coordinates. Pw is a point on the ray from the real 3D key point to the center point of the camera. It is necessary to introduce a ray detection algorithm to find the intersection point on the surface of the 3D face model. The intersection point is the real 3D key point.
[0026] t is the external parameter of the camera;
[0027] K is the intrinsic parameter of the camera;
[0028] u,v are the coordinates of the 2D keypoints in the image;
[0029] zc is artificially given depth data.
[0030] Furthermore, in step S3, the group of 3D key points are weightedly fused to obtain preliminary face key points at each viewing angle, including:
[0031] Obtaining groups of 3D key points under different viewing angles, and assigning weights to the 3D key points in each group according to the positions of the 3D key points;
[0032] The weighted fusion of each group of 3D key points is performed according to the following formula:
[0033] ,
[0034] in:
[0035] Pwi represents the 3D key points obtained from each perspective;
[0036] ai is the weight assigned to the perspective, and the weight is assigned according to the position information of the key point itself.
[0037] Further, S4, calculating the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, using the facial semantic features to calibrate the positions of the preliminary facial key points, and obtaining the final 3D key points Pw at each viewing angle, including:
[0038] The 3D model curvature K of the area near the preliminary facial key points on the 3D face model at each viewing angle is calculated according to the following formula:
[0039] ,
[0040] in:
[0041] df(X) represents the curvature of the preliminary facial key points on the 3D face model;
[0042] dN(X) represents the curvature of a point located near the preliminary facial key point on the 3D face model;
[0043] |df(X)| represents the absolute value of curvature;
[0044] k n (X) represents the 3D model curvature K of the vicinity of the preliminary facial key point on the 3D face model;
[0045] For several points near the preliminary facial key point, the 3D model curvature K of each point is calculated according to the above steps to obtain several 3D model curvatures K, and the position of the point with the largest 3D model curvature K is found, and this point is taken as the final 3D key point Pw.
[0046] Further, S4, calculating the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, using the facial semantic features to calibrate the positions of the preliminary facial key points, and obtaining the final 3D key points Pw at each viewing angle, including:
[0047] Acquire nearby point data of the preliminary facial key points on the 3D face model at each viewing angle;
[0048] Performing denoising and filtering processing on the nearby point data;
[0049] Using a curve approximate fitting method, fitting calculation is performed on the preprocessed nearby point data to generate a corresponding nearby point fitting curve;
[0050] Find a point on the nearby point fitting curve that has similar features to the preliminary face key point, and take the point as the final 3D key point Pw.
[0051] On the other hand, a key point recognition device based on a 3D face model and facial semantic features is provided, and the key point recognition device based on a 3D face model and facial semantic features is used to implement the key point recognition method based on a 3D face model and facial semantic features. The device includes:
[0052] An image acquisition unit, used for acquiring facial images from multiple perspectives;
[0053] A 2D key point recognition unit, used to extract 2D key points in face images at different viewing angles using the MVCNN method;
[0054] A three-dimensional conversion unit, used to convert the 2D key points at each viewing angle into a 3D world coordinate system, obtain the 3D key points at each viewing angle and map them onto a preset 3D face model;
[0055] The 3D fusion unit is used to group the same 3D key points under different viewing angles according to the key point sequence number, and perform weighted fusion on the group of 3D key points to obtain preliminary face key points under each viewing angle;
[0056] The 3D key point calibration unit is used to calculate the facial semantic features of the preliminary face key points on the face 3D model at each viewing angle, calibrate the positions of the preliminary face key points using the facial semantic features, and obtain the final 3D key points Pw at each viewing angle.
[0057] On the other hand, an electronic device is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the key point recognition method based on a 3D face model and facial semantic features as described above is implemented.
[0058] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned key point recognition method based on a 3D face model and facial semantic features.
[0059] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0060] The present invention combines 2D multi-view key point recognition and 3D semantic feature calibration, and can obtain relatively accurate 3D face key point recognition results at a lower development cost.
[0061] In terms of key point recognition, compared with general single-view key point detection, the multi-view detection of the present invention can eliminate the interference of face posture and light and shadow on the detection results, and in the process of merging key points detected from multiple viewpoints, it also further reduces the random errors that may exist in the CNN model.
[0062] In terms of calibration of 3D key points, compared with the traditional 3DMM statistical method, this technology uses a variety of semantic information such as curvature, gradient, local structural features, etc., and applies specific methods to specific key points for calibration, which can further improve accuracy and verifiability. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0064] Figure 1 This is a flow chart of a key point recognition method based on a 3D face model and facial semantic features provided by an embodiment of the present invention;
[0065] Figure 2 It is a schematic diagram of a process of preliminary identification of 3D key points provided by an embodiment of the present invention;
[0066] Figure 3 It is a schematic diagram of a process of 3D key point calibration and identification provided by an embodiment of the present invention;
[0067] Figure 4 It is a block diagram of a key point recognition device based on a 3D face model and facial semantic features provided by an embodiment of the present invention;
[0068] Figure 5 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0069] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0070] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0071] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0072] In the embodiments of the present invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are consistent.
[0073] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0074] The embodiment of the present invention provides a key point recognition method based on a 3D face model and facial semantic features, which can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of the key point recognition method based on the 3D face model and facial semantic features shown in the figure may include the following steps:
[0075] S1. Collect multi-view facial images and use the MVCNN method to extract 2D key points in the facial images at each view.
[0076] S2, converting the 2D key points at each viewing angle into a 3D world coordinate system, obtaining the 3D key points at each viewing angle and mapping them onto a preset 3D face model;
[0077] S3, according to the key point sequence number, the same 3D key points under different viewing angles are grouped together, and the group of 3D key points is weightedly fused to obtain preliminary face key points under each viewing angle;
[0078] S4. Calculate the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, use the facial semantic features to calibrate the positions of the preliminary facial key points, and obtain the final 3D key points Pw at each viewing angle.
[0079] The present invention combines the MVCNN method and the FAN method in the deep learning method (the MVCNN method and the FAN method are described in the technical background), maps the 2D key points collected from multiple perspectives to the 3D face model, and then calibrates the key points using the 3D model features in multiple dimensions to achieve more accurate facial key point recognition results.
[0080] The present invention is mainly divided into 2D key point recognition, mapping of 2D key points to a 3D face model, and calibrating 3D face key points on the 3D face model using facial semantic features.
[0081] The above steps will be described in detail below with reference to the accompanying drawings.
[0082] like Figure 2 As shown, the 2D key points collected from multiple perspectives are mapped to a 3D face model.
[0083] Visual cameras can be used to collect facial images from different perspectives.
[0084] First, key points are identified on facial images from multiple angles (using the CNN model), and then multiple groups of key points are projected onto the same 3D model. The same key points obtained from different perspectives are grouped together, and then the grouped key points are weightedly fused according to the perspective adjacency relationship to preliminarily obtain 3D key points.
[0085] The projection of 2D key points to 3D key points applies the ray detection algorithm, as follows:
[0086] Given a ray P(t)=O+tD, where O is the starting point of the ray, D is the direction of the ray, and t is a non-negative real number.
[0087] Given a plane N⋅(P−Q)=0, it can be expanded to N⋅P=N⋅Q, where N is the normal vector of the plane, Q is the coordinate vector of a known point on the plane, and P is the coordinate vector of any point on the plane.
[0088] Parametric equation for the intersection of a ray and a plane: N⋅(O+tD)=N⋅Q; expand and solve for t, substitute t back into the parametric equation, and you can find the point of intersection.
[0089] First, convert the 2D key points of each perspective into the 3D world coordinate system according to the following formula 1:
[0090] (Formula 1),
[0091] in:
[0092] Pw represents the 3D world coordinates, R, t are the external parameters of the camera, K is the internal parameters of the camera, u, v are the coordinates of the 2D key points in the image, and zc is the artificial depth data, set to 0.3 meters (this is the average distance from the camera to the face). At this time, the Pw obtained is just a point on the ray from the real 3D key point to the center point of the camera. It is necessary to introduce a ray detection algorithm to obtain the intersection point on the model surface. This intersection point is the real 3D key point.
[0093] After obtaining the 3D key points at each perspective, the same key points at different perspectives are grouped together according to the sequence number of the key points. Then, the sum of each group of key points is calculated according to Formula 2 to obtain the preliminary 3D key points of the face:
[0094] , (Formula 2)
[0095] in:
[0096] Pwi represents the 3D key points obtained from each perspective;
[0097] ai is the weight assigned to the perspective, and the weight is assigned according to the position information of the key point itself. For example, the weight of the eyebrow point is large in a front view at a small angle, and small in a side view, and may even be invisible at extreme angles.
[0098] According to the above method, the 2D key points are converted into 3D, and the corresponding 3D key points are mapped on the 3D face model, and the recognition and detection of the face key points are preliminarily obtained.
[0099] Although the texture information in RGB images is relatively accurate in locating key points related to edges and corners, the large number of facial key points include many key points with unclear texture features, such as cheekbones and nose bridges. Therefore, there is a certain position error in the detection results. These points need to be accurately calibrated in combination with the bone morphology and muscle direction of the face.
[0100] In order to improve its accuracy, the present invention utilizes 3D model features in multiple dimensions to calibrate key points to achieve more accurate facial key point recognition results.
[0101] This method uses multi-view 2D images to recognize facial key points, and gives facial semantic features (such as curvature, gradient, local structural features and other semantic information) to the 3D face model. These semantic information are used to further calibrate the key points, and high-precision 3D face key points are obtained at a relatively low cost.
[0102] Features that can be obtained from the 3D model include:
[0103] 1. Curvature features, reflecting the concavity and convexity of the face and its changes;
[0104] 2. Gradient changes, including gradient changes in the three spatial directions of X, Y, and Z and gradient changes in curvature;
[0105] 3. Local 3D edges and 3D corners; prominent corners and edges are generally generated by facial features, such as eye sockets and nasal septum.
[0106] Therefore, the present invention can utilize the above-mentioned various features to calibrate 3D key points.
[0107] like Figure 3 As shown, the 3D surface curvature calculation method can be used to calibrate the extreme values of 3D key points; the 3D model gradient calculation can also be used to calibrate the edge points of 3D key points; the local key structure fitting method of the face can also be used to calibrate the local structure related points. After calibration by the above methods, the 2D key points are reprojected and finally projected on the 3D model to obtain the final 3D key points Pw with very accurate positions.
[0108] The following will describe several of the calibration methods with examples.
[0109] Further, S4, calculating the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, using the facial semantic features to calibrate the positions of the preliminary facial key points, and obtaining the final 3D key points Pw at each viewing angle, including:
[0110] The 3D model curvature K of the area near the preliminary facial key points on the 3D face model at each viewing angle is calculated according to the following formula:
[0111] ,
[0112] in:
[0113] df(X) represents the curvature of the preliminary facial key points on the 3D face model;
[0114] dN(X) represents the curvature of a point located near the preliminary facial key point on the 3D face model;
[0115] |df(X)| represents the absolute value of curvature;
[0116] k n (X) represents the 3D model curvature K of the vicinity of the preliminary facial key point on the 3D face model;
[0117] For several points near the preliminary facial key point, the 3D model curvature K of each point is calculated according to the above steps to obtain several 3D model curvatures K, and the position of the point with the largest 3D model curvature K is found, and this point is taken as the final 3D key point Pw.
[0118] Although the 3D face model obtained by scanning with a high-precision structured light imaging system cannot obtain the bone structure and muscle composition of the head like a CT image, by calculating the gradient and curvature changes on the surface of the 3D model, combined with professional medical skull models and anatomical images, the 3D key points can be further calibrated. For example, the zygomatic area of the face is generally convex, and its principal curvature is greater than 1; while the eye socket area is generally concave, and the principal curvature is less than 1, which is a typical saddle-shaped face. By finding the maximum value of the principal curvature of the zygomatic area, the accurate zygomatic point can be obtained, and by calculating the lower edge convex point of the eye socket area, the infraorbital point can be located.
[0119] By using the above calculation method of the curvature K of the 3D model, the curvature change characteristics of the area near the key points of the face on the 3D face model can be calculated to reflect the local concavity and convexity of the face at that point and its changes. By ranking the calculation results of several nearby points, the point with the largest change is obtained, which is used as the feature point and the preliminary key points of the face are calibrated (by replacing the coordinate position), and the point is used as the final key point with the characteristics of the local area. This point reflects the characteristics of the local area where the preliminary point is located, and can accurately represent the 3D face key points in this area.
[0120] Further, S4, calculating the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, using the facial semantic features to calibrate the positions of the preliminary facial key points, and obtaining the final 3D key points Pw at each viewing angle, including:
[0121] Acquire nearby point data of the preliminary facial key points on the 3D face model at each viewing angle;
[0122] Performing denoising and filtering processing on the nearby point data;
[0123] Using a curve approximate fitting method, fitting calculation is performed on the preprocessed nearby point data to generate a corresponding nearby point fitting curve;
[0124] Find a point on the nearby point fitting curve that has similar features to the preliminary face key point, and take the point as the final 3D key point Pw.
[0125] Here, we use curve approximation fitting to find edge points similar to the preliminary facial key points (such as eye sockets, etc.), and calibrate the preliminary facial key points by finding edge points with similar features to the preliminary facial key points.
[0126] Specifically, the basic idea of finding similar edge points by curve fitting is: to find edge points similar to a certain point by curve fitting, first determine the type of curve to be fitted, then use known data points for fitting, and finally find edge points similar to the target point on the fitting curve. It includes the following steps:
[0127] Fitting curve selection: select a suitable fitting curve according to actual needs and data characteristics, such as polynomial curve, spline curve, etc.;
[0128] Data point acquisition: Extract the target edge point and its surrounding data points through image processing technology as input data for fitting curve;
[0129] Curve fitting: Use mathematical methods (such as the least squares method) to perform curve fitting and obtain the fitting curve equation;
[0130] Edge point positioning: Find edge points on the fitting curve that have similar features to the target point (such as grayscale value, gradient, etc. The gradient value can be calculated by the administrator using the corresponding gradient algorithm, such as Numerical method: By approximating the gradient, the commonly used methods are finite difference method and central difference method. The finite difference method uses the ratio of the difference between the function values of the neighboring points of a certain point and the step length to approximate the gradient; the central difference rule is improved on this basis, using the ratio of half the difference between the function values of two neighboring points and the step length to calculate in order to reduce the error; Analytical method: Calculate the gradient by solving the partial derivatives of the loss function with respect to the model parameters. The calculation is relatively accurate and the amount of calculation is small, which is suitable for high-dimensional space. Common analytical methods include chain rule and back propagation algorithm, the latter of which is particularly important in deep learning).
[0131] For curve approximation fitting, especially the edge point fitting of parts such as the eye socket, we can use some common curve fitting methods, such as polynomial fitting, spline curve fitting (such as B-spline or Bezier curve), or more advanced ones such as active contour models (such as Snake model) and machine learning methods (such as edge detection algorithms in convolutional neural networks).
[0132] Here are some basic steps and considerations for fitting edge points for areas like the eye socket:
[0133] Data Preprocessing:
[0134] First, it is necessary to obtain edge point data of eye sockets and other parts (preliminary facial key points). This is usually achieved through image processing technologies such as image segmentation and edge detection. The acquired edge points are denoised and filtered to improve the accuracy of fitting.
[0135] Next, select the fitting method:
[0136] Polynomial Fitting: For simpler curve shapes, you can use polynomial fitting. By adjusting the order of the polynomial, you can fit curves of different complexity;
[0137] Spline curve fitting: Spline curves (such as B-splines) have local control and can better adapt to complex curves. By adjusting the control points and the order of the spline, accurate fitting can be achieved;
[0138] Active Contour Model: Active contour models such as the Snake model can dynamically adjust the curve shape according to the energy function of the image (such as gradient, curvature, etc.), which is suitable for edge fitting of complex shapes;
[0139] Machine Learning Methods: For more complex scenarios, machine learning methods such as edge detection algorithms in convolutional neural networks (CNNs) can be used. These algorithms can learn the features in the image and automatically detect edge points.
[0140] Next, implement the fit:
[0141] According to the selected fitting method, use the corresponding algorithms and tools for fitting;
[0142] For polynomial fitting and spline curve fitting, mathematical software (such as MATLAB, Python's NumPy and SciPy libraries) can be used to implement;
[0143] For active contour models and machine learning methods, you need to use image processing libraries (such as OpenCV) and deep learning frameworks (such as TensorFlow, PyTorch);
[0144] Evaluate the fit:
[0145] Use some evaluation indicators (such as mean square error, goodness of fit, etc.) to evaluate the fitting effect;
[0146] The fitting effect can be intuitively evaluated by comparing the visual fitting results with the original data.
[0147] This method can effectively find edge points similar to the target points, providing an accurate basis for subsequent image processing and analysis.
[0148] Here, we calculate the gradient change results of facial key points on the 3D face model to search for the gradient change of facial key points in the three spatial directions of X, Y, and Z and the gradient change of curvature, and combine the gradient change characteristics to calibrate the preliminary facial key points.
[0149] In addition to the above methods, a gradient value search method can also be used, such as locating the point under the ear through a sharp change in the Z-axis gradient of the earlobe (the limit of the change can be determined by the administrator).
[0150] The present invention can be calibrated once, for example, for the recognition of zygomatic points:
[0151] 1. Given multi-view face images V1, V2, V3;
[0152] 2. Through the CNN key point detection model, obtain the 2D key points P1, P2, P3, and convert them to the 3D world coordinate system through formula 1 to obtain Pw1, Pw2, Pw3;
[0153] 3. Substitute into formula 2 to get the preliminary 3D key point Pw.
[0154] 4. Then, the curvature K of the 3D model in the area near Pw is obtained by formula 3, the position of the maximum point of curvature K is calculated, and the precise key point Pw1 is obtained, which is the final result.
[0155] It is also possible to combine the above calibration methods and perform multiple rounds of calibration. For example, first perform curvature calibration, then perform gradient calibration, and finally perform key structural feature calibration of the layout area. See above for details.
[0156] In addition to the above methods, there are other options for 3D key point recognition:
[0157] 1. Apply deep learning methods to directly mark and train key point recognition models on 3D models; mark key points on 3D models to build a model that can recognize key points. Follow the steps below:
[0158] Model preparation:
[0159] Make sure the 3D model has been correctly loaded into the corresponding 3D modeling software, such as 3dMax, etc.
[0160] Keypoint Selection:
[0161] According to the requirements, stable and distinctive points are selected on the model as key points. These points can be the edges, corners or locations where the surface changes significantly.
[0162] Key point annotation:
[0163] Enable the key point setting mode in the software to mark the selected key points. You can accurately set the position of the key points by dragging or entering coordinates.
[0164] Key Point Editing:
[0165] Make necessary edits to the marked key points, such as adjusting the position, adding or deleting key points, to ensure the accuracy and completeness of the key points.
[0166] Model Building:
[0167] Using the annotated key points, a recognition model is constructed that can recognize these key points. This usually involves combining the key points with local feature descriptors to form key point descriptors.
[0168] 2. Directly import the CT image of the face model to obtain the real skull and muscle direction information, and calibrate the key points more accurately. Import CT images to calibrate face key points:
[0169] CT image import: First, you need to import CT images containing human faces. These images can clearly show the skull structure and muscle direction, providing an accurate basis for calibrating facial key points.
[0170] Extract key information: Through image processing technology, key information such as skull shape and muscle distribution are extracted from CT images. This information is crucial for determining the location of key points on the face.
[0171] Calibrate facial key points: Calibrate the key points on the face model based on the extracted skull and muscle information. By adjusting the position of the key points, the face model is closer to the shape of a real face.
[0172] Optimize Model: After calibrating key points, further optimize the face model, such as adjusting facial texture, expression, etc., to obtain a more realistic effect.
[0173] This method leverages the accuracy of CT images to significantly improve the calibration accuracy of facial models, providing strong support for 3D face recognition, digital human production and other fields.
[0174] 3. Parameterize the 3D model and convert it into a 2D depth image. Then, combine the RGB information to train the model to learn the key points of model detection. This can be understood in conjunction with "1".
[0175] Therefore, in terms of key point recognition, compared with general single-view key point detection, multi-view detection can eliminate the interference of facial posture and light and shadow on the detection results, and in the process of merging key points detected from multiple viewpoints, it further reduces the random errors that may exist in the CNN model.
[0176] In terms of calibration of 3D key points, compared with the traditional 3DMM statistical method, this technology uses a variety of semantic information such as curvature, gradient, local structural features, etc., and applies specific methods to specific key points for calibration, which can further improve accuracy and verifiability.
[0177] Figure 4 The invention is a block diagram of a key point recognition device based on a 3D face model and facial semantic features according to an exemplary embodiment, and the device is used to implement a key point recognition method based on a 3D face model and facial semantic features. Figure 4 , the device comprises:
[0178] An image acquisition unit, used for acquiring facial images from multiple perspectives;
[0179] A 2D key point recognition unit, used to extract 2D key points in face images at different viewing angles using the MVCNN method;
[0180] A three-dimensional conversion unit, used to convert the 2D key points at each viewing angle into a 3D world coordinate system, obtain the 3D key points at each viewing angle and map them onto a preset 3D face model;
[0181] The 3D fusion unit is used to group the same 3D key points under different viewing angles according to the key point sequence number, and perform weighted fusion on the group of 3D key points to obtain preliminary face key points under each viewing angle;
[0182] The 3D key point calibration unit is used to calculate the facial semantic features of the preliminary face key points on the face 3D model at each viewing angle, calibrate the positions of the preliminary face key points using the facial semantic features, and obtain the final 3D key points Pw at each viewing angle.
[0183] Please understand the specific functions and interactions of each unit in combination with the above methods. I will not go into details here.
[0184] Figure 5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 5 As shown, the electronic device may include the above Figure 4 The key point recognition device based on the 3D face model and facial semantic features is shown. Optionally, the electronic device 410 may include a first processor 2001.
[0185] Optionally, the electronic device 410 may further include a memory 2002 and a transceiver 2003 .
[0186] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0187] Combine the following Figure 5 The components of the electronic device 410 are described in detail:
[0188] The first processor 2001 is the control center of the electronic device 410, and may be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be application specific integrated circuits (ASICs), or may be one or more integrated circuits configured to implement the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs).
[0189] Optionally, the first processor 2001 can perform various functions of the electronic device 410 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002 .
[0190] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in FIG.
[0191] In a specific implementation, as an embodiment, the electronic device 410 may also include multiple processors, such as Figure 5 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0192] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.
[0193] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently and access the first processor 2001 through the interface circuit ( Figure 5 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0194] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0195] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 5 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0196] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 5 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0197] It should be noted that Figure 5 The structure of the electronic device 410 shown in the figure does not constitute a limitation on the router, and the actual knowledge structure recognition device may include more or fewer components than those shown in the figure, or combine certain components, or arrange the components differently.
[0198] In addition, the technical effects of the electronic device 410 can refer to the technical effects of the key point recognition method based on the 3D face model and facial semantic features described in the above method embodiment, and will not be repeated here.
[0199] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0200] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0201] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0202] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0203] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0204] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0205] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0206] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0207] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0208] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0209] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0210] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0211] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A key point recognition method based on a 3D face model and facial semantic features, characterized in that: The method comprises: S1. Collect multi-view facial images and use the MVCNN method to extract 2D key points in the facial images at each view. Apply the CNN model to perform key point recognition on facial images from multiple angles, and obtain 2D key points through the CNN key point detection model. S2, converting the 2D key points at each viewing angle into a 3D world coordinate system, obtaining the 3D key points at each viewing angle and mapping them onto a preset 3D face model; S3, according to the key point sequence number, the same 3D key points under different viewing angles are grouped together, and the group of 3D key points is weightedly fused to obtain preliminary face key points under each viewing angle; S4, calculating the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, calibrating the positions of the preliminary facial key points using the facial semantic features, and obtaining the final 3D key points Pw at each viewing angle, including: Identification of zygomatic points: 1). Given multi-view face images V1, V2, V3; 2). Through the CNN key point detection model, obtain the 2D key points P1, P2, P3, convert them to the 3D world coordinate system to obtain Pw1, Pw2, Pw3; 3). According to the key point sequence number, the same 3D key points under different perspectives are divided into a group, and the group of 3D key points: Pw1, Pw2, Pw3 are weighted fused to obtain the preliminary 3D key point Pw; 4). Obtain the 3D model curvature K of the area near Pw, calculate the position of the maximum point of curvature K, and obtain the accurate key point Pw, which is the final result; calculate the 3D model curvature K of the area near the preliminary face key point on the face 3D model under each viewing angle according to the following formula: , in: df(X) represents the curvature of the preliminary facial key points on the 3D face model; dN(X) represents the curvature of a point located near the preliminary facial key point on the 3D face model; |df(X)| represents the absolute value of curvature; k n (X) represents the 3D model curvature K of the vicinity of the preliminary facial key point on the 3D face model; For several points near the preliminary facial key point, calculate the 3D model curvature K of each point according to the above steps to obtain several 3D model curvatures K, and find the position of the point with the largest 3D model curvature K, and take this point as the final 3D key point Pw; The above-mentioned calculation method of the 3D model curvature K is adopted to calculate the curvature change characteristics of the area near the facial key points on the 3D face model, so as to reflect the local convexity and concavity of the face at this point and its changes; by ranking the calculation results of several nearby points, the point with the largest change is obtained, which is used as the feature point and the preliminary facial key points are calibrated, and this point is used as the final key point with the characteristics of the local area.
2. The key point recognition method based on face 3D model and facial semantic features according to claim 1, characterized in that: In step S2, the 2D key points at each viewing angle are converted to a 3D world coordinate system according to the following formula: , in: Pw represents the 3D world coordinates. Pw is a point on the ray from the real 3D key point to the center point of the camera. It is necessary to introduce a ray detection algorithm to find the intersection point on the surface of the 3D face model. The intersection point is the real 3D key point. R,t is the external parameter of the camera; K is the intrinsic parameter of the camera; u,v are the coordinates of the 2D keypoints in the image; zc is artificially given depth data.
3. The key point recognition method based on a 3D face model and facial semantic features according to claim 1, characterized in that: Step S3, weighted fusion of the group of 3D key points to obtain preliminary face key points at each viewing angle, including: Obtaining groups of 3D key points under different viewing angles, and assigning weights to the 3D key points in each group according to the positions of the 3D key points; The weighted fusion of each group of 3D key points is performed according to the following formula: , in: Pwi represents the 3D key points obtained from each perspective; ai is the weight assigned to the perspective, and the weight is assigned according to the position information of the key point itself.
4. The key point recognition method based on face 3D model and facial semantic features according to claim 1, characterized in that: S4, calculating the facial semantic features of the preliminary facial key points at each viewing angle on the 3D face model, calibrating the positions of the preliminary facial key points using the facial semantic features, and obtaining the final 3D key points Pw at each viewing angle, including: Acquire nearby point data of the preliminary facial key points on the 3D face model at each viewing angle; Performing denoising and filtering processing on the nearby point data; Using a curve approximate fitting method, fitting calculation is performed on the preprocessed nearby point data to generate a corresponding nearby point fitting curve; Find a point on the nearby point fitting curve that has similar features to the preliminary face key point, and take the point as the final 3D key point Pw.
5. A key point recognition device based on a 3D face model and facial semantic features, the key point recognition device based on a 3D face model and facial semantic features is used to implement the key point recognition method based on a 3D face model and facial semantic features as claimed in any one of claims 1 to 4, characterized in that: The device comprises: An image acquisition unit, used for acquiring facial images from multiple perspectives; A 2D key point recognition unit, used to extract 2D key points in face images at different viewing angles using the MVCNN method; A three-dimensional conversion unit, used to convert the 2D key points at each viewing angle into a 3D world coordinate system, obtain the 3D key points at each viewing angle and map them onto a preset 3D face model; The 3D fusion unit is used to group the same 3D key points under different viewing angles according to the key point sequence number, and perform weighted fusion on the group of 3D key points to obtain preliminary face key points under each viewing angle; The 3D key point calibration unit is used to calculate the facial semantic features of the preliminary face key points on the face 3D model at each viewing angle, calibrate the positions of the preliminary face key points using the facial semantic features, and obtain the final 3D key points Pw at each viewing angle.
6. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for detecting key points on three-dimensional face scanning
CN111753644A