A method for intelligently determining three-dimensional facial landmark points

Through multi-view rendering and stacked hourglass neural network model combined with the hierarchical attention supervision module, the problem of automatic determination of three-dimensional facial marking points is solved, and fast and accurate marking points are achieved intelligent determination, suitable for orthodontic and maxillofacial surgical design.

CN114529967BActive Publication Date: 2025-07-25PEKING UNIV SCHOOL OF STOMATOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111639065.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-25
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The prior art is difficult to achieve rapid, accurate and automated in the determination of three-dimensional facial marking points, especially in the design of orthodontic and maxillofacial surgical procedures.

Method used

A multi-view rendering and stacked hourglass neural network model based on three-dimensional faces is adopted, combined with the hierarchical attention supervision module, to realize the automatic determination of three-dimensional face logo points.

Benefits of technology

It realizes the rapid and accurate determination of three-dimensional facial marking points under small sample training, which conforms to expert diagnostic strategies and meets the diagnostic analysis needs of oral clinical practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529967B_ABST
    Figure CN114529967B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for intelligently determining three-dimensional facial landmark points, which comprises the following steps: (1) Rendering multi-views of three-dimensional facial data; (2) Constructing two-dimensional heatmaps by using a multi-view stacked hourglass neural network model; (3) Automatically determining three-dimensional facial landmark points; The present invention proposes three-dimensional multi-view rendering of a human face based on a virtual camera to realize the transformation between a three-dimensional model and a two-dimensional depth image; The automatic determination of three-dimensional facial anatomical landmark points is realized by using a stacked hourglass neural network and a hierarchical attention supervision module signal, achieving an intelligent determination effect of landmark points that conforms to the expert diagnosis strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to three-dimensional facial landmarks; specifically, it relates to a method for intelligently determining three-dimensional facial landmark points. Background Art

[0002] Three-dimensional facial landmarks are important anatomical features of the face and also an important basis for facial morphology analysis. They have extensive applications in disciplines such as orthognathic surgery, orthodontics, and prosthodontics, and also play a very important role in fields such as face recognition, face tracking, expression analysis, and face modeling. With the wide application of three-dimensional scanning devices and the urgent clinical need for automatic determination of landmark points, efficient and automatic determination of facial landmark points has become a research hotspot.

[0003] The accurate determination of three-dimensional facial anatomical standard points is the basis and prerequisite for three-dimensional facial analysis, directly affecting the surgical design and treatment evaluation of orthodontics, maxillofacial surgery, and prosthodontics. Previous studies mostly used the method of manually marking points, which has a large workload and strong dependence on experience. Therefore, how to achieve automatic, accurate, and efficient determination of three-dimensional facial anatomical landmark points is a key problem that needs to be solved by both automated algorithms and intelligent algorithms.

[0004] 1 Stacked Hourglass Network Model

[0005] The stacked hourglass network model was initially used for human pose estimation. The network captures and integrates all scale information of the image and realizes pixel-level feature extraction based on the combination of downsampling and upsampling. The network is similar in shape to an hourglass and is symmetric. As Figure 1 shown, each box represents the Feature Map of each key point at this stage. In the first half of the hourglass network, each layer of the network obtains a Feature Map with gradually decreasing resolution through three residual modules and downsampling (Max Pooling) operations, and passes it to the subsequent part to obtain the lowest resolution. In the second half of the network, each layer gradually restores the high-resolution Feature Map through upsampling (Nearest Neighbor Interpolation) and a residual module. The joint features are gradually extracted through the jump layer of the hourglass network and passed to the second half of the hourglass network. Finally, the features of each scale retained by the jump layer are fused with the low-resolution features in the second half. The hourglass network combines the low-level and high-level Feature Maps of the network with its unique bottom-up and top-down structure to capture the spatial position information of the key points.

[0006] Meanwhile, this network takes into account that key points can be predicted with reference to each other. Since the heatmap represents all the key points of the input object, the heatmap contains all the mutual relationships of the key points. Therefore, the stacked structure uses the heatmap given by the first hourglass network as the input of the next hourglass network, which means that the second hourglass network can utilize the mutual relationships between the key points, thereby improving the prediction accuracy of the key points. On February 26, 2020, the paper searched on Baidu: Artificial Intelligence HourglassNet Stacked Hourglass Network (paper address: https: / / arxiv.org / abs / 1603.06937) is this kind of stacked hourglass network model. However, this stacked hourglass network model is only used for human pose estimation and not for the technical solution of the multi-view stacked hourglass neural network for three-dimensional facial landmark points in oral clinical medicine.

[0007] 2 Domestic and foreign research development trends and current situations

[0008] In previous literature, the automatic determination of facial landmark points mainly falls into geometric information analysis algorithms and machine learning algorithms. Geometric information analysis algorithms are automatically determined based on the changes in facial geometric morphology. In 2016, Katina et al. automatically determined 17 landmark points based on the curvature classification of the surface of three-dimensional facial data, but the positioning effect of this method for landmark points in areas with less obvious facial geometric features is poor; in 2017, Liang Yan et al. proposed a method for automatically determining 8 anatomical landmark points on three-dimensional facial data by combining HK curvature analysis with prior knowledge of facial geometric shapes.

[0009] In recent years, with the continuous development of deep learning, applying deep learning algorithms for face data analysis has become a research hotspot. Among them, the research on key point detection based on two-dimensional faces is relatively mature and is widely used in face recognition, expression recognition, face frontalization, etc. With the wide popularization of three-dimensional scanning devices, the research on automatically determining landmark points based on three-dimensional face data has gradually attracted extensive attention. In 2013, Sun et al. first used convolutional neural networks (CNN) to regress the key points of the original two-dimensional face image, realizing the regression of 5 landmark points on the two-dimensional face image from rough to fine. In 2017, Liu et al. cascaded six convolutional neural networks to determine facial landmark points in two-dimensional face images with extreme poses, and this method can cascade to predict three-dimensional face shapes and projection matrices. The above algorithm research is mostly applied in the field of face recognition and has not been applied to medical research and clinical practice currently.

[0010] 3 Review and summary

[0011] Through the above review of previous studies, it can be considered that three-dimensional facial landmark points are a hot topic in oral clinical practice. Establishing an intelligent algorithm that can quickly, accurately, automatically, and batch determine anatomical landmark points on a three-dimensional facial digital model is the key problem to be solved. Geometric information analysis algorithms automatically determine anatomical landmark points according to the laws of facial geometric shape changes, but the number of landmark points they determine is limited. The research on artificial intelligence algorithms is the future development direction. The present invention uses a deep learning algorithm in the field of artificial intelligence to establish a multi-view stacked hourglass neural network that can automatically determine three-dimensional facial anatomical landmark points, achieving an intelligent determination effect of landmark points that conforms to expert diagnosis strategies.

[0012] References

[0013] [1] O'Grady K, Antonyshyn O. Facial Asymmetry: Three-Dimensional Analysis Using Laser Surface Scanning[J]. Plast Reconstr Surg, 1999, 4(104): 928 - 937.

[0014] [2] Haraguchi S, Iguchi Y, Takada K. Asymmetry of the face in orthodontic patients[J]. Angle Orthod, 2008, 78(3): 421 - 426.

[0015] [3] Lee MS, Chung DH, Lee JW, et al. Assessing soft-tissue characteristics of facial asymmetry with photographs[J]. Am J Orthod Dentofacial Orthop, 2010, 138(1): 23 - 31.

[0016] [4] Guo Hongming, Bai Yuxing, Zhou Lixin, et al. Three-dimensional measurement study on the asymmetry of normal occlusal facial soft tissues in Beijing area[J]. Beijing Journal of Stomatology, 2006, 14(1): 50 - 52.

[0017] [5] Hartmann J, Meyer-Marcotty P, Benz M, et al. Reliability of a method for computing facial symmetry plane and degree of asymmetry based on 3D-data[J]. J Orofac Orthop, 2007, 68(6): 477-490.

[0018] [6] Klingenberg CP, Barluenga M, Meyer A. Shape analysis of symmetric structures: quantifying variation among individuals and asymmetry[J]. Evolution, 2002, 56(10): 1909-1920.

[0019] [7] Xiong Yuxue, Yang Huifang, Zhao Yijiao, et al. Comparison of two methods for evaluating facial three-dimensional surface data asymmetry[J]. Journal of Peking University(Health Sciences), 2015, 47(2): 340-343.

[0020] [8] Zhu Y, Zheng S, Yang G, et al. A novel method for 3D face symmetry reference plane based on weighted Procrustes analysis algorithm[J]. Bmc Oral Health, 2020, 20(1): 1-11.

[0021] [9] Zhu Yujia, Zhao Yijiao, Zheng Shengwen, et al. Construction method of three-dimensional facial symmetry reference plane based on weighted morphological analysis[J]. Journal of Peking University: Health Sciences, 2020, 53(1): 220-226.

[0022]

[10] De Momi E, Chapuis J, Pappas I, et al. Automatic extraction of the mid-facial plane for cranio-maxillofacial surgery planning[J]. Int J Oral Max Surg, 2006, 35(7): 636-642.

[0023]

[11] Benz M,Laboureux X,Maier T,et al.The Symmetry of Faces[C].Proceedings of the Vision,Modeling,and Cisualization Conference(VMV),2002:43-50.

[0024]

[12] Tian Kaiyue.Digital Orthodontic Treatment Plan Design for Mandibular Protrusion and Deviation Deformity[D].Peking University Health Science Center,2015.

[0025]

[13] Xiong Y,Zhao Y,Yang H,et al.Comparison Between InteractiveClosest Point and Procrustes Analysis for Determining the Median SagittalPlane of Three-Dimensional Facial Data[J].J Craniofac Surg,2016,27(2):441-444.

[0026]

[14] Zelditch,Leah M.Geometric morphometrics for biologists:A primer.[M].New York and London:Elsevier Academic Press,2004:293-319.

[0027]

[15] Xiong Y,Zhao Y,Yang H,et al.Comparison Between InteractiveClosest Point and Procrustes Analysis for Determining the Median SagittalPlane of Three-Dimensional Facial Data[J].J Craniofac Surg,2016,27(2):441-444.

[0028]

[16] Li M, Cole JB, Manyama M, et al. Rapid automated landmarking for morphometric analysis of three-dimensional facial scans[J]. J Anat, 2017, 230(4): 607-618.

[0029]

[17] Agbolade O, Nazri A, Yaakob R, et al. Homologous Multi-Points Warping: An Algorithm for Automatic 3D Facial Landmark[C]. 2019 IEEE International Conference on Automatic Control and Intelligent Systems (I2CACIS), 2019: 79-84.

[0030]

[18] Creusot C, Pears N, Austin J. A Machine-Learning Approach to Keypoint Detection and Landmarking on 3D Meshes[J]. Int J Comput Vision, 2013, 102(1-3): 146-179.

[0031]

[19] Su H, Maji S, Kalogerakis E, et al. Multi-view convolutional neural networks for 3d shape recognition[C]. 2015 IEEE International Conference on Computer Vision (ICCV), 2015: 945-953.

[0032]

[20] Paulsen RR, Juhl KA, Haspang TM, et al. Multi-view consensus CNN for 3D facial landmark placement[C]. 2018 Computer Vision - ACCV, 2018: 706-719. Summary of the Invention

[0033] (I) Technical Problems to be Solved

[0034] The object of the present invention is to provide a method for intelligently determining three-dimensional facial landmark points, which proposes to realize the transformation between a three-dimensional model and a two-dimensional depth image based on multi-view rendering of a three-dimensional face; adopts a stacked hourglass neural network and a hierarchical attention supervision module signal to automatically determine three-dimensional facial anatomical landmark points; and uses a deep learning algorithm in the field of artificial intelligence to establish a multi-view stacked hourglass neural network model capable of automatically determining three-dimensional facial anatomical landmark points, achieving an intelligent determination effect of landmark points that conforms to the expert diagnosis strategy.

[0035] (II) Technical solution

[0036] A method for intelligently determining three-dimensional facial landmark points of the present invention includes the following steps:

[0037] (1) Rendering multi-views of three-dimensional facial data:

[0038] Set the focus of the virtual camera at the geometric center focus of the three-dimensional facial data, and use the virtual camera to capture images at dozens of random positions at different angles of the three-dimensional face. Each view needs to include the complete two-dimensional facial image at that angle, and render the above dozens of two-dimensional views to obtain their depth images. The view rendering process is implemented based on the python open-source toolkit vtk;

[0039] (2) Constructing two-dimensional heatmaps using a multi-view stacked hourglass neural network model:

[0040] Apply the multi-view stacked hourglass neural network model trained by the algorithm to calculate the two-dimensional heatmaps of the dozens of two-dimensional views constructed in the above step (1) respectively.

[0041] (3) Automatically determining three-dimensional facial landmark points:

[0042] Realize the result of projecting the landmark point two-dimensional heatmap coordinates to the corresponding position of the three-dimensional facial data through the mapping of the virtual camera matrix, that is, realize the automatic determination of the landmark points of the three-dimensional facial data.

[0043] Among them, the focus of the algorithm virtual camera in the above step (1) is the geometric center of the face model and the origin of the coordinate system.

[0044] Among them, the multi-view stacked hourglass neural network model further includes a supervision module with a multi-level attention mechanism to further improve the ability of the representation space.

[0045] Among them, the multi-view hourglass neural network model is written by the pytorch framework, uses an initial learning rate of 0.001, 100 iteration times, and sets a batch size of 8.

[0046] (III) Beneficial effects

[0047] The advantages of the present invention are as follows:

[0048] 1. It can quickly, accurately, automatically, and batch-determine anatomical landmark points on a three-dimensional facial digital model. By using the deep learning algorithm in the field of artificial intelligence, a multi-view stacked hourglass neural network model capable of automatically determining three-dimensional facial anatomical landmark points is established, achieving an intelligent determination effect of landmark points that conforms to the expert diagnosis strategy.

[0049] 2. Under the premise that the training samples are small samples, the present invention can enhance data learning of more landmark point information through the multi-view stacked hourglass neural network model, and achieve higher accuracy in the form of hierarchical attention and supervised module signals, which can meet the diagnostic analysis needs of a large amount of data in oral clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the structural diagram of the existing stacked hourglass network model;

[0051] Figure 1 In it: c1-c7, c1a-c4a, c1b-c4b: represent the Feature Map (feature image) of each key point;

[0052] Figure 2 is the structural schematic diagram of the multi-view hourglass neural network model of the present invention;

[0053] Figure 2 In it: 1. Multi-view; 2. Multi-view stacked hourglass neural network model; 3. Determine the three-dimensional facial landmarks; 4. Hourglass neural network module; 5. Attention block; 6. Stacking. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0055] A method for intelligent determination of three-dimensional human face landmark points of the present invention includes the following steps:

[0056] (1) Render multi-views of three-dimensional facial data:

[0057] Manually adjust the geometric center of the three-dimensional facial data to the focus of the algorithm virtual camera, set the virtual camera at dozens of random positions at different angles of the three-dimensional human face to capture images, each view needs to include the complete facial two-dimensional image at that angle, and render the depth images of the above dozens of two-dimensional views. The view rendering process is implemented based on the python open-source toolkit vtk;

[0058] (2) Build a multi-view stacked hourglass neural network model:

[0059] The constructed multi-view stacked hourglass neural network model is trained with the multi-view depth images of 80 cases of 3D facial data in the training set using the stacked hourglass neural network model algorithm. The landmark points calculated by the stacked hourglass neural network model are presented in the form of 2D heatmaps, and then the 2D landmark points are projected onto the corresponding positions of the 3D facial data through the mapping of the virtual camera matrix.

[0060] (3) Automatic determination of 3D facial landmark points:

[0061] Based on the result of projecting the 2D heatmap coordinates of the landmark points onto the corresponding positions of the 3D facial data through the mapping of the virtual camera matrix, and the multi-view stacked hourglass neural network model constructed in step (2), by inputting the 3D facial data of 20 subjects outside the training set, the landmark points of the 3D facial data can be automatically determined.

[0062] The focus of the algorithm virtual camera in step (1) is the geometric center of the face model, the origin of the coordinate system. In the case of different coordinate systems in the face scan data, in order to adapt to the input of the multi-view stacked hourglass neural network model algorithm, it is necessary to adjust to the origin in the same coordinate system to ensure that there is no error in the rendering process, so that the rendered multi-views have content and no blanks appear.

[0063] The multi-view stacked hourglass neural network model also includes a supervision module with a multi-level attention mechanism to further improve the ability of the representation space. In the supervision module, the intermediate representation modules generated by each hourglass neural network are connected to form a multi-level representation module. This representation module can be regarded as a collective knowledge module extracted from different levels and scales of the hourglass neural network. Using this collective knowledge module as a supervision signal to calibrate the final representation module can obtain a better representation module. In order to include this collective knowledge module, all the intermediate representation modules are stacked, and this method of stacked representation is input into the supervision module of the attention mechanism. The supervision module of the attention mechanism is the attention block. This method of stacked representation is input into the attention block to calculate the weight information corresponding to each key point. The attention block is composed of convolutions and finally uses the sigmoid activation function to form a single-channel attention mechanism. We multiply the obtained attention mechanism (weight parameter) information by the final representation to recalibrate the representation space and teach the multi-view hourglass neural network model to pay more attention to the specific positions of the face.

[0064] The multi-view hourglass neural network model is written in the pytorch framework, with an initial learning rate of 0.001, 100 iteration times, and a batch size of 8.

[0065] In this invention, 20 subjects with no obvious facial deformities were clinically collected, and the method of this invention was used to automatically determine the three-dimensional facial landmark points. The errors of the landmark points in each facial region were within the clinically acceptable range. The error was the smallest in the nasal region, followed by the oral region, and the largest in the orbital region. The preliminary evaluation showed that it could meet the application requirements of oral clinical practice. This invention constructs a multi-view stacked hourglass neural network model, which can automatically determine the three-dimensional facial anatomical landmark points. On the premise of a small training sample, it can learn more landmark point information through multi-view enhanced data, and achieve higher accuracy in the form of hierarchical attention and supervision signals, and can meet the diagnostic analysis needs of a large amount of data in oral clinical practice.

[0066] As described above, the present invention can be more fully realized. The above description is only a relatively reasonable implementation example of the present invention. The protection scope of the present invention includes but is not limited to this. Any non-substantive variant changes based on the technical solution of the present invention by those skilled in the art are included within the scope of the present invention.

Claims

1. A method for intelligently determining three-dimensional facial landmark points, characterized in that Including the following steps: (1) Render multi-views of three-dimensional facial data: Set the focus of the virtual camera at the geometric center focus of the three-dimensional facial data. The virtual camera captures images at several random positions from different angles of the three-dimensional human face. Each view needs to include a complete two-dimensional facial image at that angle, and depth images are rendered for several of the two-dimensional images. The view rendering process is implemented based on the python open-source toolkit vtk; (2) Construct two-dimensional heatmaps using a multi-view stacked hourglass neural network model: Apply the multi-view stacked hourglass neural network model trained by the algorithm to calculate the two-dimensional heatmaps of dozens of two-dimensional images constructed in step (1) above; The multi-view stacked hourglass neural network model also includes a supervision module with a multi-level attention mechanism. In the supervision module, the intermediate representation modules generated by each hourglass neural network are connected to form a multi-level representation module. All the intermediate representation modules are stacked. The method of this stacked representation module is input into the supervision module of the attention mechanism. The supervision module of the attention mechanism is the attention block. The method of this stacked representation is input into the attention block to calculate the weight information corresponding to each key point. The attention block is composed of convolutions and finally uses the sigmoid activation function to form a single-channel attention mechanism; Multiply the obtained attention mechanism information by the final representation to recalibrate the representation space and teach the multi-view hourglass neural network model to pay more attention to the specific positions of the human face; (3) Automatically determine three-dimensional facial landmark points: Realize the result of projecting the two-dimensional heatmap coordinates of the landmark points to the corresponding positions of the three-dimensional facial data through the mapping of the virtual camera matrix, that is, automatically determine the landmark points of the three-dimensional facial data.

2. The intelligent determination method of three-dimensional face landmark points according to claim 1, characterized in that: The focus of the virtual camera in the algorithm of step (1) is the geometric center of the human face model, the origin of the coordinate system.

3. The intelligent determination method of three-dimensional face landmark points according to claim 1, characterized in that: The multi-view stacked hourglass neural network model also includes a supervision module with a multi-level attention mechanism to further improve the ability of the representation space.

4. The intelligent determination method of three-dimensional face landmark points according to claim 1, characterized in that: The multi-view stacked hourglass neural network model is written using the pytorch framework, with an initial learning rate of 0.001, 100 iteration times, and a batch size of 8 set.

Citation Information

Patent Citations

  • A method for 3D reconstruction and texture generation from single-view face based on multi-task learning

    CN109255831A

  • System for estimating a three dimensional pose of one or more persons in a scene

    US10853970B1