Methods for generating images of orthodontic treatment effects using artificial neural networks

Through artificial neural network technology, combined with feature extraction and image generation network, the problem that the existing technology is difficult to predict the treatment effect of dental orthodontics is solved, and high-quality patient appearance image generation is achieved after treatment, which improves the patient's confidence and treatment communication effect.

CN113223140BActive Publication Date: 2025-05-13HANGZHOU ZOHO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010064195.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-20
Publication Date
2025-05-13
Estimated Expiration
2040-01-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict and display patient appearance images after dental orthodontic treatment, especially in the presence of high-quality and realistic effect presentation.

Method used

Artificial neural networks are used, including feature extraction deep neural networks and image generation deep neural networks. By obtaining the patient's exposed face photos, extracting oral area masks and tooth contour features, and combining pose optimization with three-dimensional digital models, the patient's appearance images after dental orthodontic treatment are generated.

Benefits of technology

The generation of exposed face images of patients after dental orthodontic treatment that are very close to the actual situation is achieved, helping patients build confidence in treatment and promoting communication between orthodontic doctors and patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113223140B_ABST
    Figure CN113223140B_ABST
Patent Text Reader

Abstract

One aspect of the present application provides a method for generating an image of the effect of dental orthodontic treatment using an artificial neural network, comprising: obtaining a toothy face photograph of a patient before orthodontic treatment; extracting an oral area mask and a first group of tooth contour features from the toothy face photograph of the patient before orthodontic treatment using a trained feature extraction deep neural network; obtaining a first three-dimensional digital model representing the original tooth layout of the patient and a second three-dimensional digital model representing the target tooth layout of the patient; obtaining a first pose of the first three-dimensional digital model based on the first group of tooth contour features and the first three-dimensional digital model; obtaining a second group of tooth contour features based on the second three-dimensional digital model in the first pose; and generating an image of the toothy face of the patient after orthodontic treatment based on the toothy face photograph of the patient before orthodontic treatment, the mask and the second group of tooth contour features using a trained picture generation deep neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to a method for generating images of dental orthodontic treatment effects using an artificial neural network. Background Art

[0002] Nowadays, more and more people are beginning to understand that orthodontic treatment is not only good for health, but also can improve personal appearance. For patients who are not familiar with orthodontic treatment, if they can be shown the appearance of their teeth and face after the treatment is completed before the treatment, it can help them build confidence in the treatment and promote communication between orthodontists and patients.

[0003] Currently, there is no similar imaging technology that can predict the effect of dental orthodontic treatment, and the traditional technology using 3D model texture mapping often cannot meet the requirements of high-quality realistic effect presentation. Therefore, it is necessary to provide a method for generating an image of the patient's appearance after dental orthodontic treatment. Summary of the invention

[0004] One aspect of the present application provides a method for generating an image of the effect of dental orthodontic treatment using an artificial neural network, comprising: obtaining a toothy face photograph of a patient before orthodontic treatment; extracting an oral area mask and a first group of tooth contour features from the toothy face photograph of the patient before orthodontic treatment using a trained feature extraction deep neural network; obtaining a first three-dimensional digital model representing the original tooth layout of the patient and a second three-dimensional digital model representing the target tooth layout of the patient; obtaining a first pose of the first three-dimensional digital model based on the first group of tooth contour features and the first three-dimensional digital model; obtaining a second group of tooth contour features based on the second three-dimensional digital model in the first pose; and generating an image of the toothy face of the patient after orthodontic treatment based on the toothy face photograph of the patient before orthodontic treatment, the mask and the second group of tooth contour features using a trained picture generation deep neural network.

[0005] In some implementations, the image generation deep neural network may be a CVAE-GAN network.

[0006] In some embodiments, the sampling method adopted by the CVAE-GAN network may be a differentiable sampling method.

[0007] In some embodiments, the feature extraction deep neural network can be a U-Net network.

[0008] In some embodiments, the first posture is obtained based on the first set of tooth contour features and the first three-dimensional digital model using a nonlinear projection optimization method, and the second set of tooth contour features is obtained by projection based on the second three-dimensional digital model in the first posture.

[0009] In some embodiments, the method of generating an image of the effect of dental orthodontic treatment using an artificial neural network may also include: using a facial key point matching algorithm to capture a first mouth area image from a toothy face photograph of the patient before orthodontic treatment, wherein the mouth area mask and a first set of tooth contour features are extracted from the first mouth area image.

[0010] In some embodiments, the toothy facial photograph of the patient before orthodontic treatment may be a complete frontal facial photograph of the patient.

[0011] In some embodiments, the edge contour of the mask matches the inner edge contour of the lips in the toothy face photograph of the patient before orthodontic treatment.

[0012] In some embodiments, the first set of tooth contour features includes edge contour lines of teeth visible in a toothy facial photograph of the patient before orthodontic treatment, and the second set of tooth contour features includes edge contour lines of teeth when the second three-dimensional digital model is in the first posture.

[0013] In some embodiments, the tooth contour feature may be a tooth edge feature map. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other features of the present disclosure will be more fully understood through the following description and the attached claims in conjunction with the accompanying drawings. It should be understood that these drawings only depict several embodiments of the present disclosure and should not be considered as limiting the scope of the present disclosure. By using the drawings, the present disclosure will be more clearly and detailed.

[0015] Figure 1 A schematic flow chart of a method for generating an image of a patient's appearance after dental orthodontic treatment using an artificial neural network in one embodiment of the present application;

[0016] Figure 2 This is a first oral region picture in one embodiment of the present application;

[0017] Figure 3 In one embodiment of the present application, based on Figure 2 a mask generated from the first mouth region image shown;

[0018] Figure 4 In one embodiment of the present application, based on Figure 2 a first tooth edge feature map generated from the first oral region image shown;

[0019] Figure 5 This is a structural diagram of a feature extraction deep neural network in one embodiment of the present application;

[0020] Figure 5A Schematically shows an embodiment of the present application Figure 5 The structure of the convolutional layer of the feature extraction deep neural network shown;

[0021] Figure 5B Schematically shows an embodiment of the present application Figure 5 The structure of the deconvolution layer of the feature extraction deep neural network shown;

[0022] Figure 6 A second tooth edge feature diagram in one embodiment of the present application;

[0023] Figure 7 A structural diagram of a deep neural network for generating images in one embodiment of the present application; and

[0024] Figure 8 This is a picture of the second oral area in one embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following detailed description, reference is made to the accompanying drawings which form a part hereof. In the drawings, similar symbols generally indicate similar components unless the context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be employed, and other changes may be made, without departing from the spirit or scope of the subject matter described herein. It should be readily understood that various configurations, substitutions, combinations, designs of various aspects of the disclosure generally described herein and illustrated in the accompanying drawings may be made, all of which are expressly contemplated and form a part of the disclosure.

[0026] After extensive research, the inventors of this application have found that with the rise of deep learning technology, in some fields, adversarial generative network technology has been able to generate realistic images. However, in the field of orthodontics, there is still a lack of robust technology for generating images based on deep learning. After extensive design and experimental work, the inventors of this application have developed a method for generating images of the patient's appearance after orthodontic treatment using an artificial neural network.

[0027] Please refer to Figure 1 , is a schematic flowchart of a method 100 for generating an image of a patient's appearance after dental orthodontic treatment using an artificial neural network in one embodiment of the present application.

[0028] In 101, a toothy face photograph of a patient before orthodontic treatment is obtained.

[0029] Since people usually care more about their image when they smile with their teeth showing, in one embodiment, the tooth-showing facial photograph of a patient before orthodontic treatment can be a full frontal photograph of the patient's face when the patient smiles with his teeth showing, and such a photograph can more clearly reflect the difference before and after orthodontic treatment. In light of this application, it can be understood that the tooth-showing facial photograph of a patient before orthodontic treatment can also be a partial facial photograph, and the angle of the photograph can also be other angles besides the frontal angle.

[0030] In 103, a facial key point matching algorithm is used to capture a first mouth region image from a toothy face photograph of a patient before orthodontic treatment.

[0031] Compared with complete face photos, mouth area images have fewer features. Subsequent processing based only on mouth area images can simplify calculations, make artificial neural networks easier to learn, and make artificial neural networks more robust.

[0032] The facial key point matching algorithm can refer to "Displaced Dynamic Expression Regression for Real-Time Facial Tracking and Animation" by Chen Cao, Qiming Hou and Kun Zhou, published in 2014. ACM Transactions on Graphics (TOG) 33, 4 (2014), 43, and "One Millisecond Face Alignment with an Ensemble of Regression Trees" by Vahid Kazemi and Josephine Sullivan, published in Proceedings of the IEEE conference on computer vision and pattern recognition, 1867--1874, 2014.

[0033] In light of this application, it is understood that the scope of the mouth area can be freely defined. Figure 2 , is a picture of the oral area of ​​a patient before orthodontic treatment in one embodiment of the present application. Figure 2 The mouth area image includes a portion of the nose and a portion of the chin, but as mentioned above, the range of the mouth area can be reduced or expanded according to specific needs.

[0034] In 105 , a mouth region mask and a first set of tooth contour features are extracted based on the first mouth region image using a trained feature extraction deep neural network.

[0035] In one embodiment, the extent of the mouth region mask may be defined by the inner edge of the lips.

[0036] In one embodiment, the mask can be a black and white bitmap, and the undesired parts of the image can be removed through mask operation. Figure 3 , is based on an embodiment of the present application Figure 2 The mouth area mask obtained by using the mouth area image of .

[0037] The tooth contour feature may include the contour line of each tooth visible in the image, which is a two-dimensional feature. In one embodiment, the tooth contour feature may be a tooth contour feature map, which only includes the contour information of the tooth. In another embodiment, the tooth contour feature may be a tooth edge feature map, which includes not only the contour information of the tooth, but also the edge features inside the tooth, for example, the edge line of the spot on the tooth. Figure 4 , is based on an embodiment of the present application Figure 2 A tooth edge feature map obtained from a picture of the mouth area.

[0038] In one embodiment, the feature extraction neural network can be a U-Net network. Figure 5 , schematically showing the structure of the feature extraction neural network 200 in one embodiment of the present application.

[0039] The feature extraction neural network 200 may include 6 layers of convolution 201 (downsampling) and 6 layers of deconvolution 203 (upsampling).

[0040] Please refer to Figure 5A , each convolution layer 2011 (down) may include a convolution layer 2013 (conv), a ReLU activation function 2015 and a maximum pooling layer 2017 (max pool).

[0041] Please refer to Figure 5B , each deconvolution layer 2031 (up) may include a sub-pixel convolution layer 2033 (sub-pixel), a convolution layer 2035 (conv) and a ReLU activation function 2037.

[0042] In one embodiment, a training atlas for training a feature extraction neural network can be obtained by: obtaining multiple facial photos showing teeth; capturing mouth area images from these facial photos; and generating their respective mouth area masks and tooth edge feature maps based on these mouth area images using the PhotoShop Lasso annotation tool. These mouth area images and the corresponding mouth area masks and tooth edge feature maps can be used as a training atlas for training a feature extraction neural network.

[0043] In one embodiment, in order to improve the robustness of the feature extraction neural network, the training atlas may be augmented, including Gaussian smoothing, rotation, and horizontal flipping.

[0044] At 107, a first three-dimensional digital model representing the original dental arrangement of the patient is acquired.

[0045] The patient's original dental arrangement is the dental arrangement before orthodontic treatment.

[0046] In some embodiments, the three-dimensional digital model representing the original tooth arrangement of the patient can be obtained by directly scanning the patient's jaw. In other embodiments, a physical model of the patient's jaw, such as a plaster model, can be scanned to obtain a three-dimensional digital model representing the original tooth arrangement of the patient. In other embodiments, an impression of the patient's jaw can be scanned to obtain a three-dimensional digital model representing the original tooth arrangement of the patient.

[0047] In 109, a first pose of the first three-dimensional digital model matching the first set of tooth contour features is calculated using a projection optimization algorithm.

[0048] In one embodiment, the optimization objective of the nonlinear projection optimization algorithm can be expressed as equation (1):

[0049]

[0050] in, represents the sampling points on the first three-dimensional digital model, p i Represents the point on the tooth contour line in the first tooth edge feature map corresponding to it.

[0051] In one embodiment, the correspondence between the first three-dimensional digital model and the first set of tooth contour features may be calculated based on the following equation (2):

[0052]

[0053] Among them, t i and t j Represents p i and p j Tangent vector at two points.

[0054] At 111 , a second three-dimensional digital model representing a target dental configuration of the patient is acquired.

[0055] The method of obtaining a three-dimensional digital model representing a patient's target tooth layout based on a three-dimensional digital model representing the patient's original tooth layout is well known in the industry and will not be described in detail here.

[0056] In 113, the second three-dimensional digital model in the first pose is projected to obtain a second set of tooth contour features.

[0057] In one embodiment, the second set of tooth contour features includes edge contour lines of all teeth of the complete upper and lower dentitions when they are in the target tooth layout and in the first posture.

[0058] Please refer to Figure 6 , which is a second tooth edge feature diagram in one embodiment of the present application.

[0059] In 115, a toothy face picture of the patient after orthodontic treatment is generated based on the toothy face picture of the patient before orthodontic treatment, the mask, and the second set of tooth contour feature maps using a trained deep neural network for generating pictures.

[0060] In one embodiment, a CVAE-GAN network can be used as a deep neural network for generating images. Figure 7 , schematically illustrating the structure of a deep neural network 300 for generating images in one embodiment of the present application.

[0061] The deep neural network 300 for generating pictures includes a first sub-network 301 and a second sub-network 303. Part of the first sub-network 301 is responsible for processing shapes, and the second sub-network 303 is responsible for processing textures. Therefore, the toothy face photo of the patient before orthodontic treatment or the part of the masked area in the first mouth area photo can be input into the second sub-network 303, so that the deep neural network 300 for generating pictures can generate textures for the masked area in the toothy face photo of the patient after orthodontic treatment; and the mask and the second tooth edge feature map are input into the first sub-network 301, so that the deep neural network 300 for generating pictures can divide the masked area in the toothy face photo of the patient after orthodontic treatment into regions, that is, which part is teeth, which part is gums, which part is tooth gaps, which part is tongue (when the tongue is visible), etc.

[0062] The first sub-network 301 includes 6 layers of convolution 3011 (downsampling) and 6 layers of deconvolution 3013 (upsampling). The second sub-network 303 includes 6 layers of convolution 3031 (downsampling).

[0063] In one embodiment, the deep neural network 300 for generating images can use a differentiable sampling method to facilitate end-to-end training. For similar sampling methods, see "Auto-Encoding Variational Bayes" by Diederik Kingma and Max Welling published in ICLR 12 2013.

[0064] The training of the deep neural network 300 for generating images may be similar to the training of the feature extraction neural network 200 described above, and will not be described in detail here.

[0065] Inspired by this application, it can be understood that in addition to the CVAE-GAN network, networks such as cGAN, cVAE, MUNIT and CycleGAN can also be used as networks for generating images.

[0066] In one embodiment, a portion of the mask area in the toothy facial photograph of the patient before orthodontic treatment can be input into the deep neural network 300 for generating images to generate a portion of the mask area in the toothy facial image of the patient after orthodontic treatment. Then, based on the toothy facial photograph of the patient before orthodontic treatment and the portion of the mask area in the toothy facial image of the patient after orthodontic treatment, the toothy facial image of the patient after orthodontic treatment is synthesized.

[0067] In another embodiment, a portion of the mask area in the first mouth area picture can be input into the deep neural network 300 for generating pictures to generate a portion of the mask area in the toothy facial image of the patient after orthodontic treatment. Then, based on the first mouth area picture and the portion of the mask area in the toothy facial image of the patient after orthodontic treatment, a second mouth area picture is synthesized. Then, based on the toothy facial photo of the patient before orthodontic treatment and the second mouth area picture, a toothy facial image of the patient after orthodontic treatment is synthesized.

[0068] Please refer to Figure 8 , is a second mouth region image in one embodiment of the present application. The toothy face image of a patient after orthodontic treatment generated by the method of the present application is very close to the actual effect and has a high reference value. With the toothy face image of a patient after orthodontic treatment, it can effectively help the patient build confidence in the treatment and promote communication between the orthodontist and the patient.

[0069] In light of the present application, it can be understood that, although a complete facial picture of a patient after orthodontic treatment can allow the patient to better understand the treatment effect, this is not necessary. In some cases, a picture of the patient's oral area after orthodontic treatment is sufficient for the patient to understand the treatment effect.

[0070] Although various aspects and embodiments of the present application are disclosed herein, other aspects and embodiments of the present application will be apparent to those skilled in the art in light of the present application. The various aspects and embodiments disclosed herein are for illustrative purposes only and not for limiting purposes. The scope and subject matter of the present application are determined solely by the appended claims.

[0071] Likewise, various diagrams may illustrate exemplary architectures or other configurations of the disclosed methods and systems that aid in understanding the features and functions that may be included in the disclosed methods and systems. The claimed content is not limited to the exemplary architectures or configurations shown, and the desired features may be implemented with various alternative architectures and configurations. In addition, for flow charts, functional descriptions, and method claims, the order of blocks presented herein should not be limited to various embodiments that are implemented in the same order to perform the described functions, unless otherwise clearly indicated in the context.

[0072] Unless otherwise expressly noted, the terms and phrases used herein and their variations should be interpreted as open ended, not restrictive. In some instances, the appearance of broad words and phrases such as "one or more", "at least", "but not limited to", or other similar terms should not be understood as an intent or need to indicate a narrowing of the example where such broad terms may not be present.

Claims

1. A method for generating an image of dental orthodontic treatment effect using an artificial neural network, comprising: Obtain a toothy facial photograph of the patient before orthodontic treatment; Extracting a mouth region mask and a first set of tooth contour features from a toothy face photograph of the patient before orthodontic treatment using a trained feature extraction deep neural network; Acquire a first three-dimensional digital model representing the original dental configuration of the patient and a second three-dimensional digital model representing the target dental configuration of the patient; Based on the first group of tooth contour features and the first three-dimensional digital model, obtaining a first pose of the first three-dimensional digital model; obtaining a second set of tooth contour features based on the second three-dimensional digital model in the first posture; as well as A deep neural network is generated using the trained images, and based on the toothy facial photograph of the patient before orthodontic treatment, the mask and the second set of tooth contour features, an image of the toothy facial of the patient after orthodontic treatment is generated.

2. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 1, characterized in that: The image generation deep neural network is a CVAE-GAN network.

3. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 2, characterized in that: The sampling method adopted by the CVAE-GAN network is a differentiable sampling method.

4. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 1, characterized in that: The feature extraction deep neural network is a U-Net network.

5. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 1, characterized in that: The first posture is obtained based on the first group of tooth contour features and the first three-dimensional digital model using a nonlinear projection optimization method, and the second group of tooth contour features is obtained through projection based on the second three-dimensional digital model in the first posture.

6. The method for generating an image of dental orthodontic treatment effect using an artificial neural network according to any one of claims 1 to 5, characterized in that: It also includes: using a facial key point matching algorithm to capture a first mouth area picture from a toothy face photo of the patient before orthodontic treatment, wherein the mouth area mask and a first set of tooth contour features are extracted from the first mouth area picture.

7. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 6, characterized in that: The toothy face photograph of the patient before orthodontic treatment is a complete frontal face photograph of the patient.

8. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 6, characterized in that: The edge contour of the mask matches the inner edge contour of the lips in the toothy face photograph of the patient before orthodontic treatment.

9. The method for generating an image of dental orthodontic treatment effect using an artificial neural network as claimed in claim 8, characterized in that: The first set of tooth contour features includes edge contour lines of teeth visible in the toothy face photo of the patient before orthodontic treatment, and the second set of tooth contour features includes edge contour lines of teeth when the second three-dimensional digital model is in the first posture.

10. The method for generating an image of dental orthodontic treatment effect using an artificial neural network according to claim 9, characterized in that: The first and second sets of tooth contour features are tooth edge feature maps.

Citation Information

Patent Citations

  • Orthodontic accessory planning method based on oral voxel model feature extraction

    CN110428021A

  • Method of manufacturing orthodontic devices

    US10258439B1