Denoising diffusion probabilistic models for post-treatment anatomy prediction in digital oral care
Denoising diffusion probabilistic models, particularly U-Nets, are employed to accurately predict post-treatment anatomy in digital oral care by combining target dentition and anatomy representations, addressing existing challenges in precision and realism.
Patent Information
- Application Number
- PCT/IB2024/062651
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Current technologies face challenges in accurately predicting post-treatment anatomy in digital oral care, particularly in generating realistic 2D or 3D representations of a patient's smile after orthodontic or dental restorative treatment.
The use of denoising diffusion probabilistic models, specifically trained neural networks like U-Nets, to combine representations of a patient's post-treatment target dentition with their anatomy, effectively removing noise from initial representations to generate accurate and precise predictions.
This approach enhances the accuracy and data precision of predicted post-treatment anatomy, allowing for realistic and clinically valid predictions that can aid in treatment planning and patient consultation.
Smart Images

Figure IB2024062651_19062025_PF_FP_ABST
Abstract
Description
DENOISING DIFFUSION PROBABILISTIC MODELS FOR POST-TREATMENT ANATOMY PREDICTION IN DIGITAL ORAL CARERelated Documents
[0001] The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosures of each of PCT Applications with Publication Nos. WO2022123402A1, WO2021245480A1, and W02020026117A1 are incorporated herein by reference. The entire disclosure of each of the following Provisional U.S. Patent Applications is incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 264,914; 63 / 609581; 63 / 588028; 63 / 609938; and 63 / 432,627.Technical Field
[0002] This disclosure relates to configurations and training of denoising diffusion probabilistic models (e.g., based on neural networks) to improve the accuracy and data precision of 2D or 3D representations of a patient’s predicted post-treatment anatomy (e.g., the patient’s predicted smile after orthodontic or dental restorative treatment, among other treatments).Summary
[0003] The present disclosure describes systems and techniques for training and using one or more flowbased ML models (e.g., normalizing flows or denoising diffusion models) to generate 2D or 3D representations of a patient’s predicted post-treatment anatomy. Examples of flow -based ML models include denoising diffusion probabilistic models (e.g., where denoising is performed by trained neural networks such as U-Nets or other encoder-decoder structures). Denoising diffusion-based techniques are described which combine a representation of the patient’s post-treatment target dentition (e.g., 3D meshes in final setup or final occlusion poses) with a representation of the patient’s anatomy (e.g., a 2D photo or a 3D scan of the face). Of the neural networks which may be trained to perform denoising diffusion, U-Nets are an example of models that may enable improvements to data precision. Other hierarchical neural network feature extraction models can, alternatively, be used to perform denoising operations.
[0004] Techniques of this disclosure include methods of a generating a data structure describing the posttreatment anatomy of a patient (e.g., a predicted smile 530). One or more oral care arguments 502 may be provided to such methods. The one or more oral care argument 502 may describe the target output of a trained machine learning model. Training methods such as methods 222, 500 or 800 may generate one or more noisy representations of an intended output, and then use the series of one or more noisy representations to train a denoising diffusion ML model 522 to perform a denoising operation (e.g., to remove noise from initially noisyrepresentations). The output of the fully trained denoising ML model 320 (e.g., the output of the reverse pass 322) may include one or more generated denoised representations, which may be used to define one or more aspects of a patient’s post-treatment anatomy 324 (e.g., the appearance of the patient’s face in combination with the target dentition), or may be used as a part of one or more other digital oral care treatments. During the forward pass 220, a training dataset may be generated or refined by successively modifying one or more 2D or 3D representations of the patient’s dentition (e.g., the post-treatment dentition), or of the patient’s face (e.g., a 2D photo with a randomized mask applied to remove aspects of the face which include the mouth). A partially trained denoising ML model 214 (e.g., a UNet) may be trained, at least in part, using the refined training dataset. In some implementations, other types of hierarchical neural network feature extraction module (HNNFEM) may be substituted for the UNet. The resulting fully trained denoising ML model 320 may be deployed for clinical treatment of patients (e.g., to show the patient a prediction of what they’ll look like after the completion of treatment, and do so in approximately real-time). Through the course of generating the training dataset, one or more representations of the training dataset may be modified, such as through the addition of noise (e.g., Gaussian noise). For example, salt-and-pepper noise may be added to a masked photograph of the patient (or an image of the patient’s post-treatment dentition), resulting in hundreds, thousands or tens or thousands of incrementally noisier images.
[0005] In some implementations, one or more aspects of one or more representations of the training dataset may be encoded into one or more latent representations, and then noise may be added to the one or more latent representations, to produce one or more noisy latent representations. The one or more oral care arguments may include at least one of a real value, a categorical value or a natural language text value. The one or more oral care arguments may include at least one of an oral care metric, or an oral care parameter. The partially trained denoising ML model 214, or fully trained denoising ML model 320 may include one or more neural networks (e.g., an encoder-decoder structure, etc.). An encoder-decoder structure may comprise at least one encoder or at least one decoder. Non-limiting examples of an encoder-decoder structure include a UNet, a transformer, a pyramid encoder-decoder, or an autoencoder, among others. One or more 3D representations of the patient’s dentition (e.g., which may include at least one tooth) may be provided to the denoising diffusion methods of this disclosure. The one or more 3D representations of the patient’s dentition may be encoded into one or more latent representations. The one or more denoised representations (e.g., 3D oral care representations generated by fully trained diffusion ML model 320) may include either a 2D or 3D representation of the patient’s anatomy (e.g., the face) which is combined with a representation of the post-treatment dental anatomy.
[0006] The denoising diffusion techniques described herein may, in some instances, be practiced in combination with other operations in digital oral care. For example, a patient’s dentition may be scanned in a clinical environment (or at a clinical context), resulting in a pre-segmentation mesh of the teeth and gums. The mesh may undergo validation to identify any scanning defects, undergo mesh cleanup to fill holes and correct flaws, and then undergo segmentation. The segmented tooth meshes and corresponding maloccludedtransforms may be provided to an automated orthodontic setup prediction model, which may generate a predicted final setup. The predicted final setup (e.g., comprising 3D tooth meshes which are placed in their final setup poses) may describe the patient’s target dentition 302. The target dentition 302 may subsequently be provided to the smile prediction methods of this disclosure, which may combine the digital representation of the patient’s target dentition 302 (e.g., either 2D or 3D) with a digital representation of the patient’s face 300 (e.g., either 2D or 3D) using fully trained denoising diffusion ML model 320. Other combinations of the techniques described herein should also be considered within the scope of this disclosure.
[0007] Techniques of this disclosure relate to a computer-implemented method for generating predictions of post-treatment dental anatomy for a patient. Representations of the patient's pre-treatment dental anatomy, which include exposed teeth, as well as representations of the desired post-treatment dentition may be provided to the techniques. The techniques may generate noisy first representations of the intended output, which may be denoised by using a trained flow-based machine learning model (e.g., aa denoising diffusion ML model). This process yields second representations of the intended output and automatically defines various aspects of the patient’s post-treatment anatomy. Techniques of this disclosure may optimize camera parameters which may be used to register two or more dentitions, to prepare those dentitions to be provided to ML models of this disclosure. In some implementations, the optimization methods may optimize the alignment of the upper and lower 3D arches of the patient’s dentition (e.g., by optimizing translation or orientation values that control the alignments of the upper and lower arches). In addition to optimizing the alignment of the upper and lower arches, in some implementations, the camera parameters can also be optimized (e.g., using a genetic algorithm, etc.). Either camera parameters, or arch alignments (or both) may be optimized, according to particular implementations.
[0008] The methods further incorporate the use of oral care arguments and camera parameters to enhance the accuracy of the predictions. The techniques include steps for encoding pre-treatment and post-treatment representations of patient dentitions into latent representations, processing 2D images and 3D representations of teeth, and utilizing tooth image masks for refining the predicted shapes of teeth. Additionally, the methods employ optimization techniques to refine camera parameters and cost functions to ensure the precision of the generated data structures (e.g., 3D meshes or 2D images of post-treatment patient appearances). A posttreatment appearance can include aspects of the full face, partial face, and / or dentition of the patient.
[0009] Moreover, the methods may automatically generate orthodontic setups or tooth restoration designs using machine learning models and provides functionality for determining confidence scores to validate tooth identifications. The methods offers a comprehensive and efficient approach to predicting a patient’s posttreatment dental anatomy, thereby improving the quality of dental care and treatment planning.Brief Description of Drawings
[0010] FIG. 1 shows a method for preparing the patient’s face data and the patient’s final setup dentition to be used in training a denoising diffusion probabilistic model for smile prediction.
[0005] FIG. 2 shows a method for training a denoising diffusion probabilistic model to combine a representation of the patient’s post-treatment dental anatomy with a representation of the patient’s face.
[0006] FIG. 3 shows a method for using a fully trained and deployed denoising diffusion probabilistic model to combine a representation of the patient’s post treatment dental anatomy with a representation of the patient’s face.
[0007] FIG. 4 shows a method of generating a color palette which can be used to influence the color of one or more teeth in the predicted smile.
[0008] FIG. 5A-1 shows a method of training a flow-based model to generate a predicted smile (with images that show an example of dental restorative treatment).
[0009] FIG. 5A-2 shows a method of training a flow-based model to generate a predicted smile.
[0010] FIG. 5B show images with an example of predicting a smile using method for orthodontic treatment.
[0011] FIG. 6 shows a method of using a fully trained flow-based model to generate a predicted smile for orthodontic treatment.
[0012] FIG. 7 shows a method of using a fully trained flow-based model to generate a predicted smile for dental restorative treatment.
[0013] FIG. 8 shows a method of training a denoising diffusion ML model.
[0014] FIG. 9 shows a method of using a fully trained denoising diffusion ML model.
[0015] FIG. 10 shows a method of preparing a patient’s dentition for registration.
[0016] FIG. 11 shows a method of preparing a patient’s dentition for registration.
[0017] FIG. 12 shows a method of segmenting a patient’s dentition.
[0018] FIG. 13 shows a method of generating a smile mask.
[0019] FIG. 14 shows a method of applying smile mask to a patient’s dentition.
[0020] FIG. 15 shows a method of applying smile mask to a patient’s dentition.
[0021] FIG. 16 shows a method of applying smile mask to a patient’s dentition.
[0022] FIG. 17 shows a 2D image of a patient’s smile, and an associated segmentation mask.
[0023] FIG. 18 shows a 3D mesh of the patient’s dentition, a smile mask and a 2D projection of the patient’s dentition.
[0024] FIG. 19 shows subset mask images that were generated using initial or optimized camera parameters.
[0025] FIG. 20 shows 2D dentition masks.
[0026] FIG. 21 shows patient dentitions before and after registration is performed.
[0027] FIG. 22 shows a method of computing the fitness of a population member in a genetic algorithm.Detailed Description
[0028] Diffusion models may be applied to 2D image generation, or the generation of 3D representations, among others. In some instances, such implementations may take input from natural language text (e.g., natural language text, real values, categorical values, reference images or reference 3D representations which describe an intended outcome or a post-treatment anatomy). The techniques described herein expand denoising diffusion models into the digital oral care space. The techniques described herein use a denoising diffusion model to combine a patient’s post-treatment dental anatomy with a representation of the patient’s face. The combining may be conditioned, at least in part, on provided oral care arguments (e.g., which may include natural language text, integer arguments, real-valued arguments, categorical arguments, and the like). Such oral care arguments may contain one or more attributes describing an intended output from a trained machine learning model. Techniques described herein may, in some instances, take as input representations of the patient’s post-treatment dentition (e.g., 3D point clouds, or 2D views of 3D representations), which are to be used as guides for the generation of one or more post-treatment renderings of the patient’s face. The renderings can be used in treatment planning by clinicians and can help patients decide what kind of treatment they’d like to receive. Techniques described herein may, in some instances, take as input representations of the patient’s face, or of the patient’s post-treatment target dentition. In some instances, these inputs may be encoded into latent representations by autoencoders or by other encoder-decoder neural network structures (e.g., a representation generation module). A latent representation may include an information-rich and / or reduced-dimensionality form of the original data. In some instances, techniques of this disclosure may realize improved data precision through the use of such latent representations. For example, denoising diffusion ML models of this disclosure (e.g., model 214, model 320, model 522, model 608, and model 708, etc.) may generate outputs with improved accuracy when the models are trained on the latent representations, because the small sizes of the latent representations make the latent representations easier for the neural networks to encode. Stated another way, the denoising diffusion ML models are better able to learn the distribution of the data within the latent representations because of the smaller size of the latent representations (e.g., relative to the size of the data in the data’s original form).
[0029] Method 132 in FIG. 1 describes the preparation of tuples of patient data (with optional augmentation) which may be used by method 222 in FIG. 2 to train a denoising diffusion probabilistic model 214. FIG. 3 shows method 326 which uses a fully trained denoising ML model 320to combine a photo (or mesh) of the patient’s face with a digital representation of a predicted post-treatment dentition (e.g., to generate one or more predicted smiles).
[0030] The training of a partially trained denoising diffusion model 214 may involve a forward pass 220 over the input data. The input data may include text-based oral care arguments 200 (e.g., instructions for clinicians), non-text oral care arguments 202, or one or more tuples 204 (e.g., a photo of the patient’s face, amasked photo of the patient’s face, and / or a latent representation of the patient’s target dentition, among other inputs described herein). In some implementation, optional augmentation may be applied to one or more fields of the tuple 204, wherein the tuple may contain (e.g., an augmented photo of the patient’s face, an augmented masked photo of the patient’s face, or a latent representation of the augmented patient’s dentition). The forward pass 220 may add small amounts of noise to the tuple data 204, generating a Markov chain of steps 216. The inputs 200, 202, or 204 may be combined (212) and then provided to the partially trained denoising ML model 214, to condition the model 214 on the patient’s data and / or the treatment instructions. Text-based oral care arguments 200, or non-text oral care argument 202 may include treatment instructions from clinicians. The forward pass 220 may generate training data. Stated another way, the Markov chain may comprise a set of successively noisier training data examples. Partially trained denoising diffusion ML model 214 may be trained generate noise tensors which may be subtracted from inputs, in order to perform the denoising operation. Input oral care arguments 202 may influence the functioning of the denoising diffusion models 214 or 320, causing the denoising diffusion models 214 or 320 to generate output to the specification of the clinician (e.g., enabling the customization of the output that is generated by the denoising diffusion models 214 or 320). Oral care arguments 202 may include oral care parameters, oral care metrics, among other examples described herein. Oral care argument 202 may include: categorical information such as patient age or gender, or other embeddings containing medical data (e.g., pertaining to diagnoses, etc.). Text-based oral care arguments 200 may contain natural / colloquial language embeddings. In some implementations, such as with stable diffusion, the masked representation of the patient’s face may be encoded (210) into latent form. A latent encoding module may be trained to encode data into a reduced-dimensionality latent form. Examples of data which may be encoded include one or more 3D point clouds or 3D meshes describing the patient’s dentition, 2D photos of the patient’s face, 3D representations of the patient’s face, or 2D renderings of the patient’s dentition, among others. Such data may be encoded into a latent vector or latent capsule. Oral care arguments 200 may be encoded (206) in latent representations. Oral care arguments 202 may be encoded (208) into latent representations. The inputs 200, 202 or 204 (or latent representations thereof) may be combined (212) and then be used to condition the output of the partially trained denoising ML model 214. The one or more noisy representations of the patient’s face (e.g., coming from the Markov chain 216) may be provided to the partially trained denoising diffusion ML model 214, which may generate a predicted representation of the patient’s post-treatment appearance. Loss may be computed (218) between the predicted representation and a corresponding ground truth representation (e.g., the original photo of the patient, which is provided as a part of the tuple 204). The loss may be used to further train, at least in part, the partially trained denoising diffusion ML model 214. In some implementations, partially trained denoising ML model 214 may generate one or more predicted noise tensors. The one or more predicted noise tensors may be subtracted from the inputs. Stated another way, the predicted noise tensors may be removed for the inputs, or the inputs may be denoised. In some implementations, loss may be computed as the differencebetween the predicted noise tensors and corresponding ground truth noise tensors which are provided by the Markov chain 216.
[0031] The Markov chain 216 may generate a succession of increasingly noisy versions of the input data 204, which may then be used to train, at least in part, a denoising ML module 214 (e.g., which may be used in deployment as a part of the reverse pass 322). The denoising ML model 322 which is trained for use in reverse pass 322 may include one or more neural networks (e.g., U-Net, VAE, 3D SWIN transformer, pyramid encoderdecoder, or the like) and may be trained to denoise a highly noisy version of the data from the Markov chain 318. For example, the denoising ML model 320 may iteratively remove noise from a randomized data structure (e.g., a 2D image containing Gaussian noise or a 3D point cloud or mesh with randomized mesh elements). Stated another way, the initially noisy data structure may be denoised by fully trained denoising diffusion ML model 320 until the data structure converges on a final state, and is output as a predicted smile 324 (e.g., a predicted post-treatment photo of the patient). In deployment, a pre-treatment 2D photo (or 3D representation) 300 of the patient’s face may undergo latent encoding (308) and be provided to a combination module 316. A digital representation (either 3D representation or 2D representation) of the patient’s predicted post-treatment target dentition 302 (e.g., tooth meshes arranged in a final setup, or the like) may undergo latent encoding (310) or (118), and then be provided to combination module 316. Text-based oral care arguments 304 may undergo latent encoding (312), and be provided to combination module 316. Likewise, non-text based oral care arguments 306 may undergo latent encoding (316) and be provided to the combination module 316. The combination module may perform concatenations or additive operations on its inputs and provide its output to the fully trained denoising ML model 320, to condition that model on the patient’s data and / or the treatment instructions from clinicians.
[0032] In some implementations, a loss function (e.g., cross-entropy or mean squared error (MSE), among others) may be computed (218) to quantify the differences between a generated 3D oral care representation and a corresponding ground truth (or reference) 3D oral care representation. In some implementations, the loss function may be used to train, at least in part, the denoising ML model 214.
[0033] Stated another way, in some instances, the forward pass 220 of the denoising diffusion model may generate a training dataset of increasingly noisy examples of the input data. Noise may be introduced to disfigure the input data (e.g., 2D or 3D representations of the patient’s face and / or post-treatment dentition, etc.) and those noisy examples may be used, at least in part, to train a denoising diffusion machine learning model 214 to reverse of this noise-introducing process (e.g., in model deployment). The reverse pass 322 may be executed to reconstruct the pristine input data by removing the noise from a noisy example of that input data. After the denoising diffusion model (e.g., a U-Net or other HNNFEM) is trained, the denoising diffusion model is capable to generate new 3D oral care representations (e.g., a photo of the patient with post-treatment dentition, among others) by passing a noisy data example (e.g., a photo of the patient with the mouth area masked-out) through the denoising diffusion process 320 (aka the reverse process 322). Examples of hierarchical neural networkfeature extraction modules (HNNFEM) include 3D SWIN Transformer architectures, U-Nets or pyramid encoder-decoders, among others. A HNNFEM may be trained to generate multi-scale voxel (or point) embeddings of a 3D representation (or multi-scale embeddings of other mesh elements described herein). For example, a HNNFEM of one or more layers (or levels) may be trained on 3D representations of patient dentitions to generate neural network feature embeddings which encompass global, intermediate or local aspects of the 3D representation of the patient’s dentition.
[0034] Techniques of this disclosure may require a training dataset of hundreds or thousands of cohort patient cases, to ensure that the neural network is able to encode the distribution of patient cases which are likely to be encountered in clinical treatment. A cohort patient case may include a set of tooth crown meshes, a set of tooth root meshes, a photograph of the patient (e.g., with pre-treatment dentition), a representation of a posttreatment predicted dentition, or a data file containing attributes of the case (e.g., a JSON file). In some implementations, the post-treatment dentition (e.g., a final setup) can be computed automatically in nearly realtime while the patient waits in the clinician’s office. A typical example of a cohort patient case may contain up to 32 crown meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), multiple gingiva mesh (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces) or one or more JSON files which may each contain tens of thousands of values (e.g., objects, arrays, strings, real values, Boolean values or Null values).
[0035] FIG. 1 describes a method to prepare tuples 130 for use in training the denoising ML model 214. The patient’s pre-treatment face photo 100 (or a 3D representation of the patient’s face - with optional color data associated with the mesh elements) may be provided as input data. Facial landmarks (e.g., corresponding to the mouth or lower face) may be computed and one or more masks 104 may be generated (102) from the landmarks. The one or masks 104 and the face photo 100 may be augmented (106). Augmentation (106) may include one or more of the following operations on either 2D or 3D data: flips, warps, rotations, introduction of Gaussian noise, changing colors, among other operations. M augmentations may be generated (116). The augmented masks may be applied (108) to corresponding augmented versions of the patient’s face photo 100. The resulting M masked and (augmented) photos may be provided to the collection of M output tuples 130. The original (nonmasked) face photo 100 may undergo augmentation (110) and be provided to the collection of M output tuples 130. The 3D representation of the patient’s post-treatment target dentition 112 (e.g., a final setup, etc.) and (optional) oral care arguments 114 may be augmented (116), which may generate M clinically plausible augmentations of the patient’ s dentition. The present disclosure may augment the dentition of the patient in ways that are biologically plausible, and so fall within the distribution of the training dataset of cohort patient case data. In some implementations, modifications to the 3D tooth meshes can be performed using an encoder-decoder structure and a latent representation modification module (LRMM). For example, the encoder-decoder structure working in conjunction with an LRMM can make the teeth wider / narrower, longer / shorter, change color, squarecomers / rounded comers, or vary some other restoration design metrics or orthodontic metrics. In some implementations, the relative poses of the teeth (e.g., poses of the teeth relative to each other) can be varied. Furthermore, the positions or orientations of the teeth may be jittered (e.g., undergo small random changes). The M clinically plausible augmentations of the patient’s post-treatment dentition (e.g., final setup or restored tooth meshes) may undergo latent encoding (118). The Dentition Latent Representation Generation Module 118 may provide its one or more generated latent representations to the collection of M output tuples 130. The module 118 may render the patient’s 3D dentition into 2D representations (120) and then encode (128) those images into latent representations, for example, using a CLIP transformer encoder 122. In some implementations, the latent representations may be reconstructed using an MLP or encoder 124. In some implementations, the latent representations may be provided to the collection of M output tuples 130. In some implementations, the patient’s 3D dentition may be provided to a 3D representation generation module 126 (e.g., a U-Net, or others described herein), which may generate one or more latent representations which include hierarchical neural network features of the dentition. The resulting latent representation may be provided to the collection of M output tuples 130. Method 400 may generate one or more data structures which describe the color and / or surface texture of one or more teeth of the patient's dentition 404 (e.g., one or more color palettes 416). The patient's dentition 404 may be described by one or more 2D images (e.g., photographs or 2D renderings of 3D representations), or one or more 3D representations (e.g., 3D meshes, 3D surfaces, 3D point clouds, etc.). The patient's dentition (e.g., 2D image or 3D meshes) may be segmented (410), which may result in one or more segmented image masks (or mesh element labels) 414. In some implementations, oral care arguments 402 may designate one or more of the segmented teeth for inclusion in the color palette generation. The method 400 may generate (406) an initial color palette 408 (e.g., by downsampling a 2D image of the patient's dentition 404). The initial color palette 408 may, in some implementations, undergo modification (412). For example, the color palette pixels of the initial color palette 408 which correspond to the one or more segmentation masks 414 may undergo modification (e.g., increase or decrease) in whiteness, brightness, hue, saturation, to name a few attributes. Color spaces can include RGB, HSV, LAB, or the like. When the patient's dentition includes color-bearing 3D meshes (e.g., one or more 3D meshes whose individual mesh elements bear colors), the initial color palette 408 may be generated (406) by downsampling or averaging the colors of mesh elements within connected components (or within a threshold distance of each other in the mesh), among other methods. The segmented mesh element labels 414 may be used to designate one or more teeth or other aspects of the patient's dentition 404 for color palette generation. The colors of mesh elements which are designed for color palette generation may be modified (412), as described herein. The method 400 may output one or more modified color palettes 416, which may provided to techniques of this disclosure to influence the colors and / or textures of predicted smiles (e.g., predicted smile 530). In some implementations, oral care arguments 402 may influence the generation of a color palette. For example, oral care arguments 402 may designate one or more teeth for processing or analysis, according to techniques of this disclosure.
[0036] Method 500 may train a smile prediction ML to generate one or more predicted smiles 530. A predicted smile 530 may show the patient's target dentition 542 (e.g., a lateral incisor with a clinically desired shape, among other examples) integrated into the patient's original dentition data 508. The patient's original dentition data 508 may include 3D tooth meshes, 3D meshes of gums, 2D photographs of teeth (or gums), 2D renderings of 3D mesh data, or the like. The smile prediction ML model may use one or more diffusion conditioning adapters 534, as shown in diffusion conditioning module 526, to pre-process the patient’s dentition data before the dentition data are provided to denoising diffusion ML model (e.g., to perform dentition registration as described herein). Dentition registration is a digital processing operation that may align two or more dentitions, so that differences between the dentitions can be quantified and used in a cost function or loss calculation. For instance, in some implementations, registration may involve comparing aspects of a representation (e.g., 2D images, 3D models, or other representation) of a first dentition with a representation of a second dentition to determine whether the first and second dentitions are related in some way. Method 500 may include the training of a denoising diffusion ML model 522. Further detail on the training of the denoising diffusion ML model 522 is described by method 800. Method 900 describes the operational use of a fully trained denoising diffusion ML model 522 (e.g., when the denoising diffusion ML model 522 is deployed as module 608 in method 600, or as module 708 in method 700, or the like).
[0037] The method 500 may train a smile prediction ML model to generate one or more smile predictions 530 which have teeth whose poses or appearances that are influenced by one or more oral care arguments, one or more color palettes, or one or more target dentitions 542. Target dentition 542 may include 3D meshes (e.g., containing colors and / or textures, or geometric shape information), or 2D images (e.g., containing 2D image masks or detected edges with defined tooth shapes, among others) that illustrate intended modifications to make to the patients dentition. The method 500 may generate one or more predictions of the patient's full face, one or more predictions of the patient's mouth, or may include other parts of the patient's anatomy.
[0038] During model training, an 'area of interest' 512 is specified or generated (504) which circumscribes a portion of dental anatomy which is to be in-painted into a segmentation region 540 (e.g., which is described by mesh element labels in 3D or image masks in 2D). One goal of training a smile prediction ML model using training method 500 is to train the smile prediction ML model how to in-paint arbitrary objects into masked regions (e.g., to in-paint aspects of target dentition 542 into a 2D or 3D representation of the patient's mouth or the patient's face), such as the masked region shown in masked dentition 506. In some implementations, one or more teeth of the patient's original dentition 508 may undergo segmentation (510), which may generate 2D image masks (or 3D mesh element labels) 540, which may then be applied to the patient's dentition 508, resulting in masked dentition 506. In some implementations, the one or more 2D image masks (or 3D mesh element labels) 540 may undergo augmentation (528), for the purpose of training denoising diffusion ML model 522 to respond to differently shaped image masks (e.g., to generate a custom smile containing realistic-looking teeth whose shapes and / or poses have been influenced by the provided image masks). For example, when the patient’s targetdentition 712 includes one or more target image masks that define the target shapes for one or more treated teeth, the one or more target image masks may be provided to the fully trained denoising diffusion ML model 708 which can be used as a reference representation for the model 708 to generate a realistic predicted smile in which the one or more treated teeth have assumed the shapes defined by the respective one or more target image masks. The resulting teeth have color, transparency and / or specular reflections which look realistic (e.g., look consistent with the other teeth of the patient’s dentition). This method of influencing the shapes (or poses) of teeth in the predicted smile is generally used to generate teeth with clinically and / or aesthetically desired shapes (or poses) which are biologically plausible. In some instances, such as for artistic or entertainment purposes, target tooth masks which are not biologically typical may be provided to denoising diffusion ML model 708. For example, target image masks may be generated for the left and right upper cuspids which have exaggerated lengths and / or exaggerated pointiness, so as to influence the denoising diffusion ML model 708 to generate a predicted smile for a human subject where the human subject is given the teeth of a vampire or other fanciful creature (e.g., for use in a work of art, or in a work of entertainment such as a film).
[0039] In some implementations, method 500 may train a smile prediction ML model to in-paint an image of a target tooth anatomy 542 into a photograph of a patient's face in a manner that is aesthetically pleasing and / or clinically plausible (e.g., the colors, reflections, and / or shapes of the teeth are realistic, etc.). Examples of target tooth anatomy 542 include post-restoration teeth or post-treatment final setups for orthodontics.
[0040] When the fully trained smile prediction ML model is deployed (e.g., as module 608 or 708, etc.), the 'area of interest' 512 may be used to define the portion of a post-treatment target dentition 542 that is to be realistically integrated with a photograph of the patient's face, where the patient's face photo initially shows the pre-treatment dentition. In this example of model deployment, a segmentation image mask 540 is generated (510) which designates a portion of the patient's pre-treatment photo 508 into which the target dentition image 542 is to be in-painted or otherwise realistically integrated by the fully trained smile prediction ML model. When the smile prediction ML model is being trained, the 'area of interest' 512 may be used to define a portion of the patient's original dentition 508 which is to be provided to the denoising diffusion ML model 522, as a substitute for the patient's post-treatment dentition. During training, the masked dentition 506 may be provided to the denoising diffusion ML model 522 to train the denoising diffusion ML model 522 to learn the distribution of patient dentitions 508.
[0041] In some 2D implementations, diffusion conditioning adapter 534 may process the patient’s dentition data (e.g., 3D or 2D data), to clarify and / or strengthen the signal in those data, and / or prepare those data to be provided to the denoising diffusion ML model 522. For example, when the patient’s dentition includes 2D image data, diffusion conditioning adapter 534 may generate a 2D image containing at an outline of one or more target tooth shapes (e.g., generated by applying an edge detector, such as the Canny edge detector or Sobel edge detector, and then optionally modifying the contours of those detected edges). In some implementations, one or more image masks which describe the target poses or target shapes of one or more teeth may be provided todenoising diffusion ML model 522, and may control the resulting shapes or poses of respective one or more teeth in the predicted smile 530. In some implementations, diffusion conditioning module 534 may perform method 1100 to generate one or more dentition masks 1114. The one or more dentition masks 1114 may contain one or more registered projected 2D dentition mask images (i.e., one or more images that have undergone a registration processes defined herein), which may contain connected components. Each connected component may describe the target shapes and / or target poses for one or more teeth. Denoising diffusion ML model 522 may denoise the connected components, to transform those flat shapes into realistic-looking teeth with colors, textures, specular reflections, or shadows that are consistent with other teeth of the patient’s dentition, and / or are consistent with the distribution of teeth in the training dataset. The output of diffusion conditioning adapter 534 may be provided to encoder 536 to encode the output (e.g., 2D image masks, 3D representations, etc.) into one or more latent representations, and then the one or more latent representations may be provided to the denoising diffusion ML model 522 (e.g., or another flow-based ML model). In some implementations, diffusion conditioning adapter 534 may compute edges, perform registration or perform image sharpening (or perform other operations described herein), the results of which may be provided to encoder 536. The denoising diffusion ML model 522 may then generate a predicted smile 530, where one or more teeth have the shapes of the one or more tooth image masks (or tooth outlines). Stated another way, the shapes of one or more teeth in the predicted smile 530 may be influenced by the contours of the connected components in an image mask. The image masks (or edge- detected image outlines) may specify the intended shapes of one or more teeth. The method 500 may generate a predicted smile in which one or more treated teeth have the color, translucency, specular reflections, and / or texture of the other teeth of the patient (e.g., as specified by color palette 532). A color palette 532 may be generated by module 538 (e.g., according to method 400), and then be provided to diffusion conditioning module 526. The color palette 532 may specify color modifications, or modification to the whiteness of one or more teeth in the original patient dentition data 508.
[0042] In some 3D implementations, target patient dentition data 542 may include one or more 3D meshes (or other 3D representation described herein), each of which has the intended target shape (or color, texture, or pose, etc.) of a tooth. Such target patient dentition data 542 may be provided to conditioning adapter 534, which may process the patient's dentition to make the data cleaner or otherwise improve the efficient use of the dentition data by the denoising diffusion ML model 522 (e.g., by performing segmentation, mesh cleanup, downsampling, upsampling, registration between current and target dentitions, or the like). In some implementations, the target patient dentition data 542 may be provided directly to encoder 536, which may generate one or more latent representations (e.g., embedding vectors, latent vectors, latent capsules, or the like). In some instances, the 3D mesh of the patient's dentition data 508 or the target dentition data 542 (e.g., one or more tooth meshes) may have color, and / or texture on the surface of the 3D mesh. The 3D mesh of the tooth may be provided to conditioning adapter 534, and then be encoded into one or more latent representations by encoder 536. The one or more latent representations may then be provided to concatenation module 524, or be directly provided todenoising diffusion ML model 522. Denoising diffusion ML model 522 may then generate one or more smile predictions 530 (e.g. 2D or 3D predictions). When a 2D prediction is generated, the prediction may, for example, comprise a 2D image of the patient's smile where the specified one or more teeth (e.g. one or more teeth specified by oral care arguments 502) has assumed the poses or appearances specified by the corresponding teeth in target dentition 542. In some implementations, any aspects of the patients teeth (e.g., the 3D geometric information of a 3D tooth mesh, the tooth's surface color and / or texture) may be processed by the same conditioning adapter 534, and / or subsequently encoded by one or more encoders 536. In other implementations, aspects of the patients teeth (e.g., the 3D geometric information of a 3D tooth mesh, the tooth's surface color and / or texture) may be processed by separate conditioning adapters, and / or subsequently encoded by one or more encoders 536. In some implementations, a conditioning adapter 534, encoder 536, and / or a denoising diffusion ML model 522 may be trained end-to-end, so that the models are trained concurrently, sometimes using the same one or more loss functions. Other unrolled or interconnected ML models of this disclosure may also be trained in an end-to-end manner.
[0043] Various inputs (e.g., tooth transforms, 2D images, 3D meshes of crowns, 3D meshes of roots, or color palettes, or other information) pertaining to the current patient dentition data 508, target patient dentition 542, or oral care arguments 502 (e.g., which pertain to the digital oral care of the patient) may be provided to a diffusion conditioning adapter 534 (e.g., T2I adapters, among others) and then subsequently be encoded by encoder 536. Diffusion conditioning module 526 may contain one or more pairs of diffusion conditioning adapter 534 and encoder 536. In implementations when the original patient dentition 508 or target patient dentition 542 contain 2D images, edges may be generated using, for example, a Canny edge detector, a Sobel edge detector, or other another type of edge detector. The edges may reveal the contours of teeth, gums, lips, or other aspects of the patient's dentition, mouth or facial appearance. In some implementations, the edges (or 2D image masks) that describe the contours of one or more teeth may be modified to define new target shapes. The images which are generated by edge detection may be provided to aggregation module 524, which may aggregate or concatenate one or more latent representations. The concatenated latent representations may then be provided to denoising diffusion ML model 522. The denoising diffusion ML model 522 can, in some implementations, be replaced with other flow-based ML models, such as continuous normalizing flows. In some implementations, the flow-based models of this disclosure may be trained, at least in part, using flow matching. When denoising diffusion is used, denoising diffusion ML model 522 may include an encoder-decoder structure (e.g., a U-Net, Vision Transformer (ViT), or others described herein). In some implementations, other image processing filters (e.g., image sharpening, or the like) may be applied by diffusion conditioning adapters 534 to images of the patient's dentition data 508 to clarify aspects of the images and provide enhanced information about the shape and / or structure of the patient's anatomy to the denoising diffusion ML model 522. A non-limiting example of image sharpening is Unsharp Masking (USM). In some implementations, other filters may be applied by diffusion conditioning adapter 534, such as Bas Relief or Chalk & Charcoal.
[0044] Denoising diffusion ML model 522 may output a predicted smile 530, which may include modified patient dentition data integrated into a 2D or 3D representation of the patient’s face. For example, predicted smile 530 may comprise a version of the original patient dentition data which has undergone tooth whitening, tooth color alteration, or tooth texture alteration, based upon information in the color palette. In a further example, modified patient dentition data 530 may comprise a version of the original patient dentition data 508 which has undergone tooth shape modification (e.g., lengthening or shortening of one or more teeth). In yet other examples, modified patient dentition data 530 may comprise a version of the original patient dentition data 508 which has undergone tooth pose modification (e.g., as specified by an orthodontic setup included in target patient data 542).
[0045] Denoising diffusion ML model 522 may be trained, at least, in part using augmented training data. For example, an area of interest 512 of the original patient dentition data 508 may be generated or specified (504). The area of interest 512 of the dentition data may undergo (optional) augmentation (514), which results in augmented dentition data 516. The area of interest 512 or the augmented data 516 may be provided to encoder 518 (e.g., a CLIP encoder, or the like), which may generate one or more latent representations, which may, in some implementations, be provided to neural network 520 (e.g., a multilayer perceptron, among others), or directly to denoising diffusion ML model 522. The neural network 520 may, in some implementations, change the shapes of the one or more latent representations which are generated by encoder 518, so that the one or more latent representations are properly formatted to be provided to the stable diffusion process in denoising diffusion ML module 522. The output of the neural network 520 may be provided to denoising diffusion ML model 522, for the purpose of inducing denoising diffusion ML model 522 to learn how to in-paint arbitrarily shaped teeth into a masked region (e.g., masked using either a 2D image mask or using mesh element labels in a 3D mesh). An example of a masked region is shown in masked dentition 506.
[0046] In a non-limiting example, the training method 500 shows a pre-restoration tooth (e.g., a peg lateral) within the area of interest 512. After the fully trained model is deployed, the area of interest 512 would be configured to circumscribe (or otherwise designate) one or more teeth which describe the target post-restoration appearances of the patient's teeth. After deployment, the fully trained model has been trained (via training method 500) to in-paint one or more teeth of target dentition 542 (e.g., with clinically or aesthetically idealized aspects, such as target shapes, colors, textures, or whitening) into a 2D image or 3D mesh representation of the patient's face (or mouth).
[0047] In some implementations, the patient dentition data 508 may be segmented (510). The outputs 540 of segmentation may include one or more image masks (when patient dentition data 508 includes 2D images), or one or more mesh element labels (when the patient dentition data 508 includes 3D meshes). In some implementations, the image masks (or mesh element labels) may be augmented (528), and then be applied to the original patient dentition data 508, resulting in masked dentition data 506, which may then be provided to the denoising diffusion ML model 522. The purpose of the masking is to define the one or more regions of the patient's original dentition data 508 into which the post-treatment dentition 542 is to be integrated. In someinstances, original dentition data 508 can include 2D images of the patient's face and / or teeth. In some instances, original dentition data 508 can include 3D representations (e.g., 3D meshes, etc.) of the patient's teeth and / or face. According to various implementations, tooth image masks (or mesh element labels) 540, augmented tooth image masks (or augmented mesh element labels) 528, masked dentition 506, and / or the original patient dentition data 508 may be provided to the input of the denoising diffusion ML model 522. Method 800 of FIG. 8 shows additional detail for the training of the denoising diffusion ML model 522. Method 900 describes the functioning of a fully trained denoising diffusion ML model 608 or 708 in deployment.
[0048] Denoising diffusion ML model 522 may be trained, at least, in part through the calculation (826) of one or more loss values. For example, a loss may be computed (826) that quantifies the difference between a predicted noise tensor 824 and a ground truth noise tensor 816. The ground truth (or actual) noise tensor 816 may, in some implementations, be computed (814) using a pseudorandom number generator to generate Gaussian noise values (or noise values which are drawn from other distributions). The ground truth noise tensor is added to latent representation 812, and the resulting latent representation is provided to encoder-decoder structure 822. Encoder-decoder structure 822 may then determine which part of the input is noise, and / or which part of the input is a meaningful signal (e.g., which part of the input pertains to the patient’s dentition). Encoder-decoder structure 822 may then output a predicted noise tensor 824. The example in method 800 shows how stable diffusion has been customized to the task of predicting smiles. In stable diffusion the inputs undergo latent encoding (e.g., using encoder 518 or 536), and then are provided to denoising diffusion ML model 522. In some implementations, a Markov chain of increasingly noisy examples of the input data may be generated and used to train the denoising diffusion ML model 522 in how to perform the denoising operation. During each successive step in the Markov chain, more noise is added to the inputs. The inputs are then provided to encoder-decoder structure 822, which may predict a data structure describing the noise that was just added (e.g., a predicted noise tensor 824). Method 800 provides additional detail on the training on denoising diffusion ML model 522.
[0049] Oral care arguments 502 (as described herein) may, in some implementations, be provided to the method 500 to influence or customize the outputs 530. Oral care arguments 502 may include, for example, specifications of which teeth to segment (510), specifications of which teeth to treat, or specifications of which teeth are static or pontic (e.g. and may therefore be designated to not move during orthodontic treatment), among others. In some implementations, the oral care arguments 502 may specify one or more teeth which are to undergo color modification, texture modification, and / or specify the nature of the whitening (e.g., to increase or to decrease the lightness of one or more teeth). Stated another way, in some implementations, oral care arguments 502 may specify which teeth are to be whitened (or otherwise processed), the magnitude of whitening which is to be applied, or which other types of augmentations are to be performed. In some implementations, oral care argument 502 may specify changes to the shapes of one or more teeth (e.g., such as increasing or decreasing tooth crown length). In some implementations, oral care argument 502 may include one or more 2D image masks which may be used to define new shapes for one or more teeth. In some implementations, the oralcare arguments 502 may specify the magnitude of change in length of one or more teeth (e.g., to shorten the lateral incisors, or to lengthen the cuspids, to name a couple examples). Oral care arguments 502 may specify one or more oral care metrics (e.g., Arch Symmetry, Proportions of Adjacent Teeth, or others described herein), which may influence the denoising diffusion ML model 522 (or 608 or 708) to generate one or more smile predictions in which the patient's teeth show shapes and / or poses which are influenced by the one or more oral care metrics. In examples pertaining to smile prediction in 2D images, one or more image masks may be defined (e.g., via segmentation).
[0050] The one or more image masks may correspond to one or more teeth which are to be modified (e.g., lengthened, or shortened, or otherwise modified in shape or pose). An image mask may include pixels which has the value of zero (e.g., which are to be ignored), and / or pixels which have non-zero values (e.g., which are to be processed). The one or more masks may themselves be augmented (528) (e.g., by lengthening or shortening each of the one or more masks - according to the desired change in the patient's dentition), and then the one or more augmented masks 506 be provided to denoising diffusion ML model 522 to influence the model 522 to generate a predicted smile 530 in which the one or more teeth corresponding to the one or more masks are lengthened or shortened (or otherwise altered in shape) according to the one or more augmented masks 506. The resulting altered teeth look realistic, including color, shading and / or reflections which are consistent with other teeth shown in the patient's dentition 624 or 724. In some implementations, oral care arguments 502 may specify a change in length or otherwise a change in shape (e.g., as defined by an image mask, restoration design parameters, or restoration design metrics, etc.) of one or more teeth in patient dentition data 508. The oral care arguments 502 may induce denoising diffusion ML model 522 to generate a 2D predicted smile which contains the specified tooth modifications. In other implementations, the predicted smile 530 may comprise 3D data, such as when the original patient dentition data 508, and / or target patient dentition data 542 comprise 3D mesh data. In some implementations, oral care arguments 502 may specify a change in pose (e.g., as defined by orthodontic procedure parameters, or orthodontic metrics, etc.) of one or more teeth in patient dentition data 508, which may induce denoising diffusion ML model 522 to generate a predicted smile 530 which contains the specified tooth modifications. In some implementations, original patient dentition data 508 may contain 2D images or 3D meshes of the patient's face or other parts of anatomy. In some implementations, oral care arguments 502 may undergo latent encoding (556), and the resulting one or more latent representations may be provided to concatenation module 524, which may provide its output to the denoising diffusion ML model 522.
[0051] In addition to other descriptions found herein, oral care arguments 502 may: 1) specify which one or more teeth are to be treated; 2) specify one or more color palettes to be used to alter the color of one or more teeth; 3) specify an increase or decrease in whiteness to apply to one or more teeth; 4) specify an amount of staining that should be removed from one or more teeth; 5) specify a diastema which is to be generated (or modified) between two adjacent teeth; 6) specify a change to the cant or pose of a tooth (e.g., specified in units or degrees or radians of rotation about one or more axes of the tooth); and 7) specify the distance one or moreteeth are to be intruded or extruded (e.g., a distance in mm by which the upper canines are to be extruded for the final setup, etc.), among others described herein.
[0052] In some implementations, method 500 can train a flow-based smile prediction ML model which can predict the outcome of dental restoration treatment (e.g., as seen in FIG. 5A-1 and 5A-2), orthodontic treatment (e.g., visualizations for which are shown in FIG. 5B), or other types of oral care treatments as well (e.g., surgical reconstruction, etc.). Orthodontic treatment may include bracket-based treatment (e.g., through the use of indirect bonding, such as the Solventum Digital Bonding Tray), or aligner-based treatment (e.g., through the use of CLARITY Aligners), among other examples. Dental restoration treatment may include the use of dental restoration appliances (e.g., FILTER Matrix), crowns, bridges, inlays or onlays, fillings, dental implants, dentures, among others.
[0053] Whereas method 500 can be used to train flow-based smile prediction models for orthodontics, dental restoration, or other treatments, the illustration of the patient's teeth 504 in FIG. 5 A- 1 shows a non-limiting example where an area of interest is formed around a "peg lateral" tooth, which is designated as the target of dental restorative treatment. FIG. 5B shows non-limiting examples of data from method 500 when the flowbased smile prediction ML model is trained to predict the outcome of orthodontic treatment. Example 544 shows the original patient dentition data 508 (either 2D image or 3D mesh of the patient's mouth or full face) where the patient's teeth are maloccluded and in need of orthodontic treatment. Example 546 shows an image mask 540 that was generated by segmentation of the patient's face (e.g., segmentation to identify the region or area between the patient's lips). Example 548 shows the patient dentition data with mask applied 506. Example 550 shows the patient's dentition data area of interest 512 (e.g., the portion of the original patient dentition data 508 which lies underneath the image mask 540). Example 552 shows the registered output (as described herein) of an example conditioning module 534 which has processed the target dentition 542 to clarify or enhance aspects of the target dentition 542 (e.g., by performing image sharpening, edge detection, edge modification, segmentation, mask generation, mask modification, image noise removal, or the like). Registration in this context means to to align one or more dentitions with one or more other dentitions. Alignment may be determined in a number of ways. For example, alingment may be determined by comparing certain aspects (e.g., pixels) of a first image with respective aspects of a second image to determine whether those the aspects of the first and second images are substantially similar. As another example, when the target dentition 542 contains 3D mesh data, registration (as described herein) may be performed / alignment may be determined by diffusion conditioning adapter 534 so that the 3D target dentition looks natural (e.g., contains few visual artifacts) when projected into a 2D plane, for integration into the predicted smile 530. Example 554 show the completion of smile prediction for orthodontic treatment, where the patient's predicted smile 530 shows the target dentition 542 realistically integrated into the region between the patient's lips in the patient's original dentition data 508.
[0054] The method 600 may use a fully trained smile prediction ML model that operates using conditioned diffusion. In general, conditioned diffusion is conditioned on 3D mesh data of the patient's dentition or 2D imagedata of the patient's dentition, or other examples of data described herein. The method 600 may use an ML model that was trained using methods 500 and / or 800 to generate smile predictions for orthodontic treatment in real-time, such as to show the patient waiting in the treatment chair what their teeth, mouth, and / or other aspects of the patient’s face will look like at the completion of orthodontic treatment (e.g., the appearance of the patient with a final setup), or to show the patient their appearance mid-treatment (e.g., the appearance of the patient during an intermediate stage of orthodontic treatment). The method 600 may generate a full-face prediction, or a full mouth cavity prediction that realistically shows the patient's face, lips, and / or dentition (e.g., mouth, gums and / or teeth), where the predicted target dentition 612 is realistically integrated into the cavity between the parted lips.
[0055] The predicted smile 610 shows the target dentition 612 registered and integrated within the lips and surrounding facial structure in a manner that looks realistic (e.g., is free of shape, color, or texture artifacts that might tip-off the viewer that the image was artificially generated). The patient's target dentition data 612 and / or the patient’s dentition (or facial) data 624 may include 3D data or 2D data. Examples of 3D data include 3D tooth meshes, or 3D gums meshes. Examples of 2D data include 2D photos of teeth or gums, or 2D image renderings of 3D data. The patient’s target dentition data 612 may include mid-treatment or post-treatment setups for orthodontic treatment, or post-treatment target tooth shapes for dental restorative treatment, or the like. The patient's target dentition data 612 and / or the patient’s dentition data 624 (which may optionally include facial data) may be provided to diffusion conditioning module 614. Diffusion conditioning module 614 may modify the data in a way that enables the data to be used by denoising diffusion ML model 522 to generate a predicted smile 610. For example, when the patient’s dentition data 624 includes a 2D photograph of the patient’s face, and the patient’s target dentition data 612 includes 3D tooth meshes, diffusion conditioning module 614 may process the patient’s 3D tooth meshes to generate one or more 2D image masks (using techniques of this disclosure) which are aligned with the teeth in the patient’s 2D facial data 624. Diffusion conditioning module 614 may output one or more latent representations, which may then be provided to aggregation module 622. Aggregation module 622 may may provide an aggregated or concatenated latent representation of the patient’s existing and / or target dentition data to denoising diffusion ML model 608. Examples of patient's target dentition data 612 include predicted final setups (for orthodontics), or generated tooth restoration designs (e.g., of crowns or roots, or the like), among other examples. In some implementations, a color palette 632, or other information pertaining to the target color, translucency, whiteness, or specular reflections of one or more teeth may be provided to diffusion conditioning module 614, to influence the appearance of the patient's dentition in the predicted smile 610. For example, the denoising diffusion ML model 608 may change the appearance of a treatment tooth in the predicted smile 610 based, at least in part, upon one or more entries found in the color palette 632 (e.g., one or more colors). Other aspects of tooth appearance can likewise be changed by denoising diffusion ML model 608, such as texture, based at least in part on either the color palette 632, of the target dentition 612. An image mask may contain connected components, each of whichmay have an identifier (ID). Each connected component may consist of one or more contiguous pixels. The contiguous pixels of a connected component may have the same non-zero value. The pixels in the image mask which do not belong to a connected component may be set to zero. A majority of connected components describe a single tooth, but a minority of connected components describe more than one tooth. Then the facial data 624 includes 3D mesh data, the values of the mesh element labels may be associated with corresponding mesh elements in the patient's facial data 624. The mesh element labels may identify substructures of the mesh which belong to different parts of anatomy (e.g., different teeth, oral care hardware attached to teeth, gums, lips, etc.). In some implementations, the patient's facial data 624 may be cut or separated into two or more meshes (e.g., corresponding to the region between the lips and the rest of the face, etc.). The patient's facial data with mask (or mesh elements) applied (604) may be encoded (606) into one or more latent representations, and subsequently be provided to denoising diffusion ML model 608.
[0056] The patient's facial data 624 (e.g., 2D photograph with the region between the patient's lips masked- off, or 3D mesh of patient's face with one or more mesh element labels configured to designate the region between the patient's lips, or the like) may be encoded (606) into one or more latent representations, which may then be provided to denoising diffusion ML model 608 (or a continuous normalizing flow -based model, or other flow-based model). One or more optional oral care arguments 602 may be encoded (616) into one or more latent representations, which may then be provided to denoising diffusion ML model 608, to customize the generated smile prediction 610 (e.g., to customize the whiteness of the teeth in the predicted smile, among others described herein). In other implementations, oral care arguments 602 may be provided to the model 608 directly, to the concatenation module 622, or to the diffusion conditioning module 614. Oral care arguments 602 may also be provided to other modules described herein, to influence those modules to generate outputs which are customized to the clinical treatment of the patient. The denoising diffusion ML model 608 may combine the patient's dentition data 612 with the patient's facial data 624, generating one or more realistic depictions 610 of the patient's appearance during or after treatment (e.g., at an intermediate stage of orthodontic treatment, or at the final setup stage of orthodontic treatment, etc.). The predicted smile 610 may include one or more 2D images, or one or more 3D meshes, and may show the patient’s target dentition 612 cleanly registered and / or integrated into the region between the patient's lips in a manner which minimizes visual artifacts.
[0057] The patient’s target dentition data 612 may include 2D or 3D representations of the patient's midtreatment or post-treatment dentition. When the patient’s target dentition 612 contains 3D mesh data and / or the facial data 624 includes at least one 2D image of the patient’s face, diffusion conditioning adapter 620 may generate a 2D view of the patient’s 3D target dentition 612 which aligns well with the opening between the patient’s lips in the 2D facial data 624, according to techniques of this disclosure. The 2D view of the patient’s 3D target dentition 612 may then be encoded (618) into a latent representation, and be provided to concatenation module 622. Oral care arguments 602 may influence the manner in which the patient's target dentition 612 andfacial data 624 are combined (e.g., are registered). For example, oral care arguments 602 may indicate which one or more teeth are used to compute a cost function. Other examples are described herein.
[0058] Method 700 may generate teeth which match the colors, textures and / or shapes of one or more reference teeth (e.g. , which may be provided as a part of the patient's target dentition data 712), using conditioned diffusion. The patient's target dentition 712 (e.g., one or more reference teeth with clinically or aesthetically desirable shapes, colors or whiteness, etc.) may be described by one or more 2D images, or a by one or more 3D meshes (e.g., a 3D mesh which has color and / or texture defined on the surface of the tooth). Similarly to the method 600, the method 700 may use a fully trained smile prediction ML model that operates using flow -based ML models (e.g., diffusion that is conditioned on 3D mesh data of the patient's dentition or 2D image data of the patient's dentition, or other examples of data described herein). In the example of FIG. 7, the method 700 uses the fully trained smile prediction ML model 708 to predict the outcome of dental restorative treatment. The patient's 2D or 3D dentition data 724 (which may include full-face or other facial data) may undergo segmentation (726). When the patient's dentition data 724 includes 2D image data, the segmentation may generate one or more image masks 728 which cover the one or more teeth which are to undergo dental restoration treatment. When the patient's dentition data 724 includes 3D mesh data, the segmentation may generate one or more mesh element labels 728 which designate one or more mesh elements as belonging to particular teeth (e.g., a tooth which is to undergo dental restoration). The output 728 of segmentation (e.g., either image masks or mesh element labels) may be applied (730) to the patient's dentition data 724. The patient's dentition data with mask (or mesh element labels) applied 704 may undergo latent encoding (706), and then be provided to denoising diffusion ML model 708 (e.g., a fully trained denoising diffusion ML model). Oral care arguments 702 (as described herein) may, in some implementations, be provided to the method 700 to customize the predicted smile 710. In some implementations, oral care arguments 702 may first undergo latent encoding (716) before being provided to denoising diffusion ML model 708. The patient's target dentition data 712 (e.g., one or more 2D or 3D representations of teeth) may include data pertaining to one or more of a target texture, target color, target whiteness, target shape, or other target characteristic for the patient's dentition during or after the completion of dental restorative treatment. The method 700 may transfer into the patient's original dentition (or facial data) 724 one or more of these aspects of the target dentition data 712. In some implementations, one or more color palettes 732 may be provided to diffusion conditioning module 714, and may influence the denoising diffusion ML model 708 to generate one or more predicted smiles with patient dentition colors which are specified by the one or more color palettes 732. In some implementations, diffusion conditioning module 714 may include one or more pairs of diffusion conditioning adapter 720 and encoder 718. Diffusion conditioning adapter 720 may process input data (e.g., patient's original facial data 724, target dentition data 712, oral care arguments 702, a color palette 732, etc.) to enhance or clarify aspects of the signals within that input data (e.g., using mesh element labelling, image segmentation, edge detection, image sharpening, or other methods described herein), and thenprovide its outputs to encoder 718, which may generate one or more latent representations (or latent embeddings), which may be provided to aggregation (or concatenation) module 722.
[0059] Oral care arguments 702 may include oral care arguments described herein, or any of the following: 1) indication of which one or more teeth of the patient's dentition 724 which are to undergo whitening, or color change; 2) Information pertaining to the amount of increase or decrease in whitening; 3) information pertaining to the new color for one or more teeth (e.g., specified by a color palette, or other color description, etc.) of the patient's dentition 724; 4) information pertaining to the amount of specular reflections for one or more teeth of the patient's dentition 724; 5) information pertaining to the translucency of one or more teeth of the patient's dentition 724; and 6) information pertaining to a change in shape of one or more teeth of the patient's dentition 724.
[0060] In some implementations, when the patient's dentition data 724 includes 2D image data, the image mask 728 may be modified (e.g., by changing the shape of the mask), in order to alter the shapes of one or more teeth in predicted smile (or predicted dentition) 710, which is generated by the denoising diffusion ML model 708. For example, whereas the image mask 728 may initially describe the shape or contours of a mal-formed, damaged or chipped tooth, in some implementations, the image mask 728 may be modified to describe the shape of a clinically or aesthetically improved tooth (e.g., resulting in an image mask in which the target teeth have uniform shapes, free of chips, cracks, unwanted diastemas, evidence of decay, etc.). In some implementations, this modified image mask may be provided to diffusion conditioning module 714, or may undergo latent encoding (e.g., using an encoder 718) and then be provided to latent representation concatenation module 722, or directly to denoising diffusion ML model 708. In other implementations, the modified image mask may be applied (730) to the patient's dentition (or facial) data 724, resulting in the patient's dentition data with mask applied 704, which may then undergo latent encoding (706), and finally be provided to denoising diffusion ML model 708. In still other implementations, the target shapes for one or more teeth may be provided to the method 700 by patient's target dentition data 712, and those target tooth shapes in the patient's target dentition data 712 may influence the denoising diffusion ML model 708 to generate one or more predicted smiles in which the corresponding teeth have the specified target shapes.
[0061] Denoising diffusion ML model 522 may include one or more normalizing flows-based models, one or more denoising diffusion-based models, or the like. Method 800 describes the training of a denoising diffusion ML model 522 for use in clinical or aesthetic treatment of the patient. According to various implementations, various input data may be provided to denoising diffusion model 522, including patient's target dentition data 542, the patient's dentition with mask applied 806, oral care arguments 830, or mask data 804 (e.g., one or more image masks, or one or more mesh element labels), among others described herein. Any of these inputs may be provided to one or more encoders 808 to encode those inputs into one or more latent representations 812. After the patient’s target dentition 542 has been encoded into one or more latent representations, then the one or more latent representations of the patient’s target dentition 802 may be provided directly to encoder-decoder structure822 (e.g., a UNet). The one or more latent representations 812 or 802 may have a reduced dimensionality relative to the original data, and / or may be formatted so as to be suitable to be provided to encoder-decoder structure 822 (e.g., a UNet, a vision transform ViT, an autoencoder, a pyramid encoder-decoder structure, etc.). During model training, noise may be computed (814). For example, Gaussian noise may be computed (814) and be formatted into one or more actual noise tensors 816. The one or more actual noise tensors 816 may be combined, added or otherwise integrated (818) with the one or more latent representations 812. In some implementations, a concatenated latent representation of all inputs 812 is summed (818) with an actual noise tensor 816, and then provided to encoder-decoder structure 822 (e.g., a UNet that contains skip connections, and that extracts hierarchy of global neural network features, to local neural network features from the UNet' s inputs), which generates a predicted data structure describing noise (e.g., a noise tensor 824). Loss may be computed (826) between the predicted noise tensor 824 and the actual noise tensor 816. Among the other losses described herein, the loss may compute the mean squared error difference between the actual noise tensor 816 and the predicted noise tensor 824. The computed loss may be used to update (828) the weights of the encoder-decoder structure 822 (e.g., a UNet), until training is done (820). In some implementations, training may proceed (810) until loss drops below a threshold value, or until a target count of epochs have been completed.
[0062] The fully trained denoising diffusion ML model 522 is shown in method 600 as module 608, and method 700 as module 708. Method 900 describes the operational use of that fully trained model, for example, in a time-constrained setting (e.g., while the patient waits in the treatment chair). For instance, method 900 shows the operational use of the fully trained denoising diffusion ML models 608 or 708. According to various implementations, various input data may be provided to denoising diffusion models 608 or 708, the patient's dentition with mask applied 906, oral care arguments 926, or mask data 904 (e.g., one or more image masks, or one or more mesh element labels), among others described herein. Any of these inputs may be provided to one or more encoders 908 to encode those inputs into one or more latent representations 912. The denoising ML model may iterate (910) for M iterations, until a target accuracy is achieved, or until the outputs are otherwise deemed to be clinically or aesthetically suitable. In some implementations, a concatenated latent representation of all inputs 912, a latent representation of the patient’s target dentition 928, and / or a noising image (or noisy 3D representation) 902 may be provided to encoder-decoder structure 914 (e.g., a UNet), which may generate predicted noise tensor 916. Predicted noise tensor 916 may be removed (920) (e.g., by subtraction between matrices or tensors, etc.) from the latent representation 912 (e.g., a latent vector or latent embedding, etc.). When done, the denoised latent representation may be provided to decoder 922, which may reconstruct the denoised latent representation into one or more 2D images or 3D representations which are suitable for clinical treatment of the patient (e.g., for smile prediction). The resulting predicted smile 924 may be outputted for aesthetic prediction or clinical treatment planning. When the input data includes 3D mesh data (or other 3D representations, such as 3D point clouds, 3D surfaces, etc.), mesh element feature vectors may be computed (e.g., a mesh element feature vector may be computed for each 3D point, 3D vertex, 3D edge, 3D voxel, or 3Dface of a 3D representation). The mesh elements and corresponding mesh element feature vectors of the 3D input data may be provided to the neural networks of this disclosure (e.g., encoder 536, or UNet 822, among others), to improve the ability of those neural networks to encode the distribution of the 3D input data (e.g., enable the neural networks to better encode the shapes and / or structures of those 3D representations). The structure of a 3D representation may include aspects of the 3D representation such as: 1) which mesh elements are adjacent to each other; 2) which mesh elements are connected to each other; 3) which mesh elements are within a threshold Euclidean distance of each other; and / or 4) which mesh elements are within a particular count of connections of each other, or the like.
[0063] In some implementations, conditioning adapter 534 may register the patient's target dentition data 1002 (e.g., 2D or 3D data) with the patient's original dentition data 1202 (e.g., 2D or 3D data) as follows. Diffusion conditioning adapter 534 may use method 1100 to segment (and / or project) (1008) target dentition 1002 into dentition mask 1114 (e.g., registered projected 2D dentition mask image). Method 1100 may use optimized camera parameters 1104 to perform (1008) the projection operation (e.g., using raycasting). Diffusion conditioning adapter 534 may use method 1600 to apply (1602) a patient smile mask 1308 to a dentition mask 1114, resulting in dentition mask 1604 (e.g., registered projected 2D dentition mask with smile mask applied image). Diffusion conditioning adapter 534 may output dentition mask 1604 (e.g., registered projected 2D dentition mask with smile mask applied image), which may be provided to denoising diffusion ML model 522, 607, or 708. Optimized camera parameters 1104 may be generated using optimization techniques described herein.
[0064] As already discussed, to or more dentitions can be registered. In this context, for example, when the patient's target dentition 1002 contains one or more 3D representations of teeth (or gums), and the patient's original dentition contains one or more 2D dentition images 1202, techniques of this disclosure can register a 2D projection 2204 of the patient’s 3D target dentition 1002 with the patient's 2D dentition image 1202. The patient's 2D dentition photo may include one or more images of the patient's teeth (e.g., when the patient's mouth is open for a smile, etc.) In some instances, a target dentition 1002 may include 3D representations of the teeth and / or gums. In some instances, the target dentition 1002 may contain one or more teeth which are designated for dental restorative treatment, and / or one or more teeth which are designated for orthodontic treatment.
[0065] Raycasting techniques known to one skilled in the art may use camera parameters to project a 3D representation onto a 2D surface (e.g., project a 3D mesh into a 3D image). Camera parameters (e.g., azimuth, elevation, roll, focus, y-axis coordinate, x-axis coordinate, or z-axis coordinate, among others) may be computed which, in some implementations, describe the position and / or orientation of the camera that was used to generate each of the one or more 2D images. The positive y-axis may point outward from the incisors, in the direction of the patient's forward gaze. Angles may be expressed in degrees, radians, or the like.
[0066] Stated another way, in some implementations, camera parameters may be optimized using a stochastic, bounded global optimizer (e.g., differential evolution). Either gradient-based or non-gradient basedmethods may be used. In some implementations of genetic algorithms, there may be a population that includes a plurality of members (otherwise known as individuals), each of member includes at least a set of camera parameters. Each individual in the population may undergo variation operators, which may change the population member in small or large ways (e.g., by adding small amounts of noise to one or more of the intrinsic or extrinsic parameter values which are included in a population member). A population member may include a set of camera parameters which is capable of projecting the patient's 3D dentition into 2D using raycasting. Variation operators (e.g., mutation, crossover, etc.) may operate on the population member (e.g., the data structure that is being optimized). In some implementations, the population member may be encoded into a latent representation using an encoder neural network, and then the variation operators may vary that latent representation through the course of optimization. The variation operators may vary the camera parameters (or latent representations of the camera parameters) of the individual, and then the fitness (or cost) associated with those modified camera parameters may be computed.
[0067] This fitness evaluation may use the camera parameters 2202 to project (1008) the 3D meshes of the patient's target dentition 1002 into a segmented dentition mask 2204 (e.g., projected 2D dentition mask image). When projecting one or more 3D tooth meshes from 3D to 2D, segmentation is performed through raycasting. For example, a pixel in a particular connected component of segmented dentition mask 2204 (e.g., projected 2D dentition mask image) belongs to the tooth through which the ray intersects(?) during the raycasting calculation.
[0068] A fitness or cost function (e.g., MSE) may be computed (2210) to quantify the difference between the distributions of two or more dentitions (e.g., 2D or 3D dentitions). For example, cost may be computed to quantify the alignment of the dentition mask 2206 (e.g., projected 2D subset mask image) and the corresponding dentition mask 2208 (e.g., original dentition 2D subset mask image). The computed fitness is then assigned (2212) to the associated member 2202. With each population member (e.g., set of camera parameters) thusly assigned a fitness, the method may, in some implementations, eliminate one or more of the least-fit members from the population, and replace the eliminated members with new members. The new members may comprise mutated copies of the remaining members, or may be generated by combining aspects of two or more remaining members (e.g., crossover). The genetic algorithm may then evaluate the fitness of each member and iterate until a stopping criterion is met (e.g., a threshold count of iterations has transpired, an average population fitness has been achieved, or the population has produced at least one population member with a minimal fitness, etc.). In this manner, the optimization algorithm may generate a set of camera parameters which does a good job of projecting the patient's 3D target dentition 1002 into a projected 2D dentition mask image 2204 which has a good alignment with the corresponding segmented 2D dentition mask image 1208. Stated another way, dentition mask 2208 may be overlaid with dentition mask 2206, and a fitness (or cost) function may be computed. The fitness function determines how well the two dentition masks overlap.
[0069] When target dentition data 1002 contains 3D tooth meshes, raycasting may project (1008) the 3D data of those tooth meshes into a 2D image 2204, forming one or more connected components. Each connectedcomponent (or set of adjacent connected components belonging to the same tooth) may have an assigned a color or other identifying attributes which associate the connected component with a particular tooth. In some implementations, the dentition masks 2204 and 2206 each contain sets of these colored (or labelled) connected components. A segmented dentition 1208 (e.g., a segmented 2D dentition masks image) may be generated by performing segmentation (1206) on patient dentition 1202 (e.g., a 2D patient dentition image).
[0070] In some implementations, a cost function (or fitness function) may generate a zero value when the 'projected 2D dentition mask' and 'segmented 2D dentition mask' perfectly overlap. Examples of such a cost function include MSE, Huber Loss, Mean Absolute Error (MAE), Quantile Loss, Log-Cosh Loss, and Mean Absolute Percentage Error (MAPE), among others. A cost function may, in some implementations, compute the binary segmentation overlap error at each pixel location between two image masks. In some implementations, the overlap error at each pixel may be computed, at least in part, based on the distance between the pixel and the closest lip edge (e.g., the edge of smile mask 1308). In some implementations, the cost function may calculate the distance between each pixel in the dentition mask 2204 and the corresponding pixel(s) in dentition mask 1208. In some implementations, the cost function may calculate the distance between each pixel in the dentition mask 2206 and the corresponding pixel(s) in dentition mask 1208.
[0071] Method 1000 shows how initial camera parameters 1004 may be used to project (1008) the 3D target dentition 1002 into an initial 2D representation 1014 (e.g., 2D image mask that contains connected components which show the outlines of the teeth). Each connected component may have a color or other identifier. A connected component may define the 2D profiles of one or more teeth. In some implementations, a subset of the connected components in image 1014 may be selected (1016). The mask which contains that subset of connected components is the initial projected 2D subset mask image 1018. The mask image 1018 may be registered (or aligned) with the segmented 2D dentition masks image 1208 by optimizing a cost function (or fitness function). The mask image 1208 may be generated by performing 2D tooth segmentation (1206) on the 2D patient dentition image 1202. Alternatively, if the patient dentition 1202 contains 3D data, then 3D mesh segmentation (1206) may be performed. Oral care arguments 1106 may be provided to the 2D or 3D segmentation methods. Examples of oral care arguments 1106 that can influence segmentation include thresholding values, among others.
[0072] For the registration method to function optimally, camera parameters (e.g., intrinsic or extrinsic parameters) can be used, which as described elsewhere, do a good job of projecting the 3D target dentition 1002 into 2D. Optimization algorithms which can be used to perform such an optimization include: simulated annealing, genetic algorithms, differential evolution, particle swarm analysis, ant colony optimization, or other optimization algorithms (e.g., either gradient-based or non-gradient based algorithms).
[0073] One or more cost functions or fitness functions may quantify the correctness of the alignment of two or more dentitions. Cost (or fitness) functions may be iteratively evaluated to guide the operation of an optimization method, until the cost (or fitness) crosses a pre-determined threshold (or until a pre-determinedcount of iterations have transpired). The cost function may be minimized using optimization algorithms described herein. Method 2200 describes a method of using a cost (or fitness) function to determine the effectiveness or utility of a set of one or more camera parameters (e.g., described herein as a population member 2202). The 2D image 2204 (or 3D meshes 1002) of the patient's target dentition may be aligned with a segmented 2D dentition masks image 1208 of the patient’s current dentition. The masks image 1208 may be generated by performing (1206) 2D image segmentation on a photo 1202 of the patient (such as when the patient’s teeth are showing in a smile). The masks image 1208 may contain connected components that show the outlines of the teeth.
[0074] A cost function may be computed between two (or more) masks, to quantify the extent to which the masks differ. For example, a cost function may be computed to quantify the difference between mask subset 2208 (e.g., a subset of tooth connected components from the original dentition data) and mask subset 2206 (e.g., a subset of tooth connected components from the target dentition data). One of the advantages of computing a cost function between two subset images is that the operation can be performed between teeth whose identities are known with high confidence. After tooth segmentation is performed (e.g., segmentation of patient dentition 1202), tooth identities may be determined. The determination of tooth identities may lead to erroneous identifications (or errors) a certain percentage of the time. Methods of determining tooth identify (e.g., templatebased approaches, ML model-based approaches, etc.) may, in some implementations, output confidence values. For example, the central incisors may be identified in segmented dentition mask 1208 with high confidence, but a second bicuspid may be identified with low confidence (e.g., due to partial occlusion by the lips). Therefore, there is value in using teeth whose identities are known with high confidence.
[0075] In some implementations, cost may be computed between corresponding teeth between two or more image masks. For example, a cost may be computed that quantifies the alignment of the ULI connected components in dentition mask 2206 and image mask 2208. Costs may be computed for other teeth, as well. In some implementations, the cost calculation for a particular tooth may incorporate cost information from one or more other teeth in the dentition (e.g., adjacent teeth). This approach can mitigate the error introduced by erroneous tooth identifications that may result from tooth segmentation. In some non-limiting examples, the alignment cost for UL3 may be computed using a min() function, as follows: cost_UL3 = min( (cost of overlapping of mesh UL3 and photo UL3), (cost of overlapping of mesh UL2 and photo UL3), (cost of overlapping of mesh UL4 and photo UL3)). This approach can mitigate the effects of an erroneous tooth identification when, for example, the adjacent tooth is correctly identified.
[0076] The total cost of aligning the two or more image masks may be computed based on the alignment costs of one or more of the individual teeth contained within the image masks. In some non-limiting examples, oral care arguments 602 or segmentation confidence results may determine that UR3, UR1, ULI, UL3 are to be included in cost function calculation. In such examples, the total cost between the two or more image masks may be computed as (cost_UR3 + cost_URl + cost_ULl + cost_UL3)A2.
[0077] In some implementations, the subset masks 1018 and 1118 may retain teeth which are segmented and identified with high confidence. In other implementations, the subset masks 1018 and 1118 may retain teeth which are specified by oral care arguments 1106.Oral care arguments 602 may include oral care arguments described herein, or any of the following: 1) indication of which one or more teeth are to be used to register the target dentition 1002 (e.g., 2D or 3D data) with the patient's original dentition data 1202 (e.g., 2D or 3D data). A subset of the patient's teeth may be used for registration (e.g., 4 teeth, or another count of teeth), thereby saving compute resources while still enabling registration to take place. Accuracy may actually be improved, because registration may be performed on teeth which may be more easily segmented (e.g., upper central incisors, among other examples);2) information pertaining to segmentation, such as the threshold for the confidence of a tooth identity assignment that may result from tooth segmentation. Stated another way, the tooth segmentation process may assign a tooth identity to each segmented tooth in the segmentation output (e.g., 3D mesh or 2D image output). Each tooth identity assignment may have an associated confidence; 3) mouth openness threshold, which may specify a minimum width (or height) for patient smile mask 1308, or a minimum number of teeth which must be visible in patient dentition 1202 or target dentition 1002, for the patient's mouth to be considered sufficiently open for registration to be performed; 4) initial camera parameters (e.g., intrinsic or extrinsic) which may be used to project target dentition data 1002 into 2D image mask 2204 (e.g., when target dentition 1002 contains one or more 3D representations). Initial camera parameters may, in some instances, include zero values for some terms. The camera parameters may be optimized using a genetic algorithm, or other techniques described herein; and 5) other information pertaining to the registration of target dentition data 1002 and original dentition data 1202.
[0078] An optimization module (e.g., containing a genetic algorithm, or others described herein) may be executed to generate new variations of the camera parameters (e.g., population members 2202). The new variations of camera parameters (e.g., population member 2202) may be evaluated by the process described above. As a result, each variation of camera parameters is assigned a cost or fitness value. In this manner, the best variation of camera parameters may be identified (e.g., the variation of camera parameters which does the best job of projecting the 3D target dentition into a 2D image mask that aligns well with the patient’s current dentition). The population member with the highest fitness can be used as the optimal camera parameters 1104. In this manner, according to some implementations, a predicted setup or a predicted tooth restoration design (either of which may comprise 3D or 2D data) can be registered with a 2D photograph of the patient’s smile. The resulting image mask 1114 (e.g., a registered projected 2D dentition mask image) of the target dentition can be provided to denoising diffusion ML model 522 as a guide (or instruction) regarding the shapes (or poses) of the teeth which are to appear in the predicted smile.
[0079] In addition to the camera parameters, other data may also be optimized according to particular implementations. For instance, additional data may be included with the camera parameters within each population member 2202. Using a particular non-limiting example, data which describes the alignment (e.g.,translational and / or rotational parameters that adjust the relative poses of the upper and lower arches) of the upper and lower arches of the patient's 3D target dentition 1002 (or other 3D mesh dentitions) may be optimized. The optimization of these data may improve the alignment between the upper and lower arches, and therefore improve the accuracy of the registration techniques, smile prediction techniques, or other digital orthodontic operations described herein. Stated another way, the relative alignment of the patient's upper and lower arches may be apparent from a 2D photo of the patient's smile 1202. For example, when both upper and lower teeth are visible in the smile photo 1202, the relative poses of those upper and lower teeth in the 2D photograph 1202 can be used as a reference to register the 3D upper and lower aches in target dentition data 1002 (when the target data includes 3D meshes of the teeth or gums), according to the registration techniques described herein. The arch alignment technique can optimize arch alignment parameters (e.g., translation and / or rotation values) which can orient the arches relative to each other. The technique can be used to align an upper arch to a lower arch, or to align a lower arch to an upper arch. Arch alignment parameters can be defined relative to a first arch, and then be used to rotate and / or translate a second arch into alignment with the first arch.
[0080] The technique can align arches so that the arches achieve at least the level or quality of intercuspation that is present in the dentition shown in the patient's smile 1202. Acceptable intercuspation is where occlusal contacts are favorable. After the upper and lower arch become aligned with each other, the pair can then be oriented into alignment with a canonical coordinate system (e.g., X = patient right, Y = patient front, Z = patient up). For example, in some implementations, the pair of arches can be rotated so that the occlusal plane is approximately aligned with the XY plane. The pair of arches may be further rotated to orient the pair in XY plane to align with +X axis (e.g., patient right) and +Y axis (e.g., patient front midline). Other canonical alignments may also be defined.
[0081] FIG. 21 highlights the improvement in dentition alignment that is achieved using the registration techniques described herein. Example 2102 shows pre-registration misaligned dentitions. Example 2102 shows the overlap of mask image 1404 (e.g., initial projected 2D subset mask with smile mask applied image) and a corresponding subset of teeth from mask image 1208 (e.g., segmented 2D dentition masks image). The mask image 1404 was generated using initial camera parameters and raycasting. Example 2104 shows postregistration well-aligned dentitions. Example 2104 shows the overlap of mask image 1504 (e.g., registered projected 2D subset mask with smile mask applied image) and a corresponding subset of teeth from mask image 1208 (e.g., segmented 2D dentition masks image). The mask image 1504 was generated using the optimized camera parameters and raycasting to project 3D dentition 1002 into a 2D dentition, and then selecting a subset of teeth, to produce image mask 1118. The smile mask 1308 is then applied to the image mask 1118, to produce image mask 1504.
[0082] Aspects of the present disclosure can provide a technical solution to the technical problem of predicting a realistic post-treatment appearance for a patient (e.g., a smile prediction), which may involve performing dentition registration. In particular, by practicing techniques disclosed herein computing systemsspecifically adapted to perform smile prediction for aesthetic or clinical planning purposes are improved. For example, aspects of the present disclosure improve the performance of a computing system having a 3D representation of the patient’s dentition by reducing the consumption of computing resources. In particular, aspects of the present invention reduce computing resource consumption by decimating 3D representations of the patient’s dentition (e.g., reducing the counts of mesh elements used to describe aspects of the patient’s dentition) so that computing resources are not unnecessarily wasted by processing excess quantities of mesh elements. Additionally, decimating the meshes does not reduce the overall predictive accuracy of the computing system (and indeed may actually improve predictions because the input provided to the ML model after decimation is a more accurate (or better) representation of the patient’s dentition). For example, noise or other artifacts which are unimportant (and which may reduce the accuracy of the predictive models) are removed. That is, aspects of the present invention provide for more efficient allocation of computing resources while simultaneously increasing predictive accuracy of the computing system.
[0083] In a further example, aspects of the present disclosure improve the performance of a computing system which computes cost functions as a part of one or more optimization techniques. The cost functions may quantify the differences between two or more dentition masks, and techniques of this disclosure may reduce the number and identities of teeth represented in the dentition masks, and yet may still produce accurate registration results. In fact, removing teeth that have low-confidence segmentations (e.g., whose identities are know with lower confidence) from the dentition masks can actually increase the accuracy of the registration operation, because potentially erroneous data are removed.
[0084] Furthermore, aspects of the present disclosure may need to be executed in a time -constrained manner, such as when an oral care appliance must be generated for a patient immediately after intraoral scanning (e.g., while the patient waits in the clinician’s office). As such, aspects of the present disclosure are necessarily rooted in the underlying computer technology of post-treatment appearance prediction and cannot be performed by a human, even with the aid of pen and paper. For instance, implementations of the present disclosure must be capable of: 1) storing thousands or millions of mesh elements of the patient’s dentition in a manner that can be processed by a computer processor; 2) performing calculation on thousands or millions of mesh elements, e.g., to quantify aspects of the shape and or / structure of an individual tooth in the 3D representation of the patient’s dentition, 3) generating thousands or millions of successively noisy versions of the training data in a Markov chain, 4) training UNets to perform the denoising operation (which may involve computing and removing noise tensors containing thousands or millions of values); and 5) a realistic smile for a patient, which may involve performing dentition registration, and do so during the course of a short office visit.
[0085] In some implementations, representation learning may be used to train ML models of this disclosure. For example, a first machine learning (ML) model may generate one or more first latent representations of the input (e.g., the patient’s dentition, oral care arguments, etc.). Mesh element featurevectors may be computed for the 3D representations among the inputs (as described herein) and subsequently be provided to the first ML model (e.g., an encoder or other representation generation module). For example, the patient’s pre-treatment dentition may comprise one or more 3D meshes of teeth and or gums. Mesh element feature vectors may be computed for one or more mesh elements in the 3D dentition meshes, and then be provided, along with the associated mesh elements of the 3D dentition meshes, to the first ML model. Mesh elements can include vertices, edges, voxels, or faces, etc. The one of more latent representations may be provided to a second ML model, which may generate data structures related to digital oral care treatment. Examples of data structures related to digital oral care treatment include 1) orthodontic setups transforms, 2) tooth restoration designs, or 3) other 3D oral care representations described here. These representation learning techniques include ML models to perform automated setups prediction, automated restoration design generation, or the automated generation of other 3D oral care representations described herein. In some implementations, automated restoration design generation may generate one of more 3D representations of structures interior to the tooth (e.g., 3D meshes that describe the boundaries between enamel and dentin, or between dentin and pulp).
[0086] 3D oral care representations may include, but are not limited to: 1) a set of mesh element labels which may be applied to the 3D mesh elements of teeth / gums / hardware / appliance meshes (or point clouds) in the course of mesh segmentation or mesh cleanup; 2) 3D representation(s) for one or more teeth / gums / hardware / appliances for which shapes have been modified (e.g., trimmed, distorted, or filled-in) in the course of mesh segmentation or mesh cleanup; 3) one or more coordinate systems (e.g., describing one, two, three or more coordinate axes) for a single tooth or a group of teeth (such as a full arch - as with the LDE coordinate system); 4) 3D representation(s) for one or more teeth for which shapes have been modified or otherwise made suitable for use in dental restoration; 5) 3D representation(s) for one or more dental restoration appliance components; 6) one or more transforms to be applied to one or more of: dental restoration appliance library component placement relative to one or more teeth, a tooth to be placed for an orthodontic setup (either final setup or intermediate stage), a hardware element to be placed relative to one or more teeth or the like; 7) an orthodontic setup; 8) a 3D representation of a hardware element (such as facial bracket, lingual bracket, orthodontic attachment, button, hook, bite ramp, etc.) to be placed relative to one or more teeth, etc.; 8) a 3D representation of a bonding pad for a hardware element (which may be generated for a specific tooth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting-out the tooth via a Boolean operation); 9) 3D representation of a clear tray aligner (CTA); 10) the location or shape of a CTA trimline (e.g., described as either a mesh or polyline); 11) archform that describes the contours or layout of an arch of teeth (e.g., described as a 3D polyline or as a 3D mesh or surface), which may follow the incisal edges one or more teeth, which may follow the facial surfaces of one or more teeth, which may in some implementations correspond to the maloccluded arch and in other implementations correspond to the final setup arch (the effects of malocclusion on the shape of the archform may be diminished by smoothing oraveraging of the shape of the archform), which may be described by one or more control points and / or a spline; 12) 3D representation of a fixture models (e.g., depictions of teeth and gums for use in thermoforming clear tray aligners, or depictions of teeth / gums / hardware for use in thermoforming indirect bonding trays); 13) one or more latent space vectors (or latent capsules) produced by the 3D encoder stage of a 3D autoencoder which has been trained on the reconstruction of oral care meshes (e.g., a variational autoencoder which has been trained for tooth reconstruction); 14) one or more oral care metrics values (e.g., such as orthodontic metrics or restoration design generation metrics) for one or more teeth; 15) one or more landmarks (e.g., 3D points) which describe the shapes and / or geometrical attributes of one or more teeth, other dentition structures or hardware structures (e.g., to be used in orthodontic setups creation or restoration appliance component generation or placement); 16) 3D representation created by scanning (e.g., optically scanning, CT scanning or MRI scanning) a 3D printed part corresponding to one or more teeth / gums / hardware / appliances (e.g., a scanned fixture model); 17) 3D printed aligners (including optionally local thickness, reinforcing rib geometry, flap positioning, or the like); 18) 3D representation of the patient's dentition that was captured chairside by a clinician or medical practitioner (e.g., in a context where the 3D representation is validated chairside, before the patient leaves the clinic, so that errors can be detected and re-scans performed as necessary); 19) dental restoration tooth design (e.g., for veneers, crowns, bridges or dental restoration appliances); 20) 3D representations of one or more teeth for use in digital oral care treatment; 21) other 3D printed parts pertaining to oral care procedures or other fields; 22) IPR cut surfaces; 23) one or more orthodontic setups transforms associated with one or more IPR cut surfaces; 24) a (digital) pontic tooth design which may fill at least a portion of the space between teeth to allow room in an orthodontic setup for an erupting tooth to later emerge from the gums; 25) a component of a fixture model (e.g., comprising fixture model components such as interproximal webbing, block-out, bite locks, bite ramps, interproximal reinforcement, gingival ridges, torque points, power ridges, pontic tooth or dimples, among others). In some instances, 3D oral care representations may include a prediction of a patient’s smile, which may show the post-treatment teeth, lips, cheeks, nose, and / or other facial features in relation to each other. In implementations where 3D oral care representations include a prediction of a patient’s smile, the predictions may be presented in several formats, including as 2D images or as 3D representations.
[0087] Techniques of this disclosure may, in some implementations, use PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using either PointNet or PointNet++ as a basis for training) to extract local or global neural network features from a 3D point cloud or other 3D representation (e.g., a 3D point cloud describing aspects of the patient’s dentition - such as teeth or gums). Techniques of this disclosure may, in some implementations, use U-Nets to extract local or global neural network features from a 3D point cloud or other 3D representation.
[0088] 3D oral care representations are described herein as such because 3-dimensional representations are currently state of the art. Nevertheless, 3D oral care representations are intended to be used in a non-limitingfashion to encompass any representations of 3-dimensions or higher orders of dimensionality (e.g., 4D, 5D, etc.), and it should be appreciated that machine learning models can be trained using the techniques disclosed herein to operate on representations of higher orders of dimensionality.
[0089] In some instances, input data may comprise 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data pertaining to a spline (e.g., control points). An encoder-decoder structure may comprise one or more encoders, or one or more decoders. In some implementations, the encoder may take as input mesh element feature vectors for one or more of the inputted mesh elements. By processing mesh element feature vectors, the encoder is trained in a manner to generate more accurate representations of the input data. For example, the mesh element feature vectors may provide the encoder with more information about the shape and / or structure of the mesh, and therefore the additional information provided allows the encoder to make better-informed decisions and / or generate more-accurate latent representations of the mesh. Examples of encoder-decoder structures include U-Nets, autoencoders or transformers (among others). A representation generation module may comprise one or more encoder-decoder structures (or portions of encoders-decoder structures - such as individual encoders or individual decoders). A representation generation module may generate an information-rich (optionally reduced-dimensionality) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[0090] A U-Net may comprise an encoder, followed by a decoder. The architecture of a U-Net may resemble a U shape. The encoder may extract one or more global neural network features from the input 3D representation, zero or more intermediate -level neural network features, or one or more local neural network features (at the most local level as contrasted with the most global level). The output from each level of the encoder may be passed along to the input of corresponding levels of a decoder (e.g., by way of skip connections). Like the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For instance, the decoder may output a representation of the input data which may contain global, intermediate or local information about the input data. The U-Net may, in some implementations, generate an information-rich (optionally reduced-dimensionality) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[0091] An autoencoder may be configured to encode the input data into a latent form. An autoencoder may train an encoder to reformat the input data into a reduced-dimensionality latent form in between the encoder and the decoder, and then train a decoder to reconstruct the input data from that latent form of the data. A reconstruction error may be computed to quantify the extent to which the reconstructed form of the data differs from the input data. The latent form may, in some implementations, be used as an information-rich reduced- dimensionality representation of the input data which may be more easily consumed by other generative or discriminative machine learning models. In most scenarios, an autoencoder may be trained to input a 3D representation, encode that 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close facsimile of that input 3D representation at the output.
[0092] A transformer may be trained to use self-attention to generate, at least in part, representations of its input. A transformer may encode long-range dependencies (e.g., encode relationships between a large number of inputs). A transformer may comprise an encoder or a decoder. Such an encoder may, in some implementations, operate in a bi-directional fashion or may operate a self-attention mechanism. Such a decoder may, in some implementations, may operate a masked self-attention mechanism, may operate a cross-attention mechanism, or may operate in an auto-regressive manner. The self-attention operations of the transformers described herein may, in some implementations, relate different positions or aspects of an individual 3D oral care representation in order to compute a reduced-dimensionality representation of that 3D oral care representation. The crossattention operations of the transformers described herein may, in some implementations, mix or combine aspects of two (or more) different 3D oral care representations. The auto-regressive operations of the transformers described herein may, in some implementations, consume previously generated aspects of 3D oral care representations (e.g., previously generated points, point clouds, transforms, etc.) as additional input when generating a new or modified 3D oral care representation. The transformer may, in some implementations, generate a latent form of the input data, which may be used as an information-rich reduced-dimensionality representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[0093] In some implementations, an encoder-decoder structure may first be trained as an autoencoder. In deployment, one or more modifications may be made to the latent form of the input data. This modified latent form may then proceed to be reconstructed by the decoder, yielding a reconstructed form of the input data which differs from the input data in one or more intended aspects. Oral care arguments, such as oral care parameters or oral care metrics may be provided to the encoder, the decoder, or may be used in the modification of the latent form, to influence the encoder-decoder structure in generating a reconstructed form that has desired characteristics (e.g., characteristics which may differ from that of the input data).
[0094] Techniques of this disclosure may, in some instances, be trained using federated learning. Federated learning may enable multiple remote clinicians to iteratively improve a machine learning model, while protecting data privacy (e.g., the clinical data may not need to be sent “over the wire” to a third party). Data privacy is particularly important to clinical data, which is protected by applicable laws. A clinician may receive a copy of a machine learning model, use a local machine learning program to further train that ML model using locally available data from the local clinic, and then send the updated ML model back to the central hub or third party. The central hub or third party may integrate the updated ML models from multiple clinicians into a single updated ML model which benefits from the learnings of recently collected patient data at the various clinical sites. In this way, a new ML model may be trained which benefits from additional and updated patient data (possibly from multiple clinical sites), while those patient data are never actually sent to the 3rd party. Training on a local inclinic device may, in some instances, be performed when the device is idle or otherwise be performed during off-hours (e.g., when patients are not being treated in the clinic). Devices in the clinical environment for thecollection of data and / or the training of ML models for techniques described herein may include intra-oral scanners, CT scanners, Xray machines, laptop computers, servers, desktop computers or handheld devices (such as smart phones with image collection capability). In addition to federated learning techniques, in some implementations, contrastive learning may be used to train, at least in part, the ML models described herein. Contrastive learning may, in some instances, augment samples in a training dataset to accentuate the differences in samples from difference classes and / or increase the similarity of samples of the same class.
[0095] This disclosure pertains to digital oral care, which encompasses the fields of digital dentistry and digital orthodontics. This disclosure generally describes methods of processing three-dimensional (3D) representations of oral care data, and / or associated transforms. It should be understood, without loss of generality, that there are various types of 3D representations. One type of 3D representation is a 3D geometry. A 3D representation may include, be, or be part of one or more of a 3D polygon mesh, a 3D Although the term “mesh” is used frequently throughout this disclosure, the term should be understood, in some implementations, to be interchangeable with other types of 3D representations. A 3D representation may describe elements of the 3D geometry and / or 3D structure of an object, point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels - for sparse processing), or 3D representations which are described by mathematical equations.
[0096] A 3D representation may be produced using a 3D scanner, such as an intraoral scanner, a computerized tomography (CT) scanner, ultrasound scanner, a magnetic resonance imaging (MRI) machine or a mobile device which is enabled to perform stereophotogrammetry. A 3D representation may describe the shape and / or structure of a subject. A 3D representation may include one or more 3D mesh, 3D point cloud, and / or a 3D voxelized representation, among others. A 3D mesh includes edges, vertices, or faces. Though interrelated in some instances, these three types of data are distinct. The vertices are the points in 3D space that define the boundaries of the mesh. These points would alternatively be described as a point cloud but for the additional information about how the points are connected to each other, as described by the edges. An edge is described by two points and can also be referred to as a line segment. A face is described by a number of edges and vertices. For instance, in the case of a triangle mesh, a face comprises three vertices, where the vertices are interconnected to form three contiguous edges, which in turn define a face for the triangle mesh. Some meshes may contain degenerate elements, such as non-manifold mesh elements, which may be removed using conventional techniques, to the benefit of later processing. Other mesh pre-processing operations are possible in accordance with aspects of this disclosure. 3D meshes are commonly formed using triangles, but may in other implementations be formed using quadrilaterals, pentagons, or some other n-sided polygon. In some implementations, a 3D mesh may be converted to one or more voxelized geometries (i.e., comprising voxels), for example, so that sparse processing can be performed (e.g., using an encoder). In some examples, one feature vector is generated per vertex of the mesh.
[0097] A 3D mesh is a data structure which may describe the geometry or shape of an object related to oral care, including but not limited to a tooth, a hardware element, or a patient’s gum tissue. A 3D mesh may include one or more mesh elements such as one or more of vertices, edges, faces and combinations thereof. In some implementations, mesh elements may include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features may be computed for these mesh elements and be provided to the predictive models of this disclosure, with the predictive models of this disclosure providing the technical advantage of improving data precision in the form of the models of this disclosure outputting more accurate predictions.
[0098] One or more oral care meshes (e.g., teeth arranged in one or more dental arches) may be provided to techniques of this disclosure. Each of these meshes may undergo pre-processing before being provided to the predictive architectures (e.g., an encoder, decoder, pyramid encoder-decoder, or U-Net, among others). In some implementations, this pre-processing may include the encoding of the mesh into latent form. In some implementations, the pre-processing may include the conversion of the mesh into lists of mesh elements, such as vertices, edges, faces or in the case of sparse processing - voxels. For the chosen mesh element type or types, (e.g., vertices), feature vectors may be generated. In some examples, one feature vector is generated per vertex of the mesh (or per edge or per face). Each feature vector may contain a combination of spatial and / or structural features, as specified in Table 1.Table 1
[0099] Table 1 discloses non-limiting examples of mesh element features. In some implementations, color (or other visual cues / identifiers) may be considered as a mesh element feature in addition to the spatial or structural mesh element features described in Table 1. As used herein (e.g., in Table 1), a point differs from a vertex in that a point is part of a 3D point cloud, whereas a vertex is part of a 3D mesh and may have incident faces or edges. A dihedral angle (which may be expressed in either radians or degrees) may be computed as the angle (e.g., a signed angle) between two connected faces (e.g., two faces which are connected along an edge). A sign on a dihedral angle may reveal information about the convexity or concavity of a mesh surface. For example, a positively signed angle may, in some implementations, indicate a convex surface. Furthermore, a negatively signed angle may, in some implementations, indicate a concave surface. To calculate the principal curvature of a mesh vertex, directional curvatures may first be calculated to each adjacent vertex around the vertex. These directional curvatures may be sorted in circular order (e.g., 0, 49, 127, 210, 305 degrees) in proximity to the vertex normal vector and may comprise a subsampled version of the complete curvature tensor. Circular order means: sorted in by angle around an axis. The sorted directional curvatures may contribute to a linear system of equations amenable to a closed form solution which may estimate the two principal curvatures and directions, which may characterize the complete curvature tensor. Consistent with Table 1, a voxel may also have features which are computed as the aggregates of the other mesh elements (e.g., vertices, edges and faces) which either intersect the voxel or, in some implementations, are predominantly or fully contained within the voxel. Rotating the mesh may not change structural features but may change spatial features. And, as described elsewhere in thisdisclosure, the term “mesh” should be considered in a nonlimiting sense to be inclusive of 3D mesh, 3D point cloud and 3D voxelized representation.
[0100] The 3D representations of the patient’s dentition may comprise mesh elements. Mesh element features (e.g., spatial or structural mesh element features described in Table 1, such as “Curvature” or location coordinates such as “XYZ position”) may be computed for the mesh elements, and subsequently be provided to the first ML module. The mesh element features may improve the ability of a neural network (e.g., encoder 536, 618, 616 or 606, among other neural networks described herein) to encode the shape and / or structure of the teeth (or other aspects of the dentition or the face), thereby improving the accuracy of those latent representations. The latent representations of the teeth, mouth, or face (e.g., encoded as latent vectors or embedding vectors of lower dimensionality than the input data) may be concatenated using a variety of concatenation techniques. The concatenated latent vectors may, in some implementations, be concatenated with one or more oral care arguments. The concatenated latent vectors may subsequently be provided to a UNet which has been trained to denoise images or 3D representations. Stated another way, one or more latent representations of the patient’s dentition may be generated by an encoder (e.g., encoder 536), be concatenated with vectors which contain other inputs (e.g., oral care arguments), and subsequently be provided to flow-based ML model for smile prediction. In some implementations, the flow-based ML model may contain one or more denoising diffusion probabilistic models (e.g., a U-Net which may iteratively denoise an initially noisy data structure, which was trained on a series of iteratively noisier examples of data from a training dataset).
[0101] Representation generation neural networks based on autoencoders, U-Nets, transformers, other types of encoder-decoder structures, convolution and / or pooling layers, or other models may benefit from the use of mesh element features. Mesh element features may convey aspects of a 3D representation’s surface shape and / or structure to the neural network models of this disclosure. Each mesh element feature describes distinct information about the 3D representation that may not be redundantly present in other input data that are provided to the neural network. For example, a vertex curvature may quantify aspects of the concavity or convexity of the surface of a 3D representation which would not otherwise be understood by the network. Stated differently, mesh element features may provide a processed version of the structure and / or shape of the 3D representation, data that would not otherwise be available to the neural network. This processed information is often more accessible, or more amenable for encoding by the neural network. A system implementing the techniques disclosed herein has been utilized to run a number of experiments on 3D representations of teeth. For example, mesh element features have been provided to a representation generation neural network which is based on a U- Net model, and also to a representation generation model based on a variational autoencoder with continuous normalizing flows. Based on experiments, it was found that systems using a full complement of mesh element features (e.g., “XYZ” coordinates tuple, “Normal vector”, “Vertex Curvature”, Points-Pivoted, and Normals- Pivoted) were at least 3% more accurate than systems that did not. Points-Pivoted describes “XYZ” coordinates tuples that have local coordinate systems (e.g., at the centroid of the respective tooth). Normals-Pivoteddescribes “Normal Vectors” which have local coordinate systems (e.g., at the centroid of the respective tooth). Furthermore, training converges more quickly when the full complement of mesh element features are used. Stated another way, the machine learning models trained using the full complement of mesh element features tended to be more accurate more quickly (at earlier epochs) than systems which did not. For an existing system observed to have a historical accuracy rate of 91%, an improvement in accuracy of 3% reduces the actual error rate by more than 30%, which, as would be apparent to a person of ordinary skill in the art, represents a significant overall increase in the predictive accuracy of the models disclosed here.
[0102] The neural networks of the present disclosure may embody part or all of a variety of different neural network models. Examples include the U-Net architecture, multi-later perceptron (MLP), transformer, pyramid architecture, recurrent neural network (RNN), autoencoder, variational autoencoder, regularized autoencoder, conditional autoencoder, capsule network, capsule autoencoder, stacked capsule autoencoder, denoising autoencoder, sparse autoencoder, conditional autoencoder, long / short term memory (LSTM), gated recurrent unit (GRU), deep belief network (DBN), deep convolutional network (DCN), deep convolutional inverse graphics network (DCIGN), liquid state machine (LSM), extreme learning machine (ELM), echo state network (ESN), deep residual network (DRN), Kohonen network (KN), neural Turing machine (NTM), or generative adversarial network (GAN). In some implementations, an encoder structure or a decoder structure may be used. Each of these models provides one or more of its own particular advantages. For example, a particular neural networks architecture may be especially well suited to a particular ML technique. For example, UNets are particularly suited to the task of denoising a data structure, due to the ability to predict noise tensors.
[0103] Systems of this disclosure may implement end-to-end training. Some of the end-to-end trainingbased techniques of this disclosure may involve two or more neural networks (e.g., diffusion conditioning adapter 534, encoder 536 and / or a UNet included in denoising diffusion ML model 522, among others), where the two or more neural networks are trained together (i.e., the weights are updated concurrently during the processing of each batch of input oral care data). End-to-end training may, in some implementations, be applied to smile prediction by concurrently training a neural network which learns a representation of the patient's dentition, along with a neural network (e.g., a UNet) which performs a denoising operation.
[0104] According to some of the transfer learning-based implementations of this disclosure, a neural network (e.g., a U-Net) may be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task may be executed to provide one or more of the starting neural network weights for the training of another neural network that is trained to perform a second task (e.g., denoising diffusion). The first network may learn the low-level neural network features of oral care meshes and be shown to work well at the first task. The second network may exhibit faster training and / or improved performance by using the first network as a starting point in training. Certain layers may be trained to encode neural network features for the oral care meshes that were in the training dataset. These layers may thereafter be fixed (or be subjected to minor changes over the course of training) and be combined with other neural network components, such as additionallayers, which are trained for one or more oral care tasks (such as setups prediction). In this manner, a portion of a neural network for one or more of the techniques of the present disclosure (e.g., setups prediction) may receive initial training on another task, which may yield important learning in the trained network layers. This encoded learning may then be built upon with further task-specific training of another network.
[0105] Oral care arguments may include oral care parameters as disclosed herein, or other real-valued, textbased or categorical inputs which specify intended aspects of the one or more predicted smiles which are to be generated (e.g., a percentage of whiteness to increase or decrease in one or more teeth of the patient's dentition). In some instances, oral care arguments may include oral care metrics, which may describe intended aspects of the one or more 3D oral care representations which are to be generated. Oral care arguments are specifically adapted to the implementations described herein. For example, the oral care arguments may specify the intended designs (e.g., including which teeth are to be restored, or one or more intended oral care metric values which should be evident in the patient's dentition at the completion of orthodontic treatment) of smile predictions which may be generated (or modified) according to techniques described herein. In short, implementations using the specific oral care arguments disclosed herein generate more accurate smile predictions than implementations that do not use the specific oral care arguments.
[0106] Oral care arguments may include oral care parameters, preferences, oral care metrics, and other values (e.g., real values, text, enumerations, vectors, images, 3D representations, etc.) which are intended to influence the output of the generative techniques of this disclosure.
[0107] Oral care parameters are intended as instructions and / or specifications which describe the shape, structure, or other intended aspects of a 3D oral care representation which is to be generated using techniques of this disclosure. Oral care parameters may include orthodontic procedure parameters (OPP), restoration design parameters (RDP), or among others. Oral care parameters may define one or more intended aspects of a 3D oral care representation which is to be generated using techniques of this disclosure, and may be provided to an ML model to promote that ML model to generate output which may be used in the generation of oral care appliances that are suitable for the treatment of a patient. Non-limiting examples of orthodontic procedure parameters include: Teeth To Move: {AnteriorsOnly, AnteriorsAndBicuspids, FullArch}, Spacing: {CloseAllSpaces, LeaveSpecificSpaces}, Resolve Lower Crowding by IPR - Posterior Right: {Primarily, AsNeeded, None}, or Resolve Lower Crowding by IPR - Posterior Left: {Primarily, AsNeeded, None}, among others. Doctor preferences may pertain to a clinician’s past treatment practices, whereas oral care parameters may pertain to the treatment of a particular patient. Restoration Design Parameters (RDP) may, in some implementations, specify at least one value which defines at least one aspect of planned dental restoration treatment for the patient (e.g., specifying desired target attributes of a tooth which is to undergo treatment with a dental restoration appliance). Doctor Restoration Design Preferences (DROP) may, in some implementations, specify at least one typical value for an RDP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners. Restoration design parameters (RDP) may be used to encode aspects of smile designguidelines, such as parameters which pertain to the intended dimensions of a restored tooth. Nonlimiting examples of restoration design parameters include Tooth width at base (mesial to distal distance) in mm, Overall tooth shape {rectangular or ovoid, squared edges or rounded edges, etc.}, Amount of tooth display when the lips are at rest in mm, Tooth morphology - shape style guide {triangular, oval and rectangular, etc.}, Tooth morphology - Mamelon grooves {mamelon_style01, mamelon_style02, mamelon_style03, etc.}, Tooth morphology - perikymata [perikymata_style01, perikymata_style02, perikymata_style03, etc.}, among others. Other types of oral care arguments include doctor preferences (DP), restoration design preferences (RDP), or other types of preferences (e.g., preferences which pertain to the designs or specifications of fixture models or oral care appliances). Doctor preferences, restoration design preferences, or other types of preferences may define the typical treatment choices or practices of a particular clinician. Restoration design preferences are subjective to a particular clinician, and so differ from restoration design parameters. In some implementations, DP, RDP, or other preferences may be computed by unsupervised means, such as clustering, which may determine the typical values that a clinician uses in patient treatment. Oral care arguments may include oral care metrics. Oral care metrics may include orthodontic metrics (which may measure physical relationships between two or more teeth), restoration design metrics (which may measure physical relationships between two or more teeth, or may quantify physical aspects of particular tooth), or other types of metrics (which may measure aspects of an existing or generated 3D representation of oral care data). Orthodontic metrics may be used to quantify the physical arrangement of an arch of teeth for the purpose of orthodontic treatment or for other oral care treatments (e.g., fixture model generation or appliance component generation). These orthodontic metrics can measure how badly maloccluded the arch is, or conversely the metrics can measure how correctly arranged the teeth are. In some implementations, one or more orthodontic metrics may be taken from this section and incorporated into a loss computation, to quantify patterns of errors or deficiencies which may appear in predicted outputs. Within- arch orthodontic metrics includes: Alignment - A 3D tooth orientation vector may be calculated using the tooth's mesial-distal axis. Canine Overbite - A distance may be computed between the upper canine and the lower canine on a given side, and / or between the upper pre-molar and the corresponding lower pre-molar. Leveling - The difference in height between two or more neighboring teeth.Midline - May compute the position of the midline for the upper incisors and / or the lower incisors, and then may compute the distance between them.Overjet - The upper and lower central incisors may be compared along the y-axis. The difference along the y- axis may be used as the overjet score.Many other orthodontic metrics are also possible.
[0108] The following restoration design metrics (RDM) may be measured and used in the generation of crowns, dental restoration appliances, veneers (veneers are a type of dental restoration appliance), or the like, with the objective of making the resulting teeth natural looking. Symmetry is generally a preferred facet. Shade and translucency may pertain, in particular, to the generation of crowns, though some implementations of dentalrestoration appliances may also consider this information (e.g., when a succession of dental restoration appliances is used to form nested veneers with at least one inner structure). Examples of inter-tooth RDM are enumerated as follows: si) Bilateral Symmetry and / or Ratios: A measure of the symmetry between one or more teeth and one or more other teeth on opposite sides of the dental. For example, for a pair of corresponding teeth, a measure of the width of each tooth. s2) Proportions of Adjacent Teeth: Measure the width proportions of adjacent teeth as measured as a projection along an arch onto a plane (e.g., a plane that is situated in front of the patient's face). The ideal proportions for use in the final restoration design can be, for example, the so-called golden proportions. r3) Arch Discrepancies: A measure of any size discrepancies between the upper arch and lower arch, for example, pertaining to the widths of the teeth, for the purpose of dental restoration. r4) Midline: A measure of the midline of the maxillary incisors, relative to the midline of the mandibular incisors. Techniques of this disclosure may measure the midline of the maxillary incisors, relative to the midline of the nose (if data about nose location is available). r5) Proximal Contacts: A measure of the size (area, volume, circumference, etc.) of thesproximal contact between adjacent teeth. In the ideal circumstance, the teeth touch along the mesial / distal surfaces and the gums fill in gingivally to where the teeth touch.6) Embrasure: In some implementations, techniques of this disclosure may measure the size (area, volume, circumference, etc.) of an embrasure, the gap between teeth at either of the gingival or incisal edge.Examples of Intra-tooth RDM are enumerated below, continuing with the numbering of other RDM listed above.7) Length and / or Width: A measure of the length of a tooth relative to the width of that tooth. This metric may reveal, for example, that a patient has long central incisors. Width and length are defined as: a) width - mesial to distal distance; b) length - gingival to incisal distance; c) other dimensions of tooth body - the portions of tooth between the gingival region and the incisal edge.8) Tooth Morphology: A measure of the primary anatomy of the tooth shape, such as line angles, buccal contours, and / or incisal angles and / or embrasures. The frequency and / or dimensions may be measured. d9) Shade and / or Translucency: A measure of tooth shade and / or translucency. Tooth shade is often described by the Vita Classical or 3D Master shade guide. Tooth translucency is described by transmittance or a contrast ratio. Tooth shade and translucency may be evaluated (or measured) based on one or more of the following kinds of data pertaining to teeth: the incisal edge, incisal third, body and gingival third. The enamel layer translucency is general higher than the dentin or cementum layer. Shade and translucency may, in some implementations, be measured on a per-voxel (local) basis. Shade and translucency may, in some implementations, be measured on a per-area basis, such as an incisal area, tooth body area, etc. Tooth body may pertain to the portions of the tooth between the gingival region and the incisal edge.dlO) Height of Contour: A measure of the contour of a tooth. When viewed from the proximal view, all teeth have a specific contour or shape, moving from the gingival aspect to the incisal. This is referred to as the facial contour of the tooth. A patient’s dentition may include one or more 2D representations of the patient's anatomy (including mouth, gums, teeth, and / or full face), or one or more 3D representations of the patient’s teeth (e.g., and / or associated transforms), gums and / or other oral anatomy. Anorthodontic metric (OM) may, in some implementations, quantify the relative positions and / or orientations of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. A restoration design metric (RDM) may, in some implementations, quantify at least one aspect of the structure and / or shape of a 3D representation of a tooth. An orthodontic landmark (OL) may, in some implementations, locate one or more points or other structural regions of interest on a 3D representation of a tooth. An OL may, in some implementations, be provided to the generation of an orthodontic or dental appliance, such as a clear tray aligner or a dental restoration appliance. A mesh element may, in some implementations, comprise at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth that is represented by a 3D mesh, mesh elements may include at least: vertices, edges, faces and voxels. A mesh element feature may, in some implementations, quantify some aspect of a 3D representation in proximity to or in relation with one or more mesh elements, as described elsewhere in this disclosure. Orthodontic procedure parameters (OPP) may, in some implementations, specify at least one value which defines at least one aspect of planned orthodontic treatment for the patient (e.g., specifying desired target attributes of a final setup in final setups prediction). Orthodontic Doctor preferences (ODP) may, in some implementations, specify at least one typical value for an OPP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners. Restoration Design Parameters (RDP) may, in some implementations, specify at least one value which defines at least one aspect of planned dental restoration treatment for the patient (e.g., specifying desired target attributes of a tooth which is to undergo treatment with a dental restoration appliance). Doctor Restoration Design Preferences (DROP) may, in some implementations, specify at least one typical value for an RDP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners. Losses may also be used to train encoder structures, or decoder structures, among others. A KL-Divergence loss may be used, at least in part, to train one or more of the neural networks of the present disclosure, with the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior may enable a reconstruction autoencoder to produce a better reconstruction (e.g., when a latent vector representation is modified and that modified latent vector is reconstructed using a decoder, the resulting reconstruction is more likely to be a valid instance of the provided representation). There are other techniques for computing losses which may be described elsewhere in this disclosure. Such losses may be based on quantifying the difference between two or more 3D representations.
[0109] MSE loss calculation may involve the calculation of an average squared distance between two sets, vectors or datasets. MSE may be generally minimized. MSE may be applicable to a regression problem, where the prediction generated by the neural network or other machine learning model may be a real number. In some implementations, a neural network may be equipped with one or more linear activation units on the output to generate an MSE prediction. MAE loss and MAPE loss can also be used in accordance with the techniques of this disclosure.
[0110] Cross entropy may, in some implementations, be used to quantify the difference between two or more distributions. Cross entropy loss may, in some implementations, be used to train the neural networks of the present disclosure. Cross entropy loss may, in some implementations, involve comparing a predicted probability to a ground truth probability. Other names of cross entropy loss include “logarithmic loss,” “logistic loss,” and “log loss”. A small cross entropy loss may indicate a better (e.g., more accurate) model. Cross entropy loss may be logarithmic. Cross entropy loss may, in some implementations, be applied to binary classification problems. In some implementations, a neural network may be equipped with a sigmoid activation unit at the output to generate a probability prediction. In the case of multi-class classifications, cross entropy may also be used. In such a case, a neural network trained to make multi-class predictions may, in some implementations, be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for class that is to be predicted). Other loss calculation techniques which may be applied in the training of the neural networks of this disclosure include one or more of: Huber loss, Hinge loss, Categorical hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss calculation methods are described herein and may be applied to the training of any of the neural networks described in the present disclosure. Stated another way, in some instances, the forward pass 206 of the denoising diffusion model may generate a training dataset of increasingly noisy examples of the input data. Noise may be introduced to disfigure the input data (e.g., an image, a 3D point cloud, a transform, or latent representations of one or more of these inputs, etc.) and those noisy examples may be used, at least in part, to train a denoising diffusion machine learning model 210 to reverse of this noise-introducing process (e.g., in model deployment). The reverse pass 208 may be trained to reconstruct the pristine input data by removing the noise from a noisy example of that input data. This reverse pass 208 may generate a new output based upon a data structure of an initially random configuration, or may output a modified version of an initial data structure which is provided to the input of the reverse pass 208. Stated another way, after the denoising ML model 210 (e.g., a U-Net) is trained, the model is capable to generate new 3D oral care representations by passing a noisy data example (e.g., a randomly generated noisy example) through the denoising diffusion process (aka the reverse process). Furthermore, the trained denoising ML model 210 is capable to modify an initial 3D oral care representation which is provided to the input of the reverse pass 222, and output a resulting 3D oral care representation which is suitable for use in generating an oral care appliance which is customized to the treatment needs of the patient. In some implementations, the denoising ML model 210 of the reverse pass 208 may modify an existing 3D oral care representation (e.g., modify a pre-restoration tooth design). For instance, the existing 3D oral care representation (e.g., an example of optional instant patient case data 216) can be provided to the input of the reverse pass, and then undergo a succession of denoising steps by the denoising ML model 210, until modifications are complete (e.g., as measured by oral care metrics, loss functions, or a threshold number of iterations has expired, etc.). Optional instant patient case data 216 may include data pertaining to the dentition of a patient. The instant patient case data may, in some implementations, be introduced to customize the functioning of the denoising diffusionML model to the anatomy of the patient. In some implementations, the instant patient case data may be encoded 218 into a latent or embedded form. In some implementations, the instant patient case data may be provided to the denoising diffusion ML model 210. In other implementations, the instant data 216 may comprise an appliance component or a fixture model component which requires modification. In some implementations, which data structure that undergoes iterative denoising by the denoising ML model 210 may be initialized, at least in part, according to a stochastic process (e.g., by introducing random noise, normally distributed noise, or random configurations), or by aspects of the instant patient case data 216, or by a combination of the two. The instant patient case data 216 may include 3D oral care representations described herein, including tooth transforms (e.g., maloccluded transforms for one or more teeth during setups prediction), tooth meshes with transforms already applied, tooth meshes without transforms applied, one or more 3D representations of pre-restoration tooth designs, pre-segmentation dental arch mesh(es), pre-cleanup dental arch mesh(es), one or more mesh element labels (e.g., for segmentation or mesh cleanup), one or more segmented teeth for use in coordinate system prediction, one or more coordinate systems (e.g., each of which may be described by a transform), one or more 3D representations of appliance components (e.g., parting surfaces, etc.), one or more fixture model components (e.g., digital pontic teeth or interproximal webbing, etc.), or the like.
[0111] The denoising diffusion ML model 210 may be trained to generate one or more generated 3D oral care representations 220 (e.g., as defined herein) for the patient's treatment. A generated 3D oral care representation 220 (aka a denoised representation) may be generated over one or more iterations of denoising by the denoising diffusion ML model 210. Generated 3D oral care representations 220 may include setups transforms for one or more teeth, transforms for the placement of one or more appliance components (e.g., for the generation of a dental restoration appliance), one or more 3D representations of post-restoration tooth designs, one or more generated (or modified) appliance components (e.g., a parting surface or gingival ribbon for use in generating a dental restoration appliance), one or more generated (or modified) fixture model components (e.g., one or more trimlines, etc.), post-segmentation dental arch mesh(es), post-cleanup dental arch mesh(es), one or more mesh element labels for use in segmentation or mesh cleanup, one or more object masks (e.g., masks to be applied to mesh elements) for use in segmentation or mesh cleanup, one or more coordinate axes for one or more predicted coordinate systems, one or more archforms, or other of the 3D oral care representations described herein.
[0112] Techniques of this disclosure may require a training dataset of hundreds or thousands of cohort patient cases, to ensure that the neural network is able to encode the distribution of patient cases which are likely to be encountered in clinical treatment. A cohort patient case may include a set of tooth crown meshes, a set of tooth root meshes, or a data file containing attributes of the case (e.g., a JSON file). A typical example of a cohort patient case may contain up to 32 crown meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), multiple gingiva mesh (e.g., which may each contain tens of thousands of verticesor tens of thousands of faces) or one or more JSON files which may each contain tens of thousands of values (e.g., objects, arrays, strings, real values, Boolean values or Null values).
[0113] In some instances, a text encoder may encode a set of natural language instructions from the clinician (e.g., generate a text embedding). A text string may comprise tokens. An encoder for generating text embeddings may, in some implementations, apply either mean-pooling or max-pooling between the token vectors. In some instances, a transformer (e.g., BERT or Siamese BERT) may be trained to extract embeddings of text for use in digital oral care (e.g., by training the transformer on examples of clinical text, such as those given below). In some instances, such a model for generating text embeddings may be trained using transfer learning (e.g., initially trained on another corpus of text, and then receive further training on text related to digital oral care). Some text embeddings may encode text at the word level. Some text embeddings may encode text at the token level. A transformer for generating a text embedding may, in some implementations, be trained, at least in part, with a loss calculation which compares predicted outputs to ground truth outputs (e.g., softmax loss, multiple negatives ranking loss, MSE margin loss, cross-entropy loss or the like). In some instances, the non-text arguments, such as real values or categorical values, may be converted to text, and subsequently embedded using the techniques described herein. The following are examples of natural language instructions that may be issued by a clinician to the generative models described herein: “Generate a setup to set to Class I molar and canine, 2 mm overbite and add 2mm of expansion 5-575-5”, “Generate a setup to align with proclination and expansion and finish with .5 mm spaces U2-2 for future restorative”, or “Adjust the setup with no second molar movement, rotate upper first molars mesial out for Class I, level lower to a reverse curve of Spee 2 mm, advance the mandible with elastics to Class I canine.” The resulting predicted setup may then be integrated with a 2D or 3D representation of the patient's face, resulting in one or more predicted setups.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method of a generating a prediction of a post-treatment dental anatomy of a patient comprising: receiving, by one or more computer processors, one or more representations of the pre-treatment anatomy of the patient which shows one or more exposed teeth; receiving, by one or more computer processors, one or more representations of the target posttreatment dentition of the patient; generating, by the one or more computer processors, one or more first representations of the intended output; denoising, by the one or more computer processors, and using a trained flow-based machine learning model, the one or more first representations of the intended output; generating, by the one or more computer processors, one or more second representations of the intended output; automatically defining, by the one or more computer processors and using the one or more generated second representations, one or more aspects of a post-treatment anatomy of a patient.
2. The method of claim 1, wherein one or more oral care arguments describing an intended output from a trained machine learning model is received.
3. The method of claim 1, wherein the flow-based machine learning model is a denoising diffusion machine learning model.
4. The method of claim 3, wherein the denoising diffusion model includes one or more neural networks which generate one or more data structures that including noise, and the one or more data structures that specify how the noise is removed from the one or more first representations of the intended output.
5. The method of claim 3, wherein the one or more representations of the pre-treatment anatomy of the patient is encoded into one or more pre-treatment latent representations and the one or more representations of the target post-treatment dentition of the patient is encoded into one or more latent representations of the target post-treatment dentition.
6. The method of claim 5, wherein the one or more representations of the pre-treatment anatomy include one or more 2D images of the patient’s anatomy, and the one or more representations of the target post-treatment dentition include one or more 3D representations of teeth.
7. The method of claim 6, further comprising: receiving, by the one or more processors, one or more tooth image masks that define the post-treatment shapes of one or more teeth of the patient’s pre-treatment anatomy; providing, by the one or more processors, the one or more tooth image masks to the trained denoising diffusion machine learning model; denoising, by the one or more processors, the one more tooth image masks to generate one or more denoised tooth image masks; and generating, by the one or more computer processors and the trained denoising diffusion machine learning model, one representations of teeth to be included in a predicted post-treatment dentition, , wherein the shapes of one or more representations of teeth are determined, at least in part, by the shapes of the one or more denoised tooth image masks.
8. The method of claim 1, wherein the one or more representations of the target post-treatment dentition of the patient includes one or more of one or more orthodontic setups, and one or more tooth restoration designs.
9. The method of claim 8, wherein the one or more representations of the target post-treatment dentition of the patient is generated by one or more of an automated setups prediction machine learning model, or an automated restoration design generation machine learning model.
10. The method of claim 3, further comprising: computing, by the one or more processors, one or more color palettes; providing, by the one or more processors, the one or more color palettes to the trained denoising diffusion machine learning model; and generating, by the one or more computer processors and the trained denoising diffusion machine learning model, one or more second representations of the intended output, containing one or more teeth, where the colors of one or more teeth are determined, at least in part, by the one or more color palettes.
11. The method of claim 1, further comprising: receiving, by the one or more processors, one or more camera parameters; computing, by the one or more processors and the one or more camera parameters, one or more 2D representations, based at least in part, on the one or more 3D representations of teeth, where the one or more 2D representations include one or more projected image masks;generating, by the one or more processors, one or more dentition masks, wherein the one or more dentition masks are generated, at least in part, by performing tooth segmentation on the one or more 2D images of the patient’s anatomy; and registering, by the one or more processors, the one or more projected image masks with corresponding aspects of the one or more dentition image masks.
12. The method of claim 11, further comprising: receiving one or more oral care arguments, wherein the registering is based at least in part on the received oral care arguments.
13. The method of claim 12, further comprising: determining one or more confidence scores that specify an accuracy of one or more tooth identifications that result from performing tooth segmentation on the one or more 2D images of the patient’s anatomy, wherein the registering is based at least in part on the confidence scores.
14. The method of claim 11, wherein the registration involves calculating, by the one or more processors, one or more cost functions.
15. The method of claim 14, further comprising: optimizing, by the one or mor processors, the one or more camera parameters using one or more optimization modules.
16. The method of claim 15, wherein the one or more optimization modules are configured to perform at least one of simulated annealing, one or more genetic algorithms, differential evolution, particle swarm analysis, and ant colony optimization.
17. The method of claim 1, wherein the one or more first representations include data that does not correlate with any of the one or more representations of the pre-treatment anatomy or the one or more representations of the target post-treatment dentition.
18. The method of claim 17, wherein after generating the one or more second representations, the data that does not correlate with any of the one or more representations of the pre-treatment anatomy or the one or more representations of the target post-treatment dentition is not present in the one or more second representations.
19. The method of claim 14, further comprising: generating, by the one or more processors, one or more data values that describe the alignment between upper and lower arches of the patient's 3D dentition; and optimizing, by the one or more processors, the data values using one or more optimization modules.
20. A system for generating a prediction of a post-treatment dental anatomy of a patient comprising: one or more computer processors; non-transitory computer-readable storage communicatively coupled to the one or more processors having instructions stored thereon that when executed by the one or more processors cause the one or more processors to: receive one or more representations of the pre-treatment anatomy of the patient which shows one or more exposed teeth; receive one or more representations of the target post-treatment dentition of the patient; generate one or more first representations of the intended output; denoise, using a trained flow-based machine learning model, the one or more first representations of the intended output; generate one or more second representations of the intended output; automatically define, using the one or more generated second representations, one or more aspects of a post-treatment anatomy of a patient.
Citation Information
Patent Citations
Providing a simulated outcome of dental treatment on a patient
WO2020005386A1
Cited By
Denture generation method and device based on digital intelligence
CN121562451A
Orthodontic treatment prediction system and method based on multi-modal fusion and diffusion flow, electronic equipment and storage medium
CN121885164A