Denoising Diffusion Model for Digital Oral Care
Denoising diffusion models enhance the accuracy and efficiency of generating and positioning oral care items by iteratively denoising 3D representations, addressing data conversion errors and resource inefficiencies in existing methods.
Patent Information
- Application Number
- JP2025534723
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-23
- Filing Date
- 2023-12-14
- Publication Date
- 2026-02-19
AI Technical Summary
Existing technologies face challenges in accurately generating and positioning oral care items relative to 3D representations of teeth, such as dental restoration prosthetic components, due to noise and inaccuracies in data conversion between different dimensions, leading to errors and increased resource consumption.
The use of denoising diffusion probabilistic models, specifically neural networks like U-Nets and autoencoders, to iteratively denoise noisy 3D oral care representations, enabling precise generation and positioning of oral care items by encoding and decoding data through latent formats, reducing resource usage and enhancing accuracy.
This approach improves data precision and accuracy in generating 3D oral care representations, allowing for automated and efficient production of dental appliances by directly processing 3D data without intermediate conversions, thus reducing computational costs and storage requirements.
Smart Images

Figure 2026505880000001_ABST
Abstract
Description
[Technical Field]
[0001] Related literature The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosures of PCT Applications WO2022123402, WO2021245480, and WO2020026117 are incorporated herein by reference. The entire disclosures of each of the following U.S. provisional patent applications are incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 366,514; 63 / 264,914 and 63 / 432,627.
[0002] The present disclosure relates to the construction and training of denoising diffusion models (e.g., neural networks trained for that purpose) to improve the accuracy and data precision of 3D oral care representations used in dental or orthodontic treatment. Summary of the Invention
[0003] This disclosure describes systems and techniques for training and using one or more denoising diffusion probabilistic models (e.g., neural networks such as U-Nets or autoencoders) to generate 3D oral care representations. A denoising diffusion-based technique is described for positioning oral care items relative to one or more 3D representations of teeth. The positioned oral care items may include 3D representations of teeth (e.g., for generating orthodontic setups), dental restoration prosthetic components, oral care hardware (e.g., lingual brackets, labial brackets, orthodontic attachments, bite lamps, etc.). Additionally, a denoising diffusion-based technique is described for generating the geometry and / or structure of oral care items based at least in part on one or more 3D representations of teeth. Oral care items that may be generated include dental restoration tooth designs, crowns, veneers, arch forms, clear tray aligner (CTA) trim lines, and prosthetic components (e.g., generated components for use in fabricating dental restorations, etc.). Among neural networks that can be trained to perform denoising diffusion, U-Net is an example of a model that may enable improved data accuracy. The denoising diffusion model can be trained to automatically generate (or modify) a 3D oral care representation, such as a 3D point cloud (or other 3D representation described herein). In some examples, the denoising diffusion model can be trained to generate a 3D polyline (e.g., for an arch form and CTA trim line) or a set of control points (e.g., control points to which a spline can be fitted to an arch form). The denoising diffusion model can be trained to predict an arch form (e.g., taking dental arch and tooth data as input). The arch form can take the form of a surface, a 3D mesh, a 3D polyline, or a set of control points (e.g., used to define a spline). Such an arch form, in some examples, can be provided as input to a setup prediction machine learning model, such as a setup prediction neural network.The denoising diffusion model may be trained to label aspects of the 3D representation (or to generate one or more object masks that identify one or more objects in the input 3D representation) for use in segmentation or mesh cleanup of the 3D oral care representation.
[0004] The techniques of the present disclosure include methods for generating a data structure for oral care treatments. Such methods may receive as input one or more attributes describing an intended output from a trained machine learning model. Such methods may generate one or more noisy representations of the intended output and then denoise the one or more noisy representations. The output of the method may include one or more generated denoised representations, which may be used to define one or more aspects of one or more digital oral care treatments. A training dataset may be generated or refined by successively modifying one or more 3D oral care representations. An untrained (or partially trained) machine learning model may be trained, at least in part, using the refined training dataset. The resulting trained machine learning model may be deployed for clinical treatment of patients. Throughout the process of generating a training dataset, one or more representations of the training dataset may be modified, such as through the addition of noise. In some implementations, one or more aspects of one or more representations of the training dataset may be encoded into a latent format before generating one or more noisy representations. The one or more attributes may include at least one of a real value, a categorical value, or a natural language text value. The one or more attributes may include at least one of oral care metrics or oral care parameters. The trained machine learning model may include at least one neural network. The at least one neural network may include at least one encoder-decoder structure. The encoder-decoder structure may include at least one encoder or at least one decoder. Non-limiting examples of encoder-decoder structures include U-Net, Transformer, Pyramid Encoder-Decoder, or Autoencoder, among others. The denoising diffusion technique may receive one or more 3D representations of the patient's dentition (e.g., which may include at least one tooth). The one or more 3D representations of the patient's dentition may be encoded into a latent format.The one or more denoised representations (also known as generated 3D oral care representations) may include at least one of the following: one or more 3D oral care representations as defined herein, at least one label of a mesh element, at least one transformation for placing at least one tooth of the patient's dentition in a setup pose, appliance components, trim lines, arch forms, coordinate systems, or dental restoration designs (among others). The labels on the mesh elements may be used to segment the one or more 3D oral care representations (e.g., to segment at least one representation of the patient's dentition). The labels may be used for cleanup of the 3D oral care representations (e.g., the labels may be used to modify one or more 3D oral care representations, such as a representation of the patient's dentition). The transformations predicted using the techniques described herein may place the teeth in a setup pose. In some implementations, the setup may correspond to a final setup. In some implementations, the setup may correspond to an intermediate stage. In some examples, the appliance components (e.g., positioned or generated using the denoised diffusion model) may be used to generate one or more dental restoration appliances.
[0005] The denoising diffusion techniques described herein may be implemented in combination in some examples. For example, a denoising diffusion mesh cleanup may be performed on a dental arch from an intraoral scanner, followed by a denoising diffusion segmentation on the cleaned-up 3D representation of the patient's dentition. In some examples, one or more of the segmented teeth may be identified for dental restoration treatment and then provided to a denoising diffusion model for geometry generation (e.g., a geometry generation diffusion model GGDM) to generate a dental restoration tooth design. In some implementations, the segmented teeth of the dental arch may be provided to a diffusion setup model for setup prediction. Other combinations of the techniques described herein should also be considered within the scope of this disclosure.
[0006] The techniques of the present disclosure may generate a data structure for an oral care treatment using a denoising diffusion probabilistic model. The method may receive oral care arguments (or oral care attributes) describing the intended output from the trained denoising diffusion model. The method may perform a forward pass to generate a series of increasingly noisy versions of each training example. In deployment, a denoising ML model (e.g., U-Net) may be trained to iteratively denoise an initial noisy data structure. After multiple denoising iterations, the data structure may be evolved into a 3D oral care representation suitable for use in treating a patient. In other words, one or more denoised representations are generated, and based on these representations, one or more aspects of the digital oral care treatment are automatically defined (e.g., a tooth restoration design is generated or an orthodontic setup is generated, among other examples).
[0007] The method may access a training dataset of 3D oral care representations and generate a refined training dataset by modifying the representations of the training dataset (e.g., iteratively adding noise). An untrained machine learning model may be trained using the refined training dataset, and a trained machine learning model is output based on the training.
[0008] Modifying the representations of the training dataset includes adding noise to aspects of the representations. Furthermore, the method encodes aspects of the representations into a latent form before generating the noisy representations.
[0009] The attributes describing the intended output may include real-valued values, categorical values, or natural language text values. Additionally, the attributes may include oral care metrics or oral care parameters.
[0010] The trained machine learning model can include a neural network such as a U-Net, a Transformer, or an Autoencoder (etc.). The method can receive a 3D representation of the patient's dentition, and the denoised representation can include a 3D oral care representation. These representations can be encoded into a latent format (e.g., if stable diffusion is used).
[0011] The method may also be trained to label mesh elements, apply those mesh element labels to 3D representations of oral care data, and segment those 3D representations based on the labels. The 3D oral care representations can include, among other things, a representation of the patient's dentition.
[0012] Additionally, the denoised representation can include a transformation (e.g., a 4x4 transformation matrix, among others) that can be used to place the patient's teeth in a setup pose (e.g., final setup or an intermediate stage). In some implementations, the transformation may place other 3D representations of oral care data in a pose appropriate for the oral care treatment. Other 3D representations of oral care data include fixture model components, dental restoration designs, prosthetic components, fixture model components, or archforms (etc.). Additionally, the denoised representation can include a coordinate system or other 3D oral care representations described herein.
[0013] In summary, the disclosed method utilizes a denoised diffusion probabilistic model to generate a data structure for use in an oral care treatment, allowing for the automatic definition of various aspects of a digital oral care treatment based on the trained model and the denoised representation.
[0014] In other words, the disclosed method may generate a data structure for use in a digital oral care procedure. The method may define one or more attributes (or oral care arguments) that describe the intended output from a trained machine learning model. The method may generate a training dataset by generating one or more noisy representations of the intended output (e.g., a series of increasingly noisy representations) and / or training a machine learning model using one or more noisy representations of the intended output. The method may use a trained model (e.g., a hierarchical neural network feature extraction model such as U-Net) to iteratively denoise a representation (e.g., a representation initially having a random, Gaussian, or other default value or distribution). The method may generate (or modify) one or more denoised representations of the intended output after one or more iterations of denoising. The method may use the denoised representations to generate one or more oral care appliances or use the denoised representations for one or more aspects of one or more digital oral care procedures. A refined dataset may be generated by successively modifying one or more representations of the training dataset. An initially untrained machine learning model may be trained using a refined training dataset and output for use in oral care appliance generation. The refined training dataset may be generated by adding noise to aspects of one or more representations of the training dataset. In some implementations, one or more aspects of one or more representations of the training dataset may be encoded into a latent form before generating one or more noisy representations. In other words, each example of the training dataset may first undergo latent encoding (e.g., using an encoder) before a refined dataset (of iteratively noisy examples) is generated. One or more oral care attributes (or oral care arguments) may describe the shape, structure, layout, or other characteristics of the intended output. The oral care arguments may include oral care metrics and / or oral care parameters. Other oral care attributes may include practitioner preferences. The one or more oral care attributes (or oral care arguments) may include at least one of real-valued values, categorical values, or natural-language text values.The trained machine learning model may include at least one neural network, such as a hierarchical neural network feature extraction model (e.g., U-Net), an autoencoder, or a transformer (e.g., a 3D SWIN transformer). One or more 3D representations of the patient's dentition (e.g., one or more teeth of the patient) may be provided to the method. The method may denoise the 3D oral care representation (e.g., the 3D oral care representation undergoing modification). In some implementations, the 3D oral care representation (e.g., the 3D representation of the patient's dentition) may be encoded into a latent format. In some examples, the one or more denoised representations include at least one or more labels of mesh elements (e.g., for use in mesh segmentation or mesh cleanup). The method may segment the 3D representation (e.g., a mesh, a point cloud, a voxelized representation, etc.) using the one or more labels. One or more 3D representations of the oral care data may be segmented or undergo mesh cleanup (e.g., the 3D representation of the patient's dentition). Mesh cleanup may include using one or more labels to modify one or more 3D oral care representations (e.g., removing mesh elements with particular label values or modifying mesh elements with particular label values). In some implementations, the method can denoise one or more representations of transformations (e.g., transformations that may position a tooth in a setup pose—such as a final setup or intermediate stage—that may position appliance components for appliance generation, or transformations that may position appliance model components for appliance model generation). In some implementations, the method can denoise one or more 3D representations of tooth restoration designs to modify the shape and / or structure of a pre-restoration tooth to make the tooth suitable for use in digital oral care treatment (e.g., for use in generating dental restoration appliances). In some implementations, the method can denoise one or more 3D representations of appliance components to make the one or more appliance components suitable for use in digital oral care treatment (e.g., for use in generating dental restoration appliances).In some implementations, the method can denoise one or more 3D representations of fixture model components, making the one or more fixture model components suitable for use in digital oral care treatment (e.g., for use in generating a digital fixture model). The digital fixture model can be 3D printed, resulting in a physical fixture model. Orthodontic aligner trays can be thermoformed using the physical fixture model and used in the patient's orthodontic treatment. In some implementations, the method can denoise one or more 3D representations of clear tray aligner trim lines, making the one or more clear tray aligner trim lines suitable for use in digital oral care treatment (e.g., for use in trimming aligner trays from the physical fixture model). In some implementations, the method can denoise one or more 3D representations of arch forms, making the one or more arch forms suitable for use in digital oral care treatment (e.g., for use in generating oral care appliances). In some implementations, the method can denoise one or more 3D representations of coordinate systems, making the one or more coordinate systems suitable for use in digital oral care treatments (e.g., for use in generating oral care appliances). [Brief explanation of the drawings]
[0015] [Figure 1] We demonstrate how to use a denoising diffusion probabilistic model to generate (or modify) 3D oral care representations. [Figure 2] We show how to train a denoising diffusion probabilistic model to generate (or modify) 3D oral care representations. [Figure 3] We demonstrate how to generate orthodontic setup transformations using a denoising diffusion probability model. [Figure 4] 1 illustrates a method for using a hierarchical neural network feature extraction module to perform mesh element labeling (e.g., for segmentation or mesh cleanup) of a 3D representation in accordance with the disclosed denoising diffusion probability model. [Figure 5]We show a U-Net structure that can be used to extract hierarchical features from a 3D representation. [Figure 6] 1 illustrates a pyramidal encoder-decoder structure that can be used to extract hierarchical features from a 3D representation. DETAILED DESCRIPTION OF THE INVENTION
[0016] Diffusion models can be applied to, among other things, 2D image generation or 3D representation generation. In some examples, such implementations can take input from natural language text (e.g., natural language text, real-valued values, categorical values, reference images or reference 3D representations describing intended results). The techniques described herein extend diffusion models to the digital oral care space. The techniques described herein use diffusion models to generate 3D oral care representations (e.g., 3D point clouds, 3D meshes, 3D voxelized representations, 3D surfaces, etc.) based on input arguments (which may include, for example, natural language text, integer arguments, real-valued arguments, categorical arguments, etc.). Such arguments may include one or more attributes that describe the intended output from a trained machine learning model. In some examples, the techniques described herein can take an input representation (e.g., a 3D point cloud, a 2D view, or a 2D digital photograph) of a patient's dentition to be used as a guide for generating artifacts used in digital oral care. The techniques described herein can, in some examples, take input representations of appliances, appliance components, trim lines, arch forms, or other 3D oral care representations (e.g., in the form of 3D point clouds, 2D views, or 2D digital photographs) that are modified or used as templates or references for artifacts generated for use in digital oral care. In some examples, these inputs (e.g., teeth, gums, appliances, appliance components, arch forms, trim lines for the refinement of manufactured trays such as printed or thermoformed trays, segmentation masks on mesh elements, or other 3D oral care representations) can be encoded into latent (e.g., information-rich, dimensionally reduced) form by an autoencoder or other encoder-decoder neural network. In some cases, trim lines can be generated that are used to separate thermoformed trays from fixture models (e.g., fixture models used in the manufacture of either indirect bonding trays or orthodontic aligners). In some examples, the techniques of the present disclosure can achieve improved data accuracy through the use of such latent representations.These techniques are beneficial to 3M's digital oral care platform, particularly for custom smile design in the form of dental restoration design generation, appliance component generation, aligner trim line generation, set-up prediction, and more.
[0017] The denoising diffusion model (e.g., shown in FIG. 2 ) may include a forward pass 206 on input data 200 (e.g., a 3D point cloud of teeth, a transformation, or another 3D oral care representation described herein), which may generate a Markov chain of step 204 that may introduce successively more noise (e.g., Gaussian noise) to the input data 200. The forward pass 206 may generate training data. The Markov chain may generate a set of successively noisier training data examples. Input oral care arguments (or attributes) 202 may affect the function of the denoising diffusion model and cause the denoising diffusion model to generate output to a clinician's specifications (e.g., allowing for customization of the output generated by the denoising diffusion model). The oral care arguments may include oral care parameters, oral care metrics, among other examples. In some implementations, such as using stable diffusion, an optional latent encoding module 214 may encode the 3D oral care representation 200 into a latent form. The latent encoding module may be trained to encode the data into a reduced-dimensionality latent form. Examples of data that may be encoded include a transformation (e.g., a tooth transformation), a point cloud or mesh describing a tooth, or a set of mesh element labels. Such data may be encoded into a latent vector or latent capsule. Similarly, the optional latent encoding module 212, in some implementations, may encode one or more of the oral care arguments 202 into a latent format.
[0018] The Markov chain can generate a series of increasingly noisy versions of the input data 200, which can then be used to at least partially train the denoising ML module 210 (e.g., which can function as part of the reverse pass 208). The denoising ML models trained for use in the reverse pass can include one or more neural networks 210 (e.g., U-Net, VAE, 3D SWIN transformer, pyramid encoder-decoder, etc.) and can be trained to denoise noisy versions of the data (e.g., starting with a fully randomized version of the input data structure and denoising that data structure until it assumes the appearance of the appropriate generated output). The reverse pass 208 can be used during model deployment (e.g., after deployment of the trained denoising ML model 210). For example, the inverse pass 208 may start with a point cloud (or other raw data structure, such as a transform or mesh element labels) having a Gaussian distribution and, over successive applications of the inverse pass denoising ML model 210, gradually shape the point cloud (or other data structure) into a suitable example of a target 3D oral care representation (e.g., a tooth design suitable for use in creating dental restorations). In some implementations, a loss function (e.g., cross-entropy or MSE, among others) may be calculated to quantify the difference between the generated 3D oral care representation and a corresponding ground truth (or reference) 3D oral care representation. In some implementations, the loss function may be used to at least partially train the denoising ML model 210. In some implementations, oral care metrics may be calculated for the generated 3D oral care representation 220, such as when an orthodontic setup or tooth restoration design is generated. The oral care metrics may be used, at least in part, to assess the quality of the generated 3D oral care representation's fitness for use (e.g., to measure whether the generated 3D oral care representation meets the specifications of the oral care arguments 202 and is ready for use in generating oral care appliances).
[0019] Stated another way, in some examples, the denoising diffusion model's forward pass 206 may generate a training data set of increasingly noisy examples of input data. Noise may be introduced to corrupt the input data (e.g., images, 3D point clouds, transformations, or latent representations of one or more of these inputs), and these noisy examples may be used, at least in part, to train the denoising diffusion machine learning model 210 to reverse this noise-introduction process (e.g., in model deployment). The reverse pass 208 may be trained to reconstruct the original input data by removing noise from the noisy examples of that input data. After the denoising diffusion model (e.g., U-Net) is trained, the model can generate new 3D oral care representations by passing noisy data examples (e.g., randomly generated noisy examples) through the denoising diffusion process (also known as the inverse process).
[0020] The techniques of the present disclosure may require a training dataset of hundreds or thousands of cohort patient cases to ensure that the neural network can encode the distribution of patient cases likely to be encountered in clinical care. A cohort patient case may include a set of crown meshes, a set of root meshes, or a data file (e.g., a JSON file) containing case attributes. A typical example of a cohort patient case may include up to 32 crown meshes (e.g., each may contain tens of thousands of vertices or faces), up to 32 root meshes (e.g., each may contain tens of thousands of vertices or faces), multiple gingival meshes (e.g., each may contain tens of thousands of vertices or faces), or one or more JSON files, each of which may contain tens of thousands of values (e.g., objects, arrays, strings, real values, Boolean values, or null values).
[0021] In some implementations, the denoising ML model 210 of the reverse pass 208 may modify an existing 3D oral care representation (e.g., modify a pre-restoration tooth design). The existing 3D oral care representation (e.g., an example of optional instant patient case data 216) is provided to the input of the reverse pass and can then go through a series of denoising steps by the denoising ML model 210 until modification is complete (e.g., as measured by an oral care metric, a loss function, or the expiration of a threshold number of iterations). The optional instant patient case data 216 may include data regarding the patient's dentition. In some implementations, the instant patient case data may be introduced to customize the functionality of the denoising diffusion ML model to the patient's anatomy. In some implementations, the instant patient case data may be encoded in a latent or embedded format (218). In some implementations, the instant patient case data may be provided to the denoising diffusion ML model 210. In other implementations, the instant data 216 may include appliance components or fixture model components that require modification.
[0022] In some implementations, the data structure that undergoes iterative denoising by the denoising ML model 210 may be initialized, at least in part, according to a stochastic process (e.g., using random noise or random configurations), or by aspects of the instant patient case data 216, or a combination of the two. The instant patient case data 216 may include the 3D oral care representations described herein, including tooth transformations (e.g., malocclusion transformations of one or more teeth during setup prediction), tooth meshes with transformations already applied, tooth meshes without transformations applied, one or more 3D representations of a restoration anterior tooth design, a segmentation anterior arch mesh, a cleanup anterior arch mesh, one or more mesh element labels (e.g., for segmentation or mesh cleanup), one or more segmented teeth for use in coordinate system prediction, one or more coordinate systems (e.g., each of which may be described by a transformation), one or more 3D representations of appliance components (e.g., parting surfaces, etc.), one or more fixture model components (e.g., digital pontic or interproximal webbing, etc.), etc.
[0023] The denoising diffusion ML model 210 may be trained to generate one or more generated 3D oral care representations 220 (e.g., as defined herein) for treating a patient. The generated 3D oral care representations 220 (also known as denoised representations) may be generated over one or more iterations of denoising by the denoising diffusion ML model 210. The generated 3D oral care representation 220 may include one or more tooth setup transformations, transformations for the placement of one or more appliance components (e.g., for generating a dental restoration), one or more 3D representations of a post-restoration tooth design, one or more generated (or modified) appliance components (e.g., parting surfaces or gingival ribbons for use in generating a dental restoration), one or more generated (or modified) fixture model components (e.g., one or more trim lines, etc.), a post-segmentation dental arch mesh, a post-cleanup dental arch mesh, one or more mesh element labels for use in segmentation or mesh cleanup, one or more object masks (e.g., masks applied to mesh elements) for use in segmentation or mesh cleanup, one or more coordinate axes of one or more predicted coordinate systems, one or more arch forms, or others of the 3D oral care representations described herein.
[0024] The machine learning techniques described herein may receive various input data, including tooth meshes of one or both dental arches of a patient, as described herein. Tooth data, appliance components, and fixture model components may be provided in the form of 3D representations, such as meshes, point clouds, or voxelized shapes, to name a few. These data may be preprocessed, for example, by arranging the constituent mesh elements into a list and calculating an optional mesh element feature vector for each mesh element. Such vectors can provide valuable information about the shape and / or structure of the oral care mesh to the machine learning models described herein. Additional inputs, such as one or more oral care metrics, may be received as inputs to the machine learning models described herein. Oral care metrics may be used to measure one or more physical aspects of the oral care mesh (e.g., physical relationships within or between different teeth). In some examples, oral care metrics may be calculated for either or both of the example malocclusion oral care mesh and / or the example ground truth oral care mesh, which are then used in training the machine learning models described herein. Metric values may be received as input to the machine learning models described herein, or as methods for training such models, to encode the distribution of such metrics across examples in a training dataset. During training, the network may then receive metric values as input to help train the network to link the input metric values to physical aspects of the ground truth oral care mesh used in a loss calculation. Such loss calculation may quantify the difference between the predictions and the ground truth examples (e.g., between the predicted oral care mesh and the ground truth oral care mesh). By providing the network with metric values, the neural network techniques of the present disclosure may, through a process of loss calculation and subsequent backpropagation, train the neural network to encode the distribution of a given metric.In deployment, one or more oral care arguments may be defined to specify one or more aspects of the intended 3D oral care representation (e.g., a 3D mesh, polylines, 3D point cloud, or voxelized geometry), which are generated using a machine learning model (e.g., a denoising diffusion model) described herein trained for that purpose. In some implementations, the oral care arguments may be defined to specify one or more aspects of a customized vector, matrix, or any other numerical representation (e.g., describing the 3D oral care representation, such as a spline, an archform, a transformation for positioning teeth or prosthetic components relative to another 3D oral care representation, or control points for a coordinate system), which are generated using a machine learning model (e.g., a denoising diffusion model) described herein trained for that purpose. The customized vector, matrix, or other numerical representation may describe a 3D oral care representation that matches the intended outcome of the patient's treatment. The oral care arguments may include, among other things, oral care metrics or oral care parameters. The oral care arguments may specify one or more aspects of the oral care procedure, such as orthodontic setup prediction or restoration design generation, among other things. In some implementations, one or more oral care parameters corresponding to individual oral care metrics may be defined. The oral care arguments may be provided as inputs to the machine learning models described herein and interpreted as instructions to that module to generate an oral care mesh with the specified customizations, position the oral care mesh for generation of an orthodontic setup (or appliance), segment the oral care mesh, or clean up the oral care mesh, to name a few. This interplay between oral care metrics and oral care parameters may be similarly applied to the training and deployment of other predictive models in oral care.
[0025] In some implementations, the predictive model of the present disclosure may generate more accurate results by incorporating one or more of the following inputs: arch form information V, interproximal reduction (IPR) information U, tooth dimension information P, tooth gap information Q, latent capsule representation of the oral care mesh T, latent vector representation of the oral care mesh A, treatment parameters K (which may describe the clinician's intended patient treatment), doctor preferences L (which may describe typical treatment parameters selected by the doctor), tooth condition flags M (such as fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth names / dental notations R, oral care metrics S (including at least one of oral care metrics and restoration design metrics).
[0026] The systems of the present disclosure, in some examples, can be deployed in a clinical setting (such as a dental or orthodontic office) for use by a clinician (e.g., a doctor, dentist, orthodontist, nurse, hygienist, oral care technician). Such systems deployed in a clinical setting can enable a clinician to process oral care data (such as dental scans) in an office environment, or in some examples, in a "chairside" setting (when the patient is present in the clinical environment). A non-limiting list of example techniques can include segmentation (e.g., diffusion segmentation, etc.), mesh cleanup, coordinate system prediction, CTA trimline generation, restoration design generation (e.g., using a diffusion model for 3D point cloud generation, etc.), appliance component generation or placement or assembly (e.g., using a diffusion model for 3D point cloud generation, etc.), generation of other oral care meshes, oral care mesh validation, setup prediction (e.g., diffusion setup, etc.), hardware removal from tooth meshes, hardware placement on teeth, missing value imputation, clustering on oral care data, oral care mesh classification, setup comparison, metric calculation, or metric visualization. Implementation of these techniques may, in some instances, allow patient data to be processed, analyzed, and used in appliance creation by clinicians before the patient leaves the clinical setting (which may facilitate treatment planning as feedback may be received from the patient during the treatment planning process).
[0027] Some existing techniques may use a diffusion model for crown restoration. Such techniques may convert a 3D point cloud of the tooth into a 2D depth map, process the depth map using a diffusion model, and then attempt to convert the resulting modified depth map back into a 3D point cloud of the tooth. Attempting to convert between 3D and 2D representations in this manner may introduce errors. In other words, noise may be introduced when data is converted between representations of different dimensions (e.g., from 2D to 3D or from 3D to 2D). The denoising diffusion probability model of the present disclosure can be trained to directly generate a 3D tooth restoration design (or other 3D representations described herein, such as fixture model components or appliance components) without the intervening step of depth map generation. For example, the denoising ML model 210 of the present disclosure may be trained directly on a series of increasingly noisy point clouds 204 of a tooth restoration design (e.g., a pre-restoration or post-restoration tooth design). The fully trained denoising ML model 210 can then be used to directly denoise the noisy (or randomly initialized) 3D point cloud (or other 3D representation) to generate a tooth restoration design, without first converting the 3D tooth data into a 2D depth map or any other 2D representation and then back to a 3D representation. Not only do the techniques of the present disclosure generate more accurate tooth shapes and / or structures (e.g., less noisy tooth surfaces), but the techniques of the present disclosure also enable reduced resource usage. The techniques of the present disclosure avoid the computational costs (e.g., compute cycles) and storage (e.g., computer memory) requirements of converting between representations of different dimensions (e.g., as associated with depth map approaches). Stated another way, the techniques of the present disclosure have a lower resource footprint than techniques that convert between representations of different dimensions as a preprocessing step.
[0028] The disclosed system can automate operations in digital orthodontics (e.g., setup prediction, hardware placement, setup comparison), digital dentistry (e.g., restoration design generation), or a combination thereof. Some techniques can be applied to either or both digital orthodontics and digital dentistry. A non-limiting list of examples follows: segmentation, mesh cleanup, coordinate system prediction, oral care mesh verification, oral care parameter imputation, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalized flow, or denoising diffusion models), metric visualization, appliance component placement, or appliance component generation. In some examples, the disclosed system may enable a clinician or technician to process oral care data (such as a scanned dental arch). In addition to segmentation, mesh cleanup, coordinate system prediction, or verification operations, the disclosed system may enable orthodontic treatment planning, which may include setup prediction as at least one operation. The disclosed system may also enable restoration design generation (e.g., using denoising diffusion models), where designs of one or more restored teeth are generated and processed in the course of making oral care appliances. The systems of the present disclosure may enable either orthodontic or dental treatment planning, or may enable automated steps in the generation of either orthodontic or dental appliances, or both. Some appliances may enable both dental and orthodontic treatment, while other appliances may enable one or the other.
[0029] The present disclosure relates to digital oral care, encompassing the fields of digital dentistry and digital orthodontics. This disclosure generally describes methods for processing three-dimensional (3D) representations of oral care data. Without loss of generality, it should be understood that various types of 3D representations exist. One type of 3D representation is a 3D geometry. The 3D representation may include, be, or be part of one or more of a 3D polygon mesh, a 3D point cloud (e.g., derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels for sparse processing), or a 3D representation described by a mathematical formula. While the term "mesh" is frequently used throughout this disclosure, it should be understood that, in some implementations, this term is interchangeable with other types of 3D representations. The 3D representation may describe elements of an object's 3D geometry and / or 3D structure.
[0030] The dental arches S1, S2, S3, and S4 all contain the exact same tooth meshes, but these tooth meshes are transformed differently according to the following description. The first arch S1 contains a set of tooth meshes positioned (e.g., using a transformation) in the mouth with the teeth in abnormal positions and orientations. The second arch S2 contains the same set of tooth meshes from S1 positioned (e.g., using a transformation) in positions in the mouth with the teeth in their ground truth setup positions and orientations. The third arch S3 contains the same meshes as S1 and S2, positioned (e.g., using a transformation) in the mouth with the teeth in their predicted final setup poses (e.g., as predicted by one or more of the techniques of this disclosure). S4 is the counterpart of S3, with the teeth in a pose corresponding to one of several intermediate stages of orthodontic treatment using clear tray aligners.
[0031] Without loss of generality, it should be understood that the techniques of the present disclosure applied to the final setup are also applicable to intermediate staging in orthodontic treatment, particularly to geometric deep learning (GDL) setups, reinforcement learning (RL) setups, variational autoencoder (VAE) setups, capsule setups, multilayer perceptron (MLP) setups, diffusion setups, pose transfer (PT) setups, similarity setups, force-directed graph (FDG) setups, transformer setups, setup comparisons, or setup classifications. Metric visualization aspects of the present disclosure may be configured to visualize data from both the final setup and intermediate stages. The MLP setup, VAE setup, and capsule setup each fall within the scope of the autoencoder setup. Some implementations of the MLP setup may fall within the scope of the transformer setup. Representation setup refers to any of the MLP setup, VAE setup, capsule setup, and any other setup predictive machine learning model that uses an autoencoder to create a representation of at least one tooth.
[0032] Each of the disclosed setup prediction techniques can be applied to the production of clear tray aligners and / or indirect bonding trays. The setup prediction techniques may also be applicable to other products that involve the final tooth posture. The posture can include position (or location) and rotation (or orientation).
[0033] A 3D mesh is a data structure that may describe the geometry or shape of an object related to oral care, including, but not limited to, a tooth, a hardware element, or a patient's gum tissue. The 3D mesh may include one or more mesh elements, such as one or more of vertices, edges, faces, and combinations thereof. In some implementations, the mesh elements may include voxels, such as in the context of a sparse mesh processing operation. Various spatial and structural features may be calculated for these mesh elements and provided to the disclosed predictive models, which provide the technical advantage of improved data accuracy in the form of the disclosed models outputting more accurate predictions.
[0034] The patient's dentition may include one or more 3D representations of the patient's teeth (e.g., and / or associated deformities), gums, and / or other oral anatomical structures. Orthodontic metrics (OMs), in some implementations, may quantify the relative position and / or orientation of at least one 3D representation of a tooth relative to at least one other 3D representation of the tooth. Restoration design metrics (RDMs), in some implementations, may quantify at least one aspect of the structure and / or shape of the 3D representation of the tooth. Orthodontic landmarks (OLs), in some implementations, may locate one or more points or other structural regions of interest on the 3D representation of the tooth. The OLs, in some implementations, may be used to generate orthodontic or dental appliances, such as clear tray aligners or dental restorative appliances. In some implementations, mesh elements may include at least one component of the 3D representation of the oral care data. For example, in the case of teeth represented by a 3D mesh, the mesh elements may include at least vertices, edges, faces, and voxels. The mesh element features, in some implementations, may quantify some aspect of the 3D representation in proximity to or in association with one or more mesh elements, as described elsewhere in this disclosure. The orthodontic treatment parameters (OPPs), in some implementations, may specify at least one value defining at least one aspect of the patient's planned orthodontic treatment (e.g., specifying a desired target attribute of the final setup in the final setup prediction). The orthodontic doctor preferences (ODPs), in some implementations, may specify at least one typical value of the OPP, which in some examples may be derived from past cases treated by one or more oral care practitioners. The restoration design parameters (RDPs), in some implementations, may specify at least one value defining at least one aspect of the patient's planned dental restorative treatment (e.g., specifying a desired target attribute of the teeth to be treated with dental restoration appliances). The doctor restoration design preferences (DRDPs), in some implementations, may specify at least one typical value of the RDP, which in some examples may be derived from past cases treated by one or more oral care practitioners. 3D oral care expression,1) a set of mesh element labels that may be applied to 3D mesh elements of a teeth / gum / hardware / appliance mesh (or point cloud) during mesh segmentation or mesh cleanup; 2) a 3D representation of one or more teeth / gum / hardware / appliances whose shape has been modified (e.g., cropped, distorted, or filled) during mesh segmentation or mesh cleanup; 3) one or more coordinate systems (e.g., describing one, two, three, or more coordinate axes) for a single tooth or group of teeth (e.g., full arch, as well as LDE coordinate systems); 4) a 3D representation of one or more teeth whose shape has been modified or that are suitable for use in a dental restoration; 5) a 3D representation of one or more dental restoration components; 6) one or more transformations applied to one or more of the placement of dental restoration library components relative to one or more teeth, teeth to be positioned for an orthodontic setup (either final setup or intermediate stage), hardware elements to be positioned relative to one or more teeth, etc.; 7) the orthodontic setup; 8) hardware elements (facial brackets, lingual brackets, orthodontic brackets) to be positioned relative to one or more teeth, etc. 8) a 3D representation of the bonding pads of the hardware elements (this can be generated for a particular tooth by contouring the tooth perimeter, specifying the thickness to form a shell, and then subtracting the tooth via a Boolean operation); 9) a 3D representation of the clear tray aligner (CTA); 10) the location or shape of the CTA trim lines (e.g., described as either a mesh or a polyline); 11) an arch form describing the contour or layout of the dental arch (e.g., described as a 3D polyline or 1) an arch form (described as a 3D mesh or surface) that may follow the incisal edges of one or more teeth, may follow the facial surfaces of one or more teeth, and in some implementations may correspond to a maloccluded arch, and in other implementations may correspond to a final setup arch (the effect of a malocclusion on the shape of the arch form may be reduced by smoothing or averaging the shape of the arch form), and may be described by one or more control points and / or splines; 12) a 3D representation of a fixture model (e.g., a depiction of teeth and gums for use in thermoforming clear tray aligners;or tooth / gums / hardware representations for use in thermoforming indirect bonding trays); 13) one or more latent space vectors (or latent capsules) generated by the 3D encoder stage of a 3D autoencoder trained on oral care mesh reconstruction (e.g., a variational autoencoder trained on tooth reconstruction); 14) one or more oral care metric values (e.g., orthodontic metrics or restorative design generation metrics, etc.) for one or more teeth; 15) one or more landmarks (e.g., 3D points) describing the shape and / or geometric attributes of one or more teeth, other dentition structures, or hardware structures (e.g., used in orthodontic setup creation or restorative appliance component generation or placement); 16) a 3D representation created by scanning (e.g., optical scan, CT scan, or MRI scan) a 3D printed part corresponding to one or more teeth / gums / hardware / appliances (e.g., scanned appliance models); 17) a 3D printed aligner (optionally, local thickness, geometry of reinforcing ribs, flaps, etc.); positioning, etc.), 18) 3D representations of a patient's dentition captured at the chairside by a clinician or doctor (e.g., in situations where the 3D representation is validated at the chairside, so that errors can be detected and re-scans can be performed if necessary before the patient leaves the clinic), 19) dental restoration tooth designs (e.g., for veneers, crowns, bridges or dental restorations), 20) 3D representations of one or more teeth for use in digital oral care procedures, 21) other 3D printed parts related to oral care procedures or other fields, 22) IPR cutting surfaces, 23) one or more orthodontic setup transformations associated with one or more IPR cutting surfaces; 24) a (digital) pontic design that can fill at least a portion of the space between the teeth to allow room for the erupted teeth to later emerge from the gums in the orthodontic setup; 25) fixture model components (e.g., including, but not limited to, fixture model components such as interdental webbing, blockouts, occlusal locks, occlusal ramps, interdental reinforcements, gingival margins, torque points, power ridges, pontic teeth, or dimples, among others).
[0035] The techniques of the present disclosure may be advantageously combined. For example, a setup comparison tool may be used to compare the output of a GDL setup model with ground truth data, the output of an RL setup model with ground truth data, the output of a VAE setup model with ground truth data, and the output of an MLP setup model with ground truth data. With each of these setup prediction models compared to ground truth data, it may be possible to determine which model provides the best performance for a particular dataset or within a given problem domain. Furthermore, a metric visualization tool may enable a global view of the final setup and intermediate stages generated by one or more of the setup prediction models, advantageously enabling selection of the best setup prediction model. Furthermore, the metric visualization tool enables the calculation of metrics with global scope across a set of intermediate stages. These global metrics, in some implementations, may be consumed as input to a neural network for predicting setups (e.g., GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, among others). Global metrics may also be provided for FDG setups. Local metrics from the present disclosure (i.e., local metrics are metrics that can be calculated for one stage or setup of a treatment, rather than across several stages or setups) can be consumed by the neural networks herein to predict setups, in some implementations, with the advantage of improving prediction results. The metrics described in the present disclosure can be visualized, in some implementations, using a metric visualization tool.
[0036] VAE and MAE models for mesh element labeling and mesh infill can be advantageously combined with a setup prediction neural network for the purpose of cleaning up the mesh before or during the prediction process. In some implementations, the VAE for mesh element labeling can be used to flag mesh elements for further processing, such as metric calculation, removal, or modification. In some examples, such flagged mesh elements can be provided as input to the setup prediction neural network to inform the neural network about important mesh features, attributes, or geometries, with the benefit of improving the performance of the resulting setup prediction model. In some implementations, mesh infill can more closely approximate the tooth geometry, allowing for better functioning of the setup prediction model (i.e., improved accuracy of prediction due to better-defined geometry). In some examples, a neural network for classifying the setup (i.e., a setup classifier) can assist the functioning of the setup prediction neural network because the setup classifier informs the setup prediction neural network when the predicted setup is acceptable for use and can be provided to a method for aligner tray generation. Setup classifiers (e.g., GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, among others) can assist in generating the final setup and in generating intermediate stages. Furthermore, the setup classifier neural network may be combined with a metric visualization tool. In other implementations, the setup classification neural network may be combined with a setup comparison tool (e.g., the setup comparison tool may output an indication of how a setup partially generated by the setup classifier compares to a setup generated by another setup prediction method).In some implementations, the VAE for mesh element labeling may identify one or more mesh elements for use in metric calculations. The resulting metric output may be visualized by a metric visualization tool.
[0037] In some examples, the setup classifier neural network may assist the setup prediction techniques described in U.S. Patent Application No. 20210259808 (incorporated herein by reference in its entirety), or the setup prediction techniques described in PCT Application No. WO 2021245480 (incorporated herein by reference in its entirety) or PCT Application No. PCT / IB2022 / 057373 (incorporated herein by reference in its entirety). The setup classifier helps one or more of these techniques know when the predicted final setup is closest to being accurate. In some examples, the setup classifier neural network may output an indication of how far a given setup is from the final setup (i.e., a progress indicator).
[0038] In some implementations, the latent space embedding vector from the reconstruction VAE can be concatenated as input to the setup prediction neural network described in WO2021245480. The latent space vector can also be incorporated as input to other setup prediction models, including GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup, among others. The advantage is that it imparts reconstruction properties (e.g., latent vector dimensions of the tooth mesh) to the neural network, thus improving the generated setup prediction.
[0039] In some examples, various setup prediction neural networks of the present disclosure may work together to generate the setup required for orthodontic treatment. For example, a GDL setup model may generate a final setup, and an RL setup model may use the final setup as an input to generate a series of intermediate-stage setups. Alternatively, a VAE setup model (or an MLP setup model) may create a final setup that can be used by the RL setup model to generate a series of intermediate-stage setups. In some implementations, a setup prediction may be generated by one setup prediction neural network and then taken as input to another setup prediction neural network for further improvement and tuning. In some implementations, such improvement may be performed in an iterative manner.
[0040] In some implementations, a setup validation model, such as the model disclosed in U.S. Provisional Application No. 63 / 366,495, may be involved in this iterative setup prediction loop. First, a setup may be generated (e.g., using models trained for setup prediction, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, among others), and then the setup undergoes validation. If the setup passes validation, the setup may be output for use. If the setup fails validation, the setup may be sent back to one or more of the setup prediction models for correction, improvement, and / or adjustment. In some examples, the setup validation model may output an indication of what is wrong with the setup, allowing the setup generation model to create an improved version in the next iteration. This process repeats until completion.
[0041] Generally speaking, in some implementations, two or more of the following techniques of the present disclosure may be combined during the course of orthodontics and / or dental treatment: GDL setup, setup classification, reinforcement learning (RL) setup, setup comparison, autoencoder setup (VAE setup or capsule setup), VAE mesh element labeling, masked autoencoder (MAE) mesh infill, multi-layer perceptron (MLP) setup, metric visualization, imputation of missing oral care parameter values, tooth classification using latent vectors, FDG setup, pose movement setup, restoration design metric calculation, neural network techniques for dental restoration and / or orthodontics (e.g., 3D oral care representation generation or modification using a transformer), landmark-based (LB) setup, diffusion setup, tooth movement procedure imputation, capsule autoencoder segmentation, diffusion segmentation, similarity setup, oral care representation validation (e.g., using an autoencoder), coordinate system prediction, restoration design generation, or geometry generation (or modification) using a denoising diffusion model.
[0042] Oral care parameters may include one or more values specifying orthodontic treatment parameters or restoration design parameters (RDPs), as described herein. Oral care parameters may define one or more intended aspects of the 3D oral care representation and may be provided to an ML model to facilitate the ML model generating output that can be used to generate oral care appliances suitable for treating a patient. Other types of values include physician preferences and restoration design preferences, as described herein. Practitioner preferences and restoration design preferences can define a particular clinician's typical treatment choices or practices. Restoration design preferences are subjective to a particular clinician and therefore differ from restoration design parameters. In some implementations, physician preferences or restoration design preferences may be calculated by unsupervised means, such as clustering, that can determine typical values that clinicians use in treating patients. These typical values can be stored in a data store and called upon to be provided to an automated ML model as default values (e.g., default values that can be modified prior to execution of the model).
[0043] For example, one clinician may prefer one value for a restoration design parameter (RDP), while another clinician, when faced with a similar diagnosis or treatment protocol, may prefer a different value for that RDP. One example of such an RDP is a dental restoration style. In some implementations, treatment parameters and / or doctor preferences may be provided to a setup prediction model for orthodontic treatment to improve customization of the resulting orthodontic appliance. The restoration design parameters and doctor's restoration preferences may, in some implementations, be used to design tooth geometries for use in creating a dental restoration to improve customization of the appliance. In addition to oral care parameters, doctor's preferences, and doctor's restoration preferences, some implementations of the disclosed ML predictive models may also take a setup (e.g., tooth arrangement) as input in orthodontic treatment. In some such implementations, the disclosed ML predictive models may take a final setup (i.e., final arrangement of teeth) as input, such as in the case of a predictive model trained to generate intermediate stages. For simplicity, these preferences will be referred to as physician restoration preferences, but this is intended to be used in a non-limiting sense. In particular, it should be understood that these preferences may be designated by any treating professional or other appropriate medical professional and are not intended to be limited to physician preferences per se (i.e., preferences from a physician or equivalent degree holder).
[0044] An oral care professional or clinician, such as a dentist or orthodontist, can specify information about a patient's treatment in the form of a set of patient-specific treatment parameters. In some examples, the oral care professional may specify a set of general preferences (also known as doctor preferences) for use across a wide range of cases to use as default values in a set of treatment parameter specification processes. The oral care parameters, in some implementations, may be incorporated into techniques described in this disclosure, such as one or more of GDL setup, VAE setup, RL setup, setup comparison, setup classification, VAE mesh element labeling, MAE mesh infill, validation using an autoencoder, missing treatment parameter value imputation, metric visualization, or FDG setup. One or more of these models may take one or more treatment parameter vectors K and / or one or more doctor preference vectors L as input. In some implementations, one or more of these models may introduce one or more treatment parameter vectors K and / or one or more doctor preference vectors L into one or more hidden layers of a neural network. In some implementations, one or more of these models may incorporate either or both of K and L into their mathematical calculations, such as force calculations, in order to improve the ultimate customization of the resulting brace for the patient.
[0045] Some implementations of neural networks for predicting setups (such as GDL setups, VAE setups, or RL setups) can incorporate information from oral care professionals (a.k.a., doctors). This information influences the placement of teeth in the final setup, allowing the tooth positions and orientations to match, within tolerance, the specifications set by the doctor. In some implementations of GDL setup models, oral care parameters can be provided directly to the generator network as separate inputs along with the mesh data. In some implementations of GDL setups, oral care parameters can be incorporated into the feature vectors calculated for each mesh element before the mesh element is provided to the generator for processing. Some implementations of setup prediction models (e.g., VAE setups or diffusion setups) can incorporate oral care parameters into the setup prediction. In some implementations, treatment parameters K and / or doctor preference information L may be concatenated with the latent space vector C. Practitioner preferences (e.g., in the context of orthodontics) and / or practitioner restorative preferences may be expressed in terms of treatment form, or they may be based on characteristics in the treatment plan, such as final set-up characteristics (e.g., amount of occlusal or midline correction in the planned final set-up), intermediate staging characteristics (e.g., treatment duration, tooth movement protocol, or overcorrection strategy), or outcome (e.g., number of corrections / refinements).
[0046] Orthodontic treatment parameters may specify one or more of the following (possible values are indicated in {}): Some exemplary non-limiting categorical values of OPPs are described below. In some implementations, real numeric values may be specified for one or more of these OPPs. For example, an overbite OPP may specify the amount of overbite desired at setup (e.g., in millimeters) and may be received as an input to a setup prediction model to provide the setup prediction model information regarding the amount of overbite desired at setup. Some implementations may specify a numeric value for an overjet OPP or other OPP. In some implementations, one or more OPPs may be defined that correspond to one or more orthodontic metrics (OMs). In some examples, numeric values may be specified for such OPPs for the purpose of controlling the output of the setup prediction model. Teeth moving: {anterior only, anterior and premolar, full arch} Tooth movement restrictions: For each tooth, indicate whether the tooth is {not moved, missing, to be extracted, primary / erupted, intact} Overbite: {ShowResultingOverbiteAfterAlignment, MaintainInitialOverbite, CorrectOpenBite, CorrectDeepBite} Overjet: {ShowResultingOverjetAfterAlignment, MaintainInitialOverjet, ImproveResultingOverjet} Anterior / Aft (AP) Relationship Maintain: {Right, Left, Both} Improve canine relationship only: {right, left, both} Improve canine and / or molar relationship by up to 4mm: {right, left, both} Modify to Class I (canines and molars): {right, left, both} Crossbite (if present) Forward: {Don't Fix, Fix, Not Applicable} Backward: {Don't Fix, Fix, Not Applicable} Class I (canine and molar) correction: {right, left, both} Fix with Posterior IPR: {Yes, No} Class II / III Modification Simulation (requires elastic): {Yes, No} Continuous centrifugal movement (elastic recommended): {Yes, No} Include cuts for elastic?: {Yes, No} Preferred cuts for elastics: {UseButtonCutoutsOnMolarsAndHooksOnCanines, UseButtonCutoutsOnly, UseHooksOnly} Start cutting for elastic: [integer] Level Top Anterior: {Laterals0.5mmShorterThanCentral, LevelIncisalEdges, LevelGingivalMargins, AsIndicated} Interval: {CloseAllSpaces, LeaveSpecificSpaces} Preferred midline position: {SetTheUpperMidlineToIdeal, MatchTheUpperAndLowerToEachOther} Resolving upper crowding with dilation: {primarily, if needed, none} Resolving upper crowding by forward tilt: {primarily, if needed, none} Resolve upper crowding with IPR - Anterior: {Mainly, if needed, None} Resolve upper crowding with IPR - Right posterior: {primarily, if needed, none} Resolve upper crowding with IPR - Left posterior: {primarily, if needed, none} Resolving lower crowding with dilation: {primarily, if needed, never} Resolving lower crowding by forward tilt: {primarily, if necessary, none} Resolve lower crowding with IPR - Anterior: {Mainly, if needed, None} Resolve lower crowding with IPR - Right posterior: {primarily, if needed, none} Resolve lower crowding with IPR - Left posterior: {primarily, if needed, none} Arch Form Finish: {Patient's natural, as indicated} [Physicians can specify the arch form – selected from a set of options or custom designed]
[0047] Other orthodontic treatment parameters may be defined, such as those that may be used to position standardized brackets at predetermined occlusal heights on teeth. In some implementations, one or more orthodontic treatment parameters may be defined to specify at least one of secondary and tertiary rotation angles (i.e., angle and torque, respectively) to be applied to teeth, which may enable, for example, a target setup configuration in which crown landmarks are within a common occlusal threshold distance. In some implementations, one or more orthodontic treatment parameters may be defined to specify a location in global coordinates where at least one crown (or root) landmark (e.g., centroid) is located in the tooth setup arrangement. In general, oral care parameters may be defined that correspond to oral care metrics. For example, orthodontic treatment parameters may be defined that correspond to orthodontic metrics (e.g., to specify the amount of a particular metric desired to appear in the predicted setup at the input of a setup prediction model).
[0048] Physician preferences may differ from orthodontic treatment parameters in that physician preferences are related to an oral care provider and may include the mean, mode, median, minimum, or maximum (or some other statistical value) of past settings related to the oral care provider's treatment decisions for past orthodontic cases. Treatment parameters, on the other hand, are related to a particular patient and may describe the treatment needs of a particular patient. Physician preferences may be related to the physician and the physician's past treatment practices, while treatment parameters are related to the treatment of a particular patient. Physician preferences (or "treatment preferences") may specify one or more of the following (some non-limiting possible values are indicated in {}): Other possible values are found elsewhere in this disclosure.
[0049] Physician preferences may specify one or more of the following: Deep bite cases (amount of bite correction) - Final overbite: [actual value in millimeters, e.g., 0.5 mm] Options - Top Forward Barge: {Yes, No} Optional - Include mandibular canines in vertical overcorrection: {Yes, No} Midline correction in planned final setup: {MaintainInitialMidline, ImproveMidlineWithIPR, AsIndicated} Deep bite cases - Velocity inverse curve: {Yes, No} Anterior Open Bite Cases - Final Overbite: [Real value in millimeters, e.g., 2mm] Is arch expansion a priority for your case? {Yes, No} If so, specify the allowable expansion per quadrant in mm. When expanding upper molars, apply buccal root torque: {Yes, No} Is the IPR of the first Tx design acceptable? {Yes, No} Max IPR per contact: Upper front: [specify in mm] Lower front: [specify in mm] Upper and lower front: [specify in mm] Is asymmetric IPR acceptable?: {Yes, No} Final tooth position (overcorrection strategy): {ideal, overcorrection} Root Movement: {MoveRootsAsNeededToAchieveTreatmentGoals, LimitPosteriorRootMovement, LimitAllRootMovement} Final occlusal contacts: {AllContactsBalancedWhenPossible, NoOcclusalContactOnUpperIncisors, FinishWithHeavyPosteriorContacts, Other} Is asymmetric AP shift acceptable for class correction?: {Yes, No, Other} Treatment period: [stage count] Tooth movement protocols: {protocol_A, protocol_B, protocol_C}
[0050] Existing techniques attempt to move teeth toward the archform V after setup predictions have already been rendered by other workflow components, which introduces error into the resulting setup and reduces performance given the purpose of the workflow component that made the setup prediction. The present disclosure provides several improvements to these existing techniques by allowing archform information to be directly introduced into a setup prediction neural network as an input to that neural network, with technical improvements providing setup predictions that more accurately meet the patient's orthodontic treatment needs (thereby improving data accuracy). The archform information V may be provided as input to any of the GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup prediction neural networks. In some implementations, the archform information V may be provided directly to one or more internal neural network layers in one or more of those setup applications.
[0051] Additional treatment parameters may include a textual description of the patient's medical condition and intended treatment. Such textual descriptions may be analyzed via natural language processing operations including tokenization, stop-word removal, stemming, n-gram formation, text data vectorization, bag-of-word analysis, term frequency-inverse document frequency (TF-IDF) analysis, sentiment analysis, naive Bayes classification, and / or logistic regression classification. The output of such analysis techniques may be used as input to one or more of the neural networks of the present disclosure, which advantageously customizes and refines the predicted output (e.g., predicted setup or predicted mesh shape).
[0052] These additional orthodontic parameters and practitioner preferences may also be incorporated into the neural networks of the present disclosure, with benefits related to data accuracy and efficiency improving the customization of those neural networks and enabling them to predict example outputs that better match the treatment needs of individual patients.
[0053] In some implementations, the datasets used to train one or more of the neural network models of the present disclosure may be conditionally filtered on one or more of the orthodontic treatment parameters described in this section. In some examples, patient cases exhibiting outliers for one or more of these treatment parameters may be omitted from the dataset for training (or, alternatively, used to form) one or more of the neural networks of the present disclosure.
[0054] One or more treatment parameters and / or physician preferences may be provided to the neural network during training. In this manner, the neural network may be conditioned to one or more treatment parameters and / or physician preferences. Examples of such neural networks include conditional generative adversarial networks (cGANs) and / or conditional variational autoencoders (cVAEs), any of which may be used in various neural network-based applications of the present disclosure.
[0055] In some examples, tooth shape-based inputs can be provided to the neural network for setup prediction. In other examples, non-shape-based inputs, such as tooth names or designations, can be used because they relate to dental notation. In some implementations, a vector R of flags can be provided to the neural network, where a "1" value indicates that the tooth is present and a "0" value indicates that the tooth is not present in the patient case (although other values are possible). The vector R can include a one-hot vector, with each element in the vector corresponding to a tooth type, name, or designation. Identifying information about the tooth (e.g., tooth name) can be provided to the predictive neural network of the present disclosure, with the advantage of allowing the neural network to be trained to handle different teeth in a tooth-specific manner. For example, a setup prediction model can be trained to perform setup transformation predictions for a specific tooth designation (e.g., upper right central incisor or lower left canine). In the case of a mesh cleanup autoencoder (to label mesh elements or fill in missing mesh data), the autoencoder can thus be trained to provide specialized treatment for a tooth according to its tooth designation. In the case of a setup classification neural network, a list of names of the teeth present in the patient's arch may better enable the neural network to output an accurate determination of the setup classification, as tooth designations are valuable input for training such a neural network. Tooth designations / names may be defined, for example, according to the Universal Numbering System, the Palmer system, or FDI World Dental Federation Notation (ISO 3950).
[0056] In one example, if all but (up to four) wisdom teeth are present in a case, a vector R may be defined as an optional input to the setup predictive neural network of the present disclosure, where there are 0's in the vector elements corresponding to each of the wisdom teeth and 1's in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, UL1, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7
[0057] In some examples, the positions of the tooth tips may be provided to the neural network for setup prediction. In other examples, one or more vectors S of orthodontic metrics described elsewhere in this disclosure may be provided to the neural network for setup prediction. The advantage is an improved ability for the network to become trained to understand improperly occluded setup conditions and therefore be able to predict more accurate final setups or intermediate stages.
[0058] In some implementations, the neural network can take as input one or more indications of interproximal reduction (IPR) U, which can indicate the amount of enamel to be removed from a tooth during a course of orthodontic treatment (either mesial or distal). In some implementations, IPR information (e.g., the amount of IPR performed on one or more teeth, measured in millimeters, or one or more binary flags indicating whether IPR will be performed on each tooth identified by flagging) can be concatenated with the latent vector A generated by the VAE or latent capsule T autoencoder. The vectors and / or capsules resulting from such concatenation can be provided to one or more of the neural networks of the present disclosure, with technical improvements or additional benefits that enable the predictive neural network to take IPR into account. IPR is particularly relevant to setup prediction methods, which can determine tooth position and orientation at the end of treatment or during one or more stages during treatment. It is important to consider the amount of enamel to be removed prior to the predicted tooth movement.
[0059] In some implementations, one or more treatment parameters K and / or a doctor preference vector L may be introduced into the setup prediction model. In some implementations, one or more optional vectors or values of tooth position N (e.g., XYZ coordinates in either local or global tooth coordinates), tooth orientation O (e.g., pose in a transformation matrix or quaternion, Euler angles, or other forms described herein), tooth P dimensions (e.g., length, width, height, circumference, diameter, diagonal dimension, volume—any of these dimensions may be normalized relative to another tooth or teeth), distance between adjacent teeth Q. These “tooth P dimensions” may, in some cases, be used to describe the intended dimensions of the tooth for dental restoration design generation.
[0060] In some implementations, a tooth dimension P (e.g., length, width, height, or circumference) can be measured within a plane, such as a plane intersecting the tooth's centroid or a plane intersecting a center point located midway between the centroid and either the most incisal or most gingival area of the tooth. A tooth's height dimension can be measured as the distance from the gum to the incisal edge. A tooth's width dimension can be measured as the distance from the mesial to the distal area of the tooth. In some implementations, the circularity or roundness of the tooth cross section can be measured and included in the vector P. Circularity or roundness can be defined as the ratio of the radii of the inscribed and circumscribed circles.
[0061] The distance Q between adjacent teeth can be implemented in different ways (and calculated using different distance definitions, such as Euclidean distance or geodesic distance). In some implementations, the distance Q1 may be measured as the average distance between the mesh elements of two adjacent teeth. In some implementations, the distance Q2 may be measured as the distance between the centers or centroids of two adjacent teeth. In some implementations, the distance Q3 may be measured between the closest approaching mesh elements between two adjacent teeth. In some implementations, the distance Q4 may be measured between the cusp tips of two adjacent teeth. In some implementations, the teeth may be considered to be adjacent within an arch. In some implementations, the teeth may be considered to be adjacent between opposing arches. In some implementations, any of Q1, Q2, Q3, and Q4 may be divided by a term to normalize the resulting value of Q. In some implementations, the normalization term may involve one or more of tooth volume, tooth mesh element count, tooth surface area, tooth cross-sectional area (e.g., projected onto the XY plane), or some other term related to tooth size.
[0062] Other information about the patient's dentition or treatment needs (or related parameters) may be concatenated with other input vectors to one or more of the MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and / or any of the neural network models listed elsewhere in this disclosure.
[0063] The vector M may include flags that apply to one or more teeth. In some implementations, M includes at least one flag for each tooth to indicate whether the tooth is pinned. In some implementations, M includes at least one flag for each tooth to indicate whether the tooth is fixed. In some implementations, M includes at least one flag for each tooth to indicate whether the tooth is a pontic. Other additional flags are possible for teeth, such as combinations of fixed, pinned, and pontic flags. A flag set to a value indicating that a tooth should be fixed signals to the network that the tooth should not move over the course of treatment. In some implementations, the neural network loss function may be designed to penalize (and in some specific cases may severely penalize) any movement in the indicated tooth. A flag indicating that a tooth is a pontic informs the network that the tooth's gap should be maintained, but the gap is allowed to move. In some cases, M may include a flag indicating that the tooth is missing. In some implementations, the presence of one or more fixed teeth in an arch can aid in setup prediction, as the one or more fixed teeth can provide an anchor for the posture of the other teeth in the arch (i.e., provide a fixed reference for posture transformation of one or more of the other teeth in the arch). In some implementations, one or more teeth may be intentionally fixed to provide an anchor against which the other teeth can be positioned. In some implementations, a 3D representation (e.g., a mesh) corresponding to the gums may be introduced to provide a reference point against which the teeth can be moved.
[0064] Without loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U, and V described elsewhere in this disclosure may be provided to the input or intermediate layers of one or more of the predictive models of the present disclosure. In particular, these optional vectors may be provided to the MLP setup, the GDL setup, the RL setup, the VAE setup, the capsule setup, and / or the diffusion setup, advantageously allowing each model to generate a setup that better matches the patient's orthodontic treatment needs. In some implementations, such inputs may be provided, for example, by concatenating with one or more latent vectors A that are also provided to one or more of the predictive models of the present disclosure. In some implementations, such inputs may be introduced, for example, by concatenating with one or more latent capsules T that are also provided to one or more of the predictive models of the present disclosure.
[0065] In some implementations, one or more of K, L, M, N, O, P, Q, R, S, U, and V may be introduced directly into a neural network (e.g., an MLP or a transformer) in a hidden layer of the network. In some examples, one or more of K, L, M, N, O, P, Q, R, S, U, and V may be introduced directly into the internal processing of the encoder structure.
[0066] In some implementations, a setup prediction model (such as a GDL setup, a RL setup, a VAE setup, a capsule setup, a MLP setup, a PT setup, a similarity setup, and a diffusion setup) may take as input one or more latent vectors A corresponding to one or more input oral care meshes (e.g., tooth meshes, etc.). In some implementations, a setup prediction model (such as a GDL setup, a RL setup, a VAE setup, a capsule setup, a MLP setup, and a diffusion setup) may take as input one or more latent capsules T corresponding to one or more input oral care meshes (e.g., tooth meshes, etc.). In some implementations, a setup prediction method may take both A and T as inputs.
[0067] Various loss calculation techniques are generally applicable to the techniques of this disclosure (e.g., GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, setup classification, tooth classification, VAE mesh element labeling, MAE mesh infill, and treatment parameter imputation).
[0068] These losses include L1 loss, L2 loss, mean squared error (MSE) loss, and cross-entropy loss, among others. Losses can be computed and used in training neural networks such as multilayer perceptrons (MLPs), U-Net structures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, and transformer structures. Some implementations may use either triplet loss or contrastive loss, for example, in sequence learning.
[0069] The loss may also be used to train the encoder and decoder structures. The KL divergence loss may be used at least in part to train one or more of the neural networks of the present disclosure, such as the generator of a mesh reconstruction autoencoder or a GDL setup, taking advantage of the Gaussian behavior in the optimization space. This Gaussian behavior may enable the reconstruction autoencoder to generate better reconstructions (e.g., when a latent vector representation is modified and the modified latent vector is reconstructed using a decoder, the resulting reconstruction is more likely to be a valid instance of the input representation). There are other techniques for calculating losses that may be described elsewhere in this disclosure. Such losses may be based on quantifying the difference between two or more 3D representations.
[0070] The MSE loss calculation may include calculating the mean squared distance between two sets, vectors, or datasets. The MSE may be minimized generally. The MSE may be applicable to regression problems where predictions generated by neural networks or other machine learning models may be real numbers. In some implementations, the neural network may include one or more linear activation units on the output to generate the MSE predictions. Mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss may also be used in accordance with the techniques of this disclosure.
[0071] Cross-entropy, in some implementations, may be used to quantify the difference between two or more distributions. Cross-entropy loss, in some implementations, may be used to train neural networks of the present disclosure. Cross-entropy loss, in some implementations, may involve comparing predicted probabilities to ground truth probabilities. Other names for cross-entropy loss include "logarithmic loss," "logistic loss," and "logarithmic loss." A smaller cross-entropy loss may indicate a better (e.g., more accurate) model. Cross-entropy loss may be logarithmic. Cross-entropy loss, in some implementations, may be applied to binary classification problems. In some implementations, the neural network may include a sigmoid activation unit at the output to generate probability predictions. In the case of multi-class classification, cross-entropy can also be used. In such cases, a neural network trained to make multi-class predictions may include one or more softmax activation functions at the output (e.g., if there is one output node for the class to be predicted). Other loss calculation techniques that may be applied to training the neural networks of the present disclosure include one or more of Huber loss, Hinge loss, Categorical hinge loss, Cosine similarity, Poisson loss, Logcosh loss, or Mean Squared Logarithmic Error loss (MSLE). Other loss calculation methods are described herein and may be applied to training any of the neural networks described in this disclosure.
[0072] In some implementations, one or more of the neural networks of the present disclosure may be trained at least in part by a loss based on at least one of Point-wise Mesh Euclidean Distance (PMD) and Earth Mover's Distance (EMD). Some implementations may incorporate a Hausdorff distance (HD) calculation into the loss calculation. Calculating the Hausdorff distance between two or more 3D representations (such as 3D meshes) may provide one or more technical improvements in that HD not only considers the distance between the two meshes, but also considers how the meshes are oriented and the relationship between the mesh shapes in those orientations (or positions or poses). The Hausdorff distance may improve the comparison of two or more tooth meshes, such as two or more instances of a tooth mesh in different poses (e.g., a comparison between a predicted setup and a ground truth setup, which may be performed in the course of calculating a loss value for training a setup prediction neural network).
[0073] The reconstruction loss can compare the predicted output to the ground truth (or reference) output. The system of the present disclosure can calculate the reconstruction loss as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_loss=0.5*L1(all_points_target, all_points_predicted)+0.5*MSE(all_points_target, all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the ground truth data (e.g., a ground truth example of a ground truth tooth restoration design or some other type of 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the generated or predicted data (e.g., a generated example of a generated tooth restoration design or some other type of 3D oral care representation). Other implementations of the reconstruction loss may additionally (or alternatively) involve an L2 loss, a mean absolute error (MAE) loss, or a Huber loss term.
[0074] The reconstruction error may compare the reconstructed output data (e.g., generated by the reconstruction autoencoder, such as a tooth design generated for use in generating dental restorations) to the original input data (e.g., data provided to the input of the reconstruction autoencoder, such as the pre-restoration teeth). The system of the present disclosure may calculate the reconstruction error as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_error=0.5*L1(all_points_input, all_points_reconstructed)+0.5*MSE(all_points_input, all_points_reconstructed). In the above example, all_points_input is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the input data (e.g., the pre-restoration tooth design provided to the reconstruction autoencoder or another 3D oral care representation provided to the input of an ML model). In the above example, all_points_reconstructed is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the reconstructed (or generated) data (e.g., a reconstructed tooth restoration design, or another example of a generated 3D oral care representation).
[0075] In other words, reconstruction loss involves calculating the difference between the predicted output and the reference output, while reconstruction error involves calculating the difference between the reconstructed output and the original input from which the reconstructed data is derived.
[0076] The techniques of this disclosure may include operations such as 3D convolution, 3D pooling, 3D deconvolution, and 3D unpooling. 3D convolution can assist the segmentation process, for example, when downsampling a 3D mesh. 3D deconvolution undoes 3D convolution, for example, in U-Net. 3D pooling can assist the segmentation process, for example, in summarized neural network feature maps. 3D unpooling undoes 3D pooling, for example, in U-Net. These operations may be implemented by one or more layers within the predictive or generative neural networks described herein. These operations can be applied directly to mesh elements, such as mesh edges or mesh faces. These operations are invariant to changes in mesh rotation, scale, and translation, providing a technical improvement over other approaches. Generally, these operations depend on edge (or face) connectivity, and therefore remain invariant to mesh changes in 3D space as long as edge (or face) connectivity is preserved. That is, operations can be applied to an oral care mesh and produce the same output regardless of the orientation, position, or scale of the oral care mesh, which can lead to improved data accuracy. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes and can be used for tasks such as 3D shape classification or mesh element labeling (e.g., segmentation or mesh cleanup). MeshCNN performs these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.
[0077] In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on a 2D representation (such as an image). In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on a 3D representation (such as a mesh or point cloud). An intraoral scanner may capture 2D images of a patient's dentition from various perspectives. The intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data describing the patient's dentition. According to various techniques, an autoencoder (or other neural network described herein) may be trained to operate on either or both of the 2D and 3D representations.
[0078] A 2D autoencoder (comprising a 2D encoder and a 2D decoder) can be trained on 2D image data to encode the input 2D image into a latent format (such as a latent vector or latent capsule) using the 2D encoder, and then reconstruct a replica of the input 2D image using the 2D decoder. For handheld mobile applications developed for such analysis (e.g., analysis of dental anatomical structures), the 2D images can be easily captured using one or more on-board cameras. In other examples, the 2D images can be captured using an intraoral scanner configured for such functionality. Among the operations that can be used in implementations of 2D autoencoders (or other 2D neural networks) for 2D image analysis are 2D convolution, 2D pooling, and 2D reconstruction error calculation.
[0079] 2D image convolution may involve "sliding" a kernel across the 2D image, computing element-wise multiplications, and adding those element-wise multiplications to the output pixels. The output pixels resulting from each new position of the kernel are stored in an output 2D feature matrix. In some implementations, neighboring elements (e.g., pixels) may be at well-defined locations (e.g., above, below, left, and right) in a rectilinear grid.
[0080] A 2D pooling layer can be used to downsample a feature map and summarize the presence of some features within that feature map.
[0081] The 2D reconstruction error can be calculated between pixels in the input image and pixels in the reconstructed image. The mapping between pixels can be well understood (e.g., assuming both images have the same dimensions, the top pixel [23, 134] in the input image is directly compared to pixel [23, 134] in the reconstructed image).
[0082] Among the advantages provided by the 2D autoencoder-based techniques of the present disclosure is the ease of capturing 2D image data using a handheld device. In some instances, when an external data source provides data for analysis, there may be instances where only 2D image data is available. When only 2D image data is available, analysis using a 2D autoencoder is warranted.
[0083] Modern mobile devices (such as commercially available smartphones) may also have the capability to generate 3D data (e.g., using multiple cameras and stereo photogrammetry, or a single camera moved around an object to capture multiple images from different views, or both), which in some implementations may be arranged into a 3D representation, such as a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation. Analysis of a 3D representation of an object may, in some instances, provide a technical improvement over a 2D analysis of the same object. For example, a 3D representation may describe the object's geometry and / or structure with less ambiguity than a 2D representation (which may include shadows and other artifacts that complicate the depiction of depth from the object and the object's texture). In some implementations, 3D processing may enable technical improvements due to inverse optics problems that may sometimes affect 2D representations. Inverse optics problems refer to the phenomenon whereby the size of an object, the orientation of the object, and the distance between the object and the imaging device may sometimes be confused in a 2D image of that object. Any given projection of an object on the imaging sensor can be mapped to an infinite number of {size, orientation, distance} pairings. 3D representations may enable technical improvements in that they remove ambiguities introduced by inverse optics problems.
[0084] Devices configured for the dedicated purpose of 3D scanning, such as 3D intraoral scanners (or CT or MRI scanners), can generate 3D representations of objects (e.g., a patient's dentition) with significantly higher fidelity and accuracy than is possible with handheld devices. When such high-fidelity 3D data is available (e.g., in the application of oral care mesh classification or other 3D techniques described herein), the use of a 3D autoencoder provides technical improvements (such as increased data accuracy) for extracting the best possible signals from those 3D data (i.e., for obtaining signals from 3D crown meshes used in tooth classification or setup classification).
[0085] A 3D autoencoder (comprising a 3D encoder and a 3D decoder) can be trained on a 3D data representation, using the 3D encoder to encode the input 3D representation into a latent form (such as a latent vector or latent capsule), and then using the 3D decoder to reconstruct a replica of the input 3D representation. Among the operations that can be used to implement a 3D autoencoder for analysis of a 3D representation (e.g., a 3D mesh or a 3D point cloud) are 3D convolution, 3D pooling, and 3D reconstruction error calculation.
[0086] For each mesh element, a 3D convolution can be performed to aggregate local features from nearby mesh elements. Processing can be performed above and beyond techniques for 2D convolution to account for different counts and locations of neighboring mesh elements (relative to a particular mesh element). A particular 3D mesh element may have a variable number of neighbors, and those neighbors may not be found in expected locations (as opposed to pixels in 2D convolution, which may have a fixed number of neighboring pixels that may be found in known or expected locations). In some examples, the order of neighboring mesh elements may be relevant to the 3D convolution.
[0087] A 3D pooling operation may enable the combination of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling may iteratively reduce a 3D mesh to the mesh elements that are most highly relevant to a given application (e.g., a neural network trained on). Similar to 3D convolution, 3D pooling may benefit from special processing beyond that involved in 2D convolution to account for different counts and locations of neighboring mesh elements (relative to a particular mesh element). In some instances, the order of neighboring mesh elements may be less relevant for 3D pooling than for 3D convolution.
[0088] The 3D reconstruction error may be calculated using one or more of the techniques described herein, such as calculating the Euclidean distance between corresponding mesh elements or between two meshes. Other techniques are possible according to aspects of the present disclosure. The 3D reconstruction error may generally be calculated on 3D mesh elements rather than on 2D pixels as in the 2D reconstruction error. Because the 3D representation may, in some instances, have less ambiguity than the 2D representation (i.e., less ambiguity in form, shape, and / or structure), the 3D reconstruction error may allow for technical improvements over the 2D reconstruction error. In some implementations, additional processing may be required for 3D reconstruction over that of the 2D reconstruction due to the complexity of the mapping between the input mesh elements and the reconstructed mesh elements (i.e., the input mesh and the reconstructed mesh may have different mesh element counts, and the mapping between mesh elements may be less clear than the mapping between pixels in the 2D reconstruction). Technical improvements in 3D reconstruction error calculation include improved data precision.
[0089] The 3D representation may be generated using a 3D scanner, such as an intraoral scanner, a computed tomography (CT) scanner, an ultrasound scanner, a magnetic resonance imaging (MRI) machine, or a mobile device capable of performing stereophotogrammetry. The 3D representation may describe the shape and / or structure of an object. The 3D representation may include one or more of a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation, among others. A 3D mesh includes edges, vertices, or faces. Although interrelated in some instances, these three types of data are distinct. Vertices are points in 3D space that define the boundary of the mesh. These points are alternatively described as a point cloud, apart from further information about how the points are connected to each other, as described by edges. An edge is described by two points and may also be called a line segment. A face is described by several edges and vertices. For example, in the case of a triangular mesh, a face includes three vertices, which are interconnected to form three adjacent edges. Some meshes may include degenerate elements, such as non-manifold mesh elements, that can be removed for the benefit of later processing. Other mesh preprocessing operations are possible according to aspects of the present disclosure. 3D meshes are generally formed using triangles, but in other implementations, they may be formed using quadrilaterals, pentagons, or some other n-gons. In some implementations, the 3D mesh may be converted into one or more voxelized geometries (i.e., including voxels), such as when sparse processing is performed. Techniques of the present disclosure that operate on 3D meshes may receive as input meshes of one or more teeth (e.g., arranged in one or more dental arches). Each of these meshes may undergo preprocessing before being input to a predictive architecture (e.g., including at least one of an encoder, decoder, pyramidal encoder-decoder, and U-Net). This preprocessing may include converting the mesh into a list of mesh elements, such as vertices, edges, faces, or, in the case of sparse processing, voxels. A feature vector may be generated for a selected mesh element type or types (e.g., vertices). In some examples, one feature vector is generated for each vertex of the mesh.Each feature vector may include a combination of spatial and / or structural features as specified in the table below:
[0090] [Table 1]
[0091] Table 1 discloses non-limiting examples of mesh element features. In some implementations, color (or other visual cues / identifiers) may be considered a mesh element feature in addition to the spatial or structural mesh element features described in Table 1. As used herein (e.g., in Table 1), a point differs from a vertex, which is part of a 3D point cloud, but a vertex is part of a 3D mesh and may have an incident face or edge. A dihedral angle (which may be expressed in either radians or degrees) may be calculated as the angle (e.g., a signed angle) between two connected faces (e.g., two faces connected along an edge). The sign of the dihedral angle can reveal information about the convexity or concavity of the mesh surface. For example, a positively signed angle may indicate a convex surface in some implementations. Additionally, a negatively signed angle may indicate a concave surface in some implementations. To calculate the principal curvatures of a mesh vertex, first, directional curvatures may be calculated for each neighboring vertex around the vertex. These directional curvatures may be sorted into circular order (e.g., 0, 49, 127, 210, 305 degrees) in proximity to the vertex normal vector and may comprise a subsampled version of the full curvature tensor. Circular order means sorted by angle around an axis. The sorted directional curvatures can contribute to a system of linear equations amenable to a closed-form solution from which the two principal curvatures and directions that can characterize the full curvature tensor can be estimated. Consistent with Table 1, voxels may also have characteristics that are calculated as a collection of other mesh elements (e.g., vertices, edges, and faces) that either intersect with the voxel or, in some implementations, are primarily or completely contained within the voxel. Rotating a mesh may not change structural characteristics, but may change spatial characteristics. Also, as discussed elsewhere in this disclosure, the term "mesh" should be considered to include, in a non-limiting sense, 3D meshes, 3D point clouds, and 3D voxelized representations. In some implementations, apart from mesh element features, there are alternative ways to describe the geometry of a mesh, such as 3D keypoints and 3D descriptors.Examples of such 3D keypoints and 3D descriptors can be found in TONIONI A et al., "Learning to detect good 3D keypoints," Int J Comput Vis. 2018 Vol. 126, pp. 1-20. In some implementations, the 3D keypoints and 3D descriptors may describe extrema (either minima or maxima) of the surface of the 3D representation. In some implementations, one or more mesh element features may be computed, at least in part, via deep feature synthesis (DFS), for example, as described in J.M. Canter and K. Veeramachaneni, "Deep feature synthesis: Towards automating data science endeavors," 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp. 1-10, doi:10.1109 / DSAA.2015.7344858.
[0092] Mesh element feature vectors may, in some implementations, be calculated for a 3D representation provided to latent coding module 214 or 218 (e.g., a representation generation neural network). For example, when a 3D mesh of teeth is provided to latent coding module 214 or 218, mesh element feature vectors may be calculated for one or more of the mesh elements of the 3D mesh to improve the accuracy of the resulting latent representation (e.g., the latent representation generated by latent coding module 214 or 218).
[0093] Representation-generating neural networks based on autoencoders, U-Nets, transformers, 3D SWIN transformers, other types of encoder-decoder structures, convolutional and / or pooling layers, or other models can benefit from the use of mesh element features. Mesh element features can convey aspects of the surface shape and / or structure of the 3D representation to the neural network model of the present disclosure. Each mesh element feature describes distinct information about the 3D representation that may not be redundantly present in other input data provided to the neural network. For example, vertex curvature can quantify concave or convex aspects of the surface of the 3D representation that would otherwise not be understood by the network. In other words, mesh element features can provide a processed version of the structure and / or shape of the 3D representation, data that would not otherwise be available to the neural network. This processed information is often more accessible or amenable to encoding by the neural network. A system embodying the techniques disclosed herein has been utilized to perform several experiments on 3D representations of teeth. For example, mesh element features have been provided to a representation-generating neural network based on a U-Net model and also to a representation-generating model based on a variational autoencoder with continuous normalization flow. Experiments have shown that systems using the full complement of mesh element features (e.g., "XYZ" coordinate tuples, "normal vectors," "vertex curvatures," point turns, and normal turns) are at least 3% more accurate than systems that do not. Point turns describe the "XYZ" coordinate tuples with a local coordinate system (e.g., the centroid of each tooth). Normal turns describe the "normal vectors" with a local coordinate system (e.g., the centroid of each tooth). Furthermore, training converges more quickly when the full complement of mesh element features is used. In other words, machine learning models trained using the full complement of mesh element features tended to be more accurate quickly (earlier) than systems that were not trained. For an existing system observed to have a 91% historical accuracy, a 3% improvement in accuracy reduces the actual error rate by more than 30%.
[0094] Predictive models that can operate on feature vectors of the aforementioned features include, but are not limited to, GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, other denoising diffusion models, PT setup, similarity setup, tooth classification, setup classification, setup comparison, VAE mesh element labeling, MAE mesh infill, mesh reconstruction autoencoder, validation using autoencoder, mesh segmentation, coordinate system prediction, mesh cleanup, restoration design generation, appliance component generation and / or placement, or arch form prediction. Such feature vectors can be presented to the input of a predictive model. In some implementations, such feature vectors can be presented to one or more inner layers of a neural network that is part of one or more of the predictive models.
[0095] As described herein, tooth movement can be encoded in various ways to specify the position and orientation of the teeth within a setup, and can specify one or more tooth transformations to be applied to the 3D representation of the teeth. For example, according to certain implementations, tooth positions can be Cartesian coordinates of a tooth's reference origin position defined in some semantic context. Tooth orientations can be expressed as rotation matrices, unit quaternions, or other 3D rotation representations, such as Euler angles relative to a reference system (either global or local). Dimensions are real-valued 3D spatial extents, and gaps can be binary presence indicators or real-valued gap sizes between teeth, particularly if a particular tooth is missing. In some implementations, tooth rotations can be described by a 3x3 matrix (or by a matrix of other dimensions). In some implementations, tooth position and rotation information can be combined into the same transformation matrix, for example, as a 4x4 matrix, which can reflect homogeneous coordinates. In some examples, affine spatial transformation matrices can be used to describe tooth transformations, such as those describing the malocclusion posture of the teeth, the intermediate posture of the teeth, and / or the final setup posture of the teeth. Some implementations can use relative coordinates, where the setup transformation is predicted relative to the malocclusion coordinate system (e.g., the malocclusion-to-setup transformation is predicted directly instead of the setup coordinate system). Other implementations can use absolute coordinates, where the setup coordinate system is predicted directly for each tooth. In relative mode, the transformation can be calculated relative to the centroid of each tooth's mesh (relative to the global origin), which is referred to as "relative local." Some advantages of using relative local coordinates include eliminating the need for a malocclusion coordinate system (landmarking data), which may not be available for all patient case datasets. Some advantages of using absolute coordinates include simplifying data preprocessing, since the mesh data is originally represented relative to the global origin.These details regarding tooth position encoding and tooth orientation encoding may, in some implementations, also be applied to one or more of the neural network models of the present disclosure, including, but not limited to, GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, other denoising diffusion models, PT setup, similarity setup, FDG setup, setup classification, setup comparison, VAE mesh element labeling, MAE mesh infill, mesh reconstruction VAE, and validation using an autoencoder.
[0096] According to a particular implementation, convolutional layers in the various 3D neural networks described herein may use edge data to perform mesh convolution. The use of edge information ensures that the model is insensitive to different input orders of 3D elements. In addition to, or alternatively to, the use of edge data, the convolutional layers may use vertex data to perform mesh convolution. The use of vertex information is advantageous in that there are typically fewer vertices than edges or faces, and therefore, vertex-oriented processing may lead to lower processing overhead and lower computational costs. In addition to, or alternatively to, the use of edge or vertex data, the convolutional layers may use face data to perform mesh convolution. Furthermore, in addition to, or alternatively to, the use of edge, vertex, or face data, the convolutional layers may use voxel data to perform mesh convolution. The use of voxel information is advantageous in that, depending on the selected granularity, there may be significantly fewer voxels to process compared to the vertices, edges, or faces in the mesh. Low density processing (using voxels) can lead to lower processing overhead and lower computational costs (especially in terms of computer memory or RAM usage).
[0097] Representation-generating neural networks based on autoencoders, U-Nets, transformers, other types of encoder-decoder structures, convolutional and / or pooling layers, or other models can benefit from the use of oral care arguments (e.g., oral care metrics or oral care parameters). For example, oral care metrics (e.g., orthodontic metrics or restorative design metrics) may convey aspects of the shape and / or structure of a patient's dentition (e.g., the shape and / or structure of individual teeth, or the particular relationship between two or more teeth) to a neural network model of the present disclosure. Each oral care metric describes distinct information about the patient's dentition that may not be redundantly present in other input data provided to the neural network. For example, an "overbite" metric may quantify the overlap between the upper and lower central incisors along the vertical Z-axis, information that may not otherwise be readily ascertainable by conventional neural networks in some implementations. In other words, oral care metrics provide refined information about a patient's dentition, for which conventional neural networks (e.g., representation-generating neural networks) may not be adequately trained or configured to extract the oral care metrics described herein. However, neural networks specifically trained to generate oral care metrics can overcome such shortcomings, for example, by calculating losses in a manner that facilitates accurate oral care metric prediction. Mesh oral care metrics may provide a processed version of the structure and / or shape of a patient's dentition, data that may not be available to neural networks. This processed information is often more accessible or amenable to encoding by neural networks. Systems embodying the techniques disclosed herein have been utilized to perform several experiments on 3D representations of teeth. For example, oral care metrics have been provided to a representation-generating neural network based on a U-Net model.Based on experiments, systems that use oral care metrics (e.g., "overbite," "overjet," and "canine class relationship" metrics) were found to be at least 2.5% more accurate than systems that did not. Furthermore, when oral care metrics are used, training converges more quickly. In other words, machine learning models trained using oral care metrics tended to be more accurate quickly (earlier) than systems that were not trained. For an existing system observed to have a historical accuracy of 91%, a 2.5% improvement in accuracy reduces the actual error rate by nearly 30%.
[0098] Oral care arguments may include oral care parameters, or oral care metrics. Examples of oral care metrics include orthodontic metrics (OMs) and restorative design metrics (RDMs). RDMs may describe the shape and / or form of one or more 3D representations of teeth for use in dental restorations. One use case is the creation of one or more dental restorations. Another example use case is the creation of one or more veneers (such as zirconia veneers). Some RDMs may quantify the shape and / or other characteristics of teeth. Other RDMs may quantify the relationship (e.g., spatial relationship) between two or more teeth. RDMs differ from restorative design parameters (RDPs) in that restorative design metrics define the current state of a patient's dentition, while restorative design parameters serve as specifications for machine learning or other optimization models to generate the desired tooth shape and / or form. RDMs describe the current (e.g., starting or failed) shape of the teeth. Restoration design parameters specify how an oral care provider (such as a dentist or dental technician) intends the teeth to appear after completion of restorative treatment. Either or both of the RDM and RDP may be provided with a neural network or other machine learning or optimization algorithm for dental restoration purposes. In some implementations, the RDM may be calculated for the patient's pre-restoration dentition (i.e., primary implementation). In other implementations, the RDM may be calculated for the patient's post-restoration dentition. A restoration design may include one or more teeth and may be referred to as a restoration arch. Restoration design generation may include generating an improved geometry and / or structure for one or more teeth in the restoration arch.
[0099] Aspects of RDM calculation are described below. In some implementations, the RDM may be measured, for example, through locating landmarks within the teeth (or gums, hardware, and / or other elements of the patient's dentition) and measuring distances between those landmarks, or may be otherwise made relative to those landmarks. In some implementations, one or more neural networks or other machine learning models may be trained to identify or extract one or more RDMs from one or more 3D representations of the teeth (or gums, hardware, and / or other elements of the patient's dentition). Techniques of the present disclosure may use the RDMs in various ways. For example, in some implementations, one or more neural networks or other machine learning models may be trained to classify or label one or more setups, arches, dentitions, or other sets of teeth based at least in part on the RDMs. Thus, in these examples, the RDMs form part of the training data used to train these models.
[0100] Aspects of a tooth mesh reconstruction autoencoder that can be used in accordance with the techniques of the present disclosure are described below. An autoencoder for restoration design generation is disclosed in U.S. Provisional Application No. 63 / 366,514. This autoencoder (e.g., a variational autoencoder or VAE) takes as input a tooth mesh (or other 3D representation) that reflects the abnormal state (i.e., the pre-restoration tooth shape). The encoder component of the autoencoder encodes the tooth mesh into a latent form (e.g., a latent vector). Modifications can be applied to this latent vector (e.g., based on mapping the latent space through previous experiments) to alter the geometry and / or structure of the final reconstructed mesh. In some implementations, additional vectors can be included with the latent vector (e.g., through concatenation), and the resulting concatenation of vectors can be reconstructed via the decoder component of the autoencoder into a reconstructed tooth mesh that is a replica of the input tooth mesh.
[0101] The RDMs and RDPs may also be used as neural network inputs in the execution phase, according to aspects of the present disclosure. In some implementations, one or more RDMs may be concatenated with inputs to the encoder to convey specific information about the input 3D tooth representation to the encoder. In some implementations, one or more RDMs may be concatenated with latent vectors before reconstruction to provide specific information about the input 3D tooth representation to the decoder component. Furthermore, in some implementations, one or more restoration design parameters (RDPs) may be concatenated with inputs to the encoder component to provide encoder-specific information about the input 3D tooth representation. Similarly, in some implementations, one or more restoration design parameters (RDPs) may be concatenated with latent vectors before reconstruction to provide decoder-specific information about the input 3D tooth representation.
[0102] In this way, either or both of the RDM and RDP may be introduced into the functionality of an autoencoder (e.g., a tooth reconstruction autoencoder) and play a role in influencing the geometry and / or structure of the reconstructed restoration design (i.e., influencing the shape of the teeth on the output of the autoencoder). In some implementations, the variational autoencoder of U.S. Provisional Application No. 63 / 366,514 may be replaced by a capsule autoencoder (e.g., instead of encoding the tooth mesh into a latent vector, the tooth mesh is encoded into one or more latent capsules).
[0103] In some implementations, clustering or other unsupervised techniques may be performed on the RDM to cluster one or more setups, arches, rows, or other sets of teeth based on tooth restoration characteristics. Such clusters may be useful in treatment planning because they provide insight into patient categories with different treatment needs. This information may be beneficial to clinicians as they learn about possible treatment options. In some examples, best practices (such as default RDP values) may be identified for patient cases that fall into one or another cluster (e.g., as determined by a similarity measure, as in k-NN). After a new case is classified into a particular cluster, information about the associated best practice may be provided to the clinician responsible for processing the case. Such default values may, in some examples, be subject to further adjustment or modification.
[0104] Case Assignment: Such clusters can be used to gain further insight into the types of patient cases present in the dataset. Analysis of such clusters may reveal that patient treatment cases with particular RDM values (or ranges of values) may take less time to treat (or alternatively, more time to treat). Cases that take longer (or are otherwise more difficult) to treat can be assigned to experienced or senior technicians for processing. Cases that take less time to treat can be assigned to newer or less experienced processing techniques. Such assignment can be further aided by finding correlations between the RDM values of particular cases and known processing times associated with these cases.
[0105] The following RDMs can be measured and used in the creation of either or both dental restorations and veneers (veneers are a type of dental restoration) with the goal of creating a natural-looking resultant tooth. Symmetry is a generally preferred facet. There may be variations between patients based on demographic differences. The creation of dental restorations can benefit from some or all of the following RDMs. Shade and translucency may be particularly relevant to the creation of veneers, although some implementations of dental restorations may also take this information into account.
[0106] Examples of interdental RDMs are listed below.
[0107] 1) Bilateral Symmetry and / or Proportions: A measure of symmetry between one or more teeth and one or more other teeth opposite them. For example, for a pair of corresponding teeth, a measure of the width of each tooth. In one example, one tooth is of normal width and the other tooth is too narrow. In another example, both teeth are of normal width. The following is a list of attributes that can be measured for a tooth and compared to corresponding measurements for one or more corresponding teeth: a) Width—distance from mesial to distal; b) Length—distance from gingival to incisal; c) Diagonal—distance across the tooth, e.g., from the mesial gingival corner to the distal incisal corner (this measure is one of many that can be used to quantify tooth shape beyond length and width). The ratio between a and b can be calculated as a / b or b / a. Such ratios can indicate whether spatial symmetry exists (e.g., by measuring the ratio a / b on the left side, measuring the ratio a / b on the right side, and then comparing the left and right ratios). In some implementations, when spatial symmetry is "off," lengths, widths, and / or proportions may not match. Such ratios may be calculated relative to a standard in some implementations. Numerous aesthetic standards are described in dental literature. Examples include the Golden Ratio and Recurring Esthetic Dental Proportion. In some implementations, spatial symmetry may be measured on a pair of teeth, one tooth on the right side of the arch and the other tooth on the left side of the arch.
[0108] 2) Adjacent Tooth Ratio: Measure the width ratio of adjacent teeth as measured as a projection along the arch onto a plane (e.g., a plane located in front of the patient's face). The ideal ratio for use in the final restoration design can be, for example, the so-called golden ratio. The golden ratio relates to adjacent teeth such as central and lateral incisors. This metric relates to measuring these ratios as they exist in a poor dentition before restoration. The ideal golden ratio is 1.6, 1, and 0.6 for the central incisors, lateral incisors, and canines on a particular side (either left or right) of a particular arch (e.g., upper arch). If one or more of these ratio values are off (e.g., in the case of "microtooth"), the patient may desire restorative dental treatment to correct the ratios.
[0109] 3) Arch discrepancy: a measure of any size discrepancy between the upper and lower arches, e.g., with respect to tooth width, for purposes of dental restoration. For example, the techniques of the present disclosure can perform adjacent tooth width ratio measurements on the upper and lower arches. In some implementations, a Bolton analysis measurement can be performed by measuring the upper width, the lower width, and the ratio between those quantities. Arch discrepancy can be described in various implementations in absolute measurements (e.g., in mm or other suitable units) or in terms of ratios or proportions.
[0110] 4) Midline: A measure of the upper incisor midline relative to the lower incisor midline. The techniques of the present disclosure may measure the upper incisor midline relative to the nasal midline (if data regarding the location of the nose is available).
[0111] 5) Proximal Contact: A measure of the size (area, volume, circumference, etc.) of the proximal contact between adjacent teeth. In an ideal situation, the teeth contact along their mesial / distal surfaces, and the gums fill gingivally up to where the teeth contact. If the gum tissue cannot fill the space below the proximal contact, a black triangle may form. In some instances, the size of the proximal contact may become progressively shorter for teeth located further back in the arch. In an ideal scenario, the proximal contact is long enough so that an adequately sized incisor interproximal space exists and the gum tissue fills the area below the contact or gingival.
[0112] 6) Interdental Spaces: In some implementations, the techniques of the present disclosure may measure the size (area, volume, circumference, etc.) of interdental spaces, which are the gaps between teeth at either the gingival or incisal edges. In some implementations, the techniques of the present disclosure may measure the symmetry of interdental spaces on both sides of the arch. Interdental spaces are based at least in part on the length of the contact between the teeth and / or at least in part on the shape of the teeth. In some instances, the size of the interdental spaces may be progressively longer for teeth located toward the back of the arch.
[0113] Examples of endodontic RDMs are listed below, continuing the numbering of the other RDMs listed above.
[0114] 7) Length and / or Width: A measure of tooth length relative to tooth width. This metric may reveal, for example, that a patient has long central incisors. Width and length are defined as follows: a) width—the distance from mesial to distal; b) length—the distance from gingival to incisal edge; c) other dimensions of the tooth body—the portion of the tooth between the gingival area and the incisal edge. In some implementations, either or both length and width may be measured for a tooth and compared to the length and / or width of one or more teeth.
[0115] 8) Tooth Morphology: Measures of the primary anatomical structures of tooth shape, such as line angle, buccal contour, and / or incisor angle and / or interdental space. Frequency and / or dimension may be measured. In some implementations, the observed aspects of the primary tooth shape can be matched to one or more known styles. The techniques of the present disclosure can measure secondary anatomical structures of tooth shape, such as mamelon grooves. For example, frequency and / or dimension may be measured. In some implementations, the observed aspects of the secondary tooth shape can be matched to one or more known styles. In some examples, the techniques of the present disclosure can measure tertiary anatomical structures of tooth shape, such as grooves or striations. For example, frequency and / or dimension may be measured. In some implementations, the observed aspects of the tertiary tooth shape can be matched to one or more known styles.
[0116] 9) Shade and / or Translucency: A measure of tooth shade and / or translucency. Tooth shade is often described by a Vitaclassical or 3D Master Shade Guide. Tooth translucency is described by transmittance or contrast ratio. Tooth shade and translucency can be evaluated (or measured) based on one or more of the following types of tooth data: incisal edge, incisal third, body, and gingival third. The translucency of the enamel layer is generally higher than that of the dentin or cementum layer. Shade and translucency can be measured per voxel (locally) in some implementations. Shade and translucency can be measured per region, such as the incisal region, the body region, etc. The body region can refer to the portion of the tooth between the gingival region and the incisal edge.
[0117] 10) Contour Height: A measure of the tooth's contour. When viewed from a proximal view, every tooth has a particular contour or shape moving from gingival to incisal. This is called the tooth's facial contour. Each tooth has a contour height where its shape is most pronounced. This contour height varies from teeth in the anterior arch to teeth in the posterior arch. In some implementations, this measurement can take the form of fitting against a template of known dimensions and / or known proportions. In some implementations, this measurement may quantify the degree of curvature along the facial tooth surface. In some implementations, the location along the tooth contour where the height of the curvature is most pronounced is measured. This location can be measured as a distance from the gingival margin, or a distance from the incisal edge, or as a percentage along the length of the tooth.
[0118] The PCT application WO2020026117 is incorporated herein by reference in its entirety. WO2020026117 lists several examples of orthodontic metrics (OM). Further examples are disclosed herein. Orthodontic metrics can be used to quantify the physical configuration of an arch for orthodontic treatment purposes (as opposed to restoration design metrics, which are dental-related and describe the shape and / or form of one or more pre-restoration teeth for the purpose of assisting dental restorations). These orthodontic metrics can measure how severely maloccluded the arch is, or conversely, metrics can measure how correctly aligned the teeth are. In some implementations, a GDL setup model (or RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup) may incorporate one or more of these orthodontic metrics or other similar or related orthodontic metrics. In some implementations, such orthodontic metrics may be incorporated into feature vectors of mesh elements, and these per-element feature vectors are provided as inputs to a setup prediction network. In some implementations, such orthodontic metrics may be consumed directly by a generator, MLP, Transformer, or other neural network as direct inputs (as presented in one or more input vectors of real numbers S, as described elsewhere in this disclosure). The use of such orthodontic metrics in training a generator may improve the performance (i.e., accuracy) of the resulting generator, resulting in predicted transformations that position the teeth closer to the correct final setup pose than would otherwise be possible. Such orthodontic metrics may be provided to the encoder structure or by a U-Net structure (in the case of a GDL setup). Such orthodontic metrics may be consumed by an autoencoder, variational autoencoder, masked autoencoder, or regularized autoencoder (in the cases of a VAE setup, VAE mesh element labeling, and MAE mesh infill).Such orthodontic metrics may be consumed by a neural network that generates behavioral predictions as part of a reinforcement learning RL set-up model. Such orthodontic metrics may be consumed by a classifier that applies labels to set-up arches (e.g., labels such as bad, staged, or final set-up). This description is non-limiting, as orthodontic metrics may be incorporated into various techniques of the present disclosure in other ways.
[0119] The various loss calculations of the present disclosure can, in some examples, incorporate one or more orthodontic metrics, advantageously improving the accuracy of the resulting neural network. Orthodontic metrics can be used to directly compare predicted examples with corresponding ground truth examples (as is done with the metrics in the setup comparison discussion). In other examples, one or more orthodontic metrics can be obtained from this section and incorporated into the loss calculation. Such orthodontic metrics can be calculated for the predicted examples, and then orthodontic metrics can also be calculated for the ground truth examples. These two orthodontic metric results are then provided to the loss calculation, advantageously improving the performance of the resulting neural network. In some implementations, one or more orthodontic metrics related to the alignment of two or more adjacent teeth can be calculated and incorporated into the loss function, for example, to at least partially train the setup prediction neural network. In some implementations, such orthodontic metrics can influence the network to align the mesial surface of a tooth with the distal surface of an adjacent tooth. Backpropagation is an exemplary algorithm that can train a neural network using one or more loss values.
[0120] In some implementations, one or more orthodontic metrics can be used to evaluate the neural network's predicted output, such as setup prediction. Such metrics can enable the training algorithm to determine how close the predicted output is to an acceptable output, for example, in a quantified sense. In some implementations, the use of this orthodontic metric can enable the calculation of a loss value that does not entirely rely on comparison to ground truth. In some implementations, such use of an orthodontic metric can allow loss calculation and network training to proceed without requiring comparison to ground truth examples. An advantage of such an approach is that the loss can be calculated based on general principles or specifications for predicted outputs (such as setups) rather than tying the loss calculation to specific ground truth examples (which may be defined by a particular doctor, clinician, or technician, whose treatment philosophy may differ from that of other technicians or doctors). In some implementations, such orthodontic metrics can be defined based on FID (Fréchet Inception Distance) scores.
[0121] Below are descriptions of some of the orthodontic metrics used to quantify the condition of a set of teeth in an arch for orthodontic treatment purposes. These orthodontic metrics indicate the degree of malocclusion the teeth are in at a given stage of clear tray aligner treatment.
[0122] Orthodontic metrics that can be calculated using tensors can be particularly advantageous when training one of the neural networks of the present disclosure, as tensor operations can facilitate efficient computation: the more efficient (and faster) the computation, the faster training can proceed.
[0123] In some examples, error patterns may be identified in one or more predicted outputs of the ML model (e.g., transformation matrices for predicted tooth setups, labeling of mesh elements for mesh cleanup, adding mesh elements to a mesh for mesh infill purposes, classification labels for setups, classification labels for tooth meshes, etc.) One or more orthodontic metrics may be selected to be inputs to a next round of ML model training to address any patterns of errors or imperfections that may be identified in the one or more predicted outputs.
[0124] Some OMs may be defined relative to the arch form coordinate frame, the LDE coordinate system. In some implementations, points may be described using the LDE coordinate frame relative to the arch form, where L, D, and E correspond to 1) the length along the curve of the arch form, 2) the distance away from the arch form, and 3) the distance perpendicular to the L and D axes (this is sometimes called eminence), respectively.
[0125] Various OMs and other techniques of the present disclosure can calculate collisions between 3D representations (e.g., of oral care objects, such as teeth). Such collisions may be calculated as at least one of: 1) penetration distance between the 3D tooth representations; 2) number of overlapping mesh elements between the 3D tooth representations; and 3) overlapping volume between the 3D tooth representations. In some implementations, an OM may be defined to quantify collisions of two or more 3D representations of oral care structures, such as teeth. Some optimization algorithms, such as setup prediction techniques, may seek to minimize collisions between oral care structures (such as teeth).
[0126] The inter-arch orthodontic metrics are as follows:
[0127] Six metrics for comparing two or more arches are listed below: Other suitable comparison orthodontic metrics are found elsewhere in this disclosure, such as in the Setup Comparison Techniques section. 1. Rotational geodesic distance (rotation between predicted and ground truth setup examples) 2. Translation distance (gap between predicted example and ground truth setup example) 3. Normalized translation distance 4. 3D alignment error, measuring the distance between the predicted mesh elements and the ground truth mesh elements in mm. 5. Normalized 3D Alignment 6. Percent overlap by volume (% overlap) (or % overlap by mesh element) between the predicted example and the corresponding ground truth example
[0128] The orthodontic metrics within the arch are as follows: Alignment—A 3D tooth orientation vector can be calculated using the mesial-distal axis of the tooth. A 3D vector can also be calculated, which can be a tangent vector to the arch form at the tooth location. The XY components (i.e., can be a 2D vector) can then be used to compare the orientation of the arch form at the tooth location with the tooth orientation in XY space. Cosine similarity can be used to calculate the 2D orientation difference (angle) between the arch form tangent and the mesial-distal axis of the tooth. Arch symmetry—For each bilateral pair of teeth (e.g., lower left lateral incisor and / or lower right lateral incisor), the absolute difference between the X coordinate of each tooth and the X axis of a global coordinate reference frame can be calculated. This delta can indicate the arch asymmetry of a given tooth pair. The result of such a calculation can be the average X axis delta of one or more tooth pairs from the arch. This calculation can be performed relative to the Y axis with the y coordinate (and / or relative to the Z axis with the Z coordinate) in some implementations. Arch Form D Axis Difference - The D dimension difference (i.e., facial-lingual position difference) between two arch states may be calculated for one or more teeth. In some implementations, a dictionary of D direction tooth movements for each tooth may be returned, keyed by tooth UNS number. The LDE coordinate system can be used for the arch form. Arch form (lower) length ratio - The ratio between the current lower arch length and the arch length in the original maloccluded lower arch may be calculated. Arch Form (Upper) Length Ratio - The ratio between the current upper arch length and the arch length in the original maloccluded upper arch may be calculated. Arch form parallelism (full arch) - for at least one local tooth coordinate system origin in the upper arch, one or more nearest origins in the lower arch (e.g., tooth local coordinate system origins). In some implementations, the two nearest origins may be used. The straight-line distance from the upper arch point to a line formed between the origins of two teeth in the opposing (lower) arch may be calculated. The standard deviation of the set of "point-to-line" distances described above may be returned, and this set may consist of the point-to-line distances for each tooth in the arch. Archform parallelism (individual teeth)—This metric may share some calculation elements with the archform_parallelism_global orthodontic metric, except that it may input the average distance from the tooth origin to the line formed by adjacent teeth in opposing arches (e.g., a tooth in the upper arch and a corresponding tooth in the lower arch). The average distance may be calculated for one or more such tooth pairs. In some implementations, it may be calculated for all pairs of teeth. The average distance may then be subtracted from the distance calculated for each tooth pair. This OM can account for deviations of teeth from “typical” tooth parallelism in the arch. Buccolingual inclination—For at least one molar or premolar, find the corresponding tooth on the opposite side of the same arch (i.e., for a tooth on the left side of the arch, find the tooth of the same type on the right side, and vice versa). The OM may calculate an n-element list for each tooth (e.g., n may be equal to 2). This list may include, at a minimum, the tooth IDs of the teeth in each pair (e.g., LeftLowerFirstMolar and RightLowerFirstMolar=[left_tooth_idx_1,right_tooth_idx_2] in the list). Such an n-element vector may be calculated for each molar and premolar on the upper and lower arches. Buccal cusps may be identified on the molars and premolars on the left and right sides of the arch, respectively. A line may be drawn between the buccal cusps of the left tooth and the buccal cusp of the right tooth. A plane is created using this line and the z-axis of the arch. The lingual cusps may be projected onto the plane (i.e., the inclination angle may be determined at this point). By performing additional projections, the approximate vertical distance between the lingual and buccal cusps can be calculated, which can be used as the buccolingual inclination OM. Canine overbite - The upper and lower canines may be identified. The first premolar on a given side of the mouth may be identified. The distances between the upper and lower canines, and between the upper and lower premolars on a given side of the arch may be calculated. The mean (or median, or mode, or some other statistical value) may be calculated for the measured distances. The z-component of this result indicates the degree of overbite. An overbite may be calculated between any tooth on one arch and the corresponding tooth on the other arch. Canine overjet contact - The collision (e.g., collision distance) between pairs of canines on opposing arches can be calculated. Canine Overjet Contact KDE - The orthodontic metric scores for the current patient case may be taken as input, and a previously trained kernel density estimation (KDE) model or distribution may be used to convert the scores into log-likelihoods. This operation can yield information about where this patient case lies in the distribution of "typical" values. Canine Overjet—This OM may share some calculation steps with the Canine Overbite OM. In some implementations, an average distance may be calculated. In some implementations, the distance calculation may calculate the Euclidean distance of the XY components of the teeth in the upper and lower arches to obtain the overjet (i.e., as opposed to calculating the difference in the Z component, as may be done for canine overbite). The overjet may be calculated between any tooth in one arch and the corresponding tooth in the other arch. Canine class relationship (also applies to first, second, and third molars) - This OM, in some implementations, may include two functions (eg, written in Python): get_canine_landmarks(): Gets the landmarks for each tooth that can be used to calculate class relationships, and then in some implementations maps those landmarks onto a global coordinate space so that measurements can be made between teeth. class_relationship_score_by_side(): The average position of at least one landmark for at least one tooth in the lower arch may be calculated, and similarly for the upper arch. A vector may then be calculated from the upper arch landmark position to the lower arch landmark position, and finally this vector may be projected onto the lower arch to provide a quantification (e.g., as a scalar) of the amount of delta in the "arch l axis" position there. This OM may calculate how far forward or backward a tooth is positioned on the l axis relative to one or more target teeth in the opposing arch. Crossbite - The fossa of at least one maxillary molar can be located by finding the midpoint between the distal and mesial marginal ridge saddles of the tooth. The mandibular molar cusp can be located between the marginal ridges of the corresponding maxillary molar. The OM can calculate a vector from the midpoint of the maxillary molar fossa to the mandibular molar cusp. This vector can be projected onto the d-axis of the archform, resulting in a lateral measurement of the distance from the cusp to the fossa. This distance can define the magnitude of the crossbite. Edge Alignment - The OM may identify the leftmost and rightmost edges of a tooth, as well as its neighbors. The OM can then draw a vector from the leftmost edge of the tooth to the leftmost edge of the tooth's neighbor. The OM can then draw a vector from the rightmost edge of the tooth to the rightmost edge of the tooth's neighbor. The OM can then calculate the linear fit error between the two vectors. Such a calculation may involve creating two vectors. Vec_tooth=right_tooths_leftside to left_tooths_leftside Vec_neighbor = right_tooths_rightside to left_tooths_leftside Then, it may involve calculating the dot product of these two vectors and subtracting the result from 1. (i.e., EdgeAlignmentScore=1-abs(dot(Vec_tooth,Vec_neighbor))). A score of 0 may indicate perfect alignment. A score of 1 may mean perpendicular alignment. Incisor Interarch Contact KDE - May identify deviations of IncisorInterarchContact from the mean of a modeled distribution of such statistics across one or more other patient case datasets. Leveling—A measure of leveling between a tooth and its neighbors may be calculated. The OM may calculate the height difference between two or more adjacent teeth. For molars, the OM may use the midpoint between the mesial saddle ridge and the distal saddle ridge as the molar height. For non-molars, the OM may use the crown length from the gum to the apex. In some implementations, the apex may be the origin of the tooth's local coordinate space. Other implementations may place the origin elsewhere. A simple subtraction between the heights of adjacent teeth may result in a leveling delta between the teeth (e.g., by comparing the Z components). Midline - The midline position of the upper and / or lower incisors can be calculated and then the distance between them can be calculated. Posterior Arch Contact KDE - Posterior arch contact score (i.e., impaction depth or other type of impaction) can be calculated and then identified where that score lies within a predefined KDE (distribution) constructed from representative cases. Occlusal contacts - For a particular tooth from an arch, the OM may identify one or more landmarks (e.g., mesial cusp, or central cusp, etc.). A tooth transformation for that tooth is obtained. For each cusp of the current tooth, the cusp may be scored according to how well it contacts the adjacent (corresponding) tooth in the opposite arch. A vector may be found from the cusp of the tooth in question to the perpendicular intersection at the corresponding tooth in the opposing arch. The distance and / or direction (i.e., above or below) to the opposing arch may be calculated. A list containing the resulting signed distances may be returned, one for each cusp of the tooth in question. Overbite - The upper and lower central incisors can be compared along the z-axis. The difference along the z-axis can be used as the overbite score. Overjet - The upper and lower central incisors can be compared along the y-axis. The difference along the y-axis can be used as the overjet score. Molar interarch contacts - Molar interarch contact scores may be calculated or impaction measurements (such as impaction depth) may be used. Root movement d - The transformation of the tooth from the initial state to the next state can be received. The arch form axis at point L along the arch form can be calculated. The OM can return the distance moved along the d axis. This can be achieved by projecting the root pivot point onto the d axis. Root movement l - The transformation of the tooth from the initial state to the next state can be received. The arch form axis at point L along the arch form can be calculated. The OM can return the distance moved along the l axis. This can be achieved by projecting the root pivot point onto the l axis. Spacing - The spacing between each tooth and its neighbor can be calculated. The arch transformation and mesh may be received. The left and right edges of each tooth mesh may be calculated. One or more points of interest may be transformed from local coordinates to the whole arch coordinate frame. Spacing may be calculated in a plane (e.g., the XY plane) between each tooth and its neighbor to the "left". An array of one or more Euclidean distances (e.g., in the XY plane) may be returned that may represent the spacing between each tooth and its left neighbor. Torque - Torque (i.e., rotation around an axis, such as the x-axis) can be calculated. For one or more teeth, one or more rotations can be converted from Euler angles to one or more rotation matrices. A component of the rotation (such as the x-component) can be extracted and converted back to Euler angles. This x-component can be interpreted as the torque for the tooth. A list containing the torque for one or more teeth can be returned and indexed by the UNS number of the tooth. Neural networks of the present disclosure can take advantage of one or more parameter tuning operations, whereby neural network inputs and parameters are optimized to produce more data-accurate results. One parameter that can be adjusted is the neural network learning rate (which can have a value of, for example, 0.1, 0.01, 0.001, etc.). Data augmentation schemes, such as those in which "shivers" are added to the tooth mesh before being input to the neural network, can also be adjusted or optimized (i.e., small random rotations, translations, and / or scalings can be applied to vary the data set and make the neural network robust to changes in the data). A subset of the neural network model parameters available for tuning are: Learning rate (LR) decay rate (e.g., how much the LR decays during a training run) ○ Learning Rate (LR), a floating point value (e.g., 0.001) used by the optimizer.
[0129] ○ LR schedules (e.g., cosine annealing, step, exponential) ○ Voxel size (for sparse mesh processing) Dropout % (e.g., dropout that may occur in a linear encoder) ○ LR decay step size (e.g., decay every 10, 20, or 30 epochs) o Model scaling, which may increase or decrease the number of layers and / or the number of parameters per layer.
[0130] Parameter adjustment may be advantageously applied to training a neural network for predicting final setups or intermediate staging to provide technical improvements in data accuracy. Parameter adjustment may also be advantageously applied to training a neural network for mesh element labeling or mesh infill. In some examples, parameter adjustment may be advantageously applied to training a neural network for tooth reconstruction. With respect to the classifier models of the present disclosure, parameter adjustment may be advantageously applied to a neural network for classification of one or more setups (i.e., classification of one or more arrangements of teeth). An advantage of parameter adjustment is improving the data accuracy of the output of a prediction or classification model. Parameter adjustment may, in some examples, provide the advantage of obtaining the last few remaining percentage points of validation accuracy from a prediction or classification model.
[0131] Various neural network models of the present disclosure can benefit from data augmentation. Examples include models trained on 3D meshes, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, FDG setup, setup classification, setup comparison, VAE mesh element labeling, MAE mesh infill, mesh reconstruction VAE, and validation using an autoencoder. Data augmentation can increase the size of a dental arch training dataset. Data augmentation can provide additional training examples by adding random rotations, translations, and / or rescaling to copies of existing dental arches. In some implementations of the techniques of the present disclosure, data augmentation can be performed by perturbing or jittering the vertices of the mesh in a manner similar to that described in (“Equidistant and Uniform Data Augmentation for 3D Objects,” IEEE Access, Digital Object Identifier 10.1109 / ACCESS.2021.3138162). The positions of the vertices may be perturbed, for example, through the addition of Gaussian noise with 0 mean and 0.1 standard deviation. Other mean and standard deviation values are possible according to the techniques of this disclosure.
[0132] Some techniques of the present disclosure, such as setup comparison techniques and setup prediction techniques (e.g., GDL setup, MLP setup, VAE setup, etc.), can benefit from a processing step that can align (or register) dental arches (e.g., teeth can be represented by 3D point clouds or some other type of 3D representation described herein). Such a processing step can be used, for example, to register a ground truth setup arch from a patient case with a malocclusion arch from that same case, and then these malocclusion and ground truth setup arches are used to train a setup prediction neural network model. Such a step can aid in loss calculation, as the predicted arch (e.g., the arch output by the generator) can be better registered with the ground truth setup arch, a condition that can facilitate calculation of reconstruction loss, representation loss, L1 loss, L2 loss, MSE loss, and / or other types of loss described herein. In some implementations, an iterative nearest neighbor (ICP) technique can be used for such registration. ICP can minimize the squared error between corresponding entities, such as 3D representations. In some implementations, a linear least-squares calculation may be performed. In some implementations, a non-linear least-squares calculation may be performed. Various alignment models can fully or partially incorporate some of the following algorithms: Levenberg-Marquardt ICP, least-squares rigid transformation, robust rigid transformation, random sample consensus (RANSAC) ICP, k-means-based RANSAC ICP, and generalized ICP (GICP). Alignment can, in some instances, help reduce subjectivity and / or randomness that may occur in some instances in reference ground truth setup designs designed by engineers (i.e., two engineers may produce different but valid final setup outputs for the same case) or by other optimization techniques.
[0133] In the experiment, a ground truth (or reference) setup was aligned to the malocclusion (or malocclusion setup) during training of the setup prediction model. The maloccluded teeth were provided to the setup prediction model, which generated a final setup transformation for the maloccluded teeth. The loss between the resulting predicted setup and the pre-aligned ground truth setup was calculated, so that corresponding aspects of the two setups were aligned. This resulted in a more accurate loss calculation. This pre-alignment operation resulted in a 6% improvement in absolute accuracy (e.g., as measured by the ADD10 score), which corresponds to a reduction in error rate of nearly 50% compared to conventional techniques.
[0134] Because the generator network of the present disclosure may be implemented as one or more neural networks, the generator may include an activation function. When executed, the activation function outputs a decision on whether a neuron in the neural network should fire (e.g., send an output to the next layer). Some activation functions may include a binary step function or a linear activation function. Other activation functions impart nonlinear behavior to the network and include a sigmoid / logistic activation function, a Tanh (hyperbolic tangent) function, a rectified linear unit (ReLU), a leaky ReLU function, a parametric ReLU function, an exponential linear unit (ELU), a softmax function, a swish function, a Gaussian error linear unit (GELU), or a scaled exponential linear unit (SELU). Linear activation functions may be well suited to some regression applications (among others) in the output layer. A sigmoid / logistic activation function may be well suited to some binary classification applications (among other applications) at the output layer. A softmax activation function may be well suited to some multi-class classification applications (among other applications) at the output layer. A sigmoid activation function may be well suited to some multi-label classification applications (among other applications) at the output layer. A ReLU activation function may be well suited to some convolutional neural network (CNN) applications (among other applications) at the hidden layer. Tanh and / or sigmoid activation functions may be well suited to some recurrent neural network (RNN) applications, e.g., at the hidden layer (among other applications).There are several optimization algorithms that can be used in training the neural networks of the present disclosure (such as when updating neural network weights), including gradient descent (which uses first-order derivatives to determine training gradients and is commonly used in training neural networks), Newton's method (which may utilize second-order derivatives in loss calculations to find better training directions than gradient descent, but may require calculations involving Hessian matrices), and conjugate gradient methods (which may have faster convergence than gradient descent, but do not require Hessian matrix calculations that may be required by Newton's method). In some implementations, additional methods can be used to update weights in addition to or instead of the techniques described above. These further methods include the Levenberg-Marquardt method and / or simulated annealing. A backpropagation algorithm is used to transfer the results of the loss calculations back to the network so that the network weights can be adjusted and learning can proceed.
[0135] Neural networks contribute to many functions of the applications of the present disclosure, including, but not limited to, GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, tooth classification, setup classification, setup comparison, VAE mesh element labeling, MAE mesh infill, mesh reconstruction autoencoder, validation using autoencoder, oral care parameter imputation, 3D mesh segmentation (3D representation segmentation), coordinate system prediction, mesh cleanup, restoration design generation, appliance component generation and / or placement, or arch form prediction. The neural networks of the present disclosure may embody some or all of a variety of different neural network models.Examples include U-Net architecture, multi-layer perceptron (MLP), transformer, pyramid architecture, recurrent neural network (RNN), autoencoder, variational autoencoder, regularized autoencoder, conditional autoencoder, capsule network, capsule autoencoder, stacked capsule autoencoder, denoising autoencoder, sparse autoencoder, conditional autoencoder, long / short term memory (LSTM), gated recurrent unit (GRU), deep belief network (DBN), deep convolutional network (DCN), deep convolutional inverse graphics network (DCIGN), liquid state machine (LSM), extreme learning machine (ELM), echo state network (ESN), deep residual network (DRN), Kohonen network (KN), and others. network (KN), neural Turing machine (NTM), or generative adversarial network (GAN). In some implementations, an encoder or decoder structure may be used. Each of these models offers one or more of its own particular advantages. For example, certain neural network architectures may be particularly well suited to particular ML techniques. For example, autoencoders are particularly well suited for classifying 3D oral care representations due to their ability to encode the 3D oral care representations into a more easily classifiable form.
[0136] In some implementations, the neural networks of the present disclosure can be adapted to operate on 3D point cloud data (alternatively, on 3D mesh or 3D voxelized representations). Numerous neural network implementations may be applied to processing 3D representations and training predictive and / or generative models for oral care applications, including PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN, and DSG-Net. Oral care applications include, but are not limited to, setup prediction (e.g., using VAE, RL, MLP, GDL, capsule, diffusion, etc. trained for setup prediction), 3D representation segmentation, 3D representation coordinate system prediction, element labeling for 3D representation cleanup (VAE for mesh element labeling), infill of missing elements in 3D representations (MAE for mesh infill), dental restoration design generation, setup classification, appliance component generation and / or placement, arch form prediction, oral care parameter completion, setup validation, or other validation applications, and tooth 3D representation classification.
[0137] Some implementations of the techniques of this disclosure incorporate the use of autoencoders. Autoencoders that can be used in accordance with aspects of this disclosure include, but are not limited to, AtlasNet, FoldingNet, and 3D-PointCapsNet. Some autoencoders can be implemented based on PointNet.
[0138] Representation learning can be applied to the setup prediction techniques of the present disclosure by training a neural network to learn tooth representations and then using another neural network to generate tooth transformations. Some implementations may use a VAE or capsule autoencoder to generate a representation of the reconstruction characteristics of one or more meshes related to the oral care domain (including, in some examples, information about the structure of the tooth mesh). That representation (either latent vectors or latent capsules) can then be used as input to a module that generates one or more transformations of one or more teeth. These transformations, in some implementations, may place the teeth in a final setup pose. In some implementations, these transformations may bring the teeth into an intermediate staging pose. In some implementations, the transformations may be described by a 9×1 transformation vector (e.g., specifying a translation vector and a quaternion). In other implementations, the transformations may be described by a transformation matrix (e.g., a 4×4 affine transformation matrix).
[0139] In some implementations, the system of the present disclosure may implement principal component analysis (PCA) on the oral care mesh and use the resulting principal components as at least part of a representation of the oral care mesh in subsequent machine learning and / or other predictive or generative processes.
[0140] An autoencoder may be trained to generate a latent form of a 3D oral care representation. The autoencoder may include a 3D encoder (which encodes the 3D oral care representation into a latent form) and / or a 3D decoder (which reconstructs the latent form into a replica of the input 3D oral care representation). While this disclosure refers to a 3D encoder and a 3D decoder, the term 3D should be interpreted non-limitingly to encompass multidimensional modes of operation. For example, the system of this disclosure may train a multidimensional encoder and / or a multidimensional decoder.
[0141] The systems of the present disclosure may implement end-to-end training. Some of the end-to-end training-based techniques of the present disclosure may involve two or more neural networks that are trained together (i.e., the weights are updated simultaneously during the processing of each batch of input oral care data). In some implementations, end-to-end training may be applied to setup prediction by simultaneously training a neural network that learns tooth representations along with a neural network that generates tooth transformations.
[0142] According to some transfer learning-based implementations of the present disclosure, a neural network (e.g., U-Net) may be trained for a first task (e.g., coordinate system prediction). The neural network trained on the first task may be used to provide one or more starting neural network weights for training another neural network trained to perform a second task (e.g., setup prediction). The first network learns low-level neural network features of the oral care mesh and may be shown to perform well on the first task. The second network may demonstrate faster training and / or improved performance by using the first network as a starting point in training. Several layers may be trained to encode neural network features for the oral care mesh that were in the training dataset. These layers may then be fixed (or undergo slight changes over the course of training) and combined with other neural network components, such as additional layers, to be trained for one or more oral care tasks (e.g., setup prediction). In this way, a portion of a neural network for one or more of the techniques of this disclosure (e.g., setup prediction) may receive initial training on another task, which may result in significant learning in the trained network layer. This encoded learning may then be built upon with further task-specific training of another network.
[0143] The disclosed system may train ML models using representation learning. Advantages of representation learning include the fact that a generative network (e.g., a neural network predicting a transformation for use in setup prediction) can be configured to receive inputs with a known size and / or standard format, as opposed to receiving inputs with variable size or structure. Representation learning may produce improved performance over other techniques because noise in the input data may be reduced (e.g., because the representation-generative model extracts important aspects of the input representation (e.g., a mesh or point cloud) through a loss calculation or network architecture selected for that purpose). Such loss calculation methods may include the KL divergence loss, the reconstruction loss, or other losses disclosed herein. Representation learning may reduce the size of the dataset required to train the model, as the representation model learns the representation and allows the generative network to focus on learning the generation task. The result may be improved model generalization, as meaningful features of the input data (e.g., local and / or global features) are made available to the generative network. In some examples, transfer learning may first train a representation-generative model. The representation generation model (in whole or in part) can then be used to pre-train a subsequent model, such as a generative model (e.g., to generate transformation predictions). The representation generation model may benefit from employing mesh element features as input to improve its understanding of the structure and / or shape of the input 3D oral care representations in the training dataset.
[0144] According to the present disclosure, transfer learning can be used for setup prediction, as well as for other oral care applications such as mesh classification (e.g., tooth or setup classification), mesh element labeling, mesh element infill, treatment parameter completion, mesh segmentation, coordinate system prediction, restoration design generation, mesh validation (for any of the applications disclosed herein). In some implementations, a neural network trained to output predictions based on an oral care mesh can first be partially trained on one of the following publicly available datasets: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, Thingi10K dataset (particularly related to 3D printed part validation), ABC: Big CAD Model Dataset for Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Parts Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset, before being further trained on oral care data.
[0145] In some implementations, a neural network previously trained on a first dataset (either oral care data or other data) may subsequently receive further training on oral care data and be applied to an oral care application (such as setup prediction). Transfer learning can be used to further train any of the following networks: GCN (Graph Convolutional Network), PointNet, ResNet, or any of the other neural networks from the published literature listed above.
[0146] In some implementations, a first neural network may be trained to predict a tooth coordinate system (such as by using the techniques described in International Publication No. 2022123402 or U.S. Provisional Application No. 63 / 366492). A second neural network may be trained for setup prediction according to any of the setup prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein). Transfer learning may transfer at least some of the knowledge or capabilities of the first neural network to the second neural network. Thus, transfer learning may provide the second neural network with an accelerated training phase to reach convergence. In some implementations, training of the second network, after being augmented with the transferred learning, may then be completed using one or more of the techniques of the present disclosure.
[0147] The systems of the present disclosure may train ML models using representation learning. Advantages of representation learning include that a generative network (e.g., a neural network that predicts transformations for use in setup prediction) may be configured to receive inputs having a known size and / or standard format, as opposed to receiving inputs having variable size or structure. Representation learning may produce improved performance over other techniques because noise in the input data may be reduced (e.g., because the representation generative model extracts hierarchical neural network features and / or reconstruction properties of the input representation (e.g., a mesh or point cloud) through a loss calculation or network architecture selected for that purpose).
[0148] The reconstruction characteristics may include values of the latent representation (e.g., a latent vector) that describe aspects of the shape and / or structure of the 3D representation provided to the representation generation module that generated the latent representation. The weights of the encoder module of the reconstruction autoencoder may, for example, be trained to encode a 3D representation (e.g., a 3D mesh, or others described herein) into a latent vector representation (e.g., a latent vector). Stated differently, the ability to encode a large set of mesh elements (e.g., hundreds, thousands, or millions) into a latent vector (e.g., hundreds or thousands of real values, e.g., 512, 1024, etc.) may be learned by the encoder weights. Each dimension of that latent vector may contain a real number that describes some aspect of the shape and / or structure of the original 3D representation. The weights of the decoder module of the reconstruction autoencoder may be trained to reconstruct the latent vector into a close replica of the original 3D representation. Stated differently, the ability to interpret the dimensions of the latent vector and decode the values within those dimensions may be learned by the decoder. In summary, the encoder and decoder neural network modules are trained to perform a mapping of 3D representations to latent vectors, which can then be mapped back (or otherwise reconstructed) to 3D representations that are substantially similar to the original 3D representations from which the latent vectors were generated.
[0149] Returning to loss calculation, examples of loss calculations may include the KL divergence loss, the reconstruction loss, or other losses disclosed herein. Representation learning may reduce the size of the dataset required to train a model because the representation model learns the representation and allows the generative network to focus on learning the generation task. The result may be improved model generalization because meaningful neural network features (e.g., local and / or global features) of the input data are made available to the generative network. In other words, the first network can learn the representation, and the second network can make the predictive decision. By training two networks to perform their own separate tasks, each of the networks may produce more accurate results for their respective tasks than using a single network trained to both learn the representation and make the decision. In some examples, transfer learning may first train a representation-generative model. That representation-generative model (in whole or in part) can then be used to pre-train a subsequent model, such as a generative model (e.g., generating transformation predictions). The representation generation model may benefit from employing mesh element features as input to improve the ability of the second ML model to encode the structure and / or shape of the input 3D oral care representations in the training dataset.
[0150] One or more of the neural network models of the present disclosure may have an attention gate integrated therein. Attention gate integration provides an enhancement that allows the associated neural network architecture to focus resources on one or more input values. In some implementations, attention gates may be integrated with a U-Net architecture, with the advantage of allowing the U-Net to focus on certain inputs, such as input flags corresponding to teeth intended to be fixed (e.g., prevented from moving) during orthodontic treatment (or requiring other special handling). Attention gates may also be integrated with encoders or autoencoders (such as VAEs or capsule autoencoders) to improve prediction accuracy, according to aspects of the present disclosure. For example, attention gates can be used to configure machine learning models to give higher weight to aspects of the data that are more likely to be associated with correctly generated outputs. Thus, machine learning models configured with these attention gates (or mechanisms) utilize aspects of the data that are more likely to be associated with correctly generated outputs, thereby improving the ultimate prediction accuracy of those machine learning models.
[0151] The quality and composition of a neural network's training dataset can affect the performance of the neural network during its execution phase. Dataset filtering and outlier removal can remove noise from the dataset and can be advantageously applied to training neural networks for various techniques of the present disclosure (e.g., neural networks for final setup or intermediate staging prediction, mesh element labeling or mesh infill, tooth reconstruction, 3D mesh classification, etc.). And while the mechanism for achieving improvement is different from using attention gates, the end result is that this approach allows machine learning models to focus on relevant aspects of the dataset and can result in accuracy improvements similar to those achieved with attention gates.
[0152] For a neural network configured to predict a final setup, a patient case may include at least one of a set of segmented tooth meshes for the patient, a maladaptive transformation for each tooth, and / or a ground truth setup transformation for each tooth. For a neural network configured to predict a set of intermediate stage setups, a patient case may include at least one of a set of segmented tooth meshes for the patient, a maladaptive transformation for each tooth, and / or a set of ground truth intermediate stage transformations for each tooth. In some implementations, the training dataset may exclude patient cases that fall into a passive stage (i.e., a stage in which the teeth in the arch do not move). In some implementations, the dataset may exclude cases in which a passive stage exists at the end of treatment. In some implementations, the dataset may exclude cases in which overcrowding exists at the end of treatment (i.e., an oral care provider such as an orthodontist or dentist selects a final setup in which the tooth meshes overlap to some extent). In some implementations, the dataset may exclude cases with a certain level of difficulty (e.g., easy, medium, and difficult).
[0153] In some implementations, the data set may include cases with zero pinned teeth (or may include cases with at least one pinned tooth). A pinned tooth may be specified by the technician when designing treatment to stop various tools from moving that particular tooth. In some implementations, the data set may exclude cases without any fixed teeth (or conversely, when at least one tooth is fixed). A fixed tooth may be defined as a tooth that does not move during the course of treatment. In some implementations, the data set may exclude cases with no pontic teeth at all (or conversely, when at least one tooth is a pontic). A pontic may be described as a "ghost" tooth that is represented in the digital model of the arch but does not actually exist in the patient's dentition, or it may be a small or partial tooth that may benefit from future work (such as the addition of composite material through a dental restoration). The benefit of including a pontic in a patient case is to leave space in the arch as part of planning for other tooth movement during the course of orthodontic treatment. In some instances, a pontic can save space in a patient's dentition for future dental or orthodontic work, such as the placement of an implant or crown, or the application of a dental restorative device, such as for adding composite material to an existing tooth that is too small or has an undesirable shape.
[0154] In some implementations, the dataset may exclude cases where the patient does not meet an age requirement (e.g., under 12 years old). In some implementations, the dataset may exclude cases with interproximal reduction (IPR) above a certain threshold amount (e.g., greater than 1.0 mm). A dataset for training a neural network to predict clear tray aligner (CTA) setup may exclude patient cases not relevant to CTA treatment. A dataset for training a neural network to predict indirect bonding tray product setup may exclude cases not relevant to indirect bonding tray treatment. In some implementations, the dataset may exclude cases in which only certain teeth are treated. In such implementations, the dataset may include only cases in which at least one of the following is treated: anterior teeth, molars, premolars, molars, incisors, and / or canines.
[0155] The mesh comparison module can compare two or more meshes, for example, to calculate a loss function or a reconstruction error. Some implementations may involve comparing the volume and / or area of the two meshes. Some implementations may involve calculating the minimum distance between corresponding vertices / faces / edges / voxels of the two meshes. For a point in one mesh (e.g., a vertex, a midpoint on an edge, or a triangle center), the minimum distance between that point and a corresponding point in the other mesh is calculated. If the other mesh has a different number of elements or if there is no clear mapping between corresponding points in the two meshes, different approaches may be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools that can function in the mesh comparison module for the present disclosure. In some implementations, the Hausdorff distance can be calculated to quantify the difference in shape between two meshes. The open-source software tool Metro, developed by Visual Computing Lab, can also serve to quantify the difference between two meshes. The following paper, "Metro: Measuring Error on Simplified Surfaces," P. Cignoni, C. Rocchini and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, vol. 17(2), June 1998, pp. 167-174, describes the approach taken by Metro that can be adapted by the neural network applications of the present disclosure for use in mesh comparison and difference quantification.
[0156] Some techniques of the present disclosure may incorporate the operation of projecting a ray normal to the mesh surface for one or more points on a first mesh and calculating the distance before the ray impinges on a second mesh. The length of the resulting line segment may be used to quantify the distance between the meshes. According to some techniques of the present disclosure, the distance may be assigned a color based on the magnitude of the distance, and the color may be applied to the first mesh by visualization.
[0157] A patient's restorative treatment may include specifications for one or more of the following: restoration guidelines, restoration design parameters, and / or restoration rules for modifying one or more aspects of the patient's dentition. When designing a 3D restoration, whether from an aesthetic and / or technical perspective, one or more of many possible factors can be considered. For example, from an aesthetic perspective, the dental and facial midline and angulation may provide overall guidance, as may the amount of teeth seen by others when the lips are at rest and / or when smiling. After these criteria are considered, a set of "golden sections" may also inform the aesthetic design of overall tooth size. Tooth-to-tooth ratios may be configured to reflect these "golden ratios," which are 1.618:1.0:0.618 for central incisors, lateral incisors, and canines, respectively. In some implementations, real values may be specified for one or more of the RDPs and received at the input of a dental restoration design predictive model (e.g., a machine learning model for predicting the final tooth shape upon completion of the restoration design). In some implementations, one or more RDPs may be defined that correspond to one or more repair design metrics (RDMs).
[0158] Constraints based on tooth position (i.e., malocclusion) and orientation (i.e., rotation and tilt) are balanced against achieving appropriate symmetry, tooth proportions, and tooth-to-tooth ratios. Once these parameters are established, various tooth shapes can be utilized to match the overall aesthetics of the patient's face and smile. For example, tooth shapes may be approximately rectangular with squared edges or approximately oval with rounded edges. Furthermore, tooth-to-tooth ratios can be manipulated to elicit different overall aesthetics. 3D dental CAD programs often offer libraries of different tooth "styles" to choose from, providing the designer with the ability to tailor the results to best match the aesthetic and medical requirements of the practitioner and patient. In some instances, symmetry may be observed in that the left side should mirror the right side, thereby allowing symmetry to be measured.
[0159] Tooth length, width, and aesthetic width-to-length relationships may be specified for one or more teeth. In one example, the length of the maxillary central incisors may be set at 11 mm, and the aesthetic width-to-length relationship may be set at either 70% or 80%. In some examples, the lateral incisors may be 1.0 mm to 2.5 mm shorter than the central incisors. The canines may, in some cases, be 0.5 mm to 1.0 mm shorter than the central incisors. Other ratios and dimensions are possible for various teeth.
[0160] From a technical point of view, there are other considerations that may be taken into account. For example, the restoration fabricated from a given material must be thick enough to have the mechanical strength necessary for long-term use. Furthermore, the width and shape of the tooth must be designed to provide adequate contact with the adjacent teeth.
[0161] The example style options listed below are from the Las Vegas Institute (LVI) of Advanced Dental Studies standards. Other style guides are commercially or freely available.
[0162] In these and / or other examples, the neural network engine of the present disclosure may incorporate as input one or more of acceptable "golden ratio" guidelines for tooth sizes, acceptable "ideal" tooth shapes, patient preferences, practitioner preferences, and the like.
[0163] Restoration design parameters (RDPs) may be used to encode aspects of the smile design guidelines described herein, such as parameters related to the intended dimensions of restored teeth. Restoration design parameters are intended as instructions and / or specifications describing the shape and / or form one or more teeth should assume after completion of dental restoration treatment. One or more RDPs may be received by a neural network or other machine learning or optimization algorithm for dental restoration design, advantageously providing guidance to the optimization algorithm. Some neural networks may be trained for dental restoration design generation, such as some examples of GANs or autoencoders. In some examples, the dental restoration design may be used to define a target shape of one or more teeth for the generation of a dental restoration appliance. In some examples, the dental restoration design may be used to define a target tooth shape for the generation of one or more veneers.
[0164] A partial list of tooth dimensions may include length, width, height, circumference, diameter, diagonal dimension, and volume, any of which may be normalized relative to another tooth or teeth. In some implementations, one or more restoration design parameters may be defined that relate to the gap between two or more teeth and the amount of gap, if any, that the patient wants to remain after treatment (e.g., if the patient wants to maintain a small gap between their maxillary central incisors).
[0165] Additional restoration design parameters may include those specified in Table 2. If one of these parameters conflicts with another, the following order may determine priority (i.e., the first parameter in the list below is considered authoritative): If a parameter value is not specified, the parameter may be ignored. In some implementations, default values may be introduced for one or more parameters. Such default values may be determined, for example, through clustering of previous patient cases. Golden proportion guidelines may specify one or more numbers related to the width of adjacent teeth, such as {1.6, 1, 0.6}.
[0166] [Table 2]
[0167] Tooth-to-tooth ratios can be defined similarly between other tooth pairs. Ratios can be created for tooth width, height, diagonals, etc. Angular lines, incisor angles, and buccal contours can describe major aspects of a tooth's macroshape. Mamelon grooves can be vertical macrotextures on the anterior surface of a tooth, sometimes taking on a V-shape. Grooves or striations can be horizontal microtextures on a tooth. Symmetry may generally be desirable. There may be differences between male and female patients.
[0168] Parameters may be defined to encode physician restoration design preferences (DRDPs) as they relate to various use case scenarios. These use case scenarios reflect information about one or more physician treatment preferences and may directly affect one or more tooth characteristics in a dental restoration design or veneer. In addition, a DRDP may describe a value or range of values for the RDP that is preferred or customarily associated with the physician or other treating medical professional. In some examples, such values or ranges of values may be derived from past patient cases treated by that physician or medical professional. In some examples, a DRDP may be defined that is derived from an RDP (e.g., aesthetic width-to-length relationship) or from an RDM.
[0169] A machine learning model, such as one described herein (e.g., a denoising diffusion model), can be trained to generate a design for a crown or root (or both). The dental restoration design can describe the intended tooth shape at the end of the dental restoration. A neural network (such as a generative neural network) can be trained to generate a dental restoration design to be used to manufacture either a veneer (e.g., a zirconia veneer) or a dental restoration appliance. Such a model obtains as input data from a cohort of patient cases, including pre-restoration tooth meshes and corresponding ground truth examples of completed restorations (e.g., tooth meshes having restored shapes and / or structures). Such a model can be trained, at least in part, through the calculation of a loss function that can quantify the difference between the generated crown restoration design and the ground truth crown restoration design. The resulting loss can be used to update the weights of a generative neural network model (e.g., a denoising diffusion model, which can include a U-Net), thereby (at least in part) training the model. The reconstruction loss may be calculated to compare a predicted tooth mesh with a ground truth tooth mesh, or to compare a pre-restoration tooth mesh with a finished restoration design tooth mesh. The reconstruction loss may be calculated as the sum of pairwise distances between corresponding mesh elements and may be calculated to quantify the difference between two crown designs. Other losses disclosed herein may also be included in the training. The generated restoration design may be used to create veneers, which may be 3D printed.
[0170] Machine learning models, such as those described herein (e.g., denoising diffusion models), can be trained to generate components for use in creating dental restorations. Such dental restorations can be used to mold dental composite materials in a patient's mouth, which are cured (e.g., using a curing light) to ultimately produce veneers on one or more of the patient's teeth. 3M FILTEK Matrix is an example of such a product. In some examples, the machine learning model for generating the restoration components can take inputs operable to customize the shape and / or structure of the restoration components, including inputs such as oral care parameters. The one or more oral care parameters can, in some examples, be defined based on oral care metrics. Oral care metrics (e.g., orthodontic metrics or restorative design metrics) may describe the physical and / or spatial relationship between two or more teeth or may describe the physical and / or dimensional characteristics of individual teeth. Oral care parameters may be defined that are intended to provide guidance to the machine learning model regarding generating 3D oral care representations with specific physical characteristics (e.g., related to shape and / or structure). For example, physical properties may be measured using oral care metrics to which oral care parameters correspond, and such oral care parameters may be defined to customize the generation of mold parting surfaces, gingival trim meshes, or other generated appliance components to fit the patient's dental anatomy.
[0171] In some implementations, the 3D representation generation techniques described herein (e.g., denoising diffusion-based techniques) can be trained to generate custom appliance components by determining characteristics of the custom appliance components, such as their size, shape, position, and / or orientation. Examples of custom appliance components include, among others, a mold parting surface, a gingival trim surface, a shell, a facial ribbon, a lingual shelf (also called a "stiffening rib"), a door, a window, an incisal ridge, a case frame sparing, or an interdental matrix wrapping, and a spline. A spline refers to a curve that passes through multiple points or vertices, such as a piecewise polynomial parametric curve. A mold parting surface refers to a 3D mesh that bisects two sides of one or more teeth (e.g., separating the facial side of one or more teeth from the lingual side of one or more teeth). A gingival trim surface refers to a 3D mesh that trims the surrounding shell along the gingival margin. A shell refers to a body of nominal thickness. In some examples, the inner surface of the shell matches the surface of the dental arch, and the outer surface of the shell is a nominal offset of the inner surface. A surface ribbon refers to a reinforcing rib of nominal thickness offset from the shell in a superficial direction. A window refers to an opening that provides access to the tooth surface so that dental composite material can be placed on the tooth. A door refers to a structure that covers the window. An incisal ridge provides reinforcement to the incisal edge of a dental appliance and may be derived from an arch form. A case frame sparing refers to a connecting material that connects components of a dental appliance (e.g., a lingual portion of a dental appliance, a facial portion of a dental appliance, and their subcomponents) to a manufacturing case frame. In this way, the case frame sparing can tie components of a dental appliance to the case frame during manufacturing, protecting various components from damage or loss and / or reducing the risk of mixing up components. These appliance components and others are described in PCT Patent Applications WO2020240351 and WO2021240290, both of which are incorporated herein by reference in their entireties.
[0172] A mold splitting surface refers to a 3D mesh that bisects two sides of one or more teeth (e.g., by separating the facial side of one or more teeth from the lingual side of one or more teeth). A gingival trimming surface refers to a 3D mesh that trims the surrounding shell along the gingival margin. A shell refers to a body of nominal thickness. In some instances, the inner surface of the shell matches the surface of the dental arch, and the outer surface of the shell is a nominal offset of the inner surface.
[0173] A surface ribbon refers to a reinforcing rib of nominal thickness offset from the shell in a superficial direction. A window refers to an opening that provides access to the tooth surface so that dental composite material can be placed on the tooth. A door refers to a structure that covers the window. An incisal ridge provides reinforcement to the incisal edge of a dental restoration and may be derived from an arch form. A case frame sparing refers to a connecting material that connects components of a dental restoration (e.g., a lingual portion of a dental restoration, a facial portion of a dental restoration, and their subcomponents) to a manufacturing case frame. In this way, the case frame sparing can tie components of a dental restoration to the case frame during manufacturing, protecting various components from damage or loss and / or reducing the risk of component mix-up.
[0174] Additional 3D oral care representations that may be generated by a denoising diffusion model (such as a denoising diffusion model trained as described herein) include tooth restoration designs (which may include, for example, interproximal surfaces of teeth, fossae, incisal edges, cusp tips, roots, etc.).
[0175] In some implementations, a denoising diffusion model (DDM) described herein may be trained to perform 3D mesh element labeling (e.g., vertex, edge, face, voxel, or point labeling) in a 3D oral care representation. These labeled mesh elements may be used for mesh cleanup or mesh segmentation. For mesh cleanup, the labeled aspects of the scanned tooth mesh may be used for appliance erasure (removal + replacement) or to modify one or more aspects of the tooth (e.g., by smoothing) to remove aspects of attached hardware (or other aspects of the mesh that may be undesirable for certain treatments and appliance creation, such as foreign material). Mesh element features, such as those described herein, may be calculated for one or more mesh elements in a 3D representation of oral care data (e.g., a 3D representation of a patient's dentition). A vector of such mesh element features may be calculated for each mesh element and then received by a DDM trained to label mesh elements in a 3D oral care representation for the purposes of either mesh segmentation or mesh cleanup. Such mesh element features can provide valuable information about the shape and / or structure of the input mesh to the labeling DDM. For example, training data 200 can include ground truth data, including a 3D representation of the patient's pre-restoration dentition (e.g., from an intraoral scanner, CT scanner, etc.) and ground truth mesh element labels. The ground truth mesh element labels can describe a ground truth (or reference) segmentation of the patient's dentition (e.g., each mesh element occupying the lower left central incisor can have the same label, each mesh element in the gum of the upper arch can have the same label, each mesh element appearing in the upper right canine can have the same label, etc.). Either or both of the patient's dentition and the ground truth mesh element labels can undergo latent encoding (214). Either the original training data 200 or corresponding latent encoding training data can be provided to a Markov chain 204 as part of a forward pass 206.The forward pass may iteratively add noise to the mesh element labels, generating a series of tens, hundreds, or thousands of successively noisier versions of the set of ground truth mesh element labels. In some implementations, the successive addition of noise may include randomizing the values of one or more mesh element labels. In some implementations, the successive addition of noise may include adjusting the labels of neighboring mesh elements of randomly selected mesh elements (e.g., mesh elements located along tooth boundaries). In this manner, boundaries between teeth or between teeth and gums may be adjusted to become increasingly noisy or distorted. The Markov chain may generate a series of increasingly noisy sets of mesh element labels that may be used to at least partially train the denoising ML model 210. In some implementations, a 3D representation of the patient's dwelling, along with the mesh element labels, may also be provided to train the denoising ML model 210. In deployment, a reverse pass 208 may be performed. Oral care arguments 202 may be provided to customize the functionality of the segmentation operation. In deployment, instant patient case data 216 (e.g., a pre-segmented 3D representation of the patient's dentition) may be provided to the denoising ML model 210. In some implementations, the instant patient case data 216 may undergo latent encoding (218). In some implementations, an initial set of mesh element labels corresponding to the 3D representation of the patient's dentition 216 may be generated. The mesh element labels may be initialized randomly or according to a heuristic in some implementations. An example of a heuristic is that a set of random mesh elements may be selected, and neighbors of those mesh elements (e.g., neighbors with a certain distance or number of edges from the selected mesh elements) may be grouped together into a clique (e.g., each mesh element in the clique has the same mesh element label). A reverse pass 208 of the denoising ML model 210 may cause the denoising ML model 210 to iteratively refine the set of mesh element labels for the 3D representation of the patient's dentition 216.After many iterations (e.g., hundreds, thousands, or millions), the iteratively refined mesh element labels may be set to output 220. These generated mesh element labels may be used to segment the patient's retention or to perform mesh cleanup of the patient's dentition. In some examples, if the instant patient case data includes the patient's pre-segmentation dentition, the patient's dentition may be segmented through application of the generated mesh element labels 220. The generated mesh element labels may enable tooth cutting (e.g., generation of a new mesh for each individual tooth) according to the generated mesh element labels 220.
[0176] Some implementations of the DDM-based mesh cleanup techniques described herein can train a DDM to remove (or correct) common triangular mesh defects (e.g., via mesh element labeling), such as: degenerate triangles with zero surface area; redundant triangles that cover the same surface area as another triangle; non-manifold edges with more than two adjacent triangles, also known as "fins"; non-manifold vertices with two or more adjacent sequences of connected triangles (triangle fans); intersecting triangles - two triangles that penetrate each other; and spikes - consisting of multiple triangles, often conical, caused by one or more vertices displaced from the actual surface. sharp features that are missing; creases - sharp features composed of multiple triangles, often Z-shaped with small undercut areas caused by one or more vertices displaced from the actual surface; islands / small components - isolated objects in a scan that should contain only a single object (e.g., smaller objects are typically removed); small holes in the mesh surface, either from the original scan or from a previous deletion due to a defect (e.g., holes can be removed by filling the hold, for example by adding one or more mesh elements); rough boundaries - smooth boundaries are desirable for extending the gingival surface and creating the model base.
[0177] Some implementations of the DDM-based mesh cleanup techniques described herein can train a DDM model to remove (or correct) undesirable mesh aspects and / or domain-specific defects (e.g., via mesh element labeling) under certain circumstances: foreign bodies—portions of an intraoral scan outside the anatomical region of interest, e.g., non-tooth surfaces not within a certain distance from the tooth surface, or scan artifacts that do not represent actual anatomical structures; divots—concave depressions in the surface (e.g., which may be scan artifacts to be fixed or anatomical features generally left intact); undercuts—tooth sides with a lower radius than the crown, which may make physical impressions or aligners difficult to remove or place. Undercuts can be natural features or can result from damage such as abfraction. Abfraction is associated with erosion of teeth near the gum line, causing or exacerbating undercuts. Appliances that the DDM-based models of the present disclosure can process include orthodontic hardware such as attachments, brackets, wires, buttons, tongue bars, and curiel appliances, which may be present in intraoral scans. Digital removal and replacement with a synthetic tooth / gingival surface can be beneficial in some circumstances if performed before the appliance fabrication process proceeds.
[0178] Denoising diffusion may train a neural network (or other machine learning model) to iteratively remove (or refine) noise from a data structure. The techniques of the present disclosure remove noise from 3D point clouds (or other 3D representations), transformation matrices, vectors of labels (e.g., labels that can be applied to mesh elements, used for segmentation or mesh cleanup, used to define object masks), or other types of 3D oral care representations. The techniques may start with randomly generated data structures (e.g., 3D point clouds or transformation matrices), and the techniques may convert those random (or noisy) data structures into artifacts usable in digital oral care. A denoising model may be trained to perform these operations for the type of training data generated using the techniques described herein. This training data may include a series of increasingly noisy examples of the relevant data structure. A denoising neural network (U-Net is one example) may be trained to gradually convert the random (or noisy) data structures in small steps into artifacts usable in digital oral care procedures (e.g., generating oral care appliances).
[0179] Digital oral care involves many different types of 3D representations customized to the patient's anatomy. Diffusion models are particularly applicable to digital oral care due to the customizable nature of the output representations they generate.
[0180] In some implementations, the denoising diffusion model may operate on input data (e.g., a 3D oral care representation) in its original format (e.g., training on and / or generating a 3D representation such as a 3D point cloud). In other implementations, the denoising diffusion model may operate on a latent form of the input data (e.g., stable diffusion, etc.), which may preserve the 3D nature of the input data. The latent form may be information-rich (describes the structure and / or shape of the input data, such as a 3D oral care representation) and / or have a small data footprint that makes it easier for an ML model to be trained. In some examples, an autoencoder may be trained to generate the latent representation. The information-rich nature of the latent representation may be demonstrated when the latent representation is reconstructed into a close replica of the original 3D representation (e.g., as measured by the reconstruction error defined herein). In other words, the original 3D representation may include hundreds or thousands of mesh elements, which in some examples may be encoded by the autoencoder into a vector of hundreds of real values. This vector of hundreds of real values can then, in some instances, be reconstructed into a close replica of the original 3D representation. The accuracy of the reconstruction (e.g., the fidelity with which the reconstructed form matches the original form) can be verified by calculating the reconstruction error.
[0181] An oral care diffusion model (OCDM), in some implementations, may be trained to modify the 3D oral care representations described herein (e.g., appliance components, tooth restoration designs, trim lines, transformations for the 3D oral care representation—orthodontic setups or teeth within appliance components, etc.). In some examples, a template or reference 3D oral care representation may be received at input and then modified by the OCDM. In some examples, an initial tooth representation (e.g., teeth before restoration) may be received at input and then modified by the OCDM to generate teeth ready for restorative treatment (e.g., for fabrication using dental restoration appliances such as crowns, bridges, or FILTEK matrices). An OCDM, in some implementations, may be trained to generate the 3D oral care representations described herein (e.g., appliance components, tooth restoration designs, trim lines, orthodontic setups, etc.). In some implementations, the diffusion model may be trained to generate a 3D oral care representation (e.g., a point cloud or a 3D mesh representing a 3D oral care representation, e.g., a trim line, appliance components, or tooth design for a dental restoration, e.g., veneer generation, dental restoration appliance generation). Such a model may be referred to as a geometric generative diffusion neural network or geometric generative diffusion model (GGDM).
[0182] In some examples, input data for a diffusion model (e.g., a stable diffusion model) may be preprocessed using an encoder (e.g., to generate a reduced-dimensionality representation of the input data structure). Such latent representations may, in some examples, be reconstructed using a decoder. The decoder may reconstruct the latent representation into a replica of the input data (e.g., using a reconstruction autoencoder, etc.). In some examples, the decoder may reconstruct the latent representation into a modified version of the input data (e.g., when one or more aspects of the latent representation are modified). A point-voxel CNN (PVCNN) or an encoder-decoder network (e.g., a variational autoencoder, among others) may be trained to implement the encoder or decoder of the techniques described herein. A 3D oral care representation 102 may be received at the input to the oral care diffusion model (OCDM) and encoded into a latent representation (e.g., a latent vector) by an encoder. In some implementations, such as FIG. 1 , the optional encoder 104 (e.g., as a structure latent encoder that may encode shape and / or structure aspects of the input data) may encode the input 3D oral-care representation 102 into a latent representation 108 of its shape (e.g., referred to as a structure latent representation). In some implementations, such as FIG. 1 , the optional latent mesh element encoder 106 may encode the input 3D oral-care representation 102 into a latent representation 110 of mesh elements (e.g., referred to as a latent mesh element representation). Such mesh elements may include, among others, points in a point cloud, voxels in a sparse representation, or vertices / edges / faces in a 3D mesh. In some implementations, each of these two representations may be generated. In some implementations, such an encoder may incorporate mesh element features at the input to improve the latent representation generated by the encoder. Mesh element feature vectors may be calculated by mesh element feature modules 130 and 128.
[0183] In the forward process, a Markov chain of diffusion steps can be applied to generate a set of training data from such latent representations. The forward process can generate increasingly noisy versions of the input 3D oral care representation 102, which can be used in training one or more denoising diffusion neural network models described herein. This Markov chain of diffusion steps can apply Gaussian noise to the input 3D oral care representation 102 over T steps (e.g., T = 150), which can generate a series of increasingly noisy versions of the input 3D oral care representation 102, which can be used herein to train a denoising diffusion neural network. A diffusion neural network (which can be implemented, for example, as a U-Net, a variational autoencoder, or the like) can be trained to reverse the forward process, starting with the 3D oral care representation at time T and performing a denoising operation to convert the 3D oral care representation into a version corresponding to time T-1. This is referred to as the inverse process. Given a latent representation at time t, the diffusion neural network can iteratively compute a version of the latent representation at time t-1. The diffusion neural network may introduce modifications to the latent representation (e.g., to remove noise or to introduce desired geometric or structural aspects in the output). The diffusion model may consume instructions for such modifications in the form of oral care arguments. The oral care arguments may include one or more natural language text string arguments, one or more categorical arguments, one or more integer arguments, one or more real-valued arguments, one or more image arguments (e.g., reference images describing aspects of the intended results of model execution), one or more 3D representation arguments (e.g., reference 3D point clouds describing aspects outside of the intended model execution), or multiple combinations of the foregoing. Such arguments may first undergo latent encoding (e.g., using the encoder portion of an autoencoder or using a text or numeric encoder). The encoded arguments may be introduced into the diffusion neural network (e.g., by appending the arguments to input data, such as concatenating the arguments with a latent representation of the input data generated using the encoder).For example, if the diffusion neural network is implemented using a U-Net, the coded arguments may be introduced into one or more of the U-Net's successive resolution layers (e.g., a U-Net may have a set of layers corresponding to progressively lower levels of resolution, followed by a set of layers corresponding to progressively higher levels of resolution). The U-Net may use skip connections to share information from input to output at each level of resolution. The U-Net may incorporate attention layers. The coded argument information may be introduced into any or all such layers, with the advantage of guiding the generated output of the diffusion neural network. In some implementations, the denoising diffusion neural network of the present disclosure may include U-Nets, ResNets, transformers, or autoencoders (e.g., variational autoencoders), among others.
[0184] The latent representation of the shape of the 3D point cloud 102 (or other 3D representations described herein) may be referred to as a structure latent representation 108 and may encode aspects of the structure of the input 3D point cloud 102. The oral care arguments 100 (or input parameters) may include one or more attributes that describe the intended output from the trained machine learning model. The structure latent representation 108 may be denoised and / or modified by a structure latent denoising diffusion model 112. The latent representations of the mesh elements (e.g., points, etc.) are referred to as latent mesh element representations 110 (e.g., may encode aspects of the shape of the input 3D point cloud 102). The latent mesh element representations may be denoised and / or modified by a latent mesh element denoising diffusion model 114. The denoised / modified structure latent representation 116 may be provided to a final decoder module 120 (e.g., implemented as a PVCNN). The latent mesh element representation 118 may also be provided to the final decoder module 120. The final decoder module can reconstruct a generated geometry 122 (e.g., or other 3D oral care representation) that is a modified version of the input 3D oral care representation 102 or an entirely new 3D oral care representation having aspects specified by the oral care arguments 100. If the input 3D oral care representation is embodied by a 3D point cloud, the generated 3D oral care representation 122 may also be embodied by a 3D point cloud. Such a point cloud may undergo an optional subsequent processing step 124 to generate a mesh, for example, using Shape As Points (SAP) surface reconstruction (e.g., as shown in FIG. 1). If SAP is applied, a 3D mesh 126 is output. In other examples, the 3D oral care representations of the input 102 and reconstruction 122 may be embodied as voxels or other of the mesh elements described herein.
[0185] Figure 1 illustrates a method using a fully trained GGDM, which has been trained to modify an input 3D oral care representation according to (optional) arguments (e.g., as text instructions or according to oral care parameters described herein).
[0186] In some examples, a diffusion model or text-to-image (TTI) system (e.g., DALL-E 2, MidJourney, or Stable Diffusion) initially trained to generate 2D representations (e.g., 2D images) may be further trained on 3D oral care representations using transfer learning. For example, the denoising ML model 210 may be first trained on 2D data and then further trained (via transfer learning) on 3D data. Such a transfer learning model may, in some implementations, be trained to generate 3D oral care representations. In some examples, the stable diffusion model may operate against the latent representations of the trial 3D oral care representations 216 with the aim of modifying those 3D oral care representations 216.
[0187] In some non-limiting implementations, the denoising diffusion probabilistic model of the present disclosure may operate on a 3D representation of a tooth (e.g., to generate a tooth restoration design). In an example of restoration design generation, a 3D representation of a pre-restoration tooth 102 (e.g., described as a 3D point cloud, a 3D mesh, or a voxelized representation) may be provided as input and processed by the denoising diffusion probabilistic neural network to preserve information about the 3D shape and / or 3D structure of the pre-restoration tooth. In some examples, the pre-restoration tooth may undergo optional mesh element feature vector generation (128) and, optionally, encoding (106) into a latent format that preserves information about the 3D quality of the pre-restoration tooth. The resulting latent mesh element representation 110 may represent the 3D mesh elements (e.g., 3D points, voxels, edges, faces, or vertices) of the pre-restoration tooth in latent format. The latent mesh element representation 110 may be provided to a latent mesh element denoising diffusion model 114, which may generate a modified latent mesh element representation 118. In some implementations, the pre-restoration tooth may undergo optional mesh element vector generation (130) and optionally encoding (104) into a latent format that describes aspects of the 3D structure of the pre-restoration tooth in latent format 108. The latent representation 108 may be provided to a structure latent denoising diffusion model 112, which may generate a modified structure latent representation 116. Either or both of the modified structure latent representation 116 and the modified latent mesh element representation 118 may be provided to a decoder 120, which may reconstruct the representation into a reconstructed 3D oral care representation 122 (e.g., a post-restoration tooth design). If the reconstructed 3D oral care representation includes a point cloud, the point cloud may then undergo surface reconstruction (124) to generate a 3D mesh 126. The oral care arguments 100 can optionally be provided to either or both of the structural latent denoising diffusion model 112 and the latent mesh element denoising diffusion model 114 to customize the output of those models and configure them to generate output suitable for use in generating oral care appliances.Oral care arguments may include oral care parameters or metrics (e.g., restoration design metrics such as "tooth morphology" or "symmetry and / or proportions," among others described herein).
[0188] Such latent formats 110 and 108 can reduce the data size of the pre-restoration teeth and improve data accuracy while still preserving rich information about the pre-restoration tooth shape and / or structure. Neural networks can be more easily trained on input data with a smaller data footprint (e.g., requiring fewer neural network parameters or weights to encode a solution) as long as the input data is information-rich. Reconstruction autoencoders are specifically configured to generate information-rich latent representations, as indicated by the reconstruction autoencoder's ability to reconstruct a close replica of the input data. Ground truth post-restoration 3D representations of the teeth can similarly be processed fully in three dimensions. A geometry generative diffusion model (GGDM) can be trained, at least in part, by calculating a loss that quantifies the difference between a generated post-restoration tooth design (e.g., a design predicted using a denoising diffusion model as described herein) and a ground truth post-restoration tooth (e.g., which may be provided in training data for a given patient case). Such loss calculations can be performed fully in 3D, for example, using reconstruction loss calculations as described herein. The techniques described herein can generate tooth designs for use in designing crowns, veneers, bridges, or other types of restorations, or dental restoration appliances, or fixture models, etc.
[0189] A denoising diffusion model (DDM) (such as that shown in FIG. 2 ) may include one or more Markov chains. A Markov chain may, in some examples, be defined as a sequence of stochastic events in which each time point depends on the previous time point. The transition distribution on the Markov chain forward process may, in some examples, be conditioned by low-level Gaussian noise. FIG. 2 shows a denoising diffusion model trained on structural latent data. The forward pass 206 in FIG. 2 generates training data for the denoising diffusion model (e.g., in some implementations, it may be trained on latent representations of mesh elements). The training data may be used to train a denoising ML model 210. The fully trained denoising ML model 210 shown in FIG. 2 may be used by a GGDM (e.g., when the model is trained on a latent representation of a tooth, such as with restoration design generation), a diffusion setup model (e.g., when the model is trained on a latent representation of a transformation of a tooth or other type of 3D oral care representation), or other types of diffusion models to generate output based on a training dataset of 3D oral care representations.
[0190] The DDM may be trained to approximate a data distribution q(x) over the latent variable Y (e.g., the data distribution may describe a 3D oral care representation). The forward and backward processes of the denoising diffusion model may operate over many time points t, e.g., T==1000. Any of the following three operations may be involved in the denoising diffusion model: a noise scheduler used to implement the forward process (e.g., may introduce progressively more noise of a particular distribution, such as Gaussian noise, at each successive time point—e.g., isotropic Gaussian noise with zero mean and variance in all directions); a neural network (e.g., U-Net, VAE, or others described herein) used to implement the backward process (e.g., may recover, over many time points, the original distribution of the data fed to the start of the forward process); and / or an optional time point encoding method (e.g., using positional embedding). The neural network may be trained, at least in part, using a loss function based on the L1 or L2 norm, or KL divergence (among others described herein). The KL divergence loss may, in some instances, drive the minimization of the difference between the observed distribution and the corresponding generated distribution.
[0191] The forward process may, in some examples, involve further encoding of the received data (e.g., latent data), and the reverse (or inverse) process may, in some examples, involve decoding (or partial decoding) of the data processed by the forward process. In some examples, the reverse process may be performed starting with random noise to generate a sequence of data structures with progressively less noise.
[0192] At the final time point t==T, the distribution of the latent variable Y may, in some examples, be normally distributed after adding noise (e.g., Gaussian noise) using a forward transition module. The forward process may, in some examples, introduce noise into the 3D oral care representation by perturbing one or more mesh elements of the 3D oral care representation, in the example of a 3D representation of oral care data. For example, the positions of points or voxels of a point cloud may be perturbed. The mesh elements may be subjected to perturbations that may change the values of measured mesh element features (e.g., edge lengths, face areas, point or voxel positions, vertex incident edge counts, etc.). The transition module of the forward process, in some implementations, may be written as follows (where M t is q(y t )~=p(y t ) where I is the identity matrix): q(y t |y t-1 )=N(y t ;sqrt(1-M t )y t-1 ,M t I)
[0193] 2 may be implemented using a neural network such as a U-Net in some implementations. The transition module of this reverse path may be described, in some examples, by the following product from time t==1 to t==T: p(y 0:T )=p(y T )Πp W (y t-1 |y t ), where p W is Gaussian.
[0194] The U-Net can have a descending phase (in which the resolution of the input data is successively reduced to the coarsest resolution level using convolutional and downsampling pooling layers, gradually revealing the global features of the input data) and an ascent phase (which receives the results of the descending phase and restores the resolution through a series of complementary stages of deconvolution and unpooling). A skip (or residual) connection may connect corresponding stages of the descending and ascent phases. An attention module may be integrated into the U-Net. A batch normalization layer may be integrated into the U-Net.
[0195] 2 shows an exemplary DDM for setup prediction. An example of data processed (across time points t) by a denoising diffusion model for a 3D oral care representation (e.g., to denoise a latent representation of the 3D oral care representation, such as a structure latent or latent mesh element). The structure latent or latent mesh element may be described by one or more latent vectors.
[0196] A denoising diffusion neural network (e.g., U-Net) may generate a target 3D oral care representation (e.g., appliance components, dental restoration designs, fixture model components (e.g., trim lines, etc.)) by iteratively denoising a set of mesh elements (e.g., points of a 3D point cloud) until a 3D oral care representation is generated that meets the specifications of the oral care arguments supplied to the model. The oral care arguments may include oral care parameters disclosed herein or other real-valued, text-based, or categorical inputs that specify intended aspects of the one or more target 3D oral care representations to be generated. In some examples, the oral care arguments may include oral care metrics that may describe intended aspects of the one or more 3D oral care representations to be generated. In some examples, a text encoder may encode a set of natural language instructions from a clinician (e.g., generate text embeddings). The text string may include tokens. The encoder for generating the text embeddings may, in some implementations, apply either average pooling or max pooling between token vectors. In some examples, a Transformer (e.g., BERT or Siamese BERT) may be trained to extract text embeddings for use in digital oral care (e.g., by training the Transformer on clinical text examples such as those shown below). In some examples, such a model for generating text embeddings may be trained using transfer learning (e.g., first trained on another corpus of text and then receiving further training on text related to digital oral care). Some text embeddings may encode text at the word level. Some text embeddings may encode text at the token level. In some implementations, a Transformer for generating text embeddings may be trained at least in part using a loss calculation (e.g., softmax loss, multiple negative ranking loss, MSE margin loss, cross-entropy loss, etc.) that compares predicted outputs to ground truth outputs.In some examples, non-textual oral care arguments, such as real or categorical values, may be converted to text and then embedded using the techniques described herein. In some implementations, oral care arguments 202 may include natural language text instructions such as the following, which may (optionally) be encoded into a latent representation by latent encoding module 212 (which may, for example, include a text transformer as described herein, among other architectures described herein): The following are examples of natural language instructions that may be provided by a clinician to the generative model described herein to describe the intended outcome of 3D oral care representation generation using the denoising diffusion probability model of the present disclosure: 1. "Generate a setup for Class I molars and canines, 2mm overbite, and add 2mm of 5-5 / 5-5 extension." 2. "Generate a setup and align it with anterior tilt and extension, finishing with 0.5mm space U2-2 for future restoration." 3. Create a bracket setup, finish with a 2mm overbite, 2mm overjet, apply Class II elastics, and ideally apply an incisal edge of L2-2.5mm. 4. "Adjust the trim line to align with the gingival margin of all teeth except #8 and #9, where the trim line should be 1 mm apical to the gingival margin." 5. "Adjust the setup to include 0.3mm IPR at L4-4 and retract to close the space to increase the overjet." 6. "Adjust the setup without moving the second molar, rotate the upper first molar mesial for Class I, lower the level to 2 mm below the reverse speaker, and advance the mandible with elastics to the Class I canine." 7. "Create a setup, intrude the molars to allow a 2mm overbite with autorotation, apply less IPR as needed for the overjet, and expand and proclinate to allow room for alignment." 8. "Create a setup, upright the premolars for a wide arch form, then expand, close the midline space with mesial translation, and intrude the lower 2-2, 2 mm." 9. "Generate a setup, which expands the arch form and applies the distal crown tip to the lower canine until it is vertical, then apply the mesial root tip for the final 10 degrees distal tip from vertical." 10. Create a setup that will extend #24, close the space to a +1mm overjet and 2mm overbite, and anteriorly tilt the upper 2-2 with 0.5mm of space distal to #8 and #9. 11. "Generate a restoration design (alternatively a veneer design) that closes the gap between teeth #8-9 by adding width evenly mesial to both teeth." 12. "Generate a restoration design (alternatively a veneer design) that adds length to teeth #6 and #11 so that they are the same length as teeth #8 and #9. Add length to the lateral incisors so that they are 1 / 2 mm shorter than the central incisors." 13. "Generate a restoration design (alternatively a veneer design) that ensures that the incisal edges of #6-11 form an even semicircle with the incisal edges of the molars (when viewed from the incisal view)." 14. "Generate a restoration design (alternatively a veneer design) for teeth #6-11, ensuring that tooth #6 is symmetrical to #11, #7 is symmetrical to #10, and teeth #8-9 are symmetrical." 15. “Generate a restorative design (alternatively a veneer design) that closes the gap between the lateral and central incisors by making the central incisor 70% of its current length. The remaining space should be added mesial to the lateral incisor. 16. “Create a porcelain crown on #12 that replicates the shape of the first premolar, has close proximal contacts, and has two points of engagement with the lower teeth. The porcelain crown should have a facial contour that blends from the canine to the first premolar to the second premolar. 17. "Create a restoration design (alternatively a veneer design) for rotated tooth #7 to add facial bulk so that the width and shape of tooth #10 are similar." 18. "Generate a restoration design (or alternatively a veneer design) so that the facial contours and positions of teeth 6, 7, 8, 10, and 11 match the facial position of tooth 9." 19. "We generate (zirconia) veneer designs and the final color follows the shade recommended by the dentist." 20. "(Zirconia) creates a design where the margins of the veneer fit smoothly and evenly to the margins already formed on the tooth." 21. "Generate a (zirconia) veneer design where the shape and contour of the veneer matches the same tooth on the other side of the arch, ensuring all proximal contacts are closed." 22. "Produce a customized crown for implanting the upper left central incisor. The crown shape should take into account the shape of the adjacent teeth and should have a space of no more than x mm (e.g., 0.1 mm) between adjacent teeth." 23. "Create a missing upper right canine and a missing upper left first molar positioned on the arch with no more than x mm (e.g., 0.05 mm) of space / impact with adjacent teeth."
[0197] A denoising diffusion probabilistic model is a class of deep generative neural networks that can be trained to generate transformations (e.g., that can be used to correct the position or orientation of a 3D representation in 3D space, such as teeth), generate 2D images (e.g., heat maps or color images), generate 3D representations (e.g., point clouds, 3D meshes, or voxelized representations), or generate other 3D oral care representations described herein. A denoising diffusion model may be trained to generate transformations that put the teeth of an arch into a pose appropriate for orthodontic treatment (e.g., intermediate or final setup). A denoising diffusion model can be implemented using one or more encoders, one or more MLPs, one or more autoencoders, one or more U-Nets, one or more transformers (e.g., 3D SWIN transform encoders or 3D SWIN transform decoders), one or more pyramid encoder-decoders, among other machine learning models. In some implementations, a denoising diffusion model (e.g., for setup prediction, etc.) may take as input one or more oral care arguments and one or more 3D representations of oral care data, such as teeth (e.g., a full arch of segmented teeth in a malocclusion position), appliance components, or fixture model components. The input teeth may be in a malocclusion position. In some examples, a malocclusion transformation of the teeth may be provided to the denoising diffusion model. Non-limiting oral care arguments may include a doctor's treatment plan (including at least zero or more treatment parameters, zero or more doctor preferences, and a set of zero or more text samples describing the nature of the intended oral care treatment, such as a final setup or restoration design generation). The oral care arguments may include real values, categorical values, natural language instructions, among others. In some implementations, the denoising diffusion model may be trained to generate a setup, such as a final setup, that meets the specifications described by the treatment parameters and / or text. A denoising diffusion model for generating orthodontic setups may be referred to as a diffusion setup neural network or a diffusion setup model.A diffusion model conditioned on text may use a neural network to reformat and / or reduce the dimensionality of that text, for example, to generate a latent encoding or embedding of the text (e.g., using a transformer or encoder to generate text embeddings).
[0198] The denoising diffusion model may include at least one of a forward pass and a reverse pass. The forward pass of the diffusion model may generate training data by iteratively adding noise (e.g., Gaussian noise) to a received 3D oral care representation (e.g., a point cloud representation of a tooth transformation or a tooth restoration design). In deployment, the reverse pass of the diffusion model may further operate through an iterative denoising process (e.g., using a purpose-trained U-Net or other model described herein), which iteratively removes noise from the received 3D oral care representation (e.g., a tooth transformation or point cloud representation, tooth restoration design). Such tooth deformation may define tooth postures in one or more arches. The diffusion setup model may generate a 3D oral care representation, such as an orthodontic setup (e.g., for final setup or intermediate staging). During training, the DDM (e.g., a diffusion setup or other described herein) may build a training dataset by incrementally adding noise to the latent vector TA generated by the latent encoding module 214. In deployment, a DDM (e.g., a diffusion setup or others described herein) may iteratively denoise one or more latent vectors TB, such as may be generated by the latent encoding module 218. The denoising may be customized, at least in part, by oral care arguments 202, which may undergo optional encoding (212) to generate latent vectors TC. The latent vectors TA and TB include dimensionally reduced information (e.g., tooth mesh information and / or tooth transformation information) for one or more 3D oral care representations. If a latent vector TA or TB corresponds to an encoded tooth transformation, then TA or TB may, in some implementations, be conditioned on (or combined with) a latent vector A corresponding to one or more 3D representations of the teeth. A diffusion model receiving one or more 3D representations of the teeth may be trained to generate dental restoration designs (e.g., using a GGDM, etc., which may also be trained to generate other types of 3D oral care representations, such as appliance components, fixture model components, transformations, or trim lines).These latent vectors TA or TB may also be conditioned on (or combined with) one or more treatment parameters K and / or one or more physician preferences L. Such conditioning may be implemented, for example, by concatenating the latent vectors TA or TB with K, L, or any of the other oral care arguments described in this disclosure, such as M, N, O, R, S, P, Q, U, V, etc.
[0199] Within a diffusion model, the latent vector TA (e.g., potentially generated using an encoder) may be subjected to many iterations of noise throughout the course of a forward pass. A series of increasingly noisy versions of TA may be used, at least in part, to train a denoising diffusion neural network (e.g., an autoencoder, a U-Net, or others described herein). The latent vector TB may start out as Gaussian noise at the beginning of the reverse pass. Through several iterations of the diffusion model reverse pass process (e.g., using a trained U-Net, an autoencoder, or other models described herein), TA may evolve into a form that can be reconstructed (e.g., using a decoder) into a 3D oral care representation that can be used in oral care appliance generation (e.g., one or more transformations that may place one or more teeth in a setup configuration (e.g., final setup or an intermediate stage)).
[0200] In some implementations, the denoising ML model 210 may use a U-Net architecture with a ResNet block and a self-attention layer as part of the reverse pass. The ResNet block refers to a residual block in which activations from one layer in a neural network are directly transferred to subsequent, deeper layers, which has the advantage of enabling deeper networks to be trained. The self-attention layer utilizes an attention mechanism that associates different parts of a sequence (i.e., particularly intermediate-stage sequences) to compute a representation of that same sequence. Self-attention is advantageous for intermediate-stage prediction in that as the diffusion model iterates, each stage of the sequence can be updated or refined with information about other stages in the sequence. The denoising ML model 210 (e.g., for setup prediction or 3D representation generation) may be trained through gradient descent and / or backpropagation. In some implementations, losses described elsewhere in this disclosure may be used to at least partially train the setup diffusion model. The L1, L2, MSE, or other losses described herein may be used to at least partially train the denoising ML model 210.
[0201] An example of a denoising diffusion model for setup prediction is shown in FIG. 3. FIG. 3 illustrates how a denoising ML model 302 can be trained on a sequence of increasingly noisy copies of patient case data (e.g., orthodontic setups). The noise can be Gaussian in some examples. FIG. 3 illustrates a Markov chain 300 including 34 time points (although other Markov chain sizes are possible). An orthodontic setup 304 (e.g., a malocclusion setup, an intermediate setup, or a final setup) can be provided to the Markov chain 300. In some implementations, the orthodontic setup 304 can be encoded 310 into a latent form before being provided to the Markov chain 300. The Markov chain 300 has length 34, although other lengths are possible. In some examples, noise can be added to the patient data starting at T0 and progressing over multiple time steps to generate a set of training data with gradually increasing amounts of noise (e.g., a setup transformation with increasingly more noise). Each time point from T0 to T33 may include a transformation (e.g., one transformation for each tooth in the arch) and / or a 3D representation of the teeth. Data from the time points in the Markov chain (e.g., tooth transformations) may be used to train a denoising ML model 302 (e.g., an encoder-decoder network such as a pyramid encoder-decoder, a 3D SWIN transformer, an autoencoder, or a U-Net) at the center of the inverse pass of the diffusion model. In some implementations, optional oral care arguments 306 (e.g., oral care parameters or oral care metrics) may be provided to the denoising ML model 302. The optional oral care arguments may be encoded into a latent format (308) in some implementations. Training may proceed over many iterations. The transformation may comprise a 4x4 affine transformation, although other dimensions are also compatible with the diffusion model-based techniques of the present disclosure. For example, the transformation may include one or more translation vectors, one or more Euler angles, and / or one or more quaternions. The "t" input is a time representation and may control which time points are sampled during a given time step of diffusion model training (e.g., training a denoising neural network). The "t" input may be used for position encoding.In some implementations, the time t may correspond to a stage in orthodontic treatment. The use of position encoding similar to that used in the transformer may enable the model to recognize the relative position of the stage in the treatment planning sequence. Through this approach, the model may learn how to proceed through orthodontic treatment stage by stage, leveraging what the diffusion model learned from previous orthodontic treatment examples in the training dataset.
[0202] FIG. 4 illustrates an example of a mesh element labeling model based on a denoising diffusion probability model (DDPM) that can be used to enable either 3D mesh segmentation (e.g., semantic segmentation) and a 3D mesh cleanup pipeline. DDPM for 3D representation segmentation may include at least one of a forward pass 206 (e.g., which may include many steps of iteratively adding more noise to the input 3D representation and may be used in training an ML model to perform the reverse process) and a reverse pass 208 (e.g., which may use an ML model such as a neural network to iteratively denoise the noisy representation, resulting in a segmentation of the input 3D representation). DDPM for 3D representation segmentation, in some implementations, may approximate aspects of a Markov process. HNNFEM 406, described in FIG. 4, may be trained to perform the reverse pass 208. Dental arch data 400 may (optionally) be converted into a list of mesh elements (402). The mesh elements may undergo optional mesh element feature vector calculation (404). Mesh element features described herein may be provided to HNNFEM 406; for example, mesh element features include edge midpoints, edge curvatures, signed dihedral angles, edge lengths, edge normals, or others described herein. DDPM may start with a noisy representation (e.g., a noisy representation of mesh element labels) and iteratively denoise that representation to generate a denoised representation that may include one or more mesh element labels or object masks (alternatively, the noisy representation may be denoised directly into one or more segmented 3D representations, such as a 3D point cloud or a 3D mesh). Equations formalizing the forward and reverse passes are described herein.
[0203] In some implementations, DDPM for 3D representation segmentation may include the following steps: iteratively add noise (e.g., Gaussian noise) to an input 3D oral care representation to generate a series of increasingly noisy versions of the input 3D oral care representation, which can then be used to train a denoising neural network for the inverse process. For each 3D representation in the series of increasingly noisy 3D representations of oral care data: 1) use HNNFEM 406 to extract feature map hierarchical neural network features 408; 2) collect mesh element-level representations 410 by upsampling the feature maps from HNNFEM and concatenating the upsampled feature maps into a mesh element representation vector 410; 4) the mesh element representation vector 410 from step #3 may be used to train one or more ML models for mesh element labeling (e.g., to train an ensemble of neural networks 412 for mesh element labeling). The one or more ML models may vote (414) to determine the final output mesh element label 416. The one or more ML models may generate labels for one or more mesh element features. The generated set of mesh element labels may, in some implementations, include one or more object masks. The object masks may label or otherwise indicate which mesh elements belong to which objects in the scene (e.g., which mesh elements belong to the upper left central incisor, gums, or hardware elements in an arch mesh from an intraoral scanner). FIG. 4 further illustrates input of a 3D oral care representation 400 (e.g., a dental arch) to be segmented, which may (optionally) be placed into a vector of mesh elements (402). One or more mesh element features may be calculated for each mesh element (404).
[0204] The HNNFEM 406 may include one or more U-Nets 502, one or more pyramid encoder-decoders from FIG. 6, one or more 3D SWIN transform decoders, or other architectures described herein. In FIG. 5, mesh elements of the 3D representation 500 may be provided to a U-Net 502, which may extract hierarchical features using a U-shaped structure of convolution 504 / deconvolution 512 and / or pooling 506 / unpooling operations 510. The most global features are extracted at the lowest resolution (508). The vectors 414 or 514 containing the hierarchical neural network features may be assembled into a mesh element representation vector 416 and provided to an ensemble of ML models 418 (e.g., to resolve mesh element labels). Voting (420) may be performed to resolve ambiguities between mesh element labels predicted by the ensemble of ML models. The final predicted mesh element labels 422 are sent to the output.
[0205] In some implementations (e.g., when a pre-segmentation dental arch is segmented into teeth and gums), the training data may include at least a 3D mesh of the arch and a set of mesh element labels (or object masks) indicating which tooth each mesh element belongs to. This set of mesh element labels (or masks) may be subjected to the iterative addition of noise, such as in the forward pass shown in FIG. 2. The resulting set of increasingly noisy mesh element labels may be used as training data for training a denoising diffusion ML model (e.g., U-Net). Such a denoising diffusion ML model may take as input a noisy (or random) set of mesh element labels (or object masks) for an instant 3D mesh (e.g., of a pre-segmentation dental arch) and be trained to gradually denoise that set of mesh element labels (or object masks) over many iterations. The denoised set of labels (or object masks) may be output by the reverse pass of FIG. 2 and used to segment the instant mesh (or point cloud or other 3D representation). In some examples, the set of mesh element labels (or object masks) may undergo encoding using an encoder, which may provide a latent representation to the forward pass. After the reverse pass is completed, in unfolding, the denoised representation may undergo reconstruction using a decoder. The mesh element labels may be used to segment the teeth of the arch before segmentation. The mesh element labels may, in some examples, be used to label mesh elements as part of a mesh cleanup operation (e.g., labeling mesh elements for removal, smoothing, scaling, or other modification following artifact generation for use in digital oral care).
[0206] The disclosed system can train a denoising diffusion model to segment a 3D representation of oral care data, such as a 3D mesh of a patient's dentition. Figure 4 relates to 3D mesh segmentation and 3D mesh cleanup. Both of these techniques share an important attribute: labeling of 3D mesh elements. In the case of 3D mesh segmentation (e.g., as related to segmenting an oral care mesh such as teeth), various mesh elements of the mesh may be labeled according to the portion of the dental anatomy to which they belong (e.g., gum, upper right central incisor, lower left second premolar). Teeth mesh segmentation may label elements according to their membership in various teeth as specified by one or more of the dental notation systems (e.g., Palmer) mentioned herein. In some implementations, for face-lingual segmentation, each of the mesh elements may be labeled according to its membership in either the facial side of the arch or the lingual side of the arch. Other dental anatomy segmentation implementations are possible. In some implementations, oral care arguments 202, such as a decimation factor (or decimation percentage), may be provided to the denoising diffusion model for segmentation or mesh cleanup as described herein. The decimation factor may control the degree to which the pre-segmentation mesh is decimated before labeling of the mesh elements occurs.
[0207] A pre-segmentation mesh (e.g., a dental arch generated by an intraoral scanner or a CT scanner) may be received by the segmentation system. One or more mesh element feature vectors may be calculated for each mesh element. One or more mesh element features from elsewhere in this disclosure may be used to form the mesh element feature vector. A mesh element may include edges, faces, vertices, or voxels (or any combination thereof).
[0208] In some implementations, the list of mesh element feature vectors can be provided to a neural network that refines the mesh element vectors to extract local and global features from them, such as the U-Net structure shown in FIG. 4. The U-Net in HNNFEM can include 3D mesh convolution and / or subsequent 3D mesh pooling operations, which can function to reduce the resolution of the mesh and extract neural network features at increasingly global scales. After a series of such operations, once the most global neural network features have been extracted, there can be a series of 3D mesh unpooling and 3D mesh deconvolution operations that return the mesh to its original scale. After various levels of the U-Net structure generate outputs, there can be scaling-up and concatenation operations. The HNNFEM can generate one or more per-3D mesh element feature vectors, which may have been scaled up to the original mesh resolution. These feature vectors can be concatenated and provided as input to one or more ML models for mesh element classification or labeling. Any of the supervised ML models disclosed elsewhere in this disclosure, such as SVM or logistic regression, can be used for this classification. In some examples, an ensemble of ML classifiers may take mesh element representation vectors as input and generate mesh element classification labels. In some implementations, an ensemble of fully connected neural networks may be used for such classification. In some cases, a multilayer perceptron (MLP) may be used for this classification, including a linear layer and an associated ReLU activation function (e.g., which has the advantage that not all neurons fire simultaneously, allowing for a more tailored response to the input than activation functions that all fire for every evaluation), as well as a batch normalization operation. Other activation functions, such as those mentioned elsewhere in this disclosure, are also possible. Each of the ensemble of ML models may generate a predicted class label for each mesh element. A voting mechanism may then be used to combine these results and output a final class label prediction for each 3D mesh element.Some implementations may use a transformer to assist in applying labels to mesh elements.
[0209] Implementations of mesh cleanup and mesh segmentation may, in some examples, differ in the placement of ground truth mesh element labels. In the case of tooth segmentation, ground truth data may be provided for each arch mesh to be segmented (i.e., each tooth in the arch has a mesh element labeled according to which tooth it belongs to). Such losses may then be calculated for face-tongue segmentation and / or teeth-gum segmentation (as defined in U.S. Provisional Application No. 63 / 366,490). Various loss functions of this disclosure may be used to compare the mesh element labels of a predicted segmentation with the mesh element labels of the corresponding ground truth segmentation. Cross-entropy loss is one example of some candidate loss functions for this comparison. Other possible losses are disclosed elsewhere in this disclosure. In the case of mesh cleanup, ground truth data may be provided for each arch mesh to be cleaned. For example, for foreign body removal, each mesh element corresponding to a foreign body in the arch may be labeled as such, and each mesh element that does not contain a foreign body (i.e., is retained after mesh cleanup) is labeled as such. For divot removal, each mesh element corresponding to a divot in the arch may be labeled as such, and each mesh element that does not contain a divot (i.e., is retained after mesh cleanup) is labeled as such.
[0210] After the segmentation labels are applied, a process is performed that copies mesh elements of a particular label to a new mesh (a.k.a., mesh cutting), and the mesh is saved (e.g., to an electronic storage medium) for further processing. Mesh Cleanup After labeling is complete, a process is performed that removes each designated mesh element from the mesh (e.g., using classical mesh processing techniques), and then optionally a process may be performed that fills any holes that may have been created by that process (e.g., using techniques described elsewhere in this disclosure that are trained for mesh infill).
[0211] Mesh element labeling for mesh segmentation may also be achieved using a purpose-trained autoencoder, such as a variational autoencoder, as described elsewhere in this disclosure. More generally, an encoder-decoder network may be trained for mesh element labeling.
[0212] In some implementations, the denoising diffusion techniques of the present disclosure can be trained to predict arch forms. The arch forms can have shapes that describe aspects of the shape of a dental arch. The arch forms can include data structures such as one or more 3D meshes (e.g., meshes aligned with at least one coordinate axis of at least one tooth's local coordinate system), 3D point clouds, 3D polylines, or sets of control points (e.g., control points defining a spline). Such data structures can be generated, for example, using the denoising diffusion method described in FIG. 2.
[0213] In some implementations, the disclosed denoising and diffusion techniques can be trained to generate information related to IPR for use in setup prediction. The information related to IPR can include one or more of: 1) IPR cut surfaces of a particular tooth; 2) a measure of the amount of IPR applied to the side of a particular tooth; 3) a designation of whether a particular tooth will receive IPR, including an indication of whether IPR will be applied to the mesial or distal side of the tooth; or 4) a designation of one or more stages at which a particular tooth will receive IPR. Training data for such a model can be generated, at least in part, by iteratively adding noise to 1) representations of one or more IPR cut surfaces from cases in a training dataset, 2) representations of one or more teeth or other aspects of the patient's residence, or 3) one or more tooth deformations. An ML model (e.g., a U-Net) can be trained to iteratively denoise an initial representation of one or more such examples of the training data (e.g., the IPR cut surfaces can be initialized, at least in part, randomly or using default values). Upon completion of the reverse pass 208, one or more generated (or modified) data 220 related to IPR can be output. For example, IPR cutting planes can be generated by the denoising diffusion techniques described herein. In some implementations, a denoising diffusion model can be trained to render a determination of whether a target tooth is accessible for IPR. Information about IPR can be generated, for example, using the denoising diffusion method described in FIG. 2. For example, the method of FIG. 2 can predict an orthodontic setup with any associated IPR information. For each tooth, a setup transformation can be generated along with one or more associated items of IPR information (e.g., one or more IPR cutting planes, a flag indicating whether IPR should be performed, a list of conditions for performing IPR, etc.).
[0214] In some implementations, the techniques of the present disclosure can be trained to predict trim lines. The trim lines may define cutting paths for removing a thermoformed tray (e.g., a tray for orthodontic aligner treatment or an indirect bonding tray for delivering brackets to teeth for orthodontic treatment) from a fixture model. One or more polylines (or meshes or point clouds) can be used to define the trim lines. Such polylines (or meshes or point clouds) can be generated, for example, using the denoising diffusion method described in FIG. 2.
[0215] In some implementations, the denoising diffusion techniques of the present disclosure may be trained to predict one or more local Cartesian coordinate axes of a tooth (e.g., to predict one or more of the X, Y, and Z Cartesian axes of a tooth). In other implementations, the denoising diffusion techniques of the present disclosure may be trained to predict one or more arch form coordinate axes. The position along the arch form coordinate axes can include a tuple [l, d, e] relative to a reference arch form spline S that approximates the shape of the dental arch. The rotation can include a tuple [a, b, g] representing alpha, beta, and gamma rotations. Alpha describes the rotation about the l axis. Beta describes the rotation about the d axis. Gamma describes the rotation about the e axis. A complete tuple to describe the position and rotation can include [l, d, e, a, b, g]. p is a point along S with arch length l. d is the distance between the tooth origin t and the reference arch form spline S. The tooth origin t is obtained by translating a distance "d" upward along the d-axis, then a distance "e" along the e-axis. The e-axis is perpendicular to the d-axis and l-axis and can be defined as coming out of or into the page. e represents protrusion. l represents the length across the archform spline. d represents the distance from the archform spline. Such an archform coordinate system is described in commonly assigned U.S. Patent Application Publication No. 20210259808, which is incorporated herein by reference in its entirety. The coordinate system described herein can facilitate subsequent automation for patient oral care procedures, such as automated setup prediction. One or more transforms, vectors, or tuples can be used to define the coordinate system (or define a pose within the coordinate system). Such transforms (or vectors or tuples) can be generated, for example, using the denoising diffusion method described in FIG. 2.
[0216] Table 3 lists input data and generated data for some non-limiting examples of the generation techniques described herein. The denoising diffusion models described herein, in some implementations, can be trained to generate (or modify) the input data in Table 3, resulting in the generated data in Table 3.
[0217] The techniques of this disclosure may be trained to generate (or modify) a point cloud (e.g., where points can be described as 1D vectors such as (x, y, z)), a polyline (points connected in order by edges), a 3D mesh (points connected via edges to form a surface), a spline (which may be computed through a set of generated control points), a sparse voxelized representation (which may be described as a set of points corresponding to each voxel's centroid or some other landmark of the voxel, such as the voxel's boundary), one or more transformations (which may take the form of one or more 1D vectors or one or more 2D matrices, such as a 4x4 matrix), etc. In some implementations, the voxelized representation may be computed from the 3D point cloud or the 3D mesh. In some implementations, the 3D point cloud may be computed from the voxelized representation. In some implementations, the 3D mesh may be computed from the 3D point cloud.
[0218] [Table 3]
[0219] In some implementations, an automated setup prediction model (e.g., a denoising diffusion probability model) may be trained to generate setups with customized speaker profiles (e.g., speaker profiles that match the intended outcome of a patient's treatment). Such a model may be trained on cohort patient case data. One or more oral care metrics may be calculated for each case to quantify or measure aspects of that case's speaker profile. During training, one or more of such metrics may be provided to the setup prediction model to, for example, influence a model regarding the geometry and / or structure of each case's speaker profile. During deployment of the setup prediction model, that same input path to the trained neural network may be configured with one or more values as instructions to the model regarding the intended speaker profile. Such values may automatically generate setups with speaker profiles that meet the aesthetic and / or medical treatment needs of a particular patient case.
[0220] In some implementations, the occlusal metric may measure the curvature of the occlusal or incisal surfaces of teeth on either the left or right side of the arch relative to the occlusal plane. The occlusal plane may, in some instances, be calculated as a surface averaging the incisal or occlusal surfaces of the teeth (for one or both arches). In some implementations, the curvature metric may be calculated along a normal vector, such as a vector perpendicular to the occlusal plane. In other implementations, the curvature metric may be calculated along the normal vector of another plane. In some implementations, an XY plane may be defined to correspond to the occlusal plane. An orthogonal plane may be defined as a plane perpendicular to the occlusal plane, which also passes through the occlusal line segment, the occlusal line segment defined by a first endpoint that is a landmark point on the first tooth (e.g., the canine) and a second endpoint that is a landmark point on the most posterior tooth on the same side of the arch. In some implementations, the landmark point may be located along the incisal edge of the tooth or on the cusp of the tooth. In some instances, landmarking points on intermediate teeth (e.g., teeth located between the first and rearmost teeth) on either the left or right side of the arch may form a curved path that may be described by a polyline. The following is a non-limiting list of speakerb oral care metrics:
[0221] 1) Measure the vertical height between the line segment and the point. In other words, measure the distance between the line segment and the point along the z-axis. The line segment is defined by connecting the highest cusp of the most posterior tooth (lower arch) with the cusp of the first tooth (lower arch) on that side. Considering the subset of teeth between the first tooth and the rearmost tooth, the point is defined by the highest cusp of the lowest tooth in this subset. In other words, the speakerb metric can be calculated using the following four steps: i) Line: Form a line between the highest cusp of the rearmost tooth and the cusp of the first tooth. ii) Curve_Point_A: Considering the set of teeth between the rearmost tooth and the first tooth, find the highest point of the lowest tooth. iii) Curve_Point_B: Project Curve_Point_A onto the line to find the point (Curve_Point_B) along the line that is closest to Curve_Point_A. iv) Speaker: Find the height difference between Curve_Point_B and Curve_Point_A.
[0222] 2) Project one or more intermediate landmark points (e.g., points on the teeth between the first tooth and the rearmost tooth on that side of the arch) and the septum line segment onto an orthogonal plane. Calculate the septum metric by measuring the distance between the farthest of the projected intermediate points and the projected septum line segment. This provides a measure of the curvature of the arch relative to the orthogonal plane.
[0223] 3) Project one or more intermediate landmark points and the occlusal curve segment onto the occlusal plane. Calculate the occlusal curve in this plane by measuring the distance between the farthest of the projected intermediate points and the projected occlusal curve segment. This provides a measure of the curvature of the arch relative to the occlusal plane.
[0224] 4) Skip the projection and calculate distance and curvature in 3D space. Calculate the spe- cific arc by measuring the distance between the farthest midpoint and the spe- cific arc line segment. This gives us a measure of the curvature of the arch in 3D space.
[0225] 5) Calculate the slope of the projected speaker line segment on the occlusal plane.
[0226] 6) Calculate the slope of the projected speaker segment in the orthogonal plane.
[0227] Speaker metrics 5 and 6 may help the network reduce some additional degrees of freedom in defining how the patient's arches curve at the back of the mouth.
[0228] The techniques described herein (e.g., denoising diffusion probability models) can, in some implementations, be trained to generate transformations that can place a patient's teeth in a pose suitable for use in an orthodontic setup (e.g., an intermediate or final setup) according to specifications of oral care arguments that can be provided to the generative model. Optional oral care arguments 202 can be provided to a trained denoising ML model 210 (e.g., a U-Net) that performs a reverse pass 220. The oral care arguments can include oral care parameters disclosed herein or other real-valued, text-based, or categorical inputs that specify intended aspects of one or more 3D oral care representations to be generated. In some examples, the oral care arguments can include oral care metrics that can describe intended aspects of one or more 3D oral care representations to be generated. Oral care arguments are particularly adapted to the implementations described herein. For example, the oral care arguments can specify the intended design (e.g., including shape and / or structure) of a 3D oral care representation that can be generated (or modified) according to the techniques described herein. In sum, implementations that use the specific oral care arguments disclosed herein generate more accurate 3D oral care representations than implementations that do not use the specific oral care arguments. In some examples, a text encoder may encode a set of natural language instructions from a clinician (e.g., generate text embeddings). The text string may include tokens. In some implementations, the encoder for generating the text embeddings may apply either average pooling or max pooling between token vectors. In some examples, a transformer (e.g., BERT or Siamese BERT) may be trained to extract text embeddings for use in digital oral care (e.g., by training the transformer on clinical text examples such as those shown below). In some examples, such a model for generating text embeddings may be trained using transfer learning (e.g., first trained on another corpus of text and then receiving further training on text related to digital oral care). Some text embeddings may encode text at the word level.Some text embeddings may encode text at the token level. In some implementations, the transformer for generating the text embeddings may be trained, at least in part, using a loss calculation (e.g., softmax loss, multiple negative ranking loss, MSE margin loss, cross-entropy loss, etc.) that compares predicted outputs to ground truth outputs. In some examples, non-text arguments, such as real-valued or categorical values, may be converted to text and then embedded using the techniques described herein. The following are examples of natural language instructions that may be issued by a clinician to a generative model described herein: "Generate setup, set to Class I molars and canines, 2 mm overbite, add 2 mm extension 5-5 / 5-5," "Generate setup, align with anterior inclination and extension, and finish with 0.5 mm space U2-2 for future restoration," or "Adjust setup without second molar movement, rotate upper first molar mesial for Class I, lower level to 2 mm below reverse speaker, and advance mandible to Class I canine using elastic."
[0229] The techniques of this disclosure, in some implementations, may use PointNet, PointNet++, or a derived neural network (e.g., a network trained via transfer learning using either PointNet or PointNet++ as a basis for training) to extract local or global neural network features from a 3D point cloud or other 3D representation (e.g., a 3D point cloud describing aspects of a patient's dentition, such as teeth or gums). The techniques of this disclosure, in some implementations, may use U-Net to extract local or global neural network features from a 3D point cloud or other 3D representation.
[0230] 3D oral care representations are described herein as such because three-dimensional representations are currently state of the art. Nevertheless, it should be understood that 3D oral care representation is intended to be used non-limitingly to encompass any representation in three or higher dimensions (e.g., 4D, 5D, etc.), and machine learning models can be trained using the techniques disclosed herein to operate on higher dimensional representations.
[0231] In some examples, the input data may include 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data related to splines (e.g., control points). An encoder-decoder structure may include one or more encoders or one or more decoders. In some implementations, an encoder may take as input one or more mesh element feature vectors of the input mesh elements. By processing the mesh element feature vectors, the encoder is trained to generate a more accurate representation of the input data. For example, the mesh element feature vectors may provide the encoder with more information about the shape and / or structure of the mesh; thus, the additional information provided enables the encoder to make better-informed decisions and / or generate a more accurate latent representation of the mesh. Examples of encoder-decoder structures include (among others) a U-Net, an autoencoder, or a transformer. The representation generation module may include one or more encoder-decoder structures (or portions of an encoder-decoder structure, such as individual encoders or individual decoders). The representation generation module can generate information-rich (optionally dimensionally reduced) representations of the input data, which can be more easily consumed by other generative or discriminative machine learning models.
[0232] A U-Net may comprise an encoder followed by a decoder. The architecture of a U-Net may resemble a U. The encoder may extract one or more global neural network features, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level as opposed to the most global level) from the input 3D representation. The output from each level of the encoder may be passed to the input of a corresponding level of the decoder (e.g., by a skip connection). Similar to the encoder, the decoder may operate on multiple levels of neural network features from global to local. For example, the decoder may output a representation of the input data that may include global, intermediate, or local information about the input data. In some implementations, a U-Net can generate an information-rich (optionally dimensionally reduced) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[0233] An autoencoder can be configured to encode input data into a latent form. The autoencoder can train an encoder to reformat the input data between the encoder and decoder into a dimensionally reduced latent form, and then train a decoder to reconstruct the input data from that latent form of the data. A reconstruction error can be calculated to quantify the degree to which the reconstructed form of the data differs from the input data. The latent form, in some implementations, can be used as an information-rich, dimensionally reduced representation of the input data, which can be more easily consumed by other generative or discriminative machine learning models. In most scenarios, an autoencoder can be trained to input a 3D representation, encode the 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close replica of the input 3D representation as output.
[0234] A Transformer can be trained to at least partially generate a representation of its input using self-attention. The Transformer can encode long-range dependencies (e.g., encode relationships between multiple inputs). The Transformer may comprise an encoder or a decoder. Such an encoder, in some implementations, may operate in a bidirectional manner or may operate a self-attention mechanism. Such a decoder, in some implementations, may operate a masked self-attention mechanism, a mutual attention mechanism, or may operate in an autoregressive manner. The self-attention operation of the Transformers described herein, in some implementations, can associate different positions or aspects of a 3D oral care representation to compute a reduced-dimensional representation of the individual 3D oral care representation. The mutual attention operation of the Transformers described herein, in some implementations, may mix or combine aspects of two (or more) different 3D oral care representations. The autoregressive operations of the Transformers described herein, in some implementations, may consume previously generated aspects of the 3D oral care representation (e.g., previously generated points, point clouds, transformations, etc.) as additional inputs when generating new or modified 3D oral care representations. The Transformers, in some implementations, may generate a latent form of the input data, which can be used as an information-rich, reduced-dimensionality representation of the input data that can be more easily consumed by other generative or discriminative machine learning models.
[0235] In some implementations, the encoder-decoder structure may initially be trained as an autoencoder. During deployment, one or more modifications may be made to the latent form of the input data. This modified latent form may then proceed to be reconstructed by the decoder, resulting in a reconstructed form of the input data that differs from the input data in one or more intended ways. Oral care arguments, such as oral care parameters or oral care metrics, may be supplied to the encoder, decoder, or used in modifying the latent form to influence the encoder-decoder structure in generating a reconstructed form with desired properties (e.g., properties that may differ from those of the input data).
[0236] The techniques of the present disclosure may, in some examples, be trained using federated learning. Federated learning may enable multiple remote clinicians to iteratively improve machine learning models (e.g., 3D oral care representation validation, mesh segmentation, mesh cleanup, other techniques involving mesh element labeling, coordinate system prediction, non-organic object placement on teeth, appliance component generation, dental restoration design generation, techniques for placing 3D oral care representations, setup prediction, generating or modifying 3D oral care representations using autoencoders, generating or modifying 3D oral care representations using transformers, generating or modifying 3D oral care representations using diffusion models, 3D oral care representation classification, and missing value imputation) while preserving data privacy (e.g., there may be no need to send clinical data "over the wire" to a third party). Data privacy is particularly important for clinical data protected by applicable law. Clinicians may receive a copy of the machine learning model, use a local machine learning program to further train the ML model using locally available data from their local clinic, and then send the updated ML model back to a central hub or a third party. A central hub or third party can consolidate updated ML models from multiple clinicians into a single updated ML model that benefits from learning from recently collected patient data at various clinical sites. In this way, new ML models can be trained that benefit from additional and updated patient data (possibly from multiple clinical sites), without these patient data actually being transmitted to a third party. Training on local in-clinic devices may, in some instances, occur when the device is idle or otherwise during off-hours (e.g., when patients are not being treated in the clinic). Devices in clinical environments for collecting data and / or training ML models for the techniques described herein may include intraoral scanners, CT scanners, X-ray machines, laptop computers, servers, desktop computers, or handheld devices (such as smartphones with image collection capabilities).In addition to associative learning techniques, in some implementations, contrastive learning may be used to at least partially train the ML models described herein. Contrastive learning may, in some instances, augment samples in a training dataset to highlight differences between samples from different classes and / or to increase similarity between samples of the same class.
[0237] In some examples, a local coordinate system for a 3D oral care representation, such as a tooth, can be described by one or more transformations (e.g., an affine transformation matrix, a translation vector, or a quaternion). The system of the present disclosure can be trained for coordinate system prediction using historical cohort patient case data. The historical patient data can include at least one or more tooth meshes or one or more ground truth tooth coordinate systems. A machine learning model, such as a U-Net, an encoder, an autoencoder, a pyramid encoder-decoder, a transformer, or a convolutional and / or pooling layer, can be trained for coordinate system prediction. Representation learning can determine a representation of the tooth (e.g., encoding a mesh or point cloud into a latent representation using a U-Net, an encoder, a transformer, a convolutional and / or pooling layer, etc.), and then predict a transformation of that representation that defines the representation's local coordinate system (e.g., including one or more coordinate axes) (e.g., using a trained multilayer perceptron, a transformer, an encoder, a transformer, etc.). When a coordinate system is predicted for a tooth mesh, the mesh convolution techniques described herein can exploit the invariance to rotation, translation, and / or scale of the tooth mesh to produce predictions that techniques that are not invariant to rotation, translation, and / or scale of the tooth mesh cannot produce. Pose translation techniques can be trained for coordinate system prediction in a manner that predicts tooth transformations. Reinforcement learning techniques can be trained for coordinate system prediction in a manner that predicts tooth transformations.
[0238] A machine learning model, such as a U-Net, encoder, autoencoder, pyramid encoder-decoder, transformer, or convolutional and / or pooling layer, can be trained as part of the method for hardware (or appliance component) placement. Representation learning can train a first module to determine an embedding representation of the 3D oral care representation (e.g., using an autoencoder, or encoding a mesh or point cloud into a latent format using blocks of U-Net, encoder, transformer, convolutional and / or pooling layers, etc.). The representation can include a dimensionally reduced and / or information-rich version of the input 3D oral care representation. In some implementations, generation of the representation can be aided by calculation of a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some implementations, the representation can be calculated for the hardware element (or appliance component). Such representations are suitable for being provided to a second module that may perform generation tasks such as transformation prediction (e.g., transformations for positioning a 3D oral care representation relative to another 3D oral care representation, such as for positioning a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such transformations may include affine transformation matrices, translation vectors, quaternions, etc. Machine learning models that may be trained to predict transformations for positioning hardware elements (or appliance components) relative to elements of the patient's dentition include MLPs, transformers, encoders, etc. The disclosed system may be trained for 3D oral care appliance placement using historical cohort patient case data. The historical patient data may include at least one or more ground truth transformations and one or more 3D oral care representations (such as tooth meshes or other elements of the patient's dentition). When a U-Net (among other neural networks) is trained to generate a representation of a tooth mesh, the mesh convolution and / or mesh pooling techniques described herein leverage the invariance of that tooth mesh to rotation, translation, and / or scaling to generate predictions that techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh cannot produce.Postural locomotion techniques can be trained for placement of hardware or prosthetic components. Reinforcement learning techniques can be trained for placement of hardware or prosthetic components.
[0239] The denoising diffusion techniques described herein (e.g., the method of FIG. 2 ) may, in some implementations, be trained to fill in missing aspects of a 3D representation of a patient's dentition (e.g., of a segmented tooth, or of a digital fixture model or digital fixture model component). In some examples, a 3D scan of a tooth crown (or a fixture model or other 3D oral care representation described herein) may be missing portions of one or more surfaces (e.g., due to occlusion by surrounding teeth during an intraoral scan). The denoising diffusion model may generate training data by iteratively degrading or adding noise to the 3D representation of the tooth (e.g., by removing patches or portions of the crown or root surface). A sequence of progressively noisier versions of the 3D representation of the tooth from the training dataset may be generated. Noise may be added to the point cloud / mesh / voxelized representation of the tooth by, among other methods: 1) adding random translations to one or more mesh elements; 2) applying small random scaling or distortions to aspects of the tooth; or 3) removing mesh elements (e.g., creating holes). This sequence of progressively noisier versions of the teeth may then be used to train a denoising diffusion neural network (e.g., a U-Net or other encoder-decoder structure such as a pyramidal encoder-decoder, transformer, or autoencoder) to reconstruct the noisy teeth (e.g., in small steps). A set of noisy mesh elements (e.g., initialized from teeth with missing data) may be provided to the denoising diffusion neural network, which may be trained to iteratively remove noise and / or restore the tooth surfaces. After multiple iterations of progressive restoration, restored teeth (or other 3D oral care representations) are output by the denoising diffusion model and may be used in subsequent digital oral care processing to generate oral care appliances.
[0240] The fixture model components can be generated or modified using the disclosed denoising diffusion techniques (e.g., the method of FIG. 2). The digital fixture model can include a 3D representation of the patient's dentition, with optional fixture model components attached to the dentition. The disclosed 3D representation generation techniques (e.g., denoising diffusion techniques for generating 3D point clouds, 3D voxelized representations, or other representations) can be trained to generate (or modify) aspects of the digital fixture model to include treatment-enhanced fixture model components. The fixture model components can include 3D representations (e.g., 3D point clouds, 3D meshes, or voxelized representations) of one or more of the following non-limiting items: 1) Interdental webbing—can fill spaces between teeth or smooth gaps between teeth to ensure aligner removability. 2) Blockout—this can be added to the fixture model to remove overhangs that may interfere with thermoforming of plastic trays or to ensure aligner removability. 3) Occlusal Block - An occlusal feature on a molar or premolar intended to support an open bite. 4) Occlusal Ramp - A lingual feature on an incisor or canine intended to support an open bite. 5) Interproximal Reinforcement - A structure on the exterior of an oral care appliance (e.g., an aligner tray) that may extend from the gingival margin of a first tooth on the labial side of the appliance body, along the interproximal region between the first and second teeth, to the gingival margin of a second tooth on the lingual side of the appliance body. The effect of the interproximal reinforcement on the appliance body in the interproximal region may be stiffer than the labial and lingual surfaces of the first shell. This allows the aligner to more firmly grip the teeth on both sides of the reinforcement. 6) Gingival Margin - A structure that may extend along the gingival margin of a tooth in a mesial-distal direction to strengthen the engagement between the aligner and a given tooth. 7) Torque Point - A structure that may increase the force delivered to a given tooth at a specific location. 8) Power Ridges - Structures that can increase the force delivered to a given tooth at specific locations. 9) Dimples - Structures that can increase the force delivered to a given tooth at specific locations. 10) Digital Pontic - Structures that can leave space open or reserve space in the arch, such as for partially erupted teeth.In an aligner, a physical pontic is a tooth pocket that does not cover the tooth when the aligner is attached to the tooth. The tooth pocket may be filled with tooth-colored wax, silicone, or composite material to provide a more aesthetic appearance. 11) Power Bar - A blockout added to a toothless space to provide strength and support to the tray. The power bar can fill the void. An abutment or healing cap can be blocked out with the power bar. 12) Trim Line - A digital path along the digital fixture model that can roughly follow the contour of the gums (e.g., biased 1 or 2 mm toward the gums). After 3D printing, the trim line can define a path along which the clear aligner can be cut or separated from the physical fixture model. 13) Undercut Fill - Material added to the fixture model to avoid the formation of cavities between the height of the fixture model's contour and another boundary (e.g., the gums, or a plane that lies beneath the physical fixture model after 3D printing).
[0241] The techniques of the present disclosure may also generate (or modify) other shapes that penetrate tooth pockets or reinforce oral care appliances (e.g., orthodontic aligner trays). The techniques of the present disclosure, in some implementations, can determine the location, size, and / or shape of fixture model components to generate a desired treatment result.
[0242] Interdental webbing can be generated using the techniques of the present disclosure, for example, during the fixture model quality control stage of orthodontic aligner fabrication. Interdental webbing is material (e.g., which may include mesh elements such as vertices / edges / faces, among others) that smooths or securely fills the interdental areas of the fixture model. Interdental webbing is additional material added to the interdental areas of the teeth in a digital fixture model (e.g., which may be 3D printed and rendered in physical form) to reduce the tendency of an aligner, retainer, mounting template, or bonding tray to lock onto the physical fixture model during the orthodontic aligner thermoforming process. Interdental webbing can improve the function of the aligner tray by improving the tray's ability to slide off the fixture model after thermoforming or by improving the tray's fit to the patient's teeth (e.g., making it easier to insert or remove the tray from the teeth).
[0243] Blockouts (which may include mesh elements such as vertices / edges / faces, among others) can be generated using the techniques of the present disclosure. Blockouts are material that can be added to undercut areas of a digital fixture model so that an aligner, retainer, mounting template, or bonding tray, 3D printed mold, or other oral care appliance does not lock onto the physical fixture model during the orthodontic aligner thermoforming process. Blockouts may also be generated to fill portions of the digital fixture model with undercuts so that the thermoformed aligner tray does not grab onto the undercut. Blockouts can improve the function of the aligner tray by improving the tray's ability to slide off the fixture model after thermoforming. Blockouts can facilitate later thermoforming (e.g., help prevent the aligner tray from sticking onto the physical fixture model).
[0244] A pontic design can be generated using the techniques of the present disclosure. A pontic is a digital 3D representation of a tooth that can serve as a placeholder in the arch. In an aligner, the pontic can function as a tooth pocket that can be filled with a tooth-colored material (e.g., wax) to improve aesthetics. The pontic can leave a space open in the arch during orthodontic setup generation. When automated setup prediction is performed, transformations for intermediate stages can be generated. One or more pontic teeth can be defined to hold an open space for a missing tooth or to act as a placeholder as a space closes or opens in the arch during predicted setup staging, or an unerupted tooth can erupt into the space where the pontic exists. A pocket is created for the tooth to erupt into.
[0245] A digital pontic may be placed in (or generated in) an arch (e.g., during setup generation or fixture model generation) to reserve space in the arch for a missing or extracted tooth (e.g., to prevent adjacent teeth from encroaching on that space over the course of successive intermediate stages of orthodontic treatment). In some implementations, a digital pontic may be used when a space (e.g., a space to be held open) is at least a threshold dimension (e.g., 4 mm wide, among others). In some examples, a UL4-UR4 or LL4-LR4 digital pontic may be placed (or generated or modified) when space is available or when there is a partially erupted tooth in the space. If there is a partially erupted tooth in the space, the digital pontic may be placed over the erupted tooth to maintain space for the erupted tooth to erupt. The pontic may be placed (or created or modified) to be inside the gums, to minimize (or avoid) heavy occlusal contacts (e.g., contact between the chewing surfaces of the upper or lower arch), or to cover erupted teeth (if present), among other conditions.
[0246] Aspects of the present disclosure can provide a technical solution to the technical problem of generating a 3D oral care representation using a denoising diffusion model. The 3D oral care representation that can be generated includes mesh element labels for segmentation or mesh cleanup, transformations for use in setup generation (or fixture model or appliance generation), coordinate systems (e.g., for use as local coordinate systems for teeth), 3D point clouds (or other 3D representations described herein) describing teeth (or appliance components or fixture model components), or others described herein. In particular, implementing the techniques disclosed herein improves computing systems specifically adapted to generate 3D oral care representations for use in oral care appliance generation. For example, aspects of the present disclosure improve the performance of computing systems with 3D representations of a patient's dentition by reducing the consumption of computing resources. In particular, aspects of the present disclosure reduce computing resource consumption by encoding inputs (e.g., 3D representations of teeth, mesh element labels, tooth transformations, coordinate systems, 3D representations of appliance components, 3D representations of fixture model components, or other 3D oral care representations described herein) into reduced-dimensionality latent representations (e.g., using the encoder portion of a reconstruction autoencoder). This method can potentially reduce thousands of mesh elements (each of which may be described by several real numbers) to latent vectors of potentially hundreds of real numbers, so that computing resources are not unnecessarily wasted by processing excessive amounts of data. In addition, encoding input data into latent representations does not reduce the overall predictive accuracy of the computing system (in fact, it may actually improve predictions because the input provided to the ML model after encoding is a more accurate (or better) representation of the patient's dentition). For example, the latent representation may remove noise or other artifacts that are inconsequential (and that may reduce the accuracy of the predictive model) and provide an extracted set of values that accurately describe the 3D oral care representation being encoded. That is, aspects of the present disclosure provide for more efficient allocation of computing resources in a manner that improves the accuracy of the underlying system.
[0247] Furthermore, aspects of the present disclosure may need to be performed in a time-constrained manner, such as when an oral care appliance must be generated for a patient immediately after an intraoral scan (e.g., while the patient is waiting in a clinician's office). Thus, aspects of the present disclosure are necessarily rooted in the underlying computer technology of generating 3D oral care representations using denoising diffusion models, and cannot be performed by humans even with the aid of pen and paper. For example, implementations of the present disclosure should be able to 1) store thousands or millions of mesh elements of a patient's dentition (or appliance components or fixtur...
Claims
1. 1. A method for generating a data structure for an oral care treatment, comprising: receiving, by one or more computer processors, one or more attributes describing an intended output from the trained machine learning model; generating, by the one or more computer processors, one or more noisy representations of the intended output; denoising, by the one or more computer processors, the one or more noisy representations of the intended output using a trained machine learning model; and generating, by the one or more computer processors, one or more denoised representations of the intended output; and automatically defining, by the one or more computer processors, one or more aspects of one or more digital oral care treatments using the one or more generated denoised representations.
2. accessing a training dataset; and generating, by the one or more computer processors, a refined training data set by successively modifying one or more representations of the training data set; training, by the one or more computer processors, an untrained machine learning model using the refined training dataset; and The method of claim 1 , further comprising: outputting the trained machine learning model based on the training.
3. The method of claim 2 , wherein the modifying comprises adding noise to aspects of the one or more representations of the training data set.
4. 3. The method of claim 2, further comprising encoding, by the one or more computer processors, one or more aspects of one or more representations of the training data set into a latent form prior to generating the one or more noisy representations.
5. The method of claim 1 , wherein the one or more attributes include at least one of a real-valued value, a categorical value, or a natural language text value.
6. The method of claim 4 , wherein the one or more attributes include at least one of an oral care metric or an oral care parameter.
7. The method of claim 1 , wherein the trained machine learning model comprises at least one neural network.
8. The method of claim 7 , wherein the at least one neural network comprises at least one of a U-Net, a Transformer, or an Autoencoder.
9. The method of claim 1 , wherein the method receives one or more 3D representations of the patient's dentition including at least one tooth.
10. The method of claim 9 , wherein the one or more denoised representations include at least one 3D oral care representation.
11. The method of claim 10 , wherein the one or more 3D representations of the patient's dentition are encoded in a latent format.
12. The method of claim 10 , wherein the one or more denoised representations include at least one label of a mesh element.
13. The method of claim 12 , further comprising segmenting one or more 3D representations of oral-care data using the at least one label.
14. The method of claim 13 , wherein the one or more 3D oral care representations include at least one representation of the patient's dentition.
15. The method of claim 12 , further comprising modifying one or more 3D representations of oral-care data using the at least one label.
16. The method of claim 15 , wherein the one or more 3D oral care representations include at least one representation of the patient's dentition.
17. The method of claim 10 , wherein the one or more denoised representations include at least one transformation for placing at least one tooth of the patient's dentition in a setup pose.
18. The method of claim 17 , wherein the setup corresponds to a final setup.
19. The method of claim 10 , wherein the 3D oral care representation includes one or more of a fixture model component, a dental restoration design, an appliance component, a fixture model component, and an arch form.
20. The method of claim 10 , wherein the one or more denoised representations include one or more coordinate systems.