Denoising diffusion model for digital oral care

The 3D oral care representation of dental or orthodontic treatments is directly handled in 3D space through the denoising diffusion probability model, which solves the problems of insufficient data accuracy and waste of resources in the prior art, and achieves more efficient tooth shape and structure generation.

CN120322210APending Publication Date: 2025-07-15SOLVENTUM INTELLECTUAL PROPERTIES CO
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202380085976.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-23
Filing Date
2023-12-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Prior art In dental or orthodontic treatment, the accuracy of 3D oral care representations and data accuracy are insufficient, especially when converting between representations of different dimensions, it is easy to introduce noise, resulting in increased computing resources and storage requirements.

Method used

Using denoising diffusion probability models, such as U-Net neural networks, generate and modify 3D oral care representations through training and using denoising diffusion technology, directly processed in 3D space, avoiding the intermediate steps of reconversion from 2D to 3D, and directly generating or modifying dental care products such as dental restoration designs.

Benefits of technology

Improve the data accuracy of 3D oral care representation, reduces computing resources and storage requirements, achieves more accurate tooth shape and structure generation, and reduces computing costs and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120322210A_ABST
    Figure CN120322210A_ABST
Patent Text Reader

Abstract

Systems and techniques are disclosed for automatically generating a data structure for oral care processing using a trained machine learning model. The method involves receiving one or more attributes describing an expected output from a trained machine learning model. One or more noisy representations of the expected output are generated by the one or more computer processors. These noisy representations are then de-noised using a trained machine learning model, which is also performed by the computer processor. A de-noised representation of the expected output is generated, providing more accurate and more reliable data. Based on the generated de-noised representation, one or more aspects of the digital oral care process are automatically defined by the computer processor. These systems and techniques can create an integrated data structure for oral care processing, thereby enhancing processing planning and decision making processes in the field of oral healthcare.
Need to check novelty before this filing date? Find Prior Art

Description

Related Literature

[0001] The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosure of each of the PCT applications with publication numbers WO2022123402A1, WO2021245480A1, and WO2020026117A1 is incorporated herein by reference. The entire disclosure of each of the following provisional U.S. patent applications is incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 366,514; 63 / 264,914; and 63 / 432,627. Technical Field

[0002] The present disclosure relates to the configuration and training of denoising diffusion models (e.g., neural networks trained for that purpose) to improve the accuracy and data precision of 3D oral care representations to be used in dental or orthodontic procedures. Summary of the Invention

[0003] The present disclosure describes systems and techniques for training and using one or more denoising diffusion probability models (e.g., neural networks such as U-Net or autoencoders) to generate 3D oral care representations. Denoising diffusion-based techniques are described for placing oral care articles relative to one or more 3D representations of teeth. Oral care articles to be placed may include: 3D representations of teeth (e.g., for generating orthodontic setups), prosthodontic appliance components, oral care hardware (e.g., tongue brackets, lip brackets, orthodontic attachments, bite ramps, etc.), and the like. Additionally, denoising diffusion-based techniques are described for generating the geometry and / or structure of oral care articles at least partially based on one or more 3D representations of teeth. Oral care articles that may be generated include: prosthodontic tooth designs, crowns, veneers, dental arch forms, clear tray aligner (CTA) trim lines, and appliance components (e.g., generation components such as those used in forming prosthodontic appliances). Among neural networks that can be trained to perform denoising diffusion, U-Net is an example of a model that can improve data accuracy. The denoising diffusion model can be trained to automatically generate (or modify) 3D oral care representations, such as 3D point clouds (or other 3D representations described herein). In some cases, the denoising diffusion model can be trained to generate 3D polylines (e.g., for dental arch forms and CTA trim lines) or sets of control points (e.g., control points through which a spline can be fit for a dental arch form). The denoising diffusion model can be trained to predict a dental arch form (e.g., taking dental arch and tooth data as input). The dental arch form can be in the form of a surface, 3D mesh, 3D polyline, or a set of control points (e.g., for defining a spline). In some cases, such a dental arch form can be given as input to a setup prediction machine learning model (such as a setup prediction neural network). The denoising diffusion model can be trained to label aspects of the 3D representation (or generate one or more object masks identifying one or more objects within the input 3D representation) for use in segmenting or mesh cleaning of 3D oral care representations.

[0004] The techniques of the present disclosure include methods for generating data structures for oral care processes. Such methods may receive, as input, one or more attributes that describe the expected output from a trained machine learning model. Such methods may generate one or more noisy representations of the expected output and then denoise the one or more noisy representations. The output of the method may include one or more generated denoised representations that can be used to define one or more aspects of a digital oral care process. A training dataset may be generated or refined by successively modifying one or more 3D oral care representations. An untrained (or partially trained) machine learning model may be trained using at least in part the refined training dataset. The resulting trained machine learning model may be deployed for clinical treatment of a patient. By the process of generating the training dataset, one or more representations of the training dataset may be modified, such as by adding noise. In some embodiments, one or more aspects of one or more representations of the training dataset may be encoded into a latent form prior to generating the one or more noisy representations. The one or more attributes may include at least one of real-valued, categorical, or natural language text values. The one or more attributes may include at least one of an oral care metric or an oral care parameter. The trained machine learning model may include at least one neural network. The at least one neural network may include at least one encoder-decoder structure. The encoder-decoder structure may include at least one encoder or at least one decoder. Non-limiting examples of encoder-decoder structures include U-Net, transformers, pyramid encoder-decoders, or autoencoders, etc. A denoising diffusion technique may receive one or more 3D representations of a patient dentition (e.g., which may include at least one tooth). The one or more 3D representations of the patient dentition may be encoded into a latent form. The one or more denoised representations (also referred to as generated 3D oral care representations) may include at least one of the following: one or more 3D oral care representations as defined herein, at least one label for a mesh element, at least one transformation for placing at least one tooth of the patient dentition into a set pose, an appliance component, a trim line, an arch form, a coordinate system, or a tooth restoration design, etc. Labels on the mesh elements may be used to segment one or more 3D oral care representations (e.g., segment at least one representation of the patient dentition). Labels may be used to clean up the 3D oral care representations (e.g., the labels may be used to modify one or more 3D oral care representations, such as a representation of the patient dentition). The transformation predicted using the techniques described herein may place the tooth into the set pose. In some embodiments, the set may correspond to a final set. In some embodiments, the set may correspond to an intermediate stage. In some cases, the appliance component (e.g., placed or generated using a denoising diffusion model) may be used to generate one or more dental restoration appliances.

[0005] In some cases, the denoising diffusion techniques described herein can be practiced in combination. For example, denoising diffusion mesh cleaning can be performed on an arch from an intraoral scanner, and then denoising diffusion segmentation can be performed on the cleaned 3D representation of the patient's dentition. In some cases, one or more segmented teeth can be identified for dental restoration processing and subsequently provided to a denoising diffusion model for geometry generation (e.g., Geometric Generation Diffusion Model GGDM) to generate a dental restoration tooth design. In some embodiments, the segmented teeth of the arch can be provided to a diffusion setup model for setup prediction. Other combinations of the techniques described herein should also be considered within the scope of the present disclosure.

[0006] The techniques of the present disclosure can use a denoising diffusion probability model to generate data structures for oral care processing. The method can receive oral care parameters (or oral care attributes) that describe the expected output from a trained denoising diffusion model. The method can perform a forward pass to generate a series of increasingly noisy versions of each training example. In deployment, a denoising ML model (e.g., U-Net) can be trained to iteratively denoise an initially noisy data structure. After multiple denoising iterations, the data structure can evolve into a 3D oral care representation suitable for use in treating a patient. In other words, one or more denoised representations are generated, and based on these representations, one or more aspects of a digital oral care process are automatically defined (e.g., generating a restoration design for a tooth, or generating an orthodontic setup, and other examples).

[0007] The method can access a training dataset of 3D oral care representations and generate a refined training dataset by modifying the representations of the training dataset (e.g., iteratively adding noise). The refined training dataset can be used to train an untrained machine learning model, and a trained machine learning model is based on the training output.

[0008] Modifying the representation of the training dataset involves adding noise to aspects of the representation. In addition, the method encodes aspects of the representation into a latent form before generating the noisy representation.

[0009] The attributes that describe the expected output can include real-valued, categorical, or natural language text values. Additionally, the attributes can include oral care metrics or oral care parameters.

[0010] The trained machine learning model can include a neural network, such as a U-Net, a transformer, or an autoencoder, etc. The method can receive a 3D representation of a patient's dentition, and the denoised representation can include a 3D oral care representation. These representations can be encoded into a latent form (e.g., when using stable diffusion).

[0011] The method can also be trained to label grid elements and apply those grid element labels to 3D representations of oral care data to segment those 3D representations based on the labels. The 3D oral care representations can include representations of a patient's dentition, as well as other representations described herein.

[0012] In addition, the denoised representation can include a transformation (e.g., a 4×4 transformation matrix, etc.) that can be used to place the patient's teeth into a set pose (e.g., for a final set or an intermediate stage). In some embodiments, the transformation can place other 3D representations of oral care data into a pose suitable for oral care processing. Other 3D representations of oral care data include fixture model components, dental restoration designs, appliance components, fixture model components, or dental arch morphology, etc. Additionally, the denoised representation can include a coordinate system or other 3D oral care representations described herein.

[0013] Generally speaking, the disclosed method utilizes a denoising diffusion probability model to generate data structures for use in oral care processing, thereby allowing aspects of a digital oral care process to be automatically defined based on the trained model and the denoised representation.

[0014] In other words, the method of the present disclosure can generate data structures for use in digital oral care processes. The method can define one or more attributes (or oral care variables) that describe the expected output from a trained machine learning model. The method can generate a training data set by generating one or more noisy representations of the expected output (e.g., a series of increasingly noisy representations) and / or training a machine learning model using one or more noisy representations of the expected output. The method can use the trained model (e.g., a hierarchical neural network feature extraction model such as a U-Net) to iteratively denoise a representation (e.g., a representation initially having random, Gaussian, or other default values or distributions). After one or more denoising iterations, the method can generate (or modify) one or more denoised representations of the expected output. The method can use the denoised representations to generate one or more oral care appliances or otherwise use the denoised representations for one or more aspects of one or more digital oral care processes. A refined data set can be generated by continuously modifying one or more representations of the training data set. An initially untrained machine learning model can be trained using the refined training data set and output for use in oral care appliance generation. The refined training data set can be generated by adding noise to aspects of one or more representations of the training data set. In some embodiments, one or more aspects of one or more representations of the training data set can be encoded into a latent form before generating one or more noisy representations. In other words, each example of the training data set can first undergo latent encoding (e.g., using an encoder) before generating the refined data set (of increasingly noisy examples). One or more oral care attributes (or oral care variables) can describe the shape, structure, layout, or other characteristics of the expected output. Oral care variables can include oral care metrics and / or oral care parameters. Other oral care attributes can include doctor preferences. One or more oral care attributes (or oral care variables) can include at least one of real-valued, categorical values, or natural language text values. The trained machine learning model can include at least one neural network, such as a hierarchical neural network feature extraction model (e.g., a U-Net), an autoencoder, or a transformer (e.g., a 3D SWIN transformer). One or more 3D representations of a patient's dentition (e.g., one or more teeth of a patient) can be provided to the method. The method can denoise a 3D oral care representation (e.g., a 3D oral care representation to undergo modification). In some embodiments, a 3D oral care representation (e.g., a 3D representation of a patient's dentition) can be encoded into a latent form. In some cases, one or more denoised representations include at least one or more labels of mesh elements (e.g., used in mesh segmentation or mesh cleaning). The method can use one or more labels to segment a 3D representation (e.g., a mesh, point cloud, voxelized representation, etc.). One or more 3D representations of oral care data can be segmented or undergo mesh cleaning (e.g., a 3D representation of a patient's dentition).Mesh cleaning may involve using one or more tags to modify one or more 3D oral care representations (e.g., removing mesh elements with a specific tag value or modifying mesh elements with a specific tag value). In some embodiments, the method may denoise one or more representations of the transformation (e.g., a transformation that places teeth into a set pose, such as a final set or an intermediate stage, places appliance components for appliance generation, or places fixture model components for fixture model generation). In some embodiments, the method may denoise one or more 3D representations of a tooth restoration design to modify the shape and / or structure of the pre-restoration tooth and make the tooth suitable for use in a digital oral care process (e.g., for use in the generation of a dental restoration appliance). In some embodiments, the method may denoise one or more 3D representations of appliance components and make one or more appliance components suitable for use in a digital oral care process (e.g., for use in the generation of a dental restoration appliance). In some embodiments, the method may denoise one or more 3D representations of fixture model components and make one or more fixture model components suitable for use in a digital oral care process (e.g., for use in the generation of a digital fixture model). The digital fixture model may be 3D printed to obtain a physical fixture model. An orthodontic appliance tray may be thermoformed using the physical fixture model and used in a patient's orthodontic treatment. In some embodiments, the method may denoise one or more 3D representations of a transparent tray aligner trim line and make one or more transparent tray aligner trim lines suitable for use in a digital oral care process (e.g., for trimming an appliance tray from a physical fixture model). In some embodiments, the method may denoise one or more 3D representations of an arch form and make one or more arch forms suitable for use in a digital oral care process (e.g., for use in the generation of an oral care appliance). In some embodiments, the method may denoise one or more 3D representations of a coordinate system and make one or more coordinate systems suitable for use in a digital oral care process (e.g., for use in the generation of an oral care appliance). BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Illustrates a method of using a denoising diffusion probability model to generate (or modify) a 3D oral care representation.

[0016] Figure 2 Illustrates a method of training a denoising diffusion probability model to generate (or modify) a 3D oral care representation.

[0017] Figure 3 Illustrates a method of using a denoising diffusion probability model to generate an orthodontic setup transformation.

[0018] Figure 4A method of performing mesh element labeling (e.g., for segmentation or mesh cleaning) of 3D representations using a hierarchical neural network feature extraction module of a denoising diffusion probability model according to the present disclosure is shown.

[0019] Figure 5 A U-Net structure that can be used to extract hierarchical features from 3D representations is shown.

[0020] Figure 6 A pyramid encoder-decoder structure that can be used to extract hierarchical features from 3D representations is shown. Detailed Description

[0021] Diffusion models can be applied to 2D image generation or 3D representation generation, etc. In some cases, such implementations can obtain inputs from natural language text (e.g., natural language text describing expected results, real-valued, categorical values, reference images, or reference 3D representations). The techniques described herein extend diffusion models to the digital oral care field. The techniques described herein use a diffusion model to generate 3D oral care representations (e.g., as 3D point clouds, 3D meshes, 3D voxelized representations, 3D surfaces, etc.) based on input variables (e.g., which can include natural language text, integer variables, real-valued variables, categorical variables, etc.). Such variables can include as input one or more attributes describing the expected output from a trained machine learning model. In some cases, the techniques described herein can take as input a representation of a patient's dentition (e.g., 3D point cloud, 2D view, or 2D digital photograph), which will be used as a guide for generating artifacts to be used in digital oral care. In some cases, the techniques described herein can take as input a representation of an appliance, appliance component, trim line, dental arch form, or other 3D oral care representation (e.g., in the form of a 3D point cloud, 2D view, or 2D digital photograph), which will be modified or used as a template or reference for generating artifacts to be used in digital oral care. In some cases, these inputs (e.g., teeth, gums, appliances, appliance components, dental arch forms, trim lines for trays used for refining manufacturing such as printed or thermoformed trays, segmentation masks on mesh elements, or other 3D oral care representations) can be encoded by an autoencoder or by other encoder-decoder neural networks into a latent (e.g., information-rich dimension-reduced) form. In some cases, trim lines for cutting thermoformed trays from a jig model (e.g., a jig model used in the manufacture of indirect bonding trays or orthodontic appliances) can be generated. In some cases, the techniques of the present disclosure can achieve improved data accuracy by using such latent representations. These techniques are valuable for 3M's digital oral care platform, particularly for customized smile design in the form of tooth restoration design generation, appliance component generation, appliance trim line generation, setup prediction, etc.

[0022] A denoising diffusion model (e.g., as shown in Figure 2 ) may involve a forward pass 206 of input data 200 (e.g., a 3D point cloud of teeth, a transformation, or another 3D oral care representation described herein), which may generate a Markov chain that may continuously introduce more noise (e.g., Gaussian noise) to the input data 200 in step 204. The forward pass 206 may generate training data. The Markov chain may generate a set of increasingly noisy training data examples. Input oral care variables (or attributes) 202 may affect the functionality of the denoising diffusion model, such that the denoising diffusion model generates an output that is customized for a clinician (e.g., enables customization of the output generated by the denoising diffusion model). Oral care variables may include oral care parameters, oral care metrics, etc. In some embodiments, such as with Stable Diffusion, an optional latent encoding module 214 may encode the 3D oral care representation 200 into a latent form. The latent encoding module may be trained to encode data into a latent form with reduced dimensionality. Examples of data that may be encoded include transformations (e.g., tooth transformations, etc.), point clouds or meshes describing teeth, a set of mesh element labels. Such data may be encoded into latent vectors or latent capsules. Similarly, in some embodiments, an optional latent encoding module 212 may encode one or more oral care variables 202 into a latent form.

[0023] A Markov chain can generate a series of increasingly noisy patterns of the input data 200, which can then be used to at least partially train the denoising ML module 210 (e.g., which can be used as part of the backpropagation 208). The denoising ML model trained to be used in the backpropagation can include one or more neural networks 210 (e.g., U-Net, VAE, 3D SWIN Transformer, pyramid encoder-decoder, etc.), and can be trained to denoise highly noisy patterns of the data (e.g., starting from a completely randomized pattern of the input data structure and denoising the data structure until the data structure presents aspects suitable for generating the output). The backpropagation 208 can be used during model deployment (e.g., after the deployment of the trained denoising ML model 210). For example, the backpropagation 208 can start from a point cloud (or other raw data structure, such as a transformation or mesh element labels) with a Gaussian distribution, and when the backpropagation is continuously applied, the denoising ML model 210 gradually shapes the point cloud (or other data structure) into a suitable example of the target 3D oral care representation (e.g., a tooth design suitable for use in creating a dental restoration appliance). In some specific implementations, a loss function (e.g., cross-entropy or MSE, etc.) can be calculated to quantify the difference between the generated 3D oral care representation and the corresponding ground truth (or reference) 3D oral care representation. In some specific implementations, the loss function can be used to at least partially train the denoising ML model 210. In some specific implementations, such as when generating an orthodontic setting or a dental restoration design, an oral care metric can be calculated on the generated 3D oral care representation 220. The oral care metric can be used at least in part to evaluate the quality of the health using the generated 3D oral care representation (e.g., for measuring whether the generated 3D oral care representation meets the specifications of the oral care variable 202 and is ready to be used in generating an oral care appliance).

[0024] In other words, in some cases, the forward pass 206 of the denoising diffusion model can generate a training dataset of increasingly noisy examples of the input data. Noise can be introduced to corrupt the input data (e.g., an image, 3D point cloud, transformation, or a latent representation of one or more of these inputs, etc.), and those noisy examples can be used at least in part to train the denoising diffusion machine learning model 210 to invert the noise introduction process (e.g., during model deployment). The backpropagation 208 can be trained to reconstruct the original input data by removing the noise from the noisy examples of the input data. After training the denoising diffusion model (e.g., U-Net), the model is capable of generating new 3D oral care representations by passing the noisy data examples (e.g., randomly generated noisy examples) through the denoising diffusion process (also known as the reverse process).

[0025] The techniques of the present disclosure may require a training dataset of hundreds or thousands of queued patient cases to ensure that the neural network can encode the distribution of patient cases that may be encountered in clinical processing. Queued patient cases may include a set of crown meshes, a set of root meshes, or a data file (e.g., a JSON file) including case attributes. Typical examples of queued patient cases may include up to 32 crown meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), multiple gingiva meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), or one or more JSON files (each of which may include tens of thousands of values (e.g., objects, arrays, strings, real values, boolean values, or null values)).

[0026] In some specific implementations, the denoising ML model 210 of the backpropagation 208 may modify an existing 3D oral care representation (e.g., modify a pre-restoration tooth design). The existing 3D oral care representation (e.g., an example of the optional immediate patient case data 216) may be provided as an input to the backpropagation and then undergo a series of denoising steps by the denoising ML model 210 until the modification is complete (e.g., as measured by an oral care metric, a loss function, or an expiration of a threshold number of iterations). The optional immediate patient case data 216 may include data related to the dentition of the patient. In some specific implementations, the immediate patient case data may be introduced to customize the functionality of the denoising diffusion ML model for the patient's anatomy. In some specific implementations, the immediate patient case data may be encoded 218 into a latent or embedded form. In some specific implementations, the immediate patient case data may be provided to the denoising diffusion ML model 210. In other specific implementations, the immediate data 216 may include appliance components or fixture model components that need to be modified.

[0027] In some specific implementations, a data structure undergoing iterative denoising by the denoising ML model 210 may be initialized by: at least partially according to a random process (e.g., using random noise or a random configuration); or by aspects of the immediate patient case data 216; or by a combination of both. The immediate patient case data 216 may include the 3D oral care representations described herein, including tooth transformations (e.g., setting the malocclusion transformation of one or more teeth during prediction), tooth meshes to which the transformation has been applied, tooth meshes to which the transformation has not been applied, one or more 3D representations of the pre-restoration tooth design, the anterior dental arch mesh before segmentation, the anterior dental arch mesh after cleaning, one or more mesh element labels (e.g., for segmentation or mesh cleaning), one or more segmented teeth used in coordinate system prediction, one or more coordinate systems (e.g., each of which may be described by a transformation), one or more 3D representations of appliance components (e.g., parting surfaces, etc.), one or more fixture model components (e.g., digital pontics or interdental marginal strips, etc.).

[0028] The denoising diffusion ML model 210 can be trained to generate one or more generated 3D oral care representations 220 (e.g., as defined herein) for patient treatment. The generated 3D oral care representations 220 (also referred to as denoised representations) can be generated by the denoising diffusion ML model 210 in one or more denoising iterations. The generated 3D oral care representations 220 can include a setup transformation of one or more teeth, a transformation of the placement of one or more appliance components (e.g., for generating a dental restoration appliance), one or more 3D representations of a post-restoration tooth design, one or more generated (or modified) appliance components (e.g., a parting surface or gingival band used in generating a dental restoration appliance), one or more generated (or modified) fixture model components (e.g., one or more trimming lines, etc.), a segmented dental arch mesh, a cleaned dental arch mesh, one or more mesh element labels used in segmentation or mesh cleaning, one or more object masks used in segmentation or mesh cleaning (e.g., a mask to be applied to a mesh element), one or more coordinate axes of one or more predicted coordinate systems, one or more dental arch morphologies, or other 3D oral care representations as described herein.

[0029] The machine learning techniques described herein can receive various input data as described herein, including dental meshes of one or both dental arches of a patient. For example, dental data, appliance components, fixture model components can be provided in the form of 3D representations such as meshes, point clouds, or voxelized geometries. This data can be preprocessed, for example, by arranging the constituent mesh elements into a list and calculating optional mesh element feature vectors for each mesh element. Such vectors can impart valuable information about the shape and / or structure of the oral care mesh to the machine learning models described herein. Additional inputs can be received as inputs to the machine learning models described herein, such as one or more oral care metrics. Oral care metrics can be used to measure one or more physical aspects of the oral care mesh (e.g., physical relationships within or between teeth). In some cases, oral care metrics can be calculated for either or both of the malocclusion oral care mesh examples and / or the ground truth oral care mesh examples and then used in the training of the machine learning models described herein. Metric values can be received as inputs to the machine learning models described herein as a way to train the model or those models to encode the distribution of such metrics over several examples in the training dataset. During training, the network can then receive the metric values as inputs to help train the network to link the metric values of the input to the physical aspects of the ground truth oral care mesh used in the loss calculation. This loss calculation can quantify the difference between the prediction and the ground truth example (e.g., between the predicted oral care mesh and the ground truth oral care mesh). By providing the metric values to the network, the neural network techniques of the present disclosure can learn to train the neural network to encode the distribution of a given metric through the process of loss calculation and subsequent backpropagation. In deployment, one or more oral care variables can be defined to specify one or more aspects of an expected 3D oral care representation (e.g., 3D mesh, polyline, 3D point cloud, or voxelized geometry) that will be generated using the machine learning models (e.g., denoising diffusion models) trained for this purpose described herein. In some embodiments, oral care variables can be defined to specify one or more aspects of a custom vector, matrix, or any other numerical representation (e.g., to describe a 3D oral care representation such as control points for a spline, dental arch morphology, transformation or coordinate system for placing teeth or appliance components relative to another 3D oral care representation) that will be generated using the machine learning models (e.g., denoising diffusion models) trained for this purpose described herein. The custom vector, matrix, or other numerical representation can describe a 3D oral care representation that conforms to the expected outcome of patient treatment. Oral care variables can include oral care metrics or oral care parameters, etc. Oral care variables can specify one or more aspects of an oral care protocol, such as orthodontic setup prediction or prosthetic design generation, etc.In some specific implementations, one or more oral care parameters corresponding to respective oral care metrics may be defined. Oral care variables may be provided as input to the machine learning models described herein and may be provided as instructions to the module to generate an oral care mesh with specified customization, place an oral care mesh for generating an orthodontic setting (or appliance), segment an oral care mesh, or clean an oral care mesh, to name just a few examples. This interaction between oral care metrics and oral care parameters may also apply to the training and deployment of other predictive models in oral care.

[0030] In some specific implementations, the predictive models of the present disclosure may obtain more accurate results by incorporating one or more of the following inputs: arch form information V, interproximal reduction (IPR) information U, tooth size information P, diastema information Q, latent capsule representation T of an oral care mesh, latent vector representation A of an oral care mesh, protocol parameter K (which may describe the clinician's expected treatment of the patient), doctor preference L (which may describe typical protocol parameters selected by the doctor), flag M regarding tooth status (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name / dental symbol R, oral care metric S (including at least one of an oral care metric and a prosthetic design metric).

[0031] In some cases, the systems of the present disclosure may be deployed in a clinical environment (such as a dental or orthodontic clinic) for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians). Such systems deployed in a clinical environment may enable clinicians to process oral care data (such as dental scans) in a clinical environment or in some cases in a "chairside" environment (when the patient is in the clinical environment). A non-limiting list of examples of techniques may include: segmentation (e.g., such as diffusion segmentation), mesh cleaning, coordinate system prediction, CTA trim line generation, prosthetic design generation (e.g., such as using a diffusion model for 3D point cloud generation), appliance component generation or placement or assembly (e.g., such as using a diffusion model for 3D point cloud generation), generation of other oral care meshes, verification of oral care meshes, setting prediction (e.g., such as diffusion setting), removal of hardware from a tooth mesh, placement of hardware on a tooth, estimation of missing values, clustering of oral care data, oral care mesh classification, setting comparison, metric calculation, or metric visualization. In some cases, the execution of these techniques may enable patient data to be processed, analyzed, and used by clinicians in appliance creation before the patient leaves the clinical environment (which may facilitate treatment planning as feedback may be received from the patient during the treatment planning process).

[0032] Some existing technologies can use diffusion models for dental crown restoration. Such technologies can convert the 3D point cloud of a tooth into a 2D depth map, use the diffusion model to process the depth map, and then attempt to convert the resulting modified depth map back into the 3D point cloud of the tooth. By attempting to convert between 3D and 2D representations in this way, errors may be introduced. In other words, noise may be introduced when data is converted between representations of different dimensions (e.g., from 2D to 3D, or from 3D to 2D). The denoising diffusion probability model of the present disclosure can be trained to directly generate 3D dental restoration designs (or other 3D representations described herein, such as fixture model components or appliance components) without the intermediate step of depth map generation. For example, the denoising ML model 210 of the present disclosure can be directly trained on a series of increasingly noisy point clouds 204 of dental restoration designs (e.g., pre-restoration or post-restoration tooth designs). Then, the fully trained denoising ML model 210 can be used to directly denoise a noisy (or randomly initialized) 3D point cloud (or other 3D representation) to generate a dental restoration design without first converting the 3D dental data into a 2D depth map or any other 2D representation and then converting back to a 3D representation. The techniques of the present disclosure not only generate more accurate tooth shapes and / or structures (e.g., less noisy tooth surfaces), but the techniques of the present disclosure are also capable of reducing resource usage. The techniques of the present disclosure avoid the computational costs (e.g., computational loops) and storage (e.g., computer memory) requirements for converting between representations of different dimensions (e.g., as regarding depth map methods). In other words, the techniques of the present disclosure have a lower resource footprint than techniques that convert between representations of different dimensions as a preprocessing step.

[0033] The systems of the present disclosure can automate operations in digital orthodontics (e.g., setup prediction, hardware placement, setup comparison), digital dentistry (e.g., prosthetic design generation), or combinations thereof. Some techniques can be applied to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleaning, coordinate system prediction, oral care mesh validation, estimation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows, or denoising diffusion models), metric visualization, appliance component placement, or appliance component generation, etc. In some cases, the systems of the present disclosure can enable clinicians or technicians to process oral care data (such as scanned dental arches). In addition to segmentation, mesh cleaning, coordinate system prediction, or validation operations, the systems of the present disclosure can also implement orthodontic treatment planning, which may involve setup prediction as at least one operation. The systems of the present disclosure can also implement prosthetic design generation (e.g., using a denoising diffusion model), where one or more restored tooth designs are generated and processed during the creation of an oral care appliance. The systems of the present disclosure can implement either or both of orthodontic or dental treatment planning, or can automate steps in the generation of either or both of orthodontic or dental appliances. Some appliances can implement both dental and orthodontic treatments, while other appliances can implement one or the other.

[0034] The present disclosure relates to digital oral care encompassing the fields of digital dentistry and digital orthodontics. The present disclosure generally describes methods for processing three-dimensional (3D) representations of oral care data. It should be understood that there are various types of 3D representations without loss of generality. One type of 3D representation is 3D geometry. The 3D representation can include, be one or more of the following, or be a part thereof: 3D polygon meshes, 3D point clouds (e.g., such as derived from 3D meshes), 3D voxelized representations (e.g., a collection of voxels for sparse processing), or 3D representations described by mathematical equations. Although the term "mesh" is frequently used throughout the present disclosure, in some specific implementations, this term should be understood to be interchangeable with other types of 3D representations. The 3D representation can describe the 3D geometry and / or elements of the 3D structure of an object.

[0035] The dental arches S1, S2, S3, and S4 all include exactly the same dental meshes, which are transformed differently according to the following description. The first dental arch S1 includes a set of dental meshes that are arranged (e.g., using transformation) in their positions in the oral cavity, where the teeth are in malposed positions and orientations. The second dental arch S2 includes the same set of dental meshes from S1 that are arranged (e.g., using transformation) in their positions in the oral cavity, where the teeth are in a reference true set position and orientation. The third dental arch S3 includes the same meshes as S1 and S2 that are arranged (e.g., using transformation) in their positions in the oral cavity, where the teeth are in a predicted final set pose (e.g., as predicted by one or more techniques of the present disclosure). S4 is the counterpart of S3, where the teeth are in a pose corresponding to one of several intermediate stages of orthodontic treatment with a clear tray appliance.

[0036] It should be understood that, without loss of generality, the techniques of the present disclosure applied to the final set are also applicable to intermediate gradings in orthodontic treatment, specifically geometric deep learning (GDL) settings, reinforcement learning (RL) settings, variational autoencoder (VAE) settings, capsule settings, multi-layer perceptron (MLP) settings, diffusion settings, pose transfer (PT) settings, similarity settings, force-directed graph (FDG) settings, transformer settings, setting comparison, or setting classification. The metric visualization aspect of the present disclosure can also be configured to visualize data from both the final set and intermediate stages. The MLP setting, VAE setting, and capsule setting each fall within the scope of the autoencoder setting. Some specific implementations of the MLP setting can fall within the scope of the transformer setting. A representation setting refers to any one of the MLP setting, VAE setting, capsule setting, and any other setting prediction machine learning model that uses an autoencoder to create a representation of at least one tooth.

[0037] Each of the setting prediction techniques of the present disclosure is applicable to the manufacture of clear tray appliances and / or indirectly bonded trays. The setting prediction techniques can also be applicable to other products that also involve the final tooth pose. The pose can include position (or location) and rotation (or orientation).

[0038] A 3D mesh is a data structure that can describe the geometry and / or shape of an object related to oral care, which includes but is not limited to teeth, hardware elements, or the gingival tissue of a patient. The 3D mesh can include one or more mesh elements, such as vertices, edges, faces, and combinations thereof. In some specific implementations, the mesh elements can include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features can be calculated for these mesh elements and provided to the prediction model of the present disclosure, and the prediction model of the present disclosure provides a technical advantage of improving data accuracy in the form of a more accurate prediction of the model output.

[0039] The dentition of a patient may include one or more 3D representations of the patient's teeth (e.g., and / or associated transducers), gums, and / or other oral anatomical structures. In some embodiments, an orthodontic metric (OM) may quantify the relative position and / or orientation of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. In some embodiments, a restorative design metric (RDM) may quantify at least one aspect of the structure and / or shape of a 3D representation of a tooth. In some embodiments, an orthodontic landmark (OL) may locate one or more points or other regions of interest structures on a 3D representation of a tooth. In some embodiments, an OL may be used in the generation of orthodontic or prosthetic appliances, such as clear tray aligners or prosthetic restoration appliances. In some embodiments, a mesh element may include at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth represented by a 3D mesh, the mesh elements may at least include: vertices, edges, faces, and voxels. In some embodiments, mesh element features may quantify some aspects of the 3D representation that are proximate to or associated with one or more mesh elements, as described elsewhere in the present disclosure. In some embodiments, an orthodontic procedure parameter (OPP) may specify at least one value that defines at least one aspect of a patient's planned orthodontic treatment (e.g., specifying desired target attributes of a final setting in a final setting prediction). In some embodiments, an orthodontist preference (ODP) may specify at least one typical value of an OPP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. In some embodiments, a restorative design parameter (RDP) may specify at least one value that defines at least one aspect of a patient's planned prosthetic restoration treatment (e.g., specifying desired target attributes of a tooth to be treated with a prosthetic restoration appliance). In some embodiments, a doctor restorative design preference (DRDP) may specify at least one typical value of an RDP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. The 3D oral care representation may include but is not limited to: 1) a set of mesh element labels that may be applied to 3D mesh elements of a tooth / gum / hardware / appliance mesh (or point cloud) during mesh segmentation or mesh cleaning; 2) one or more 3D representations of teeth / gums / hardware / appliances whose shapes have been modified (e.g., trimmed, deformed, or filled) during mesh segmentation or mesh cleaning; 3) one or more coordinate systems (e.g., describing one, two, three, or more coordinate axes) for a single tooth or a group of teeth (such as a full dental arch, e.g., the LDE coordinate system); 4) 3D representations of one or more teeth whose shapes have been modified or otherwise made suitable for use in prosthetic restoration; 5) 3D representations of one or more prosthetic restoration appliance components;6) One or more transformations to be applied to one or more of the following: placement of prosthetic appliance library components relative to one or more teeth, teeth to be placed for an orthodontic setting (final setting or intermediate stage), hardware elements to be placed relative to one or more teeth, etc.; 7) Orthodontic settings; 8) 3D representations of hardware elements (such as face bows, tongue bows, orthodontic attachments, buttons, hooks, occlusal ramps, etc.) placed relative to one or more teeth, etc.; 8) 3D representations of bonding pads for hardware elements (which can be generated for a specific tooth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting the tooth through a Boolean operation); 9) 3D representations of clear tray appliances (CTAs); 10) The position or shape of CTA trim lines (e.g., described as a grid or polyline); 11) Arch forms (e.g., described as 3D polylines or 3D grids or surfaces) that describe the contour or layout of the dental arch, which can follow the incisal edges of one or more teeth, which can follow the facials of one or more teeth, which in some specific implementations can correspond to malocclusion dental arches and in other specific implementations correspond to final setting dental arches (the effect of malocclusion on the shape of the arch form can be reduced by smoothing or averaging the shape of the arch form), and which can be described by one or more control points and / or splines; 12) 3D representations of jig models (e.g., depictions of teeth and gums used in thermoformed clear tray appliances, or depictions of teeth / gums / hardware used in thermoformed indirect bonding trays); 13) One or more latent space vectors (or latent capsules) generated by the 3D encoder stage of a 3D autoencoder (e.g., a variational autoencoder trained for tooth reconstruction) that has been trained on the reconstruction of an oral care mesh; 14) One or more oral care metrics for one or more teeth (e.g., such as orthodontic metrics or prosthetic design generation metrics); 15) One or more landmarks (e.g., 3D points) that describe the shape and / or geometric properties of one or more teeth, other dentition structures, or hardware structures (e.g., to be used for orthodontic setting creation or prosthetic appliance component generation or placement); 16) 3D representations created by scanning (e.g., optical scanning, CT scanning, or MRI scanning) 3D printed parts (e.g., scanned jig models) corresponding to one or more teeth / gums / hardware / appliances; 17) 3D printed appliances (optionally including local thickness, reinforcement rib geometry, tab positioning, etc.); 18) 3D representations of a patient's dentition captured by a clinician or healthcare practitioner at the chairside (e.g., in an environment where the 3D representation is verified at the chairside, before the patient leaves the clinic, such that errors can be detected and re-scanning can be performed as needed); 19) Prosthetic tooth designs (e.g., for veneers, crowns, bridges, or prosthetic appliances); 20) 3D representations of one or more teeth used in digital oral care processes; 21) Other 3D printed parts belonging to oral care procedures or other fields; 22) IPR cutting surfaces;23) One or more orthodontic setting transformations associated with one or more IPR cutting surfaces; 24) (Digital) pontic design that can fill at least a portion of the space between teeth to create space for erupting teeth in an orthodontic setting and then emerge from the gum; or 25) Components of a jig model (e.g., including jig model components such as interdental bands, sealants, bite locks, bite ramps, interdental reinforcements, gingival ridges, torque points, power ridges, pontics, or pits, etc.).;

[0040] The techniques of the present disclosure can be advantageously combined. For example, a setting comparison tool can be used to compare the output of a GDL setting model with reference ground truth data, compare the output of an RL setting model with reference ground truth data, compare the output of a VAE setting model with reference ground truth data, and compare the output of an MLP setting model with reference ground truth data. By comparing each of these setting prediction models with reference ground truth data, it can be determined which model achieves the best performance on a certain dataset or within a given problem domain. Additionally, a metric visualization tool can enable a global view of the final settings and intermediate stages produced by one or more of the setting prediction models, with the advantage of being able to select the best setting prediction model. Furthermore, the metric visualization tool enables the calculation of metrics with a global scope within a set of intermediate stages. In some specific implementations, these global metrics can be consumed as inputs to a neural network for predicting settings (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, etc.). The global metrics can also be provided to FDG settings. In some specific implementations, local metrics from the present disclosure (i.e., local metrics are metrics that can be calculated for one stage or setting of a process rather than within several stages or settings) can be consumed by the neural networks herein for predicting settings, with the advantage of improving the prediction results. In some specific implementations, the metrics described in the present disclosure can be visualized using a metric visualization tool.

[0041] VAE and MAE models for mesh element tagging and mesh filling can be advantageously combined with a setup prediction neural network for mesh cleaning before or during the prediction process. In some specific implementations, the VAE for mesh element tagging can be used to tag mesh elements for further processing, such as metric calculation, removal, or modification. In some cases, such tagged mesh elements can be provided as input to the setup prediction neural network to inform the neural network of important mesh features, properties, or geometries, with the advantage of improving the performance of the resulting setup prediction model. In some specific implementations, mesh filling can make the geometry of the teeth closer to complete, enabling the setup prediction model to function better (i.e., improving the correctness of the prediction due to the better-formed geometry). In some cases, a neural network for classifying setups (i.e., a setup classifier) can help the setup prediction neural network function because the setup classifier tells the setup prediction neural network when a predicted setup can be accepted for use and can be provided to methods for generating orthodontic trays. Setup classifiers (e.g., GDL setups, RL setups, VAE setups, capsule setups, MLP setups, diffusion setups, PT setups, similarity setups, and FDG setups, etc.) can help generate the final setup and also help generate intermediate stages. Additionally, the setup classifier neural network can be combined with a metric visualization tool. In other specific implementations, the setup classification neural network can be combined with a setup comparison tool (e.g., the setup comparison tool can output an indication of how a setup generated, in part, by the setup classifier compares to a setup generated by another setup prediction method). In some specific implementations, the VAE for mesh element tagging can identify one or more mesh elements used in metric calculation. The resulting metric output can be visualized by a metric visualization tool.

[0042] In some examples, the setup classifier neural network can assist with the setup prediction techniques described in U.S. Patent Application No. US20210259808A1, the entire contents of which are incorporated herein by reference, or PCT Application Publication No. WO2021245480A1, the entire contents of which are incorporated herein by reference, or PCT Application No. PCT / IB2022 / 057373, the entire contents of which are incorporated herein by reference. The setup classifier will help one or more of those techniques know when the predicted final setup is closest to being correct. In some cases, the setup classifier neural network can output an indication of how far a given setup is from the final setup (i.e., a progress indicator).

[0043] In some specific implementations, the latent space embedding vectors from the reconstructed VAE can be cascaded with the inputs of the setup prediction neural network described in WO2021245480A1. The latent space vectors can also be combined as inputs into other setup prediction models: GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup, etc. The advantage is to endow the neural network with reconstruction characteristics (e.g., the latent vector dimension of the tooth mesh), thereby improving the generated setup prediction.

[0044] In some examples, the various setup prediction neural networks of the present disclosure can work together to generate the setups required for orthodontic treatment. For example, the GDL setup model can generate the final setup, and the RL setup model can use the final setup as an input to generate a series of intermediate stage setups. Alternatively, the VAE setup model (or MLP setup model) can create the final setup, which can be used by the RL setup model to generate a series of intermediate stage setups. In some specific implementations, the setup prediction can be generated by one setup prediction neural network and then used as an input for another setup prediction neural network for further improvement and adjustment. In some specific implementations, such improvements can be performed in an iterative manner.

[0045] In some specific implementations, a setup verification model may be involved in this iterative setup prediction loop, such as the model disclosed in U.S. Provisional Application No. US63 / 366495. First, setups can be generated (e.g., using models trained for setup prediction, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, etc.), and then the setups are verified. If the setup passes the verification, the setup can be output for use. If the setup fails the verification, the setup can be sent back to one or more of the setup prediction models for correction, improvement, and / or adjustment. In some cases, the setup verification model can output an indication of what is wrong with the setup, so that the setup generation model can be improved in the next iteration. This process iterates until completion.

[0046] Generally, in some embodiments, two or more of the following techniques of the present disclosure may be combined during orthodontic and / or dental treatment: GDL setting, setting classification, reinforcement learning (RL) setting, setting comparison, autoencoder setting (VAE setting or capsule setting), VAE grid element labeling, masked autoencoder (MAE) grid filling, multi-layer perceptron (MLP) setting, metric visualization, estimation of missing oral care parameter values, tooth classification using latent vectors, FDG setting, pose transfer setting, prosthetic design metric calculation, neural network techniques for dental restoration and / or orthodontics (e.g., generation or modification of 3D oral care representations using transformers), landmark-based (LB) setting, diffusion setting, estimation of tooth movement protocols, capsule autoencoder segmentation, diffusion segmentation, similarity setting, validation of oral care representations (e.g., using autoencoders), coordinate system prediction, prosthetic design generation or geometry generation (or modification) using denoising diffusion models.

[0047] Oral care parameters may include one or more values specifying orthodontic protocol parameters or prosthetic design parameters (RDPs), as described herein. Oral care parameters may define one or more expected aspects of a 3D oral care representation and may be provided to an ML model to facilitate the ML model in generating an output that can be used to generate an oral care appliance suitable for treating a patient. Other types of values include doctor preferences and prosthetic design preferences, as described herein. Doctor preferences and prosthetic design preferences may define typical treatment choices or practices of a particular clinician. Prosthetic design preferences are subjective for a particular clinician and thus different from prosthetic design parameters. In some embodiments, doctor preferences or prosthetic design preferences may be calculated by unsupervised means such as clustering, so that typical values used by the clinician in patient treatment can be determined. Those typical values may be stored in a data store and invoked to be provided to an automated ML model as default values (e.g., default values that can be modified before executing the model).

[0048] For example, when faced with a similar diagnostic or treatment scenario, one clinician may prefer one value of a restoration design parameter (RDP), while another clinician may prefer a different value of that RDP. An example of such an RDP is a dental restoration style. In some embodiments, protocol parameters and / or clinician preferences can be provided to a setup prediction model for orthodontic treatment for the purpose of improving the customization of the resulting orthodontic appliance. In some embodiments, restoration design parameters and clinician restoration preferences can be used to design the tooth geometry used in creating a dental restoration appliance for the purpose of improving the customization of that appliance. In addition to oral care parameters, clinician preferences, and clinician restoration preferences, some embodiments of the ML prediction models of the present disclosure in orthodontic treatment can also take as input a setup (e.g., the arrangement of teeth). In some such embodiments, the ML prediction models of the present disclosure can take as input the final setup (i.e., the final arrangement of teeth), such as in the case of a prediction model that is trained to generate intermediate stage predictions. For simplicity, these preferences are referred to as clinician restoration preferences, but are intended to be used in a non-limiting sense. Specifically, it should be understood that these preferences can be specified by any practitioner or other appropriate healthcare professional and are not intended to be limited to clinician preferences per se (i.e., preferences from a person with an MD or equivalent degree).

[0049] An oral care professional or clinician (such as a dentist or orthodontist) can specify information regarding a patient's treatment in the form of a patient-specific set of protocol parameters. In some cases, the oral care professional can specify a set of general preferences (also referred to as clinician preferences) for a large number of cases to be used as default values during the process of specifying that set of protocol parameters. In some embodiments, oral care parameters can be incorporated into the techniques described in the present disclosure, such as one or more of GDL setup, VAE setup, RL setup, setup comparison, setup classification, VAE grid element tagging, MAE grid filling, validation using an autoencoder, estimation of missing protocol parameter values, metric visualization, or FDG setup. One or more of these models can take as input one or more protocol parameter vectors K and / or one or more clinician preference vectors L. In some embodiments, one or more of these models can introduce one or more protocol parameter vectors K and / or one or more clinician preference vectors L into a hidden layer of a neural network. In some embodiments, one or more of these models can introduce either or both of K and L into a mathematical calculation, such as a force calculation, for the purpose of improving that calculation and the final customization of the resulting appliance to the patient.

[0050] Some embodiments of neural networks for prediction settings such as GDL settings, VAE settings, or RL settings may incorporate information from oral care professionals (aka doctors). This information may affect the arrangement of teeth in the final setting, bringing the position and orientation of the teeth into compliance with the specifications set by the doctor and within tolerances. In some embodiments of the GDL setting model, oral care parameters may be fed directly as separate inputs along with the mesh data to the generator network. In some embodiments of the GDL setting, oral care parameters may be incorporated into the feature vectors calculated for each mesh element before the mesh elements are provided to the generator for processing. Some embodiments of setting prediction models (e.g., VAE settings or diffusion settings) may incorporate oral care parameters into the setting prediction. In some embodiments, the protocol parameter K and / or doctor preference information L may be concatenated with the latent space vector C. Doctor preferences (e.g., in an orthodontic context) and / or doctor restoration preferences may be indicated in a processed form, or they may be based on characteristics in the treatment plan, such as final setting characteristics (e.g., the amount of bite correction or midline correction in the planned final setting), intermediate staging characteristics (e.g., treatment duration, tooth movement plan, or overcorrection strategy), or outcomes (e.g., the number of corrections / improvements).

[0051] Orthodontic protocol parameters may specify one or more of the following (possible values are shown in {}). Non-limiting classification values for some example OPPs are described below. In some embodiments, real values may be specified for one or more of these OPPs. For example, the overbite OPP may specify the desired amount of overbite in the setting (e.g., in millimeters) and may be received as an input to the setting prediction model to provide information to the setting prediction model about the desired amount of overbite in the setting. Some embodiments may specify numerical values for the overjet OPP or other OPPs. In some embodiments, one or more OPPs corresponding to one or more orthodontic metrics (OMs) may be defined. In some cases, numerical values may be specified for such OPPs for the purpose of controlling the output of the setting prediction model. Teeth to be moved: {Anterior teeth only, Anterior teeth and canines, Full dental arch} Tooth movement restrictions: For each tooth, indicate whether the tooth is {Not moving, Missing, To be extracted, Native / Erupted, Debridement} Overbite: {Resulting overbite after showing alignment, Maintain initial overbite, Calibrate bite opening, Correct deep bite} Overjet: {Resulting overjet after showing alignment, Maintain initial overjet, Improve resulting overjet} Anterior / Posterior (AP) relationship Retention: {Right, Left, Both} Improve canine relationship only: {Right, Left, Both} Increase the canine and / or molar relationship by 4 mm: {right, left, both} Correction of Class I (canine and molar): {right, left, both} Interdigitation (if present) Anterior teeth: {do not correct, correct, not applicable} Posterior teeth: {do not correct, correct, not applicable} Correction of Class I (canine and molar): {right, left, both} Correction with posterior IPR: {yes, no} Class II / III correction simulation (elastic required): {yes, no} Sequential distal movement (elastic recommended): {yes, no} Is an incision for the elastic included?: {yes, no} Preferred incision for the elastic: {use button incisions on molars and hooks on canines, use only button incisions, use only hooks} Stage to start cutting the elastic: [integer] Levelling of the upper anterior teeth: {lateral 0.5 mm shorter than central, level incisal edge, level gingival edge, as indicated} Spacing: {close all spaces, leave specific spaces open} Preferred midline position: {set the upper midline to the ideal value, match the upper and lower to each other} Resolve upper crowding by expansion: {primarily, as needed, none} Resolve upper crowding by proclination: {primarily, as needed, none} Resolve upper crowding by IPR - anterior teeth: {primarily, as needed, none} Resolve upper crowding by IPR - right posterior teeth: {primarily, as needed, none} Resolve upper crowding by IPR - left posterior teeth: {primarily, as needed, none} Resolve lower crowding by expansion: {primarily, as needed, none} Resolve lower crowding by proclination: {primarily, as needed, none} Resolve lower crowding by IPR - anterior teeth: {primarily, as needed, none} Resolve lower crowding by IPR - right posterior teeth: {primarily, as needed, none} Resolve lower crowding by IPR - left posterior teeth: {primarily, as needed, none} Trim the arch form: {patient's natural arch form, as indicated} [The doctor can specify the arch form - select from a set of options or a custom design]

[0052] Other orthodontic protocol parameters may be defined, such as those that can be used to place standardized brackets at a specified occlusal height on teeth. In some specific implementations, one or more orthodontic protocol parameters may be defined to specify at least one of the second- and third-order rotation angles (i.e., angle and torque, respectively) to be applied to teeth, which can achieve a target setup arrangement where, for example, the crown landmarks are within a threshold distance of a common occlusal plane. In some specific implementations, one or more orthodontic protocol parameters may be defined to specify a position in global coordinates where at least one landmark (e.g., centroid) of the crown (or root) will be placed in the setup arrangement of the teeth. Generally speaking, oral care parameters corresponding to oral care metrics may be defined. For example, orthodontic protocol parameters corresponding to orthodontic metrics may be defined (e.g., to specify the amount of a particular metric expected to appear in a predicted setup at the input of a setup prediction model).

[0053] Doctor preferences may differ from orthodontic protocol parameters because doctor preferences are related to the oral care provider and may include the mean, mode, median, minimum, or maximum (or some other statistical value) of past setups associated with the oral care provider's treatment decisions for past orthodontic cases. On the other hand, protocol parameters may be related to a specific patient and describe the requirements of treating that specific patient. Doctor preferences may be related to the doctor and the doctor's past treatment practices, while protocol parameters may be related to treating a specific patient. Doctor preferences (or "treatment preferences") may specify one or more of the following (some non-limiting possible values are shown in {}). Other possible values can be found elsewhere in this disclosure.

[0054] Doctor preferences may specify one or more of the following. Deep bite case (amount of bite correction) - final overbite: [real value, in millimeters, e.g., 0.5mm] Option - intrusion of maxillary anterior teeth: {yes, no} Option - include vertical overcorrection of lower canines: {yes, no} Midline correction in the planned final setup: {maintain initial midline, improve midline with IPR, as indicated} Deep bite case - speed reversal curve: {yes, no} Anterior bite opening case - final overbite: [real value, in millimeters, e.g., 2mm] Is arch expansion a priority for your case?: {yes, no} If so, specify the acceptable expansion per quadrant in mm. When upper molars are expanded, apply buccal root torque: {yes, no} Is IPR of the first Tx design acceptable?: {yes, no} Maximum IPR per contact: Upper anterior teeth: [Specified in mm] Lower anterior teeth: [Expressed in mm] Upper and lower anterior teeth: [Specified in mm] Is asymmetric IPR acceptable?: {Yes, No} Final tooth position (overcorrection strategy): {Ideal, Overcorrection} Root movement: {Move roots as needed to achieve treatment goals, Limit posterior root movement, Limit all root movement} Final occlusal contact: {Balance all contacts as much as possible, No occlusal contact on upper incisors, End with heavy posterior contact, Other} For Class corrections, is asymmetric AP shift acceptable?: {Yes, No, Other} Treatment duration: [Phase count] Tooth movement plan: {Plan_A, Plan_B, Plan_C}

[0055] The prior art has attempted to move teeth towards arch form V after a setup prediction has already been presented by other workflow components, which introduces error into the resulting setup and degrades performance due to the purpose of the workflow component that performs the setup prediction. The present disclosure provides several improvements over these prior arts by enabling arch form information to be directly introduced into a setup prediction neural network as an input to the neural network, where the technical improvement is to provide a setup prediction that more accurately meets the orthodontic treatment needs of the patient (thereby improving data accuracy). The arch form information V can be provided as an input to any of the GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup prediction neural networks. In some specific implementations, the arch form information V can be directly provided to one or more internal neural network layers in one or more of those setup applications.

[0056] Additional protocol parameters can include a text description of the patient's medical condition and the expected treatment. Such text descriptions can be analyzed via natural language processing operations, including tokenization, stop word removal, stemming, n-gram formation, text data vectorization, bag-of-words analysis, term frequency-inverse document frequency (TF-IDF) analysis, sentiment analysis, naive Bayes classification, and / or logistic regression classification. The output of such analysis techniques can be used as an input to one or more of the neural networks of the present disclosure, with the advantage of customizing and improving the prediction output (e.g., predicted setup or predicted grid geometry).

[0057] These additional orthodontic parameters and doctor preferences can also be incorporated into the neural networks of the present disclosure, which has the advantages related to improving the customization of those neural networks and enabling those neural networks to predict output examples that better match the treatment needs of individual patients in terms of data accuracy and efficiency.

[0058] In some embodiments, a dataset for training one or more of the neural network models of the present disclosure may be conditionally filtered according to one or more of the orthodontic protocol parameters described in this section. In some cases, patient cases exhibiting outliers for one or more of these protocol parameters may be omitted from the dataset (alternatively used to form the dataset) for training one or more of the neural networks of the present disclosure.

[0059] During training, one or more protocol parameters and / or doctor preferences may be provided to the neural network. In this way, the neural network may be conditioned on one or more protocol parameters and / or doctor preferences. Examples of such neural networks include conditional generative adversarial networks (cGANs) and / or conditional variational autoencoders (cVAEs), either of which may be used for various neural network-based applications of the present disclosure.

[0060] In some cases, a shape-based input may be provided to the neural network for setting predictions. In other cases, non-shape-based inputs may be used, such as tooth names or nomenclature, as it relates to dental notation. In some embodiments, a vector R of flags may be provided to the neural network, where a value of "1" indicates the presence of a tooth and a value of "0" indicates the absence of the tooth in the patient case (although other values are possible). The vector R may include one-hot vectors, where each element in the vector corresponds to a tooth type, name, or nomenclature. Identification information about the tooth (e.g., the name of the tooth) may be provided to the prediction neural network of the present disclosure, which advantageously enables the neural network to be trained to handle different teeth in a tooth-specific manner. For example, the setting prediction model may learn to make setting transformation predictions for specific tooth names (e.g., the upper right central incisor or the lower left canine, etc.). In the case of a mesh cleaning autoencoder (for labeling mesh elements or for filling in missing mesh data), the autoencoder may be trained in this way to provide specialized processing to a tooth based on the tooth's nomenclature. In the case of a setting classification neural network, a list of tooth names present in the patient's dental arch may better enable the neural network to output an accurate determination of the setting classification, as tooth nomenclature is a valuable input for training such a neural network. For example, tooth naming / names may be defined according to the Universal Numbering System, the Palmer Notation Method, or the FDI World Dental Federation notation (ISO 3950).

[0061] In one example, in the presence of all teeth except (at most four) wisdom teeth, the vector R can be defined as an optional input to the set prediction neural network of the present disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, UL1, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7.

[0062] In some cases, the position of the cusp tips can be provided to the neural network for set prediction. In other cases, one or more vectors S of orthodontic metrics described elsewhere in the present disclosure can be provided to the neural network for set prediction. The advantage is an improved ability of the network to be trained to understand the state of the malocclusion setting and thus be able to predict a more accurate final or intermediate stage setting.

[0063] In some embodiments, the neural network can take as input one or more indications of interproximal reduction (IPR) U, which can indicate the amount of enamel to be removed from the teeth (from the mesial or from the distal) during the course of orthodontic treatment. In some embodiments, the IPR information (e.g., the amount of IPR to be performed on one or more teeth, measured in millimeters, or one or more binary markers indicating whether IPR is to be performed on each tooth identified by the marker) can be concatenated with the latent vector A generated by a VAE or a latent capsule T autoencoder. The vector and / or capsule resulting from such concatenation can be provided to one or more of the neural networks of the present disclosure, which has the technical improvement or additional advantage of enabling the prediction neural network to take into account IPR. IPR is particularly relevant to a set prediction method that can determine the position and pose of teeth at the end of treatment or during one or more stages during treatment. It is important to consider the amount of enamel to be removed before the predicted tooth movement.

[0064] In some embodiments, one or more protocol parameters K and / or doctor preference vectors L can be introduced into the set prediction model. In some embodiments, one or more optional vectors or values include: tooth position N (e.g., XYZ coordinates in local or global coordinates of the tooth), tooth orientation O (e.g., pose, such as in a transformation matrix or quaternion, Euler angles or other forms described herein), tooth size P (e.g., length, width, height, perimeter, radius, diagonal measurement, volume, any size can be normalized relative to one or more other teeth), distance Q between adjacent teeth. In some cases, these "tooth sizes P" can be used to describe the expected size of teeth for dental restoration design generation.

[0065] In some specific implementations, the tooth dimension P (e.g., length, width, height, or perimeter) can be measured in a plane, such as a plane intersecting the centroid of the tooth, or a plane intersecting a central point that is the midpoint between the centroid of the tooth and the most incisal extent or the most gingival extent. The tooth height dimension can be measured as the distance from the gingiva to the incisal edge. The tooth width dimension can be measured as the distance from the mesial extent to the distal extent of the tooth. In some specific implementations, the roundness or circularity of the tooth cross-section can be measured and included in the vector P. The roundness or circularity can be defined as the ratio of the radii of the inscribed circle and the circumscribed circle.

[0066] The distance Q between adjacent teeth can be implemented in different ways (and calculated using different distance definitions, such as Euclidean or geodesic). In some specific implementations, the distance Q1 can be measured as the average distance between the mesh elements of two adjacent teeth. In some specific implementations, the distance Q2 can be measured as the distance between the centers or centroids of two adjacent teeth. In some specific implementations, the distance Q3 can be measured between the closest mesh elements between two adjacent teeth. In some specific implementations, the distance Q4 can be measured between the tooth tips of two adjacent teeth. In some specific implementations, teeth can be considered adjacent within an arch. In some specific implementations, teeth can also be considered adjacent between opposing arches. In some specific implementations, any one of Q1, Q2, Q3, and Q4 can be divided by a term to normalize the resulting value of Q. In some specific implementations, the normalization term can involve one or more of the following: the volume of the tooth, the count of mesh elements in the tooth, the surface area of the tooth, the cross-sectional area of the tooth (e.g., as projected onto the XY plane), or some other term related to the tooth dimension.

[0067] Other information regarding the patient's dentition or treatment needs (or related parameters) can be concatenated with other input vectors to one or more of an MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and / or any neural network model listed elsewhere in this disclosure.

[0068] Vector M may include markers applied to one or more teeth. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is pinned. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is fixed. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is a pontic. Other and additional markers are possible for teeth, such as combinations of fixed, pinned, and pontic markers. A marker set to a value indicating that a tooth should be fixed is a signal that the tooth should not move during processing and is sent to the network. In some embodiments, the neural network loss function can be designed to penalize any movement in the indicated teeth (and in some cases, may be severely penalized). A marker indicating that a tooth is a pontic notifies the network to maintain the diastema, although movement of the gap is allowed. In some cases, M may include a marker indicating tooth loss. In some embodiments, the presence of one or more fixed teeth in the dental arch can assist in setting the prediction because the one or more fixed teeth can provide an anchor for the posture of other teeth in the dental arch (i.e., can provide a fixed reference for the posture transformation of one or more other teeth in the dental arch). In some embodiments, one or more teeth can be intentionally fixed in order to provide an anchor to which other teeth can be positioned. In some embodiments, a 3D representation (such as a mesh) corresponding to the gingiva can be introduced to provide a reference point according to which the teeth can move.

[0069] Without loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U, and V described elsewhere in this disclosure can also be provided as inputs to one or more of the prediction models of this disclosure or fed into their intermediate layers. Specifically, these optional vectors can be provided to the MLP setup, GDL setup, RL setup, VAE setup, capsule setup, and / or diffusion setup, with the advantage of enabling the corresponding models to generate setups that better meet the orthodontic treatment needs of the patient. In some embodiments, such inputs can be provided, for example, by concatenating with one or more latent vectors A that are also provided to one or more of the prediction models of this disclosure. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent capsules T that are also provided to one or more of the prediction models of this disclosure.

[0070] In some embodiments, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into a neural network (such as an MLP or a transformer) in the hidden layer of the network. In some cases, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into the internal processing of the encoder structure.

[0071] In some specific implementations, a prediction model (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, PT settings, similarity settings, and diffusion settings) can take as input one or more latent vectors A corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a prediction model (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, and diffusion settings) can take as input one or more latent capsules T corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a prediction method can take both A and T as inputs.

[0072] A variety of loss calculation techniques generally apply to the techniques of the present disclosure (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, setting classification, tooth classification, VAE mesh element labeling, MAE mesh filling, and estimation of protocol parameters).

[0073] These losses include L1 loss, L2 loss, mean squared error (MSE) loss, cross-entropy loss, etc. The losses can be calculated and used to train neural networks such as multi-layer perceptrons (MLPs), U-Net architectures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer architectures, etc. For example, in the learning of sequences, some specific implementations can use triplet loss or contrastive loss.

[0074] Losses can also be used to train encoder architectures and decoder architectures. KL divergence loss can be used at least in part to train one or more neural networks of the present disclosure, such as a grid reconstruction autoencoder or a generator in GDL settings, which has the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior can enable the reconstruction autoencoder to produce better reconstructions (e.g., when modifying the latent vector representation and using the decoder to reconstruct the modified latent vector, the resulting reconstruction is more likely to be a valid instance of the input representation). There are other techniques for calculating losses that can be described elsewhere in the present disclosure. Such losses can be based on quantifying the difference between two or more 3D representations.

[0075] MSE loss calculation can involve the calculation of the average squared distance between two sets, vectors, or data sets. MSE can generally be minimized. MSE can be applied to regression problems, where the predictions generated by a neural network or other machine learning model can be real numbers. In some specific implementations, the neural network can be equipped with one or more linear activation units on the output to generate MSE predictions. According to the techniques of the present disclosure, mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used.

[0076] In some specific implementations, cross-entropy can be used to quantify the difference between two or more distributions. In some specific implementations, cross-entropy loss can be used to train the neural networks of the present disclosure. In some specific implementations, cross-entropy loss can involve comparing predicted probabilities with ground-truth probabilities. Other names for cross-entropy loss include "log loss", "logistic loss", and "log loss". A small cross-entropy loss can indicate a better (e.g., more accurate) model. Cross-entropy loss can be logarithmic. In some specific implementations, cross-entropy loss can be applied to binary classification problems. In some specific implementations, a neural network can be equipped with a sigmoid activation unit at the output to generate probability predictions. In the case of multi-class classification, cross-entropy can also be used. In this case, in some specific implementations, a neural network trained to make multi-class predictions can be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for each class to be predicted). Other loss calculation techniques that can be applied in the training of the neural networks of the present disclosure include one or more of the following: Huber loss, hinge loss, classification hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss calculation methods are described herein and can be applied to the training of any neural network described in the present disclosure.

[0077] In some specific implementations, one or more neural networks of the present disclosure can be trained, at least in part, by a loss based on at least one of the following: pointwise mesh Euclidean distance (PMD) and Earth Mover's Distance (EMD). Some specific implementations can incorporate Hausdorff distance (HD) calculation into the loss calculation. Calculating the Hausdorff distance between two or more 3D representations (such as 3D meshes) can provide one or more technical improvements because HD not only considers the distance between two meshes, but also the way those meshes are oriented and the relationship between the mesh shapes in those orientations (or positions or poses). Hausdorff distance can improve the comparison of two or more tooth meshes, such as two or more instances of tooth meshes in different poses (e.g., comparison of a predicted setting with a ground-truth setting, which can be performed during the process of calculating the loss value for training a setting prediction neural network).

[0078] The reconstruction loss can compare the predicted output with the ground truth (or reference) output. The systems of the present disclosure can calculate the reconstruction loss as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_loss = 0.5 * L1(all_points_target, all_points_predicted) + 0.5 * MSE(all_points_target, all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the ground truth data (e.g., a ground truth tooth restoration design, or some other ground truth example of a 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the generated or predicted data (e.g., a generated tooth restoration design, or some other generated example of a 3D type of oral care representation). Other specific implementations of the reconstruction loss can additionally (or alternatively) involve an L2 loss, a mean absolute error (MAE) loss, or a Huber loss term.

[0079] The reconstruction error can compare the reconstructed output data (e.g., generated by a reconstruction autoencoder, such as a tooth design that has been generated for generating a dental restoration appliance) with the initial input data (e.g., the data input to the reconstruction autoencoder, such as a pre - restoration tooth). The systems of the present disclosure can calculate the reconstruction error as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_error = 0.5 * L1(all_points_input, all_points_reconstructed) + 0.5 * MSE(all_points_input, all_points_reconstructed). In the above example, all_points_input is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the input data (e.g., a pre - restoration tooth design input to the reconstruction autoencoder, or some other 3D oral care representation input to an ML model). In the above example, all_points_reconstructed is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the reconstructed (or generated) data (e.g., a reconstructed tooth restoration design, or some other example of a generated 3D oral care representation).

[0080] In other words, the reconstruction loss involves calculating the difference between the predicted output and the reference output, while the reconstruction error involves calculating the difference between the reconstructed output and the initial input from which the reconstructed data is derived.

[0081] The techniques of the present disclosure may include operations such as 3D convolution, 3D pooling, 3D transposed convolution, and 3D unpooling. 3D convolution may assist in segmentation processing, for example, when downsampling a 3D mesh. 3D transposed convolution, for example, performs the inverse operation of 3D convolution in a U-Net. 3D pooling may assist in segmentation processing, for example, in a generalized neural network feature map. 3D unpooling, for example, performs the inverse operation of 3D pooling in a U-Net. These operations may be implemented by one or more layers in a predictive or generative neural network described herein. These operations may be applied directly to mesh elements, such as mesh edges or mesh faces. These operations provide a technical improvement over other methods because these operations are invariant to mesh rotation, scaling, and translation changes. Generally speaking, these operations depend on edge (or face) connectivity, so as long as the edge (or face) connectivity is maintained, these operations will not be affected by mesh changes in 3D space. That is, these operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position, or scale of the oral care mesh, which may improve data accuracy. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes and can be used for tasks such as 3D shape classification or mesh element labeling (e.g., for segmentation or mesh cleaning). MeshCNN performs these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.

[0082] In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 2D representations, such as images. In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 3D representations, such as meshes or point clouds. An intraoral scanner may capture 2D images of a patient's dentition from various angles. The intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data depicting the patient's dentition. According to various techniques, an autoencoder (or other neural network described herein) may be trained to operate on either or both of 2D and 3D representations.

[0083] A 2D autoencoder (including a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or latent capsule) using the 2D encoder and then reconstruct a copy of the input 2D image using the 2D decoder. For a handheld mobile application that has been developed for such analysis (e.g., for the analysis of dental anatomy), 2D images may be easily captured using one or more on-board cameras. In other examples, 2D images may be captured using an intraoral scanner configured for such a function. Operations that may be used in implementations of a 2D autoencoder (or other 2D neural network) for 2D image analysis are 2D convolution, 2D pooling, and 2D reconstruction error calculation.

[0084] 2D image convolution may involve the "sliding" of a kernel across a 2D image and the calculation of element-wise multiplications, as well as summing these element-wise multiplications into output pixels. The output pixels generated from each new position of the kernel are saved into an output 2D feature matrix. In some specific implementations, adjacent elements (e.g., pixels) may be in well-defined positions in a straight-line grid (e.g., above, below, left, and right).

[0085] A 2D pooling layer can be used to downsample a feature map and summarize the presence of certain features in that feature map.

[0086] A 2D reconstruction error can be calculated between the pixels of an input image and a reconstructed image. The mapping between the pixels can be well understood (e.g., directly comparing the upper pixel [23,134] of the input image with the pixel [23,134] of the reconstructed image, assuming the two images have the same dimensions).

[0087] One of the advantages provided by the 2D autoencoder-based techniques of the present disclosure is the ease of capturing 2D image data with a handheld device. In some cases where an external data source provides data for analysis, there may be instances where only 2D image data is available. When only 2D image data is available, it is necessary to use a 2D autoencoder for analysis.

[0088] Modern mobile devices (such as commercially available smartphones) may also have the ability to generate 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera that moves around an object to capture multiple images from different views, or both), and this 3D data can be arranged into a 3D representation, such as a 3D mesh, 3D point cloud, and / or 3D voxelized representation in some specific implementations. In some cases, the analysis of the 3D representation of an object may provide a technical improvement over the 2D analysis of the same object. For example, the 3D representation can describe the geometry and / or structure of an object with less ambiguity than a 2D representation (which may include shadows and other artifacts that complicate the depiction of the depth and texture of the object). In some specific implementations, 3D processing can achieve a technical improvement due to the inverse optics problem, which affects 2D representations in some cases. The inverse optics problem refers to the phenomenon that in some cases, the size of an object, the orientation of the object, and the distance between the object and the imaging device may be combined in the 2D image of the object. Any given projection of an object on an imaging sensor can map to an infinite count of {size, orientation, distance} pairings. The 3D representation achieves a technical improvement because it removes the ambiguity introduced by the inverse optics problem.

[0089] Devices configured for a dedicated purpose with 3D scanning, such as 3D intraoral scanners (or CT scanners or MRI scanners), can generate 3D representations of an object (e.g., a patient's dentition) with significantly higher fidelity and precision than a handheld device might have. When such high-fidelity 3D data is available (e.g., in oral care mesh classification or the application of other 3D techniques described herein), the use of 3D autoencoders provides technical improvements (such as increased data precision) to extract the best possible signal from those 3D data (i.e., obtain a signal from the 3D crown mesh used in tooth classification or setup classification).

[0090] A 3D autoencoder (including a 3D encoder and a 3D decoder) can be trained on 3D data to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder, and then reconstruct a copy of the input 3D representation using the 3D decoder. Operations that can be used to implement a 3D autoencoder for analyzing 3D representations (e.g., 3D meshes or 3D point clouds) are 3D convolution, 3D pooling, and 3D reconstruction error calculation.

[0091] For each mesh element, 3D convolution can be performed to aggregate local features from nearby mesh elements. Processing can be performed on top of and in addition to techniques used for 2D convolution to account for different counts and positions of adjacent mesh elements (relative to a particular mesh element). A particular 3D mesh element can have a variable neighbor count, and those neighbors can be absent from their expected positions (unlike pixels in 2D convolution, which can have a fixed adjacent pixel count present in known or expected positions). In some cases, the order of adjacent mesh elements can be relevant to 3D convolution.

[0092] 3D pooling operations can enable the combination of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling can iteratively reduce a 3D mesh to the mesh elements most highly relevant to a given application (e.g., the neural network that has been trained for it). Similar to 3D convolution, 3D pooling can benefit from special processing in addition to that required in 2D convolution to account for different counts and positions of adjacent mesh elements (relative to a particular mesh element). In some cases, the order of adjacent mesh elements may be less relevant to 3D pooling than to 3D convolution.

[0093] The 3D reconstruction error can be calculated using one or more of the techniques described herein, such as calculating the Euclidean distance between corresponding mesh elements, between two meshes. According to aspects of the present disclosure, other techniques are possible. The 3D reconstruction error can generally be calculated on 3D mesh elements rather than 2D pixels of the 2D reconstruction error. The 3D reconstruction error can achieve a technical improvement over the 2D reconstruction error because in some cases, the 3D representation can have less ambiguity (i.e., less ambiguity in form, shape, and / or structure) than the 2D representation. In some specific implementations, due to the complexity of the mapping between the input mesh elements and the reconstructed mesh elements (i.e., the input mesh and the reconstructed mesh may have different mesh element counts, and there may be a less clear mapping between the mesh elements compared to the mapping between pixels in 2D reconstruction), additional processing may be required for 3D reconstruction over and above 2D reconstruction. Technical improvements in 3D reconstruction error calculation include increased data precision.

[0094] A 3D scanner, such as an intraoral scanner, a computed tomography (CT) scanner, an ultrasound scanner, a magnetic resonance imaging (MRI) machine, or a mobile device capable of performing photogrammetry, can be used to generate a 3D representation. The 3D representation can describe the shape and / or structure of an object. The 3D representation can include one or more of a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation. A 3D mesh includes edges, vertices, or faces. Although in some cases these three types of data are related to each other, they are distinct. A vertex is a point in 3D space that defines the boundary of a mesh. These points could alternatively be described as a point cloud without additional information about how the points are connected to each other (as described by edges). An edge is described by two points and can also be referred to as a line segment. A face is described by multiple edges and vertices. For example, in the case of a triangular mesh, a face includes three vertices that are interconnected to form three consecutive edges. Some meshes can include degenerate elements, such as non-manifold mesh elements, which can be removed to benefit subsequent processing. According to aspects of the present disclosure, other mesh preprocessing operations are also possible. 3D meshes are typically formed using triangles, but in other embodiments quadrilaterals, pentagons, or some other n-sided polygon can be used. In some embodiments, such as in the case of performing sparse processing, a 3D mesh can be converted into one or more voxelized geometries (i.e., including voxels). The techniques of the present disclosure operating on a 3D mesh can receive one or more tooth meshes (e.g., arranged in one or more dental arches) as input. Each of these meshes can be preprocessed before being input into a prediction architecture (e.g., including at least one of an encoder, a decoder, a pyramid encoder-decoder, and a U-Net). Such preprocessing can include converting the mesh into a list of mesh elements such as vertices, edges, faces, or into voxels in the case of sparse processing. For one or more selected types of mesh elements (e.g., vertices), a feature vector can be generated. In some examples, a feature vector is generated for each vertex of the mesh. Each feature vector can include a combination of spatial features and / or structural features, as specified in the following table:

[0095] Table 1 discloses non - limiting examples of mesh element features. In some specific implementations, in addition to the spatial or structural mesh element features described in Table 1, color (or other visual cues / identifiers) can also be considered mesh element features. As used herein (e.g., in Table 1), a point is different from a vertex in that a point is part of a 3D point cloud, while a vertex is part of a 3D mesh and can have incident faces or edges. A dihedral angle (which can be expressed in radians or degrees) can be calculated as the angle (e.g., a signed angle) between two connected faces (e.g., two faces connected along an edge). The sign on the dihedral angle can reveal information about the convexity or concavity of the mesh surface. For example, in some specific implementations, a positively signed angle can indicate a convex surface. Additionally, in some specific implementations, a negatively signed angle can indicate a concave surface. To calculate the principal curvatures of a mesh vertex, the directional curvatures of each adjacent vertex around that vertex can first be calculated. These directional curvatures can be sorted in a circular order (e.g., 0 degrees, 49 degrees, 127 degrees, 210 degrees, 305 degrees) near the vertex normal vector and can include a subsampled form of the full curvature tensor. Circular order means sorting by angle around an axis. The sorted directional curvatures can contribute to a system of linear equations that admits a closed - form solution, which can estimate the two principal curvatures and directions, which can characterize the full curvature tensor. Consistent with Table 1, a voxel can also have features calculated as an aggregation of other mesh elements (e.g., vertices, edges, and faces) that either intersect the voxel or, in some specific implementations, are mainly or entirely contained within the voxel. Rotating a mesh may not change the structural features but may change the spatial features. And, as described elsewhere in this disclosure, the term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non - limiting sense. In some specific implementations, in addition to mesh element features, there are alternative ways to describe the geometry of a mesh (such as 3D key points and 3D descriptors). Examples of such 3D key points and 3D descriptors can be found in "TONIONI A et al., 'Learning to detect good 3D keypoints.', Int J Comput.Vis. Vol. 126, pp. 1 - 20, 2018". In some specific implementations, 3D key points and 3D descriptors can describe the extrema (minima or maxima) of the surface of a 3D representation.In some embodiments, one or more mesh element features may be computed at least in part via deep feature synthesis (DFS), such as that described in: J.M. Kanter and K. Veeramachaneni, “Deep feature synthesis: Towards automating data science endeavors”, 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp. 1-10, doi: 10.1109 / DSAA.2015.7344858.

[0096] In some embodiments, a mesh element feature vector may be computed for a 3D representation provided to a latent encoding module 214 or 218 (e.g., representing a generative neural network). For example, when a 3D mesh of a tooth is provided to latent encoding module 214 or 218, a mesh element feature vector may be computed for one or more of the mesh elements of the 3D mesh to improve the accuracy of the resulting latent representation (e.g., the latent representation generated by latent encoding module 214 or 218).

[0097] Neural networks for representation generation based on autoencoders, U-Nets, transformers, 3D SWIN transformers, other types of encoder-decoder architectures, convolutional and / or pooling layers, or other models can benefit from the use of mesh element features. Mesh element features can convey aspects of the surface shape and / or structure of a 3D representation to the neural network models of the present disclosure. Each mesh element feature describes different information about the 3D representation, which may not redundantly exist in other input data provided to the neural network. For example, vertex curvature can quantify aspects of the concavity or convexity of the surface of a 3D representation that the network may not otherwise understand. In other words, mesh element features can provide a processed form of the structure and / or shape of a 3D representation that data that would otherwise not be available to the neural network. This processed information is generally more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, mesh element features have been provided to a representation generation neural network based on a U-Net model and also to a representation generation model based on a variational autoencoder with continuous normalizing flows. Based on the experiments, it was found that systems using a full complement of mesh element features (e.g., "XYZ" coordinate tuples, "normal vectors", "vertex curvature", point pivots, and normal pivots) were at least 3% more accurate than systems that did not use mesh element features. A point pivot describes an "XYZ" coordinate tuple with a local coordinate system (e.g., at the centroid of the corresponding tooth). A normal pivot describes a "normal vector" with a local coordinate system (e.g., at the centroid of the corresponding tooth). Additionally, when using a full complement of mesh element features, training converges faster. In other words, machine learning models trained using a full complement of mesh element features tend to be faster and more accurate (at an earlier epoch) than systems that do not. For an existing system with a historical accuracy rate of 91%, a 3% increase in accuracy reduces the actual error rate by more than 30%.

[0098] Prediction models that can operate on the feature vectors of the above-described features include, but are not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, other denoising diffusion models, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE mesh element tagging, MAE mesh filling, mesh reconstruction autoencoders, validation using autoencoders, mesh segmentation, coordinate system prediction, mesh cleaning, restoration design generation, appliance component generation and / or placement, or dental arch form prediction. Such feature vectors can be presented to the inputs of the prediction models. In some specific implementations, such feature vectors can be presented to one or more internal layers of a neural network that is part of one or more of those prediction models.

[0099] As described herein, tooth movement can specify one or more tooth transformations, which can be encoded in various ways to specify tooth positions and orientations within a setting and applied to a 3D representation of the teeth. For example, according to a particular embodiment, the tooth position can be Cartesian coordinates of a tooth canonical origin position defined in certain semantic contexts. The tooth orientation can be represented as a rotation matrix, a unit quaternion, or other 3D rotation representations such as Euler angles relative to a reference frame (global or local). The dimensions are real-valued 3D spatial extents, and the gaps can be binary existence indicators or real-valued gap sizes between teeth, especially in cases where some teeth are missing. In some embodiments, tooth rotation can be described by a 3×3 matrix (or a matrix of other dimensions). In some embodiments, the tooth position and rotation information can be combined into the same transformation matrix (e.g., combined as a 4×4 matrix), which can reflect homogeneous coordinates. In some cases, an affine space transformation matrix can be used to describe tooth transformations, such as transformations that describe the malocclusion pose of the teeth, the intermediate pose of the teeth, and / or the final setting pose of the teeth. Some embodiments can use relative coordinates, where the setting transformation is predicted relative to a malocclusion coordinate system (i.e., predicting the malocclusion-to-setting transformation rather than directly predicting the setting coordinate system). Other embodiments can use absolute coordinates, where the setting coordinate system is predicted directly for each tooth. In the relative mode, the transformation can be calculated relative to the centroid of each tooth mesh (relative to the global origin), which is referred to as "relative local." Some advantages of using relative local coordinates include eliminating the need for a malocclusion coordinate system (landmark data), which may not be applicable to all patient case datasets. Some advantages of using absolute coordinates include simplifying data preprocessing, since the mesh data is initially represented relative to the global origin. In some embodiments, these details regarding tooth position encoding and tooth orientation encoding can also be applied to one or more of the neural network models of the present disclosure, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, other denoising diffusion models, PT settings, similarity settings, FDG settings, setting classification, setting comparison, VAE mesh element tagging, MAE mesh filling, mesh reconstruction VAE, and validation using autoencoders.

[0100] According to a particular embodiment, the convolutional layers in the various 3D neural networks described herein can perform mesh convolutions using edge data. The use of edge information ensures that the model is insensitive to different input orders of 3D elements. In addition to or separate from using edge data, the convolutional layer can perform mesh convolutions using vertex data. The advantage of using vertex information is that vertices are typically fewer than edges or faces, so vertex-oriented processing can result in lower processing overhead and lower computational costs. In addition to or separate from using edge data or vertex data, the convolutional layer can perform mesh convolutions using face data. Further, in addition to or separate from using edge data, vertex data, or face data, the convolutional layer can perform mesh convolutions using voxel data. The advantage of using voxel information is that, depending on the chosen granularity, there may be far fewer voxels to process compared to vertices, edges, or faces in the mesh. Sparse processing (using voxels) can result in lower processing overhead and lower computational costs (especially in terms of computer memory or RAM usage).

[0101] Neural networks for representation generation based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional layers, and / or pooling layers, or other models can benefit from the use of oral care variables (e.g., oral care metrics or oral care parameters). For example, oral care metrics (e.g., orthodontic metrics or prosthodontic design metrics) can convey aspects of the shape and / or structure of a patient's dentition (e.g., the shape and / or structure of a single tooth, or a particular relationship between two or more teeth) to the neural network models of the present disclosure. Each oral care metric describes different information about the patient's dentition, which may not be redundantly present in other input data provided to the neural network. For example, an "overbite" metric can quantify the overlap between the maxillary central incisors and the mandibular central incisors along the vertical Z-axis, which may not be easily determinable by a traditional neural network in some embodiments. In other words, oral care metrics provide refined information about the patient's dentition that traditional neural networks (e.g., representation generation neural networks) may not be adequately trained or configured to extract as described herein. However, a neural network specifically trained to generate oral care metrics can overcome this shortcoming because, for example, the loss can be computed in a manner that promotes accurate oral care metric prediction. Mesh oral care metrics can provide a processed form of the structure and / or shape of a patient's dentition, data that may not otherwise be available to the neural network. This processed information is typically more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, oral care metrics have been provided to a representation generation neural network based on a U-Net model. Based on the experiments, it was found that systems using oral care metrics (e.g., "overbite," "overjet," and "canine class relationship" metrics) were at least 2.5% more accurate than systems that did not use oral care metrics. Additionally, training converged faster when oral care metrics were used. In other words, machine learning models trained with oral care metrics tend to be faster and more accurate (at earlier epochs) than systems that do not. For an existing system that observed a historical accuracy of 91%, a 2.5% increase in accuracy reduced the actual error rate by nearly 30%.

[0102] Oral care variables may include oral care parameters or oral care metrics. Examples of oral care metrics include orthodontic metrics (OM) and restorative design metrics (RDM). The RDM may describe the shape and / or form of one or more 3D representations of teeth used in dental restorations. One example use case is in creating one or more dental restoration appliances. Another example use case is in creating one or more veneers (such as zirconia veneers). Some RDMs may quantify the shape and / or other characteristics of teeth. Other RDMs may quantify the relationship (e.g., spatial relationship) between two or more teeth. The RDM differs from a restorative design parameter (RDP) in that the restorative design metric defines the current state of a patient's dentition, while the restorative design parameter is used as a specification for a machine learning or other optimization model to generate a desired tooth shape and / or form. The RDM describes the current (e.g., as-is or maloccluded) shape of the teeth. The restorative design parameter specifies the expected appearance of the teeth after a restorative treatment by an oral care provider (such as a dentist or dental technician). For the purpose of dental restorations, a neural network or other machine learning or optimization algorithm may be provided to either or both of the RDM and the RDP. In some embodiments, the RDM may be calculated with respect to a patient's pre-restorative dentition (i.e., primary embodiment). In other embodiments, the RDM may be calculated with respect to a patient's post-restorative dentition. The restorative design may include one or more teeth and may be referred to as a restorative arch. The restorative design generation may involve generating an improved geometry and / or structure of one or more teeth in the restorative arch.

[0103] Aspects of RDM calculation are described below. In some embodiments, the RDM may be measured, for example, by locating landmarks in the teeth (or gums, hardware, and / or other elements of the patient's dentition) and measurements of distances between those landmarks, or otherwise with respect to those landmarks. In some embodiments, one or more neural networks or other machine learning models may be trained to identify or extract one or more RDMs from one or more 3D representations of teeth (or gums, hardware, and / or other elements of the patient's dentition). The techniques of the present disclosure may use the RDM in various ways. For example, in some embodiments, one or more neural networks or other machine learning models may be trained to classify or label one or more settings, arches, dentitions, or other groups of teeth at least in part based on the RDM. Thus, in these examples, the RDM forms part of the training data for training these models.

[0104] Aspects of a dental mesh reconstruction autoencoder that can be used in accordance with the techniques of the present disclosure are described below. An autoencoder for prosthetic design generation is disclosed in U.S. Provisional Application No. US63 / 366514. The autoencoder (e.g., a variational autoencoder or VAE) takes as input a dental mesh (or other 3D representation) that reflects a malocclusion state (i.e., the tooth shape before prosthetics). The encoder component of the autoencoder encodes the dental mesh into a latent form (e.g., a latent vector). To alter the geometry and / or structure of the final reconstructed mesh, a modification can be applied to the latent vector (e.g., based on a mapping of the latent space through prior experiments). In some embodiments, additional vectors can be included with the latent vector (e.g., by concatenation), and the resulting vector concatenation can be reconstructed by the decoder component of the autoencoder into a reconstructed dental mesh that is a replica of the input dental mesh.

[0105] In accordance with aspects of the present disclosure, RDMs and RDPs can also be used as neural network inputs during the execution phase. In some embodiments, to tell the encoder specific information about the input 3D dental representation, one or more RDMs can be concatenated with the input to the encoder. In some embodiments, to provide the decoder component with specific information about the input 3D dental representation, one or more RDMs can be concatenated with the latent vector before reconstruction. Additionally, in some embodiments, to provide the encoder with specific information about the input 3D dental representation, one or more prosthetic design parameters (RDPs) can be concatenated with the input to the encoder component. Similarly, in some embodiments, to provide the decoder with specific information about the input 3D dental representation, one or more prosthetic design parameters (RDPs) can be concatenated with the latent vector before reconstruction.

[0106] In this way, either or both of RDMs and RDPs can be introduced into the functionality of the autoencoder (e.g., a dental reconstruction autoencoder) and used to affect the geometry and / or structure of the reconstructed prosthetic design (i.e., affect the shape of the teeth on the output of the autoencoder). In some embodiments, the variational autoencoder of U.S. Provisional Application No. US63 / 366514 can be replaced by a capsule autoencoder (e.g., instead of encoding the dental mesh into a latent vector, encoding the dental mesh into one or more latent capsules).

[0107] In some embodiments, clustering or other unsupervised techniques may be performed on the RDMs to cluster one or more sets, arches, dentitions, or other groups of teeth based on the restorative characteristics of the teeth. Such clustering may be useful in treatment planning as the clustering provides insights into categories of patients with different treatment needs. This information may be instructive to clinicians as they are aware of the possible treatment options. In some cases, best practices (such as default RDP values) may be identified for patient cases that fall into one or another cluster (e.g., as determined by a similarity metric, such as in k-NN). After classifying a new case into a particular cluster, information about the associated best practices may be provided to the clinician responsible for treating that case. In some cases, such default values may undergo further adjustment or modification.

[0108] Case assignment: Such clustering can be used to gain further insights into the types of patient cases present in a dataset. Analysis of such clustering may reveal that patient treatment cases with certain RDM values (or value ranges) may require less treatment time (or alternatively more treatment time). Cases that require more time to treat (or are otherwise more difficult) can be assigned to experienced or senior technicians for treatment. Cases that take less time to treat can be assigned to newer or less experienced technicians for treatment. This assignment can be further aided by finding a correlation between the RDM values of certain cases and the known treatment durations associated with those cases.

[0109] The following RDMs can be measured and used to create either or both of a dental restorative appliance and a veneer (a veneer is a type of dental restorative appliance) with the aim of making the resulting teeth look natural. Symmetry is generally a preferred aspect. There may be differences between patients based on demographic differences. Generation of the dental restorative appliance may benefit from some or all of the following RDMs. Hue and translucency may particularly relate to the creation of veneers, although some embodiments of the dental restorative appliance may also consider this information.

[0110] Examples of interdental RDMs are enumerated below.

[0111] 1) Symmetry and / or ratio: A measure of the symmetry between one or more teeth on opposite sides of a tooth and one or more other teeth. For example, for a pair of corresponding teeth, a measure of the width of each tooth. In one case, one tooth has a normal width while the other is too narrow. In another case, both teeth have normal widths. The following is a list of properties that can be measured for a tooth and compared to corresponding measurements of one or more corresponding teeth: a) Width - mesial to distal distance; b) Length - gingival to incisal distance; c) Diagonal - distance across the tooth, such as from the mesial gingival corner to the distal incisal corner (this measure is one of many measures that can be used to quantify the shape of a tooth other than length and width). A ratio between a and b can be calculated, such as a / b or b / a. Such a ratio can indicate whether there is spatial symmetry (e.g., by measuring the ratio a / b on the left side and measuring the ratio a / b on the right side and then comparing the left and right ratios). In some specific embodiments, the length, width, and / or ratio may not match when the spatial symmetry is "off". In some specific embodiments, such a ratio can be calculated relative to a standard. Many aesthetic standards are available in the dental literature. Examples include the golden ratio and the cyclic aesthetic dental ratio. In some specific embodiments, spatial symmetry can be measured on a pair of teeth, where one tooth is on the right side of the dental arch and the other tooth is on the left side of the dental arch.

[0112] 2) Proportion of adjacent teeth: Measuring the width proportion of adjacent teeth, such as measured as a projection onto a plane along the dental arch (e.g., a plane located in front of the patient's face). The ideal proportion used in the final restoration design can be, for example, the so-called golden ratio. The golden ratio is relevant to adjacent teeth, such as the central incisor and the lateral incisor. This measure involves the measurement of these proportions as they exist in a pre-restorative malocclusion. For the central incisor, lateral incisor, and canine, the ideal golden ratio on a particular side (left or right) of a particular dental arch (e.g., the upper dental arch) is 1.6, 1, 0.6. If one or more of these proportion values deviate (e.g., in the case of "peg lateral incisors"), the patient may desire a dental restoration procedure to correct the proportion.

[0113] 3) Dental arch difference: A measure of any dimensional difference between the upper dental arch and the lower dental arch, such as related to the width of the teeth, for dental restoration purposes. For example, the techniques of the present disclosure can perform adjacent tooth width proportion measurements in the upper and lower dental arches. In some specific embodiments, Bolton analysis measurements can be performed by measuring the upper width, the lower width, and the ratio between these quantities. In various specific embodiments, the dental arch difference can be described in absolute measurement values (e.g., in mm or other suitable units) or as a proportion or ratio.

[0114] 4) Midline: The measurement of the midline of the maxillary incisors relative to the midline of the mandibular incisors. The techniques of the present disclosure may measure the midline of the maxillary incisors relative to the midline of the nose (if data on the position of the nose is available).

[0115] 5) Proximal contact: The measurement of the size (area, volume, perimeter, etc.) of the proximal contact between adjacent teeth. Ideally, the teeth contact along the mesial / distal surfaces, and the gingiva fills in along the gingival to where the teeth contact. If the gingival tissue fails to fill the space below the proximal contact, a black triangle may form. In some cases, for teeth located closer to the posterior of the dental arch, the size of the proximal contact may gradually become shorter. In an ideal scenario, the proximal contact will be long enough such that there is an appropriately sized incisal embrasure and the gingival tissue fills the area below the contact or gingival to the contact.

[0116] 6) Embrasure: In some embodiments, the techniques of the present disclosure may measure the size (area, volume, perimeter, etc.) of the embrasure, i.e., the gap between teeth at the gingival or incisal margin. In some embodiments, the techniques of the present disclosure may measure the symmetry between the embrasures on opposite sides of the dental arch. The embrasure is at least partially based on the length of the contact between teeth, and / or at least partially based on the shape of the teeth. In some cases, for teeth located closer to the posterior of the dental arch, the size of the embrasure may gradually become longer.

[0117] Examples of intra-tooth RDM are listed below, continuing the numbering of the other RDM listed above.

[0118] 7) Length and / or width: The measurement of the length of a tooth relative to the width of that tooth. This measurement may reveal, for example, that a patient has long central incisors. The width and length are defined as: a) width - mesial to distal distance; b) length - gingival to incisal distance; c) other dimensions of the tooth body - the portion of the tooth between the gingival region and the incisal edge. In some embodiments, either or both of the length and width of a tooth may be measured and compared to the length and / or width of one or more teeth.

[0119] 8) Tooth morphology: Measurements of the major anatomical structures of tooth shape, such as line angles, facial contours, and / or incisal angles and / or embrasures. Frequency and / or dimensions may be measured. In some embodiments, the observed primary tooth shape aspects may be matched to one or more known patterns. The techniques of the present disclosure may measure secondary anatomical structures of tooth shape, such as incisal tubercle grooves. For example, frequency and / or dimensions may be measured. In some embodiments, the observed secondary tooth shape aspects may be matched to one or more known patterns. In some examples, the techniques of the present disclosure may measure tertiary anatomical structures of tooth shape, such as enamel striae or striations. For example, frequency and / or dimensions may be measured. In some embodiments, the observed tertiary tooth shape aspects may be matched to one or more known patterns.

[0120] 9) Hue and / or translucency: Measurements of tooth hue and / or translucency. Tooth hue is typically described by the Vita Classical or 3D Master hue guides. Tooth translucency is described by transmittance or contrast ratio. Tooth hue and translucency may be evaluated (or measured) based on one or more of the following types of data related to the tooth: incisal edge, incisal third, body, and gingival third. The translucency of the enamel layer is typically higher than that of the dentin or cementum layer. In some embodiments, hue and translucency may be measured on a per-voxel (local) basis. In some embodiments, hue and translucency may be measured on a per-region basis, such as the incisor region, tooth body region, etc. The tooth body may refer to the portion of the tooth between the gingival region and the incisal edge.

[0121] 10) Contour height: Measurement of the tooth contour. When viewed from the proximal view, all teeth have a specific contour or shape moving from the gingival surface to the incisal edge. This is referred to as the facial contour of the tooth. In each tooth, there is a contour height where the shape is most pronounced. This contour height varies from the teeth in the front of the dental arch to the teeth in the back of the dental arch. In some embodiments, the measurement may take the form of fitting a template to known dimensions and / or known ratios. In some embodiments, the measurement may quantify the degree of curvature along the facial tooth surface. In some embodiments, the position where the curvature along the tooth contour is most pronounced is measured. This position may be measured as the distance from the gingival margin or the distance from the incisal edge, or as a percentage along the tooth length.

[0122] The entire text of the PCT application with publication number WO2020026117A1 is incorporated herein by reference. WO2020026117A1 lists some examples of orthodontic metrics (OM). Additional examples are disclosed herein. Orthodontic metrics can be used to quantify the physical arrangement of the dental arch for orthodontic treatment purposes (as opposed to prosthetic design metrics, which relate to dentistry and describe the shape and / or form of one or more pre-prosthetic teeth for supporting prosthetic dental purposes). These orthodontic metrics can measure the degree of malocclusion of the dental arch, or conversely, these metrics can measure the degree of correct arrangement of the teeth. In some specific implementations, the GDL setting model (or RL setting, VAE setting, capsule setting, MLP setting, diffusion setting, PT setting, similarity setting, and FDG setting) can incorporate one or more of these orthodontic metrics, or other similar or related orthodontic metrics. In some specific implementations, such an orthodontic metric can be incorporated into the feature vector of a mesh element, where these element-based feature vectors are provided as inputs to the setting prediction network. In some specific implementations, such an orthodontic metric can be directly used as a direct input by a generator, MLP, transformer, or other neural network (such as presented in one or more input vectors of real numbers S, as described elsewhere in this disclosure). Using such an orthodontic metric in the training of the generator can improve the performance (i.e., correctness) of the resulting generator, thereby producing a predicted transformation that places the teeth closer to the correct final setting pose than other possible cases. Such an orthodontic metric can be provided to an encoder structure or provided by a U-Net structure (in the case of the GDL setting). Such an orthodontic metric can be consumed by an autoencoder, variational autoencoder, masked autoencoder, or regularized autoencoder (in the case of the VAE setting, VAE mesh element labeling, MAE mesh filling). Such an orthodontic metric can be used by a neural network that predicts generation actions as part of an enhanced learning RL setting model. Such an orthodontic metric can be used by a classifier that applies labels to the set dental arch (such as labels for malocclusion, grading, or final setting). This description is non-limiting because orthodontic metrics can also be incorporated into the various techniques of this disclosure in other ways.

[0123] In some examples, the various loss calculations of the present disclosure can be combined with one or more orthodontic metrics, which has the advantage of improving the correctness of the resulting neural network. The orthodontic metrics can be used to directly compare the predicted examples with the corresponding ground truth examples (such as by using the metrics set in the comparison description). In other examples, one or more orthodontic metrics can be obtained from this part and incorporated into the loss calculation. Such orthodontic metrics can be calculated on the predicted examples, and then the orthodontic metrics will also be calculated on the ground truth examples. Then the results of these two orthodontic metrics are provided to the loss calculation, which has the advantage of improving the performance of the resulting neural network. In some specific implementations, one or more orthodontic metrics related to the alignment of two or more adjacent teeth can be calculated and incorporated into the loss function, for example, to at least partially train the setting prediction neural network. In some specific implementations, such orthodontic metrics can influence the network to align the mesial surface of a tooth with the distal surface of an adjacent tooth. Backpropagation is an example algorithm through which a neural network can be trained using one or more loss values.

[0124] In some specific implementations, one or more orthodontic metrics can be used to evaluate the predicted output of a neural network, such as setting prediction. Such metrics can enable the training algorithm to determine how close the predicted output is to an acceptable output, for example, in a quantitative sense. In some specific implementations, this use of orthodontic metrics can enable the calculation of loss values that do not solely depend on comparison with the ground truth. In some specific implementations, this use of orthodontic metrics can enable the loss calculation and network training to continue without the need to compare with ground truth examples. The advantage of this method is that the loss can be calculated based on general principles or specifications of the predicted output (such as settings), rather than associating the loss calculation with a specific ground truth example (which may have been defined by a specific doctor, clinician, or technician, whose treatment concept may be different from that of other technicians or doctors). In some specific implementations, such orthodontic metrics can be defined based on the FID (Fréchet Inception Distance) score.

[0125] The following is a description of some orthodontic metrics used to quantify the state of a set of teeth in an arch for orthodontic treatment. These orthodontic metrics indicate the degree of malocclusion of the teeth at a given stage of clear aligner treatment.

[0126] When training one of the neural networks of the present disclosure, it may be particularly advantageous to use orthodontic metrics calculated by tensor operations because tensor operations can facilitate efficient calculations. The more efficient (and faster) the calculations are, the faster the training can proceed.

[0127] In some examples, error patterns may be identified in one or more prediction outputs of the ML model (e.g., transformation matrices for predicting tooth settings, markings of mesh elements for mesh cleaning, addition of mesh elements to a mesh for mesh filling purposes, classification labels for settings, classification labels for tooth meshes, etc.). One or more orthodontic metrics may be selected to be input for the next round of ML model training to address any error or defect patterns that may be identified in the one or more prediction outputs.

[0128] Some OMs may be defined relative to the dental arch morphology coordinate system (LDE coordinate system). In some specific implementations, points may be described using the LDE coordinate system relative to the dental arch morphology, where L, D, and E respectively correspond to: 1) the length along the curve of the dental arch morphology, 2) the distance from the dental arch morphology, and 3) the distance in a direction perpendicular to the L-axis and the D-axis (which may be referred to as Eminence).

[0129] Various OMs and other techniques of the present disclosure may calculate conflicts between 3D representations (e.g., of oral care objects such as teeth). Such conflicts may be calculated as at least one of the following: 1) the penetration distance between 3D tooth representations, 2) the count of overlapping mesh elements between 3D tooth representations, and 3) the overlapping volume between 3D tooth representations. In some specific implementations, an OM may be defined to quantify the conflicts of two or more 3D representations of oral care structures (such as teeth). Some optimization algorithms (such as setting prediction techniques) may seek to minimize the conflicts between oral care structures (such as teeth).

[0130] Inter-arch orthodontic metrics are as follows.

[0131] Six (6) metrics for comparing two or more dental arches are listed below. Other suitable comparative orthodontic metrics are found elsewhere in the present disclosure, such as in the section on setting comparison techniques. 1. Rotational geodesic distance (rotation between the predicted example and the ground truth setting example) 2. Translation distance (gap between the predicted example and the ground truth setting example) 3. Normalized translation distance 4. 3D alignment error, which measures the distance between the predicted mesh elements and the ground truth mesh elements, in mm. 5. Normalized 3D alignment 6. Percentage of volume overlap (%) (alternatively % overlap of mesh elements) between the predicted example and the corresponding ground truth example

[0132] Intra-arch orthodontic metrics are as follows.

[0133] Alignment - The mesial - distal axis of the tooth can be used to calculate the 3D tooth orientation vector. A 3D vector that can be the tangent vector of the dental arch form at the tooth position can also be calculated. Then, the XY components (i.e., which can be 2D vectors) can be used to compare the orientation of the dental arch form at the tooth position with the tooth orientation in the XY space. Cosine similarity can be used to calculate the 2D orientation difference (angle) between the tangent of the dental arch form and the mesial - distal axis of the tooth.

[0134] Dental arch symmetry - For each pair of left and right teeth (e.g., the left lower lateral incisor and / or the right lower lateral incisor), the absolute difference between the X coordinate of each tooth and the X - axis of the global coordinate reference system can be calculated. This increment can indicate the dental arch asymmetry of a given tooth pair. The result of such a calculation can be the average X - axis increment from one or more tooth pairs of the dental arch. In some specific embodiments, this calculation can be performed with respect to the Y - axis having a Y coordinate (and / or with respect to the Z - axis having a Z coordinate).

[0135] D - dimensional difference of dental arch form - The D - dimensional difference (i.e., the positional difference in the facial - lingual direction) between two dental arch states of one or more teeth can be calculated. In some specific embodiments, a dictionary of D - direction tooth movements for each tooth, with the tooth UNS number as the key, can be returned. The LDE coordinate system with respect to the dental arch form can be used.

[0136] Ratio of the length of the lower dental arch form - The ratio between the current length of the lower dental arch and the length of the dental arch when it was in the initial malocclusion of the lower dental arch can be calculated.

[0137] Ratio of the length of the upper dental arch form - The ratio between the current length of the upper dental arch and the length of the dental arch when it was in the initial malocclusion of the upper dental arch can be calculated.

[0138] Parallelism of the dental arch form (whole dental arch) - For at least one origin of the local tooth coordinate system in the upper dental arch, one or more nearest origins (e.g., the origin of the tooth local coordinate system) in the lower dental arch. In some specific embodiments, two nearest origins can be used. The straight - line distance from a point in the upper dental arch to the line formed between the origins of two teeth in the opposite (lower) dental arch can be calculated. The standard deviation of the set of the above - mentioned "point - to - line" distances, where the set can consist of the point - to - line distances of each tooth in the dental arch, can be returned.

[0139] Arch form parallelism (single tooth) - This metric may share some computational elements with the global orthodontic metric of arch form parallelism, except that this metric can input the mean distance from the tooth origin to the line formed by adjacent teeth in the opposing arch (e.g., a tooth in the upper arch and the corresponding tooth in the lower arch). The mean distance can be calculated for one or more such tooth pairs. In some specific implementations, the mean distance can be calculated for all tooth pairs. Then, the mean distance can be subtracted from the distances calculated for each tooth pair. This OM can produce the deviation of the tooth from the "typical" tooth parallelism in the arch.

[0140] Buccolingual inclination - For at least one molar or premolar, find the corresponding tooth on the opposite side of the same arch (i.e., for a tooth on the left side of the arch, find the same type of tooth on the right side, and vice versa). This OM can calculate an n-element list (e.g., n can be equal to 2) for each tooth. The list can at least include the tooth IDs of the teeth in each pair of teeth (e.g., LeftLowerFirstMolar and RightLowerFirstMolar in the list = [left_tooth_idx_1, right_tooth_idx_2]). Such n-element vectors can be calculated for each molar and each premolar in the upper and lower arches. The buccal cusp can be identified on each side of the left and right sides of the arch for molars and premolars. Draw a line between the buccal cusp of the left tooth and the buccal cusp of the right tooth. Use this line and the z-axis of the arch form to make a plane. The lingual cusp can be projected onto this plane (i.e., at this point, the inclination angle can be determined). By performing additional projections, the approximate perpendicular distance between the lingual cusp and the buccal cusp can be calculated. This distance can be used as the buccolingual inclination OM.

[0141] Canine overbite - The upper and lower canines can be identified. The first premolar on a given side of the mouth can be identified. On a given side of the arch, the distance between the upper and lower canines can be calculated, and the distance between the upper and lower first premolars can also be calculated. An average value (or median, or mode, or some other statistical value) can be calculated for the measured distances. The z-component of the result indicates the degree of overbite. The overbite can be calculated between any tooth in one arch and the corresponding tooth in the other arch.

[0142] Canine cross-contact - The conflict (e.g., conflict distance) between the pairs of canines on the opposing arches can be calculated.

[0143] Canine cross-contact KDE - The orthodontic metric score of the current patient case can be used as input, and this score can be converted to a log-likelihood using a previously trained kernel density estimation (KDE) model or distribution. This operation can produce information about where the patient case is located in the distribution of "typical" values.

[0144] Canine Overjet - This OM can share some calculation steps with the canine overbite OM. In some specific implementations, an average distance can be calculated. In some specific implementations, the distance calculation can calculate the Euclidean distance of the XY components of a tooth in the upper dental arch and a tooth in the lower dental arch to produce the overjet (i.e., as opposed to calculating the difference in the Z component, as can be performed for canine overbite). The overjet can be calculated between any tooth in one dental arch and the corresponding tooth in the other dental arch.

[0145] Canine Class Relationship (also applicable to the first, second, and third molars) - In some specific implementations, this OM can include two functions (e.g., written in Python). get_canine_landmarks(): Obtain the landmarks for each tooth, which can be used to calculate the class relationship, and then in some specific implementations, map these landmarks onto the global coordinate space so that measurements can be made between teeth. class_relationship_score_by_side(): The average position of at least one landmark on at least one tooth in the lower dental arch can be calculated, and this value can be calculated for the upper dental arch. Then, a vector from the upper dental arch landmark position to the lower dental arch landmark position can be calculated, and finally, this vector can be projected onto the lower dental arch to produce a quantification (e.g., as a scalar) of the amount of increment in the "dental arch l-axis" position. This OM can calculate how far a tooth is positioned in front of or behind one or more teeth of interest in the opposite dental arch along the l-axis.

[0146] Interdigitation - By finding the midpoint between the distal marginal ridge saddle and the mesial marginal ridge saddle of a tooth, the fossa in at least one upper molar can be located. The cusp of the lower molar can be located between the marginal ridges of the corresponding upper molar. This OM can calculate the vector from the midpoint of the upper molar fossa to the cusp of the lower molar. This vector can be projected onto the d-axis of the dental arch form, thereby producing a transverse measurement of the distance from the cusp to the fossa. This distance can define the interdigitation magnitude.

[0147] Side Alignment - This OM can identify the leftmost and rightmost sides of a tooth and can identify the leftmost and rightmost sides of the adjacent teeth of that tooth. The OM can then draw a vector from the leftmost side of the tooth to the leftmost side of the adjacent tooth of that tooth. The OM can then draw a vector from the rightmost side of the tooth to the rightmost side of the adjacent tooth of that tooth. The OM can then calculate the linear fitting error between the two vectors. This calculation can involve producing two vectors: Vec_tooth = right_tooths_leftside to left_tooths_leftside Vec_neighbor = right_tooths_rightside to left_tooths_leftside Then it may involve calculating the dot product of these two vectors and subtracting the result from 1. (That is, edge alignment score = 1 - abs(dot(Vec_tooth, Vec_neighbor))). A score of 0 may indicate perfect alignment. A score of 1 may mean perpendicular alignment.

[0148] Inter-incisor arch contact KDE - may identify the deviation of the inter-incisor arch contact from the mean of the modeled distribution of this statistical information in the dataset of one or more other patient cases.

[0149] Levelling - may calculate a measure of the levelling between a tooth and its adjacent teeth. This OM may calculate the height difference between two or more adjacent teeth. For molars, this OM may use the midpoint between the mesial saddle ridge and the distal saddle ridge as the height of the molar. For non - molars, this OM may use the crown length from the gum to the tip. In some specific implementations, the tip may be the origin of the local coordinate space of the tooth. Other specific implementations may place the origin at other positions. A simple subtraction between the heights of adjacent teeth may yield the levelling increment between the teeth (e.g., by comparing the Z components).

[0150] Midline - may calculate the position of the midline of the upper incisors and / or lower incisors, and then may calculate the distance between them.

[0151] Inter - molar arch contact KDE - may calculate the inter - molar arch contact score (i.e., depth of conflict or other type of conflict), and then may identify the position of this score in a predefined KDE (distribution) constructed from representative cases.

[0152] Occlusal contact - For a specific tooth from an arch, this OM may identify one or more landmarks (e.g., mesial cusp or central cusp, etc.). Obtain the tooth transformation of this tooth. For each cusp on the current tooth, the cusp may be scored according to the degree of contact of the cusp with the adjacent (corresponding) tooth in the opposite arch. A vector from the cusp of the tooth under discussion to the vertical intersection point in the corresponding tooth of the opposite arch may be found. The distance and / or direction (i.e., up or down) to the opposite arch may be calculated. A list including the resulting signed distances, one for each cusp on the tooth under discussion, may be returned.

[0153] Overbite - may compare the upper central incisor and the lower central incisor along the z - axis. The difference along the z - axis may be used as the overbite score.

[0154] Overjet - may compare the upper central incisor and the lower central incisor along the y - axis. The difference along the y - axis may be used as the overjet score.

[0155] Intermolar Arch Contact - The contact fraction between molars can be calculated and conflict metrics (such as conflict depth) can be used.

[0156] Root Movement d - The tooth transformation for the initial state and the next state can be received. The arch form axis at point L along the arch form can be calculated. This OM can return the distance moved along the d axis. This can be achieved by projecting the root pivot point onto the d axis.

[0157] Root Movement l - The tooth transformation for the initial state and the next state can be received. The arch form axis at point L along the arch form can be calculated. This OM can return the distance moved along the l axis. This can be achieved by projecting the root pivot point onto the l axis.

[0158] Spacing - The spacing between each tooth and its adjacent tooth can be calculated. The transformation and mesh for the arch can be received. The left and right sides of each tooth mesh can be calculated. One or more points of interest can be transformed from local coordinates to the global arch coordinate system. The spacing can be calculated in the plane (e.g., the XY plane) between each tooth and its "left - hand" adjacent tooth. An array of one or more Euclidean distances (e.g., such as in the XY plane) can be returned, which can represent the spacing between each tooth and its left - hand adjacent tooth.

[0159] Torque - The torque (i.e., rotation about an axis such as the x - axis) can be calculated. For one or more teeth, one or more rotations can be converted from Euler angles to one or more rotation matrices. The components of the rotation (such as the x - component) can be extracted and converted back to Euler angles. This x - component can be interpreted as the torque of the tooth. A list including the torque of one or more teeth can be returned, and this list can be indexed by the UNS number of the teeth.

[0160] The neural network of the present disclosure can utilize one or more benefits of parameter tuning operations, thereby optimizing the input and parameters of the neural network to produce more data - accurate results. One parameter that can be tuned is the neural network learning rate (e.g., which can have values such as 0.1, 0.01, 0.001, etc.). The data augmentation scheme can also be tuned or optimized, such as a scheme that adds "shiver" to the tooth mesh before inputting to the neural network (i.e., small random rotations, translations, and / or scalings can be applied to change the dataset and make the neural network robust to data variations).

[0161] A subset of neural network model parameters that can be used for tuning is as follows: ○ Learning Rate (LR) decay rate (e.g., how much the LR decays during a training run) ○ Learning Rate (LR). A floating - point value used by the optimizer (e.g., 0.001). ○ LR Scheduling (e.g., cosine annealing, step, exponential) ○ Voxel size (for the case of sparse grid processing operations) ○ Dropout % (e.g., dropout that can be performed in a linear encoder) ○ LR decay step size (e.g., decay every 10 or 20 or 30 epochs) ○ Model scaling, which can increase or decrease the layer count and / or the parameter count per layer.

[0162] Parameter tuning can be advantageously applied to the training of neural networks to predict a final setting or an intermediate grading, thereby providing a technical improvement in data-oriented accuracy. Parameter tuning can also be advantageously applied to the training of neural networks for mesh element labeling or for mesh filling. In some examples, parameter tuning can be advantageously applied to the training of neural networks for tooth reconstruction. In terms of the classifier model of the present disclosure, parameter tuning can be advantageously applied to neural networks for the classification of one or more settings (i.e., the classification of one or more arrangements of teeth). The advantage of parameter tuning is to improve the data accuracy of the output of the prediction model or the classification model. In some cases, parameter tuning can provide the advantage of obtaining the last remaining few percentage points of validation accuracy from the prediction or classification model.

[0163] Various neural network models of the present disclosure can benefit from data augmentation. Examples include models trained on 3D meshes, such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, FDG settings, setting classification, setting comparison, VAE mesh element labeling, MAE mesh filling, mesh reconstruction VAE, and validation using autoencoders. Data augmentation can increase the size of the training dataset for the dental arch. Data augmentation can provide additional training examples by adding random rotations, translations, and / or rescaling to copies of the existing dental arch. In some specific implementations of the techniques of the present disclosure, data augmentation can be performed by perturbing or jittering the vertices of the mesh in a manner similar to that described in (“Equidistant and Uniform Data Augmentation for 3D Objects”, IEEE Access, Digital Object Identifier 10.1109 / ACCESS.2021.3138162). The position of the vertices can be perturbed by adding Gaussian noise, e.g., having a zero mean and a standard deviation of 0.1. According to the techniques of the present disclosure, other mean and standard deviation values are possible.

[0164] Some techniques of the present disclosure, such as setting comparison techniques and setting prediction techniques (e.g., such as GDL settings, MLP settings, VAE settings, etc.), can benefit from processing steps that can align (or register) dental arches (e.g., where teeth can be represented by 3D point clouds or some other type of 3D representation described herein). Such processing settings can be used, for example, to register a baseline ground truth set dental arch from a patient case with a malocclusion dental arch from the same case, and then these malocclusion dental arches and the baseline ground truth set dental arch can be used to train a setting prediction neural network model. Such steps can assist in loss calculation because the predicted dental arch (e.g., the dental arch output by the generator) can be better aligned with the baseline ground truth set dental arch, which is a condition that can facilitate the calculation of reconstruction loss, representation loss, L1 loss, L2 loss, MSE loss, and / or other types of losses described herein. In some specific implementations, the iterative closest point (ICP) technique can be used for such registration. ICP can minimize the squared error between corresponding entities such as 3D representations. In some specific implementations, linear least squares calculations can be performed. In some specific implementations, non-linear least squares calculations can be performed. Various registration models can incorporate in whole or in part portions of the following algorithms: Levenberg-Marquardt ICP, least squares rigid transformation, robust rigid transformation, random sample consensus (RANSAC) ICP, K-means based RANSAC ICP, and generalized ICP (GICP). In some cases, registration can help reduce subjectivity and / or randomness, which in some cases can occur in the design of a reference baseline ground truth set designed by a technician (i.e., two technicians may produce different but valid final setting outputs for the same case) or other optimization techniques.

[0165] In experiments, during the training of the setting prediction model, the ground truth (or reference) setting is registered to the malocclusion (or malocclusion setting). The malocclusion teeth are provided to the setting prediction model that generates the final setting transformation for the malocclusion teeth. The loss between the resulting predicted setting and the pre-registered ground truth setting is calculated such that the corresponding aspects of the two settings will be aligned. The result is a more accurate loss calculation. This pre-registration operation results in a 6% increase in absolute accuracy (e.g., as measured by the ADD10 score), which is equivalent to a nearly 50% reduction in the error rate compared to conventional techniques.

[0166] Since the generator network of the present disclosure can be implemented as one or more neural networks, the generator may include activation functions. When executed, the activation function outputs a determination as to whether a neuron in the neural network will fire (e.g., send an output to the next layer). Some activation functions may include: the binary step function or the linear activation function. Other activation functions impart non-linear behavior to the neural network, including: the sigmoid / logistic activation function, the Tanh (hyperbolic tangent) function, the rectified linear unit (ReLU), the leaky ReLU function, the parametric ReLU function, the exponential linear unit (ELU), the softmax function, the swish function, the Gaussian error linear unit (GELU), or the scaled exponential linear unit (SELU). The linear activation function may be well-suited for some regression applications (and other applications) in the output layer. In the output layer, the sigmoid / logistic activation function may be well-suited for certain binary classification applications (and other applications). The sigmoid activation function may be well-suited for some multi-class classification applications (and other applications) in the output layer. In the output layer, the sigmoid activation function may be well-suited for some multi-label classification applications (and other applications). The ReLU activation function may be well-suited for some convolutional neural network (CNN) applications (and other applications) in the hidden layer. The Tanh and / or sigmoid activation functions may be well-suited for some recurrent neural network (RNN) applications (and other applications) in, for example, the hidden layer. There are various optimization algorithms that can be used to train the neural networks of the present disclosure (such as updating the neural network weights), including gradient descent (which uses the first derivative to determine the training gradient and is commonly used in the training of neural networks), Newton's method (which may use the second derivative in the loss calculation to find a better training direction than gradient descent but may require calculations involving the Hessian matrix), and the conjugate gradient method (which may converge faster than gradient descent but does not require the Hessian matrix calculations that Newton's method may require). In some specific implementations, in addition to or instead of the above techniques, additional methods may be employed to update the weights. These additional methods include the Levenberg-Marquardt method and / or simulated annealing. The backpropagation algorithm is used to convey the results of the loss calculation back into the network so that the network weights can be adjusted for learning.

[0167] Neural networks contribute to the functionality of the applications of the present disclosure, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element labeling, MAE grid filling, grid reconstruction autoencoders, verification using autoencoders, estimation of oral care parameters, 3D grid segmentation (3D representation segmentation), coordinate system prediction, grid cleaning, restoration design generation, appliance component generation and / or placement, or dental arch morphology prediction. The neural networks of the present disclosure may embody some or all of various different neural network models. Examples include U-Net architectures, multi-layer perceptrons (MLPs), transformers, pyramid architectures, recurrent neural networks (RNNs), autoencoders, variational autoencoders, regularized autoencoders, conditional autoencoders, capsule networks, capsule autoencoders, stacked capsule autoencoders, denoising autoencoders, sparse autoencoders, conditional autoencoders, long / short-term memory (LSTM), gated recurrent units (GRUs), deep belief networks (DBNs), deep convolutional networks (DCNs), deep convolutional inverse graphics networks (DCIGNs), liquid state machines (LSMs), extreme learning machines (ELMs), echo state networks (ESNs), deep residual networks (DRNs), Kohonen networks (KNs), neural Turing machines (NTMs), or generative adversarial networks (GANs). In some specific implementations, an encoder structure or a decoder structure may be used. Each of these models offers one or more of its own specific advantages. For example, a particular neural network architecture may be particularly suitable for a specific ML technique. For example, autoencoders are particularly suitable for the classification of 3D oral care representations due to their ability to transform 3D oral care representations into a form that is easier to classify.

[0168] In some specific implementations, the neural networks of the present disclosure may be suitable for operating on 3D point cloud data (alternatively on 3D meshes or 3D voxelized representations). Many neural network specific implementations can be applied to the processing of 3D representations and can be applied to training prediction and / or generation models for oral care applications, including: PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN, and DSG-Net. Oral care applications include, but are not limited to: setting prediction (e.g., using VAEs, RLs, MLPs, GDLs, capsules, diffusion, etc. trained for setting prediction), 3D representation segmentation, 3D representation coordinate system prediction, element tagging for 3D representation cleaning (VAE for mesh element tagging), filling of missing elements in 3D representations (MAE for mesh filling), dental restoration design generation, setting classification, appliance component generation and / or placement, dental arch form prediction, estimation of oral care parameters, setting verification or other verification applications, and 3D representation classification of teeth.

[0169] Some specific implementations of the techniques of the present disclosure incorporate the use of autoencoders. Autoencoders that can be used in accordance with aspects of the present disclosure include, but are not limited to: AtlasNet, FoldingNet, and 3D-PointCapsNet. Some autoencoders can be implemented based on PointNet.

[0170] Representation learning can be applied to the setting prediction techniques of the present disclosure by training a neural network to learn a representation of a tooth and then using another neural network to generate a transformation of the tooth. Some specific implementations can use a VAE or a capsule autoencoder to generate a representation of the reconstructed features (in some cases, including information about the structure of the tooth mesh) of one or more meshes relevant to the oral care field. Then, this representation (latent vector or latent capsule) can be used as an input to a module that generates one or more transformations of one or more teeth. In some specific implementations, these transformations can place the tooth in a final set pose. In some specific implementations, these transformations can place the tooth in an intermediate graded pose. In some specific implementations, the transformation can be described by a 9×1 transformation vector (e.g., specifying a translation vector and a quaternion). In other specific implementations, the transformation can be described by a transformation matrix (e.g., a 4×4 affine transformation matrix).

[0171] In some specific implementations, the system of the present disclosure may perform principal component analysis (PCA) on an oral care mesh and use the resulting principal components as at least a part of the representation of the oral care mesh in subsequent machine learning and / or other predictive or generative processes.

[0172] An autoencoder can be trained to generate a latent form of a 3D oral care representation. The autoencoder can include a 3D encoder (which encodes the 3D oral care representation into a latent form) and / or a 3D decoder (which reconstructs the latent form into a copy of the input 3D oral care representation). Although the present disclosure relates to a 3D encoder and a 3D decoder, the term 3D should be interpreted in a non-limiting manner to cover multi-dimensional operation modes. For example, the system of the present disclosure can train a multi-dimensional encoder and / or a multi-dimensional decoder.

[0173] The system of the present disclosure can implement end-to-end training. Some end-to-end training-based techniques of the present disclosure can involve two or more neural networks, where the two or more neural networks are trained together (i.e., the weights are updated simultaneously during the processing of each batch of input oral care data). In some specific implementations, end-to-end training can be applied to pose prediction by simultaneously training a neural network that learns the representation of teeth and a neural network that can generate tooth transformations.

[0174] According to some transfer learning-based specific implementations of the present disclosure, a neural network (e.g., a U-Net) can be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task can be executed to provide one or more initial neural network weights for training another neural network, which is trained to perform a second task (e.g., pose prediction). The first network can learn the low-level neural network features of the oral care mesh and is shown to perform well in the first task. By using the first network as a starting point for training, the second network can exhibit faster training and / or improved performance. Certain layers can be trained to encode the neural network features of the oral care mesh in the training dataset. These layers can then be fixed (or undergo minor changes during the training process) and combined with other neural network components (such as additional layers), which are trained for one or more oral care tasks (such as pose prediction). In this way, a part of the neural network for one or more techniques of the present disclosure (e.g., pose prediction) can receive initial training for another task, which can result in important learning in the trained network layers. Then, this encoded learning can be built upon by further task-specific training of another network.

[0175] The systems of the present disclosure can utilize representation learning to train ML models. Advantages of representation learning include the fact that, as opposed to receiving inputs with variable sizes or structures, the generation network (e.g., a neural network used for predicting transformations in a setup prediction) can be configured to receive inputs with known sizes and / or standard formats. Representation learning can yield performance superior to other techniques because noise in the input data can be reduced (e.g., because the representation generation model extracts important aspects of the input representation (e.g., a mesh or point cloud) via a loss calculation or a network architecture selected for that purpose). Such loss calculation methods can include KL divergence loss, reconstruction loss, or other losses disclosed herein. Representation learning can reduce the size of the dataset required to train a model because the representation model learns a representation such that the generation network can focus on learning the generation task. Since meaningful features of the input data (e.g., local and / or global features) are available to the generation network, the result can be improved model generalization. In some cases, transfer learning can first train a representation generation model. Then this representation generation model (either wholly or in part) can be used to pre-train subsequent models, such as generation models (e.g., generation transformation predictions). The representation generation model can benefit from using mesh element features as inputs to improve the understanding of the structure and / or shape of the input 3D oral care representations in the training dataset.

[0176] According to the present disclosure, transfer learning can be used for setup prediction, as well as for other oral care applications, such as mesh classification (e.g., tooth or setup classification), mesh element labeling, mesh element filling, protocol parameter estimation, mesh segmentation, coordinate system prediction, restoration design generation, mesh validation (for any of the applications disclosed herein). In some embodiments, a neural network trained to output predictions based on an oral care mesh can first be partially trained on one of the following publicly available datasets before further training on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, Thingi10K dataset (which is particularly relevant for 3D printed component validation), ABC: Large CAD Model Dataset for Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Component Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.

[0177] In some specific implementations, a neural network previously trained on a first dataset (oral care data or other data) can subsequently receive further training on oral care data and be applied to oral care applications (such as setting prediction). Transfer learning can be used to further train any one of the following networks: GCN (Graph Convolutional Network), PointNet, ResNet, or any other neural network from the published literature listed above.

[0178] In some specific implementations, a first neural network can be trained to predict the coordinate system of teeth (such as by using the techniques described in WO2022123402A1 or U.S. Provisional Application No. US63 / 366492). According to any one of the setting prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein), a second neural network can be trained for setting prediction. Transfer learning can transfer at least a part of the knowledge or ability of the first neural network to the second neural network. Thus, transfer learning can provide an accelerated training phase for the second neural network to reach convergence. In some specific implementations, the training of the second network can be completed after being enhanced with the transferred learning and then using one or more techniques of the present disclosure.

[0179] The system of the present disclosure can utilize representation learning to train an ML model. The advantages of representation learning include that, in contrast to receiving inputs with variable sizes or structures, a generative network (e.g., a neural network used for predicting transformations in setting prediction) can be configured to receive inputs with known sizes and / or standard formats. Representation learning can produce performance superior to other techniques because noise in the input data can be reduced (e.g., because the representation generation model extracts hierarchical neural network features and / or reconstruction characteristics of the input representation (e.g., grid or point cloud) through loss calculation or the network architecture selected for this purpose).

[0180] The reconstruction characteristics may include values in a latent representation (e.g., a latent vector) that describe aspects of the shape and / or structure of the 3D representation provided to the representation generation module that generates the latent representation. For example, the weights of the encoder module of a reconstruction autoencoder may be trained to encode a 3D representation (e.g., a 3D mesh or others described herein) into a latent vector representation (e.g., a latent vector). In other words, the ability to encode a large set of mesh elements (e.g., hundreds, thousands, or millions) into a latent vector (e.g., hundreds or thousands of real values, e.g., 512, 1024, etc.) can be learned through the weights of the encoder. Each dimension of the latent vector may include real numbers that describe some aspects of the shape and / or structure of the initial 3D representation. The weights of the decoder module of the reconstruction autoencoder may be trained to reconstruct the latent vector into a close replica of the initial 3D representation. In other words, the decoder can learn the ability to interpret the dimensions of the latent vector and decode the values within those dimensions. Generally speaking, the encoder and decoder neural network modules are trained to perform a mapping of the 3D representation to a latent vector, and then the latent vector can be mapped back (or otherwise reconstructed) to a 3D representation that is substantially similar to the initial 3D representation for which the latent vector was generated.

[0181] Returning to the loss calculation, examples of the loss calculation may include KL divergence loss, reconstruction loss, or other losses disclosed herein. Representation learning can reduce the size of the dataset required to train a model because the representation model learns a representation such that the generation network can focus on learning the generation task. Since meaningful neural network features (e.g., local and / or global features) of the input data are available to the generation network, the result can be improved model generalization. In other words, the first network can learn a representation, and the second network can make prediction decisions. By training two networks to perform their own separate tasks, each network can generate more accurate results for its corresponding task than a single network that is trained to both learn a representation and make decisions. In some cases, transfer learning may first train a representation generation model. Then this representation generation model (either wholly or in part) can be used to pre-train subsequent models, such as a generation model (e.g., generation transformation prediction). The representation generation model can benefit from using mesh element features as input to improve the ability of the second ML module to encode the structure and / or shape of the input 3D oral care representation in the training dataset.

[0182] One or more neural network models of the present disclosure may have attention gates integrated therein. The attention gate integration provides an enhancement that enables the associated neural network architecture to focus resources on one or more input values. In some specific implementations, the attention gate may be integrated with a U-Net architecture, which has the advantage of enabling the U-Net to focus on certain inputs, such as input landmarks corresponding to teeth that are intended to be fixed (e.g., prevented from moving) during orthodontic treatment (or in cases where other special treatment is required). According to aspects of the present disclosure, the attention gate may also be integrated with an encoder or with an autoencoder (such as a VAE or a capsule autoencoder) to improve prediction accuracy. For example, the attention gate may be used to configure a machine learning model to give higher weights to aspects of the data that are more likely to be related to the correctly generated output. Thus, and because the machine learning models configured with these attention gates (or mechanisms) utilize aspects of the data that are more likely to be related to the correctly generated output, the final prediction accuracy of those machine learning models is improved.

[0183] The quality and composition of the training dataset for a neural network can affect the performance of the neural network during its execution phase. Dataset screening and outlier removal can be advantageously applied to the training of neural networks for various techniques of the present disclosure (e.g., for predictions of final settings or intermediate gradings, for neural networks for mesh element labeling or for mesh filling, for tooth reconstruction, for 3D mesh classification, etc.), because dataset screening and outlier removal can remove noise from the dataset. Although the mechanisms for achieving the improvement are different from using attention gates, the end result is that the method allows the machine learning model to focus on the relevant aspects of the dataset and can lead to an improvement in accuracy similar to the improvement achieved with attention gates.

[0184] In the case of a neural network configured to predict a final setting, the patient case may include at least one of a set of segmented tooth meshes of the patient, the malocclusion transformation of each tooth, and / or the ground truth setting transformation of each tooth. In the case of a neural network predicting a set of intermediate stage settings, the patient case may include at least one of a set of segmented tooth meshes of the patient, the malocclusion transformation of each tooth, and / or a set of ground truth intermediate stage transformations of each tooth. In some specific implementations, the training dataset may exclude patient cases in the contact passive phase (i.e., the phase where the teeth of the dental arch do not move). In some specific implementations, the dataset may exclude cases where there is a passive phase at the end of the treatment. In some specific implementations, the dataset may exclude cases where there is overcrowding at the end of the treatment (i.e., cases where an oral care provider such as an orthodontist or a dentist has selected a final setting where the tooth meshes overlap to some extent). In some specific implementations, the dataset may exclude cases of a specific difficulty level (or levels) (e.g., easy, medium, and difficult).

[0185] In some specific implementations, the dataset may include cases with zero pinned teeth (or may include cases with at least one pinned tooth). A technician can specify the pinned teeth during their design process to prevent various tools from moving that particular tooth. In some specific implementations, the dataset may exclude cases with no fixed teeth (conversely, where at least one tooth is fixed). Fixed teeth can be defined as teeth that should not move during the process. In some specific implementations, the dataset may exclude cases with no pontic teeth (conversely, cases where at least one tooth is a pontic). Pontic teeth can be described as "ghost" teeth, which are represented in the digital model of the dental arch but do not actually exist in the patient's dentition, or where there may be small teeth or partial teeth that could benefit from future work, such as adding composite materials through a prosthodontic appliance. The advantage of including pontic teeth in a patient case is to leave space in the dental arch as part of the plan for the movement of other teeth during orthodontic treatment. In some cases, pontic teeth can save space in the patient's dentition for future dental or orthodontic work, such as installing implants or crowns, or applying prosthodontic appliances, such as adding composite materials to existing teeth that are too small or have an undesirable shape.

[0186] In some specific implementations, the dataset may exclude cases where the patient does not meet the age requirement (e.g., less than 12 years old). In some specific implementations, the dataset may exclude cases where interproximal reduction (IPR) exceeds a certain threshold amount (e.g., greater than 1.0 mm). The dataset for training a neural network to predict the settings of a clear tray appliance (CTA) may exclude patient cases unrelated to CTA treatment. The dataset for training a neural network to predict the settings of an indirectly bonded tray product may exclude cases unrelated to indirectly bonded tray treatment. In some specific implementations, the dataset may exclude cases where only certain teeth are treated. In such specific implementations, the dataset may only include cases where at least one of the following is treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and / or canines.

[0187] The mesh comparison module can compare two or more meshes, for example, for the calculation of a loss function or for the calculation of a reconstruction error. Some specific implementations may involve the comparison of the volumes and / or areas of two meshes. Some specific implementations may involve calculating the minimum distance between corresponding vertices / faces / edges / voxels of two meshes. For a point in one mesh (e.g., a vertex, the midpoint on an edge, or the center of a triangle), calculate the minimum distance between that point and the corresponding point in the other mesh. In cases where the other mesh has a different number of elements or there is no clear mapping between corresponding points of the two meshes, different methods may be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools that can be used in the mesh comparison module of the present disclosure. In some specific implementations, the Hausdorff distance can be calculated to quantify the shape difference between two meshes. The open-source software tool Metro developed by the Visual Computing Lab can also be used to quantify the difference between two meshes. The following paper describes the method adopted by Metro, which can be modified by the neural network application of the present disclosure for mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces", P. Cignoni, C. Rocchini, and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, Vol. 17(2), June 1998, pp. 167-174.

[0188] Some techniques of the present disclosure may combine the following operations: for one or more points on a first mesh, project a ray perpendicular to the mesh surface and calculate the distance before the ray impinges on a second mesh. The length of the resulting line segment can be used to quantify the distance between the meshes. According to some techniques of the present disclosure, a color can be assigned to the distance based on the magnitude of the distance, and the color can be applied to the first mesh by means of visualization.

[0189] The prosthetic treatment of a patient can involve the specification of one or more of the following: prosthetic guidelines, prosthetic design parameters, and / or prosthetic rules for modifying one or more aspects of the patient's dentition. One or more of many possible factors may be considered when designing a 3D prosthesis, whether from an aesthetic perspective and / or from a technical perspective. For example, from an aesthetic perspective, the dental and facial midlines and angles can provide overall guidance, such as the amount of teeth visible to others when the lips are at rest and / or smiling. After considering these criteria, a set of "golden ratios" can also describe the aesthetic design of overall tooth size. The tooth-to-tooth ratios can be configured to reflect these "golden ratios," which are 1.618:1.0:0.618 for the central incisors, lateral incisors, and canines, respectively. In some specific implementations, the actual values can be specified for one or more of the RDPs and received at the input of a dental restoration design prediction model (e.g., a machine learning model that predicts the final tooth shape at the completion of the restoration design). In some specific implementations, one or more RDPs can be defined corresponding to one or more prosthetic design metrics (RDMs).

[0190] The constraints based on tooth position (i.e., malocclusion) and orientation (i.e., rotation and tilt) are balanced against the attempt to achieve proper symmetry, tooth proportions, and tooth-to-tooth ratios. After establishing these parameters, various tooth shapes can be utilized to match the overall aesthetics of the patient's face and smile. For example, the tooth shapes can be generally rectangular with square edges, or they can be generally oval with rounded edges. Additionally, the tooth-to-tooth ratios can be manipulated to achieve different overall aesthetics. 3D dental CAD programs typically provide a library of different tooth "styles" to choose from and give the designer the ability to adjust the results to best match the aesthetic and medical requirements of the doctor and patient. In some examples, symmetry may be observed, as the left side should mirror the right side, and symmetry can thus be measured.

[0191] The tooth length, width, and aesthetic relationship between the width and length can be specified for one or more teeth. In one example, the length of the maxillary central incisor can be set to 11 mm, and the aesthetic relationship between the width and length can be set to 70% or 80%. In some examples, the lateral incisors can be 1.0 mm to 2.5 mm shorter than the central incisors. In some cases, the canines can be 0.5 mm to 1.0 mm shorter than the central incisors. Other ratios and measurements are possible for various teeth.

[0192] From a technical perspective, there are other considerations to take into account. For example, a prosthesis made of a given material must have sufficient thickness to have the necessary mechanical strength for long-term use. Additionally, the width and shape of the teeth must be designed to provide proper contact with adjacent teeth.

[0193] The example style options in the following list are from the LVI standard, from the Las Vegas Institute for Advanced Dental Studies (LVI). Other style guides are available commercially or for free.

[0194] In these and / or other examples, the neural network engine of the present disclosure can be combined with one or more of the accepted "golden ratio" guidelines for tooth size, the accepted "ideal" tooth shape, patient preferences, practitioner preferences, etc. as inputs.

[0195] Restorative design parameters (RDPs) can be used to encode aspects of the smile design guidelines described herein, such as parameters related to the expected size of the restored tooth. The restorative design parameters are intended to serve as an indication and / or specification of the shape and / or form that one or more teeth should present after a dental restoration procedure. One or more RDPs can be received by a neural network or other machine learning or optimization algorithm for dental restoration design, with the advantage of providing guidance for the optimization algorithm. Some neural networks can be trained for dental restoration design generation, such as some examples of GANs or autoencoders. In some cases, the dental restoration design can be used to define the target shape of one or more teeth for generating dental restoration appliances. In some cases, the dental restoration design can be used to define the target tooth shape for generating one or more veneers.

[0196] A partial list of tooth dimensions can include length, width, height, perimeter, diameter, diagonal measurement, volume, and any of these dimensions can be normalized compared to another tooth or teeth. In some specific implementations, one or more restorative design parameters can be defined that relate to the gap between two or more teeth and the size of the gap that the patient wishes to retain after treatment (if any) (e.g., such as when the patient wishes to retain a small gap between the maxillary central incisors).

[0197] Additional restorative design parameters can include the parameters specified in Table 2. In the case where one parameter conflicts with another parameter, the following order can determine the precedence (i.e., let the first parameter in the following list be considered authoritative). If a parameter value is not specified, the parameter can be ignored. In some specific implementations, default values can be introduced for one or more parameters. For example, such default values can be determined by clustering previous patient cases. The golden ratio guidelines can specify one or more numbers related to the width of adjacent teeth, such as: {1.6, 1, 0.6}.

[0198] Tooth-to-tooth ratios can also be defined between other tooth pairs. The ratios can be made with respect to tooth width, height, diagonal, etc. Angular lines, chamfer angles, and buccal contours can describe the main aspects of the macroscopic shape of the tooth. The marginal ridge groove can be a vertical macroscopic texture on the front of the tooth and can sometimes be V-shaped. Striations or perikymata can be horizontal micro-textures on the tooth. Symmetry is usually desired. There may be differences between male and female patients.

[0199] Parameters can be defined to encode doctor restoration design preferences (DRDP) related to various use case scenarios. These use case scenarios can reflect information about the treatment preferences of one or more doctors and directly affect the properties of one or more teeth in a dental restoration design or veneer. Additionally, DRDP can describe the RDP values or ranges of values that are habitually involved in the preferences or habits of doctors or other treating healthcare professionals. In some cases, such values or ranges of values can be derived from historical patient cases treated by that doctor or healthcare professional. In some cases, DRDP can be defined from RDP (e.g., the aesthetic relationship such as width to length) or from RDM.

[0200] Machine learning models (such as those described herein) (e.g., denoising diffusion models) can be trained to generate designs for crowns or roots (or both). A dental restoration design can describe the expected tooth shape at the end of a dental restoration. A neural network (such as a generative neural network) can be trained to generate dental restoration designs that will be used to generate veneers (e.g., zirconia veneers) or dental restoration appliances. Such models take as input data from a cohort of patient cases, including pre-restoration tooth meshes and corresponding ground truth examples of the completed restorations (e.g., tooth meshes with restoration shape and / or structure). Such models can be trained at least in part by computing a loss function that can quantify the difference between the generated crown restoration design and the ground truth crown restoration design. The resulting loss can be used to update the weights of the generative neural network model (e.g., a denoising diffusion model, which can include a U-Net), thereby (at least in part) training the model. A reconstruction loss can be computed to compare the predicted tooth mesh with the ground truth tooth mesh or the pre-restoration tooth mesh with the completed restoration design tooth mesh. The reconstruction loss can be computed as the sum of the pairwise distances between corresponding mesh elements and can be computed to quantify the difference between two crown designs.

[0201] Other losses disclosed herein can also be in the training. The generated restoration designs can be used to create veneers. For example, the veneers can be 3D printed.

[0202] Machine learning models (such as those described herein) (e.g., denoising diffusion models) can be trained to generate components used in creating dental restoration appliances. Such dental restoration appliances can be used to shape dental composites in a patient's oral cavity while curing the composites (e.g., using a curing light), ultimately producing veneers on one or more of the patient's teeth. 3M FILTEK Matrix is an example of such a product. In some cases, the machine learning model used to generate appliance components can take inputs that are operable to customize the shape and / or structure of the appliance components, including inputs such as oral care parameters. In some cases, one or more oral care parameters can be defined based on oral care metrics. Oral care metrics (e.g., orthodontic metrics or prosthodontic design metrics) can describe the physical and / or spatial relationships between two or more teeth, or can describe the physical and / or dimensional characteristics of a single tooth. Oral care parameters can be defined that are intended to provide guidance to the machine learning model regarding generating a 3D oral care representation having specific physical characteristics (e.g., related to shape and / or structure). For example, physical characteristics can be measured using the oral care metrics corresponding to the oral care parameters. Such oral care parameters can be defined to customize the generation of mold parting surfaces, gingival trim meshes, or other generated appliance components to adapt those appliance components to the patient's dental anatomy.

[0203] In some specific implementations, the 3D representation generation techniques described herein (e.g., denoising diffusion-based techniques) can be trained to generate custom appliance components by determining the characteristics of the custom appliance components, such as the size, shape, position, and / or orientation of the custom appliance components. Examples of custom appliance components include mold parting surfaces, gingival trimming surfaces, shells, facings, lingual shelves (also referred to as "reinforcing ribs"), doors, windows, incisal ridges, outer shell frame spare parts or interdental matrix wraps, splines, etc. A spline refers to a curve that passes through multiple points or vertices, such as a piecewise polynomial parametric curve. A mold parting surface refers to a 3D mesh that bisects the two sides of one or more teeth (e.g., separates the facial side of one or more teeth from the lingual side of one or more teeth). A gingival trimming surface refers to a 3D mesh that trims the shell along the gingival margin. A shell refers to a body with a nominal thickness. In some examples, the inner surface of the shell matches the surface of the dental arch, and the outer surface of the shell is a nominal offset of the inner surface. A facing refers to a reinforcing rib with a nominal thickness that is offset from the face of the shell. A window refers to an orifice that provides access to the tooth surface so that dental composite materials can be placed on the tooth. A door refers to a structure that covers the window. An incisal ridge provides reinforcement at the incisal edge of the dental appliance and can be obtained from the dental arch morphology. An outer shell frame spare part refers to a connecting material that couples the components of the dental appliance (e.g., the lingual part of the dental appliance, the facial part of the dental appliance, and its sub-components) to the manufactured outer shell frame. In this way, the outer shell frame spare part can bind the components of the dental appliance to the shell frame during manufacturing, protect the individual components from damage or loss, and / or reduce the risk of mixing components. These appliance components and other components are described in PCT patent applications WO2020240351A1 and WO2021240290A1, both of which are incorporated herein by reference in their entireties.

[0204] A mold parting surface refers to a 3D mesh that bisects the two sides of one or more teeth (e.g., by separating the facial side of one or more teeth from the lingual side of one or more teeth). A gingival trimming surface refers to a 3D mesh that trims the shell along the gingival margin. A shell refers to a body with a nominal thickness. In some examples, the inner surface of the shell matches the surface of the dental arch, and the outer surface of the shell is a nominal offset of the inner surface.

[0205] The facial surface refers to a reinforcing rib of a nominal thickness offset from the surface of the housing. The window refers to an orifice that provides access to the tooth surface such that a dental composite material can be placed on the tooth. The door refers to a structure that covers the window. The incisal ridge provides reinforcement at the incisal edge of the dental restoration appliance and can be obtained from the dental arch form. The outer shell frame spare part refers to a connecting material that couples components of the dental restoration appliance (e.g., the lingual portion of the dental restoration appliance, the facial portion of the dental restoration appliance, and its subassemblies) to the manufactured outer shell frame. In this way, the outer shell frame spare part can bind the components of the dental restoration appliance to the housing frame during manufacturing, protect the individual components from damage or loss, and / or reduce the risk of mixing components.

[0206] Additional 3D oral care representations that can be generated by a denoising diffusion model (such as the denoising diffusion model trained as described herein) include dental restoration designs (e.g., which can include interproximal tooth surfaces, fossae, incisal edges, tooth tips, tooth roots, etc.).

[0207] In some embodiments, the denoising diffusion model (DDM) described herein can be trained to perform 3D mesh element labeling in a 3D oral care representation (e.g., labeling vertices, edges, faces, voxels, or points). Those labeled mesh elements can be used for mesh cleaning or mesh segmentation. In the case of mesh cleaning, the labeled aspects of the scanned tooth mesh can be used for appliance erasure (remove + replace) or for modifying (e.g., by smoothing) one or more aspects of the tooth to remove aspects of the attached hardware (or other aspects of the mesh that may not be needed for certain processing and appliance creation, such as foreign materials). Mesh element features, such as those described herein, can be computed for one or more mesh elements in the 3D representation of the oral care data (e.g., the 3D representation of the patient's dentition). A vector of such mesh element features can be computed for each mesh element and then received by the DDM that has been trained to label mesh elements in the 3D oral care representation for mesh segmentation or mesh cleaning purposes. Such mesh element features can impart valuable information about the shape and / or structure of the input mesh to the labeling DDM. For example, the training data 200 can include a 3D representation of the patient's pre-restorative dentition (e.g., from an intraoral scanner, a CT scanner, etc.), as well as ground truth data including ground truth mesh element labels. The ground truth mesh element labels can describe the ground truth (or reference) segmentation of the patient's dentition (e.g., each mesh element that occupies the lower left central incisor can have the same label, each mesh element in the gingiva of the upper dental arch can have the same label, each mesh element that appears in the upper right canine can have the same label, etc.). Either or both of the patient's dentition and the ground truth mesh element labels can undergo latent encoding (214). The initial training data 200 or the training data corresponding to the latent encoding can be provided to the Markov chain 204 as part of the forward pass 206. The forward pass can iteratively add noise to the mesh element labels, thereby generating a series of dozens, hundreds, or thousands of successive noisy versions of a set of ground truth mesh element labels. In some embodiments, the successive addition of noise can include randomizing the values of one or more mesh element labels. In some embodiments, the successive addition of noise can include adjusting the labels of the neighbors of a randomly selected mesh element (e.g., a mesh element placed along the boundary of a tooth). In this way, the boundaries between teeth or between teeth and the gingiva can be adjusted to become increasingly noisy or distorted. The Markov chain can generate a series of increasingly noisy sets of mesh element labels, which can be used to at least partially train the denoising ML model 210. In some embodiments, the 3D representation of the patient's dentition can also be provided to the training of the denoising ML model 210 along with the mesh element labels. In deployment, a backward pass 208 can be performed. An oral care variable 202 can be provided to customize the functionality of the segmentation operation. In deployment, the immediate patient case data 216 (e.g., the pre-segmentation 3D representation of the patient's dentition) can be provided to the denoising ML model 210.In some specific implementations, the immediate patient case data 216 may undergo potential encoding (218). In some specific implementations, a set of initial mesh element labels corresponding to the 3D representation of the patient dentition 216 may be generated. In some specific implementations, the mesh element labels may be initialized randomly or according to heuristics. An example of a heuristic is to select a set of random mesh elements, and the neighbors of those mesh elements (e.g., neighbors having a specific distance or number of edges from the selected mesh elements) may be grouped together in small cliques (e.g., each mesh element in a clique has the same mesh element label). The backpropagation 208 of the denoising ML model 210 may cause the denoising ML model 210 to iteratively refine the set of mesh element labels with respect to the 3D representation of the patient dentition 216. After multiple iterations (e.g., hundreds, thousands, or millions), the iteratively refined mesh element labels may be set as the output 220. These generated mesh element labels can be used to segment the patient dentition or perform mesh cleaning of the patient dentition. In some examples, when the immediate patient case data contains the pre-segmented dentition of the patient, the patient dentition can be segmented by applying the generated mesh element labels 220. According to the generated mesh element labels 220, the generated mesh element labels can enable tooth cutting (e.g., generating a new mesh for each different tooth).

[0208] Some specific implementations of the DDM-based mesh cleaning techniques described herein may train the DDM to remove (or modify) general triangular mesh defects (e.g., via mesh element marking), such as: degenerate triangles with zero surface area; redundant triangles that cover the same surface area as another triangle; non-manifold edges with more than two adjacent triangles, also known as "flaps"; non-manifold vertices with more than one adjacent sequence of connected triangles (triangle fans); intersecting triangles - where two triangles pass through each other; spikes - sharp features composed of multiple triangles, typically conical, caused by one or more vertices deviating from the actual surface; folds - sharp features composed of multiple triangles, typically Z-shaped with a small undercut area, caused by one or more vertices deviating from the actual surface; islands / sub-components - disconnected objects where only a single object should be included in the scan (e.g., typically smaller objects are deleted); small holes in the mesh surface, either from the original scan or from deletions due to previous defects (e.g., the holes can be removed by filling the holes, e.g., by adding one or more mesh elements); rough boundaries - smooth boundaries are beneficial for extending the gingival surface and creating a model base.

[0209] Some specific implementations of the DDM-based mesh cleaning techniques described herein can train a DDM model to remove (or modify) aspects of a mesh that are not needed and / or are domain-specific defects in certain cases (e.g., via mesh element tagging), such as: extraneous material—the portion of an intraoral scan outside of the anatomical region of interest, e.g., non-tooth surfaces not within a certain distance of a tooth surface, or scan artifacts that do not represent actual anatomical structures; concavities—recesses in a surface (e.g., which may be scan artifacts to be repaired or anatomical features that are generally left intact); undercuts—a tooth side that is smaller than the crown radius, and thus a physical impression or appliance may be difficult to remove or place. Undercuts can be natural features or caused by damage such as internal fractures. Internal fractures are associated with erosion of the tooth near the gum line, which causes or exacerbates undercuts. Appliances that can be processed by the DDM-based models of the present disclosure include orthodontic hardware such as attachments, brackets, wires, buttons, lingual bars, Carriere appliances, etc., and may be present in an intraoral scan. In some cases, it may be beneficial to perform digital removal and replacement with synthetic tooth / gum surfaces prior to the appliance creation step.

[0210] Denoising diffusion can train a neural network (or other machine learning model) to iteratively remove noise (or refine) from a data structure. The techniques of the present disclosure remove noise from a 3D point cloud (or other 3D representation), from a transformation matrix, a label vector (e.g., a label applied to a mesh element and used for segmentation or mesh cleaning, and which can be used to define an object mask), or from other kinds of 3D oral care representations. These techniques can start with a randomly generated data structure (e.g., a 3D point cloud or transformation matrix), and these techniques can transform those random (or noisy) data structures into artifacts that can be used for digital oral care. Due to the kinds of training data generated using the techniques described herein, denoising models can be trained to perform these operations. The training data can include a series of increasingly noisy examples of the relevant data structures. A denoising neural network (U-Net is an example) can be trained to gradually transform a random (or noisy) data structure into an artifact that can be used for digital oral care processing (e.g., the generation of oral care appliances) in small steps.

[0211] Digital oral care involves many different kinds of 3D representations customized to a patient's anatomy. Due to the customizable nature of generating the output representations, diffusion models are particularly well-suited for digital oral care.

[0212] In some specific implementations, the denoising diffusion model can operate on input data in its initial data format (e.g., 3D oral care representation) (e.g., training on the 3D representation and / or generating a 3D representation such as a 3D point cloud, etc.). In other specific implementations, the denoising diffusion model can operate on the latent form of the input data (e.g., such as in Stable Diffusion), which can preserve the 3D nature of the input data. The latent form can be information-rich (describing the structure and / or shape of the input data such as the 3D oral care representation) and / or have a small data footprint, which is easier for an ML model to train on. In some cases, an autoencoder can be trained to generate the latent representation. When the latent representation is reconstructed as a close facsimile of the initial 3D representation (e.g., as measured by the reconstruction error defined herein), the information-rich nature of the latent representation can be demonstrated. In other words, the initial 3D representation can include hundreds or thousands of mesh elements, which in some cases can be encoded by the autoencoder into a few hundred real-valued vectors. In some cases, these few hundred real-valued vectors can then be reconstructed as a close facsimile of the initial 3D representation. The accuracy of the reconstruction (e.g., the fidelity with which the reconstructed form matches the initial form) can be verified by computing the reconstruction error.

[0213] In some specific implementations, an oral care diffusion model (OCDM) can be trained to modify the 3D oral care representations described herein (e.g., appliance components, tooth restoration designs, trim lines, transformations for 3D oral care representations such as teeth or appliance components in an orthodontic setting, etc.). In some cases, a template or reference 3D oral care representation can be received at the input and then modified by the OCDM. In some cases, an initial tooth representation (e.g., pre-restoration teeth) can be received at the input and then modified by the OCDM to generate teeth ready for a restoration process (e.g., for manufacturing a crown, bridge, or dental restoration appliance such as FILTEK Matrix). In some specific implementations, the OCDM can be trained to generate the 3D oral care representations described herein (e.g., appliance components, tooth restoration designs, trim lines, orthodontic settings, etc.). In some specific implementations, the diffusion model can be trained to generate 3D oral care representations (e.g., a point cloud or 3D mesh representing a 3D oral care representation such as a trim line, appliance components or tooth designs for dental restorations, such as the generation of veneers, the generation of dental restoration appliances). Such a model can be referred to as a geometric generation diffusion neural network or geometric generation diffusion model (GGDM).

[0214] In some cases, an encoder can be used to preprocess the input data of a diffusion model (e.g., the Stable Diffusion model) (e.g., to generate a representation with reduced dimensionality of the input data structure). In some cases, a decoder can be used to reconstruct such a latent representation. The decoder can reconstruct the latent representation into a facsimile of the input data (e.g., such as with an autoencoder). In some cases, the decoder can reconstruct the latent representation into a modified form of the input data (e.g., when one or more aspects of the latent representation are modified). A point-voxel CNN (PVCNN) or an encoder-decoder network (e.g., such as a variational autoencoder) can be trained to implement the encoder or decoder of the techniques described herein. The 3D oral care representation 102 can be received at the input of an oral care diffusion model (OCDM) and encoded by an encoder into a latent representation (e.g., a latent vector). In some embodiments, such as in Figure 1 in, an optional encoder 104 (e.g., as a structural latent encoder that can encode aspects of the shape and / or structure of the input data) can encode the input 3D oral care representation 102 into a latent representation 108 of that shape (e.g., referred to as a structural latent representation). In some embodiments, such as in Figure 1 in, an optional latent grid element encoder 106 can encode the input 3D oral care representation 102 into a latent representation 110 of grid elements (e.g., referred to as a latent grid element representation). Such grid elements can include points in a point cloud, voxels in a sparse representation, or vertices / edges / faces in a 3D mesh, etc. In some embodiments, each of these two representations can be generated. In some embodiments, such encoders can incorporate grid element features at the input to improve the latent representation generated by the encoder. The grid element feature vectors can be calculated by grid element feature modules 130 and 128.

[0215] During the forward process, a Markov chain of diffusion steps can be applied to generate a training dataset from such a latent representation. The forward process can generate increasingly noisy versions of the input 3D oral care representation 102, which can be used to train one or more denoising diffusion neural network models described herein. This Markov chain of diffusion steps can apply Gaussian noise to the input 3D oral care representation 102 over T steps (e.g., T = 150), which can produce a series of increasingly noisy versions of the input 3D oral care representation 102 now available for training the denoising diffusion neural network. A diffusion neural network (e.g., which can be implemented using a U-Net, variational autoencoder, etc.) can be trained to invert the forward process, starting from the 3D oral care representation at time T, and perform denoising operations to transform the 3D oral care representation into a version corresponding to time T-1. This is referred to as the reverse process. Given the latent representation at time t, the diffusion neural network can iteratively compute the version of the latent representation at time t-1. The diffusion neural network can introduce changes to the latent representation (e.g., to remove noise or introduce desired geometric or structural aspects in the output). The diffusion model can consume instructions regarding such changes in the form of oral care variables. Oral care variables can include one or more natural language text string variables, one or more categorical variables, one or more integer variables, one or more real-valued variables, one or more image variables (e.g., reference images depicting aspects of the expected outcome of the model's execution), one or more 3D representation variables (e.g., reference 3D point clouds depicting aspects of the expected outcome of the model's execution), or combinations of multiple of the foregoing variables. Such variables can first undergo latent encoding (e.g., using the encoder portion of an autoencoder or using a text or numerical encoder). The encoded variables can be introduced into the diffusion neural network (e.g., by appending the variables to the input data, such as concatenating the variables with the latent representation of the input data generated using the encoder). For example, when implementing the diffusion neural network using a U-Net, the encoded variables can be introduced into one or more successive resolution layers of the U-Net (e.g., the U-Net can have a set of layers corresponding to increasingly lower resolution levels, followed by a set of layers corresponding to increasingly higher resolution levels). The U-Net can employ skip connections to share information from the input to the output at each resolution level. The U-Net can incorporate attention layers. The encoded variable information can be introduced into any or all such layers, with the advantage of guiding the generated output of the diffusion neural network. In some specific implementations, the denoising diffusion neural network of the present disclosure can include a U-Net, ResNet, transformer, or autoencoder (e.g., variational autoencoder), etc.

[0216] A potential representation of the shape of the 3D point cloud 102 (or other 3D representations described herein) is referred to as a structural latent representation 108, which can encode aspects of the structure of the input 3D point cloud 102. The oral care variable 100 (or input parameter) can include one or more attributes that describe the expected output from a trained machine learning model. The structural latent representation 108 can undergo denoising and / or modification by the structural latent denoising diffusion model 112. The latent representation of a mesh element (e.g., such as a point) is referred to as a latent mesh element representation 110 (e.g., which can encode aspects of the shape of the input 3D point cloud 102). The latent mesh element representation can undergo denoising and / or modification by the latent mesh element denoising diffusion model 114. The denoised / modified structural latent representation 116 can be provided to the final decoder module 120 (e.g., implemented as a PVCNN). The latent mesh element representation 118 can also be provided to the final decoder module 120. The final decoder module can reconstruct the generated geometry 122 (e.g., or other 3D oral care representation), which is a modified form of the input 3D oral care representation 102 or a completely new 3D oral care representation with aspects specified by the oral care parameter 100. In the case where the input 3D oral care representation is embodied by 3D point cloud points, the generated 3D oral care representation 122 can be embodied by 3D point cloud points. Such a point cloud can undergo an optional subsequent processing step 124 to generate a mesh, for example, using Shape as Points (SAP) surface reconstruction (e.g., as Figure 1 shown). If SAP is applied, the output 3D mesh 126 is produced. In other cases, the input 102 and the reconstructed 122 3D oral care representation can be embodied as voxels or other mesh elements described herein.

[0217] Figure 1 A method using a fully trained GGDM is shown. The GGDM has been trained to modify an input 3D oral care representation according to (optional) variables (e.g., as text instructions or according to the oral care parameters described herein).

[0218] In some cases, transfer learning can be used to further train an initially trained diffusion model or text-to-image (TTI) system (e.g., such as DALL-E 2, MidJourney, or Stable Diffusion) that generates 2D representations (such as 2D images) on 3D oral care representations. For example, the denoising ML model 210 can first be trained on 2D data and then further trained on 3D data (via transfer learning). In some embodiments, such a transfer learning model can be trained to generate 3D oral care representations. In some cases, the Stable Diffusion model can operate on the latent representations of the trial 3D oral care representations 216 to modify those 3D oral care representations 216.

[0219] In some non-limiting embodiments, the denoising diffusion probability model of the present disclosure can operate on a 3D representation of a tooth (e.g., to generate a dental restoration design). In an example of restoration design generation, a 3D representation of the pre-restoration tooth 102 (e.g., described by a 3D point cloud, 3D mesh, or voxelized representation) can be provided as an input and processed by a denoising diffusion probability neural network in a manner that preserves information about the 3D shape and / or 3D structure of the pre-restoration tooth. In some cases, the pre-restoration tooth can undergo optional mesh element feature vector generation (128) and optionally be encoded into a latent form (106) that preserves information about the 3D quality of the pre-restoration tooth. The resulting latent mesh element representation 110 can represent the 3D mesh elements (e.g., 3D points, voxels, edges, faces, or vertices) of the pre-restoration tooth in latent form. The latent mesh element representation 110 can be provided to a latent mesh element denoising diffusion model 114 that can generate a modified latent mesh element representation 118. In some embodiments, the pre-restoration tooth can undergo optional mesh element vector generation (130) and optionally be encoded into a latent form (104) that describes aspects of the 3D structure of the pre-restoration tooth in latent form 108. The latent representation 108 can be provided to a structural latent denoising diffusion model 112 that can generate a modified structural latent representation 116. Either or both of the modified structural latent representation 116 and the modified latent mesh element representation 118 can be provided to a decoder 120 that can reconstruct these representations into a reconstructed 3D oral care representation 122 (e.g., a post-restoration tooth design). When the reconstructed 3D oral care representation includes a point cloud, the point cloud can subsequently undergo surface reconstruction (124) to generate a 3D mesh 126. The oral care variable 100 can optionally be provided to either or both of the structural latent denoising diffusion model 112 and the latent mesh element denoising diffusion model 114 to customize the output of those models and configure those models to generate outputs suitable for use in generating oral care appliances. The oral care variable can include oral care parameters or oral care metrics (e.g., restoration design metrics such as "tooth morphology" or "left-right symmetry and / or ratio", and other metrics described herein).

[0220] Such latent forms 110 and 108 can reduce the data size of the pre-restoration tooth, resulting in increased data precision while still retaining rich information about the shape and / or structure of the pre-restoration tooth. Neural networks can be more easily trained on input data with a smaller data footprint (e.g., requiring fewer neural network parameters or weights to encode the solution), provided that the data is information-rich. The reconstruction autoencoder is specifically configured to generate information-rich latent representations, as shown by its ability to reconstruct a close approximation of the input data. The ground truth post-restoration 3D representation of the tooth can also be processed entirely in three dimensions. The Geometry Generation Diffusion Model (GGDM) can be trained at least in part by computing a loss that quantifies the difference between the generated post-restoration tooth design (e.g., a design predicted using a denoising diffusion model as described herein) and the ground truth post-restoration tooth (e.g., which can be provided in the training data for a given patient case). As described herein, such loss computation can be performed entirely in 3D, for example, using reconstruction loss computation. The techniques described herein can generate tooth designs for crowns, veneers, bridges, or for other types of restorations, or tooth designs used in designing dental restoration appliances or fixture models, and so on.

[0221] The denoising diffusion model (DDM) (e.g., such as the DDM shown in Figure 2 ) can involve one or more Markov chains. In some cases, a Markov chain can be defined as a sequence of random events, where each time point depends on the previous time point. In some cases, the transition distribution on the Markov chain forward process can be modulated by low-level Gaussian noise. Figure 2 A denoising diffusion model trained on structural latent data is shown. Figure 2 The forward pass 206 of generates training data for the denoising diffusion model (e.g., in some embodiments, it can be trained on the latent representation of grid elements). The training data can be used to train the denoising ML model 210. Figure 2 The fully trained denoising ML model 210 shown can be used by the GGDM (e.g., where the model is trained on the latent representation of the tooth, such as for restoration design generation), by a diffusion setup model (e.g., where the model is trained on the latent representation of the transformation of the tooth or other types of 3D oral care representations), or by other types of diffusion models to generate an output based on a training dataset of 3D oral care representations.

[0222] The DDM can be trained to approximate the data distribution q(x) over the latent variable Y (e.g., where the data distribution can describe a 3D oral care representation). The forward and backward processes of the denoising diffusion model can operate at many time points t, e.g., T = 1000. Any one of the following three operations may involve the denoising diffusion model: a noise scheduler for implementing the forward process (e.g., which may gradually introduce more noise of a specific distribution to each successive time point, such as Gaussian noise, such as isotropic Gaussian noise with zero mean and variance in all directions), a neural network for implementing the backward process (e.g., a U-Net, VAE, or other neural network described herein) (e.g., which may recover the initial distribution of the data provided to the start of the forward process at many time points), and / or an optional time point encoding method (e.g., using positional embeddings). The neural network can be trained at least in part using a loss function based on the L1 or L2 norm or KL divergence (and others described herein). In some cases, the KL divergence loss can drive the minimization of the difference between the observed distribution and the corresponding generated distribution.

[0223] In some cases, the forward process can include further encoding of the received data (e.g., latent data), and in some cases, the backward (or reverse) process can include decoding (or partial decoding) of the data processed by the forward process. In some cases, the backward process can be executed starting from random noise, and a sequence of data structures with gradually decreasing noise can be generated.

[0224] At the last time point t = T, in some cases, after adding noise (e.g., Gaussian noise) using the forward transition module, the distribution of the latent variable Y can be a normal distribution. In some cases, in an example of a 3D representation of oral care data, the forward process can introduce noise into the 3D oral care representation by perturbing one or more mesh elements of the 3D oral care representation. For example, the positions of point cloud points or voxels may be perturbed. The mesh elements may undergo perturbations that may change the values of the measured mesh element features (e.g., the length of an edge, the area of a face, the position of a point or voxel, the count of incident edges of a vertex, etc.). In some specific implementations, the transition module of the forward process can be described as (where M t is a variance distribution such that q(y t ) ~= p(y t ), and I is the identity matrix): q(y t |y t-1 ) = N(y t ; sqrt(1 - M t )y t-1 , M t I)

[0225] In some specific implementations, Figure 2 the reverse pass in Figure 2 can be implemented using a neural network such as U-Net. In some cases, the transformation module for this reverse pass can be described by the following product from time t = 1 to t = T: p(y 0:T ) = p(y T ) Π p W (y t-1 |y t ), where p W is Gaussian.

[0226] The U-Net can have a down phase (where the resolution of the input is successively reduced using convolutional and downsampling pooling layers until the coarsest resolution level, which can reveal increasingly global features of the input) and an up phase (which receives the results of the down phase and restores the resolution in a series of complementary deconvolution and unpooling phases). Skip (or residual) connections can connect corresponding phases of the down phase and the up phase. Attention modules can be integrated into the U-Net. Batch normalization layers can be integrated into the U-Net.

[0227] Figure 2 An example DDM for setting predictions is shown. An example of data processed (within a time point t) by a denoising diffusion model for 3D oral care representation (e.g., for denoising a latent representation such as a structural latent or latent grid element of a 3D oral care representation). The structural latent or latent grid element can be described by one or more latent vectors.

[0228] A denoising diffusion neural network (e.g., U-Net) can generate a target 3D oral care representation (e.g., an appliance component, a dental restoration design, a fixture model component, such as a trim line, etc.) by iteratively denoising a set of grid elements (e.g., points of a 3D point cloud) until a 3D oral care representation is generated that meets the specifications of the oral care arguments provided to the model. Oral care arguments can include oral care parameters as disclosed herein, or other real-valued, text-based, or categorical inputs that specify expected aspects of one or more target 3D oral care representations to be generated. In some cases, oral care arguments can include oral care metrics that can describe expected aspects of one or more 3D oral care representations to be generated. In some cases, a text encoder can encode a set of natural language instructions from a clinician (e.g., generate a text embedding). The text string can include tokens. In some embodiments, the encoder used to generate the text embedding can apply average pooling or max pooling between the token vectors. In some cases, a transformer (e.g., BERT or Siamese BERT) can be trained to extract embeddings of text used in digital oral care (e.g., by training the transformer on examples of clinical text such as those given below). In some cases, such a model used to generate text embeddings can be trained using transfer learning (e.g., initially trained on another text corpus and then receiving further training on text related to digital oral care). Some text embeddings can encode text at the word level. Some text embeddings can encode text at the token level. In some embodiments, the transformer used to generate the text embedding can be trained at least in part using a loss calculation that compares the predicted output to the ground truth output (e.g., softmax loss, multi-negative example ranking loss, MSE margin loss, cross-entropy loss, etc.). In some cases, non-text oral care arguments (such as real-valued or categorical values) can be converted to text and subsequently embedded using the techniques described herein. In some embodiments, the oral care argument 202 can include natural language text instructions such as the following and can be (optionally) encoded into a latent representation by a latent encoding module 212 (e.g., which can include a text transformer as described herein, as well as other architectures described herein). The following are examples of natural language instructions that can be provided by a clinician to the generation model described herein to describe the expected results of 3D oral care representation generation using the denoising diffusion probability model of the present disclosure: 1. "Generate settings for Class I molars and canines with 2mm overbite and add 2mm of extension 5-5 / 5-5." 2. "Generate settings to conform to the proclination and extension and leave 0.5mm of space for U2-2 for future restoration." 3. "Generate bracket settings to be completed with 2mm overbite and 2mm overjet, apply Class II elastics and apply an L2 - 2.5mm incision to the ideal state." 4. "Adjust the trim line to conform to the gingival margin on all teeth except teeth #8, 9, where the trim line should be 1mm apical to the gingival margin." 5. "Adjust the settings to include 0.3mm IPR for L4 - 4, retract to close the space for increased overjet." 6. "Adjust the settings, no second molar movement, rotate the first upper molar mesially outwards for Class I, reduce to a reverse curve of Spee 2mm, advance the mandible to Class I canines with elastics." 7. "Generate settings to intrude posterior teeth to allow a 2mm overbite with spin, apply lower IPR as needed for overjet, expand and tip forward to achieve space alignment." 8. "Generate settings to upright and then expand the premolars to form a wide arch form, close the midline space by mesial translation, intrude the lower 2 - 2.2mm." 9. "Generate settings, expand the arch form, apply the distal crown tip to the lower canine until vertical, then apply the mesial root tip to obtain a final 10 - degree distal tip deviating from vertical." 10. "Generate settings, extract tooth #24, retract and close the space to +1mm overjet and 2mm overbite, tip the upper 2 - 2 forward, with a 0.5mm space distal to teeth #8, 9." 11. "Generate a restoration design (alternatively, a veneer design) to close the diastema between teeth #8 - 9 by uniformly increasing the width on the mesial surfaces of the two teeth." 12. "Generate a restoration design (alternatively, a veneer design), add length to teeth #6 and #11 so that they have the same length as teeth #8 and #9. Add length to the lateral incisors so that they are 1 / 2mm shorter than the centrals." 13. "Generate a restoration design (alternatively, a veneer design), ensure that the incisal edges of teeth #6 - 11 form a uniform semi - circle with the incisal edges of the posterior teeth (when viewed from the incisal view)." 14. "Generate a restoration design (alternatively, a veneer design), for teeth #6 - 11, ensure that tooth #6 is symmetric with #11, #7 is symmetric with #10, and teeth #8 - 9 are symmetric." 15. "Generate a restoration design (alternatively, a veneer design), close the diastema between the lateral incisors and the central incisors by making the width of the central incisors 70% of the current length. Any remaining space should be added to the mesial surface of the lateral incisors." 16. "Generate a porcelain crown on tooth #12, replicating the shape of the first premolar, having a tight proximal contact and having two occlusal points with the lower teeth. The porcelain crown should have a facial contour that blends from the canine to the first premolar to the second premolar." 17. "Generate a restoration design (alternatively, a veneer design) for the rotated tooth #7, which increases the facial volume such that the width and shape of the tooth are similar to #10." 18. "Generate a restoration design (alternatively, a veneer design) where the facial contours and positions of teeth #6, 7, 8, 10, 11 match the facial positions of tooth #9." 19. "Generate a (zirconia) veneer design where the final color is the shade recommended by the dentist." 20. "Generate a (zirconia) design where the edges of the veneer are smooth and fit evenly into the edges formed on the tooth." 21. "Generate a (zirconia) veneer design where the veneer shape and contour match the same tooth in the opposite side of the dental arch, ensuring that all proximal contacts are closed." 22. "Generate a custom crown for the left upper central incisor to be implanted. The crown shape should take into account the shapes of adjacent teeth and the space between adjacent teeth should not exceed xmm (e.g.: 0.1mm)." 23. "Generate the missing right upper canine and the missing left upper first molar such that they are placed on the dental arch and the space / conflict with adjacent teeth does not exceed xmm (e.g.: 0.05mm)."

[0229] Denoising diffusion probability models are a class of deep generative neural networks that can be trained to generate transformations (e.g., that can be used to modify the position or orientation of a 3D representation (such as a tooth) in 3D space), generate 2D images (e.g., heatmaps or color images), generate 3D representations (e.g., such as point clouds, 3D meshes, or voxelized representations), or other 3D oral care representations as described herein. The denoising diffusion model can be trained to generate a transformation that positions the teeth in the dental arch into a pose suitable for orthodontic treatment (e.g., an intermediate stage or a final setting). The denoising diffusion model can be implemented using one or more encoders, one or more MLPs, one or more autoencoders, one or more U-Nets, one or more transformers (e.g., a 3D SWIN transformer encoder or a 3D SWIN transformer decoder), one or more pyramid encoder-decoders, and other machine learning models. In some embodiments, the denoising diffusion model (e.g., such as for set prediction) can take as input one or more oral care variables and also one or more 3D representations of oral care data, such as teeth (e.g., a full dental arch of segmented teeth in a malocclusion pose), appliance components, or fixture model components. The input teeth may be in a malocclusion pose. In some cases, the malocclusion transformation of the teeth can be provided to the denoising diffusion model. Non-limiting oral care variables can include a doctor's treatment plan (including at least: a set of zero or more protocol parameters, zero or more doctor preferences, and zero or more text samples that describe the nature of the expected oral care treatment (such as a final setting or prosthetic design generation)). Oral care variables can include real-valued, categorical values, natural language instructions, etc. In some embodiments, the denoising diffusion model can be trained to generate a setting that meets the specifications dictated by the protocol parameters and / or text description, such as a final setting. A denoising diffusion model for generating orthodontic settings can be referred to as a diffusion setting neural network or a diffusion setting model. A text-conditioned diffusion model can use a neural network to reformat and / or reduce the dimensionality of the text, e.g., to generate a latent encoding or latent embedding of the text (e.g., using a transformer or encoder to generate a text embedding).

[0230] The denoising diffusion model may include at least one of a forward pass and a reverse pass. The forward pass of the diffusion model may generate training data by iteratively adding noise (e.g., Gaussian noise) to a received 3D oral care representation (e.g., a point cloud representation of a tooth transformation or a tooth restoration design). In deployment, the reverse pass of the diffusion model may further operate through an iterative denoising process (e.g., such as using a U-Net trained for this purpose or other models described herein), which iteratively removes noise from the received 3D oral care representation (e.g., a point cloud representation of a tooth transformation or a tooth restoration design). Such tooth transformations may define the pose of teeth in one or more dental arches. The diffusion setup model may generate 3D oral care representations, such as orthodontic setups (e.g., for final setup or intermediate grading). During training, the DDM (e.g., diffusion setup or other setups described herein) may build a training data set by incrementally adding noise to a latent vector TA generated by the latent encoding module 214. In deployment, the DDM (e.g., diffusion setup or other setups described herein) may iteratively denoise one or more latent vectors TB, e.g., which may be generated by the latent encoding module 218. The denoising may be at least partially customized by the oral care variable 202, which may undergo optional encoding (212) to generate a latent vector TC. The latent vectors TA and TB contain dimension-reduced information about one or more 3D oral care representations (e.g., tooth mesh information and / or tooth transformation information). When the latent vector TA or TB corresponds to an encoded tooth transformation, in some embodiments, TA or TB may be conditioned on (or combined with) the latent vector A of one or more 3D representations of the teeth. The diffusion model receiving one or more 3D representations of teeth may be trained to generate tooth restoration designs (e.g., such as using GGDM, which may also be trained to generate other types of 3D oral care representations, such as appliance components, jig model components, transformations, or trim lines). These latent vectors TA or TB may also be conditioned on (or combined with) one or more protocol parameters K and / or one or more doctor preferences L. For example, such conditioning may be achieved by concatenating the latent vector TA or TB with K, L, or any other oral care variable described in this disclosure (such as M, N, O, R, S, P, Q, U, V).

[0231] Inside the diffusion model, a latent vector TA (e.g., which may have been generated using an encoder) may undergo multiple noise iterations during the forward pass. A series of increasingly noisy versions of TA can be used to at least partially train a denoising diffusion neural network (e.g., such as an autoencoder, a U-Net, or other neural networks described herein). The latent vector TB can start as Gaussian noise at the beginning of the backward pass. Through several iterations of the diffusion model backward pass (e.g., using a trained U-Net, autoencoder, or other model described herein), TA can evolve into a form that can be reconstructed (e.g., using a decoder) into a 3D oral care representation that can be used in oral care appliance generation (e.g., one or more transformations of placing one or more teeth into a setup configuration, such as a final setup or an intermediate stage).

[0232] In some embodiments, the denoising ML model 210 can use a U-Net architecture with ResNet blocks and self-attention layers as part of the backward pass. A ResNet block refers to a residual block where the activation of one layer in a neural network is directly forwarded to a subsequent deeper layer, which has the advantage of being able to train deeper networks. The self-attention layer utilizes an attention mechanism that makes different parts of a sequence (i.e., a sequence at an intermediate stage, etc.) relevant in order to compute a representation of that same sequence. Self-attention is beneficial for intermediate hierarchical prediction because as the diffusion model iterates, information about other stages in the sequence can be utilized to update or refine each stage of the sequence. The denoising ML model 210 (e.g., for setup prediction or 3D representation generation) can be trained by gradient descent and / or backpropagation. In some embodiments, the losses described elsewhere in this disclosure can be used to at least partially train the setup diffusion model. The L1, L2, MSE, or other losses described herein can be used to at least partially train the denoising ML model 210.

[0233] Figure 3 An example of a denoising diffusion model for setup prediction is shown. Figure 3 Shows how a denoising ML model 302 can be trained on a sequence of increasingly noisy copies of patient case data (e.g., orthodontic setups). In some cases, the noise can be Gaussian noise. Figure 3A Markov chain 300 is shown that includes 34 time points (although other Markov chain sizes are possible). An orthodontic setting 304 (e.g., a malocclusion setting, an intermediate setting, or a final setting) can be provided to the Markov chain 300. In some embodiments, the orthodontic setting 304 can be encoded (310) into a latent form before being provided to the Markov chain 300. Although the Markov chain 300 has a length of 34, other lengths are possible. In some cases, noise can be added to the patient data starting from T0 and continuing for multiple time steps, resulting in a training data set with a gradually increasing amount of noise (e.g., setting transformations with increasing amounts of noise). Each time point from T0 to T33 can include a transformation (e.g., one transformation per tooth in the dental arch) and / or a 3D representation of the teeth. Data from the time points in the Markov chain (e.g., tooth transformations) can be used to train a denoising ML model 302 (e.g., an encoder-decoder network, such as a pyramid encoder-decoder, a 3D SWIN transformer, an autoencoder, or a U-Net) at the center of the reverse pass of a diffusion model. In some embodiments, an optional oral care variable 306 (e.g., an oral care parameter or an oral care metric) can be provided to the denoising ML model 302. In some embodiments, the optional oral care variable can be encoded (308) into a latent form. Training can be performed through multiple iterations. The transformation can include a 4×4 affine transformation, but other dimensions are also compatible with the diffusion model-based techniques of the present disclosure. For example, the transformation can include one or more translation vectors, one or more Euler angles, and / or one or more quaternions. The "t" input is a time representation that can control which time point is sampled during a given time step of diffusion model training (e.g., training of a denoising neural network). The input "t" can be used for positional encoding. In some embodiments, the time point t can correspond to a stage in an orthodontic treatment. The use of positional encoding, similar to its use in a transformer, can enable the model to know the relative position of a stage in the treatment plan sequence. In this way, the model can learn how to perform orthodontic treatment from stage to stage, thus leveraging what has been learned from previous orthodontic treatment examples in the training data set by the diffusion model.

[0234] Figure 4An example of a mesh element labeling model is illustrated that is based on a denoising diffusion probability model (DDPM) and can be used to implement either 3D mesh segmentation (e.g., semantic segmentation) and a 3D mesh cleaning pipeline. The DDPM for 3D representation segmentation can include at least one of a forward pass 206 (e.g., which can involve many steps of iteratively adding more noise to the input 3D representation, which can be used in training an ML model to perform the reverse process) and a reverse pass 208 (e.g., which can use an ML model such as a neural network to iteratively denoise the noisy representation, resulting in a segmentation of the input 3D representation). In some embodiments, the DDPM for 3D representation segmentation can approximate aspects of a Markov process. Figure 4 The HNNFEM 406 described in Figure 4 can be trained to perform the reverse pass 208. The dental arch data 400 can optionally be transformed (402) into a list of mesh elements. The mesh elements can undergo an optional mesh element feature vector calculation (404). The mesh element features described herein can be provided to the HNNFEM 406, such as mesh element features including: edge midpoints, edge curvature, signed dihedral angles, edge lengths, edge normals, or other features described herein. The DDPM can begin operating on a noisy representation (e.g., a noisy representation of mesh element labels) and can iteratively denoise the representation to produce a denoised representation that can include one or more mesh element labels or object masks (alternatively, the noisy representation can be directly denoised into one or more segmented 3D representations, such as a 3D point cloud or a 3D mesh). Formulas that formalize the forward pass and the reverse pass are described herein.

[0235] In some specific implementations, the DDPM for 3D representation segmentation may include the following steps: iteratively add noise (e.g., Gaussian noise) to the input 3D oral care representation to generate a series of increasingly noisy versions of the input 3D oral care representation that can subsequently be used in a denoising neural network for training the reverse process. For each 3D representation in a series of increasingly noisy 3D representations of oral care data: 1) Use HNNFEM 406 to extract hierarchical neural network features 408 of the feature map. 2) Collect the grid element-level representation 410 by upsampling the feature maps from HNNFEM and concatenating those upsampled feature maps into a grid element representation vector 410. 4) The grid element representation vector 410 from step #3 can be used to train one or more ML models for grid element labeling (e.g., train an ensemble of neural networks 412 for grid element labeling). One or more ML models can vote (414) to determine the final output of the grid element label 416. One or more ML models can generate labels for one or more grid element features. In some specific implementations, the generated set of grid element labels may include one or more object masks. The object masks can label or otherwise indicate which grid elements belong to which object in the scene (e.g., which grid elements belong to the upper left central incisor, belong to the gingiva, or belong to the hardware elements in the dental arch grid from an intraoral scanner). Figure 4 Also shown is an input (e.g., dental arch) of the 3D oral care representation 400 to be segmented, which can be (optionally) arranged (402) into a vector of grid elements. One or more grid element features can be calculated (404) for each grid element.

[0236] HNNFEM 406 can include one or more U-Nets 502, one or more pyramid encoder-decoders from Figure 6 one or more 3D SWIN Transformer decoders, or other architectures described herein. In Figure 5 a grid element of the 3D representation 500 can be provided to the U-Net 502, which can use a U-shaped structure of convolutional 504 / transposed convolutional 512 and / or pooling 506 / transposed pooling operations 510 to extract hierarchical features. The most globalized features are extracted at the lowest resolution (508). A vector 414 or 514 containing hierarchical neural network features can be assembled into a grid element representation vector 416 and provided to an ensemble of ML models 418 (e.g., for resolving grid element labels). Voting (420) can be performed to resolve ambiguities between grid element labels predicted by the ensemble of ML models. The final predicted grid element label 422 is sent to the output.

[0237] In some specific implementations (e.g., when the pre-segmented dental arch is segmented into teeth and gums), the training data may include at least a 3D mesh of the dental arch and a set of mesh element labels (or object masks) indicating which tooth each mesh element belongs to. This set of mesh element labels (or masks) may undergo iterative noise addition, such as in the Figure 2 forward pass shown. The resulting set of increasingly noisy mesh element labels can be used as training data to train a denoising diffusion ML model (e.g., such as a U-Net). Such a denoising diffusion ML model can be trained to take as input a set of noisy (or random) mesh element labels (or object masks) of an immediate 3D mesh (e.g., of a pre-segmented dental arch) and gradually denoise this set of mesh element labels (or object masks) over multiple iterations. This set of denoised labels (or object masks) can be output by the Figure 2 backward pass and used to segment the immediate mesh (or point cloud or other 3D representation). In some cases, this set of mesh element labels (or object masks) can be encoded using an encoder, which can provide a latent representation to the forward pass. After completing the backward pass, in deployment, a decoder can be used to reconstruct the denoised representation. The mesh element labels can be used to segment the teeth of the pre-segmented dental arch. In some cases, the mesh element labels can be used to mark mesh elements as part of a mesh cleaning operation (e.g., marking mesh elements for removal, smoothing, scaling, or other modifications based on the generation of artifacts used in digital oral care).

[0238] The systems of the present disclosure can train denoising diffusion models to segment 3D representations of oral care data, such as 3D meshes of a patient's dentition. Figure 4 It relates to 3D mesh segmentation and 3D mesh cleaning. These two techniques share an important property, namely the labeling of 3D mesh elements. In the case of 3D mesh segmentation (e.g., related to the segmentation of an oral care mesh such as a tooth), various mesh elements of the mesh can be labeled according to which part of the dental anatomy they belong to (e.g., gum, upper right central incisor, lower left second bicuspid). Dental mesh segmentation can label elements according to their membership in various teeth, as specified by one or more of the dental notation systems mentioned herein (e.g., Palmer). In some specific implementations, for facio-lingual segmentation, each mesh element can be labeled according to its membership in the facial side of the dental arch or the lingual side of the dental arch. Other specific implementations of dental anatomy segmentation are possible. In some specific implementations, an oral care variable 202 can be provided to the denoising diffusion model for segmentation or mesh cleaning as described herein, such as a downsampling factor (or downsampling percentage). The downsampling factor can control the degree to which the pre-segmented mesh is downsampled before mesh element labeling occurs.

[0239] A pre-segmentation mesh (e.g., an arch generated by an intraoral scanner or a CT scanner) can be received by a segmentation system. Feature vectors for mesh elements can be computed, one or more vectors per mesh element. One or more mesh element features from elsewhere in the present disclosure can be used to form the feature vectors of the mesh elements. Mesh elements can include edges, faces, vertices, or voxels (or any combination thereof).

[0240] In some embodiments, a list of mesh element feature vectors can be provided to a neural network that refines the mesh element vectors to extract local and global features from those feature vectors, such as Figure 4 the U-Net architecture shown. The U-Net within the HNNFEM can include 3D mesh convolutions and / or subsequent 3D mesh pooling operations that can be used to reduce the resolution of the mesh and extract neural network features at increasingly global scales. After a series of such operations, after the most global neural network features have been extracted, there may be a series of 3D mesh unpooling and 3D mesh deconvolution operations that return the mesh to the initial scale. After outputs have been generated at various levels of the U-Net architecture, there may be upsampling and concatenation operations. The HNNFEM can generate one or more 3D mesh element-level feature vectors that may have been upsampled to the initial mesh resolution. These feature vectors can be concatenated and provided as input to one or more ML models for mesh element classification or labeling. Any supervised ML model disclosed elsewhere in the present disclosure can be used for such classification, such as an SVM or logistic regression. In some examples, an ensemble of ML classifiers can take the mesh element representation vectors as input and generate mesh element classification labels. In some embodiments, an ensemble of fully connected neural networks can be used for such classification. In some cases, a multi-layer perceptron (MLP) can be used for such classification, including linear layers and an associated ReLU activation function (e.g., which has the advantage that not all neurons fire simultaneously, enabling a more customized response to the input than an activation function that fires for every evaluation) and batch normalization operations. Other activation functions are possible, such as those mentioned elsewhere in the present disclosure. Each of the ML models in the ensemble can generate a predicted class label for each mesh element. A voting mechanism can then be employed to combine these results and output the final class label prediction for each 3D mesh element. Some embodiments can use transformers to assist in applying labels to the mesh elements.

[0241] In some cases, the specific implementation of mesh cleaning and mesh segmentation may differ in terms of the arrangement of ground truth mesh element labels. In the case of tooth segmentation, ground truth data can be given for each dental arch mesh to be segmented (i.e., each tooth in the dental arch has mesh elements labeled according to the tooth to which the mesh element belongs). Further, such a loss can be calculated for facio-lingual segmentation and / or for tooth-gum segmentation (as defined in U.S. Provisional Application No. US63 / 366490). The various loss functions of the present disclosure can be used to compare the mesh element labels of the predicted segmentation with the mesh element labels of the corresponding ground truth segmentation. Cross-entropy loss is an example of several candidate loss functions for this comparison. Other possible losses are disclosed elsewhere in the present disclosure. In the case of mesh cleaning, ground truth data can be given for each dental arch mesh to be cleaned. For example, in the case of removing foreign material, each mesh element corresponding to the foreign material in the dental arch can be so labeled, and each mesh element that does not include foreign material (i.e., that will be retained after mesh cleaning) can be so labeled. In the case of removing depressions, each mesh element corresponding to the depressions in the dental arch can be so labeled, and each mesh element that does not include depressions (i.e., that will be retained after mesh cleaning) can be so labeled.

[0242] After applying the segmentation labels, a process is performed to copy the mesh elements of a specific label into a new mesh (also referred to as mesh cutting), and the mesh is saved (e.g., saved to an electronic storage medium) for further processing. After completing the mesh cleaning labeling, a process is performed to remove each specified mesh element from the mesh (e.g., using classical mesh processing techniques), and then optionally a process can be performed to fill any holes that may have been created by this process (e.g., using techniques trained for mesh filling as described elsewhere in the present disclosure).

[0243] The mesh element labeling for mesh segmentation can also be done using an autoencoder trained for this purpose, such as a variational autoencoder, as described elsewhere in the present disclosure. More generally, an encoder-decoder network can be trained for mesh element labeling.

[0244] In some embodiments, the denoising diffusion techniques of the present disclosure can be trained to predict the dental arch morphology. The dental arch morphology can have a shape that describes various aspects of the shape of the dental arch. The dental arch morphology can include data structures such as one or more 3D meshes (e.g., meshes aligned with at least one coordinate axis of the local coordinate system of at least one tooth), 3D point clouds, 3D polylines, or a set of control points (e.g., control points defining a spline). For example, the Figure 2 denoising diffusion method described in

[0245] In some specific implementations, the denoising diffusion technique of the present disclosure can be trained to generate information related to IPR for use in setup prediction. The information related to IPR can include one or more of the following: 1) the IPR cutting surface of a specific tooth, 2) a measure of the IPR magnitude to be applied to one side of a specific tooth, 3) a designation of whether a specific tooth is to undergo IPR, including an indication of whether the IPR is to be applied to the mesial side or the distal side of the tooth, or 4) a designation of one or more stages that a specific tooth is to undergo IPR. The training data for such a model can be generated at least in part by iteratively adding noise to: 1) a representation of one or more IPR cutting surfaces from cases in the training dataset; 2) a representation of one or more teeth or other aspects of the patient's dentition; or 3) one or more tooth transformations. An ML model (e.g., a U-Net) can be trained to iteratively denoise an initial representation of one or more such examples of the training data (e.g., the IPR cutting surface can be at least partially randomly initialized or initialized with default values). When the backpropagation 208 is completed, one or more generated (or modified) data 220 related to IPR can be output. For example, the IPR cutting surface can be generated by the denoising diffusion technique described herein. In some specific implementations, the denoising diffusion model can be trained to determine whether a target tooth is suitable for IPR. For example, the denoising diffusion method described in Figure 2 can be used to generate information about IPR. For example, Figure 2 's method can predict orthodontic setups with optional associated IPR information. For each tooth, a setup transformation can be generated, as well as one or more associated items of IPR information (e.g., one or more IPR cutting surfaces, a flag indicating whether IPR is to be performed, a list of the status of performing IPR, etc.).

[0246] In some specific implementations, the techniques of the present disclosure can be trained to predict trimming lines. The trimming lines can define cutting paths to remove thermoformed trays from a fixture model (e.g., a tray for orthodontic appliance processing or an indirect bonding tray for delivering brackets to teeth for orthodontic treatment). One or more polylines (or meshes or point clouds) can be used to define the trimming lines. For example, the denoising diffusion method described in Figure 2 can be used to generate such polylines (or meshes or point clouds).

[0247] In some embodiments, the denoising diffusion technique of the present disclosure can be trained to predict one or more local orthogonal coordinate axes of a tooth (e.g., one or more of the X, Y, and Z orthogonal axes of a tooth). In other embodiments, the denoising diffusion technique of the present disclosure can be trained to predict one or more dental arch form coordinate axes. The position along a dental arch form coordinate axis can include a tuple [l, d, e] relative to a reference dental arch form spline S that approximates the shape of the dental arch. The rotation can include a tuple [a, b, g] that represents rotations of α, β, and γ. α describes the rotation about the l axis. β describes the rotation about the d axis. γ describes the rotation about the e axis. The full tuple that describes the position and rotation can include [l, d, e, a, b, g]. p is the point along S at arc length l. d is the distance between the tooth origin t and the reference dental arch form spline S. The tooth origin t is obtained by translating upward a distance “d” along the d axis and then translating a distance “e” along the e axis. The e axis is perpendicular to the d axis and the l axis and can be defined as the direction out of or into the page. e represents height. l represents the length across the dental arch form spline. d represents the distance from the dental arch form spline. This dental arch form coordinate system is described in U.S. Patent Application Publication No. US20210259808A1 by the same applicant, the entire content of which is incorporated herein by reference. The coordinate systems described herein can facilitate subsequent automation of a patient's oral care procedures, such as automated setting prediction. One or more transformations, vectors, or tuples can be used to define the coordinate system (or define a pose within the coordinate system). For example, the denoising diffusion method described in Figure 2 can be used to generate such transformations (or vectors or tuples).

[0248] Table 3 describes the input data and the generated data for several non-limiting examples of the generation techniques described herein. In some embodiments, the denoising diffusion model described herein can be trained to generate (or modify) the input data in Table 3 to obtain the generated data in Table 3.

[0249] The techniques of the present disclosure can be trained to generate (or modify) point clouds (e.g., where points can be described as 1D vectors, such as (x, y, z)), polylines (points connected in order by edges), 3D meshes (points connected by edges to form faces), splines (which can be calculated from a set of generated control points), sparse voxelized representations (which can be described as a set of points corresponding to the centroid of each voxel or some other landmark of the voxel, such as the boundary of the voxel), one or more transformations (which can take the form of one or more 1D vectors or one or more 2D matrices, such as a 4×4 matrix), etc. In some embodiments, a voxelized representation can be calculated from a 3D point cloud or a 3D mesh. In some embodiments, a 3D point cloud can be calculated from a voxelized representation. In some embodiments, a 3D mesh can be calculated from a 3D point cloud.

[0250] In some specific implementations, an automated setting prediction model (e.g., a denoising diffusion probability model) can be trained to generate settings with a customized spee curve (e.g., a spee curve that conforms to the expected outcome of patient treatment). Such a model can be trained on cohort patient case data. One or more oral care metrics can be calculated on each case to quantify or measure aspects of the spee curve of that case. During training, one or more of such metrics can be provided to the setting prediction model, e.g., to influence the model regarding the geometry and / or structure of the spee curve for each case. When deploying the setting prediction model, the same input path to the trained neural network can be configured with one or more values as instructions for the model regarding the expected spee curve. Such values can automatically generate settings with a spee curve that meets the aesthetic and / or medical treatment needs of a specific patient case.

[0251] In some specific implementations, the spee curve metric can measure the curvature of the occlusal or incisal surface of teeth on the left or right side of the dental arch relative to the occlusal plane. In some cases, the occlusal plane can be calculated as the surface that averages the incisal or occlusal surfaces of the teeth (for one or both dental arches). In some specific implementations, the curvature metric can be calculated along a normal vector (such as a vector perpendicular to the occlusal plane). In other specific implementations, the curvature metric can be calculated along the normal vector of another plane. In some specific implementations, the XY plane can be defined as corresponding to the occlusal plane. The orthogonal plane can be defined as the plane orthogonal to the occlusal plane that also passes through the spee curve segment, where the spee curve segment is defined by a first endpoint and a second endpoint, the first endpoint being a landmark point on the first tooth (e.g., canine) and the second endpoint being a landmark point on the last tooth on the same side of the dental arch. In some specific implementations, the landmark points can be located along the incisal edge of the tooth or on the cusp of the tooth. In some cases, the landmark points of the intermediate teeth (e.g., teeth located between the first tooth and the last tooth) on the left or right side of the dental arch can form a curved path, such as can be described by a polyline. The following is a non-limiting list of spee curve oral care metrics.

[0252] 1) Measure the perpendicular height between a line segment and a point. In other words, measure the distance between the line segment and the point along the z-axis. The line segment is defined by connecting the highest cusp of the last tooth (in the lower dental arch) and the cusp of the first tooth (in the lower dental arch) on that side. Given a subgroup of teeth between the first and last teeth, the point is defined by the highest cusp of the lowest tooth in that subgroup. In other words, the following 4 steps can be used to calculate the Spee curve metric. i) Line: Form a line between the highest cusp on the last tooth and the cusp of the first tooth. ii) Curve_Point_A: Given the set of teeth between the last and first teeth, find the highest point of the lowest tooth. iii) Curve_Point_B: Project Curve_Point_A onto the line to find the point on the line closest to Curve_Point_A (Curve_Point_B). iv) Spee curve: Find the height difference between Curve_Point_B and Curve_Point_A.

[0253] 2) Project one or more intermediate landmark points (e.g., points on the teeth that are between the first and last teeth on that side of the dental arch) and the Spee curve segment onto an orthogonal plane. Calculate the Spee curve metric by measuring the distance between the farthest projected intermediate point and the projected Spee curve segment. This yields a measure of the curvature of the dental arch relative to the orthogonal plane.

[0254] 3) Project one or more intermediate landmark points and the Spee curve segment onto the occlusal plane. Calculate the Spee curve on this plane by measuring the distance between the farthest projected intermediate point and the projected Spee curve segment. This yields a measure of the curvature of the dental arch relative to the occlusal plane.

[0255] 4) Skip the projection and calculate the distance and curvature in 3D space. Calculate the Spee curve by measuring the distance between the farthest intermediate point and the Spee curve segment. This yields a measure of the curvature of the dental arch in 3D space.

[0256] 5) Calculate the slope of the projected Spee curve segment on the occlusal plane.

[0257] 6) Calculate the slope of the projected Spee curve segment on the orthogonal plane.

[0258] The Spee curve metrics 5 and 6 can help the network reduce some more degrees of freedom when defining how the patient's dental arch curves at the back of the mouth.

[0259] The techniques described herein (e.g., denoising diffusion probability models) can be trained to generate a transformation that can place a patient's teeth into a pose suitable for use in an orthodontic setting (e.g., intermediate or final setting) according to a specification of oral care parameters that may be provided to the generative model in some embodiments. An optional oral care argument 202 can be provided to a denoising ML model 210 (e.g., a U-Net) that is trained to perform a reverse pass 220. The oral care argument can include oral care parameters as disclosed herein, or other real-valued, text-based, or categorical inputs that specify an expected aspect of one or more 3D oral care representations to be generated. In some cases, the oral care argument can include an oral care metric that can describe an expected aspect of one or more 3D oral care representations to be generated. The oral care argument is particularly applicable to the embodiments described herein. For example, the oral care argument can specify an expected design (e.g., including shape and / or structure) of a 3D oral care representation that can be generated (or modified) according to the techniques described herein. In short, embodiments that use the specific oral care arguments disclosed herein generate more accurate 3D oral care representations than embodiments that do not use specific oral care arguments. In some cases, a text encoder can encode a set of natural language instructions from a clinician (e.g., generate a text embedding). The text string can include tokens. In some embodiments, the encoder used to generate the text embedding can apply average pooling or max pooling between token vectors. In some cases, a transformer (e.g., BERT or Siamese BERT) can be trained to extract an embedding of text used in digital oral care (e.g., by training the transformer on examples of clinical text such as those given below). In some cases, such a model for generating text embeddings can be trained using transfer learning (e.g., initially trained on another text corpus and then receiving further training on text related to digital oral care). Some text embeddings can encode text at the word level. Some text embeddings can encode text at the token level. In some embodiments, the transformer used to generate the text embedding can be trained at least in part using a loss calculation that compares a predicted output to a ground truth output (e.g., softmax loss, multi-negative example ranking loss, MSE margin loss, cross-entropy loss, etc.). In some cases, non-text arguments (such as real-valued or categorical values) can be converted to text and subsequently embedded using the techniques described herein.Examples of natural language instructions that can be issued by a clinician to the generative model described herein are: "Generate settings to set as Class I molars and canines, 2mm overbite and add 2mm of expansion 5-5 / 5-5", "Generate settings to conform to proclination and expansion and leave 0.5mm of space U2-2 for future restoration", or "Adjust settings without second molar movement, rotate the upper first molar mesially outwards for Class I, reduce the height to a reverse curve of Spee 2mm, and advance the mandible to Class I canines with elastics".

[0260] In some specific implementations, the techniques of the present disclosure can extract local or global neural network features from 3D point clouds or other 3D representations (e.g., 3D point clouds describing aspects of a patient's dentition such as teeth or gums) using PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using PointNet or PointNet++ as a basis for training). In some specific implementations, the techniques of the present disclosure can extract local or global neural network features from 3D point clouds or other 3D representations using U-Net.

[0261] 3D oral care representations are described herein because 3-dimensional representations are current state of the art. However, 3D oral care representations are intended to be used in a non-limiting manner to cover any representation of 3 dimensions or higher order dimensions (e.g., 4D, 5D, etc.), and it should be understood that the techniques disclosed herein can be used to train machine learning models to operate on representations of higher order dimensions.

[0262] In some cases, the input data can include 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data related to splines (e.g., control points). The encoder-decoder structure can include one or more encoders, or one or more decoders. In some specific implementations, the encoder can take as input the mesh element feature vectors of one or more input mesh elements in the input mesh. The encoder is trained in a manner that generates a more accurate representation of the input data by processing the mesh element feature vectors. For example, the mesh element feature vectors can provide the encoder with more information about the shape and / or structure of the mesh, and thus the additional information provided allows the encoder to make more informed decisions and / or generate a more accurate latent representation of the mesh. Examples of encoder-decoder structures include U-Net, autoencoders, or transformers, etc. The representation generation module can include one or more encoder-decoder structures (or parts of encoder-decoder structures, such as individual encoders or individual decoders). The representation generation module can generate an information-rich (optionally dimension-reduced) representation of the input data that can be more easily consumed by other generative or discriminative machine learning models.

[0263] The U-Net may include an encoder followed by a decoder. The architecture of the U-Net may be similar to a U shape. The encoder may extract one or more global neural network features, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level compared to the most global level) from the input 3D representation. The output of each level from the encoder may be passed to the input of the corresponding level of the decoder (e.g., via skip connections). Similar to the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For example, the decoder may output a representation of the input data that may include global, intermediate, or local information about the input data. In some embodiments, the U-Net may generate an informative (optionally dimension-reduced) representation of the input data that may be more easily consumed by other generative or discriminative machine learning models.

[0264] An autoencoder may be configured to encode input data into a latent form. The autoencoder may train the encoder to reformulate the input data into a dimension-reduced latent form between the encoder and the decoder, and then train the decoder to reconstruct the input data from this latent form of the data. A reconstruction error may be computed to quantify the extent to which the reconstructed form of the data differs from the input data. In some embodiments, the latent form may be used as an informative dimension-reduced representation of the input data that may be more easily consumed by other generative or discriminative machine learning models. In most scenarios, the autoencoder may be trained on an input 3D representation, encode the 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close replica of the input 3D representation as an output.

[0265] The transformer can be trained to generate a representation of its input using self-attention at least in part. The transformer can encode long-range dependencies (e.g., encode relationships between a large number of inputs). The transformer can include an encoder or a decoder. In some embodiments, such an encoder can operate in a bidirectional manner or can operate a self-attention mechanism. In some embodiments, such a decoder can operate a masked self-attention mechanism, can operate a cross-attention mechanism, or can operate in an autoregressive manner. In some embodiments, the self-attention operation of the transformer described herein can be related to different positions or aspects of a single 3D oral care representation in order to compute a dimensionally reduced representation of the 3D oral care representation. In some embodiments, the cross-attention operation of the transformer described herein can mix or combine aspects of two (or more) different 3D oral care representations. In some embodiments, the autoregressive operation of the transformer described herein can consume previously generated aspects (e.g., previously generated points, point clouds, transforms, etc.) of a 3D oral care representation as additional inputs when generating a new or modified 3D oral care representation. In some embodiments, the transformer can generate a latent form of the input data that can be an information-rich dimensionally reduced representation of the input data and that can be more readily consumed by other generative or discriminative machine learning models.

[0266] In some embodiments, the encoder-decoder architecture can first be trained as an autoencoder. In deployment, one or more modifications can be made to the latent form of the input data. Then, the modified latent form can continue to be reconstructed by the decoder, resulting in a reconstructed form of the input data that is different from the input data in one or more desired aspects. Oral care variables such as oral care parameters or oral care metrics can be provided to the encoder, the decoder, or can be used to modify the latent form in order to influence the encoder-decoder architecture when generating a reconstructed form with desired characteristics (e.g., characteristics that may be different from those of the input data).

[0267] In some cases, federated learning can be used to train the techniques of the present disclosure. Federated learning can enable multiple remote clinicians to iteratively improve machine learning models (e.g., validation of 3D oral care representations, mesh segmentation, mesh cleaning, other techniques involving labeling mesh elements, coordinate system prediction, placement of non-organic objects on teeth, appliance component generation, dental restoration design generation, techniques for placing 3D oral care representations, setting prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, estimation of missing values), while protecting data privacy (e.g., clinical data may not need to be sent "over the network" to a third party). Data privacy is particularly important for clinical data protected by applicable laws. A clinician can receive a copy of the machine learning model, use a local machine learning program to further train the ML model using locally available data from the local clinic, and then send the updated ML model back to a central hub or a third party. The central hub or third party can integrate the updated ML models from multiple clinicians into a single updated ML model that benefits from learning from patient data recently collected at various clinical sites. In this way, a new ML model can be trained that benefits from additional and updated patient data (possibly from multiple clinical sites), while that patient data is never actually sent to a third party. In some cases, training on devices within a local clinic can be performed when the device is idle or otherwise during non-working hours (e.g., when patients are not being treated at the clinic). Devices in a clinical environment for collecting data and / or training an ML model for the techniques described herein can include intraoral scanners, CT scanners, X-ray machines, laptop computers, servers, desktop computers, or handheld devices (such as a smartphone with image collection capabilities). In addition to federated learning techniques, in some embodiments, contrastive learning can be used to at least partially train the ML models described herein. In some cases, contrastive learning can augment samples in a training dataset to emphasize differences between different class samples and / or increase the similarity of samples within the same class.

[0268] In some cases, a local coordinate system for 3D oral care representations (such as teeth) can be described by one or more transformations (e.g., an affine transformation matrix, a translation vector, or a quaternion). The systems of the present disclosure can be trained to predict coordinate systems using past cohort patient case data. The past patient data can include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems. Machine information models (such as U-Nets, encoders, autoencoders, pyramid encoder-decoders, transformers, or convolutional layers and / or pooling layers) can be trained for coordinate system prediction. Representation learning can determine a representation of a tooth (e.g., encoding a mesh or point cloud into a latent representation, e.g., using a U-Net, an encoder, a transformer, or convolutional layers and / or pooling layers, etc.), and then predict a transformation for that representation (e.g., using a trained multi-layer perceptron, transformer, encoder, transformer, etc.), which defines a local coordinate system for that representation (e.g., including one or more coordinate axes). In the case of predicting a coordinate system for a tooth mesh, the mesh convolution techniques described herein can utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth. Reinforcement learning techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth.

[0269] Machine information models such as U-Net, encoder, autoencoder, pyramid encoder-decoder, transformer, or convolutional layers and / or pooling layers can be trained as part of a method for hardware (or appliance component) placement. Representation learning can train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder or using blocks of U-Net, encoder, transformer, convolutional layers, and / or pooling layers). The representation can include a dimensionally reduced form and / or an information-rich form of the input 3D oral care representation. In some embodiments, generation of the representation can be assisted by computing a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some embodiments, a representation can be computed for a hardware element (or appliance component). Such a representation is adapted to be provided to a second module that can perform a generation task such as transform prediction (e.g., a transform for placing a 3D oral care representation relative to another 3D oral care representation, such as a representation for placing a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such a transform can include an affine transform matrix, a translation vector, or a quaternion, etc. Machine learning models that can be trained to predict a transform for placing a hardware element (or appliance component) relative to an element of a patient dentition include MLP, transformer, encoder, etc. The system of the present disclosure can be trained for 3D oral care appliance placement using 3D oral care appliance placement using past cohort patient case data. Past patient data can include at least: one or more ground truth transforms and one or more 3D oral care representations (such as tooth meshes or other elements of a patient dentition). In the case where U-Net (and other neural networks) are trained to generate a representation of a tooth mesh, the mesh convolution and / or mesh pooling techniques described herein utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be performed for hardware or appliance component placement. Reinforcement learning techniques can be performed for hardware or appliance component placement.

[0270] In some embodiments, the denoising diffusion techniques described herein (e.g., Figure 2The method in ) can be trained to fill in the missing aspects of the 3D representation of a patient's dentition (e.g., segmented teeth or digital fixture models or digital fixture model components). In some examples, a 3D scan of a dental crown (or a fixture model or other 3D oral care representation described herein) may be missing some parts of one or more surfaces (e.g., due to occlusion by surrounding teeth during an intraoral scan). The denoising diffusion model can generate training data by iteratively degrading or adding noise to the 3D representation of the tooth (e.g., by removing patches or portions of the surface of the dental crown or root). A progressive sequence of noisy versions of the 3D representation of the tooth can be generated from the training dataset. Noise can be added to the point cloud / mesh / voxelized representation of the tooth by: 1) adding a random translation to one or more mesh elements; 2) applying a small random scale or distortion to aspects of the tooth; or 3) by removing mesh elements (e.g., creating holes); and other methods. Then, this progressive sequence of noisy versions of the tooth can be used to train a denoising diffusion neural network (e.g., a U-Net or other encoder-decoder architecture such as a pyramid encoder-decoder, a transformer, or an autoencoder) to reconstruct (e.g., in small steps) the noisy tooth. A set of noisy mesh elements (e.g., initialized from a tooth with missing data) can be provided to the denoising diffusion neural network, and the denoising diffusion neural network can be trained to iteratively remove noise and / or repair the tooth surface. After multiple incremental repair iterations, the repaired tooth (or other 3D oral care representation) can be output by the denoising diffusion model and used in subsequent digital oral care processes to generate oral care appliances.

[0271] The denoising diffusion techniques of the present disclosure can be used (e.g., Figure 2to generate or modify a jig model component. The digital jig model can include a 3D representation of a patient's dentition, to which optional jig model components are attached. The 3D representation generation techniques of the present disclosure (e.g., denoising diffusion techniques that generate a 3D point cloud, a 3D voxelized representation, or other representations) can be trained to generate (or modify) aspects of the digital jig model to include processed enhanced jig model components. The jig model components can include 3D representations (e.g., 3D point clouds, 3D meshes, or voxelized representations) of one or more of the following non-limiting items: 1) Interproximal sidebands - which can fill or smooth the space between teeth to ensure appliance removability. 2) Undercuts - which can be added to the jig model to remove overhangs that may interfere with the thermoforming of a plastic tray or to ensure appliance removability. 3) Bite blocks - occlusal features on molars or premolars that are designed to support occlusal opening. 4) Bite ramps - lingual features on incisors and canines that are designed to support occlusal opening. 5) Interproximal reinforcements - structures on the outside of an oral care appliance (e.g., an orthodontic tray) that can extend from a first gingival edge of the appliance body on the labial side of the appliance body between a first tooth and a second tooth in an interproximal region to a second gingival edge of the appliance body on the lingual side of the appliance body. The effect of the interproximal reinforcement on the appliance body in the interproximal region can be harder than the labial and lingual surfaces of the first shell. This can allow the orthodontic appliance to grip the teeth on either side of the reinforcement more firmly. 6) Gingival ridges - structures that can extend along the gingival edge of a tooth in the mesial-distal direction to enhance the engagement between the orthodontic appliance and the given tooth. 7) Torque points - structures that can enhance the force delivered to a given tooth at a specified location. 8) Power ridge structures - structures that can enhance the force delivered to a given tooth at a specified location. 9) Pits - structures that can enhance the force delivered to a given tooth at a specified location. 10) Digital pontics - which can keep a space open or reserved in the dental arch for a partially erupted tooth, etc. In an orthodontic appliance, a physical pontic is a pocket that does not cover a tooth when the orthodontic appliance is installed on the teeth. The pocket can be filled with tooth-colored wax, silicone, or composite material to provide a more aesthetic appearance. 11) Power bars - blocks added in an edentulous space to provide strength and support to the tray. The power bar can fill a void. A abutment or healing cap can be blocked with the power bar. 12) Trim lines - digital paths along the digital jig model that can generally follow the contour of the gingiva (e.g., can be offset 1 or 2 mm in the gingival direction). After 3D printing, the trim line can define the path along which a clear tray-type orthodontic appliance can be cut or separated from the physical jig model. 13) Undercut fill - material added to the jig model to avoid forming a cavity between the contour height of the jig model and another boundary (e.g., the gingiva or a plane below the physical jig model after 3D printing).

[0272] The techniques of the present disclosure may also generate (or modify) other geometries that invade the gingival sulcus or enhance oral care appliances (e.g., orthodontic appliance trays). In some specific implementations, the techniques of the present disclosure may determine the position, size, and / or shape of a jig model component to produce a desired treatment result.

[0273] The techniques of...

Claims

1. A method for generating a data structure for an oral care treatment, the method comprising: Receiving, by one or more computer processors, one or more attributes that describe an expected output from a trained machine learning model; Generating, by the one or more computer processors, one or more noisy representations of the expected output; Denoising, by the one or more computer processors and using the trained machine learning model, the one or more noisy representations of the expected output; Generating, by the one or more computer processors, one or more denoised representations of the expected output; Automatically defining, by the one or more computer processors and using one or more generated denoised representations, one or more aspects of one or more digital oral care treatments.

2. The method according to claim 1, the method further comprising: Accessing a training data set; Generating, by the one or more computer processors, a refined training data set by continuously modifying one or more representations of the training data set; Training, by the one or more computer processors, an untrained machine learning model using the refined training data set; And Outputting the trained machine learning model based on the training.

3. The method according to claim 2, wherein the modification comprises adding noise to aspects of the one or more representations of the training data set.

4. The method according to claim 2, the method further comprising, before generating the one or more noisy representations, encoding, by the one or more computer processors, one or more aspects of one or more representations of the training data set into a latent form.

5. The method according to claim 1, wherein the one or more attributes comprise at least one of a real-valued, categorical value, or natural language text value.

6. The method according to claim 4, wherein the one or more attributes comprise at least one of an oral care metric or an oral care parameter.

7. The method according to claim 1, wherein the trained machine learning model comprises at least one neural network.

8. The method according to claim 7, wherein the at least one neural network comprises at least one of a U-Net, a transformer, or an autoencoder.

9. The method according to claim 1, wherein the method receives one or more 3D representations of a patient's dentition, the patient's dentition comprising at least one tooth.

10. The method according to claim 9, wherein the one or more denoised representations comprise at least one 3D oral care representation.

11. The method according to claim 10, wherein the one or more 3D representations of the patient's dentition are encoded into a latent form.

12. The method according to claim 10, wherein the one or more denoised representations comprise at least one label of a mesh element.

13. The method according to claim 12, the method further comprising using the at least one label to segment one or more 3D representations of oral care data.

14. The method according to claim 13, wherein the one or more 3D oral care representations comprise at least one representation of a patient's dentition.

15. The method according to claim 12, the method further comprising modifying one or more 3D representations of oral care data using the at least one tag.

16. The method according to claim 15, wherein the one or more 3D oral care representations include at least one representation of a patient's dentition.

17. The method according to claim 10, wherein the one or more denoised representations include at least one transformation to place at least one tooth of a patient's dentition in a set pose.

18. The method according to claim 17, wherein the set corresponds to a final set.

19. The method according to claim 10, wherein the 3D oral care representation includes one or more of the following: a jig model component, a tooth restoration design, an appliance component, a jig model component, and an arch form.

20. The method according to claim 10, wherein the one or more denoised representations include one or more coordinate systems.

Citation Information

Patent Citations

  • Method for automated generation of orthodontic treatment final setups

    US20210259808A1

  • Method for automated generation of orthodontic treatment final setups

    WO2020026117A1

  • Automated creation of tooth restoration dental appliances

    WO2020240351A1

  • Neural network-based generation and placement of tooth restoration dental appliances

    WO2021240290A1

  • System to generate staged orthodontic aligner treatment

    WO2021245480A1