Neural network techniques for appliance creation in digital oral care
Generating 3D oral care representations through encoder-decoder structure and transformer neural networks solves the problems of wasted computing resources and insufficient accuracy in the prior art, and achieves efficient and accurate dental and orthodontic treatments.
Patent Information
- Application Number
- CN202380085879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-25
- Filing Date
- 2023-12-14
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to accurately generate 3D oral care representations in dental or orthodontic treatment, especially dental restoration device components and orthodontic accessories, resulting in waste of computing resources and insufficient prediction accuracy.
Using encoder-decoder structure and transformer neural network, 3D oral care representations are generated or modified by training machine learning models, including dental restoration designs and orthodontic attachments, and using potential representation encoding and decoding techniques to improve data accuracy and computational efficiency.
It improves the accuracy and computing efficiency of 3D oral care representation, reduces the consumption of computing resources, and can quickly generate oral care devices in a clinical environment to meet clinical needs.
Smart Images

Figure CN120380469A_ABST
Abstract
Description
Related Literature
[0001] The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosure of each of the PCT applications with publication numbers WO2022123402A1, WO2021240290A1, WO2020240351A1, WO2021245480A1, and WO2020026117A1 is incorporated herein by reference. The entire disclosure of each of the following provisional U.S. patent applications is incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 366,514; and 63 / 264,914. Technical Field
[0002] The present disclosure relates to the configuration and training of machine learning models to improve the accuracy and data precision of 3D oral care representations to be used in dental or orthodontic procedures. Summary of the Invention
[0003] The present disclosure describes systems and techniques for training and using one or more machine learning models, such as neural networks, to generate 3D oral care representations. Neural network-based techniques are described for placing an oral care article relative to one or more 3D representations of teeth. The oral care article to be placed can include: a dental restoration appliance component, oral care hardware (e.g., tongue brackets, lip brackets, orthodontic attachments, occlusal ramps, etc.), and the like. Additionally, neural network-based techniques are described for generating the geometry and / or structure of an oral care article based at least in part on one or more 3D representations of teeth. Oral care articles that can be generated include: dental restoration appliance components, dental restoration tooth designs, crowns, veneers, etc. In a neural network that can be trained to generate 3D oral care representations, a transformer is an example of a model that can improve data accuracy. A transformer can be trained to automatically generate 3D oral care representations, such as 3D meshes or 3D point clouds. Additional examples of 3D oral care representations that can be generated include, but are not limited to: arch forms, clear tray aligner (CTA) trim lines, and appliance components (e.g., generated components such as those used in creating a dental restoration appliance). In some cases, a transformer can be trained to generate 3D polylines (e.g., for arch forms and CTA trim lines) or sets of control points (e.g., control points through which a spline can be fit for an arch form). Due to the transformer's ability to process long data sequences, the transformer can produce results that improve data accuracy and data precision compared to prior art for 3D mesh / 3D polyline generation, each of which may potentially include hundreds or even thousands of mesh elements. A neural network, such as a transformer, can be trained to predict an arch form (e.g., taking arch and tooth data as input). The arch form can be in the form of a surface, 3D mesh, 3D polyline, or a set of control points (e.g., for defining a spline). In some cases, such an arch form can be given as input to a setting prediction machine learning model, such as a setting prediction neural network).
[0004] The techniques of the present disclosure can train an encoder-decoder architecture to generate (or modify) 3D oral care representations suitable for oral care appliance generation (e.g., dental restoration designs, IPR cutting surfaces, appliance components, or others disclosed herein). The encoder-decoder architecture can include at least one encoder or at least one decoder. Non-limiting examples of encoder-decoder architectures include 3D U-Net, transformers, pyramid encoder-decoders, or autoencoders, etc. Non-limiting examples of autoencoders include variational autoencoders, regularized autoencoders, masked autoencoders, or capsule autoencoders. In some particular implementations, the generation techniques described herein can include aspects derived from denoising diffusion models (e.g., neural networks that can be trained to iteratively denoise one or more 3D oral care representations for use in appliance generation). In some particular implementations, the generation techniques described herein can train one or more neural networks to use mathematical operations associated with continuous normalizing flows (e.g., using neural networks that can be trained in one form and then inverted for use in inference).
[0005] The techniques of the present disclosure can be trained to generate a three-dimensional (3D) representation of oral care data that can be used in oral care processes. An input 3D representation of a patient's dentition may undergo a latent representation encoding (e.g., encoded into a latent representation having a lower dimensional order than the input data). A first machine learning (ML) module can be trained to perform the latent encoding. The first ML module can provide its output to a trained second ML module, which can include one or more transformer encoders or one or more transformer decoders, and the one or more transformer encoders or one or more transformer decoders can generate a second latent representation at least partially based on the first latent representation. Additionally, oral care arguments can be provided to the second ML module to customize the output of the second ML module. The second latent representation can be reconstructed into a 3D oral care representation using a decoder, and the reconstructed representation can be output for use in oral care appliance generation. In some embodiments, a trained latent representation modification module (LRMM) can be used to modify one or more aspects of the first latent space representation (e.g., in response to one or more oral care arguments). The modified first latent representation can be reconstructed into a representation that has been customized for use in treating a particular patient. The customized representation can be used for oral care appliance generation (e.g., prosthetic appliances, orthodontic appliances, etc.). The method can generate a representation of appliance components, a representation of an arch form, a representation of a patient's dentition with at least one generated fixture model component, a representation of at least one tooth in a pre-restoration state, a representation of at least one tooth in a post-restoration state, or other representations described herein. The input 3D representation of the patient's dentition can include one or more mesh elements. Mesh element features can be calculated for the mesh elements and subsequently provided to the first ML module (or the second ML module) to improve the accuracy of the resulting latent representation. In some embodiments, the first ML module can include one or more of the following: an encoder, a U-Net, a pyramid encoder-decoder, a 3D SWIN transformer, one or more convolutional layers, or one or more pooling layers. In some embodiments, either or both of the trained first ML module or the trained second ML module can be trained according to a transfer learning paradigm. Similarly, according to the transfer learning paradigm, either or both of the first ML module or the second ML module can be used to at least partially train one or more other ML models used in digital oral care. The method can be deployed in a clinical environment and can be performed near real-time during a session with a patient. In some embodiments, the trained first ML module can be configured to generate one or more hierarchical neural network features that can be based at least in part on one or more aspects of at least one of the shape or structure of the input 3D representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1Shows a transducer that can be configured to generate orthodontic setup transformations.
[0007] Figure 2 Shows a method for enhancing training data used in training a machine learning (ML) model of the present disclosure.
[0008] Figure 3 Shows a method for training a machine learning (ML) model to generate (or modify) a 3D representation of oral care data (e.g., an oral care mesh).
[0009] Figure 4 Shows a method for using a trained ML model to generate (or modify) a 3D representation of oral care data (e.g., an oral care mesh).
[0010] Figure 5 Shows a method for training an ML model to generate an arch form.
[0011] Figure 6 Shows a method for training a reconstruction autoencoder that can be trained to reconstruct an oral care mesh (e.g., an arch form, teeth, dental restoration design, etc.).
[0012] Figure 7 Shows a method for using a trained reconstruction autoencoder to reconstruct an oral care mesh (e.g., an arch form, teeth, dental restoration design, etc.).
[0013] Figure 8 Shows a method for training a reconstruction autoencoder.
[0014] Figure 9 Shows a latent space in which the loss includes a reconstruction loss but not a KL divergence loss.
[0015] Figure 10 Shows a latent space in which the loss includes both a reconstruction loss and a KL divergence loss.
[0016] Figure 11 Shows a recursive inference (RI) model for 3D oral care representation generation (or modification).
[0017] Figure 12 Shows a method for using a trained ML model comprising one or more transducers to generate (or modify) a 3D representation of oral care data (e.g., a dental restoration design, a fixture model component, an appliance component, etc.).
[0018] Figure 13 Shows a U-Net structure that can be used to extract hierarchical features from a 3D representation.
[0019] Figure 14Shows a pyramid encoder-decoder structure that can be used to extract hierarchical features from 3D representations.
[0020] Figure 15 Shows a fixture model component of the interproximal marginal band.
[0021] Figure 16 Shows a fixture model component of the blockout.
[0022] Figure 17 Shows a fixture model component of the digital pontic tooth.
[0023] Figure 18 Shows a fixture model component of the interproximal reinforcement structure.
[0024] Figure 19 Shows a fixture model component of the gingival ridge structure.
[0025] Figure 20 Shows the visualization of the reconstruction error of the teeth.
[0026] Figure 21 Shows a 3D representation of the dental arch form that can be generated (or modified) using the methods of the present disclosure.
[0027] Figure 22 Shows a method for training a latent representation modification module (LRMM).
[0028] Figure 23 Shows a method for using a fully trained LRMM. Detailed Description
[0029] The machine learning techniques described herein can receive various input data as described herein, including a dental mesh of one or both dental arches of a patient. The dental data can be presented in the form of a 3D representation such as a mesh, a point cloud, or a voxelized geometry. This data can be pre-processed, for example, by arranging the constituent mesh elements into a list and calculating an optional mesh element feature vector for each mesh element. Such vectors can impart valuable information about the shape and / or structure of the oral care mesh to the machine learning models described herein. Additional inputs can be received as input to the machine learning models described herein, such as one or more oral care metrics. Oral care metrics can be used to measure one or more physical aspects of the oral care mesh (e.g., physical relationships within or between teeth). In some cases, oral care metrics can be calculated for either or both of a malocclusion oral care mesh example and / or a ground truth oral care mesh example and then used in the training of the machine learning models described herein. Metric values can be received as input to the machine learning models described herein, thereby as a way to train the model or those models to encode the distribution of such metrics over a number of examples in the training dataset. During training, the network can then receive the metric values as input to help train the network to link the metric values of the input to the physical aspects of the ground truth oral care mesh used in the loss calculation. This loss calculation can quantify the difference between the prediction and the ground truth example (e.g., between the predicted oral care mesh and the ground truth oral care mesh). By providing the metric values to the network, the neural network techniques of the present disclosure can learn to train the neural network to encode the distribution of a given metric through the process of loss calculation and subsequent backpropagation. In deployment, one or more oral care parameters (protocol parameters or prosthetic design parameters) can be defined to specify one or more aspects of the expected oral care mesh that will be generated using the machine learning models trained for this purpose described herein.
[0030] One or more oral care variables can be defined to specify one or more aspects of an expected 3D oral care representation (e.g., 3D mesh, polyline, 3D point cloud, or voxelized geometry), which will be generated using a machine learning model trained for this purpose as described herein (e.g., a 3D representation generation model using a transformer). In some specific implementations, oral care variables can be defined to specify one or more aspects of a custom vector, matrix, or any other numerical representation (e.g., to describe a 3D oral care representation, such as control points for splines, dental arch morphology, transformation or coordinate system for placing teeth or appliance components relative to another 3D oral care representation), which will be generated using a machine learning model trained for this purpose as described herein (e.g., a 3D representation generation model using a transformer). The custom vector, matrix, or other numerical representation can describe a 3D oral care representation that conforms to the expected outcome of patient treatment. Oral care variables can include oral care metrics or oral care parameters, etc. Oral care variables can specify one or more aspects of an oral care protocol, such as orthodontic setup prediction or prosthetic design generation, etc. In some specific implementations, one or more oral care parameters corresponding to respective oral care metrics can be defined. Oral care variables can be provided as input to the machine learning model described herein and can be provided as instructions to the module to generate an oral care mesh with specified customization, place an oral care mesh for generating an orthodontic setup (or appliance), segment an oral care mesh, or clean an oral care mesh, generate or modify a 3D representation of oral care data, to name just a few examples. This interaction between oral care metrics and oral care parameters can also apply to the training and deployment of other prediction models in oral care.
[0031] In some specific implementations, the prediction model of the present disclosure can obtain more accurate results by combining one or more of the following inputs: dental arch morphology information V, interproximal reduction (IPR) information U, tooth size information P, diastema information Q, latent capsule representation T of the oral care mesh, latent vector representation A of the oral care mesh, protocol parameter K (which can describe the clinician's expected treatment of the patient), doctor preference L (which can describe typical protocol parameters selected by the doctor), flag M regarding tooth status (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name / dental symbol R, oral care metric S (including at least one of oral care metrics and prosthetic design metrics).
[0032] In some cases, the systems of the present disclosure may be deployed in a clinical environment, such as a dental or orthodontic clinic, for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians). Such systems deployed in a clinical environment may enable clinicians to process oral care data (such as dental scans) in the clinical environment or, in some cases, in a "chairside" environment (when the patient is in the clinical environment). A non-limiting list of examples of techniques may include: segmentation, mesh cleaning, coordinate system prediction, CTA trim line generation, restoration design generation, appliance component generation or placement or assembly, generation of other oral care meshes, verification of oral care meshes, setup prediction, removal of hardware from tooth meshes, placement of hardware on teeth, estimation of missing values, clustering of oral care data, oral care mesh classification, setup comparison, metric calculation, or metric visualization. In some cases, the execution of these techniques may enable patient data to be processed, analyzed, and used by clinicians in appliance generation before the patient leaves the clinical environment (which may facilitate treatment planning as feedback can be received from the patient during the treatment planning process).
[0033] The systems of the present disclosure may automate operations in digital orthodontics (e.g., setup prediction, hardware placement, setup comparison), digital dentistry (e.g., restoration design generation), or combinations thereof. Some techniques may be applied to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleaning, coordinate system prediction, oral care mesh verification, estimation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows, or denoising diffusion models), metric visualization, appliance component placement, or appliance component generation, etc. In some cases, the systems of the present disclosure may enable clinicians or technicians to process oral care data (such as scanned dental arches). In addition to segmentation, mesh cleaning, coordinate system prediction, or verification operations, the systems of the present disclosure may also implement orthodontic treatment planning, which may involve setup prediction as at least one operation. The systems of the present disclosure may also implement restoration design generation, where one or more restored tooth designs are generated and processed during the creation of an oral care appliance. The systems of the present disclosure may implement either or both of orthodontic or dental treatment planning, or may automate steps in the generation of either or both of orthodontic or dental appliances. Some appliances may implement both dental and orthodontic treatment, while other appliances may implement one or the other.
[0034] Aspects of the present disclosure may provide a technical solution to the technical problem of using 3D representations of a patient's dentition (and / or appliance components or fixture model components) and / or a transformer neural network to generate 3D oral care representations for use in oral care appliance generation. Specifically, by practicing the techniques disclosed herein, a computing system particularly adapted to perform 3D oral care representation generation for use in generating oral care appliances is improved. For example, aspects of the present disclosure improve the performance of a computing system having a 3D representation of a patient's dentition by reducing the consumption of computing resources. Specifically, aspects of the present disclosure reduce computing resource consumption by decimating the 3D representation of the patient's dentition (e.g., reducing the count of mesh elements used to describe aspects of the patient's dentition), such that computing resources are not wasted unnecessarily due to processing an excessive number of mesh elements. Additionally, decimating the mesh does not degrade the overall prediction accuracy of the computing system (and may actually improve prediction because the input provided to the ML model after decimation is a more accurate (or better) representation of the patient's dentition). For example, unimportant (and potentially accuracy-degrading) noise or other artifacts are removed. That is, aspects of the present disclosure provide a more efficient allocation of computing resources in a manner that improves the accuracy of the underlying system.
[0035] In addition, aspects of the present disclosure may need to be performed in a time-limited manner, such as when an oral care appliance must be generated for a patient immediately after an intraoral scan (e.g., when the patient is waiting in a clinician's office). Thus, aspects of the present disclosure must be rooted in the underlying computer technology of using a transformer neural network for 3D oral care representation generation and cannot be performed by a human, even with the aid of pen and paper. For example, specific implementations of the present disclosure must be able to: 1) store thousands or millions of mesh elements of a patient's dentition in a manner that can be processed by a computer processor; 2) perform computations on the thousands or millions of mesh elements, such as to quantify aspects of the shape and / or structure of an individual tooth in the 3D representation of the patient's dentition; and 3) generate a 3D oral care representation for use in oral care appliance generation (e.g., an orthodontic appliance tray, a dental restoration appliance, an indirect bonding tray for orthodontic treatment, etc.) based on a transformer neural network, and do so during a short patient visit.
[0036] The present disclosure relates to digital oral care encompassing the fields of digital dentistry and digital orthodontics. The present disclosure generally describes methods for processing three-dimensional (3D) representations of oral care data. It should be understood that, without loss of generality, there are various types of 3D representations. One type of 3D representation is 3D geometry. The 3D representation can include, be one or more of, or be a part of: a 3D polygon mesh, a 3D point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels for sparse processing), or a 3D representation described by mathematical equations. Although the term "mesh" is frequently used throughout the present disclosure, in some specific implementations, the term should be understood to be interchangeable with other types of 3D representations. The 3D representation can describe the 3D geometry and / or elements of the 3D structure of an object.
[0037] The dental arches S1, S2, S3, and S4 all include exactly the same tooth meshes, but these tooth meshes are transformed differently according to the following description. The first dental arch S1 includes a set of tooth meshes that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a malpositioned position and orientation. The second dental arch S2 includes the same set of tooth meshes from S1 that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a reference true set position and orientation. The third dental arch S3 includes the same meshes as S1 and S2 that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a predicted final set pose (e.g., as predicted by one or more techniques of the present disclosure). S4 is a counterpart of S3 where the teeth are in a pose corresponding to one of several intermediate stages of orthodontic treatment with a clear tray aligner.
[0038] It should be understood that, without loss of generality, the techniques of the present disclosure applied to the final set are also applicable to intermediate gradings in orthodontic treatment, specifically geometric deep learning (GDL) settings, reinforcement learning (RL) settings, variational autoencoder (VAE) settings, capsule settings, multi-layer perceptron (MLP) settings, diffusion settings, pose transfer (PT) settings, similarity settings, force-directed graph (FDG) settings, transformer settings, setting comparison, or setting classification. The metric visualization aspect of the present disclosure can also be configured to visualize data from both the final set and intermediate stages. The MLP setting, VAE setting, and capsule setting each fall within the scope of the autoencoder setting. Some specific implementations of the MLP setting can fall within the scope of the transformer setting. A representation setting refers to any one of the MLP setting, VAE setting, capsule setting, and any other setting prediction machine learning model that uses an autoencoder to create a representation of at least one tooth.
[0039] Each of the setup prediction techniques of the present disclosure is applicable to the manufacture of clear tray aligners and / or indirectly bonded trays. The setup prediction techniques may also be applicable to other products that also involve the final tooth position. The position may include location (or positioning) and rotation (or orientation).
[0040] A 3D mesh is a data structure that can describe the geometry and / or shape of an object related to oral care, the object including but not limited to teeth, hardware elements, or the gingival tissue of a patient. The 3D mesh may include one or more mesh elements, such as one or more of vertices, edges, faces, and combinations thereof. In some specific implementations, the mesh elements may include voxels, such as in the context of a sparse mesh processing operation. Various spatial and structural features can be calculated for these mesh elements and provided to the prediction model of the present disclosure, and the prediction model of the present disclosure provides a technical advantage of improved data accuracy in the form of a more accurate prediction of the model output of the present disclosure.
[0041] The dentition of a patient may include one or more 3D representations of the patient's teeth (e.g., and / or associated transducers), gums, and / or other oral anatomical structures. In some embodiments, an orthodontic metric (OM) may quantify the relative position and / or orientation of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. In some embodiments, a restorative design metric (RDM) may quantify at least one aspect of the structure and / or shape of a 3D representation of a tooth. In some embodiments, an orthodontic landmark (OL) may locate one or more points or other regions of interest structures on a 3D representation of a tooth. In some embodiments, an OL may be used in the generation of orthodontic or prosthetic appliances, such as clear tray aligners or prosthetic restorative appliances. In some embodiments, a mesh element may include at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth represented by a 3D mesh, the mesh elements may at least include: vertices, edges, faces, and voxels. In some embodiments, mesh element features may quantify some aspects of the 3D representation that are proximate to or associated with one or more mesh elements, as described elsewhere in the present disclosure. In some embodiments, an orthodontic procedure parameter (OPP) may specify at least one value that defines at least one aspect of a patient's planned orthodontic treatment (e.g., specifying desired target attributes of a final setting in a final setting prediction). In some embodiments, an orthodontist preference (ODP) may specify at least one typical value of an OPP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. In some embodiments, a restorative design parameter (RDP) may specify at least one value that defines at least one aspect of a patient's planned prosthetic restorative treatment (e.g., specifying desired target attributes of a tooth to be treated with a prosthetic restorative appliance). In some embodiments, a doctor restorative design preference (DRDP) may specify at least one typical value of an RDP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. The 3D oral care representation may include but is not limited to: 1) a set of mesh element labels that may be applied to 3D mesh elements of a tooth / gum / hardware / appliance mesh (or point cloud) during the process of mesh segmentation or mesh cleaning; 2) one or more 3D representations of teeth / gums / hardware / appliances whose shapes have been modified (e.g., trimmed, deformed, or filled) during the process of mesh segmentation or mesh cleaning; 3) one or more coordinate systems (e.g., describing one, two, three, or more coordinate axes) for a single tooth or a group of teeth (such as a full dental arch, e.g., the LDE coordinate system); 4) 3D representations of one or more teeth whose shapes have been modified or otherwise made suitable for use in prosthetic restorations; 5) 3D representations of one or more prosthetic restorative appliance components;6) One or more transformations to be applied to one or more of the following: placement of prosthetic appliance library components relative to one or more teeth, teeth to be placed for an orthodontic setting (final setting or intermediate stage), hardware elements to be placed relative to one or more teeth, etc.; 7) Orthodontic settings; 8) 3D representations of hardware elements (such as facebows, lingual bows, orthodontic attachments, buttons, hooks, occlusal ramps, etc.) placed relative to one or more teeth, etc.; 8) 3D representations of bonding pads for hardware elements (which can be generated for specific teeth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting the tooth through a Boolean operation); 9) 3D representations of clear tray appliances (CTAs); 10) The position or shape of CTA trim lines (e.g., described as a grid or polyline); 11) Arch forms (e.g., described as 3D polylines or 3D grids or surfaces) that describe the contour or layout of the dental arch, which can follow the incisal edges of one or more teeth, which can follow the facial surfaces of one or more teeth, which in some specific implementations can correspond to malocclusion dental arches and in other specific implementations correspond to final setting dental arches (the impact of malocclusion on the shape of the arch form can be reduced by smoothing or averaging the shape of the arch form), and which can be described by one or more control points and / or splines; 12) 3D representations of jig models (e.g., depictions of teeth and gums used in thermoformed clear tray appliances, or depictions of teeth / gums / hardware used in thermoformed indirect bonding trays); 13) One or more latent space vectors (or latent capsules) generated by the 3D encoder stage of a 3D autoencoder (e.g., a variational autoencoder trained for tooth reconstruction) that has been trained on the reconstruction of an oral care mesh; 14) One or more oral care metrics for one or more teeth (e.g., such as orthodontic metrics or prosthetic design generation metrics); 15) One or more landmarks (e.g., 3D points) that describe the shape and / or geometric properties of one or more teeth, other dentition structures, or hardware structures (e.g., to be used in orthodontic setting creation or prosthetic appliance component generation or placement); 16) 3D representations created by scanning (e.g., optical scanning, CT scanning, or MRI scanning) 3D printed parts (such as scanned jig models) corresponding to one or more teeth / gums / hardware / appliances; 17) 3D printed appliances (optionally including local thickness, reinforcement rib geometry, tab positioning, etc.); 18) 3D representations of a patient's dentition captured by a clinician or healthcare practitioner at the chairside (e.g., in an environment where the 3D representation is verified at the chairside, before the patient leaves the clinic, such that errors can be detected and re-scanning can be performed as needed); 19) Prosthetic dental designs (e.g., for veneers, crowns, bridges, or prosthetic appliances); 20) 3D representations of one or more teeth used in digital oral care processes; 21) Other 3D printed parts belonging to oral care procedures or other fields;22) IPR cutting surfaces, such as IPR cutting planes that may define a portion of enamel to be removed from a tooth; 23) one or more orthodontic setting transformations associated with one or more IPR cutting surfaces; 24) (digital) pontic design that may fill at least a portion of the space between teeth to create space for erupting teeth in an orthodontic setting and then emerge from the gums; or 25) components of a jig model (e.g., including jig model components such as interdental straps, sealed concavities, bite locks, bite ramps, interdental reinforcements, gingival ridge structures, torque points, power ridges, pontics or pits, etc.).;
[0042] The techniques of this disclosure may require a training data set of hundreds or thousands of cohort patient cases to ensure that the neural network can encode the distribution of patient cases that may be encountered in clinical treatment. Cohort patient cases may include a set of crown meshes, a set of root meshes, or a data file (e.g., a JSON file) including case attributes. Typical examples of cohort patient cases may include up to 32 crown meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), multiple gingival meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), or one or more JSON files (each of which may include tens of thousands of values (e.g., objects, arrays, strings, real values, boolean values, or null values)).
[0043] The techniques of the present disclosure can be advantageously combined. For example, a setup comparison tool can be provided to compare the output of the GDL setup model with ground truth data, compare the output of the RL setup model with ground truth data, compare the output of the VAE setup model with ground truth data, and compare the output of the MLP setup model with ground truth data. By comparing each of these setup prediction models with the ground truth data, it can be determined which model achieves the best performance on a certain dataset or within a given problem domain. Additionally, a metric visualization tool can enable a global view of the final setup and intermediate stages produced by one or more of the setup prediction models, with the advantage of being able to select the best setup prediction model. Moreover, the metric visualization tool enables the calculation of metrics with a global scope within a set of intermediate stages. In some specific implementations, these global metrics can be consumed as inputs to a neural network for predicting setups (e.g., GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, etc.). The global metrics can also be provided to the FDG setup. In some specific implementations, local metrics from the present disclosure (i.e., local metrics are metrics that can be calculated for one stage or setup of processing rather than within several stages or setups) can be consumed by the neural networks herein for predicting setups, with the advantage of improving the prediction results. In some specific implementations, the metrics described in the present disclosure can be visualized using the metric visualization tool.
[0044] VAE and MAE models for mesh element tagging and mesh filling can be advantageously combined with a setup prediction neural network for mesh cleanup before or during the prediction process. In some specific implementations, the VAE for mesh element tagging can be used to tag mesh elements for further processing, such as metric calculation, removal, or modification. In some cases, such tagged mesh elements can be provided as input to the setup prediction neural network to inform the neural network of important mesh features, properties, or geometries, with the advantage of improving the performance of the resulting setup prediction model. In some specific implementations, mesh filling can make the geometry of the teeth closer to complete, enabling the setup prediction model to function better (i.e., improving the correctness of the prediction due to the better-formed geometry). In some cases, a neural network for classifying setups (i.e., a setup classifier) can assist the setup prediction neural network in functioning because the setup classifier tells the setup prediction neural network when a predicted setup is acceptable for use and can be provided to a method for generating an orthodontic tray. Setup classifiers (e.g., GDL setups, RL setups, VAE setups, capsule setups, MLP setups, diffusion setups, PT setups, similarity setups, and FDG setups, etc.) can help generate the final setup and also help generate intermediate stages. Additionally, the setup classifier neural network can be combined with a metric visualization tool. In other specific implementations, the setup classification neural network can be combined with a setup comparison tool (e.g., the setup comparison tool can output an indication of how a setup generated, in part, by the setup classifier compares to a setup generated by another setup prediction method). In some specific implementations, the VAE for mesh element tagging can identify one or more mesh elements used in metric calculation. The resulting metric output can be visualized by a metric visualization tool.
[0045] In some examples, the setup classifier neural network can assist the setup prediction techniques described in U.S. Patent Application No. US20210259808A1, the entire content of which is incorporated herein by reference, or PCT Application Publication No. WO2021245480A1, the entire content of which is incorporated herein by reference, or PCT Application No. PCT / IB2022 / 057373, the entire content of which is incorporated herein by reference. The setup classifier will help one or more of those techniques know when the predicted final setup is closest to being correct. In some cases, the setup classifier neural network can output an indication of how far a given setup is from the final setup (i.e., a progress indicator).
[0046] In some specific implementations, the latent space embedding vectors from the reconstructed VAE can be cascaded with the input of the setup prediction neural network described in WO2021245480A1. The latent space vectors can also be combined as inputs into other setup prediction models: GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup, etc. The advantage is to endow the neural network with reconstruction characteristics (e.g., the latent vector dimension of the dental mesh), thereby improving the generated setup prediction.
[0047] In some examples, the various setup prediction neural networks of the present disclosure can work together to generate the setups required for orthodontic treatment. For example, the GDL setup model can generate the final setup, and the RL setup model can use the final setup as an input to generate a series of intermediate stage setups. Alternatively, the VAE setup model (or MLP setup model) can create the final setup, which can be used by the RL setup model to generate a series of intermediate stage setups. In some specific implementations, the setup prediction can be generated by one setup prediction neural network and then used as an input to another setup prediction neural network for further improvement and adjustment. In some specific implementations, such improvements can be performed in an iterative manner.
[0048] In some specific implementations, a setup verification model may be involved in this iterative setup prediction loop, such as the model disclosed in U.S. Provisional Application No. US63 / 366495. First, a setup can be generated (e.g., using a model trained for setup prediction, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, etc.), and then the setup is verified. If the setup passes the verification, the setup can be output for use. If the setup does not pass the verification, the setup can be sent back to one or more of the setup prediction models for correction, improvement, and / or adjustment. In some cases, the setup verification model can output an indication of what is wrong with the setup, such that the setup generation model can be improved in the next iteration. The process is iterated until completion.
[0049] Generally, in some specific implementations, two or more of the following techniques of the present disclosure can be combined during orthodontic and / or dental treatment: GDL setting, setting classification, reinforcement learning (RL) setting, setting comparison, autoencoder setting (VAE setting or capsule setting), VAE grid element labeling, masked autoencoder (MAE) grid filling, multi-layer perceptron (MLP) setting, metric visualization, estimation of missing oral care parameter values, tooth classification using latent vectors, FDG setting, pose transfer setting, prosthetic design metric calculation, neural network techniques for dental restoration and / or orthodontics (e.g., generation or modification of 3D oral care representation using transformers), landmark-based (LB) setting, diffusion setting, estimation of tooth movement protocol, capsule autoencoder segmentation, diffusion segmentation, similarity setting, verification of oral care representation (e.g., using autoencoders), coordinate system prediction, prosthetic design generation or generation (or modification) of 3D oral care representation using diffusion models.
[0050] In some cases, coordinate system prediction can be used in combination with the techniques of the present disclosure. The pose transfer technique can be trained for coordinate system prediction in the form of predicting the transformation of teeth. The reinforcement learning technique can be trained for coordinate system prediction in the form of predicting the transformation of teeth.
[0051] In some cases, a shape-based input of teeth can be provided to the neural network for setting prediction. In other cases, a non-shape-based input, such as tooth name or nomenclature, can be used as it is related to dental notation. In some specific implementations, a vector R of flags can be provided to the neural network, where a "1" value indicates the presence of a tooth and a "0" value indicates the absence of a tooth in the patient case (although other values are also possible). The vector R can include one-hot vectors, where each element in the vector corresponds to a tooth type, name, or nomenclature. Identification information about the teeth (e.g., the name of the teeth) can be provided to the prediction neural network of the present disclosure, which has the advantage of enabling the neural network to be trained to handle different teeth in a tooth-specific manner. For example, the setting prediction model can learn to predict setting transformation for a specific tooth name (e.g., the upper right central incisor or the lower left canine, etc.). In the case of a grid cleaning autoencoder (for labeling grid elements or for filling missing grid data), the autoencoder can be trained in this way to provide specialized processing to the tooth according to the tooth's nomenclature. In the case of a setting classification neural network, a list of tooth names present in the patient's dental arch can better enable the neural network to output an accurate determination of setting classification, as tooth nomenclature is a valuable input for training such a neural network. For example, tooth naming / name can be defined according to the Universal Numbering System, the Palmer Quadrant System, or the FDI World Dental Federation notation (ISO 3950).
[0052] In one example, in the presence of all teeth except (at most four) wisdom teeth, the vector R can be defined as an optional input to the set prediction neural network of the present disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, UL1, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7.
[0053] In some cases, the position of the cusps can be provided to the neural network for set prediction. In other cases, one or more vectors S of orthodontic metrics described elsewhere in the present disclosure can be provided to the neural network for set prediction. The advantage is that the network's ability to be trained to understand the state of the malocclusion set is improved, and thus it can predict a more accurate final set or intermediate stage.
[0054] In some specific embodiments, the neural network can take as input one or more indications of interproximal reduction (IPR) U, which can indicate the amount of enamel to be removed from the teeth (from the mesial or from the distal) during the course of orthodontic treatment. In some specific embodiments, the IPR information (e.g., the amount of IPR to be performed on one or more teeth, measured in millimeters, or one or more binary markers indicating whether IPR is to be performed on each tooth identified by a label) can be concatenated with the latent vector A generated by the VAE or the latent capsule T autoencoder. The vector and / or capsule resulting from such concatenation can be provided to one or more of the neural networks of the present disclosure, which has the technical improvement or additional advantage of enabling the prediction neural network to take IPR into account. IPR is particularly relevant to the set prediction method, which can determine the position and pose of the teeth at the end of the treatment or during one or more stages of the treatment. It is important to consider the amount of enamel to be removed before the predicted tooth movement.
[0055] In some specific embodiments, one or more protocol parameters K and / or doctor preference vectors L can be introduced into the set prediction model. In some specific embodiments, one or more optional vectors or values include: tooth position N (e.g., XYZ coordinates in local or global coordinates of the tooth), tooth orientation O (e.g., pose, such as in a transformation matrix or quaternion, Euler angles or other forms described herein), tooth size P (e.g., length, width, height, perimeter, radius, diagonal measurement, volume, any size can be normalized relative to one or more other teeth), distance Q between adjacent teeth. In some cases, these "tooth sizes P" can be used to describe the expected size of the teeth for dental restoration design generation.
[0056] In some specific implementations, a tooth dimension P, such as length, width, height, or perimeter, can be measured in a plane, such as a plane intersecting the centroid of the tooth, or a plane intersecting a central point located at the midpoint between the centroid of the tooth and the most incisal extent or the most gingival extent. The tooth height dimension can be measured as the distance from the gingiva to the incisal edge. The tooth width dimension can be measured as the distance from the mesial extent to the distal extent of the tooth. In some specific implementations, the roundness or circularity of the tooth cross-section can be measured and included in the vector P. The roundness or circularity can be defined as the ratio of the radii of the inscribed circle and the circumscribed circle.
[0057] The distance Q between adjacent teeth can be implemented in different ways (and calculated using different distance definitions, such as Euclidean or geodesic). In some specific implementations, the distance Q1 can be measured as the average distance between the mesh elements of two adjacent teeth. In some specific implementations, the distance Q2 can be measured as the distance between the centers or centroids of two adjacent teeth. In some specific implementations, the distance Q3 can be measured between the closest mesh elements between two adjacent teeth. In some specific implementations, the distance Q4 can be measured between the tooth tips of two adjacent teeth. In some specific implementations, teeth can be considered adjacent within an arch. In some specific implementations, teeth can also be considered adjacent between opposing arches. In some specific implementations, any one of Q1, Q2, Q3, and Q4 can be divided by a term to normalize the resulting value of Q. In some specific implementations, the normalization term can involve one or more of the following: the volume of the tooth, the count of mesh elements in the tooth, the surface area of the tooth, the cross-sectional area of the tooth (e.g., as projected onto the XY plane), or some other term related to the tooth dimension.
[0058] Other information regarding the patient's dentition or treatment needs (or related parameters) can be concatenated with other input vectors to one or more of an MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and / or any neural network model listed elsewhere in this disclosure.
[0059] The vector M may include markers applied to one or more teeth. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is pinned. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is fixed. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is a pontic. Other and additional markers are possible for teeth, such as combinations of fixed, pinned, and pontic markers. A marker set to a value indicating that a tooth should be fixed is a signal that the tooth should not move during processing and is sent to the network. In some embodiments, the neural network loss function can be designed to penalize any movement in the indicated teeth (and in some cases, may be severely penalized). A marker indicating that a tooth is a pontic notifies the network to maintain the diastema, although the gap is allowed to move. In some cases, M may include a marker indicating tooth loss. In some embodiments, the presence of one or more fixed teeth in the dental arch can help set the prediction because the one or more fixed teeth can provide an anchor for the posture of other teeth in the dental arch (i.e., can provide a fixed reference for the posture transformation of one or more other teeth in the dental arch). In some embodiments, one or more teeth can be intentionally fixed in order to provide an anchor to which other teeth can be positioned. In some embodiments, a 3D representation (such as a mesh) corresponding to the gingiva can be introduced to provide a reference point according to which the teeth can move.
[0060] Without loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U, and V described elsewhere in this disclosure can also be provided as an input to one or more of the prediction models of this disclosure or fed into an intermediate layer thereof. Specifically, these optional vectors can be provided to the MLP setup, GDL setup, RL setup, VAE setup, capsule setup, and / or diffusion setup, with the advantage of enabling the corresponding models to generate setups that better meet the orthodontic treatment needs of the patient. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent vectors A that are also provided to one or more of the prediction models of this disclosure. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent capsules T that are also provided to one or more of the prediction models of this disclosure.
[0061] In some embodiments, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into a neural network (e.g., an MLP or a transformer) in the hidden layer of the network. In some cases, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into the internal processing of an encoder structure.
[0062] In some specific implementations, a prediction model (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, PT settings, similarity settings, and diffusion settings) can take as input one or more latent vectors A corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a prediction model (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, and diffusion settings) can take as input one or more latent capsules T corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a prediction method can take both A and T as input.
[0063] A variety of loss calculation techniques generally apply to the techniques of the present disclosure (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, setting classification, tooth classification, VAE mesh element labeling, MAE mesh filling, and estimation of protocol parameters).
[0064] These losses include L1 loss, L2 loss, mean squared error (MSE) loss, cross-entropy loss, etc. The losses can be calculated and used to train neural networks such as multi-layer perceptrons (MLPs), U-Net architectures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer architectures, etc. For example, in the learning of sequences, some specific implementations can use triplet loss or contrastive loss.
[0065] Losses can also be used to train encoder architectures and decoder architectures. KL divergence loss can be at least partially used to train one or more neural networks of the present disclosure, such as a grid reconstruction autoencoder or the generator of GDL settings, which has the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior can enable the reconstruction autoencoder to produce better reconstructions (e.g., when modifying the latent vector representation and using the decoder to reconstruct the modified latent vector, the resulting reconstruction is more likely to be a valid instance of the input representation). There are other techniques for calculating losses that can be described elsewhere in the present disclosure. Such losses can be based on quantifying the difference between two or more 3D representations.
[0066] MSE loss calculation can involve the calculation of the mean squared distance between two sets, vectors, or data sets. MSE can generally be minimized. MSE can be applied to regression problems where the predictions generated by a neural network or other machine learning model can be real numbers. In some specific implementations, the neural network can be equipped with one or more linear activation units on the output to generate MSE predictions. According to the techniques of the present disclosure, mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used.
[0067] In some specific implementations, cross-entropy can be used to quantify the difference between two or more distributions. In some specific implementations, cross-entropy loss can be used to train the neural networks of the present disclosure. In some specific implementations, cross-entropy loss can involve comparing predicted probabilities with ground truth probabilities. Other names for cross-entropy loss include "log loss", "logistic loss", and "log loss". A small cross-entropy loss can indicate a better (e.g., more accurate) model. Cross-entropy loss can be logarithmic. In some specific implementations, cross-entropy loss can be applied to binary classification problems. In some specific implementations, a neural network can be equipped with sigmoid activation units at the output to generate probability predictions. In the case of multi-class classification, cross-entropy can also be used. In this case, in some specific implementations, a neural network trained to make multi-class predictions can be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for each class to be predicted). Other loss calculation techniques that can be applied in the training of the neural networks of the present disclosure include one or more of the following: Huber loss, hinge loss, classification hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss calculation methods are described herein and can be applied to the training of any neural network described in the present disclosure.
[0068] In some specific implementations, one or more neural networks of the present disclosure can be trained, at least in part, by a loss based on at least one of the following: pointwise mesh Euclidean distance (PMD) and Earth Mover's Distance (EMD). Some specific implementations can incorporate Hausdorff distance (HD) calculation into the loss calculation. Calculating the Hausdorff distance between two or more 3D representations, such as 3D meshes, can provide one or more technical improvements because HD not only considers the distance between two meshes, but also the way those meshes are oriented and the relationship between the mesh shapes in those orientations (or positions or poses). Hausdorff distance can improve the comparison of two or more tooth meshes, such as two or more instances of tooth meshes in different poses (e.g., comparing a predicted setting with a ground truth setting, which can be done during the process of calculating the loss value for training a setting prediction neural network).
[0069] The reconstruction loss can compare the predicted output with the ground truth (or reference) output. The systems of the present disclosure can calculate the reconstruction loss as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_loss = 0.5 * L1(all_points_target, all_points_predicted) + 0.5 * MSE(all_points_target, all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the ground truth data (e.g., a ground truth dental restoration design, or some other ground truth example of a 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the generated or predicted data (e.g., a generated dental restoration design, or some other generated example of a 3D type of oral care representation). Other specific implementations of the reconstruction loss can additionally (or alternatively) involve an L2 loss, a mean absolute error (MAE) loss, or a Huber loss term.
[0070] The reconstruction error can compare the reconstructed output data (e.g., generated by a reconstruction autoencoder, such as a tooth design that has been generated for generating a dental restoration appliance) with the initial input data (e.g., the data input to the reconstruction autoencoder, such as the tooth before restoration). The systems of the present disclosure can calculate the reconstruction error as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_error = 0.5 * L1(all_points_input, all_points_reconstructed) + 0.5 * MSE(all_points_input, all_points_reconstructed). In the above example, all_points_input is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the input data (e.g., a pre-restoration tooth design input to the reconstruction autoencoder, or some other 3D oral care representation input to an ML model). In the above example, all_points_reconstructed is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the reconstructed (or generated) data (e.g., a reconstructed dental restoration design, or some other example of a generated 3D oral care representation).
[0071] In other words, the reconstruction loss involves calculating the difference between the predicted output and the reference output, while the reconstruction error involves calculating the difference between the reconstructed output and the initial input from which the reconstructed data is derived.
[0072] The entire content of the following paper is incorporated herein by reference: "Attention Is All You Need"; Ashish Vaswani, Noam Shazeer, Niki Parmar, Niki Parmar, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin; NIPS 2017. The neural network-based models of the present disclosure may provide additional advantages in specific implementations where they are integrated with a neural network architecture known as a "transformer". Figure 1 An example implementation of the transformer architecture is shown.
[0073] Prior to recently developed models such as transformer models, RNN-type models represented the state of the art for natural language processing (NLP). An example application of NLP is to generate new text based on previous words or text. Due to the important property of the transformer model with multi-head attention characteristics, the transformer then provides a significant improvement over GRU, LSTM and other such RNN-based NLP techniques. In some specific implementations, the NLP concept of multi-head attention can describe the relationship between each word in a sentence (or paragraph or document or document corpus) and each other word in the sentence (or paragraph or document or document corpus). These relationships can be generated by a multi-head attention module and can be encoded in vector form. The vector can describe how each word in a sentence (or paragraph or document or document corpus) should pay attention to each other word in the sentence (or paragraph or document or document corpus). RNN, LSTM and GRU models process sequences, such as sentences, one word at a time from the beginning to the end of the sequence. In addition, the model can only consider a given subset of sentences (called a window) when making predictions. However, in some cases, transformer-based models can take into account the entire previous text by processing the sequence as a whole in a single step. Transformers, RNNs, LSTMs, and GRU models can all be adapted for use in predictive models in digital dentistry and digital orthodontics, particularly in setting prediction tasks. In some implementations, an exemplary transformer model for use with 3D meshes and 3D transforms in setting predictions (or other oral care techniques) can be adapted based on a bidirectional encoder representation (BERT) from a transformer and / or a generative pre-trained (GPT) model. For example, a GPT (or BERT) model can first be trained on other data, such as text or document data, and then used in transfer learning. This transfer learning process can receive a previously trained GPT or BERT model and then further train it using data including a 3D oral care representation. Such transfer learning can be performed to train oral care models, such as: segmentation, mesh cleaning, coordinate system prediction, setting prediction, validation of 3D oral care representations, transformation prediction for placement of oral care meshes (e.g., teeth, hardware, appliance components, fixture model components), dental restoration design generation (or other 3D oral care representation generation, such as appliance components, fixture models, or dental arch morphology), classification of 3D oral care representations, estimation of missing oral care parameters, clustering of clinicians or clustering of clinician preferences, etc.
[0074] Oral care data may include one or more of the following (or combinations thereof): 3D representations of teeth (e.g., meshes, point clouds, or voxels), portions of a tooth mesh (such as a subset of mesh elements), tooth transformations (such as in the form of matrices, vectors, and / or quaternions, or combinations thereof), transformations for appliance components, transformations for fixture model components, and mesh coordinate system definitions (such as represented by a transformation, e.g., a transformation matrix) and / or other 3D oral care representations described herein.
[0075] A transducer can be trained to generate a transformation to position a tooth into a set pose (or place an appliance component used in appliance generation or a fixture model component used in fixture model generation). Some embodiments can operate in an offline prediction environment, and some embodiments operate in an online reinforcement learning (RL) environment. In some embodiments, the transducer can initially be trained in an offline environment and then undergo further fine-tuning training in an online environment. In an offline prediction environment, the transducer can be trained from a dataset of queued patient case data. In an online RL environment, the transducer can be trained from, for example, a physical model or a CAD model. The transducer can learn from static data, such as a transformation (e.g., a trajectory transducer). In some embodiments, the transformation can provide a mapping from a malocclusion to a set (e.g., receive a transformation matrix as input and generate a transformation matrix as output). Some embodiments of the transducer can be trained to process 3D representations, such as 3D meshes, 3D point clouds, or voxels (e.g., using a decision transducer) and take a geometry (e.g., a mesh, point cloud, voxel, etc.) as input and output a transformation. The decision transducer can be coupled to a representation generation module that encodes a representation of a patient's dentition (e.g., teeth), such as a VAE, U-Net, encoder, transducer encoder, pyramid encoder-decoder, or a simple dense or fully connected network, or combinations thereof. In some embodiments, the representation generation module (e.g., a VAE, U-Net, encoder, pyramid encoder-decoder, or dense network for generating a tooth representation) can be trained to generate a representation on one or more teeth. The representation generation module can be trained on all teeth in both dental arches, only teeth within the same dental arch (upper or lower), only anterior teeth, only posterior teeth, or some other subset of teeth. In some embodiments, such a model can be trained on each individual tooth (e.g., the right upper canine), such that the model is trained or otherwise configured to generate a highly accurate representation of an individual tooth. In some embodiments, an encoder structure can encode such a representation. In some embodiments, the decision transducer can learn in an online environment, an offline environment, or both. The online decision transducer can be trained (e.g., using RL techniques) to output actions, states, and / or rewards. In some embodiments, the transformation can be discretized to allow for segmented or stepwise actions.
[0076] In some specific implementations, the transformer can be trained to handle the embedding of the dental arch (i.e., predict the transformations of multiple teeth simultaneously) for prediction settings. In some specific implementations, the embeddings of single teeth can be cascaded into a sequence and then input into the transformer. The VAE can be trained to perform this embedding operation, the U-Net can be trained to perform such embedding, or a simple dense or fully connected network can be trained, or a combination of these operations can be employed. In some specific implementations, the transformer-based techniques of the present disclosure can predict the actions of single teeth or the actions of multiple teeth (e.g., predict the transformations of each of the multiple teeth).
[0077] The 3D mesh transformer can include a transformer encoder structure (which can encode oral care data) and can be followed by a transformer decoder structure. The 3D mesh transformer encoder can encode the oral care data into a latent representation, which can be combined with attention information (e.g., by concatenating the vector of attention information to the latent representation). In some specific implementations, the attention information can help the decoder focus on relevant oral care data during the decoding process (e.g., focus on tooth order or mesh element connectivity), such that the transformer decoder can generate a useful output for the 3D mesh transformer (e.g., an output that can be used in generating an oral care appliance). Either or both of the transformer encoder or the transformer decoder can generate the latent representation. The decoder can be used to reconstruct the output of the transformer decoder (or the transformer encoder) into, for example, one or more tooth transformations for settings, one or more mesh element labels for segmentation, a coordinate system transformation used in coordinate system generation, or one or more points of a point cloud or voxels or other mesh elements for another 3D representation. The transformer can include one or more of the following modules: a multi-head attention module, a feed-forward module, a normalization module, a linear module, and a softmax module, as well as a convolutional model for latent vector compression and / or representation.
[0078] The encoder can be stacked one or more times to further encode oral care data and enable learning of different representations of the oral care data (e.g., different latent representations). These representations can be embedded with attention information (which can affect the decoder's focus on relevant parts of the latent representation of the oral care data) and provided to the decoder in a sequential form (e.g., as a concatenation of latent representations, such as latent vectors). In some embodiments, the encoded output of the encoder (e.g., the latent representation) can be used by downstream processing steps in the generation of the oral care appliance. For example, the generated latent representation can be reconstructed into a transformation (e.g., for placing teeth in a setting, or placing appliance components or fixture model components) or can be reconstructed into a 3D representation (e.g., a 3D point cloud, a 3D mesh, or other representations disclosed herein). In other words, the latent representation generated by the transducer (e.g., including sequentially encoded attention information) can be provided to a decoder that has been configured to reconstruct the latent representation into a specific data structure required for a particular domain area. Sequentially encoded attention information can include attention information that has undergone processing by multiple multi-head attention modules within the transducer encoder or the transducer decoder, to name just one example. Additionally, data from a specific domain can be used to compute a loss for that domain. The loss computation can train the transducer decoder to accurately reconstruct the latent representation into an output data structure related to the specific domain.
[0079] For example, when the decoder generates a transformation for an orthodontic setting, the decoder can be configured with an output that describes, for example, 16 real values including a 4×4 transformation matrix (other data structures for describing the transformation are possible). In other words, the latent output generated by the transducer encoder (or the transducer decoder) can be used to predict the set-up tooth transformation for one or more teeth to place those teeth in a set-up position (e.g., the final set-up or an intermediate stage). Such a transducer encoder (or transducer decoder) can be trained at least in part using a reconstruction loss (or a representation loss, and other losses described herein) function that can compare the predicted transformation to a ground truth (or reference) transformation.
[0080] In another example, when the decoder generates a transformation for a tooth coordinate system, the decoder can be configured with an output that describes, for example, 16 real values including a 4×4 transformation matrix (other data structures for describing the transformation are possible). In other words, the latent output generated by the transducer encoder (or the transducer decoder) can be used to predict the local coordinate system for one or more teeth. Such a transducer encoder (or transducer decoder) can be trained at least in part using a representation loss (or a reconstruction loss, and other losses described herein) function that can compare the predicted coordinate system to a ground truth (or reference) coordinate system.
[0081] In another example, when the decoder produces a 3D point cloud (or other 3D representation, such as a 3D mesh, voxelized representation, etc.), the decoder can be configured to output a description of, for example, one or more 3D points (e.g., including XYZ coordinates). In other words, the latent output generated by the transformer encoder (or transformer decoder) can be used to predict mesh elements for generating (or modifying) the 3D representation. Such a transformer encoder (or transformer decoder) can be trained using at least in part a reconstruction loss (or L1, L2, or MSE loss, and other losses described herein) function that compares the predicted 3D representation to a ground truth (or reference) 3D representation.
[0082] In another example, when the decoder generates mesh element labels for 3D representation segmentation or 3D representation cleaning, the decoder can be configured to output a description of, for example, one or more labels of the mesh elements. In other words, the latent output generated by the transformer encoder (or transformer decoder) can be used to predict mesh element labels for mesh segmentation or mesh cleaning. Such a transformer encoder (or transformer decoder) can be trained using at least in part a cross-entropy loss (or other losses described herein) function that compares the predicted mesh element labels to ground truth (or reference) mesh element labels.
[0083] Multi-head attention and transformers can be advantageously applied to setting generation problems. Multi-head attention is a module in a 3D transformer encoder network that calculates attention weights for the provided oral care data and produces an output vector with encoded information about how each example of the oral care data should attend to each other piece of oral care data in the dental arch. Attention weights are a quantification of the relationship between pairs of oral care data.
[0084] A 3D representation of oral care data (e.g., including voxels, point clouds, or 3D meshes composed of vertices, faces, or edges) can be provided to a transducer. The 3D representation can describe a patient's dentition, a fixture model (or components of a fixture model), an appliance (or components of an appliance), etc. In some specific implementations, a transducer decoder (or transducer encoder) can be equipped with multi-head attention. Multi-head attention can enable the transducer decoder (or transducer encoder) to attend to different parts of the 3D representation of oral care data. For example, multi-head attention can enable the transducer to attend to grid elements within a local neighborhood (or clique), or to attend to global dependencies between grid elements (or cliques). For example, multi-head attention can enable a transducer for setting prediction (e.g., a transducer-based setting prediction model) to generate a transformation of a tooth and, when generating the transformation, to attend to each of the other teeth in the dental arch substantially simultaneously. In other words, the transformation of each tooth can be generated based on the pose of one or more other teeth in the dental arch, resulting in a more accurate transformation (e.g., a transformation that more closely conforms to a ground truth or reference transformation). In an example of 3D representation generation (e.g., generation of a 3D point cloud), a transducer model can be trained to generate a dental restoration design. Multi-head attention can enable the transducer to attend to multiple parts of a tooth (or to attend to the surfaces of adjacent teeth) as the tooth undergoes the generation process. For example, a transducer for restoration design generation can generate grid elements for the incisal edge of an incisor while at least substantially simultaneously attending to grid elements of the mesial, distal, facial, or lingual surfaces of the incisor. The result can be the generation of grid elements to form an incisal edge of a tooth that seamlessly merges with the adjacent surfaces of the tooth. This use of multi-head attention results in a more accurate modeling of the distribution of the training dataset compared to techniques that do not apply multi-head attention.
[0085] In some specific implementations of the present disclosure, one or more attention vectors can be generated that describe how aspects of oral care data interact with other aspects of oral care data associated with a dental arch. In some specific implementations, one or more attention vectors can be generated to describe how one or more parts of a tooth T1 interact with one or more parts of teeth T2, T3, T4, etc. A portion of a mesh can be described as a set of mesh elements as defined herein. In some specific implementations, the interacting parts of tooth T1 and tooth T2 can be determined, in part, by computing mesh correspondences as described herein. Any of these models (RNNs, GRUs, LSTMs, and transducers) can be advantageously applied to the task of setting transformation prediction, such as in the models described herein. Transducers can be particularly advantageous because transducers can enable the generation of transformations of multiple teeth or even an entire dental arch at once, rather than generating them individually, as may occur in some other models such as encoder architectures. In other specific implementations, an attention-free transducer can be used to make predictions based on oral care data.
[0086] A specific implementation of setting the neural network model by the GDL may include a representation generation module (e.g., including a U-Net structure, an autoencoder encoder, a transformer encoder, another type of encoder-decoder structure, or an encoder, etc.), and this representation generation module may provide its output to a module that is trained to generate a tooth transformer (e.g., a set of fully connected layers with optional skip connections, or an encoder structure) to generate predictions of the transformation of each single tooth. In some specific implementations, the skip connection may connect the output of a specific layer in the neural network to the input of another layer (e.g., a layer that is not immediately adjacent to the initial layer) in the neural network. The transformation generation module (e.g., an encoder) may process the transformation predictions one tooth at a time. Other specific implementations may replace this encoder structure with a transformer (e.g., a transformer encoder or a transformer decoder) that can process all predictions of all teeth substantially simultaneously.
[0087] In other words, the transformer can be configured to receive a much larger number of input values than some other neural network models (e.g., than a typical MLP). This is because the transformer can accommodate an increasing number of inputs, and the predictions corresponding to those inputs can be generated substantially simultaneously. The representation generation module (e.g., a U-Net structure) may provide its output to the transformer, and the transformer can generate the setting transformation for all several teeth at once, with the technical advantage of improved accuracy (because the transformation of each tooth is generated based on the transformation of each adjacent or nearby tooth, resulting in fewer conflicts and better consistency with the processing target). The transformer can be trained to output a transformation, such as a transformation encoded by a 4×4 matrix (or some other size), a quaternion, a translation vector, Euler angles, or some other form. The transformation can place the tooth into a set pose, place the fixture model component into a pose suitable for fixture model generation, or place the appliance component into a pose suitable for appliance generation (e.g., a dental restoration appliance, a clear tray aligner, etc.). In some specific implementations, the transformation can define a coordinate system for aspects of the patient's dentition, such as a tooth mesh (e.g., the local coordinate system of the teeth). In some specific implementations, a neural network may be first used to encode the input to the transformer (e.g., a latent representation or an embedding may be generated), such as one or more linear layers and / or one or more convolutional layers. In some specific implementations, the transformer may be first trained on an offline dataset and then trained using an auxiliary actor-critic network, so that online reinforcement learning can be achieved.
[0088] In some specific implementations, a transformer can achieve large model capacity and / or implement an attention mechanism (e.g., the ability to attend to and respond to certain inputs). The attention mechanism found within a transformer (e.g., multi-head attention) enables relationships within a sequence to be encoded into neural network features. Relationships within a sequence can be encoded, for example, by associating sequence numbers (e.g., 1, 2, 3, etc.) with each tooth in a dental arch, or by associating sequence numbers with each mesh element in a 3D representation (e.g., of a tooth). In a specific implementation where latent vectors of teeth are provided to a transformer, relationships within a sequence can be encoded, for example, by associating sequence numbers (e.g., 1, 2, 3, etc.) with each element in the latent vector.
[0089] A transformer can be scaled by increasing the number of attention heads and / or by increasing the number of transformer layers. In other words, one or more aspects of a transformer can be independently trained to handle discrete tasks and later combined to allow the resulting transformer to perform all the tasks for which its individual components have been trained, without degrading the prediction accuracy of the neural network. Scaling a convolutional network can be more difficult because the model may have lower extensibility or may have lower interchangeability.
[0090] Performing convolutions as described herein can result in rotation- and translation-invariant systems and techniques, which leads to improved generalization because a convolutional model may not need to consider the way the input data is rotated or translated. A transformer configured as described herein can be permutation-invariant because relationships within a sequence can be encoded into neural network features.
[0091] In some specific implementations for generating or modifying 3D oral care representations, a transformer can be combined with a convolutional-based neural network, such as by vertically stacking convolutional layers and attention layers. Stacking transformer blocks with convolutional blocks enables the resulting structure to have the translational invariance of convolutions and also the permutation invariance of transformers. Such stacking can improve model capacity and / or model generalization. CoAtNet is an example of a network architecture that combines convolutional and attention-based elements and can be applied to the processing of oral care data. In some cases, transfer learning can be used, at least in part, to train a network for modifying or generating 3D oral care representations from CoAtNet (or another model that combines convolution and self-attention / transformer).
[0092] The techniques of the present disclosure may include operations such as 3D convolution, 3D pooling, 3D deconvolution, and 3D upsampling. 3D convolution, for example, may assist in segmentation processing when downsampling a 3D mesh. 3D deconvolution, for example, performs an inverse operation to 3D convolution in a U-Net. 3D pooling may assist in segmentation processing, for example, in a generalized neural network feature map. 3D upsampling, for example, performs an inverse operation to 3D pooling in a U-Net. These operations may be implemented by one or more layers in a predictive or generative neural network as described herein. These operations may be applied directly to mesh elements such as mesh edges or mesh faces. These operations provide a technical improvement over other methods because these operations are invariant to mesh rotation, scaling, and translation changes. Generally speaking, these operations depend on edge (or face) connectivity, so as long as the edge (or face) connectivity is maintained, these operations are not affected by mesh changes in 3D space. That is, these operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position, or scale of the oral care mesh, which may improve data accuracy. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes and can be used for tasks such as 3D shape classification or mesh element labeling (e.g., for segmentation or mesh cleaning). MeshCNN performs these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.
[0093] In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 2D representations such as images. In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 3D representations such as meshes or point clouds. An intraoral scanner may capture 2D images of a patient's dentition from various angles. The intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data depicting the patient's dentition. According to various techniques, an autoencoder (or other neural network described herein) may be trained to operate on either or both of 2D and 3D representations.
[0094] A 2D autoencoder (including a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or latent capsule) using the 2D encoder and then reconstruct a copy of the input 2D image using the 2D decoder. For a handheld mobile application that has been developed for such analysis (e.g., for analysis of dental anatomy), 2D images may be easily captured using one or more onboard cameras. In other examples, 2D images may be captured using an intraoral scanner configured for such a function. Operations that may be used in implementations of a 2D autoencoder (or other 2D neural network) for 2D image analysis are 2D convolution, 2D pooling, and 2D reconstruction error calculation.
[0095] 2D image convolution can involve the "sliding" of a kernel across a 2D image and the calculation of element-wise multiplications, as well as summing these element-wise multiplications into output pixels. The output pixels generated from each new position of the kernel are saved into an output 2D feature matrix. In some specific implementations, adjacent elements (e.g., pixels) can be in well-defined positions in a straight-line grid (e.g., above, below, left, and right). 2D Pooling :
[0096] A 2D pooling layer can be used to downsample a feature map and summarize the presence of certain features in that feature map.
[0097] A 2D reconstruction error can be computed between the pixels of an input image and a reconstructed image. The mapping between the pixels can be well understood (e.g., directly comparing the upper pixels [23,134] of the input image with the pixels [23,134] of the reconstructed image, assuming the two images have the same dimensions).
[0098] One of the advantages provided by the 2D autoencoder-based techniques of the present disclosure is the ease of capturing 2D image data with a handheld device. In some cases where an external data source provides data for analysis, there can be instances where only 2D image data is available. When only 2D image data is available, it is necessary to use a 2D autoencoder for analysis.
[0099] Modern mobile devices (such as commercially available smartphones) can also have the ability to generate 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera that moves around an object to capture multiple images from different views, or both), and this 3D data can be arranged into a 3D representation, such as a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation in some specific implementations. In some cases, the analysis of a 3D representation of an object can provide a technical improvement over a 2D analysis of the same object. For example, a 3D representation can describe the geometry and / or structure of an object with less ambiguity than a 2D representation (which can include shadows and other artifacts that complicate the depiction of the depth and texture of the object). In some specific implementations, 3D processing can achieve a technical improvement due to the inverse optics problem, which affects 2D representations in some cases. The inverse optics problem refers to the phenomenon that, in some cases, the size of an object, the orientation of the object, and the distance between the object and the imaging device may be combined in a 2D image of the object. Any given projection of an object on an imaging sensor can map to an infinite count of {size, orientation, distance} pairings. 3D representations achieve a technical improvement because they remove the ambiguity introduced by the inverse optics problem.
[0100] Devices configured for a specific purpose with 3D scanning, such as 3D intraoral scanners (or CT scanners or MRI scanners), can generate 3D representations of an object (e.g., a patient's dentition) that have significantly higher fidelity and accuracy than those that a handheld device might have. When such high-fidelity 3D data is available (e.g., in oral care mesh classification or the application of other 3D techniques described herein), the use of 3D autoencoders provides technical improvements (such as increased data accuracy) to extract the best possible signal from those 3D data (i.e., obtain a signal from the 3D crown meshes used in tooth classification or setup classification).
[0101] A 3D autoencoder (including a 3D encoder and a 3D decoder) can be trained on 3D data to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder, and then reconstruct a facsimile of the input 3D representation using the 3D decoder. Operations that can be used to implement a 3D autoencoder for analyzing 3D representations (e.g., 3D meshes or 3D point clouds) are 3D convolution, 3D pooling, and 3D reconstruction error calculation.
[0102] For each mesh element, 3D convolution can be performed to aggregate local features from nearby mesh elements. Processing can be performed on top of and in addition to techniques used for 2D convolution to account for different counts and positions of adjacent mesh elements (relative to a particular mesh element). A particular 3D mesh element can have a variable neighbor count, and those neighbors can be absent from expected positions (unlike pixels in 2D convolution, which can have a fixed adjacent pixel count present in known or expected positions). In some cases, the order of adjacent mesh elements can be relevant to 3D convolution.
[0103] 3D pooling operations can enable the combination of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling can iteratively reduce a 3D mesh to the mesh elements that are most highly relevant to a given application (e.g., the neural network that has been trained for it). Similar to 3D convolution, 3D pooling can benefit from special processing in addition to that required in 2D convolution to account for different counts and positions of adjacent mesh elements (relative to a particular mesh element). In some cases, the order of adjacent mesh elements may be less relevant to 3D pooling than to 3D convolution.
[0104] The 3D reconstruction error can be calculated using one or more of the techniques described herein, such as calculating the Euclidean distance between corresponding mesh elements, between two meshes. According to aspects of the present disclosure, other techniques are possible. The 3D reconstruction error can generally be calculated on 3D mesh elements rather than 2D pixels of the 2D reconstruction error. The 3D reconstruction error can achieve a technical improvement over the 2D reconstruction error because, in some cases, the 3D representation can have less ambiguity (i.e., less ambiguity in form, shape, and / or structure) than the 2D representation. In some specific implementations, due to the complexity of the mapping between the input mesh elements and the reconstructed mesh elements (i.e., the input mesh and the reconstructed mesh may have different mesh element counts, and there may be a less clear mapping between the mesh elements compared to the mapping between pixels in 2D reconstruction), additional processing may be required for 3D reconstruction over and above 2D reconstruction. Technical improvements in 3D reconstruction error calculation include increased data precision.
[0105] A 3D scanner, such as an intraoral scanner, a computed tomography (CT) scanner, an ultrasound scanner, a magnetic resonance imaging (MRI) machine, or a mobile device capable of performing photogrammetry, can be used to generate a 3D representation. The 3D representation can describe the shape and / or structure of an object. The 3D representation can include one or more of a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation, etc. A 3D mesh includes edges, vertices, or faces. Although in some cases these three types of data are related to each other, they are different. A vertex is a point in 3D space that defines the boundary of the mesh. These points could alternatively be described as a point cloud without additional information about how the points are connected to each other (as described by the edges). An edge is described by two points and can also be referred to as a line segment. A face is described by multiple edges and vertices. For example, in the case of a triangular mesh, a face includes three vertices that are interconnected to form three consecutive edges. Some meshes can include degenerate elements, such as non-manifold mesh elements, which can be removed to benefit subsequent processing. According to aspects of the present disclosure, other mesh preprocessing operations are also possible. 3D meshes are typically formed using triangles, but in other implementations, quadrilaterals, pentagons, or some other n-sided polygon can be used. In some implementations, such as in the case of performing sparse processing, a 3D mesh can be converted into one or more voxelized geometries (i.e., including voxels). The techniques of the present disclosure operating on a 3D mesh can receive one or more tooth meshes (e.g., arranged in one or more dental arches) as input. Each of these meshes can be preprocessed before being input into a prediction architecture (e.g., including at least one of an encoder, a decoder, a pyramid encoder-decoder, and a U-Net). Such preprocessing can include converting the mesh into a list of mesh elements such as vertices, edges, faces, or into voxels in the case of sparse processing. For one or more selected types of mesh elements (e.g., vertices), a feature vector can be generated. In some examples, a feature vector is generated for each vertex of the mesh. Each feature vector can include a combination of spatial features and / or structural features, as specified in the following table:
[0106] Table 1 discloses non-limiting examples of mesh element features. In some specific implementations, in addition to the spatial or structural mesh element features described in Table 1, color (or other visual cues / identifiers) may also be considered mesh element features. As used herein (e.g., in Table 1), a point differs from a vertex in that a point is part of a 3D point cloud, while a vertex is part of a 3D mesh and may have incident faces or edges. A dihedral angle (which may be expressed in radians or degrees) can be calculated as the angle (e.g., a signed angle) between two connected faces (e.g., two faces connected along an edge). The sign on the dihedral angle can reveal information about the convexity or concavity of the mesh surface. For example, in some specific implementations, a positively signed angle may indicate a convex surface. Additionally, in some specific implementations, a negatively signed angle may indicate a concave surface. To calculate the principal curvatures of a mesh vertex, the directional curvatures of each adjacent vertex around that vertex can be calculated first. These directional curvatures can be sorted in a circular order (e.g., 0 degrees, 49 degrees, 127 degrees, 210 degrees, 305 degrees) near the vertex normal vector and may include a subsampled form of the full curvature tensor. Circular order means sorting by angle around an axis. The sorted directional curvatures can contribute to a system of linear equations that admits a closed-form solution, which can estimate the two principal curvatures and directions, which can characterize the full curvature tensor. Consistent with Table 1, a voxel may also have features that are calculated as an aggregation of other mesh elements (e.g., vertices, edges, and faces) that either intersect the voxel or, in some specific implementations, are mainly or fully contained within the voxel. Rotating a mesh may not change the structural features but may change the spatial features. And, as described elsewhere in this disclosure, the term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non-limiting sense. In some specific implementations, in addition to mesh element features, there are alternative ways to describe the geometry of a mesh (such as 3D key points and 3D descriptors). Examples of such 3D key points and 3D descriptors can be found in "TONIONI A et al., 'Learning to detect good 3D keypoints.', Int J Comput. Vis. Vol. 126, pp. 1-20, 2018". In some specific implementations, 3D key points and 3D descriptors can describe the extrema (minima or maxima) of the surface of a 3D representation.In some embodiments, one or more mesh element features may be computed at least in part via deep feature synthesis (DFS), such as described in: J.M. Kanter and K. Veeramachaneni, “Deep feature synthesis: Towards automating data science endeavors”, 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp. 1-10, doi: 10.1109 / DSAA.2015.7344858.
[0107] A neural network for representation generation based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional and / or pooling layers, or other models may benefit from the use of mesh element features. Mesh element features may convey aspects of the surface shape and / or structure of a 3D representation to the neural network model of the present disclosure. Each mesh element feature describes different information about the 3D representation that may not redundantly exist in other input data provided to the neural network. For example, vertex curvature may quantify aspects of the concavity or convexity of the surface of a 3D representation that the network would not otherwise understand. In other words, mesh element features may provide a processed form of the structure and / or shape of a 3D representation; data that would otherwise not be available to the neural network. This processed information is generally more accessible, or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, mesh element features have been provided to a neural network for representation generation based on a U-Net model, and also to a representation generation model based on a variational autoencoder with continuous normalizing flows. Based on the experiments, it was found that a system using a full complement of mesh element features (e.g., “XYZ” coordinate tuples, “normal vectors”, “vertex curvature”, point pivots, and normal pivots) was at least 3% more accurate than a system that did not use mesh element features. A point pivot describes an “XYZ” coordinate tuple with a local coordinate system (e.g., at the centroid of the corresponding tooth). A normal pivot describes a “normal vector” with a local coordinate system (e.g., at the centroid of the corresponding tooth). Additionally, when using a full complement of mesh element features, training converges more quickly. In other words, a machine learning model trained using a full complement of mesh element features tends to be faster and more accurate (at an earlier epoch) than a system that does not. For an existing system that observes a historical accuracy of 91%, a 3% increase in accuracy reduces the actual error rate by more than 30%.
[0108] Prediction models that can operate on the feature vectors of the above features include, but are not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoders, validation using autoencoders, grid segmentation, coordinate system prediction, grid cleaning, repair design generation, appliance component generation and / or placement, and dental arch morphology prediction. Such feature vectors can be presented to the input of the prediction model. In some specific implementations, such feature vectors can be presented to one or more internal layers of a neural network that is part of one or more of those prediction models.
[0109] As described herein, tooth movement specifies one or more tooth transformations, which can be encoded in various ways to specify the position and orientation of teeth within a setting and applied to a 3D representation of the teeth. For example, according to a particular implementation, the tooth position can be the Cartesian coordinates of the tooth canonical origin position defined in some semantic context. The tooth orientation can be represented as a rotation matrix, a unit quaternion, or other 3D rotation representations such as Euler angles relative to a reference frame (global or local). The dimensions are real-valued 3D spatial extents, and the gaps can be binary existence indicators or real-valued gap sizes between teeth, especially in cases where some teeth are missing. In some implementations, tooth rotation can be described by a 3×3 matrix (or a matrix of other dimensions). In some implementations, the tooth position and rotation information can be combined into the same transformation matrix (e.g., combined into a 4×4 matrix), which can reflect homogeneous coordinates. In some cases, an affine space transformation matrix can be used to describe tooth transformations, such as transformations that describe the malocclusion pose of the teeth, the intermediate pose of the teeth, and / or the final setting pose of the teeth. Some implementations can use relative coordinates, where the setting transformation is predicted relative to the malocclusion coordinate system (i.e., predicting the malocclusion-to-setting transformation rather than directly predicting the setting coordinate system). Other implementations can use absolute coordinates, where the setting coordinate system is predicted directly for each tooth. In the relative mode, the transformation can be calculated relative to the centroid of each tooth mesh (relative to the global origin), which is referred to as "relative local". Some advantages of using relative local coordinates include eliminating the need for a malocclusion coordinate system (landmark data), which may not be applicable to all patient case datasets. Some advantages of using absolute coordinates include simplifying data preprocessing, since the mesh data is initially represented relative to the global origin. In some implementations, these details regarding tooth position encoding and tooth orientation encoding can also be applied to one or more of the neural network models of the present disclosure, including but not limited to: GDL setting, RL setting, VAE setting, capsule setting, MLP setting, diffusion setting, PT setting, similarity setting, FDG setting, setting classification, setting comparison, VAE mesh element labeling, MAE mesh filling, mesh reconstruction VAE, and validation using an autoencoder.
[0110] According to a particular implementation, the convolutional layers in the various 3D neural networks described herein may use edge data to perform mesh convolutions. The use of edge information ensures that the model is insensitive to different input orders of 3D elements. In addition to or separate from using edge data, the convolutional layer may use vertex data to perform mesh convolutions. The advantage of using vertex information is that vertices are typically fewer than edges or faces, so vertex-oriented processing can result in lower processing overhead and lower computational costs. In addition to or separate from using edge data or vertex data, the convolutional layer may use face data to perform mesh convolutions. Further, in addition to or separate from using edge data, vertex data, or face data, the convolutional layer may use voxel data to perform mesh convolutions. The advantage of using voxel information is that, depending on the selected granularity, there may be far fewer voxels to process compared to vertices, edges, or faces in the mesh. Sparse processing (using voxels) may result in lower processing overhead and lower computational costs (especially in terms of computer memory or RAM usage).
[0111] Examples of oral care metrics include orthodontic metrics (OM) and restorative design metrics (RDM). RDM may describe the shape and / or form of one or more 3D representations of teeth used in dental restorations. One example use case is in creating one or more dental restorative appliances. Another example use case is in creating one or more veneers (such as zirconia veneers). Some RDM may quantify the shape and / or other characteristics of teeth. Other RDM may quantify the relationship between two or more teeth (e.g., spatial relationship). RDM differs from restorative design parameters (RDP) in that restorative design metrics define the current state of a patient's dentition, while restorative design parameters are used as specifications for a machine learning or other optimization model to generate a desired tooth shape and / or form. RDM describes the current (e.g., starting or maloccluded condition) shape of teeth. Restorative design parameters specify the expected appearance of teeth after a restorative process by an oral care provider (such as a dentist or dental technician). For the purpose of dental restorations, a neural network or other machine learning or optimization algorithm may be provided to either or both of RDM and RDP. In some implementations, RDM may be calculated with respect to a patient's pre-restorative dentition (i.e., the primary implementation). In other implementations, RDM may be calculated with respect to a patient's post-restorative dentition. The restorative design may include one or more teeth and may be referred to as a restorative arch. Restorative design generation may involve generating an improved geometry and / or structure of one or more teeth in the restorative arch.
[0112] Aspects of RDM calculations are described below. In some specific implementations, RDM can be measured, for example, by locating landmarks in the teeth (or gums, hardware, and / or other elements of the patient's dentition) and measurements of the distances between these landmarks, or otherwise with respect to these landmarks. In some specific implementations, one or more neural networks or other machine learning models can be trained to identify or extract one or more RDMs from one or more 3D representations of the teeth (or gums, hardware, and / or other elements of the patient's dentition). The techniques of the present disclosure can use RDMs in various ways. For example, in some specific implementations, one or more neural networks or other machine learning models can be trained to classify or label one or more settings, dental arches, dentitions, or other groups of teeth at least in part based on RDMs. Thus, in these examples, RDMs form part of the training data for training these models.
[0113] Aspects of a dental mesh autoencoder that can be used in accordance with the techniques of the present disclosure are described below. An autoencoder for prosthetic design generation is disclosed in U.S. Provisional Application No. US63 / 366514. The autoencoder (e.g., a variational autoencoder or VAE) takes as input a dental mesh (or other 3D representation) that reflects a malocclusion state (i.e., the tooth shape before prosthetics). The encoder component of the autoencoder encodes the dental mesh into a latent form (e.g., a latent vector). To alter the geometry and / or structure of the final reconstructed mesh, modifications can be applied to the latent vector (e.g., based on a mapping of the latent space through prior experiments). In some specific implementations, additional vectors can be included with the latent vector (e.g., by concatenation), and the resulting vector concatenation can be reconstructed by the decoder component of the autoencoder into a reconstructed dental mesh that is a copy of the input dental mesh.
[0114] In accordance with aspects of the present disclosure, RDMs and RDPs can also be used as neural network inputs during the execution phase. In some specific implementations, to tell the encoder specific information about the input 3D dental representation, one or more RDMs can be concatenated with the input to the encoder. In some specific implementations, to provide the decoder component with specific information about the input 3D dental representation, one or more RDMs can be concatenated with the latent vector before reconstruction. Additionally, in some specific implementations, to provide the encoder with specific information about the input 3D dental representation, one or more prosthetic design parameters (RDPs) can be concatenated with the input to the encoder component. Similarly, in some specific implementations, to provide the decoder with specific information about the input 3D dental representation, one or more prosthetic design parameters (RDPs) can be concatenated with the latent vector before reconstruction.
[0115] In this way, either or both of the RDM and RDP can be incorporated into the functionality of an autoencoder (e.g., a tooth reconstruction autoencoder) and used to affect the geometry and / or structure of the reconstruction restoration design (i.e., affect the shape of the tooth on the output of the autoencoder). In some embodiments, the variational autoencoder of U.S. Provisional Application No. US63 / 366514 can be replaced by a capsule autoencoder (e.g., instead of encoding a tooth mesh into a latent vector, encoding the tooth mesh into one or more latent capsules).
[0116] In some embodiments, clustering or other unsupervised techniques can be performed on the RDM to cluster one or more sets, arches, dentitions, or other groups of teeth based on the restorative characteristics of the teeth. Such clustering can be useful in treatment planning as it provides insights into categories of patients with different treatment needs. This information can be instructive to clinicians as they are aware of the possible treatment options. In some cases, best practices (such as default RDP values) can be identified for patient cases falling into one or another cluster (e.g., as determined by a similarity metric, such as in k-NN). After classifying a new case into a specific cluster, information about the relevant best practices can be provided to the clinician responsible for treating that case. In some cases, such default values may undergo further adjustment or modification.
[0117] Case assignment: Such clustering can be used to gain further insights into the types of patient cases present in a dataset. Analysis of such clustering can reveal that patient treatment cases with certain RDM values (or ranges of values) may require less treatment time (or alternatively more treatment time). Cases that require more time to treat (or are otherwise more difficult) can be assigned to experienced or senior technicians for treatment. Cases that take less time to treat can be assigned to newer or less experienced technicians for treatment. This assignment can be further aided by finding a correlation between the RDM values of certain cases and the known treatment durations associated with those cases.
[0118] The following RDMs can be measured and used in creating either or both of dental restorative appliances and veneers (veneers being a type of dental restorative appliance) with the aim of making the resulting teeth look natural. Symmetry is generally a preferred aspect. There may be differences between patients based on demographic variations. The generation of dental restorative appliances can benefit from some or all of the following RDMs. Hue and translucency can be particularly relevant to the creation of veneers, although some embodiments of dental restorative appliances may also consider this information. Examples of interdental RDMs are described below.
[0119] 1) Left - right symmetry and / or ratio: A measure of symmetry between one or more teeth on opposite sides of a tooth relative to one or more other teeth. For example, for a pair of corresponding teeth, a measure of the width of each tooth. In one case, one tooth has a normal width while the other tooth is too narrow. In another case, both teeth have normal widths. The following is a list of properties that can be measured for a tooth and compared to corresponding measurements of one or more corresponding teeth: a) Width - mesial - distal distance; b) Length - gingival - incisal distance; c) Diagonal - distance across the tooth, e.g., from the mesial gingival corner to the distal incisal corner (this measure is one of many measures that can be used to quantify the shape of a tooth other than length and width). A ratio between a and b can be calculated, such as a / b or b / a. Such a ratio can indicate whether there is spatial symmetry (e.g., by measuring the ratio a / b on the left side and measuring the ratio a / b on the right side and then comparing the left ratio and the right ratio). In some specific embodiments, the length, width, and / or ratio may not match when spatial symmetry is "off". In some specific embodiments, such a ratio can be calculated relative to a standard. Many aesthetic standards are available in dental literature. Examples include the golden ratio and the cyclic aesthetic dental ratio. In some specific embodiments, spatial symmetry can be measured on a pair of teeth, where one tooth is on the right side of the dental arch and the other tooth is on the left side of the dental arch.
[0120] 2) Proportion of adjacent teeth: Measuring the width proportion of adjacent teeth, such as measured as a projection onto a plane along the dental arch (e.g., a plane located in front of the patient's face). The ideal proportion used in the final restoration design can be, for example, the so - called golden ratio. The golden ratio is relevant to adjacent teeth, such as the central incisor and the lateral incisor. This measure involves the measurement of these proportions as they exist in a pre - restorative malocclusion. For the central incisor, lateral incisor, and canine, the ideal golden ratio on a particular side (left or right) of a particular dental arch (e.g., the upper dental arch) is 1.6, 1, 0.6. If one or more of these proportion values deviate (e.g., in the case of a "peg - shaped lateral incisor"), the patient may desire a dental restoration procedure to correct the proportion.
[0121] 3) Dental arch difference: A measure of any dimensional difference between the upper dental arch and the lower dental arch, e.g., related to the width of the teeth, for dental restoration purposes. For example, the techniques of the present disclosure can perform adjacent tooth width proportion measurements in the upper and lower dental arches. In some specific embodiments, Bolton analysis measurements can be made by measuring the upper width, the lower width, and the ratio between these quantities. In various specific embodiments, the dental arch difference can be described in absolute measurement values (e.g., in mm or other suitable units) or as a proportion or ratio.
[0122] 4) Midline: The measurement of the midline of the maxillary incisors relative to the midline of the mandibular incisors. The techniques of the present disclosure may measure the midline of the maxillary incisors relative to the midline of the nose (if data on the position of the nose is available).
[0123] 5) Proximal contact: The measurement of the size (area, volume, perimeter, etc.) of the proximal contact between adjacent teeth. Ideally, the teeth contact along the mesial / distal surfaces, and the gingiva fills in along the gingival to where the teeth contact. If the gingival tissue fails to fill the space below the proximal contact, a black triangle may form. In some cases, for teeth located closer to the posterior of the dental arch, the size of the proximal contact may gradually become shorter. In an ideal scenario, the proximal contact will be long enough such that there is an appropriately sized incisal embrasure and the gingival tissue fills the area below the contact or gingival to the contact.
[0124] 6) Embrasure: In some embodiments, the techniques of the present disclosure may measure the size (area, volume, perimeter, etc.) of the embrasure, i.e., the gap between teeth at the gingival or incisal margin. In some embodiments, the techniques of the present disclosure may measure the symmetry between embrasures on opposite sides of the dental arch. The embrasure is at least partially based on the length of the contact between teeth, and / or at least partially based on the shape of the teeth. In some cases, for teeth located closer to the posterior of the dental arch, the size of the embrasure may gradually become longer.
[0125] Examples of intra-tooth RDM are listed below, continuing the numbering of the other RDMs listed above.
[0126] 7) Length and / or width: The measurement of the length of a tooth relative to the width of that tooth. This measurement may reveal, for example, that the patient has long central incisors. The width and length are defined as: a) width - mesial to distal distance; b) length - gingival to incisal distance; c) other dimensions of the tooth body - the portion of the tooth between the gingival region and the incisal edge. In some embodiments, either or both of the length and width of a tooth may be measured and compared to the length and / or width of one or more teeth.
[0127] 8) Tooth morphology: Measurements of the major anatomical structures of tooth shape, such as line angles, facial contours, and / or incisal angles and / or embrasures. Frequency and / or dimensions can be measured. In some specific embodiments, the observed primary tooth shape aspects can be matched to one or more known patterns. The techniques of the present disclosure can measure secondary anatomical structures of tooth shape, such as incisal marginal grooves. For example, frequency and / or dimensions can be measured. In some specific embodiments, the observed secondary tooth shape aspects can be matched to one or more known patterns. In some examples, the techniques of the present disclosure can measure tertiary anatomical structures of tooth shape, such as enamel striae or striations. For example, frequency and / or dimensions can be measured. In some specific embodiments, the observed tertiary tooth shape aspects can be matched to one or more known patterns.
[0128] 9) Hue and / or translucency: Measurements of tooth hue and / or translucency. Tooth hue is typically described by the Vita Classical or 3D Master shade guides. Tooth translucency is described by transmittance or contrast ratio. Tooth hue and translucency can be evaluated (or measured) based on one or more of the following types of data related to the tooth: incisal edge, incisal third, body, and gingival third. The translucency of the enamel layer is typically higher than that of the dentin or cementum layer. In some specific embodiments, hue and translucency can be measured on a per-voxel (local) basis. In some specific embodiments, hue and translucency can be measured on a per-region basis, such as the incisor region, tooth body region, etc. The tooth body can refer to the portion of the tooth between the gingival region and the incisal edge.
[0129] 10) Contour height: Measurement of tooth contour. When viewed from the proximal view, all teeth have a specific contour or shape moving from the gingival surface to the incisal edge. This is called the facial contour of the tooth. In each tooth, there is a contour height where the shape is most prominent. This contour height varies from the teeth in the front of the dental arch to the teeth in the back of the dental arch. In some specific embodiments, the measurement can take the form of fitting a template to known dimensions and / or known ratios. In some specific embodiments, the measurement can quantify the degree of curvature along the facial tooth surface. In some specific embodiments, the position where the curvature along the tooth contour is most prominent is measured. This position can be measured as the distance from the gingival margin or the distance from the incisal edge, or as a percentage along the tooth length.
[0130] Neural networks for representation generation based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional layers, and / or pooling layers, or other models can benefit from the use of oral care variables (e.g., oral care metrics or oral care parameters). For example, oral care metrics (e.g., orthodontic metrics or prosthodontic design metrics) can convey aspects of the shape and / or structure of a patient's dentition (e.g., the shape and / or structure of a single tooth, or a particular relationship between two or more teeth) to the neural network models of the present disclosure. Each oral care metric describes different information about the patient's dentition, which may not redundantly exist in other input data provided to the neural network. For example, an "overbite" metric can quantify the overlap between the upper central incisor and the lower central incisor along the vertical Z-axis, which may not be easily determined by a conventional neural network in some embodiments. In other words, oral care metrics provide refined information about the patient's dentition that a conventional neural network (e.g., a representation generation neural network) may not be sufficiently trained or configured to extract. However, a neural network specifically trained to generate oral care metrics can overcome this shortcoming because, for example, the loss can be computed in a way that facilitates accurate oral care metric prediction. Mesh oral care metrics can provide a processed form of the structure and / or shape of a patient's dentition, data that would otherwise not be available to the neural network. This processed information is generally more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, oral care metrics have been provided to a representation generation neural network based on a U-Net model. Based on the experiments, it was found that systems using oral care metrics (e.g., "overbite", "overjet", and "canine class relationship" metrics) were at least 2.5% more accurate than systems that did not use oral care metrics. Additionally, training converged more quickly when oral care metrics were used. In other words, machine learning models trained with oral care metrics tend to be faster and more accurate (at an earlier epoch) than systems that do not. For an existing system that observed a historical accuracy of 91%, a 2.5% increase in accuracy reduced the actual error rate by almost 30%.
[0131] The entire text of the PCT application with publication number WO2020026117A1 is incorporated herein by reference. WO2020026117A1 lists some examples of orthodontic metrics (OM). Additional examples are disclosed herein. Orthodontic metrics can be used to quantify the physical arrangement of dental arches for orthodontic treatment purposes (as opposed to prosthodontic design metrics, which relate to dentistry and describe the shape and / or form of one or more pre-prosthetic teeth for supporting prosthodontic purposes). These orthodontic metrics can measure the degree of malocclusion of the dental arch, or conversely, these metrics can measure the degree of correct arrangement of the teeth. In some specific implementations, the GDL setting model (or RL setting, VAE setting, capsule setting, MLP setting, diffusion setting, PT setting, similarity setting, and FDG setting) can incorporate one or more of these orthodontic metrics, or other similar or related orthodontic metrics. In some specific implementations, such orthodontic metrics can be incorporated into the feature vectors of grid elements, where these element-based feature vectors are provided as inputs to the setting prediction network. In some specific implementations, such orthodontic metrics can be directly used as direct inputs by a generator, MLP, transformer, or other neural network (such as presented in one or more input vectors of real numbers S, as described elsewhere in this disclosure). Using such orthodontic metrics in the training of the generator can improve the performance (i.e., correctness) of the resulting generator, thereby producing a predicted transformation that places the teeth closer to the correct final setting pose than other possible scenarios. Such orthodontic metrics can be used by an encoder structure or by a U-Net structure (in the case of the GDL setting). Such orthodontic metrics can be provided by an autoencoder, variational autoencoder, masked autoencoder, or regularized autoencoder (in the case of the VAE setting, VAE grid element labeling, MAE grid filling). Such orthodontic metrics can be used by a neural network that generates action predictions as part of a reinforcement learning RL setting model. Such orthodontic metrics can be used by a classifier that applies labels to the set dental arch (e.g., labels such as misaligned, graded, or final setting). This description is non-limiting as orthodontic metrics can also be incorporated into the various techniques of this disclosure in other ways.
[0132] In some examples, the various loss calculations of the present disclosure may be combined with one or more orthodontic metrics, which has the advantage of improving the accuracy of the resulting neural network. Orthodontic metrics can be used to directly compare predicted examples with corresponding ground truth examples (such as by using the metrics set in the comparison description). In other examples, one or more orthodontic metrics can be obtained from this part and incorporated into the loss calculation. Such orthodontic metrics can be calculated on the predicted examples, and then the orthodontic metrics will also be calculated on the ground truth examples. The results of these two orthodontic metrics will then be used in the loss calculation, which has the advantage of improving the performance of the resulting neural network. In some embodiments, one or more orthodontic metrics related to the alignment of two or more adjacent teeth can be calculated and incorporated into the loss function, for example, to at least partially train the setting prediction neural network. In some embodiments, such orthodontic metrics can influence the network to align the mesial surface of a tooth with the distal surface of an adjacent tooth. Backpropagation is an example algorithm by which a neural network can be trained using one or more loss values.
[0133] In some embodiments, one or more orthodontic metrics can be used to evaluate the predicted output of a neural network, such as a setting prediction. Such metrics can enable the training algorithm to determine how close the predicted output is to an acceptable output, for example, in a quantitative sense. In some embodiments, this use of orthodontic metrics can enable the calculation of loss values that do not depend entirely on comparison with the ground truth. In some embodiments, this use of orthodontic metrics can enable the loss calculation and network training to continue without the need to compare with ground truth examples. The advantage of this method is that the loss can be calculated based on general principles or specifications of the predicted output (such as settings), rather than associating the loss calculation with a specific ground truth example (which may have been defined by a specific doctor, clinician, or technician, whose treatment concept may be different from that of other technicians or doctors). In some embodiments, such orthodontic metrics can be defined based on the FID (Frechet Inception Distance) score.
[0134] The following is a description of some orthodontic metrics for quantifying the state of a set of teeth in an arch for orthodontic treatment. These orthodontic metrics indicate the degree of malocclusion of the teeth at a given stage of treatment with a clear aligner appliance.
[0135] When training one of the neural networks of the present disclosure, it may be particularly advantageous to use orthodontic metrics calculated using tensor operations because tensor operations can facilitate efficient calculations. The more efficient (and faster) the calculations, the faster the training can proceed.
[0136] In some examples, error patterns may be identified in one or more prediction outputs of an ML model (e.g., a transformation matrix for predicting a tooth setup, markings of mesh elements for mesh cleanup, addition of mesh elements to a mesh for mesh filling purposes, classification labels for a setup, classification labels for a tooth mesh, etc.). One or more orthodontic metrics may be selected to be inputs for the next round of ML model training to address any error or defect patterns that may be identified in the one or more prediction outputs.
[0137] Some OMs may be defined relative to an arch form coordinate system (LDE coordinate system). In some embodiments, points may be described using the LDE coordinate system relative to the arch form, where L, D, and E respectively correspond to: 1) the length along the curve of the arch form, 2) the distance from the arch form, and 3) the distance in a direction perpendicular to the L-axis and the D-axis (which may be referred to as Eminence).
[0138] The various OMs and other techniques of the present disclosure may compute conflicts between 3D representations (e.g., of oral care objects such as teeth). Such conflicts may be computed as at least one of the following: 1) the penetration distance between 3D tooth representations, 2) the count of overlapping mesh elements between 3D tooth representations, and 3) the overlapping volume between 3D tooth representations. In some embodiments, an OM may be defined to quantify the conflict of two or more 3D representations of oral care structures (such as teeth). Some optimization algorithms (such as setup prediction techniques) may seek to minimize the conflicts between oral care structures (such as teeth). Inter-arch orthodontic metrics are as follows.
[0139] Six (6) metrics for comparing two or more arches are listed below. Other suitable comparative orthodontic metrics are found elsewhere in the present disclosure, such as in the section on setup comparison techniques. 1. Rotational geodesic distance (rotation between a predicted example and a ground truth setup example) 2. Translation distance (gap between a predicted example and a ground truth setup example) 3. Normalized translation distance 4. 3D alignment error, which measures the distance between a predicted mesh element and a ground truth mesh element, in millimeters. 5. Normalized 3D alignment 6. Percentage of volume overlap (% overlap) of a predicted example and the corresponding ground truth example (alternatively % overlap of mesh elements)
[0140] Intra-arch orthodontic metrics are as follows. Alignment- The mesial-distal central axis of the tooth can be used to calculate the 3D tooth orientation vector. A 3D vector that can be the tangential vector of the dental arch form at the tooth position can also be calculated. Then, the XY components (i.e., which can be 2D vectors) can be used to compare the orientation of the dental arch form at the tooth position with the tooth orientation in the XY space. Cosine similarity can be used to calculate the 2D orientation difference (angle) between the tangent of the dental arch form and the mesial-distal central axis of the tooth. Dental Arch Symmetry - For each pair of left and right teeth (e.g., the left lower lateral incisor and / or the right lower lateral incisor), the absolute difference between the X coordinate of each tooth and the X-axis of the global coordinate reference system can be calculated. This increment can indicate the dental arch asymmetry of the given tooth pair. The result of this calculation can be the average X-axis increment from one or more tooth pairs of the dental arch. In some specific implementations, this calculation can be performed relative to the Y-axis having a Y coordinate (and / or relative to the Z-axis having a Z coordinate). D-Axis Difference of Dental Arch Shape - The D-dimensional difference (i.e., the positional difference in the buccal-lingual direction) between the two dental arch states of one or more teeth can be calculated. In some specific implementations, a dictionary of the D-direction tooth movement of each tooth can be returned, where the tooth UNS number is used as the key. The LDE coordinate system relative to the dental arch form can be used. Length Ratio of Dental Arch Shape (Lower) - The ratio between the current length of the lower dental arch and the length of the dental arch when it is in the initial malocclusion of the lower dental arch can be calculated. Length Ratio of Dental Arch Shape (Upper) - The ratio between the current length of the upper dental arch and the length of the dental arch when it is in the initial malocclusion of the upper dental arch can be calculated. Parallelism of Dental Arch Shape (Full Dental Arch) - For at least one local tooth coordinate system origin in the upper dental arch, one or more nearest origins (e.g., tooth local coordinate system origins) in the lower dental arch. In some specific implementations, two nearest origins can be used. The straight-line distance from the upper dental arch point to the line formed between the origins of two teeth in the opposite (lower) dental arch can be calculated. The standard deviation of the set of the above "point-to-line" distances can be returned, where the set can be composed of the point-to-line distances of each tooth in the dental arch. Parallelism of Dental Arch Shape (Single Tooth) - This metric can share some calculation elements with the global orthodontic metric of the dental arch form parallelism, except that this metric can input the mean distance from the tooth origin to the line formed by adjacent teeth in the opposite dental arch (e.g., one tooth in the upper dental arch and the corresponding tooth in the lower dental arch). The mean distance can be calculated for one or more such tooth pairs. In some specific implementations, the mean distance can be calculated for all tooth pairs. Then, the mean distance can be subtracted from the distances calculated for each tooth pair. This OM can produce the deviation of the tooth from the "typical" tooth parallelism in the dental arch. Buccolingual Inclination- For at least one molar or premolar, find the corresponding tooth on the opposite side of the same dental arch (i.e., for a tooth on the left side of the arch, find the same type of tooth on the right side, and vice versa). This OM can calculate an n-element list for each tooth (e.g., n can be equal to 2). The list can include at least the tooth IDs of the teeth in each pair of teeth (e.g., LeftLowerFirstMolar and RightLowerFirstMolar in the list = [left_tooth_idx_1, right_tooth_idx_2]). Such an n-element vector can be calculated for each molar and each premolar in the upper and lower dental arches. Identify the buccal cusp on each molar and each premolar on each side of the left and right sides of the dental arch. Draw a line between the buccal cusp of the left tooth and the buccal cusp of the right tooth. Use this line and the z-axis of the dental arch morphology to create a plane. The lingual cusp can be projected onto this plane (i.e., at this point, the inclination angle can be determined). By performing an additional projection, the approximate perpendicular distance between the lingual cusp and the buccal cusp can be calculated. This distance can be used as the buccolingual inclination OM. Canine Overbite - The upper and lower canines can be identified. The first premolar on a given side of the mouth can be identified. On a given side of the dental arch, the distance between the upper and lower canines can be calculated, and the distance between the upper and lower first premolars can also be calculated. An average value (or median, or mode, or some other statistical value) can be calculated for the measured distances. The z-component of this result indicates the degree of overbite. Overbite can be calculated between any tooth in one dental arch and the corresponding tooth in the other dental arch. Canine Overjet Contact - The conflict (e.g., conflict distance) between the pairs of canines on the opposing dental arches can be calculated. Canine Overjet Contact KDE - The orthodontic metric score of the current patient case can be taken as input, and this score can be converted to a log-likelihood using a previously trained kernel density estimation (KDE) model or distribution. This operation can produce information about where the patient case lies in the distribution of "typical" values. Canine Overjet - This OM can share some calculation steps with the canine overbite OM. In some specific implementations, the average distance can be calculated. In some specific implementations, the distance calculation can calculate the Euclidean distance of the XY components of a tooth in the upper dental arch and a tooth in the lower dental arch to produce coverage (i.e., as opposed to calculating the difference in the Z component, as can be performed for canine overbite). Coverage can be calculated between any tooth in one dental arch and the corresponding tooth in the other dental arch. Canine Category Relationship (Also Applicable to First, Second, and Third Molars) - In some specific implementations, this OM can include two functions (e.g., written in Python). get_canine_landmarks(): Obtain the landmarks for each tooth. These landmarks can be used to calculate class relationships. Then, in some specific implementations, these landmarks are mapped to a global coordinate space so that measurements can be made between teeth. class_relationship_score_by_side(): Can calculate the average position of at least one landmark on at least one tooth in the lower dental arch, and can calculate this value for the upper dental arch. Then, a vector can be calculated from the upper dental arch landmark position to the lower dental arch landmark position, and finally, this vector is projected onto the lower dental arch to produce a quantification (e.g., as a scalar) of the amount of increment in the "dental arch l-axis" position. This OM can calculate how far a tooth is positioned in front of or behind one or more teeth of interest in the opposing dental arch along the l-axis. Interdigitation - By finding the midpoint between the distal marginal ridge saddle and the mesial marginal ridge saddle of a tooth, the fossa in at least one upper molar can be located. The cusp of the lower molar can be located between the marginal ridges of the corresponding upper molar. This OM can calculate the vector from the midpoint of the upper molar fossa to the cusp of the lower molar. This vector can be projected onto the d-axis of the dental arch form, thereby producing a lateral measurement of the distance from the cusp to the fossa. This distance can define the overbite magnitude. Edge Alignment - This OM can identify the leftmost and rightmost sides of a tooth, and can identify the leftmost and rightmost sides of the adjacent teeth of that tooth. The OM can then draw a vector from the leftmost side of a tooth to the leftmost side of the adjacent tooth of that tooth. The OM can then draw a vector from the rightmost side of a tooth to the rightmost side of the adjacent tooth of that tooth. The OM can then calculate the linear fitting error between the two vectors. This calculation can involve generating two vectors: Vec_tooth = right_tooths_leftside to left_tooths_leftside Vec_neighbor = right_tooths_rightside to left_tooths_leftside Then it can involve calculating the dot product of these two vectors and subtracting the result from 1. (That is, side alignment score = 1 - abs(dot(Vec_tooth, Vec_neighbor))). A score of 0 can indicate perfect alignment. A score of 1 can mean perpendicular alignment. Incisor Interarch Contact KDE - The deviation of the incisor interarch contact from the mean of the modeled distribution of this statistical information can be identified in the dataset of one or more other patient cases. Leveling - It can calculate a measure of leveling between a tooth and its adjacent teeth. This OM can calculate the height difference between two or more adjacent teeth. For molars, this OM can use the midpoint between the mesial and distal saddle ridges as the height of the molar. For non - molars, this OM can use the crown length from the gingiva to the tip. In some specific embodiments, the tip can be the origin of the local coordinate space of the tooth. Other specific embodiments can place the origin at other positions. A simple subtraction between the heights of adjacent teeth can produce the leveling increment between the teeth (e.g., by comparing the Z - components). Midline - It can calculate the position of the mid - line of the upper incisors and / or lower incisors, and then the distance between them can be calculated. Molar Interarch Contact KDE - It can calculate the inter - molar arch contact fraction (i.e., interference depth or other types of interference), and then the position of this fraction in a predefined KDE (distribution) constructed from representative cases can be identified. Occlusal Contact - For a specific tooth from an arch, this OM can identify one or more landmarks (e.g., mesial cusp or central cusp, etc.). Obtain the tooth transformation of this tooth. For each cusp on the current tooth, the cusp can be scored according to the degree of contact between the cusp and the adjacent (corresponding) tooth in the opposite arch. A vector from the cusp of the tooth under discussion to the vertical intersection point in the corresponding tooth of the opposite arch can be found. The distance and / or direction (i.e., up or down) to the opposite arch can be calculated. A list including the resulting signed distances, one for each cusp on the tooth under discussion, can be returned. Overbite - It can compare the upper and lower central incisors along the z - axis. The difference along the z - axis can be used as the overbite fraction. Overjet - It can compare the upper and lower central incisors along the y - axis. The difference along the y - axis can be used as the overjet fraction. Molar Interarch Contact - It can calculate the contact fraction between molars and can use interference measurements (such as interference depth). Root Movement d - It can receive the tooth transformations of the initial state and the next state. It can calculate the arch - form axis at point L along the arch form. This OM can return the distance moved along the d - axis. This can be achieved by projecting the root pivot point onto the d - axis. Root Movement l - It can receive the tooth transformations of the initial state and the next state. It can calculate the arch - form axis at point L along the arch form. This OM can return the distance moved along the l - axis. This can be achieved by projecting the root pivot point onto the l - axis. Spacing- The spacing between each tooth and its adjacent teeth can be calculated. Transformations and meshes for the dental arch can be received. The left and right sides of each tooth mesh can be calculated. One or more points of interest can be transformed from local coordinates to the global dental arch coordinate system. The spacing can be calculated in a plane (e.g., the XY plane) between each tooth and its adjacent tooth on the "left side". An array of one or more Euclidean distances (e.g., such as in the XY plane) can be returned, which can represent the spacing between each tooth and its adjacent tooth on the left side. Torque - Torque (i.e., rotation about an axis such as the x-axis) can be calculated. For one or more teeth, one or more rotations can be converted from Euler angles to one or more rotation matrices. The components of the rotation (such as the x-component) can be extracted and converted back to Euler angles. The x-component can be interpreted as the torque of the tooth. A list including the torque of one or more teeth can be returned, and this list can be indexed by the UNS number of the teeth.
[0141] The neural network of the present disclosure can utilize one or more benefits of parameter tuning operations, thereby optimizing the input and parameters of the neural network to produce more data-precise results. One parameter that can be tuned is the neural network learning rate (e.g., it can have values such as 0.1, 0.01, 0.001, etc.). Data augmentation schemes can also be tuned or optimized, such as a scheme of adding "shiver" to the tooth mesh before input to the neural network (i.e., small random rotations, translations, and / or scalings can be applied to change the dataset and make the neural network robust to data variations). The subset of neural network model parameters available for tuning is as follows: ○ Learning rate (LR) decay rate (e.g., how much the LR decays during a training run) ○ Learning rate (LR). A floating-point value used by the optimizer (e.g., 0.001). ○ LR scheduling (e.g., cosine annealing, step, exponential) ○ Voxel size (for the case of sparse mesh processing operations) ○ Dropout % (e.g., dropout that can be performed in a linear encoder) ○ LR decay step size (e.g., decay every 10 or 20 or 30 epochs) ○ Model scaling, which can increase or decrease the layer count and / or the parameter count per layer.
[0142] Parameter tuning can be advantageously applied to the training of neural networks to predict final settings or intermediate gradings, thereby providing technical improvements in data-oriented accuracy. Parameter tuning can also be advantageously applied to the training of neural networks for mesh element tagging or for mesh filling. In some examples, parameter tuning can be advantageously applied to the training of neural networks for tooth reconstruction. In terms of the classifier model of the present disclosure, parameter tuning can be advantageously applied to neural networks for the classification of one or more settings (i.e., the classification of one or more arrangements of teeth). The advantage of parameter tuning is to improve the data accuracy of the output of the prediction model or classification model. In some cases, parameter tuning can provide the advantage of obtaining the last remaining few percentage points of validation accuracy from the prediction or classification model.
[0143] Various neural network models of the present disclosure can benefit from data augmentation. Examples include models trained on 3D meshes, such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, FDG settings, setting classification, setting comparison, VAE mesh element tagging, MAE mesh filling, mesh reconstruction VAE, and validation using autoencoders. Figure 2 is a method illustrating the data augmentation method of the present disclosure. Such as by Figure 2 Data augmentation by the method shown can increase the size of the training dataset of the dental arch. Data augmentation can provide additional training examples by adding random rotations, translations, and / or rescaling to copies of the existing dental arch. In some specific implementations of the technology of the present disclosure, data augmentation can be performed by perturbing or jittering the vertices of the mesh in a manner similar to that described in (“Equidistant and Uniform Data Augmentation for 3D Objects”, IEEE Access, Digital Object Identifier 10.1109 / ACCESS.2021.3138162). The position of the vertices can be perturbed by adding Gaussian noise, for example, with a zero mean and a standard deviation of 0.1. According to the technology of the present disclosure, other mean and standard deviation values are possible.
[0144] Figure 2 Illustrates a data augmentation method to which the system of the present disclosure can be applied to 3D oral care representations. Non-limiting examples of 3D oral care representations are a tooth mesh or a set of tooth meshes. Tooth data 200 (e.g., a 3D mesh) is received at the input. The system of the present disclosure can generate a copy (202) of the tooth data 200. In Figure 2 an example, the system of the present disclosure can apply one or more random rotations to the tooth data 200 (204). In Figure 2In an example, the system of the present disclosure can apply random translation to the tooth data 200 (206). The system of the present disclosure can apply a random scaling operation to the tooth data 200 (208). The system of the present disclosure can apply random perturbations to one or more mesh elements of the tooth data 200 (210). The system of the present disclosure can output enhanced tooth data 212 formed by Figure 2 the method of.
[0145] Some techniques of the present disclosure, such as setting comparison techniques and setting prediction techniques (e.g., such as GDL settings, MLP settings, VAE settings, etc.), can benefit from processing steps that can align (or register) dental arches (e.g., where teeth can be represented by 3D point clouds or some other type of 3D representation described herein). Such a processing setting can be used, for example, to register a reference ground truth set dental arch from a patient case with a malocclusion dental arch from the same case, and then use these malocclusion dental arches and the reference ground truth set dental arch for training a setting prediction neural network model. Such steps can help with loss calculation because the predicted dental arch (e.g., the dental arch output by the generator) can be better aligned with the reference ground truth set dental arch, which is a condition that can facilitate the calculation of reconstruction loss, representation loss, L1 loss, L2 loss, MSE loss, and / or other types of losses described herein. In some specific implementations, the iterative closest point (ICP) technique can be used for such registration. ICP can minimize the squared error between corresponding entities such as 3D representations. In some specific implementations, linear least squares calculations can be performed. In some specific implementations, non-linear least squares calculations can be performed. Various registration models can incorporate in whole or in part portions of the following algorithms: Levenberg-Marquardt ICP, least squares rigid transformation, robust rigid transformation, random sample consensus (RANSAC) ICP, K-means based RANSAC ICP, and generalized ICP (GICP). In some cases, registration can help reduce subjectivity and / or randomness, which in some cases can occur in the design of a reference ground truth set designed by a technician (i.e., two technicians may produce different but valid final setting outputs for the same case) or other optimization techniques.
[0146] Since the generator network of the present disclosure can be implemented as one or more neural networks, the generator may include activation functions. When executed, an activation function outputs a determination as to whether a neuron in the neural network will fire (e.g., send an output to the next layer). Some activation functions may include the binary step function or the linear activation function. Other activation functions impart non-linear behavior to the network, including: the sigmoid / logistic activation function, the Tanh (hyperbolic tangent) function, the rectified linear unit (ReLU), the leaky ReLU function, the parametric ReLU function, the exponential linear unit (ELU), the softmax function, the swish function, the Gaussian error linear unit (GELU), or the scaled exponential linear unit (SELU). The linear activation function may be well-suited for some regression applications (and other applications) in the output layer. In the output layer, the sigmoid / logistic activation function may be well-suited for certain binary classification applications (and other applications). The sigmoid activation function may be well-suited for some multi-class classification applications (and other applications) in the output layer. In the output layer, the sigmoid activation function may be well-suited for some multi-label classification applications (and other applications). The ReLU activation function may be well-suited for some convolutional neural network (CNN) applications (and other applications) in the hidden layer. The Tanh and / or sigmoid activation functions may be well-suited for some recurrent neural network (RNN) applications (and other applications) in, for example, the hidden layer. There are a variety of optimization algorithms that can be used to train the neural networks of the present disclosure (such as updating neural network weights), including gradient descent (which uses first-order derivatives to determine the training gradient and is commonly used in the training of neural networks), Newton's method (which may use second-order derivatives in the loss calculation to find a better training direction than gradient descent but may require calculations involving the Hessian matrix), and the conjugate gradient method (which may converge faster than gradient descent but does not require the Hessian matrix calculations that Newton's method may require). In some specific implementations, in addition to or instead of the above techniques, additional methods may be employed to update the weights. These additional methods include the Levenberg-Marquardt method and / or simulated annealing. The backpropagation algorithm is used to distribute the results of the loss calculation back into the network so that the network weights can be adjusted and learning can occur.
[0147] Neural networks contribute to the functionality of the applications of the present disclosure, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoders, verification using autoencoders, estimation of oral care parameters, 3D grid segmentation (3D representation segmentation), coordinate system prediction, grid cleaning, restoration design generation, appliance component generation and / or placement, or dental arch form prediction. The neural networks of the present disclosure can embody parts or all of various different neural network models. Examples include U-Net architectures, multi-layer perceptrons (MLPs), transformers, pyramid architectures, recurrent neural networks (RNNs), autoencoders, variational autoencoders, regularized autoencoders, conditional autoencoders, capsule networks, capsule autoencoders, stacked capsule autoencoders, denoising autoencoders, sparse autoencoders, conditional autoencoders, long / short-term memory (LSTM), gated recurrent units (GRUs), deep belief networks (DBNs), deep convolutional networks (DCNs), deep convolutional inverse graphics networks (DCIGNs), liquid state machines (LSMs), extreme learning machines (ELMs), echo state networks (ESNs), deep residual networks (DRNs), Kohonen networks (KNs), neural Turing machines (NTMs), or generative adversarial networks (GANs). In some specific implementations, an encoder structure or a decoder structure can be used. Each of these models offers one or more of its own specific advantages. For example, a particular neural network architecture may be particularly suitable for a particular ML technique. For example, autoencoders are particularly suitable for the classification of 3D oral care representations due to their ability to transform 3D oral care representations into a form that is easier to classify.
[0148] In some specific implementations, the neural networks of the present disclosure may be suitable for operating on 3D point cloud data (alternatively, on 3D meshes or 3D voxelized representations). Many neural network specific implementations can be applied to the processing of 3D representations and can be applied to training prediction and / or generation models for oral care applications, including: PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN, and DSG-Net. Oral care applications include but are not limited to: setting prediction (e.g., using VAEs, RLs, MLPs, GDLs, capsules, diffusion, etc. trained for setting prediction), 3D representation segmentation, 3D representation coordinate system prediction, element tagging for 3D representation cleaning (VAEs for mesh element tagging), filling missing elements in 3D representations (MAEs for mesh filling), dental restoration design generation, setting classification, appliance component generation and / or placement, dental arch form prediction, estimation of oral care parameters, setting verification or other verification applications, and 3D dental representation classification.
[0149] Some specific implementations of the techniques of the present disclosure incorporate the use of autoencoders. Autoencoders that can be used in accordance with aspects of the present disclosure include but are not limited to: AtlasNet, FoldingNet, and 3D-PointCapsNet. Some autoencoders can be implemented based on PointNet.
[0150] Representation learning can be applied to the setting prediction techniques of the present disclosure by training a neural network to learn a representation of a tooth and then using another neural network to generate a transformation of the tooth. Some specific implementations can use a VAE or a capsule autoencoder to generate a representation of the reconstructed features (in some cases, including information about the structure of a tooth mesh) of one or more meshes relevant to the field of oral care. Then, this representation (latent vector or latent capsule) can be used as an input to a module that generates one or more transformations of one or more teeth. In some specific implementations, these transformations can place the teeth in a final setting pose. In some specific implementations, these transformations can place the teeth in an intermediate graded pose. In some specific implementations, the transformation can be described by a 9×1 transformation vector (e.g., specifying a translation vector and a quaternion). In other specific implementations, the transformation can be described by a transformation matrix (e.g., a 4×4 affine transformation matrix).
[0151] In some embodiments, the systems of the present disclosure may perform principal component analysis (PCA) on the oral care grid and use the resulting principal components as at least part of the representation of the oral care grid in subsequent machine learning and / or other predictive or generative processes.
[0152] An autoencoder may be trained to generate a latent form of a 3D oral care representation. The autoencoder may include a 3D encoder that encodes the 3D oral care representation into a latent form and / or a 3D decoder that reconstructs the latent form into a copy of the input 3D oral care representation. Although the present disclosure refers to a 3D encoder and a 3D decoder, the term 3D should be interpreted in a non-limiting manner to cover multi-dimensional operating modes. For example, the systems of the present disclosure may train a multi-dimensional encoder and / or a multi-dimensional decoder.
[0153] The systems of the present disclosure may implement end-to-end training. Some end-to-end training techniques of the present disclosure may involve two or more neural networks, where the two or more neural networks are trained together (i.e., the weights are updated simultaneously during the processing of each batch of input oral care data). In some embodiments, end-to-end training may be applied to pose prediction by simultaneously training a neural network that learns the representation of teeth and a neural network that can generate tooth transformations.
[0154] According to some transfer learning embodiments of the present disclosure, a neural network (e.g., a U-Net) may be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task may be executed to provide one or more initial neural network weights for training another neural network that is trained to perform a second task (e.g., pose prediction). The first network may learn low-level neural network features of the oral care grid and is shown to perform well in the first task. By using the first network as a starting point for training, the second network may exhibit faster training and / or improved performance. Certain layers may be trained to encode the neural network features of the oral care grid in the training dataset. These layers may then be fixed (or undergo minor changes during training) and combined with other neural network components (such as additional layers) that are trained for one or more oral care tasks (such as pose prediction). In this way, a portion of the neural network for one or more techniques of the present disclosure (e.g., pose prediction) may receive initial training on another task, which may result in significant learning in the trained network layers. This encoded learning may then be built upon by further task-specific training of another network.
[0155] According to the present disclosure, transfer learning can be used for setting prediction and for other oral care applications such as mesh classification (e.g., tooth or setting classification), mesh element labeling, mesh element filling, procedure parameter estimation, mesh segmentation, coordinate system prediction, restoration design generation, mesh validation (for any of the applications disclosed herein). In some specific implementations, a neural network trained to output predictions based on an oral care mesh can be partially trained first on one of the following publicly available datasets before being further trained on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, Thingi10K dataset (which is particularly relevant for 3D printed component validation), ABC: Large CAD Model Dataset for Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Component Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.
[0156] In some specific implementations, a neural network previously trained on a first dataset (oral care data or other data) can subsequently receive further training on oral care data and be applied to oral care applications such as setting prediction. Transfer learning can be used to further train any one of the following networks: GCN (Graph Convolutional Network), PointNet, ResNet, or any other neural network from the published literature listed above.
[0157] In some specific implementations, a first neural network can be trained to predict the coordinate system of a tooth (such as by using the techniques described in WO2022123402A1 or U.S. Provisional Application No. US63 / 366492). According to any one of the setting prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein), a second neural network can be trained for setting prediction. Transfer learning can assign at least a portion of the knowledge or capabilities of the first neural network to the second neural network. Thus, transfer learning can provide an accelerated training phase for the second neural network to reach convergence. In some specific implementations, the training of the second network can be completed after being enhanced with transfer learning and then using one or more techniques of the present disclosure.
[0158] The systems of the present disclosure may utilize representation learning to train ML models. Advantages of representation learning include that, as opposed to receiving inputs with variable sizes or structures, generative networks (e.g., neural networks used for predicting transformations in settings of prediction) may be configured to receive inputs with known sizes and / or standard formats. Representation learning may yield performance superior to other techniques because noise in the input data may be reduced (e.g., because the representation generation model extracts hierarchical neural network features and / or reconstruction properties of the input representation (e.g., mesh or point cloud) via loss calculation or a network architecture selected for that purpose).
[0159] The reconstruction properties may include values in a latent representation (e.g., latent vector) that describe aspects of the shape and / or structure of the 3D representation provided to the representation generation module that generated the latent representation. For example, the weights of the encoder module of a reconstruction autoencoder may be trained to encode a 3D representation (e.g., 3D mesh or others described herein) into a latent vector representation (e.g., latent vector). In other words, the ability to encode a large set of mesh elements (e.g., hundreds, thousands, or millions) into a latent vector (e.g., hundreds or thousands of real values, e.g., 512, 1024, etc.) may be learned via the weights of the encoder. Each dimension of the latent vector may include a real number that describes some aspect of the shape and / or structure of the initial 3D representation. The weights of the decoder module of the reconstruction autoencoder may be trained to reconstruct the latent vector into a close replica of the initial 3D representation. In other words, the decoder may learn the ability to interpret the dimensions of the latent vector and decode the values within those dimensions. Generally speaking, the encoder and decoder neural network modules are trained to perform a mapping of the 3D representation to the latent vector, which may then be mapped back (or otherwise reconstructed) to a 3D representation that is substantially similar to the initial 3D representation for which the latent vector was generated.
[0160] Returning to loss calculation, examples of loss calculation can include KL divergence loss, reconstruction loss, or other losses disclosed herein. Representation learning can reduce the size of the dataset required to train a model because the representation model learns a representation such that the generation network can focus on learning the generation task. Since meaningful neural network features of the input data (e.g., local and / or global features) are available to the generation network, the result can be improved model generalization. In other words, the first network can learn a representation and the second network can make prediction decisions. By training two networks to perform their own separate tasks, each network can generate more accurate results for its corresponding task than a single network trained to both learn a representation and make decisions. In some cases, transfer learning can first train a representation generation model. Then that representation generation model (either wholly or in part) can be used to pre-train subsequent models, such as a generation model (e.g., generation transformation prediction). The representation generation model can benefit from using grid element features as input to improve the ability of the second ML module to encode the structure and / or shape of the input 3D oral care representation in the training dataset.
[0161] One or more neural network models of the present disclosure can have attention gates integrated therein. Attention gate integration provides an enhancement that enables the associated neural network architecture to focus resources on one or more input values. In some embodiments, the attention gate can be integrated with a U-Net architecture, which has the advantage of enabling the U-Net to focus on certain inputs, such as input landmarks corresponding to teeth that are intended to be fixed (e.g., to prevent movement) during an orthodontic procedure (or in cases where other special handling is required). In accordance with aspects of the present disclosure, the attention gate can also be integrated with an encoder or with an autoencoder (such as a VAE or capsule autoencoder) to improve prediction accuracy. For example, the attention gate can be used to configure a machine learning model to give higher weight to aspects of the data that are more likely to be related to the correctly generated output. In this way, and because the machine learning models configured with these attention gates (or mechanisms) utilize aspects of the data that are more likely to be related to the correctly generated output, the final prediction accuracy of those machine learning models is improved.
[0162] The quality and composition of the training dataset for a neural network can affect the performance of the neural network during its execution phase. Dataset screening and outlier removal can be advantageously applied to the training of neural networks for various techniques of the present disclosure (e.g., for predictions for final settings or intermediate gradings, for neural networks for grid element labeling or for grid filling, for tooth reconstruction, for 3D grid classification, etc.) because dataset screening and outlier removal can remove noise from the dataset. Although the mechanisms for achieving the improvement are different from using attention gates, the end result is that the method allows the machine learning model to focus on the relevant aspects of the dataset and can lead to an improvement in accuracy similar to that achieved with attention gates.
[0163] In the case of a neural network configured to predict a final setting, a patient case may include at least one of a set of segmented tooth meshes of the patient, a malalignment transformation of each tooth, and / or a ground truth setting transformation of each tooth. In the case of a neural network predicting a set of intermediate stage settings, a patient case may include at least one of a set of segmented tooth meshes of the patient, a malalignment transformation of each tooth, and / or a set of ground truth intermediate stage transformations of each tooth. In some embodiments, the training dataset may exclude patient cases of the contact passive phase (i.e., the phase where the teeth of the dental arch do not move). In some embodiments, the dataset may exclude cases where there is a passive phase at the end of the process. In some embodiments, the dataset may exclude cases where there is overcrowding at the end of the process (i.e., the case where an oral care provider such as an orthodontist or dentist has selected a final setting where the tooth meshes overlap to some extent). In some embodiments, the dataset may exclude cases of a particular difficulty level (or levels) (e.g., easy, medium, and difficult).
[0164] In some embodiments, the dataset may include cases with zero pinned teeth (or may include cases with at least one pinned tooth). A person skilled in the art may specify the pinned teeth when designing the process to prevent various tools from moving that particular tooth. In some embodiments, the dataset may exclude cases with no fixed teeth (conversely, where at least one tooth is fixed). Fixed teeth may be defined as teeth that should not move during the process. In some embodiments, the dataset may exclude cases with no pontic teeth (conversely, cases where at least one tooth is a pontic). Pontic teeth may be described as "ghost" teeth, which are represented in the digital model of the dental arch but do not actually exist in the patient's dentition, or where there may be small teeth or partial teeth that may benefit from future work (such as adding composite materials through a prosthodontic appliance). The advantage of including pontic teeth in a patient's case is to leave space in the dental arch as part of the plan for the movement of other teeth during orthodontic treatment. In some cases, pontic teeth may save space in the patient's dentition for future dental or orthodontic work, such as installing implants or crowns, or applying prosthodontic appliances, such as adding composite materials to existing teeth that are too small or have an undesirable shape.
[0165] In some specific implementations, the dataset may exclude cases where the patient does not meet the age requirement (e.g., less than 12 years old). In some specific implementations, the dataset may exclude cases where the interproximal reduction (IPR) exceeds a certain threshold amount (e.g., greater than 1.0 mm). The dataset for training a neural network to predict the settings of a clear tray appliance (CTA) may exclude patient cases unrelated to CTA treatment. The dataset for training a neural network to predict the settings of an indirectly bonded tray product may exclude cases unrelated to indirectly bonded tray treatment. In some specific implementations, the dataset may exclude cases where only certain teeth are treated. In such specific implementations, the dataset may include only cases where at least one of the following is treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and / or canines.
[0166] The mesh comparison module may compare two or more meshes, for example, for the calculation of a loss function or for the calculation of a reconstruction error. Some specific implementations may involve the comparison of the volumes and / or areas of two meshes. Some specific implementations may involve calculating the minimum distance between corresponding vertices / faces / edges / voxels of two meshes. For a point in one mesh (e.g., a vertex, the midpoint on an edge, or the center of a triangle), the minimum distance between that point and the corresponding point in the other mesh is calculated. In cases where the other mesh has a different number of elements or there is no clear mapping between corresponding points of the two meshes, different methods may be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools that can be used in the mesh comparison module of the present disclosure. In some specific implementations, the Hausdorff distance may be calculated to quantify the shape difference between two meshes. The open-source software tool Metro developed by the Visual Computing Lab can also be used in quantifying the difference between two meshes. The following paper describes the method employed by Metro, which can be modified by the neural network applications of the present disclosure for mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces", P. Cignoni, C. Rocchini, and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, Vol. 17(2), June 1998, pp. 167-174.
[0167] Some techniques of the present disclosure may combine the following operations: for one or more points on a first mesh, project a ray perpendicular to the mesh surface and calculate the distance before the ray impinges on a second mesh. The length of the resulting line segment can be used to quantify the distance between the meshes. According to some techniques of the present disclosure, a color may be assigned to the distance based on the magnitude of the distance, and the color may be applied to the first mesh by means of visualization.
[0168] Oral care parameters may include one or more values specifying orthodontic protocol parameters or restoration design parameters (RDPs), as described herein. Oral care parameters may define one or more expected aspects of a 3D oral care representation and may be provided to an ML model to facilitate the ML model generating an output that can be used to generate an oral care appliance suitable for treating a patient. Other types of values include doctor preferences and restoration design preferences, as described herein. Doctor preferences and restoration design preferences may define typical treatment choices or practices of a particular clinician. Restoration design preferences are subjective to a particular clinician and thus are different from restoration design parameters. In some embodiments, doctor preferences or restoration design preferences may be calculated by unsupervised means such as clustering, such that typical values used by the clinician in patient treatment can be determined. Those typical values may be stored in a data store and invoked to be provided to an automated ML model as default values (e.g., default values that can be modified prior to executing the model).
[0169] For example, when faced with a similar diagnosis or treatment scenario, one clinician may prefer one value of a restoration design parameter (RDP), while another clinician may prefer a different value of the RDP. An example of such an RDP is a dental restoration style. In some embodiments, protocol parameters and / or doctor preferences may be provided to a setup prediction model for orthodontic treatment for the purpose of improving the customization of the resulting orthodontic appliance. In some embodiments, restoration design parameters and doctor restoration preferences may be used to design the tooth geometry used in creating a dental restoration appliance for the purpose of improving the customization of the appliance. In addition to oral care parameters, doctor preferences, and doctor restoration preferences, some embodiments of the ML prediction models of the present disclosure in orthodontic treatment may also take a setup (e.g., the arrangement of teeth) as an input. In some such embodiments, the ML prediction models of the present disclosure may take a final setup (i.e., the final arrangement of teeth) as an input, such as in the case of a prediction model trained to generate intermediate-stage predictions. For simplicity, these preferences are referred to as doctor restoration preferences, but are intended to be used in a non-limiting sense. Specifically, it should be understood that these preferences may be specified by any treating or other appropriate healthcare professional and are not intended to be limited to doctor preferences per se (i.e., preferences from a person with a Doctor of Medicine or equivalent degree).
[0170] Oral care professionals or clinicians, such as dentists or orthodontists, can specify information regarding patient handling in the form of a patient-specific set of protocol parameters. In some cases, the oral care professional may specify a set of general preferences (also known as doctor preferences) for a large number of cases to be used as default values during the process of specifying the set of protocol parameters. In some specific implementations, oral care parameters can be incorporated into the techniques described in this disclosure, such as one or more of GDL settings, VAE settings, RL settings, setting comparison, setting classification, VAE grid element tagging, MAE grid filling, validation using autoencoders, estimation of missing protocol parameter values, metric visualization, or FDG settings. One or more of these models can take as input one or more protocol parameter vectors K and / or one or more doctor preference vectors L. In some specific implementations, one or more of these models can introduce one or more protocol parameter vectors K and / or one or more doctor preference vectors L into the hidden layer of a neural network. In some specific implementations, one or more of these models can introduce either or both of K and L into a mathematical calculation, such as a force calculation, for the purpose of improving the calculation and the resulting final customization of the appliance to the patient.
[0171] Some specific implementations of neural networks for predicting settings, such as GDL settings, VAE settings, or RL settings, can incorporate information from oral care professionals (aka doctors). This information can affect the arrangement of teeth in the final setting, bringing the position and orientation of the teeth into compliance with the specifications set by the doctor and within tolerances. In some specific implementations of the GDL setting model, oral care parameters can be provided directly to the generator network as a separate input along with the grid data. In some specific implementations of GDL settings, oral care parameters can be incorporated into the feature vector calculated for each grid element before the grid element is input to the generator for processing. Some specific implementations of the VAE setting model can incorporate oral care parameters into the setting prediction. In some specific implementations, the protocol parameter K and / or the doctor preference information L can be concatenated with the latent space vector C. Doctor preferences (e.g., in an orthodontic context) and / or doctor restoration preferences can be indicated in a treatment form, or they can be based on characteristics in the treatment plan, such as final setting characteristics (e.g., the amount of bite correction or midline correction in the planned final setting), intermediate staging characteristics (e.g., treatment duration, tooth movement scenario, or overcorrection strategy), or outcomes (e.g., the number of corrections / improvements).
[0172] The prosthetic treatment of a patient can involve the specification of one or more of the following: prosthetic guidelines, prosthetic design parameters, and / or prosthetic rules for modifying one or more aspects of the patient's dentition. One or more of many possible factors may be considered when designing a 3D prosthesis, whether from an aesthetic perspective and / or from a technical perspective. For example, from an aesthetic perspective, the dental and facial midlines and angles can provide overall guidance, such as the amount of teeth visible to others when the lips are at rest and / or smiling. After considering these criteria, a set of "golden ratios" can also account for the aesthetic design of overall tooth size. The tooth-to-tooth ratio can be configured to reflect these "golden ratios", which are 1.618:1.0:0.618 for the central incisors, lateral incisors, and canines, respectively. In some particular implementations, actual values can be specified for one or more of the RDPs and received at the input of a dental prosthetic design prediction model (e.g., a machine learning model that predicts the final tooth shape when the prosthetic design is complete). In some particular implementations, one or more RDPs can be defined corresponding to one or more prosthetic design metrics (RDMs).
[0173] Constraints based on tooth position (i.e., malocclusion) and orientation (i.e., rotation and tilt) are balanced against the attempt to achieve proper symmetry, tooth proportion, and tooth-to-tooth ratio. After establishing these parameters, various tooth shapes can be utilized to match the overall aesthetics of the patient's face and smile. For example, the tooth shapes can be generally rectangular with squared edges, or they can be generally oval with rounded edges. Additionally, the tooth-to-tooth ratio can be manipulated to achieve different overall aesthetics. 3D dental CAD programs typically provide a library of different tooth "styles" to choose from and the ability for the designer to adjust the results to best match the aesthetic and medical requirements of the doctor and patient. In some examples, symmetry may be observed, as the left side should mirror the right side, and symmetry can thus be measured.
[0174] The tooth length, width, and aesthetic relationship of width to length can be specified for one or more teeth. In one example, the length of the maxillary central incisor can be set to 11 mm, and the aesthetic relationship of width to length can be set to 70% or 80%. In some examples, the lateral incisor can be 1.0 mm to 2.5 mm shorter than the central incisor. In some cases, the canine can be 0.5 mm to 1.0 mm shorter than the central incisor. Other ratios and measurements are possible for various teeth.
[0175] From a technical perspective, there are other considerations to take into account. For example, a prosthesis made of a given material must have sufficient thickness to have the necessary mechanical strength for long-term use. Additionally, the width and shape of the teeth must be designed to provide proper contact with adjacent teeth.
[0176] The example style options in the following list are from the LVI Standard, from the Las Vegas Institute for Advanced Dental Studies (LVI). Other style guides are available commercially or for free.
[0177] In these and / or other examples, the neural network engine of the present disclosure can be combined with one or more of the accepted "golden ratio" guidelines for tooth size, the accepted "ideal" tooth shape, patient preferences, practitioner preferences, etc. as inputs.
[0178] Restorative design parameters (RDPs) can be used to encode aspects of the smile design guidelines described herein, such as parameters related to the expected size of the restored teeth. The restorative design parameters are intended to serve as an indication and / or specification of the shape and / or form that one or more teeth should present after completion of a dental restoration process. One or more RDPs can be received by a neural network or other machine learning or optimization algorithms for dental restoration design, which has the advantage of providing guidance for the optimization algorithm. Some neural networks can be trained for dental restoration design generation, such as some examples of GANs or autoencoders. In some cases, the dental restoration design can be used to define the target shape of one or more teeth for generating dental restoration appliances. In some cases, the dental restoration design can be used to define the target tooth shape for generating one or more veneers.
[0179] A partial list of tooth dimensions can include length, width, height, perimeter, diameter, diagonal measurement, volume, and any of these dimensions can be normalized relative to another tooth or multiple teeth. In some specific implementations, one or more restorative design parameters can be defined that relate to the gap between two or more teeth and the size of the gap (if any) that the patient wishes to retain after treatment (e.g., such as when the patient wishes to retain a small gap between the maxillary central incisors).
[0180] Additional restorative design parameters can include the parameters specified in the following table. In the case where one parameter contradicts another parameter, the following order can determine the precedence (i.e., let the first parameter in the following list be considered authoritative). If a parameter value is not specified, the parameter can be ignored. In some specific implementations, default values can be introduced for one or more parameters. For example, such default values can be determined by clustering previous patient cases. The golden ratio guidelines can specify one or more numbers related to the width of adjacent teeth, such as: {1.6, 1, 0.6}.
[0181] Tooth-to-tooth ratios can also be defined between other tooth pairs. The ratios can be made with respect to tooth width, height, diagonal, etc. Angular lines, chamfer angles, and buccal contours can describe the main aspects of the macroscopic tooth shape. Marginal ridge grooves can be vertical macroscopic textures on the front of the tooth and can sometimes be V-shaped. Striations or perikymata can be horizontal microtextures on the tooth. Symmetry may often be desired. There may be differences between male and female patients.
[0182] Parameters can be defined to encode doctor restoration design preferences (DRDP) related to various use case scenarios. These use case scenarios can reflect information about the handling preferences of one or more doctors and directly affect the properties of one or more teeth in a dental restoration design or veneer. Additionally, DRDP can describe the RDP values or ranges of values that are habitually involved in the preferences or habits of doctors or other treating healthcare professionals. In some cases, such values or ranges of values can be derived from historical patient cases treated by that doctor or healthcare professional. In some cases, DRDP can be defined from RDP (e.g., the aesthetic relationship such as width to length) or from RDM. Non-limiting examples of RDP are described in Table 2.
[0183] Machine learning models (such as those described herein) can be trained to generate designs for crowns or roots (or both). A dental restoration design can describe the expected tooth shape at the end of a dental restoration. Neural networks (such as generative neural networks) can be trained to generate dental restoration designs that will be used to generate veneers (e.g., zirconia veneers) or dental restoration appliances. Such models take as input data from past patient cases, including pre-restoration tooth meshes and corresponding ground truth examples of the completed restorations (e.g., tooth meshes with the restored shape and / or structure). Such models can be trained at least in part by computing a loss function that can quantify the difference between the generated crown restoration design and the ground truth crown restoration design. The resulting loss can be used to update the weights of the generative neural network model (e.g., a transformer), thereby (at least in part) training the model. A reconstruction loss can be computed to compare the predicted tooth mesh with the ground truth tooth mesh or the pre-restoration tooth mesh with the completed restoration design tooth mesh. The reconstruction loss can be computed as the sum of the pairwise distances between corresponding mesh elements and can be computed to quantify the difference between two crown designs.
[0184] Other losses disclosed herein can also be in the training. Transformers can improve data precision and may be particularly suitable for generating restoration designs for crowns since such meshes can include a large number of mesh elements. Transformers have proven to be good at handling long sequence data and other large datasets. The generated restoration designs can be used to create veneers. For example, the veneers can be 3D printed.
[0185] Machine learning models, such as those described herein, can be trained to generate components for use in creating dental restoration appliances. Such dental restoration appliances can be used to shape a dental composite material in a patient's oral cavity while curing the composite material (e.g., using a curing light), ultimately creating a veneer on one or more of the patient's teeth. 3M ® Filtek ™ Matrix is an example of such a product. In some cases, a machine learning model for generating appliance components can take inputs that can be used to customize the shape and / or structure of the appliance components, including inputs such as oral care parameters. In some cases, one or more oral care parameters can be defined based on oral care metrics. Oral care metrics (e.g., orthodontic metrics or prosthodontic design metrics) can describe the physical and / or spatial relationship between two or more teeth, or can describe the physical and / or dimensional characteristics of a single tooth. Oral care parameters can be defined to provide guidance to the machine learning model for generating a 3D oral care representation with specific physical characteristics (e.g., related to shape and / or structure). For example, physical characteristics can be measured using oral care metrics corresponding to the oral care parameters. Such oral care parameters can be defined to customize the generation of mold parting surfaces, gingival trim meshes, or other generated appliance components to adapt those appliance components to the patient's dental anatomy.
[0186] In some embodiments, the 3D representation generation techniques described herein (e.g., transformer-based techniques) can be trained to generate customized appliance components by determining characteristics of the customized appliance components, such as the size, shape, location, and / or orientation of the customized appliance components. Examples of customized appliance components include mold parting surfaces, gingival trim surfaces, shells, facings, lingual shelves (also referred to as "ribs"), doors, windows, incisal ridges, outer shell frame spare parts, or interdental matrix wraps, among others.
[0187] A mold parting surface refers to a 3D mesh that divides the two sides of one or more teeth (e.g., by separating the facial side of one or more teeth from the lingual side of one or more teeth). A gingival trim surface refers to a 3D mesh that trims the shell along the gingival margin. A shell is a body of a nominal thickness. In some examples, the inner surface of the shell matches the surface of the dental arch, and the outer surface of the shell is a nominal offset of the inner surface.
[0188] The facial surface refers to a reinforcing rib of nominal thickness offset from the face of the housing. The window refers to an orifice providing access to the tooth surface such that dental composite material can be placed on the tooth. The door refers to a structure covering the window. The incisal ridge provides reinforcement at the incisal edge of the dental restoration appliance and can be obtained from the dental arch form. The outer shell frame spare part refers to a connecting material that couples components of the dental restoration appliance (e.g., the lingual portion of the dental restoration appliance, the facial portion of the dental restoration appliance, and its subassemblies) to the fabricated outer shell frame. In this way, the outer shell frame spare part can bind the components of the dental restoration appliance to the housing frame during fabrication, protect the individual components from damage or loss, and / or reduce the risk of mixing components.
[0189] Additional 3D oral care representations that can be generated by a transducer (such as a transducer trained as described herein) include interproximal tooth surfaces and tooth roots.
[0190] In some cases, the transducers described herein can be trained to perform 3D mesh element labeling (e.g., labeling vertices, edges, faces, voxels, or points) in 3D oral care representations. Those labeled mesh elements can be used for mesh cleaning or mesh segmentation. In the case of mesh cleaning, the labeled aspects of the scanned tooth mesh can be used for appliance erasure (remove + replace) or for modifying (e.g., by smoothing) one or more aspects of the tooth to remove aspects of attached hardware (or other aspects of the mesh that may not be needed for certain processing and appliance creation, such as foreign material). Mesh element features, such as those described herein, can be calculated for one or more mesh elements in the 3D oral care representation. A vector of such mesh element features can be calculated for each mesh element and then received by a transducer that has been trained to label mesh elements in the 3D oral care representation for mesh segmentation or mesh cleaning purposes. Such mesh element features can impart valuable information about the shape and / or structure of the input mesh to the labeling transducer.
[0191] Some specific implementations of the transformer-based mesh cleaning techniques described herein can train the transformer to remove (or modify) general triangular mesh defects, such as: degenerate triangles with zero surface area; redundant triangles that cover the same surface area as another triangle; non-manifold edges with more than two adjacent triangles, also known as "flaps"; non-manifold vertices with more than one adjacent sequence of connected triangles (triangle fans); intersecting triangles - where two triangles pass through each other; spikes - sharp features composed of multiple triangles, usually conical, caused by one or more vertices deviating from the actual surface; folds - sharp features composed of multiple triangles, usually Z-shaped with small recessed areas, caused by one or more vertices deviating from the actual surface; islands / sub-components - disconnected objects where only a single object should be included in the scan (e.g., usually smaller objects are removed); small holes in the mesh surface, either from the initial scan or from deletions due to previous defects (e.g., the holes can be removed by filling the holes, e.g., by adding one or more mesh elements); rough boundaries - smooth boundaries are beneficial for extending the gingival surface and creating the model base.
[0192] Some specific implementations of the transformer-based mesh cleaning techniques described herein can train the transformer model to remove (or modify) aspects of the mesh that are not needed and / or are domain-specific defects in some cases, such as: extraneous material - parts of an intraoral scan outside the anatomical region of interest, e.g., non-tooth surfaces that are not within a certain distance of the tooth surface or scan artifacts that do not represent the actual anatomical structure; depressions - recesses in the surface (e.g., which can be scan artifacts that should be repaired or anatomical features that are usually left intact); undercuts - tooth sides that are smaller than the crown radius, so physical impressions or appliances may be difficult to remove or place. Undercuts can be natural features or caused by damage such as internal fractures. Internal fractures are associated with erosion of the tooth near the gum line, which causes or exacerbates undercuts. Appliances that can be handled by the transformer-based models of the present disclosure include orthodontic hardware, such as attachments, brackets, wires, buttons, lingual bars, Carriere appliances, etc., and can be present in intraoral scans. In some cases, it may be beneficial to perform digital removal and replacement with synthetic tooth / gum surfaces before the appliance creation step.
[0193] In some specific implementations, the neural networks of the present disclosure trained for the placement and / or generation of oral care appliance components (e.g., such as dental restoration appliance {DRA} components) can operate on the patient's post-restoration dentition. In other specific implementations, such neural networks can operate on the patient's pre-restoration dentition.
[0194] Aspects of the mesh element features are described below. In some specific implementations, either or both of the neural network of the first component for generating a dental restoration appliance (DRA) and the neural network of the second component for placing the DRA can input one or more mesh element features, with the advantage of improving the processing accuracy of the neural network regarding the input 3D representation (i.e., such as teeth, gums, hardware, and / or a third component). Mesh element features are described elsewhere in the present disclosure. Mesh element features can also be used for placing brackets and / or attachments. Mesh element features can also be used to generate 3D representations of veneers and / or crowns.
[0195] Aspects of the present disclosure use representation learning (RL) for DRA component placement. The RL-based techniques of the present disclosure can use any one of VAE, capsule autoencoders, or U-Net to create representations of teeth and library components. In some specific implementations, the RL-based techniques of the present disclosure can use a downstream encoder, a transformer, or a multi-layer perceptron (MLP) network to generate a transformation for placing a DRA component (or a bracket or an attachment) relative to one or more teeth.
[0196] Aspects of the present disclosure can also employ the use of generative models. In some specific implementations, the systems of the present disclosure can use autoencoders to generate components for the DRA. The input 3D representation of a tooth can be encoded into a latent form A (a latent vector or a latent capsule), and modifications can be applied to A. These modifications can be made according to previous experiments mapping out the latent space. In some specific implementations, the reconstruction output of such a model can be a new DRA component. In some specific implementations, when an existing DRA component and a tooth are provided to the autoencoder, the reconstruction output can include a modified DRA component (e.g., having an improved shape and fit relative to the tooth). Generative models can also be used to generate 3D representations of veneers and / or crowns and other appliances.
[0197] Some neural networks of the present disclosure can use a restoration design metric (RDM) as an input to the neural network for placing DRA components. Similarly, the RDM can be used as an input to the neural network for generating DRA components. The neural networks of the present disclosure that use RDM data as an input are such that the neural network can obtain knowledge about the geometry and / or structure of the patient's dentition through the RDM.
[0198] Figure 3Describes methods for training a machine learning model to modify or generate predictive 3D oral care representations (e.g., training a neural network to generate or modify oral care appliance components, trimming lines, dental arch morphology, crown restoration designs, etc.). An oral care mesh 300 (e.g., representing a patient's teeth) can be received at the input. Optional oral care metrics (310) can be computed, including orthodontic metrics that can describe the physical relationships between teeth (e.g., related to the position and orientation of teeth relative to other teeth or the gums) and / or restorative design metrics that can describe physical aspects within the teeth. The tooth mesh 300, any optional oral care metrics 302 computed on those teeth, and other optional inputs can be provided to a representation generation module 308. Optional oral care parameters or doctor preferences 332 can be provided to the representation generation module 308 or the generator module 320 to customize the output of those modules. Other optional inputs 304 can include a template oral care mesh (e.g., an appliance or appliance component, such as a parting surface) or a customized oral care mesh that may require further customization (e.g., a mold parting surface generated according to the techniques of WO2021240290A1 or WO2020240351A1). A ground truth 3D oral care representation 306 can be received at the input of the method and used in loss calculations (e.g., according to the loss calculation techniques described herein). In the case of appliance component generation, the input can include the patient's teeth 300, and optionally a template appliance component 304 (which can help influence the generator 320 to generate a predictive appliance component) and a ground truth appliance component. Such appliance components can include any one of a mold parting surface, a gingival trimming surface, a housing, a face plate, a lingual shelf (also known as a "reinforcing rib"), a door, a window, an incisal ridge, a housing frame spare part, an interdental matrix wrapper, etc. Optional mesh element features can be computed for each mesh element in 300, 302, and / or 304, and then representations can be generated for these oral care meshes. Representations can be generated using an autoencoder 312 that produces a latent vector or latent capsule. Representations can be generated using a U-Net 314 that produces an embedding vector. Representations can also be generated using a pyramid encoder-decoder or by using an MLP that includes convolutional and pooling layers 316 (e.g., with a convolutional kernel size of 5 and average pooling). According to aspects of the present disclosure, other representations are possible. The representations can be concatenated (318) and received by a generator module 320, which can generate one or more predictive 3D oral care representations (e.g., using any one of the following in combination with the mesh element feature vectors: a transformer, an autoencoder, PolyGen, or a neural network trained from PolyGen via transfer learning). A loss can be computed (324) between the generated oral care mesh 322 and the corresponding ground truth oral care mesh 306. After the loss drops below a threshold, training can be considered complete (326).During the process of training the generator, for example, backpropagation can be used to feedback the loss to update the weights of the generator (328). The output of method 330 can include a trained machine learning model (e.g., a neural network) for generating a predicted oral care mesh.
[0199] Figure 4 A method for generating a predicted 3D oral care representation (e.g., generating an oral care appliance component) for a deployed machine learning model is described. An oral care mesh 400 (e.g., a patient's teeth) can be received at the input. Optional oral care metrics (402) can be computed. The tooth mesh 400, any optional oral care metrics 402 computed on those teeth, and other optional inputs can be received by a representation generation module 406. The other optional inputs 404 can include a template oral care mesh or a customized oral care mesh that may require further customization. Optional oral care parameters or doctor preferences 422 can be provided to the representation generation module 406 or the generator module 418 to customize the output of those modules. Optional mesh element features can be computed (408) for each mesh element in 400, 402, and / or 404, after which a representation can be generated for these oral care meshes. A representation can be generated using an autoencoder 410 that produces a latent vector or latent capsules. A representation can be generated using a U-Net 412 that produces an embedding vector. A representation can also be generated using a pyramid encoder-decoder or by using an MLP that includes convolutional and pooling layers 414 (e.g., having a convolutional kernel size of 5 and average pooling). According to aspects of the present disclosure, other representations are possible. The representations can be concatenated (416) and received by a generator module 418, which can generate one or more predicted 3D oral care representations (e.g., using an autoregressive generative neural network model such as PolyGen or a neural network trained from PolyGen via transfer learning). The output 420 can include one or more predicted 3D oral care representations (e.g., a mesh describing a mold parting surface).
[0200] In some specific implementations, Figure 3 the method in can be trained to modify one or more input 3D oral care representations, such as an input appliance component, an input tooth design (e.g., a pre-restoration or ongoing restoration design), a trim line, an arch form, or other types of 3D oral care representations. In some cases, a transformer such as an autoregressive transformer (e.g., PolyGen) can be trained to modify 3D oral care representations. Such modifications may require operations such as adding or estimating one or more mesh elements, removing one or more mesh elements, point cloud completion, transforming one or more mesh elements (e.g., modifying the position and / or orientation of one or more mesh elements), etc. In some specific implementations, Figure 3One or more of the neural networks (such as a transformer from a generator module) can be trained at least in part by transfer learning. In Figure 3 One or more of the neural networks trained therein can then be used to at least partially train another neural network (such as a neural network for some aspects of digital oral care automation) according to the transfer learning paradigm. Grid element feature vectors can be calculated for one or more grid elements of Figure 3 and Figure 4 one or more inputs, which can enable an improved understanding of those input grids or point clouds.
[0201] Figure 3 illustrates a training method, and Figure 4 illustrates a deployment method. These methods can involve using neural networks to generate oral care appliances or oral care appliance components. In some optional scenarios, the neural networks of the present disclosure can further customize or improve existing appliances or appliance components, in which case, an appliance or appliance component having an initial configuration can be received as input data for the neural network. In some embodiments, the training method in Figure 3 can be executed to generate components of a dental restoration appliance (e.g., such as a mold parting surface, a gingival trimming surface, a housing, a face strip, a lingual shelf (also referred to as a "rib"), a door, a window, an incisal ridge, a housing frame spare part, an interdental matrix wrapper, etc.). A spline refers to a curve passing through a plurality of points or vertices, such as a piecewise polynomial parametric curve. A mold parting surface refers to a 3D mesh that bisects both sides of one or more teeth (e.g., separating the facial side of one or more teeth from the lingual side of one or more teeth). A gingival trimming surface refers to a 3D mesh that trims the housing along the gingival margin. A housing is a body of a nominal thickness. In some examples, the inner surface of the housing matches the surface of the dental arch, and the outer surface of the housing is a nominal offset of the inner surface. A face strip is a rib of a nominal thickness offset from the face of the housing. A window refers to an orifice that provides access to the tooth surface so that dental composite can be placed on the tooth. A door refers to a structure that covers the window. An incisal ridge provides reinforcement at the incisal edge of the dental appliance and can be obtained from the dental arch morphology. A housing frame spare part refers to a connecting material that couples the components of the dental appliance (e.g., the lingual part of the dental appliance, the facial part of the dental appliance, and their sub-components) to a fabricated housing frame. In this way, the housing frame spare part can bind the components of the dental appliance to the housing frame during manufacturing, protect the individual components from damage or loss, and / or reduce the risk of mixing components. These appliance components and other components are described in PCT patent applications WO2020240351A1 and WO2021240290A1, both of which are incorporated herein by reference in their entireties.
[0202] can be in Figure 3At the input of the training method, a 3D representation of a patient's teeth (such as a 3D mesh) is received, as well as an associated ground truth appliance or appliance component (such as a ground truth mold parting surface) that can be generated by an automated model (e.g., techniques such as WO2020240351A1 or WO2021240290A1) and can be modified or revised by an expert technician or other healthcare practitioner or clinician. In some specific implementations, Figure 3 the training method can be enhanced by calculating one or more oral care metrics on the received teeth. Oral care metrics include orthodontic metrics (OM) and dental restoration metrics (DRM). Orthodontic metrics can describe the relationship between two or more teeth, and dental restoration metrics can describe aspects of the shape and / or structure of a single tooth (and in some cases can describe the relationship between two or more teeth). These oral care metrics (described elsewhere) can help the representation module create a representation of the teeth. The representation of the teeth can reduce the size or amount of data required to describe the shape and / or structure of the teeth, while retaining much of the information about the shape and / or structure of the teeth, thus providing a technical improvement of the present disclosure based on reducing the use of computing resources. The representation of the teeth can be more easily consumed by a machine learning model (such as a generator module) in this reduced size and compact form.
[0203] The neural network of the present disclosure can generate a representation of a tooth or other oral care mesh (such as the received appliance or appliance component). The tooth mesh can be reconfigured for one or more lists of mesh elements (e.g., vertices, faces, edges, or voxels). For each mesh element, an optional mesh element feature vector can be calculated according to the mesh element feature description provided elsewhere in the present disclosure. This mesh element feature can help the neural network encode the tooth mesh into a reduced size representation. According to aspects of the present disclosure, an autoencoder (such as a variational autoencoder or a capsule autoencoder) can be trained to reconstruct an oral care mesh (e.g., such as a tooth or an appliance component). The trained reconstruction autoencoder can use a 3D encoder stage to encode the mesh into a latent vector (or latent capsule). This latent vector (or latent capsule) can be used as a representation of the oral care mesh.
[0204] According to various techniques of the present disclosure, an alternative to the autoencoder includes a U-Net neural network structure. The U-Net can encode the oral care mesh into an embedding vector, which can then be used as a tooth representation. In some specific implementations, one or more layers including convolutional kernels and pooling operations can be trained to perform the encoding task. For example, a convolutional kernel of size five (5) can be combined with an average pooling operation to achieve encoding the oral care mesh into a representation suitable for being received by the generator module.
[0205] The generator module may receive representations of teeth, appliance components, and / or any other oral care mesh. In some embodiments, these representations may be concatenated before being received by the generator module. The generator module may include a neural network or some other machine learning model. In some embodiments, a multi-layer perceptron (MLP) may be trained to receive the concatenated oral care representations and output a mesh corresponding to the appliance or appliance component. The output layer of the encoder may be designed to output appliance components, such as mold parting surfaces. The mold parting surfaces may be described using 3D meshes and may include a large number of mesh elements. The output layer of the generator module may accommodate the output of dozens or even hundreds of generated mesh elements, and a transformer-based model may be particularly suitable for this case. In some embodiments, a transformer may be used in place of the MLP. Other embodiments may use a 3D encoder in place of the MLP or transformer. In some implementations, the generator module may include a transformer configured to generate or modify at least some aspects of the received representation (e.g., a reformatted representation such as output by a representation generation module). In some cases, one or more other neural networks (e.g., a set of fully connected layers) may follow the transformer, and the one or more other neural networks may further reformat or rearrange the transformer output, such as rearranging the output into one or more 3D meshes or point clouds that can subsequently be output by the generator module.
[0206] The output of the generator module may include a mesh that can be received by the loss calculation module, such as a mold parting surface (or other appliance or appliance component or other type of oral care representation). The loss calculation module may also receive a ground truth mold parting surface (or other appliance or appliance component) and proceed to quantify the difference between the predicted mesh and the ground truth mesh. The loss may be used to at least partially train the generator module. The loss may be used to modify the weights of the generator module (in the case where the generator module is a neural network) via a backpropagation algorithm. After the loss drops below a specified threshold (possibly as indicated over multiple consecutive iterations), the training method is considered complete. At least one of KL divergence loss and reconstruction loss (e.g., which may involve vertex-to-vertex distance calculation) may be used to train the generator. Additional losses that may be used include normalized L1 and L2 distances, chamfer loss, and MSE loss (e.g., normalized MSE loss). In some cases, once the accuracy of the generator rises above a specified threshold, the training method is considered complete. The ADD score may be used to calculate the accuracy. The ADD score may measure the percentage of patient cases that are closer than a threshold distance to each other (e.g., measured using the L2 distance or another distance measurement technique described herein). In some embodiments, the threshold distance may be adaptively determined based on the size of the ground truth mesh. After training is complete (e.g., upon achieving a steady state or convergence), the trained model may be output.
[0207] The Generator Module (GM) may include an autoregressive machine learning model that is trained to predict a customized 3D oral care representation of a new patient based on examples of 3D oral care representations from past patients. In some embodiments, a U-Net may be trained to autoregressively generate an oral care mesh. In other cases, a Transformer (e.g., PolyGen) may be trained to autoregressively generate an oral care mesh. Such a Transformer may be trained to model a distribution over examples of a particular type of oral care mesh, examples of which are described elsewhere in this disclosure. Such a Transformer may include one or more neural networks, each of which is trained to model a distribution over a particular mesh element. For example, a first neural network may be trained to model an unconditional distribution over mesh vertices, and a second neural network model may be trained to conditionally model a distribution over mesh faces (or alternatively, mesh edges, or in the case of sparse processing, voxels). The vertices of the predicted oral care mesh are estimated first, and then the faces (or alternatively, edges) connecting those vertices are estimated based on the set of estimated vertices received. The first model (for estimating vertices) may receive a tooth mesh as input, where the mesh for the hardware is attached to the teeth and / or the mesh for the appliance or appliance component.
[0208] Such an appliance or appliance component may represent a template that serves as a starting point for oral care mesh generation. In other cases, such an appliance or appliance component may be pre-customized and subject to further modification and / or refinement from the Transformer. The first model (for estimating vertices) may include a Transformer decoder, such as described by (Vaswani, Ashish & Shazeer, Noam & Parmar, Niki & Uszkoreit, Jakob & Jones, Llion & Gomez, Aidan & Kaiser, Lukasz & Polosukhin, Illia, “Attention is all you need”, 2017). The second model may model the conditional distribution over faces and may assemble the vertices estimated by the first model into a realistic 3D mesh.
[0209] GM can be trained to generate different types of oral care meshes, such as appliance components (e.g., model parting surfaces), clear tray aligner trim lines (e.g., those implemented using polylines or meshes), or dental arch forms (e.g., those implemented using meshes or a set of control points, where associated spline curves pass through those control points), tooth restoration designs (e.g., the target tooth shape and / or structure expected upon completion of a tooth restoration process, which can be used to create a tooth restoration appliance), crown designs, or veneer designs (e.g., zirconia veneers). To generate a given type of 3D oral care representation, many examples of training data of that (or a related) type of 3D oral care representation can be introduced into the training dataset and the validation dataset. A test set can be kept available for evaluating the accuracy of the predicted representation.
[0210] In terms of dental arch form prediction, a generative machine learning model such as a neural network can be defined to predict the dental arch form. The dental arch form is defined elsewhere in this disclosure. In some specific implementations, the dental arch form can be defined at least in part by one or more of a polyline and a set of control points (where such control points describe one or more splines or other curves that can be fitted to the control points). In some specific implementations, the ML model can be trained at least in part by using a loss function that quantifies the difference between the predicted dental arch form and a reference dental arch form (such as a ground truth dental arch form previously configured by an expert or through an optimization algorithm). Such an ML model for dental arch form prediction can be trained at least in part based on data from past patient cases, where the patient cases include at least one of the following: a set of segmented tooth meshes of a patient, orthotic transformations of one or more teeth (e.g., each tooth), setup transformations of one or more teeth (e.g., each tooth), or a ground truth dental arch form (which may exhibit ideal dental arch form characteristics in some cases). The predicted dental arch form can be used in patient treatment, such as being provided to the neural network of this disclosure. The neural network for predicting the dental arch form can be based at least in part on at least one aspect of at least one neural network or neural network feature disclosed elsewhere in this disclosure.
[0211] The dental arch form can describe the contour of the dental arch and, in some cases, can describe the smooth, average, or idealized arrangement of teeth in the dental arch. In some cases, the dental arch form can be (at least approximately) aligned with the target arrangement of teeth used in orthodontic treatment and, in some cases, can be received as an input to a setup prediction machine learning model, such as a setup prediction neural network for final setup or intermediate staging prediction. In some cases, the dental arch form can be aligned with at least one of the incisal edges of teeth, the gingival margins of the dental arch, or one or more coordinate systems of one or more teeth.
[0212] Some systems of the present disclosure use RL models for the dental arch morphology prediction function. In some specific implementations, RL is used to train a dental arch morphology prediction machine learning model (e.g., including one or more neural networks). The representation learning model may include a first module and a second module. The first module may be trained to generate a representation of the received 3D oral care representation (e.g., teeth and / or gums), and the second module may be trained to receive those representations and generate one or more dental arch morphologies. Figure 5 An example method for such specific implementations is shown. That is, Figure 5 Illustrates a specific implementation of the dental arch morphology prediction machine learning model of the present disclosure. The tooth mesh of the patient's dental arch 500 can be provided to the representation generation module 502, which can provide a latent representation of the patient's teeth to the dental arch morphology prediction ML module 506 (e.g., a transformer decoder, followed by an MLP, or other structures described herein). The dental arch morphology prediction ML module 506 can generate one or more predicted dental arch morphologies, and the one or more predicted dental arch morphologies can be provided to the loss calculation module 508. A loss can be calculated by comparing the predicted dental arch morphology with the corresponding ground truth dental arch morphology 504. Calculating the loss can be used to train (510) the dental arch morphology prediction ML module 506. Output the trained model 512.
[0213] In some cases, Figure 3 The generator module (e.g., which may include at least one transformer, such as a generative autoregressive transformer) can be trained to generate new tooth restoration designs or to modify the shape and / or structure of existing tooth restoration designs. Either or both of the representation generation module and / or the generator module can take oral care parameters such as restoration design parameters (RDP) and / or doctor's restoration design preferences (DRDP) as inputs. Such oral care parameters can enable the generator module to incorporate clinical instructions from doctors / dentists / healthcare practitioners to improve the customization of the resulting restoration design. Either or both of the representation generation module and / or the generator module can take as input one or more tooth meshes (crowns and / or roots) from past patient cohort cases.
[0214] One or more mesh element features can be calculated for one or more mesh elements of one or more teeth. Such mesh element features can improve the ability of either or both of the representation and / or generator modules to understand the shape and / or structure of the input mesh (e.g., the pre-restoration tooth mesh received at the input). An autoregressive transformer can be trained to generate those aspects of the tooth anatomy based on the distribution of aspects of the shape and / or structure of the tooth anatomy found in the training dataset of a cohort of patient cases. In some cases, the transformer can be trained to assign the following five levels of aspects of tooth design to the generated restoration design (e.g., crown or root). The advantages provided by this technique are to generate tooth restoration designs that reflect one or more of the following physical properties: 0 - tooth contour, e.g., as projected onto a plane in front of the face; 1 - main tooth shape (main anatomy); 2 - surface vertical and horizontal macrotexture or striations or marginal tubercle grooves (secondary anatomy); 3 - surface horizontal microtexture, e.g., enamel striae of Retzius (tertiary anatomy); 4 - volumetric representation of the internal structure of the tooth (dentin, enamel, etc.).
[0215] In some embodiments, the transformer can achieve a large model capacity and / or implement an attention mechanism (e.g., the ability to focus resources on certain inputs and respond to them). The transformer can consume a large training dataset and has the advantage that the transformer can continue to grow in model capacity as the size of the training dataset increases. In contrast, various prior neural network models can stagnate in model capacity as the size of the training dataset increases.
[0216] In some embodiments, a convolutional-based neural network can achieve fast model convergence during training and improve model generalization. In some embodiments for generating or modifying 3D oral care representations, the transformer can be combined with a convolutional-based neural network, such as by vertically stacking convolutional layers and attention layers. Such stacking can improve efficiency, model capacity, and / or model generalization. CoAtNet is an example of a network architecture that combines convolutional and attention-based elements and can be applied to the processing of oral care data. In some cases, transfer learning can be used at least in part to train a network for modifying or generating 3D oral care representations from CoAtNet (or another model that combines convolution and self-attention / transformer).
[0217] Table 3 describes the input data and generated data for several non-limiting examples of the generation techniques described herein. An encoder-decoder structure such as an autoencoder or a transformer can be trained to generate (or modify) a point cloud as described herein. In some embodiments, such a model can be trained to generate (or modify) the input data in Table 3 to obtain the generated data in Table 3.
[0218] The techniques of the present disclosure can be trained to generate (or modify) point clouds (e.g., where points can be described as 1D vectors, such as (x, y, z)), polylines (points connected in sequence by edges), meshes (points connected via edges to form faces), splines (which can be calculated through a set of generated control points), sparse voxelized representations (which can be described as a set of points corresponding to the centroid of each voxel or some other landmark of the voxel (such as the boundary of the voxel)), transformations (which can take the form of one or more 1D vectors or one or more 2D matrices (such as a 4×4 matrix)), etc. In some specific implementations, a voxelized representation can be calculated from a 3D point cloud or a 3D mesh. In some specific implementations, a 3D point cloud can be calculated from a voxelized representation. In some specific implementations, a 3D mesh can be calculated from a 3D point cloud.
[0219] The first module can be trained to generate a 3D representation of one or more teeth suitable for consumption by a second module, where the second module is trained to output one or more predicted dental arch morphologies. In some specific implementations, one or more layers including convolutional kernels (e.g., having a kernel size of 5 or some other size) and pooling operations (e.g., average pooling, max pooling, or some other pooling method) can be trained to create a representation of one or more teeth in the first module. In some specific implementations, one or more U-Nets can be trained to generate a representation of one or more teeth in the first module. In some specific implementations, one or more autoencoders can be trained to generate a representation of one or more teeth in the first module (e.g., where the 3D encoder of the autoencoder is trained to encode one or more dental 3D representations into one or more latent representations, such as latent vectors or latent capsules, where such latent representations can be reconstructed via the 3D decoder of the autoencoder into a facsimile of one or more input tooth meshes). In some specific implementations, one or more 3D encoder structures can be trained to create a representation of one or more teeth in the first module. According to aspects of the present disclosure, other techniques for encoding representations are also possible.
[0220] A representation of one or more teeth may be provided to a second module that has been trained to output one or more dental arch forms, such as an encoder structure, a multi-layer perceptron (MLP), a transformer, an autoencoder (e.g., a variational autoencoder or a capsule autoencoder). In some embodiments, the dental arch form may include n control points (e.g., n = 6), such as those represented by a 3×6 array (including XYZ coordinates for 6 control points). The model may fit a spline to the n control points to enhance the description of the dental arch form. In some embodiments, the dental arch form may include a polyline, a 3D mesh, or a 3D surface. The second module may be trained at least in part by calculating one or more loss values, such as an L1 loss, an L2 loss, an MSE loss, a reconstruction loss, or one or more of the other loss calculation methods found elsewhere in the present disclosure. Such a loss function may quantify the difference between one or more generated representations of the dental arch form and one or more reference representations of the dental arch form (e.g., a ground truth dental arch form known to have good functionality).
[0221] In some embodiments, either or both of the first module and the second module may take optional inputs, including one or more of the following: tooth position and / or orientation information O, orthodontic protocol parameters K, orthodontist preferences L, tooth type information, orthodontic metrics S, IPR information U, a label associated with a medical condition or medical diagnosis information for one or more teeth, etc. The advantage of incorporating one or more of the above inputs is to improve the customization and / or functionality of the dental arch form and the suitability of the dental arch form used in creating oral care appliances, such as clear tray appliances or indirect bonding trays for orthodontic treatment, resulting in a technical improvement in data accuracy.
[0222] Some embodiments that set up a prediction neural network may take the dental arch form as an input, which has the advantage of improving the customization and / or functionality of the resulting prediction setup (e.g., the final setup or an intermediate stage). Other setup prediction methods may also benefit from the use of the dental arch form, such as other ML-based setup prediction methods or non-ML-based setup prediction methods. Some embodiments of landmark-based setups may benefit from using the dental arch form in the prediction of the setup.
[0223] In some cases, the dental arch form may be aligned with at least one of the incisal edges of the teeth, the gingival margins of the dental arch, or one or more coordinate systems of one or more teeth. In some cases, the dental arch form may (at least approximately) describe the target arrangement of the teeth. In some cases, the dental arch form may describe the average or smooth arrangement of the teeth in the dental arch.
[0224] The dental arch shape can be described by a set of control points, where each control point corresponds to some aspect of a tooth (e.g., corresponding to the centroid of the tooth, corresponding to the origin of the local coordinate system of the tooth, corresponding to landmarks located within the tooth anatomy, etc.). A spline can be fitted through the set of control points (e.g., a B-spline or NURBS surface), thereby approximating at least some aspects of the contour of the patient's dental arch. In some cases, aspects of the set of control points can be combined with aspects of the local coordinate axes of one or more teeth. For example, in the case where the control points correspond to the teeth of the mandibular arch, the Z-axis of the local tooth coordinate system can point downward (e.g., in the gingival direction) relative to the patient's mouth. Line segments can be defined between points from multiple control points and points along the negative Z-axis of the tooth, such that the first endpoint of the line segment is near the control point and the second endpoint of the line segment is near the root of the tooth. A 3D triangular mesh can be defined from the first and second endpoints of these line segments. The teeth can be modeled as being able to slide along the upper and lower boundaries of the dental arch shape mesh (e.g., which can define a surface and / or volume), as if the teeth were attached to a set of rails and constrained by the rails defined by the dental arch shape surface / volume. Additionally, by requiring that the mesial-distal axis of the local coordinate system of each tooth must always be tangent to the spline / occlusal surface of the rails, the teeth are further constrained on the "rails", thereby ensuring that the teeth are fully constrained on the rails / dental arch shape surface in all degrees of freedom. As long as the teeth move along the rails, the teeth can move in the mesial or distal direction. The movement of a tooth can enable or require the movement of one or more other teeth (e.g., adjacent teeth to the tooth that is moving) along the rails. In some cases, as long as the teeth remain attached to the rails, random perturbations can be applied to the teeth along the rails to adjust their position and orientation. A machine learning model can be trained to move the teeth along the rails, such as a neural network conforming to the architecture described herein.
[0225] Figure 21 Two dental arch shape meshes are shown, one dental arch shape mesh for each dental arch. As shown, Figure 21 includes a depiction of an upper dental arch shape mesh 2100 and a lower dental arch shape mesh 2102. Each mesh describes the "dental arch shape" of the dental arch, which can describe aspects of the dental arch (e.g., such as shape, structure, and / or curve). Each control point 2108 can correspond to a point associated with a tooth. Each control point 2108 has one or more associated line segments 2106 placed along the Z-axis of the corresponding tooth local coordinate system. At the end of each of these line segments is a Z-axis point 2104. For each of the upper dental arch 2100 and the lower dental arch 2102, a 3D triangular mesh is formed by the control points 2108 and the Z-axis points 2104. In some cases, one or more points 2110 can be defined that can be placed along a spline line fitted through the control points 2108 (one example of which is shown in Figure 21In some cases, one or more points 2112 can be defined that can be placed along a spline fitted through the Z-axis point 2104 (which can be interpreted as another set of control points), an example of which is shown in Figure 21 In some specific implementations, Figure 21 the points shown in can include a 3D triangle mesh or other 3D representation.
[0226] In some cases, a machine learning model such as a neural network can be trained to adjust aspects of one or both guides (or some other aspect of the 3D dental arch form mesh), such as the shape of one or more guides. These adjustments can be guided to produce favorable orthodontic treatment outcomes, such as aligning teeth into a posture suitable for a setting (e.g., an intermediate stage or a final setting). In some cases, transformations such as those described herein can be trained to effect a change in the shape of the dental arch form mesh. In some cases, an autoencoder (such as a reconstruction autoencoder, an example of which is a variational autoencoder optionally utilizing normalizing flows) can be trained to effect a change in the shape of the dental arch form mesh. The reconstruction autoencoder can be trained on a dataset of queued patient cases, where each patient case includes at least one of the following: a set of tooth meshes, the malocclusion transformation for each tooth (to place the tooth in a malocclusion posture), an intermediate transformation, and an approved final setting transformation. The dental arch form mesh can be constructed relative to a malocclusion dental arch, referred to herein as a malocclusion dental arch form mesh (or malocclusion dental arch form). The dental arch form mesh can be constructed relative to an approved final setting dental arch form, referred to herein as a final setting dental arch form mesh (or set dental arch form). The dental arch form mesh can be constructed relative to an intermediate stage dental arch form, referred to herein as an intermediate stage dental arch form mesh (or graded dental arch form).
[0227] An autoencoder (such as Figure 6 or Figure 7An autoencoder, such as a variational autoencoder (VAE), can be trained to encode 3D oral care representations (e.g., dental arch form control points, 3D dental arch form meshes, appliance components, polyline CTA trim lines, dental restoration designs, etc.) into a latent space vector A that can exist in an informative low-dimensional latent space. This latent space vector A may be particularly suitable for subsequent processing by digital oral care applications, such as modification of the 3D oral care representation, because A enables efficient manipulation of complex oral care mesh data. Such an autoencoder can be trained to reconstruct the latent space vector A back into a replica of the input 3D oral care representation (e.g., a mesh of the dental arch form). In some embodiments, the latent space vector A can be modified strategically to cause a change in the reconstructed mesh. In some cases, the reconstructed mesh can be a 3D oral care representation (e.g., a dental arch form) with a changed and / or improved shape, such as would be suitable for use in an oral care appliance (such as a dental restoration appliance, such as 3M
[0228] An autoencoder, such as a variational autoencoder (VAE), can be trained to encode 3D oral care representations (e.g., dental arch form control points, 3D dental arch form meshes, appliance components, polyline CTA trim lines, dental restoration designs, etc.) into a latent space vector A that can exist in an informative low-dimensional latent space. This latent space vector A may be particularly suitable for subsequent processing by digital oral care applications, such as modification of the 3D oral care representation, because A enables efficient manipulation of complex oral care mesh data. Such an autoencoder can be trained to reconstruct the latent space vector A back into a replica of the input 3D oral care representation (e.g., a mesh of the dental arch form). In some embodiments, the latent space vector A can be modified strategically to cause a change in the reconstructed mesh. In some cases, the reconstructed mesh can be a 3D oral care representation (e.g., a dental arch form) with a changed and / or improved shape, such as would be suitable for use in an oral care appliance (such as a dental restoration appliance, such as 3M ® Filtek ™Designs used in Matrix, veneer, or clear tray orthodontic appliances). The term "mesh" should be considered to include 3D meshes, 3D point clouds, or 3D voxelized representations in a non - restrictive sense.
[0229] The 3D oral care representation reconstruction VAE can advantageously utilize a loss function, non - linearities (also known as neural network activation functions), and / or a solver. Examples of loss functions can include one or more of the following: mean absolute error (MAE), mean squared error (MSE), L1 loss, L2 loss, KL divergence, entropy, and / or reconstruction loss. Such loss functions enable each generated prediction to be compared in a quantified manner with the corresponding ground truth, resulting in one or more loss values that can be used to at least partially train one or more neural networks. Examples of solvers can include one or more of the following: dopri5, bdf, rk4, midpoint, adams, explicit_adams, and / or fixed_adams. The solver enables the neural network to solve systems of equations and the corresponding unknown variables. Examples of non - linearities can include one or more of the following: tanh, relu, softplus, elu, swish, square, and / or identity. The activation function can be used to introduce non - linear behavior into the neural network in a way that enables the neural network to better represent the training data. The loss can be calculated by the process of training the neural network via backpropagation. Neural network layers such as one or more of the following can be used: ignore, concat, concat_v2, squash, concatsquash, scale, and / or concatscale.
[0230] In some specific embodiments, the dental arch form mesh reconstruction VAE model can be trained on examples of ground truth dental arch form meshes from a cohort of patient cases. In some specific embodiments, the appliance component mesh (e.g., parting surface or gingival trim mesh) reconstruction VAE model can be trained on examples of ground truth appliance components from a cohort of patient cases. Figure 6 A method of training such a VAE for reconstructing 3D oral care meshes is shown. According to Figure 6 The training aspect shown, the loss between the output G and the ground truth GT can be calculated using the VAE loss calculation method described herein. Backpropagation can be used to train E1 and D1 with such a loss.
[0231] Figure 7 A trained reconstruction VAE for 3D oral care meshes in deployment is shown. According to Figure 7The method in deployment shows that the oral care grid reconstruction VAE reconstructs the 3D dental arch form grid in deployment, where one or more aspects of the latent vector A have been altered to achieve improvements in aspects of the reconstructed oral care grid (e.g., improvements in the shape and / or structure of the reconstructed dental arch form grid).
[0232] Aspects of model training according to the techniques of the present disclosure are described below. A 3D oral care representation reconstruction autoencoder (e.g., where a variational autoencoder (VAE) and a capsule autoencoder are non-limiting examples) can be trained to encode teeth into a dimension-reduced form referred to as a latent space vector. The reconstruction VAE can be trained on example grids of a particular 3D oral care grid of interest (e.g., a 3D dental arch form grid). The input grid can be received by the VAE, deconstructed into a latent space vector using a 3D encoder, and then reconstructed into a facsimile of the input grid using a 3D decoder. An advantage of this method is that the encoder E1 can be trained to encode an oral care grid (e.g., a grid of a dental arch form, a dental appliance, a tooth, a gum, or other parts of the anatomical structure) into a dimension-reduced form that can be used in the training and deployment of an ML model for oral care grid modification. This dimension-reduced form of the oral care grid can be modified and then reconstructed into a reconstructed grid with one or more aspects that have been altered to improve performance (e.g., the shape of the dental arch form grid can be changed so that the dental arch form grid is more suitable for use in appliance creation). The shape of the oral care grid can be changed to obtain a technical improvement in data accuracy.
[0233] The reconstructed oral care grid can be compared with the input oral care grid, for example, using a reconstruction error that quantifies the difference between the grids. This reconstruction error can be calculated using the Euclidean distance between corresponding grid elements of the two grids. There are other methods of calculating this error that can also be derived from materials described elsewhere in the present disclosure.
[0234] In some embodiments, one or more grids provided to the grid reconstruction VAE can first be converted to a vertex list (or point cloud) before being provided to the encoder E1. This way of processing the input to E1 can be beneficial for a single grid input (such as in a tooth grid classification task) or a group of multiple teeth (such as in a setup classification task). The input grids do not need to be connected. Grid element feature vectors can be calculated for one or more grid elements of the input 3D oral care representation. Such grid element feature vectors can provide valuable information about the shape and / or structure of the input grid to the reconstruction autoencoder (e.g., a variational autoencoder optionally utilizing a normalizing flow).
[0235] Aspects of the architecture of the models of the present disclosure are described below. An encoder E1 can be trained to encode an oral care mesh into a latent space vector A (or “3D oral care representation vector”). During a 3D oral care representation modification task, the encoder E1 can arrange an input oral care mesh into a mesh element vector F that can be encoded into the latent space vector A. This latent space vector A can be a dimensionally reduced representation of F that describes important geometric properties of F. The latent space vector A can be provided to a decoder D1 to be restored to full resolution or near full resolution along with a desired geometric change. The restored full resolution or near full resolution mesh can be described by G, which can then be arranged into a reconstructed output mesh.
[0236] The performance of the mesh reconstruction VAE can be measured using reconstruction error calculation. In some examples, the reconstruction error can be calculated as an element-to-element distance between two meshes using, for example, the Euclidean distance. According to various embodiments of the techniques of the present disclosure, other distance measurements are possible, such as cosine distance, Manhattan distance, Minkowski distance, Chebyshev distance, Jaccard distance (e.g., intersection over union of the meshes), Hausdorff distance (e.g., distance across the surface), and Sorensen-Dice distance.
[0237] In some embodiments, the performance of the mesh reconstruction VAE can be verified via a reconstruction error map and / or other key performance indicators. The latent space vectors of one or more input oral care meshes can be plotted (e.g., in 2D) and compared using UMAP or t-SNE dimensionality reduction techniques to select the best available separability between classes of oral care meshes, indicating that the model is aware of strong geometric differences between different classes and strong similarities within a class. This will be illustrated by distinct non-overlapping clusters in the resulting UMAP / t-SNE map.
[0238] In some cases, the latent vectors corresponding to oral care meshes can be used as part of a classifier to classify the mesh (e.g., to identify tooth types or detect errors in the mesh or mesh arrangement, such as during a validation operation). The latent vectors and / or computed mesh features (such as the spatial and / or structural mesh features described herein) can be provided to a supervised machine learning model to classify the mesh. A non-limiting list of possible supervised ML models can be found elsewhere in the present disclosure.
[0239] Figure 6 Shown is a method by which the system of the present disclosure can implement training of an autoencoder for reconstructing an oral care mesh. Figure 7 Specific examples illustrate the training of a variational autoencoder (VAE) for reconstructing a 3D dental arch form mesh. Figure 8Additional steps in the training of a reconstruction autoencoder according to the technology of the present disclosure are described. An oral care mesh 800 (e.g., a 3D dental arch form mesh) can be provided to the input of the method. The system of the present disclosure can perform a registration step (804) to align the oral care mesh with a template example 802 of this type of oral care mesh (e.g., using the iterative closest point technique), where the technical enhancement is to improve the accuracy of the oral care mesh correspondence calculation and the data precision at 806. The system of the present disclosure can calculate the correspondence between the oral care mesh and the corresponding template oral care mesh, where the technical improvement is to adjust the oral care mesh to be ready to be provided to the reconstruction autoencoder. The dataset of the prepared oral care meshes is divided into a training set, a validation set, and a holdout test set (810), and then used to train the reconstruction autoencoder (812), which is described herein as an oral care mesh VAE, an oral care mesh reconstruction VAE, or more generally as a reconstruction autoencoder. The oral care mesh reconstruction VAE can include a 3D encoder that encodes the oral care mesh into a latent form (e.g., a latent vector A), and a subsequent 3D decoder that reconstructs the oral care mesh into a copy of the input oral care mesh. The oral care mesh reconstruction VAE of the present disclosure can be trained using a combination of a reconstruction loss and a KL divergence loss and optionally other loss functions described herein. The output of the method is a trained oral care mesh reconstruction VAE 814.
[0240] One of the steps that may occur in VAE training data preprocessing is the calculation of mesh correspondence. The correspondence between the mesh elements of the input mesh and the mesh elements of a reference or template mesh with a known structure can be calculated. The purpose of the mesh correspondence calculation may be to find the matching points between the surfaces of the input mesh and the template (reference) mesh. The mesh correspondence can generate a point-to-point correspondence between the input mesh and the template mesh by mapping each vertex from the input mesh to at least one vertex in the template mesh. The correspondence between the mesh elements of the input mesh and the mesh elements of a reference or template mesh with a known structure can be calculated. In an example of reconstructing a dental arch form mesh, the range of entries in the vector can correspond to the part of the dental arch form near the upper left first molar; another range of elements can correspond to the lower right central incisor; and so on. In some specific implementations, an input vector (e.g., a landmark vector) can be provided to the autoencoder, and this input vector can define or otherwise affect which type of oral care mesh the autoencoder may have received as input. The data precision improvement of this method of using mesh correspondence in mesh reconstruction is to reduce sampling errors, improve alignment, and enhance mesh generation quality. Further details regarding the use of mesh correspondence in the autoencoder model of the present disclosure can be found elsewhere in the present disclosure.
[0241] In some specific implementations, during the calculation of the mesh correspondence, an Iterative Closest Point (ICP) algorithm can be run between the input oral care mesh and the template oral care mesh. The correspondence can be calculated to establish vertex-to-vertex relationships (between the input oral care mesh and the reconstructed oral care mesh) for use in calculating the reconstruction error.
[0242] Aspects of the training data for the models of the present disclosure are described below. According to a particular implementation, the training data (e.g., for dental arch morphology) can be generalized to the dental arch or a larger oral care representation, or can be more specific to a particular tooth within the larger oral care representation. In cases where more specific training data is utilized, the specific training data can be presented as an oral care mesh template. For example, one oral care mesh template can be specific to one or more oral care mesh types. In some specific implementations, an oral care mesh template can be generated that is the average of many examples of a certain type of oral care mesh. In some specific implementations, an oral care mesh template can be generated that is the average of many examples of more than one oral care mesh type.
[0243] In some specific implementations, the preprocessing process can involve one or more of the following steps: registration for aligning the oral care mesh with the template oral care mesh (e.g., using ICP); and calculation of the mesh correspondence (i.e., to generate a mesh element-to-mesh element correspondence between the input oral care mesh and the template oral care mesh).
[0244] The encoder component E1 of a fully trained mesh reconstruction autoencoder (e.g., for a 3D dental arch form mesh) can generate a latent vector A. The latent vector A can be a dimension-reduced representation of the input mesh (e.g., a 3D dental arch form mesh). In some specific implementations, the latent vector A can be a vector of 128 real numbers (or in other examples consistent with this disclosure, some other size, such as 256 real numbers, 512 real numbers, etc.). The decoder D1 of the fully trained mesh reconstruction autoencoder may be able to take the latent vector A as input and reconstruct a close replica of the input oral care mesh with a low reconstruction error. In some specific implementations, the latent vector A can be modified to effect a change in the shape of the reconstructed oral care mesh output from the decoder D1. This modification can be made after first mapping out the latent space to gain insight into the effects of making specific changes. There can be various loss functions that can be used in the training of E1 and D1, which can involve terms related to the reconstruction error and / or KL divergence between distributions (e.g., in some cases, to minimize the distance between the latent space distribution and a multi-dimensional Gaussian distribution). One purpose of the reconstruction error term is to compare the predicted reconstructed oral care mesh with the corresponding ground truth reconstructed oral care mesh. One purpose of the KL divergence term can be to make the latent space more Gaussian and thus improve the quality of the reconstructed mesh (i.e., especially in cases where the latent space vector can be modified, change the shape of the output mesh, such as modifying the dental arch form, modifying appliance components, modifying CTA trim lines, modifying dental restoration designs, etc.).
[0245] In some specific implementations, the fully trained mesh reconstruction autoencoder can modify the latent vector A in a way that changes one or more characteristics of the reconstructed mesh. If only the reconstruction error is used to calculate the loss L and the latent vector A is changed, then in some use case scenarios, the reconstructed mesh may reflect the expected output form (e.g., by becoming a recognizable dental arch form mesh). However, in other use case scenarios, the output of the reconstructed mesh may not conform to the expected output form (e.g., may not be a recognizable dental arch form mesh).
[0246] In Figure 9 the point P1 corresponds to the initial form of the latent space vector A. Figure 9The latent space is an example of a loss that combines reconstruction loss but not KL divergence loss. Point P2 corresponds to a different location in the latent space that can be sampled as a result of modifying the latent vector A, but where the reconstructed oral care mesh from P2 produces a low-quality output (e.g., an output that does not look like a recognizable or otherwise suitable dental arch-shaped mesh). Point P3 corresponds to yet another different location in the latent space that can be sampled as a result of a different set of modifications to the latent vector A, but where the reconstructed oral care mesh from P3 provides a good output (e.g., having an appearance suitable for use in creating a final setup, such as a dental arch-shaped mesh design that can be used with a setup prediction neural network). In cases where the loss L only involves reconstruction error, the subset of the latent space that can be sampled to obtain latent space vectors P3 that produce valid reconstructed oral care meshes may be irregular and difficult to predict.
[0247] In some specific implementations, the loss calculation can incorporate a normalizing flow, for example, by combining a KL divergence term. Figure 10 An example of a latent space is illustrated where the loss includes both reconstruction loss and KL divergence loss. If the loss is improved by incorporating a KL divergence term, the quality of the latent space can be significantly enhanced. In this new scenario, the latent space may become more Gaussian (as Figure 10 shown), where the latent supervector A corresponds to a point P4 near the center of a multi-dimensional Gaussian curve. Changes can be made to the latent supervector A to obtain a point P5 located near P4, where the resulting reconstructed mesh is likely to reflect the desired properties (e.g., is likely to be a valid dental arch-shaped mesh). Introducing the KL divergence term into the loss can increase the reliability of the process of modifying latent space vectors A and obtaining valid reconstructed oral care meshes. In some specific implementations, like the capsule autoencoder, the training model of the present disclosure can use latent capsules instead of latent vectors and can modify and reconstruct the latent capsules for mesh reconstruction according to aspects of the present disclosure.
[0248] Figure 20 Depicts the reconstruction error of teeth that have been reconstructed by a tooth reconstruction autoencoder (e.g., trained at least in part using the mesh reconstruction autoencoder loss calculation method described herein). Figure 20 Shows a "reconstruction error map" in millimeters (mm). It should be noted that the reconstruction error at the tooth tip is less than 50 micrometers, and the reconstruction error on most of the tooth surface is much less than 50 micrometers. Compared to a typical tooth with a size of 1.0 cm, an error rate of 50 micrometers (or less) means that the tooth surface is reconstructed with an error rate of less than 0.5%.
[0249] For a given domain (e.g., oral care mesh modification), a latent space can be mapped such that a change to latent space vector A can result in a relatively good reconstructed mesh. The latent space can be systematically mapped by generating latent vectors with preselected or predetermined value variations (e.g., by experimenting with different combinations of 128 values in an example latent vector). In some cases, a grid search of values can be performed, providing the advantage of efficiently exploring the latent space. Once the latent space is mapped, the shape of the oral care mesh can be modified by incrementally moving (or "pushing") the values in one or more elements of the latent vector value towards the portion of the mapped latent space that has been found to correspond to the desired oral care mesh characteristics. Using KL divergence augmentation in the loss calculation increases the likelihood of reconstructing the modified latent vector as a valid example of the input 3D representation (e.g., 3D dental arch form mesh).
[0250] In some embodiments, the modification of the latent vector can be performed via an ML model (such as one of the neural network models or other ML models disclosed elsewhere in the present disclosure). In some embodiments, the neural network can be trained to operate within the latent space representation of such a vector A that includes the oral care mesh. The mapping of the latent space of A can represent a previously generated mapping based on applying controlled adjustments to the trial latent vectors and observing the resulting changes in the resulting reconstructed oral care meshes (e.g., after the modified A has been reconstructed back into one or more complete meshes). In some cases, the mapping of the latent space can follow an organized search pattern, such as in the case of implementing a grid search.
[0251] In some specific implementations, the oral care reconstruction VAE of the present disclosure may employ a single input of an oral care mesh name / type / designation R. R can be used to influence the VAE to output an oral care mesh of a specified type (e.g., an arch form mesh or an appliance component such as a parting surface). This can be achieved by generating a latent vector A' used in reconstructing a suitable oral care mesh. In some specific implementations, this latent vector A' can be "instantly" sampled or generated from an existing or previous mapping of the latent vector space. Such a mapping can be performed to provide an understanding of which parts of the latent vector space correspond to different shapes, structures, and / or geometries of oral care meshes. For example, among the 128 real values in an example of the latent vector A' (although it should be understood that other sizes are possible according to the present disclosure), certain elements of those vector elements, and in some cases, certain value ranges of those vector elements, can be determined to correspond to a specific type / name / designation of an oral care mesh and / or an oral care mesh having certain shapes or other desired characteristics. The model for oral care mesh generation can also be applied to the generation of oral care hardware, appliances, and / or appliance components (such as for orthodontic treatment). The model can also be trained to generate dental anatomical structures such as dental crowns and / or dental roots. The model can also be trained to generate other types on non-oral care meshes.
[0252] The system of the present disclosure can calculate the reconstruction error as a combination of the L1 loss and the MSE loss, as shown in the following pseudocode line: reconstruction_error = 0.5*L1(all_points_target,all_points_predicted) +0.5*MSE(all_points_target,all_points_predicted). In the above example, all_points_target can include a point cloud corresponding to a ground truth tooth restoration design (or a ground truth example of some other 3D oral care representation). The "all_points_predicted" variable can include a point cloud corresponding to a generated example of a tooth restoration design (or a generated example of some other kind of 3D oral care representation).
[0253] The following is a possible experiment to gain insight into the connection between changes in the latent vector and the resulting impact on the reconstructed tooth mesh.
[0254] For each dimension of the latent space (e.g., for each unit in a latent vector of, for example, 128 units (although latent vectors of 64 units, 512 units, 1024 units, and other different sizes are possible)), take five (5) data points (0, 1, 2, 3, 4), where point 2 is centered at the mean of that dimension (i.e., value 0, at the center of a Gaussian distribution). Execute the trained encoder-decoder structure (e.g., a transformer trained for tooth reconstruction, or a VAE or capsule autoencoder) to reconstruct the tooth mesh using a latent vector of all zeros. This is the "default" or "average" tooth for this experiment. Then generate 2 samples on each side of the distribution of that dimension (from data points 0, 1, 3, 4 above). These can be located, for example, 1, 1.5, or 2 standard deviations from the center of the distribution of that dimension on each side (positive and negative), enabling the experimenter to gain insight into which aspects of the reconstructed tooth correspond to that dimension of the latent vector. Any one of these dimensions may affect multiple aspects of the reconstructed tooth. In some cases, linear algebra can be used to isolate independent features by identifying the vector dimensions that most significantly affect specific aspects of the shape and / or structure of the reconstructed tooth.
[0255] In some cases, the trained encoder-decoder structure of the present disclosure (e.g., a reconstruction autoencoder or a transformer) can cluster the reconstructed teeth to gain and / or provide insight into their relationships and help understand the connection between changes in the latent vector and the characteristics of the reconstructed teeth. If mesh (A) is the input to a reconstruction autoencoder (or transformer), and mesh (B) is the reconstructed tooth generated by the tooth reconstruction autoencoder (or transformer). By generating latent codes for both A and B, the vector moving from A to B can be calculated (e.g., B - A). A large set of such mapping vectors can be compiled and used to identify the dominant subvectors responsible for pushing points in the latent vector space towards "generalized" restored tooth characteristics. Then such mapping vectors can be added to the latent vector of the tooth to generate a reconstructed tooth mesh with the desired shape and / or structural characteristics. In some cases, the observer of the present disclosure can provide views of three (3) or five (5) meshes at a time and store images of these teeth in various orientations for offline analysis and comparison. Although this experiment uses five (5) data points, any other number of data points can be used in other experiments.
[0256] Although the foregoing example methods discuss modifications to the latent vectors of teeth, it should be understood that, without loss of generality, the methods can also be applied to other 3D representations within the scope of the present disclosure. For example, the latent vectors of any of the following can be mapped: dental restoration designs, other aspects of a patient's dentition, fixture models or fixture model components, appliances, oral care appliance components, etc. In some embodiments, the latent vectors (or other latent representations) can undergo these modifications as part of a Latent Representation Modification Module (LRMM). Oral care variables can be provided to the LRMM to influence the module's modification of the latent vectors, thereby influencing the resulting generated (or modified) 3D representation.
[0257] The autoregressive generative machine learning models of the present disclosure can be trained to estimate one or more missing or incomplete aspects of 3D oral care representations. For example, the models can be trained to add one or more mesh elements to a 3D oral care representation to fill holes, fill rough edges or boundaries, or complete missing portions of the shape or structure of the 3D oral care representation. The transformers of the present disclosure can also be trained for this autoregressive behavior. For example, a transformer can be trained to fill one or more missing mesh elements (e.g., vertices, points, edges, faces, or voxels) in a 3D representation (e.g., a mesh) of a dental crown, which may have holes or other missing aspects due to an intraoral scanning process (e.g., a portion of the dental crown may have been blocked by adjacent teeth or obscured by hardware in the mouth). This mesh completion (or mesh filling or mesh element estimation) technique can be integrated with appliance generation methods, such as to clean up the dental crown (or root) mesh after segmentation and prior to further appliance generation steps (e.g., coordinate system prediction, appliance component generation and / or placement, setting prediction, restoration design generation, etc.). In accordance with aspects of the present disclosure, the generative machine learning models (e.g., a transformer trained for mesh element filling) can utilize a training dataset of past cohort patient case data (e.g., dental mesh data or some other type of 3D oral care representation) to estimate the joint distribution over the mesh elements.
[0258] The techniques described herein can be trained to generate 3D oral care representations (e.g., dental restoration designs, appliance components, and other examples of 3D oral care representations described herein). Such 3D representations can include point clouds, polylines, meshes, voxels, etc. Such 3D oral care representations can be generated according to the requirements of oral care arguments that may be provided to the generative model in some embodiments. Oral care arguments can include oral care parameters as disclosed herein, or other real-valued, text-based, or categorical inputs that specify the desired aspects of one or more 3D oral care representations to be generated. In some cases, oral care arguments can include oral care metrics that can describe the desired aspects of one or more 3D oral care representations to be generated. Oral care arguments are particularly applicable to the embodiments described herein. For example, oral care arguments can specify the desired design (e.g., including shape and / or structure) of 3D oral care representations that can be generated (or modified) according to the techniques described herein. In short, embodiments that use specific oral care arguments disclosed herein generate more accurate 3D oral care representations than embodiments that do not use specific oral care arguments. In some cases, a text encoder can encode a set of natural language instructions from a clinician (e.g., generate a text embedding). The text string can include tokens. In some embodiments, the encoder used to generate the text embedding can apply average pooling or max pooling between token vectors. In some cases, a transformer (e.g., BERT or Siamese BERT) can be trained to extract embeddings of text used in digital oral care (e.g., by training the transformer on examples of clinical text such as those given below). In some cases, such models used to generate text embeddings can be trained using transfer learning (e.g., initially trained on another text corpus and then receiving further training on text related to digital oral care). Some text embeddings can encode text at the word level. Some text embeddings can encode text at the token level. In some embodiments, the transformer used to generate the text embedding can be trained at least in part using a loss calculation that compares the predicted output to the ground truth output (e.g., softmax loss, multi-negative example ranking loss, MSE margin loss, cross-entropy loss, etc.). In some cases, non-text arguments (such as real-valued or categorical values) can be converted to text and subsequently embedded using the techniques described herein. The following are examples of natural language instructions that can be issued by a clinician to the generative model described herein: "Generate a restoration design (alternatively, veneer design) that closes the diastema between teeth #8-9 by uniformly adding width to the mesial surfaces of the two maxillary central incisors", "Generate a restoration design (alternatively, veneer design) that ensures that the incisal edges of #6-11 form a uniform semi-circle with the incisal edges of the posterior teeth (when viewed from the incisal view)", or "Generate a custom crown for the left maxillary central incisor to be implanted".The shape of the dental crown should take into account the shape of adjacent teeth and the space between adjacent teeth should not exceed xmm (e.g., 0.1mm).
[0259] In some specific implementations, the techniques of the present disclosure can use PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using PointNet or PointNet++ as a basis for training) to extract local or global neural network features from 3D point clouds or other 3D representations (e.g., 3D point clouds that describe aspects of a patient's dentition such as teeth or gums). In some specific implementations, the techniques of the present disclosure can use U-Net to extract local or global neural network features from 3D point clouds or other 3D representations.
[0260] 3D oral care representations are described herein because 3-dimensional representations are current state of the art. However, 3D oral care representations are intended to be used in a non-limiting manner to cover any representation of 3 dimensions or higher order dimensions (e.g., 4D, 5D, etc.), and it should be understood that the techniques disclosed herein can be used to train machine learning models to operate on representations of higher order dimensions.
[0261] In some cases, the input data can include 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data related to splines (e.g., control points). The encoder-decoder structure can include one or more encoders, or one or more decoders. In some specific implementations, the encoder can take as input the mesh element feature vectors of one or more input mesh elements in the input mesh. The encoder is trained in a manner that generates a more accurate representation of the input data by processing the mesh element feature vectors. For example, the mesh element feature vectors can provide the encoder with more information about the shape and / or structure of the mesh, and thus the additional information provided allows the encoder to make more informed decisions and / or generate a more accurate latent representation of the mesh. Examples of encoder-decoder structures include U-Net, autoencoders, or transformers, etc. The representation generation module can include one or more encoder-decoder structures (or parts of encoder-decoder structures, such as individual encoders or individual decoders). The representation generation module can generate an information-rich (optionally dimension-reduced) representation of the input data that can be more readily consumed by other generative or discriminative machine learning models.
[0262] The U-Net may include an encoder, followed by a decoder. The architecture of the U-Net may be similar to a U shape. The encoder may extract one or more global neural network features, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level compared to the most global level) from the input 3D representation. The output of each level from the encoder may be passed to the input of the corresponding level of the decoder (e.g., via skip connections). Similar to the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For example, the decoder may output a representation of the input data that may include global, intermediate, or local information about the input data. In some specific implementations, the U-Net may generate an information-rich (optionally dimension-reduced) representation of the input data that may be more easily consumed by other generative or discriminative machine learning models.
[0263] An autoencoder may be configured to encode input data into a latent form. The autoencoder may train the encoder to reformulate the input data into a dimension-reduced latent form between the encoder and the decoder, and then train the decoder to reconstruct the input data from this latent form of the data. A reconstruction error may be computed to quantify the extent to which the reconstructed form of the data differs from the input data. In some specific implementations, the latent form may be used as an information-rich dimension-reduced representation of the input data that may be more easily consumed by other generative or discriminative machine learning models. In most scenarios, the autoencoder may be trained to take an input 3D representation, encode the 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close replica of the input 3D representation as an output.
[0264] The transducer can be trained to generate a representation of its input using self-attention at least in part. The transducer can encode long-range dependencies (e.g., encoding relationships between a large number of inputs). The transducer can include an encoder or a decoder. In some embodiments, such an encoder can operate in a bidirectional manner or can operate a self-attention mechanism. In some embodiments, such a decoder can operate a masked self-attention mechanism, can operate a cross-attention mechanism, or can operate in an autoregressive manner. In some embodiments, the self-attention operation of the transducer described herein can be related to different positions or aspects of a single 3D oral care representation in order to compute a dimensionally reduced representation of the 3D oral care representation. In some embodiments, the cross-attention operation of the transducer described herein can mix or combine aspects of two (or more) different 3D oral care representations. In some embodiments, the autoregressive operation of the transducer described herein can consume previously generated aspects of a 3D oral care representation (e.g., previously generated points, point clouds, transforms, etc.) as additional inputs when generating a new or modified 3D oral care representation. In some embodiments, the transducer can generate a latent form of the input data that can serve as an information-rich dimensionally reduced representation of the input data and that can be more readily consumed by other generative or discriminative machine learning models.
[0265] In some embodiments, an encoder-decoder architecture can first be trained as an autoencoder. In deployment, one or more modifications can be made to the latent form of the input data. Then, the modified latent form can continue to be reconstructed by the decoder, resulting in a reconstructed form of the input data that is different from the input data in one or more desired aspects. Oral care variables such as oral care parameters or oral care metrics can be provided to the encoder, the decoder, or can be used to modify the latent form in order to influence the encoder-decoder architecture when generating a reconstructed form having desired characteristics (e.g., characteristics that may be different from the characteristics of the input data).
[0266] In some cases, federated learning can be used to train the techniques of the present disclosure. Federated learning can enable multiple remote clinicians to iteratively improve machine learning models (e.g., validation of 3D oral care representations, mesh segmentation, mesh cleaning, other techniques involving tagging mesh elements, coordinate system prediction, placement of non-organic objects on teeth, appliance component generation, dental restoration design generation, techniques for placing 3D oral care representations, setting prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, estimation of missing values), while protecting data privacy (e.g., clinical data may not need to be sent "over the network" to a third party). Data privacy is particularly important for clinical data protected by applicable laws. A clinician can receive a copy of the machine learning model, use a local machine learning program to further train the ML model using locally available data from a local clinic, and then send the updated ML model back to a central hub or a third party. The central hub or third party can integrate the updated ML models from multiple clinicians into a single updated ML model that benefits from the learning of patient data recently collected at various clinical sites. In this way, a new ML model can be trained that benefits from additional and updated patient data (possibly from multiple clinical sites), while that patient data is never actually sent to a third party. In some cases, training on devices within a local clinic can be performed when the device is idle or otherwise during non-working hours (e.g., when patients are not being treated at the clinic). Devices in a clinical environment for collecting data and / or training an ML model for the techniques described herein can include intraoral scanners, CT scanners, X-ray machines, laptop computers, servers, desktop computers, or handheld devices (such as a smartphone with image collection capabilities). In addition to federated learning techniques, in some embodiments, contrastive learning can be used to at least partially train the ML models described herein. In some cases, contrastive learning can augment samples in a training dataset to emphasize differences between samples of different classes and / or increase the similarity of samples of the same class.
[0267] In some cases, a local coordinate system for 3D oral care representations (such as teeth) can be described by one or more transformations (e.g., an affine transformation matrix, a translation vector, or a quaternion). The systems of the present disclosure can be trained for coordinate system prediction using the coordinate systems of past cohort patient cases. The past patient data can include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems. Machine information models (such as U-Net, an encoder, an autoencoder, a pyramid encoder-decoder, a transformer, or convolutional layers and / or pooling layers) can be trained for coordinate system prediction. Representation learning can determine a representation of the teeth (e.g., encoding a mesh or point cloud into a latent representation, such as using U-Net, an encoder, a transformer, or convolutional layers and / or pooling layers, etc.), and then predict a transformation for that representation (e.g., using a trained multi-layer perceptron, a transformer, an encoder, a transformer, etc.), which defines a local coordinate system for that representation (e.g., including one or more coordinate axes). In the case of predicting the coordinate system for a tooth mesh, the mesh convolution techniques described herein can utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be trained for coordinate system prediction in the form of predicting the transformation of a tooth. Reinforcement learning techniques can be trained for coordinate system prediction in the form of predicting the transformation of a tooth.
[0268] Machine information models (such as U-Net, encoders, autoencoders, pyramid encoder-decoders, transformers, or convolutional layers and / or pooling layers) can be trained as part of a method for hardware (or appliance component) placement. Representation learning can train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder or using blocks of U-Net, encoders, transformers, convolutional layers, and / or pooling layers, etc.). The representation can include a dimensionally reduced form and / or an information-rich form of the input 3D oral care representation. In some embodiments, the generation of the representation can be assisted by calculating a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some embodiments, a representation can be calculated for a hardware element (or appliance component). Such a representation is adapted to be provided to a second module, which can perform a generation task such as transform prediction (e.g., a transform for placing a 3D oral care representation relative to another 3D oral care representation, such as a representation for placing a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such a transform can include an affine transformation matrix, a translation vector, or a quaternion, etc. Machine learning models that can be trained to predict a transform for placing a hardware element (or appliance component) relative to elements of a patient dentition include MLP, transformers, encoders, etc. The systems of the present disclosure can be trained for 3D oral care appliance placement using past cohort patient case data. The past patient data can at least include: one or more ground truth transforms and one or more 3D oral care representations (such as tooth meshes or other elements of a patient dentition). In the case where U-Net (and other neural networks) is trained to generate a representation of a tooth mesh, the mesh convolution and / or mesh pooling techniques described herein utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be performed for hardware or appliance component placement. Reinforcement learning techniques can be performed for hardware or appliance component placement.
[0269] The techniques of the present disclosure can be trained to generate point clouds (e.g., where points can be described as 1D vectors, such as (x, y, z)), polylines (points connected in sequence by edges), meshes (points connected by edges to form faces), splines (which can be calculated through a set of generated control points), sparse voxelized representations (which can be described as a set of points corresponding to the centroid of each voxel or some other landmark of the voxel (such as the boundary of the voxel)), transforms (which can take the form of one or more 1D vectors or one or more 2D matrices (such as a 4×4 matrix)), etc. In some embodiments, a voxelized representation can be calculated from a 3D point cloud or 3D mesh. In some embodiments, a 3D point cloud can be calculated from a voxelized representation. In some embodiments, a 3D mesh can be calculated from a 3D point cloud.
[0270] This disclosure presents a Recursive Inference (RI) model for 3D oral care representation generation (or modification). The RI model can be trained on a training dataset of a specific type of 3D oral care representation, and examples of specific types of 3D oral care representations are described herein. Then, the RI model can be used in deployment to generate (or modify) examples of that specific type of 3D oral care representation. 3D oral care representations that can be generated or modified (based on the corresponding training data) include: 1) dental restoration designs; 2) fixture model designs (e.g., including one or more fixture model components or one or more aspects of a patient's dentition); 3) oral care appliance components (e.g., generated components); 4) tooth setup transformations; 5) transformations of prefabricated library components for oral care appliances; 6) dental arch morphology; 7) clear aligner trimming lines (e.g., for trimming an aligner tray from a printed fixture model); 8) a set of mesh element labels used in segmentation or mesh cleaning; or other 3D oral care representations described herein. For example, an RI model for generating (or modifying) a dental restoration design can be trained on dental restoration design data. Additionally, an RI model for generating (or modifying) a digital fixture model design can be trained on dental restoration design data (e.g., which includes aspects of the dentition, or one or more fixture model components, as described herein). In another example, an RI model for generating (or modifying) tooth transformations for setup prediction can be trained on tooth mesh and transformation data (e.g., malocclusion transformations of teeth, and ground truth reference transformations for the final setup or intermediate stages, which will be used in loss calculations).
[0271] Dental arch morphology can be described by control points (with splines) or polylines. Trimming lines can be described by polylines (or control points and / or splines). Dental restoration designs, fixture models, or generated appliance components can be described by 3D meshes, 3D point clouds, 3D voxelized representations, etc. Transformations can be described by transformation matrices or vectors, translation vectors, quaternions, or other data structures described herein.
[0272] The techniques of this disclosure provide technical improvements over existing systems and techniques. For example, the techniques of this disclosure provide technical improvements to the technical problems of generating (or modifying) 3D oral care representations used in generating oral care appliances, specifically improvements to the introduction of mesh element features, oral care metrics, or oral care parameters.
[0273] In addition, in some specific implementations, the techniques of the present disclosure can be trained to generate other types of data representations not envisioned by existing systems and techniques, including transformations, coordinate system axes (e.g., for other aspects of teeth or a patient's dentition), or mesh element labels as two examples. The techniques of the present disclosure can be trained on oral care data (e.g., a mesh, point cloud, or voxelized representation depicting dental anatomy or appliance components, a transformation that positions teeth or appliance components in a pose suitable for clinical handling, a mesh element label that can be defined for use with a segmentation or mesh cleaning operation; or other examples of 3D oral care representations described herein) to generate a 3D oral care representation suitable for generating an oral care appliance.
[0274] During training or deployment, the techniques of the present disclosure can take as input a 3D oral care representation that can be encoded into a latent form or latent representation by a first ML module (e.g., an encoder). In some specific implementations, the first ML module can be trained at least in part by computing a reconstruction loss (e.g., cross-entropy loss), which can compare the generated output to a ground truth reference. When processing mesh, voxel, or point cloud data (or data describing other 3D representations) by the techniques of the present disclosure, technical improvements are achieved by computing mesh element feature vectors for one or more mesh elements. Such mesh element feature vectors can be computed by the mesh element feature module 1102. Such mesh element feature vectors can be provided to the first ML module and can improve the accuracy or fidelity of the latent form generated by the first ML module. The improved accuracy of the latent form can enable subsequent generation steps to output improved generated (or modified) results (e.g., a 3D representation of a tooth for restoration, a set of transformations used in orthodontic setup generation, one or more coordinate system axes, or mesh element labels used in segmentation or mesh cleaning).
[0275] Further technical improvements are achieved by the techniques of the present disclosure by using oral care parameters collectively referred to as oral care variables (e.g., which may specify customization features of an expected 3D oral care representation to be generated or modified) or oral care metrics (e.g., which may quantify or measure physical aspects of one or more teeth; which may quantify the shape and / or structure of an individual tooth or appliance component; or which may quantify the pose and / or physical arrangement between two or more teeth or appliance components). Oral care variables are particularly applicable to the specific implementations described herein. For example, an oral care variable may specify an expected design (e.g., including shape and / or structure) of a 3D oral care representation that may be generated (or modified) according to the techniques described herein. In short, specific implementations using the specific oral care variables disclosed herein generate more accurate 3D oral care representations than specific implementations that do not use specific oral care variables. An oral care metric 1108 may be calculated for a training example (e.g., a set of teeth of a patient) and provided to either a model training method or a model deployment method of the RI model, thereby generating an enhanced feature grid AFG t . Such oral care metrics improve the enhanced feature grid AFG by quantifying specific key aspects of the input 3D oral care representation 1100 (e.g., by quantifying the shape of one or more teeth or appliance components, or by measuring the special relationship between one or more teeth or appliance components) t , and the generator module 1124 may use these key aspects to generate a customized output suitable for use in clinical treatment (e.g., generating a customized tooth restoration design, a customized appliance component, a customized orthodontic setup transformation, coordinate axes, etc.). The oral care metric 1108 may be provided to the generator module 1124 at training time to teach the generator module 1124 information about the shape and / or structure of ...
Claims
1. A method for generating a three-dimensional (3D) representation of oral care data for an oral care treatment, the method comprising: Receiving, by a processing circuit of a computing device, an input 3D representation of a patient's dentition; Executing, by the processing circuit, a trained first ML module to encode the input 3D representation of the patient's dentition into a first latent representation having a lower dimensional order than the input 3D representation of the patient's dentition; Executing, by the processing circuit, a trained second ML module, the trained second ML module including at least one of a trained transformer encoder model or a trained transformer decoder model, to: Generate a second latent representation using the first latent representation of the input 3D representation of the patient's dentition; Reconstructing, by the decoder, the second latent representation into a reconstructed 3D oral care representation; and Outputting, by the processing circuit, the reconstructed 3D representation of the oral care data.
2. The method according to claim 1, the method further comprising using the reconstructed 3D representation of the oral care data to generate one or more oral care appliances.
3. The method according to claim 1, the method further comprising: Receiving, by the processing circuit of the computing device, one or more oral care variables; Influencing the generation by the one or more oral care variables.
4. The method according to claim 1, the method further comprising: Executing, by the processing circuit, a trained latent representation modification module (LRMM); Modifying, by the LRMM, one or more aspects of the first latent representation; and Using the one or more modified aspects of the first latent representation to generate the reconstructed 3D tooth representation.
5. The method according to claim 4, the method further comprising: Receiving, by the processing circuit of the computing device, the one or more oral care variables; Influencing the operation of the LRMM by the one or more oral care variables.
6. The method according to claim 2, wherein the oral care appliance is a dental restoration appliance.
7. The method according to claim 2, wherein the oral care appliance is an orthodontic appliance.
8. The method according to claim 1, wherein the reconstructed 3D representation of the oral care data is a representation of an appliance component.
9. The method according to claim 1, wherein the reconstructed 3D representation of the oral care data is a representation of an arch form.
10. The method according to claim 1, wherein the reconstructed 3D representation of the oral care data is a representation of the patient's dentition having at least one generated fixture model component.
11. The method according to claim 1, wherein the input 3D representation of the oral care data includes a representation of at least one tooth in a pre-restoration state.
12. The method according to claim 11, wherein the reconstructed 3D representation of the oral care data is a representation of at least one tooth in a post-restoration state.
13. The method according to claim 1, wherein the input 3D representation of the patient's dentition comprises one or more mesh elements, and the method further comprises calculating, by the processing circuitry, one or more mesh element features.
14. The method according to claim 13, the method further comprising providing, by the processing circuitry, the one or more mesh element features as an input to at least one of the first ML module or the second ML module.
15. The method according to claim 1, wherein the first ML module comprises one or more of the following: an encoder, a U-Net, a pyramid encoder-decoder, a 3D SWIN transformer, one or more convolutional layers, or one or more pooling layers.
16. The method according to claim 1, the method further comprising training a machine learning (ML) model using at least one of the trained first ML module or the trained second ML module according to a transfer learning paradigm.
17. The method according to claim 1, wherein at least one of the trained first ML module or the trained second ML module is trained according to a transfer learning paradigm.
18. The method according to claim 1, wherein the computing device is deployed in a clinical environment, and wherein the method is performed near real-time during a session with the patient.
19. The method according to claim 15, wherein the trained first ML module is configured to generate one or more hierarchical neural network features, and the method further comprises generating the one or more hierarchical neural network features based at least in part on one or more aspects of at least one of the shape or structure of the input 3D representation.
20. A computing device for generating a three-dimensional (3D) representation of oral care data for oral care treatment, the computing device comprising: Interface hardware configured to receive an input three-dimensional (3D) representation of the patient's dentition; Processing circuitry configured to: Execute a trained first ML module to encode the input 3D representation of the patient's dentition into a first latent representation having a lower dimensional order than the input 3D representation of the patient's dentition; Execute, by the processing circuitry, a trained second ML module, the trained second ML module comprising at least one of a trained transformer encoder model or a trained transformer decoder model, to: Generate a second latent representation using the first latent representation of the input 3D representation of the patient's dentition; Reconstruct the second latent representation into a reconstructed 3D oral care representation by the decoder; and Output, by the processing circuitry, the reconstructed 3D representation of the oral care data.
Citation Information
Patent Citations
Method for automated generation of orthodontic treatment final setups
US20210259808A1
Method for automated generation of orthodontic treatment final setups
WO2020026117A1
Automated creation of tooth restoration dental appliances
WO2020240351A1
Neural network-based generation and placement of tooth restoration dental appliances
WO2021240290A1
System to generate staged orthodontic aligner treatment
WO2021245480A1