Auto-encoder for verifying 3D oral care representations
Verifying 3D oral care representations through the encoder-decoder structure of training machine learning models solves the problem of verification failure and improves the accuracy and efficiency of device generation.
Patent Information
- Application Number
- CN202380086160.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2023-12-14
- Publication Date
- 2025-07-18
AI Technical Summary
Existing technologies are difficult to effectively verify whether 3D oral care representation is suitable for creating oral care appliances, resulting in the possible verification failure.
Using machine learning models, especially encoder-decoder structures such as reconstruction of autoencoders, determine whether they meet the distribution of the benchmark real examples by training and verifying 3D oral care representations, thereby determining their suitability.
It improves the accuracy and reliability of 3D oral care representation, ensures that there is no verification failure during the generation of oral care appliances, reduces computing resource consumption and improves processing efficiency.
Smart Images

Figure CN120345003A_ABST
Abstract
Description
Related Literature
[0001] The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosure of each of the PCT applications with publication numbers WO2022123402A1, WO2021245480A1, and WO2020026117A1 is incorporated herein by reference. The entire disclosure of each of the following provisional U.S. patent applications is incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 366,514; and 63 / 264,914. Technical Field
[0002] The present disclosure relates to the configuration and training of machine learning models to improve the accuracy of automatically verifying 3D oral care representations using one or more autoencoders to determine whether those 3D oral care representations are suitable for use in creating oral care appliances such as dental restorative appliances or orthodontic processing appliances (e.g., clear tray aligners or indirect bonding trays). Summary of the Invention
[0003] The present disclosure describes systems and techniques for training and using one or more machine learning models such as neural networks to verify 3D oral care representations as acceptable (or unacceptable) for use in creating oral care appliances. For example, the techniques of the present disclosure can train an encoder-decoder architecture, such as a reconstruction autoencoder (e.g., a variational autoencoder optionally utilizing normalizing flows), to reconstruct a particular type of 3D oral care representation. The reconstruction autoencoder can be trained on a benchmark ground truth example of the 3D oral care representation until the reconstruction autoencoder becomes capable of reconstructing a trial example of that type of 3D oral care representation. When reconstructing a trial 3D oral care representation that falls within the distribution of examples in the training dataset, the reconstruction autoencoder can produce a low reconstruction error. A high reconstruction error can indicate that the trial 3D oral care representation does not fall within such a training distribution, which can result in a failed verification output. If a trial 3D oral care representation passes verification (e.g., falls within the distribution of examples from the training dataset), the verification techniques described herein can emit an output indicating that the trial 3D oral care representation is suitable for use in creating oral care appliances (e.g., clear tray aligners, indirect bracket bonding trays, dental restorative appliances, etc.).
[0004] The encoder-decoder structure may include at least one encoder or at least one decoder. Non-limiting examples of the encoder-decoder structure include 3D U-Net, transformers, pyramid encoder-decoders, or autoencoders, etc. In some specific implementations, the verification techniques described herein may incorporate aspects derived from a denoising diffusion model (e.g., a neural network that can be trained to iteratively denoise one or more 3D oral care representations, such as 3D oral care representations with random initialization or Gaussian noise). In some specific implementations, the verification techniques described herein may use one or more neural networks trained to use mathematical operations associated with continuous normalizing flows (e.g., neural networks that can be trained in one form and then inverted for use in inference). Non-limiting examples of autoencoders include variational autoencoders, regularized autoencoders, masked autoencoders, or capsule autoencoders.
[0005] According to the techniques of the present disclosure, a machine learning (ML) model, such as an autoencoder, can be trained on examples of 3D oral care representations, where ground truth data is provided to the ML model, and a loss function is used to quantify the difference between the predicted examples and the ground truth examples. The loss value can then be used to update the verified ML model (e.g., update the weights of the neural network). Such verification techniques can determine whether a trial 3D oral care representation is acceptable or suitable for use in creating an oral care appliance. In some cases, "acceptable" indicates that the trial 3D oral care representation conforms to the distribution of the ground truth examples used in training the ML verification model. In some cases, "acceptable" indicates that the trial 3D oral care representation is correctly shaped (or structured) or correctly positioned with respect to one or more aspects of the dental anatomy.
[0006] In an example of the generated appliance component (e.g., a dental restoration appliance component, such as a mold parting surface), the autoencoder-based verification technique can determine whether the component intersects with the correct landmarks or other parts of the tooth anatomy (e.g., the incisal edge and the tooth tip for a mold parting surface). The determination that the verification technique of the present disclosure can output includes one or more of the following: whether the CTA trim line intersects the gingiva in a manner that reflects the distribution of the ground truth, whether the library component is correctly placed relative to one or more target teeth (e.g., a snap fixture placed relative to a posterior tooth or a center fixture placed relative to an incisor) or correctly placed relative to one or more landmarks on the target tooth. Other non-limiting examples include: determining whether a hardware element is placed on the face of a tooth with an edge that reflects the distribution of the ground truth examples; whether the mesh element markings for a segmentation (or mesh cleanup) operation conform to the distribution of the labels in the ground truth examples; and / or whether the shape and / or structure of the dental restoration tooth design conforms to the distribution of the tooth designs in the ground truth training examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 Shows a method for enhancing training data used in training a machine learning (ML) model of the present disclosure.
[0008] Figure 2 Shows a method for training a capsule autoencoder.
[0009] Figure 3 Shows a method for training a tooth reconstruction autoencoder.
[0010] Figure 4 Shows a method for using a deployed and fully trained tooth reconstruction autoencoder.
[0011] Figure 5 Shows a reconstructed tooth mesh reconstructed using a reconstruction autoencoder according to the techniques of the present disclosure.
[0012] Figure 6 Shows a reconstructed tooth mesh reconstructed using a reconstruction autoencoder according to the techniques of the present disclosure.
[0013] Figure 7 Shows a visualization of the reconstruction error of a tooth.
[0014] Figure 8 Shows the reconstruction error values of several tooth reconstructions.
[0015] Figure 9 Shows a method for training a reconstruction autoencoder.
[0016] Figure 10 Shows a non - limiting example code for a reconstruction autoencoder.
[0017] Figure 11 Shows an example of a reconstructed 3D representation according to the techniques of the present disclosure.
[0018] Figure 12 Shows a latent space in which the loss includes reconstruction loss but not KL divergence loss.
[0019] Figure 13 Shows a latent space in which the loss includes both reconstruction loss and KL divergence loss.
[0020] Figure 14 Shows a method for using a fully trained reconstruction autoencoder to validate a 3D oral care representation. Detailed Description
[0021] Multiple techniques in digital oral care can benefit from using a first module (e.g., an autoencoder neural network) that has been trained to reconstruct a 3D oral care representation (e.g., trained to reconstruct a tooth mesh including crowns, roots, and / or attached artifacts). A 3D encoder can be trained to encode the oral care mesh into a latent form, and a 3D decoder can be trained to reconstruct the latent form into a replica of the received oral care mesh, where the techniques disclosed herein can be used to measure the resulting reconstruction error. The first module can create a representation. A second module can use the representation for prediction. There can be one or more instances of the first module, and there can be one or more instances of the second module.
[0022] Techniques are described herein that can leverage an autoencoder trained for oral care mesh reconstruction, which provides the advantage of encoding a potentially complex oral care mesh into a latent form (e.g., such as a latent vector or latent capsule) that can have a reduced dimension and can be ingested by an instance of a second module (e.g., a prediction model for mesh cleanup, pose prediction, tooth restoration design generation, classification of 3D representations, validation of 3D representations, or pose comparison) for prediction purposes. Although the dimension of the latent form can be reduced relative to the received oral care mesh, information about the reconstruction characteristics of the received oral care mesh can be retained. This latent representation of the initial oral care mesh can be received as an input to the prediction model of the second module, thus providing the advantage of improved accuracy and data precision compared to other techniques. In some embodiments, the latent representation can be modified according to the techniques of the present disclosure to enable the prediction model of the second module to customize the output data. The advantage of computing the reconstruction error on the reconstructed oral care mesh is to verify that the reconstructed oral care mesh is a replica of the received oral care mesh (e.g., where one or more dimensions or other aspects of the reconstructed oral care mesh are measured to be within a threshold reconstruction error of the received oral care mesh). In some embodiments, the first module can also be trained to produce other kinds of representations, such as those generated by a neural network that performs convolutional and / or pooling operations (e.g., a network with a convolutional kernel of size 5 that also performs average pooling, or a network such as a U-Net).
[0023] Either or both of the first module and / or the second module may receive various input data as described herein, including a dental mesh of one or both dental arches of a patient. The dental data may be presented in the form of a 3D representation such as a mesh, a point cloud. Such data may be pre-processed, for example, by arranging the constituent mesh elements into a list and calculating an optional mesh element feature vector for each mesh element. Such feature vectors may provide valuable information about the shape and / or structure of the oral care mesh to either or both of the first module and / or the second module. For example, the first module that generates the representation may receive the vertices of a 3D mesh (or 3D point cloud) and calculate a mesh element feature vector for each vertex. In addition to the other optional mesh element features described herein, such feature vectors may also include the XYZ coordinates of each vertex. Additional inputs may be received at an entry point of either or both of the first module and / or the second module, such as one or more oral care metrics. Oral care metrics may be used to measure one or more physical aspects of the oral care mesh (e.g., physical relationships within or between teeth). In some cases, oral care metrics may be calculated for either or both of an orthodontic oral care mesh example and a ground truth oral care mesh example and then used in the training of either or both of the first module and the second module. Metric values may be received as an input to either or both of the first module and the second module to train the underlying model of that particular module to encode the distribution of such metrics across a number of examples in the training dataset. During training, the network may then receive the metric value as an input to help train the network to link the metric value of the input to the physical aspects of the ground truth oral care mesh used in the loss calculation. Such loss calculation may quantify the difference between the prediction and the ground truth example (e.g., between a predicted oral care mesh and a ground truth oral care mesh). By providing network data that describes the metric values, the techniques of the present disclosure may train the network to encode the distribution of a given metric through the process of loss calculation and subsequent backpropagation. In deployment, one or more oral care parameters (protocol parameters or prosthetic design parameters) may be defined to specify one or more aspects of an expected oral care mesh that will be generated using either or both of the first module and / or the second module that have been trained for that purpose. In some embodiments, an oral care parameter corresponding to an oral care metric may be defined, which may be received as an input to either or both of the deployed first module and / or the deployed second module and be treated as an instruction to that module to generate an oral care mesh with the specified customization. This interaction between oral care metrics and oral care parameters may also apply to the training and deployment of other predictive models in oral care.
[0024] In some specific implementations, the prediction model of the present disclosure can obtain more accurate results by combining one or more of the following inputs: arch form information V, interproximal reduction (IPR) information U, tooth size information P, diastema information Q, latent capsule representation T of the oral care mesh, latent vector representation A of the oral care mesh, protocol parameter K (which can describe the clinician's expected treatment of the patient), doctor preference L (which can describe the typical protocol parameters selected by the doctor), flag M regarding tooth status (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name / dental symbol R, oral care metric S (including at least one of oral care metrics and prosthetic design metrics).
[0025] In some cases, the system of the present disclosure can be deployed in a clinical environment (such as a dental or orthodontic clinic) for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians). Such a system deployed in a clinical environment can enable clinicians to process oral care data (such as tooth scans) in a clinical environment or in some cases in a "chairside" environment (when the patient is in the clinical environment). A non-limiting list of examples of techniques can include: segmentation, mesh cleaning, coordinate system prediction, CTA trim line generation, prosthetic design generation, appliance component generation or placement or assembly, generation of other oral care meshes, verification of oral care meshes, setup prediction, removal of hardware from tooth meshes, placement of hardware on teeth, estimation of missing values, clustering of oral care data, oral care mesh classification, setup comparison, metric calculation, or metric visualization. In some cases, the execution of these techniques can enable patient data to be processed, analyzed, and used by clinicians in appliance generation before the patient leaves the clinical environment (which can facilitate treatment planning as feedback can be received from the patient during the treatment planning process).
[0026] The systems of the present disclosure can automate operations in digital orthodontics (e.g., setup prediction, hardware placement, setup comparison), digital dentistry (e.g., prosthetic design generation), or combinations thereof. Some techniques can be applied to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleaning, coordinate system prediction, oral care mesh validation, estimation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows, or denoising diffusion models), metric visualization, appliance component placement, or appliance component generation, etc. In some cases, the systems of the present disclosure can enable a clinician or technician to process oral care data (such as a scanned dental arch). In addition to segmentation, mesh cleaning, coordinate system prediction, or validation operations, the systems of the present disclosure can also implement an orthodontic treatment plan, which may involve setup prediction as at least one operation. The systems of the present disclosure can also implement prosthetic design generation, in which one or more restored tooth designs are generated and processed during the creation of an oral care appliance. The systems of the present disclosure can automate steps in either or both of orthodontic or dental treatment plans, or can automate steps in the generation of either or both of orthodontic or dental appliances. Some appliances can implement both dental and orthodontic treatment, while other appliances can implement one or the other.
[0027] The techniques of the present disclosure may require a training dataset of hundreds or thousands of cohort patient cases to ensure that a neural network can encode the distribution of patient cases that may be encountered in clinical treatment. Cohort patient cases can include a set of crown meshes, a set of root meshes, or a data file (e.g., a JSON file) including case attributes. Typical examples of cohort patient cases can include up to 32 crown meshes (e.g., each of which can include tens of thousands of vertices or faces), up to 32 root meshes (e.g., each of which can include tens of thousands of vertices or faces), multiple gingival meshes (e.g., each of which can include tens of thousands of vertices or faces), or one or more JSON files (each of which can include tens of thousands of values (e.g., objects, arrays, strings, real values, boolean values, or null values)).
[0028] Aspects of the present disclosure may provide a technical solution to the technical problem of validating 3D oral care representations used in oral care appliance generation (e.g., orthodontic appliance trays, dental restoration appliances, indirect bonding trays, etc.) using a trained autoencoder. Specifically, by practicing the techniques disclosed herein, a computing system specifically adapted to validate 3D oral care representations for oral care appliance generation is improved. For example, aspects of the present disclosure improve the performance of a computing system having a 3D representation of a patient's dentition by reducing the consumption of computing resources. Specifically, aspects of the present disclosure reduce computing resource consumption by decimating the 3D representation of the patient's dentition (e.g., reducing the count of mesh elements used to describe aspects of the patient's dentition), such that computing resources are not wasted unnecessarily due to processing an excessive number of mesh elements. Additionally, decimating the mesh does not degrade the overall prediction accuracy of the computing system (and may actually improve prediction because the input provided to the ML model after decimation is a more accurate (or better) representation of the patient's dentition). For example, unimportant (and potentially accuracy-degrading) noise or other artifacts are removed. That is, aspects of the present disclosure provide a more efficient allocation of computing resources in a manner that improves the accuracy of the underlying system.
[0029] In addition, aspects of the present disclosure may need to be performed in a time-limited manner, such as when an oral care appliance must be generated for a patient immediately after an intraoral scan (e.g., when the patient is waiting in a clinician's office). Thus, aspects of the present disclosure must be rooted in the underlying computer technology of setting transformation prediction for oral care appliance generation and cannot be performed by a human even with the aid of pen and paper. For example, specific implementations of the present disclosure must be able to: 1) store thousands or millions of mesh elements of a patient's dentition in a manner that can be processed by a computer processor; 2) perform computations on thousands or millions of mesh elements, such as to quantify aspects of the shape and / or structure of an individual tooth in a 3D representation of the patient's dentition; 3) encode the thousands or millions of mesh elements into a latent representation; 4) reconstruct the latent representation into a 3D representation comprising thousands or millions of mesh elements; 4) compute a reconstruction error; and 5) validate one or more 3D oral care representations based on a trained machine learning model and do so during the course of a short patient visit.
[0030] The present disclosure relates to digital oral care covering the fields of digital dentistry and digital orthodontics. The present disclosure generally describes methods for processing three-dimensional (3D) representations of oral care data. It should be understood that while not losing generality, there are various types of 3D representations. One type of 3D representation is 3D geometry. The 3D representation can include, be one or more of, or be part of: a 3D polygon mesh, a 3D point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels for sparse processing), or a 3D representation described by mathematical equations. Although the term "mesh" is frequently used throughout the present disclosure, in some specific implementations, this term should be understood to be interchangeable with other types of 3D representations. The 3D representation can describe the 3D geometry and / or elements of the 3D structure of an object.
[0031] The dental arches S1, S2, S3, and S4 all include exactly the same dental meshes, but these dental meshes are transformed differently according to the following description. The first dental arch S1 includes a set of dental meshes that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in malposition and misorientation. The second dental arch S2 includes the same set of dental meshes from S1 that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a reference true-setting position and orientation. The third dental arch S3 includes the same meshes as S1 and S2 that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a predicted final-setting pose (e.g., as predicted by one or more techniques of the present disclosure). S4 is the counterpart of S3 where the teeth are in a pose corresponding to one of several intermediate stages of orthodontic treatment with a clear tray aligner.
[0032] It should be understood that while not losing generality, the techniques of the present disclosure applied to the final setting are also applicable to intermediate gradings in orthodontic treatment, specifically geometric deep learning (GDL) settings, reinforcement learning (RL) settings, variational autoencoder (VAE) settings, capsule settings, multi-layer perceptron (MLP) settings, diffusion settings, pose transfer (PT) settings, similarity settings, force-directed graph (FDG) settings, transformer settings, setting comparison, or setting classification. The metric visualization aspect of the present disclosure can also be configured to visualize data from both the final setting and intermediate stages. The MLP setting, VAE setting, and capsule setting each fall within the scope of the autoencoder setting. Some specific implementations of the MLP setting can fall within the scope of the transformer setting. A representation setting refers to any one of the MLP setting, VAE setting, capsule setting, and any other setting prediction machine learning model that uses an autoencoder to create a representation of at least one tooth.
[0033] Each of the setup prediction techniques of the present disclosure is applicable to the manufacture of clear tray aligners and / or indirectly bonded trays. The setup prediction techniques may also be applicable to other products that also involve final tooth positions. The position may include location (or positioning) and rotation (or orientation).
[0034] A 3D mesh is a data structure that can describe the geometry and / or shape of an object related to oral care, which object includes but is not limited to teeth, hardware elements, or the gingival tissue of a patient. The 3D mesh may include one or more mesh elements, such as one or more of vertices, edges, faces, and combinations thereof. In some specific implementations, the mesh elements may include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features can be calculated for these mesh elements and provided to the prediction models of the present disclosure, and the prediction models of the present disclosure provide the technical advantage of improved data accuracy in the form of more accurate predictions of the model outputs of the present disclosure.
[0035] The patient dentition may include one or more 3D representations of the patient's teeth (e.g., and / or associated transducers), gums, and / or other oral anatomical structures. In some embodiments, an orthodontic metric (OM) may quantify the relative position and / or orientation of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. In some embodiments, a restorative design metric (RDM) may quantify at least one aspect of the structure and / or shape of a 3D representation of a tooth. In some embodiments, an orthodontic landmark (OL) may locate one or more points or other regions of interest structures on a 3D representation of a tooth. In some embodiments, the OL may be used in the generation of orthodontic or prosthetic appliances, such as clear tray aligners or prosthetic restorative appliances. In some embodiments, a mesh element may include at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth represented by a 3D mesh, the mesh elements may at least include: vertices, edges, faces, and voxels. In some embodiments, mesh element features may quantify some aspects of the 3D representation that are proximal to or associated with one or more mesh elements, as described elsewhere in the present disclosure. In some embodiments, an orthodontic procedure parameter (OPP) may specify at least one value that defines at least one aspect of a patient's planned orthodontic treatment (e.g., specifying desired target attributes of a final setting in a final setting prediction). In some embodiments, an orthodontist preference (ODP) may specify at least one typical value of the OPP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. In some embodiments, a restorative design parameter (RDP) may specify at least one value that defines at least one aspect of a patient's planned prosthetic restorative treatment (e.g., specifying desired target attributes of a tooth to be treated with a prosthetic restorative appliance). In some embodiments, a doctor restorative design preference (DRDP) may specify at least one typical value of the RDP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. The 3D oral care representation may include but is not limited to: 1) a set of mesh element labels that may be applied to 3D mesh elements of a tooth / gum / hardware / appliance mesh (or point cloud) during the process of mesh segmentation or mesh cleaning; 2) one or more 3D representations of teeth / gums / hardware / appliances whose shapes have been modified (e.g., trimmed, deformed, or filled) during the process of mesh segmentation or mesh cleaning; 3) one or more coordinate systems (e.g., describing one, two, three, or more coordinate axes) for a single tooth or a group of teeth (such as a full dental arch, e.g., the LDE coordinate system); 4) 3D representations of one or more teeth whose shapes have been modified or otherwise made suitable for use in prosthetic restorations; 5) 3D representations of one or more prosthetic restorative appliance components;6) One or more transformations to be applied to one or more of the following: placement of prosthetic appliance library components relative to one or more teeth, teeth to be placed for an orthodontic setting (final setting or intermediate stage), hardware elements to be placed relative to one or more teeth, etc.; 7) Orthodontic settings; 8) 3D representations of hardware elements (such as facebows, lingual bows, orthodontic attachments, buttons, hooks, occlusal ramps, etc.) placed relative to one or more teeth, etc.; 8) 3D representations of bonding pads for hardware elements (which can be generated for specific teeth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting the tooth through a Boolean operation); 9) 3D representations of clear tray appliances (CTAs); 10) The position or shape of CTA trim lines (e.g., described as a mesh or polyline); 11) Arch forms (e.g., described as 3D polylines or 3D meshes or surfaces) that describe the contour or layout of the dental arch, which can follow the incisal edges of one or more teeth, which can follow the facial surfaces of one or more teeth, which in some specific implementations can correspond to malocclusion dental arches and in other specific implementations correspond to final setting dental arches (the effect of malocclusion on the shape of the arch form can be reduced by smoothing or averaging the shape of the arch form), and which can be described by one or more control points and / or splines; 12) 3D representations of jig models (e.g., depictions of teeth and gums used in thermoformed clear tray appliances, or depictions of teeth / gums / hardware used in thermoformed indirect bonding trays); 13) One or more latent space vectors (or latent capsules) generated by the 3D encoder stage of a 3D autoencoder (e.g., a variational autoencoder trained for tooth reconstruction) that has been trained on the reconstruction of an oral care mesh; 14) One or more oral care metrics for one or more teeth (e.g., such as orthodontic metrics or prosthetic design generation metrics); 15) One or more landmarks (e.g., 3D points) that describe the shape and / or geometric properties of one or more teeth, other dentition structures, or hardware structures (e.g., to be used for orthodontic setting creation or prosthetic appliance component generation or placement); 16) 3D representations created by scanning (e.g., optical scanning, CT scanning, or MRI scanning) 3D printed parts (such as scanned jig models) corresponding to one or more teeth / gums / hardware / appliances; 17) 3D printed appliances (optionally including local thickness, reinforcement rib geometry, flap positioning, etc.); 18) 3D representations of a patient's dentition captured by a clinician or healthcare practitioner at the chairside (e.g., in an environment where the 3D representation is verified at the chairside before the patient leaves the clinic, such that errors can be detected and rescanning can be performed as needed); 19) Prosthetic tooth designs (e.g., for veneers, crowns, bridges, or prosthetic appliances); 20) 3D representations of one or more teeth used in digital oral care processes; 21) Other 3D printed parts belonging to oral care procedures or other fields; 22) IPR cutting surfaces;23) One or more orthodontic setting transformations associated with one or more IPR cutting surfaces; 24) (Digital) pontic design that can fill at least a portion of the space between teeth to create space for erupting teeth in an orthodontic setting and then emerge from the gums; or 25) Components of a jig model (e.g., including jig model components such as interdental bands, occlusal recesses, bite locks, bite ramps, interdental reinforcements, gingival ridges, torque points, power ridges, pontics, or pits, etc.).;
[0036] In some cases, a shape-based input of a tooth can be provided to a neural network for setting prediction. In other cases, a non-shape-based input, such as tooth name or nomenclature, can be used as it relates to dental notation. In some specific implementations, a vector R of flags can be provided to the neural network, where a "1" value indicates the presence of a tooth and a "0" value indicates the absence of a tooth in a patient case (although other values are possible). The vector R can include one-hot vectors, where each element in the vector corresponds to a tooth type, name, or nomenclature. Identification information about a tooth (e.g., the name of the tooth) can be provided to the prediction neural network of the present disclosure, which has the advantage of enabling the neural network to be trained to handle different teeth in a tooth-specific manner. For example, a setting prediction model can learn to predict setting transformations for a specific tooth name (e.g., the upper right central incisor or the lower left canine, etc.). In the case of a mesh autoencoder (for labeling mesh elements or for filling in missing mesh data), the autoencoder can be trained in this way to provide specialized processing to a tooth based on the tooth's nomenclature. In the case of a setting classification neural network, a list of tooth names present in a patient's dental arch can better enable the neural network to output an accurate determination of setting classification because tooth nomenclature is a valuable input for training such a neural network. For example, tooth nomenclature / names can be defined according to a universal numbering system, the Palmer quadrant system, or the FDI World Dental Federation notation (ISO 3950).
[0037] In one example, in the case where all teeth except (at most four) wisdom teeth are present, the vector R can be defined as an optional input to the setting prediction neural network of the present disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, UL1, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7.
[0038] In some cases, the position of the cusp tip can be provided to the neural network for setting predictions. In other cases, one or more vectors S of the orthodontic metrics described elsewhere in this disclosure can be provided to the neural network for setting predictions. The advantage is that the network's ability to be trained to understand the state of the malocclusion setting is improved, and thus it can predict a more accurate final setting or intermediate stage.
[0039] In some specific implementations, the neural network can take as input one or more indications of interproximal reduction (IPR) U, which can indicate the amount of enamel to be removed from a tooth (from mesial or from distal) during the course of orthodontic treatment. In some specific implementations, the IPR information (e.g., the amount of IPR to be performed on one or more teeth, measured in millimeters, or one or more binary flags indicating whether IPR is to be performed on each tooth identified by a marker) can be concatenated with the latent vector A generated by a VAE or a latent capsule T autoencoder. The vector and / or capsule resulting from such concatenation can be provided to one or more of the neural networks of this disclosure, which has the technical improvement or additional advantage of enabling the prediction neural network to consider IPR. IPR is particularly relevant to a setting prediction method that can determine the position and pose of teeth at the end of treatment or during one or more stages during treatment. It is important to consider the amount of enamel to be removed before the predicted tooth movement.
[0040] In some specific implementations, one or more protocol parameters K and / or doctor preference vectors L can be introduced into the setting prediction model. In some specific implementations, one or more optional vectors or values include: tooth position N (e.g., XYZ coordinates in local or global coordinates of the tooth), tooth orientation O (e.g., pose, such as in a transformation matrix or quaternion, Euler angles or other forms described herein), tooth size P (e.g., length, width, height, perimeter, radius, diagonal measurement, volume, and any dimension can be normalized relative to one or more other teeth), distance Q between adjacent teeth. In some cases, these "tooth sizes P" can be used to describe the expected size of a tooth for dental restoration design generation.
[0041] In some specific implementations, the tooth size P (e.g., such as length, width, height, or perimeter) can be measured in a plane, such as a plane intersecting the centroid of the tooth, or a plane intersecting a central point located at the midpoint between the centroid of the tooth and the most incisal extent or the most gingival extent. The tooth height dimension can be measured as the distance from the gingiva to the incisal edge. The tooth width dimension can be measured as the distance from the mesial extent to the distal extent of the tooth. In some specific implementations, the roundness or circularity of the tooth cross-section can be measured and included in the vector P. The roundness or circularity can be defined as the ratio of the radii of the inscribed circle and the circumscribed circle.
[0042] The distance Q between adjacent teeth can be implemented in different ways (and calculated using different distance definitions, such as Euclidean or geodesic). In some specific implementations, the distance Q1 can be measured as the average distance between the mesh elements of two adjacent teeth. In some specific implementations, the distance Q2 can be measured as the distance between the centers or centroids of two adjacent teeth. In some specific implementations, the distance Q3 can be measured between the closest mesh elements between two adjacent teeth. In some specific implementations, the distance Q4 can be measured between the tooth tips of two adjacent teeth. In some specific implementations, teeth can be considered adjacent within an arch. In some specific implementations, teeth can also be considered adjacent between opposing arches. In some specific implementations, any one of Q1, Q2, Q3, and Q4 can be divided by a term to normalize the resulting value of Q. In some specific implementations, the normalization term can involve one or more of the following: the volume of the tooth, the count of mesh elements in the tooth, the surface area of the tooth, the cross-sectional area of the tooth (e.g., as projected onto the XY plane), or some other term related to the tooth size.
[0043] Other information regarding the patient's dentition or treatment needs (or related parameters) can be concatenated with other input vectors to one or more of an MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and / or any neural network model listed elsewhere in this disclosure.
[0044] The vector M may include markers applied to one or more teeth. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is pinned. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is fixed. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is a pontic. Other and additional markers are possible for teeth, such as combinations of fixed, pinned, and pontic markers. A marker set to a value indicating that a tooth should be fixed is a signal that the tooth should not move during processing and is sent to the network. In some embodiments, the neural network loss function can be designed to penalize any movement in the indicated teeth (and in some cases, may be severely penalized). A marker indicating that a tooth is a pontic notifies the network to maintain the diastema, although movement of the space is allowed. In some cases, M may include a marker indicating tooth loss. In some embodiments, the presence of one or more fixed teeth in the dental arch can assist in setting the prediction because the one or more fixed teeth can provide an anchor for the posture of other teeth in the dental arch (i.e., can provide a fixed reference for the posture transformation of one or more other teeth in the dental arch). In some embodiments, one or more teeth can be intentionally fixed in order to provide an anchor to which other teeth can be positioned. In some embodiments, a 3D representation (such as a mesh) corresponding to the gingiva can be introduced to provide a reference point relative to which the teeth can move.
[0045] Without loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U, and V described elsewhere in the present disclosure can also be introduced into the input of one or more of the prediction models of the present disclosure or into an intermediate layer thereof. Specifically, these optional vectors can be introduced into an MLP setting, a GDL setting, an RL setting, a VAE setting, a capsule setting, and / or a diffusion setting, with the advantage of enabling the corresponding model to output settings that better meet the orthodontic treatment needs of the patient. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent vectors A that are also provided to one or more of the prediction models of the present disclosure. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent capsules T that are also provided to one or more of the prediction models of the present disclosure.
[0046] In some embodiments, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into a neural network (such as an MLP or a transformer) in the hidden layer of the network. In some cases, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into the internal processing of an encoder structure.
[0047] In some specific implementations, a prediction model (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, PT settings, similarity settings, and diffusion settings) can take as input one or more latent vectors A corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a prediction model (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, and diffusion settings) can take as input one or more latent capsules T corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a prediction method can take both A and T as input.
[0048] A variety of loss calculation techniques generally apply to the techniques of the present disclosure (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, setting classification, tooth classification, VAE grid element labeling, MAE grid filling, and estimation of protocol parameters).
[0049] These losses include L1 loss, L2 loss, mean squared error (MSE) loss, cross-entropy loss, etc. Losses can be calculated and used to train neural networks such as multi-layer perceptrons (MLPs), U-Net architectures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer architectures, etc. For example, in the learning of sequences, some specific implementations can use triplet loss or contrastive loss.
[0050] Losses can also be used to train encoder architectures and decoder architectures. KL divergence loss can be at least partially used to train one or more neural networks of the present disclosure, such as a grid reconstruction autoencoder or a generator in GDL settings, which has the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior can enable the reconstruction autoencoder to produce better reconstructions (e.g., when modifying the latent vector representation and using the decoder to reconstruct the modified latent vector, the resulting reconstruction is more likely to be a valid instance of the input representation). There are other techniques for calculating losses that can be described elsewhere in the present disclosure. Such losses can be based on quantifying the difference between two or more 3D representations.
[0051] MSE loss calculation can involve the calculation of the average squared distance between two sets, vectors, or data sets. MSE can generally be minimized. MSE can be applied to regression problems where the predictions generated by a neural network or other machine learning model can be real numbers. In some specific implementations, a neural network can be equipped with one or more linear activation units on the output to generate MSE predictions. According to the techniques of the present disclosure, mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used.
[0052] In some specific implementations, cross - entropy can be used to quantify the difference between two or more distributions. In some specific implementations, cross - entropy loss can be used to train the neural networks of the present disclosure. In some specific implementations, cross - entropy loss can involve comparing predicted probabilities with ground - truth probabilities. Other names for cross - entropy loss include "log loss", "logistic loss", and "log loss". A small cross - entropy loss can indicate a better (e.g., more accurate) model. Cross - entropy loss can be logarithmic. In some specific implementations, cross - entropy loss can be applied to binary classification problems. In some specific implementations, a neural network can be equipped with sigmoid activation units at the output to generate probability predictions. In the case of multi - class classification, cross - entropy can also be used. In this case, in some specific implementations, a neural network trained to make multi - class predictions can be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for each class to be predicted). Other loss - calculation techniques that can be applied in the training of the neural networks of the present disclosure include one or more of the following: Huber loss, hinge loss, classification hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean - squared logarithmic error loss (MSLE). Other loss - calculation methods are described herein and can be applied to the training of any neural network described in the present disclosure.
[0053] In some specific implementations, one or more neural networks of the present disclosure can be trained, at least in part, by a loss based on at least one of the following: point - wise mesh Euclidean distance (PMD) and Earth Mover's Distance (EMD). Some specific implementations can incorporate Hausdorff distance (HD) calculation into the loss calculation. Calculating the Hausdorff distance between two or more 3D representations (such as 3D meshes) can provide one or more technical improvements because HD not only considers the distance between two meshes, but also the way those meshes are oriented and the relationship between the mesh shapes in those orientations (or positions or poses). Hausdorff distance can improve the comparison of two or more tooth meshes, such as two or more instances of tooth meshes in different poses (e.g., comparison of a predicted setting with a ground - truth setting, which can be performed during the process of calculating the loss value for training a setting - prediction neural network).
[0054] The reconstruction loss can compare the predicted output with the ground truth (or reference) output. The systems of the present disclosure can calculate the reconstruction loss as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_loss = 0.5 * L1(all_points_target, all_points_predicted) + 0.5 * MSE(all_points_target, all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the ground truth data (e.g., a ground truth tooth restoration design, or a ground truth example of some other 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the generated or predicted data (e.g., a generated tooth restoration design, or a generated example of some other type of 3D oral care representation). Other specific implementations of the reconstruction loss can additionally (or alternatively) involve an L2 loss, a mean absolute error (MAE) loss, or a Huber loss term.
[0055] The reconstruction error can compare the reconstructed output data (e.g., generated by a reconstruction autoencoder, such as a tooth design that has been generated for use in a dental restoration appliance) with the initial input data (e.g., the data input to the reconstruction autoencoder, such as a pre-restoration tooth). The systems of the present disclosure can calculate the reconstruction error as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_error = 0.5 * L1(all_points_input, all_points_reconstructed) + 0.5 * MSE(all_points_input, all_points_reconstructed). In the above example, all_points_input is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the input data (e.g., a pre-restoration tooth design input to the reconstruction autoencoder, or another 3D oral care representation input to an ML model). In the above example, all_points_reconstructed is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the reconstructed (or generated) data (e.g., a reconstructed tooth restoration design, or another example of a generated 3D oral care representation).
[0056] In other words, the reconstruction loss involves calculating the difference between the predicted output and the reference output, while the reconstruction error involves calculating the difference between the reconstructed output and the initial input from which the reconstructed data is derived.
[0057] The techniques of the present disclosure may include operations such as 3D convolution, 3D pooling, 3D deconvolution, and 3D upsampling. 3D convolution, for example, may assist in segmentation processing when downsampling a 3D mesh. 3D deconvolution, for example, performs an inverse operation of 3D convolution in a U-Net. 3D pooling may assist in segmentation processing, for example, in a generalized neural network feature map. 3D upsampling, for example, performs an inverse operation of 3D pooling in a U-Net. These operations may be implemented by one or more layers in a predictive or generative neural network as described herein. These operations may be directly applied to mesh elements such as mesh edges or mesh faces. These operations provide a technical improvement over other methods because these operations are invariant to mesh rotation, scaling, and translation changes. Generally speaking, these operations depend on edge (or face) connectivity, so as long as the edge (or face) connectivity is maintained, these operations are not affected by mesh changes in 3D space. That is, these operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position, or scale of the oral care mesh, which may improve data accuracy. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes and can be used for tasks such as 3D shape classification or mesh element labeling (e.g., for segmentation or mesh cleaning). MeshCNN performs these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.
[0058] In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 2D representations such as images. In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 3D representations such as meshes or point clouds. An intraoral scanner may capture 2D images of a patient's dentition from various angles. The intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data depicting the patient's dentition. According to various techniques, an autoencoder (or other neural network described herein) may be trained to operate on either or both of 2D and 3D representations.
[0059] A 2D autoencoder (including a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or latent capsule) using the 2D encoder and then reconstruct a facsimile of the input 2D image using the 2D decoder. For a handheld mobile application that has been developed for such analysis (e.g., for the analysis of dental anatomy), 2D images may be easily captured using one or more onboard cameras. In other examples, 2D images may be captured using an intraoral scanner configured for such a function. Operations that may be used in implementations of a 2D autoencoder (or other 2D neural network) for 2D image analysis are 2D convolution, 2D pooling, and 2D reconstruction error calculation.
[0060] 2D Convolution :
[0061] 2D image convolution can involve the "sliding" of a kernel across a 2D image and the calculation of element-wise multiplications, as well as summing these element-wise multiplications into output pixels. The output pixels generated from each new position of the kernel are saved into an output 2D feature matrix. In some specific implementations, adjacent elements (e.g., pixels) can be in well-defined positions in a straight grid (e.g., above, below, left, and right).
[0062] 2D Pooling :
[0063] A 2D pooling layer can be used to downsample a feature map and summarize the presence of certain features in that feature map.
[0064] A 2D reconstruction error can be computed between the pixels of an input image and a reconstructed image. The mapping between pixels can be well understood (e.g., directly comparing the upper pixels [23,134] of the input image with the pixels [23,134] of the reconstructed image, assuming the two images have the same dimensions).
[0065] One of the advantages provided by the 2D autoencoder-based techniques of the present disclosure is the ease of capturing 2D image data with a handheld device. In some cases where an external data source provides data for analysis, there can be instances where only 2D image data is available. When only 2D image data is available, it is necessary to use a 2D autoencoder for analysis.
[0066] Modern mobile devices (such as commercially available smartphones) can also have the ability to generate 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera that moves around an object to capture multiple images from different views, or both), and this 3D data can be arranged in a 3D representation, such as a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation in some specific implementations. In some cases, the analysis of a 3D representation of an object can provide a technical improvement over a 2D analysis of the same object. For example, a 3D representation can describe the geometry and / or structure of an object with less ambiguity than a 2D representation (which can include shadows and other artifacts that complicate the depiction of the depth and texture of the object). In some specific implementations, 3D processing can achieve a technical improvement due to the inverse optics problem, which affects 2D representations in some cases. The inverse optics problem refers to the phenomenon that in some cases, the size of an object, the orientation of the object, and the distance between the object and the imaging device may be merged in a 2D image of the object. Any given projection of an object on an imaging sensor can map to an infinite count of {size, orientation, distance} pairings. 3D representations achieve a technical improvement because 3D representations remove the ambiguity introduced by the inverse optics problem.
[0067] Devices configured for a dedicated purpose with 3D scanning, such as 3D intraoral scanners (or CT scanners or MRI scanners), can generate 3D representations of an object (e.g., a patient's dentition) that have significantly higher fidelity and precision than what a handheld device might have. When such high-fidelity 3D data is available (e.g., in oral care mesh classification or in the application of other 3D techniques described herein), the use of a 3D autoencoder provides technical improvements (such as increased data precision) to extract the best possible signal from those 3D data (i.e., obtain a signal from the 3D crown mesh used in tooth classification or setup classification).
[0068] A 3D autoencoder (including a 3D encoder and a 3D decoder) can be trained on 3D data to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder and then reconstruct a replica of the input 3D representation using the 3D decoder. Operations that can be used to implement a 3D autoencoder for analyzing 3D representations (e.g., 3D meshes or 3D point clouds) are 3D convolution, 3D pooling, and 3D reconstruction error calculation.
[0069] For each mesh element, 3D convolution can be performed to aggregate local features from nearby mesh elements. Processing can be performed on top of and in addition to techniques used for 2D convolution to account for the different counts and positions of adjacent mesh elements (relative to a particular mesh element). A particular 3D mesh element can have a variable neighbor count, and those neighbors can be absent from expected positions (unlike pixels in 2D convolution, which can have a fixed adjacent pixel count present in known or expected positions). In some cases, the order of adjacent mesh elements can be relevant to 3D convolution.
[0070] The 3D pooling operation can enable the combination of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling can iteratively reduce a 3D mesh to the mesh elements that are most highly relevant to a given application (e.g., for which a neural network has been trained). Similar to 3D convolution, 3D pooling can benefit from special processing in addition to the processing required in 2D convolution to account for the different counts and positions of adjacent mesh elements (relative to a particular mesh element). In some cases, the order of adjacent mesh elements may be less relevant to 3D pooling than to 3D convolution.
[0071] The 3D reconstruction error can be calculated using one or more of the techniques described herein, such as calculating the Euclidean distance between corresponding mesh elements, between two meshes. According to aspects of the present disclosure, other techniques are possible. The 3D reconstruction error can generally be calculated on 3D mesh elements rather than 2D pixels of the 2D reconstruction error. The 3D reconstruction error can achieve a technical improvement over the 2D reconstruction error because, in some cases, the 3D representation can have less ambiguity (i.e., less ambiguity in form, shape, and / or structure) than the 2D representation. In some specific implementations, due to the complexity of the mapping between the input mesh elements and the reconstructed mesh elements (i.e., the input mesh and the reconstructed mesh may have different mesh element counts, and there may be a less clear mapping between mesh elements compared to the mapping between pixels in 2D reconstruction), additional processing may be required for 3D reconstruction over and above 2D reconstruction. Technical improvements in 3D reconstruction error calculation include increased data accuracy.
[0072] A 3D scanner, such as an intraoral scanner, a computed tomography (CT) scanner, an ultrasound scanner, a magnetic resonance imaging (MRI) machine, or a mobile device capable of performing photogrammetry, can be used to generate a 3D representation. The 3D representation can describe the shape and / or structure of an object. The 3D representation can include one or more of a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation, etc. A 3D mesh includes edges, vertices, or faces. Although in some cases these three types of data are related to each other, they are different. A vertex is a point in 3D space that defines the boundary of the mesh. These points would alternatively be described as a point cloud, without additional information about how the points are connected to each other (as described by edges). An edge is described by two points and can also be referred to as a line segment. A face is described by multiple edges and vertices. For example, in the case of a triangular mesh, a face includes three vertices that are interconnected to form three consecutive edges. Some meshes may include degenerate elements, such as non-manifold mesh elements, which can be removed to benefit subsequent processing. According to aspects of the present disclosure, other mesh preprocessing operations are also possible. 3D meshes are typically formed using triangles, but in other embodiments quadrilaterals, pentagons, or some other n-sided polygon can be used. In some embodiments, such as in the case of performing sparse processing, a 3D mesh can be converted into one or more voxelized geometries (i.e., including voxels). The techniques of the present disclosure operating on a 3D mesh can receive one or more tooth meshes (e.g., arranged in one or more dental arches) as input. Each of these meshes can be preprocessed before being input into a prediction architecture (e.g., including at least one of an encoder, a decoder, a pyramid encoder-decoder, and a U-Net). Such preprocessing can include converting the mesh into a list of mesh elements such as vertices, edges, faces, or into voxels in the case of sparse processing. For one or more selected types of mesh elements (e.g., vertices), a feature vector can be generated. In some examples, a feature vector is generated for each vertex of the mesh. Each feature vector can include a combination of spatial features and / or structural features, as specified in the following table:
[0073] Table 1 discloses non-limiting examples of mesh element features. In some embodiments, in addition to the spatial or structural mesh element features described in Table 1, color (or other visual cues / identifiers) may also be considered mesh element features. As used herein (e.g., in Table 1), a point differs from a vertex in that a point is part of a 3D point cloud, while a vertex is part of a 3D mesh and may have incident faces or edges. A dihedral angle (which may be expressed in radians or degrees) can be calculated as the angle (e.g., a signed angle) between two connected faces (e.g., two faces connected along an edge). The sign on the dihedral angle can reveal information about the convexity or concavity of the mesh surface. For example, in some embodiments, a positively signed angle may indicate a convex surface. Additionally, in some embodiments, a negatively signed angle may indicate a concave surface. To calculate the principal curvatures of a mesh vertex, the directional curvatures of each adjacent vertex around that vertex can be calculated first. These directional curvatures can be sorted in a circular order (e.g., 0 degrees, 49 degrees, 127 degrees, 210 degrees, 305 degrees) near the vertex normal vector and may include a subsampled form of the full curvature tensor. Circular order means sorting by angle around an axis. The sorted directional curvatures can contribute to a system of linear equations that admits a closed-form solution, which can estimate the two principal curvatures and directions, which can characterize the full curvature tensor. Consistent with Table 1, a voxel may also have features calculated as an aggregation of other mesh elements (e.g., vertices, edges, and faces) that either intersect the voxel or, in some embodiments, are primarily or entirely contained within the voxel. Rotating a mesh may not change the structural features but may change the spatial features. And, as described elsewhere in this disclosure, the term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non-limiting sense. In some embodiments, in addition to mesh element features, there are alternative ways to describe the geometry of a mesh (such as 3D key points and 3D descriptors). Examples of such 3D key points and 3D descriptors can be found in "TONIONI A et al., "Learning to detect good 3D keypoints.", Int J Comput.Vis. Vol. 126, pp. 1-20, 2018". In some embodiments, 3D key points and 3D descriptors can describe the extrema (minima or maxima) of the surface of a 3D representation.In some embodiments, one or more mesh element features may be computed at least in part via deep feature synthesis (DFS), such as described in: J.M. Kanter and K. Veeramachaneni, “Deep feature synthesis: Towards automating data science endeavors”, 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp. 1-10, doi: 10.1109 / DSAA.2015.7344858.
[0074] Neural networks that generate representations based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional and / or pooling layers, or other models may benefit from the use of mesh element features. Mesh element features may convey aspects of the surface shape and / or structure of a 3D representation to the neural network models of the present disclosure. Each mesh element feature describes different information about the 3D representation that may not redundantly exist in other input data provided to the neural network. For example, vertex curvature can quantify aspects of the concavity or convexity of the surface of a 3D representation that the network would not otherwise understand. In other words, mesh element features may provide a processed form of the structure and / or shape of a 3D representation; data that would otherwise not be available to the neural network. This processed information is generally more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, mesh element features have been provided to a neural network that generates representations based on a U-Net model and also to a representation generation model based on a variational autoencoder with continuous normalizing flows. Based on the experiments, it was found that systems using a full complement of mesh element features (e.g., “XYZ” coordinate tuples, “normal vectors”, “vertex curvature”, point pivots, and normal pivots) were at least 3% more accurate than systems that did not use mesh element features. A point pivot describes an “XYZ” coordinate tuple with a local coordinate system (e.g., at the centroid of the corresponding tooth). A normal pivot describes a “normal vector” with a local coordinate system (e.g., at the centroid of the corresponding tooth). Additionally, when using a full complement of mesh element features, training converges more quickly. In other words, machine learning models trained using a full complement of mesh element features tend to be faster and more accurate (at an earlier epoch) than systems that do not. For an existing system that observes a historical accuracy of 91%, a 3% increase in accuracy reduces the actual error rate by more than 30%.
[0075] Prediction models that can operate on the feature vectors of the above features include, but are not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoders, validation using autoencoders, grid segmentation, coordinate system prediction, grid cleaning, repair design generation, appliance component generation and / or placement, and dental arch form prediction. Such feature vectors can be presented as inputs to the prediction models. In some specific implementations, such feature vectors can be presented to one or more internal layers of a neural network that is part of one or more of those prediction models.
[0076] The neural networks of the present disclosure can leverage one or more benefits of parameter tuning operations to optimize the inputs and parameters of the neural networks to produce more data-precise results. One parameter that can be tuned is the neural network learning rate (e.g., it can have values such as 0.1, 0.01, 0.001, etc.). Data augmentation schemes can also be tuned or optimized, such as a scheme of adding "shiver" to the dental mesh before inputting it to the neural network (i.e., small random rotations, translations, and / or scalings can be applied to change the dataset and make the neural network robust to changes in the data).
[0077] A subset of the neural network model parameters that can be used for tuning is as follows: ○ Learning rate (LR) decay rate (e.g., how much the LR decays during a training run) ○ Learning rate (LR). A floating-point value used by the optimizer (e.g., 0.001). ○ LR scheduling (e.g., cosine annealing, step, exponential) ○ Voxel size (for the case of sparse grid processing operations) ○ Dropout % (e.g., dropout that can be performed in a linear encoder) ○ LR decay step size (e.g., decay every 10 or 20 or 30 epochs) ○ Model scaling, which can increase or decrease the layer count and / or the parameter count per layer.
[0078] Parameter tuning can be advantageously applied to the training of neural networks to predict final settings or intermediate gradings, thereby providing technical improvements towards data accuracy. Parameter tuning can also be advantageously applied to the training of neural networks for mesh element labeling or for mesh filling. In some examples, parameter tuning can be advantageously applied to the training of neural networks for tooth reconstruction. In terms of the classifier model of the present disclosure, parameter tuning can be advantageously applied to neural networks for the classification of one or more settings (i.e., the classification of one or more arrangements of teeth). The advantage of parameter tuning is to improve the data accuracy of the output of the prediction model or classification model. In some cases, parameter tuning can provide the advantage of obtaining the last remaining few percentage points of validation accuracy from the prediction or classification model.
[0079] Various neural network models of the present disclosure can benefit from data augmentation. Examples include models trained on 3D meshes, such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, FDG settings, setting classification, setting comparison, VAE mesh element labeling, MAE mesh filling, mesh reconstruction VAE, and validation using autoencoders. Figure 1 is a flowchart illustrating the data augmentation method of the present disclosure. Such as by Figure 1 Data augmentation of the method shown can increase the size of the training dataset of the dental arch. Data augmentation can provide additional training examples by adding random rotations, translations, and / or rescaling to copies of the existing dental arch. In some specific implementations of the techniques of the present disclosure, data augmentation can be performed by perturbing or jittering the vertices of the mesh in a manner similar to that described in ("Equidistant and Uniform Data Augmentation for 3D Objects", IEEE Access, Digital Object Identifier 10.1109 / ACCESS.2021.3138162). The position of the vertices can be perturbed by adding Gaussian noise, for example with a zero mean and a standard deviation of 0.1. According to the techniques of the present disclosure, other mean and standard deviation values are possible.
[0080] Since the generator network of the present disclosure can be implemented as one or more neural networks, the generator may include activation functions. When executed, an activation function outputs a determination as to whether a neuron in the neural network will fire (e.g., send an output to the next layer). Some activation functions may include: a binary step function or a linear activation function. Other activation functions impart non-linear behavior to the neural network, including: sigmoid / logistic activation function, Tanh (hyperbolic tangent) function, rectified linear unit (ReLU), leaky ReLU function, parametric ReLU function, exponential linear unit (ELU), softmax function, swish function, Gaussian error linear unit (GELU), or scaled exponential linear unit (SELU). The linear activation function may be well-suited for some regression applications (and other applications) in the output layer. In the output layer, the sigmoid / logistic activation function may be well-suited for certain binary classification applications (and other applications). The sigmoid activation function may be well-suited for some multi-class classification applications (and other applications) in the output layer. In the output layer, the sigmoid activation function may be well-suited for some multi-label classification applications (and other applications). The ReLU activation function may be well-suited for some convolutional neural network (CNN) applications (and other applications) in the hidden layer. The Tanh and / or sigmoid activation functions may be well-suited for some recurrent neural network (RNN) applications (and other applications) in, for example, the hidden layer. There are a variety of optimization algorithms that can be used to train the neural networks of the present disclosure (such as updating neural network weights), including gradient descent (which uses first-order derivatives to determine the training gradient and is commonly used in the training of neural networks), Newton's method (which may use second-order derivatives in loss calculations to find a better training direction than gradient descent but may require calculations involving the Hessian matrix), and conjugate gradient method (which may converge faster than gradient descent but does not require the Hessian matrix calculations that Newton's method may require). In some specific implementations, in addition to or instead of the above techniques, additional methods may be employed to update the weights. These additional methods include the Levenberg-Marquardt method and / or simulated annealing. The backpropagation algorithm is used to convey the results of the loss calculation back into the network so that the network weights can be adjusted for learning.
[0081] Neural networks contribute to the implementation of the functions of the applications of the present disclosure, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoder, verification using autoencoder, estimation of oral care parameters, 3D grid segmentation (3D representation segmentation), coordinate system prediction, grid cleaning, restoration design generation, appliance component generation and / or placement or dental arch form prediction. The neural networks of the present disclosure can embody parts or all of various different neural network models. Examples include U-Net architecture, multi-layer perceptron (MLP), transformer, pyramid architecture, recurrent neural network (RNN), autoencoder, variational autoencoder, regularized autoencoder, conditional autoencoder, capsule network, capsule autoencoder, stacked capsule autoencoder, denoising autoencoder, sparse autoencoder, conditional autoencoder, long / short-term memory (LSTM), gated recurrent unit (GRU), deep belief network (DBN), deep convolutional network (DCN), deep convolutional inverse graphics network (DCIGN), liquid state machine (LSM), extreme learning machine (ELM), echo state network (ESN), deep residual network (DRN), Kohonen network (KN), neural Turing machine (NTM) or generative adversarial network (GAN). In some specific implementations, an encoder structure or a decoder structure can be used. Each of these models offers one or more of its own specific advantages. For example, a specific neural network architecture may be particularly suitable for a specific ML technique. For example, autoencoders are particularly suitable for the classification of 3D oral care representations due to their ability to transform 3D oral care representations into a form that is easier to classify.
[0082] In some specific implementations, the neural networks of the present disclosure may be suitable for operating on 3D point cloud data (alternatively, on 3D meshes or 3D voxelized representations). Many neural network specific implementations can be applied to the processing of 3D representations and can be applied to training prediction and / or generation models for oral care applications, including: PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution, and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN, and DSG-Net. Oral care applications include but are not limited to: setup prediction (e.g., using VAEs, RLs, MLPs, GDLs, capsules, diffusion, etc. trained for setup prediction), 3D representation segmentation, 3D representation coordinate system prediction, element tagging for 3D representation cleaning (VAE for mesh element tagging), filling of missing elements in 3D representations (MAE for mesh filling), dental restoration design generation, setup classification, appliance component generation and / or placement, dental arch form prediction, estimation of oral care parameters, setup verification or other verification applications, and 3D representation classification of teeth.
[0083] Some specific implementations of the techniques of the present disclosure incorporate the use of autoencoders. Autoencoders that can be used in accordance with aspects of the present disclosure include but are not limited to: AtlasNet, FoldingNet, and 3D-PointCapsNet. Some autoencoders can be implemented based on PointNet.
[0084] Representation learning can be applied to the setup prediction techniques of the present disclosure by training a neural network to learn a representation of a tooth and then using another neural network to generate a transformation of the tooth. Some specific implementations can use a VAE or a capsule autoencoder to generate a representation of the reconstructed features of one or more meshes relevant to the oral care field (in some cases, including information about the structure of a tooth mesh). Then, this representation (latent vector or latent capsule) can be used as the input to a module that generates one or more transformations of one or more teeth. In some specific implementations, these transformations can place the tooth into a final setup pose. In some specific implementations, these transformations can place the tooth into an intermediate graded pose. In some specific implementations, the transformation can be described by a 9×1 transformation vector (e.g., specifying a translation vector and a quaternion). In other specific implementations, the transformation can be described by a transformation matrix (e.g., a 4×4 affine transformation matrix).
[0085] In some specific implementations, the system of the present disclosure can perform principal component analysis (PCA) on an oral care mesh and use the resulting principal components as at least part of a representation of the oral care mesh in subsequent machine learning and / or other predictive or generative processing.
[0086] An autoencoder can be trained to generate a latent form of a 3D oral care representation. The autoencoder can include a 3D encoder (which encodes the 3D oral care representation into a latent form) and / or a 3D decoder (which reconstructs the latent form into a copy of the input 3D oral care representation). Although the present disclosure refers to a 3D encoder and a 3D decoder, the term 3D should be interpreted in a non-limiting manner to cover multi-dimensional operation modes. For example, the system of the present disclosure can train a multi-dimensional encoder and / or a multi-dimensional decoder.
[0087] The system of the present disclosure can implement end-to-end training. Some end-to-end training-based techniques of the present disclosure can involve two or more neural networks, where the two or more neural networks are trained together (i.e., the weights are updated simultaneously during the processing of each batch of input oral care data). In some specific implementations, end-to-end training can be applied to pose prediction by simultaneously training a neural network that learns a representation of teeth and a neural network that can generate tooth transformations.
[0088] According to some transfer learning-based specific implementations of the present disclosure, a neural network (e.g., a U-Net) can be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task can be executed to provide one or more initial neural network weights for training another neural network, which is trained to perform a second task (e.g., pose prediction). The first network can learn low-level neural network features of the oral care mesh and is shown to perform well in the first task. By using the first network as a starting point for training, the second network can exhibit faster training and / or improved performance. Certain layers can be trained to encode the neural network features of the oral care mesh in the training dataset. These layers can then be fixed (or undergo minor changes during the training process) and combined with other neural network components (such as additional layers), which are trained for one or more oral care tasks (such as pose prediction). In this way, a part of the neural network for one or more techniques of the present disclosure (e.g., pose prediction) can receive initial training for another task, which can result in important learning in the trained network layers. Then, this encoded learning can be built upon by further task-specific training of another network.
[0089] According to the present disclosure, transfer learning can be used for setting predictions and for other oral care applications such as mesh classification (e.g., tooth or setting classification), mesh element labeling, mesh element filling, protocol parameter estimation, mesh segmentation, coordinate system prediction, prosthetic design generation, mesh validation (for any application disclosed herein). In some specific implementations, a neural network trained to output predictions based on an oral care mesh can be partially trained first on one of the following publicly available datasets before being further trained on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, Thingi10K dataset (which is particularly relevant for 3D printed component validation), ABC: A large CAD model dataset for geometric deep learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Component Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.
[0090] In some specific implementations, a neural network previously trained on a first dataset (oral care data or other data) can subsequently receive further training on oral care data and be applied to oral care applications such as setting predictions. Transfer learning can be used to further train any one of the following networks: GCN (Graph Convolutional Network), PointNet, ResNet, or any other neural network from the published literature listed above.
[0091] In some specific implementations, a first neural network can be trained to predict the coordinate system of teeth (such as by using the techniques described in WO2022123402A1 or U.S. Provisional Application No. US63 / 366492). According to any one of the setting prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein), a second neural network can be trained for setting predictions. Transfer learning can transfer at least a portion of the knowledge or capabilities of the first neural network to the second neural network. Thus, transfer learning can provide an accelerated training phase for the second neural network to reach convergence. In some specific implementations, the training of the second network can be completed after being enhanced with transfer learning and then using one or more techniques of the present disclosure.
[0092] The system of the present disclosure can utilize representation learning to train an ML model. Advantages of representation learning include that, as opposed to receiving inputs with variable sizes or structures, the generative network (e.g., the neural network used for predicting transformations in a setting prediction) can be configured to receive inputs with known sizes and / or standard formats. Representation learning can yield performance superior to other techniques because noise in the input data can be reduced (e.g., because the representation generation model extracts hierarchical neural network features and / or the reconstruction characteristics of the input representation (e.g., mesh or point cloud) through loss calculation or the network architecture selected for that purpose).
[0093] The reconstruction characteristics can include values in a latent representation (e.g., a latent vector) that describe aspects of the shape and / or structure of the 3D representation provided to the representation generation module that generated the latent representation. For example, the weights of the encoder module of a reconstruction autoencoder can be trained to encode a 3D representation (e.g., a 3D mesh or others described herein) into a latent vector representation (e.g., a latent vector). In other words, the ability to encode a large set of mesh elements (e.g., hundreds, thousands, or millions) into a latent vector (e.g., hundreds or thousands of real values, e.g., 512, 1024, etc.) can be learned through the weights of the encoder. Each dimension of the latent vector can include a real number that describes some aspect of the shape and / or structure of the initial 3D representation. The weights of the decoder module of the reconstruction autoencoder can be trained to reconstruct the latent vector into a close replica of the initial 3D representation. In other words, the decoder can learn the ability to interpret the dimensions of the latent vector and decode the values within those dimensions. Generally speaking, the encoder and decoder neural network modules are trained to perform a mapping of the 3D representation to a latent vector, and then the latent vector can be mapped back (or otherwise reconstructed) to a 3D representation that is substantially similar to the initial 3D representation for which the latent vector was generated.
[0094] Returning to loss calculation, examples of loss calculation can include KL divergence loss, reconstruction loss, or other losses disclosed herein. Representation learning can reduce the size of the dataset required to train a model because the representation model learns a representation such that the generative network can focus on learning the generation task. Since meaningful neural network features of the input data (e.g., local and / or global features) are available to the generative network, the result can be improved model generalization. In other words, the first network can learn a representation and the second network can make prediction decisions. By training two networks to perform their own separate tasks, each network can generate more accurate results for its corresponding task than a single network trained to both learn a representation and make decisions. In some cases, transfer learning can first train a representation generation model. Then that representation generation model (either wholly or in part) can be used to pre-train subsequent models, such as generative models (e.g., generative transformation prediction). The representation generation model can benefit from using grid element features as input to improve the ability of the second ML module to encode the structure and / or shape of the input 3D oral care representation in the training dataset.
[0095] One or more neural network models of the present disclosure can have attention gates integrated therein. Attention gate integration provides an enhancement that enables the associated neural network architecture to focus resources on one or more input values. In some specific implementations, the attention gate can be integrated with the U-Net architecture, which has the advantage of enabling the U-Net to focus on certain inputs, such as input landmarks corresponding to teeth that are intended to be fixed (e.g., to prevent movement) during an orthodontic procedure (or in cases where other special handling is required). According to aspects of the present disclosure, the attention gate can also be integrated with an encoder or with an autoencoder (such as a VAE or a capsule autoencoder) to improve prediction accuracy. For example, the attention gate can be used to configure a machine learning model to give higher weights to aspects of the data that are more likely to be relevant to the correctly generated output. Thus, and because the machine learning models configured with these attention gates (or mechanisms) utilize aspects of the data that are more likely to be relevant to the correctly generated output, the final prediction accuracy of those machine learning models is improved.
[0096] The quality and composition of the training dataset for a neural network can affect the performance of the neural network during its execution phase. Dataset screening and outlier removal can be advantageously applied to the training of neural networks for various techniques of the present disclosure (e.g., for predictions for final settings or intermediate gradings, for neural networks for grid element labeling or for grid filling, for tooth reconstruction, for 3D grid classification, etc.) because dataset screening and outlier removal can remove noise from the dataset. Although the mechanisms for achieving the improvement are different from using attention gates, the end result is that the method allows the machine learning model to focus on the relevant aspects of the dataset and can lead to an improvement in accuracy similar to the improvement achieved with attention gates.
[0097] In the case of a neural network configured to predict a final setting, a patient case may include at least one of a set of segmented tooth meshes of the patient, the malalignment transformation of each tooth, and / or the ground truth setting transformation of each tooth. In the case of a neural network predicting a set of intermediate stage settings, a patient case may include at least one of a set of segmented tooth meshes of the patient, the malalignment transformation of each tooth, and / or a set of ground truth intermediate stage transformations of each tooth. In some embodiments, the training dataset may exclude patient cases in the contact passive phase (i.e., the phase where the teeth of the dental arch do not move). In some embodiments, the dataset may exclude cases where there is a passive phase at the end of the process. In some embodiments, the dataset may exclude cases where there is overcrowding at the end of the process (i.e., cases where an oral care provider such as an orthodontist or dentist has selected a final setting where the tooth meshes overlap to some extent). In some embodiments, the dataset may exclude cases of a particular difficulty level (or levels) (e.g., easy, medium, and difficult).
[0098] In some embodiments, the dataset may include cases with zero pinned teeth (or may include cases with at least one pinned tooth). A person skilled in the art may specify the pinned teeth when designing the process to prevent various tools from moving that particular tooth. In some embodiments, the dataset may exclude cases with no fixed teeth (conversely, where at least one tooth is fixed). Fixed teeth may be defined as teeth that should not move during the process. In some embodiments, the dataset may exclude cases with no pontic teeth (conversely, cases where at least one tooth is a pontic). Pontic teeth may be described as "ghost" teeth that are represented in the digital model of the dental arch but do not actually exist in the patient's dentition, or where there may be small teeth or partial teeth that may benefit from future work such as adding composite materials through prosthodontic appliances. The advantage of including pontic teeth in a patient's case is to leave space in the dental arch as part of the plan for the movement of other teeth during orthodontic treatment. In some cases, pontic teeth may save space in the patient's dentition for future dental or orthodontic work such as installing implants or crowns, or applying prosthodontic appliances such as adding composite materials to existing teeth that are too small or have an undesirable shape.
[0099] In some specific implementations, the dataset may exclude cases where the patient does not meet the age requirement (e.g., less than 12 years old). In some specific implementations, the dataset may exclude cases where the interproximal reduction (IPR) exceeds a certain threshold amount (e.g., greater than 1.0 mm). The dataset used to train the neural network to predict the settings of a clear tray appliance (CTA) may exclude patient cases unrelated to CTA processing. The dataset used to train the neural network to predict the settings of an indirectly bonded tray product may exclude cases unrelated to indirectly bonded tray processing. In some specific implementations, the dataset may exclude cases that only treat certain teeth. In such specific implementations, the dataset may include only cases where at least one of the following has been treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and / or canines.
[0100] The techniques of the present disclosure can be advantageously combined. For example, a setting comparison tool can be used to compare the output of the GDL setting model with the ground truth data, compare the output of the RL setting model with the ground truth data, compare the output of the VAE setting model with the ground truth data, and compare the output of the MLP setting model with the ground truth data. By comparing each of these setting prediction models with the ground truth data, it can be determined which model achieves the best performance on a certain dataset or within a given problem domain. Additionally, a metric visualization tool can enable a global view of the final settings and intermediate stages generated by one or more of the setting prediction models, with the advantage of being able to select the best setting prediction model. Moreover, the metric visualization tool enables the calculation of metrics with a global scope within a set of intermediate stages. In some specific implementations, these global metrics can be consumed as inputs to a neural network for predicting settings (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, etc.). The global metrics can also be provided to the FDG settings. In some specific implementations, local metrics from the present disclosure (i.e., local metrics are metrics that can be calculated for one stage or setting of the processing rather than within several stages or settings) can be consumed by the neural networks herein for predicting settings, with the advantage of improving the prediction results. In some specific implementations, the metrics described in the present disclosure can be visualized using the metric visualization tool.
[0101] VAE and MAE models for mesh element tagging and mesh filling can be advantageously combined with a setup prediction neural network for mesh cleaning before or during the prediction process. In some specific implementations, the VAE for mesh element tagging can be used to flag mesh elements for further processing, such as metric calculation, removal, or modification. In some cases, such tagged mesh elements can be provided as input to the setup prediction neural network to inform the neural network of important mesh features, attributes, or geometries, with the advantage of improving the performance of the resulting setup prediction model. In some specific implementations, mesh filling can make the geometry of the teeth closer to complete, enabling the setup prediction model to function better (i.e., improving the correctness of the prediction due to the better-formed geometry). In some cases, a neural network for classifying setups (i.e., a setup classifier) can assist the setup prediction neural network in functioning because the setup classifier tells the setup prediction neural network when a predicted setup is acceptable for use and can be provided to a method for generating an orthodontic tray. Setup classifiers (e.g., GDL setups, RL setups, VAE setups, capsule setups, MLP setups, diffusion setups, PT setups, similarity setups, and FDG setups, etc.) can help generate the final setup and also help generate intermediate stages. Additionally, the setup classifier neural network can be combined with a metric visualization tool. In other specific implementations, the setup classification neural network can be combined with a setup comparison tool (e.g., the setup comparison tool can output an indication of how a setup generated, in part, by the setup classifier compares to a setup generated by another setup prediction method). In some specific implementations, the VAE for mesh element tagging can identify one or more mesh elements used in metric calculation. The resulting metric output can be visualized by a metric visualization tool.
[0102] In some examples, the setup classifier neural network can assist the setup prediction techniques described in U.S. Patent Application No. US20210259808A1, the entire content of which is incorporated herein by reference, or PCT Application Publication No. WO2021245480A1, the entire content of which is incorporated herein by reference, or the setup prediction techniques described in PCT Application No. PCT / IB2022 / 057373, the entire content of which is incorporated herein by reference. The setup classifier will help one or more of those techniques know when the predicted final setup is closest to being correct. In some cases, the setup classifier neural network can output an indication of how far a given setup is from the final setup (i.e., a progress indicator).
[0103] In some embodiments, the latent space embedding vectors from the reconstructed VAE can be cascaded with the inputs of the setup prediction neural network described in WO2021245480A1. The latent space vectors can also be combined as inputs into other setup prediction models: GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup, etc. The advantage is to endow the neural network with reconstruction characteristics (e.g., the latent vector dimension of the dental mesh), thereby improving the generated setup prediction.
[0104] In some examples, the various setup prediction neural networks of the present disclosure can work together to generate the setups required for orthodontic treatment. For example, the GDL setup model can generate the final setup, and the RL setup model can use the final setup as an input to generate a series of intermediate stage setups. Alternatively, the VAE setup model (or the MLP setup model) can create the final setup, which can be used by the RL setup model to generate a series of intermediate stage setups. In some embodiments, the setup prediction can be generated by one setup prediction neural network and then used as an input for another setup prediction neural network for further refinement and adjustment. In some embodiments, such refinement can be performed iteratively.
[0105] In some embodiments, a setup verification model may be involved in this iterative setup prediction loop, such as the model disclosed in U.S. Provisional Application No. US63 / 366495. First, setups can be generated (e.g., using models trained for setup prediction, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, etc.), and then the setups are verified. If the setup passes the verification, the setup can be output for use. If the setup does not pass the verification, the setup can be sent back to one or more of the setup prediction models for correction, improvement, and / or adjustment. In some cases, the setup verification model can output an indication of what is wrong with the setup, such that the setup generation model can be improved in the next iteration. The process iterates until completion.
[0106] Generally, in some specific implementations, two or more of the following techniques of the present disclosure can be combined during orthodontic and / or dental treatment: GDL setting, setting classification, reinforcement learning (RL) setting, setting comparison, autoencoder setting (VAE setting or capsule setting), VAE grid element labeling, masked autoencoder (MAE) grid filling, multi-layer perceptron (MLP) setting, metric visualization, estimation of missing oral care parameter values, tooth classification using latent vectors, FDG setting, pose transfer setting, prosthetic design metric calculation, neural network techniques for dental restoration and / or orthodontics (e.g., 3D oral care representation generation or modification using transformers), landmark-based (LB) setting, diffusion setting, estimation of tooth movement protocols, capsule autoencoder segmentation, diffusion segmentation, similarity setting, verification of oral care representations (e.g., using autoencoders), coordinate system prediction, prosthetic design generation or 3D oral care representation generation or modification using denoising diffusion models.
[0107] Some autoencoder-based specific implementations of the present disclosure use capsule autoencoders to automate processing steps in the creation of oral care appliances (e.g., for orthodontic treatment or dental restoration). The advantage of using a capsule autoencoder that has been trained on oral care data is to utilize latent space technology, which reduces the dimensionality of oral care grid data, thereby refining this data, making the signals in the data stronger and easier to use by downstream processing modules, whether these downstream modules can be other autoencoders, decoders, other neural networks, or other types of ML models (such as the supervised and unsupervised models described elsewhere in the present disclosure). Capsule autoencoders were initially applied in the 2D domain to perform object recognition in 2D images, where capsules were trained to create models of the objects to be recognized. This method is capable of recognizing objects in 2D images even if the objects are imaged from new views that do not exist in the training dataset. Later research extended capsule autoencoders to the 3D point cloud domain, such as in "3D Point Capsule Networks" in the conference of CVPR 2019, the entire content of which is incorporated herein by reference.
[0108] This disclosure extends the results of this research to apply capsule autoencoders to the digital oral care domain, dealing with 3D point clouds, 3D meshes, and 3D voxelized representations. In a particular embodiment, the term "mesh" herein should be considered interchangeable with 3D point clouds and 3D voxelized representations. A 3D autoencoder can encode one or more 3D geometries (point clouds or meshes) into latent capsules that encode the reconstruction properties of the input 3D representation. These latent capsules exist in two or more dimensions and describe the features of the input mesh (or point cloud) and the likelihood of those features. A set of latent capsules is contrasted with latent vectors that can be produced by a variational autoencoder (VaE), which can be encoded as 1D vectors. One contribution of this technology is to advantageously apply capsule autoencoders to the digital oral care space, with the technical advantage of data-oriented accuracy in improving prediction results.
[0109] Specific examples of applications include segmentation of 3D oral care geometries, setup prediction (both final and intermediate stages), mesh cleaning of 3D oral care geometries (e.g., for both labeling of mesh elements and filling of missing mesh elements), tooth classification (e.g., according to standard dental notation schemes), setup classification (e.g., as malocclusion, grading, and final setup), and automated dental restoration design generation.
[0110] One or more latent capsules that describe the input 3D representation (e.g., oral care geometries such as point clouds and / or meshes representing an unsegmented dental arch, segmented teeth such as arranged in a malocclusion setup, teeth attached with hardware, teeth without attached hardware, etc.) can be provided to a capsule decoder to reconstruct a replica of the input 3D representation. The replica can be compared with the input 3D representation by calculating a reconstruction error, thus demonstrating the information-rich nature of the latent capsules (i.e., the latent capsules describe sufficient reconstruction properties of the input mesh such that the mesh can be reconstructed from the latent capsules). A low reconstruction error (e.g., below a predetermined loss threshold) indicates a successful reconstruction. Some of the applications disclosed herein use these information-rich latent capsules for further processing (e.g., such as setup prediction, mesh segmentation, coordinate system prediction, labeling of mesh elements for mesh cleaning, filling of missing mesh elements or holes in the mesh, classification of setups, classification of oral care meshes, verification of setups, and other verification tools). Some of the applications disclosed herein make one or more changes to the latent capsules, such as to achieve a change in the reconstructed mesh, which can then be output for further use (e.g., to create a dental restoration appliance).
[0111] Figure 2 Illustrated is a training method for a capsule autoencoder for reconstructing an oral care mesh (or point cloud). Figure 2A capsule autoencoder pipeline for mesh reconstruction is shown, which is mainly applied to oral care meshes in the non-limiting examples described herein, but can also be applied to other healthcare meshes or personal safety meshes, such as meshes related to the design, shape, function, and / or use of personal protective equipment (such as disposable respirators). This deployment method omits two modules on the bottom. This training method covers the entire diagram. The potential capsule T can be a dimension-reduced form of the input oral care mesh and can be used as an input for other processing.
[0112] Some prior arts rely on inputting 3D point cloud data into a capsule autoencoder. The techniques of the present disclosure extend the input geometry to include 3D mesh data and 3D voxelized representations. In some cases, an input point cloud or mesh (such as containing oral care data) can be rearranged into one or more vectors of mesh elements. Such a vector can be N×3 (in the case of representing the XYZ coordinates of points or vertices). Such a vector can be N×3 (in the case of representing a mesh face, each mesh face can be defined by 3 indices, each index being indexed into a list of vertices / points). Such a vector can be N×2 (in the case of representing a mesh edge, each mesh edge can be defined by 2 indices, each index being indexed into a list of vertices / points). Such a vector can be N×3 (in the case of representing a voxel, each voxel has an XYZ position, such as the centroid, where the length×width×height of each voxel is known).
[0113] In some examples according to aspects of the present disclosure, a neural network such as an MLP can be used to extract features from a list of N×3 mesh element inputs, resulting in a list of N×128 feature vectors, one feature vector for each mesh element. In some cases, a vector of one or more computed mesh element features (as defined elsewhere in the present disclosure) can be calculated for one or more of the N input mesh elements. In some specific implementations, these mesh element features can be used instead of the features generated by the MLP. In some specific implementations, each mesh element can be given a feature that is a mixture of the features generated by the MLP and the computed mesh element features. In this case, the layer dimension can be enhanced to N×(128 + aug_len), where aug_len is the length of the enhancement vector composed of the computed mesh element features. For ease of discussion and without loss of generality, this layer will be referred to hereinafter simply as N×128.
[0114] The length "aug_len" can vary according to the specific implementation, depending on which mesh elements are analyzed and which mesh element features are selected for use. In some cases, information from more than one type of mesh element can be introduced along with the N×128 vector (e.g., point / vertex information can be combined with face information, point / vertex information can be combined with edge information, or point / vertex information can be combined with voxel information). Depending on the various applications, the analysis of different kinds of oral care meshes may require one type of mesh element or another, or a specific set of mesh features.
[0115] The N×128 layer can be passed to a set of subsequent convolutional layers, each of which has been trained to have its own parameter values. The purpose of each of these individual convolutional layers can be to encode a single mesh element capsule. The output of each convolutional layer in the convolutional layer can be max-pooled to a size of 1024 elements. The count of these convolutional layers can be a power of two (e.g., 8, 16, 32, 64). In some specific implementations, there can be 32 such convolutional layers, each of which outputs a 1024-element vector from the max-pooling operation. These 32 max-pooling output vectors can be concatenated to form a layer that can be 1024×32, called the primary mesh element capsule (PMEC). The dynamic routing module encodes these PMECs into one or more latent capsules, each of which can have a square size (e.g., 16×16, 32×32, 64×64, or 128×128). Non-square sizes are also possible.
[0116] In some specific implementations, the dynamic routing module can enable the output of the latent capsules to be routed to the appropriate neural network layer in the subsequent processing module of the capsule autoencoder. The dynamic routing module uses unsupervised techniques (e.g., clustering and / or other unsupervised techniques) to arrange the output of the set of max-pooled feature maps into one or more stacked latent capsules. These latent capsules summarize the feature information from the input 3D representation (e.g., one or more tooth meshes or point clouds) as well as the likelihood information associated with each capsule. These stacked capsules contain sufficient information about the input 3D representation to reconstruct the 3D representation via the capsule decoder module.
[0117] A mesh of mesh elements (i.e., such as points / vertices, edges, faces, or voxels) can be generated by a mesh patch module. In this example, points will be used for the mesh elements. In some specific implementations, the mesh can include randomly arranged points. In other specific implementations, the mesh can reflect a regular and / or linear arrangement of points. The points in each of these mesh patches are the "raw materials" that can form the reconstructed 3D representation.
[0118] Latent capsules (e.g., having dimensions 128×128) can be replicated β times, and prior to being input into one or more MLPs, each of these β latent capsules can be sequentially appended with each of the grid patches of a randomly generated grid of grid elements (e.g., points / vertices). In some examples, such an MLP can include a fully connected layer having the following dimensions: {64 - 64 - 32 - 16 - 3}. The goal of this operation is to customize the grid elements to a particular local region of the 3D representation that may be to be reconstructed. The decoder iterates, generating additional random grid patches and outputting more random portions of the reconstructed 3D representation (i.e., as point cloud patches). These point cloud patches are accumulated until the reconstruction loss drops below a target threshold. One or more of the reconstruction losses (as defined herein) and the KL divergence loss can be used to calculate the reconstruction loss.
[0119] An autoencoder, such as a variational autoencoder (VAE), can be trained to encode 3D grid data in a latent space vector A that can exist in an information-rich low-dimensional latent space. This latent space vector A may be particularly suitable for subsequent processing in digital oral care applications (e.g., such as mesh cleaning, mesh segmentation, mesh validation, mesh classification, setup classification, setup prediction, and restorative design generation), as A enables efficient manipulation of the high-dimensional dental mesh data. Such a VAE can be trained to reconstruct the latent space vector A back into a facsimile of the input grid (or a transformation or other data structure that describes the 3D oral care representation). In some embodiments, the latent space vector A can be strategically modified in order to result in a change to the reconstructed grid (or other data structure). In some cases, the reconstructed grid can be a dental mesh having an altered and / or improved shape, such as would be suitable for use in the design of a dental restorative appliance (such as a 3M FILTEK Matrix or veneer). The term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non-limiting sense.
[0120] The dental reconstruction VAE can advantageously utilize loss functions, non - linearities (also known as neural network activation functions), and / or solvers not mentioned in the prior art. Examples of loss functions can include: mean absolute error (MAE), mean squared error (MSE), L1 loss, L2 loss, KL divergence, entropy, and reconstruction loss. Such loss functions enable each generated prediction to be compared in a quantifiable manner with the corresponding ground truth, resulting in one or more loss values that can be used to at least partially train one or more neural networks. Examples of solvers can include: dopri5, bdf, rk4, midpoint, adams, explicit_adams, and fixed_adams. Solvers can enable neural networks to solve systems of equations and the corresponding unknown variables. Examples of non - linearities can include: tanh, relu, softplus, elu, swish, square, and identity. Activation functions can be used to introduce non - linear behavior into the neural network in a way that enables the neural network to better represent the training data. The loss can be calculated through the process of training the neural network via backpropagation. Neural network layers such as the following can be used: ignore, concat, concat_v2, squash, concatsquash, scale, and concatscale.
[0121] In some specific embodiments, the dental reconstruction VAE model can be trained on patient cases of teeth in malocclusion or, alternatively, in local coordinates. Figure 3 A method of training such a VAE is shown.
[0122] According to Figure 3 As shown in the grid reconstruction VAE training, the 3D oral care representation F can be provided to the encoder E1 (along with optional tooth type information R), which can generate the latent vector A. The latent vector A can be reconstructed into the reconstructed 3D oral care representation G. A loss can be calculated between the reconstructed 3D oral care representation G and the ground truth 3D oral care representation GT (e.g., using the VAE loss calculation method or other loss calculation methods described herein). Backpropagation can be used to train E1 and D1 with such a loss.
[0123] Figure 4 A trained grid reconstruction VAE in deployment is shown. In Figure 4In the case of, the Mesh Reconstruction VAE is shown reconstructing a dental mesh in a reconstruction deployment. R is an optional input, particularly in the case of dental mesh classification, when such information R is not yet available (depending on a particular implementation, since the dental mesh classification neural network is trained to generate the dental type information R as an output). In some implementations, R can be used to improve other techniques, such as mesh element labeling techniques, mesh reconstruction techniques, oral care mesh classification techniques (e.g., such as tooth classification or appliance classification), etc.
[0124] Figure 5 and Figure 6 shows a reconstructed dental mesh. Figure 5 An example of an input dental mesh is illustrated on the left, and the output reconstructed dental mesh is illustrated on the right. Figure 6 Another example of an input dental mesh is illustrated on the left, and the corresponding output reconstructed dental mesh is illustrated on the right. Figure 5 The use case of Figure 6 is different from Figure 7 shows a depiction of the reconstruction error from Figure 6 the reconstructed teeth shown, called a reconstruction error map. Figure 7 The reconstruction error in the above results is depicted in a form called a "reconstruction error map", where the unit is millimeters (mm). It should be noted that the reconstruction error at the tooth tip is less than 50 microns, and the reconstruction error on most of the tooth surface is much less than 50 microns. Compared with a typical tooth of size 1.0 cm, an error rate of 50 microns (or less) means reconstructing the tooth surface with an error rate of less than 0.5%. In other words, the implementations described herein achieve a very low error rate. Figure 8 is a bar chart where each bar represents a single tooth and represents the average absolute distance of all vertices involved in the reconstruction of that tooth in the data used to evaluate the mesh reconstruction model.
[0125] A dental mesh autoencoder, an example of which is a variational autoencoder (VAE), can be trained to encode teeth into a reduced-dimensional form called a latent space vector. The reconstruction VAE can be trained on example dental meshes. The dental mesh can be received by the VAE, deconstructed into a latent space vector using a 3D encoder, and then reconstructed into a copy of the input mesh using a 3D decoder. The prior art for setting prediction lacks this deconstruction / reconstruction method. One advantage of this method is that the encoder E1 can be trained to encode a dental mesh (or a mesh of a dental appliance, gums, or other body part or anatomical structure) into a reduced-dimensional form that can be used to train and deploy any set of powerful setting prediction methods (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, and diffusion settings, etc.). This reduced-dimensional form of the teeth can enable the setting prediction neural network to more effectively encode the reconstruction characteristics of the teeth and better learn to place the teeth into a pose suitable for the final setting or intermediate stage, thus providing a technical improvement in terms of data accuracy and resource occupancy.
[0126] The reconstructed mesh can be compared to the input mesh, for example, using a reconstruction error that quantifies the difference between the meshes (as described elsewhere in this disclosure). This reconstruction error can be calculated using the Euclidean distance between corresponding mesh elements of the two meshes. There are also other methods of calculating this error that can be derived from materials described elsewhere in this disclosure. Figure 7 and Figure 8 An example reconstruction error according to the techniques described herein is shown.
[0127] In some embodiments, one or more meshes provided to the mesh reconstruction VAE can first be converted into a vertex list (or point cloud) before being provided to the encoder E1. This way of processing the input to E1 can be beneficial for a single mesh input (such as in a dental mesh classification task) or a set of multiple teeth (such as in a setting classification task). The input meshes do not need to be connected.
[0128] The encoder E1 can be trained to encode a dental mesh into a latent space vector A (or "tooth representation vector"). During a prosthetic design task, the encoder E1 can arrange the input dental mesh into a mesh element vector F and encode it into a latent space vector A. This latent space vector A can be a reduced-dimensional representation of F that describes the important geometric properties of F. The latent space vector A can be provided to the decoder D1 to be restored to full resolution or near full resolution together with the desired geometric changes. The restored full-resolution mesh or near full-resolution mesh can be described by G, which can then be arranged into the output mesh.
[0129] In some embodiments, such as in prosthetic design generation, tooth names, tooth nomenclature, and / or tooth type R may be concatenated with the latent vector A as a means of conditioning the VAE with such information to improve the VAE's ability to respond to a particular tooth type or nomenclature.
[0130] The performance of the mesh reconstruction VAE can be measured using reconstruction error calculation. In some examples, the reconstruction error can be calculated as the element-to-element distance between two meshes using, for example, the Euclidean distance. According to various embodiments of the techniques of the present disclosure, other distance measurements are possible, such as cosine distance, Manhattan distance, Minkowski distance, Chebyshev distance, Jaccard distance (e.g., intersection over union of the meshes), Hausdorff distance (e.g., distance across a surface), and Sorensen-Dice distance.
[0131] In some embodiments, the performance of the mesh reconstruction VAE can be verified via a reconstruction error map and / or other key performance indicators. The latent space vectors of one or more input tooth meshes can be plotted (e.g., in 2D) using UMAP or t-SNE dimensionality reduction techniques and compared to select the best available separability between classes of teeth (molars, premolars, and / or incisors, etc.), thereby indicating that the model is aware of strong geometric differences between different classes and strong similarities within a class. This will be illustrated by distinct non-overlapping clusters in the resulting UMAP / t-SNE plot.
[0132] In some cases, the latent vector corresponding to the mesh can be used as part of a classifier to classify it. For example, classification can be performed to identify the tooth type or to detect errors in the mesh (or arrangement of meshes), such as in a validation operation. The latent vector and / or computed mesh element features (such as the spatial and / or structural mesh features described herein) can be provided to a supervised machine learning model to classify the mesh. An incomplete list of possible supervised ML models is found elsewhere in the present disclosure.
[0133] In some embodiments, the reconstruction VAE can be trained to reconstruct any arbitrary tooth type. In other embodiments, the reconstruction VAE can be trained to reconstruct a specific tooth type (e.g., first molar or central incisor).
[0134] Figure 9Describes the training of a grid reconstruction VAE. In some specific implementations, the grid reconstruction VAE can be used to encode a dental mesh (or other 3D oral care representation) into a latent representation (e.g., a latent vector) A. In other words, the encoder part of the VAE can encode a 3D oral care representation into a latent representation. The VAE can also be trained to encode other types of 3D representations (e.g., set transformations, mesh element labels, or meshes depicting gums, fixture model components, oral care hardware such as brackets and / or attachments, dental restoration appliance components, other parts of the anatomical structure, etc.) into the latent vector A.
[0135] In some specific implementations, the latent representation can be reconstructed (e.g., using an autoencoder decoder or other decoders described herein) into a reconstructed form of the initial 3D oral care representation (e.g., a reconstructed tooth). A reconstruction error can be calculated between the initial version and the reconstructed version of the 3D oral care representation. Validation can be performed by comparing the reconstruction error with one or more thresholds. When the measured reconstruction error exceeds the threshold, the validation method can produce a "failed" result (e.g., the 3D oral care representation fails validation). Otherwise, the validation method can produce a "passed" result (e.g., the 3D oral care representation is suitable for use in generating an oral care appliance).
[0136] In other specific implementations, the latent representation can be provided to a second ML module. The second ML module (e.g., a Gaussian process, SVM, neural network, or another discriminative machine learning model) can be trained to validate the latent representation (and, by extension, to classify the initial 3D oral care representation from which the latent representation was generated). The validation can generate a determination as to whether the 3D oral care representation (e.g., a dental mesh, one or more transformations, one or more mesh element labels, etc.) is suitable for use in generating an oral care appliance.
[0137] Figure 9 Illustrates a method by which the system of the present disclosure can implement training of a reconstruction autoencoder for reconstructing a 3D representation of a patient's dentition. Thus, Figure 9 Provides further details regarding the training of the crown reconstruction VAE of the present disclosure. Figure 9Specific examples illustrate the training of a variational autoencoder (VAE) for reconstructing a dental mesh 900. For each tooth in a patient case (908), the system of the present disclosure can generate a watertight mesh by merging the crown mesh of the tooth with the corresponding root mesh such that vertices on the open edge of the crown mesh match vertices on the open edge of the root mesh (902). The system of the present disclosure can perform a registration step (904) to align the dental mesh with a template dental mesh (e.g., using the iterative closest point technique or by applying an inverse orthonormal transform to the tooth), with a technical enhancement of improving the accuracy of mesh correspondence calculation and data precision at 906. The system of the present disclosure can calculate the correspondence between the dental mesh and the corresponding template dental mesh, where the technical improvement is to adjust the dental mesh to be ready to be provided to the reconstruction autoencoder. The dataset of the prepared dental meshes is divided into a training set, a validation set, and a held-out test set (910) and then used to train the reconstruction autoencoder (912), described herein as a dental VAE, a dental reconstruction VAE, or more generally as a reconstruction autoencoder. The dental VAE can include a 3D encoder that encodes the dental mesh into a latent form (e.g., latent vector A), and a subsequent 3D decoder that reconstructs the tooth into a replica of the input dental mesh. The dental VAE of the present disclosure can be trained using a combination of a reconstruction loss and a KL divergence loss and optionally other loss functions described herein. The output of the method is a trained dental VAE 914.
[0138] Figure 10 illustrates non-limiting code implementing an example 3D encoder and an example 3D decoder for a mesh reconstruction VAE. In Figure 10 a specific example, the code is source code in Python for the encoder and decoder. These specific implementations can include: convolutional operations, batch normalization operations, linear neural network layers, Gaussian operations, and continuous normalizing flows (CNFs), among others.
[0139] One of the steps that may occur during VAE training data preprocessing is the calculation of mesh correspondences. Correspondences can be calculated between the mesh elements of the input mesh and the mesh elements of a reference or template mesh with a known structure. The purpose of the mesh correspondence calculation may be to find matching points between the surfaces of the input mesh and the template (reference) mesh. Mesh correspondences can be generated to create a point-to-point correspondence between the input mesh and the template mesh by mapping each vertex from the input mesh to at least one vertex in the template mesh. Correspondences can be calculated between the mesh elements of the input mesh and the mesh elements of a reference or template mesh with a known structure. In one example, the range of entries in a vector may correspond to the mesial lingual cusp tip; another range of elements may correspond to the distal lingual cusp tip; another range of elements may correspond to the mesial surface of the tooth; another range of elements may correspond to the lingual surface of the tooth, and so on. In the case of a dental mesh reconstruction autoencoder (such as a VAE), in some embodiments, the autoencoder may be trained only on a subset of teeth (e.g., only molars or only the left upper first molar). In other embodiments, the autoencoder may be trained on a larger subset of the teeth in the mouth or on all teeth. In some embodiments, an input vector (e.g., a landmark vector) may be provided to the autoencoder, which may define or otherwise influence which type of dental mesh the autoencoder may have received as input. The data accuracy improvement of this method is the mesh correspondence in mesh reconstruction to reduce sampling error, improve alignment, and improve mesh generation quality. Further details regarding the use of mesh correspondences in the autoencoder model of the present disclosure are found elsewhere in the present disclosure.
[0140] In some embodiments, during the calculation of the mesh correspondence, an Iterative Closest Point (ICP) algorithm may be run between the input dental mesh and the template dental mesh. Correspondences can be calculated to establish vertex-to-vertex relationships (between the input dental mesh and the reconstructed dental mesh) for use in calculating the reconstruction error.
[0141] In some embodiments, during the calculation of the mesh correspondence, an inverse rigid transformation may be applied to at least approximately align the input dental mesh and the template dental mesh. In some embodiments, both the ICP and the inverse rigid transformation may be applied.
[0142] According to a particular implementation, training data can be generalized to one or more dental arches (e.g., as well as other 3D or larger oral care representations), or can be more specific to particular teeth within a dental arch (e.g., as well as other 3D oral care representations). In cases where more specific training data is utilized, the specific training data can be presented as a tooth template. For example, the tooth template can be specific to one or more tooth types (e.g., the right lower central incisor). In some implementations, a tooth template can be generated that is an average of many examples of a certain type of tooth (such as the average of the lower first molar). In some implementations, a tooth template can be generated that is an average of many examples of more than one tooth type (such as the average of the first and second bicuspids from both the upper and lower dental arches).
[0143] In some implementations, the preprocessing process can involve one or more of the following steps: generating a watertight mesh (e.g., ensuring that the boundaries of the root mesh cleanly seal the boundaries of the crown mesh), registration to align the tooth mesh with a template mesh (e.g., using ICP or an inverse rigid transformation), and computing mesh correspondences (i.e., generating a mesh element to mesh element correspondence between the input tooth mesh and the template tooth mesh).
[0144] In Figure 11 the left side (labeled "Training Data (ICP)") shows the tooth mesh (in the form of a 3D point cloud) after the preprocessing steps are completed, where the preprocessing uses ICP for registration. The right side shows two things: the output of the tooth reconstruction VAE (in the left column) and the corresponding ground truth tooth 3D representation. Also in this case, the 3D representation of each tooth is represented by a point cloud. Figure 11 The output shown was generated at epoch 849 of the reconstruction VAE training.
[0145] The above description mainly relates to processing meshes, point clouds, and / or voxel data into latent space vectors as a means of reducing the dimensionality of that data and enhancing the signal-to-noise ratio of that data, such that an ML classifier can make decisions based on that data. Applications include but are not limited to VAE settings, MLP settings, MAE mesh filling, VAE mesh element labeling, VAE for tooth mesh classification, and some examples of classification settings. The reconstruction autoencoder trained based on the above materials is also related to validation operations, such as segmentation validation, coordinate system validation, mesh cleaning validation, repair design validation, fixture model validation, clear tray aligner (CTA) trim line validation, setting validation, oral care appliance component validation (either or both of placement and generation), and hardware (brackets, attachments, etc.) placement validation, to name just a few examples.
[0146] Other types of data :
[0147] Autoencoders of the present disclosure, such as VAEs or capsule autoencoders, can process other types of oral care data, such as text data, categorical data, spatio-temporal data, real-time data, and / or real number vectors, such as those found in protocol parameters. The data can be qualitative or quantitative. The data can be nominal or ordinal. The data can be discrete or continuous. The data can be structured, unstructured, or semi-structured. The autoencoders of the present disclosure can also encode such data into latent space vectors (or latent capsules) for later reconstruction. Those latent vectors / latent capsules can be used for prediction and / or classification. For example, through the calculation of reconstruction error and / or the labeling of data elements, the reconstruction can be used for model validation and for validation applications.
[0148] The latent vector A (e.g., for a tooth mesh) that can be generated by the encoder E1 in a fully trained mesh reconstruction autoencoder can be a reduced-dimensional representation of the input mesh (e.g., a tooth mesh). In some specific embodiments, the latent vector A can be a vector of 128 real numbers (or some other size, such as 256 or 512). The decoder D1 of the fully trained mesh reconstruction autoencoder may be able to take the latent vector A as input and reconstruct a close approximation of the input tooth mesh with a low reconstruction error. In some specific embodiments, the latent vector A can be modified in order to achieve a change in the shape of the reconstructed mesh generated by the decoder D2. Such modification can be performed after first mapping out the latent space to gain insight into the effects of making specific changes. There are various loss functions that can be used in the training of E1 and D1, which can involve terms related to reconstruction loss and / or KL divergence between distributions (e.g., in some cases, to minimize the distance between the latent space distribution and a multi-dimensional Gaussian distribution). One purpose of the reconstruction loss term is to compare the predicted reconstructed 3D representation of a tooth with the corresponding ground truth reconstructed 3D representation of the tooth. One purpose of the KL divergence term is to make the latent space more Gaussian and thus improve the quality of the reconstructed mesh (i.e., especially in cases where the latent space vector can be modified, to change the shape of the output mesh, such as splitting a 3D mesh or performing tooth design generation for use in generating dental restoration appliances).
[0149] In some specific embodiments, the latent vector A can be modified in order to change the characteristics of the reconstructed mesh (such as in the generation of a tooth restoration tooth design mesh). If only the reconstruction loss is used to calculate the loss L and the latent vector A is changed, then in some use case scenarios, the reconstructed mesh can reflect the expected output form (e.g., a recognizable tooth). However, in other use case scenarios, the output of the reconstructed mesh may not conform to the expected output form (e.g., not a recognizable tooth).
[0150] Figure 12 Illustrates a latent space in which the loss includes reconstruction loss but does not include KL divergence loss. InFigure 12 In this case, point P1 corresponds to the initial form of the latent space vector A. Point P2 corresponds to a different position in the latent space, which can be sampled as a result of modifying the latent vector A, but the grid reconstructed from P2 may not give a good output (e.g., it may not look like a recognizable or otherwise suitable tooth). Point P3 corresponds to yet another different position in the latent space, which can be sampled as a result of a different set of modifications to the latent vector A, and the grid reconstructed from P3 can give a good output (e.g., having an appearance with a tooth design suitable for use in generating dental restoration appliances). In the case where the loss only involves the reconstruction loss, the subset of the latent space that can be sampled to obtain the latent space vector P3 that produces a valid reconstructed grid may be irregular or difficult to predict.
[0151] In some specific embodiments, the loss calculation can incorporate normalizing flows, for example, by combining a KL divergence term. Thus, including a KL divergence term as described herein enables training involving flow normalization. If the loss is improved by combining the KL divergence term, the quality of the latent space can be significantly enhanced. In this new scenario, the latent space may become more Gaussian (as Figure 13 shown), and the latent hypervector A corresponds to a point P4 near the center of the multi-dimensional Gaussian curve. In Figure 13 this case, the loss in the latent space includes both the reconstruction loss and the KL divergence loss. Changes can be made to the latent hypervector A to obtain a point P5 near P4, where the resulting reconstructed grid is likely to reflect the desired properties (e.g., it is likely to be a valid tooth). Introducing the KL divergence term into the loss can make the process of modifying the latent space vector A and obtaining a valid reconstructed grid more reliable. In some specific embodiments, similar to the capsule autoencoder, the latent vector can be replaced with a latent capsule, which can be modified and then reconstructed. In some specific embodiments, the autoencoder framework can be adapted for the segmentation of dental meshes. Additionally, in some specific embodiments, the autoencoder framework can be adapted for the task of dental coordinate system prediction. In some specific embodiments, a grid reconstruction autoencoder for coordinate system prediction can compress dental data into a latent vector form and then provide the latent vector as an input to a second ML module (e.g., an MLP), which may have been trained for coordinate system prediction (e.g., for coordinate system prediction on a mesh, the goal of which is to define a local coordinate system for the mesh, such as a dental mesh).
[0152] For a given domain (e.g., dental restoration design generation, MAE dental filling or setup design, etc.), a latent space can be mapped such that a change to the latent space vector A results in a reasonably well-reconstructed mesh. The latent space can be systematically mapped by generating latent vectors with carefully chosen value variations (e.g., by experimenting with different combinations of 128 values in an example latent vector). In some cases, a grid search of values can be performed, which has the advantage of effectively exploring the latent space. When the latent space is mapped, the shape of the mesh can be modified by nudging the values in one or more elements of the latent vector value towards the part of the mapped latent space that has been found to correspond to the desired dental characteristics. Using KL divergence augmentation in the loss calculation increases the likelihood of reconstructing the modified latent vector as a valid example of the input 3D oral care representation (e.g., 3D dental mesh).
[0153] In the case of restoration design generation, the mesh can correspond to at least some parts of a tooth. The latent vector A can be altered such that the resulting reconstructed dental mesh can have characteristics that meet the specifications set by the restoration design parameters. The neural network for dental restoration design generation is described in U.S. Provisional Application No. US63 / 366514, the entire disclosure of which is incorporated herein by reference.
[0154] A dental setup can be designed at least in part by modifying the latent vector corresponding to one or more teeth (e.g., each tooth described as a 3D point cloud, voxel, or mesh) to be placed in a setup configuration. The mesh can be encoded as the latent vector A, which then undergoes modification to adjust the pose of the resulting tooth pose. Then, the modified latent vector A' can be reconstructed into one or more meshes describing the setup. This technique can be used to design a final setup configuration or an intermediate stage configuration, etc.
[0155] In some embodiments, the modification of the latent vector can be performed via an ML model (such as one of the neural network models or other ML models disclosed elsewhere in this disclosure). In some embodiments, the neural network can be trained to operate within the latent space of such a vector A of the setup mesh. The mapping of the latent space of A may be pre-generated by making controlled adjustments to the trial latent vectors and observing the resulting changes in the setup configuration (i.e., after the modified A is reconstructed back into one or more complete meshes of the dental arch). In some cases, the mapping of the latent space can follow an organized search pattern, such as in a grid search.
[0156] In some specific implementations, the dental reconstruction VAE can take a single input of tooth name / type / designation R, which can command the VAE to output a tooth mesh of a specified type. This can be achieved by generating a latent vector A' used in reconstructing the appropriate tooth mesh. In some specific implementations, this latent vector A' can be "instantly" sampled or generated from a previous mapping of the latent vector space. This mapping can be performed to understand which parts of the latent vector space correspond to different shapes, structures, and / or geometries of teeth. For example, among the 128 real values in an example of the latent vector A' (other sizes are possible), certain elements of those vector elements, and perhaps, certain value ranges of those vector elements, can be determined to correspond to a specific type / name / designation of a tooth and / or a tooth having certain shapes or other desired characteristics. The model for tooth mesh generation can also be applied to the generation of oral care hardware, appliances, and appliance components (such as for orthodontic treatment). The model can also be trained to generate other types of anatomical structures. The model can also be trained to generate other types on non-oral care meshes.
[0157] The mesh comparison module can compare two or more meshes, for example, for the calculation of a loss function or for the calculation of reconstruction error. Some specific implementations can involve the comparison of the volumes and / or areas of two meshes. Some specific implementations can involve calculating the minimum distance between corresponding vertices / faces / edges / voxels of two meshes. For a point in one mesh (e.g., a vertex, the midpoint on an edge, or the center of a triangle), the minimum distance between that point and the corresponding point in the other mesh is calculated. In cases where the other mesh has a different number of elements or there is no clear mapping between the corresponding points of the two meshes, different methods can be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools that can play a role in the mesh comparison module of the present disclosure. In some specific implementations, the Hausdorff distance can be calculated to quantify the shape difference between two meshes. The open-source software tool Metro developed by the Visual Computing Lab can also play a role in quantifying the difference between two meshes. The following paper describes the method adopted by Metro, which can be modified by the neural network application of the present disclosure for mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces", P. Cignoni, C. Rocchini, and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, Vol. 17(2), June 1998, pp. 167-174.
[0158] Some techniques of the present disclosure may involve the following operations: for one or more points on a first grid, emitting rays perpendicular to the grid surface and calculating the distance before the rays impinge on a second grid. The length of the resulting line segment can be used to quantify the distance between the grids. According to some techniques of the present disclosure, a color can be assigned to the distance based on the magnitude of the distance, and the color can be applied to the first grid by means of visualization.
[0159] The entire disclosure of U.S. Provisional Application No. US63 / 352,850 is incorporated herein by reference. Techniques for verifying the correctness of oral care representations, such as 3D oral care grids, are described below. In some embodiments, an autoencoder, such as a variational autoencoder (VAE) or a capsule autoencoder, can be used to verify the correctness of an oral care grid. In some cases, a grid can be generated by scanning a 3D printed object. In some cases, the grid can correspond to a patient's dentition (e.g., including teeth, gums, hardware, etc.) and be used to create a dental or orthodontic appliance. It should be understood that, without loss of generality, the verification of a "3D grid" is intended to cover the verification of other forms of 3D representations, such as 3D point clouds and 3D voxelized representations (such as those that can be used in sparse processing).
[0160] The general principle is that an autoencoder, such as a VAE or a capsule autoencoder, can deconstruct, compress, and / or encode the reconstruction characteristics of an input grid into a latent space vector (e.g., a vector of length 128, although 2D and n - dimensional vectors and data structures are also possible). The input grid can include thousands, tens of thousands, or millions of grid elements and is compressed into a small data structure (e.g., a vector of m floating - point values, where m can be equal to 128 or some other number). Then the latent space vector can be provided to the decoder part of the autoencoder, which may have been trained to reconstruct the vector as a grid.
[0161] This autoencoder can be trained on many examples of correctly formed 3D representations, such as, for example, a 3D mesh of a tooth, a dental arch, or a complete fixture model including a base attached to a dentition. Through training, the autoencoder can become effective and efficient in deconstructing and subsequently reconstructing the type of mesh represented by the training examples. When training the autoencoder, providing training examples that are reasonably cohesive in form, structure, shape, layout, features, attributes, and / or other characteristics can provide one or more benefits in terms of future performance at the execution stage. If the introduced mesh significantly deviates from the characteristics of the mesh on which the autoencoder was trained, the autoencoder may not be able to accurately reconstruct the mesh. There may be a high reconstruction error (e.g., as calculated using the per-vertex Euclidean distance). Such a high reconstruction error can flag the mesh processed by the autoencoder as incorrect and in need of remediation. In some embodiments, the VAE verification engine can output an indication of which error occurred, e.g., what problem may have occurred with the received 3D representation (e.g., in the case of a crown mesh, there may be a hardware issue). A notification can be generated that the mesh should be corrected or modified to meet the expectations for use in designing and / or manufacturing dental and / or orthodontic appliances. In some embodiments, the generated notification can be displayed to the user of the system to initiate a remediation measure. In other embodiments, the system can use the generated notification to automatically remediate the identified error. For example, in the case of generating a 3D mesh by scanning a 3D printed object, it may be necessary to reprint the 3D printed part. In some cases, it may be necessary to correct the 3D part before reprinting.
[0162] The autoencoder-based verification techniques of the present disclosure can be advantageously applied to any one of the following non-exhaustive list of technologies: segmentation verification, coordinate system verification, mesh cleaning verification, in-chairside intraoral dental scan verification, clear tray aligner (CTA) setup verification, bracket / attachment placement verification, verification of custom oral care appliances (e.g., verifying the shape or placement of a dental restoration appliance component), prosthetic design generation verification, fixture model verification, and CTA trim line verification. Other meshes associated with digital oral care can also be verified using such techniques (e.g., involving an autoencoder, VAE, calculating a latent space vector from a 3D mesh, reconstructing the latent space vector into a 3D mesh, and calculating a reconstruction error).
[0163] For example, in the case of a dental setup, one or more teeth of the setup (e.g., an entire dental arch or a portion thereof) can be run through the deconstruct / latent vector / reconstruction phases of an autoencoder. Such an autoencoder can be trained on an exemplary setup (e.g., where the tooth poses are suitable for use in creating an orthodontic appliance). The reconstruction error of the output can be measured and used to determine whether the input dental arch represents a well-formed setup and is suitable for use in designing and manufacturing one or more orthodontic appliances such as clear tray aligners.
[0164] The following table describes the types of mesh data for which each validation encoder-decoder structure is trained to deconstruct / reconstruct. The encoder-decoder structure may include at least one encoder or at least one decoder. Non-limiting examples of encoder-decoder structures include U-Net, transformers, or autoencoders, among others.
[0165] In some examples, an oral care appliance or appliance component (e.g., such as for a dental restoration appliance) may be verified. An example of such a component is a mold parting surface. The mold parting surface may be combined with one or more teeth via a Boolean operation to divide the tooth. The tooth that has been cut by the parting surface may be encoded into a latent space form (e.g., using an autoencoder that has been trained to reconstruct that type of mesh). The latent space form (latent vector or latent capsule) may be classified by an ML classifier that has been trained to perform such a task.
[0166] In some embodiments, each of the above-described validation applications in this section may use a capsule autoencoder, such as Figure 2 the capsule autoencoder shown. An input 3D representation (such as that described in Table 1) is provided to the capsule encoder portion of the capsule autoencoder and may be encoded into one or more latent capsules. The one or more latent capsules may be reconstructed by the capsule decoder portion of the capsule autoencoder, thereby producing a facsimile of the input 3D representation. A reconstruction error may then be computed between the input 3D representation and the reconstructed 3D representation. A low reconstruction error (e.g., of a portion of the 3D representation, such as a subset of mesh elements, or of the entire 3D representation) may indicate that the input 3D representation has the general properties, character, geometric attributes, class, and / or category of the 3D representations used to train the capsule autoencoder. A high reconstruction error (e.g., of a portion of the 3D representation, such as a subset of mesh elements, or of the entire 3D representation) indicates that the input 3D representation does not have the general properties, character, geometric attributes, class, and / or category of the 3D representations used to train the capsule autoencoder. The reconstruction error may be computed using reconstruction loss, KL divergence loss, and / or a combination of both. Other possible losses include L1 loss, L2 loss, MSE loss, or any loss described elsewhere in this disclosure.
[0167] In some embodiments, one or more of the optional inputs may be provided to the autoencoder for mesh validation, including but not limited to: tooth size information P, tooth gap information Q, tooth position information N, tooth orientation information O, tooth name / label R for each available tooth, and / or orthodontic metric S. Figure 14 These inputs to the autoencoder for mesh validation (or validation of other 3D representations described herein) are shown. Thus, Figure 14A deployment method for validating an oral care mesh using an autoencoder is shown. In some specific implementations, a capsule autoencoder may be used instead of an autoencoder. In some specific implementations, arch shape information may be provided to the autoencoder. Other 3D oral care representations may also be validated using a reconstruction autoencoder trained to reconstruct that particular type of 3D oral care representation (e.g., labels on mesh elements or transformation matrices, etc.) by a comparable method.
[0168] An encoder-decoder structure including encoder E1 1402 and decoder D1 1406 can be trained as a reconstruction autoencoder to reconstruct a particular type of 3D oral care representation (e.g., a reconstructed mesh, a set of mesh element labels, a transformation, or others described herein). In deployment, Figure 14 The method shown can perform the validation of a trial 3D oral care representation 1400 (e.g., a tooth mesh, an appliance component mesh, a jig model mesh, a set of mesh element labels, a set of tooth transformations, a set of transformations for an appliance component, or others described herein). The trial 3D oral care representation can be provided to encoder E1 1402, which can generate a latent representation 1404. The latent representation 1404 can be provided to decoder D1 1406, which can reconstruct the latent representation 1404 into a reconstructed 3D oral care representation 1408 that is a close facsimile of the trial 3D oral care representation 1400. A reconstruction error can be calculated (1410) between the reconstructed 3D oral care representation 1408 and the trial 3D oral care representation 1400. When the reconstruction error is higher than a threshold (1412) for at least some portions of the reconstructed 3D oral care representation 1408, the method outputs (1416) an indication that the trial 3D oral care representation 1400 fails validation. When the reconstruction error is lower than the threshold (1412) for all portions of the reconstructed 3D oral care representation 1408, the method outputs an indication (1414) that the trial 3D oral care representation 1400 passes validation.
[0169] In some specific implementations, the validation application described herein can be incorporated into a source code testing framework as described in WO2022123402A1. The entire disclosure of PCT patent application WO2022123402A1 is incorporated herein by reference.
[0170] In some specific implementations, the techniques of the present disclosure can use PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using PointNet or PointNet++ as a basis for training) to extract local or global neural network features from 3D point clouds or other 3D representations (e.g., 3D point clouds depicting aspects of a patient's dentition such as teeth or gums). In some specific implementations, the techniques of the present disclosure can use U-Net to extract local or global neural network features from 3D point clouds or other 3D representations.
[0171] 3D oral care representations are described in this way herein because 3-dimensional representations are current state of the art. However, 3D oral care representations are intended to be used in a non-limiting manner to cover any representation of 3 dimensions or higher order dimensions (e.g., 4D, 5D, etc.), and it should be understood that the techniques disclosed herein can be used to train machine learning models to operate on representations of higher order dimensions.
[0172] In some cases, the input data can include 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data related to splines (e.g., control points). The encoder-decoder structure can include one or more encoders, or one or more decoders. In some specific implementations, the encoder can take as input the mesh element feature vectors of one or more of the input mesh elements. The encoder is trained in a manner that generates a more accurate representation of the input data by processing the mesh element feature vectors. For example, the mesh element feature vectors can provide the encoder with more information about the shape and / or structure of the mesh, and thus the additional information provided allows the encoder to make more informed decisions and / or generate a more accurate latent representation of the mesh. Examples of encoder-decoder structures include U-Net, autoencoders, or transformers, among others. The representation generation module can include one or more encoder-decoder structures (or portions of encoder-decoder structures, such as individual encoders or individual decoders). The representation generation module can generate an information-rich (optionally dimension-reduced) representation of the input data that can be more readily consumed by other generative or discriminative machine learning models.
[0173] The U-Net may include an encoder, followed by a decoder. The architecture of the U-Net may be similar to a U shape. The encoder may extract one or more global neural network features, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level compared to the most global level) from the input 3D representation. The output of each level from the encoder may be passed to the input of the corresponding level of the decoder (e.g., via skip connections). Similar to the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For example, the decoder may output a representation of the input data that may include global, intermediate, or local information about the input data. In some specific implementations, the U-Net may generate an informative (optionally dimension-reduced) representation of the input data that can be more easily consumed by other generative or discriminative machine learning models.
[0174] An autoencoder may be configured to encode input data into a latent form. The autoencoder may train the encoder to reformulate the input data into a dimension-reduced latent form between the encoder and the decoder, and then train the decoder to reconstruct the input data from this latent form of the data. A reconstruction error may be computed to quantify the extent to which the reconstructed form of the data differs from the input data. In some specific implementations, the latent form may be used as an informative dimension-reduced representation of the input data that can be more easily consumed by other generative or discriminative machine learning models. In most scenarios, the autoencoder may be trained on an input 3D representation, encode the 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close replica of the input 3D representation as output.
[0175] The transducer can be trained to generate a representation of its input using self-attention at least in part. The transducer can encode long-range dependencies (e.g., encoding relationships among a large number of inputs). The transducer can include an encoder or a decoder. In some embodiments, such an encoder can operate in a bidirectional manner or can operate a self-attention mechanism. In some embodiments, such a decoder can operate a masked self-attention mechanism, can operate a cross-attention mechanism, or can operate in an autoregressive manner. In some embodiments, the self-attention operation of the transducer described herein can be related to different positions or aspects of a single 3D oral care representation in order to compute a dimensionally reduced representation of the 3D oral care representation. In some embodiments, the cross-attention operation of the transducer described herein can mix or combine aspects of two (or more) different 3D oral care representations. In some embodiments, the autoregressive operation of the transducer described herein can consume previously generated aspects of a 3D oral care representation (e.g., previously generated points, point clouds, transforms, etc.) as additional inputs when generating a new or modified 3D oral care representation. In some embodiments, the transducer can generate a latent form of the input data that can serve as an information-rich dimensionally reduced representation of the input data, which can be more readily consumed by other generative or discriminative machine learning models.
[0176] In some embodiments, the encoder-decoder architecture can first be trained as an autoencoder. In deployment, one or more modifications can be made to the latent form of the input data. Then, the modified latent form can continue to be reconstructed by the decoder, resulting in a reconstructed form of the input data that is different from the input data in one or more desired aspects. Oral care variables such as oral care parameters or oral care metrics can be provided to the encoder, the decoder, or can be used to modify the latent form in order to affect the encoder-decoder architecture when generating a reconstructed form with desired characteristics (e.g., characteristics that may be different from those of the input data).
[0177] In some cases, federated learning can be used to train the techniques of the present disclosure. Federated learning can enable multiple remote clinicians to iteratively improve machine learning models (e.g., validation of 3D oral care representations, mesh segmentation, mesh cleaning, other techniques involving labeling mesh elements, coordinate system prediction, placement of non-organic objects on teeth, appliance component generation, dental restoration design generation, techniques for placing 3D oral care representations, setting prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, estimation of missing values), while protecting data privacy (e.g., clinical data may not need to be sent "over the network" to a third party). Data privacy is particularly important for clinical data protected by applicable laws. A clinician can receive a copy of the machine learning model, use a local machine learning program to further train the ML model using locally available data from a local clinic, and then send the updated ML model back to a central hub or a third party. The central hub or third party can integrate the updated ML models from multiple clinicians into a single updated ML model that benefits from the learning of patient data recently collected at various clinical sites. In this way, a new ML model can be trained that benefits from additional and updated patient data (possibly from multiple clinical sites), while that patient data is never actually sent to a third party. In some cases, training on devices within a local clinic can be performed when the device is idle or otherwise during non-working hours (e.g., when patients are not being treated at the clinic). Devices in a clinical environment for collecting data and / or training an ML model for the techniques described herein can include intraoral scanners, CT scanners, X-ray machines, laptop computers, servers, desktop computers, or handheld devices (such as smartphones with image collection capabilities). In addition to federated learning techniques, in some embodiments, contrastive learning can be used to at least partially train the ML models described herein. In some cases, contrastive learning can augment the samples in a training dataset to emphasize differences between different class samples and / or increase the similarity of samples within the same class.
[0178] In some cases, a local coordinate system for 3D oral care representations (such as teeth) can be described by one or more transformations (e.g., an affine transformation matrix, a translation vector, or a quaternion). The systems of the present disclosure can be trained for coordinate system prediction using the coordinate systems of past cohort patient case data. The past patient data can include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems. Machine information models (such as U-Net, an encoder, an autoencoder, a pyramid encoder-decoder, a transformer, or convolutional layers and / or pooling layers) can be trained for coordinate system prediction. Representation learning can determine a representation of a tooth (e.g., encoding a mesh or point cloud into a latent representation, such as using a U-Net, an encoder, a transformer, or convolutional layers and / or pooling layers, etc.), and then predict a transformation for that representation (e.g., using a trained multi-layer perceptron, a transformer, an encoder, a transformer, etc.), where the transformation defines a local coordinate system for that representation (e.g., including one or more coordinate axes). In the case of predicting a coordinate system for a tooth mesh, the mesh convolution techniques described herein can utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth. Reinforcement learning techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth.
[0179] Machine information models (such as U-Net, encoders, autoencoders, pyramid encoder-decoders, transformers, or convolutional layers and / or pooling layers) can be trained as part of a method for hardware (or appliance component) placement. Representation learning can train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder or using blocks of U-Net, encoders, transformers, convolutional layers, and / or pooling layers, etc.). The representation can include a dimension-reduced form and / or an information-rich form of the input 3D oral care representation. In some embodiments, the generation of the representation can be assisted by computing a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some embodiments, a representation can be computed for a hardware element (or appliance component). Such a representation is adapted to be provided to a second module, which can perform a generation task, such as transform prediction (e.g., a transform for placing a 3D oral care representation relative to another 3D oral care representation, such as a representation for placing a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such a transform can include an affine transformation matrix, a translation vector, or a quaternion, etc. Machine learning models that can be trained to predict a transform for placing a hardware element (or appliance component) relative to elements of a patient dentition include MLP, transformers, encoders, etc. The systems of the present disclosure can be trained for 3D oral care appliance placement using 3D oral care appliance placement using past cohort patient case data. Past patient data can include at least: one or more ground truth transforms and one or more 3D oral care representations (such as tooth meshes or other elements of a patient dentition). In the case where U-Net (and other neural networks) is trained to generate a representation of a tooth mesh, the mesh convolution and / or mesh pooling techniques described herein utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be performed for hardware or appliance component placement. Reinforcement learning techniques can be performed for hardware or appliance component placement.
Claims
1. A method for validating a three-dimensional (3D) oral care representation, the method comprising: Receiving, by a processing circuit of a computing device, a 3D oral care representation associated with a digital oral care process of a patient; Providing, by the processing circuit, the 3D oral care representation as an execution-phase input to a trained autoencoder network; and Executing, by the processing circuit, the trained autoencoder network to: Generate a reconstructed 3D oral care representation that is a copy of the 3D oral care representation; and Output an indication of the suitability of the 3D oral care representation for use in generating an oral care appliance for the patient.
2. The method of claim 1, wherein the 3D oral care representation comprises one or more mesh elements.
3. The method of claim 1, wherein the trained autoencoder comprises a multi-dimensional encoder that is trained to encode the 3D oral care representation into a latent space representation having a lower-order dimension than the 3D oral care representation.
4. The method of claim 3, wherein at least one mesh element feature is calculated for at least one of the mesh elements and provided to the trained autoencoder.
5. The method of claim 3, wherein the trained autoencoder comprises a multi-dimensional decoder configured to reconstruct the latent space representation into the reconstructed 3D oral care representation.
6. The method of claim 1, further comprising calculating, by the processing circuit, a reconstruction error that quantifies a difference between the 3D oral care representation and the reconstructed 3D oral care representation.
7. The method of claim 6, further comprising forming a suitability indication by comparing an absolute value of the reconstruction error with a threshold error value.
8. The method of claim 7, further comprising forming a suitability indication corresponding to determining that the 3D oral care representation is not suitable for use in oral care appliance generation based on determining that the absolute value of the reconstruction error is greater than the threshold error value.
9. The method of claim 7, further comprising forming a suitability indication corresponding to determining that the 3D oral care representation is suitable for use in the oral care appliance generation based on determining that the absolute value of the reconstruction error is less than or equal to the threshold error value.
10. The method of claim 8, further comprising: Generating, by the processing circuit, one or more indications of at least one modification to be made to the 3D oral care representation; and And Generating, by the processing circuit, a modified 3D oral care representation based at least in part on the one or more indications.
11. The method of claim 1, wherein the computing device is deployed in a clinical environment and wherein the method is performed near real-time during a session with the patient.
12. The method according to claim 1, wherein the 3D oral care representation comprises at least one of the following: a mesh element label, a 3D representation of teeth, an appliance component, a fixture model component, a transformation, an orthodontic setting, and an axis of a coordinate system of teeth.
13. The method according to claim 12, wherein the transformation is for at least one of teeth, an appliance component, and a fixture model component.
14. The method according to claim 12, wherein the 3D oral care representation comprises at least one of a hardware element, a fixture model component, and an appliance component that are proximate to aspects of the patient's dentition.
15. A computing device for validating a three-dimensional (3D) oral care representation, the computing device comprising: interface hardware configured to receive a 3D oral care representation associated with a digital oral care process of a patient; processing circuitry configured to: provide the 3D oral care representation as an execution-phase input to a trained autoencoder network; and execute the trained autoencoder network to: form a reconstructed 3D oral care representation that is a copy of the 3D oral care representation; and output an indication of the suitability of the 3D oral care representation for use in generating an oral care appliance for the patient.
16. The computing device according to claim 15, further comprising a reconstruction error calculated by the processing circuitry, the reconstruction error quantifying a difference between the 3D oral care representation and the reconstructed 3D oral care representation.
17. The computing device according to claim 16, further comprising forming a suitability indication by comparing an absolute value of the reconstruction error with a threshold error value.
18. The computing device according to claim 17, further comprising forming a suitability indication corresponding to determining that the 3D oral care representation is not suitable for use in oral care appliance generation based on determining that the absolute value of the reconstruction error is greater than the threshold error value.
19. The computing device according to claim 17, further comprising forming a suitability indication corresponding to determining that the 3D oral care representation is suitable for use in the oral care appliance generation based on determining that the absolute value of the reconstruction error is less than or equal to the threshold error value.
20. The computing device according to claim 18, further comprising: generating, by the processing circuitry, one or more indications of at least one modification to be made to the 3D oral care representation; and generating, by the processing circuitry, a modified 3D oral care representation based at least in part on the one or more indications.
Citation Information
Patent Citations
Method for automated generation of orthodontic treatment final setups
US20210259808A1
Method for automated generation of orthodontic treatment final setups
WO2020026117A1
System to generate staged orthodontic aligner treatment
WO2021245480A1
Automated processing of dental scans using geometric deep learning
WO2022123402A1