Geometric deep learning for final setting and intermediate grading in transparent tray appliances

By training neural network models to predict the teeth movement of transparent tray-type orthodontic devices, combined with oral care variables and doctor preferences, the problem of insufficient accuracy in the generation of CTA devices in the prior art is solved, and more efficient customization and accurate orthodontic treatment is achieved.

CN120345033APending Publication Date: 2025-07-18SOLVENTUM INTELLECTUAL PROPERTIES CO

Patent Information

Application Number
CN202380086034.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-19
Filing Date
2023-12-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art uses machine learning to generate transparent tray-type orthodontic devices (CTA) devices, and it is difficult to effectively automate the production of customized orthodontic treatments.

Method used

Using a neural network model, by receiving a digital representation of the patient's teeth, the training generator network predicts the final setup and intermediate stages of tooth movement, uses the loss function to quantify the difference and modify the model, combining oral care variables and doctor preferences to generate an accurate tooth movement plan.

Benefits of technology

It improves the generation accuracy and customization capabilities of transparent tray-type orthopedic devices, reduces computing resource consumption, and adapts to the rapid processing needs of the clinical environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345033A_ABST
    Figure CN120345033A_ABST
Patent Text Reader

Abstract

Systems and techniques for generating settings for an orthodontic alignment process are disclosed. The method involves receiving a digital representation of a patient's tooth and at least one value related to customization of an orthodontic treatment. A prediction of one or more tooth movements for setting is formed by executing a generator network comprising one or more neural networks. The generator network is further trained based on the formed prediction by performing operations including predicting the tooth movement, quantizing a difference between the predicted tooth movement and a reference tooth movement, generating a loss value based on the quantized difference, and modifying the generator network based on the loss value, to form a modified generator network. These systems and techniques enable efficient generation of settings for orthodontic alignment processes, thereby improving the accuracy and customization of the process.
Need to check novelty before this filing date? Find Prior Art

Description

Related Literature

[0001] The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosure of each of the PCT applications with publication numbers WO2022123402A1, WO2021245480A1, and WO2020026117A1 is incorporated herein by reference. The entire disclosure of each of the following provisional U.S. patent applications is incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 366,514; and 63 / 264,914. Technical Field

[0002] This disclosure relates to the configuration and training of neural networks to improve the accuracy of automatically generated clear tray aligner (CTA) devices used in orthodontic treatment. Summary of the Invention

[0003] Some prior art has attempted to use machine learning to generate CTA devices, but with mixed results. Accordingly, there is a need for better machine learning models and training methods to improve systems for automating CTA production.

[0004] This disclosure describes systems and techniques for training and using one or more machine learning models, such as neural networks, to produce intermediate and final settings of a CTA in a manner customized to a patient's treatment needs. Such a neural network is referred to herein as a "settings prediction neural network" or simply a "settings prediction model". The final setting (also referred to as the final setup) is the target configuration of a 3D tooth representation, such as a 3D tooth mesh, of the teeth as they appear at the end of treatment. The intermediate setting (also referred to as the "intermediate stage" or "intermediate grading") describes the configuration of the teeth during one of several stages of treatment after the teeth have moved out of their malocclusion posture (e.g., position and / or orientation) and before the teeth reach their final setting posture. In some embodiments, the final setting can be used to at least partially generate one or more intermediate stages. Each stage can be used to generate a clear tray aligner. Such an aligner can incrementally move the patient's teeth from an initial or malocclusion posture to a final posture represented by the final setting.

[0005] In a first aspect, a first computer-implemented method for generating a setup for an orthodontic alignment treatment is described, the method comprising the steps of: receiving, by one or more computer processors, a first digital representation of a patient's teeth; using, by one or more computer processors and for predicting one or more tooth movements of a final setup, a generator that is a machine learning model, such as comprising one or more neural networks (e.g., 3D encoder, 3D decoder, 3D U-Net, MLP, transformer, auto-decoder, pyramid encoder-decoder, neural network with attention layers, and other neural networks disclosed herein), the one or more neural networks having been initially trained to predict one or more tooth movements of a final setup; further training a setup prediction model by the one or more computer processors based on the using, and wherein the training of the setup prediction model is modified by performing operations that include predicting, by the generator, one or more tooth movements of a final setup based on the first digital representation of the patient's teeth; computing a loss function that quantifies the difference between the predicted tooth movements and reference tooth movements; and using the loss to modify the setup prediction model.

[0006] The first aspect may optionally include additional features. For example, the method may generate, by the one or more processors, an output state of the final setting. The method may determine, by the one or more computer processors, a difference between the one or more predicted tooth movements and the one or more reference tooth movements. The determined difference between the one or more predicted tooth movements and the one or more reference tooth movements may be used to modify the training of the generator. Modifying the training of the generator may include adjusting one or more weights of a neural network of the generator. The method may generate, by the one or more computer processors, one or more lists of mesh elements specifying the first digital representation of the patient's teeth. At least one of the one or more lists may specify one or more edges in the first digital representation of the patient's teeth. At least one of the one or more lists may specify one or more polygon faces in the digital representation of the patient's teeth. At least one of the one or more lists may specify one or more vertices (e.g., such as derived from a 3D mesh) in the first digital representation of the patient's teeth. At least one of the one or more lists may specify one or more points (e.g., such as derived from a 3D point cloud) in the first digital representation of the patient's teeth. In some cases, the 3D point cloud may include a plurality of vertices extracted from a 3D mesh. At least one of the one or more lists may specify one or more voxels (e.g., such as derived from a sparse representation) in the first digital representation of the patient's teeth. The method may calculate, by the one or more computer processors, one or more mesh element features. In terms of edges, the one or more mesh element features may include edge endpoints, edge curvature, edge normal vectors, edge movement vectors, edge normalized lengths, vertices, faces of the associated three-dimensional representation, voxels, and combinations thereof. Other mesh element features for edges are disclosed herein. Mesh element features for each of vertices, points, faces, and voxels are also disclosed herein. The method may generate, by the one or more computer processors, a digital representation predicting the position and orientation of the patient's teeth based on the one or more predicted tooth movements. The prediction of tooth movement may include a transformation (e.g., such as an affine transformation matrix, a translation vector, a quaternion, or one or more of the Euler angles). The setting prediction model may predict each of tooth position and tooth orientation information. In some non-limiting examples, the network may predict orientation and position information almost simultaneously. The setting prediction model may predict the setting transformation of each tooth in the dental arch to place each tooth in a final setting pose. The method may generate, by the one or more computer processors, a digital representation of the patient's teeth based on the one or more reference tooth movements. In some non-limiting embodiments, the generator of the setting prediction model may be trained at least in part with the assistance of a discriminator.The discriminator can determine whether the representation of the one or more tooth movements predicted by the generator can be distinguished from the representation of one or more reference tooth movements, and may include the following steps: receiving the representation of the one or more tooth movements predicted by the generator, the representation of the one or more reference tooth movements, and the first digital representation of the patient's teeth; comparing the representation of the one or more tooth movements predicted by the generator and the representation of the one or more reference tooth movements, wherein the comparison is at least partially based on the first digital representation of the patient's teeth; and determining, by the one or more computer processors, the probability that the representation of the one or more tooth movements predicted by the generator is the same as the representation of the one or more reference tooth movements.

[0007] In a second aspect, a second computer-implemented method for generating an orthodontic alignment treatment setup is related to intermediate hierarchical prediction. The intermediate hierarchy of teeth from a malocclusion stage to a final stage requires determining accurate individual tooth movements in such a way that the teeth do not conflict with each other, the teeth move towards their final state, and the teeth follow an optimal and preferably short trajectory. Since each tooth has six degrees of freedom and an average dental arch has approximately fourteen teeth, finding the optimal tooth trajectory from an initial stage to a final stage is a large and complex problem.

[0008] The second computer-implemented method is customized for the treatment needs of a patient (e.g., as specified by a clinician who may include a technician or a healthcare professional) and is described as including the following steps: receiving, by one or more computer processors, a first digital representation of the patient's teeth and a representation of a final setup; using, by one or more computer processors and for determining predictions of one or more tooth movements for one or more intermediate stages, a generator that is a machine learning model included in a setup prediction machine learning model, such as a neural network, such as including one or more neural networks (e.g., a 3D encoder, a 3D decoder, a multi-layer perceptron (MLP), an encoder-decoder structure, and other neural networks disclosed herein) and that has been initially trained to predict one or more tooth movements for one or more intermediate stages; further training, by one or more computer processors, the setup prediction model based on the use, wherein the training of the setup prediction model is modified by performing operations that include predicting, by the generator, one or more tooth movements for at least one intermediate stage based on the first digital representation of the patient's teeth; calculating a loss function that quantifies the difference between predicted tooth movements and reference tooth movements; and using the loss to modify the setup prediction model. The second aspect may also include one or more of the optional features described above with reference to the first aspect.

[0009] A 3D representation of a patient's teeth can be provided to a first ML module (e.g., a U-Net architecture, a pyramid encoder-decoder architecture, or a transformer architecture such as a 3D SWIN transformer), which can provide a latent representation of the patient's teeth (e.g., including hierarchical neural network features) to a second ML module (e.g., an MLP, a transformer, an encoder, or other architectures described herein). The second ML module may optionally include one or more coordinate normalization layers. During the execution of the first ML module, operations including mesh pooling, mesh unpooling, mesh convolution, or mesh deconvolution may be applied to the 3D representation of the patient's teeth. These operations can act in a manner invariant to at least one of rotation, scaling, or translation of the patient's teeth. The 3D representation can include any one of a 3D mesh, a 3D point cloud, a 3D surface, or a voxelized representation. The 3D representation of the teeth can include one or more mesh elements. Mesh element feature vectors can be calculated for the mesh elements in the 3D representation of the patient's dentition. These mesh element feature vectors can be provided to the first ML module to improve the accuracy of the latent representation generated by the first ML module. The mesh element feature vectors can include spatial mesh element features, structural mesh element features, or color-based mesh element features. An oral care variable value (e.g., related to customization of an orthodontic treatment for the patient) can be provided to either the first ML module or the second ML module.

[0010] The oral care variables that can be provided to the second ML module can include oral care protocol parameters or oral care metrics to customize the output. Doctor preferences can be provided to the second ML module to customize the output. In some embodiments, information about interproximal enamel reduction of one or more teeth can be provided to the second ML module. In some embodiments, case classification information can be provided to the second ML module. In some embodiments, information about anteroposterior displacement can be provided to the second ML module.

[0011] The second ML module can generate a setup transformation for the patient's teeth (e.g., which can describe tooth movement). The first ML module or the second ML module can be trained at least in part by a loss function that quantifies the difference between one or more predicted transformations and one or more corresponding ground truth transformations. A loss value can be calculated that quantifies the difference between a predicted setup and a predetermined ground truth setup. Example loss functions can calculate at least one pairwise distance between at least one aspect of a 3D tooth representation of a tooth pose as represented in a predicted setup and a corresponding aspect of the representation of the corresponding tooth in a reference pose in the ground truth setup. In some embodiments, the accuracy of loss calculation can be improved by registering the predicted setup with the corresponding ground truth setup.

[0012] In some specific implementations, the second ML module may include a generator that is at least partially trained by a discriminator. The orthodontic setup prediction method described herein may be used in combination with other digital oral care treatment methods for patient treatment (e.g., a machine learning model for predicting restorative tooth designs, or a machine learning model for generating at least one component or placing at least one component to generate an oral care appliance). The tooth transformation predicted by the second ML module may be used in the generation of an orthodontic appliance (e.g., a thermoformed or 3D printed clear tray aligner (CTA)). One or more binary flags may be provided to the second ML module to indicate that one or more teeth are fixed, pinned, bridged, extracted, implanted, or missing.

[0013] In some specific implementations, the first ML module and the second ML module may be trained using representation learning. In some specific implementations, the first ML module and the second ML module may be trained end-to-end. In some specific implementations, the first ML module and the second ML module may be trained using transfer learning (e.g., based on a neural network first trained on coordinate system prediction). In some specific implementations, the first ML module and the second ML module may be used as a basis for training a third neural network module using transfer learning (e.g., using the first ML module or the second ML module as an initial state to train a third neural network). In some specific implementations, the first ML module or the second ML module may include an attention mechanism (e.g., as used in a transformer).

[0014] In some specific implementations, the first ML module and / or the second ML module may be trained to use sparse processing (e.g., using voxels). In some cases, according to the techniques of the present disclosure, the training data may be augmented. In some specific implementations, tooth movement may be encoded using relative local tooth transformations. In some specific implementations, tooth movement may be encoded using absolute tooth transformations. In some specific implementations, one or more arch forms of the patient may be provided to the second ML module.

[0015] In some cases, a proximal stripping operation may be performed on at least one 3D representation of the patient's teeth before performing automatic setup prediction. In some specific implementations, information related to IPR may be provided to the second ML module to affect the generated setup transformation. The IPR cutting surface may be provided to the second ML module. The designation of which teeth should receive IPR (and optionally at which stage of orthodontic treatment) may be provided to the second ML module.

[0016] Tooth transformation can take the form of at least one of the following: transformation matrix, translation vector, quaternion, or at least Euler angles. In some specific implementations, the tooth transformation can apply a rotation to the patient's teeth, where the pivot point of the rotation is at one of the following locations: the centroid of the dental crown, the apex of the tooth root, the origin of the malocclusion transformation, or a point along the dental arch morphology close to the tooth.

[0017] In some cases, the training dataset can be filtered to remove abnormal cases, such as filtering based on at least one oral care metric. In some cases, the techniques of the present disclosure can be performed in a clinical environment, such as a clinic or a doctor's office. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A method for enhancing training data used in training a machine learning (ML) model of the present disclosure is shown.

[0019] Figure 2 An overview of some of the setup prediction methods described herein is shown.

[0020] Figure 3 A setup prediction method using a denoising diffusion probability model is shown.

[0021] Figure 4 A method for setup prediction for what is referred to as a similarity setup is shown.

[0022] Figure 5 A method for training a setup prediction model to generate setups customized for a patient's treatment needs is shown.

[0023] Figure 6 An example generator implementation for a setup prediction model is shown.

[0024] Figure 7 A method for generating orthodontic setup transformations using a U-Net and an encoder is shown.

[0025] Figure 8 A method for generating orthodontic setup transformations using a U-Net and a multi-layer perceptron (MLP) is shown.

[0026] Figure 9 A method for generating orthodontic setup transformations using a U-Net and a transformer encoder (or transformer decoder) is shown.

[0027] Figure 10 A transformer that can be configured to generate orthodontic setup transformations is shown. DETAILED DESCRIPTION

[0028] The techniques of the present disclosure can train an encoder-decoder structure (e.g., U-Net) to generate a transformation to place a 3D representation of oral care data (e.g., teeth, appliance components, fixture model components, etc.) into a pose suitable for oral care appliance generation (e.g., placing a patient's teeth into a set pose used in orthodontic treatment). The encoder-decoder structure can include at least one encoder or at least one decoder. Non-limiting examples of the encoder-decoder structure include 3D U-Net, transformers (e.g., 3D SWIN transformer), pyramid encoder-decoder, or autoencoder, etc.

[0029] Techniques for automatic prediction of setups are described herein that can provide the advantage of improved accuracy compared to the prior art, enabling novice clinicians to be trained in the generation of effective setups, enabling the generation of customized setups (e.g., conforming to clinician specifications), and providing technical improvements in enhancing data precision in the formulation of these setups.

[0030] The setting prediction model of the present disclosure can receive various input data. As described herein, the input data can include a dental mesh representing one or both dental arches of a patient. The dental data can be presented in the form of a 3D representation such as a 3D mesh, 3D point cloud, voxelized representation, etc. Such data can be preprocessed, for example, by arranging the constituent mesh elements into a list and calculating optional mesh element feature vectors for each mesh element. Such vectors can impart valuable information about the shape and / or structure of the teeth to the setting prediction neural network. Additional inputs can enable the setting prediction neural network to better understand the distribution of the provided data (e.g., the dental mesh), thereby providing a technical improvement that can be customized according to the specific medical / dental needs of the patient when deploying the setting prediction model. For example, one or more oral care metrics can be calculated. Oral care metrics can be used to measure one or more physical aspects of the setting (e.g., physical relationships within or between teeth). In some cases, orthodontic metrics can be calculated for a ground truth setting and then used for the training of a machine learning model (e.g., the setting prediction model). The metric values can be received at the input of the setting prediction model as a way to train the model to encode the distribution of such metrics over a number of examples in the training dataset. For example, an "overbiteleft" metric can be calculated for a setting (e.g., at least one of a malocclusion setting and an approved setting) received by the setting prediction model. During training, the network can then receive the metric value as an input to help train the network to link the provided metric value to the physical aspects of the received setting (e.g., to learn the distribution of the possible values of the metric over the examples in the training dataset). A metric can be calculated for a malocclusion setting, and the metric value, along with the malocclusion transformation and / or the dental mesh, is provided as an input to the network during training. Optionally or alternatively, a metric can be calculated for an approved setting, and the metric can be provided as an input to the network during training, along with the approved setting transformation and / or the dental mesh (e.g., for application during loss calculation time). Such loss calculation can quantify the difference between the prediction and the ground truth example (e.g., between the predicted setting and the ground truth setting). By providing the metric value to the network during training, the network can learn to encode the distribution of the metric through the process of loss calculation and subsequent backpropagation. The technical improvement provided by the setting prediction techniques described herein is to customize orthodontic treatment for patients. Oral care parameters can enable clinicians to customize specific desired aspects of the size, scale, and other physical aspects of the predicted setting. For example, in deployment, one or more oral care parameters (protocol parameters or prosthetic design parameters) can be defined and provided as part of the execution-phase input to the trained setting prediction model to specify one or more aspects of the expected setting during execution runtime.In some embodiments, protocol parameters corresponding to oral care metrics (e.g., such as the remaining overbite metric described above) may be defined, which may be received at the input of the deployed setting prediction neural network and considered as instructions to the setting prediction neural network to generate a setting with a specified number of metrics (e.g., remaining overbite). The setting prediction model may be particularly suitable for generating a setting with a specified value when the specified value of the protocol parameter falls within the distribution of the corresponding metric values that appear in the training dataset. Other protocol parameters corresponding to other orthodontic metrics may also be defined and provided as an indication to the setting prediction model for assigning a quantity of the relevant metric to the predicted setting. This interaction between oral care metrics and oral care parameters may also apply to the training and deployment of other prediction models in oral care.

[0031] To effectively train the setting prediction neural network, aspects of the present disclosure relate to forming training data having a distribution that describes the types of settings that the setting prediction neural network is configured to produce. For example, to produce a final setting with an overbite of approximately 2.0 mm, one approach is to use ground truth training data with an overbite of approximately 2.0 mm. This approach can yield a clean training signal and can produce useful results, and an alternative approach can enable the network to learn to account for differences in overbite among the various ground truth training samples in the training dataset. The overbite metric can be calculated for the malocclusion dental arches of the training samples (patient cases). This overbite value can be received at training as an input to the setting prediction neural network along with the malocclusion tooth data and used as a signal to the neural network regarding the magnitude of the overbite present in the malocclusion dental arch. The network thus learns that different cases have different overbite magnitudes and can encode the distribution of possible overbite magnitudes, which can then be assigned to the predicted setting. At deployment, the trained neural network can receive the malocclusion tooth data as an input and can also receive an input indicating the magnitude of the overbite desired in the predicted setting (e.g., or some other oral care metric) (e.g., in the form of a protocol parameter that has been defined for that purpose). This approach enables the setting prediction neural network to account for differences in the distribution of the training dataset without excluding patient cases from the training dataset (e.g., as might be done in the case of screening the training dataset), which has the beneficial effect of enabling the deployed setting prediction neural network to customize the predicted setting according to the specifications of the clinician using the setting prediction model. Other orthodontic metrics (e.g., those disclosed herein) may also be calculated according to this technique. Corresponding protocol parameters (e.g., those disclosed herein or those defined to correspond to a specific metric) may be provided to the trained network to enable customization of the output setting prediction. In addition to setting prediction, other techniques disclosed herein may also be trained by using oral care metrics and protocol parameters received as inputs to the prediction model.

[0032] The setting prediction neural network of the present disclosure can be trained at least in part by calculating one or more loss values (e.g., the reconstruction loss or other loss values described herein). Such loss values can quantify the difference between the predicted setting and the corresponding ground truth setting. In some cases, these settings can be registered with each other (e.g., using Iterative Closest Point (ICP) or Singular Value Decomposition (SVD)) before calculating the loss to reduce noise and improve the accuracy of the resulting trained setting prediction neural network. This registration can alternatively or additionally be performed between the malocclusion setting and the corresponding ground truth setting, with the advantage of reducing noise in the loss measurement and improving the accuracy of the trained network.

[0033] The setting prediction neural network can calculate the transformation of each tooth to move the tooth into a posture suitable for the end of orthodontic treatment (e.g., the final setting). The posture of a tooth can include a change in position in 3D space and can also include a change in orientation (e.g., relative to one or more coordinate axes, such as local coordinate axes with the origin at the centroid of the dental crown). The transformation can achieve the change in orientation by pivoting the tooth mesh relative to a pivot point or the tooth origin. The pivot point can be selected to be within the centroid of the dental crown. Alternative options include at the apex of the tooth root tip, the origin of the malocclusion transformation, or at a point along the dental arch form close to the tooth.

[0034] In some embodiments, the setting prediction neural network can be conditionally trained based on Interproximal Enamel Reduction (IPR) information. IPR can be applied to teeth to enable greater packing of the teeth in the final setting. The setting model can be trained to consider the amount of IPR (e.g., the millimeters offset from either or both the mesial and distal sides of the tooth) and / or the IPR cutting plane (which can be used in combination with mesh boolean operations to remove material from either or both the mesial and distal sides of the tooth). For example, the IPR cutting plane can be used to modify one or more tooth meshes of one or more patient cases used to train the setting prediction model. This step improves the accuracy of setting prediction model training by increasing data precision because material is removed from the teeth that could otherwise cause conflicts between teeth in the final setting (and cause noise in the training data). After the trained setting prediction model is deployed, IPR can be applied to a test patient case to modify the shape of the teeth before receiving the case as input to the setting prediction model. In some cases, IPR can be applied to one or more tooth meshes of a patient case before calculating orthodontic metrics.

[0035] During orthodontics, anteroposterior (AP) displacement may involve sagittal displacement of the mandible (lower dental arch), thereby moving the mandible forward or backward. The application of AP displacement can improve the class relationship of teeth. The class can describe the malocclusion of the patient. Possible classes include: Class 1, Class 2, or Class 3. Elastic members can assist in the displacement of the mandible. Such elastic members can be attached to hardware on the teeth, such as buttons. In some cases, the setup prediction model of the present disclosure can directly receive the AP displacement transformation as input, which can improve the data accuracy of the resulting model. In some cases, before receiving patient case data as input to the setup prediction model of the present disclosure, the AP displacement transformation can be first applied to the patient case data.

[0036] In some specific implementations, the prediction model of the present disclosure can obtain more accurate results by combining one or more of the following inputs: arch form information V, interproximal reduction (IPR) information U, tooth size information P, diastema information Q, latent capsule representation T of the oral care mesh, latent vector representation A of the oral care mesh, protocol parameter K (which can describe the expected treatment of the patient by the healthcare professional), doctor preference L (which can describe the typical protocol parameters selected by the doctor), flag M regarding the tooth status (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name / dental symbol R, oral care metric S (including at least one of oral care metrics and prosthetic design metrics).

[0037] In some cases, the system of the present disclosure can be deployed in a clinical environment (such as a dental or orthodontic clinic) for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians). Such a system deployed in a clinical environment can enable clinicians to process oral care data (such as tooth scans) in a clinical environment or in some cases in a "chairside" environment (when the patient is in the clinical environment). A non-limiting list of examples of techniques can include: segmentation, mesh cleaning, coordinate system prediction, CTA trim line generation, prosthetic design generation, appliance component generation or placement or assembly, generation of other oral care meshes, verification of oral care meshes, setup prediction, removal of hardware from tooth meshes, placement of hardware on teeth, estimation of missing values, clustering of oral care data, classification of oral care meshes, setup comparison, metric calculation, or metric visualization. In some cases, the execution of these techniques can enable patient data to be processed, analyzed, and used by clinicians in appliance creation before the patient leaves the clinical environment (which can facilitate treatment planning as feedback can be received from the patient during the treatment planning process).

[0038] The systems of the present disclosure can automate operations in digital orthodontics (e.g., setup prediction, hardware placement, setup comparison), digital dentistry (e.g., prosthetic design generation), or combinations thereof. Some techniques can be applied to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleaning, coordinate system prediction, oral care mesh validation, estimation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows, or denoising diffusion models), metric visualization, appliance component placement, or appliance component generation, etc. In some cases, the systems of the present disclosure can enable clinicians or technicians to process oral care data (such as scanned dental arches). In addition to segmentation, mesh cleaning, coordinate system prediction, or validation operations, the systems of the present disclosure can also implement orthodontic treatment planning, which may involve setup prediction as at least one operation. The systems of the present disclosure can also implement prosthetic design generation, in which one or more restored tooth designs are generated and processed during the creation of an oral care appliance. The systems of the present disclosure can implement either or both of orthodontic or dental treatment planning, or can implement automated steps in the generation of either or both of orthodontic or dental appliances. Some appliances can implement both dental and orthodontic treatment, while other appliances can implement one or the other.

[0039] Final setups, intermediate stages, or combinations or sequences thereof can be used in the design and manufacture of orthodontic appliances such as clear tray aligners (CTAs). Fixture models can be generated from setups (or stages) that can be 3D printed. Clear plastic trays can be thermoformed over such fixture models. The thermoformed trays are cut from the fixture models (e.g., by following the CTA trim lines) to complete the aligner trays. In some cases, digital fixture models can be used to form a 3D oral care representation of the aligner tray, which can then be directly 3D printed.

[0040] The techniques of the present disclosure may require training datasets of hundreds or thousands of cohort patient cases to ensure that neural networks can encode the distribution of patient cases that may be encountered in clinical treatment. Cohort patient cases can include a set of crown meshes, a set of root meshes, or data files (e.g., JSON files) including case attributes. Typical examples of cohort patient cases can include up to 32 crown meshes (e.g., each of which can include tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., each of which can include tens of thousands of vertices or tens of thousands of faces), multiple gingival meshes (e.g., each of which can include tens of thousands of vertices or tens of thousands of faces), or one or more JSON files (each of which can include tens of thousands of values (e.g., objects, arrays, strings, real values, boolean values, or null values)).

[0041] In some specific implementations, setting up the prediction model (e.g., a setting prediction model such as using U-Net) may include aspects derived from a denoising diffusion model (e.g., a neural network that can be trained to iteratively denoise one or more setting transformations (such as transformations initialized randomly or using Gaussian noise)). In some specific implementations, the setting prediction model (e.g., a setting prediction model such as using U-Net) may at least partially use one or more neural networks to generate setting transformations, where the one or more neural networks are trained to use a neural network that has been trained with continuous normalizing flow (e.g., a neural network that can be trained in one form and then inverted for use during inference).

[0042] Aspects of the present disclosure may provide a technical solution to the technical problem of predicting orthodontic settings (e.g., intermediate or final settings for generating an appliance tray) used in oral care appliance generation using a 3D representation of a patient's dentition. Specifically, by practicing the techniques disclosed herein, a computing system specifically adapted to perform setting transformation prediction for oral care appliance generation is improved. For example, aspects of the present disclosure improve the performance of a computing system having a 3D representation of a patient's dentition by reducing the consumption of computing resources. Specifically, aspects of the present disclosure reduce computing resource consumption by subsampling the 3D representation of the patient's dentition (e.g., reducing the count of mesh elements used to describe aspects of the patient's dentition) such that computing resources are not wasted unnecessarily due to processing excessive mesh elements. Additionally, subsampling the mesh does not reduce the overall prediction accuracy of the computing system (and may actually improve prediction because the input provided to the ML model after subsampling is a more accurate (or better) representation of the patient's dentition). For example, unimportant (and potentially accuracy-reducing) noise or other artifacts are removed. That is, aspects of the present disclosure provide a more efficient allocation of computing resources in a manner that improves the accuracy of the underlying system.

[0043] In addition, aspects of the present disclosure may need to be performed in a time-limited manner, such as when an oral care appliance must be generated for a patient immediately after an intraoral scan (e.g., when the patient is waiting in a clinician's office). Accordingly, aspects of the present disclosure must be rooted in underlying computer technologies for predicting setups for oral care appliance generation and cannot be performed by a human, even with the aid of pen and paper. For example, specific implementations of the present disclosure must be able to: 1) store thousands or millions of mesh elements of a patient's dentition in a manner that can be processed by a computer processor; 2) perform calculations on the thousands or millions of mesh elements, such as to quantify aspects of the shape and / or structure of individual teeth in a 3D representation of a patient's dentition; and 3) predict orthodontic setups to be used in oral care appliance generation (e.g., orthodontic setup transformations customized to a patient's treatment needs by providing oral care metrics or oral care parameters to a machine learning model), and do so during the course of a short clinic visit.

[0044] The present disclosure relates to digital oral care encompassing the fields of digital dentistry and digital orthodontics. The present disclosure generally describes methods for processing three-dimensional (3D) representations and / or associated transformations of oral care data. It should be understood that, without loss of generality, there are various types of 3D representations. One type of 3D representation is 3D geometry. A 3D representation can include, be one or more of, or be part of: a 3D polygon mesh, a 3D point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels for sparse processing), or a 3D representation described by mathematical equations. Although the term "mesh" is used frequently throughout the present disclosure, in some specific implementations, the term should be understood to be interchangeable with other types of 3D representations. A 3D representation can describe elements of the 3D geometry and / or 3D structure of an object.

[0045] Arches S1, S2, S3, and S4 all include exactly the same tooth meshes, which are transformed differently according to the following description. The first arch S1 includes a set of tooth meshes that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a malposed position and orientation. The second arch S2 includes the same set of tooth meshes from S1 that are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a reference true setup position and orientation. The third arch S3 includes the same meshes as S1 and S2, which are arranged (e.g., using a transformation) in their positions in the oral cavity where the teeth are in a predicted final setup pose (e.g., as predicted by one or more techniques of the present disclosure). S4 is a counterpart of S3 where the teeth are in a pose corresponding to one of several intermediate stages of orthodontic treatment with a clear tray appliance.

[0046] It should be understood that, without loss of generality, the techniques of the present disclosure applied to the final setting are also applicable to intermediate gradings in orthodontic treatment, specifically geometric deep learning (GDL) settings, reinforcement learning (RL) settings, variational autoencoder (VAE) settings, capsule settings, multi-layer perceptron (MLP) settings, diffusion settings, pose transfer (PT) settings, similarity settings, force directed graph (FDG) settings, transformer settings, setting comparison, or setting classification. The metric visualization aspect of the present disclosure can also be configured to visualize data from both the final setting and intermediate stages. The MLP setting, VAE setting, and capsule setting each fall within the scope of the autoencoder setting. Some specific implementations of the MLP setting can fall within the scope of the transformer setting. Figure 2 A non-limiting selection of models that can be trained for setting prediction is shown. A representation setting refers to any one of the MLP setting, VAE setting, capsule setting, and any other setting prediction machine learning model that uses an autoencoder to create a representation of at least one tooth.

[0047] Each of the setting prediction techniques of the present disclosure is applicable to the manufacture of clear tray appliances and / or indirectly bonded trays. The setting prediction techniques can also be applicable to other products that also involve the final tooth pose. The pose can include position (or location) and rotation (or orientation).

[0048] A 3D mesh is a data structure that can describe the geometry and / or shape of an object related to oral care, which object includes but is not limited to teeth, hardware elements, or the gingival tissue of a patient. The 3D mesh can include one or more mesh elements, such as vertices, edges, faces, and combinations thereof. In some specific implementations, the mesh elements can include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features can be calculated for these mesh elements and provided to the prediction model of the present disclosure, and the prediction model of the present disclosure provides a technical advantage of improving data accuracy in the form of more accurate predictions of the model output.

[0049] The dentition of a patient may include one or more 3D representations of the patient's teeth (e.g., and / or associated transducers), gums, and / or other oral anatomical structures. In some embodiments, an orthodontic metric (OM) may quantify the relative position and / or orientation of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. In some embodiments, a restorative design metric (RDM) may quantify at least one aspect of the structure and / or shape of a 3D representation of a tooth. In some embodiments, an orthodontic landmark (OL) may locate one or more points or other regions of interest structures on a 3D representation of a tooth. In some embodiments, the OL may be provided for the generation of an orthodontic or prosthetic appliance, such as a clear tray aligner or a prosthetic appliance. In some embodiments, a mesh element may include at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth represented by a 3D mesh, the mesh elements may at least include: vertices, edges, faces, and voxels. In some embodiments, mesh element features may quantify some aspects of the 3D representation that are proximal to or associated with one or more mesh elements, as described elsewhere in the present disclosure. In some embodiments, an orthodontic procedure parameter (OPP) may specify at least one value that defines at least one aspect of a patient's planned orthodontic treatment (e.g., specifying desired target properties of a final setup in a final setup prediction). In some embodiments, an orthodontist preference (ODP) may specify at least one typical value of the OPP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. In some embodiments, a restorative design parameter (RDP) may specify at least one value that defines at least one aspect of a patient's planned prosthetic treatment (e.g., specifying desired target properties of a tooth to be treated with a prosthetic appliance). In some embodiments, a doctor restorative design preference (DRDP) may specify at least one typical value of the RDP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. The 3D oral care representation may include, but is not limited to: 1) a set of mesh element labels that may be applied to 3D mesh elements of a tooth / gum / hardware / appliance mesh (or point cloud) during the process of mesh segmentation or mesh cleaning; 2) one or more 3D representations of teeth / gums / hardware / appliances whose shapes have been modified (e.g., trimmed, deformed, or filled) during the process of mesh segmentation or mesh cleaning; 3) one or more coordinate systems (e.g., describing one, two, three, or more coordinate axes) for a single tooth or a group of teeth (such as a full dental arch, e.g., the LDE coordinate system); 4) 3D representations of one or more teeth whose shapes have been modified or otherwise made suitable for use in prosthetics; 5) 3D representations of one or more prosthetic appliance components;6) One or more transformations to be applied to one or more of the following: placement of prosthetic appliance library components relative to one or more teeth, teeth to be placed for an orthodontic setting (final setting or intermediate stage), hardware elements to be placed relative to one or more teeth, etc.; 7) Orthodontic setting; 8) 3D representation of hardware elements (such as facebow, tongue crib, orthodontic attachments, buttons, hooks, occlusal ramps, etc.) placed relative to one or more teeth, etc.; 8) 3D representation of a bonding pad for a hardware element (which can be generated for a specific tooth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting the tooth through a Boolean operation); 9) 3D representation of a clear tray appliance (CTA); 10) Position or shape of the CTA trim line (e.g., described as a grid or polyline); 11) Arch form describing the contour or layout of the dental arch (e.g., described as a 3D polyline or 3D grid or surface), which can follow the incisal edge of one or more teeth, which can follow the facial aspect of one or more teeth, which in some specific implementations can correspond to a malocclusion arch and in other specific implementations corresponds to a final setting arch (the effect of malocclusion on the shape of the arch form can be reduced by smoothing or averaging the shape of the arch form), which can be described by one or more control points and / or splines; 12) 3D representation of a jig model (e.g., depiction of teeth and gums used in a thermoformed clear tray appliance, or depiction of teeth / gums / hardware used in a thermoformed indirect bonding tray); 13) One or more latent space vectors (or latent capsules) generated by the 3D encoder stage of a 3D autoencoder (e.g., a variational autoencoder trained for tooth reconstruction) that has been trained on the reconstruction of an oral care mesh; 14) One or more oral care metrics for one or more teeth (e.g., such as orthodontic metrics or prosthetic design generation metrics); 15) One or more landmarks (e.g., 3D points) that describe the shape and / or geometric properties of one or more teeth, other dentition structures, or hardware structures (e.g., to be used for orthodontic setting creation or prosthetic appliance component generation or placement); 16) 3D representation created by scanning (e.g., optical scanning, CT scanning, or MRI scanning) a 3D printed part (e.g., a scanned jig model) corresponding to one or more teeth / gums / hardware / appliances; 17) 3D printed appliance (optionally including local thickness, reinforcement rib geometry, tab positioning, etc.); 18) 3D representation of a patient's dentition captured by a clinician or healthcare practitioner at the chairside (e.g., in an environment where the 3D representation is verified at the chairside before the patient leaves the clinic, such that errors can be detected and rescan performed as needed); 19) Prosthetic tooth design (e.g., for veneers, crowns, bridges, or prosthetic appliances); 20) 3D representation of one or more teeth used in digital oral care processing; 21) Other 3D printed parts belonging to an oral care protocol or other fields; 22) IPR cutting surface;23) One or more orthodontic setting transformations associated with one or more IPR cutting surfaces; 24) (Digital) pontic design that can fill at least a portion of the space between teeth to create space for erupting teeth in an orthodontic setting and then emerge from the gums; or 25) Components of a jig model (e.g., including jig model components such as interdental bands, occlusal locks, occlusal ramps, interdental reinforcements, gingival ridges, torque points, power ridges, pontics, or pits, etc.).;

[0050] The techniques of the present disclosure can be advantageously combined. For example, a setting comparison tool can be used to compare the output of a GDL setting model with reference ground truth data, compare the output of an RL setting model with reference ground truth data, compare the output of a VAE setting model with reference ground truth data, and compare the output of an MLP setting model with reference ground truth data. By comparing each of these setting prediction models with reference ground truth data, it can be determined which model achieves the best performance on a certain dataset or within a given problem domain. Additionally, a metric visualization tool can enable a global view of the final settings and intermediate stages generated by one or more of the setting prediction models, with the advantage of being able to select the best setting prediction model. Moreover, the metric visualization tool enables the calculation of metrics with a global scope within a set of intermediate stages. In some specific implementations, these global metrics can be consumed as inputs to a neural network for predicting settings (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, etc.). The global metrics can also be provided to the FDG settings. In some specific implementations, local metrics from the present disclosure (i.e., local metrics are metrics that can be calculated for one stage or setting of a process rather than within several stages or settings) can be consumed by the neural networks herein for predicting settings, with the advantage of improving the prediction results. In some specific implementations, the metrics described in the present disclosure can be visualized using the metric visualization tool.

[0051] VAE and MAE models for mesh element tagging and mesh filling can be advantageously combined with a setup prediction neural network for mesh cleaning before or during the prediction process. In some specific implementations, the VAE for mesh element tagging can be used to tag mesh elements for further processing, such as metric calculation, removal, or modification. In some cases, such tagged mesh elements can be provided as input to the setup prediction neural network to inform the neural network of important mesh features, attributes, or geometries, with the advantage of improving the performance of the resulting setup prediction model. In some specific implementations, mesh filling can make the geometry of the teeth closer to complete, enabling the setup prediction model to function better (i.e., improving the correctness of the prediction due to the better-formed geometry). In some cases, a neural network for classifying setups (i.e., a setup classifier) can assist the setup prediction neural network in functioning because the setup classifier tells the setup prediction neural network when a predicted setup is acceptable for use and can be provided to methods for generating orthodontic trays. Setup classifiers (e.g., GDL setups, RL setups, VAE setups, capsule setups, MLP setups, diffusion setups, PT setups, similarity setups, and FDG setups, etc.) can help generate the final setup and also help generate intermediate stages. Additionally, the setup classifier neural network can be combined with a metric visualization tool. In other specific implementations, the setup classification neural network can be combined with a setup comparison tool (e.g., the setup comparison tool can output an indication of how a setup generated in part by the setup classifier compares to a setup generated by another setup prediction method). In some specific implementations, the VAE for mesh element tagging can identify one or more mesh elements used in metric calculation. The resulting metric output can be visualized by a metric visualization tool.

[0052] In some examples, the setup classifier neural network can assist the setup prediction techniques described in U.S. Patent Application No. US20210259808A1, the entire content of which is incorporated herein by reference, or PCT Application Publication No. WO2021245480A1, the entire content of which is incorporated herein by reference, or PCT Application No. PCT / IB2022 / 057373, the entire content of which is incorporated herein by reference. The setup classifier will help one or more of those techniques know when the predicted final setup is closest to being correct. In some cases, the setup classifier neural network can output an indication of how far a given setup is from the final setup (i.e., a progress indicator).

[0053] In some specific implementations, the latent space embedding vectors from the reconstructed VAE can be cascaded with the inputs of the setup prediction neural network described in WO2021245480A1. The latent space vectors can also be incorporated as inputs into other setup prediction models: GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup, etc. The advantage is to endow the neural network with reconstruction characteristics (e.g., the latent vector dimension of the dental mesh), thereby improving the generated setup prediction.

[0054] In some examples, the various setup prediction neural networks of the present disclosure can work together to generate the setups required for orthodontic treatment. For example, the GDL setup model can generate the final setup, and the RL setup model can use the final setup as an input to generate a series of intermediate stage setups. Alternatively, the VAE setup model (or the MLP setup model) can create the final setup, which can be used by the RL setup model to generate a series of intermediate stage setups. In some specific implementations, the setup prediction can be generated by one setup prediction neural network and then used as an input to another setup prediction neural network for further improvement and adjustment. In some specific implementations, such improvements can be performed in an iterative manner.

[0055] In some specific implementations, a setup verification model may be involved in this iterative setup prediction loop, such as the model disclosed in U.S. Provisional Application No. US63 / 366495. First, a setup can be generated (e.g., using a model trained for setup prediction, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, etc.), and then the setup is verified. If the setup passes the verification, the setup can be output for use. If the setup fails the verification, the setup can be sent back to one or more of the setup prediction models for correction, improvement, and / or adjustment. In some cases, the setup verification model can output an indication of what is wrong with the setup, such that the setup generation model can be improved in a next iteration. The process is iterated until completion.

[0056] Generally, in some specific implementations, two or more of the following techniques of the present disclosure can be combined during orthodontic and / or dental treatment: GDL setting, setting classification, reinforcement learning (RL) setting, setting comparison, autoencoder setting (VAE setting or capsule setting), VAE grid element labeling, masked autoencoder (MAE) grid filling, multi-layer perceptron (MLP) setting, metric visualization, estimation of missing oral care parameter values, tooth classification using latent vectors, FDG setting, pose transfer setting, prosthetic design metric calculation, neural network techniques for dental restoration and / or orthodontics (e.g., generation or modification of 3D oral care representation using transformers), landmark-based (LB) setting, diffusion setting, estimation of tooth movement protocol, capsule autoencoder segmentation, diffusion segmentation, similarity setting, validation of oral care representation (e.g., using autoencoders), coordinate system prediction, prosthetic design generation or generation or modification of 3D oral care representation using diffusion models.

[0057] Oral care parameters can include one or more values that specify protocol parameters or are otherwise related to orthodontic treatment. Oral care parameters can additionally or alternatively include one or more values that specify prosthetic design parameters or are otherwise related to digital dentistry or digital oral care.

[0058] Other types of oral care parameters that can be used include doctor preferences (which are used in orthodontic treatment). Another oral care parameter is referred to as doctor prosthetic preference and is related to digital dentistry. For example, when faced with similar diagnoses or treatment options, one clinician may prefer one value of the prosthetic design parameter (RDP), while another clinician may prefer a different value of the RDP. An example of such an RDP is the dental restoration style. In some specific implementations, protocol parameters and / or doctor preferences can be provided to a setting prediction model for orthodontic treatment for the purpose of improving the customization of the resulting orthodontic appliance. In some specific implementations, prosthetic design parameters and doctor prosthetic preferences can be used to design the tooth geometry used in creating a dental restoration appliance for the purpose of improving the customization of the appliance. In addition to oral care parameters, doctor preferences, and doctor prosthetic preferences, some specific implementations of the ML prediction model of the present disclosure in orthodontic treatment can also take settings (e.g., the arrangement of teeth) as input. In some such specific implementations, the ML prediction model of the present disclosure can take the final setting (i.e., the final arrangement of teeth) as input, such as in the case of a prediction model trained to generate predictions for intermediate stages. For simplicity, these preferences are referred to as doctor prosthetic preferences, but they are intended to be used in a non-limiting sense. Specifically, it should be understood that these preferences can be specified by any practitioner or other appropriate healthcare professional and are not intended to be limited to doctor preferences per se (i.e., preferences from a person with a Doctor of Medicine or equivalent degree).

[0059] Oral care professionals or clinicians, such as dentists or orthodontists, may specify information regarding patient handling in the form of a patient-specific set of protocol parameters. In some cases, the oral care professional may specify a set of general preferences (also referred to as doctor preferences) for a large number of cases to be used as default values during the process of specifying the set of protocol parameters. In some specific implementations, oral care parameters may be incorporated into the techniques described in this disclosure, such as one or more of GDL settings, VAE settings, RL settings, setting comparison, setting classification, VAE grid element tagging, MAE grid filling, verification using autoencoders, estimation of missing protocol parameter values, metric visualization, or FDG settings. One or more of these models may take as input one or more protocol parameter vectors K and / or one or more doctor preference vectors L. In some specific implementations, one or more of these models may introduce one or more protocol parameter vectors K and / or one or more doctor preference vectors L into the hidden layer of a neural network. In some specific implementations, one or more of these models may introduce either or both of K and L into a mathematical calculation, such as a force calculation, for the purpose of improving the calculation and the resulting final customization of the appliance for the patient.

[0060] Some specific implementations of neural networks for predicting settings, such as GDL settings, VAE settings, or RL settings, may incorporate information from an oral care professional (aka doctor). This information may affect the arrangement of teeth in the final setting, bringing the position and orientation of the teeth into compliance with the norms set by the doctor and within tolerances. In some specific implementations of the GDL setting model, oral care parameters may be provided directly to the generator network as a separate input along with the grid data. In some specific implementations of GDL settings, oral care parameters may be incorporated into the feature vectors calculated for each grid element prior to the grid elements being input to the generator for processing. Some specific implementations of the VAE setting model may incorporate oral care parameters into the setting prediction. In some specific implementations, the protocol parameter K and / or doctor preference information L may be concatenated with the latent space vector C. Doctor preferences (e.g., in an orthodontic context) and / or doctor restoration preferences may be indicated in a treatment form, or they may be based on characteristics in the treatment plan, such as final setting characteristics (e.g., the amount of bite correction or midline correction in the planned final setting), intermediate staging characteristics (e.g., treatment duration, tooth movement scenario, or overcorrection strategy), or outcomes (e.g., the number of corrections / improvements).

[0061] Orthodontic procedure parameters may specify one or more of the following (possible values are shown in {}). Some non-limiting classification values of some example OPPs are described below. In some specific implementations, real values may be specified for one or more of these OPPs. For example, the overbite OPP may specify the amount of overbite desired in the setting (e.g., in millimeters), and may be received as an input to a setting prediction model to provide information about the desired amount of overbite in the setting to the setting prediction model. Some specific implementations may specify numerical values for the overjet OPP or other OPPs. In some specific implementations, one or more OPPs corresponding to one or more orthodontic metrics (OMs) may be defined. In some cases, numerical values may be specified for such OPPs for the purpose of controlling the output of the setting prediction model. Teeth to be moved: {Anterior teeth only, Anterior teeth and canines, Full dental arch} Tooth movement restrictions: For each tooth, indicate whether the tooth is {Not moving, Missing, To be extracted, Native / Erupted, Debonded} Overbite: {Show resulting overbite after alignment, Maintain initial overbite, Calibrate bite opening, Correct deep bite} Overjet: {Show resulting overjet after alignment, Maintain initial overjet, Improve resulting overjet} Anterior / Posterior (AP) relationship Retention: {Right, Left, Both} Improve canine relationship only: {Right, Left, Both} Increase canine and / or molar relationship to 4mm: {Right, Left, Both} Correction of Class I (canine and molar): {Right, Left, Both} Interdigitation (if present) Anterior teeth: {Not corrected, Corrected, Not applicable} Posterior teeth: {Not corrected, Corrected, Not applicable} Correction of Class I (canine and molar): {Right, Left, Both} Correction with posterior tooth IPR: {Yes, No} Class II / III correction simulation (requires elastics): {Yes, No} Sequential distal movement (recommended with elastics): {Yes, No} Does it include incisions for elastics? : {Yes, No} Preferred incisions for elastics: {Use button incisions on molars and hooks on canines, Use only button incisions, Use only hooks} Stage to start cutting elastics: [Integer] Levelling of upper anterior teeth: {0.5mm shorter at sides than central, Level incisal edge, Level gingival edge, As indicated} Spacing: {Close all spaces, popularize specific spaces} Preferred midline position: {Set the upper midline to an ideal value and match the upper and lower parts to each other} Resolve upper crowding by expansion: {Mainly, as needed, none} Resolve upper crowding by proclination: {Mainly, as needed, none} Resolve upper crowding by IPR - anterior teeth: {Mainly, as needed, none} Resolve upper crowding by IPR - right posterior teeth: {Mainly, as needed, none} Resolve upper crowding by IPR - left posterior teeth: {Mainly, as needed, none} Resolve lower crowding by expansion: {Mainly, as needed, none} Resolve lower crowding by proclination: {Mainly, as needed, none} Resolve lower crowding by IPR - anterior teeth: {Mainly, as needed, none} Resolve lower crowding by IPR - right posterior teeth: {Mainly, as needed, none} Resolve lower crowding by IPR - left posterior teeth: {Mainly, as needed, none} Modify the arch form: {The patient's natural arch form, as indicated} [The doctor can specify the arch form - select from a set of options or a custom design]

[0062] Other orthodontic protocol parameters can be defined, such as those that can be used to place standardized brackets at a specified occlusal height on the teeth. In some specific implementations, one or more orthodontic protocol parameters can be defined to specify at least one of the second - order and third - order rotation angles (i.e., angle and torque, respectively) to be applied to the teeth, which can achieve a target setting arrangement where the crown landmarks, for example, are located within a threshold distance of a common occlusal plane. In some specific implementations, one or more orthodontic protocol parameters can be defined to specify a position in global coordinates at which at least one landmark (e.g., centroid) of the crown (or root) will be placed in the tooth setting arrangement. Generally speaking, oral care parameters corresponding to oral care metrics can be defined. For example, orthodontic protocol parameters corresponding to orthodontic metrics can be defined (e.g., to specify the amount of a specific metric that is expected to appear in the predicted setting at the input of a setting prediction model).

[0063] Doctor preferences may differ from orthodontic protocol parameters, as doctor preferences are related to oral care providers and may include the mean, mode, median, minimum, or maximum (or some other statistical value) of past settings associated with the oral care provider's treatment decisions for past orthodontic cases. On the other hand, protocol parameters may be related to a specific patient and describe the needs of the specific patient's treatment. Doctor preferences may be related to the doctor and the doctor's past treatment practices, while protocol parameters may be related to the treatment of a specific patient. Doctor preferences (or "treatment preferences") may specify one or more of the following (possible values are shown in {}).

[0064] Orthodontist preferences may specify one or more of the following (other possible values are found elsewhere in this disclosure). Deep bite case (bite correction amount) - final overbite: [real value, in millimeters, e.g., 0.5mm] Option - Intrude upper anterior teeth: {Yes, No} Option - Include vertical overcorrection of lower canines: {Yes, No} Midline correction in the planned final setting: {Maintain initial midline, Improve midline with IPR, As indicated} Deep bite case - Velocity reversal curve: {Yes, No} Anterior bite opening case - final overbite: [real value, in millimeters, e.g., 2mm] Is arch expansion a priority for your case?: {Yes, No} If yes, specify the acceptable expansion per quadrant in mm. When upper molars are expanded, apply buccal root torque: {Yes, No} Is IPR of the first Tx design acceptable?: {Yes, No} Maximum IPR per contact: Upper anterior teeth: [Specify in mm] Lower anterior teeth: [Expressed in mm] Upper and lower anterior teeth: [Specify in mm] Is asymmetric IPR acceptable?: {Yes, No} Final tooth position (overcorrection strategy): {Ideal, Overcorrection} Root movement: {Move roots as needed to achieve treatment goals, Limit posterior root movement, Limit all root movement} Final occlusal contact: {Balance all contacts as much as possible, No occlusal contact on upper incisors, End with heavy posterior contacts, Other} For class correction, is asymmetric AP shift acceptable?: {Yes, No, Other} Treatment duration: [Phase count] Tooth movement plan: {Plan_A, Plan_B, Plan_C}

[0065] The prior art has attempted to move teeth towards the dental arch form V after setting predictions have been presented through other workflow components, which introduces errors into the resulting settings and degrades performance due to the purpose of the workflow components performing the setting predictions. The present disclosure provides several improvements over these prior arts by enabling dental arch form information to be directly introduced into the neural network as an input to the setting prediction neural network, and the technical improvement is to provide a setting prediction that more accurately meets the orthodontic treatment needs of the patient (thereby improving data accuracy). The dental arch form information V can be provided as an input to any of the GDL setting, RL setting, VAE setting, capsule setting, MLP setting, and diffusion setting prediction neural networks. In some specific implementations, the dental arch form information V can be directly provided to one or more internal neural network layers in one or more of those setting applications.

[0066] Additional protocol parameters can include a textual description of the patient's medical condition and the expected treatment. Such textual descriptions can be analyzed via natural language processing operations, including tokenization, stop word removal, stemming, n-gram formation, text data vectorization, bag-of-words analysis, term frequency-inverse document frequency (TF-IDF) analysis, sentiment analysis, naive Bayes classification, and / or logistic regression classification. The output of such analysis techniques can be used as an input to one or more of the neural networks of the present disclosure, and the advantage is to customize and improve the prediction output (e.g., predicted settings or predicted grid geometries).

[0067] These additional orthodontic parameters and doctor preferences can also be incorporated into the neural networks of the present disclosure, which have the advantages related to data accuracy and efficiency of improving the customization of those neural networks and enabling those neural networks to predict outputs that better match the treatment needs of individual patients.

[0068] In some specific implementations, the dataset used to train one or more of the neural network models of the present disclosure can be conditionally filtered according to one or more of the orthodontic protocol parameters described in this section. In some cases, patient cases that exhibit outliers of one or more of these protocol parameters can be omitted from the dataset (alternatively used to form the dataset) used to train one or more of the neural networks of the present disclosure.

[0069] During training, one or more protocol parameters and / or doctor preferences can be provided to the neural network. In this way, the neural network can be conditioned on one or more protocol parameters and / or doctor preferences. Examples of such neural networks include conditional generative adversarial networks (cGANs) and / or conditional variational autoencoders (cVAEs), either of which can be used in various neural network-based applications of the present disclosure.

[0070] In some cases, an input based on tooth shape can be provided to a neural network for setting prediction. In other cases, non-shape-based inputs, such as tooth names or nomenclature, can be used, as it relates to dental notation. In some embodiments, a vector R of flags can be provided to the neural network, where a "1" value indicates the presence of a tooth and a "0" value indicates the absence of a tooth in the patient case (although other values are possible). The vector R can include one-hot vectors, where each element in the vector corresponds to a tooth type, name, or nomenclature. Identification information about the tooth (e.g., the name of the tooth) can be provided to the prediction neural network of the present disclosure, which has the advantage of enabling the neural network to be trained to handle different teeth in a tooth-specific manner. For example, the setting prediction model can learn to predict setting transformations for specific tooth names (e.g., the upper right central incisor or the lower left canine, etc.). In the case of a mesh cleaning autoencoder (for labeling mesh elements or for filling in missing mesh data), the autoencoder can be trained in this way to provide specialized processing to a tooth based on the tooth's nomenclature. In the case of a setting classification neural network, a list of tooth names present in the patient's dental arch can better enable the neural network to output an accurate determination of the setting classification, as tooth nomenclature is a valuable input for training such a neural network. For example, tooth naming / names can be defined according to a universal numbering system, the Palmer quadrant system, or the FDI World Dental Federation notation (ISO 3950).

[0071] In one example, in the case where all teeth except (up to four) wisdom teeth are present, the vector R can be defined as an optional input to the setting prediction neural network of the present disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth, and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, UL1, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7.

[0072] In some cases, the position of the cusp can be provided to the neural network for setting prediction. In other cases, one or more vectors S of orthodontic metrics described elsewhere in the present disclosure can be provided to the neural network for setting prediction. The advantage is that the network's ability to be trained to understand the state of the malocclusion setting is improved, and thus it can predict a more accurate final setting or intermediate stage.

[0073] In some specific embodiments, the neural network can take one or more indications of interproximal reduction (IPR) U as inputs, so as to indicate the amount of enamel to be removed from a tooth (from the mesial or from the distal) during a process orthodontic treatment. In some specific embodiments, IPR information (e.g., the amount of IPR to be performed on one or more teeth, measured in millimeters, or one or more binary flags indicating whether IPR is to be performed on each tooth identified by a marker) can be concatenated with the latent vector A generated by a VAE or a latent capsule autoencoder. The vector and / or capsule generated by such concatenation can be provided to one or more of the neural networks of the present disclosure, which has the technical improvement or additional advantage of enabling the prediction neural network to consider IPR. IPR is particularly relevant to the setting prediction method, which can determine the position and pose of teeth at the end of the treatment or during one or more stages during the treatment. It is very important to consider the amount of enamel to be removed before the predicted tooth movement.

[0074] In some specific embodiments, one or more protocol parameters K and / or doctor preference vectors L can be introduced into the setting prediction model. In some specific embodiments, one or more optional vectors or values include: tooth position N (e.g., XYZ coordinates in local or global coordinates of the tooth), tooth orientation O (e.g., pose, such as in a transformation matrix or quaternion, Euler angles or other forms described herein), tooth size P (e.g., length, width, height, perimeter, radius, diagonal measurement, volume, and any size can be normalized compared to one or more other teeth), distance Q between adjacent teeth. In some cases, these "tooth sizes P" can be used to describe the expected size of the tooth for dental restoration design generation.

[0075] In some specific embodiments, tooth size P such as length, width, height or perimeter can be measured in a plane, such as a plane intersecting the centroid of the tooth, or a plane intersecting a center point located at the midpoint between the centroid of the tooth and the most incisal range or the most gingival range. The tooth height dimension can be measured as the distance from the gum to the incisal edge. The tooth width dimension can be measured as the distance from the mesial range to the distal range of the tooth. In some specific embodiments, the roundness or circularity of the tooth cross-section can be measured and included in the vector P. The roundness or circularity can be defined as the ratio of the radii of the inscribed circle and the circumscribed circle.

[0076] The distance Q between adjacent teeth can be implemented in different ways (and calculated using different distance definitions, such as Euclidean or geodesic). In some specific implementations, the distance Q1 can be measured as the average distance between the mesh elements of two adjacent teeth. In some specific implementations, the distance Q2 can be measured as the distance between the centers or centroids of two adjacent teeth. In some specific implementations, the distance Q3 can be measured between the closest mesh elements between two adjacent teeth. In some specific implementations, the distance Q4 can be measured between the tooth tips of two adjacent teeth. In some specific implementations, teeth can be considered adjacent within an arch. In some specific implementations, teeth can also be considered adjacent between opposing arches. In some specific implementations, any one of Q1, Q2, Q3, and Q4 can be divided by a term to normalize the resulting value of Q. In some specific implementations, the normalization term can involve one or more of the following: the volume of the tooth, the count of mesh elements in the tooth, the surface area of the tooth, the cross-sectional area of the tooth (e.g., as projected onto the XY plane), or some other term related to the tooth size.

[0077] Other information regarding the patient's dentition or treatment needs (or related parameters) can be concatenated with other input vectors into one or more of an MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and / or any neural network model listed elsewhere in this disclosure.

[0078] The vector M may include markers applied to one or more teeth. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is pinned. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is fixed. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is a pontic. M may also include at least one marker for each tooth to indicate whether the tooth has been extracted or implanted. Other and additional markers are also possible for teeth, such as combinations of fixed, pinned, and pontic markers. A marker set to a value indicating that a tooth should be fixed is a signal that the tooth should not move during processing and is sent to the network. In some embodiments, the neural network loss function can be designed to penalize (and in some cases, severely penalize) any movement in the indicated teeth. A marker indicating that a tooth is a pontic notifies the network to maintain the diastema, although movement of the gap is allowed. In some cases, M may include a marker indicating tooth loss. In some embodiments, the presence of one or more fixed teeth in the dental arch can assist in setting the prediction because the one or more fixed teeth can provide an anchor for the posture of the other teeth in the dental arch (i.e., can provide a fixed reference for the posture transformation of one or more other teeth in the dental arch). In some embodiments, one or more teeth may be intentionally fixed in order to provide an anchor to which other teeth can be positioned. In some embodiments, a 3D representation (such as a mesh) corresponding to the gingiva can be introduced to provide a reference point according to which the teeth can move.

[0079] Without loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U, and V described elsewhere in this disclosure may also be provided as inputs to one or more of the prediction models of this disclosure or fed into their intermediate layers. Specifically, these optional vectors can be provided to the MLP setup, GDL setup, RL setup, VAE setup, capsule setup, and / or diffusion setup, with the advantage of enabling the corresponding models to generate setups that better meet the orthodontic treatment needs of the patient. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent vectors A that are also provided to one or more of the prediction models of this disclosure. In some embodiments, such inputs can be introduced, for example, by concatenating with one or more latent capsules T that are also provided to one or more of the prediction models of this disclosure.

[0080] In some embodiments, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into a neural network (such as an MLP or a transformer) in the hidden layer of the network. In some cases, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into the internal processing of an encoder structure.

[0081] In some specific implementations, a set of prediction models (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, PT settings, similarity settings, and diffusion settings) can take as input one or more latent vectors A corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a set of prediction models (such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, and diffusion settings) can take as input one or more latent capsules T corresponding to one or more input oral care meshes (e.g., such as tooth meshes). In some specific implementations, a set of prediction methods can take both A and T as input.

[0082] Some embodiments of the disclosed setup prediction neural networks (e.g., GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, or FDG setup or other setup prediction network architectures) may employ additional inputs to assist in setup prediction. Some of these inputs may reflect the geometric properties of one or more teeth or the entire dental arch. In some embodiments, the dental arch form or the dental arch curve may be provided to the setup prediction neural network, and the technical improvement is to help the setup prediction neural network find a suitable set of final setup poses for the teeth in a patient case (the technical improvement involves reducing the resource occupancy of both through more effective positioning capabilities and / or data precision for positioning more relevant final setup forms). The dental arch form or the dental arch curve may be encoded as a spline, B-spline, non-uniform rational B-spline (NURBS), polynomial spline, non-polynomial spline, parabola, hyperbola, or other parametric curve. Such a curve may be calculated as the average of multiple exemplars (such as exemplary final setups). Another non-limiting example of the dental arch form is the beta curve. In the case of the VAE setup, the dental arch information may be provided to the encoder E2 together with E and D as additional inputs. In the case of the GDL setup neural network, the dental arch information may be provided to the generator as an additional input of a list of mesh elements and associated network mesh element feature vectors. In some embodiments, the dental arch form may be described by one or more 3D representations (such as a 3D mesh, a set of 3D control points, and / or as a 3D polyline). In some embodiments, a Frenet frame may be overlaid on the dental arch form. The Frenet frame may locally describe the coordinate system corresponding to each point along the dental arch form. In some embodiments, such a coordinate system may be a right-handed coordinate system (or alternatively, in other embodiments, a left-handed coordinate system). In some embodiments, such a coordinate system may be determined at least in part by at least one of the tangent to the dental arch form at the point and the curvature of the dental arch form. In some embodiments, the point may be described using the LDE coordinate system relative to the dental arch form, where L, D, and E respectively correspond to: 1) the length of the curve along the dental arch form, 2) the distance from the dental arch form, and 3) the distance in a direction perpendicular to the L-axis and the D-axis (which may be referred to as the Eminence). Other geometric inputs may also assist in training the setup prediction neural network. In some embodiments, case classification information (e.g., class 1, class 2, or class 3), information about the AP shift, or information about the IPR may be received as inputs to the setup prediction model. Some definitions of case classification may describe the relationship between the cusp of the maxillary canine and one or more teeth of the mandibular arch. For example, a class 1 case may include the cusp of the maxillary canine that occludes between the corresponding mandibular canine and the first premolar. A class 2 case may include the cusp of the maxillary canine that occludes in front of the embrasure between the corresponding mandibular canine and the first premolar.Class 3 cases may include the cusp of the upper canine tooth that occludes posterior to the embrasure between the lower canine and the first premolar.

[0083] A variety of loss calculation techniques generally apply to the techniques of the present disclosure (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, setting classification, tooth classification, VAE grid element labeling, MAE grid filling, and estimation of protocol parameters).

[0084] These losses include L1 loss, L2 loss, mean squared error (MSE) loss, cross-entropy loss, etc. The losses can be calculated and used to train neural networks such as multi-layer perceptrons (MLPs), U-Net architectures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer architectures, etc. For example, in the learning of sequences, some embodiments may use triplet loss or contrastive loss.

[0085] Losses can also be used to train encoder architectures and decoder architectures. KL divergence loss can be at least partially used to train one or more neural networks of the present disclosure, such as a grid reconstruction autoencoder or a generator in a GDL setting, which has the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior can enable the reconstruction autoencoder to produce better reconstructions (e.g., when modifying the latent vector representation and using the decoder to reconstruct the modified latent vector, the resulting reconstruction is more likely to be a valid instance of the provided representation). There are other techniques for calculating losses that may be described elsewhere in the present disclosure. Such losses can be based on quantifying the difference between two or more 3D representations.

[0086] MSE loss calculation can involve the calculation of the mean squared distance between two sets, vectors, or data sets. MSE can generally be minimized. MSE can be applied to regression problems where the predictions generated by a neural network or other machine learning model can be real numbers. In some embodiments, the neural network can be equipped with one or more linear activation units on the output to generate MSE predictions. According to the techniques of the present disclosure, mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used.

[0087] In some specific implementations, cross - entropy can be used to quantify the difference between two or more distributions. In some specific implementations, cross - entropy loss can be used to train the neural networks of the present disclosure. In some specific implementations, cross - entropy loss can involve comparing predicted probabilities with ground - truth probabilities. Other names for cross - entropy loss include "log loss", "logistic loss", and "log loss". A small cross - entropy loss can indicate a better (e.g., more accurate) model. Cross - entropy loss can be logarithmic. In some specific implementations, cross - entropy loss can be applied to binary classification problems. In some specific implementations, a neural network can be equipped with a sigmoid activation unit at the output to generate probability predictions. In the case of multi - class classification, cross - entropy can also be used. In this case, in some specific implementations, a neural network trained to make multi - class predictions can be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for each class to be predicted). Other loss - calculation techniques that can be applied in the training of the neural networks of the present disclosure include one or more of the following: Huber loss, hinge loss, classification hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss - calculation methods are described herein and can be applied to the training of any neural network described in the present disclosure.

[0088] In some specific implementations, one or more neural networks of the present disclosure can be trained, at least in part, by a loss based on at least one of the following: point - wise mesh Euclidean distance (PMD) and Earth Mover's Distance (EMD). Some specific implementations can incorporate Hausdorff distance (HD) calculation into the loss calculation. Calculating the Hausdorff distance between two or more 3D representations, such as 3D meshes, can provide one or more technical improvements because HD not only considers the distance between two meshes but also the way those meshes are oriented and the relationship between the mesh shapes in those orientations (or positions or poses). The Hausdorff distance can improve the comparison of two or more tooth meshes, such as two or more instances of tooth meshes in different poses (e.g., comparison of a predicted setting with a ground - truth setting, which can be performed during the process of calculating the loss value for training a setting - prediction neural network).

[0089] The reconstruction loss can compare the predicted output with the ground truth (or reference) output. The systems of the present disclosure can calculate the reconstruction loss as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_loss = 0.5 * L1(all_points_target, all_points_predicted) + 0.5 * MSE(all_points_target, all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the ground truth data (e.g., a ground truth tooth restoration design, or a ground truth example of some other 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the generated or predicted data (e.g., a generated tooth restoration design, or a generated example of some other 3D type of oral care representation). Other specific implementations of the reconstruction loss can additionally (or alternatively) involve the L2 loss, the mean absolute error (MAE) loss, or a Huber loss term.

[0090] The entire content of the following paper is incorporated herein by reference in its entirety: “Attention Is All You Need”; Ashish Vaswani, Noam Shazeer, Niki Parmar, Niki Parmar, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin; NIPS 2017. The neural network-based models of the present disclosure can provide additional advantages in specific implementations where they are integrated with a neural network architecture known as a “transformer”. Figure 10 An example implementation of the transformer architecture is shown.

[0091] Prior to recently developed models such as transformer models, RNN-type models represented the state of the art for natural language processing (NLP). An example application of NLP is to generate new text based on previous words or text. Due to the important property of the transformer model with multi-head attention characteristics, the transformer then provides a significant improvement over GRU, LSTM and other such RNN-based NLP techniques. In some specific implementations, the NLP concept of multi-head attention can describe the relationship between each word in a sentence (or paragraph or document or document corpus) and each other word in the sentence (or paragraph or document or document corpus). These relationships can be generated by a multi-head attention module and can be encoded in vector form. The vector can describe how each word in a sentence (or paragraph or document or document corpus) should pay attention to each other word in the sentence (or paragraph or document or document corpus). RNN, LSTM and GRU models process sequences, such as sentences, one word at a time from the beginning to the end of the sequence. In addition, the model can only consider a given subset of sentences (called a window) when making predictions. However, in some cases, transformer-based models can take into account the entire previous text by processing the sequence as a whole in a single step. Transformers, RNNs, LSTMs, and GRU models can all be adapted for use in predictive models in digital dentistry and digital orthodontics, particularly in setting prediction tasks. In some implementations, an exemplary transformer model for use with 3D meshes and 3D transforms in setting predictions (or other oral care techniques) can be adapted based on a bidirectional encoder representation (BERT) from a transformer and / or a generative pre-trained (GPT) model. For example, a GPT (or BERT) model can first be trained on other data, such as text or document data, and then used in transfer learning. This transfer learning process can receive a previously trained GPT or BERT model and then further train it using data including a 3D oral care representation. Such transfer learning can be performed to train oral care models, such as: segmentation, mesh cleaning, coordinate system prediction, setting prediction, validation of 3D oral care representations, transformation prediction for placement of oral care meshes (e.g., teeth, hardware, appliance components, fixture model components), dental restoration design generation (or other 3D oral care representation generation, such as appliance components, fixture models, or dental arch morphology), classification of 3D oral care representations, estimation of missing oral care parameters, clustering of clinicians or clustering of clinician preferences, etc.

[0092] Oral care data may include one or more of the following (or combinations thereof): a 3D representation of teeth (e.g., a mesh, point cloud, or voxel), a portion of a tooth mesh (such as a subset of mesh elements), a tooth transformation (such as in the form of a matrix, vector, and / or quaternion, or combinations thereof), a transformation for an appliance component, a transformation for a fixture model component, and a mesh coordinate system definition (such as represented by a transformation, e.g., a transformation matrix) and / or other 3D oral care representations described herein.

[0093] A transducer can be trained to generate a transformation to position teeth into a set pose (or place an appliance component used in appliance generation or a fixture model component used in fixture model generation). Some embodiments can operate in an offline prediction environment, and some embodiments operate in an online reinforcement learning (RL) environment. In some embodiments, the transducer can initially be trained in an offline environment and then undergo further fine-tuning training in an online environment. In the offline prediction environment, the transducer can be trained from a dataset of queued patient case data. In the online RL environment, the transducer can be trained from, for example, a physical model or a CAD model. The transducer can learn from static data, such as a transformation (e.g., a trajectory transducer). In some embodiments, the transformation can provide a mapping from a malocclusion to a set (e.g., receive a transformation matrix as input and generate a transformation matrix as output). Some embodiments of the transducer can be trained to process 3D representations, such as 3D meshes, 3D point clouds, or voxels (e.g., using a decision transducer), taking a geometry (e.g., a mesh, point cloud, voxel, etc.) as input and outputting a transformation. The decision transducer can be coupled to a representation generation module that encodes a representation of a patient's dentition (e.g., teeth), such as a VAE, U-Net, encoder, transducer encoder, pyramid encoder-decoder, or a simple dense or fully connected network, or combinations thereof. In some embodiments, the representation generation module (e.g., a VAE, U-Net, encoder, pyramid encoder-decoder, or dense network for generating a tooth representation) can be trained to generate a representation on one or more teeth. The representation generation module can be trained on all teeth in both dental arches, only teeth within the same dental arch (upper or lower), only anterior teeth, only posterior teeth, or some other subset of teeth. In some embodiments, such a model can be trained on each individual tooth (e.g., the right upper canine), such that the model is trained or otherwise configured to generate a highly accurate representation of an individual tooth. In some embodiments, an encoder structure can encode such a representation. In some embodiments, the decision transducer can learn in an online environment, an offline environment, or both. The online decision transducer can be trained (e.g., using RL techniques) to output actions, states, and / or rewards. In some embodiments, the transformation can be discretized to allow for segmented or stepwise actions.

[0094] In some embodiments, the transformer can be trained to handle the embedding of the dental arch (i.e., predict the transformation of multiple teeth simultaneously) for prediction settings. In some embodiments, the embedding of single teeth can be cascaded into a sequence and then input to the transformer. The VAE can be trained to perform this embedding operation, the U-Net can be trained to perform such embedding, or a simple dense or fully connected network can be trained, or a combination of these operations can be employed. In some embodiments, the transformer-based techniques of the present disclosure can predict the movement of single teeth or the movement of multiple teeth (e.g., predict the transformation of each of multiple teeth).

[0095] The 3D mesh transformer can include a transformer encoder structure (which can encode oral care data) and can be followed by a transformer decoder structure. The 3D mesh transformer encoder can encode the oral care data into a latent representation, which can be combined with attention information (e.g., by concatenating the vector of attention information to the latent representation). In some embodiments, the attention information can help the decoder focus on relevant oral care data during the decoding process (e.g., focus on tooth order or mesh element connectivity), such that the transformer decoder can generate a useful output for the 3D mesh transformer (e.g., an output that can be used in generating an oral care appliance). Either or both of the transformer encoder or the transformer decoder can generate the latent representation. The decoder can be used to reconstruct the output of the transformer decoder (or the transformer encoder) into, for example, one or more tooth transformations for settings, one or more mesh element labels for segmentation, a coordinate system transformation used in coordinate system generation, or one or more points of a point cloud or voxel or other mesh element for another 3D representation. The transformer can include one or more of the following modules: a multi-head attention module, a feed-forward module, a normalization module, a linear module, and a softmax module, as well as a convolutional model for latent vector compression and / or representation.

[0096] The encoder can be stacked one or more times to further encode the oral care data and enable learning of different representations of the oral care data (e.g., different latent representations). These representations can be embedded with attention information (which can affect the decoder's focus on relevant parts of the latent representation of the oral care data) and fed to the decoder in a sequential form (e.g., as a concatenation of latent representations such as latent vectors). In some embodiments, the encoded output of the encoder (e.g., the latent representation) can be used by downstream processing steps in the generation of the oral care appliance. For example, the generated latent representation can be reconstructed into a transformation (e.g., for placing teeth in a setting or placing appliance components or fixture model components) or can be reconstructed into a 3D representation (e.g., a 3D point cloud, a 3D mesh, or other representations disclosed herein). In other words, the latent representation generated by the transducer (e.g., including sequentially encoded attention information) can be provided to a decoder that has been configured to reconstruct the latent representation into a specific data structure required for a particular domain region. Sequentially encoded attention information can include attention information that has undergone processing by multiple multi-head attention modules within the transducer encoder or the transducer decoder, to name just one example. Additionally, data from a specific domain can be used to compute a loss for that domain. The loss computation can train the transducer decoder to accurately reconstruct the latent representation into an output data structure related to the specific domain.

[0097] For example, when the decoder generates a transformation for an orthodontic setting, the decoder can be configured with an output describing, for example, 16 real values including a 4×4 transformation matrix (other data structures for describing the transformation are possible). In other words, the latent output generated by the transducer encoder (or the transducer decoder) can be used to predict the set tooth transformation for one or more teeth to place those teeth in a set position (e.g., a final setting or an intermediate stage). Such a transducer encoder (or transducer decoder) can be trained at least in part using a reconstruction loss (or a representation loss, and other losses described herein) function that can compare the predicted transformation to a ground truth (or reference) transformation.

[0098] In another example, when the decoder generates a transformation for a tooth coordinate system, the decoder can be configured with an output describing, for example, 16 real values including a 4×4 transformation matrix (other data structures for describing the transformation are possible). In other words, the latent output generated by the transducer encoder (or the transducer decoder) can be used to predict the local coordinate system for one or more teeth. Such a transducer encoder (or transducer decoder) can be trained at least in part using a representation loss (or a reconstruction loss, and other losses described herein) function that can compare the predicted coordinate system to a ground truth (or reference) coordinate system.

[0099] In another example, when the decoder produces a 3D point cloud (or other 3D representation, such as a 3D mesh, a voxelized representation, etc.), the decoder can be configured to output a description of, for example, one or more 3D points (e.g., including XYZ coordinates). In other words, the latent output generated by the transformer encoder (or transformer decoder) can be used to predict mesh elements for generating (or modifying) the 3D representation. Such a transformer encoder (or transformer decoder) can be trained using at least in part a reconstruction loss (or L1, L2, or MSE loss, and other losses described herein) function that can compare the predicted 3D representation with a ground truth (or reference) 3D representation.

[0100] In another example, when the decoder generates mesh element labels for 3D representation segmentation or 3D representation cleaning, the decoder can be configured to output a description of, for example, the labels of one or more mesh elements. In other words, the latent output generated by the transformer encoder (or transformer decoder) can be used to predict mesh element labels for mesh segmentation or mesh cleaning. Such a transformer encoder (or transformer decoder) can be trained using at least in part a cross-entropy loss (or other losses described herein) function that can compare the predicted mesh element labels with ground truth (or reference) mesh element labels.

[0101] Multi-head attention and transformers can be advantageously applied to pose generation problems. Multi-head attention is a module in a 3D transformer encoder network that is used to compute attention weights for the provided oral care data and produce an output vector with encoded information about how each example of the oral care data should attend to each other piece of oral care data in the dental arch. Attention weights are a quantification of the relationships between pairs of oral care data.

[0102] A 3D representation of oral care data (e.g., including voxels, point clouds, or 3D meshes composed of vertices, faces, or edges) can be provided to a transformer. The 3D representation can depict a patient's dentition, an appliance model (or components of the appliance model), an instrument (or components of the instrument), etc. In some embodiments, the transformer decoder (or transformer encoder) can be equipped with multi-head attention. The multi-head attention can enable the transformer decoder (or transformer encoder) to attend to different parts of the 3D representation of the oral care data. For example, the multi-head attention can enable the transformer to attend to grid elements within a local neighborhood (or clique), or attend to global dependencies between grid elements (or cliques). For example, the multi-head attention can enable a transformer for setting prediction (e.g., a transformer-based setting prediction model) to generate a transformation of a tooth and, when generating the transformation, attend to each of the other teeth in the dental arch substantially simultaneously. In other words, the transformation of each tooth can be generated based on the pose of one or more other teeth in the dental arch, resulting in a more accurate transformation (e.g., a transformation that more closely conforms to a ground truth or reference transformation). In an example of 3D representation generation (e.g., generation of a 3D point cloud), the transformer model can be trained to generate a dental restoration design. The multi-head attention can enable the transformer to attend to multiple parts of a tooth (or attend to the surfaces of adjacent teeth) as the tooth undergoes the generation process. For example, a transformer for restoration design generation can generate grid elements for an incisal edge of an incisor while at least substantially simultaneously attending to grid elements of the mesial, distal, facial, or lingual surfaces of the incisor. The result can be the generation of grid elements to form an incisal edge of the tooth that seamlessly merges with the adjacent surfaces of the tooth. This use of multi-head attention results in a more accurate modeling of the distribution of the training dataset compared to techniques that do not apply multi-head attention.

[0103] In some embodiments of the present disclosure, one or more attention vectors can be generated that describe how aspects of the oral care data interact with other aspects of the oral care data associated with the dental arch. In some embodiments, one or more attention vectors can be generated to describe how one or more parts of a tooth T1 interact with one or more parts of teeth T2, T3, T4, etc. A portion of a mesh can be described as a set of grid elements as defined herein. In some embodiments, the interacting parts of tooth T1 and tooth T2 can be determined, in part, by computing mesh correspondences as described herein. Any of these models (RNNs, GRUs, LSTMs, and transformers) can be advantageously applied to the task of setting transformation prediction, such as in the models described herein. Transformers can be particularly advantageous because they can enable the generation of transformations of multiple teeth or even an entire dental arch at once, rather than generating them individually, as may occur in some other models such as encoder architectures. In other embodiments, an attention-free transformer can be used to make predictions based on the oral care data.

[0104] A specific implementation of setting up a neural network model by the GDL may include a representation generation module (e.g., including a U-Net structure, an autoencoder encoder, a transformer encoder, another type of encoder-decoder structure, or an encoder, etc.), which can provide its output to a module that is trained to generate a tooth transformer (e.g., a set of fully connected layers with optional skip connections, or an encoder structure) to generate predictions of the transformations of each individual tooth. In some specific implementations, the skip connections can connect the output of a specific layer in the neural network to the input of another layer (e.g., a layer that is not immediately adjacent to the initial layer) in the neural network. The transformation generation module (e.g., the encoder) can process the transformation predictions one tooth at a time. Other specific implementations can replace this encoder structure with a transformer (e.g., a transformer encoder or a transformer decoder) that can process all the predictions of all the teeth substantially simultaneously. In other words, the transformer can be configured to receive a much larger number of input values than some other neural network models (e.g., than a typical MLP). This is because the transformer can accommodate an increased number of inputs, and the predictions corresponding to those inputs can be generated substantially simultaneously. The representation generation module (e.g., the U-Net structure) can provide its output to the transformer, and the transformer can generate the setting transformation for all several teeth at once, with the technical advantage of improved accuracy (because the transformation of each tooth is generated based on the transformations of each adjacent or nearby tooth, resulting in fewer conflicts and better consistency with the processing goal). The transformer can be trained to output a transformation, such as a transformation encoded by a 4×4 matrix (or some other size), a quaternion, a translation vector, Euler angles, or some other form. The transformation can place the tooth into a set pose, place the fixture model component into a pose suitable for fixture model generation, or place the appliance component into a pose suitable for appliance generation (e.g., a dental restoration appliance, a clear tray aligner, etc.). In some specific implementations, the transformation can define a coordinate system for aspects of the patient's dentition, such as a tooth mesh (e.g., the local coordinate system of the teeth). In some specific implementations, a neural network can be first used to encode the input to the transformer (e.g., a latent representation or an embedding can be generated), such as one or more linear layers and / or one or more convolutional layers. In some specific implementations, the transformer can be first trained on an offline dataset and then trained using an auxiliary actor-critic network, so that online reinforcement learning can be achieved.

[0105] In some specific implementations, the transformer can achieve large model capacity and / or implement an attention mechanism (e.g., the ability to notice and respond to certain inputs). The attention mechanism found within the transformer (e.g., multi-head attention) enables relationships within a sequence to be encoded into neural network features. The relationships within the sequence can be encoded, for example, by associating serial numbers (e.g., 1, 2, 3, etc.) with each tooth in the dental arch, or by associating serial numbers with each mesh element in a 3D representation (e.g., of a tooth). In a specific implementation where latent vectors of teeth are provided to the transformer, the relationships within the sequence can be encoded, for example, by associating serial numbers (e.g., 1, 2, 3, etc.) with each element in the latent vector.

[0106] The transformer can be scaled by increasing the number of attention heads and / or by increasing the number of transformer layers. In other words, one or more aspects of the transformer can be independently trained to handle discrete tasks and later combined to allow the resulting transformer to perform all the tasks for which its individual components have been trained, without degrading the prediction accuracy of the neural network. Scaling convolutional networks may be more difficult because the model may have lower extensibility or may have lower interchangeability.

[0107] Convolution has the ability to be rotation and translation invariant, which leads to improved generalization because a convolutional model may not need to consider the way the input data is rotated or translated. The transformer has the ability to be permutation invariant because relationships within a sequence can be encoded into neural network features.

[0108] In some specific implementations for generating or modifying 3D oral care representations, the transformer can be combined with a convolutional-based neural network, such as by vertically stacking convolutional layers and attention layers. Stacking transformer blocks with convolutional blocks enables the resulting structure to have the translational invariance of convolution and also the permutation invariance of the transformer. Such stacking can improve model capacity and / or model generalization. CoAtNet is an example of a network architecture that combines convolutional and attention-based elements and can be applied to the processing of oral care data. In some cases, transfer learning can be used at least in part to train a network for modifying or generating 3D oral care representations from CoAtNet (or another model that combines convolution and self-attention / transformer).

[0109] The techniques of the present disclosure may include operations such as 3D convolution, 3D pooling, 3D deconvolution, and 3D upsampling. 3D convolution, for example, may assist in segmentation processing when downsampling a 3D mesh. 3D deconvolution, for example, undoes or reverses 3D convolution in a U-Net. 3D pooling may assist in segmentation processing, for example, in a generalized neural network feature map. 3D upsampling, for example, reverses 3D pooling in a U-Net. These operations may be implemented by one or more layers in a predictive or generative neural network as described herein. These operations may be applied directly to mesh elements such as mesh edges or mesh faces. These operations provide a technical improvement over other methods because these operations are invariant to mesh rotation, scaling, and translation changes. Generally speaking, these operations depend on edge (or face) connectivity, so as long as edge (or face) connectivity is maintained, these operations are not affected by mesh changes in 3D space. That is, these operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position, or scale of the oral care mesh, which may improve data accuracy. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes and can be used for tasks such as 3D shape classification or mesh element labeling (e.g., for segmentation or mesh cleaning). MeshCNN performs these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.

[0110] In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 2D representations such as images. In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 3D representations such as meshes or point clouds. An intraoral scanner may capture 2D images of a patient's dentition from various angles. The intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data depicting the patient's dentition. According to various techniques, an autoencoder (or other neural network described herein) may be trained to operate on either or both of 2D and 3D representations.

[0111] A 2D autoencoder (including a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or latent capsule) using the 2D encoder and then reconstruct a copy of the input 2D image using the 2D decoder. In the case of a handheld mobile application that has been developed for such analysis (e.g., for the analysis of dental anatomy), 2D images may be easily captured using one or more onboard cameras. In other examples, 2D images may be captured using an intraoral scanner configured for such a function. Operations that may be used in implementations of a 2D autoencoder (or other 2D neural network) for 2D image analysis are 2D convolution, 2D pooling, and 2D reconstruction error calculation.

[0112] 2D image convolution can involve the "sliding" of a kernel across a 2D image and the calculation of element-wise multiplications, as well as summing these element-wise multiplications into output pixels. The output pixels generated from each new position of the kernel are saved into an output 2D feature matrix. In some specific implementations, adjacent elements (e.g., pixels) can be in well-defined positions in a straight-line grid (e.g., above, below, to the left, and to the right).

[0113] A 2D pooling layer can be used to downsample a feature map and summarize the presence of certain features in that feature map.

[0114] A 2D reconstruction error can be computed between the pixels of an input image and a reconstructed image. The mapping between the pixels can be well understood (e.g., directly comparing the upper pixel [23,134] of the input image with the pixel [23,134] of the reconstructed image, assuming the two images have the same dimensions).

[0115] One of the advantages provided by the 2D autoencoder-based techniques of the present disclosure is the ease of capturing 2D image data with a handheld device. In some cases where an external data source provides data for analysis, there can be instances where only 2D image data is available. When only 2D image data is available, it is necessary to use a 2D autoencoder for analysis.

[0116] Modern mobile devices (such as commercially available smartphones) can also have the ability to generate 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera that moves around an object to capture multiple images from different views, or both), and this 3D data can be arranged in a 3D representation, such as a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation, in some specific implementations. In some cases, the analysis of a 3D representation of an object can provide a technical improvement over a 2D analysis of the same object. For example, a 3D representation can describe the geometry and / or structure of an object with less ambiguity than a 2D representation (which can include shadows and other artifacts that complicate the depiction of the depth and texture of the object). In some specific implementations, 3D processing can achieve a technical improvement due to the inverse optics problem, which affects 2D representations in some cases. The inverse optics problem refers to the phenomenon where, in some cases, the size of an object, the orientation of the object, and the distance between the object and the imaging device can be combined in a 2D image of the object. Any given projection of an object on an imaging sensor can map to an infinite count of {size, orientation, distance} pairings. The 3D representation achieves a technical improvement because it removes the ambiguity introduced by the inverse optics problem.

[0117] Devices configured for a dedicated purpose with 3D scanning, such as a 3D intraoral scanner (or a CT scanner or an MRI scanner), can generate a 3D representation of an object (e.g., a patient's dentition) that has a significantly higher fidelity and accuracy than what a handheld device might have. When such high-fidelity 3D data is available (e.g., in oral care mesh classification or in the application of other 3D techniques described herein), the use of a 3D autoencoder provides a technical improvement (such as increased data accuracy) to extract the best possible signal from those 3D data (i.e., obtain a signal from the 3D crown mesh used in tooth classification or setup classification).

[0118] A 3D autoencoder (including a 3D encoder and a 3D decoder) can be trained on 3D data to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder, and then reconstruct a copy of the input 3D representation using the 3D decoder. Operations that can be used to implement a 3D autoencoder for analyzing a 3D representation (e.g., a 3D mesh or a 3D point cloud) are 3D convolution, 3D pooling, and 3D reconstruction error calculation.

[0119] For each mesh element, 3D convolution can be performed to aggregate local features from nearby mesh elements. Processing can be performed on top of and in addition to techniques used for 2D convolution to account for different counts and positions of adjacent mesh elements (relative to a particular mesh element). A particular 3D mesh element can have a variable neighbor count, and those neighbors may not be in the expected positions (unlike pixels in 2D convolution, which can have a fixed adjacent pixel count that exists in known or expected positions). In some cases, the order of adjacent mesh elements can be relevant to 3D convolution.

[0120] 3D pooling operations can enable the combination of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling can iteratively reduce a 3D mesh to the mesh elements that are most highly relevant to a given application (e.g., for which a neural network has been trained). Similar to 3D convolution, 3D pooling can benefit from special processing in addition to the processing required in 2D convolution to account for different counts and positions of adjacent mesh elements (relative to a particular mesh element). In some cases, the order of adjacent mesh elements may be less relevant to 3D pooling than to 3D convolution.

[0121] The 3D reconstruction error can be calculated using one or more of the techniques described herein, such as calculating the Euclidean distance between corresponding mesh elements, between two meshes. According to aspects of the present disclosure, other techniques are possible. The 3D reconstruction error can generally be calculated on 3D mesh elements rather than 2D pixels of the 2D reconstruction error. The 3D reconstruction error can achieve a technical improvement over the 2D reconstruction error because, in some cases, the 3D representation can have less ambiguity (i.e., less ambiguity in form, shape, and / or structure) than the 2D representation. In some specific implementations, due to the complexity of the mapping between the input mesh elements and the reconstructed mesh elements (i.e., the input mesh and the reconstructed mesh may have different mesh element counts, and there may be a less clear mapping between the mesh elements compared to the mapping between pixels in 2D reconstruction), additional processing may be required for 3D reconstruction over and above 2D reconstruction. Technical improvements in 3D reconstruction error calculation include increased data accuracy.

[0122] A 3D scanner, such as an intraoral scanner, a computed tomography (CT) scanner, an ultrasound scanner, a magnetic resonance imaging (MRI) machine, or a mobile device capable of performing stereophotogrammetry, can be used to generate a 3D representation. The 3D representation can describe the shape and / or structure of an object. The 3D representation can include one or more of a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation, etc. A 3D mesh includes edges, vertices, or faces. Although in some cases these three types of data are interrelated, they are distinct. A vertex is a point in 3D space that defines the boundary of the mesh. These points could alternatively be described as a point cloud without additional information about how the points are connected to each other (as described by the edges). An edge is described by two points and can also be referred to as a line segment. A face is described by multiple edges and vertices. For example, in the case of a triangular mesh, a face includes three vertices that are interconnected to form three consecutive edges. Some meshes can include degenerate elements, such as non-manifold mesh elements, which can be removed to benefit subsequent processing. According to aspects of the present disclosure, other mesh preprocessing operations are also possible. 3D meshes are typically formed using triangles, but in other embodiments, quadrilaterals, pentagons, or some other n-sided polygon can be used. In some embodiments, such as in the case of performing sparse processing, a 3D mesh can be converted into one or more voxelized geometries (i.e., including voxels). The techniques of the present disclosure operating on the 3D mesh can receive one or more tooth meshes (e.g., arranged in one or more dental arches) as input. Each of these meshes can be preprocessed before being input into a prediction architecture (e.g., including at least one of an encoder, a decoder, a pyramid encoder-decoder, and a U-Net). Such preprocessing can include converting the mesh into a list of mesh elements such as vertices, edges, faces, or into voxels in the case of sparse processing. For one or more selected types of mesh elements (e.g., vertices), a feature vector can be generated. In some examples, a feature vector is generated for each vertex of the mesh. Each feature vector can include a combination of spatial features and / or structural features, as specified in the following table:

[0123] Table 1 discloses non - limiting examples of mesh element features. In some specific implementations, in addition to the spatial or structural mesh element features described in Table 1, color (or other visual cues / identifiers) may also be considered mesh element features. As used herein (e.g., in Table 1), a point differs from a vertex in that a point is part of a 3D point cloud, while a vertex is part of a 3D mesh and may have incident faces or edges. A dihedral angle (which may be expressed in radians or degrees) can be calculated as the angle (e.g., a signed angle) between two connected faces (e.g., two faces connected along an edge). The sign on the dihedral angle can reveal information about the convexity or concavity of the mesh surface. For example, in some specific implementations, a positively signed angle may indicate a convex surface. Additionally, in some specific implementations, a negatively signed angle may indicate a concave surface. To calculate the principal curvatures of a mesh vertex, the directional curvatures of each adjacent vertex around that vertex can be calculated first. These directional curvatures can be sorted in a circular order (e.g., 0 degrees, 49 degrees, 127 degrees, 210 degrees, 305 degrees) near the vertex normal vector and may include a subsampled form of the full curvature tensor. Circular order means sorting by angle around an axis. The sorted directional curvatures can contribute to a system of linear equations that admits a closed - form solution, which can estimate the two principal curvatures and directions, which can characterize the full curvature tensor. Consistent with Table 1, a voxel may also have features that are calculated as an aggregation of other mesh elements (e.g., vertices, edges, and faces) that either intersect the voxel or, in some specific implementations, are mainly or entirely contained within the voxel. Rotating a mesh may not change the structural features but may change the spatial features. And, as described elsewhere in this disclosure, the term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non - limiting sense. In some specific implementations, in addition to mesh element features, there are alternative ways to describe the geometry of a mesh (such as 3D key points and 3D descriptors). Examples of such 3D key points and 3D descriptors can be found in "TONIONI A et al., 'Learning to detect good 3D keypoints.', Int J Comput. Vis. Vol. 126, pp. 1 - 20, 2018". In some specific implementations, 3D key points and 3D descriptors can describe the extrema (minima or maxima) of the surface of a 3D representation.In some embodiments, one or more mesh element features may be computed at least in part via deep feature synthesis (DFS), such as described in: J.M. Kanter and K. Veeramachaneni, "Deep feature synthesis: Towards automating data science endeavors", 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp. 1-10, doi: 10.1109 / DSAA.2015.7344858.

[0124] Neural networks that generate representations based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional and / or pooling layers, or other models may benefit from the use of mesh element features. Mesh element features may convey aspects of the surface shape and / or structure of a 3D representation to the neural network models of the present disclosure. Each mesh element feature describes different information about the 3D representation that may not redundantly exist in other input data provided to the neural network. For example, vertex curvature may quantify aspects of the concavity or convexity of the surface of a 3D representation that the network would not otherwise understand. In other words, mesh element features may provide a processed form of the structure and / or shape of a 3D representation that would otherwise not be available to the neural network. This processed information is generally more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, mesh element features have been provided to a neural network that generates representations based on a U-Net model and also to a model that generates representations based on a variational autoencoder with continuous normalizing flows. Based on the experiments, it was found that systems that use a full complement of mesh element features (e.g., "XYZ" coordinate tuples, "normal vectors", "vertex curvature", point pivots, and normal pivots) are at least 3% more accurate than systems that do not use mesh element features. A point pivot describes an "XYZ" coordinate tuple with a local coordinate system (e.g., at the centroid of the corresponding tooth). A normal pivot describes a "normal vector" with a local coordinate system (e.g., at the centroid of the corresponding tooth). Additionally, when using a full complement of mesh element features, training converges more quickly. In other words, machine learning models trained using a full complement of mesh element features tend to be faster and more accurate (at earlier epochs) than systems that do not. For an existing system that observes a historical accuracy of 91%, a 3% increase in accuracy reduces the actual error rate by more than 30%.

[0125] Prediction models that can operate on the feature vectors of the above features include, but are not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoder, verification using autoencoder, grid segmentation, coordinate system prediction, grid cleaning, repair design generation, appliance component generation and / or placement, and dental arch form prediction. Such feature vectors can be presented to the input of the prediction model. In some specific implementations, such feature vectors can be presented to one or more internal layers of a neural network, which is part of one or more of those prediction models.

[0126] As described herein, tooth movement specifies one or more tooth transformations, which can be encoded in various ways to specify the tooth positions and orientations within a setting and applied to a 3D representation of the teeth. For example, according to a particular implementation, the tooth position can be the Cartesian coordinates of the tooth canonical origin position defined in some semantic context. The tooth orientation can be represented as a rotation matrix, a unit quaternion, or another 3D rotation representation, such as Euler angles relative to a reference frame (global or local). The dimensions are real-valued 3D spatial extents, and the gaps can be binary existence indicators or real-valued gap sizes between teeth, especially in cases where some teeth are missing. In some implementations, tooth rotation can be described by a 3×3 matrix (or a matrix of other dimensions). In some implementations, the tooth position and rotation information can be combined into the same transformation matrix (e.g., combined into a 4×4 matrix), which can reflect homogeneous coordinates. In some cases, an affine space transformation matrix can be used to describe tooth transformations, such as the transformation describing the malocclusion pose of the teeth, the intermediate pose of the teeth, and / or the final setting pose of the teeth. Some implementations can use relative coordinates, where the setting transformation is predicted relative to the malocclusion coordinate system (i.e., predicting the malocclusion-to-setting transformation rather than directly predicting the setting coordinate system). Other implementations can use absolute coordinates, where the setting coordinate system is predicted directly for each tooth. In the relative mode, the transformation can be calculated relative to the centroid of each tooth mesh (relative to the global origin), which is referred to as "relative local". Some advantages of using relative local coordinates include eliminating the need for a malocclusion coordinate system (landmark data), which may not be applicable to all patient case datasets. Some advantages of using absolute coordinates include simplifying data preprocessing because the mesh data is initially represented relative to the global origin. In some implementations, these details regarding tooth position encoding and tooth orientation encoding can also be applied to one or more of the neural network models of the present disclosure, including but not limited to: GDL setting, RL setting, VAE setting, capsule setting, MLP setting, diffusion setting, PT setting, similarity setting, FDG setting, setting classification, setting comparison, VAE mesh element labeling, MAE mesh filling, mesh reconstruction VAE, and validation using an autoencoder.

[0127] According to a particular specific implementation, the convolutional layers in the various 3D neural networks described herein may perform mesh convolutions using edge data. The use of edge information ensures that the model is insensitive to different input orders of 3D elements. In addition to or separate from using edge data, the convolutional layers may perform mesh convolutions using vertex data. The advantage of using vertex information is that vertices are typically fewer than edges or faces, and thus vertex-oriented processing may result in lower processing overhead and lower computational costs. In addition to or separate from using edge data or vertex data, the convolutional layers may perform mesh convolutions using face data. Furthermore, in addition to or separate from using edge data, vertex data, or face data, the convolutional layers may perform mesh convolutions using voxel data. The advantage of using voxel information is that, depending on the selected granularity, there may be far fewer voxels to process compared to vertices, edges, or faces in the mesh. Sparse processing (using voxels) may result in lower processing overhead and lower computational costs (especially in terms of computer memory or RAM usage).

[0128] Neural networks for representation generation based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional layers, and / or pooling layers or other models can benefit from the use of oral care variables (e.g., oral care metrics or oral care parameters). For example, oral care metrics (e.g., orthodontic metrics or prosthodontic design metrics) can convey aspects of the shape and / or structure of a patient's dentition (e.g., the shape and / or structure of a single tooth, or a particular relationship between two or more teeth) to the neural network models of the present disclosure. Each oral care metric describes different information about the patient's dentition, which may not redundantly exist in other input data provided to the neural network. For example, the "overbite" metric can quantify the overlap between the upper central incisor and the lower central incisor along the vertical Z-axis, which may not be easily determined by traditional neural networks in some embodiments. In other words, oral care metrics provide refined information about the patient's dentition that traditional neural networks (e.g., representation generation neural networks) may not be sufficiently trained or configured to extract as described herein. However, a neural network specifically trained to generate oral care metrics can overcome this shortcoming because, for example, the loss can be calculated in a manner that promotes accurate oral care metric prediction. The mesh oral care metrics can provide a processed form of the structure and / or shape of the patient's dentition, data that would otherwise not be available to the neural network. This processed information is generally more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, oral care metrics have been provided to a representation generation neural network based on a U-Net model. Based on the experiments, it was found that systems using oral care metrics (e.g., "overbite," "overjet," and "canine class relationship" metrics) were at least 2.5% more accurate than systems that did not use oral care metrics. Additionally, training converged faster when oral care metrics were used. In other words, machine learning models trained with oral care metrics tend to be faster and more accurate (at an earlier epoch) than systems that do not. For a system that was already 91% accurate, a 2.5% increase in accuracy reduced the actual error rate by almost 30%.

[0129] The entire text of the PCT application with publication number WO2020026117A1 is incorporated herein by reference. WO2020026117A1 lists some examples of orthodontic metrics (OM). Additional examples are disclosed herein. Orthodontic metrics can be used to quantify the physical arrangement of the dental arch for orthodontic treatment purposes (as opposed to prosthodontic design metrics, which relate to dentistry and describe the shape and / or form of one or more pre-prosthetic teeth for prosthodontic support purposes). These orthodontic metrics can measure the degree of malocclusion of the dental arch, or conversely, these metrics can measure the degree of correct tooth arrangement. In some embodiments, the GDL setting model (or RL setting, VAE setting, capsule setting, MLP setting, diffusion setting, PT setting, similarity setting, and FDG setting) can incorporate one or more of these orthodontic metrics, or other similar or related orthodontic metrics. In some embodiments, such orthodontic metrics can be incorporated into the feature vector of a mesh element, where these element-based feature vectors are fed as inputs to a setting prediction network. In some embodiments, such orthodontic metrics can be directly used as direct inputs by a generator, MLP, transformer, or other neural network (such as presented in one or more input vectors of real numbers S, as described elsewhere in this disclosure). Using such orthodontic metrics in the training of a generator can improve the performance (i.e., correctness) of the resulting generator, thereby producing a predicted transformation that places the teeth closer to the correct final setting pose than other possible scenarios. Such orthodontic metrics can be used by an encoder structure or by a U-Net structure (in the case of the GDL setting). Such orthodontic metrics can be consumed by an autoencoder, variational autoencoder, masked autoencoder, or regularized autoencoder (in the case of the VAE setting, VAE mesh element labeling, MAE mesh filling). Such orthodontic metrics can be used by a neural network that generates action predictions as part of a reinforcement learning RL setting model. Such orthodontic metrics can be used by a classifier that applies labels to the set dental arch (such as labels for misalignment, grading, or final setting). This description is non-limiting as orthodontic metrics can also be incorporated into the various techniques of this disclosure in other ways.

[0130] In some examples, the various loss calculations of the present disclosure may be combined with one or more orthodontic metrics, which has the advantage of improving the accuracy of the resulting neural network. Orthodontic metrics can be used to directly compare predicted examples with corresponding ground truth examples (such as by using the metrics set in the comparison description). In other examples, one or more orthodontic metrics can be obtained from this part and incorporated into the loss calculation. Such orthodontic metrics can be calculated on the predicted examples, and then the orthodontic metrics will also be calculated on the ground truth examples. The results of these two orthodontic metrics will then be used in the loss calculation, which has the advantage of improving the performance of the resulting neural network. In some specific implementations, one or more orthodontic metrics related to the alignment of two or more adjacent teeth can be calculated and incorporated into the loss function, for example, to at least partially train the setting prediction neural network. In some specific implementations, such orthodontic metrics can help the network align the mesial surface of one tooth with the distal surface of an adjacent tooth. Backpropagation is an exemplary algorithm by which one or more loss values can be used to train a neural network.

[0131] In some specific implementations, one or more orthodontic metrics can be used to evaluate the prediction output of a neural network, such as setting prediction. Such metrics can enable the training algorithm to determine how close the prediction output is to an acceptable output, for example, in a quantitative sense. In some specific implementations, this use of orthodontic metrics can enable the calculation of loss values that do not solely depend on comparison with the ground truth. In some specific implementations, this use of orthodontic metrics can enable the loss calculation and network training to continue without the need to compare with ground truth examples. The advantage of this method is that the loss can be calculated based on general principles or specifications of the prediction output (such as settings), rather than associating the loss calculation with a specific ground truth example (which may have been defined by a specific doctor, clinician, or technician, and whose treatment concept may be different from that of other technicians or doctors). In some specific implementations, such orthodontic metrics can be defined based on the FID (Frechet Inception Distance) score.

[0132] The following is a description of some orthodontic metrics for quantifying the state of a set of teeth in an arch for orthodontic treatment. These orthodontic metrics indicate the degree of malocclusion of the teeth at a given stage of clear aligner treatment.

[0133] When training one of the neural networks of the present disclosure, the use of orthodontic metrics calculated by tensor operations may be particularly advantageous because tensor operations can facilitate efficient calculations. The more efficient (and faster) the calculations are, the faster the training can proceed.

[0134] In some examples, error patterns can be identified in one or more prediction outputs of an ML model (e.g., a transformation matrix for predicting a tooth setup, a labeling of mesh elements for mesh cleanup, adding mesh elements to a mesh for mesh filling purposes, a classification label for a setup, a classification label for a tooth mesh, etc.). One or more orthodontic metrics can be selected to be input for the next round of ML model training to address any error or defect patterns that can be identified in the one or more prediction outputs.

[0135] Some OMs can be defined relative to an arch form coordinate system (LDE coordinate system). In some specific implementations, points can be described using the LDE coordinate system relative to the arch form, where L, D, and E respectively correspond to: 1) the length along the curve of the arch form, 2) the distance from the arch form, and 3) the distance in a direction perpendicular to the L-axis and the D-axis (which can be referred to as Eminence).

[0136] Various OMs and other techniques of the present disclosure can calculate conflicts between 3D representations (e.g., of oral care objects such as teeth). Such conflicts can be calculated as at least one of the following: 1) the penetration distance between 3D tooth representations, 2) the count of overlapping mesh elements between 3D tooth representations, and 3) the overlapping volume between 3D tooth representations. In some specific implementations, an OM can be defined to quantify the conflicts of two or more 3D representations of an oral care structure (such as teeth). Some optimization algorithms (such as setup prediction techniques) can seek to minimize the conflicts between oral care structures (such as teeth).

[0137] Inter-arch orthodontic metrics are as follows.

[0138] Six (6) metrics for comparing two or more arches are listed below. Other suitable comparative orthodontic metrics are found elsewhere in the present disclosure, such as in the section on setup comparison techniques. 1. Rotational geodesic distance (rotation between the predicted example and the ground truth setup example) 2. Translation distance (gap between the predicted example and the ground truth setup example) 3. Normalized translation distance 4. 3D alignment error, which measures the distance between the predicted mesh elements and the ground truth mesh elements, in millimeters. 5. Normalized 3D alignment 6. Percentage of overlap by volume (% overlap) of the predicted example and the corresponding ground truth example (alternatively % overlap by mesh elements)

[0139] Intra-arch orthodontic metrics are as follows.

[0140] Alignment - The mesial - distal axis of the tooth can be used to calculate the 3D tooth orientation vector. A 3D vector that can be the tangent vector of the dental arch form at the tooth position can also be calculated. Then, the XY components (i.e., which can be 2D vectors) can be used to compare the orientation of the dental arch form at the tooth position with the tooth orientation in the XY space. Cosine similarity can be used to calculate the 2D orientation difference (angle) between the tangent of the dental arch form and the mesial - distal axis of the tooth.

[0141] Dental arch symmetry - For each pair of left - right teeth (e.g., the left - lower lateral incisor and / or the right - lower lateral incisor), the absolute difference between the X - coordinate of each tooth and the X - axis of the global coordinate reference system can be calculated. This increment can indicate the dental arch asymmetry of a given tooth pair. The result of such a calculation can be the average X - axis increment from one or more tooth pairs of the dental arch. In some specific implementations, this calculation can be performed with respect to the Y - axis (and / or with respect to the Z - axis having a Z - coordinate).

[0142] D - axis difference of dental arch form - The D - dimensional difference (i.e., the positional difference in the facial - lingual direction) between two dental arch states of one or more teeth can be calculated. In some specific implementations, a dictionary of D - direction tooth movements for each tooth, with the tooth UNS number as the key, can be returned. The LDE coordinate system with respect to the dental arch form can be used.

[0143] Ratio of the length of the lower dental arch form - The ratio between the current length of the lower dental arch and the length of the dental arch when it was in the initial malocclusion of the lower dental arch can be calculated.

[0144] Ratio of the length of the upper dental arch form - The ratio between the current length of the upper dental arch and the length of the dental arch when it was in the initial malocclusion of the upper dental arch can be calculated.

[0145] Parallelism of the dental arch form (entire dental arch) - For at least one origin of the local tooth coordinate system in the upper dental arch, one or more nearest origins (e.g., the origin of the tooth local coordinate system) in the lower dental arch. In some specific implementations, two nearest origins can be used. The straight - line distance from a point in the upper dental arch to the line formed between the origins of two teeth in the opposite (lower) dental arch can be calculated. The standard deviation of the set of the above - mentioned "point - to - line" distances, where the set can be composed of the point - to - line distances of each tooth in the dental arch, can be returned.

[0146] Arch form parallelism (single tooth) - This metric may share some computational elements with the global orthodontic metric of arch form parallelism, except that this metric can input the mean distance from the tooth origin to the line formed by adjacent teeth in the opposing arch (e.g., a tooth in the upper arch and the corresponding tooth in the lower arch). The mean distance can be calculated for one or more such tooth pairs. In some specific implementations, the mean distance can be calculated for all tooth pairs. Then, the mean distance can be subtracted from the distances calculated for each tooth pair. This OM can produce the deviation of the tooth from the "typical" tooth parallelism in the arch.

[0147] Buccolingual inclination - For at least one molar or premolar, find the corresponding tooth on the opposite side of the same arch (i.e., for a tooth on the left side of the arch, find the same type of tooth on the right side, and vice versa). This OM can calculate an n-element list (e.g., n can be equal to 2) for each tooth. The list can at least include the tooth IDs of the teeth in each pair of teeth (e.g., LeftLowerFirstMolar and RightLowerFirstMolar in the list = [left_tooth_idx_1, right_tooth_idx_2]). Such an n-element vector can be calculated for each molar and each premolar in the upper and lower arches. Identify the buccal cusp on each molar and each premolar on each side of the left and right sides of the arch. Draw a line between the buccal cusp of the left tooth and the buccal cusp of the right tooth. Use this line and the z-axis of the arch form to make a plane. The lingual cusp can be projected onto this plane (i.e., at this point, the inclination angle can be determined). By performing additional projections, the approximate perpendicular distance between the lingual cusp and the buccal cusp can be calculated. This distance can be used as the buccolingual inclination OM.

[0148] Canine overbite - The upper and lower canines can be identified. The first premolar on a given side of the mouth can be identified. On a given side of the arch, the distance between the upper and lower canines can be calculated, and the distance between the upper and lower first premolars can also be calculated. The average value (or median, or mode, or some other statistical value) can be calculated for the measured distances. The z-component of the result indicates the degree of overbite. The overbite can be calculated between any tooth in one arch and the corresponding tooth in the other arch.

[0149] Canine crossbite - The conflict (e.g., the conflict distance) between the canine pairs on the opposing arches can be calculated.

[0150] Canine crossbite KDE - The orthodontic metric score of the current patient case can be input, and the score can be converted to a log-likelihood using a previously trained kernel density estimation (KDE) model or distribution. This operation can produce information about where the patient case is located in the distribution of "typical" values.

[0151] Canine Overjet - This OM may share some calculation steps with the canine overbite OM. In some specific implementations, an average distance may be calculated. In some specific implementations, the distance calculation may calculate the Euclidean distance of the XY components of a tooth in the upper dental arch and a tooth in the lower dental arch to produce the overjet (i.e., as opposed to calculating the difference in the Z component, as may be performed for canine overbite). The overjet can be calculated between any tooth in one dental arch and the corresponding tooth in the other dental arch.

[0152] Canine Class Relationship (also applicable to first, second, and third molars) - In some specific implementations, this OM may include two functions (e.g., written in Python). get_canine_landmarks(): Obtain the landmarks of each tooth, which can be used to calculate the class relationship, and then in some specific implementations, map these landmarks onto a global coordinate space such that measurements can be made between teeth. class_relationship_score_by_side(): The average position of at least one landmark on at least one tooth in the lower dental arch can be calculated, and this value can be calculated for the upper dental arch. Then a vector from the upper dental arch landmark position to the lower dental arch landmark position can be calculated, and finally this vector can be projected onto the lower dental arch to produce a quantification (e.g., as a scalar) of the amount of increment in the "arch l-axis" position. This OM can calculate how far a tooth is positioned in front of or behind one or more teeth of interest in the opposite dental arch along the l-axis.

[0153] Interdigitation - By finding the midpoint between the distal marginal ridge saddle and the mesial marginal ridge saddle of a tooth, the fossa in at least one upper molar can be located. The cusp of the lower molar can be positioned between the marginal ridges of the corresponding upper molar. This OM can calculate the vector from the midpoint of the upper molar fossa to the cusp of the lower molar. This vector can be projected onto the d-axis of the dental arch form, thereby producing a lateral measurement of the distance from the cusp to the fossa. This distance can define the interdigitation magnitude.

[0154] Side Alignment - This OM can identify the leftmost and rightmost sides of a tooth, and can identify the leftmost and rightmost sides of the adjacent teeth of that tooth. The OM can then draw a vector from the leftmost side of the tooth to the leftmost side of the adjacent tooth of that tooth. The OM can then draw a vector from the rightmost side of the tooth to the rightmost side of the adjacent tooth of that tooth. The OM can then calculate the linear fitting error between the two vectors. This calculation may involve producing two vectors: Vec_tooth = right_tooths_leftside to left_tooths_leftside Vec_neighbor = right_tooths_rightside to left_tooths_leftside Then it may involve calculating the dot product of these two vectors and subtracting the result from 1. (That is, edge alignment score = 1 - abs(dot(Vec_tooth, Vec_neighbor))). A score of 0 may indicate perfect alignment. A score of 1 may mean perpendicular alignment.

[0155] Incisor arch contact KDE - can identify the deviation of the incisor arch contact from the mean of the modeled distribution of such statistical information in the dataset of one or more other patient cases.

[0156] Levelling - can calculate a measure of the levelling between a tooth and its adjacent tooth. This OM can calculate the height difference between two or more adjacent teeth. For molars, this OM can use the midpoint between the mesial and distal saddle ridges as the height of the molar. For non - molars, this OM can use the crown length from the gum to the tip. In some specific embodiments, the tip can be the origin of the local coordinate space of the tooth. Other specific embodiments can place the origin in other positions. A simple subtraction between the heights of adjacent teeth can produce the levelling increment between the teeth (e.g., by comparing the Z components).

[0157] Midline - can calculate the position of the midline of the upper incisors and / or lower incisors, and then can calculate the distance between them.

[0158] Molar arch contact KDE - can calculate the molar arch contact score (i.e., conflict depth or other type of conflict), and then can identify the position of this score in a predefined KDE (distribution) constructed from representative cases.

[0159] Occlusal contact - For a specific tooth from an arch, this OM can identify one or more landmarks (e.g., mesial cusp or central cusp, etc.). Obtain the tooth transformation of this tooth. For each cusp on the current tooth, the cusp can be scored according to the degree of contact between the cusp and the adjacent (corresponding) tooth in the opposite arch. A vector from the cusp of the tooth under discussion to the vertical intersection point in the corresponding tooth of the opposite arch can be found. The distance and / or direction (i.e., up or down) to the opposite arch can be calculated. A list including the resulting signed distances, one for each cusp on the tooth under discussion, can be returned.

[0160] Overbite - can compare the upper and lower central incisors along the z - axis. The difference along the z - axis can be used as the overbite score.

[0161] Overjet - can compare the upper and lower central incisors along the y - axis. The difference along the y - axis can be used as the overjet score.

[0162] Inter-molar arch contact - The contact fraction between molars can be calculated and conflict metrics (such as conflict depth) can be used.

[0163] Root movement d - The tooth transformation for the initial and next states can be received. The arch form axis at point L along the arch form can be calculated. This OM can return the distance moved along the d axis. This can be achieved by projecting the root pivot point onto the d axis.

[0164] Root movement l - The tooth transformation for the initial and next states can be received. The arch form axis at point L along the arch form can be calculated. This OM can return the distance moved along the l axis. This can be achieved by projecting the root pivot point onto the l axis.

[0165] Spacing - The spacing between each tooth and its adjacent tooth can be calculated. The transformation and mesh for the arch can be received. The left and right sides of each tooth mesh can be calculated. One or more points of interest can be transformed from local coordinates to the global arch coordinate system. The spacing can be calculated in a plane (e.g., the XY plane) between each tooth and its "left" adjacent tooth. An array of one or more Euclidean distances (e.g., such as in the XY plane) can be returned, which can represent the spacing between each tooth and its left adjacent tooth.

[0166] Torque - The torque (i.e., rotation about an axis such as the x axis) can be calculated. For one or more teeth, one or more rotations can be converted from Euler angles to one or more rotation matrices. The components of the rotation (such as the x component) can be extracted and converted back to Euler angles. This x component can be interpreted as the torque of the tooth. A list including the torque of one or more teeth can be returned, and this list can be indexed by the UNS number of the teeth.

[0167] The neural network of the present disclosure can utilize one or more benefits of parameter tuning operations, thereby optimizing the input and parameters of the neural network to produce more data-precise results. One parameter that can be tuned is the neural network learning rate (e.g., which can have values such as 0.1, 0.01, 0.001, etc.). The data augmentation scheme can also be tuned or optimized, such as a scheme of adding "shiver" to the tooth mesh before inputting to the neural network (i.e., small random rotations, translations, and / or scalings can be applied to change the dataset and make the neural network robust to data variations).

[0168] A subset of neural network model parameters available for tuning is as follows: ○ Learning rate (LR) decay rate (e.g., how much the LR decays during the training run) ○ Learning rate (LR). A floating-point value used by the optimizer (e.g., 0.001). ○ LR schedule (e.g., cosine annealing, step, exponential) ○ Voxel size (for the case of sparse grid processing operations) ○ Dropout % (e.g., dropout that can be performed in a linear encoder) ○ LR decay step size (e.g., decay every 10 or 20 or 30 epochs) ○ Model scaling, which can increase or decrease the layer count and / or the parameter count per layer.

[0169] Parameter tuning can be advantageously applied to the training of neural networks to predict final settings or intermediate gradings, thereby providing technical improvements towards data accuracy. Parameter tuning can also be advantageously applied to the training of neural networks for grid element labeling or for grid filling. In some examples, parameter tuning can be advantageously applied to the training of neural networks for tooth reconstruction. For the classifier models of the present disclosure, parameter tuning can be advantageously applied to neural networks for the classification of one or more settings (i.e., the classification of one or more arrangements of teeth). The advantage of parameter tuning is to improve the data accuracy of the output of the prediction model or classification model. In some cases, parameter tuning can provide the advantage of obtaining the last remaining few percentage points of validation accuracy from the prediction or classification model.

[0170] Some techniques of the present disclosure, such as setup comparison techniques and setup prediction techniques (e.g., such as GDL setup, MLP setup, VAE setup, etc.), may benefit from processing steps that can align (or register) dental arches (e.g., where teeth can be represented by 3D point clouds or some other type of 3D representation described herein). Such processing setups can be used, for example, to register a baseline ground truth setup dental arch from a patient case with a malocclusion dental arch from the same case, and then use these malocclusion and baseline ground truth setup dental arches to train a setup prediction neural network model. Such steps can assist in loss calculation because the predicted dental arch (e.g., the dental arch output by the generator) can be better aligned with the baseline ground truth setup dental arch, which is a condition that can facilitate the calculation of reconstruction loss, representation loss, L1 loss, L2 loss, MSE loss, and / or other types of losses described herein. In some specific implementations, the iterative closest point (ICP) technique can be used for such registration. ICP can minimize the squared error between corresponding entities such as 3D representations. In some specific implementations, linear least squares calculations can be performed. In some specific implementations, non-linear least squares calculations can be performed. Various registration models can incorporate in whole or in part portions of the following algorithms: Levenberg-Marquardt ICP, least squares rigid transformation, robust rigid transformation, random sample consensus (RANSAC) ICP, K-means based RANSAC ICP, and generalized ICP (GICP). In some cases, registration can help reduce subjectivity and / or randomness, which may occur, in some cases, in the design of a reference baseline ground truth setup designed by a technician (i.e., two technicians may produce different but valid final setup outputs for the same case) or other optimization techniques.

[0171] In experiments, during the training of the setup prediction model, the ground truth (or reference) setup is registered to the malocclusion (or malocclusion setup). The malocclusion teeth are provided to the setup prediction model that generates the final setup transformation for the malocclusion teeth. A loss is calculated between the resulting predicted setup and the pre-registered ground truth setup such that corresponding aspects of the two setups will be aligned. The result is a more accurate loss calculation. This pre-registration operation results in a 6% increase in absolute accuracy (e.g., as measured by the ADD10 score), which is equivalent to an almost 50% reduction in the error rate compared to conventional techniques.

[0172] Various neural network models of the present disclosure can benefit from data augmentation. Examples include models trained on 3D meshes, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, FDG setup, setup classification, setup comparison, VAE mesh element labeling, MAE mesh filling, mesh reconstruction VAE, and validation using autoencoders. Such as by Figure 1Data augmentation of the workflow shown can increase the size of the training dataset for the dental arch. Data augmentation can provide additional training examples by adding random rotations, translations, and / or rescaling to copies of the existing dental arch. In some specific implementations of the techniques of the present disclosure, data augmentation can be performed by perturbing or jittering the vertices of the mesh in a manner similar to that described in ("Equidistant and Uniform DataAugmentation for 3D Objects", IEEE Access, Digital Object Identifier 10.1109 / ACCESS.2021.3138162). The positions of the vertices can be perturbed by adding Gaussian noise, for example, having a zero mean and a standard deviation of 0.1. According to the techniques of the present disclosure, other mean and standard deviation values are possible. Other data augmentation techniques not disclosed in the prior art may involve augmentation using a neural network. For example, an encoder-decoder architecture (e.g., which may have been trained to be used as a reconstruction autoencoder during training) can be used for data augmentation. For example, a 3D oral care representation (e.g., one or more tooth meshes, or one or more tooth transformations) can be provided to the encoder portion of such an encoder-decoder and encoded into a latent representation. In some specific implementations, the techniques of the present disclosure can compute correspondences (e.g., mesh correspondences or correspondences between elements of the transformation) between a 3D oral care representation and a template 3D oral care representation (e.g., teeth having a standard or average shape). Such correspondence computations can normalize the data, thereby improving the accuracy of the latent representation generated by the encoder (e.g., as observed during the training of the systems and techniques of the present disclosure). For example, it has been observed that systems that do not utilize correspondence computations generate latent representations of teeth, which results in reconstructions with a higher error rate. The latent representation (e.g., a latent vector, such as a latent vector of size 512 or 1024 representing a tooth mesh) may undergo a target modification (e.g., using an ML model trained for that purpose) and then be reconstructed using the decoder portion of the encoder-decoder architecture. In some specific implementations, the target modification can be based on a mapping of the latent space, such as can be performed through a series of experiments in which the latent vector is changed and the effect on the subsequent reconstructed 3D oral care representation is recorded (e.g., recorded in a table or other data storage). The reconstructed 3D oral care representation (e.g., a reconstructed tooth mesh or a reconstructed transformation, or other examples of 3D oral care representations disclosed herein) can have enhanced characteristics. For example, the reconstructed 3D oral care representation can have an enhanced (or modified) shape and / or structure relative to the initial version of the 3D oral care representation (e.g., the reconstructed tooth can have a different shape, or the reconstructed transformation can place the object in a slightly different pose).The reconstructed 3D oral care representation can be used as an augmented data sample and can be output for use in training the ML models of the present disclosure.

[0173] Figure 1 A data augmentation workflow to which the system of the present disclosure can be applied to 3D oral care representations is shown. Non-limiting examples of 3D oral care representations are a tooth mesh or a set of tooth meshes. Tooth data 100 (e.g., a 3D mesh) is received at the input. The system of the present disclosure can generate a copy (102) of the tooth data 100. In Figure 1 an example, the system of the present disclosure can apply one or more random rotations to the tooth data 100 (104). In Figure 1 an example, the system of the present disclosure can apply a random translation to the tooth data 100 (106). The system of the present disclosure can apply a random scaling operation to the tooth data 100 (108). The system of the present disclosure can apply a random perturbation to one or more mesh elements of the tooth data 100 (110). The system of the present disclosure can output the augmented tooth data 112 formed by Figure 1 the method.

[0174] Since the generator network of the present disclosure can be implemented as one or more neural networks, the generator may include activation functions. When executed, the activation function outputs a determination as to whether a neuron in the neural network will fire (e.g., send an output to the next layer). Some activation functions may include: the binary step function or the linear activation function. Other activation functions endow the neural network with non-linear behavior, including: the sigmoid / logistic activation function, the Tanh (hyperbolic tangent) function, the rectified linear unit (ReLU), the leaky ReLU function, the parametric ReLU function, the exponential linear unit (ELU), the softmax function, the swish function, the Gaussian error linear unit (GELU), or the scaled exponential linear unit (SELU). The linear activation function may be well-suited for some regression applications (and other applications) in the output layer. In the output layer, the sigmoid / logistic activation function may be well-suited for certain binary classification applications (and other applications). The sigmoid activation function may be well-suited for some multi-class classification applications (and other applications) in the output layer. In the output layer, the sigmoid activation function may be well-suited for some multi-label classification applications (and other applications). The ReLU activation function may be well-suited for some convolutional neural network (CNN) applications (and other applications) in the hidden layer. The Tanh and / or sigmoid activation functions may be well-suited for some recurrent neural network (RNN) applications (and other applications) in the hidden layer, for example. There are various optimization algorithms that can be used to train the neural networks of the present disclosure (such as updating neural network weights), including gradient descent (which uses first-order derivatives to determine the training gradient and is commonly used in the training of neural networks), Newton's method (which may use second-order derivatives in loss calculations to find a better training direction than gradient descent, but may require calculations involving the Hessian matrix), and conjugate gradient method (which may converge faster than gradient descent but does not require the Hessian matrix calculations that Newton's method may require). In some specific implementations, in addition to or instead of the above techniques, additional methods may be employed to update the weights. These additional methods include the Levenberg-Marquardt method and / or simulated annealing. The backpropagation algorithm is used to convey the results of the loss calculation back into the network so that the network weights can be adjusted for learning.

[0175] Neural networks contribute to the implementation of the functions of the applications of the present disclosure, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoder, verification using autoencoders, estimation of oral care parameters, 3D grid segmentation (3D representation segmentation), coordinate system prediction, grid cleaning, restoration design generation, appliance component generation and / or placement, or dental arch form prediction. The neural networks of the present disclosure can embody parts or all of various different neural network models. Examples include U-Net architecture, multi-layer perceptron (MLP), transformer, pyramid architecture, recurrent neural network (RNN), autoencoder, variational autoencoder, regularized autoencoder, conditional autoencoder, capsule network, capsule autoencoder, stacked capsule autoencoder, denoising autoencoder, sparse autoencoder, conditional autoencoder, long / short-term memory (LSTM), gated recurrent unit (GRU), deep belief network (DBN), deep convolutional network (DCN), deep convolutional inverse graphics network (DCIGN), liquid state machine (LSM), extreme learning machine (ELM), echo state network (ESN), deep residual network (DRN), Kohonen network (KN), neural Turing machine (NTM), or generative adversarial network (GAN). In some specific implementations, an encoder structure or a decoder structure can be used. Each of these models offers one or more of its own specific advantages. For example, a specific neural network architecture may be particularly suitable for a specific ML technique. For example, autoencoders are particularly suitable for the classification of 3D oral care representations due to their ability to transform 3D oral care representations into a form that is easier to classify.

[0176] In some specific implementations, the neural networks of the present disclosure may be suitable for operating on 3D point cloud data (alternatively on 3D meshes or 3D voxelized representations). Many neural network specific implementations can be applied to the processing of 3D representations and can be applied to training predictive and / or generative models for oral care applications, including: PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN, and DSG-Net. Oral care applications include but are not limited to: setup prediction (e.g., using VAEs, RLs, MLPs, GDLs, capsules, diffusion, etc. trained for setup prediction), 3D representation segmentation, 3D representation coordinate system prediction, element tagging for 3D representation cleaning (VAE for mesh element tagging), filling of missing elements in 3D representations (MAE for mesh filling), dental restoration design generation, setup classification, appliance component generation and / or placement, dental arch form prediction, estimation of oral care parameters, setup verification or other verification applications, and 3D representation classification of teeth.

[0177] Some specific implementations of the techniques of the present disclosure incorporate the use of autoencoders. Autoencoders that can be used in accordance with aspects of the present disclosure include but are not limited to: AtlasNet, FoldingNet, and 3D-PointCapsNet. Some autoencoders can be implemented based on PointNet.

[0178] Representation learning can be applied to the setup prediction techniques of the present disclosure by training a neural network to learn a representation of a tooth and then using another neural network to generate a transformation of the tooth. Some specific implementations can use a VAE or a capsule autoencoder to generate a representation of the reconstructed features of one or more meshes relevant to the field of oral care (in some cases, including information about the structure of a tooth mesh). Then, this representation (latent vector or latent capsule) can be used as an input to a module that generates one or more transformations of one or more teeth. In some specific implementations, these transformations can place the teeth in a final setup pose. In some specific implementations, these transformations can place the teeth in an intermediate hierarchical pose. In some specific implementations, the transformation can be described by a 9×1 transformation vector (e.g., specifying a translation vector and a quaternion). In other specific implementations, the transformation can be described by a transformation matrix (e.g., a 4×4 affine transformation matrix).

[0179] In some embodiments, the systems of the present disclosure may perform principal component analysis (PCA) on an oral care mesh and use the resulting principal components as at least part of a representation of the oral care mesh in subsequent machine learning and / or other predictive or generative processing.

[0180] The systems of the present disclosure may implement end-to-end training. Some end-to-end training techniques of the present disclosure may involve two or more neural networks, where the two or more neural networks are trained together (i.e., weights are updated simultaneously during processing of each batch of input oral care data). In some embodiments, end-to-end training may be applied to pose prediction by simultaneously training a neural network that learns a representation of teeth and a neural network that can generate a tooth transformation.

[0181] According to some transfer learning embodiments of the present disclosure, a neural network (e.g., a U-Net) may be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task may be executed to provide one or more starting neural network weights for training another neural network that is trained to perform a second task (e.g., pose prediction). The first network may learn low-level neural network features of an oral care mesh and is shown to perform well in the first task. By using the first network as a starting point for training, the second network may exhibit faster training and / or improved performance. Certain layers may be trained to encode neural network features of the oral care mesh in the training dataset. These layers may then be fixed (or undergo minor changes during training) and combined with other neural network components (such as additional layers) that are trained for one or more oral care tasks (such as pose prediction). In this way, a portion of the neural network for one or more techniques of the present disclosure (e.g., pose prediction) may receive initial training on another task, which may result in significant learning in the trained network layers. This encoded learning may then be built upon by further task-specific training of another network.

[0182] According to the present disclosure, transfer learning can be used for setting prediction, as well as for other oral care applications, such as mesh classification (e.g., tooth or setting classification), mesh element labeling, mesh element filling, protocol parameter estimation, mesh segmentation, coordinate system prediction, restoration design generation, mesh verification (for any application disclosed herein). In some specific implementations, a neural network trained to output predictions based on an oral care mesh can be partially trained first on one of the following publicly available datasets before being further trained on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, Thingi10K dataset (which is particularly relevant for 3D printed component verification), ABC: Large CAD Model Dataset for Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Component Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.

[0183] In some specific implementations, a neural network previously trained on a first dataset (oral care data or other data) can subsequently receive further training on oral care data and be applied to oral care applications (such as setting prediction). Transfer learning can be used to further train any one of the following networks: GCN (Graph Convolutional Network), PointNet, ResNet, or any other neural network from the published literature listed above.

[0184] In some specific implementations, a first neural network can be trained to predict the coordinate system of teeth (such as by using the techniques described in WO2022123402A1 or U.S. Provisional Application No. US63 / 366492). According to any one of the setting prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein), a second neural network can be trained for setting prediction. Transfer learning can transfer at least a portion of the knowledge or capabilities of the first neural network to the second neural network. Thus, transfer learning can provide an accelerated training phase for the second neural network to reach convergence. In some specific implementations, the training of the second network can be completed using one or more techniques of the present disclosure after being enhanced with the transferred learning.

[0185] The system of the present disclosure may utilize representation learning to train an ML model. Advantages of representation learning include the fact that, as opposed to receiving inputs of variable size or structure, the generation network (e.g., a neural network used to predict a transformation in a setup prediction) can be configured to receive inputs of known size and / or standard format. Representation learning may yield performance superior to other techniques because noise in the input data can be reduced (e.g., because the representation generation model extracts hierarchical neural network features and / or reconstruction properties of the input representation (e.g., mesh or point cloud) through loss calculation or a network architecture chosen for that purpose).

[0186] The reconstruction properties may include values in a latent representation (e.g., a latent vector) that describe aspects of the shape and / or structure of the 3D representation provided to the representation generation module that generates the latent representation. For example, the weights of the encoder module of a reconstruction autoencoder may be trained to encode a 3D representation (e.g., a 3D mesh or others described herein) into a latent vector representation (e.g., a latent vector). In other words, the ability to encode a large set of mesh elements (e.g., hundreds, thousands, or millions) into a latent vector (e.g., hundreds or thousands of real values, e.g., 512, 1024, etc.) can be learned through the weights of the encoder. Each dimension of the latent vector may include a real number that describes some aspect of the shape and / or structure of the initial 3D representation. The weights of the decoder module of the reconstruction autoencoder may be trained to reconstruct the latent vector into a close replica of the initial 3D representation. In other words, the decoder can learn the ability to interpret the dimensions of the latent vector and decode the values within those dimensions. Generally speaking, the encoder and decoder neural network modules are trained to perform a mapping of the 3D representation to a latent vector, which can then be mapped back (or otherwise reconstructed) to a 3D representation that is substantially similar to the initial 3D representation for which the latent vector was generated.

[0187] Returning to loss calculation, examples of loss calculation can include KL divergence loss, reconstruction loss, or other losses disclosed herein. Representation learning can reduce the size of the dataset required to train a model because the representation model learns a representation such that the generative network can focus on learning the generation task. Since meaningful neural network features of the input data (e.g., local and / or global features) are available to the generative network, the result can be improved model generalization. In other words, the first network can learn a representation and the second network can make a prediction decision. By training two networks to perform their own separate tasks, each network can generate more accurate results for its corresponding task than a single network trained to both learn a representation and make a decision. In some cases, transfer learning can first train a representation generation model. Then that representation generation model (either wholly or in part) can be used to pre-train subsequent models, such as a generative model (e.g., a generative transformation prediction). The representation generation model can benefit from taking grid element features as input to improve the ability of the second ML module to encode the structure and / or shape of the input 3D oral care representation in the training dataset.

[0188] One or more neural network models of the present disclosure can have attention gates integrated therein. Attention gate integration provides an enhancement that enables the associated neural network architecture to focus resources on one or more input values. In some specific implementations, the attention gate can be integrated with a U-Net architecture, the advantage of which is that it enables the U-Net to focus on certain inputs, such as input landmarks corresponding to teeth that are intended to be fixed (e.g., to prevent movement) during an orthodontic treatment (or in cases where other special handling is required). In accordance with aspects of the present disclosure, the attention gate can also be integrated with an encoder or with an autoencoder (such as a VAE or a capsule autoencoder) to improve prediction accuracy. For example, the attention gate can be used to configure a machine learning model to give higher weights to aspects of the data that are more likely to be relevant to the correctly generated output. In this way, and because the machine learning models configured with these attention gates (or mechanisms) utilize aspects of the data that are more likely to be relevant to the correctly generated output, the final prediction accuracy of those machine learning models is improved.

[0189] The quality and composition of the training dataset for a neural network can affect the performance of the neural network during its execution phase. Dataset screening and outlier removal can be advantageously applied to the training of neural networks for various techniques of the present disclosure (e.g., for predictions for final settings or intermediate gradings, for neural networks for grid element labeling or for grid filling, for tooth reconstruction, for 3D grid classification, etc.) because dataset screening and outlier removal can remove noise from the dataset. Although the mechanisms for achieving the improvement are different from using attention gates, the end result is that the method allows the machine learning model to focus on the relevant aspects of the dataset and can lead to an improvement in accuracy similar to that achieved with attention gates.

[0190] In the case of a neural network configured to predict a final setting, a patient case may include at least one of a set of segmented tooth meshes of the patient, a malalignment transformation of each tooth, and / or a ground truth setting transformation of each tooth. In the case of a neural network predicting a set of intermediate stage settings, a patient case may include at least one of a set of segmented tooth meshes of the patient, a malalignment transformation of each tooth, and / or a set of ground truth intermediate stage transformations of each tooth. In some embodiments, the training dataset may exclude patient cases in the contact passive phase (i.e., the phase where the teeth of the dental arch do not move). In some embodiments, the dataset may exclude cases where there is a passive phase at the end of the process. In some embodiments, the dataset may exclude cases where there is overcrowding at the end of the process (i.e., cases where an oral care provider such as an orthodontist or dentist has selected a final setting where the tooth meshes overlap to some extent). In some embodiments, the dataset may exclude cases of a particular difficulty level (or levels) (e.g., easy, medium, and hard).

[0191] In some embodiments, the dataset may include cases with zero pinned teeth (or may include cases with at least one pinned tooth). A person skilled in the art may specify the pinned teeth when designing the process to prevent various tools from moving that particular tooth. In some embodiments, the dataset may exclude cases with no fixed teeth (conversely, where at least one tooth is fixed). Fixed teeth may be defined as teeth that should not move during the process. In some embodiments, the dataset may exclude cases with no pontic teeth (conversely, cases where at least one tooth is a pontic). Pontic teeth may be described as "ghost" teeth that are represented in the digital model of the dental arch but do not actually exist in the patient's dentition, or where there may be small teeth or partial teeth that may benefit from future work such as adding composite material through a prosthetic appliance. The advantage of including pontic teeth in a patient's case is to leave space in the dental arch as part of the plan for the movement of other teeth during orthodontic treatment. In some cases, pontic teeth may save space in the patient's dentition for future dental or orthodontic work such as installing implants or crowns, or applying prosthetic appliances such as adding composite material to existing teeth that are too small or have an undesirable shape.

[0192] In some specific implementations, the dataset may exclude cases where the patient does not meet the age requirement (e.g., less than 12 years old). In some specific implementations, the dataset may exclude cases where the interproximal reduction (IPR) exceeds a certain threshold amount (e.g., greater than 1.0 mm). The dataset for training a neural network to predict the settings of a clear tray appliance (CTA) may exclude patient cases that are not relevant to the CTA treatment. The dataset for training a neural network to predict the settings of an indirectly bonded tray product may exclude cases that are not relevant to the indirectly bonded tray treatment. In some specific implementations, the dataset may exclude cases where only certain teeth are treated. In such specific implementations, the dataset may include only cases where at least one of the following is treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and / or canines.

[0193] The mesh comparison module can compare two or more meshes, for example, for the calculation of a loss function or for the calculation of a reconstruction error. Some specific implementations may involve the comparison of the volumes and / or areas of two meshes. Some specific implementations may involve calculating the minimum distance between corresponding vertices / faces / edges / voxels of two meshes. For a point in one mesh (e.g., a vertex, the midpoint on an edge, or the center of a triangle), calculate the minimum distance between that point and the corresponding point in the other mesh. In cases where the other mesh has a different number of elements or there is no clear mapping between the corresponding points of the two meshes, different methods may be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools that can be utilized in the mesh comparison module of the present disclosure. In some specific implementations, the Hausdorff distance can be calculated to quantify the shape difference between two meshes. The open-source software tool Metro developed by the Visual Computing Lab can also play a role in quantifying the difference between two meshes. The following paper describes the method employed by Metro, which can be modified by the neural network applications of the present disclosure for use in mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces", P. Cignoni, C. Rocchini, and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, Vol. 17(2), June 1998, pp. 167-174.

[0194] Some techniques of the present disclosure may combine the following operations: for one or more points on a first mesh, project a ray perpendicular to the mesh surface and calculate the distance before the ray impinges on a second mesh. The length of the resulting line segment can be used to quantify the distance between the meshes. According to some techniques of the present disclosure, a color can be assigned to the distance based on the magnitude of the distance, and the color can be applied to the first mesh by means of visualization.

[0195] The setting prediction techniques described herein can generate a transformation for placing a tooth into a set pose. Such a predicted transformation may require the position and orientation of the tooth, which is a significant improvement over the prior art that uses one neural network to generate a position prediction and another neural network to generate an orientation prediction. In setting prediction, the predicted position and predicted orientation influence each other. Generating the predicted position and predicted orientation substantially simultaneously provides an improvement in prediction accuracy as compared to generating the predicted position and predicted orientation separately (e.g., predicting one without the benefit of the other).

[0196] The MLP setting, VAE setting, and capsule setting models of the present disclosure improve upon the prior art by adding (among other things) a latent space input (the latent space vector A of the oral care mesh or the latent capsule T of the oral care mesh). Existing setting prediction techniques do not train a reconstruction autoencoder to generate a representation of the tooth and thus cannot verify the correctness of their output. The advantage of using a reconstruction autoencoder to generate a tooth representation is that the latent representation (e.g., A or T) can be reconstructed by the reconstruction autoencoder. A reconstruction error can be computed (as described herein) to demonstrate the correctness of the latent encoding (e.g., demonstrating that the latent representation correctly describes the shape and / or structure of the tooth). Results with a high reconstruction error can be excluded from downstream (e.g., further or additional) processing, which results in a more accurate system overall. Either or both of A and T can be reconstructed (via a decoder) as a copy of the input oral care 3D representation (e.g., the input tooth mesh). One or more latent space vectors A (or latent capsules T) can be provided to the MLP setting model. One or more latent space vectors A (or latent capsules T) can also be provided to the VAE setting model. One or more latent capsules T (or latent vectors A) can also be provided to the capsule autoencoder setting model.

[0197] The latent space vector A (or latent capsule T) of a dental mesh, which may include thousands of interconnected mesh elements, describes the reconstruction characteristics of the dental mesh in a compact form, such as a vector of length N (e.g., where in one example N = 128). The latent space vector A (or latent capsule T) can be reconstructed into a close replica of the input dental mesh through the operation of a decoder trained for this task. The latent space vector A (or latent capsule T) is powerful because although A (or T) is relatively extremely compact, A (or T) describes sufficient characteristics of the input oral care mesh (e.g., dental mesh) to enable such a reconstruction of the oral care mesh (e.g., dental mesh). In some embodiments, the latent space vector A (or latent capsule T) can be used as an additional input to the prediction or generation models of the present disclosure. The latent space vector A (or latent capsule T) can be used as an additional input to at least one of the MLP, encoder, transformer, regularized autoencoder, or VAE of the present disclosure. The latent space vector A (or latent capsule T) can be used as an input to the GDL setting model described in the present disclosure. Additionally, the latent space vector A (or latent capsule T) can be used as an input to the RL setting model described in the present disclosure. The advantage of training a setting prediction neural network to take the latent space vector A (or latent capsule T) as an input is to provide the network with information about the reconstruction characteristics of the dental mesh. The reconstruction characteristics can include information about the local and / or global properties of the mesh. The reconstruction characteristics can include information about the mesh structure. In some cases, it can include information about the shape. The perception of these reconstruction characteristics can better enable the trained setting prediction model to predict the final setting or intermediate grading, thus providing a technical improvement in higher data accuracy. Another advantage of using the latent space vector A (or latent capsule T) is the size of the vector. If the input mesh and pose data are presented in a compact form, such as a vector of 128 real values, the neural network can more resource-efficiently encode the understanding of those data, rather than the input of the complete mesh, which may include thousands of mesh elements. The latent representation of the mesh (or meshes) can provide a more favorable signal-to-noise ratio than the initial form of the mesh or those meshes, thus enhancing the ability of subsequent ML models, such as neural networks or SVMs, to form predictions, draw inferences, and / or otherwise generate outputs (such as transformations or meshes) based on the input mesh.

[0198] Figure 2 Shows how some of the various setting prediction models can take as input 1) a dental mesh, 2) a latent space vector (or latent capsule) representing a dimension-reduced form of the dental mesh or a dental transformation.

[0199] Diffusion models are a class of deep generative neural networks that can be trained to generate transformations (e.g., that can be used to modify the position or orientation of a 3D representation in 3D space), generate images, generate 3D representations (e.g., such as point clouds, 3D meshes, or voxelized representations), or other 3D oral care representations. Diffusion models can be implemented using one or more encoders, one or more MLPs, one or more autoencoders, one or more U-Nets, and other machine learning models. In some specific implementations, the diffusion model can take as input a doctor's treatment plan (including at least a set of one or more protocol parameters, zero or more doctor preferences, and zero or more text samples describing the nature of the expected oral care treatment, such as a final setting or a restoration design generation), and also take as input one or more 3D representations of teeth (e.g., a full dental arch of segmented teeth in a malocclusion pose). In some specific implementations, the diffusion model can be trained to generate a setting that meets the specifications described by the protocol parameters and / or the text, such as a final setting. Such a model can be referred to as a diffusion setting neural network or a diffusion setting model. A text-conditioned diffusion model can use a neural network to reformat and / or reduce the dimension of the text, e.g., to generate a latent encoding or latent embedding of the text (e.g., using a transformer or an encoder).

[0200] The diffusion model can include at least one of a forward pass and a backward pass. The forward pass of the diffusion model can generate training data by iteratively adding noise (e.g., Gaussian noise) to the received 3D oral care representation (e.g., a point cloud representation of a tooth transformation or a tooth restoration design). The backward pass of the diffusion model can also operate through an iterative denoising process (e.g., such as using a U-Net trained for this purpose), which iteratively removes noise from the received 3D oral care representation (e.g., a point cloud representation of a tooth transformation or a tooth restoration design). Such tooth transformations can define the pose of the teeth in one or more dental arches. The diffusion setting model can generate settings (such as a final setting or an intermediate stage). This iterative denoising of the diffusion setting model can operate on one or more latent vectors TA trained on the tooth transformation information. The latent vector TA includes dimension-reduced information about one or more tooth transformations. In some specific implementations, these latent vectors TA can be conditioned on the latent vector A of one or more 3D representations of teeth. The diffusion model that receives one or more 3D representations of teeth can be trained to generate tooth restoration designs (e.g., such as using GGDM, which can also be trained to generate other types of 3D oral care representations, such as appliance components, transformations, or trim lines). These latent vectors TA can also be conditioned on one or more orientation parameters K and / or one or more doctor preferences L. For example, such conditioning can be achieved by concatenating the latent vector TA with K, L, or any other model input described in this disclosure (such as M, N, O, R, S, P, Q, U, V).

[0201] Inside the diffusion model, a latent vector TA (e.g., which may have been generated using an encoder) may undergo multiple iterations of noise during the forward pass. A series of increasingly noisy versions of TA can be used to at least partially train a denoising diffusion neural network (e.g., such as an autoencoder or U-Net). The latent vector TA can start as Gaussian noise at the beginning of the backward pass. Through several iterations of the backward pass of the diffusion model (e.g., using a trained U-Net or autoencoder), TA can evolve into a form that can be reconstructed (e.g., using a decoder) into one or more transformations that can place one or more teeth into a set configuration (e.g., a final set or an intermediate stage).

[0202] In some embodiments, the set diffusion model can use a U-Net architecture with ResNet blocks and self-attention layers as part of the backward pass. ResNet blocks refer to residual blocks where the activation of one layer in a neural network is directly forwarded to subsequent deeper layers, which has the advantage of being able to train deeper networks. The self-attention layer utilizes an attention mechanism that makes different parts of a sequence (i.e., a sequence at an intermediate stage) relevant in order to compute a representation of that same sequence. Self-attention is beneficial for intermediate hierarchical prediction because as the diffusion model iterates, information about other stages in the sequence can be used to update or refine each stage of the sequence. The diffusion model for set transformation prediction can be trained via gradient descent and / or backpropagation. In some embodiments, the losses described elsewhere in this disclosure can be used to at least partially train the set diffusion model. In some embodiments, either or both of the L1 loss and the L2 loss can play a role in the loss calculation.

[0203] Example of a diffusion model for set prediction: Figure 3Shows how to train a denoising diffusion setup model on a patient case consisting of 34 stages. Transformations (T0 to T33) 300 from different stages in the sequence are iteratively used to train an autoencoder (or U-Net) 302 at the center of the reverse pass of the diffusion model. The transformation may include a 4×4 affine transformation, but it should be understood that other dimensions are also compatible with the diffusion model-based techniques of the present disclosure. In some cases, the transformation may include one or more translation vectors, one or more Euler angles, and / or one or more quaternions. The “t” input is a time representation that controls which time point (i.e., stage) is sampled during a given time step of the diffusion model training (e.g., training of the denoising neural network). The input “t” is used for positional encoding. The use of positional encoding is similar to that used in transformers, enabling the model to know the relative position of the stages in the treatment plan sequence. By this method, the model can learn how to perform orthodontic treatment from stage to stage, thus leveraging what the diffusion model has learned from previous orthodontic treatment examples in the training dataset.

[0204] The similarity setup prediction method involves searching for case data that is similar to one or more trial cases. The trial malocclusion setup is designated as S4, and the corresponding predicted final setup is designated as S5. The data repository of patient case data includes malocclusion arches and corresponding ground truth final setup arches. An example malocclusion arch from the data repository is designated as S6, and the corresponding ground truth final setup is designated as S7. The final setup for S4 is predicted by extracting ground truth final setup data from one or more similar cases from the data repository.

[0205] The data repository can store data in a structured or unstructured form. Example data repositories can be any one or more of a relational database management system, an online analytical processing database, a table, a network file share, a cloud storage drive, or a folder on a hard drive or any other suitable structure for storing data.

[0206] This setting technique may involve searching a data repository of patient cases (each patient case including a corresponding upper dental arch and / or lower dental arch) to find the k patient cases S6 that are most similar to the test case S4. This similarity metric can take many different forms. One or more of the following methods and / or operations may be performed in the calculation of the set similarity (also known as arch similarity), which attempts to compare two or more dental arches (e.g., S4 and S6), such as to find similar malocclusion arches. Once it is determined that S6 is similar to S4, S7 can be used to assist in the calculation of S5. In some specific implementations, when k = 1, S5 can be set equal to S7. In other cases where k = 1, S5 can be set to a modified form of S7 (possibly modified using the output from one or more other setting prediction methods, such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, and FDG settings). In cases where k > 1, then S5 can be set to the average of the k S7 arches (alternatively, methods other than taking the average can be used to combine the k S7 arches). In some specific implementations, K-Nearest Neighbor (k-NN) can be used to compare S4 with S6 and / or quantify the difference between S4 and S6.

[0207] The similarity between S4 and S6 can be calculated by one or more different methods. The following techniques can be used alone or in combination with each other. In some specific implementations, one or more of the metrics from elsewhere in this disclosure can be calculated to quantify the difference between S4 and S6, such that the similarity of S4 and S6 can be determined and used for the similarity setting. In some specific implementations, tooth identity can be used to establish the similarity between S4 and S6 (e.g., in cases where the search for S6 is restricted to include only one or more of the teeth that appear in S4). For example, if S4 includes all teeth except the 3rd molar (i.e., S4 includes the standard 28 teeth), then in some specific implementations, the search for S6 can be limited to cases that include these 28 teeth.

[0208] In some specific implementations, the similarity between S4 and S6 can be calculated at least in part by at least one of the pointwise mesh Euclidean distance (PMD), earth mover's distance (EMD), and chamfer distance (CD), where the similarity is then used for the similarity setting. PMD, EMD, and / or CD can be used to compare the mesh elements (such as vertices) of S4 and S6.

[0209] In some specific implementations, the dental arch forms of S4 and S6 can be compared during the similarity comparison process. A dental arch can be represented using a set of nodes (or control points), a set of vertices and / or edges, curves, B-splines, NURBS surfaces, 3D meshes, or another geometry that describes the dental arch curve. PMD, EMD, and / or CD can be used to compare the representations of the dental arch forms of S4 and S6 such that the similarity between S4 and S6 can be determined and used for similarity settings. In some specific implementations, 3D key points and / or 3D descriptors (such as those described elsewhere in this disclosure) can be calculated for S4 and S6 and used to determine the similarity match used in the similarity setting prediction techniques of this disclosure. In some cases, the setting comparison tools described herein can compare S4 and S6 to determine the similarity match used in the similarity setting. In some specific implementations, one of the methods described elsewhere in this disclosure can compare S4 and S6 to determine the similarity match used in the similarity setting. In some specific implementations, the overlap between two or more teeth at the closure of dental arches S4 and S6 (e.g., between the central incisors of S4 and S6, or between the upper left second bicuspid of S4 and the upper left second bicuspid of S6) can be calculated. The greater the overlap, the greater the similarity match. The overlap can be calculated as the overlap in 3D volume, the overlap in surface area, the overlap in the shadow projected onto a plane (such as the XY plane), etc. In some specific implementations, the depth features between two or more teeth between S4 and S6 can be calculated, and the depth features can be used by a setting classifier neural network to find one or more S6s similar to S4. These depth features can include at least one feature from the following depth feature types: pose features, mesh element features, and 3D key points / 3D descriptors.

[0210] In some specific implementations, to identify j standard setting configurations, the metrics described elsewhere in this disclosure can be used to cluster patient cases. In some cases, the data of the patient cases from one of these j clusters can be averaged or otherwise combined to form one or more standard or typical setting examples for that cluster. In the case where the input setting S4 matches well with this cluster, such an averaged or typical cluster setting can be used for the output S7.

[0211] Each tooth in the dental arch has a local coordinate system. A similarity setting may be advantageous for standard methods of assigning a coordinate system to each tooth. In some embodiments, the same automated and / or machine learning methods can be used to apply a coordinate system to each tooth in each patient case in a data repository, with the advantage of standardizing the coordinate system. This standardization can be enhanced by establishing standard guidelines for applying the coordinate system to the teeth. The more standardized the application of the local coordinate system, the better the results of the similarity setting. If both S4 and S6 have a standard local coordinate system (i.e., a coordinate system applied to the teeth in the same way), then the transformation that can transform S6 into S7 will be valid in the transformation of S4 and S5.

[0212] In some embodiments, a latent space representation can be used to facilitate similarity comparison. In some cases, the meshes of dental arches S4 and S6 can be encoded into one or more latent space forms, such as latent vectors (via a variational autoencoder) or latent capsules (via a capsule autoencoder). In some embodiments, the latent space form of the dental arch mesh can be used to find a match between S4 and S6. In some embodiments, the latent space form of the dental arch mesh can be used for the clustering applications described above (e.g., the dental arch can be encoded into a latent space form and then clustered for the purpose of creating clusters that can be used for similarity setting). L2 distance, L1 norm, PMD, EMD, and / or CD can be used to compare two or more latent vectors (or latent capsules) in order to find a match. Generally speaking, such distances will be minimized. Figure 4 An overview of the similarity setting is shown.

[0213] The step of averaging over the k most similar cases may require calculating the average of one or more tooth transformations. For example, in the case where k = 2 and the two S6s in the set have the same 28 teeth, the bad-to-final setting transformation for each tooth can be averaged between each S6 in the set. For example, the bad-to-final setting transformation of the upper left canine of the first S6 can be averaged with the bad-to-final setting transformation of the upper left canine of the second S6, and the resulting average transformation can be used as the bad-to-final setting transformation in the output S5. Other methods of combining and / or averaging a set of k settings are also possible.

[0214] The GDL setting can use a U-Net architecture, an encoder structure, or a pyramid encoder-decoder structure to predict the final setting (or intermediate stage) for orthodontic treatment using a clear tray appliance or an indirect bonding tray.

[0215] Figure 5 An example training method for the ML techniques of the present disclosure is shown, which can be trained to generate transformations for 3D oral care representations. Figure 5The method can train the model to predict the setting transformation of the dental arch, predict the local coordinate system of the teeth, predict the transformation for placing the hardware component on the teeth, predict the transformation for placing the appliance or appliance component on the teeth, or place some other oral care mesh relative to another oral care mesh. In Figure 5 In the example method shown, an oral care mesh 500 (e.g., including a dental arch with multiple segmented teeth in a poor posture) can be received as an input. The tooth mesh can undergo processing to organize the mesh elements into a list and, in the case of sparse processing, optionally be converted to voxels by a mesh preprocessor module 502. Optionally, a mesh element feature vector can be calculated for each mesh element by a mesh element feature module 504. An optional input 524 (described elsewhere in the present disclosure) can be provided to a generator 506, which has the advantage of enhancing the ability of the generator 506 to customize the output. A predicted transformation 508 is generated by the generator 506 and compared with a corresponding ground truth transformation 510 (e.g., numerically comparing the predicted tooth setting with the ground truth tooth setting). The resulting loss G1 is fed back (512) for use in training the generator 506, for example, by backpropagation. In this way, the generator 506 can be trained to predict the setting transformation (or other transformations for other types of 3D oral care representations present in the training data, such as appliance components or fixture model components). The generator 506 can be implemented using at least in part an encoder, a U-Net, a transformer (e.g., including at least one of a transformer encoder and a transformer decoder, such as a GPT2 decoder), an autoencoder, a pyramid encoder-decoder, or a multi-layer perceptron (e.g., 4 fully connected layers with optional skip connections). In some specific implementations, additional training of the generator 506 can be achieved by using a discriminator 520, which can be trained using a loss D 518 to distinguish between the predicted setting 516 and the ground truth setting 514. The discriminator 520 can output a loss G2 522, which in some specific implementations can be combined with the loss G1 (512) and used to train the generator 506, the technical enhancement being to improve the data accuracy and precision.

[0216] Figure 6Shows an example implementation of generator 606. In the case of setting predictions, the malocclusion setting 602's tooth mesh is provided to module 604 (e.g., which can include a U-Net, encoder, transformer encoder, transformer decoder, or pyramid encoder-decoder, etc.), which converts the tooth mesh 602 into a form 606 (e.g., a representation including hierarchical neural network features such as global, intermediate, or local features), which form can be further processed by module 608 that can extract mesh elements (e.g., edges, faces, vertices, points, or voxels) on a per-tooth basis. Each mesh element can optionally compute a mesh element feature vector (612, 610, 614, etc.). These mesh elements 608 (and optionally associated mesh feature vectors) can be received by module 616 (e.g., an encoder or a set of fully connected layers) that can generate transformation predictions 618. In this method, some implementations can operate on edges, while other implementations can operate on other mesh elements (and their corresponding mesh element feature vectors). In other variants of this method, module 600 can use module 604 to generate an embedding vector, which is directly provided to module 616, resulting in transformation predictions 618. Such transformation predictions can be used for setting predictions.

[0217] Figure 5 Is an example architecture for GDL settings. The dental arch in a malocclusion configuration is taken as input (in some cases, the malocclusion transformation can be provided separately). Other optional inputs can be customized according to the patient's treatment needs, including protocol parameter K, doctor preference L, flag M (such as indicating fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name / nomenclature information R, one or more orthodontic metrics S, tooth size information P, distance between adjacent teeth Q, and IPR information U (such as the amount of mesial and / or distal IPR applied to each tooth, in mm). In some implementations, either or both of the latent vector A and latent capsule T can be provided to the generator. In some implementations, the dental arch form information V can be provided to the generator.

[0218] Figure 6A non-limiting example generator architecture for GDL settings is shown. The dental arch in a malocclusion configuration is taken as input (in some cases, the malocclusion transformation can be provided separately). Other optional inputs include the protocol parameter K, the doctor's preference L, the flag M (such as indicating fixed or pinned teeth), the tooth position information N, the tooth orientation information O, the tooth name / nomenclature information R, one or more orthodontic metrics S, the tooth size information P, the distance between adjacent teeth Q, and the IPR information U. In some specific embodiments, E3 can be replaced by a U-Net or a pyramid encoder-decoder. In some specific embodiments, E4 can be replaced by a transformer or a series of fully connected layers. In some specific embodiments, the encoder, U-Net, or pyramid encoder-decoder for generating tooth representations can be trained to generate representations on one or more teeth. Such a model can be trained on all teeth in both dental arches, only the teeth within the same dental arch (upper or lower dental arch), only the anterior teeth, only the posterior teeth, or some other subset of teeth. In some specific embodiments, such a model can be trained on each single tooth (e.g., the right upper canine), such that the model becomes very good at generating representations of single teeth. In some specific embodiments, the dental arch morphology information V can be provided to the generator.

[0219] The architecture can include the following elements. The maloccluded dental arch (or malocclusion dental arch) S1 is input into the model. For each tooth, a mesh feature vector is calculated for each mesh element such as a vertex or an edge or a face. In this example, vertex mesh elements are considered. Spatial features such as XYZ coordinates can be concatenated to calculate the feature vector. The mesh elements (and the associated feature vectors) of all teeth in the dental arch are concatenated and provided to a U-Net. The U-Net encodes the local information and the global information for each mesh element. Then, the output of the U-Net is concatenated with the spatial features and provided to an encoder structure. Although the information of all mesh elements of the entire dental arch is provided to the U-Net simultaneously, the encoder structure processes the mesh elements of each tooth separately. The output of the encoder structure is the transformation of the input tooth (e.g., as a 3×3 matrix). Other types of outputs are possible. The 3×3 matrix is converted to a 4×3 matrix. Two rows of the 3×3 matrix are used for two vectors, which can be used to calculate an effective coordinate system with 3 orthogonal axes using the Gram-Schmidt process (a process for orthogonalizing a set of vectors in an inner product space), and the remaining row corresponds to the translation vector of the tooth. The result is a 4×3 transformation matrix. In some specific embodiments, the transformation matrix can be converted to a 4×4 matrix by appending zeros. Other specific embodiments can encode the rotation transformation in other forms such as quaternions or Euler angles. The transformation (including rotation transformation, translation transformation, and their combination) can be defined relative to a local (tooth) coordinate system or relative to a global (entire dental arch) coordinate system. The term "mesh" should be considered in a non-limiting sense to include 3D representations such as 3D meshes, 3D point clouds, and 3D voxelized representations.

[0220] In some specific implementations, an encoder structure can be used to predict the transformation of each tooth in the dental arch for prediction settings. The encoder structure can be combined with a decoder structure, where the output of one or more levels of the encoder structure is input to one or more inputs of the decoder structure. The U-Net architecture is the result of the combination of an encoder structure and a decoder structure, as Figure 7 , Figure 8 or Figure 9 shown. In some specific implementations, the U-Net architecture is followed by an encoder structure that outputs the transformation of the teeth. This transformation is applied to the tooth mesh to move the tooth to the predicted setting position. In some specific implementations, the U-Net architecture is followed by a transformer that outputs the transformation of one or more teeth. The advantage of the transformer is that, taking into account the physical interactions between the teeth, the transformer can generate multiple teeth at once. In some specific implementations, the transformer can generate the transformation for all the teeth in the dental arch at once, taking into account the interactions between the teeth. In some specific implementations, the generator includes a U-Net and is trained at least in part using a discriminator neural network. In other specific implementations, the generator includes an encoder structure.

[0221] The advantage of the U-Net is to extract the local and global information of each mesh element in the tooth mesh. The advantage of the encoder structure is to extract the local and global information of the entire tooth mesh rather than each individual mesh element. Another advantage of the U-Net is to perform feature selection on the features that will be consumed by the encoder that consumes the output of the U-Net.

[0222] In other specific implementations, such as in Figure 6 , Figure 6 the first encoder 604 can directly output the tooth transformation of several teeth of one or more dental arches. In other specific implementations, Figure 6 the first encoder 604 in Figure 5 can be replaced by a transformer encoder (or transformer decoder) that can directly output the tooth transformation of several teeth of one or more dental arches. In other specific implementations, the transformer can be trained to act as a preprocessor for tooth data and be trained to generate an embedding vector (or alternatively, a latent vector), and then the embedding vector can be provided to another machine learning model that can be trained to output one or more setting transformations. In some specific implementations, the mesh preprocessor module (which, in some specific implementations, can convert the mesh elements of the input 3D oral care representation into a list that is convenient for the generator to consume) and the mesh element feature module from Figure 6 can also be applied to the input of the network in Figure 5 and Figure 6The methods in each of them can implement sparse processing and operate on voxelized grid data. In this case, the grid preprocessor can convert the input grid data into voxels.

[0223] In some specific implementations, the generator can adopt the general pyramid encoder-decoder structure shown in U.S. Provisional Patent Application No. 63 / 264,914, the entire disclosure of which is incorporated herein by reference in its entirety.

[0224] Grid pooling downsamples the count of grid elements in a grid while aggregating the features associated with each grid element into features that contain information from other grid elements (i.e., the grid elements contributing to the pooling). Grid pooling is typically local in a layer of a neural network, thus involving a small neighborhood around each grid element. Global grid pooling involves the entire tooth rather than a small neighborhood. Grid unpooling reverses the grid pooling process by unpacking the information contained in each downsampled grid element into information that can be applied to multiple grid elements in a higher (or initial) resolution grid.

[0225] Grid convolution uses a kernel to aggregate information from the local neighborhood of grid elements. Grid convolution can use the kernel to learn.

[0226] The coordinate normalization layer is a layer in a neural network that can be used to generate normalized position information of grid elements, which can then be concatenated with other features and fed into another layer in the neural network. For an example of coordinate normalization, see (“An intriguing failing of convolutional neural networks and the CoordConv solution” Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, Jason Yosinski. NeurIPS 2018), the entire disclosure of which is incorporated herein by reference.

[0227] The GDL setting generator can include one or more elements, including U-Net, encoder, decoder, fully connected layer, transformer, autoencoder, embedding vector (or latent vector), or pyramid encoder-decoder. Figures 7 to 9 Some non-limiting generator designs are shown. Figure 7 The U-Net+EV+encoder architecture is shown. Figure 8 The U-Net+EV+MLP architecture is shown. Figure 9 The U-Net+EV+transformer architecture is shown.

[0228] Some specific implementations of a 3D U-net (e.g., a U-net for GDL setup prediction) can incorporate dropout layers into either or both of the encoder and decoder components, which is an improvement over the prior art because dropout enables the U-net to counteract the effects of overfitting. Some specific implementations can also utilize normalization layers.

[0229] The embedding vector (EV) output by the U-Net encodes both global and local information from the input mesh. The advantage of the EV is that all these global and local mesh characteristics are encoded in a concise data structure that is easily consumable by subsequent ML models, such as for classification (i.e., in the case of a mesh element labeling model that segments the mesh and labels the mesh elements with class labels) or output generation (i.e., in the case of feeding the output of the U-Net structure into an MLP, transformer, or encoder for setup prediction model to generate one or more transformations to place one or more teeth into a setup pose, for a final setup or an intermediate stage).

[0230] The spatial and / or structural features described elsewhere in this disclosure can be applied to the input for GDL setup, just like other setup prediction models. These features can be provided as input to the generator and / or discriminator. The features can also be provided to one or more layers inside the generator or discriminator.

[0231] In some specific implementations, spatial features are concatenated onto the list of mesh elements provided to the generator. The spatial features are calculated based on the provided coordinates. The spatial features can include the XYZ vertices of the mesh. The vertices can be normalized to have a normal distribution, or can be normalized to be incorporated onto a unit sphere, among other examples. The advantage of this type of spatial feature is that the spatial features can be injected into selected layers of the network, not just at the input layer. The input mesh features are concatenated with the spatial features, and the resulting vector of mesh element features enters the U-Net for processing. The spatial features are also concatenated with the output of the U-Net for input into the encoder.

[0232] One or a group of tooth meshes can be converted into a voxelized form before being provided to the generator. In some specific implementations, for the purposes of this disclosure, a voxel can be considered a type of mesh element (along with vertices, faces, and edges). This conversion can be formed by a mesh preprocessor module (see Figure 5 ), which reformats the mesh data for input into the generator (e.g., rearranges the mesh elements into a list).

[0233] One advantage of using an automatic differentiation engine is support for handling disconnected meshes. Some of the automatic differentiation engines incorporated by the techniques of this disclosure may use 3D convolution techniques that borrow from and extend conventional techniques for convolution in 2D images. The automatic differentiation engine is used to overcome challenges in 3D processing (i.e., large memory consumption) by leveraging sparsity in the data representation (e.g., a point cloud spanning the surface of an object rather than the volume of the object). The techniques of this disclosure advantageously apply the voxelization techniques of one or more automatic differentiation engines to the field of digital oral care. The voxelization-based resource-saving advantages used in the techniques of this disclosure include the following: 1) saving the RAM required to describe two dental arches, thereby enabling the entireties of the two dental arches to be processed together, enabling arch interactions to be considered; and 2) reducing the time required to train the network.

[0234] The techniques of this disclosure use an automatic differentiation engine for voxelization and sparse processing. Spatial information is included in the input features. It is important that those spatial features are not distorted by the way voxelization is performed. When voxelization is performed, the features of different vertices are aggregated to the voxel centroid. Each element has a feature assigned to it, and these features are aggregated to the voxel centroid, which may cause distortion of the spatial features. The splat() and sparse() functions from the automatic differentiation engine perform different interpolations. Splat() scales the feature relative to the centroid of the voxel, which may distort the spatial position information. Sparse() does not have this problem. Sparse() uses a different method to assign mesh element features to voxels. Sparse() averages all the features of the mesh elements within a voxel and assigns that average to the centroid. Sparse() produces better experimental results than Splat().

[0235] A feature vector can be calculated for a voxel in a manner similar to a mesh element. Just as each mesh element has a feature vector, each voxel can have a feature vector that accompanies the voxel when the voxel is provided to one of the neural networks of this disclosure. The feature vector can be formed as a result of averaging the feature vectors of the mesh elements that fall within the voxel. Such features can include the spatial features and structural features described herein. In some embodiments, each voxel can have voxel-specific computed features, such as the volume of the voxel or a quantification of the distribution of the mesh elements contained within the voxel.

[0236] Teeth can also be represented as finite elements, such as the finite elements used in finite element analysis (FEA). Teeth, for example, can be used to train and deploy one or more of the neural networks of this disclosure.

[0237] In some specific implementations, one or more of the other neural network models of the present disclosure may also benefit from sparse grid processing operations, which involve converting one or more grids into voxels, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, setting comparison, setting classification, VAE grid element tagging, MAE grid filling, verification using an autoencoder, or estimation of missing protocol parameter values. The advantages of this are lower memory requirements and potentially faster processing. This can also improve the prediction results of the corresponding model, as voxelization can enable the model to isolate and / or identify important features of the input grid and focus the model training on the encoding of these features.

[0238] Figure 5 The generator or discriminator in may be trained using at least one or more of loss G1, loss G2, and / or loss D. Loss G1 and loss G2 may involve the calculation of one or more of the following: representation loss, reconstruction loss, L1 loss, L2 loss, MSE loss, smooth L1 loss, etc. In some specific implementations, a normalization operation may be applied. For example, in some specific implementations, the loss may be normalized with respect to the size of the tooth (e.g., as measured by the average L1 or L2 distance of the grid elements from the center of the grid).

[0239] Figure 5 The generator in may include one or more encoder structures, one or more decoders, one or more U-Net structures, one or more multi-layer perceptrons (MLPs), or one or more transformers. In some specific implementations, the generator may include an encoder structure that outputs a transformation. In some specific implementations, the generator may include a U-Net structure that generates an output consumed by the encoder structure, which outputs a transformation. In some specific implementations, the generator may include a U-Net structure that generates an output consumed by the transformer, which outputs one or more transformations.

[0240] Some specific implementations may incorporate chamfer distance (CD) calculations into the loss, which measure the squared distance between each point in a set of grid elements and its nearest neighbor in another set of grid elements. In some specific implementations, the chamfer distance can implement a differentiable method for comparing groups of grid elements (such as groups of vertices). Some specific implementations may incorporate earth mover's distance (EMD) into the loss, which measures the squared distance between two sets of grid elements. Some specific implementations may incorporate pointwise mesh Euclidean distance (PMD) calculations into the loss calculation. Some specific implementations may incorporate Hausdorff distance (HD) calculations into the loss calculation. Figure 5 The discriminator in may include an encoder structure or another type of classifier. Loss D may include the calculation of cross-entropy loss, etc.

[0241] It is shown that the loss has two components. One component is related to rotation, while the other component is related to the translation of the teeth. Each component is directly calculated on the coordinate system representation (e.g., 3×3 rotation matrix and 3×1 translation vector). The 4×3 matrix is formed by the 3×3 matrix and the 1×3 matrix. The distance is calculated as the difference between the ground truth transformation and the predicted transformation.

[0242] The predicted 3×3 matrix is subtracted from the ground truth 3×3 rotation matrix, and the L1 norm or L2 norm of the result is calculated as the rotation loss. The predicted 1×3 matrix is also subtracted from the 1×3 ground truth translation vector, and the L1 norm or L2 norm of the result is similarly calculated as the translation loss. The total loss is the weighted average or sum of the rotation loss and the translation loss.

[0243] It is shown that the loss produces good rotation predictions, but there are one or more challenges in translation predictions. The reconstruction loss overcomes this limitation of the representation loss. First, the predicted transformation is applied to the malaligned teeth, and then the ground truth setting transformation is applied to the malaligned teeth. The calculated loss is the average distance between the grid elements of these two transformation grids. The L1 norm or L2 norm can be used to calculate the distance, or a combination of both can be used. The distance between the grid elements of the ground truth setting transformation grid and the predicted transformation grid can be normalized later. The normalization process may require dividing the calculated distance value by a term such as the size of the ground truth grid. In some specific implementations, the size can be defined by the grid diameter. In other specific implementations, the size can be defined by the average distance of the grid elements from the grid centroid (e.g., calculated by the L1 norm or L2 norm of each grid element from the grid centroid of the grid). In one specific implementation, the L1 distance between the grid elements is calculated, and the distance is normalized using the L1 size of the grid. Next, the L2 distance value between the grid elements is calculated, and the distance is normalized using the L2 size of the grid. Subsequently, the results of these normalized L1 and L2 calculations are added and output as the loss. In some specific implementations, the L2 loss can be replaced by the MSE loss. Similar calculations are performed by the comparison tool to quantify the difference between the ground truth setting transformation pattern of the teeth and the predicted transformation pattern of the teeth. The advantage of the reconstruction loss is that both the rotation loss and the translation loss are trained, and neither the rotation loss nor the translation loss is superior to the other.

[0244] In some specific implementations, the loss G1 corresponding to each patient case is the average of the losses of single teeth in that patient case. In other specific implementations, the loss is calculated as a weighted average of the losses of single teeth, where some teeth have a greater weight than others. In other specific implementations, the loss is calculated as the maximum loss of any tooth in that patient case. In some specific implementations, different transformations may be predicted for each mesh element or for a subgroup of mesh elements. In other specific implementations, a coordinate system may be predicted for the entire tooth. Combinations are also possible. Some specific implementations of the loss G1 calculate a unified loss, where the loss is calculated on a single transformation for each tooth, rather than calculating the loss on many transformations in the tooth (i.e., one transformation per tooth element before global pooling). Before global pooling, there may be one transformation for each tooth mesh element (such as a vertex, edge, or face). In the unified loss, there is one transformation for the entire tooth, and the loss is calculated based on that one loss. There are relative losses and absolute losses. "Absolute" is a relative alternative form of local. That is, in "absolute", the local coordinate system of each tooth is predicted. Techniques based on absolute loss may benefit from the network's knowledge of the malocclusion local coordinate system. In contrast to the "absolute" loss, the "relative" loss predicts the transformation that maps the tooth from the malocclusion to the set. "Local" means calculating the relative transformation by assuming that the malocclusion tooth is located at the origin. Some specific implementations may combine two or more of the loss calculation strategies described herein, such as the representation loss and the reconstruction loss.

[0245] In some specific implementations, the ADD score can be used to measure the accuracy of the generator G. The number or percentage of dental mesh elements in the predicted transformed dental mesh that are less than a threshold distance (i.e., 1 / 10 of the dental size distance) from the reference ground truth transformed dental mesh is measured. The distance between corresponding mesh elements (e.g., between vertices in the predicted transformed dental mesh and corresponding vertices in the reference ground truth transformed dental mesh) is calculated. The distance can be calculated using the L1 norm, the L2 norm, or by another method. In some specific implementations, the distance value can be normalized, for example, by dividing the distance value by the dental mesh size (such as the dental diameter). If the normalized value is less than the threshold (e.g., 0.1), the predicted transformed dental mesh element is considered to be close enough to the corresponding reference ground truth transformed dental mesh element. In some specific implementations, the training accuracy (or validation accuracy) is the percentage of dental mesh elements in the predicted transformed dental mesh that fall within a specified threshold distance of the corresponding mesh elements in the reference ground truth transformed dental mesh. In some specific implementations, the training accuracy (or validation accuracy) is the count or percentage of teeth in the dental arch, where a threshold percentage (e.g., 90%) of the dental mesh elements in the predicted transformed dental mesh fall within a specified threshold distance of the corresponding mesh elements in the reference ground truth transformed dental mesh. The advantage of using the ADD score is that the ADD score can incorporate the geometry of the dental mesh while providing a standard for comparing the predicted setup transformation and the reference ground truth setup transformation.

[0246] One or more of the other neural networks and / or predictive ML models of the present disclosure may also provide a technical improvement in specific implementations, where they are enhanced by calculating the ADD score as a measure of progress during model training. Examples of such models of the present disclosure include, but are not limited to: GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, setup comparison, setup classification, VAE mesh element tagging, MAE mesh filling, validation using an autoencoder, or estimation of missing protocol parameter values. The FDG setup can also benefit from calculating the ADD score as a measure of progress during model training.

[0247] There are other methods for accurately calculating the generator G. In other specific implementations, the training and validation accuracy can be calculated as the percentage of overlap between two or more corresponding dental meshes. For example, if 90% of the two meshes overlap (e.g., by volume, by element count, or by surface area), the prediction can be considered accurate. In these examples, a correct prediction is one in which the predicted teeth overlap the reference ground truth setup teeth by at least a threshold amount (e.g., by volume, by element count, or by surface area). A technical improvement based on data precision is to make the loss calculation consider the physical boundaries of the teeth.

[0248] Some specific implementations of the generator G (e.g., involving U-Net, etc.) may use a coordinate normalization layer, such that coordinates and other position information can be processed in a different way from how other features are processed. There are two sets of information: position information and other features. During the execution of the neural network, these two sets of information can be processed differently. The position information is processed separately from other features because the position information may be distorted due to the voxelization process. There are two ways to input the position information into the generator network G.

[0249] NoCoordNormLayer: The position information is provided to the network together with other features appended to the feature vector associated with each grid element.

[0250] WithCoordNormLayer (also known as CoordNorm): The position information is not provided to the network together with other features. Instead, the position information is first provided to a normalization layer, where the position information is normalized / standardized in 3D space and then concatenated with other input features or some intermediate network levels.

[0251] Experiments have found that NoCoordNormLayer and WithCoordNormLayer have comparable validation accuracies, although the training accuracies do show differences. The training accuracy of the coordinate normalization layer is higher.

[0252] In some specific implementations, an embedding vector (EV) may be a vector that represents the output of a generative machine learning module (e.g., a first machine learning model) and is a feature to be used during a prediction transformation. The generative machine learning module may include one or more U-Net architectures, one or more 3D SWIN transformers, one or more pyramid encoder-decoder structures, or other networks that can be trained to extract hierarchical neural network features from a 3D representation of a tooth (e.g., or other 3D representations for which a transformation is to be generated). The embedding vector may include a dimension-reduced form of the tooth. This dimension-reduced form of the tooth can enable the setup prediction neural network to more effectively encode the reconstruction characteristics of the tooth and better learn to place the tooth in a pose suitable for the final setup or intermediate stage, thereby providing technical improvements in both data accuracy and resource consumption.

[0253] In addition, a reduced - dimensional representation of the teeth can be provided to a second ML module that can generate a predicted setup transformation. The low dimension can offer many advantages. For example, training a machine - learning model on data samples (e.g., from a training dataset) with variable sizes (e.g., one sample has a different size from another sample) can be very error - prone, and the resulting machine - learning model generates less accurate prediction outputs. Additionally, training a machine - learning model on data samples larger than a certain size may result in a less accurate model because the model cannot encode the distribution of the large data samples. Both of these problems exist in typical datasets of cohort patient case data. The standard size and low - dimensional nature of the latent vectors described herein address both of these problems, resulting in a more accurate machine - learning model (e.g., the second ML module that can be trained to generate setup transformations or perform classification).

[0254] In some specific implementations, the EV can be provided to a second machine - learning module (e.g., an encoder or MLP) that has been trained to generate predictions for the final setup or grading transformation of one or more teeth. In some specific implementations, the EV can be provided to one or more global average pooling layers that have been trained to generate predictions for the final setup or grading transformation of a tooth. In other specific implementations, the EV can be received by a multi - layer perceptron (MLP) that has been trained to generate predictions for the final setup or grading transformation of a tooth. In one specific implementation, the MLP includes four (4) fully - connected neural network layers. In some specific implementations, the EV can include local and / or global spatial information about one or more input meshes. In some specific implementations, the EV can include structural information about one or more input meshes. In some cases, there may be an EV for each tooth grid element. In other cases, there may be an EV for the entire tooth mesh. An EV or a set of EVs can be used by, for example, an encoder structure (alternatively, an MLP or a transformer) to predict the transformation that moves the tooth to a setup pose (for an intermediate stage or for the final setup). The size of the EV is a parameter that can be adjusted during the process of training the setup - prediction neural network. Possible sizes can include, but are not limited to, powers of two: 2, 4, 8, 16, 32, 64, 128, 256, 512, etc.

[0255] In some specific implementations, the U-Net can output an EV for each mesh element. In an example where there are 28 teeth in a patient case and each tooth contains 1000 vertices (a number chosen for simplicity), the U-Net can input an embedding vector (e.g., of size 9) for each of the 28 × 1000 = 28000 mesh elements received at the input of the U-Net. These sets of embedding vectors (28,000 embedding vectors in this specific example) can be concatenated and fed into a second ML model, which can map those embedding vectors to one or more vectors of transformation values for one or more teeth (e.g., the second ML model can generate the transformation of the 28 teeth of the patient to place those teeth in a set pose).

[0256] In some cases, the embedding vector EV at the output of the U-Net structure can be used for classification. In some specific implementations, the EV can be used for tooth mesh classification (e.g., tooth type, tooth health). In some specific implementations, the EV can be used to classify the full dental arch for setup classification purposes (e.g., malocclusion, grading, final setup).

[0257] The techniques described herein can be trained to generate transformations that can place a patient's teeth into a posture suitable for use in an orthodontic setting (e.g., an intermediate or final setting) according to oral care parameters that may be provided to the generation model in some embodiments. Oral care arguments can include oral care parameters as disclosed herein, or other real-valued, text-based, or categorical inputs that specify expected aspects of one or more 3D oral care representations to be generated. In some cases, oral care arguments can include oral care metrics that can describe expected aspects of one or more 3D oral care representations to be generated. Oral care arguments are particularly applicable to the embodiments described herein. For example, an oral care argument can specify an expected design (e.g., including shape and / or structure) of a 3D oral care representation that can be generated (or modified) according to the techniques described herein. In short, embodiments that use the specific oral care arguments disclosed herein generate more accurate 3D oral care representations than embodiments that do not use specific oral care arguments. In some cases, a text encoder can encode a set of natural language instructions from a clinician (e.g., generate a text embedding). The text string can include tokens. In some embodiments, the encoder used to generate the text embedding can apply average pooling or max pooling between token vectors. In some cases, a transformer (e.g., BERT or Siamese BERT) can be trained to extract embeddings of text used in digital oral care (e.g., by training the transformer on examples of clinical text such as those given below). In some cases, such a model used to generate text embeddings can be trained using transfer learning (e.g., initially trained on another text corpus and then receiving further training on text related to digital oral care). Some text embeddings can encode text at the word level. Some text embeddings can encode text at the token level. In some embodiments, the transformer used to generate the text embedding can be trained at least in part using a loss calculation that compares a predicted output to a ground truth output (e.g., softmax loss, multi-negative example ranking loss, MSE margin loss, cross-entropy loss, etc.). In some cases, non-text arguments (such as real-valued or categorical values) can be converted to text and subsequently embedded using the techniques described herein. The following are examples of natural language instructions that can be issued by a clinician to the generation model described herein: "Generate a setting for Class I molars and canines, 2 mm overbite and add 2 mm of expansion 5-5 / 5-5", "Generate a setting to conform to proclination and expansion and leave 0.5 mm of space U2-2 for future restoration", or "Adjust the setting without second molar movement, rotate the upper first molar mesially outwards for Class I, lower the height to an inverted curve of Spee 2 mm, and advance the mandible to Class I canines with elastics".

[0258] In some cases, a local coordinate system for 3D oral care representations (such as teeth) can be described by one or more transformations (e.g., an affine transformation matrix, a translation vector, or a quaternion). The systems of the present disclosure can be trained for coordinate system prediction using the coordinate systems of past cohort patient case data. The past patient data can include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems. Machine information models (such as U-Net, an encoder, an autoencoder, a pyramid encoder-decoder, a transformer, or convolutional layers and / or pooling layers) can be trained for coordinate system prediction. Representation learning can determine a representation of a tooth (e.g., encoding a mesh or point cloud into a latent representation, such as using U-Net, an encoder, a transformer, or convolutional layers and / or pooling layers, etc.), and then predict a transformation for that representation (e.g., using a trained multi-layer perceptron, a transformer, an encoder, a transformer, etc.), the transformation defining a local coordinate system for that representation (e.g., including one or more coordinate axes). In the case of predicting a coordinate system for a tooth mesh, the mesh convolution techniques described herein can utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth. Reinforcement learning techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth.

[0259] Machine information models (such as U-Net, encoders, autoencoders, pyramid encoder-decoders, transformers, or convolutional layers and / or pooling layers) can be trained as part of a method for hardware (or appliance component) placement. Representation learning can train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder or using blocks of U-Net, encoder, transformer, convolutional layer, and / or pooling layer, etc.). The representation can include a dimensionally reduced form and / or an information-rich form of the input 3D oral care representation. In some embodiments, the generation of the representation can be assisted by calculating a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some embodiments, a representation can be calculated for a hardware element (or appliance component). Such a representation is adapted to be provided to a second module that can perform a generation task, such as transform prediction (e.g., a transform for placing a 3D oral care representation relative to another 3D oral care representation, such as a representation for placing a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such a transform can include an affine transformation matrix, a translation vector, or a quaternion, etc. Machine learning models that can be trained to predict a transform for placing a hardware element (or appliance component) relative to an element of a patient's dentition include MLP, transformers, encoders, etc. The systems of the present disclosure can be trained for 3D oral care appliance placement using past cohort patient case data. The past patient data can include at least: one or more ground truth transforms and one or more 3D oral care representations (such as tooth meshes or other elements of a patient's dentition). In cases where U-Net (and other neural networks) are trained to generate a representation of a tooth mesh, the mesh convolution and / or mesh pooling techniques described herein utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be performed for hardware or appliance component placement. Reinforcement learning techniques can be performed for hardware or appliance component placement.

[0260] In some embodiments, the techniques of the present disclosure can extract local or global neural network features from a 3D point cloud or other 3D representation (e.g., a 3D point cloud describing aspects of a patient's dentition such as teeth or gums) using PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using PointNet or PointNet++ as a basis for training). In some embodiments, the techniques of the present disclosure can extract local or global neural network features from a 3D point cloud or other 3D representation using U-Net.

[0261] 3D oral care representations are described herein as such because 3D representations are current state of the art. However, 3D oral care representations are intended to be used in a non-limiting manner to cover any representation in 3D or higher-order dimensions (e.g., 4D, 5D, etc.), and it should be understood that the techniques disclosed herein can be used to train machine learning models to operate on representations in higher-order dimensions.

[0262] In some cases, the input data may include 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data related to splines (e.g., control points). The encoder-decoder structure may include one or more encoders, or one or more decoders. In some specific implementations, the encoder may take as input the mesh element feature vectors of one or more input mesh elements in the input mesh. The encoder is trained in a way that processes the mesh element feature vectors to generate a more accurate representation of the input data. For example, the mesh element feature vectors may provide the encoder with more information about the shape and / or structure of the mesh, and thus the additional information provided allows the encoder to make more informed decisions and / or generate a more accurate latent representation of the mesh. Examples of encoder-decoder structures include U-Net, autoencoders, or transformers, among others. The representation generation module may include one or more encoder-decoder structures (or parts of encoder-decoder structures, such as individual encoders or individual decoders). The representation generation module may generate an information-rich (optionally dimension-reduced) representation of the input data that can be more easily consumed by other generative or discriminative machine learning models.

[0263] U-Net may include an encoder followed by a decoder. The architecture of U-Net may be similar to a U shape. The encoder may extract one or more global neural network features, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level compared to the most global level) from the input 3D representation. The output of each level from the encoder may be passed to the input of the corresponding level of the decoder (e.g., via skip connections). Similar to the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For example, the decoder may output a representation of the input data that may include global, intermediate, or local information about the input data. In some specific implementations, U-Net may generate an information-rich (optionally dimension-reduced) representation of the input data that can be more easily consumed by other generative or discriminative machine learning models.

[0264] An autoencoder can be configured to encode input data into a latent form. The autoencoder can train an encoder to reformulate the input data into a dimensionally reduced latent form between the encoder and the decoder, and then train the decoder to reconstruct the input data from this latent form of the data. A reconstruction error can be computed to quantify the degree to which the reconstructed form of the data differs from the input data. In some embodiments, the latent form can be used as an information-rich dimensionally reduced representation of the input data, which can be more readily consumed by other generative or discriminative machine learning models. In most scenarios, the autoencoder can be trained to take a 3D representation as input, encode this 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close approximation of the input 3D representation at the output.

[0265] A transformer can be trained to generate a representation of its input using self-attention at least in part. The transformer can encode long-range dependencies (e.g., encode relationships between a large number of inputs). The transformer can include an encoder or a decoder. In some embodiments, such an encoder can operate in a bidirectional manner or can operate a self-attention mechanism. In some embodiments, such a decoder can operate a masked self-attention mechanism, can operate a cross-attention mechanism, or can operate in an autoregressive manner. In some embodiments, the self-attention operation of the transformer described herein can be related to different positions or aspects of a single 3D oral care representation in order to compute a dimensionally reduced representation of the 3D oral care representation. In some embodiments, the cross-attention operation of the transformer described herein can mix or combine aspects of two (or more) different 3D oral care representations. In some embodiments, the autoregressive operation of the transformer described herein can consume previously generated aspects of a 3D oral care representation (e.g., previously generated points, point clouds, transforms, etc.) as additional inputs when generating a new or modified 3D oral care representation. In some embodiments, the transformer can generate a latent form of the input data that can be used as an information-rich dimensionally reduced representation of the input data, which can be more readily consumed by other generative or discriminative machine learning models.

[0266] In some embodiments, an encoder-decoder architecture can first be trained as an autoencoder. In deployment, one or more modifications can be made to the latent form of the input data. Then, this modified latent form can continue to be reconstructed by the decoder, resulting in a reconstructed form of the input data that differs from the input data in one or more desired aspects. Oral care variables such as oral care parameters or oral care metrics can be provided to the encoder, the decoder, or can be used to modify the latent form in order to influence the encoder-decoder architecture when generating a reconstructed form with desired characteristics (e.g., characteristics that may differ from the characteristics of the input data).

[0267] In some cases, federated learning can be used to train the techniques of the present disclosure. Federated learning can enable multiple remote clinicians to iteratively improve machine learning models (e.g., validation of 3D oral care representations, mesh segmentation, mesh cleaning, other techniques involving labeling mesh elements, coordinate system prediction, placement of non-organic objects on teeth, appliance component generation, dental restoration design generation, techniques for placing 3D oral care representations, setting prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, estimation of missing values), while protecting data privacy (e.g., clinical data may not need to be sent "over the network" to a third party). Data privacy is particularly important for clinical data protected by applicable laws. A clinician can receive a copy of the machine learning model, use a local machine learning program to further train the ML model using locally available data from a local clinic, and then send the updated ML model back to a central hub or third party. The central hub or third party can integrate the updated ML models from multiple clinicians into a single updated ML model that benefits from the learning of patient data recently collected at various clinical sites. In this way, a new ML model can be trained that benefits from additional and updated patient data (possibly from multiple clinical sites), while that patient data is never actually sent to a third party. In some cases, training on devices within a local clinic can be performed when the device is idle or otherwise during non-working hours (e.g., when patients are not being treated at the clinic). Devices in a clinical environment for collecting data and / or training an ML model for the techniques described herein can include intraoral scanners, CT scanners, X-ray machines, laptop computers, servers, desktop computers, or handheld devices (such as a smartphone with image collection capabilities). In addition to federated learning techniques, in some embodiments, contrastive learning can be used to at least partially train the ML models described herein. In some cases, contrastive learning can augment the samples in a training dataset to emphasize the differences between samples of different classes and / or increase the similarity of samples of the same class.

[0268] In some embodiments, the automated setup prediction model can be trained to generate setups with a customized SPEE curve (e.g., an SPEE curve that conforms to the expected outcome of patient treatment). Such a model can be trained on cohort patient case data. One or more oral care metrics can be calculated on each case to quantify or measure aspects of the SPEE curve of that case. During training, one or more of such metrics can be provided to the setup prediction model, for example to influence the model regarding the geometry and / or structure of the SPEE curve for each case. When deploying the setup prediction model, the same input path to the trained neural network can be configured with one or more values as instructions for the model regarding the expected SPEE curve. Such values can automatically generate setups with an SPEE curve that meets the aesthetic and / or medical treatment needs of a specific patient case.

[0269] In some embodiments, the SPEE curve metric can measure the curvature of the occlusal or incisal surface of teeth on the left or right side of the dental arch relative to the occlusal plane. In some cases, the occlusal plane can be calculated as the surface that averages the incisal or occlusal surfaces of the teeth (for one or both dental arches). In some embodiments, the curvature metric can be calculated along a normal vector (such as a vector perpendicular to the occlusal plane). In other embodiments, the curvature metric can be calculated along the normal vector of another plane. In some embodiments, the XY plane can be defined as corresponding to the occlusal plane. The orthogonal plane can be defined as the plane orthogonal to the occlusal plane that also passes through an SPEE curve segment, where the SPEE curve segment is defined by a first endpoint and a second endpoint, the first endpoint being a landmark point on the first tooth (e.g., a canine) and the second endpoint being a landmark point on the last tooth on the same side of the dental arch. In some embodiments, the landmark points can be located along the incisal edge of the tooth or on the cusp of the tooth. In some cases, the landmark points of the intermediate teeth (e.g., the teeth located between the first and last teeth) on the left or right side of the dental arch can form a curved path, such as can be described by a polyline. The following is a non-limiting list of SPEE curve oral care metrics.

[0270] 1) Measure the vertical height between a line segment and a point. In other words, measure the distance between the line segment and the point along the z-axis. The line segment is defined by connecting the highest cusp of the last tooth (in the lower dental arch) and the cusp of the first tooth (in the lower dental arch) on that side. Given a subgroup of teeth between the first tooth and the last tooth, the point is defined by the highest cusp of the lowest tooth in that subgroup. In other words, the following 4 steps can be used to calculate the Spee curve metric. i) Line: Form a line between the highest cusp on the last tooth and the cusp of the first tooth. ii) Curve_Point_A: Given the set of teeth between the last tooth and the first tooth, find the highest point of the lowest tooth. iii) Curve_Point_B: Project Curve_Point_A onto the line to find the point on the line closest to Curve_Point_A (Curve_Point_B). iv) Spee curve: Find the height difference between Curve_Point_B and Curve_Point_A.

[0271] 2) Project one or more intermediate landmark points (e.g., points on the teeth that are between the first tooth and the last tooth on that side of the dental arch) and the Spee curve segment onto an orthogonal plane. Calculate the Spee curve metric by measuring the distance between the farthest projected intermediate point and the projected Spee curve segment. This yields a measurement of the curvature of the dental arch relative to the orthogonal plane.

[0272] 3) Project one or more intermediate landmark points and the Spee curve segment onto the occlusal plane. Calculate the Spee curve on this plane by measuring the distance between the farthest projected intermediate point and the projected Spee curve segment. This yields a measurement of the curvature of the dental arch relative to the occlusal plane.

[0273] 4) Skip the projection and calculate the distance and curvature in 3D space. Calculate the Spee curve by measuring the distance between the farthest intermediate point and the Spee curve segment. This yields a metric of the curvature of the dental arch in 3D space.

[0274] 5) Calculate the slope of the projected Spee curve segment on the occlusal plane.

[0275] 6) Calculate the slope of the projected Spee curve segment on the orthogonal plane.

[0276] The Spee curve metrics 5 and 6 can help the network reduce some more degrees of freedom when defining how the patient's dental arch curves at the back of the mouth.

[0277] Example :

[0278] Example 1. A method of generating a setup for an orthodontic alignment process, the method comprising: receiving, by a processing circuit of a computing device, a digital representation of a patient's teeth; Receive, by the processing circuit, at least one value related to customization of orthodontic treatment of a patient; form, by the processing circuit by executing a generator network, a prediction of one or more tooth movements for a setting, the generator network including one or more neural networks initially trained to predict the one or more tooth movements for the setting; and further train, by the processing circuit based on the formed prediction, the generator network to modify the generator network by performing operations including: using the generator network to predict, based on the digital representation of the patient's teeth, the one or more tooth movements for the setting, wherein the one or more tooth movements are described by at least one of a position or an orientation; using the generator network to quantify a difference between a representation of the one or more tooth movements predicted by the generator network and a representation of one or more reference tooth movements; generating a loss value based on the quantified difference; and modifying the generator network at least in part based on the loss value to form a modified generator network.

[0279] Example 2. The method according to Example 1, wherein the generator is used in combination with a machine learning model for predicting a prosthetic tooth design.

[0280] Example 3. The method according to Example 1, wherein the generator is used in combination with a machine learning model for generating at least one component or placing at least one component to create an oral care appliance.

[0281] Example 4. The method according to Example 1, wherein at least one transformation predicted by the generator is used in the generation of the orthodontic appliance.

[0282] Example 5. The method according to Example 4, wherein the orthodontic appliance is a clear tray aligner (CTA).

[0283] Example 6. The method according to Example 5, wherein the CTA is thermoformed.

[0284] Example 7. The method according to Example 5, wherein the CTA is 3D printed.

[0285] Example 8. The method according to Example 1, wherein the generator includes at least one attention mechanism.

[0286] Example 9. The method according to Example 1, wherein during the process of generating the representation of the patient's teeth, at least one of mesh pooling, mesh unpooling, mesh convolution, or mesh deconvolution is applied to the digital representation of the patient's teeth.

[0287] Example 10. The method according to Example 1, wherein the generator uses sparse processing.

[0288] Example 11. The method according to Example 1, wherein the generator is trained at least in part using end-to-end training.

[0289] Example 12. The method according to Example 1, wherein the generator is trained at least in part using representation learning.

[0290] Example 13. The method according to Example 1, wherein the generator is trained at least in part on setup data that has undergone data augmentation.

[0291] Example 14. The method according to Example 1, wherein at least one of the one or more tooth movements is encoded using a relative local tooth transformation.

[0292] Example 15. The method according to Example 1, wherein at least one of the one or more tooth movements is encoded using an absolute tooth transformation.

[0293] Example 16. The method according to Example 1, wherein the generator takes as input information related to the dental arch form.

[0294] Example 17. The method according to Example 1, wherein the generator is trained at least in part by transfer learning.

[0295] Example 18. The method according to Example 17, wherein the generator is trained at least in part by transfer learning using a neural network that is first trained on coordinate system prediction.

[0296] Example 19. The method according to Example 1, wherein the generator that has undergone at least partial training is subsequently used to train a neural network for another technique using transfer learning.

[0297] Example 20. The method according to Example 1, wherein the processing circuitry further applies at least one interproximal enamel reduction operation to at least one tooth of the digital representation of the patient's teeth, and then provides the at least one tooth to the generator network.

[0298] Example 21. The method according to Example 9, wherein the representation of the patient's teeth is invariant to at least one of rotation, scaling, or translation.

[0299] Example 22. The method according to Example 1, wherein the one or more tooth movements are achieved by one or more transformations.

[0300] Example 23. The method according to Example 22, wherein the one or more transformations take the form of at least one of the following: a transformation matrix, a translation vector, a quaternion, or at least Euler angles.

[0301] Example 24. The method according to Example 22, wherein the one or more transformations are generated by at least one of the following: an MLP, a transformer, or an encoder in the generator.

[0302] Example 25. The method according to Example 1, wherein the generator includes one or more coordinate normalization layers.

[0303] Example 26. The method according to Example 1, wherein the rotation to be applied to at least one tooth of the patient's dentition has a pivot point located at one of the following positions: the centroid of the crown, the apex of the root tip, the origin of the malocclusion transformation, or a point along the arch form close to the tooth.

Claims

1. A method for generating a setup for orthodontic alignment, the method comprising: Receiving, by a processing circuit of a computing device, a digital representation of a patient's teeth; Receiving, by the processing circuit, at least one value related to customization of an orthodontic treatment for the patient; Forming, by the processing circuit, a prediction of one or more tooth movements for the setup by executing a generator network, the generator network including one or more neural networks initially trained to predict the one or more tooth movements for the setup; And Further training, by the processing circuit, the generator network based on the formed prediction to modify the generator network by performing operations including: Using the generator network to predict the one or more tooth movements for the setup based on the digital representation of the patient's teeth, wherein the one or more tooth movements are described by at least one of position or orientation; Using the generator network to quantify a difference between a representation of the one or more tooth movements predicted by the generator network and a representation of one or more reference tooth movements; Generating a loss value based on the quantified difference; And Modifying, at least in part, the generator network based on the loss value to form a modified generator network.

2. The method according to claim 1, wherein the digital representation includes a plurality of grid elements and corresponding grid element feature vectors associated with at least one of the plurality of grid elements.

3. The method according to claim 2, the method further comprising calculating, by the processing circuit, a corresponding spatial feature corresponding to at least one of the corresponding grid element feature vectors.

4. The method according to claim 2, the method further comprising calculating, by the processing circuit, a corresponding structural feature corresponding to at least one of the corresponding grid element feature vectors.

5. The method according to claim 1, the method further comprising providing, by the processing circuit, at least one oral care metric as an input to the generator network.

6. The method according to claim 1, the method further comprising providing, by the processing circuit, at least one orthodontic protocol parameter as an input to the generator network.

7. The method according to claim 1, the method further comprising providing, by the processing circuit, at least one doctor preference as an input to the generator network.

8. The method according to claim 1, wherein the one or more neural networks of the generator network include at least one of an encoder structure, a decoder structure, an autoencoder structure, a U-Net structure, a pyramid encoder-decoder structure, or a transformer structure.

9. The method according to claim 1, the method further comprising providing, by the processing circuit, a binary flag as an input to the generator network.

10. The method according to claim 9, wherein the value of the binary flag indicates at least one of: whether a corresponding tooth is fixed, pinned, a pontic, extracted, implanted, or missing.

11. The method according to claim 1, wherein the digital representation comprises at least one of a 3D mesh, a 3D point cloud, or a voxelized representation.

12. The method according to claim 1, the method further comprising calculating a loss value by the processing circuit by comparing the predicted settings with predetermined ground truth settings, wherein the loss value calculation is at least one pairwise distance between at least one aspect of a 3D dental representation of a tooth pose represented in the predicted settings and a corresponding aspect of a representation of a corresponding tooth of a reference pose in the ground truth settings.

13. The method according to claim 1, wherein the initial training is based on a training dataset, and the training dataset is filtered based on at least one oral care metric.

14. The method according to claim 1, wherein the settings associated with the tooth movement represent a predicted setting of the patient, and the method further comprises registering the predicted setting with a corresponding ground truth setting.

15. The method according to claim 1, wherein the digital representation represents a malocclusion setting associated with the patient, and the method further comprises registering the malocclusion setting with a corresponding ground truth setting.

16. The method according to claim 1, the method further comprising the processing circuit providing information about interproximal enamel reduction as an input to the generator network.

17. The method according to claim 1, the method further comprising the processing circuit providing information about anteroposterior displacement as an input to the generator network.

18. The method according to claim 1, the method further comprising the processing circuit providing case classification as an input to the generator network.

19. The method according to claim 1, wherein the computing device is deployed in a clinical environment, and wherein the method is performed near real-time during a session with a patient.

20. A computing device for generating settings for an orthodontic alignment process, the computing device comprising: Interface hardware configured to: Receive a digital representation of a patient's teeth; And Receive at least one value related to customization of an orthodontic treatment for the patient; And A processing circuit configured to: Form a prediction of one or more tooth movements for the settings by executing a generator network, the generator network comprising one or more neural networks initially trained to predict the one or more tooth movements for the settings; And Further train the generator network based on the formed prediction to modify the generator network, wherein for further training the generator network, the processing circuit is configured to: Use the generator network to predict the one or more tooth movements for the settings based on the digital representation of the patient's teeth, wherein the one or more tooth movements are described by at least one of position or orientation; Use the generator network to quantify the difference between the representation of the one or more tooth movements predicted by the generator network and the representation of one or more reference tooth movements; Generate a loss value based on the quantified differences; and Modify the generator network at least in part based on the loss value to form a modified generator network.

Citation Information

Patent Citations

  • Method for automated generation of orthodontic treatment final setups

    US20210259808A1

  • Method for automated generation of orthodontic treatment final setups

    WO2020026117A1

  • System to generate staged orthodontic aligner treatment

    WO2021245480A1

  • Automated processing of dental scans using geometric deep learning

    WO2022123402A1

Cited By

  • Oral cavity scanning data occlusion contact point analysis method adopting deep learning

    CN121812179A

  • Method for analyzing occlusal contact points of oral scan data using deep learning

    CN121812179B