Auto-encoder for processing of 3D representations in digital oral care

The abnormal and missing parts in 3D oral care representation are handled through the autoencoder network, which solves the detection and processing problems in the prior art and achieves more efficient oral care device generation.

CN120345004APending Publication Date: 2025-07-18SOLVENTUM INTELLECTUAL PROPERTIES CO
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380086181.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-28
Filing Date
2023-12-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and handle abnormalities in 3D oral care representations, such as dental hardware and missing parts, resulting in inaccuracy and inefficiency in the generation of oral care appliances.

Method used

Using an autoencoder network, especially a variational autoencoder and a masked autoencoder, the model is trained to identify and process exceptions and missing parts in the 3D grid, and the grid is cleaned and rebuilt through the encoder-decoder structure.

Benefits of technology

Improves the accuracy and efficiency of 3D oral care representation, enables automatic detection and repair of tooth hardware, ensuring the accuracy and speed of oral care appliance generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345004A_ABST
    Figure CN120345004A_ABST
Patent Text Reader

Abstract

Systems and techniques for encoding and reconstructing a three-dimensional (3D) representation of oral care data are disclosed. The method involves receiving an input 3D representation of oral care data and an encoder that utilizes processing circuitry to execute a trained auto-encoder network. The encoder encodes the input 3D representation into a potential spatial representation having a lower dimension. The processing circuitry then executes a decoder of the trained auto-encoder network to reconstruct the potential spatial representation to generate an output 3D representation that is very similar to the initial input. To quantify the accuracy of the reconstruction, the processing circuitry calculates a reconstruction error that measures a difference between at least one mesh element of the input 3D representation and a corresponding mesh element of the output 3D representation. Efficient encoding and reconstruction of 3D representations of oral care data is achieved, facilitating improved analysis, diagnosis and processing plans in oral care applications.
Need to check novelty before this filing date? Find Prior Art

Description

Related Literature

[0001] The entire disclosure of PCT Application No. PCT / IB2022 / 057373 is incorporated herein by reference. The entire disclosure of each of the PCT applications with publication numbers WO2022123402A1, WO2021245480A1, and WO2020026117A1 is incorporated herein by reference. The entire disclosure of each of the following provisional U.S. patent applications is incorporated herein by reference: 63 / 432,627; 63 / 366,492; 63 / 366,495; 63 / 352,850; 63 / 366,490; 63 / 366,494; 63 / 370,160; 63 / 366,507; 63 / 352,877; 63 / 366,514; 63 / 366,498; 63 / 366,514; and 63 / 264,914. Technical Field

[0002] The present disclosure relates to the configuration and training of neural networks to improve the accuracy of automated cleaning and refinement operations for 3D oral care representations such as 3D triangle meshes using autoencoders. Summary of the Invention

[0003] The present disclosure describes systems and techniques for training and using one or more machine learning models, such as neural networks, to perform cleaning operations on 3D oral care representations, such as 3D triangle meshes. Such cleaning operations can include techniques for detecting anomalies in the 3D mesh, techniques for removing anomalies from the 3D mesh, and techniques for replacing or filling in missing portions of the 3D mesh (e.g., portions of the 3D mesh that create holes or rough edges by removing anomalous aspects of the mesh). These techniques can be trained to identify (e.g., using an encoder-decoder structure, such as a variational autoencoder) mesh elements that require further processing (such as removal). These techniques can be trained to fill in missing portions of the mesh (e.g., using an encoder-decoder structure, such as a masked autoencoder). The encoder-decoder structure can be trained to perform such mesh cleaning techniques, which can, in some embodiments, add one or more mesh elements to a trial 3D mesh, can, in some embodiments, remove one or more mesh elements from a trial 3D mesh, or can, in some embodiments, transform (e.g., translate, rotate, smooth, etc.) one or more mesh elements in a trial 3D mesh. The encoder-decoder structure can include at least one encoder or at least one decoder. Non-limiting examples of encoder-decoder structures include 3D U-Net, transformers, pyramid encoder-decoders, or autoencoders, etc. Non-limiting examples of autoencoders include variational autoencoders, regularized autoencoders, masked autoencoders, or capsule autoencoders.

[0004] An example of anomaly detection is described below. An encoder-decoder structure can be trained to reconstruct a 3D triangular mesh of a specific type of 3D oral care representation (e.g., a tooth, including a crown and / or a root). In some embodiments, such an encoder-decoder structure can be trained to perform the reconstruction of a specific tooth type in order to become proficient at reconstructing that specific tooth type (e.g., the upper right first molar or the lower left central incisor). After training is complete, the encoder-decoder structure can be deployed for use in digital oral care (e.g., for oral care appliance generation). A test crown mesh can be introduced into the input of a reconstruction encoder-decoder structure (e.g., a variational autoencoder that has been trained to reconstruct 3D oral care representations, examples of which are disclosed herein, optionally utilizing continuous normalizing flows) and encoded into a latent form that may subsequently undergo reconstruction. When the reconstruction autoencoder is trained to reconstruct a specific type of 3D oral care representation such as a crown, a reconstruction error can be computed for use in anomaly detection. For example, the reconstruction autoencoder can be trained to reconstruct a specific crown type. The reconstruction error can be computed to quantify the difference between the test input crown and the reconstructed crown. In the case where the reconstruction error is below a threshold, the anomaly detection model can conclude that the test input crown mesh is from the distribution of crown meshes used to train the reconstruction autoencoder. That is, if a normal upper right central incisor is presented to the input of an autoencoder that has been trained to reconstruct the upper right central incisor from a dataset of cohort patient cases, it can be reasonably expected that the reconstruction error will be less than the threshold.

[0005] However, if the input crown mesh is not from the distribution of crown meshes used to train the reconstruction autoencoder (e.g., the crown mesh is fitted with hardware such as orthodontic brackets), it can be expected that the reconstruction error will be higher than a threshold, thereby signaling the presence of an anomaly. The reconstruction error can be calculated for local portions of the reconstructed tooth mesh, even at the granularity level of mesh elements (e.g., vertices, points, faces, edges, or voxels), so that one or more sub-sections of the reconstructed tooth mesh can be marked as anomalous. In some embodiments, one or more anomalous aspects (e.g., mesh element features) may undergo modification (e.g., using mesh processing techniques known to those skilled in the art, such as mesh element removal or smoothing). In the case of orthodontic treatment, a patient's teeth can be scanned by an intraoral scanner, thereby generating 3D points. The 3D points (e.g., a point cloud) can be converted into a 3D mesh. The 3D mesh of the scanned dental arch can undergo segmentation, thereby generating a 3D mesh for each crown, which can then be used in the generation of an oral care appliance such as a clear aligner tray. Sometimes, when performing such an intraoral scan, the patient may already have hardware attached to the teeth. If the hardware is removed, in some cases, appliance generation (e.g., the generation of a clear tray aligner) can proceed more smoothly. The anomaly detection techniques described herein (e.g., an anomaly detection autoencoder) can be trained to detect hardware on a patient's teeth so that the hardware can subsequently be removed from the 3D mesh of the patient's teeth (e.g., using an automated process). In some embodiments, such an automated anomaly detection protocol can be performed with higher data accuracy than by a technician (e.g., in a manufacturing environment) or a clinician (e.g., "chairside" in a clinic, such that the results of the anomaly detection can be used as part of a mesh cleaning operation that is performed as part of an automated appliance creation operation that can be performed in the clinician's office, e.g., in the case where the appliance will be 3D printed in the doctor's office "the same day").

[0006] An example of 3D mesh filling is as follows. A masked encoder-decoder structure, such as a masked autoencoder, can be trained on examples of 3D oral care representations such as crowns and / or roots. For each training example, a mask (e.g., a randomly generated mask) can be applied that can label one or more aspects (e.g., one or more mesh elements) of the input 3D mesh. The masked autoencoder can be trained for reconstruction. In some embodiments, the masked autoencoder can ignore the masked aspects of the input 3D oral care representation (e.g., the labeled mesh elements). The masked autoencoder can be trained using many such masked examples of the 3D oral care representation (e.g., 3D meshes of crowns and / or roots) and become capable of reconstructing the input 3D oral care representation regardless of the masked aspects of the input 3D oral care representation (e.g., masked or hidden mesh elements). This process of masking the input training examples can be used to augment the training dataset to some extent. In some embodiments, the training of the autoencoder can continue until the reconstruction error of the reconstructed mesh drops below a threshold. The reconstruction error can be calculated for the entire 3D oral care representation or one or more parts of the 3D oral care representation (e.g., the reconstruction error can label or annotate a set of mesh elements as corresponding to anomalies in the reconstructed tooth mesh, such as hardware, foreign material, depressions, undercuts, internal fractures, lingual bars, etc.).

[0007] The method of the present disclosure can train an ML model (e.g., a masked autoencoder) to fill in missing aspects of a 3D representation (e.g., missing mesh elements). The masked autoencoder can be trained at least in part on data of a masked 3D representation. The masked autoencoder can be trained to act as a type of reconstruction autoencoder (e.g., which has been configured to estimate missing data). In some embodiments, a variational autoencoder (VAE) can be trained for such an estimation. A 3D representation can be masked by replacing at least one aspect of the 3D representation in a training dataset with a masked token (e.g., replacing at least one coordinate of a mesh element of the 3D representation with a masked token). For example, when a masked autoencoder is trained to fill in a missing part of a 3D mesh of a tooth (e.g., a part occluded during an intraoral scan), training examples can be generated by replacing at least one coordinate of at least one mesh element (e.g., a mesh describing other aspects of the tooth or dentition) of the input 3D representation of the oral care data with a masked token. The masked autoencoder can be trained to reconstruct the input 3D representation without considering the missing data (e.g., as a reconstruction autoencoder that fills in the missing data according to the distribution of the training dataset). In other words, the masked autoencoder can be trained to reconstruct the masked aspects of the input 3D representation of the oral care data (e.g., the masked mesh elements of the tooth). Mesh elements can include vertices, edges, faces, voxels, or points. For example, the input 3D representation of the oral care data can be a 3D mesh or a 3D point cloud. In some embodiments, a mask (e.g., which relates to the use of masked tokens) can be applied to the input 3D oral care representation before the training data is provided to the masked autoencoder. In some embodiments, the mask can be generated randomly. In some embodiments, the masked token can be provided to the masked autoencoder, and the masked autoencoder can be signaled that a masked mesh element exists in the structure of the input 3D oral care representation, but one or more of the mesh element features in the corresponding mesh element feature vector (e.g., XYZ coordinates) of the mesh element are not available to the autoencoder. In some embodiments, modifying the 3D oral care representation can include multiple mesh elements masked in consecutive blocks. In some embodiments, the input 3D representation of the oral care data represents teeth (e.g., a tooth mesh, etc.). In some embodiments, training the masked autoencoder can involve training the masked autoencoder based on the distribution of a dataset associated with the input 3D oral care mesh. In some embodiments, the masked autoencoder can include a multi-dimensional encoder configured to encode the input 3D oral care representation into a latent space representation and a multi-dimensional decoder configured to reconstruct the latent space representation into a copy of the input 3D oral care representation. In some embodiments, a reconstruction error can be calculated to quantify the difference between the input 3D oral care representation and the copy of the input 3D oral care representation. In some embodiments, the reconstruction error can be associated with at least one of a reconstruction loss calculation or a KL divergence calculation.In some specific implementations, the additional input for the masked autoencoder may include at least one of the following: (i) one or more vectors P that include at least one value related to at least one method for calculating the dimensions of at least one tooth, or (ii) one or more vectors R of at least one of tooth name, nomenclature, tooth type, and tooth classification.

[0008] The method of the present disclosure can train a reconstruction autoencoder (e.g., a variational autoencoder network) including one or more encoders and one or more decoders. The encoder of the trained reconstruction autoencoder network can encode an input 3D oral care representation into a latent space representation. The decoder of the trained autoencoder network can reconstruct the latent space representation to form an output 3D oral care representation that is a copy of the input 3D oral care representation. In some embodiments, a reconstruction error can be calculated to quantify the difference between at least one mesh element of the input 3D oral care representation and the corresponding at least one mesh element of the output 3D oral care representation. Anomaly detection can be performed based on the reconstruction error. The input 3D oral care representation can include at least one of a 3D mesh, a point cloud, or a voxelized representation. In some embodiments, the mesh element (and the corresponding mesh element) is at least one of a corresponding edge, a corresponding vertex, a corresponding face, a corresponding voxel, or a corresponding point of a point cloud. When the reconstruction error exceeds a predetermined threshold for one or more mesh elements in the input 3D representation of the oral care data, at least one mesh element or the corresponding at least one mesh element can receive a classification label. For example, when a tooth mesh with attached hardware is provided to the trained reconstruction autoencoder, the reconstruction autoencoder can generate a reconstructed tooth mesh with a high reconstruction error (e.g., the reconstructed mesh elements corresponding to the attached hardware can be labeled with a high reconstruction error because those mesh elements are not extracted from the distribution of the training dataset that the reconstruction autoencoder was trained to reconstruct). In some embodiments, calculating the reconstruction error can involve calculating the distance between at least one aspect of the input 3D representation of the oral care data and the corresponding aspect in the output 3D representation of the oral care data. In some embodiments, the trained autoencoder network can be trained according to a paradigm including continuous normalizing flows. The input 3D representation of the oral care data can include one or more 3D meshes or one or more 3D point clouds describing one or more teeth of the input 3D oral care representation to the trained autoencoder network. Additional inputs to the trained autoencoder can include: (i) one or more vectors P that include at least one value related to at least one method of calculating the size of at least one tooth, or (ii) one or more vectors R of at least one of tooth name, nomenclature, tooth type, and tooth classification. The autoencoder can be trained at least in part by calculating a loss that is associated with at least one of a term associated with reconstruction loss calculation or a term associated with KL divergence loss calculation. In some cases, the method of the present disclosure can be deployed in a clinical environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A method for enhancing training data used in training a machine learning (ML) model of the present disclosure is shown.

[0010] Figure 2Shows a method for training a capsule autoencoder.

[0011] Figure 3 Shows a method for training a tooth reconstruction autoencoder.

[0012] Figure 4 Shows a method for using a deployed tooth reconstruction autoencoder.

[0013] Figure 5 Shows a reconstructed tooth mesh that has been reconstructed using a reconstruction autoencoder according to the techniques of the present disclosure.

[0014] Figure 6 Shows a reconstructed tooth mesh that has been reconstructed using a reconstruction autoencoder according to the techniques of the present disclosure.

[0015] Figure 7 Shows a visualization of the reconstruction error of a tooth.

[0016] Figure 8 Shows the reconstruction error values for several tooth reconstructions.

[0017] Figure 9 Shows a method for training a reconstruction autoencoder.

[0018] Figure 10 Shows a non - limiting example code of a reconstruction autoencoder.

[0019] Figure 11 Shows an example of a reconstructed 3D representation according to the techniques of the present disclosure.

[0020] Figure 12 Shows a latent space in which the loss includes reconstruction loss but not KL - divergence loss.

[0021] Figure 13 Shows a latent space in which the loss includes both reconstruction loss and KL - divergence loss.

[0022] Figure 14 Shows a U - Net structure used in a denoising diffusion probability model for 3D representation segmentation.

[0023] Figure 15 Shows a method for training a capsule autoencoder to segment 3D representations.

[0024] Figure 16 Describes techniques for mesh element labeling and / or mesh cleaning.

[0025] Figure 17 Shows a method for performing anomaly detection in a 3D representation of a patient's dentition.

[0026] Figure 18A method of training a masked autoencoder to fill in missing aspects of a 3D representation is shown.

[0027] Figure 19 A method of using a trained masked autoencoder to fill in missing aspects of a 3D representation is shown.

[0028] Figure 20 A method of training a masked autoencoder to fill in missing aspects of a 3D representation is shown.

[0029] Figure 21 A method of training a masked capsule autoencoder to fill in missing aspects of a 3D representation is shown.

[0030] Figure 22 A U-Net structure that can be used to extract hierarchical features from a 3D representation is shown.

[0031] Figure 23 A pyramid encoder-decoder structure that can be used to extract hierarchical features from a 3D representation is shown. Detailed Description

[0032] Multiple techniques in digital oral care can benefit from the use of a first module (e.g., an autoencoder neural network) that has been trained to reconstruct a 3D oral care representation (e.g., trained to reconstruct a tooth mesh including crowns, roots, and / or attachments). The 3D encoder can be trained to encode the oral care mesh into a latent form, and the 3D decoder can be trained to reconstruct the latent form into a facsimile of the received oral care mesh, where the techniques disclosed herein can be used to measure the resulting reconstruction error. The first module can create a representation. The second module can use the representation for prediction. There can be one or more instances of the first module, and there can be one or more instances of the second module.

[0033] This document describes techniques that utilize an autoencoder trained for oral care mesh reconstruction, which offers the advantage of encoding a potentially complex oral care mesh into a latent form (e.g., such as a latent vector or latent capsule), which can have a reduced dimension and can be ingested by an instance of a second module (e.g., a prediction model for mesh cleaning, pose prediction, tooth restoration design generation, classification of 3D representations, validation of 3D representations, or pose comparison) for prediction purposes. Although the dimension of the latent form can be reduced relative to the received oral care mesh, information about the reconstruction characteristics of the received oral care mesh can be preserved. This latent representation of the initial oral care mesh can be received as an input to the prediction model of the second module, thus offering the advantage of improved accuracy and data precision compared to other techniques. In some specific implementations, the latent representation can be modified according to the techniques of the present disclosure to enable the prediction model of the second module to customize the output data. The advantage of calculating a reconstruction error on the reconstructed oral care mesh is to verify that the reconstructed oral care mesh is a replica of the received oral care mesh (e.g., where one or more dimensions or other aspects of the reconstructed oral care mesh are measured to be within a threshold reconstruction error of the received oral care mesh). In some specific implementations, the first module can also be trained to produce other kinds of representations, such as those generated by a neural network that performs convolutional and / or pooling operations (e.g., a network with a convolutional kernel of size 5 that also performs average pooling, or a network such as a U-Net).

[0034] Either or both of the first module and / or the second module may receive various input data as described herein, including a dental mesh of one or both dental arches of a patient. The dental data may be presented in the form of a 3D representation such as a mesh or a point cloud. Such data may be preprocessed, for example, by arranging the constituent mesh elements into a list and calculating an optional mesh element feature vector for each mesh element. Such feature vectors may provide valuable information about the shape and / or structure of the oral care mesh to either or both of the first module and / or the second module. For example, the first module that generates the representation may receive the vertices of the 3D mesh (or 3D point cloud) and calculate a mesh element feature vector for each vertex. In addition to the other optional mesh element features described herein, such a feature vector may also include the XYZ coordinates of each vertex. Additional inputs may be received at an entry point of either or both of the first module and / or the second module, such as one or more oral care metrics. Oral care metrics may be used to measure one or more physical aspects of the oral care mesh (e.g., physical relationships within or between teeth). In some cases, the oral care metrics may be calculated for either or both of a malocclusion oral care mesh example and a benchmark oral care mesh example and then used in the training of either or both of the first module and / or the second module. The metric values may be received as an input to either or both of the first module and / or the second module, thereby serving as a way to train the underlying model of that particular module to encode the distribution of such metrics over a number of examples in the training dataset. During training, the network may then receive the metric values as an input to help train the network to link the metric values of the input to the physical aspects of the ground truth oral care mesh used in the loss calculation. Such a loss calculation may quantify the difference between the prediction and the ground truth example (e.g., between the predicted oral care mesh and the ground truth oral care mesh). By providing network data that describes the metric values, the techniques of the present disclosure may train the network to encode the distribution of a given metric through the process of loss calculation and subsequent backpropagation. In deployment, one or more oral care parameters (protocol parameters or prosthetic design parameters) may be defined to specify one or more aspects of the expected oral care mesh that will be generated by either or both of the first module and / or the second module that has been trained for that purpose. In some embodiments, an oral care parameter corresponding to an oral care metric may be defined, which may be received as an input to either or both of the deployed first module and / or the deployed second module and be regarded as an instruction to that module to generate an oral care mesh with a specified customization. This interaction between oral care metrics and oral care parameters may also apply to the training and deployment of other predictive models in oral care.

[0035] In some specific implementations, the prediction model of the present disclosure can obtain more accurate results by combining one or more of the following inputs: arch form information V, interproximal reduction (IPR) information U, tooth size information P, diastema information Q, latent capsule representation T of the oral care mesh, latent vector representation A of the oral care mesh, protocol parameter K (which can describe the clinician's expected treatment of the patient), doctor preference L (which can describe the typical protocol parameters selected by the doctor), flag M regarding tooth status (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name / dental symbol R, oral care metric S (including at least one of oral care metrics and prosthetic design metrics).

[0036] Some specific implementations of the autoencoder-based mesh cleaning technique described herein can be trained to remove (or modify) common triangular mesh defects such as: degenerate triangles with zero surface area; redundant triangles that cover the same surface area as another triangle; non-manifold edges with more than two adjacent triangles, also known as "flaps"; non-manifold vertices with more than one adjacent connected triangle sequence (triangle fan); intersecting triangles (where two triangles pass through each other); spikes - sharp features composed of multiple triangles, typically conical, caused by one or more vertices deviating from the actual surface; folds (sharp features composed of multiple triangles, typically Z-shaped with small undercut regions, caused by one or more vertices deviating from the actual surface); islands / sub-components, which represent disconnected objects where the scan should contain only a single object (e.g., typically smaller objects are deleted); small holes in the mesh surface, either from the initial scan or from deletions due to previous defects (e.g., the holes can be removed by filling the holes, e.g., by adding one or more mesh elements); rough boundaries - smooth boundaries are beneficial for extending the gingival surface and creating the model base.

[0037] Some specific implementations of the autoencoder-based mesh cleaning techniques described herein can be trained to remove (or modify) aspects of the mesh that are not needed and / or are domain-specific defects in certain situations, such as: foreign materials (parts of an intraoral scan outside the anatomical region of interest, e.g., non-tooth surfaces not within a certain distance of a tooth surface, or scan artifacts that do not represent actual anatomical structures); concavities - recesses in the surface (e.g., which can be scan artifacts that should be repaired or anatomical features that are normally left intact); undercuts (tooth sides that are smaller than the crown radius, so that physical impressions or appliances may be difficult to remove or place). Undercuts can be natural features or caused by damage such as internal fractures. Other features for which the cleaning operations of the present disclosure may be needed include internal fractures (tooth erosion near the gingival line, causing or exacerbating undercuts); appliances - orthodontic hardware such as attachments, brackets, wires, buttons, lingual bars, Carriere appliances, etc., which may be present in an intraoral scan. In some cases, digital removal and replacement with synthetic tooth / gum surfaces may be needed before any subsequent appliance creation steps are carried out.

[0038] In some cases, the systems of the present disclosure can be deployed in a clinical environment (such as a dental or orthodontic clinic) for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians). Such systems deployed in a clinical environment can enable clinicians to process oral care data (such as tooth scans) in a clinical environment or in some cases in a "chairside" environment (when the patient is in the clinical environment). A non-limiting list of examples of techniques can include: segmentation, mesh cleaning, coordinate system prediction, CTA trim line generation, restoration design generation, appliance component generation or placement or assembly, generation of other oral care meshes, verification of oral care meshes, setup prediction, removal of hardware from tooth meshes, placement of hardware on teeth, estimation of missing values, clustering of oral care data, oral care mesh classification, setup comparison, metric calculation or metric visualization. In some cases, the execution of these techniques can enable patient data to be processed, analyzed, and used by clinicians in appliance creation before the patient leaves the clinical environment (this can facilitate treatment planning as feedback can be received from the patient during the treatment planning process).

[0039] The systems of the present disclosure can automate operations in digital orthodontics (e.g., setup prediction, hardware placement, setup comparison), digital dentistry (e.g., prosthetic design generation), or combinations thereof. Some techniques can be applied to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleaning, coordinate system prediction, oral care mesh validation, estimation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows, or denoising diffusion models), metric visualization, appliance component placement, or appliance component generation, etc. In some cases, the systems of the present disclosure can enable clinicians or technicians to process oral care data (such as scanned dental arches). In addition to segmentation, mesh cleaning, coordinate system prediction, or validation operations, the systems of the present disclosure can also implement orthodontic treatment planning, which may involve setup prediction as at least one operation. The systems of the present disclosure can also implement prosthetic design generation, in which one or more restored tooth designs are generated and processed during the creation of an oral care appliance. The systems of the present disclosure can implement either or both of orthodontic or dental treatment planning, or can implement automated steps in the generation of either or both of orthodontic or dental appliances. Some appliances can implement both dental and orthodontic treatments, while other appliances can implement one or the other.

[0040] Aspects of the present disclosure can provide technical solutions to technical problems of using one or more encoder-decoder architectures to label one or more mesh elements of a 3D representation of a patient's dentition (e.g., used in segmentation or cleaning of the 3D representation, or for anomaly detection). That is, by practicing the techniques disclosed herein, computing systems specifically adapted to perform mesh element labeling for anomaly detection of a 3D representation of a patient's dentition are improved. Additionally, techniques for filling holes or missing mesh elements in a 3D representation of a patient's dentition are improved. For example, aspects of the present disclosure improve the performance of a computing system having a 3D representation of a patient's dentition by reducing the consumption of computing resources. Specifically, aspects of the present disclosure reduce computing resource consumption by subsampling the 3D representation of the patient's dentition (e.g., reducing the count of mesh elements describing aspects of the patient's dentition) such that computing resources are not wasted unnecessarily due to processing an excessive number of mesh elements. Additionally, subsampling the mesh does not reduce the overall prediction accuracy of the computing system (and can actually improve prediction because the input provided to the ML model after subsampling is a more accurate (or better) representation of the patient's dentition). For example, unimportant (and potentially accuracy-degrading) noise or other artifacts are removed. That is, aspects of the present disclosure provide a more efficient allocation of computing resources in a manner that improves the accuracy of the underlying system.

[0041] In addition, aspects of the present disclosure may need to be performed in a time-limited manner, such as when an oral care appliance must be generated for a patient immediately after an intraoral scan (e.g., when the patient is waiting in a clinician's office). Accordingly, aspects of the present disclosure are necessarily rooted in underlying computer technologies for mesh element labeling and filling in missing portions of a patient's dentition (e.g., including hundreds of thousands of mesh elements) via an encoder-decoder architecture (e.g., a variational autoencoder or a masked autoencoder), and cannot be performed by a human even with the aid of pen and paper. For example, specific implementations of the present disclosure must be able to: 1) store thousands or millions of mesh elements of a patient's dentition in a manner processable by a computer processor; 2) perform calculations on thousands or millions of mesh elements, such as to quantify aspects of the shape and / or structure of an individual tooth in a 3D representation of a patient's dentition; 3) encode a patient's dentition into a latent form; 4) reconstruct the latent form into one or more 3D representations; and 5) fill in missing portions of a patient's dentition (e.g., involving generating hundreds or thousands of mesh elements); or 6) perform anomaly detection in a 3D representation of a patient's dentition, and do so during the course of a short clinic visit.

[0042] The present disclosure relates to digital oral care encompassing the fields of digital dentistry and digital orthodontics. The present disclosure generally describes methods of processing three-dimensional (3D) representations of oral care data. It should be understood that while not losing generality, there are various types of 3D representations. One type of 3D representation is 3D geometry. A 3D representation may include, be one or more of, or be a part of: a 3D polygon mesh, a 3D point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels for sparse processing), or a 3D representation described by a mathematical equation. Although the term "mesh" is frequently used throughout the present disclosure, in some specific implementations, the term should be understood to be interchangeable with other types of 3D representations. A 3D representation may describe elements of the 3D geometry and / or 3D structure of an object.

[0043] The dental arches S1, S2, S3, and S4 all include exactly the same dental meshes, which are transformed differently according to the following description. The first dental arch S1 includes a set of dental meshes that are arranged (e.g., using a transformation) in their positions in the oral cavity, where the teeth are in a malposition and orientation. The second dental arch S2 includes the same set of dental meshes from S1 that are arranged (e.g., using a transformation) in their positions in the oral cavity, where the teeth are in a reference true set position and orientation. The third dental arch S3 includes the same meshes as S1 and S2 that are arranged (e.g., using a transformation) in their positions in the oral cavity, where the teeth are in a predicted final set pose (e.g., as predicted by one or more techniques of the present disclosure). S4 is the counterpart of S3, where the teeth are in a pose corresponding to one of several intermediate stages of orthodontic treatment with a clear aligner appliance.

[0044] It should be understood that, without loss of generality, the techniques of the present disclosure applied to the final set are also applicable to intermediate gradings in orthodontic treatment, specifically geometric deep learning (GDL) settings, reinforcement learning (RL) settings, variational autoencoder (VAE) settings, capsule settings, multi-layer perceptron (MLP) settings, diffusion settings, pose transfer (PT) settings, similarity settings, force-directed graph (FDG) settings, transformer settings, setting comparison, or setting classification. The metric visualization aspect of the present disclosure can also be configured to visualize data from both the final set and intermediate stages. The MLP setting, VAE setting, and capsule setting each fall within the scope of the autoencoder setting. Some specific implementations of the MLP setting can fall within the scope of the transformer setting. A representation setting refers to any one of the MLP setting, VAE setting, capsule setting, and any other setting prediction machine learning model that uses an autoencoder to create a representation of at least one tooth.

[0045] Each of the setting prediction techniques of the present disclosure is applicable to the manufacture of clear aligner appliances and / or indirectly bonded trays. The setting prediction techniques can also be applicable to other products that also involve the final tooth pose. The pose can include position (or location) and rotation (or orientation).

[0046] A 3D mesh is a data structure that can describe the geometry and / or shape of an object related to oral care, which object includes but is not limited to teeth, hardware elements, or the gingival tissue of a patient. The 3D mesh can include one or more mesh elements, such as vertices, edges, faces, and one or more of their combinations. In some specific implementations, the mesh elements can include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features can be calculated for these mesh elements and provided to the prediction model of the present disclosure, and the prediction model of the present disclosure provides a technical advantage of improved data accuracy in the form of a more accurate prediction of the model output of the present disclosure.

[0047] The dentition of a patient may include one or more 3D representations of the patient's teeth (e.g., and / or associated transducers), gums, and / or other oral anatomical structures. In some embodiments, an orthodontic metric (OM) may quantify the relative position and / or orientation of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. In some embodiments, a restorative design metric (RDM) may quantify at least one aspect of the structure and / or shape of a 3D representation of a tooth. In some embodiments, an orthodontic landmark (OL) may locate one or more points or other regions of interest structures on a 3D representation of a tooth. In some embodiments, an OL may be used in the generation of orthodontic or prosthetic appliances, such as clear tray aligners or prosthetic restorative appliances. In some embodiments, a mesh element may include at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth represented by a 3D mesh, the mesh elements may at least include: vertices, edges, faces, and voxels. In some embodiments, mesh element features may quantify some aspects of the 3D representation that are proximal to or associated with one or more mesh elements, as described elsewhere in the present disclosure. In some embodiments, an orthodontic procedure parameter (OPP) may specify at least one value that defines at least one aspect of a patient's planned orthodontic treatment (e.g., specifying desired target attributes of a final setup in a final setup prediction). In some embodiments, an orthodontist preference (ODP) may specify at least one typical value of an OPP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. In some embodiments, a restorative design parameter (RDP) may specify at least one value that defines at least one aspect of a patient's planned prosthetic restorative treatment (e.g., specifying desired target attributes of a tooth to be treated with a prosthetic restorative appliance). In some embodiments, a doctor restorative design preference (DRDP) may specify at least one typical value of an RDP, which in some cases may be derived from past cases that have been treated by one or more oral care practitioners. The 3D oral care representation may include but is not limited to: 1) a set of mesh element labels that may be applied to 3D mesh elements of a tooth / gum / hardware / appliance mesh (or point cloud) during mesh segmentation or mesh cleaning; 2) one or more 3D representations of teeth / gums / hardware / appliances whose shapes have been modified (e.g., trimmed, deformed, or filled) during mesh segmentation or mesh cleaning; 3) one or more coordinate systems (e.g., describing one, two, three, or more coordinate axes) for a single tooth or a group of teeth (such as a full dental arch, e.g., the LDE coordinate system); 4) 3D representations of one or more teeth whose shapes have been modified or otherwise made suitable for use in prosthetic restorations; 5) 3D representations of one or more prosthetic restorative appliance components;6) One or more transformations to be applied to one or more of the following: placement of prosthetic appliance library components relative to one or more teeth, teeth to be placed for an orthodontic setting (final setting or intermediate stage), hardware elements to be placed relative to one or more teeth, etc.; 7) Orthodontic settings; 8) 3D representations of hardware elements (such as facebows, lingual bows, orthodontic attachments, buttons, hooks, occlusal ramps, etc.) placed relative to one or more teeth, etc.; 8) 3D representations of bonding pads for hardware elements (which can be generated for specific teeth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting the tooth through a Boolean operation); 9) 3D representations of clear tray appliances (CTAs); 10) The position or shape of CTA trim lines (e.g., described as a grid or polyline); 11) Arch forms (e.g., described as 3D polylines or 3D grids or surfaces) that describe the contour or layout of the dental arch, which can follow the incisal edges of one or more teeth, which can follow the facial surfaces of one or more teeth, which in some specific implementations can correspond to malocclusion dental arches and in other specific implementations correspond to final setting dental arches (the influence of malocclusion on the shape of the arch form can be reduced by smoothing or averaging the shape of the arch form), which can be described by one or more control points and / or splines; 12) 3D representations of jig models (e.g., depictions of teeth and gums used in thermoformed clear tray appliances, or depictions of teeth / gums / hardware used in thermoformed indirect bonding trays); 13) One or more latent space vectors (or latent capsules) generated by the 3D encoder stage of a 3D autoencoder (e.g., a variational autoencoder trained for tooth reconstruction) that has been trained on the reconstruction of an oral care mesh; 14) One or more oral care metrics for one or more teeth (e.g., such as orthodontic metrics or prosthetic design generation metrics); 15) One or more landmarks (e.g., 3D points) that describe the shape and / or geometric properties of one or more teeth, other dentition structures, or hardware structures (e.g., to be used for orthodontic setting creation or prosthetic appliance component generation or placement); 16) 3D representations created by scanning (e.g., optical scanning, CT scanning, or MRI scanning) 3D printed parts (such as scanned jig models) corresponding to one or more teeth / gums / hardware / appliances; 17) 3D printed appliances (optionally including local thickness, reinforcement rib geometry, tab positioning, etc.); 18) 3D representations of a patient's dentition captured by a clinician or healthcare practitioner at the chairside (e.g., in an environment where the 3D representation is verified at the chairside, before the patient leaves the clinic, such that errors can be detected and rescanning can be performed as needed); 19) Prosthetic tooth designs (e.g., for veneers, crowns, bridges, or prosthetic appliances); 20) 3D representations of one or more teeth used in digital oral care processes; 21) Other 3D printed parts belonging to oral care procedures or other fields; 22) IPR cutting surfaces;23) One or more orthodontic setting transformations associated with one or more IPR cutting surfaces; 24) (Digital) pontic design that can fill at least a portion of the space between teeth to create space for erupting teeth in an orthodontic setting and then emerge from the gums; or 25) Components of a jig model (e.g., including jig model components such as interdental bands, occlusal locks, occlusal ramps, interdental reinforcements, gingival ridges, torque points, power ridges, pontics, or pits, etc.).;

[0048] The techniques of the present disclosure may require a training dataset of hundreds or thousands of cohort patient cases to ensure that the neural network can encode the distribution of patient cases that may be encountered in clinical treatment. Cohort patient cases may include a set of crown meshes, a set of root meshes, or a data file (e.g., a JSON file) including case attributes. Typical examples of cohort patient cases may include up to 32 crown meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), multiple gingival meshes (e.g., each of which may include tens of thousands of vertices or tens of thousands of faces), or one or more JSON files (each of which may include tens of thousands of values (e.g., objects, arrays, strings, real values, boolean values, or null values)).

[0049] The techniques of the present disclosure can be advantageously combined. For example, a setting comparison tool can be used to compare the output of a GDL setting model with ground truth data, compare the output of an RL setting model with ground truth data, compare the output of a VAE setting model with ground truth data, and compare the output of an MLP setting model with ground truth data. By comparing each of these setting prediction models with ground truth data, it can be determined which model achieves the best performance on a certain dataset or within a given problem domain. Additionally, a metric visualization tool can provide a global view of the final settings and intermediate stages generated by one or more of the setting prediction models, with the advantage of being able to select the best setting prediction model. Moreover, the metric visualization tool enables the calculation of metrics with a global scope within a set of intermediate stages. In some specific implementations, these global metrics can be consumed as inputs to a neural network for predicting settings (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, etc.). The global metrics can also be provided to FDG settings. In some specific implementations, local metrics from the present disclosure (i.e., local metrics are metrics that can be calculated for one stage or setting of the treatment rather than within several stages or settings) can be consumed by the neural networks herein for predicting settings, with the advantage of improving the prediction results. In some specific implementations, the metrics described in the present disclosure can be visualized using the metric visualization tool.

[0050] VAE and MAE models for mesh element tagging and mesh filling can be advantageously combined with a setup prediction neural network for mesh cleanup before or during the prediction process. In some specific implementations, the VAE for mesh element tagging can be used to tag mesh elements for further processing, such as metric calculation, removal, or modification. In some cases, such tagged mesh elements can be provided as input to the setup prediction neural network to inform the neural network of important mesh features, properties, or geometries, with the advantage of improving the performance of the resulting setup prediction model. In some specific implementations, mesh filling can make the geometry of the teeth closer to complete, enabling the setup prediction model to function better (i.e., improving the correctness of the prediction due to the better-formed geometry). In some cases, a neural network for classifying setups (i.e., a setup classifier) can assist the setup prediction neural network in functioning because the setup classifier tells the setup prediction neural network when a predicted setup is acceptable for use and can be provided to a method for generating an orthodontic tray. Setup classifiers (e.g., GDL setups, RL setups, VAE setups, capsule setups, MLP setups, diffusion setups, PT setups, similarity setups, and FDG setups, etc.) can help generate the final setup and also help generate intermediate stages. Additionally, the setup classifier neural network can be combined with a metric visualization tool. In other specific implementations, the setup classification neural network can be combined with a setup comparison tool (e.g., the setup comparison tool can output an indication of how a setup generated partly by the setup classifier compares to a setup generated by another setup prediction method). In some specific implementations, the VAE for mesh element tagging can identify one or more mesh elements used in metric calculation. The resulting metric output can be visualized by a metric visualization tool.

[0051] In some examples, the setup classifier neural network can assist the setup prediction techniques described in U.S. Patent Application No. US20210259808A1, the entire content of which is incorporated herein by reference, or PCT Application Publication No. WO2021245480A1, the entire content of which is incorporated herein by reference, or the setup prediction techniques described in PCT Application No. PCT / IB2022 / 057373, the entire content of which is incorporated herein by reference. The setup classifier will help one or more of those techniques know when the predicted final setup is closest to being correct. In some cases, the setup classifier neural network can output an indication of how far a given setup is from the final setup (i.e., a progress indicator).

[0052] In some specific implementations, the latent space embedding vectors from the reconstructed VAE can be cascaded with the inputs of the setup prediction neural network described in WO2021245480A1. The latent space vectors can also be combined as inputs into other setup prediction models: GDL setup, RL setup, VAE setup, capsule setup, MLP setup, and diffusion setup, etc. The advantage is to endow the neural network with reconstruction characteristics (e.g., the latent vector dimension of the dental mesh), thereby improving the generated setup prediction.

[0053] In some examples, the various setup prediction neural networks of the present disclosure can work together to generate the setups required for orthodontic treatment. For example, the GDL setup model can generate the final setup, and the RL setup model can use the final setup as an input to generate a series of intermediate stage setups. Alternatively, the VAE setup model (or the MLP setup model) can create the final setup, which can be used by the RL setup model to generate a series of intermediate stage setups. In some specific implementations, the setup prediction can be generated by one setup prediction neural network and then used as an input for another setup prediction neural network for further improvement and adjustment. In some specific implementations, such improvements can be performed in an iterative manner.

[0054] In some specific implementations, a setup verification model may be involved in this iterative setup prediction loop, such as the model disclosed in U.S. Provisional Application No. US63 / 366495. First, setups can be generated (e.g., using models trained for setup prediction, such as GDL setup, RL setup, VAE setup, capsule setup, MLP setup, diffusion setup, PT setup, similarity setup, and FDG setup, etc.), and then the setups are verified. If the setup passes the verification, the setup can be output for use. If the setup does not pass the verification, the setup can be sent back to one or more of the setup prediction models for correction, improvement, and / or adjustment. In some cases, the setup verification model can output an indication of what is wrong with the setup, such that the setup generation model can be improved in the next iteration. The process is iterated until completion.

[0055] Generally, in some specific implementations, two or more of the following techniques of the present disclosure can be combined during orthodontic and / or dental treatment: GDL setting, setting classification, reinforcement learning (RL) setting, setting comparison, autoencoder setting (VAE setting or capsule setting), VAE grid element labeling, masked autoencoder (MAE) grid filling, multi-layer perceptron (MLP) setting, metric visualization, estimation of missing oral care parameter values, tooth classification using latent vectors, FDG setting, pose transfer setting, prosthetic design metric calculation, neural network techniques for dental restoration and orthodontics (e.g., generation or modification of 3D oral care representations using transformers), landmark-based (LB) setting, diffusion setting, estimation of tooth movement protocols, capsule autoencoder segmentation, diffusion segmentation, similarity setting, validation of oral care representations (e.g., using autoencoders), coordinate system prediction, prosthetic design generation or generation or modification of 3D oral care representations using denoising diffusion models.

[0056] In some cases, a shape-based input can be provided to a neural network for setting prediction. In other cases, a non-shape-based input, such as tooth name or nomenclature, can be used as it is related to dental notation. In some specific implementations, a vector R of flags can be provided to the neural network, where a "1" value indicates the presence of a tooth and a "0" value indicates the absence of a tooth in a patient case (although other values are possible). The vector R can include one-hot vectors, where each element in the vector corresponds to a tooth type, name, or nomenclature. Identification information about the teeth (e.g., the name of the tooth) can be provided to the prediction neural network of the present disclosure, which has the advantage of enabling the neural network to be trained to handle different teeth in a tooth-specific manner. For example, a setting prediction model can learn to make setting transformation predictions for a specific tooth name (e.g., the upper right central incisor or the lower left canine, etc.). In the case of a grid cleaning autoencoder (for labeling grid elements or for filling missing grid data), the autoencoder can be trained in this way to provide specialized processing to a tooth based on the tooth's nomenclature. In the case of a setting classification neural network, a list of tooth names present in a patient's dental arch can better enable the neural network to output an accurate determination of the setting classification, as tooth nomenclature is a valuable input for training such a neural network. For example, tooth nomenclature / name can be defined according to a universal numbering system, a Palmer quadrant system, or an FDI World Dental Federation notation (ISO 3950).

[0057] In one example, in the case where all teeth are present except (at most four) wisdom teeth, the vector R can be defined as an optional input to the setting prediction neural network of the present disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth, and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, UL1, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7.

[0058] In some cases, the position of the cusp can be provided to the neural network for setting prediction. In other cases, one or more vectors S of the orthodontic metrics described elsewhere in the present disclosure can be provided to the neural network for setting prediction. The advantage is that the ability of the network to be trained to understand the state of the malocclusion setting is improved, and thus it can predict a more accurate final setting or intermediate stage.

[0059] In some specific embodiments, the neural network can take as input one or more indications of interproximal reduction (IPR) U, which can indicate the amount of enamel to be removed from a tooth (from mesial or from distal) during the process of orthodontic treatment. In some specific embodiments, the IPR information (e.g., the amount of IPR to be performed on one or more teeth, measured in millimeters, or one or more binary markers indicating whether IPR is to be performed on each tooth identified by a label) can be concatenated with the latent vector A generated by a VAE or a latent capsule autoencoder. The vector and / or capsule resulting from such concatenation can be provided to one or more of the neural networks of the present disclosure, which has the technical improvement or additional advantage of enabling the prediction neural network to consider the IPR. IPR is particularly relevant to the setting prediction method, which can determine the position and posture of teeth at the end of the treatment or during one or more stages of the treatment. It is very important to consider the amount of enamel to be removed before the predicted tooth movement.

[0060] In some specific embodiments, one or more protocol parameters K and / or doctor preference vectors L can be introduced into the setting prediction model. In some specific embodiments, one or more optional vectors or values include: tooth position N (e.g., XYZ coordinates in local or global coordinates of the tooth), tooth orientation O (e.g., pose, such as in a transformation matrix or quaternion, Euler angles or other forms described herein), tooth size P (e.g., length, width, height, perimeter, radius, diagonal measurement, volume, any size can be normalized compared to another tooth or teeth), distance Q between adjacent teeth. In some cases, these "tooth sizes P" can be used to describe the expected size of the tooth for dental restoration design generation.

[0061] In some specific embodiments, the tooth dimension P (e.g., length, width, height, or perimeter) can be measured in a plane, such as a plane intersecting the centroid of the tooth, or a plane intersecting a central point that is the midpoint between the centroid of the tooth and the most incisal extent or the most gingival extent. The tooth height dimension can be measured as the distance from the gum to the incisal edge. The tooth width dimension can be measured as the distance from the mesial extent to the distal extent of the tooth. In some specific embodiments, the roundness or circularity of the tooth cross-section can be measured and included in the vector P. The roundness or circularity can be defined as the ratio of the radii of the inscribed circle and the circumscribed circle.

[0062] The distance Q between adjacent teeth can be implemented in different ways (and calculated using different distance definitions, such as Euclidean or geodesic). In some specific embodiments, the distance Q1 can be measured as the average distance between the mesh elements of two adjacent teeth. In some specific embodiments, the distance Q2 can be measured as the distance between the centers or centroids of two adjacent teeth. In some specific embodiments, the distance Q3 can be measured between the closest mesh elements between two adjacent teeth. In some specific embodiments, the distance Q4 can be measured between the tooth tips of two adjacent teeth. In some specific embodiments, teeth can be considered adjacent within an arch. In some specific embodiments, teeth can also be considered adjacent between opposing arches. In some specific embodiments, any one of Q1, Q2, Q3, and Q4 can be divided by a term to normalize the resulting value of Q. In some specific embodiments, the normalization term can involve one or more of the following: the volume of the tooth, the count of mesh elements in the tooth, the surface area of the tooth, the cross-sectional area of the tooth (e.g., as projected onto the XY plane), or some other term related to the tooth dimension.

[0063] Other information regarding the patient's dentition or treatment needs (or related parameters) can be concatenated with other input vectors to one or more of an MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and / or any neural network model listed elsewhere in this disclosure.

[0064] The vector M may include markers applied to one or more teeth. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is pinned. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is fixed. In some embodiments, M includes at least one marker for each tooth to indicate whether the tooth is a pontic. Other and additional markers are possible for teeth, such as combinations of fixed, pinned, and pontic markers. A marker set to a value indicating that a tooth should be fixed is a signal that the tooth should not move during processing and is sent to the network. In some embodiments, the neural network loss function can be designed to penalize any movement in the indicated teeth (and in some cases, may be severely penalized). A marker indicating that a tooth is a pontic notifies the network to maintain the diastema, although movement of the gap is allowed. In some cases, M may include a marker indicating tooth loss. In some embodiments, the presence of one or more fixed teeth in the dental arch can help set the prediction because the one or more fixed teeth can provide an anchor for the posture of other teeth in the dental arch (i.e., can provide a fixed reference for the posture transformation of one or more other teeth in the dental arch). In some embodiments, one or more teeth may be intentionally fixed in order to provide an anchor to which other teeth can be positioned. In some embodiments, a 3D representation (such as a mesh) corresponding to the gingiva can be introduced to provide a reference point according to which the teeth can move.

[0065] Without loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U, and V described elsewhere in this disclosure may also be provided as an input to one or more of the prediction models of this disclosure or fed into an intermediate layer thereof. Specifically, these optional vectors can be provided to the MLP setup, GDL setup, RL setup, VAE setup, capsule setup, and / or diffusion setup, with the advantage of enabling the corresponding model to output a setup that better meets the orthodontic treatment needs of the patient. In some embodiments, such an input can be provided, for example, by concatenating with one or more latent vectors A that are also provided to one or more of the prediction models of this disclosure. In some embodiments, such an input can be provided, for example, by concatenating with one or more latent capsules T that are also provided to one or more of the prediction models of this disclosure.

[0066] In some embodiments, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into a neural network (such as an MLP or a transformer) in the hidden layer of the network. In some cases, one or more of K, L, M, N, O, P, Q, R, S, U, and V can be directly introduced into the internal processing of an encoder structure.

[0067] In some specific implementations, a prediction model (such as a GDL setting, an RL setting, a VAE setting, a capsule setting, an MLP setting, a PT setting, a similarity setting, and a diffusion setting) can take as input one or more latent vectors A corresponding to one or more input oral care meshes (e.g., such as a tooth mesh). In some specific implementations, a prediction model (such as a GDL setting, an RL setting, a VAE setting, a capsule setting, an MLP setting, and a diffusion setting) can take as input one or more latent capsules T corresponding to one or more input oral care meshes (e.g., such as a tooth mesh). In some specific implementations, a prediction method can take both A and T as input.

[0068] Various loss calculation techniques generally apply to the techniques of the present disclosure (e.g., GDL setting, RL setting, VAE setting, capsule setting, MLP setting, diffusion setting, PT setting, similarity setting, setting classification, tooth classification, VAE mesh element labeling, MAE mesh filling, and estimation of procedure parameters).

[0069] These losses include L1 loss, L2 loss, mean squared error (MSE) loss, cross-entropy loss, etc. The losses can be calculated and used to train neural networks, such as multi-layer perceptrons (MLP), U-Net architectures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer architectures, etc. For example, in the learning of sequences, some specific implementations can use triplet loss or contrastive loss.

[0070] Losses can also be used to train encoder architectures and decoder architectures. KL divergence loss can be at least partially used to train one or more neural networks of the present disclosure, such as a grid reconstruction autoencoder or a generator in a GDL setting, which has the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior can enable the reconstruction autoencoder to produce better reconstructions (e.g., when modifying the latent vector representation and using the decoder to reconstruct the modified latent vector, the resulting reconstruction is more likely to be a valid instance of the input representation). There are other techniques for calculating losses that can be described elsewhere in the present disclosure. Such losses can be based on quantifying the difference between two or more 3D representations.

[0071] MSE loss calculation can involve the calculation of the average squared distance between two sets, vectors, or data sets. MSE can generally be minimized. MSE can be applied to regression problems, where the predictions generated by a neural network or other machine learning model can be real numbers. In some specific implementations, the neural network can be equipped with one or more linear activation units on the output to generate MSE predictions. According to the techniques of the present disclosure, mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used.

[0072] In some specific implementations, cross entropy can be used to quantify the difference between two or more distributions. In some specific implementations, cross entropy loss can be used to train the neural network of the present disclosure. In some specific implementations, cross entropy loss can involve comparing predicted probabilities with ground truth probabilities. Other names for cross entropy loss include "log loss", "logistic loss", and "log loss". A small cross entropy loss can indicate a better (e.g., more accurate) model. Cross entropy loss can be logarithmic. In some specific implementations, cross entropy loss can be applied to binary classification problems. In some specific implementations, the neural network can be equipped with sigmoid activation units at the output to generate probability predictions. In the case of multi-class classification, cross entropy can also be used. In this case, in some specific implementations, the neural network trained to make multi-class predictions can be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for each class to be predicted). Other loss calculation techniques that can be applied in the training of the neural network of the present disclosure include one or more of the following: Huber loss, hinge loss, classification hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss calculation methods are described herein and can be applied to the training of any neural network described in the present disclosure.

[0073] In some specific implementations, one or more neural networks of the present disclosure can be trained at least in part by a loss based on at least one of the following: pointwise mesh Euclidean distance (PMD) and Earth Mover's Distance (EMD). Some specific implementations can incorporate Hausdorff distance (HD) calculation into the loss calculation. Calculating the Hausdorff distance between two or more 3D representations (such as 3D meshes) can provide one or more technical improvements because HD not only considers the distance between two meshes, but also the way those meshes are oriented and the relationship between the mesh shapes in those orientations (or positions or poses). The Hausdorff distance can improve the comparison of two or more tooth meshes, such as two or more instances of tooth meshes in different poses (e.g., comparison of a predicted setting with a ground truth setting, which can be performed during the process of calculating the loss value for training a setting prediction neural network).

[0074] The reconstruction loss can compare the predicted output with the ground truth (or reference) output. The systems of the present disclosure can calculate the reconstruction loss as a combination of the L1 loss and the MSE loss, as shown in the following line of pseudocode: reconstruction_loss = 0.5 * L1(all_points_target, all_points_predicted) + 0.5 * MSE(all_points_target, all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the ground truth data (e.g., a ground truth dental restoration design, or some other ground truth example of a 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to the generated or predicted data (e.g., a generated dental restoration design, or some other generated example of a 3D type of oral care representation). Other specific implementations of the reconstruction loss can additionally (or alternatively) involve an L2 loss, a mean absolute error (MAE) loss, or a Huber loss term.

[0075] The entire contents of the following paper are hereby incorporated by reference in their entirety: "Attention Is All You Need"; Ashish Vaswani, Noam Shazeer, Niki Parmar, Niki Parmar, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin; NIPS 2017. The neural network-based models of the present disclosure can provide additional advantages in specific implementations where they are integrated with a neural network architecture known as a "transformer".

[0076] Prior to recently developed models such as transformer models, RNN-type models represented the state of the art for natural language processing (NLP). An example application of NLP is to generate new text based on previous words or text. Due to the important property of the transformer model with multi-head attention characteristics, the transformer then provides a significant improvement over GRU, LSTM and other such RNN-based NLP techniques. In some specific implementations, the NLP concept of multi-head attention can describe the relationship between each word in a sentence (or paragraph or document or document corpus) and each other word in the sentence (or paragraph or document or document corpus). These relationships can be generated by a multi-head attention module and can be encoded in vector form. The vector can describe how each word in a sentence (or paragraph or document or document corpus) should pay attention to each other word in the sentence (or paragraph or document or document corpus). RNN, LSTM and GRU models process sequences, such as sentences, one word at a time from the beginning to the end of the sequence. In addition, the model can only consider a given subset of sentences (called a window) when making predictions. However, in some cases, transformer-based models can take into account the entire previous text by processing the sequence as a whole in a single step. Transformers, RNNs, LSTMs, and GRU models can all be adapted for use in predictive models in digital dentistry and digital orthodontics, particularly in setting prediction tasks. In some implementations, an exemplary transformer model for use with 3D meshes and 3D transforms in setting predictions (or other oral care techniques) can be adapted based on a bidirectional encoder representation (BERT) from a transformer and / or a generative pre-trained (GPT) model. For example, a GPT (or BERT) model can first be trained on other data, such as text or document data, and then used in transfer learning. This transfer learning process can receive a previously trained GPT or BERT model and then further train it using data including a 3D oral care representation. Such transfer learning can be performed to train oral care models, such as: segmentation, mesh cleaning, coordinate system prediction, setting prediction, validation of 3D oral care representations, transformation prediction for placement of oral care meshes (e.g., teeth, hardware, appliance components, fixture model components), dental restoration design generation (or other 3D oral care representation generation, such as appliance components, fixture models, or dental arch morphology), classification of 3D oral care representations, estimation of missing oral care parameters, clustering of clinicians or clustering of clinician preferences, etc.

[0077] Oral care data may include one or more of the following (or combinations thereof): 3D representations of teeth (e.g., meshes, point clouds, or voxels), portions of a tooth mesh (such as a subset of mesh elements), tooth transformations (such as in the form of matrices, vectors, and / or quaternions, or combinations thereof), transformations for appliance components, transformations for fixture model components, and mesh coordinate system definitions (such as represented by a transformation, e.g., a transformation matrix) and / or other 3D oral care representations described herein.

[0078] A transducer can be trained to generate a transformation to position a tooth into a set pose (or position an appliance component used in appliance generation or a fixture model component used in fixture model generation). Some embodiments may operate in an offline prediction environment, and some embodiments operate in an online reinforcement learning (RL) environment. In some embodiments, the transducer may initially be trained in an offline environment and then undergo further fine-tuning training in an online environment. In the offline prediction environment, the transducer can be trained from a dataset of queued patient case data. In the online RL environment, the transducer can be trained from, for example, a physical model or a CAD model. The transducer can learn from static data, such as transformations (e.g., a trajectory transducer). In some embodiments, the transformation can provide a mapping from a malocclusion to a set (e.g., receive a transformation matrix as an input and generate a transformation matrix as an output). Some embodiments of the transducer can be trained to process 3D representations, such as 3D meshes, 3D point clouds, or voxels (e.g., using a decision transducer), taking a geometry (e.g., a mesh, point cloud, voxel, etc.) as an input and outputting a transformation. The decision transducer can be coupled to a representation generation module that encodes a representation of a patient's dentition (e.g., teeth), such as a VAE, U-Net, encoder, transducer encoder, pyramid encoder-decoder, or a simple dense or fully connected network, or combinations thereof. In some embodiments, the representation generation module (e.g., a VAE, U-Net, encoder, pyramid encoder-decoder, or dense network for generating a tooth representation) can be trained to generate a representation on one or more teeth. The representation generation module can be trained on all teeth in both dental arches, only teeth within the same dental arch (upper or lower), only anterior teeth, only posterior teeth, or some other subset of teeth. In some embodiments, such a model can be trained on each individual tooth (e.g., the right upper canine), such that the model is trained or otherwise configured to generate a highly accurate representation of an individual tooth. In some embodiments, an encoder structure can encode such a representation. In some embodiments, the decision transducer can learn in an online environment, an offline environment, or both. The online decision transducer can be trained (e.g., using RL techniques) to output actions, states, and / or rewards. In some embodiments, the transformation can be discretized to allow for segmented or stepwise actions.

[0079] In some specific implementations, the transformer can be trained to handle the embedding of the dental arch (i.e., simultaneously predict the transformations of multiple teeth) for prediction settings. In some specific implementations, the embeddings of single teeth can be cascaded into a sequence and then input into the transformer. The VAE can be trained to perform this embedding operation, the U-Net can be trained to perform such an embedding, or a simple dense or fully connected network can be trained, or a combination of these operations can be employed. In some specific implementations, the transformer-based techniques of the present disclosure can predict the movements of single teeth or can predict the movements of multiple teeth (e.g., predict the transformations of each of multiple teeth).

[0080] The 3D mesh transformer can include a transformer encoder structure (which can encode oral care data) and can be followed by a transformer decoder structure. The 3D mesh transformer encoder can encode the oral care data into a latent representation, which can be combined with attention information (e.g., by concatenating the vector of the attention information to the latent representation). In some specific implementations, the attention information can help the decoder focus on relevant oral care data during the decoding process (e.g., focus on tooth order or mesh element connectivity), such that the transformer decoder can generate a useful output for the 3D mesh transformer (e.g., an output that can be used in generating an oral care appliance). Either or both of the transformer encoder or the transformer decoder can generate the latent representation. The decoder can be used to reconstruct the output of the transformer decoder (or the transformer encoder) into, for example, one or more tooth transformations for settings, one or more mesh element labels for segmentation, a coordinate system transformation used in coordinate system generation, or one or more points of a point cloud or voxels or other mesh elements for another 3D representation. The transformer can include one or more of the following modules: a multi-head attention module, a feed-forward module, a normalization module, a linear module, and a softmax module, as well as a convolutional model for latent vector compression and / or representation.

[0081] The encoder can be stacked one or more times to further encode oral care data and enable learning of different representations of the oral care data (e.g., different latent representations). These representations can be embedded with attention information (which can affect the decoder's focus on relevant parts of the latent representation of the oral care data) and provided to the decoder in a sequential form (e.g., as a concatenation of latent representations, such as latent vectors). In some specific implementations, the encoded output of the encoder (e.g., the latent representation) can be used by downstream processing steps in the generation of the oral care appliance. For example, the generated latent representation can be reconstructed into a transformation (e.g., for placing teeth in a setting, or placing appliance components or fixture model components), or can be reconstructed into a 3D representation (e.g., a 3D point cloud, a 3D mesh, or other representations disclosed herein). In other words, the latent representation generated by the transformer (e.g., including sequentially encoded attention information) can be provided to a decoder that has been configured to reconstruct the latent representation into a specific data structure required for a particular domain region. Sequentially encoded attention information can include attention information that has undergone processing by multiple multi-head attention modules within the transformer encoder or the transformer decoder, to name just one example. Additionally, data from a specific domain can be used to compute a loss for that domain. The loss computation can train the transformer decoder to accurately reconstruct the latent representation into an output data structure relevant to the specific domain.

[0082] For example, when the decoder generates a transformation for an orthodontic setting, the decoder can be configured with an output that describes, for example, 16 real values including a 4×4 transformation matrix (other data structures for describing the transformation are possible). In other words, the latent output generated by the transformer encoder (or the transformer decoder) can be used to predict the set tooth transformation for one or more teeth to place those teeth in a set position (e.g., the final setting or an intermediate stage). Such a transformer encoder (or transformer decoder) can be trained at least in part using a reconstruction loss (or a representation loss, and other losses described herein) function that can compare the predicted transformation to a ground truth (or reference) transformation.

[0083] In another example, when the decoder generates a transformation for a tooth coordinate system, the decoder can be configured with an output that describes, for example, 16 real values including a 4×4 transformation matrix (other data structures for describing the transformation are possible). In other words, the latent output generated by the transformer encoder (or the transformer decoder) can be used to predict the local coordinate system for one or more teeth. Such a transformer encoder (or transformer decoder) can be trained at least in part using a representation loss (or a reconstruction loss, and other losses described herein) function that can compare the predicted coordinate system to a ground truth (or reference) coordinate system.

[0084] In another example, when the decoder generates a 3D point cloud (or other 3D representation, such as a 3D mesh, voxelized representation, etc.), the decoder can be configured to output a description of, for example, one or more 3D points (e.g., including XYZ coordinates). In other words, the latent output generated by the transformer encoder (or transformer decoder) can be used to predict the mesh elements for generating (or modifying) the 3D representation. Such a transformer encoder (or transformer decoder) can be trained at least in part using a reconstruction loss (or L1, L2, or MSE loss, and other losses described herein) function that compares the predicted 3D representation to a ground truth (or reference) 3D representation.

[0085] In another example, when the decoder generates mesh element labels for 3D representation segmentation or 3D representation cleaning, the decoder can be configured to output a description of, for example, the labels of one or more mesh elements. In other words, the latent output generated by the transformer encoder (or transformer decoder) can be used to predict the mesh element labels for mesh segmentation or mesh cleaning. Such a transformer encoder (or transformer decoder) can be trained at least in part using a cross-entropy loss (or other losses described herein) function that compares the predicted mesh element labels to ground truth (or reference) mesh element labels.

[0086] Multi-head attention and transformers can be advantageously applied to setting generation problems. Multi-head attention is a module in a 3D transformer encoder network that calculates attention weights for the provided oral care data and produces an output vector with encoded information on how each example of the oral care data should attend to each other piece of oral care data in the dental arch. Attention weights are a quantification of the relationships between pairs of oral care data.

[0087] A 3D representation of oral care data (e.g., including voxels, point clouds, or 3D meshes composed of vertices, faces, or edges) can be provided to a transformer. The 3D representation can depict a patient's dentition, a fixture model (or components of a fixture model), an appliance (or components of an appliance), etc. In some embodiments, the transformer decoder (or transformer encoder) can be equipped with multi-head attention. Multi-head attention can enable the transformer decoder (or transformer encoder) to attend to different parts of the 3D representation of oral care data. For example, multi-head attention can enable the transformer to attend to grid elements within a local neighborhood (or clique), or attend to global dependencies between grid elements (or cliques). For example, multi-head attention can enable a transformer for setting prediction (e.g., a transformer-based setting prediction model) to generate a transformation of a tooth and, when generating the transformation, substantially simultaneously attend to each of the other teeth in the dental arch. In other words, the transformation of each tooth can be generated based on the pose of one or more other teeth in the dental arch, resulting in a more accurate transformation (e.g., a transformation that more closely conforms to a ground truth or reference transformation). In an example of 3D representation generation (e.g., generation of a 3D point cloud), the transformer model can be trained to generate a dental restoration design. Multi-head attention can enable the transformer to attend to multiple parts of a tooth (or attend to the surfaces of adjacent teeth) as the tooth undergoes the generation process. For example, a transformer for restoration design generation can generate grid elements for the incisal edge of an incisor while at least substantially simultaneously attending to grid elements of the mesial, distal, facial, or lingual surfaces of the incisor. The result can be the generation of grid elements to form an incisal edge of the tooth that seamlessly merges with the adjacent surfaces of the tooth. This use of multi-head attention results in a more accurate modeling of the distribution of the training dataset compared to techniques that do not apply multi-head attention.

[0088] In some embodiments of the present disclosure, one or more attention vectors can be generated that describe how aspects of oral care data interact with other aspects of oral care data associated with the dental arch. In some embodiments, one or more attention vectors can be generated to describe how one or more parts of tooth T1 interact with one or more parts of tooth T2, tooth T3, tooth T4, etc. A part of a mesh can be described as a set of grid elements as defined herein. In some embodiments, the interacting parts of tooth T1 and tooth T2 can be determined, in part, by computing mesh correspondences as described herein. Any of these models (RNNs, GRUs, LSTMs, and transformers) can be advantageously applied to the task of setting transformation prediction, such as in the models described herein. Transformers can be particularly advantageous because transformers can enable the generation of transformations of multiple teeth or even an entire dental arch at once, rather than generating them individually, as may occur in some other models such as encoder architectures. In other embodiments, an attentionless transformer can be used to make predictions based on oral care data.

[0089] A particular implementation of setting the neural network model by the GDL may include a representation generation module (e.g., including a U-Net structure, an autoencoder encoder, a transformer encoder, another type of encoder-decoder structure, or an encoder, etc.), and this representation generation module may provide its output to a module that is trained to generate a tooth transformer (e.g., a set of fully connected layers with optional skip connections, or an encoder structure) to generate predictions of the transformation of each individual tooth. In some particular implementations, the skip connection may connect the output of a particular layer in the neural network to the input of another layer (e.g., a layer that is not immediately adjacent to the initial layer) in the neural network. The transformation generation module (e.g., the encoder) may process the transformation predictions one tooth at a time. Other particular implementations may replace this encoder structure with a transformer (e.g., a transformer encoder or a transformer decoder) that can process all the predictions of all teeth substantially simultaneously. In other words, the transformer can be configured to receive a much larger number of input values than some other neural network models (e.g., than a typical MLP). This is because the transformer can accommodate an increasing number of inputs, and the predictions corresponding to those inputs can be generated substantially simultaneously. The representation generation module (e.g., a U-Net structure) may provide its output to the transformer, and the transformer may generate the setting transformation for all several teeth at once, with the technical advantage of improved accuracy (since the transformation of each tooth is generated based on the transformations of each adjacent or nearby tooth, resulting in fewer conflicts and better consistency with the processing target). The transformer can be trained to output a transformation, such as a transformation encoded by a 4×4 matrix (or some other size), a quaternion, a translation vector, Euler angles, or some other form. The transformation can place the tooth into a setting pose, place the fixture model component into a pose suitable for fixture model generation, or place the appliance component into a pose suitable for appliance generation (e.g., a dental restoration appliance, a clear tray orthodontic appliance, etc.). In some particular implementations, the transformation can define a coordinate system for aspects of the patient's dentition, such as a tooth mesh (e.g., the local coordinate system of the teeth). In some particular implementations, a neural network may first be used to encode the input to the transformer (e.g., a latent representation or embedding may be generated), such as one or more linear layers and / or one or more convolutional layers. In some particular implementations, the transformer may first be trained on an offline dataset and then trained using an auxiliary actor-critic network, so that online reinforcement learning can be achieved.

[0090] In some specific implementations, a transformer can achieve large model capacity and / or implement an attention mechanism (e.g., the ability to attend to and respond to certain inputs). The attention mechanism found within a transformer (e.g., multi-head attention) enables relationships within a sequence to be encoded into neural network features. Relationships within a sequence can be encoded, for example, by associating sequence numbers (e.g., 1, 2, 3, etc.) with each tooth in a dental arch, or by associating sequence numbers with each mesh element in a 3D representation (e.g., of a tooth). In a specific implementation where latent vectors of teeth are provided to the transformer, relationships within a sequence can be encoded, for example, by associating sequence numbers (e.g., 1, 2, 3, etc.) with each element in the latent vector.

[0091] A transformer can be scaled by increasing the number of attention heads and / or by increasing the number of transformer layers. In other words, one or more aspects of a transformer can be independently trained to handle discrete tasks and later combined to allow the resulting transformer to perform all the tasks for which its individual components have been trained, without degrading the prediction accuracy of the neural network. Scaling a convolutional network can be more difficult because the model may have lower extensibility or may have lower interchangeability.

[0092] Convolution has the ability to be rotation and translation invariant, which leads to improved generalization because a convolutional model may not need to consider the way the input data is rotated or translated. A transformer has the ability to be permutation invariant because relationships within a sequence can be encoded into neural network features.

[0093] In some specific implementations for generating or modifying 3D oral care representations, a transformer can be combined with a convolutional-based neural network, such as by vertically stacking convolutional layers and attention layers. Stacking transformer blocks with convolutional blocks enables the resulting structure to have the translational invariance of convolution and also the permutation invariance of a transformer. Such stacking can improve model capacity and / or model generalization. CoAtNet is an example of a network architecture that combines convolutional and attention-based elements and can be applied to the processing of oral care data. In some cases, transfer learning can be used at least in part to train a network for modifying or generating 3D oral care representations from CoAtNet (or another model that combines convolution and self-attention / transformer).

[0094] The techniques of the present disclosure may include operations such as 3D convolution, 3D pooling, 3D transposed convolution, and 3D unpooling. 3D convolution, for example, may assist in segmentation processing when downsampling a 3D mesh. 3D transposed convolution, for example, performs the inverse operation of 3D convolution in a U-Net. 3D pooling may assist in segmentation processing, for example, in a generalized neural network feature map. 3D unpooling, for example, performs the inverse operation of 3D pooling in a U-Net. These operations may be implemented by one or more layers in a predictive or generative neural network as described herein. These operations may be applied directly to mesh elements such as mesh edges or mesh faces. These operations provide a technical improvement over other methods because these operations are invariant to mesh rotation, scaling, and translation changes. Generally speaking, these operations depend on edge (or face) connectivity, so as long as the edge (or face) connectivity is maintained, these operations are not affected by mesh changes in 3D space. That is, these operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position, or scale of the oral care mesh, which may improve data accuracy. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes and can be used for tasks such as 3D shape classification or mesh element labeling (e.g., for segmentation or mesh cleaning). MeshCNN performs these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.

[0095] In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 2D representations such as images. In some implementations of the techniques of the present disclosure, a neural network may be trained to operate on 3D representations such as meshes or point clouds. An intraoral scanner may capture 2D images of a patient's dentition from various angles. The intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data depicting the patient's dentition. According to various techniques, an autoencoder (or other neural network described herein) may be trained to operate on either or both of 2D and 3D representations.

[0096] A 2D autoencoder (including a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or latent capsule) using the 2D encoder and then reconstruct a facsimile of the input 2D image using the 2D decoder. For a handheld mobile application that has been developed for such analysis (e.g., for the analysis of dental anatomy), 2D images may be easily captured using one or more onboard cameras. In other examples, 2D images may be captured using an intraoral scanner configured for such a function. Operations that may be used in implementations of a 2D autoencoder (or other 2D neural network) for 2D image analysis are 2D convolution, 2D pooling, and 2D reconstruction error calculation.

[0097] 2D Convolution :

[0098] 2D image convolution may involve the "sliding" of a kernel across a 2D image and the calculation of element-wise multiplications, as well as summing these element-wise multiplications into output pixels. The output pixels generated from each new position of the kernel are saved into an output 2D feature matrix. In some specific implementations, adjacent elements (e.g., pixels) may be in well-defined positions in a straight-line grid (e.g., above, below, left, and right).

[0099] 2D Pooling :

[0100] A 2D pooling layer can be used to downsample a feature map and summarize the presence of certain features in that feature map.

[0101] 2D Reconstruction Error :

[0102] A 2D reconstruction error can be computed between the pixels of an input image and a reconstructed image. The mapping between pixels can be well understood (e.g., directly comparing the upper pixels [23,134] of the input image with the pixels [23,134] of the reconstructed image, assuming the two images have the same dimensions).

[0103] One of the advantages provided by the 2D autoencoder-based techniques of the present disclosure is the ease of capturing 2D image data with a handheld device. In some cases where an external data source provides data for analysis, there may be instances where only 2D image data is available. When only 2D image data is available, it is necessary to use a 2D autoencoder for analysis.

[0104] Modern mobile devices (such as commercially available smartphones) may also have the ability to generate 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera that moves around an object to capture multiple images from different views, or both), and this 3D data can be arranged in a 3D representation, such as a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation in some specific implementations. In some cases, the analysis of a 3D representation of an object can provide a technical improvement over a 2D analysis of the same object. For example, a 3D representation can describe the geometry and / or structure of an object with less ambiguity than a 2D representation (which may include shadows and other artifacts that complicate the depiction of the depth and texture of the object). In some specific implementations, 3D processing can achieve a technical improvement due to the inverse optics problem, which affects 2D representations in some cases. The inverse optics problem refers to the phenomenon that, in some cases, the size of an object, the orientation of the object, and the distance between the object and the imaging device may be combined in a 2D image of the object. Any given projection of an object on an imaging sensor can map to an infinite count of {size, orientation, distance} pairings. 3D representations achieve a technical improvement because they remove the ambiguity introduced by the inverse optics problem.

[0105] Devices configured for a specific purpose with 3D scanning, such as 3D intraoral scanners (or CT scanners or MRI scanners), can generate 3D representations of an object (e.g., a patient's dentition), which have a significantly higher fidelity and accuracy than what a handheld device might have. When such high-fidelity 3D data is available (e.g., in oral care mesh classification or in the application of other 3D techniques described herein), the use of 3D autoencoders provides technical improvements (such as increased data accuracy) to extract the best possible signal from those 3D data (i.e., obtain a signal from the 3D crown meshes used in tooth classification or setup classification).

[0106] A 3D autoencoder (including a 3D encoder and a 3D decoder) can be trained on 3D data to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder, and then reconstruct a facsimile of the input 3D representation using the 3D decoder. Operations that can be used to implement a 3D autoencoder for analyzing 3D representations (e.g., 3D meshes or 3D point clouds) are 3D convolution, 3D pooling, and 3D reconstruction error calculation.

[0107] For each mesh element, 3D convolution can be performed to aggregate local features from nearby mesh elements. Processing can be performed on top of and in addition to techniques used for 2D convolution to account for the different counts and positions of adjacent mesh elements (relative to a specific mesh element). A particular 3D mesh element can have a variable neighbor count, and those neighbors can be absent from expected positions (unlike pixels in 2D convolution, which can have a fixed adjacent pixel count present in known or expected positions). In some cases, the order of adjacent mesh elements can be relevant to 3D convolution.

[0108] 3D Pooling :

[0109] The 3D pooling operation can enable the combination of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling can iteratively reduce a 3D mesh to the mesh elements that are most highly relevant to a given application (e.g., for which a neural network has been trained). Similar to 3D convolution, 3D pooling can benefit from special processing in addition to that required in 2D convolution to account for the different counts and positions of adjacent mesh elements (relative to a specific mesh element). In some cases, the order of adjacent mesh elements may be less relevant to 3D pooling than to 3D convolution.

[0110] 3D Reconstruction Error :

[0111] The 3D reconstruction error can be calculated using one or more of the techniques described herein, such as calculating the Euclidean distance between corresponding mesh elements, between two meshes. According to aspects of the present disclosure, other techniques are possible. The 3D reconstruction error can generally be calculated on 3D mesh elements rather than 2D pixels of the 2D reconstruction error. The 3D reconstruction error can achieve a technical improvement over the 2D reconstruction error because in some cases, the 3D representation can have less ambiguity (i.e., less ambiguity in form, shape, and / or structure) than the 2D representation. In some specific implementations, due to the complexity of the mapping between the input mesh elements and the reconstructed mesh elements (i.e., the input mesh and the reconstructed mesh may have different mesh element counts, and there may be a less clear mapping between mesh elements compared to the mapping between pixels in 2D reconstruction), additional processing may be required for 3D reconstruction over and above 2D reconstruction. Technical improvements in 3D reconstruction error calculation include improved data accuracy.

[0112] The 3D representation of the mesh element feature vector can be generated using a 3D scanner such as an intraoral scanner, a computed tomography (CT) scanner, an ultrasound scanner, a magnetic resonance imaging (MRI) machine, or a mobile device capable of performing stereophotogrammetry. The 3D representation can describe the shape and / or structure of an object. The 3D representation can include one or more of a 3D mesh, a 3D point cloud, and / or a 3D voxelized representation, etc. The 3D mesh includes edges, vertices, or faces. Although in some cases these three types of data are interrelated, they are different. A vertex is a point in 3D space that defines the boundary of the mesh. These points would alternatively be described as a point cloud without additional information about how the points are connected to each other (as described by the edges). An edge is described by two points and can also be referred to as a line segment. A face is described by multiple edges and vertices. For example, in the case of a triangular mesh, a face includes three vertices that are interconnected to form three consecutive edges. Some meshes can include degenerate elements such as non-manifold mesh elements, which can be removed to benefit subsequent processing. According to aspects of the present disclosure, other mesh preprocessing operations are also possible. The 3D mesh is typically formed using triangles, but in other embodiments, quadrilaterals, pentagons, or some other n-sided polygon can be used. In some embodiments, such as in the case of performing sparse processing, the 3D mesh can be converted into one or more voxelized geometries (i.e., including voxels). The techniques of the present disclosure operating on the 3D mesh can receive one or more tooth meshes (e.g., arranged in one or more dental arches) as input. Each of these meshes can be preprocessed before being input into a prediction architecture (e.g., including at least one of an encoder, a decoder, a pyramid encoder-decoder, and a U-Net). Such preprocessing can include converting the mesh into a list of mesh elements such as vertices, edges, faces, or into voxels in the case of sparse processing. For one or more selected types of mesh elements (e.g., vertices), a feature vector can be generated. In some examples, a feature vector is generated for each vertex of the mesh. Each feature vector can include a combination of spatial features and / or structural features, as specified in the following table:

[0113] Table 1 discloses non-limiting examples of mesh element features. In some specific implementations, in addition to the spatial or structural mesh element features described in Table 1, color (or other visual cues / identifiers) may also be considered a mesh element feature. As used herein (e.g., in Table 1), a point differs from a vertex in that a point is part of a 3D point cloud, while a vertex is part of a 3D mesh and may have incident faces or edges. A dihedral angle (which may be expressed in radians or degrees) can be calculated as the angle (e.g., a signed angle) between two connected faces (e.g., two faces connected along an edge). The sign on the dihedral angle can reveal information about the convexity or concavity of the mesh surface. For example, in some specific implementations, a positively signed angle may indicate a convex surface. Additionally, in some specific implementations, a negatively signed angle may indicate a concave surface. To calculate the principal curvatures of a mesh vertex, the directional curvatures of each adjacent vertex around that vertex can be calculated first. These directional curvatures can be sorted in a circular order (e.g., 0 degrees, 49 degrees, 127 degrees, 210 degrees, 305 degrees) near the vertex normal vector and may include a subsampled form of the full curvature tensor. Circular order means sorting by angle around an axis. The sorted directional curvatures can contribute to a system of linear equations that admits a closed-form solution, which can estimate the two principal curvatures and directions, which can characterize the full curvature tensor. Consistent with Table 1, a voxel may also have features that are calculated as an aggregation of other mesh elements (e.g., vertices, edges, and faces) that either intersect the voxel or, in some specific implementations, are mainly or fully contained within the voxel. Rotating a mesh may not change the structural features but may change the spatial features. And, as described elsewhere in this disclosure, the term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non-limiting sense. In some specific implementations, in addition to mesh element features, there are alternative ways to describe the geometry of a mesh (such as 3D key points and 3D descriptors). Examples of such 3D key points and 3D descriptors can be found in "TONIONI A et al., 'Learning to detect good 3D keypoints.', Int J Comput. Vis. Vol. 126, pp. 1 - 20, 2018". In some specific implementations, 3D key points and 3D descriptors can describe the extrema (minima or maxima) of the surface of a 3D representation.In some embodiments, one or more mesh element features may be computed at least in part via deep feature synthesis (DFS), such as described in: J.M. Kanter and K. Veeramachaneni, “Deep feature synthesis: Towards automating datascience endeavors”, 2015 IEEE International Conference on Data Science andAdvanced Analytics (DSAA), 2015, pp. 1-10, doi: 10.1109 / DSAA.2015.7344858.

[0114] A neural network that generates representations based on autoencoders, U-Nets, transformers, other types of encoder-decoder architectures, convolutional and / or pooling layers, or other models may benefit from the use of mesh element features. Mesh element features may convey aspects of the surface shape and / or structure of a 3D representation to the neural network models of the present disclosure. Each mesh element feature describes different information about the 3D representation that may not redundantly exist in other input data provided to the neural network. For example, vertex curvature may quantify aspects of the concavity or convexity of the surface of a 3D representation that the network would not otherwise understand. In other words, mesh element features may provide a processed form of the structure and / or shape of a 3D representation; data that would otherwise not be available to the neural network. This processed information is generally more accessible or more suitable for encoding by the neural network. Systems implementing the techniques disclosed herein have been used to run multiple experiments on 3D representations of teeth. For example, mesh element features have been provided to a neural network that generates representations based on a U-Net model and also to a representation generation model based on a variational autoencoder with continuous normalizing flows. Based on the experiments, it was found that systems that use a full complement of mesh element features (e.g., “XYZ” coordinate tuples, “normal vectors”, “vertex curvature”, point pivots, and normal pivots) are at least 3% more accurate than systems that do not use mesh element features. A point pivot describes an “XYZ” coordinate tuple with a local coordinate system (e.g., at the centroid of the corresponding tooth). A normal pivot describes a “normal vector” with a local coordinate system (e.g., at the centroid of the corresponding tooth). Additionally, when using a full complement of mesh element features, training converges more quickly. In other words, machine learning models trained using a full complement of mesh element features tend to be faster and more accurate (at earlier epochs) than systems that do not. For an existing system that observes a historical accuracy of 91%, a 3% increase in accuracy reduces the actual error rate by more than 30%.

[0115] Prediction models that can operate on the feature vectors of the above features include, but are not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element marking, MAE grid filling, grid reconstruction autoencoders, validation using autoencoders, grid segmentation, coordinate system prediction, grid cleaning, restorative design generation, appliance component generation and placement, and dental arch form prediction. Such feature vectors can be presented as inputs to the prediction models. In some specific implementations, such feature vectors can be presented to one or more internal layers of a neural network that is part of one or more of those prediction models.

[0116] According to a specific implementation, the convolutional layers in the various 3D neural networks described herein can use edge data to perform mesh convolutions. The use of edge information ensures that the model is insensitive to different input orders of 3D elements. In addition to or separate from using edge data, the convolutional layer can use vertex data to perform mesh convolutions. The advantage of using vertex information is that vertices are generally fewer than edges or faces, so vertex-oriented processing can result in lower processing overhead and lower computational costs. In addition to or separate from using edge data or vertex data, the convolutional layer can use face data to perform mesh convolutions. Furthermore, in addition to or separate from using edge data, vertex data, or face data, the convolutional layer can use voxel data to perform mesh convolutions. The advantage of using voxel information is that depending on the selected granularity, the number of voxels to be processed may be much fewer compared to vertices, edges, or faces in the mesh. Sparse processing (using voxels) may result in lower processing overhead and lower computational costs (especially in terms of computer memory or RAM usage).

[0117] The neural networks of the present disclosure can leverage one or more benefits of parameter tuning operations to optimize the inputs and parameters of the neural network to produce more data-precise results. One parameter that can be tuned is the neural network learning rate (e.g., it can have values such as 0.1, 0.01, 0.001, etc.). The data augmentation scheme can also be tuned or optimized, such as a scheme that adds "shiver" to the dental mesh before input to the neural network (i.e., small random rotations, translations, and / or scalings can be applied to change the dataset and make the neural network robust to data variations).

[0118] The subset of neural network model parameters that can be used for tuning is as follows: ○ Learning rate (LR) decay rate (e.g., how much the LR decays during a training run) ○ Learning rate (LR). A floating-point value used by the optimizer (e.g., 0.001). ○LR scheduling (e.g., cosine annealing, step, exponential) ○Voxel size (for the case of sparse grid processing operations) ○Dropout % (e.g., dropout that can be performed in a linear encoder) ○LR decay step size (e.g., decay every 10 or 20 or 30 epochs) ○Model scaling, which can increase or decrease the layer count and / or the parameter count per layer.

[0119] Parameter tuning can be advantageously applied to the training of neural networks to predict final settings or intermediate gradings, thereby providing technical improvements towards data accuracy. Parameter tuning can also be advantageously applied to the training of neural networks for mesh element labeling or for mesh filling. In some examples, parameter tuning can be advantageously applied to the training of neural networks for tooth reconstruction. In terms of the classifier models of the present disclosure, parameter tuning can be advantageously applied to neural networks for the classification of one or more settings (i.e., the classification of one or more arrangements of teeth). The advantage of parameter tuning is to improve the data accuracy of the output of the prediction model or classification model. In some cases, parameter tuning can provide the advantage of obtaining the last remaining few percentage points of validation accuracy from the prediction or classification model.

[0120] Various neural network models of the present disclosure can benefit from data augmentation. Examples include models trained on 3D meshes, such as GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, FDG settings, setting classification, setting comparison, VAE mesh element labeling, MAE mesh filling, mesh reconstruction VAE, and validation using autoencoders. Such as through Figure 1 Data augmentation by the method shown can increase the size of the training dataset for the dental arch. Data augmentation can provide additional training examples by adding random rotations, translations, and / or rescaling to copies of the existing dental arch. In some specific implementations of the techniques of the present disclosure, data augmentation can be performed by perturbing or jittering the vertices of the mesh in a manner similar to that described in (“Equidistant and Uniform Data Augmentation for 3D Objects”, IEEE Access, Digital Object Identifier 10.1109 / ACCESS.2021.3138162). The position of the vertices can be perturbed by adding Gaussian noise, for example, with a zero mean and a standard deviation of 0.1. According to the techniques of the present disclosure, other mean and standard deviation values are possible.

[0121] Figure 1Shows a data augmentation method to which the system of the present disclosure can be applied for 3D oral care representation. A non-limiting example of a 3D oral care representation is a tooth mesh or a set of tooth meshes. Tooth data 100 (e.g., a 3D mesh) is received at the input. The system of the present disclosure can generate a copy (102) of the tooth data 100. In Figure 1 an example, the system of the present disclosure can apply one or more random rotations to the tooth data 100 (104). In Figure 1 an example, the system of the present disclosure can apply a random translation to the tooth data 100 (106). The system of the present disclosure can apply a random scaling operation to the tooth data 100 (108). The system of the present disclosure can apply a random perturbation to one or more mesh elements of the tooth data 100 (110). The system of the present disclosure can output the augmented tooth data 112 formed by the Figure 1 method.

[0122] Since the generator network of the present disclosure can be implemented as one or more neural networks, the generator may include activation functions. When executed, the activation functions output a determination as to whether a neuron in the neural network will fire (e.g., send an output to the next layer). Some activation functions may include: the binary step function or the linear activation function. Other activation functions impart non-linear behavior to the neural network, including: the sigmoid / logistic activation function, the Tanh (hyperbolic tangent) function, the rectified linear unit (ReLU), the leaky ReLU function, the parametric ReLU function, the exponential linear unit (ELU), the softmax function, the swish function, the Gaussian error linear unit (GELU), or the scaled exponential linear unit (SELU). The linear activation function may be well-suited for some regression applications (and other applications) in the output layer. In the output layer, the sigmoid / logistic activation function may be well-suited for certain binary classification applications (and other applications). The sigmoid activation function may be well-suited for some multi-class classification applications (and other applications) in the output layer. In the output layer, the sigmoid activation function may be well-suited for some multi-label classification applications (and other applications). The ReLU activation function may be well-suited for some convolutional neural network (CNN) applications (and other applications) in the hidden layer. The Tanh and / or sigmoid activation functions may be well-suited for some recurrent neural network (RNN) applications (and other applications) in, for example, the hidden layer. There are a variety of optimization algorithms that can be used to train the neural networks of the present disclosure (such as updating neural network weights), including gradient descent (which uses the first derivative to determine the training gradient and is commonly used for the training of neural networks), the Newton method (which may use the second derivative in the loss calculation to find a better training direction than gradient descent but may require calculations involving the Hessian matrix), and the conjugate gradient method (which may converge faster than gradient descent but does not require the Hessian matrix calculations that the Newton method may require). In some specific implementations, in addition to or instead of the above techniques, additional methods may be employed to update the weights. These additional methods include the Levenberg-Marquardt method and / or simulated annealing. The backpropagation algorithm is used to convey the results of the loss calculation back into the network so that the network weights can be adjusted for learning.

[0123] Neural networks contribute to the functionality of the applications of the present disclosure, including but not limited to: GDL settings, RL settings, VAE settings, capsule settings, MLP settings, diffusion settings, PT settings, similarity settings, tooth classification, setting classification, setting comparison, VAE grid element labeling, MAE grid filling, grid reconstruction autoencoders, verification using autoencoders, estimation of oral care parameters, 3D grid segmentation (3D representation segmentation), coordinate system prediction, grid cleaning, restoration design generation, appliance component generation and placement, or dental arch form prediction. The neural networks of the present disclosure can embody parts or all of various different neural network models. Examples include U-Net architectures, multi-layer perceptrons (MLPs), transformers, pyramid architectures, recurrent neural networks (RNNs), autoencoders, variational autoencoders, regularized autoencoders, conditional autoencoders, capsule networks, capsule autoencoders, stacked capsule autoencoders, denoising autoencoders, sparse autoencoders, conditional autoencoders, long / short-term memory (LSTM), gated recurrent units (GRUs), deep belief networks (DBNs), deep convolutional networks (DCNs), deep convolutional inverse graphics networks (DCIGNs), liquid state machines (LSMs), extreme learning machines (ELMs), echo state networks (ESNs), deep residual networks (DRNs), Kohonen networks (KNs), neural Turing machines (NTMs), or generative adversarial networks (GANs). In some specific implementations, encoder structures or decoder structures can be used. Each of these models offers one or more of its own specific advantages. For example, a particular neural network architecture may be particularly suitable for a specific ML technique. For example, autoencoders are particularly suitable for the classification of 3D oral care representations due to their ability to transform 3D oral care representations into a form that is easier to classify.

[0124] In some specific implementations, the neural networks of the present disclosure may be suitable for operating on 3D point cloud data (alternatively, on 3D meshes or 3D voxelized representations). Many neural network specific implementations can be applied to the processing of 3D representations and can be applied to training predictive and / or generative models for oral care applications, including: PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN, and DSG-Net. Oral care applications include, but are not limited to: setting prediction (e.g., using VAEs, RLs, MLPs, GDLs, capsules, diffusion, etc. trained for setting prediction), 3D representation segmentation, 3D representation coordinate system prediction, element tagging for 3D representation cleaning (VAEs for mesh element tagging), filling of missing elements in 3D representations (MAEs for mesh filling), dental restoration design generation, setting classification, appliance component generation and / or placement, dental arch form prediction, estimation of oral care parameters, setting verification or other verification applications, and 3D representation classification of teeth.

[0125] Some specific implementations of the techniques of the present disclosure incorporate the use of autoencoders. Autoencoders that can be used in accordance with aspects of the present disclosure include, but are not limited to: AtlasNet, FoldingNet, and 3D-PointCapsNet. Some autoencoders can be implemented based on PointNet.

[0126] Representation learning can be applied to the setting prediction techniques of the present disclosure by training a neural network to learn a representation of a tooth and then using another neural network to generate a transformation of the tooth. Some specific implementations can use a VAE or a capsule autoencoder to generate a representation of the reconstructed features of one or more meshes relevant to the oral care field (in some cases, including information about the structure of a tooth mesh). Then, this representation (latent vector or latent capsule) can be used as the input to a module that generates one or more transformations of one or more teeth. In some specific implementations, these transformations can place the tooth into a final setting pose. In some specific implementations, these transformations can place the tooth into an intermediate graded pose. In some specific implementations, the transformation can be described by a 9×1 transformation vector (e.g., specifying a translation vector and a quaternion). In other specific implementations, the transformation can be described by a transformation matrix (e.g., a 4×4 affine transformation matrix).

[0127] In some specific implementations, the systems of the present disclosure may perform principal component analysis (PCA) on an oral care mesh and use the resulting principal components as at least a part of the representation of the oral care mesh in subsequent machine learning and / or other predictive or generative processes.

[0128] In accordance with aspects of the present disclosure, an autoencoder may be trained to generate a latent form of a 3D oral care representation. For example, the autoencoder may include a 3D encoder that encodes the 3D oral care representation into a latent form and a 3D decoder that reconstructs the latent form into a copy of the input 3D oral care representation. Although the present disclosure refers to a 3D encoder and a 3D decoder, the term 3D should be interpreted in a non-limiting manner to cover multi-dimensional operating modes. For example, the systems of the present disclosure may train and deploy multi-dimensional encoders and / or multi-dimensional decoders.

[0129] The systems of the present disclosure may implement end-to-end training. Some end-to-end training-based techniques of the present disclosure may involve two or more neural networks, where the two or more neural networks are trained together (i.e., weights are updated simultaneously during the processing of each batch of input oral care data). In some specific implementations, end-to-end training may be applied to pose prediction by simultaneously training a neural network that learns the representation of teeth and a neural network that can generate tooth transformations.

[0130] In accordance with some transfer learning-based specific implementations of the present disclosure, a neural network (e.g., a U-Net) may be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task may be executed to provide one or more initial neural network weights for training another neural network that is trained to perform a second task (e.g., pose prediction). The first network may learn low-level neural network features of the oral care mesh and is shown to perform well in the first task. By using the first network as a starting point for training, the second network may exhibit faster training and / or improved performance. Certain layers may be trained to encode the neural network features of the oral care mesh in the training dataset. These layers may thereafter be fixed (or undergo minor changes during training) and combined with other neural network components (such as additional layers) that are trained for one or more oral care tasks (such as pose prediction). In this way, a part of the neural network for one or more techniques of the present disclosure (e.g., pose prediction) may receive initial training for another task, which may result in significant learning in the trained network layers. Then, this encoded learning may be built upon by further task-specific training of another network.

[0131] According to the present disclosure, transfer learning can be used for setting prediction and for other oral care applications such as mesh classification (e.g., tooth or setting classification), mesh element labeling, mesh element filling, protocol parameter estimation, mesh segmentation, coordinate system prediction, prosthetic design generation, mesh validation (for any application disclosed herein). In some specific implementations, a neural network trained to output a prediction based on an oral care mesh can first be partially trained on one of the following publicly available datasets before further training on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, Thingi10K dataset (which is particularly relevant for 3D printed component validation), ABC: Large CAD Model Dataset for Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Component Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.

[0132] In some specific implementations, a neural network previously trained on a first dataset (oral care data or other data) can subsequently receive further training on oral care data and be applied to an oral care application such as setting prediction. Transfer learning can be used to further train any one of the following networks: GCN (Graph Convolutional Network), PointNet, ResNet, or any other neural network from the published literature listed above.

[0133] In some specific implementations, a first neural network can be trained to predict the coordinate system of a tooth (such as by using the techniques described in WO2022123402A1 or U.S. Provisional Application No. US63 / 366492). According to any one of the setting prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein), a second neural network can be trained for setting prediction. Transfer learning can transfer at least a portion of the knowledge or capabilities of the first neural network to the second neural network. Thus, transfer learning can provide an accelerated training phase for the second neural network to reach convergence. In some specific implementations, the training of the second network can be completed after being enhanced with transferred learning and then using one or more techniques of the present disclosure.

[0134] The system of the present disclosure can utilize representation learning to train an ML model. Advantages of representation learning include that, as opposed to receiving inputs with variable sizes or structures, the generation network (e.g., a neural network used for predicting transformations in a setup prediction) can be configured to receive inputs with known sizes and / or standard formats. Representation learning can yield performance superior to other techniques because noise in the input data can be reduced (e.g., because the representation generation model extracts hierarchical neural network features and / or reconstruction characteristics of the input representation (e.g., a mesh or point cloud) through loss calculation or a network architecture selected for that purpose).

[0135] The reconstruction characteristics can include values in a latent representation (e.g., a latent vector) that describe aspects of the shape and / or structure of the 3D representation provided to the representation generation module that generated the latent representation. For example, the weights of the encoder module of a reconstruction autoencoder can be trained to encode a 3D representation (e.g., a 3D mesh or others described herein) into a latent vector representation (e.g., a latent vector). In other words, the ability to encode a large set of mesh elements (e.g., hundreds, thousands, or millions) into a latent vector (e.g., hundreds or thousands of real values, e.g., 512, 1024, etc.) can be learned through the weights of the encoder. Each dimension of the latent vector can include a real number that describes some aspect of the shape and / or structure of the initial 3D representation. The weights of the decoder module of the reconstruction autoencoder can be trained to reconstruct the latent vector into a close replica of the initial 3D representation. In other words, the decoder can learn the ability to interpret the dimensions of the latent vector and decode the values within those dimensions. Generally speaking, the encoder and decoder neural network modules are trained to perform a mapping of the 3D representation to a latent vector, and then the latent vector can be mapped back (or otherwise reconstructed) to a 3D representation that is substantially similar to the initial 3D representation for which the latent vector was generated.

[0136] Returning to loss calculation, examples of loss calculation can include KL divergence loss, reconstruction loss, or other losses disclosed herein. Representation learning can reduce the size of the dataset required to train a model because the representation model learns a representation such that the generation network can focus on learning the generation task. Since meaningful neural network features of the input data (e.g., local and / or global features) are available to the generation network, the result can be improved model generalization. In other words, the first network can learn a representation and the second network can make a prediction decision. By training two networks to perform their own separate tasks, each network can generate more accurate results for its corresponding task than a single network trained to both learn a representation and make a decision. In some cases, transfer learning can first train a representation generation model. Then that representation generation model (either wholly or in part) can be used to pre-train subsequent models, such as a generation model (e.g., generation transformation prediction). The representation generation model can benefit from using grid element features as input to improve the ability of the second ML module to encode the structure and / or shape of the input 3D oral care representation in the training dataset.

[0137] One or more neural network models of the present disclosure can have attention gates integrated therein. Attention gate integration provides an enhancement that enables the associated neural network architecture to focus resources on one or more input values. In some embodiments, the attention gate can be integrated with a U-Net architecture, which has the advantage of enabling the U-Net to focus on certain inputs, such as input landmarks corresponding to teeth that are intended to be fixed (e.g., prevented from moving) during an orthodontic procedure (or in cases where other special handling is required). In accordance with aspects of the present disclosure, the attention gate can also be integrated with an encoder or with an autoencoder (such as a VAE or capsule autoencoder) to improve prediction accuracy. For example, the attention gate can be used to configure a machine learning model to give higher weight to aspects of the data that are more likely to be related to the correctly generated output. Thus, and because the machine learning models configured with these attention gates (or mechanisms) utilize aspects of the data that are more likely to be related to the correctly generated output, the final prediction accuracy of those machine learning models is improved.

[0138] The quality and composition of the training dataset for a neural network can affect the performance of the neural network during its execution phase. Dataset filtering and outlier removal can be advantageously applied to the training of neural networks for various techniques of the present disclosure (e.g., for predictions for final settings or intermediate gradings, for neural networks for grid element labeling or for grid filling, for tooth reconstruction, for 3D grid classification, etc.) because dataset filtering and outlier removal can remove noise from the dataset. Although the mechanisms for achieving the improvement are different from using attention gates, the end result is that the method allows the machine learning model to focus on the relevant aspects of the dataset and can lead to an improvement in accuracy similar to the improvement achieved with attention gates.

[0139] In the case of a neural network configured to predict a final setting, a patient case may include at least one of a set of segmented tooth meshes of the patient, an orthotic transformation of each tooth, and / or a ground truth setting transformation of each tooth. In the case of a neural network predicting a set of intermediate stage settings, a patient case may include at least one of a set of segmented tooth meshes of the patient, an orthotic transformation of each tooth, and / or a set of ground truth intermediate stage transformations of each tooth. In some embodiments, the training dataset may exclude patient cases in the contact passive phase (i.e., the phase where the teeth of the dental arch do not move). In some embodiments, the dataset may exclude cases where there is a passive phase at the end of the process. In some embodiments, the dataset may exclude cases where there is overcrowding at the end of the process (i.e., cases where an oral care provider such as an orthodontist or dentist has selected a final setting where the tooth meshes overlap to some extent). In some embodiments, the dataset may exclude cases of a particular difficulty level (or levels) (e.g., easy, medium, and hard).

[0140] In some embodiments, the dataset may include cases with zero pinned teeth (or may include cases with at least one pinned tooth). A technician may specify the pinned teeth when designing the process to prevent various tools from moving that particular tooth. In some embodiments, the dataset may exclude cases with no fixed teeth (conversely, where at least one tooth is fixed). Fixed teeth may be defined as teeth that should not move during the process. In some embodiments, the dataset may exclude cases with no pontic teeth (conversely, cases where at least one tooth is a pontic). Pontic teeth may be described as "ghost" teeth that are represented in the digital model of the dental arch but do not actually exist in the patient's dentition, or where there may be small teeth or partial teeth that may benefit from future work such as adding composite materials via prosthetic appliances. The advantage of including pontic teeth in a patient case is to leave space in the dental arch as part of the plan for the movement of other teeth during orthodontic treatment. In some cases, pontic teeth may save space in the patient's dentition for future dental or orthodontic work such as installing implants or crowns, or applying prosthetic appliances such as adding composite materials to existing teeth that are too small or have an undesirable shape.

[0141] In some specific implementations, the dataset may exclude cases where the patient does not meet the age requirement (e.g., less than 12 years old). In some specific implementations, the dataset may exclude cases where the interproximal reduction (IPR) exceeds a certain threshold amount (e.g., greater than 1.0 mm). The dataset for training a neural network to predict the settings of a clear tray appliance (CTA) may exclude patient cases that are not relevant to CTA processing. The dataset for training a neural network to predict the settings of an indirectly bonded tray product may exclude cases that are not relevant to indirectly bonded tray processing. In some specific implementations, the dataset may exclude cases where only certain teeth are treated. In such specific implementations, the dataset may include only cases where at least one of the following is treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and / or canines.

[0142] Some specific implementations of the present disclosure that are based on autoencoders use capsule autoencoders to automate processing steps in the creation of oral care appliances (e.g., for orthodontic treatment or dental restoration). The advantage of using a capsule autoencoder that has been trained on oral care data is to utilize latent space technology, which reduces the dimensionality of oral care mesh data and thus refines this data, making the signals in the data stronger and easier to use by downstream processing modules, regardless of whether those downstream modules can be other autoencoders, decoders, other neural networks, or other types of ML models (e.g., the supervised and unsupervised models described elsewhere in the present disclosure). Capsule autoencoders were initially applied in the 2D domain to perform object recognition in 2D images, where capsules were trained to create models of the objects to be recognized. This method is capable of recognizing objects in 2D images even if the objects are imaged from new views that do not exist in the training dataset. Later research extended capsule autoencoders to the domain of 3D point clouds, such as in the "3D Point Capsule Network" in the conference of CVPR 2019, which is incorporated herein by reference in its entirety.

[0143] The present disclosure extends the results of this research to apply capsule autoencoders to the digital oral care domain, thereby processing 3D point clouds, 3D meshes, and 3D voxelized representations. In a specific implementation, the term "mesh" hereinafter should be considered interchangeable with 3D point clouds and 3D voxelized representations. A 3D autoencoder can encode one or more 3D geometries (point clouds or meshes) into latent capsules that encode the reconstruction characteristics of the input 3D representation. These latent capsules exist in two or more dimensions and describe the features of the input mesh (or point cloud) and the likelihood of those features. A set of latent capsules is contrasted with latent vectors that can be generated by a variational autoencoder (VaE) and encoded into 1D vectors. One of the contributions of this technology is to advantageously apply capsule autoencoders to the digital oral care space, with the technical advantage of data-oriented precision that improves prediction results.

[0144] Specific examples of applications include segmentation of 3D oral care geometries, setup prediction (both final and intermediate stages), mesh cleaning of 3D oral care geometries (e.g., for both marking of mesh elements and filling of missing mesh elements), tooth classification (e.g., according to a standard dental notation scheme), setup classification (e.g., as malocclusion, grading, and final setup), and automated dental restoration design generation.

[0145] One or more potential capsules representing the input 3D representation (e.g., an oral care geometry such as a point cloud and / or mesh representing an unsegmented dental arch, segmented teeth (such as arranged in a malocclusion setup), teeth attached with hardware, teeth without attached hardware, etc.) can be provided to a capsule decoder to reconstruct a replica of the input 3D representation. The replica can be compared with the input 3D representation by calculating a reconstruction error, thereby demonstrating the information-rich nature of the potential capsules (i.e., the potential capsules describe sufficient reconstruction characteristics of the input mesh such that the mesh can be reconstructed from the potential capsule). A low reconstruction error (e.g., below a predetermined loss threshold) indicates a successful reconstruction. Some of the applications disclosed herein use such information-rich potential capsules for further processing (e.g., such as setup prediction, mesh segmentation, coordinate system prediction, marking of mesh elements for mesh cleaning, filling of missing mesh elements or holes in the mesh, classification of setups, classification of oral care meshes, verification of setups, and other verification tools). Some of the applications disclosed herein make one or more changes to the potential capsules, such as to effect a change in the reconstructed mesh, which can then be output for further use (e.g., to create a dental restoration appliance).

[0146] Figure 2 An example training method for a capsule autoencoder for reconstructing an oral care mesh (or point cloud) of the present disclosure is illustrated. Figure 2 A capsule autoencoder method for mesh reconstruction is shown, which is mainly applied to oral care meshes in the non-limiting examples described herein, but can also be applied to other healthcare meshes or personal safety meshes, such as meshes related to the design, shape, function, and / or use of personal protective equipment such as disposable respirators. The deployment method omits two modules at the bottom. The training method covers the entire illustration. The potential capsule T can be a dimension-reduced form of the input oral care mesh and can be used as an input for other processing.

[0147] Some prior arts rely on inputting 3D point cloud data into a capsule autoencoder. The techniques of the present disclosure expand the input geometry to include 3D mesh data and 3D voxelized representations. In some cases, an input point cloud or mesh (such as containing oral care data) can be rearranged into one or more vectors of mesh elements. Such a vector can be N×3 (in the case of representing the XYZ coordinates of points or vertices). Such a vector can be N×3 (in the case of representing mesh faces, where each mesh face can be defined by 3 indices, each index will index into a list of vertices / points). Such a vector can be N×2 (in the case of representing mesh edges, where each mesh edge can be defined by 2 indices, each index can be indexed into a list of vertices / points). Such a vector can be N×3 (in the case of representing voxels, where each voxel has an XYZ position, such as the centroid, and the length×width×height of each voxel is known).

[0148] In some examples according to aspects of the present disclosure, a neural network such as an MLP can be used to extract features from a list of N×3 mesh element inputs, resulting in a list of N×128 feature vectors, one feature vector per mesh element. In some cases, one or more vectors for calculating mesh element features (as defined elsewhere in the present disclosure) can be calculated for one or more of the N input mesh elements. In some embodiments, these mesh element features can be used in place of the features generated by the MLP. In some embodiments, a feature can be given for each mesh element, which is a mixture of the features generated by the MLP and the calculated mesh element features. In this case, the layer dimension can be enhanced to N×(128 + aug_len), where aug_len is the length of the augmented vector composed of the calculated mesh element features. For the sake of discussion and without loss of generality, this layer will be referred to hereinafter simply as N×128.

[0149] The length "aug_len" can vary according to the embodiment, depending on which mesh elements are analyzed and which mesh element features are selected for use. In some cases, information from more than one type of mesh element can be introduced together with the N×128 vector (for example, point / vertex information can be combined with face information, point / vertex information can be combined with edge information, or point / vertex information can be combined with voxel information). Depending on various applications, the analysis of different types of oral care meshes may require one type of mesh element or another, or a specific set of mesh features.

[0150] An N×128 layer can be passed to a set of subsequent convolutional layers, each of which has been trained to have its own parameter values. The purpose of each of these independent convolutional layers can be to encode individual grid element capsules. The output of each of these convolutional layers can be max-pooled to a size of 1024 elements. The count of these convolutional layers can be a power of two (e.g., 8, 16, 32, 64). In some embodiments, there can be 32 such convolutional layers, each of which outputs a 1024-element vector from the max-pooling operation. These 32 max-pooling output vectors can be concatenated to form a layer that can be 1024×32, called the primary mesh element capsule (PMEC). The dynamic routing module encodes these PMECs into one or more latent capsules, each of which can have a square size (e.g., 16×16, 32×32, 64×64, or 128×128). Non-square sizes are also possible.

[0151] In some embodiments, the dynamic routing module can enable the output of the latent capsules to be routed to the appropriate neural network layer in the subsequent processing module of the capsule autoencoder. The dynamic routing module uses unsupervised techniques (e.g., clustering and / or other unsupervised techniques) to arrange the output of a set of max-pooled feature maps into one or more stacked latent capsules. These latent capsules summarize the feature information from the input 3D representation (e.g., one or more tooth meshes or point clouds) as well as the likelihood information associated with each capsule. These stacked capsules contain sufficient information about the input 3D representation to reconstruct the 3D representation via the capsule decoder module.

[0152] A mesh of grid elements (i.e., such as points / vertices, edges, faces, or voxels) can be generated by the grid patch module. In this example, points will be used for the grid elements. In some embodiments, the mesh can include randomly arranged points. In other embodiments, the mesh can reflect a regular and / or linear arrangement of points. The points in each of these grid patches are the "raw materials" that can form the reconstructed 3D representation.

[0153] Latent capsules (e.g., having dimensions 128×128) can be replicated β times, and prior to being input into one or more MLPs, each of these β latent capsules can sequentially append each of the grid patches of randomly generated grid elements (e.g., points / vertices). In some examples, such an MLP can include a fully connected layer having the following dimensions: {64-64-32-16-3}. The goal of this operation is to customize the grid elements to a specific local region of the 3D representation that may be reconstructed. The decoder performs iterations to generate additional random grid patches and output more random parts of the reconstructed 3D representation (i.e., as point cloud patches). These point cloud patches are accumulated until the reconstruction loss drops below a target threshold. The reconstruction loss can be calculated using one or more of the reconstruction loss (as defined herein) and the KL divergence loss.

[0154] An autoencoder such as a variational autoencoder (VAE) can be trained to encode 3D grid data in a latent space vector A, which can exist in an information-rich low-dimensional latent space. This latent space vector A may be particularly suitable for subsequent processing of digital oral care applications (e.g., such as mesh cleaning, mesh segmentation, mesh verification, mesh classification, setup classification, setup prediction, and repair design generation), because A is capable of effectively manipulating high-dimensional tooth mesh data. Such a VAE can be trained to reconstruct the latent space vector A back into a replica of the input grid (or a transformation or other data structure that describes a 3D oral care representation). In some embodiments, the latent space vector A can be modified strategically to cause a change to the reconstructed grid (or other data structure). In some cases, the reconstructed grid can be a tooth mesh with a changed and / or improved shape, such as would be suitable for use in the design of a dental restoration appliance such as 3M FILTEK Matrix or a veneer. The term "mesh" should be considered in a non-limiting sense to include 3D meshes, 3D point clouds, and 3D voxelized representations.

[0155] Aspects of a tooth mesh reconstruction autoencoder (e.g., a variational autoencoder optionally utilizing normalizing flows) that can be used in accordance with the techniques of the present disclosure are described below. A continuous normalizing flow (CNF) can include a series of invertible mappings that transform probability distributions. The CNF can be implemented by a series of blocks in the decoder of the autoencoder. Such blocks can restrict complex probability distributions, enabling the decoder to learn to map simple distributions to more complex distributions and back, resulting in a technical improvement related to data accuracy such that the tooth shape distribution after reconstruction is more representative of the tooth shape distribution in the training dataset. The invertibility of the CNF provides a technical advantage of resource reduction that improves mathematical efficiency during training, resulting in a technical improvement related to resource usage.

[0156] The dental reconstruction VAE can advantageously utilize loss functions, non - linearities (also known as neural network activation functions), and / or solvers not mentioned in the prior art. Examples of loss functions can include: Mean Absolute Error (MAE), Mean Squared Error (MSE), L1 loss, L2 loss, KL divergence, entropy, and reconstruction loss. Such loss functions enable each generated prediction to be compared with the corresponding ground truth in a quantifiable manner, resulting in one or more loss values that can be used to at least partially train one or more neural networks. Examples of solvers can include: dopri5, bdf, rk4, midpoint, adams, explicit_adams, and fixed_adams. Solvers can enable neural networks to solve systems of equations and corresponding unknown variables. Examples of non - linearities can include: tanh, relu, softplus, elu, swish, square, and identity. Activation functions can be used to introduce non - linear behavior into neural networks in a way that enables the neural networks to better represent the training data. The loss can be calculated through the process of training the neural network via backpropagation. Neural network layers such as ignore, concat, concat_v2, squash, concatsquash, scale, and concatscale can be used.

[0157] In some specific implementations, the dental reconstruction VAE model can be trained on patient cases with malocclusion or alternatively in local coordinates. Figure 3 A method of training such a VAE is shown.

[0158] According to Figure 3 As shown in the grid reconstruction VAE training, a 3D oral care representation F can be provided to an encoder E1 (along with optional tooth type information R) that can generate a latent vector A. The latent vector A can be reconstructed into a reconstructed 3D oral care representation G. A loss between the reconstructed 3D oral care representation G and a ground truth 3D oral care representation GT can be calculated (e.g., using the VAE loss calculation method or other loss calculation methods described herein). Backpropagation can be used to train E1 and D1 with such a loss.

[0159] Figure 4 A trained grid reconstruction VAE in deployment is shown. Figure 4Shown is a grid reconstruction VAE that reconstructs a dental mesh in a deployment. R is an optional input, particularly in the case of dental mesh classification, where such information R is not yet available (according to some specific embodiments, since the dental mesh classification neural network is trained to generate dental type information R as an output). In some specific embodiments, R can be used to improve other techniques, such as mesh element tagging techniques, mesh reconstruction techniques, oral care mesh classification techniques (e.g., such as tooth classification or setting classification), etc.

[0160] Figure 5 and Figure 6 shows a reconstructed dental mesh. Figure 5 Illustrates an example of an input dental mesh (left) and an output reconstructed dental mesh (right). Figure 6 Illustrates another example of an input dental mesh (left) and an output reconstructed dental mesh (right). Figure 6 The use case shown is different from Figure 5 the use case shown.

[0161] Figure 7 shows the reconstruction error from Figure 6 the reconstructed teeth shown, which is called a reconstruction error map. That is, Figure 7 shows the reconstruction error in the above results in a form called a "reconstruction error map". In Figure 7 it, the unit is millimeters (mm). It should be noted that the reconstruction error at the tooth tip is less than 50 microns, and the reconstruction error on most of the tooth surface is much less than 50 microns. Compared with a typical tooth with a size of 1.0 cm, an error rate of 50 microns (or less) means reconstructing the tooth surface with an error rate of less than 0.5%.

[0162] Figure 8 is a bar chart, where each bar or bin represents a single tooth and represents the mean absolute distance of all vertices involved in the reconstruction of that tooth in the data used to evaluate the mesh reconstruction model.

[0163] A dental mesh autoencoder, where a variational autoencoder (VAE) is an example, can be trained to encode a tooth into a reduced-dimension form called a latent space vector. A reconstruction VAE can be trained on example dental meshes. The dental mesh can be received by the VAE, deconstructed into a latent space vector using a 3D encoder, and then reconstructed into a copy of the input mesh using a 3D decoder. The prior art for setting prediction lacks this deconstruction / reconstruction method. One advantage of this method is that the encoder E1 can be trained to encode a dental mesh (or a mesh of a dental appliance, gum, or other body part or anatomical structure) into a reduced-dimension form that can be used to train and deploy any set of powerful setting prediction methods (e.g., GDL settings, RL settings, VAE settings, capsule settings, MLP settings, and diffusion settings, etc.). This reduced-dimension form of the tooth can enable a setting prediction neural network to more effectively encode the reconstruction characteristics of the tooth and better learn to place the tooth into a pose suitable for the final setting or intermediate stage, thereby providing a technical improvement in terms of data accuracy and resource occupancy.

[0164] The reconstructed mesh can be compared to the input mesh, for example, using a reconstruction error that quantifies the difference between the meshes (as described elsewhere in this disclosure). This reconstruction error can be calculated using the Euclidean distance between corresponding mesh elements of the two meshes. There are also other methods of calculating this error that can be derived from materials described elsewhere in this disclosure. Figure 7 and Figure 8 An example reconstruction error according to the techniques described herein is shown.

[0165] In some embodiments, one or more meshes provided to the mesh reconstruction VAE can first be converted to a vertex list (or point cloud) before being provided to the encoder E1. This way of processing the input to E1 can be beneficial for a single mesh input (such as in a dental mesh classification task) or a set of multiple teeth (such as in a setting classification task). The input meshes do not need to be connected.

[0166] The encoder E1 can be trained to encode a dental mesh into a latent space vector A (or "tooth representation vector"). During a prosthetic design task, the encoder E1 can arrange the input dental mesh into a mesh element vector F and encode it into a latent space vector A. This latent space vector A can be a reduced-dimension representation of F that describes the important geometric properties of F. The latent space vector A can be provided to the decoder D1 to be restored to full resolution or near full resolution along with the desired geometric changes. The restored full-resolution mesh or near full-resolution mesh can be described by G, which can subsequently be arranged into the output mesh.

[0167] In some embodiments, such as in prosthetic design generation, tooth names, tooth nomenclature, and / or tooth type R may be concatenated with the latent vector A as a means of conditioning the VAE on such information to improve the VAE's ability to respond to a particular tooth type or nomenclature.

[0168] The performance of the mesh reconstruction VAE can be measured using reconstruction error calculation. In some examples, the reconstruction error can be calculated as the element-to-element distance between two meshes using, for example, the Euclidean distance. According to various embodiments of the techniques of the present disclosure, other distance measurements are possible, such as cosine distance, Manhattan distance, Minkowski distance, Chebyshev distance, Jaccard distance (e.g., intersection over union of the meshes), Hausdorff distance (e.g., distance across the surface), and Sorensen-Dice distance.

[0169] In some embodiments, the performance of the mesh reconstruction VAE can be verified via a reconstruction error map and / or other key performance indicators. The latent space vectors of one or more input tooth meshes can be plotted (e.g., in 2D) using UMAP or t-SNE dimensionality reduction techniques and compared to select the best available separability between tooth classes (molars, premolars, incisors, etc.), indicating that the model is aware of strong geometric differences between different classes and strong similarities within a class. This will be illustrated by distinct non-overlapping clusters in the resulting UMAP / t-SNE map.

[0170] In some cases, the latent vector corresponding to a mesh can be used as part of a classifier to classify the mesh. For example, classification can be performed to identify the tooth type or detect errors in the mesh or mesh arrangement, such as in a validation operation. The latent vector and / or computed mesh element features (such as the spatial and / or structural mesh features described herein) can be provided to a supervised machine learning model to classify the mesh. An incomplete list of possible supervised ML models is found elsewhere in the present disclosure.

[0171] In some embodiments, the reconstruction VAE can be trained to reconstruct any arbitrary tooth type. In other embodiments, the reconstruction VAE can be trained to reconstruct a specific tooth type (e.g., first molar or central incisor).

[0172] Figure 9Describes the training of a mesh reconstruction VAE, which in some specific implementations can be used to encode a dental mesh (or other 3D oral care representation) into a latent representation (e.g., a latent vector) A. The VAE can also be trained to encode other types of 3D representations (e.g., meshes describing gums, retainer model components, oral care hardware such as brackets and / or attachments, dental restoration appliance components, other parts of the anatomical structure, etc.) into the latent vector A. In some specific implementations, a decoder that can generate a reconstructed mesh (e.g., a dental mesh) or other reconstructed representation can be used to reconstruct the latent representation. The reconstructed mesh can include filling materials (e.g., gaps, holes, or incomplete aspects can be filled with mesh elements). In some specific implementations, the latent representation can be reconstructed, and subsequently, a reconstruction error can be calculated. When the reconstruction error of one or more aspects of the reconstructed mesh (e.g., the reconstruction error close to one or more mesh elements) exceeds a threshold error, the technique can generate an indication that an anomaly has been detected. For example, when a dental mesh contains hardware (e.g., a bracket), the mesh elements corresponding to the bracket can be marked with a reconstruction error exceeding the threshold. The technique can generate an indication that hardware (or other abnormal materials) is attached to the tooth. The reconstruction error reflects how well the mesh is reconstructed. When the reconstruction error exceeds the threshold, either or both of the initial mesh and the reconstructed mesh can be considered defective. Figure 9 Provides further details regarding the training of a crown reconstruction VAE.

[0173] Figure 9 Illustrates a method by which the system of the present disclosure can implement training a reconstruction autoencoder for reconstructing a 3D representation of a patient's dentition. Figure 9Specific examples illustrate the training of a variational autoencoder (VAE) for reconstructing a dental mesh 900. For each tooth (908) in the patient case 900, the system of the present disclosure can generate a watertight mesh (902) by merging the crown mesh of the tooth with the corresponding root mesh such that the vertices on the open edges of the crown mesh match the vertices on the open edges of the root mesh. The system of the present disclosure can perform a registration step (904) to align the dental mesh with a template dental mesh (e.g., using the iterative closest point technique or by applying an inverse non-rigid transformation to the tooth), where the technique enhancement is to improve the accuracy of mesh correspondence calculation and data precision at 906. The system of the present disclosure can calculate the correspondence between the dental mesh and the corresponding template dental mesh, where the technique improvement is to adjust the dental mesh to be ready to be provided to the reconstruction autoencoder. The dataset of the prepared dental meshes is divided into a training set, a validation set, and a held-out test set (910), and then used to train the reconstruction autoencoder (912), described herein as a dental VAE, a dental reconstruction VAE, or more generally as a reconstruction autoencoder. The dental VAE can include a 3D encoder that encodes the dental mesh into a latent form (e.g., latent vector A), and a subsequent 3D decoder that reconstructs the tooth as a replica of the input dental mesh. The dental VAE of the present disclosure can be trained using a combination of a reconstruction loss and a KL divergence loss and optionally other loss functions described herein. The output of the method is a trained dental VAE 914.

[0174] Figure 10 Illustrates non-limiting code implementing an example 3D encoder and an example 3D decoder for mesh reconstruction VAE. Figure 10 Illustrates source code (written in Python) corresponding to the encoder and decoder. These specific implementations can include: convolutional operations, batch normalization operations, linear neural network layers, Gaussian operations, and continuous normalizing flows (CNF), etc.

[0175] One of the steps that may occur in VAE training data preprocessing is the calculation of mesh correspondence. The correspondence between the mesh elements of the input mesh and the mesh elements of a reference or template mesh with a known structure can be calculated. The purpose of mesh correspondence calculation may be to find matching points between the surfaces of the input mesh and the template (reference) mesh. Mesh correspondence can generate a point-to-point correspondence between the input mesh and the template mesh by mapping each vertex from the input mesh to at least one vertex in the template mesh. The correspondence between the mesh elements of the input mesh and the mesh elements of a reference or template mesh with a known structure can be calculated. In one example, the range of entries in a vector can correspond to the mesial lingual cusp tip; another range of elements can correspond to the distal lingual cusp tip; another range of elements can correspond to the mesial surface of the tooth; another range of elements can correspond to the lingual surface of the tooth, and so on. In the case of a dental mesh reconstruction autoencoder (such as a VAE), in some embodiments, the autoencoder can be trained only on a subset of teeth (e.g., only molars or only the upper left first molar). In other embodiments, the autoencoder can be trained on a larger subset of teeth in the mouth or all teeth. In some embodiments, an input vector (e.g., a landmark vector) can be provided to the autoencoder, which can define or otherwise influence which type of dental mesh the autoencoder may have received as input. The improvement in data accuracy of this method is to use mesh correspondence in mesh reconstruction to reduce sampling error, improve alignment, and enhance mesh generation quality. Further details regarding the use of mesh correspondence in the autoencoder model of the present disclosure are found elsewhere in the present disclosure.

[0176] In some embodiments, during the calculation of mesh correspondence, an Iterative Closest Point (ICP) algorithm can be run between the input dental mesh and the template dental mesh. Correspondences can be calculated to establish vertex-to-vertex relationships (between the input dental mesh and the reconstructed dental mesh) for calculating the reconstruction error.

[0177] In some embodiments, during the calculation of mesh correspondence, an inverse rigid transformation can be applied to at least approximately align the input dental mesh and the template dental mesh. In some embodiments, both ICP and inverse rigid transformation can be applied.

[0178] According to a particular implementation, training data can be generalized to one or more dental arches (e.g., in other 3D oral care representations), or can be more specific to particular teeth within a dental arch (e.g., in other 3D oral care representations). In cases where more specific training data is utilized, the specific training data can be presented as a tooth template. For example, the tooth template can be specific to one or more tooth types (e.g., the right lower central incisor). In some implementations, a tooth template can be generated that is an average of many examples of a certain type of tooth (such as the average of the first lower molar). In some implementations, a tooth template can be generated that is an average of many examples of more than one tooth type (such as the average of the first and second bicuspids from both the upper and lower dental arches).

[0179] In some implementations, the preprocessing procedure can involve one or more of the following steps: generating an impermeable mesh (e.g., ensuring that the boundaries of the root mesh cleanly seal the boundaries of the crown mesh), registration to align the tooth mesh with the template mesh (e.g., using ICP or an inverse rigid transformation), and calculating mesh correspondence (i.e., to generate mesh element to mesh element correspondence between the input tooth mesh and the template tooth mesh).

[0180] Figure 11 Illustrates a tooth reconstruction generated after the 849th training epoch of a tooth reconstruction autoencoder. In Figure 11 , the left side (labeled "Training Data (ICP)") shows the tooth mesh (in the form of a 3D point cloud) after the preprocessing step is completed, where the preprocessing uses ICP for registration. The right side shows two things: the output of the tooth reconstruction VAE (in the left column) and the corresponding ground truth tooth 3D representation. Also in this case, the 3D representation of each tooth is represented by a point cloud. This output was generated at epoch 849 of the reconstruction VAE training.

[0181] The above description mainly relates to processing meshes, point clouds, and / or voxel data into latent space vectors as a means of reducing the dimensionality of that data and enhancing the signal-to-noise ratio of that data, such that an ML classifier can make decisions based on that data. Applications include but are not limited to VAE settings, MLP settings, MAE mesh filling, VAE mesh element labeling, VAE for tooth mesh classification, and some examples of setting classification. The reconstruction autoencoder trained based on the above materials is also related to validation operations, such as segmentation validation, coordinate system validation, mesh cleaning validation, restoration design validation, fixture model validation, clear tray aligner (CTA) trim line validation, setting validation, oral care appliance component validation (either or both of placement and generation), and hardware (brackets, attachments, etc.) placement validation, to name just a few examples.

[0182] Autoencoders of the present disclosure (such as VAEs or capsule autoencoders) can process other types of oral care data, such as text data, categorical data, spatio-temporal data, real-time data, and / or real number vectors, such as those found in protocol parameters. The data can be qualitative or quantitative. The data can be nominal or ordinal. The data can be discrete or continuous. The data can be structured, unstructured, or semi-structured. The autoencoders of the present disclosure can also encode such data into latent space vectors (or latent capsules) for later reconstruction. Those latent vectors / latent capsules can be used for prediction and / or classification. The reconstruction can be used for model validation and for validation applications, such as by calculation of reconstruction error and / or labeling of data elements.

[0183] A latent vector A (e.g., for a tooth mesh) that can be generated by an encoder E1 in a fully trained mesh reconstruction autoencoder can be a reduced-dimensional representation of an input mesh (e.g., a tooth mesh). In some embodiments, the latent vector A can be a vector of 128 real numbers (or some other size, such as 256 or 512). The decoder D1 of the fully trained mesh reconstruction autoencoder may be able to take the latent vector A as input and reconstruct a close approximation of the input tooth mesh with a low reconstruction error. In some embodiments, the latent vector A can be modified to effect a change in the shape of the reconstructed mesh generated by the decoder D2. Such modification can be done after first mapping out the latent space to gain insight into the effect of making a particular change. There can be various loss functions that can be used in the training of E1 and D1, which can involve terms related to reconstruction loss and / or KL divergence between distributions (e.g., in some cases, to minimize the distance between the latent space distribution and a multi-dimensional Gaussian distribution). One purpose of the reconstruction loss term is to compare the predicted reconstructed 3D representation of a tooth with the corresponding ground truth reconstructed 3D representation of the tooth. One purpose of the KL divergence term is to make the latent space more Gaussian and thus improve the quality of the reconstructed mesh (i.e., especially in cases where the latent space vector can be modified to change the shape of the output mesh, such as splitting a 3D mesh or performing tooth design generation for use in generating dental restoration appliances).

[0184] In some embodiments, the latent vector A can be modified to change the characteristics of the reconstructed mesh (such as with the generation of a dental restoration tooth design mesh). If only the reconstruction loss is used to calculate the loss L and the latent vector A is changed, then in some use case scenarios, the reconstructed mesh can reflect the expected output form (e.g., is a recognizable tooth). However, in other use case scenarios, the output of the reconstructed mesh may not conform to the expected output form (e.g., is not a recognizable tooth).

[0185] Figure 12 Illustrates a latent space where the loss includes reconstruction loss but not KL divergence loss. InFigure 12 In this case, point P1 corresponds to the initial form of the latent space vector A. Point P2 corresponds to a different position in the latent space, which can be sampled as a result of modifying the latent vector A, but the mesh reconstructed from P2 may not produce a good output (for example, it may not look like a recognizable or otherwise suitable tooth). Point P3 corresponds to yet another different position in the latent space, which can be sampled as a result of a different set of modifications to the latent vector A, and the mesh reconstructed from P3 can produce a good output (for example, having an appearance suitable for a tooth design used in generating dental restoration appliances). In the case where the loss only involves the reconstruction loss, the subset of the latent space that can be sampled to obtain the latent space vector P3 that produces a valid reconstructed mesh may be irregular or difficult to predict.

[0186] Figure 13 An example of a latent space where the loss includes both the reconstruction loss and the KL divergence loss is illustrated. In some embodiments, the loss calculation can incorporate a KL divergence term. If the loss is improved by incorporating the KL divergence term, the quality of the latent space can be significantly enhanced. In this new scenario, the latent space may become more Gaussian (as Figure 13 shown), and the latent supervector A corresponds to a point P4 near the center of the multi-dimensional Gaussian curve. Changes can be made to the latent supervector A to obtain a point P5 near P4, where the resulting reconstructed mesh is likely to reflect the desired properties (for example, is likely to be a valid tooth). Introducing the KL divergence term into the loss can make the process of modifying the latent space vector A and obtaining a valid reconstructed mesh more reliable. In some embodiments, similar to the capsule autoencoder, the latent vector can be replaced with a latent capsule, which can be modified and then reconstructed. In some embodiments, the autoencoder framework may be adapted for the segmentation of tooth meshes. Additionally, in some embodiments, the autoencoder framework may be adapted for the task of tooth coordinate system prediction. In some embodiments, a mesh reconstruction autoencoder for coordinate system prediction can compress tooth data into a latent vector form and then provide the latent vector as input to a second ML module (e.g., an MLP) that has been trained for coordinate system prediction (for example, for predicting the coordinate system on a mesh, the goal of which is to define a local coordinate system for the mesh such as a tooth mesh).

[0187] For a given domain (e.g., dental restoration design generation, MAE dental filling or placement design, etc.), a latent space can be mapped such that a change to the latent space vector A results in a reasonably well-reconstructed mesh. The latent space can be systematically mapped by generating latent vectors with carefully chosen value variations (e.g., by experimenting with different combinations of 128 values in an example latent vector). In some cases, a grid search of values can be performed, which has the advantage of effectively exploring the latent space. In the case where the latent space is mapped, the shape of the mesh can be modified by pushing the values in one or more elements of the latent vector value towards the part of the mapped latent space that has been found to correspond to the desired dental characteristics. Using KL divergence augmentation in the loss calculation increases the likelihood of modifying the latent vector into a valid example of reconstructing the input 3D oral care representation (e.g., 3D dental mesh).

[0188] In the case of restoration design generation, the mesh can correspond to at least some parts of the teeth. The latent vector A can be changed such that the resulting reconstructed dental mesh can have characteristics that meet the specifications set by the restoration design parameters. The neural network for dental restoration design generation is described in U.S. Provisional Application No. US63 / 366514, the entire disclosure of which is incorporated herein by reference.

[0189] A dental setup can be designed at least in part by modifying the latent vector corresponding to one or more teeth of one or more dental arches to be placed in the setup configuration (e.g., each tooth is described as a 3D point cloud, voxel, or mesh). The mesh can be encoded into a latent vector A, which then undergoes modification to adjust the pose of the resulting tooth pose. The modified latent vector A' can then be reconstructed into one or more meshes describing the setup. Such techniques can be used to design the final setup configuration or intermediate stage configurations, etc.

[0190] In some embodiments, the modification of the latent vector can be performed via an ML model (such as one of the neural network models or other ML models disclosed elsewhere in the present disclosure). In some embodiments, the neural network can be trained to operate within the latent space representation of such a vector A of the setup mesh. The mapping of the latent space of A may have been previously generated by making controlled adjustments to the trial latent vectors and observing the resulting changes to the setup configuration (e.g., after the modified A has been reconstructed back into one or more complete meshes of the dental arch). In some cases, the mapping of the latent space can follow an organized search pattern, such as in a grid search.

[0191] In some specific implementations, the dental reconstruction VAE can take a single input of tooth name / type / designation R, which can command the VAE to output a dental mesh of a specified type. This can be achieved by generating a latent vector A' used in reconstructing the appropriate dental mesh. In some specific implementations, this latent vector A' can be "instantly" sampled or generated from a previous mapping of the latent vector space. This mapping can be performed to understand which parts of the latent vector space correspond to different shapes, structures, and / or geometries of teeth. For example, among the 128 real values in the example latent vector A' (other sizes are possible), it may have been determined that certain elements of those vector elements and certain value ranges may correspond to specific types / names / designations of teeth and / or teeth having certain shapes or other desired characteristics. The model for dental mesh generation can also be applied to the generation of oral care hardware, appliances, and appliance components (such as for orthodontic treatment). The model can also be trained to generate other types of anatomical structures. The model can also be trained to generate other types on non-oral care meshes.

[0192] The mesh comparison module can compare two or more meshes, for example, for the calculation of a loss function or for the calculation of reconstruction error. Some specific implementations can involve the comparison of the volumes and / or areas of two meshes. Some specific implementations can involve calculating the minimum distance between corresponding vertices / faces / edges / voxels of two meshes. For a point in one mesh (e.g., a vertex, the midpoint on an edge, or the center of a triangle), the minimum distance between that point and the corresponding point in the other mesh is calculated. In cases where the other mesh has a different number of elements or there is no clear mapping between corresponding points of the two meshes, different methods can be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools that can play a role in the mesh comparison module of the present disclosure. In some specific implementations, the Hausdorff distance can be calculated to quantify the shape difference between two meshes. The open-source software tool Metro developed by the Visual Computing Lab can also play a role in quantifying the difference between two meshes. The following paper describes the method adopted by Metro, which can be modified by the neural network applications of the present disclosure for mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces", P. Cignoni, C. Rocchini, and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, Vol. 17(2), June 1998, pp. 167-174.

[0193] Some techniques of the present disclosure may involve the following operations: for one or more points on a first grid, shooting rays perpendicular to the grid surface and calculating the distance before the rays are incident on a second grid. The length of the resulting line segment can be used to quantify the distance between the grids. According to some techniques of the present disclosure, a color can be assigned to the distance based on the magnitude of the distance, and the color can be applied to the first grid by means of visualization.

[0194] Mesh element marking techniques have several advantageous applications for the processing of 3D oral care representations (such as 3D oral care meshes), with the ultimate goal of creating oral care appliances. Specifically, mesh element marking is advantageous for the segmentation of 3D oral care representations. Mesh element marking is also advantageous for the task of 3D oral care mesh cleaning (3D representation cleaning), which in some specific implementations may involve the marking of mesh elements, the removal of the marked mesh elements, and the repair or filling of any resulting holes or rough boundaries. The following sections discuss mesh segmentation and mesh cleaning.

[0195] Figure 14 An example of a mesh element marking model is illustrated, which is based on a denoising diffusion probability model (DDPM) and can be used to implement either 3D mesh segmentation (e.g., semantic segmentation) or 3D mesh cleaning methods. The DDPM for 3D representation segmentation may include at least one of a forward pass (e.g., which may involve adding more noise iteratively to the input 3D representation in many steps, which can be used in training an ML model to perform the reverse process) and a reverse path (e.g., which may use an ML model such as a neural network to iteratively denoise the noisy representation, resulting in the segmentation of the input 3D representation). In some specific implementations, the DDPM for 3D representation segmentation may approximate aspects of a Markov process. A U-Net (e.g., the U-Net shown as part of the method in Figure 22 can be trained to perform the reverse process, which may start operating on the noisy representation and iteratively denoise the representation to produce a denoised representation that may include one or more mesh element labels (alternatively, the noisy representation can be directly denoised into one or more segmented 3D representations, such as 3D point clouds or 3D meshes).

[0196] In some specific implementations, the DDPM for 3D representation segmentation may include the following steps: 1) Iteratively add noise (e.g., Gaussian noise) to the input 3D oral care representation to generate a series of increasingly noisy versions of the input 3D oral care representation that can subsequently be used in the denoising neural network during the training of the reverse process. 2) Use the U-Net in Figure 37 (a U-shaped method of convolutional / transposed convolutional and / or pooling / unpooling operations) to extract feature maps from the series of increasingly noisy 3D oral care representations in Setup #1. 3) Collect grid element-level representations by upsampling the feature maps from the U-Net in step #2 and concatenating those upsampled feature maps. 4) The concatenated grid element-level feature vectors from step #3 can be used to train one or more ML models for grid element labeling (e.g., to train an ensemble of neural networks for grid element labeling). Figure 14 The hierarchical neural network feature module (HNNFM) in may include one or more of: a U-Net (see Figure 22 ), a pyramid encoder-decoder structure (see Figure 23 ), or a 3D SWIN Transformer, as well as other architectures.

[0197] This section describes methods for segmenting 3D oral care representations such as 3D oral care meshes. Figure 14 It involves 3D mesh segmentation and 3D mesh cleaning. These two techniques share an important property, namely the labeling of 3D mesh elements. In the case of 3D mesh segmentation (e.g., related to the segmentation of an oral care mesh such as teeth), the various mesh elements of the mesh can be labeled according to the part of the dental anatomy to which the mesh element belongs (e.g., gum, upper right central incisor, lower left second bicuspid). Tooth mesh segmentation can label elements according to membership in various teeth, as specified by one or more of the dental notation systems mentioned herein (e.g., Palmer notation). In some specific implementations, for facio-lingual segmentation, each mesh element can be labeled according to membership in the facial side of the dental arch or the lingual side of the dental arch. Other specific implementations of dental anatomy segmentation are also possible.

[0198] A pre-segmented mesh (e.g., an arch generated by an intraoral scanner or a CT scanner) can be received by the segmentation system. Grid element feature vectors can be computed, one vector for each grid element. One or more grid element features from elsewhere in this disclosure can be used to form the feature vectors of the grid elements. Grid elements can include edges, faces, vertices, and voxels (or any combination thereof).

[0199] In some specific implementations, a list of grid element feature vectors can be provided to a neural network that refines the grid element feature vectors to extract local and global features from those feature vectors, such as Figure 22The U-Net structure shown in []. The U-Net may include 3D grid convolution and / or subsequent 3D grid pooling operations, which can be used to reduce the resolution of the grid and extract neural network features at increasingly global scales. After a series of such operations, after the most global neural network features have been extracted, there may be a series of 3D grid unpooling and 3D grid deconvolution operations that return the grid to the initial scale. After outputs have been generated at various levels of the U-Net structure, there may be upsampling and concatenation operations. F1, F2, F3, and F4 can be 3D grid element-level feature vectors that may have been upsampled to the initial grid resolution. These feature vectors are concatenated and provided as input to one or more ML models for grid element classification. Any supervised ML model disclosed elsewhere in this disclosure can be used for such classification, such as SVM and logistic regression. In some examples, an ensemble of ML classifiers can take the grid element representation vector as input and generate grid element classification labels. In some embodiments, an ensemble of fully connected neural networks can be used for such classification. In some cases, a multi-layer perceptron (MLP) can be used for such classification, including linear layers and associated ReLU activation functions (e.g., which has the advantage that not all neurons fire simultaneously, enabling a more customized response to the input than activation functions that fire for every evaluation) and batch normalization operations. Other activation functions are also possible, such as those mentioned elsewhere in this disclosure. Each of the ML models in the ensemble can output a predicted class label for each grid element. A voting mechanism can then be employed to combine these results and output the final class label prediction for each 3D grid element. Some embodiments can use transformers to assist in applying labels to grid elements.

[0200] The mesh cleaning and mesh segmentation implementations differ in terms of the arrangement of the ground truth mesh element labels. In the case of tooth segmentation, ground truth data can be given for each dental arch mesh to be segmented (i.e., each tooth in the dental arch has mesh elements labeled according to the tooth to which the mesh element belongs). Further, such a loss can be calculated for the facial-lingual segmentation and / or for the tooth-gum segmentation (as defined in U.S. Provisional Application No. US63 / 366490). Various loss functions of the present disclosure can be used to compare the predicted segmentation with the mesh element labels of the ground truth segmentation. Cross-entropy loss is an exemplary choice among several candidate loss functions for this comparison. Other possible losses are disclosed elsewhere in the present disclosure. In the case of mesh cleaning, ground truth data can be given for each dental arch mesh to be cleaned. For example, in the case of removing foreign materials, each mesh element corresponding to the foreign materials in the dental arch can be so labeled, and each mesh element that does not include foreign materials (i.e., that will be retained after mesh cleaning) can be so labeled. In the case of removing depressions, each mesh element corresponding to the depressions in the dental arch can be so labeled, and each mesh element that does not include depressions (i.e., that will be retained after mesh cleaning) can be so labeled.

[0201] After applying the segmentation labels, a process is performed to copy the mesh elements of a particular label into a new mesh (also referred to as tooth cutting), and the mesh is saved (e.g., saved to an electronic storage medium) for further processing. After completing the mesh cleaning labeling, a process is performed to remove each designed mesh element from the mesh (e.g., using classical mesh processing techniques), and then optionally a process can be performed to fill any holes that may have been created by this process (e.g., using techniques trained for mesh filling as described elsewhere in the present disclosure).

[0202] The mesh element labeling for mesh segmentation can also be done using an autoencoder trained for that purpose, such as a variational autoencoder, as described elsewhere in the present disclosure.

[0203] This portion of the present disclosure describes methods for segmenting 3D oral care representations such as 3D oral care meshes. Figure 15An example training method of a capsule autoencoder for 3D representation segmentation (e.g., segmenting a 3D mesh) is illustrated. Oral care arguments 1504 can be provided to affect the functionality of the segmentation method, including: 1) a minimum triangle area threshold, 2) a maximum / minimum angle between adjacent faces, 3) a smoothness argument affecting the size of mesh elements, etc. A pre-segmented 3D representation 1500 (e.g., a 3D mesh) can be arranged as a vector of mesh elements that can be provided to a mesh element feature module 1502. The mesh elements (along with optional mesh element feature vectors) can be provided to a representation generation module 1506, which in some embodiments can compute hierarchical neural network features of the 3D representation. The generated latent representation can be provided to a latent capsule generation module 1508, which can assemble latent capsules 1512 from one or more latent vectors. The latent capsules can be provided to one or more 3D representation sub-segment generation modules 1514, which can be trained to generate mesh element labels for corresponding portions of the 3D representation by means of a mesh patch module 1510. The sub-segments can be provided to one or more sub-segment recombination layers 1516, which can generate the final mesh element labels of the 3D representation (or the final segmented 3D representation of the teeth) 1518.

[0204] 3D representation segmentation can be implemented using a capsule autoencoder that has been trained for segmentation, such as Figure 15The method shown. A loss can be calculated (1520) to compare a predicted segmentation (e.g., a grid element label or a separately segmented tooth mesh) with a corresponding ground truth segmentation. In some embodiments, a cross-entropy loss can be calculated and the network (1522) can be updated via backpropagation. Other losses described herein can be used for training. The techniques of the present disclosure can train and use a capsule autoencoder to segment an oral care mesh (such as the raw unsegmented 3D mesh data of an arch output from an intraoral scanner). This segmentation can be achieved via capsule associations that can leverage dynamic routing, where each potential capsule can be trained to identify sub-parts of the 3D representation, and a final sub-segment reassembly layer 1516 (e.g., one or more fully connected layers with optional skip connections, one or more transformer encoders, or one or more transformer decoders, etc.) can organize these sub-parts into an entire set of segmentations of grid elements. Each capsule can correspond to a portion of the arch mesh (alternatively, a portion of the arch point cloud). Each capsule can have an associated label and can apply that label to certain sets of grid elements. Such a capsule autoencoder (e.g., a 3D-PointCapsNet network) can be trained on examples of patient case data where the raw unsegmented 3D mesh data of the arch may be available and the ground truth segmentations of those same datasets may also be available. Other types of oral care meshes can also be segmented. In some cases, transfer learning techniques can be applied to assist the training process and transfer encodings from other previous neural networks to a new neural network that may be the subject of training. Using the techniques of the present disclosure, an oral care mesh can be segmented at each capsule segmentation. Possible segmentation methods include, but are not limited to: tooth segmentation, facial-lingual segmentation, tooth-gum segmentation. In some embodiments, the SegCaps network can assist in such 3D representation segmentation methods.

[0205] The capsule encoder (of the first part of the method) can encode a pre-segmented mesh into one or more potential capsules. In some embodiments, the potential capsules can be data structures of at least two dimensions, while in some embodiments, the latent vector A can be 1D. The capsule decoder can reconstruct these potential capsules into a facsimile of the input 3D mesh (i.e., or other 3D representation), where sub-parts of the 3D mesh may be segmented (or labeled). For example, different sets of grid elements can be labeled based on their membership in different anatomical parts of the teeth and / or gums. The sub-segment reassembly layer 1516 (which can be an MLP) can combine these labeled groupings of grid elements into larger cohesive groupings of grid elements that can correspond to the teeth and / or gums. The sub-segment reassembly layer 1516 can bring these smaller parts of the 3D mesh together into the labeled groupings. Once the grid elements of each tooth may have been labeled, tooth cutting (in accordance with U.S. Provisional Application No. US63 / 366490) can be performed, and a segmented tooth mesh can be output for subsequent processing and use.

[0206] Capsules can be randomly initialized by randomly sampling from the grid patch module. As the training of the segmentation implementation proceeds, these randomly initialized capsules may become dedicated to certain types of structures and / or shapes in the oral care grid of the training dataset. In the case of tooth segmentation, the capsules may become dedicated to 3D oral care representation segments such as cusp tips, occlusal surfaces, lingual surfaces, incisal edges, gingival margins, gingival surfaces, mesial (or distal) tooth surfaces, low-curvature tooth surfaces, high-curvature tooth surfaces, convex tooth surfaces, concave tooth surfaces, etc.

[0207] At Figure 15 At the end of the method in, there may be one or more layers that may have been trained to combine capsules into larger 3D oral care representation parts such as segments of a tooth crown or an entire tooth crown. These end layers are referred to as the sub-segment recombination layer 1516. In some implementations, these sub-segment recombination layer 1516 may use only a portion of the training data that is used to train the rest of the capsule autoencoder. The sub-segment recombination layer 1516 may provide grid element results (i.e., point / vertex cloud, edges, faces, or voxels) of one or more capsules and combine these grid elements into meaningful and / or recognizable parts such as a tooth crown or a gingiva. In this way, the trained capsule autoencoder can be dedicated to performing segmentation of 3D representations (e.g., point cloud, mesh, voxelized representation, etc.). The training of the sub-segment recombination layer 1516 may use cross-entropy loss and backpropagation. Other losses described herein are possible. The intersection over union (IOU) accuracy is a metric that can be used to measure the accuracy of the resulting segmentation. The IOU can be defined as the overlapping grid area of the predicted grid and the ground truth grid divided by the total combined area of the two grids.

[0208] This disclosure describes advantageous new techniques for mesh processing in digital oral care using geometric deep learning (GDL) models that have been trained on oral care meshes. 3D meshes of teeth (e.g., those generated by an intraoral scanner or by scanning 3D representations of teeth, such as a jig model) can be processed during the creation of dental and orthodontic appliances. An important step in this processing involves removing abnormal materials from the mesh, which is an activity called mesh cleaning, where errors or non-standard aspects of one or more tooth meshes can be corrected to prepare for subsequent processing and appliance creation.

[0209] The 3D representation may include a 3D mesh, a 3D point cloud, or a 3D voxelized representation. Additionally, the term "mesh" should be considered to include 3D meshes, 3D point clouds, and 3D voxelized representations in a non-limiting sense.

[0210] The present disclosure features two related and advantageous techniques in which an autoencoder can be used to perform steps in a mesh cleaning method. The first technique can use an autoencoder to label mesh elements (e.g., mesh elements to be removed). This first technique can be performed for anomaly detection, such as detecting oral care hardware on teeth. The second technique can use an autoencoder to reconstruct missing parts of the mesh (e.g., fill holes left after removing mesh elements). Some specific implementations of the first technique can use a variational autoencoder (VAE). Some specific implementations of the second technique can use a masked autoencoder (MAE). Figure 16 Illustrates how these two techniques can be used consistently. The first technique and the second technique can also be combined with other types of unsupervised machine learning models, such as the examples described elsewhere in the present disclosure.

[0211] Figure 16 An example of mesh cleaning using an autoencoder is illustrated. One or more meshes 1600 that may need to be modified (or "cleaned") are received as input into a neural network (1602) that has been trained for mesh element labeling. According to the techniques of the present disclosure, a mesh reconstruction autoencoder can be trained to label anomalous and mesh elements that need to be modified or removed. Such a reconstruction autoencoder (e.g., including a 3D encoder followed by a 3D decoder) can be trained to encode the received mesh 1600 into a latent form and then reconstruct it as a reconstructed mesh, where a reconstruction error is computed on various aspects of the reconstructed mesh (e.g., to compare those aspects such as mesh elements to the input mesh), thereby labeling aspects of the reconstructed mesh that deviate from the input mesh by more than a threshold. Then, any additional aspects (e.g., mesh elements) of the input mesh that need to be modified or removed can be identified using optional mesh processing techniques (1604) known to those skilled in the art (KTOSITA). A mask 1608 can be generated that labels aspects of the mesh 1600 that need further processing by the mesh element labeling method 1602. Mesh processing techniques (1606) can be applied to modify the mesh 1600, in some cases utilizing the mask 1608. In some specific implementations, a masked autoencoder (1610) (which has been trained according to the description contained herein) can be used to fill holes, smooth boundaries, and / or generally repair damage to the mesh caused by the operation (1606). The resulting cleaned mesh 1612 is output for further processing (e.g., for oral care appliance generation).

[0212] The correspondence between features is typically useful in data analysis applications where it is helpful to identify similar features. The calculation of correspondences is frequently done in computer vision, especially for object tracking, but also for other applications. The use of correspondences enables data such as 3D data to be sorted in a more structured format, such that similar parts of one 3D representation mesh 1 can be more easily compared to corresponding parts of another 3D representation mesh 2, for example, by a neural network. When a neural network is consuming and understanding input data, depending on the network architecture used, it may be beneficial to place all mesh data in order, as the network can then focus more on the data structure without having to compare each part of the data to every other possible part of the data. Instead, the neural network (or other ML model) can observe how similar parts of the data change relative to each other and learn what the typical variations in the data are in the process. In this case, the ordering of similar features helps to advance the learning process.

[0213] The first part of the process is to identify corresponding similar features in each data sample (e.g., each tooth within an arch, or each example of the same tooth type within a larger dataset). These corresponding features can be mesh elements representing all distal margins or, for example, canine cusps. Typically, a neural network takes an ordered feature vector as input. Since 3D structures are objects composed of unordered elements or features, it can be difficult to structure the data in a way that the neural network can easily understand the data. By performing a correspondence calculation on these features as a preprocessing step, the techniques of the present disclosure can then arrange the input vectors of our network in such a way that all similar features (e.g., margins or cusps) appear in the same position in the input vector, regardless of what data sample (e.g., tooth) they come from.

[0214] This preprocessing can also have the advantage of reducing the complexity of the neural network, as the preprocessing only requires less work to gain some initial understanding of the data.

[0215] One of the various advantages of the VAE architecture is the fact that the VAE can be stably trained. The VAE does not suffer from negative effects such as mode collapse. The VAE can use multiple data samples to construct a latent representation of the data. This can force the model to consider various modes of the data, thus avoiding mode collapse.

[0216] The KL divergence loss can cause the data to exhibit at least some of the characteristics of Gaussian behavior when rendered in latent form. The KL divergence term of the model can encourage the distribution of one or more elements of the latent vector to be Gaussian and can penalize the model (via this KL divergence loss) for any deviation from the Gaussian distribution.

[0217] In some specific implementations, the autoencoder can be trained on an unlabeled dataset in an unsupervised manner. This feature enables a large amount of unlabeled data to be used for training without the need for expensive and sometimes inaccurate labeling, thereby achieving a more robust resulting trained neural network model due to the larger dataset available for model training.

[0218] Figure 16 The grid element labeling method 1602 shows a variational autoencoder (VAE) that can be trained at least in part using an example of a 3D grid for grid element labeling, where the example of the 3D grid already has ground truth grid element labels associated with the grid elements of the 3D grid. Such an autoencoder can be used for, for example, anomaly detection. The VAE can generate grid element labels for one or more grid elements of a 3D representation so that the 3D representation can undergo, for example, mesh segmentation or mesh cleaning.

[0219] In some specific implementations, these grid element labels can be used for 3D mesh segmentation (such as tooth segmentation, face-tongue segmentation, and tooth-gum segmentation). In some specific implementations, the autoencoder can be used for 3D mesh segmentation. In some specific implementations, the variational autoencoder can be used to encode 3D mesh data into a latent space vector that can contain grid element-level latent features. In some specific implementations, a GAN inversion operation (especially an operation with semantic perception attributes) can be performed on the embedding vector generated by a 3D segmentation U-Net structure to obtain grid element-level latent features (or attributes). Such grid element-level latent features (e.g., neural network features) can be provided to an ML model for labeling or classifying grid elements, such as in Figure 14 In some specific implementations, a 3D denoising diffusion probability model (DDPM), a 3D adaptation of swapping assignments between multiple views (SwAV), or a 3D masked autoencoder (MAE) can be used for 3D mesh segmentation. In some specific implementations, 3D mesh segmentation can be accomplished by the manipulation of a latent vector A. The complete pre-segmented mesh can be provided to an encoder E1, thereby generating the latent vector A. The latent vector A can be edited to isolate one or more teeth (in the case of tooth segmentation). The part of the dental arch that does not correspond to the one or more teeth can be removed from the mesh via these latent vector edits. Then, the modified latent vector can be reconstructed into one or more meshes, thereby producing one or more segmented teeth as output.

[0220] In some specific implementations, these mesh element labels can be used for anomaly detection as part of a mesh cleaning operation. A variational autoencoder can be trained to be used as an anomaly detection autoencoder for detecting anomalies by labeling mesh elements. Anomalous mesh elements can be identified by identifying mesh elements with a reconstruction error higher than a threshold. The reconstruction error can be calculated as a result of performing a variational autoencoder (VAE) on the input mesh. Although this specific implementation describes the use of a VAE, it should be understood that, without loss of generality, other types of autoencoders can replace the VAE (e.g., capsule autoencoders and / or stacked autoencoders). A mesh can be provided to the VAE, which can decompose the mesh into a latent space vector by performing an encoder structure. The latent space vector can at least partially correspond to a lower-dimensional representation of the input mesh. Then, the VAE can restore the lower-dimensional representation of the mesh to a higher-dimensional (or original-dimensional) representation of the mesh by performing a decoder structure. The encoder and decoder structures of the VAE can be trained on a dataset of accurately segmented tooth meshes such that the encoder and decoder can become proficient at decomposing and subsequently reconstructing the original tooth mesh. If, after deploying the VAE, a tooth mesh that is not original (i.e., contains anomalies such as brackets, attachments, or other pieces of oral care hardware) is input to the VAE, the VAE may have difficulty reconstructing the tooth mesh (e.g., the VAE may not be able to accurately reconstruct the mesh elements near the anomalous object). Because the encoder and / or decoder have not been trained to deconstruct and / or reconstruct the hardware (or other anomalies such as foreign materials, depressions, undercuts, internal fractures, or lingual bars), the reconstructed tooth mesh may not be very similar to the input tooth mesh (at least in the local area of the anomalous object). The similarity between the input tooth mesh and the reconstructed tooth mesh can be measured using the reconstruction error. The reconstruction error can be calculated in a variety of ways, including but not limited to the Euclidean distance between vertex A in the input mesh and the corresponding vertex in the reconstructed mesh, or the closest point from the input mesh to the surface of the reconstructed mesh. Each mesh element in the input mesh can be assigned a reconstruction error. Mesh elements with a reconstruction error higher than a specified threshold (e.g., 50 µm, 100 µm, or some similar measurement depending on the needs of various applications) can be considered to be poorly reconstructed and thus labeled as anomalous. Otherwise, the mesh element is labeled as non-anomalous. In some specific implementations, the labels of one or more mesh elements can be provided to a mesh classifier ML model, which can classify the mesh at least partially based on the labels of one or more mesh elements. In some specific implementations, the mesh classifier ML model can identify the type or classification of the detected anomaly. Such a mesh classifier can consist, for example, of an encoder structure, a U-Net implemented by the MinkowskiEngine toolkit, a U-Net using MeshCNN features, or the original dual mesh convolutional network developed by MIT / UTH Zurich. A mask or vector of mesh element labels can be generated by the "AI Mesh Element Labeling" module.This mask can be used by subsequent modules to specify mesh elements that can benefit from further processing. Some specific implementations may employ a decision tree model (or other ML models disclosed herein) to select which further processing steps to perform on one or more dental meshes. Other suitable classifiers are disclosed elsewhere in this disclosure.

[0221] Figure 17 An example implementation of a VAE for dental mesh reconstruction is shown, where the original dental mesh has been reconstructed by the VAE. Figure 17 An example implementation of a VAE for dental mesh reconstruction is shown, where the VAE has attempted to reconstruct an abnormal dental mesh (i.e., a dental mesh with attached orthodontic brackets). Figure 17 The autoencoder is trained for mesh element anomaly detection. In this example, the autoencoder reconstructs the mesh elements of the dental mesh with a relatively low reconstruction error, but has difficulty reconstructing the mesh elements of the attached orthodontic brackets (i.e., reconstructs the mesh elements corresponding to the brackets with a high reconstruction error). The dental mesh elements corresponding to the brackets are marked as abnormal. The VAE can generate a mask that labels these abnormal dental mesh elements.

[0222] In some specific implementations, the nature of the abnormal material can be determined. For example, some parts of the input dental mesh can be isolated (e.g., by eliminating the well-reconstructed parts). The remaining parts of the input dental mesh can represent the parts that did not become well-reconstructed in this example. Such parts can be provided to a mesh classifier to identify the nature of the anomaly (e.g., to determine whether the anomaly is a bracket or something else).

[0223] In some specific implementations, the autoencoder for the mesh element labeling technique (e.g., the VAE for mesh element labeling) can be conditioned on optional inputs elsewhere in this disclosure, such as a vector P containing dental size information, one or more latent vectors B, and / or information R related to tooth names, nomenclature, tooth types, and / or tooth classifications. The model can be conditioned on such optional inputs by concatenating such inputs with one or more input vectors (i.e., the inputs to the encoder) and / or by concatenating such inputs with one or more latent vectors generated by the encoder in the autoencoder (the VAE for mesh element labeling) for mesh element labeling.

[0224] The VAE model can be trained using examples of pristine tooth meshes (i.e., tooth meshes without foreign objects, hardware, foreign materials, pits, undercuts, internal fractures, lingual bars, or other non - pristine geometric features). In some specific implementations, the VAE can be trained using examples named by a single tooth (e.g., the right lower central incisor). The resulting VAE can then be deployed for use in labeling mesh elements in the right lower central incisor. In some specific implementations, the VAE can be trained on pristine examples of a subgroup of teeth (e.g., incisors and canines). Then, the resulting VAE can be specifically deployed for use in labeling mesh elements in incisors and canines. In some specific implementations, the VAE can be trained only on molars or on the entire set of teeth. The key is that the training dataset has no anomalies such that when such anomalies are encountered in deployment, the resulting deployed VAE may not accurately reconstruct these anomalies. Anomalous structures that are inaccurately reconstructed (i.e., the mesh elements of those structures) can then be identified by computing the reconstruction error and / or marked for further processing. There are other techniques, such as MLP setups and VAE setup techniques, that can benefit from a tooth reconstruction auto - encoder trained on a single tooth type (or a subgroup of tooth types). Such techniques can benefit from training a tooth reconstruction auto - encoder on one or a small number of tooth types because the resulting latent vectors (or latent capsules) can have an improved ability to represent the 3D representation of the input. Due to the increased accuracy of the tooth representation, the resulting latent vectors or latent capsules may be more favorable for use in predicting ML models in the training setup.

[0225] Similar to the VAE for mesh element tagging, the capsule autoencoder can be trained to encode a mesh (or other 3D representation) into one or more latent capsules using a capsule encoder structure and then reconstruct one or more latent capsules into a replica of the input mesh (or other input 3D representation) using a capsule decoder structure. Such an autoencoder can be trained to encode and reconstruct a typical (or nominal) oral care mesh (e.g., a typical or healthy crown mesh). When in deployment, an atypical mesh (or other 3D representation) is provided to the autoencoder, and the autoencoder can reconstruct a mesh that is different from the input mesh in certain ways, thereby flagging the atypical nature of the input mesh. In the case where the input mesh is an oral care mesh with attached hardware (such as a crown with orthodontic brackets attached), the capsule autoencoder can reconstruct the mesh where the mesh elements corresponding to the brackets are not well-formed and / or do not resemble the brackets. In such a case, a reconstruction error can be computed between the input mesh and the reconstructed mesh, and it can be revealed which mesh elements are not well-reconstructed. These mesh elements can be flagged or annotated as abnormal and may require further processing, such as having the abnormal mesh elements removed and the resulting holes filled, similar to the VAE application for hardware removal. In some embodiments, the abnormal portion of the mesh can be provided to a mesh classification neural network, which can classify the nature of the abnormality based at least in part on the shape and / or structure of the abnormal portion of the mesh. In some embodiments, one or more mesh element labels can be provided to the mesh classification neural network to assist the neural network in classifying the abnormality.

[0226] Figure 16 The surface reconstruction method 1610 describes an autoencoder that can perform mesh surface reconstruction. In some examples, mesh surface reconstruction may be beneficial after marked mesh elements (e.g., mesh elements corresponding to hardware, or mesh elements corresponding to other mesh abnormalities such as foreign material, depressions, undercuts, internal fractures, or lingual bars) have been removed from the mesh. In some examples, the input mesh data for an object such as a crown may be incomplete because adjacent teeth, hardware, or other objects obstruct the intraoral scanner during data collection. In some embodiments, surface reconstruction can involve filling or estimating missing mesh elements. The techniques of the present disclosure can train and deploy masked autoencoders to perform such filling or estimating tasks.

[0227] Regions for masking mesh elements can be defined for two or more consecutive mesh elements. Depending on various specific implementations, the mask can include one or more regions of masking mesh elements. In some examples, mesh elements such as vertices in a mesh or points in a point cloud can be masked. In some specific implementations, masking can involve overwriting one or more of the mesh element feature vector values associated with the mesh element with masking tokens (e.g., overwriting the XYZ coordinates of the mesh element). A masked autoencoder can be trained to consume these masked mesh elements (e.g., vertices whose XYZ coordinates have been overwritten with masking token values) and reconstruct the missing coordinate data. The autoencoder can have information indicating that the masked mesh elements are part of the mesh and information indicating which mesh elements are neighbors (e.g., preserving connectivity information). However, the autoencoder is denied knowledge of the XYZ coordinates of the mesh element positions (or denied knowledge of other mesh element features) by the masking token values.

[0228] The autoencoder can be trained to reconstruct the input mesh in a way that fills in or estimates the missing parts of the input mesh according to the distribution of mesh examples in the training dataset. In one example, a masked reconstruction autoencoder can be trained to reconstruct the right lower central incisor based on a dataset of thousands of right lower central incisors, where randomly generated masks are applied to one or more of those training examples. The mask applies masking tokens to aspects of the training example (e.g., overwriting the XYZ coordinates of one or more mesh elements with masking token values). In some cases, it can be seen that the application of this mask enhances the training dataset. Over many iterations of training, the reconstruction autoencoder becomes trained to accurately reconstruct the left lower central incisor, even though these patches of mesh elements have been masked or obfuscated.

[0229] Mesh surface reconstruction can involve the process of regenerating one or more new sub-meshes within a larger mesh or modifying the positions and / or orientations of existing mesh elements (such as vertices) in anomalous portions of the mesh such that those mesh elements form the surface of a more desirable-shaped 3D representation, e.g., by modifying or replacing the vertices (or other mesh elements) of attachment structures in a tooth mesh such that these vertices now take the form of a smooth tooth surface (i.e., thereby effectively removing the 3D representation of the attachment from the mesh). One such approach is to implement a masked autoencoder (MAE), whereby a neural network (e.g., a reconstruction autoencoder) can be trained in an unsupervised manner to "mask" portions of the surface of a tooth or other oral care mesh (such as one or more dental arches, which may be provided directly from an intraoral scanner, before or after segmentation) at the input and reconstruct the entire unmasked input data at the output (see Figure 20 ) Figure 20An example training method for a masked autoencoder to fill in missing grid elements in a grid is illustrated. Once trained, labeled information (e.g., from reconstruction error calculations) can be used to inform the network about which vertices should be regenerated or repositioned (also known as deployed) during inference. The encoder-decoder structure of the MAE largely follows the same structure as other members of the autoencoder family, but masking of certain sub-elements within the input vector can also be considered. In this case, the input data can optionally be registered and / or correspondences applied given some template grid such that vertices can be consistently ordered across data samples, enabling more effective training and masking of appropriate grid elements. In another specific implementation, the MAE architecture can be modified to include aspects of other (generative) network architectures such as generative adversarial neural networks (GANs) to achieve more representative or accurate filling of labeled grid vertices. In the case of a GAN, the generator can be trained to perform grid surface reconstruction and the discriminator can be partially trained to train the generator.

[0230] The dental mesh M1 can be provided to a module (such as Figure 20 shown). This example describes how the MAE can be trained to fill in missing elements of a dental mesh, although the technique can be applied to any other oral care mesh (i.e., gums, other anatomical structures, brackets / attachments or other hardware such as prosthetic appliance components) or a mesh in general form. The elements of the tooth M1 (e.g., vertices, edges, faces or voxels) can be enumerated into a vector M2. This vector can be compared with a masking vector M3 and used to set the affected elements of M2 to a specified value (e.g., setting masked elements to zero). The masking vector can be provided to the module, for example, by the mesh element labeling method 1602 described herein. The M3 masked M2 vector can be provided to the encoder ME1, resulting in a latent vector MA. The latent vector MA can be provided to the decoder MD1, resulting in a reconstructed mesh M4. M4 can be compared with a ground truth vector M5 via a loss function such as mean squared error (MSE) or other functions described herein. In some specific implementations, the MSE can be calculated only on the reconstructed grid elements (the grid elements masked at the input by M3). In other specific implementations, the MSE can be calculated on other grid elements. Other loss functions are also possible, such as those disclosed elsewhere in this disclosure.

[0231] In some specific implementations, grid correspondences can be calculated to give the input vector M2 a better structure. The grid elements of M2 can be rearranged corresponding to one or more template grids. This operation is used to standardize the ordering of the grid elements in M2 and better enable the MAE to learn the structure of certain types of meshes (e.g., teeth or other forms of anatomical structures, or appliances or hardware).

[0232] The grid correspondence calculation procedure is beneficial because the correspondence can assign a known order to the vectors of the provided grid elements. This order can make the training of the encoder or decoder faster and more accurate. When the input data vectors have a consistent structure, the training of the encoder-decoder can converge at an earlier time, as opposed to when the input data vectors have a variable structure. The correspondence can provide a better structure for the input data, enabling the MAE model to focus more on the structure of the input tooth grid and use less network encoding to understand how to identify the structure on its own.

[0233] In some specific implementations, the autoencoder for grid reconstruction (MAE for grid filling) can be conditioned on optional inputs from elsewhere in the present disclosure, such as a vector P related to tooth size, one or more latent vectors B, and / or information R related to tooth name, nomenclature, tooth type, and / or tooth classification. The model can be conditioned on such optional inputs by concatenating such inputs with one or more input vectors (i.e., the inputs to the encoder) and / or by concatenating such inputs with one or more latent vectors output by the encoder in the autoencoder for grid reconstruction (MAE for grid filling).

[0234] In some cases, the autoencoder for grid filling can take as input one or more grid element features of one or more grid elements (as described herein). Such grid elements can improve the data quality of the latent encoding of the grid produced by the autoencoder (e.g., the encoder part of the autoencoder). Such grid element features can improve the ability of the encoder to encode aspects of the structure and / or shape of the received grid.

[0235] The training dataset may consist of an entire reasonably well-formed dental mesh (or other 3D oral care representation). The mesh should be understood to encompass other 3D representations described herein (e.g., 3D point clouds, voxels, surfaces, etc.). The dental mesh of the training dataset can be replicated and the copy can be modified to generate new data samples for training, thereby enhancing the training dataset. In some specific implementations, mesh processing techniques can be used to cut out parts of the tooth surface (or otherwise mark parts to be ignored, such as with masking tokens). For example, random mesh elements can be selected, and all mesh elements within a threshold distance (e.g., geodesic distance or Euclidean distance) can be cut out or otherwise marked as being ignored. In some specific implementations, unsupervised methods can be used to randomly cut out parts of the tooth surface (or mark parts to be ignored, such as with masking tokens) and train the model to reconstruct the missing parts of the mesh elements. In other specific implementations, parts of consecutive mesh elements can be marked with masking token values and then consumed by a masked autoencoder. The masking token values can be used to overwrite one or more values of the mesh element feature vector (e.g., XYZ coordinates or other mesh element features) of the mesh element to be masked. The masked autoencoder can access data indicating the presence / absence of one or more masked mesh elements in the mesh structure. In some specific implementations, information about the position of the mesh elements may not be given to the masked autoencoder (e.g., the X, Y, or Z coordinates can be removed). Some specific implementations may conform to one or more operations from the following list: 1. Perform registration on each tooth in the dataset to align each tooth with a known orientation. 2. Perform correspondence for each tooth in the dataset to obtain an ordered vector M2 of all mesh elements (e.g., vertices) related to a (predefined) tooth template. 3. Divide the vector M2 into equally sized chunks. 4. Randomly select one or more portions of the vector M2 to become a mask M3. Mark all mesh elements of each selected portion as "masked" (e.g., using a masking token). In some specific implementations, the randomly selected mesh elements can be designated as "seeds" for consecutive portions of mesh elements. Additional mesh elements can be added to the consecutive portion by traversing the edges of the mesh (e.g., in a depth-first or breadth-first manner, or a combination thereof). 5. Designate the masked mesh elements to be ignored / neglected by the masked autoencoder (e.g., by setting one or more elements of the associated mesh element feature vector of those elements to zero, null, or a masking token value). This is now the input vector to the masked autoencoder model at training time. 6. After the encoder of the masked autoencoder receives the input of the unmasked continuous part of the grid elements and generates an embedding vector as output, the masked grid elements are reinserted into their positions in the vector and influence the decoder to reconstruct the vector. The result of the reconstruction is a grid with filled data in the masked grid elements (e.g., grid elements with masked token values). The loss can be quantified Figure 20 to which the decoder MD1 reconstructs the latent vector MA into a reconstructed facsimile M4 of the input grid M1 (which, in some specific implementations, can be designated as the ground truth M5).

[0236] In some cases, the MAE network can be trained (e.g., using transfer learning) on examples of 3D oral care representations such as teeth. In some specific implementations, after training (e.g., for a classification task), the encoder output (e.g., which can include the latent embedding) can be used as a representation from which to classify the content of the input data. The embedding can be used to efficiently train another network (e.g., a set of fully connected layers) to classify the underlying data. In this case, the decoder can be set aside after training, and the trained encoder can be deployed for use during inference.

[0237] In some specific implementations, the full list (see steps 1-6 above) can be performed during both training and inference. In this approach, regions of the input grid that are problematic or need modification can be masked, and the role of the decoder can be to reconstruct / replace the masked regions, thereby fixing any defects that may be present in the input data (e.g., chips in teeth, breaks in appliances, or missing grid elements due to occlusion during intraoral scanning). In this case, during training, random parts of the teeth can be masked, allowing the model to learn to fill the masked regions in an unsupervised manner, and during inference, the damaged / corrupted regions can instead be masked in order to fill or complete those regions.

[0238] Figure 18 A training method of a masked autoencoder (MAE) for filling in missing aspects of a 3D oral care representation is shown. A 3D oral care representation is received at the input. Although this particular example focuses on teeth, according to aspects of the present disclosure, other types of 3D oral care representations described herein can be used. A mask can be randomly generated, where each element of the mask corresponds to a grid element and indicates whether that grid element should have some information hidden from the masked autoencoder (e.g., by covering elements of the grid element feature vector with a masked token). Information about the presence of the masked grid elements and / or information about which other grid elements are neighbors (e.g., by edges) can be provided to the MAE. However, in some specific implementations, certain grid element features can be retained from the MAE (e.g., the XYZ coordinates of the grid element, such as vertices or points).

[0239] A mask can be applied to the input 3D oral care representation to overwrite one or more elements of the mesh element feature vector of each mesh element (e.g., the XYZ coordinates of the affected mesh elements can be overwritten with masking tokens). The masked mesh elements can then be provided to the MAE, which can use the encoder to encode the masked data into a latent form (such as a latent vector) and reconstruct the latent vector into a copy of the input 3D oral care representation. A loss calculation can be performed to quantify the difference between the input form and the reconstructed form of the 3D oral care representation. A reconstruction loss and / or a KL divergence loss (or alternatively, one or more other losses disclosed herein) can be computed. Such losses can be used to at least partially train the encoder or decoder of the masked autoencoder. The loss calculation helps train the autoencoder to fill in any missing aspects of a trial input 3D oral care representation according to the distribution of 3D oral care representations in the training dataset. Many input 3D oral care representations (e.g., tooth meshes) from cohort patient cases can be used for training. During training, the MAE learns the distribution of the shapes and / or structures of typical teeth. In some embodiments, the MAE can be trained for a specific tooth type (e.g., the left lower central incisor), which has the advantage of improving data accuracy and the ability of the MAE to accurately reconstruct teeth of that type. The MAE implements a form of data augmentation because each input training sample can be augmented with a different mask, making the autoencoder robust to a variety of input data (e.g., different kinds of anomalies or missing mesh elements in the input 3D mesh or point cloud).

[0240] Figure 19 A deployment method for filling in missing aspects of a 3D oral care representation using a trained method for a masked autoencoder is shown. In other words, Figure 19 An example deployment of a masked autoencoder for reconstructing an oral care mesh in accordance with aspects of the present disclosure is illustrated. A trial input 3D oral care representation such as a tooth is provided to the input. The trial tooth may have some missing aspects, e.g., because a portion of the tooth was occluded during an intraoral scan. The tooth mesh elements are provided to the MAE, which encodes the tooth data into a latent form and reconstructs the latent form into a repaired, enhanced, or filled form of the input 3D oral care representation. The filling, enhancing, or repairing operation utilizes what the MAE learned during training about the distribution of tooth shapes and / or structures.

[0241] In some specific implementations, the MAE for mesh reconstruction can assist the tooth mesh segmentation operation by cleaning the results of dental mesh segmentation. After each crown is excised from the original unsegmented dental arch, the crown can benefit from the operation of the MAE. For example, the MAE can fill in missing surfaces that are missed by an intraoral scanner due to occlusion by adjacent teeth (e.g., along the mesial or distal surfaces of the tooth). In some examples, sharp edges on the tooth mesh can indicate missing data. In such cases, the MAE operation will be used as a means to improve the tooth segmentation output after tooth segmentation is completed.

[0242] In some specific implementations, the operational (deployed) MAE can take as input a mesh and a mask that defines the regions of the mesh to be changed / replaced by the MAE model. Then, the MAE can proceed to fill in the missing mesh elements as defined by the mask. Data augmentation can be employed to expand the dataset available for training the MAE. Many training data examples can be generated from the same tooth mesh by varying the mask applied to the tooth. A random subgroup of tooth mesh elements can be selected and masked. The resulting mesh and mask can be added to the training dataset. The same tooth mesh can be combined with different random masks, and the resulting pairs are also added to the training dataset. This can be done at runtime, with the advantage of avoiding the need to store large datasets. Instead, in some specific implementations, the dataset can be generated on the fly, and those samples can be discarded immediately after being used for that training batch, and so on. In some examples, the randomly selected subgroup of tooth mesh elements can be contiguous to represent a contiguous portion of the tooth mesh surface. In other examples, the randomly selected subgroup of tooth mesh elements can correspond to more than one contiguous region of the tooth mesh surface.

[0243] In some specific implementations, the MAE can be trained to modify the position and / or orientation of one or more teeth in the dental arch for orthodontic setup purposes. In some cases, the teeth may have been previously segmented. Through an algorithm or some other process, some teeth in the dental arch can be designated for movement and other teeth are designated for non - movement (i.e., fixed).

[0244] In some specific implementations, the MAE can be used to modify elements of one or more tooth transformations (e.g., a 4×4 transformation matrix or a vector transformation, etc.). Such transformations can be used to place the tooth, appliance component, or fixture model component into a pose suitable for appliance generation. For example, the tooth can be placed into an orthodontic setup pose. For example, the MAE can be trained to fill in the elements of the transformation matrix (or other transformation). In some cases, all transformation elements can be filled, and in other cases, only a subgroup of the transformation elements can be filled by the techniques of the present disclosure.

[0245] In some cases, one or more teeth can be designed not to move during orthodontic treatment. The MAE can be trained to manipulate the teeth to move in a way that keeps the anchored teeth from moving. Benefits can include minimizing the movement of the patient's teeth, which can also provide better anchorage for the appliance to move the non-moving teeth.

[0246] Figure 21 An example training method for masking capsule autoencoder to fill in missing grid elements in a grid according to one or more aspects of the present disclosure is illustrated. In some specific implementations, such as in Figure 21 , the MAE for grid filling can replace the encoder structure and the decoder structure with a capsule encoder structure and a capsule decoder structure, respectively. Other specific implementations can be replicated from the foregoing description.

[0247] In some specific implementations, a trained neural network can be completed for a grid (or point cloud or voxel) with the aid of skip connections. The skip connections can connect the output of the encoder to the input of the corresponding decoder at each resolution level in a series of resolution levels (e.g., such as in a U-Net). The skip connections between the encoder and the decoder can facilitate the training / learning of the representation of the input 3D oral care representation and generally can facilitate the information flow around the network.

[0248] In some specific implementations, the generation implementations described herein can include one or more hierarchical feature extraction modules (e.g., modules that extract global, intermediate, or local neural network features from 3D representations such as point clouds). Examples of hierarchical neural network feature extraction modules (HNNFEMs) include 3D SWIN transformer architectures, U-Nets, or pyramid encoder-decoders, etc. The HNNFEM can be trained to generate multi-scale voxel (or point or other grid element) embeddings of the 3D representation (or multi-scale embeddings of other grid elements described herein). For example, one or more layers (or levels) of the HNNFEM can be trained on a 3D representation of a patient's dentition to generate neural network feature embeddings that cover global, intermediate, or local aspects of the 3D representation of the patient's dentition.

[0249] In some specific implementations, the techniques of the present disclosure can use PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using PointNet or PointNet++ as a basis for training) to extract local or global neural network features from 3D point clouds or other 3D representations (e.g., 3D point clouds that describe aspects of a patient's dentition such as teeth or gums). In some specific implementations, the techniques of the present disclosure can use a U-Net to extract local or global neural network features from 3D point clouds or other 3D representations.

[0250] 3D oral care representations are described herein as such because 3D representations are current state of the art. However, 3D oral care representations are intended to be used in a non-limiting manner to cover any representation of 3D or higher order dimensions (e.g., 4D, 5D, etc.), and it should be understood that the techniques disclosed herein can be used to train machine learning models to operate on representations of higher order dimensions.

[0251] In some cases, the input data may include 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data related to splines (e.g., control points). The encoder-decoder structure may include one or more encoders, or one or more decoders. In some specific implementations, the encoder may take as input the mesh element feature vectors of one or more input mesh elements in the input mesh. The encoder is trained in a manner that processes the mesh element feature vectors to generate a more accurate representation of the input data. For example, the mesh element feature vectors may provide the encoder with more information about the shape and / or structure of the mesh, and thus the additional information provided allows the encoder to make more informed decisions and / or generate a more accurate latent representation of the mesh. Examples of encoder-decoder structures include U-Net, autoencoders, or transformers, among others. The representation generation module may include one or more encoder-decoder structures (or portions of encoder-decoder structures, such as individual encoders or individual decoders). The representation generation module may generate an information-rich (optionally dimension-reduced) representation of the input data that can be more readily consumed by other generative or discriminative machine learning models.

[0252] U-Net may include an encoder followed by a decoder. The architecture of U-Net may be similar to a U shape. The encoder may extract one or more global neural network features, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level compared to the most global level) from the input 3D representation. The output of each level from the encoder may be passed to the input of the corresponding level of the decoder (e.g., via skip connections). Similar to the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For example, the decoder may output a representation of the input data that may include global, intermediate, or local information about the input data. In some specific implementations, U-Net may generate an information-rich (optionally dimension-reduced) representation of the input data that can be more readily consumed by other generative or discriminative machine learning models.

[0253] An autoencoder can be configured to encode input data into a latent form. The autoencoder can train an encoder to reformulate the input data into a latent form with reduced dimensions between the encoder and the decoder, and then train the decoder to reconstruct the input data from this latent form of the data. A reconstruction error can be computed to quantify the extent to which the reconstructed form of the data differs from the input data. In some specific implementations, the latent form can be used as an information-rich reduced-dimension representation of the input data, which can be more readily consumed by other generative or discriminative machine learning models. In most scenarios, the autoencoder can be trained to take a 3D representation as input, encode the 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close replica of the input 3D representation as output.

[0254] A transformer can be trained to generate a representation of its input using self-attention at least in part. The transformer can encode long-range dependencies (e.g., encode the relationships between a large number of inputs). The transformer can include an encoder or a decoder. In some specific implementations, such an encoder can operate in a bidirectional manner or can operate a self-attention mechanism. In some specific implementations, such a decoder can operate a masked self-attention mechanism, can operate a cross-attention mechanism, or can operate in an autoregressive manner. In some specific implementations, the self-attention operation of the transformer described herein can be related to different positions or aspects of a single 3D oral care representation in order to compute a reduced-dimension representation of the 3D oral care representation. In some specific implementations, the cross-attention operation of the transformer described herein can mix or combine aspects of two (or more) different 3D oral care representations. In some specific implementations, the autoregressive operation of the transformer described herein can consume previously generated aspects (e.g., previously generated points, point clouds, transformations, etc.) of a 3D oral care representation as additional inputs when generating a new or modified 3D oral care representation. In some specific implementations, the transformer can generate a latent form of the input data that can be used as an information-rich reduced-dimension representation of the input data, which can be more readily consumed by other generative or discriminative machine learning models.

[0255] In some specific implementations, an encoder-decoder architecture can first be trained as an autoencoder. In deployment, one or more modifications can be made to the latent form of the input data. Then, the modified latent form can continue to be reconstructed by the decoder, thereby producing a reconstructed form of the input data that differs from the input data in one or more expected aspects. Oral care variables such as oral care parameters or oral care metrics can be provided to the encoder, the decoder, or can be used to modify the latent form in order to affect the encoder-decoder architecture when generating a reconstructed form with desired characteristics (e.g., characteristics that may be different from those of the input data).

[0256] In some cases, federated learning can be used to train the techniques of the present disclosure. Federated learning can enable multiple remote clinicians to iteratively improve a machine learning model (e.g., verification of 3D oral care representations, mesh segmentation, mesh cleaning, other techniques involving labeling mesh elements, coordinate system prediction, placement of non-organic objects on teeth, appliance component generation, dental restoration design generation, techniques for placing 3D oral care representations, setting prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, estimation of missing values), while protecting data privacy (e.g., clinical data may not need to be sent "over the network" to a third party). Data privacy is particularly important for clinical data protected by applicable laws. A clinician can receive a copy of the machine learning model, use a local machine learning program to further train the ML model using locally available data from the local clinic, and then send the updated ML model back to a central hub or a third party. The central hub or third party can integrate the updated ML models from multiple clinicians into a single updated ML model that benefits from learning on patient data recently collected at individual clinical sites. In this way, a new ML model can be trained that benefits from additional and updated patient data (which may come from multiple clinical sites), while that patient data is never actually sent to a third party. In some cases, training on devices within a local clinic can be performed when the device is idle or otherwise during non-working hours (e.g., when patients are not being treated at the clinic). Devices in a clinical environment for collecting data and / or training an ML model for the techniques described herein can include intraoral scanners, CT scanners, X-ray machines, laptops, servers, desktop computers, or handheld devices (such as smartphones with image collection capabilities). In addition to federated learning techniques, in some embodiments, contrastive learning can be used to at least partially train the ML models described herein. In some cases, contrastive learning can augment samples in a training dataset to emphasize differences between samples of different classes and / or increase the similarity of samples of the same class.

[0257] In some cases, a local coordinate system for 3D oral care representations (such as teeth) can be described by one or more transformations (e.g., an affine transformation matrix, a translation vector, or a quaternion). The systems of the present disclosure can be trained for coordinate system prediction using the coordinate systems of past cohort patient case data. The past patient data can include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems. Machine information models (such as U-Net, an encoder, an autoencoder, a pyramid encoder-decoder, a transformer, or convolutional layers and / or pooling layers) can be trained for coordinate system prediction. Representation learning can determine a representation of a tooth (e.g., encoding a mesh or point cloud into a latent representation, e.g., using U-Net, an encoder, a transformer, or convolutional layers and / or pooling layers, etc.), and then predict a transformation for that representation (e.g., using a trained multi-layer perceptron, a transformer, an encoder, a transformer, etc.), where the transformation defines a local coordinate system for that representation (e.g., including one or more coordinate axes). In the case of predicting a coordinate system for a tooth mesh, the mesh convolution techniques described herein can utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh. Pose transfer techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth. Reinforcement learning techniques can be trained for coordinate system prediction in the form of predicting a transformation for a tooth.

[0258] Machine information models (such as U-Net, encoder, autoencoder, pyramid encoder-decoder, transformer, or convolutional and / or pooling layers) can be trained as part of a method for hardware (or appliance component) placement. Representation learning can train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder or using blocks of U-Net, encoder, transformer, convolutional layer, and / or pooling layer, etc.). The representation can include a dimension-reduced form and / or an information-rich form of the input 3D oral care representation. In some embodiments, generation of the representation can be assisted by computing a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some embodiments, a representation can be computed for a hardware element (or appliance component). Such a representation is adapted to be provided to a second module, which can perform a generation task, such as transformation prediction (e.g., a transformation for placing a 3D oral care representation relative to another 3D oral care representation, such as a representation for placing a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such a transformation can include an affine transformation matrix, a translation vector, or a quaternion, etc. Machine learning models that can be trained to predict a transformation for placing a hardware element (or appliance component) relative to an element of a patient dentition include MLP, transformer, encoder, etc. The systems of the present disclosure can be trained for 3D oral care appliance placement using 3D oral care appliance placement of past cohort patient case data. Past patient data can include at least: one or more ground truth transformations and one or more 3D oral care representations (such as tooth meshes or other elements of a patient dentition). In the case where a U-Net (and other neural networks) is trained to generate a representation of a tooth mesh, the mesh convolution and / or mesh pooling techniques described herein utilize invariance to rotation, translation, and / or scaling of the tooth mesh to generate predictions that cannot be generated by techniques that are not invariant to rotation, translation, and / or scaling of the tooth mesh.

[0259] The techniques described herein can be trained to generate or modify 3D oral care representations (e.g., dental restoration designs, appliance components, and other examples of 3D oral care representations described herein). For example, these techniques can generate or modify mesh element tags for 3D representations, enabling those 3D representations to undergo mesh segmentation, mesh cleaning, or mesh modification. Such 3D representations can include point clouds, polylines, meshes, voxels, etc. Such 3D oral care representations can be generated according to the requirements of oral care variables that may be provided to the generation model in some embodiments. Oral care variables can include oral care parameters as disclosed herein, or other real-valued, text-based, or categorical inputs that specify expected aspects of one or more 3D oral care representations to be generated. In some cases, oral care variables can include oral care metrics that can describe expected aspects of one or more 3D oral care representations to be generated (e.g., to describe the expected aspects of mesh element tags generated for use in mesh cleaning). Oral care variables are particularly applicable to the embodiments described herein. For example, oral care variables can specify the expected design (e.g., including shape and / or structure) of 3D oral care representations that can be generated (or modified) according to the techniques described herein. In short, embodiments that use specific oral care variables disclosed herein generate more accurate 3D oral care representations than embodiments that do not use specific oral care variables. In some cases, a text encoder can encode a set of natural language instructions from a clinician (e.g., generate a text embedding). The text string can include tokens. In some embodiments, the encoder used to generate the text embedding can apply average pooling or max pooling between token vectors. In some cases, a transformer (e.g., BERT or Siamese BERT) can be trained to extract embeddings of text used in digital oral care (e.g., by training the transformer on examples of clinical text such as those given below). In some cases, such models used to generate text embeddings can be trained using transfer learning (e.g., initially trained on another text corpus and then receiving further training on text related to digital oral care). Some text embeddings can encode text at the word level. Some text embeddings can encode text at the token level. In some embodiments, the transformer used to generate the text embedding can be trained at least in part using a loss calculation that compares a predicted output to a ground truth output (e.g., softmax loss, multi-negative example ranking loss, MSE margin loss, cross-entropy loss, etc.). In some cases, non-text variables (such as real-valued or categorical values) can be converted to text and then embedded using the techniques described herein.Examples of natural language instructions that can be issued by a clinician to the generative model described herein are: "For each crown in the dental arch, fill in the missing portions of the tooth (e.g., portions that may have been missing during an intraoral scan)", or "Mark abnormal mesh elements of the dental arch, i.e., abnormal mesh elements that may correspond to errors in the scan or clinical diagnoses". Example :

[0260] Example 1. A method, the method comprising: Receiving an input 3D representation of oral care data by a processing circuit of a computing device; Modifying the input 3D representation of the oral care data by the processing circuit at least in part by replacing at least one coordinate of a mesh element of the input 3D representation of the oral care data with a masking token to form a modified 3D representation of the oral care data; Providing the modified 3D representation of the oral care data as training data to a masked autoencoder model by the processing circuit; and Training the masked autoencoder model by the processing circuit to reconstruct a copy of the input 3D representation of the oral care data.

[0261] Example 2. The method according to Example 1, wherein the mesh element is at least one of a vertex or a point.

[0262] Example 3. The method according to Example 1, wherein the input 3D representation of the oral care data is a 3D mesh.

[0263] Example 4. The method according to Example 1, wherein the input 3D representation of the oral care data is a 3D point cloud.

[0264] Example 5. The method according to Example 1, wherein the masking token signals to the masked autoencoder the presence of the mesh element in the structure of the input 3D representation of the oral care data, and wherein the coordinate of the mesh element is not available to the autoencoder.

[0265] Example 6. The method according to Example 1, wherein the modified 3D representation of the oral care data includes a plurality of mesh elements masked in consecutive blocks.

[0266] Example 7. The method according to Example 1, the method further comprising applying a mask to the input 3D representation of the oral care data before providing the training data to the masked autoencoder.

[0267] Example 8. The method according to Example 1, wherein the input 3D representation of the oral care data represents teeth.

[0268] Example 9. The method according to Example 8, wherein training the masked autoencoder includes training the masked autoencoder based on distribution information associated with the input 3D representation of the oral care data.

[0269] Example 10. The method according to Example 7, the method further comprising randomly generating the mask.

[0270] Example 11. The method according to Example 1, wherein the masked autoencoder includes a multi-dimensional encoder configured to encode the input 3D representation of the oral care data into a latent space representation and a multi-dimensional decoder configured to reconstruct the latent space representation into a facsimile of the input 3D representation of the oral care data.

[0271] Example 12. The method according to Example 1, the method further comprising calculating a reconstruction error to quantify the difference between the input 3D representation of the oral care data and the facsimile of the input 3D representation of the oral care data.

[0272] Example 13. The method according to Example 12, wherein calculating the reconstruction error includes calculating a term associated with at least one of a reconstruction loss calculation or a KL divergence calculation.

[0273] Example 14. The method according to Example 1, the method further comprising providing at least one of the following as additional input data to the masked autoencoder: (i) one or more vectors P, the vectors P including at least one value related to at least one method of calculating the size of at least one tooth, or (ii) one or more vectors R of at least one of tooth name, nomenclature, tooth type, and tooth classification.

[0274] Example 15. The method according to Example 1, wherein the masked autoencoder is a reconstruction autoencoder.

[0275] Example 16. The method according to Example 1, wherein the reconstruction autoencoder is a variational autoencoder (VAE).

[0276] Example 17. A method, the method comprising: receiving, by a processing circuit of a computing device, an input 3D representation of oral care data; providing, by the processing circuit, the 3D representation of the oral care data as training data to a trained masked autoencoder model as an execution-phase input; executing, by the processing circuit, the trained masked autoencoder model to fill one or more missing mesh elements of the input 3D representation of the oral care data, thereby forming a reconstructed 3D representation of the oral care data; and Output the reconstructed oral care representation.

[0277] Example 18. The method according to Example 17, the method further comprising identifying, by the processing circuit, one or more missing mesh elements of the input 3D representation of the oral care data.

[0278] Example 19. The method according to Example 17, wherein the one or more missing mesh elements include at least one of the following: one or more vertices, one or more faces, one or more edges, one or more points, or one or more voxels.

[0279] Example 20. The method according to Example 17, wherein the input 3D representation of the oral care data is a 3D mesh.

[0280] Example 21. The method according to Example 17, wherein the input 3D representation of the oral care data is a 3D point cloud.

[0281] Example 22. The method according to Example 17, wherein the input 3D representation of the oral care data represents teeth.

[0282] Example 23. The method according to Example 17, wherein the masked autoencoder is a variational autoencoder (VAE).

[0283] Example 24. The method according to Example 17, wherein the computing device is deployed in a clinical environment.

[0284] Example 25. The method according to Example 24, wherein the method is executed in the clinical environment.

[0285] Example 26. A method, the method comprising: Receiving, by a processing circuit of a computing device, an input 3D representation of oral care data; Providing, by the processing circuit, the input 3D representation of the oral care data as training data to a completion model, wherein the completion model includes at least one of the following: at least one encoder, at least one decoder, or at least one skip connection between the encoder and the decoder; and Training, by the processing circuit, the completion neural network model to reconstruct a facsimile of the input 3D representation of the oral care data, wherein one or more missing aspects of the input 3D representation of the oral care data have been filled.

Claims

1. A method, the method comprising: Receiving, by a processing circuit of a computing device, an input three-dimensional (3D) representation of oral care data; Executing, by the processing circuit, an encoder of a trained autoencoder network to encode the input 3D representation of the oral care data into a latent space representation having a lower-order dimension than a first 3D representation of the oral care data; Executing, by the processing circuit, a decoder of the trained autoencoder network to reconstruct the latent space representation, thereby forming an output 3D representation of the oral care data, the output 3D representation of the oral care data being a replica of the input 3D representation of the oral care data; And Calculating, by the processing circuit, a reconstruction error to quantify a difference between at least one mesh element of the input 3D representation of the oral care data and a corresponding at least one mesh element of the output 3D representation of the oral care data.

2. The method according to claim 1, wherein the input 3D representation of the oral care data comprises at least one of a 3D mesh, a point cloud, or a voxelized representation.

3. The method according to claim 1, the method further comprising performing, by the processing circuit, anomaly detection based on the reconstruction error.

4. The method according to claim 1, wherein each of the at least one mesh element and the corresponding at least one mesh element comprises at least one of a corresponding edge, a corresponding vertex, a corresponding face, a corresponding voxel, or a corresponding point of a point cloud.

5. The method according to claim 1, the method further comprising: Determining that the reconstruction error exceeds a predetermined threshold; And In response to determining that the reconstruction error exceeds the predetermined threshold, assigning, by the processing circuit, a classification label to the at least one mesh element or the corresponding at least one mesh element.

6. The method according to claim 1, wherein the trained autoencoder network comprises a trained variational autoencoder (VAE) network.

7. The method according to claim 1, wherein calculating the reconstruction error comprises calculating a distance between at least one aspect of the input 3D oral care representation and a corresponding aspect in the output 3D representation of the oral care data.

8. The method according to claim 7, wherein the trained autoencoder network is trained according to a paradigm comprising normalizing flows.

9. The method according to claim 1, the method further comprising providing, by the processing circuit, additional execution phase inputs to the trained autoencoder network, the additional execution phase inputs comprising one or more 3D meshes or one or more 3D point clouds of one or more teeth describing the input 3D representation of the oral care data.

10. The method according to claim 1, the method further comprising providing, by the processing circuit, at least one of the following as additional input data to the autoencoder model: (i) one or more vectors P, the vectors P comprising at least one value related to at least one method of calculating the dimensions of at least one tooth, or (ii) one or more vectors R of at least one of tooth name, nomenclature, tooth type, and tooth classification.

11. The method according to claim 1, the method further comprising calculating, by the processing circuit, a loss associated with at least one of the following: a term associated with reconstruction loss calculation, or a term associated with KL divergence loss calculation.

12. The method according to claim 1, wherein the computing device is deployed in a clinical environment, and wherein the method is performed near real-time during a session with a patient.

13. A device, the device comprising: interface hardware configured to receive an input three-dimensional (3D) representation of oral care data; a processing circuit configured to: execute an encoder of a trained autoencoder network to encode the input 3D representation of oral care data into a latent space representation; execute a decoder of the trained autoencoder network to reconstruct the latent space representation, thereby forming an output 3D representation of oral care data, the output 3D representation of oral care data being a replica of the input 3D representation of oral care data; and calculate a reconstruction error to quantify a difference between at least one mesh element of the input 3D representation of oral care data and a corresponding at least one mesh element of the output 3D representation of oral care data; and a memory unit configured to store the reconstruction error.

14. The device according to claim 13, wherein the input 3D representation of oral care data comprises at least one of a 3D mesh, a point cloud, or a voxelized representation.

15. The device according to claim 13, wherein the processing circuit is further configured to perform anomaly detection based on the reconstruction error.

16. The device according to claim 13, wherein each of the at least one mesh element and the corresponding at least one mesh element comprises at least one of a corresponding edge, a corresponding vertex, a corresponding face, a corresponding voxel, or a corresponding point of a point cloud.

17. The device according to claim 13, wherein the processing circuit is further configured to: determine that the reconstruction error exceeds a predetermined threshold; and in response to determining that the reconstruction error exceeds the predetermined threshold, assign a classification label to the at least one mesh element or the corresponding at least one mesh element by the processing circuit.

18. The device according to claim 13, wherein the trained autoencoder network comprises a trained variational autoencoder (VAE) network.

19. The device according to claim 13, wherein, in order to calculate the reconstruction error, the processing circuit is configured to calculate a distance between at least one aspect of the input 3D representation of oral care data and a corresponding aspect of the output 3D representation of oral care data.

20. The device according to claim 13, wherein the processing circuit is further configured to provide an additional execution stage input to the trained autoencoder network, the additional execution stage input comprising one or more 3D meshes or one or more 3D point clouds of one or more teeth describing the input 3D oral care representation.

Citation Information

Patent Citations

  • Method for automated generation of orthodontic treatment final setups

    US20210259808A1

  • Method for automated generation of orthodontic treatment final setups

    WO2020026117A1

  • System to generate staged orthodontic aligner treatment

    WO2021245480A1

  • Automated processing of dental scans using geometric deep learning

    WO2022123402A1