Geometric Deep Learning for Setup and Staging in Clear Tray Aligners

JP2024542691A5Pending Publication Date: 2025-12-05SOLVENTUM INTELLECTUAL PROPERTIES CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024532455
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-03
Filing Date
2022-11-29
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing automated systems for generating clear tray aligners (CTAs) in orthodontic treatment face challenges in accurately determining optimal tooth trajectories due to the complexity of tooth movements, often resulting in incomplete alignment and the need for human intervention to correct errors.

Method used

The use of generative adversarial networks (GANs) to train neural networks for predicting tooth movements, incorporating geometric deep learning techniques to improve the accuracy of CTA setups by distinguishing between predicted and reference tooth movements, and adjusting neural network weights based on discrepancies.

Benefits of technology

This approach enhances the precision and automation of CTA generation, reducing computational overhead and the need for human correction, resulting in more accurate digital representations of intermediate and final tooth alignments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and methods are described for training and using a generative adversarial network (GAN) to produce intermediate stages and a final setup for a clear tray aligner (CTA), the systems including: one or more computer processors receiving a first digital representation of a patient's teeth; the one or more computer processors using a neural network included in the GAN, a generator trained to predict the one or more tooth movements, to determine a prediction of the one or more tooth movements; and the one or more processors generating an output state including at least one of a final setup and one or more intermediate stages.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to the construction and training of neural networks to improve the accuracy of automatically generated clear tray aligner (CTA) devices used in orthodontic treatment. [Background technology]

[0002] Intermediate staging of teeth from the malocclusion stage to the final stage requires determining precise individual tooth movements so that the teeth do not collide with each other, the teeth move towards their final state, and the teeth follow optimal and preferably short trajectories. Since each tooth has six degrees of freedom and the average dental arch has about 14 teeth, finding optimal tooth trajectories from the initial to the final stage is a large and complex task.

[0003] Previous approaches to automating the fabrication of CTAs have involved the use of certain rules or metrics to quantify the condition of a set of teeth for reconditioning using one or more CTA devices. Other approaches have attempted to use machine learning techniques to generate CTA devices with mixed results. As a result, there is a need for better machine learning models and training techniques to improve systems that automate the fabrication of CTAs. Summary of the Invention

[0004] This disclosure describes systems and techniques for training and using generative adversarial networks (GANs) to create intermediate and final setups for CTA. In a first aspect, a first computer-implemented method of generating a setup for orthodontic alignment treatment is described, comprising: one or more computer processors receiving a first digital representation of a patient's teeth; the one or more computer processors using a generator, a neural network included in a generative adversarial network (GAN), initially trained to predict the one or more tooth movements for the final setup, to determine a prediction of one or more tooth movements for the final setup; and the one or more computer processors further training the GAN based on the using step, wherein training the GAN comprises: the generator predicting the one or more tooth movements for the final setup based on the first digital representation of the patient's teeth; and a classifier, which is also a neural network configured to distinguish between the predicted tooth movements and reference tooth movements and is also part of the GAN, is modified by performing operations including: determining whether the representation of the one or more tooth movements predicted by the generator is distinguishable from the representation of the one or more reference tooth movements; and modifying one of the neural networks for at least one of the generator and the classifier based on a determination of the classifier.

[0005] The first aspect may optionally include additional features. For example, the method may generate, by one or more processors, an output state for a final setup. The method may determine, by one or more computer processors, a difference between the one or more predicted tooth movements and the one or more reference tooth movements. The determined difference between the one or more predicted tooth movements and the one or more reference tooth movements may be used to modify the training of the generator. Modifying the training of the generator may include adjusting one or more weights of a neural network of the generator. The method may generate, by one or more computer processors, one or more lists specifying elements of the first digital representation of the patient's teeth. At least one of the one or more lists may specify one or more edges in the first digital representation of the patient's teeth. At least one of the one or more lists may specify one or more polygonal faces in the digital representation of the patient's teeth. At least one of the one or more lists may specify one or more vertices in the first digital representation of the patient's teeth. The method may calculate, by one or more computer processors, one or more mesh features. The one or more mesh features can include edge endpoints, edge curvatures, edge normal vectors, edge movement vectors, edge normalized lengths, vertices, faces of the associated three-dimensional representation, voxels, and combinations thereof. The method can generate, by one or more computer processors, a digital representation that predicts the positions and orientations of the patient's teeth based on the one or more predicted tooth movements. The method can generate, by one or more computer processors, a digital representation of the patient's teeth based on the one or more reference tooth movements.Determining by the identifier whether the one or more tooth movement representations predicted by the generator are distinguishable from the one or more reference tooth movement representations may include receiving the one or more tooth movement representations predicted by the generator, the one or more reference tooth movement representations, and a first digital representation of the patient's teeth; comparing the one or more tooth movement representations predicted by the generator to the one or more reference tooth movement representations, where the comparison is based at least in part on the first digital representation of the patient's teeth; and determining, by the one or more computer processors, a probability that the one or more tooth movement representations predicted by the generator are the same as the one or more reference tooth movement representations.

[0006] In a second aspect, the method includes the steps of one or more computer processors receiving a first digital representation of the patient's teeth and a representation of the final set-up; the one or more computer processors using a generator, a neural network included in a generative adversarial network (GAN), initially trained to predict the one or more tooth movements for the one or more intermediate stages, to determine predictions of the one or more tooth movements for the one or more intermediate stages; and the one or more computer processors further training the GAN based on the using step, the training of the GAN comprising the generator determining at least one of the following predictions based on the first digital representation of the patient's teeth: A second computer-implemented method of generating a setup for an orthodontic alignment treatment is described, comprising: predicting one or more tooth movements for at least one intermediate stage; and a classifier, which is also a neural network and part of a GAN configured to distinguish between the predicted tooth movements and a reference tooth movement, is modified by performing operations including: determining whether the representation of the one or more tooth movements predicted by the generator is distinguishable from the representation of the one or more reference tooth movements; and modifying one of the neural networks for at least one of the generator and the classifier based on the determination of the classifier. The second aspect may also include one or more of the optional features described above with reference to the first aspect.

[0007] In a third aspect, a third computer-implemented method of generating a setup for orthodontic alignment treatment is described, comprising: one or more computer processors receiving a first digital representation of the patient's teeth; the one or more computer processors using a generator, which is a neural network included in a generative adversarial network (GAN) and trained to predict the one or more tooth movements, to determine a prediction of the one or more tooth movements; and the one or more computer processors generating an output state comprising at least one of a final setup and one or more intermediate stages, wherein the GAN has been trained using operations including: the generator predicts the one or more tooth movements based on the first digital representation of the patient's teeth; and a classifier, which is also a neural network configured to distinguish between the predicted tooth movements and reference tooth movements, and which is also part of the GAN, determining whether the representation of the one or more tooth movements predicted by the generator is distinguishable from the representation of the one or more reference tooth movements; and modifying one of the neural networks for at least one of the generator and the classifier based on a determination of the classifier. The third aspect may also include one or more of the optional features discussed above with reference to the first aspect. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an example approach that can be used to train a machine learning model used to determine the final setup of a CTA. [Diagram 2] FIG. 2 is an exemplary visualization of a workflow performed using the technique shown in FIG. 1. [Diagram 3] FIG. 2 is another diagram of the approach shown in FIG. 1. [Figure 4] FIG. 1 illustrates an exemplary approach that can be used to train a machine learning model used to determine intermediate staging of CTA. [Diagram 5]FIG. 5 is an exemplary visualization of a workflow performed using the technique shown in FIG. 4. [Figure 6] FIG. 5 is another diagram of the approach shown in FIG. [Figure 7] FIG. 2 is an expanded view of the approach shown in FIG. 1, focusing on aspects of the approach that use geometric deep learning. [Figure 8] FIG. 5 illustrates an example workflow using the U-Net architecture for the generator shown in either FIG. 1 or FIG. 4. [Figure 9] FIG. 9 illustrates an exemplary U-Net architecture shown in FIG. [Figure 10] FIG. 10 illustrates an example workflow 1000 for the generator 110 shown in either FIG. 1 or FIG. [Figure 11] FIG. 11 illustrates an exemplary pyramid encoder-decoder shown in FIG. [Figure 12] FIG. 11 illustrates the exemplary encoder shown in FIGS. 8 and 10. [Figure 13] FIG. 2 illustrates an example processing unit that operates in accordance with the techniques of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Clear tray aligners, or CTAs, are a series of orthodontic forms used to reposition and / or reorient any number of a patient's teeth over the course of treatment. Once the patient's teeth have been fitted into one tray or form, the existing tray can be replaced with the next tray in the sequence to achieve the desired result. CTAs can be made of a variety of materials, but as the name suggests, they are generally transparent, so the trays can be worn all day to achieve the desired effect without being overly cosmetically disturbed. As used herein, "final set-up" refers to the target arrangement of teeth that corresponds to the final CTA in the sequence (i.e., representing the final desired alignment of the patient's teeth). "Intermediate stages" are used herein to identify intermediate arrangements of teeth that correspond to other trays in the sequence that are used to reach the final set-up.

[0010] Automated tools have been developed for the creation of digital final setups and intermediate stages that can be used to generate physical trays representing the intermediate and final setups. In one known embodiment, a landmark-based, anatomical structure-based approach was used that attempts to quantify the condition of a set of teeth according to certain metrics or rules to determine the digital representations of the intermediate and final setups. In another embodiment, neural networks were used for the generation of digital representations for the final setups and intermediate stages.

[0011] However, research with these solutions has revealed areas that need improvement. One area that needs improvement is that the models implemented by the above techniques may not be able to generate all of the tooth movements that would be required to position the patient's teeth as desired. For example, it has been observed that the calculated tooth movements may continue to cause one or more teeth to overlap, to cite one example. This may ultimately result in an arrangement of aligners that does not provide the desired cosmetic results, since certain teeth may still overlap even after the final setup is worn by the patient. The digital representation may need to be processed a further number of times to address identified shortcomings in the trained model, resulting in further computational overhead, which unnecessarily consumes computational resources. Even then, it may not be possible to finalize the digital representation, and thus human intervention may be required to correct the resulting output of these systems and techniques. In other words, an advantage of the present disclosure provides a better trained system that results in a more accurate digital representation of the underlying system and improved automation, to cite two examples. In essence, a system implementing the disclosed techniques will be better trained, will generate more accurately generated digital representations of each intermediate stage and final setup, and will do so in a shorter duration.

[0012] FIG. 1 is an exemplary method 100 that can be used to train a machine learning model used to determine the final setup of the CTA. As described in more detail below, the method 100 can be implemented on computer hardware to achieve a desired result. A receiving module 102 receives patient case data. Generally, the patient case data represents a digital representation of a patient's mouth. As shown, the patient case data received by the module 102 can be received as either a maloccluded dental arch 106 (e.g., a three-dimensional ("3D") mesh representing the patient's upper and lower arches of teeth), a maloccluded dental arch 106 as shown, but positioned in an "occlusal" position 104 where the upper and lower arches are engaged with one another, or a combination of the two. According to certain implementations, the 3D mesh representing the occlusal position 104 can also include a 3D mesh geometry of the patient's gingival tissue (i.e., gums) in addition to mesh data of the patient's teeth. The illustrative portion of FIG. 1 is presented as being agnostic to the type of training being performed (i.e., training for a final setup or training for an intermediate stage). However, when analyzing those transformations with reference to mesh transformation training, it should be understood that the method 100 is intended to operate with the purpose of training a neural network to more accurately and quickly generate the final setup. Intermediate stage training is described with reference to FIG.

[0013] It should be understood that, according to certain implementations, the occlusion position geometry 104 and the maloccluded dental arch geometry 106 may include or be otherwise defined by the same or similar 3D geometries, but arranged in a particular configuration. That is, in some circumstances, the occlusion position geometry 104 includes the same underlying mesh data arranged or otherwise rendered to represent an occlusal configuration as contained within the maloccluded dental arch geometry 106. Thus, for example, if the receiving module 102 receives only the maloccluded dental arch geometry 106, the receiving module 102 may automatically generate the occlusion position geometry 104. Conversely, if the receiving module 102 receives only the occlusion position geometry 104, the receiving module 102 may automatically generate the maloccluded dental arch geometry 106. As used herein, "3D mesh" and "3D geometry" are used interchangeably to refer to a 3D digital representation. That is, without loss of generality, it should be understood that there are various types of 3D representations. One type of 3D representation may include a 3D mesh, a 3D point cloud, a voxelized geometry (i.e., a collection of voxels), or other representations described by mathematical formulas. Although the term "mesh" is used frequently throughout this disclosure, it should be understood that the term is, in some implementations, interchangeable with other types of 3D representations.

[0014] Aspects of the various setup prediction embodiments described herein are applicable to the manufacture of clear tray aligners and indirect bonding trays. The various setup prediction embodiments may also be applicable to other products involving final tooth postures. The postures include at least one of position (or location) and rotation (or orientation).

[0015] A 3D mesh is a data structure that describes the geometry (or shape) and structure of an object, such as a tooth, a hardware element, or a patient's gum tissue. A 3D mesh is composed of mesh elements, such as vertices, edges, and faces. In some implementations, the mesh elements may include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features may be calculated for these mesh elements and input into the predictive models of the present disclosure, advantageously improving the ability of these models to make accurate predictions.

[0016] The mesh feature module 108 can use the patient case data received by the receiving module 102 to calculate a number of features associated with the 3D meshes 104 and 106. In general, the method 100 is most interested in optimizing the 3D geometry associated with the patient's teeth, and less interested in optimizing the 3D geometry associated with the patient's gingival tissue. As a result, the mesh feature module 108 is configured to calculate features for each tooth present in the corresponding 3D geometry. According to certain implementations, the mesh feature module 108 can calculate one or more of edge midpoints, edge curvatures, edge normal vectors, edge normalization vectors, edge translation vectors, and other information for each tooth in the 3D meshes 104 and 106. According to certain implementations, the mesh feature module 108 may or may not be utilized. That is, it should be understood that the calculation of any of the edge midpoints, edge curvatures, edge normal vectors, and edge translation vectors for each tooth in the 3D meshes 104 and 106 is optional. One advantage of using the mesh feature module 108 is that systems that utilize the mesh feature module 108 can be trained more quickly and accurately, yet the technique 100 nevertheless performs better than existing techniques that do not use the mesh feature module 108. Another advantage of using 3D meshes over conventional techniques is that errors introduced by mapping 2D results into 3D space and back are not present in the present disclosure. Thus, operating directly in 3D improves the underlying accuracy of the machine learning model and the results that are generated.

[0017] A 3D mesh includes edges, vertices, and faces. Although interrelated, these three types of data are separate. Vertices are points in 3D space that define the boundary of the mesh. These points are described as a point cloud, without further information about how the points are connected to each other, i.e., the edges. An edge is composed of two points and can also be called a line segment. A face is composed of an edge and a vertex. In the case of a triangular mesh, a face includes three vertices that are interconnected to form three adjacent edges. Some meshes may include degenerate elements, such as non-manifold geometry, that must be removed before processing can proceed. Other mesh preprocessing operations are possible. 3D meshes are generally formed using triangles, but in other implementations, they may be formed using quadrilaterals, pentagons, or some other n-gons. In some implementations, the 3D mesh may be converted to one or more voxelized geometries (i.e., including voxels), such as when sparse processing is performed.

[0018] Techniques of the present disclosure operating on 3D meshes may receive as input one or more meshes of teeth (e.g., arranged in one or more dental arches). Each of these meshes typically undergoes pre-processing before being input to a predictive architecture (e.g., including at least one of an encoder, a decoder, a pyramidal encoder-decoder, and a U-Net). This pre-processing includes converting the mesh into a list of mesh elements, such as vertices, edges, faces, or, in the case of low-density processing, voxels. For the selected type(s) of mesh elements (e.g., vertices), feature vectors are generated. In some examples, one feature vector is generated for each vertex of the mesh. Each feature vector may include a combination of spatial and structural features, as specified in the table below.

[0019] [Table 1]

[0020] Consistent with the above description, a voxel may also have features calculated as a collection of other mesh elements (e.g., vertices, edges, and faces) that either intersect with the voxel or, in some implementations, are primarily or completely contained within the voxel. Rotating a mesh does not change structural features, but may change spatial features. Also, as already explained, the term mesh should be considered in a non-limiting sense to include 3D meshes, 3D point clouds, and 3D voxelized geometries. In some implementations, apart from mesh element features, there are alternative ways to describe the geometry of a mesh, such as 3D keypoints and 3D descriptors. Examples of such 3D keypoints and 3D descriptors can be found in "TONIONI A et al., "Learning to detect good 3D keypoints." Int J Comput Vis. 2018 Vol. 126, pages 1-20." In some implementations, the 3D keypoints and 3D descriptors may describe extrema (either minimum or maximum) of the mesh surface.

[0021] The technique 100 also leverages a generative adversarial network ("GAN") to achieve certain aspects of the improvements. In general, a GAN is a machine learning model in which two neural networks "compete" with each other to provide predictions, which are evaluated, and the evaluation of the two models is used to improve the training of each. As shown in FIG. 1, the two neural networks of the GAN are a generator 110 and a classifier 134. The generator 110 receives inputs (e.g., one or more of the occlusal position 104, the maloccluded dental arch 106, and mesh features determined by the mesh feature module 108). The generator 110 uses the received inputs to determine predicted tooth movements 112 for each tooth mesh. In some implementations, the generator 110 may also receive random noise, which may include garbage data or other information that may be used to intentionally attempt to confuse the generator 110. The manner in which the generator 110 determines the predicted tooth movements 112 is described in more detail below in FIGS. 7-12.

[0022] As described herein, tooth movements can be encoded in various ways to specify the position and orientation of the teeth in the setup, specifying one or more tooth transformations to be applied to the 3D representation of the teeth. For example, according to certain implementations, the tooth positions can be Cartesian coordinates of a reference origin position of the teeth defined in some semantic context. The tooth orientation can be expressed as a rotation matrix, a unit quaternion, or another 3D rotation representation such as Euler angles with respect to a reference system (either global or local). The dimensions are real-valued 3D spatial extents, and the gaps can be binary presence indicators or real-valued gap sizes between the teeth, especially in the case where a particular tooth is missing. In some implementations, the tooth rotations can be described by a 3×3 matrix (or by a matrix of other dimensions). In some implementations, the tooth position and rotation information can be combined into the same transformation matrix, for example as a 4×4 matrix that can reflect homogeneous coordinates. In some cases, an affine space transformation matrix can be used to describe the tooth transformations, for example transformations describing the malocclusion posture of the teeth, the intermediate posture of the teeth, and / or the final setup posture of the teeth. Some implementations may use relative coordinates, where the setup transformation is predicted relative to the malocclusion coordinate system (i.e., the malocclusion to setup transformation is predicted directly instead of the setup coordinate system). Other implementations may use absolute coordinates, where the setup coordinate system is predicted directly for each tooth. In relative mode, the transformation may be calculated relative to the centroid of the mesh for each tooth (vs. the global origin), which is referred to as "relative local". Some advantages of using relative local coordinates include eliminating the need for a malocclusion coordinate system (landmarking data), which may not be available for all patient case data sets. Some advantages of using absolute coordinates include simplifying data pre-processing, since the mesh data is originally represented relative to the global origin.

[0023] After the predicted tooth movements 112 are determined by the generator 110, the generator 110 can be trained. For example, in one embodiment, each predicted tooth movement 112 is compared to a corresponding ground truth tooth movement 114 for each tooth mesh. For example, a predicted tooth movement 112 for a canine corresponding to number 27 in the international tooth numbering system is compared to a ground truth tooth movement 114 for the same canine. A ground truth tooth movement is a tooth movement that has been verified as a correct tooth movement for a particular tooth mesh. In some embodiments, the ground truth tooth movements 114 are specified by a human user, such as a dentist or other healthcare provider. In other embodiments, the ground truth tooth movements 114 can be automatically generated based on patient case data or other information provided to a system implementing the method 100.

[0024] The difference between the predicted tooth movements 112 and the ground truth tooth movements 114 can be used to calculate one or more loss values ​​G1 116. For example, G1 116 can represent a regression loss between the predicted tooth movements 112 and the ground truth tooth movements 114. That is, according to one embodiment, the loss G1 116 reflects the rate at which the predicted tooth movements 112 deviate from the ground truth tooth movements 115. That said, the generator loss G1 116 can be an L2 loss, a smooth L1 loss, or some other type of loss. According to a particular embodiment, the L1 loss can be:

[0025]

number

[0026]

number

[0027] Referring again to the predicted tooth movements 112, each predicted tooth movement 112 is represented by one or more transformations to the respective tooth mesh. For example, in one embodiment, each predicted tooth movement 112 is represented by a six-element rotation vector transformation and a three-element translation vector. In this embodiment, the six-element rotation vector represents one or more rotations performed on each tooth to correct its rotation within the 3D geometry, and the three-element translation vector describes each tooth's respective position within the 3D geometry using X, Y, and Z coordinates. In another embodiment, each predicted tooth movement 112 is represented by a seven-element vector, i.e., four elements describing a quaternion rotation and three elements describing a position using X, Y, and Z coordinates. By generating both rotation and translation predictions as part of determining the predictive tooth movements, further advantages can be realized over existing systems. For example, it has been observed that generating both translational and rotational predictions together in the same transform improves accuracy over systems that attempt to combine separately or otherwise predict the translational and rotational predictions separately after either the translational or rotational prediction has been determined.

[0028] Using mesh transformers 118 and 126, technique 100 then transforms the tooth meshes corresponding to 3D meshes 104 and 106 using the predicted tooth movements 112 and ground truth tooth movements 114, respectively. That is, the respective transformations are applied to the 3D geometry to modify the 3D geometry to correspond to the specified movements. For example, with reference to predicted tooth movements 112, in an embodiment represented by a 7-element vector, the tooth mesh is rotated using the specified quaternion rotation of the predicted tooth movements 112 of that tooth mesh in the 3D meshes 104 and 106, and the X, Y, and Z coordinates of the mesh are modified to be equal to the X, Y, and Z coordinates of the predicted tooth movements 112 of that tooth mesh in the 3D meshes 104 and 106. Similarly, ground truth tooth movements 114 can be applied to the 3D meshes 104 and 106 to generate ground tooth movements for each tooth in the 3D meshes 104 and 106.

[0029] These transformations result in modified 3D geometries corresponding to the predicted tooth movement representation 120 and the ground truth tooth movement representation 128. According to certain implementations, both the predicted tooth movement representation 120 and the ground truth tooth movement representation 128 may include occlusion position 3D geometries 124 and 132, respectively, and maloccluded dental arch 3D geometries 122 and 130, respectively. That is, the predicted tooth movement representation 120 may be represented by the occlusion position mesh 104 and one or more maloccluded dental arch meshes 106, which correspond to changes in the occlusion position mesh 124 and the maloccluded dental arch mesh 122, as specified by the predicted tooth movement transform 112. Similarly, the ground truth tooth movement representation 128 may be represented by the occlusion position mesh 132 and one or more maloccluded dental arch meshes 132, as specified by the ground truth tooth movement transform 114.

[0030] Additionally, the predicted tooth movement representation 120 and the ground truth tooth movement representation 128 can be flagged or otherwise annotated to indicate whether the representations correspond to the ground truth transformations. For example, in one implementation, the predicted tooth movement representation 120 is assigned a value of "false" to indicate that it does not correspond to the ground truth tooth movement 114, and the ground truth tooth movement representation 128 is assigned a value of "true."

[0031] According to certain embodiments, the representations 120 and 128 are provided as inputs to the identifier 134. Furthermore, according to certain embodiments, the mesh geometries 104 and 106 are also provided to the identifier 134. However, information regarding the representations 120 and 128 and the meshes 104 and 106 may also be provided to the identifier 134 in other ways. In particular, the identifier 134 need not receive the transformed meshes (i.e., the representations 120 and 128). Instead, the identifier 134 may receive the starting mesh geometries 104 and 106 and the transformations 112 and 114. According to another embodiment, instead of the transformations 112 and 114, the identifier 134 may receive a list of one or more translations to be applied to each element in the meshes 104 and 106. That is, the identifier 134 may receive various representations of data corresponding to the meshes 104 and 106, the transformations 112 and 114, and the representations 120 and 128. In general, the identifier 134 is configured to determine when the input is generated from the predicted tooth movements 112 or when the input is generated from the ground truth tooth movement representation 128. For example, in one implementation, the identifier 134 can output a “false” indication when the identifier 134 determines that the input is generated from the predicted tooth movements 112 and can output a “true” indication when the input is generated from the ground truth tooth movements 114.

[0032] The discriminator 134 may be initially trained in a variety of ways. For example, the discriminator 134 may be configured as an encoder (a particular type of neural network), which may be configured to perform validation in some circumstances as described herein. For example, an initial encoder included in the discriminator 134 may be set with random edge weights. Using backpropagation, the encoder, and therefore the discriminator 134, may be continuously refined by modifying the values ​​of the weights to allow the discriminator 134 to more accurately determine which inputs should be identified as "true" ground truth representations and which inputs should be identified as "false" ground truth representations. In other words, the discriminator 134 may be initially trained, but as the method 100 is performed, the discriminator 134 continues to evolve / train. As with the generator 110, each time the method 100 is performed, the accuracy of the discriminator improves. As will be appreciated by those skilled in the art, improvements to the classifier 134 reach a limit where the accuracy of the classifier 134 does not statistically improve, at which point training of the classifier 134 is considered complete.

[0033] After the identifier 134 generates an output, the method 100 then compares the output of the identifier 134 to the input to determine whether the identifier correctly distinguished between the predicted tooth movement representation 120 and the ground truth tooth movement representation 128. For example, the output of the identifier 134 may be compared to the annotations of the representations. If the output and the annotations match, the identifier 134 correctly predicted the type of input that the identifier 134 received. Conversely, if the output and the annotations do not match, the identifier 134 did not correctly predict the type of input that the identifier 134 received. In some implementations, similar to the generator 110, the identifier 134 may also receive random noise to intentionally confuse the identifier 134.

[0034] Additionally, according to certain implementations, the classifier 134 may generate additional values ​​that may be used to train aspects of a system implementing the method 100. In one example, the classifier 134 may generate a classifier loss value 136 that reflects how accurately the classifier 134 determined whether the input corresponds to the predicted tooth movement representation 120 and / or the ground truth tooth movement representation 128. According to certain implementations, the classifier loss 136 is larger when the classifier 134 is less accurate in its predictions and is smaller when the classifier 134 is more accurate in its predictions. In another example, the classifier 134 may generate a generator loss value G2 138. According to certain implementations, the generator loss value G2 138 is not directly inverse to the classifier loss 136, but generally exhibits an inverse relationship to the classifier loss 136. That is, when the discriminator loss 136 is large, the generator loss G2 138 is small, and when the discriminator loss 136 is small, the generator loss G2 138 is large. In some implementations, the discriminator loss 136 may be determined using a binary cross-entropy loss function calculated for both the "true" and "false" models. In some implementations, the generator loss may be composed of two losses: 1) the first loss is the generator loss G2 138 as determined by the discriminator (hence, binary cross-entropy may be used), and 2) the second loss may be implemented by, for example, the l1 norm or mean square error that measures the difference between the desired output and the actual output of the generator 110, as specified by the generator loss G1 116.

[0035] In other words, as shown in FIG. 1, the generator loss G2 138 can be added to the generator loss G1 116 using an addition operation 140. The sum of the generator losses G1 116 and G2 138 can then be provided to the generator 110 for training the generator 110. It should be understood, however, that the calculation of the generator loss G1 116 is not necessary for training the GAN. In some implementations, it may be possible to train either the generator 110 or the discriminator 134 using only the combination of the generator loss G2 138 and the discriminator loss 136. However, as with other optional aspects of the present disclosure, the use of the generator loss G1 116 can be utilized to train the generator 134 more quickly and generate more accurate predictions. Additional aspects of the method 100 will become apparent as part of the description of the subsequent figures.

[0036] FIG. 2 is an exemplary visualization of a workflow 200 performed using the method 100 shown in FIG. 1. As should be understood by the description of FIG. 1, in a first step 202 of the workflow 200, initial 3D position and orientation data (e.g., one or more of the occlusal position 104, the maloccluded dental arch 106) is received in the form of one or more 3D geometric shapes, and the method 100 calculates final position and orientation information in step 206 of the workflow 200. Additional steps 204a-204n are also shown in the workflow 200. However, in general, these steps of the workflow are intentionally omitted when determining the final setup as described with reference to FIG. 1. Instead, steps 204a-204n can be used to generate intermediate stages, which are described in more detail below with reference to FIGS. 4-6.

[0037] 3 is a different view of the method 100 shown in FIG. 1 (referred to herein as method 300). According to certain implementations, method 300 first accesses patient data, such as patient data received by module 102. In step 302, the processor executing method 300 may optionally generate random noise. The processor then provides the patient case data and any random noise to a generator 110. As described above with reference to FIG. 1, the processor executes instructions that cause the trained generator 110 to generate predicted tooth movements 112 that can be used to determine a predicted tooth movement representation 120.

[0038] The system performing the method 300 may also access one or more ground truth transforms in step 304 and select one or more sample ground truth transforms 114 that correspond to the selected patient case data received by the receiving module 102. As described above with reference to FIG. 1, the one or more sample ground truth transforms 114 may be used to generate the ground truth tooth movement representations 128. The system performing the method 300 may then provide any of the patient case data received by the receiving module 102, the predicted tooth movement representations 120, and the ground truth tooth movement representations 128 to the identifier 134.

[0039] Next, as described with reference to method 100, in step 306, the identifier 134 determines whether the input corresponds to a ground truth transformation by providing a probability that the input is a true or false ground truth transformation. In some implementations, the probability returned by the identifier 134 may be in the range of 0 to 1. That is, the identifier 134 may provide a value closer to 0 to indicate a low probability that the input is true (i.e., corresponds to the predicted tooth movement 112) or a value closer to 1 to indicate a high probability that the input is true (i.e., corresponds to the ground truth tooth movement 114).

[0040] As discussed above with reference to FIG. 1, the output of the classifier 134 may be used to train both the classifier 134 and the generator 110 .

[0041] 4 is an exemplary method 400 that can be used to train a machine learning model used to determine the intermediate staging of a CTA. As shown, aspects of the method 400 are similar to the method 100. For example, the method 400 utilizes a receiving module 402 that receives patient case data. The receiving module 402 operates similarly to the receiving module 102, e.g., the receiving module 402 can receive data corresponding to the bite positioning geometry and the maloccluded dental arch geometry 106. The receiving module differs from the receiving module 102 in that the receiving module 402 is also configured to receive an end tooth transformation 404 corresponding to the final set-up. According to certain implementations, the end tooth transformation 404 can be predefined or provided as a result of performing the method 100.

[0042] Method 400 also employs a mesh feature module 108 that can use the patient case data received by the receiving module 102 and calculate several features related to the 3D meshes 104 and 106, as described above with reference to FIG. 1. Similar to method 100, method 100 also leverages a generative adversarial network ("GAN") to achieve certain aspects of the improvements as described throughout this disclosure. However, method 400 employs a generator 411 and a classifier 435 that are used differently than the generator 110 and classifier 134 as described above with reference to FIG. 1. For example, instead of the generator 411 receiving inputs (e.g., one or more of the occlusal position 104, the maloccluded dental arch 106, and the mesh features determined by the mesh feature module 108) and generating predicted tooth movements for the final setup, the generator 411 uses the received inputs to determine predicted intermediate tooth movements 406 for each tooth mesh. According to some implementations, the predicted intermediate stage tooth movement 406 can be used to determine one or more of the following values: 1) which direction the tooth is moving, 2) how far the tooth is located toward the final state for the current stage, and 3) how the tooth is rotating. However, other aspects of the generator 411 are the same. For example, in some implementations, the generator 411 can also receive random noise, which may include garbage data or other information that may be used to intentionally attempt to confuse the generator 411. As a result, it should be understood that in many aspects of the disclosure described herein, the generator 110 and the generator 411 can be used interchangeably. Similarly, the identifier 134 and the identifier 435 can be used interchangeably.

[0043] After the predicted intermediate tooth movements 406 are determined by the generator 411, the generator 411 can be trained. For example, in one implementation, each predicted intermediate tooth movement 406 is compared to a corresponding ground truth intermediate tooth movement 408 for each tooth mesh. The comparisons performed as part of the method 400 are the same as the method 100 described with reference to FIG.

[0044] Similarly, the difference between the predicted intermediate tooth movements 406 and the ground truth intermediate tooth movements 408 may be used to calculate one or more loss values ​​G1 116, as described above with respect to method 100. Similarly, the loss values ​​G1 116 may be provided to the generator 411 to further train the generator 411, as described in connection with method 100, for example, by modifying one or more weights in a neural network of the generator 411 to train the underlying model and improve the model's ability to generate predicted intermediate tooth movements 406 that reflect or substantially reflect the ground truth intermediate tooth movements 408.

[0045] Referring again to the predicted intermediate tooth movements 406, each predicted intermediate tooth movement 406 is represented by one or more transformations to the respective tooth mesh. For example, in one embodiment, each predicted intermediate tooth movement 406 is represented by a six-element rotation vector transformation and a three-element translation vector. In this embodiment, the six-element rotation vector represents one or more rotations performed on each tooth to correct its rotation within the 3D geometry, and the three-element translation vector describes the respective position of each tooth within the 3D geometry using X, Y, and Z coordinates. In another embodiment, each predicted intermediate tooth movement 406 is represented by a seven-element vector, i.e., four elements to describe the quaternion rotation and three elements to describe the position using X, Y, and Z coordinates.

[0046] Using mesh transformers 118 and 126, technique 400 then transforms the tooth meshes corresponding to 3D meshes 104 and 106 using predicted intermediate tooth movement 406 and ground truth intermediate tooth movement 408, respectively. That is, each transformation is applied to the 3D geometry to modify the 3D geometry to correspond to the specified movement. For example, with reference to predicted intermediate tooth movement 406, in an embodiment represented by a 7-element vector, the tooth mesh is rotated using the specified quaternion rotation of predicted intermediate tooth movement 406 relative to its tooth mesh in 3D meshes 104 and 106, and the X, Y, and Z coordinates of the mesh are modified to be equal to the X, Y, and Z coordinates of predicted intermediate tooth movement 406 relative to its tooth mesh in 3D meshes 104 and 106. Similarly, the ground truth intermediate tooth movements 114 can be applied to the 3D meshes 104 and 106 to generate ground tooth movements for each tooth in the 3D meshes 104 and 106.

[0047] These transformations result in modified 3D geometries corresponding to the predicted intermediate tooth movement representation 410 and the ground truth intermediate tooth movement representation 418. According to certain implementations, both the predicted intermediate tooth movement representation 410 and the ground truth intermediate tooth movement representation 418 can include occlusion position 3D geometries 414 and 422, respectively, and maloccluded dental arch 3D geometries 412 and 420, respectively. That is, the predicted intermediate tooth movement representation 410 can be represented by the occlusion position mesh 104 and one or more maloccluded dental arch meshes 106, which correspond to the changes in the occlusion position mesh 414 and the maloccluded dental arch mesh 412, as specified by the predicted intermediate tooth movement transform 406. Similarly, the ground truth intermediate tooth movement representation 418 can be represented by the occlusion position mesh 422 and one or more maloccluded dental arch meshes 420, as specified by the ground truth intermediate tooth movement transform 408.

[0048] Additionally, the predicted intermediate tooth movement representation 410 and the ground truth intermediate tooth movement representation 418 can be flagged or otherwise annotated to indicate whether the representations correspond to the ground truth transformations. For example, in one implementation, the predicted intermediate tooth movement representation 410 is assigned a value of "false" to indicate that it does not correspond to the ground truth intermediate tooth movement 408, and the ground truth intermediate tooth movement representation 418 is assigned a value of "true."

[0049] The representations 410 and 418 are provided as inputs to the identifier 134. Additionally, according to certain implementations, the mesh geometries 104 and 106 are also provided to the identifier 435. However, information regarding the representations 410 and 418 and the meshes 104 and 106 may be provided to the identifier 435 in other ways. In particular, the identifier 435 need not receive the transformed meshes (i.e., the representations 410 and 418). Instead, the identifier 435 may receive the starting mesh geometries 104 and 106 and the transformations 406 and 408. According to another implementation, instead of the transformations 406 and 408, the identifier 435 may receive a list of one or more translations to be applied to each element in the meshes 104 and 106. That is, similar to the identifier 134, the identifier 435 may receive various representations of data corresponding to the meshes 104 and 106, the transformations 406 and 408, and the representations 410 and 418. According to method 400, the identifier 435 is configured to determine when the input is generated from predicted intermediate tooth movements 406 or when the input is generated from ground truth intermediate tooth movements 408. For example, in one embodiment, the identifier 435 can output a “false” indication when the identifier 435 determines that the input is generated from predicted intermediate tooth movements 408 and can output a “true” indication when the input is generated from ground truth intermediate tooth movements 406.

[0050] The identifier 435 is otherwise substantially similar to the identifier 435 described with reference to FIG. 1. For example, after the identifier 435 generates an output, the method 400 then compares the output of the identifier 435 to the input to determine whether the identifier correctly distinguished between the predicted tooth movement representation 410 and the ground truth tooth movement representation 418. For example, the output of the identifier 435 can be compared to the annotations of the representations. If the output and the annotations match, the identifier 435 correctly predicted the type of input that the identifier 435 received. Conversely, if the output and the annotations do not match, the identifier 435 did not correctly predict the type of input that the identifier 435 received. In some implementations, like the generator 411, the identifier 435 may also receive random noise to intentionally confuse the identifier 435.

[0051] Additionally, according to certain implementations, the classifier 435 may generate additional values ​​that may be used to train aspects of a system implementing the method 400. In one example, the classifier 435 may generate a classifier loss value 136 that reflects how accurately the classifier 435 determined whether the input corresponds to the predicted intermediate tooth movement representation 410 and / or the ground truth intermediate tooth movement representation 418. According to certain implementations, the classifier loss 136 is larger when the classifier 435 is less accurate in its prediction and is smaller when the classifier 435 is more accurate in its prediction. In another example, the classifier 435 may generate a generator loss value G2 138. According to certain implementations, the generator loss value G2 138 is not directly inverse to the classifier loss 136, but generally exhibits an inverse relationship to the classifier loss 136. That is, when the discriminator loss 136 is large, the generator loss G2 138 is small, and when the discriminator loss 136 is small, the generator loss G2 138 is large. In some implementations, the discriminator loss 136 may be determined using a binary cross-entropy loss function calculated for both the "true" and "false" models. In some implementations, the generator loss may be composed of two losses: 1) the first loss is the generator loss G2 138 as determined by the discriminator (hence, binary cross-entropy may be used), and 2) the second loss may be implemented by, for example, the l1 norm or mean square error that measures the difference between the desired output and the actual output of the generator 110, as specified by the generator loss G1 116.

[0052] 4, the generator loss G2 138 may be added to the generator loss G1 116 using an addition operation 140. The sum of the generator losses G1 116 and G2 138 may be provided to the generator 411 for the purpose of training the generator 411. Additional aspects of the method 400 will become apparent as part of the description of subsequent figures.

[0053] Figure 5 is an example visualization of a workflow 500 performed using the method 400 shown in Figure 4. As should be understood by the description of Figure 4, in a first step 502 of the workflow 500, both initial and final 3D position and orientation data (e.g., one or more of the occlusal position 104, the maloccluded dental arch 106) are received in the form of one or more 3D geometric shapes, and the method 400 calculates intermediate position and orientation information in steps 204a-204n in the workflow to generate n intermediate stages for the CTA.

[0054] 6 is a different view of the method 400 shown in FIG. 4 (referred to herein as method 600). According to certain implementations, the method 300 first accesses patient data, such as patient data 402. In step 302, the processor executing the method 600 may optionally generate random noise. The processor then provides the patient case data, the final tooth setup 404, and the optional random noise to a generator 411. As described above with reference to FIG. 4, the processor executes instructions that cause the trained generator 411 to generate predicted intermediate tooth movements 406 that can be used to determine a predicted intermediate tooth movement representation 410.

[0055] The system performing method 600 may also access one or more ground truth transforms in step 304 and select one or more sample ground truth intermediate transforms 408 that correspond to the selected patient case data received by module 102. As described above with reference to Figure 4, the one or more sample ground truth transforms 408 may be used to generate the ground truth tooth movement representations 418. The system performing method 600 may then provide any of the patient case data received by receiving module 402, the predicted tooth movement representations 410, and the ground truth tooth movement representations 418 to the identifier 435.

[0056] Next, as described with reference to method 400, in step 306, the identifier 435 determines whether the input corresponds to a ground truth transformation by providing a probability that the input is a true ground truth transformation or a false ground truth transformation. In some implementations, the probability returned by the identifier 435 may be in the range of 0 to 1. That is, the identifier 435 may provide a value closer to 0 to indicate a low probability that the input is true (i.e., corresponds to the predicted intermediate tooth movement 406), or a value closer to 1 to indicate a high probability that the input is true (i.e., corresponds to the ground truth intermediate tooth movement 408).

[0057] As described above with reference to FIG. 4, the output of the classifier 435 can be used to train both the classifier 435 and the generator 411.

[0058] 7 is an expanded view 700 of the method 100 shown in FIG. 1, focusing on aspects of the method 100 that use geometric deep learning. According to certain implementations, geometric information about the 3D meshes 104 and 106 may be identified or otherwise determined by a mesh converter 702. Although not included in FIG. 1, many implementations of the method 100 are contemplated to utilize a mesh converter 702, as doing so provides various benefits to the method 100 related to both improving the predictive quality of the output of the generator 110 and improving training for a GAN based on the output of the generator 110, as described above.

[0059] As shown in Fig. 7, the receiving module 102 can provide the 3D bite position geometry 104 and the 3D maloccluded dental arch geometry 106 to the mesh converter 702. Generally, according to a well-established definition, the 3D geometries 104 and 106 are defined by a set of vertices, where each pair of vertices specifies an edge of a 3D polygon, and the set of edges can specify one or more faces (or surfaces) of the 3D geometry. Thus, according to a particular implementation, this allows the 3D mesh converter 702 to decompose the 3D meshes 104 and 106 into their respective component parts.

[0060] Stated another way, the 3D mesh converter 702 may extract or generate various geometric features from the 3D meshes 104 and 106, and then these converted mesh data are used as input data to the generator 110. For example, the 3D mesh converter 702 may generate one or more of: one or more mesh edge lists 704, one or more mesh face lists 706, and one or more mesh vertex lists 708.

[0061] By providing this additional information to the generator 110, several advantages can be realized. For example, providing this information to the generator 110 allows the generator 110 to generate more accurate predicted tooth movements 112. This allows the training of a system implementing the method 100 to be improved, since both the training of the generator 110 and the training of the classifier 134 are based at least in part on the quality of the predicted tooth movements 112. In short, implementing the mesh converter 702 as part of the method 100 can reduce the number of training epochs that the neural networks configured in the generator 110 and the classifier 134 must undergo, while also improving accuracy. In other words, by using the mesh converter 702 as part of the method 100, a system performing the method 100 can conserve the computational resources involved in the training process, while improving the model trained as described.

[0062] Furthermore, the description of the zoom 700 should not be considered limiting. For example, the zoom 700 is shown and described in relation to the method 100 presented in FIG. 1, but it should be understood that the zoom 700 can also be used as part of the method 400 shown in FIG. 4. For example, the receiving module 102 of FIG. 7 can be replaced with the receiving module 402 shown in FIG. 4. This allows, for example, the method 400 to achieve the same improved computing resource utilization and model accuracy as described above in relation to the method 100. The only difference is that the method 400 generates predictions of intermediate stages, whereas the method 100 generates predictions of the final setup. In other words, replacing the module 102 with the module 402 in FIG. 7 also causes the generator 411 to be replaced instead of the generator 110, which generates the predicted intermediate tooth movements 406 instead of the predicted tooth movements 112. However, despite these configuration changes, the method 400 can still leverage the improvements from the geometric deep learning described above.

[0063] 8-12 show certain aspects of the generator 110 according to certain implementations. In these illustrated implementations, the generator 110 may be configured as at least one of a first 3D encoder, a 3D U-Net encoder-decoder, or a 3D pyramid encoder-decoder followed by a second 3D encoder (optionally replaced with a multi-layer perceptron (MLP)). The generator may be implemented as one or more neural networks, and therefore may include activation functions. The activation functions determine whether or not neurons in the neural network fire (e.g., send output to the next layer). Some activation functions may include binary step functions and linear activation functions. Other activation functions give the network nonlinear behavior and include sigmoid / logistic activation functions, Tanh (hyperbolic tangent) function, rectified linear unit (ReLU), leaky ReLU function, parametric ReLU function, exponential linear unit (ELU), softmax function, swish function, Gaussian error linear unit (GELU), and scaled exponential linear unit (SELU). Linear activation functions may be well suited for some regression applications at the output layer (among other applications). Sigmoid / logistic activation functions may be well suited for some binary classification applications at the output layer (among other applications). Softmax activation functions may be well suited for some multi-class classification applications at the output layer (among other applications). Sigmoid activation functions may be well suited for some multi-label classification applications at the output layer (among other applications). The ReLU activation function may be well suited for some convolutional neural network (CNN) applications (among other applications) in the hidden layer.Tanh and / or sigmoid activation functions may be well suited in some recurrent neural network (RNN) applications (among other applications), e.g., in hidden layers.

[0064] There are several optimization algorithms that may be used in training the neural network of the present disclosure, including gradient descent (which uses first derivatives to determine training gradients and is commonly used in training neural networks), Newton's method (which may utilize second derivatives in loss calculations to find better training directions than gradient descent, but may require calculations involving Hessian matrices), and conjugate gradient methods (which may have faster convergence than gradient descent, but do not require Hessian matrix calculations that may be required by Newton's method). A backpropagation algorithm is used to transfer the results of the loss calculations back to the network so that the network weights can be adjusted and learning can proceed.

[0065] In some implementations, the neural networks of the present disclosure can be adapted to operate on 3D point cloud data (alternatively, on 3D meshes or 3D voxelized geometry). Numerous neural network implementations may be applied to process 3D representations and to train predictive and / or generative models for oral care applications, including PointNet, PointNet++, SO-Net, spherical convolution, Monte Carlo convolution and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN and DSG-Net.

[0066] As mentioned above, each tooth mesh 104 and 106 includes several mesh elements such as edges, faces and vertices. In some implementations, the edges included in the meshes 104 and 106 may be more useful in generating accurate predictions, but operations can also be performed on faces and vertices. When edges are used to make predictions, feature vectors are calculated for each edge mesh element. The feature vectors may include various 3D geometric representations, such as the 3D coordinates of the vertices, or the curvature and midpoint of the edges. Other features are possible. According to certain implementations, the output of the encoder-decoder structure maintains the same resolution as the input (i.e., the input and output have the same number of elements). In general, the encoder-decoder structure (either a U-Net architecture or a pyramid architecture) serves to extract high-dimensional features from the tooth meshes, for example, by converting one or more tooth meshes into a representation (which may include either or both local and global information about the tooth mesh) that a second encoder can use to generate a tooth transformation for either the final setup or an intermediate stage.

[0067] Further, although not explicitly shown, an additional embodiment of FIGS. 8-12 replaces the U-Net Encoder-Decoder 806 or the Pyramid Encoder-Decoder 1004 with an encoder (such as the Encoder 814 shown in FIGS. 8 and 10). According to this embodiment, the generator 110 operates on a lower resolution mesh. That is, a first encoder (not shown, but replacing either the U-Net Encoder-Decoder 806 or the Pyramid Encoder-Decoder 1004 in FIGS. 8 and 10, respectively) coarsens the resolution of the input geometry received by the Encoder 814. This provides certain advantages, including reducing the memory consumption of the generator 110. Nonetheless, to achieve these improvements, additional processing may be performed, including, but not limited to, maintaining a list of tooth labels for each element (i.e., for each edge, face, or vertex in the 3D geometry).

[0068] Figure 8 illustrates an example workflow 800 for the generator 110 shown in either Figure 1 or Figure 4 and using the U-Net architecture. In step 802 of the workflow, the input is processed to modify the input into a data format that can be extracted as edge elements. In general, step 802 takes mesh data having a feature vector of a first size and provides it to a machine learning model, which generates a feature vector of a second size that corresponds to the mesh data.

[0069] As shown in FIG. 8, the mesh data in step 804 may be any combination of tooth meshes 104 and 106. These meshes 104 and 106 may include thousands or tens of thousands of mesh elements, such as edges, faces, vertices, and / or voxels. For one or more mesh elements, one or more mesh feature vectors may be calculated. The mesh elements and any associated mesh feature vectors may be input to a generator. Typically, each of the mesh elements in the meshes 104 and 106 may be described by a feature vector having a variable size depending on the feature. For example, when describing a point, the mesh element may be described by a three-channel vector, where the three channels describe the X, Y, and Z coordinates of the position in three-dimensional space. When describing an edge, the mesh element may be represented by two integers, one for each vertex that defines the edge, where each integer is an index into an array of vertices that make up the mesh. When describing a face, the mesh element may be represented by three integers, one for each vertex that defines the face. When describing voxels, mesh elements may be represented by cubic volumes of space. In some implementations, a list of vertices of the 3D mesh may be fed into the open source MinkowskiEngine toolkit, which may convert the vertices to voxels for sparce processing. In some implementations, a mesh feature vector may be calculated for one or more mesh elements. A mesh feature is a quantity that describes an attribute (e.g., geometric and / or structural attribute) of the mesh at the location of a particular mesh element. In some implementations, only the mesh elements are input to the generator. In some implementations, each mesh element is accompanied by an associated mesh element feature vector, such as the feature vector described in connection with Table 1 above. Additionally, the feature vector may include additional information such as mesh curvature information (an additional three channels) and edge normal vector information (an additional three channels), for a total of nine channels. Still other feature vector configurations are possible, with corresponding numbers of channels.The 3D mesh described thus far is just one of several types of 3D representations that can be used to describe teeth. Other forms of 3D representations include 3D point clouds and voxelized representations.

[0070] The U-Net architecture used in step 806 works by first reducing the resolution of the input tooth mesh 804, and then restoring the simplified tooth mesh (i.e., low-resolution mesh) to the original resolution. This operation allows information about adjacent teeth (or about the entire dental arch) to be captured and integrated into the feature computation. In step 808 of workflow 800, a feature vector is computed for each element (e.g., each edge) in the high-dimensional space (e.g., 128 channels).

[0071] In step 810 of the workflow, for each tooth, an element with high dimensional features is extracted from the output of the encoder-decoder structure. This generates n tooth edges 812a-812n that are provided to another encoder in step 814 of the workflow 800. The encoder is trained by backpropagation with the high dimensional features of a given tooth to predict the tooth movement. Backpropagation is a well-established technique for training neural networks and is known to those skilled in the art.

[0072] The output of the encoder in step 814 of workflow 800 is the predicted tooth movements 112 that are applied to the teeth to move them to desired positions (either for the final setup described with reference to FIGS. 1-3 or for the intermediate stages described with reference to FIGS. 4-6). The encoder in step 814 of workflow 800 is trained via backpropagation to output transformations for the teeth, regardless of the identity of the teeth. In these embodiments shown in the figures, the same encoder is trained to process each of the teeth present in each of the two dental arch shapes 106. In other embodiments, the encoder may be trained to provide a specific tooth or a specific set of teeth. This latter embodiment would be reflected in workflow 800 as multiple encoders in step 814, rather than just the one shown.

[0073] FIG. 9 illustrates an exemplary U-Net architecture 900 shown in FIG. 8. In general, the U-Net architecture uses several pooling layers, such as pooling layers 904a and 904b. The pooling layers associated with the convolution layers, such as convolution layers 902a, 902b, 908a, 908b, and 910, downsample or reduce the mesh input. For example, downsampling of information in 3D space can take a 3×3×3 set of information and combine it into a single 1×1×1 representation. In the context of 3D mesh information, for example, the four neighbors of a given edge are combined into a single edge at the next resolution level. The mesh resolution (mesh surface area) after downsampling is reduced by a factor of four.

[0074] According to certain implementations, the convolution layers 902a, 902b, 908b, 908b, and 910 can use edge data to perform mesh convolution. The use of edge information ensures that the model is not sensitive to different input orders of the 3D elements. In addition to or in addition to using edge data, the convolution layers 902a, 902b, 908b, 908b, and 910 can use vertex data to perform mesh convolution. The use of vertex information is advantageous in that there are typically fewer vertices than edges or faces, and thus vertex-oriented processing can lead to lower processing overhead and lower computational costs.

[0075] In addition to or in addition to using edge or vertex data, the convolution layers 902a, 902b, 908b, 908b, and 910 can use face data to perform mesh convolution. Furthermore, in addition to or in addition to using edge, vertex, or face data, the convolution layers 902a, 902b, 908b, 908b, and 910 can use voxel data to perform mesh convolution. The use of voxel information is advantageous in that, depending on the selected granularity, there may be significantly fewer voxels to process compared to the vertices, edges, or faces in a mesh. Low-density processing (using voxels) may lead to lower processing overhead and lower computational costs (especially in terms of computer memory or RAM usage).

[0076] As described above with reference to FIG. 8, the purpose of the U-Net architecture 900 is to compute a high-dimensional feature vector for an input mesh (which may include either or both local and global information for one or more tooth meshes). For example, according to certain implementations, the U-Net architecture 900 computes a feature vector for each mesh element (e.g., a 128-element feature vector for each edge). This vector exists in a high-dimensional space that may represent the local geometry of the edge within the local tooth context, and may also represent the global geometry of the two dental arches. The high-dimensional features of the elements within each tooth are used by the encoder to predict tooth movement. The accuracy of the tooth movement prediction is aided by this combination of local and global information. The combination of local and global information allows the U-Net architecture 900 to take into account geometric constraints. For example, during the course of CTA treatment, it is undesirable for teeth to collide in 3D space. The combination of local and global information allows the U-Net architecture 900 to generate transformations that reduce or eliminate the occurrence of collisions, thus resulting in higher accuracy compared to the prior art. Stated another way, one advantage of using mesh element features to train a machine learning model (such as U-Net architecture 900) over conventional approaches is that the mesh element features provide additional information about at least one of the geometry and structure of the tooth mesh, which improves the resulting representation(s) generated from the trained U-Net architecture.

[0077] The U-Net architecture 900 involves pooling and unpooling operations that aid in the process of extracting mesh element neighborhood information. Each successive pooling layer helps the model learn adjacent geometric information by decreasing the resolution relative to the previous layer. Each successive unpooling layer helps the model extend this summarized neighborhood information back to a higher resolution. The sequence of unpooling layers followed by a sequence of pooling layers enables efficient and accurate training of the U-Net, allowing it to output features for each element that contain both local and global geometric information.

[0078] Although FIG. 9 is shown with a total of nine layers, it should be understood that the U-Net architecture 900 may be composed of any number of convolutional layers, any number of pooling layers, and any number of unpooling layers to achieve desired results.

[0079] FIG. 10 illustrates an example workflow 1000 for the generator 110 shown in either FIG. 1 or FIG. 4. Workflow 1000 is similar to workflow 800 illustrated and described in FIG. 8. For example, both workflows 800 and 1000 generate predicted tooth movements 112. Furthermore, according to certain embodiments, the encoder in step 814 of workflow 1000 may be replaced with multiple encoders, as described above with reference to FIG. 8. Thus, for the sake of brevity, each element of workflow 1000 will not be described, and instead, only the differences between workflow 800 and workflow 1000 will be noted.

[0080] Specifically, in step 1002, a pyramid encoder-decoder is used in step 1004 instead of the U-Net architecture used in step 806. As expected, the pyramid encoder-decoder used in step 1004 operates differently than the U-Net architecture used in step 806. For example, the input elements of each tooth mesh (e.g., edge elements identified in step 804 of workflow 1000) are passed through an encoder structure to generate multiple layers of features in a pyramid. Each successive layer of the encoder has fewer elements, but the elements reveal higher dimensional information about the tooth mesh in the feature vector. In other words, each successive layer in the pyramid architecture is configured to reveal higher dimensional information about the tooth mesh. Furthermore, an interpolation step is performed at each layer to convert the features from the series of lower resolutions back to the input resolution of the original tooth mesh. The interpolated features from the multiple layers are concatenated and further processed to become high dimensional features for each mesh element as the output of the pyramid encoder-decoder.

[0081] The output of the pyramid encoder architecture generated in step 1004 is used by the remainder of the workflow 1000 in a manner similar to how the output of the U-Net architecture generated in step 806 is used by the remainder of the workflow 800. Importantly, similar to workflow 800, the end result of workflow 1000 is a predicted tooth movement 112. This allows the above-described techniques 100 and 400 to be agnostic to the implementation type of generators 110 and 411. This flexibility provides various advantages, including but not limited to the ability to explore the accuracy of differently trained U-Net and pyramid architecture based generators without having to reconfigure the entire system. This may, for example, allow a system implementing the techniques 100 and 400 to use one generator 110 or 411 trained on one type of patient case data received by module 102 or 402 and another generator 110 or 411 trained on a different type of patient case data without disruption or performance degradation of the system. In some implementations, the generators 110 or 411 can be trained to generate tooth movements for all types of teeth (e.g., incisors, canines, bicuspids, molars, etc.). In other implementations, one generator 110 or 411 may be trained only for anterior teeth (e.g., incisors and canines) and another generator 110 or 411 may be trained only for molar teeth (e.g., bicuspids and molars). The advantage of this latter approach is that accuracy is improved, since each of the two generators 110 or 411 is tuned to generate transformations for specific teeth with their own specific geometry.

[0082] The U-Net structure in step 806 involves high computer memory usage due to the fine-grained representation of neighborhood geometric information learned for each mesh element. The advantage of the U-Net structure in step 806 is a very accurate prediction of tooth movement commensurate with the fine-grained data used for the calculation. The pyramid encoder structure in step 1004 can be used as an alternative when lower memory requirements exist (such as when the computing environment cannot handle the fine-grained data involved in using the U-Net structure in step 806). Further memory savings can be realized by implementing the alternative structures described above that replace the U-Net architecture 900 or pyramid architecture in step 1004 with an encoder.

[0083] FIG. 11 illustrates an exemplary pyramid encoder-decoder 1100 shown in FIG. 10. As previously described, the pyramid encoder-decoder has successive layers 1104a-1104n. The manner in which the pyramid architecture 1100 is used during step 1004 of the workflow 1000 was also previously described. In particular, in a first step 1102, the pyramid architecture 1100 generates successive layers of lower mesh resolution. This step reveals high dimensional information about the tooth mesh, which is contained in the feature vector that the pyramid architecture 1100 generates as output.

[0084] Next, in step 1106, the pyramid architecture 1100 uses interpolation to increase the resolution of each successive layer of mesh elements (e.g., edges, faces, or vertices). For example, an encoder included in the pyramid architecture 1100 includes successive layers 1110a-1110n, whereby each mesh element is downsampled to extract information about the mesh at each of a succession of resolutions. Each successive layer attributes additional feature channels to each of the included mesh elements. Each successive layer includes information about a larger portion of the tooth to which the mesh element belongs, or even about adjacent teeth, or about an entire dental arch, or even about two entire dental arches. Interpolation is performed to facilitate the process of concatenating global mesh information from the lower resolution layers with local mesh information from the higher resolution layers. As shown, this interpolation results in layers 1110a-1110n. Finally, in step 1108, the pyramid architecture 1100 concatenates elements of the input mesh with elements of successive layers to generate the output of the architecture 1100. According to a particular implementation, the final concentration vectors are all of the same resolution (ie, the same number of elements).

[0085] Figure 12 illustrates an example encoder 814 as shown in Figures 8 and 10. As illustrated, the encoder 814 is a collection of convolutional layers 1202a, 1202b, and 1206, and pooling layers 1204a and 1204b. Although Figure 12 is shown with a total of five layers, it should be understood that the encoder 814 may be composed of any number of layers.

[0086] 13 illustrates an example processing unit 1302 that operates according to the techniques of this disclosure. The processing unit 1302 provides a hardware environment for training one or more of the neural networks described above. For example, the processing unit 1302 may execute the techniques 100 and / or 400 to train the neural networks 110 and 134.

[0087] In this example, the processing unit includes processing circuitry, which may include one or more processors 1304 and memory 1306, which in some examples provide a computer platform for executing an operating system 1316, which may be, for example, a real-time multitasking operating system or other type of operating system. The operating system 1316, in turn, provides a multitasking operating environment for executing one or more software components, such as applications 1318. The processor 1304 is coupled to one or more I / O interfaces 1314 that provide an I / O interface for communicating with devices such as keyboards, controllers, display devices, image capture devices, other computing systems, etc. Further, the one or more I / O interfaces 1314 may include one or more wired or wireless network interface controllers (NICs) for communicating with a network. Additionally, the processor 1304 may be coupled to an electronic display 1308.

[0088] In this example, the processing unit includes processing circuitry, which may include one or more processors 1304 and memory 1306, which in some examples provide a computer platform for running an operating system 1316, which may be, for example, a real-time multitasking operating system or other type of operating system. The operating system 1316, in turn, provides a multitasking operating environment for running one or more software components, such as applications 1318. The processor 1304 is coupled to one or more I / O interfaces 1314 that provide an I / O interface for communicating with devices such as keyboards, controllers, display devices, image capture devices, other computing systems, etc. Further, the one or more I / O interfaces 1314 may include one or more wired or wireless network interface controllers (NICs) for communicating with a network. Additionally, the processor 1304 may be coupled to an electronic display 1308.

[0089] In some examples, the processor 1304 and the memory 1306 may be separate, individual components. In other examples, the memory 1306 may be an on-chip memory co-located with the processor 1304 in a single integrated circuit. There may be multiple instances of a processing circuit (e.g., multiple processors 1304 and / or memory 1306) in the processing unit 1302 to facilitate running applications in parallel. The multiple instances may be of the same type, e.g., a multi-processor system or a multi-core processor. The multiple instances may be of different types, e.g., a multi-core processor with multiple associated graphics processor units (GPUs). In some examples, the processor 1304 may be implemented as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or equivalent discrete or integrated logic circuits, or a combination of any of the foregoing devices or circuits.

[0090] The architecture of the processing unit 1302 shown in FIG. 13 is shown for illustrative purposes only. The processing unit 1302 should not be limited to the illustrated exemplary architecture. In other examples, the processing unit 1302 may be configured in various ways. The processing unit 1302 may be implemented as any suitable computing system (e.g., at least one server computer, workstation, mainframe, appliance, cloud computing system, and / or other computing system) that may be capable of performing the operations and / or functions described in accordance with at least one aspect of the present disclosure. By way of example, the processing unit 1302 may represent (or be implemented through) at least one virtualized compute instance (e.g., virtual machine or container) of a data center, cloud computing system, server farm, and / or server cluster. In some examples, the processing unit 1302 includes at least one computing device, each computing device having a memory 1306 and at least one processor 1304.

[0091] The storage unit 1334 may be configured to store information (e.g., the geometries 104 and 106, or the transforms 114 or 408) within the processing unit 1302 during operation. The storage unit 1334 may include a computer-readable storage medium or a computer-readable storage device. In some examples, the storage unit 1334 includes at least a short-term memory or a long-term memory. The storage unit 1334 may include, for example, a form of random access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), magnetic disk, optical disk, flash memory, magnetic disk, optical disk, flash memory, or electrically programmable memory (EPROM) or electrically erasable and programmable memory (EEPROM).

[0092] In some examples, the storage unit 1334 is used to store program instructions for execution by the processor 1304. The storage unit 1334 may be used by software or applications running on the processing unit 1302 to store information during program execution and to store results of program execution. For example, the storage unit 1334 may store the neural network configurations 110 and 134 as they are being trained using the methods 100 and 400, respectively.

[0093] Although this specification describes many specific embodiment details, these should not be interpreted as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to a particular embodiment. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, although features may be described above as acting in a particular combination and initially claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be subject to a subcombination or a variation of the subcombination.

[0094] Similarly, although operations are shown in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in sequential order, or that all of the illustrated operations be performed, to achieve desired results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, it should be understood that the separation of various system modules and components in the embodiments described above does not require such separation in all embodiments, and that the components and systems described may generally be integrated together in a single system or distributed across multiple systems.

[0095] Certain embodiments of the subject matter have been described above. Other embodiments are within the scope of the following claims.

Claims

1. 1. A computer-implemented method for generating a setup for orthodontic alignment treatment, comprising: one or more computer processors receiving a first digital representation of a patient's teeth, the first digital representation including a plurality of mesh elements and a respective mesh element feature vector associated with each mesh element in the plurality of mesh elements, the mesh element feature vector encoding at least one of one or more spatial features and one or more structural features of a corresponding mesh element; the one or more computer processors use a generator including one or more neural networks initially trained to predict one or more tooth movements for a set-up to determine a prediction of one or more tooth movements for a set-up; the one or more computer processors further training the generator based on the using, wherein the training of the neural network includes: the generator predicting one or more tooth movements for a setup based on the first digital representation of the patient's teeth, the one or more tooth movements being described by at least one of a position and an orientation; said generator quantifying a difference between the representation of the one or more tooth movements predicted by said generator and one or more reference tooth movement representations; generating a loss value based on said quantifying; and modifying the generator based at least in part on the loss value.

2. The computer-implemented method of claim 1 , wherein the mesh elements include at least one of vertices, edges, faces, and voxels.

3. The computer-implemented method of claim 1 , wherein the mesh features include at least one of spatial features and structural features.

4. The computer-implemented method of claim 1 , further comprising the one or more processors generating an output describing one or more transformations applied to one or more teeth.

5. The computer-implemented method of claim 4 , wherein the setup is an intermediate setup.