A method for generating orthodontic treatment plans based on reinforcement learning

Through a deep neural network based on multi-agent reinforcement learning, a dental orthodontic treatment solution is generated, which solves the problems of low efficiency, high cost and tooth collisions of traditional methods, and achieves the effects of automation, high efficiency and collision avoidance.

CN119139042BActive Publication Date: 2025-05-09HANGZHOU ZOHO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411651369.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-05-09
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

Traditional dental orthodontic treatments are inefficient and costly, and lack a perfect solution to avoid tooth collisions.

Method used

Using a deep neural network based on multi-agent reinforcement learning, features are extracted and dental orthodontic treatment plans are generated by obtaining the initial and target three-dimensional digital model of the tooth jaw. The method includes determining the collision dependencies between teeth, performing topological ordering, and deciding on the orthodontic steps of each tooth based on the sorting results.

Benefits of technology

An automated generation of dental orthodontic treatment plan is realized, which improves generation efficiency, reduces costs, and effectively avoids tooth collisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119139042B_ABST
    Figure CN119139042B_ABST
Patent Text Reader

Abstract

One aspect of the present application provides a computer-implemented method for generating an orthodontic treatment plan for teeth, which includes: obtaining a first three-dimensional digital model representing an initial tooth layout of a jaw and a first digital data set representing a target tooth layout of the jaw; using a trained first deep neural network to perform feature extraction on the first three-dimensional digital model; and using a trained second deep neural network to generate an orthodontic treatment plan for the jaw based on the extracted features and the first digital data set, wherein the second deep neural network is a deep neural network based on multi-agent reinforcement learning, and the generation of each correction step includes: based on the current tooth layout, determining the collision dependency relationship between each tooth and other teeth; based on the collision dependency relationship, topologically sorting the teeth of the jaw; and using the second deep neural network to decide the action of the current correction step for each tooth according to the topological sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to a method for generating an orthodontic treatment plan based on reinforcement learning, and more particularly to a method for generating an orthodontic treatment plan by utilizing topological sorting to avoid tooth collisions. Background Art

[0002] Nowadays, polymer-based shell appliances are becoming more and more popular due to their advantages of beauty, convenience, and ease of cleaning. For orthodontic treatment of single-jaw (maxillary or mandibular) teeth, a set of shell appliances usually includes more than a dozen or even dozens of successive shell appliances, which are used to gradually reposition the teeth from an initial layout to a target layout, wherein between the initial layout and the target layout, there are N successive intermediate layouts from a first intermediate layout to a final intermediate layout.

[0003] The series of successive shell appliances are produced based on a three-dimensional digital model representing the series of successive tooth arrangements, which is often referred to as an orthodontic treatment plan for the jaw.

[0004] Typically, an orthodontic treatment plan for a single-jaw tooth is generated based on the initial layout and target layout of the teeth. Currently, this process requires manual intervention, and professionals involved in developing an orthodontic treatment plan need to undergo a lot of relevant training. Therefore, the traditional method of generating an orthodontic treatment plan is inefficient and costly. In addition, the existing method of generating an orthodontic treatment plan does not yet have a comprehensive solution for avoiding tooth collisions.

[0005] In view of the above, it is necessary to provide a new method for generating orthodontic treatment plans. Summary of the invention

[0006] One aspect of the present application provides a computer-implemented method for generating an orthodontic treatment plan for teeth, which includes: obtaining a first three-dimensional digital model representing an initial tooth layout of a jaw and a first digital data set representing a target tooth layout of the jaw; using a trained first deep neural network to perform feature extraction on the first three-dimensional digital model; and using a trained second deep neural network to generate an orthodontic treatment plan for the jaw based on the extracted features and the first digital data set, wherein the second deep neural network is a deep neural network based on multi-agent reinforcement learning, and in the process of generating the orthodontic treatment plan, each tooth is regarded as an agent, and the orthodontic treatment plan is an orthodontic treatment plan using a shell-shaped dental brace, which includes a series of successive correction steps for gradually positioning the jaw from the initial tooth layout to the target tooth layout, wherein generating each correction step includes: determining the collision dependency relationship between each tooth and other teeth based on the current tooth layout; topologically sorting the teeth of the jaw based on the collision dependency; and using the second deep neural network to decide the action of the current correction step for each tooth based on the topological sorting.

[0007] In some embodiments, the topological sorting of the teeth includes: establishing a graph theory model based on the collision dependency relationship between each tooth and other teeth; performing strong connected components and shrinking points on the graph theory model to obtain at least multiple strongly connected component blocks; topologically sorting the nodes in each of the strongly connected component blocks; topologically sorting the strongly connected blocks; and obtaining the topological sorting of the teeth based on the topological sorting of the nodes in each strongly connected component block and the topological sorting between the strongly connected component blocks.

[0008] In some embodiments, in the graph theory model, the weight of the edge between two teeth is calculated based on the number of types of actions that one tooth depends on the other tooth. If an action taken by the first tooth will collide with the second tooth, then the action of the first tooth depends on the second tooth.

[0009] In some embodiments, when deciding the action of the current correction step for each tooth, if the tooth collides with other teeth after taking an action, the action is marked as infeasible.

[0010] In some embodiments, generating each of the correction steps also includes: after using the second deep neural network to decide the action of a tooth in the current correction step, applying the action of the tooth to the model environment, and based on this, making a decision on the action of the next tooth in the topological sorting in the current correction step.

[0011] In some embodiments, the computer-implemented method for generating an orthodontic treatment plan also includes: calculating a single movement of each tooth based on the initial tooth layout and the target tooth layout, and the deep neural network based on multi-agent reinforcement learning generates an orthodontic treatment plan for the dental jaw based on the extracted features, the first digital data set, and the calculated single movement of each tooth.

[0012] In some embodiments, the jaw includes maxillary and mandibular teeth.

[0013] In some embodiments, the computer-implemented method for generating an orthodontic treatment plan further includes: uniformly sampling M points on each tooth of the first three-dimensional digital model to obtain a point cloud X of the jaw, and the first deep neural network extracts features from the point cloud X.

[0014] In some embodiments, the first deep neural network includes a first MLP module and a second MLP module, wherein the first MLP module is used for feature extraction, and the second MLP module is used to reconstruct the point cloud of the dental jaw based on the features extracted by the first MLP when training the first deep neural network.

[0015] In some embodiments, at each correction step of the orthodontic treatment plan, the optional actions of each agent include: remaining still, performing only one translation, performing only one rotation, and performing one translation and one rotation.

[0016] In some embodiments, the constraints imposed by the second deep neural network on the agent include: a single move amount constraint and a split constraint.

[0017] In some embodiments, in the training of the second deep neural network, the reward function used includes a reward value calculated based on whether any constraints are violated, whether an agent chooses to translate or rotate, whether an agent continues to perform the same translation or rotation as the previous correction step, whether an agent reaches the target posture, and whether all agents reach the target posture.

[0018] Another aspect of the present application provides a computer system for generating an orthodontic treatment plan, which includes a processor and a storage device, wherein the storage device stores a computer program for generating an orthodontic treatment plan, and when the computer program is run, the processor will execute the method for generating an orthodontic treatment plan. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other features of the present application will be further described below in conjunction with the accompanying drawings and detailed descriptions thereof. It should be understood that these drawings only illustrate several exemplary embodiments according to the present application and should not be considered as limiting the scope of protection of the present application. Unless otherwise specified, the drawings are not necessarily to scale and similar reference numerals represent similar components.

[0020] Figure 1 A schematic flowchart of a method for generating an orthodontic treatment plan executed by a computer in one embodiment of the present application;

[0021] Figure 2 The basic structure of a feature extraction deep neural network in one embodiment of the present application is schematically shown;

[0022] Figure 3A The graph theory model based on the collision dependency relationship between teeth 11 to 16 in one example is schematically shown;

[0023] Figure 3B Schematically shows Figure 3A The graph model shown is the result after strong connected components and shrinking points;

[0024] Figure 4 Schematically shows the Figure 3A and Figure 3B The topological order obtained by the graph theory model shown predicts the flow of actions of each tooth in the current frame; and

[0025] Figure 5 The basic structure of a deep neural network based on multi-agent reinforcement learning in one embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0026] The following detailed description refers to the drawings that form part of this specification. The illustrative embodiments mentioned in the specification and drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present application. Under the guidance of the present application, those skilled in the art will understand that many other embodiments can be adopted and various changes can be made to the described embodiments without departing from the subject matter and scope of protection of the present application. It should be understood that the various aspects of the present application described and illustrated herein can be arranged, replaced, combined, separated and designed according to many different configurations, and these different configurations are within the scope of protection of the present application.

[0027] One aspect of the present application provides a computer-implemented method for generating an orthodontic treatment plan, which uses a deep neural network based on multi-agent reinforcement learning to automatically generate an orthodontic treatment plan based on a three-dimensional digital model representing an initial layout and a target layout of a jaw. In the process of generating each frame of the orthodontic treatment plan, the method of the present application performs a topological sorting based on the dependence of the movement of each tooth on other teeth, and the deep neural network determines the action taken by each tooth in the current frame according to the sorting.

[0028] In one embodiment, the orthodontic treatment plan may be a dental action table, which includes the actions taken by each tooth in each correction step in the process of positioning the dental jaw from the initial layout to the target layout. It can be understood that based on the dental action table and the initial layout or the target layout, a series of successive intermediate layouts between the initial layout and the target layout can be generated, and the orthodontic treatment plan may be the series of successive intermediate layouts. It can be understood that based on the three-dimensional digital model representing the initial layout or the target layout and the series of successive intermediate layouts, a three-dimensional digital model representing a series of successive intermediate layouts between the initial layout and the target layout can be generated, and the orthodontic treatment plan may also be a three-dimensional digital model representing the series of successive intermediate layouts.

[0029] Another aspect of the present application provides a computer system for generating an orthodontic treatment plan, which includes a processor and a storage device, wherein the storage device stores a computer program for generating an orthodontic treatment plan, and when the computer program is run, the processor will execute the method for generating an orthodontic treatment plan.

[0030] Orthodontic treatment is the process of repositioning the patient's teeth from an initial layout to a target layout, and the orthodontic treatment plan includes the sequence of various actions of each tooth in the process of repositioning the teeth from the initial layout to the target layout. It can be understood that the target layout is the tooth layout that the orthodontic treatment expects to achieve, and the initial layout can be the patient's tooth layout before the orthodontic treatment, or it can be the patient's current tooth layout based on which the orthodontic plan is generated using the method of the present application.

[0031] Please refer to Figure 1 , which is a schematic flowchart of a method 100 for generating an orthodontic treatment plan executed by a computer in one embodiment of the present application.

[0032] In one embodiment, the method of the present application can be used to generate an orthodontic treatment plan for the upper and lower teeth at one time based on the data of the upper and lower teeth. In another embodiment, the method of the present application can also be used to generate an orthodontic treatment plan for only a single jaw tooth (upper or lower teeth). The method of the present application is described in detail below by taking the generation of an orthodontic treatment plan for the upper and lower teeth as an example.

[0033] In 101 , first and second three-dimensional digital models are acquired.

[0034] The first three-dimensional digital model is a three-dimensional digital model representing the initial layout of the upper and lower teeth, and the second three-dimensional digital model is a three-dimensional digital model representing the target layout of the teeth. The second three-dimensional digital model is used as a reference for determining whether a complete orthodontic treatment plan has been generated (i.e., all teeth are in the target position at the last correction step). The second three-dimensional digital model can also be replaced by the position data of each tooth under the target tooth layout, and both can be summarized by a digital data set representing the target tooth layout.

[0035] In one embodiment, the three-dimensional digital model representing the initial layout of the teeth can be obtained by directly scanning the patient's jaw. In another embodiment, the three-dimensional digital model representing the initial layout of the teeth can be obtained by scanning a physical model of the patient's jaw, such as a plaster model. In another embodiment, the three-dimensional digital model representing the initial layout of the teeth can be obtained by scanning an impression of the patient's jaw.

[0036] In one embodiment, after the three-dimensional digital model representing the initial layout of the teeth is obtained, it can be segmented so that the teeth in the three-dimensional digital model are independent of each other, so that each tooth in the three-dimensional digital model can be moved individually.

[0037] In one embodiment, a three-dimensional digital model representing a target layout of the teeth can be obtained based on the segmented three-dimensional digital model representing the initial layout of the teeth. In one embodiment, the segmented three-dimensional digital model representing the initial layout of the teeth can be manually operated to move each tooth to a target position to obtain a three-dimensional digital model representing the target layout of the teeth. The position of the teeth includes the position and posture of the teeth. In another embodiment, a computer can be used to automatically move the teeth to a target position based on the segmented three-dimensional digital model representing the initial layout of the teeth to obtain a three-dimensional digital model representing the target layout of the teeth.

[0038] In 103, feature extraction is performed on the first three-dimensional digital model.

[0039] Since the number of facets in the 3D digital model of each tooth is huge, which may reach tens of thousands or even hundreds of thousands of facets, if these facets are directly used as the features of the tooth, the amount of data will be too large. Therefore, its features in high-dimensional space can be extracted to represent the morphological information of the tooth.

[0040] In one embodiment, for each tooth of the first three-dimensional digital model, M points are uniformly sampled on its grid to obtain a point cloud X of the three-dimensional digital model. Then, a trained feature extraction deep neural network is used to extract N*d-dimensional global features based on the point cloud X. Wherein, M is a predetermined number. In one embodiment, M can be determined according to specific circumstances and requirements. For example, M can be 4000. N is the number of teeth in the first three-dimensional digital model.

[0041] In one embodiment, the feature extraction deep neural network may be a network based on a multilayer perceptron (MLP for short).

[0042] Please refer to Figure 2 , schematically illustrating the basic structure of a feature extraction deep neural network 200 in one embodiment of the present application.

[0043] The feature extraction deep neural network 200 includes an input module 201 , a first MLP module 203 , a maximum pooling layer 205 , a first output module 207 , a second MLP module 209 and a second output module 211 .

[0044] The input module 201 inputs the point cloud X into the first MLP module 203. For each tooth, the first MLP module 203 extracts M*d-dimensional features based on the point cloud X. Then, the maximum pooling layer 205 performs maximum pooling on the M*d-dimensional features of all teeth to obtain N*d-dimensional global features, which are output by the first output module 207. The N*d-dimensional global features are the input data required for the subsequent generation of orthodontic treatment plans.

[0045] When training the feature extraction neural network 200, the second MLP module 209 is used to reconstruct the point cloud Y based on the N*d-dimensional global features, and the Chamfer Distance between the two point clouds X and Y is used as the loss function. After iterative training, the parameters of the first deep neural network 200 are obtained. In one embodiment, the loss function can be expressed by the following equation (1):

[0046] Equation (1)

[0047] In 105, an orthodontic treatment plan is generated based on the global features using a trained multi-agent reinforcement learning based deep neural network.

[0048] The orthodontic treatment plan using a shell-shaped dental appliance includes a series of successive tooth layouts, each of which can be called a frame, representing the target tooth layout to be achieved in the corresponding treatment step, that is, the target posture of each tooth to be achieved in the corresponding treatment step. The tooth posture includes the position and posture information of the tooth.

[0049] In the method of the present application, each tooth is treated as an intelligent agent, and an orthodontic treatment plan is generated using a multi-agent reinforcement learning method.

[0050] In one embodiment, in each frame, each agent can choose one of the following actions: remain motionless, perform only one translation, perform only one rotation, and perform one translation and one rotation. Then, the problem to be solved by the multi-agent reinforcement learning method is to reasonably select the actions of each agent in each frame so that all agents meet the preset constraints and the total number of frames is as small as possible. In one embodiment, in addition to meeting the above conditions, the same agent can also be made to move as continuously as possible so that the moving path is as beautiful as possible.

[0051] In order to avoid collisions between teeth as much as possible, all teeth are topologically sorted according to the dependency of each tooth's action on other teeth before predicting each frame using the deep neural network based on multi-agent reinforcement learning. Then, the action of each tooth in the current frame is predicted based on the tooth order obtained by the topological sorting. When a first tooth collides with a second tooth after executing the first action, it is considered that the first action of the first tooth depends on the second tooth, that is, in order to avoid collision, the second tooth needs to be moved before the first tooth executes the first action.

[0052] In one embodiment, for the convenience of calculation, only four components can be set for the action of the tooth, which are the translation components along the x, y, and z axes of the local coordinate system and a rotation component. Each component has two options: action or no action, so there are 2 4 A total of 16 actions.

[0053] For each tooth, since its target position is known, the direction of its optional movement is also known.

[0054] In one embodiment, the FDI tooth position representation method can be used to number the teeth, and the following description is based on this tooth numbering method. It can be understood that in addition to the FDI tooth position representation method, other tooth numbering methods can also be used.

[0055] First, a graph theory model is established based on the dependency of each tooth action on other teeth.

[0056] If the first tooth will collide with the second tooth after performing an action, then an edge from the second tooth to the first tooth is connected in the graph model. The weight of the edge is the number of types of actions that cause collision. From the above, we can see that there are 16 types of actions, so the value range of the edge weight is [1, 16]. If the first tooth will not collide with the second tooth when performing any action, then the edge from the second tooth to the first tooth is not connected in the graph model.

[0057] Please refer to Figure 3A , schematically shows a graph model based on the collision dependency relationship between teeth 11 to 17 in an example. The numbers in the circles in the figure represent the tooth numbers, and the numbers on the edges connecting the circles represent the edge weights.

[0058] Next, the graph model is subjected to strongly connected components and point shrinkage.

[0059] Please refer to Figure 3B , schematically showing Figure 3A The graph model shown is the result after strong connected components and shrinking points.

[0060] After the strongly connected components, four strongly connected component blocks 401, 403, 405 and 407 are obtained.

[0061] There is only one node in the strongly connected component block 401, so the topological order in the strongly connected component block 401 is {11}.

[0062] There are three nodes in the strongly connected component block 403, and the node weights of each node are W(12)=10, W(13)=7, and W(14)=4. The node weight is the sum of the weights of the incoming edges of each node, which describes the number of actions that the tooth depends on other teeth. The smaller the node weight, the higher the priority of the tooth movement. Therefore, tooth No. 12 is prioritized, and node 12 is deleted to obtain a directed acyclic graph (DAG). Therefore, the topological order in the strongly connected component block 403 is {12,14,13}.

[0063] There are only two nodes in the strongly connected component block 405, and the node weight is 5. Therefore, any topological order can be adopted, such as {15, 16}.

[0064] There is only one node in the strongly connected component block 407, so the topological order in the strongly connected component block 401 is {17}.

[0065] Then, the strongly connected component blocks are topologically sorted according to the collision dependencies between them.

[0066] In some cases, one or more strongly connected component blocks may have no collision dependencies with other strongly connected component blocks. In this case, the disconnected strongly connected component blocks can be topologically sorted in any way. For example, Figure 3B The strongly connected component block 407 in has no dependency relationship with other strongly connected component blocks. Therefore, the final topological order can be {11, 12, 14, 13, 15, 16, 17} or {17, 11, 12, 14, 13, 15, 16}.

[0067] Next, the deep neural network based on multi-agent reinforcement learning can be used to predict the actions of each tooth in the current frame according to the topological sorting.

[0068] In the process of predicting the action of each tooth in the current frame, for tooth T i , simulate all possible actions of the tooth. If an action collides with other teeth after execution, then the action is marked as infeasible. Then, the reinforcement learning algorithm is used to make decisions and select the tooth T from the selected actions. i The action in the current frame.

[0069] Then, the tooth T i The action in the current frame is applied to the simulation environment to predict the action of the next tooth in the topological sequence in the current frame, until the actions of all teeth in the current frame are obtained, and the design of the current frame or the current correction step is obtained.

[0070] Please refer to Figure 4 , schematically showing the Figure 3A and Figure 3B The topological order {11, 12, 14, 13, 15, 16, 17} obtained by the graph theory model shown predicts the action process of each tooth in the current frame.

[0071] Since in the process of designing each frame, the action of each tooth in the current frame is determined according to the topological sorting determined based on the collision dependency relationship between teeth, the collision between teeth caused by the movement of teeth can be avoided as much as possible.

[0072] In one embodiment, the following constraints can be set for the movement of teeth so that the deep neural network based on multi-agent reinforcement learning can produce a more ideal orthodontic treatment plan.

[0073] Movement constraint: that is, the change component and total change of the same tooth in 6 degrees of freedom (3 translational degrees of freedom and 3 rotational degrees of freedom) between two frames cannot exceed the corresponding thresholds, namely the translation threshold and the rotation threshold;

[0074] Collision constraint: that is, in each frame, the collision amount between teeth cannot exceed the collision threshold;

[0075] Split constraints: extension and twisting cannot be performed at the same time, and adduction and depression cannot be performed at the same time;

[0076] Support and restraint: Since the forces between objects are mutual, there must be sufficient support to exert the force to move the teeth.

[0077] It can be understood that the setting of constraints is not limited to the above schemes, and constraints can be set according to specific needs and situations.

[0078] The core idea of ​​reinforcement learning is trial-and-error: the agent iteratively optimizes based on the feedback information obtained through interaction with the environment. In the field of multi-agent reinforcement learning, the problem to be solved is usually described as a Markov decision process.

[0079] In one embodiment, a partially observable Markov decision process (S, N, O, U, T, P, R, ) to describe the problem.

[0080] Here, N represents the number of agents.

[0081] S represents a finite state set, the state of the current frame is a vector, where The features representing the i-th agent include the d-dimensional vector representing the morphological information of the agent extracted in the previous step, the current position of the agent, the single movement of the agent (i.e., the step length of each translation or rotation of the agent), and other information.

[0082] In one embodiment, the single movement amount of an agent can be calculated based on the posture difference of the corresponding tooth between the first and second three-dimensional digital models and the threshold of the single translation and single rotation of the tooth. For example, assuming that the threshold of the single translation of a tooth is 0.2 mm, if the posture difference of the tooth between the first and second three-dimensional digital models is 1.1 mm, then the single translation amount of the tooth can be minimized without exceeding the threshold. In this example, the single translation amount of the tooth is 1.1 / 6=0.18 mm.

[0083] In another embodiment, the same single movement amount may be preset for all teeth, for example, a single translation amount of 0.2 mm. By moving the teeth according to the given single movement amount, the teeth may reach a position close to the target position. If the difference between the two is less than a given threshold, it can be considered that the tooth has reached the target position.

[0084] O represents a finite set of observations, represents the observation of the ith agent in the current frame, where Represents the features of the six adjacent agents (the agent itself, two left and right neighbors, and three opposite neighbors), including the d-dimensional vector representing the morphological information of the agent extracted in the previous step, the current position of the agent, the single movement amount of the agent, and other information.

[0085] U represents the joint action set, Represents the action of the ith agent in the current frame (i.e., the four actions mentioned above), and the selection of actions is subject to the splitting constraints.

[0086] T represents the trajectory of this Monte Carlo sampling.

[0087] P represents the state transfer equation.

[0088] R represents the reward function used when training the reinforcement learning model. In the current frame s, after the agent selects the joint action u, a reward is given .

[0089] In one embodiment, rewards may be set as follows:

[0090] 1) If any agent violates any constraint, the Monte Carlo sampling ends and rewards are given ;

[0091] 2) If an agent performs a translation or rotation, give it a reward ;

[0092] 3) If an agent translates or rotates continuously (translation and rotation are calculated separately), give a reward 0.01;

[0093] 4) If an agent reaches the target position, it will be rewarded ;

[0094] 5) If all agents reach the target position, the Monte Carlo sampling ends and rewards are given .

[0095] 6) Calculate the corresponding reward value based on the force on the anchorage tooth, as follows.

[0096] For each correction step (also called a frame), the reinforcement learning algorithm predicts a movement amount disp (also called a design amount) for each tooth, including the mobile tooth and the anchor tooth. In one embodiment, the movement amount can be expressed by the translation along the x, y and z axes of the tooth local coordinate system and the rotation around the x, y and z axes of the tooth local coordinate system, as shown in the following equation (2):

[0097] Equation (2)

[0098] Among them, T stands for translation and R stands for rotation.

[0099] Next, the biomechanical calculation model can be used to calculate the forces and moments borne by tooth i in various directions based on the movement of tooth i, as shown in the following equation (3):

[0100] Equation (3)

[0101] In one embodiment, a finite element analysis method may be used to obtain a simulation data set, and the biomechanical calculation model may be trained using the simulation data set.

[0102] Then, the teeth can be divided into mobile teeth and support teeth according to the amount of movement. In one embodiment, teeth whose amount of movement and rotation is less than a preset threshold can be classified as support teeth. The support teeth are basically immobile in the current correction step and provide support for the mobile teeth. The preset thresholds may include a translation threshold and a rotation threshold. For example, the translation threshold may be 0.05 mm, and the rotation threshold may be 0.5°.

[0103] Next, according to the optimal force range [F min ,F max ] and the following equation (4) calculates a reward value R for the tooth:

[0104] Equation (4)

[0105] Among them, R penalty is the penalty factor, which is a constant less than 0.

[0106] In one embodiment, the optimal force ranges for different types of teeth may be different. For example, different optimal force ranges may be set for different types of teeth, for example, different optimal force ranges may be set for molars, premolars, canines, lateral incisors, and central incisors. In one embodiment, the optimal force ranges may be set for different teeth by statistical data.

[0107] In another embodiment, the force range may be replaced by a preset force threshold. When the force on an anchoring tooth is greater than the corresponding force threshold, the reward given is smaller than the reward given when the force on the anchoring tooth is less than the force threshold.

[0108] In the above embodiments, the reward value is calculated only for the force on the tooth, but not for the moment on the tooth. However, based on the present application, it can be understood that the reward value can be calculated for the force and moment on a tooth, respectively, and then the total reward value for the force on the tooth can be obtained by adding the two together.

[0109] After calculating the reward of the current frame, it can be normalized according to the following equation (5):

[0110] Equation (5)

[0111] in, represents the maximum possible value of the reward, is the normalization factor.

[0112] represents the discount factor, and .

[0113] After the problem is modeled, the parameters of the reinforcement learning model can be solved using algorithms in the field of multi-agent reinforcement learning, such as the QMix algorithm. It is understood that in addition to the QMix algorithm, any other applicable algorithms can also be used, such as IQL (Implicit Q-learning), VDN (Value-Decomposition Networks), COMA (Counterfactual Multi-Agent Policy Gradients), and Reinforce.

[0114] In one embodiment, for the observation of agent i, a recurrent neural network can be used to estimate its probability under each possible action. In one embodiment, a Gated Recurrent Unit (a type of recurrent neural network, referred to as GRU) can be used to estimate the agent i under each possible action. function.

[0115] Please refer to Figure 5 , schematically illustrating the basic structure of a deep neural network 300 based on multi-agent reinforcement learning in one embodiment of the present application.

[0116] The deep neural network 300 based on multi-agent reinforcement learning includes an observation module 301, a first MLP module 303, a GRU module 305, a second MLP module 307 and an output module 309.

[0117] The observation module 301 is the observation of the environment by the agent, which is used as the input of the network to predict the next action. The first MLP module 303 reduces the dimension of the input features to reduce the complexity of the GRU module 305. The GRU module 305 is used to estimate the agent i under each possible action. Function. Module h returns the output of GRU to its input. The second MLP module 307 is used to perform regression prediction on the output of GRU module 305 to obtain the value of each action.

[0118] In getting After all, the global Function, as expressed in the following equation (6):

[0119] Equation (6)

[0120] Wherein, W and b in equation (6) can be predicted by the first MLP module 303 and the second MLP module 307 respectively, and T represents the transpose of the matrix.

[0121] Equation (7)

[0122] In equation (7), TD target Represents the target, Represents the discount factor.

[0123] Then the loss function of the deep neural network 300 based on multi-agent reinforcement learning can be expressed by the following equation (8):

[0124] Equation (8)

[0125] In the training of the deep neural network 300 based on multi-agent reinforcement learning, two identical networks were used, one of which produced Q tot , and the other produces Q' tot The difference between these two networks is that the tot The parameters of the network are lagged to generate Q' tot The network frames.

[0126] After a large number of Monte Carlo samplings, a series of intelligent agent action trajectories are obtained and put into the experience pool. Each time, some trajectories are randomly sampled from the experience pool as samples for training to obtain the parameters of the required model.

[0127] Finally, after reasoning, we get the value function Q of all possible actions selected by all agents: tot (s, a), take the action a corresponding to the maximum value function value as the decision value, and obtain a frame of the orthodontic treatment plan.

[0128] In light of the present application, it can be understood that in addition to the recurrent network, any other suitable neural network can be used to generate orthodontic treatment plans, such as CommNet and G2ANet.

[0129] In light of the present application, it can be understood that in addition to the modeling methods in the above embodiments, any other applicable modeling methods may be adopted, for example, other reasonable reward functions, and other reasonable state and observation designs.

[0130] The method for generating an orthodontic treatment plan of the present application is automatically executed by a computer, which can save a lot of manpower.

[0131] Although various aspects and embodiments of the present application are disclosed herein, other aspects and embodiments of the present application will be apparent to those skilled in the art in light of the present application. The various aspects and embodiments disclosed herein are for illustrative purposes only and not for limiting purposes. The scope and subject matter of the present application are determined solely by the appended claims.

[0132] Likewise, various diagrams may illustrate exemplary architectures or other configurations of the disclosed methods and systems that aid in understanding the features and functions that may be included in the disclosed methods and systems. The claimed content is not limited to the exemplary architectures or configurations shown, and the desired features may be implemented with various alternative architectures and configurations. In addition, for flow charts, functional descriptions, and method claims, the order of blocks presented herein should not be limited to various embodiments that are implemented in the same order to perform the described functions, unless otherwise clearly indicated in the context.

[0133] Unless otherwise expressly noted, the terms and phrases used herein and their variations should be interpreted as open ended, not restrictive. In some instances, the appearance of broad words and phrases such as "one or more", "at least", "but not limited to", or other similar terms should not be understood as an intent or need to indicate a narrowing of the example where such broad terms may not be present.

Claims

1. A computer-implemented method for generating an orthodontic treatment plan, comprising: Acquire a first three-dimensional digital model representing an initial tooth arrangement of a dental jaw and a first digital data set representing a target tooth arrangement of the dental jaw; Using the trained first deep neural network to perform feature extraction on the first three-dimensional digital model; as well as Using a trained second deep neural network, an orthodontic treatment plan for the jaw is generated based on the extracted features and the first digital data set, wherein the second deep neural network is a deep neural network based on multi-agent reinforcement learning, and in the process of generating the orthodontic treatment plan, each tooth is regarded as an agent, and the orthodontic treatment plan is an orthodontic treatment plan using a shell-shaped dental appliance, which includes a series of successive correction steps for gradually positioning the jaw from an initial tooth layout to a target tooth layout, wherein generating each correction step includes: determining a collision dependency relationship between each tooth and other teeth based on the current tooth layout; topologically sorting the teeth of the jaw based on the collision dependency relationship; and using the second deep neural network to decide the action of the current correction step for each tooth according to the topological sorting; Among them, the topological sorting of the teeth includes: establishing a graph theory model based on the collision dependency relationship between each tooth and other teeth, in which the weight of the edge between two teeth is calculated based on the number of action types that one tooth depends on the other tooth, and if the first tooth takes an action that will collide with the second tooth, then the action of the first tooth depends on the second tooth; performing strong connected components and shrinking points on the graph theory model to obtain at least multiple strong connected component blocks; topologically sorting the nodes in each of the strong connected component blocks; topologically sorting the strongly connected component blocks; and obtaining the topological sorting of the teeth based on the topological sorting of the nodes in each strongly connected component block and the topological sorting between the strongly connected component blocks.

2. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: When deciding the action for the current correction step for each tooth, if the tooth collides with other teeth after taking an action, the action is marked as infeasible.

3. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: Generating each of the correction steps also includes: after using the second deep neural network to decide the action of a tooth in the current correction step, applying the action of the tooth to the model environment, and based on this, making a decision on the action of the next tooth in the topological sorting in the current correction step.

4. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: It also includes: calculating the single movement amount of each tooth based on the initial tooth layout and the target tooth layout, and the deep neural network based on multi-agent reinforcement learning generates an orthodontic treatment plan for the dental jaw based on the extracted features, the first digital data set and the calculated single movement amount of each tooth.

5. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: The jaw includes maxillary and mandibular teeth.

6. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: It also includes: uniformly sampling M points on each tooth of the first three-dimensional digital model to obtain a point cloud X of the jaw, and the first deep neural network extracts features from the point cloud X.

7. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: The first deep neural network includes a first MLP module and a second MLP module, wherein the first MLP module is used for feature extraction, and the second MLP module is used to reconstruct the point cloud of the dental jaw based on the features extracted by the first MLP when training the first deep neural network.

8. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: In each correction step of the orthodontic treatment plan, the optional actions of each agent include: keeping still, performing only one translation, performing only one rotation, and performing one translation and one rotation.

9. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: The constraints imposed by the second deep neural network on the agent include: a single movement amount constraint and a splitting constraint.

10. The computer-implemented method for generating an orthodontic treatment plan according to claim 1, wherein: In the training of the second deep neural network, the reward function used includes a reward value calculated based on whether any constraints are violated, whether an agent chooses to translate or rotate, whether an agent continues to perform the same translation or rotation as the previous correction step, whether an agent reaches the target posture, and whether all agents reach the target posture.

11. A computer system for generating an orthodontic treatment plan, comprising a processor and a storage device, wherein: The storage device stores a computer program for generating an orthodontic treatment plan. When the computer program is run, the processor will execute the method for generating an orthodontic treatment plan as claimed in claim 1.

Citation Information

Patent Citations

  • Computer-implemented method for generating orthodontic treatment regimen

    CN118675698A

  • Direct fractional step method for generating tooth arrangement

    US20160175068A1