Multi-dental implant positioning planning method based on distributed reinforcement learning
Optimizing multi-dental implant positioning through distributed reinforcement learning and bitter fish optimization algorithms has solved the problems of unstable planning results and difficult to dynamically optimize in the existing technology, and efficient and accurate multi-dental implant positioning is achieved, improving the stability and aesthetic effect of the implant.
Patent Information
- Application Number
- CN202510370612.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-dental implant positioning technology relies on doctor experience, poor stability and repetition of planning results, and the digital planning method lacks dynamic optimization mechanism. It is difficult for traditional machine learning methods to achieve efficient and accurate automatic multi-dental implant planning in complex oral environments.
A distributed reinforcement learning method is adopted to establish a multi-agent reinforcement learning model. Each dental implant is used as an independent agent. It is trained by defining the state space, action space and reward function, combined with the bitter fish optimization algorithm for dynamic perturbation, optimize the position and angle of the implant, and use biomechanical and aesthetic evaluation to ensure the feasibility of the planning scheme.
It improves the intelligence level of implant planning, reduces the risk of postoperative complications, enhances the long-term stability and biomechanical performance of the implant, optimizes the occlusal balance and aesthetic effect, and shortens the planning time.
Smart Images

Figure CN120260943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dental technology, and in particular, to a multi-dental implant positioning and planning method based on distributed reinforcement learning. Background Art
[0002] With the progress of oral medicine technology, dental implants have become a widely adopted restoration method for toothless patients. However, the precise positioning and planning of multi-dental implants remain a complex and challenging issue, involving multiple factors such as oral anatomical structures, biomechanical stability, and aesthetic coordination. Currently, clinically, the positioning of implants mainly relies on doctors' experience and traditional digital technologies. However, existing technologies still have obvious limitations in terms of precision and intelligence.
[0003] Currently, the positioning of multi-dental implants mainly relies on doctors to manually plan the implant positions based on the patient's oral medical imaging data. Doctors need to comprehensively consider factors such as the alveolar bone morphology, bone density distribution, occlusion relationship, and aesthetic requirements, and formulate implant plans in combination with clinical experience. Although this manual planning method ensures personalized treatment to a certain extent, due to differences in doctors' experience levels, the stability and repeatability of the planning results are poor. In addition, doctors need to perform a large amount of image data analysis and evaluation during the planning process, which takes a long time and is easily affected by subjective judgment, resulting in an increase in the uncertainty of implant position and angle selection.
[0004] To improve the accuracy of multi-dental implant positioning, some computer-aided design and computer-aided manufacturing technologies have been applied to implant surgery planning. For example, a three-dimensional model of the patient's oral cavity is established based on digital imaging data, and the implant implantation is guided by a surgical guide. However, the static planning method cannot fully consider dynamic factors such as bone remodeling and biomechanical adaptation during the operation, resulting in risks to the long-term stability of the implants after surgery. In addition, existing digital planning methods usually rely on manual interaction for parameter adjustment, which requires a high level of computer operation ability for doctors and is difficult to achieve high automation and intelligence.
[0005] In recent years, machine learning and artificial intelligence technologies have begun to be applied to the fields of medical image analysis and surgery planning. For example, deep learning technologies have been used for the automatic segmentation and feature extraction of oral medical images to assist doctors in implant planning. However, traditional machine learning methods often rely on a large amount of labeled data for training and are difficult to achieve adaptive optimization in complex and variable oral anatomical environments. In addition, existing automatic planning methods usually only optimize single-tooth implants and are difficult to balance biomechanical stability, aesthetic effects, and occlusion balance under the complex spatial constraints of multi-dental implants.
[0006] In summary, the existing multi - dental implant positioning technologies have obvious deficiencies in the following aspects: relying on doctors' subjective experience, the stability and repeatability of the planning scheme are poor; the existing digital planning methods lack a dynamic optimization mechanism and are difficult to adapt to complex oral biomechanical environments; traditional machine learning methods are limited by data dependence and optimization capabilities and are difficult to achieve efficient and accurate automatic planning of multi - dental implants. Therefore, there is an urgent need for a new method to improve the intelligent level of multi - dental implant positioning planning, improve accuracy and stability to meet clinical needs. Summary of the Invention
[0007] An object of the present invention is to propose a multi - dental implant positioning and planning method based on distributed reinforcement learning. The present invention effectively improves the clinical feasibility of the implant planning scheme, reduces the risk of postoperative complications, and improves the stability of long - term use of implants.
[0008] A multi - dental implant positioning and planning method based on distributed reinforcement learning according to an embodiment of the present invention includes the following steps:
[0009] S1. Obtain the medical image data of the patient, including computed tomography data and magnetic resonance imaging data, and establish a three - dimensional virtual oral model of the patient's oral cavity based on the medical image data. The three - dimensional virtual oral model includes an alveolar bone region model, a tooth arrangement model, a bone density distribution model, and a bite mechanics model;
[0010] S2. Based on the three - dimensional virtual oral model, establish a multi - agent reinforcement learning model under a distributed reinforcement learning framework. Each candidate position of the to - be - implanted tooth in the multi - agent reinforcement learning model is used as an independent reinforcement learning agent. Define the state space, action space, and reward function of each reinforcement learning agent. Among them, the state space includes the spatial coordinate information, local bone density distribution information, and bite mechanics parameter information of each candidate position. The action space includes the spatial coordinate adjustment direction and angle change information of each candidate position. The reward function is jointly composed of a biomechanical stability index, an aesthetic optimization index, and a bite balance index;
[0011] S3. Based on the reinforcement learning agents, adopt a training strategy of centralized batch training and distributed execution in a distributed computing environment, use a multi - node parallel architecture to train the multi - agent reinforcement learning model, and update the state information, action information, and reward information of each reinforcement learning agent in real time to obtain a preliminary training result of the multi - agent reinforcement learning model;
[0012] S4. During the training process of the multi-agent reinforcement learning model, the bitter fish optimization algorithm is introduced as a dynamic perturbation mechanism to randomly perturb the reinforcement learning agents with an obvious convergence trend of the reward function during training. When the change of the reward function of any reinforcement learning agent tends to stable convergence, the bitter fish optimization algorithm automatically performs random perturbation on the state parameters of the agent. The perturbed state parameters re-participate in the reinforcement learning training process and recalculate the reward function, continuously optimizing the agent's decision-making strategy until the global optimization condition is met, and obtaining the globally optimal training result of the multi-agent reinforcement learning model;
[0013] S5. Import the multi-tooth implant position and angle scheme output by the multi-agent reinforcement learning model training result into the biomechanical simulation analysis system, and conduct biomechanical simulation analysis and aesthetic evaluation on the multi-tooth implant scheme based on the patient's three-dimensional virtual oral model, evaluate the biomechanical stability, aesthetic effect and occlusal balance of the scheme, confirm that the scheme meets the safety and effectiveness of clinical application, and obtain the verified multi-tooth implant positioning and planning scheme;
[0014] S6. Generate a surgical guide design file for preoperative guidance based on the multi-tooth implant positioning and planning scheme.
[0015] Optionally, the specific content in step S1 includes:
[0016] S11. Obtain the oral medical image data of the patient, including computed tomography data and magnetic resonance imaging data;
[0017] S12. Perform spatial registration and data fusion on the oral medical image data respectively to obtain the fused oral medical image dataset D fusion ;
[0018] S13. Based on the oral medical image data, extract the alveolar bone region characteristics, tooth arrangement characteristics, bone density distribution characteristics and occlusal contact surface characteristics in the patient's oral cavity through medical image segmentation algorithms;
[0019] S14. According to the extracted feature information, establish a model including the alveolar bone region model M bone , tooth arrangement model M teeth , bone density distribution model BD density (x, y, z) and occlusal mechanics model F bite . Among them, the alveolar bone region model is constructed according to the alveolar bone boundary characteristics of the patient individual, the tooth arrangement model is constructed based on the tooth row position and direction characteristics of the patient individual, the bone density distribution model is constructed by converting the Hounsfield unit value HU(x, y, z) into a density value, and the occlusal mechanics model is constructed from the occlusal contact area and bite force distribution data;
[0020] S15. Based on the alveolar bone region model, tooth arrangement model, bone density distribution model, and occlusal mechanics model, perform three-dimensional spatial registration and fusion to construct a three-dimensional virtual oral model M of the patient's oral cavity oral3D :
[0021]
[0022] Among them, μ BD and σ BD respectively represent the mean and standard deviation of the bone density data. α and β are weight factors that respectively adjust the contribution degrees of the alveolar bone region model and the tooth arrangement model in the global model construction. Λ(·) and Ω(·) are transformation functions for normalizing and preprocessing the alveolar bone region model M bone and the tooth arrangement model M teeth . tanh(·) is the hyperbolic tangent function. BD density (x, y, z) - μ BD represents the deviation between the local bone density and the average bone density. σ BD is the normalization factor. sin(·) is the sine function. κ is the proportional scaling constant of the occlusal mechanics data. Ψ{·} is the three-dimensional spatial registration and fusion operator
[0023] Optionally, the specific steps in S2 include:
[0024] S21. Based on the three-dimensional virtual oral model M of the patient's oral cavity oral3D establish a multi-reinforcement learning agent reinforcement learning model under a distributed reinforcement learning framework. Each candidate position for dental implantation is used as an independent reinforcement learning agent Agent i , where i represents the i-th candidate implantation position
[0025] S22. Define the state space S i of each reinforcement learning agent Agent i :
[0026] S i = {P i (x, y, z), BD density,i (x, y, z), F bite,i (x, y, z)}
[0027] Among them, P i (x, y, z) is the spatial coordinate information, representing the three-dimensional spatial coordinates of the corresponding candidate implantation position of the reinforcement learning agent Agent i in the three-dimensional virtual oral model M oral3D . BD density,i (x, y, z) is the local bone density distribution information, used to characterize the local bone density environment of the candidate implantation position, Fbite,i (x, y, z) is the occlusal mechanics parameter information, which is used to characterize the occlusal mechanics distribution characteristics of the candidate implant position;
[0028] S23. Define each reinforcement learning agent i The action space A i :
[0029] A i ={Δd i ,Δθ i};
[0030] Where, Δd i The three-dimensional vector information representing the spatial coordinate adjustment direction of each candidate implant position represents the candidate implant position in the three-dimensional virtual oral model M oral3D The adjustment of the spatial coordinates in i The implant angle change information of each candidate implantation position is represented, which characterizes the adjustment of the implant in the spatial implantation direction;
[0031] S24. Define each reinforcement learning agent i The reward function R i , the reward function R i The biomechanical stability index R bio,i , aesthetic optimization index R esth,i and occlusal balance index R occlu,i Common composition:
[0032] R i =γ1·R bio,i +γ2·R esth,i +γ3·R occlu,i ;
[0033] Among them, R bio,i It is a biomechanical stability index used to evaluate reinforcement learning agents i In the 3D virtual oral model M oral3D The stability of local bone stress distribution after implantation at the candidate position, R esth,i It is an aesthetic optimization index used to evaluate reinforcement learning agents i In the 3D virtual oral model M oral3D Candidate position and tooth arrangement model M teeth The spatial coordination between occlu,i It is a bite balance indicator used to evaluate the reinforcement learning agent i In the 3D virtual oral model M oral3D The balance of occlusal mechanics distribution after implantation in the corresponding candidate position.
[0034] Optionally, the step S3 specifically includes:
[0035] S31. In a distributed computing environment, all reinforcement learning agents are trained according to the centralized batch training strategy. i The state space S i , action space A i With the reward function R i Information constitutes training batch B t :
[0036] B t ={(S i ,A i ,R i )|i=1,2,…,N};
[0037] Where N represents the total number of reinforcement learning agents to be trained;
[0038] S32. Using a multi-node parallel architecture, each distributed computing node j is trained according to the training batch B t Calculate the local loss function L j (B t ), and use the gradient descent method to update the multi-agent reinforcement learning model parameters:
[0039]
[0040] in, represents the parameter set of the multi-agent reinforcement learning model in the jth node at the tth iteration, η is the learning rate, Indicates that based on training batch B t The calculated local gradient;
[0041] S33. Update the model parameters of each node through the distributed communication mechanism Perform global aggregation to form a global multi-agent reinforcement learning model parameter set Θ t+1 :
[0042]
[0043] Where Φ(·) is the global parameter aggregation operator, and M is the total number of distributed computing nodes participating in the training;
[0044] S34. Repeat steps S31 to S33 until the global loss function L(Θ t ) satisfies the preset convergence condition, that is, L(Θ t )≤∈, where ∈ is the preset threshold, or reaches the preset maximum number of iterations T max , where T max is the preset maximum number of iterations, thereby obtaining the preliminary training results Θ of the multi-agent reinforcement learning model* .
[0045] Optionally, the specific steps in step S4 include:
[0046] S41. During the training process of the multi-agent reinforcement learning model, for each reinforcement learning agent Agent i the value of the reward function is dynamically monitored and evaluated, and the variance of the reward function change in its last K iterations is calculated in real time
[0047]
[0048] wherein, represents the degree of fluctuation of the reward function of the reinforcement learning agent Agent i in multiple consecutive iterations, is the average value of the reward function in the last K iterations, and T is the current iteration number;
[0049] S42. When the reinforcement learning agent Agent i meets the convergence condition of the reward function change variance where δ σ is a preset convergence determination threshold, the bitter fish optimization algorithm automatically triggers a targeted adaptive dynamic perturbation mechanism, and adaptively adjusts the perturbation amplitude ζ according to the reward function variance i :
[0050]
[0051] where ζ i is the dynamic adaptive perturbation amplitude of the reinforcement learning agent Agent i , ζ max is the maximum preset value of the perturbation amplitude, and λ1 is the perturbation amplitude decay coefficient;
[0052] S43. Based on the perturbation amplitude ζ i combined with the bone density distribution information BD i corresponding to the candidate planting position of the reinforcement learning agent Agent density,i (x, y, z) and the occlusal mechanics parameter information F bite,i (x, y, z), calculate the bone density-occlusal mechanics coupling direction perturbation vector that adapts to the characteristics of the complex oral environment to intelligently guide the reinforcement learning agent Agent i to optimize and explore in the direction where the bone density gradient changes significantly and the occlusal mechanics distribution is more balanced:
[0053]
[0054] Among them, and respectively represent the gradient directions of the bone mineral density distribution and the occlusal mechanical parameters at the candidate planting positions corresponding to the intelligent agent Agent i × represents the cross product;
[0055] S44. Perform a deep optimization perturbation operation on the original state parameter S i of the reinforcement learning intelligent agent Agent i to obtain the state parameter after optimization perturbation
[0056]
[0057] where S i is the state parameter before perturbation, represents a normal distribution random vector with a mean of 0 and a covariance of the identity matrix I, represents the optimization perturbation vector calculated based on the bone mineral density-occlusal mechanics coupling direction;
[0058] S45. Replace the original state parameter S with the state parameter after optimization perturbation i and re-invest it into the training process of the reinforcement learning model, and recalculate the reward function after perturbation Continuously iteratively optimize the decision-making strategy of the intelligent agent until the global loss function L(Θ t ) reaches the preset convergence threshold ∈, and obtain the training result of the global optimal multi-agent reinforcement learning model.
[0059] Optionally, the specific steps in step S5 include:
[0060] S51. Import the multi-dental implant position and angle scheme * output by the global optimal multi-agent reinforcement learning model training result Θ into the biomechanical simulation analysis system, where is the three-dimensional space coordinate position of the i-th implant output by the reinforcement learning intelligent agent Agent i , and is the spatial implantation angle of the corresponding implant;
[0061] S52. Based on the patient's three-dimensional virtual oral model M oral3D , establish a multi-dental implant biomechanical finite element simulation model M FEA using the finite element analysis method:
[0062]
[0063] where Γ{·} is the finite element modeling operator, which is used to perform finite element modeling on the three-dimensional virtual oral model Moral3D Precisely implant the position and angle of the implant, and automatically generate the corresponding meshed finite element simulation model;
[0064] S53. Based on the finite element simulation model M FEA Apply the load conditions, and evaluate the biomechanical stability evaluation function E after implanting multiple dental implants through calculation bio Simulate the patient's real occlusal behavior:
[0065]
[0066] where V i represents the finite element analysis volume region of the alveolar bone around the i-th implant, and σ stress,i (x, y, z) represents the stress field distribution function of the local bone tissue region where the i-th implant is located, and E bio is the total stress index of biomechanical stability, characterizing the overall stress distribution of the implant scheme under the occlusal load;
[0067] S54. Evaluate the aesthetic effect of the implant scheme in the three-dimensional virtual oral model M oral3D and calculate the aesthetic evaluation index E esthetic :
[0068]
[0069] where is the position coordinate of the i-th implant optimized by reinforcement learning, represents the ideal implant position coordinate given by the tooth arrangement model M teeth ;
[0070] S55. Evaluate the occlusal balance of the implant scheme in the three-dimensional virtual oral model M oral3D and calculate the occlusal balance index E occlusal :
[0071]
[0072] where represents the occlusal load borne by the local position after implanting the i-th implant, is the average value of the occlusal loads of all implants, and this index is used to quantify the balance degree of the implant scheme in the overall occlusal load distribution;
[0073] S56. Synthesize the total stress index E of biomechanical stability bio , the aesthetic evaluation index E esthetic and the occlusal balance index E occlusal, confirm that all indicators meet the preset clinical safety and effectiveness threshold conditions, and finally obtain a multi-tooth implant positioning planning scheme verified by biomechanical simulation and aesthetic evaluation.
[0074] The beneficial effects of the present invention are as follows:
[0075] (1) By establishing a multi-agent reinforcement learning model, the present invention takes the position of each tooth to be implanted as an independent agent, and defines its state space, action space and reward function. During the reinforcement learning process, each agent can adaptively learn based on factors such as alveolar bone density and occlusal mechanical properties, and can be efficiently trained through a distributed computing architecture to dynamically adjust the spatial position and angle of the implant, ensuring the best balance among biomechanical stability, aesthetic optimization and occlusal balance, and avoiding the uncertainty and inefficiency brought by manual adjustment by doctors.
[0076] (2) The present invention introduces a bitter fish optimization algorithm to dynamically perturb the agents with an obvious convergence trend of the reward function during the reinforcement learning training process. By recalculating the reward function after perturbation, it prompts the reinforcement learning agents to jump out of the local optimal solution and improve the global optimization ability. The bitter fish optimization algorithm can adaptively adjust the perturbation amplitude according to the change of the variance of the reward function during the training process, enabling the reinforcement learning agents to better explore the best implantation scheme of the implant in a complex bone density environment.
[0077] (3) The present invention combines a biomechanical simulation analysis system, imports the multi-tooth implant position and angle scheme output by the reinforcement learning agents into the biomechanical simulation system, and conducts biomechanical analysis and aesthetic evaluation based on the patient's three-dimensional oral model. During the simulation process, the finite element analysis method is used to calculate the stress distribution of the multi-tooth implant under the action of occlusal force, and through aesthetic evaluation indicators and occlusal balance indicators, it is ensured that the final positioning scheme meets the clinical safety and functional requirements, effectively improving the clinical feasibility of the implant planning scheme, reducing the risk of postoperative complications, and improving the long-term use stability of the implant. Description of the Drawings
[0078] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0079] Figure 1 is a flowchart of a multi-tooth implant positioning planning method based on distributed reinforcement learning proposed by the present invention. Detailed Embodiments
[0080] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0081] Reference Figure 1 , a multi-dental implant positioning and planning method based on distributed reinforcement learning, comprising the following steps:
[0082] S1. Obtain the medical image data of the patient, including computed tomography data and magnetic resonance imaging data, and establish a three-dimensional virtual oral cavity model of the patient's oral cavity based on the medical image data. The three-dimensional virtual oral cavity model includes an alveolar bone region model, a tooth arrangement model, a bone density distribution model, and a bite mechanics model;
[0083] S2. Based on the three-dimensional virtual oral cavity model, establish a multi-agent reinforcement learning model under the distributed reinforcement learning framework. Each candidate position of the dental implant to be implanted is used as an independent reinforcement learning agent. Define the state space, action space, and reward function of each reinforcement learning agent. Among them, the state space includes the spatial coordinate information, local bone density distribution information, and bite mechanics parameter information of each candidate position. The action space includes the spatial coordinate adjustment direction and angle change information of each candidate position. The reward function is jointly composed of a biomechanical stability index, an aesthetic optimization index, and a bite balance index;
[0084] S3. Based on the reinforcement learning agent, adopt a training strategy of centralized batch training and distributed execution in a distributed computing environment, use a multi-node parallel architecture to train the multi-agent reinforcement learning model, and update the state information, action information, and reward information of each reinforcement learning agent in real time to obtain a preliminary training result of the multi-agent reinforcement learning model;
[0085] S4. During the training process of the multi-agent reinforcement learning model, introduce the bitter fish optimization algorithm as a dynamic perturbation mechanism to randomly perturb the reinforcement learning agents with an obvious convergence trend of the reward function during training. When the change of the reward function of any reinforcement learning agent tends to be stable and convergent, the bitter fish optimization algorithm automatically randomly perturbs the state parameters of the agent. The perturbed state parameters re-participate in the reinforcement learning training process and recalculate the reward function, continuously optimizing the agent's decision-making strategy until the global optimization condition is met, and obtaining the globally optimal training result of the multi-agent reinforcement learning model;
[0086] S5. Import the multi-dental implant position and angle plan output from the training result of the multi-agent reinforcement learning model into the biomechanical simulation analysis system, and conduct biomechanical simulation analysis and aesthetic evaluation on the multi-dental implant plan based on the patient's three-dimensional virtual oral model to evaluate the biomechanical stability, aesthetic effect and occlusal balance of the plan, confirm that the plan meets the safety and effectiveness of clinical application, and obtain the verified multi-dental implant positioning and planning plan;
[0087] S6. Generate a surgical guide design file for preoperative guidance based on the multi-dental implant positioning and planning plan.
[0088] In this embodiment, step S1 specifically includes:
[0089] S11. Obtain the oral medical image data of the patient, including computed tomography data and magnetic resonance imaging data;
[0090] S12. Perform spatial registration and data fusion on the oral medical image data respectively to obtain the fused oral medical image dataset D fusion ;
[0091] S13. Extract the alveolar bone region features, tooth arrangement features, bone density distribution features and occlusal contact surface features in the patient's oral cavity based on the oral medical image data through medical image segmentation algorithms;
[0092] S14. According to the extracted feature information, establish a model including the alveolar bone region model M bone , tooth arrangement model M teeth , bone density distribution model BD density (x, y, z) and occlusal mechanics model F bite , where the alveolar bone region model is constructed based on the alveolar bone boundary characteristics of the patient's individual, the tooth arrangement model is constructed based on the tooth row position and direction characteristics of the patient's individual, the bone density distribution model is constructed by converting the Hounsfield unit value HU(x, y, z) into a density value, and the occlusal mechanics model is constructed from the occlusal contact area and bite force distribution data;
[0093] S15. Based on the alveolar bone region model, tooth arrangement model, bone density distribution model and occlusal mechanics model, perform three-dimensional spatial registration and fusion to construct the three-dimensional virtual oral model M of the patient's oral cavity oral3D :
[0094]
[0095] Among them, μ BD and σ BDrespectively represent the mean and standard deviation of the bone density data, α and β are weight factors that respectively adjust the contribution degrees of the alveolar bone region model and the tooth arrangement model in the construction of the global model, Λ(·) and Ω(·) are respectively transformation functions for normalizing and preprocessing the alveolar bone region model M bone , the tooth arrangement model M teeth , tanh(·) is the hyperbolic tangent function, BD density (x,y,z)-μ BD represents the deviation between the local bone density and the average bone density, σ BD is the normalization factor, sin(·) is the sine function, κ is the scaling constant of the occlusal mechanics data, and Ψ{·} is the three-dimensional space registration and fusion operator.
[0096] In this embodiment, step S2 specifically includes:
[0097] S21. Based on the three-dimensional virtual oral model M of the patient's oral cavity oral3D establish a multi-reinforcement learning agent reinforcement learning model under the distributed reinforcement learning framework, and each candidate position for implanting teeth is used as an independent reinforcement learning agent Agent i , where i represents the i-th candidate implant position;
[0098] S22. Define the state space S i of each reinforcement learning agent Agent i :
[0099] S i ={P i (x,y,z), BD density,i (x,y,z), F bite,i (x,y,z)};
[0100] Among them, P i (x,y,z) is the spatial coordinate information, indicating the three-dimensional spatial coordinates of the corresponding candidate implant position in the three-dimensional virtual oral model M i , BD oral3D (x,y,z) is the local bone density distribution information, used to characterize the local bone density environment of the candidate implant position, and F density,i (x,y,z) is the occlusal mechanics parameter information, used to characterize the occlusal mechanics distribution characteristics borne by the candidate implant position; bite,i (x,y,z) is the occlusal mechanics parameter information, used to characterize the occlusal mechanics distribution characteristics borne by the candidate implant position;
[0101] S23. Define the action space A i of each reinforcement learning agent Agent i :
[0102] A i ={Δdi , Δθ i};
[0103] Among them, Δd i represents the three-dimensional vector information of the adjustment direction of the spatial coordinates of each candidate implant position, characterizing the adjustment of the spatial coordinates of the candidate implant position in the three-dimensional virtual oral model M oral3D , and Δθ i represents the implant angle change information of each candidate implant position, characterizing the adjustment of the implant in the spatial implantation direction;
[0104] S24. Define the reward function R i of each reinforcement learning agent Agent i , and the reward function R i is composed of the biomechanical stability index R bio,i , the aesthetic optimization index R esth,i and the occlusion balance index R occlu,i :
[0105] R i = γ1·R bio,i + γ2·R esth,i + γ3·R occlu,i ;
[0106] Among them, R bio,i is the biomechanical stability index, used to evaluate the stability of the local bone stress distribution after implanting an implant at the corresponding candidate position by the reinforcement learning agent Agent i in the three-dimensional virtual oral model M oral3D , R esth,i is the aesthetic optimization index, used to evaluate the spatial coordination between the candidate position and the tooth arrangement model M i in the three-dimensional virtual oral model M oral3D , and R teeth is the occlusion balance index, used to evaluate the balance of the occlusion mechanical distribution after implanting an implant at the corresponding candidate position by the reinforcement learning agent Agent occlu,i in the three-dimensional virtual oral model M i . oral3D
[0107] In this embodiment, step S3 specifically includes:
[0108] S31. In the distributed computing environment, according to the centralized batch training strategy, form the training batch B i by combining the state space S i , the action space A i and the reward function R i information of all reinforcement learning agents Agent t :
[0109] B t ={(S i ,A i ,R i )∣i=1,2,…,N};
[0110] Among them, N represents the total number of reinforcement learning agents to be trained;
[0111] S32. Using a multi-node parallel architecture, each distributed computing node j calculates the local loss function L t (B j (B t ) and updates the multi-agent reinforcement learning model parameters using the gradient descent method:
[0112]
[0113] Among them, represents the parameter set of the multi-agent reinforcement learning model in the j-th node at the t-th iteration, η is the learning rate, represents the local gradient calculated based on the training batch B t ;
[0114] S33. Through the distributed communication mechanism, the updated model parameters of each node are globally aggregated to form the global multi-agent reinforcement learning model parameter set Θ t+1 :
[0115]
[0116] Among them, Φ(·) is the global parameter aggregation operator, and M is the total number of distributed computing nodes participating in the training;
[0117] S34. Repeat steps S31 to S33 until the global loss function L(Θ t ) satisfies the preset convergence condition, that is, L(Θ t ) ≤ ∈, where ∈ is the preset threshold, or reaches the preset maximum number of iterations T max , where T max is the preset maximum number of iterations, so as to obtain the preliminary training result Θ * of the multi-agent reinforcement learning model.
[0118] In this embodiment, step S4 specifically includes:
[0119] S41. During the training of the multi-agent reinforcement learning model, for each reinforcement learning agent Agent i 's reward function value Perform dynamic monitoring and evaluation, and calculate the variance of the change in the reward function in its most recent K iterations in real time
[0120]
[0121] where characterizes the degree of fluctuation of the reward function of the reinforcement learning agent Agent i in multiple consecutive iterations, is the average value of the reward function in the most recent K iterations, and T is the current iteration number;
[0122] S42. When the reinforcement learning agent Agent i meets the convergence condition of the variance of the change in the reward function at this time, where δ σ is a preset convergence determination threshold, and the bitter fish optimization algorithm automatically triggers a targeted adaptive dynamic perturbation mechanism, and adaptively adjusts the perturbation amplitude ζ according to the variance of the reward function i :
[0123]
[0124] where ζ i is the dynamic adaptive perturbation amplitude of the reinforcement learning agent Agent i , ζ max is the maximum preset value of the perturbation amplitude, and λ1 is the perturbation amplitude attenuation coefficient;
[0125] S43. Based on the perturbation amplitude ζ i combined with the bone density distribution information BD i corresponding to the candidate planting position of the reinforcement learning agent Agent density,i (x, y, z) and the occlusal mechanics parameter information F bite,i (x, y, z), calculate the bone density-occlusal mechanics coupling direction perturbation vector that adapts to the characteristics of the complex oral environment for intelligently guiding the reinforcement learning agent Agent i to optimize and explore in the direction where the bone density gradient changes significantly and the occlusal mechanics distribution is more balanced:
[0126]
[0127] where and respectively represent the gradient directions of the bone density distribution and the occlusal mechanics parameters at the candidate planting position corresponding to the agent Agent i , and × represents the cross product;
[0128] S44. For the reinforcement learning agent Agent iOriginal state parameter S i Perform a deep optimization perturbation operation to obtain the state parameter after optimization perturbation
[0129]
[0130] Among them, S i is the state parameter before perturbation, represents a normal distribution random vector with a mean of 0 and a covariance of the identity matrix I, represents the optimized perturbation vector calculated based on the bone density-occlusal mechanics coupling direction;
[0131] S45. Use the state parameter after optimization perturbation to replace the original state parameter S i and re-input it into the training process of the reinforcement learning model, and recalculate the reward function after perturbation Continuously iterate to optimize the agent's decision-making strategy until the global loss function L(Θ t ) reaches the preset convergence threshold ∈ to obtain the training result of the global optimal multi-agent reinforcement learning model.
[0132] In this embodiment, step S5 specifically includes:
[0133] S51. Import the multi-dental implant position and angle scheme * output by the training result Θ of the global optimal multi-agent reinforcement learning model into the biomechanical simulation analysis system, where is the three-dimensional spatial coordinate position of the i-th implant output by the reinforcement learning agent Agent i , and is the spatial implantation angle of the corresponding implant;
[0134] S52. Based on the patient's oral three-dimensional virtual oral model M oral3D , use the finite element analysis method to establish a multi-dental implant biomechanical finite element simulation model M FEA :
[0135]
[0136] Among them, Γ{·} is a finite element modeling operator, which is used to accurately implant the implant position and angle scheme in the three-dimensional virtual oral model M oral3D and automatically generate the corresponding meshed finite element simulation model;
[0137] S53. Apply load conditions based on the finite element simulation model M FEA and calculate the biomechanical stability evaluation function E after implanting the multi-dental implant bio Simulate the real occlusal behavior of the patient:
[0138]
[0139] Among them, V i represents the finite element analysis volume area of the alveolar bone around the i-th implant, and σ stress,i (x, y, z) represents the stress field distribution function of the local bone tissue area where the i-th implant is located, and E bio is the total stress index of biomechanical stability, characterizing the overall stress distribution of the implant scheme under the action of occlusal load;
[0140] S54. Evaluate the aesthetic effect of the implant scheme in the three-dimensional virtual oral model M oral3D and calculate the aesthetic evaluation index E esthetic :
[0141]
[0142] Among them, is the position coordinate of the i-th implant obtained by reinforcement learning optimization, represents the ideal implant position coordinate given by the tooth arrangement model M teeth ;
[0143] S55. Evaluate the occlusal balance of the implant scheme in the three-dimensional virtual oral model M oral3D and calculate the occlusal balance index E occlusal :
[0144]
[0145] Among them, represents the occlusal load borne by the local position after implanting the i-th implant, is the average value of the occlusal loads of all implants, and this index is used to quantify the balance degree of the implant scheme in the overall occlusal load distribution;
[0146] S56. Synthesize the total stress index E bio of biomechanical stability, the aesthetic evaluation index E esthetic and the occlusal balance index E occlusal , confirm that each index meets the preset clinical safety and effectiveness threshold conditions, and finally obtain a multi-tooth implant positioning planning scheme verified by biomechanical simulation and aesthetic evaluation.
[0147] Example 1:
[0148] At 9:30 am on March 15, 2024, a well-known oral medical center received a 57-year-old male patient, Mr. Wang. Due to long-term periodontal disease, Mr. Wang had four missing teeth on the left upper jaw (tooth positions 14, 15, 16, 17), which caused difficulty in eating. At the same time, due to alveolar bone resorption, his face collapsed, affecting his appearance. He hoped to repair the missing teeth through dental implants. However, due to the uneven alveolar bone density and complex occlusion relationship, after the doctor's evaluation, it was considered that the traditional manual planning method was difficult to ensure the long-term stability of the implant, and it was recommended to adopt a multi-tooth implant positioning and planning method based on distributed reinforcement learning to optimize the implant placement position.
[0149] At 10:00 am on the same day, Mr. Wang underwent CBCT (Cone Beam Computed Tomography) and MRI (Magnetic Resonance Imaging) examinations. The doctor obtained his oral three-dimensional image data. After the scan data was uploaded to the medical center's image analysis system, the system automatically performed medical image segmentation and extracted the patient's alveolar bone characteristics, bone density distribution, tooth arrangement information, and key parameters of the occlusal contact surface.
[0150] The system analysis results showed that:
[0151] The bone density in the areas of tooth positions 14 and 15 was relatively high (850 HU), and the implant conditions were good.
[0152] The bone density in the areas of tooth positions 16 and 17 was relatively low (520 HU), and there was a risk of stress concentration.
[0153] The patient's occlusion relationship was uneven, and the force on the left side was 20% higher than that on the right side. It was necessary to optimize the implant angle.
[0154] Based on the image data, the system automatically generated a three-dimensional virtual oral model, completed spatial registration and fusion calculations, and constructed Mr. Wang's personalized oral anatomical structure.
[0155] At 11:00 am, the system entered the reinforcement learning model training stage. The method of the present invention adopted a distributed multi-agent reinforcement learning framework, in which:
[0156] The position of each tooth to be implanted was regarded as an independent reinforcement learning agent, representing the implant positions of tooth positions 14, 15, 16, and 17 respectively.
[0157] The state space included three-dimensional coordinate information, bone density distribution, and occlusal mechanics data variables.
[0158] The action space was defined as the spatial displacement adjustment (Δd) and angle adjustment (Δθ) of the implant.
[0159] In the initial stage of training, the reinforcement learning model randomly explored multiple implant plans, but some agents fell into local optima. For example:
[0160] After the 5000th iteration of training, the reward function fluctuation of the implant position for the agent at tooth position 16 decreased, and it fell into a local optimum. There was still stress concentration in the area with low bone density.
[0161] To avoid the local optimum problem, the system introduced the bitter fish optimization algorithm and randomly perturbed the agents with an obvious convergence trend of the reward function during the training process. After the 8000th iteration, the agent at tooth position 16 adjusted the implant angle from 42° to 35° and moved backward by 0.8 mm at the same time, enabling the implant to obtain more stable support in the area with higher bone density.
[0162] After 48 hours of training and optimization, the reinforcement learning model finally converged, and the system generated the optimal multi-tooth implant positioning plan:
[0163] Tooth position 14: Implant depth 10.5 mm, angle 30°;
[0164] Tooth position 15: Implant depth 11.2 mm, angle 28°;
[0165] Tooth position 16: Implant depth 10.8 mm, angle 35° (after bitter fish optimization);
[0166] Tooth position 17: Implant depth 9.7 mm, angle 33°;
[0167] The system then conducted a finite element biomechanical simulation analysis to simulate the stress distribution under different biting forces after the patient's surgery. The analysis results showed that:
[0168] Compared with the traditional method, the maximum stress value decreased by 21% (from 85.3 MPa to 67.2 MPa), and the bone stress distribution was more uniform.
[0169] The balance of the biting force increased by 15%, reducing the impact of occlusal bias on the implant.
[0170] After the doctor confirmed the plan, the system automatically generated the surgical guide plate design file to ensure accurate implantation during the operation.
[0171] At 10:00 on March 17, 2024, Mr. Wang underwent an implant surgery. The doctor completed the implantation of four implants under local anesthesia according to the surgical guide plate generated by the reinforcement learning optimization. The operation process was smooth. The implantation time was reduced by 45% compared with the traditional method, the surgical wound was smaller, and there was no problem of implant offset during the operation.
[0172] Review at 3 months after surgery:
[0173] CBCT showed no obvious absorption of the alveolar bone around the implant, and the implant was stable.
[0174] The patient reported comfortable occlusion and normal recovery of chewing function.
[0175] Aesthetic score (0 - 10 points): 9.2 points (7.8 points for the traditional method).
[0176] To verify the applicability of the method of the present invention in different cases, we selected 50 similar cases in the past 6 months and compared the traditional method with the method of the present invention. The results are as follows:
[0177] Index Traditional method Method of the present invention Improvement amplitude Planning time 2.5 hours 23 minutes -85% Peak stress (MPa) 85.3 67.2 -21% Tooth arrangement coordination score (0 - 10 points) 7.8 9.2 +18% Occlusal balance (Eocclusal value) 1.35 1.15 -15% Implant stability score (0 - 10 points) 7.5 9.1 +21%
[0178] The method of the present invention has been significantly improved in terms of accuracy, stability, and efficiency. It not only improves the long-term stability of implants but also optimizes biomechanical properties and aesthetic effects. At the same time, it greatly shortens the planning time. In actual clinical applications, the implant plan optimized by reinforcement learning reduces the stress concentration phenomenon compared with the traditional method and reduces the risk of postoperative bone resorption and complications.
[0179] This embodiment proves the advantages of the multi-tooth implant positioning and planning method based on distributed reinforcement learning of the present invention compared with the traditional method in terms of efficiency, accuracy, and biomechanical optimization through real patient cases. Especially in cases with uneven bone density and complex occlusion relationships, this method can dynamically adjust the position and angle of implants to ensure long-term stability after surgery, providing an efficient and intelligent optimization plan for oral implant surgery.
[0180] The present invention establishes a multi-agent reinforcement learning model, takes the position of each tooth to be implanted as an independent agent, and defines its state space, action space, and reward function. During the reinforcement learning process, each agent can adaptively learn based on factors such as alveolar bone density and occlusal mechanical characteristics. Through an efficient training with a distributed computing architecture, it can dynamically adjust the spatial position and angle of the implant to ensure the best balance among biomechanical stability, aesthetic optimization, and occlusal balance, avoiding the uncertainty and inefficiency brought by manual adjustment by doctors.
[0181] The present invention introduces the bitter fish optimization algorithm, which dynamically perturbs the agents with an obvious convergence trend of the reward function during the reinforcement learning training process. By recalculating the reward function after perturbation, it prompts the reinforcement learning agents to jump out of the local optimal solution and improve the global optimization ability. The bitter fish optimization algorithm can adaptively adjust the perturbation amplitude according to the change of the variance of the reward function during the training process, enabling the reinforcement learning agents to better explore the best implantation plan of the implant in a complex bone density environment.
[0182] Combined with a biomechanical simulation analysis system, the present invention imports the multi-dental implant position and angle scheme output by the reinforcement learning agent into the biomechanical simulation system, and conducts biomechanical analysis and aesthetic evaluation based on the patient's three-dimensional oral model. During the simulation process, the finite element analysis method is used to calculate the stress distribution of the multi-dental implant under the action of the biting force, and through the aesthetic evaluation index and the occlusal balance index, it is ensured that the final positioning scheme meets the clinical safety and functional requirements, effectively improving the clinical feasibility of the implant planning scheme, reducing the risk of postoperative complications, and improving the long-term use stability of the implant.
[0183] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered within the protection scope of the present invention.
Claims
1. A multi-dental implant positioning and planning method based on distributed reinforcement learning, characterized in that It includes the following steps: S1. Obtain the medical image data of the patient, and establish a three-dimensional virtual oral cavity model of the patient's oral cavity based on the medical image data; S2. Based on the three-dimensional virtual oral cavity model, establish a multi-agent reinforcement learning model under a distributed reinforcement learning framework. Each candidate position of the dental implant to be implanted is used as an independent reinforcement learning agent, and the state space, action space, and reward function of each reinforcement learning agent are defined; S3. Based on the reinforcement learning agents, adopt a training strategy of centralized batch training and distributed execution in a distributed computing environment, and use a multi-node parallel architecture to train the multi-agent reinforcement learning model to obtain a preliminary training result of the multi-agent reinforcement learning model; S4. During the training process of the multi-agent reinforcement learning model, introduce the bitter fish optimization algorithm as a dynamic perturbation mechanism, randomly perturb the reinforcement learning agents with an obvious convergence trend of the reward function during training, and the perturbed state parameters re-participate in the reinforcement learning training process and recalculate the reward function. Continuously optimize the agent decision-making strategy until the global optimization condition is met, and obtain the globally optimal training result of the multi-agent reinforcement learning model; S5. Import the multi-dental implant position and angle scheme output by the training result of the multi-agent reinforcement learning model into the biomechanical simulation analysis system, and conduct biomechanical simulation analysis and aesthetic evaluation on the multi-dental implant scheme based on the patient's three-dimensional virtual oral cavity model to obtain a verified multi-dental implant positioning and planning scheme; S6. Generate a surgical guide design file for preoperative guidance based on the multi-dental implant positioning and planning scheme.
2. The multi-dental implant positioning and planning method based on distributed reinforcement learning according to claim 1, characterized in that Specifically included in step S1 are: S11. Obtain the oral medical image data of the patient, including computed tomography data and magnetic resonance imaging data; S12. Perform spatial registration and data fusion on the oral medical image data respectively to obtain the fused oral medical image dataset D fusion ; S13. Based on the oral medical image data, extract the alveolar bone region characteristics, tooth arrangement characteristics, bone density distribution characteristics, and occlusal contact surface characteristics in the patient's oral cavity through a medical image segmentation algorithm; S14. Based on the extracted feature information, establish a model M including the alveolar bone region model bone , the tooth arrangement model M teeth , the bone density distribution model BD density (x, y, z) and the occlusal mechanics model F bite , where the alveolar bone region model is constructed based on the alveolar bone boundary characteristics of the patient's individual, the tooth arrangement model is constructed based on the tooth row position and direction characteristics of the patient's individual, the bone density distribution model is constructed by converting the Hounsfield unit value HU(x, y, z) into a density value, and the occlusal mechanics model is constructed from the occlusal contact area and bite force distribution data; S15. Based on the alveolar bone region model, tooth arrangement model, bone density distribution model, and occlusal mechanics model, perform three-dimensional spatial registration and fusion to construct a three-dimensional virtual oral model M of the patient's oral cavity oral3D : where, μ BD and σ BD represent the mean and standard deviation of the bone density data respectively, α and β are weight factors that respectively adjust the contribution degrees of the alveolar bone region model and the tooth alignment model in the global model construction, Λ(·) and Ω(·) are transformation functions for normalizing and preprocessing the alveolar bone region model M bone , the tooth alignment model M teeth respectively, tanh(·) is the hyperbolic tangent function, BD density (x, y, z) - μ BD represents the deviation between the local bone density and the average bone density, σ BD is the normalization factor, sin(·) is the sine function, κ is the scaling constant of the occlusal mechanics data, and Ψ{·} is the three-dimensional space registration and fusion operator.
3. The multi-dental implant positioning and planning method based on distributed reinforcement learning according to claim 2, characterized in that Specifically included in step S2 are: S21. Based on the three-dimensional virtual oral model M of the patient's oral cavity oral3D Establish a multi-reinforcement learning agent reinforcement learning model under a distributed reinforcement learning framework, where each candidate position for dental implantation is used as an independent reinforcement learning agent Agent i , where i represents the i-th candidate implantation position; S22. Define each reinforcement learning agent i 's state space S i : S i = {P i (x, y, z), BD density,i (x, y, z), F bite,i (x, y, z)}; Among them, P i (x, y, z) is the spatial coordinate information, representing the reinforcement learning agent i The corresponding candidate implant position in the three-dimensional virtual oral model M oral3D The three-dimensional spatial coordinates in, BD density,i (x, y, z) is the local bone density distribution information, used to characterize the local bone density environment of the candidate implant position, F bite,i (x, y, z) is the occlusal mechanical parameter information, used to characterize the occlusal mechanical distribution characteristics borne by the candidate implant position; S23. Define each reinforcement learning agent i 's action space A i : A i = {Δd i , Δθ i}; Among them, Δd i represents three-dimensional vector information of the adjustment direction of the spatial coordinates of each candidate implantation position, characterizing the adjustment of the spatial coordinates of the candidate implantation position in the three-dimensional virtual oral cavity model M oral3D ; Δθ i represents the implant angle change information of each candidate implantation position, characterizing the adjustment of the implant in the spatial implantation direction; S24. Define each reinforcement learning agent i 's reward function R i , the reward function R i is composed of a biomechanical stability index R bio,i , an aesthetic optimization index R esth,i and a bite balance index R occlu,i together: R i = γ1·R bio,i + γ2·R esth,i + γ3·R occlu,i ; Among them, R bio,i is a biomechanical stability index used to evaluate the stability of the local bone stress distribution after implanting an implant at the corresponding candidate position in the three-dimensional virtual oral model M i by the reinforcement learning agent Agent oral3D ; R esth,i is an aesthetic optimization index used to evaluate the spatial coordination between the candidate position in the three-dimensional virtual oral model M i and the tooth arrangement model M oral3D by the reinforcement learning agent Agent teeth ; R occlu,i is a bite balance index used to evaluate the balance of the bite mechanical distribution after implanting an implant at the corresponding candidate position in the three-dimensional virtual oral model M i by the reinforcement learning agent Agent oral3D .
4. A multi-dental implant positioning and planning method based on distributed reinforcement learning according to claim 1, characterized in that Specifically included in step S3 are: S31. In a distributed computing environment, according to the centralized batch training strategy, all reinforcement learning agents Agent i 's state space S i , action space A i and reward function R i information to form a training batch B t : B t = {(S i , A i , R i ) | i = 1, 2, …, N}; Where N represents the total number of reinforcement learning agents to be trained; S32. Using a multi-node parallel architecture, each distributed computing node j calculates the local loss function L t for the training batch B j (B t ) and updates the multi-agent reinforcement learning model parameters using the gradient descent method: Among them, represents the set of parameters of the multi-agent reinforcement learning model in the j-th node at the t-th iteration, η is the learning rate, represents the local gradient calculated based on the training batch B t obtained by calculation; S33. Aggregate the updated model parameters of each node through a distributed communication mechanism to form a global multi-agent reinforcement learning model parameter set Θ t+1 : Where Φ(·) is a global parameter aggregation operator, and M is the total number of distributed computing nodes participating in training; S34. Repeat steps S31 to S33 until the global loss function L(Θ t ) satisfies the preset convergence condition, that is, L(Θ t ) ≤ ∈, where ∈ is the preset threshold, or reaches the preset maximum number of iterations T max , where T max is the preset maximum number of iterations, so as to obtain the preliminary training result Θ * of the multi-agent reinforcement learning model.
5. A multi-dental implant positioning and planning method based on distributed reinforcement learning according to claim 4, characterized in that Specifically included in step S4 are: S41. During the training process of the multi-agent reinforcement learning model, for each reinforcement learning agent Agent i the value of the reward function is dynamically monitored and evaluated, and the variance of the change in the reward function in its most recent K iterations is calculated in real time Among them, characterizes the reinforcement learning agent Agent i the degree of fluctuation of the reward function in multiple consecutive iterations, is the average value of the reward function in the most recent K iterations, and T is the current iteration number; S42. When the reinforcement learning agent Agent i satisfies the convergence condition of the variance of the reward function At this time, where δ σ is a preset convergence determination threshold, the bitter fish optimization algorithm automatically triggers a targeted adaptive dynamic perturbation mechanism, and adaptively adjusts the perturbation amplitude ζ according to the variance of the reward function i : Among them, ζ i is the reinforcement learning agent Agent i 's dynamic adaptive perturbation amplitude, ζ max is the maximum preset value of the perturbation amplitude, and λ1 is the perturbation amplitude decay coefficient; S43. Based on the perturbation amplitude ζ i Combined with the reinforcement learning agent Agent i The bone density distribution information BD corresponding to the candidate planting positions density,i (x, y, z) and the occlusal mechanics parameter information F bite,i (x, y, z), calculate the bone density-occlusal mechanics coupling direction perturbation vector that adapts to the characteristics of the complex oral environment Used to intelligently guide the reinforcement learning agent Agent i Optimize and explore in the direction where the bone density gradient changes significantly and the occlusal mechanics distribution is more balanced: Among them, and respectively represent the gradient directions of the bone density distribution and the occlusal mechanical parameters at the candidate planting positions corresponding to the intelligent agent Agent i × represents the cross product; S44. For the reinforcement learning agent Agent i with the original state parameter S i perform a deep optimization perturbation operation to obtain the state parameter after optimization perturbation Among them, S i is the state parameter before perturbation, represents a normal distribution random vector with a mean of 0 and a covariance of the identity matrix I, represents an optimized perturbation vector calculated based on the bone density-occlusal mechanics coupling direction; S45. Replace the original state parameter S with the optimized and perturbed state parameter i Re - input it into the training process of the reinforcement learning model and recalculate the perturbed reward function Continuously iterate to optimize the agent's decision - making strategy until the global loss function L(Θ t ) reaches the preset convergence threshold ∈, and obtain the training result of the global optimal multi - agent reinforcement learning model. 6. The multi-dental implant positioning and planning method based on distributed reinforcement learning according to claim 5, wherein Specifically included in step S5 are: S51. Import the multi-dental implant position and angle scheme * output by the training result Θ of the global optimal multi-agent reinforcement learning model into the biomechanical simulation analysis system, where is the three-dimensional space coordinate position of the i-th implant output by the reinforcement learning agent Agent i , and is the spatial implantation angle of the corresponding implant. S52. Based on the three-dimensional virtual oral model M of the patient's oral cavity oral3D , a finite element simulation model M of the biomechanics of multiple dental implants is established by using the finite element analysis method FEA :[[]]END]] where Γ{·} is a finite element modeling operator for accurately implanting the implant position and angle scheme in the three-dimensional virtual oral model M oral3D and automatically generating the corresponding meshed finite element simulation model; S53. Based on the finite element simulation model M FEA Apply loading conditions and evaluate the biomechanical stability evaluation function E after implanting multiple dental implants through calculation bio Simulate the patient's real occlusal behavior: Among them, V i represents the finite element analysis volume region of the alveolar bone around the i-th implant, and σ stress,i (x, y, z) represents the stress field distribution function of the local bone tissue region where the i-th implant is located. E bio is the total stress index of biomechanical stability, characterizing the overall stress distribution of the implant scheme under occlusal load; S54. Evaluate the aesthetic effect of the implant plan in the three-dimensional virtual oral model M oral3D and calculate the aesthetic evaluation index E esthetic : Among them, is the position coordinate of the i-th implant optimized by reinforcement learning, represents the ideal implant position coordinates given by the tooth arrangement model M teeth ; S55. Evaluate the occlusal balance of the implant plan in the three-dimensional virtual oral model M oral3D and calculate the occlusal balance index E occlusal : Among them, represents the occlusal load borne by the local position after the implantation of the i-th implant, is the average value of the occlusal loads of all implants, and this index is used to quantify the balance degree of the implant scheme in the overall occlusal load distribution; S56. Comprehensive Biomechanical Stability Total Stress Index E bio , Aesthetic Evaluation Index E esthetic and Occlusal Balance Index E occlusal , confirm that each index meets the preset clinical safety and effectiveness threshold conditions, and finally obtain a multi-tooth implant positioning planning scheme verified by biomechanical simulation and aesthetic evaluation.
Citation Information
Cited By
Retention method of single-crown screw retention implant for dental implant repair
CN121081146A
Patient oral cavity state self-service monitoring method based on deep learning
CN121337278A
Multifunctional intelligent steaming oven control system and intelligent steaming oven
CN121857883A