Diffusion model dental crown design method based on multi-view consistency and shape prior

Through the diffusion model crown design method based on multi-view consistency and shape priors, the problems of insufficient personalization and cumbersome operation of traditional crown design methods are solved, and efficient, personalized and high-quality crown design is achieved.

CN120107476APending Publication Date: 2025-06-06ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510182084.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Due to insufficient personalization and cumbersome operation, traditional crown design methods are difficult to meet the needs of modern dental science, and the design process is time-consuming and labor-intensive.

Method used

Using a diffusion model crown design method based on multi-view consistency and shape priors, the three-dimensional dental neural radiation field model is trained by iterating the basic framework updated by the data set, multi-view image editing and neural radiation field fine-tuning are carried out to ensure that the generated dental model has 3D consistency and no collision conflicts.

Benefits of technology

It improves the efficiency and quality of crown design, enhances the personalization and practicality of the design, and improves the functionality and accuracy of the design through the multi-view correlation calculation module and the implicit model of denoising and diffusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107476A_ABST
    Figure CN120107476A_ABST
Patent Text Reader

Abstract

The invention discloses a diffusion model dental crown design method based on multi-view consistency and shape prior, and relates to the technical field of computer vision, and the method comprises the following steps: training a three-dimensional dental nerve radiation field model through an input Shining3D dental crown design data set, and further rendering to obtain a multi-view RGB image; the hidden space representation of the multi-view image is obtained through an encoder, and the measurement of the consistency between views is determined through a multi-view correlation calculation module; modeling is carried out on joint image distribution of multiple views, consistency measurement, obtained through a multi-view correlation calculation module, between the views is introduced into joint probability distribution modeling, and multi-view images are edited through a multi-view fractional distillation sampling method. According to the method, the basic framework updated by the iterative data set is adopted, the model framework is divided into two stages of multi-view editing and nerve radiation field fine tuning, missing teeth with 3D consistency and without collision conflict are automatically generated, and the efficiency and quality of dental crown design are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a multi-view Figure 1 Diffusion model crown design method based on consistency and shape prior. Background Art

[0002] 3D model editing is an important task in computer vision. Traditional methods often rely on manual modeling or rule-based algorithms, which are inefficient and difficult to meet complex tasks and personalized needs. In recent years, with the development of diffusion model technology, 3D editing methods based on diffusion models have gradually emerged. They achieve efficient and controllable 3D model editing through neural network learning and simulation of data distribution. This technology has been widely used in film production, game development, computer-aided design, and crown design.

[0003] Crown design is the core link in dental restoration. By designing a reasonable crown, patients can restore missing or damaged teeth in their mouths to improve the chewing experience. However, traditional crown design methods are difficult to meet the needs of modern dentistry due to lack of personalization and cumbersome operations. Although computer-aided design and manufacturing such as CAD and CAM technologies have improved design efficiency, a large amount of manual interaction is still required in the design process, and the design quality is limited by the doctor's technical level and subjective judgment. In addition, since the condition of each patient is different, the crown design needs to be highly customized for the individual patient, so this task is still time-consuming and labor-intensive.

[0004] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention

[0005] In view of the problems in the related art, the present invention proposes a multi-view Figure 1 A diffusion model crown design method based on consistency and shape prior is proposed to overcome the above-mentioned technical problems existing in the existing related technologies.

[0006] To this end, the specific technical solution adopted by the present invention is as follows:

[0007] Based on multi-view Figure 1 A diffusion model crown design method based on consistency and shape prior is proposed. The method adopts the basic framework of iterative dataset update, and each iteration process includes the following steps:

[0008] S1. Train the 3D dental nerve radiation field model by inputting the Shining3D crown design dataset, and further render it to obtain multi-view RGB images. Obtain the latent space representation of the multi-view images through the encoder, and determine the consistency measure between views through the multi-view correlation calculation module.

[0009] S2. Model the joint image distribution of multiple views, introduce the consistency measure between views obtained by the multi-view correlation calculation module into the joint probability distribution modeling, and edit the multi-view images through the multi-view score distillation sampling method;

[0010] S3, learning the prior knowledge of tooth shape through the denoising diffusion implicit model, and fine-tuning the three-dimensional teeth based on the model to avoid collision conflicts in the tooth model;

[0011] S4. For the constructed three-dimensional tooth model, conduct age-based chewing force assessment and perform stress analysis on the crown.

[0012] As a preferred embodiment, the S1 comprises the following steps:

[0013] S11. Train the 3D tooth nerve radiation field model according to the input Shining3D crown design dataset, obtain the initial nerve radiation field parameters, and render 40 RGB images of a given viewing angle on each sample through the rendering function;

[0014] S12, the rendered RGB image is processed by the encoder to obtain the corresponding latent space representation, and the target viewpoint image with the same content as the initial viewpoint image is generated using the Wonder3D model through the viewpoint conditional image conversion in the multi-view correlation calculation module:

[0015] Among them, the viewpoint conditional image conversion is implemented using the Wonder3D model. This method generates a target viewpoint image with consistent content with the initial viewpoint image, which is defined as follows:

[0016]

[0017] Among them, z i,I is the image x i,e The latent space representation at the i-th iteration is, is the transformation matrix of the camera matrix of the target viewpoint v and the reference viewpoint e;

[0018] S13. Determine a measure of consistency between views by using an energy function in a multi-view correlation calculation module.

[0019] As a preferred implementation, the S11 comprises the following steps:

[0020] S111, randomly selected 3D scanned dental plaster models from dental hospital patients to form a crown design dataset, where each plaster model contains 1416 meshes;

[0021] S112. The dental crown design dataset includes a training set, a test set, and a validation set. The training set, the test set, and the validation set contain 1150, 133, and 133 samples, respectively. Each sample is subjected to instance segmentation to classify the teeth, and a tooth is selected. In each sample, 60 images of the selected tooth are cropped, and 40 images contain the selected tooth.

[0022] S113, training the 3D tooth nerve radiation field model according to the input Shining3D crown design data set, and obtaining the initial nerve radiation field parameter θ 0 , and on each sample through the rendering function g(θ 0 ,c) Render 40 RGB images of a given view angle x 0 .

[0023] As a preferred implementation, in S13, the energy function is used to model the consistency between different views, and the L2 reconstruction loss is used to calculate the consistency measure between the reference viewpoint image and the target viewpoint image. The specific algorithm formula is:

[0024]

[0025] in, is the latent space representation of the set of rendered images of the 3D dental nerve radiation field at the i-th iteration The latent space representation of the reference view selected in , is the latent space representation of the target view, and t represents the time step of the diffusion model.

[0026] As a preferred embodiment, S2 comprises the following steps:

[0027] S21, modeling the joint image distribution of multiple views, and introducing the consistency measure between views obtained by the multi-view correlation calculation module into the joint probability distribution modeling. Specifically, the joint image distribution of multiple views is modeled and expressed as The algorithm formula is:

[0028]

[0029] in, represents the latent space representation of the image set rendered by the neural radiance field at the i-th iteration, represents the corresponding camera matrix set, c I and c T They represent reference images and text prompts used to control image editing, respectively, and V represents a viewpoint set;

[0030] S22, use a pre-trained 2D diffusion model to represent the latent space z obtained from S1 i,tNoise prediction network Predict the noise at time t, and then minimize the image distribution sampled from the 2D pre-trained diffusion model And from the rendering function g(θ i ,c) The image distribution of the rendered image obtained in the forward diffusion process at time t The KL divergence between is used to optimize the parameter θ of the neural radiation field. The algorithm formula is:

[0031]

[0032] S23. Edit the multi-view image by a multi-view score distillation sampling method, wherein the multi-view score distillation function formula is:

[0033]

[0034] As a preferred embodiment, the multi-view score distillation sampling method is improved from the original score distillation sampling method:

[0035] where the original fractional distillation sampling defines the initial neural radiation field by the parameter θ 0 Indicates that, in the i-th iteration, through the rendering function g(θ i ,c) Get RGB image x i,0 ;

[0036] Where c is the camera parameter, the RGB image x i,0 After being processed by the encoder ε(.), the corresponding latent space representation z is obtained i,0 =ε(x i,0 ), using a pre-trained 2D diffusion model to represent the input latent space z i,t Through its noise prediction network Predict the noise at time t, where c i,I and c T Respectively represent reference images and textual hints used to control image editing;

[0037] Finally, we minimize the image distribution p sampled from the 2D pre-trained diffusion model. i,t (z i,t c I ,c T ) and from the rendering function g(θ i ,c) The image distribution of the rendered image obtained in the forward diffusion process at time t The KL divergence between them is used to optimize θ, and the formula is as follows:

[0038]

[0039] The model is trained using the following fractional distillation function formula:

[0040]

[0041] Among them, θ i is the neural radiation field model parameter at the i-th iteration, ω(i,t) is the hyperparameter controlling the loss weight, c is the camera parameter of the current view, is the noise predicted by the noise prediction network, ε i is the real noise.

[0042] As a preferred embodiment, the tooth shape prior knowledge is obtained by learning a denoising diffusion implicit model, which includes the following steps:

[0043] The tooth model in the dental design dataset is voxelized into cubes, where the value x of each cube 0 Indicates that the cube is free space x 0 = -1 or occupied space x 0 =1;

[0044] Based on the training steps of the denoising diffusion model, real noise ε is added to the cube in the forward process, and the noise is predicted using the U-Net noise prediction network with parameter π in the reverse process. The real noise ε and the predicted noise The mean square error of is used as the loss function of the 3D voxel denoising diffusion model to update the noise prediction network. The algorithm formula is:

[0045]

[0046] The three-dimensional tooth nerve radiation field is fine-tuned using a diffusion model that contains prior information about the tooth shape to ensure that the generated tooth model is aesthetically pleasing and conflict-free. During the fine-tuning process, the density of the cube b in the three-dimensional tooth nerve radiation field at the i-th iteration is defined as σ i,b , and binarize it into density d by comparing it with the threshold ρ i,b,t = ±1;

[0047] D i,b,t Input the pre-trained 3D voxel denoising diffusion model to get the denoised cubemap density When the predicted density is inconsistent with the density generated by the neural radiation field, the neural radiation field is optimized by shape prior loss, and the algorithm formula is:

[0048]

[0049] in, is the binarization result of the denoised prediction value of the cube at the i-th iteration, and ω is a hyperparameter, which indicates the upper limit of the occupied space density;

[0050] Fine-tuning of the 3D dental nerve radiation field via appearance loss L app LPIPS loss L LPIPS and shape prior loss L prior The specific formula is:

[0051] L all =L app +λ 1 L LPIPS +λ 2 L prior ;

[0052] Among them, the appearance loss L app The LPIPS loss L is calculated by the pixel-wise L1 norm between the local region in the rendered image and the corresponding patched region. LPIPS It is used to measure the perceptual similarity between the rendering result and the target, 1 and λ 2 It is a coefficient used to balance various losses.

[0053] As a preferred embodiment, the S4 comprises the following steps:

[0054] S41, for the constructed three-dimensional tooth model, perform a crown usability assessment based on the preset age groups, and perform a stress analysis on the currently constructed three-dimensional tooth model in combination with the crown material;

[0055] S42. Based on the stress distribution results, the fatigue life of the crown in long-term use is predicted by Miner linear cumulative damage theory. The algorithm formula is:

[0056]

[0057] Among them, n d represents the dth stress cycle number, N f,d Represents the fatigue life under the corresponding stress level.

[0058] As a preferred implementation, the S41 includes the following steps:

[0059] S411, dividing into different age groups and setting average masticatory force standard values ​​based on clinical statistical data, wherein the age groups include 12-18 years old, 19-60 years old, and over 60 years old;

[0060] S412. By using the finite element analysis method and combining the patient's age, the average chewing force standard value of the corresponding age group is applied as a load to the three-dimensional tooth model. Based on Hooke's law in elastic mechanics, within the elastic range, stress is proportional to strain. Based on the corresponding elastic modulus, Poisson's ratio, and allowable stress defined by the crown material, the stress distribution and deformation of the crown under the chewing force are calculated by using finite element software.

[0061] S413. Based on the maximum principal stress obtained by finite element analysis, the three-dimensional tooth model is evaluated in combination with the current allowable stress. When the maximum principal stress is less than or equal to the allowable stress, it means that the crown strength of the current three-dimensional tooth model is qualified. When the maximum principal stress is greater than the allowable stress, it means that the crown strength of the current three-dimensional tooth model is insufficient. It is necessary to further adjust the crown geometry through the topology optimization tool and re-perform finite element analysis until the maximum principal stress is less than or equal to the allowable stress.

[0062] The beneficial effects of the present invention are:

[0063] 1. The present invention proposes a novel diffusion model crown design framework, adopts a basic framework for iterative data set updating, and divides the model framework into two stages: multi-view editing and neural radiation field fine-tuning. It automatically generates missing teeth with 3D consistency and no collision conflicts, so as to improve the efficiency and quality of crown design and enhance practicality.

[0064] 2. The present invention sets a multi-view correlation calculation module to explore the distribution relationship between multiple views based on the energy function, combines this module to improve the editing consistency of the diffusion model, and uses this module to extend the fractional distillation sampling to multi-view fractional distillation sampling to achieve 2D consistent editing, thereby improving the efficiency and effect of neural radiation field optimization and enhancing the functionality of the method;

[0065] 3. The present invention enhances the accuracy of the generated tooth model by using a denoising diffusion implicit model to explicitly learn the prior knowledge of tooth shape in the occupied space and introduces shape prior loss to enhance the consistency between the tooth prior and the generated teeth;

[0066] 4. The present invention performs stress analysis on the constructed three-dimensional tooth model based on age, further evaluates the theoretical strength of the crown of the generated three-dimensional tooth model, assists the user to adjust the model to meet actual clinical use needs, and enhances the applicability of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0068] Figure 1 According to an embodiment of the present invention, Figure 1 Flow chart of the diffusion model crown design method based on consistency and shape prior;

[0069] Figure 2 According to an embodiment of the present invention, Figure 1 Schematic diagram of the framework of the diffusion model crown design method based on consistency and shape prior;

[0070] Figure 3 It is a comparison chart of the results of the present invention according to an embodiment of the present invention, the DreamFusion method, and the NeRFiller method on the Shining3D crown design data set. DETAILED DESCRIPTION

[0071] To further illustrate each embodiment, the present invention provides drawings, which are part of the disclosure of the present invention and are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, ordinary technicians in the field should be able to understand other possible implementations and advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0072] According to an embodiment of the present invention, a multi-view based Figure 1 Diffusion model crown design method based on consistency and shape prior.

[0073] The present invention is further described with reference to the accompanying drawings and specific embodiments:

[0074] Embodiment 1:

[0075] like Figure 1-2 As shown, according to an embodiment of the present invention, Figure 1 The diffusion model crown design method based on consistency and shape prior adopts the basic framework of iterative dataset updating, including the following steps:

[0076] S1. Train the 3D dental nerve radiation field model by inputting the Shining3D crown design dataset, and further render it to obtain multi-view RGB images. Obtain the latent space representation of the multi-view images through the encoder, and determine the consistency measure between views through the multi-view correlation calculation module.

[0077] It should be noted that the present invention adopts a basic framework of iterative dataset updating and divides the model framework into two stages: multi-view editing and neural radiation field fine-tuning.

[0078] S11. Train the 3D tooth nerve radiation field model according to the input Shining3D crown design dataset, obtain the initial nerve radiation field parameters, and render 40 RGB images of a given viewing angle on each sample through the rendering function;

[0079] S111, randomly selected 3D scanned dental plaster models from dental hospital patients to form a crown design dataset, where each plaster model contains 1416 meshes;

[0080] S112. The dental crown design dataset includes a training set, a test set, and a validation set. The training set, the test set, and the validation set contain 1150, 133, and 133 samples, respectively. Each sample is subjected to instance segmentation to classify the teeth, and a tooth is selected. In each sample, 60 images of the selected tooth are cropped, and 40 images contain the selected tooth.

[0081] It should be noted that instance segmentation can not only identify different target categories in an image, but also distinguish different individual instances belonging to the same category. In the scene of tooth images, instance segmentation can accurately separate each tooth from the entire oral image and mark the specific position and outline of the tooth.

[0082] S113, training the 3D tooth nerve radiation field model according to the input Shining3D crown design data set, and obtaining the initial nerve radiation field parameter θ 0 , and on each sample through the rendering function g(θ 0 ,c) Render 40 RGB images of a given view angle x 0 ;

[0083] S12, the rendered RGB image is processed by the encoder to obtain the corresponding latent space representation, and the target viewpoint image with the same content as the initial viewpoint image is generated using the Wonder3D model through the viewpoint conditional image conversion in the multi-view correlation calculation module:

[0084] Among them, the viewpoint conditional image conversion is implemented using the Wonder3D model. This method generates a target viewpoint image with consistent content with the initial viewpoint image, which is defined as follows:

[0085]

[0086] Among them, z i,I is the image x i,e The latent space representation at the i-th iteration is, is the transformation matrix of the camera matrix of the target viewpoint v and the reference viewpoint e;

[0087] S13, determining a measure of consistency between views by using an energy function in a multi-view correlation calculation module;

[0088] It should be noted that the rendered RGB image x i,0 After being processed by the encoder ε(.), the corresponding latent space representation z is obtained i,0 =ε(x i,0 ), and then use the Wonder3D model to generate a target view image with consistent content with the initial view image. Specifically, when the Wonder3D model is integrated with the framework of the present invention, a random view is first selected from the 3D rendered image set as a reference view image, and then the camera transformation matrix and the reference view between the reference view image and the target view image are input into the Wonder3D model to obtain the target view image, and the consistency measure between the views is determined by calculating the reconstruction loss between the generated target view image and the reference view;

[0089] The rendering image set at the i-th iteration is defined as The latent space representation of the rendered image set is Viewpoint {1,...,v,...,V} and camera parameter matrix Then, the designed energy function is used to model the consistency between different views, where the above Wonder3D model is used to transform from the reference view to the target view. The specific algorithm formula is:

[0090]

[0091] in, is the latent space representation of the set of rendered images of the 3D dental nerve radiation field at the i-th iteration The latent space representation of the reference view selected in , is the latent space representation of the target view, t represents the time step of the diffusion model, t is no more than 50, is the transformation matrix of the camera matrix of the target viewpoint v and the reference viewpoint e, W(.) is the viewpoint conditional image transformation model Wonder3D;

[0092] After obtaining the target viewpoint image through Wonder3D, we choose to use L2 reconstruction loss to calculate the consistency measure between the reference viewpoint image and the target viewpoint image. L2 reconstruction loss is based on the calculation method of mean square error, which is particularly suitable for the reconstruction task of images, signals or other numerical data. Finally, the latent space representation of the rendered image set of the neural radiation field is The consistency measures of all viewpoints in are added together to get the total consistency measure value.

[0093] Embodiment 2:

[0094] S2. Model the joint image distribution of multiple views, introduce the consistency measure between views obtained by the multi-view correlation calculation module into the joint probability distribution modeling, and edit the multi-view images through the multi-view score distillation sampling method;

[0095] Specifically, the joint image distribution of multiple views is modeled and expressed as At the same time, the consistency measure between views obtained by the multi-view correlation calculation module is introduced into the joint probability distribution modeling, as shown below:

[0096]

[0097] in, represents the latent space representation of the image set rendered by the neural radiance field at the i-th iteration, represents the corresponding camera matrix set, c I and c T They represent the reference image and text prompt used to control image editing. The text prompt is the category of the teeth to be generated, such as "generate a incisor". V represents the viewpoint set. At the same time, the energy function is converted to Mapped to a non-negative value such that It complies with the basic property of probability distribution, that is, the probability must be non-negative;

[0098] At the same time, the original energy function It is expressed as: when the consistency between the reference viewpoint image and the target viewpoint image is stronger, the consistency metric value is smaller, that is, the consistency between images is inversely proportional to the final consistency metric value. Here, an exponential function is used to convert the multi-view image into Figure 1 Strong and weak relationships and multi-viewing Figure 1 The consistency metric value changes from the original inverse relationship to a direct relationship, and its value is always between 0 and 1, which is convenient for modeling joint image distribution;

[0099] Then a pre-trained 2D diffusion model is used to represent the latent space z obtained from S1 i,t Noise prediction network Predict the noise at time t, and then minimize the image distribution sampled from the 2D pre-trained diffusion model And from the rendering function g(θ i ,c) The image distribution of the rendered image obtained in the forward diffusion process at time t The KL divergence between is used to optimize the parameter θ of the neural radiation field. The algorithm formula is:

[0100]

[0101] And the model is trained through the multi-view score distillation function formula:

[0102]

[0103] Among them, θ i is the neural radiation field model parameter at the iteration, ω(i,t) is the hyperparameter that controls the loss weight, c is the camera parameter of the current view, is the noise predicted by the noise prediction network, The neural radiation field model uses the Adam gradient descent method to generate real noise. The learning rate of the geometric encoding module is 0.01, the learning rate of the density module and the feature module is 0.001, the optimization times are 10,000 times, and the image is rendered with a resolution of 64x64 for the first 3,000 times, and then with a resolution of 256x256. The batch size of the model is 4.

[0104] Embodiment 3:

[0105] S3, learning the prior knowledge of tooth shape through the denoising diffusion implicit model, and fine-tuning the three-dimensional teeth based on the model to avoid collision conflicts in the tooth model;

[0106] Specifically, during the training process of the 3D voxel denoising diffusion model, the 3D crown mesh in the Shining3D crown design dataset is voxelized into 32 3 The cube is randomly rotated and expanded to adjust the thickness of the cube. The value x of each cube 0 Indicates that the cube is free space x 0 = -1 or occupied space x 0 =1, free space refers to the area not occupied by any object, represented as blank area, and occupied space refers to the area occupied by objects, such as the tooth surface and internal structure;

[0107] Then follow the training steps of the denoising diffusion model, add real noise ε to the cube in the forward process, and predict the noise using the U-Net noise prediction network with parameters . U-Net has three downsampling layers, and the number of channels in each downsampling layer is gradually doubled. The initial learning rate of the 3D voxel denoising diffusion model is set to 0.00005, and the starting step value of the noise scheduling hyperparameter is 0.00085, which is linearly increased to 0.0120. This can ensure a smooth transition of noise addition and make the generation process more stable and continuous. The model batch size is 4, and the Adam gradient descent method is used. A total of 20,000 time steps are trained;

[0108] The real noise ε and the predicted noise The mean square error of is used as the loss function of the 3D voxel denoising diffusion model to update the noise prediction network. The formula is as follows:

[0109]

[0110] In the inference process of the 3D voxel denoising diffusion model, the density of cube b in the 3D dental nerve radiation field at the i-th iteration is defined as σ i,b , and binarize it to density d by comparing it with the threshold ρ = 0.01 i,b,t = ±1, d i,b,t Input the pre-trained 3D voxel denoising diffusion model to get the denoised cubemap density When the predicted density is inconsistent with the density generated by the neural radiation field, the neural radiation field is optimized by shape prior loss, as follows:

[0111]

[0112] in, It is the binarization result of the denoised prediction value of the cube at the i-th iteration. ω is a hyperparameter, which represents the upper limit of the occupied space density. It is used to limit the maximum optimization range of the NeRF output density value, thereby ensuring that the model does not over-predict the occupied space. The density of free space is usually close to 0, and the area slightly above 0 can be more safely regarded as occupied space. Therefore, ω=0.01 is usually set. A small ω value can completely screen out the density of free space during voxelization. At the same time, it limits the density optimization range of occupied space when fine-tuning the neural radiation field using shape prior loss, making the model more stable.

[0113] When u i,b = 1, which means that the model predicts that the cube is a free space. The shape prior loss can directly affect the density value σ of NeRF. i,b This penalty will reduce the density of free space and ensure that NeRF does not generate unnecessary objects in free space. i,b = 0, which means the model predicts that the cube is occupied space. At this time, if the density σ i,b <ω, then by max(ω-σ i,b ,0) Increase the density and ensure that the density of the occupied space is large enough. If the density σ i,b ≥ω, then the shape prior loss is 0, indicating that there is no need to further optimize this area. This can prevent the density of the occupied space from increasing infinitely, resulting in unreasonable geometric shapes or artifacts generated by the model;

[0114] Fine-tuning of the 3D dental nerve radiation field via appearance loss L app LPIPS loss LLPIPS and shape prior loss L prior The specific formula is:

[0115] L all =L app +λ 1 L LPIPS +λ 2 L prior ;

[0116] Among them, the appearance loss L app The LPIPS loss L is calculated by the pixel-wise L1 norm between the local region in the rendered image and the corresponding patched region. LPIPS It is used to measure the perceptual similarity between the rendering result and the target, 1 =1.5 and λ 2 =3.0 is the coefficient used to balance various losses;

[0117] To evaluate the model, the Frichett embedding distance, peak signal-to-noise ratio, and average learned perceptual patch similarity are calculated for the generated multi-view images and the corresponding views rendered by the fine-tuned neural radiance field. These evaluation criteria are defined as follows:

[0118]

[0119] Among them, μ x and μ y are the mean of edited image and rendered image respectively, ∑ x and∑ y is the covariance matrix of the mean of the edited image and the eigenvector of the rendered image;

[0120]

[0121] in, is the square of the maximum possible pixel value of the edited image I, and R is the rendered image of the neural radiance field fine-tuning model;

[0122]

[0123] Here, x and y refer to the edited image and the rendered image, respectively, l refers to the number of layers in AlexNet, and H l and W l Refers to the height and width of the feature map of layer l, h and w refer to the index of the height and width of the feature map, respectively. l refers to the weights of layer l, and φ() refers to the pre-trained AlexNet.

[0124] Embodiment 4:

[0125] S4. For the constructed three-dimensional tooth model, conduct age-based chewing force assessment and stress analysis on the crown;

[0126] S41, for the constructed three-dimensional tooth model, perform a crown usability assessment based on the preset age groups, and perform a stress analysis on the currently constructed three-dimensional tooth model in combination with the crown material;

[0127] S411, dividing into different age groups and setting average chewing force standard values ​​based on clinical statistical data, wherein the age groups include 12-18 years old, 19-60 years old, and over 60 years old;

[0128] S412. By using the finite element analysis method and combining the patient's age, the average chewing force standard value of the corresponding age group is applied as a load to the three-dimensional tooth model. Based on Hooke's law in elastic mechanics, within the elastic range, stress is proportional to strain. Based on the corresponding elastic modulus, Poisson's ratio, and allowable stress defined by the crown material, the stress distribution and deformation of the crown under the chewing force are calculated by using finite element software.

[0129] It should be noted that the allowable stress requires conducting mechanical property tests such as tension and compression on the crown material to determine the material's ultimate stress, including yield strength or tensile strength. During the experiment, the standard specimen is loaded through a material testing machine and the stress-strain curve is recorded to determine the ultimate stress value and set the allowable stress.

[0130] S413, based on the maximum principal stress obtained by finite element analysis, the three-dimensional tooth model is evaluated in combination with the current allowable stress. When the maximum principal stress is less than or equal to the allowable stress, it means that the crown strength of the current three-dimensional tooth model is qualified. When the maximum principal stress is greater than the allowable stress, it means that the crown strength of the current three-dimensional tooth model is insufficient, and it is necessary to further adjust the crown geometry through the topology optimization tool, and re-perform the finite element analysis until the maximum principal stress is less than or equal to the allowable stress;

[0131] S42. Based on the stress distribution results, the fatigue life of the crown in long-term use is predicted by Miner's linear cumulative damage theory. The algorithm formula is:

[0132]

[0133] Among them, n d represents the dth stress cycle number, N f,d Represents the fatigue life under the corresponding stress level.

[0134] It should be noted that by analyzing the fatigue life under the corresponding stress level, auxiliary medical personnel can further adjust the crown design parameters, including thickness and curvature, based on the fatigue life analysis results to extend the service life.

[0135] In order to verify the multi-view Figure 1 The generation effect of the diffusion model crown design method based on consistency and shape prior is compared with the results of the DreamFusion method and the NeRFiller method on the Shining3D crown design dataset. The comparison results are shown in Figure 2. Figure 3 As shown, the results of the present invention are better than those of the DreamFusion method and the NeRFiller method.

[0136] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A diffusion model crown design method based on multi-view consistency and shape prior, characterized in that: This method adopts the basic framework of iterative dataset update, and each iteration process includes the following steps: S1. Train the 3D dental nerve radiation field model by inputting the Shining3D crown design dataset, and further render it to obtain multi-view RGB images. Obtain the latent space representation of the multi-view images through the encoder, and determine the consistency measure between views through the multi-view correlation calculation module. S2. Model the joint image distribution of multiple views, introduce the consistency measure between views obtained by the multi-view correlation calculation module into the joint probability distribution modeling, and edit the multi-view images through the multi-view score distillation sampling method; S3, learning the prior knowledge of tooth shape through the denoising diffusion implicit model, and fine-tuning the three-dimensional teeth based on the model to avoid collision conflicts in the tooth model; S4. For the constructed three-dimensional tooth model, conduct age-based chewing force assessment and perform stress analysis on the crown.

2. The method for designing a dental crown based on a diffusion model based on multi-view consistency and shape prior according to claim 1, characterized in that: The S1 comprises the following steps: S11. Train the 3D tooth nerve radiation field model according to the input Shining3D crown design dataset, obtain the initial nerve radiation field parameters, and render 40 RGB images of a given viewing angle on each sample through the rendering function; S12, the rendered RGB image is processed by the encoder to obtain the corresponding latent space representation, and the target viewpoint image with the same content as the initial viewpoint image is generated using the Wonder3D model through the viewpoint conditional image conversion in the multi-view correlation calculation module: Among them, the viewpoint conditional image conversion is implemented using the Wonder3D model. This method generates a target viewpoint image with consistent content with the initial viewpoint image, which is defined as follows: Among them, z i,I is the image x i,e The latent space representation at the i-th iteration is, is the transformation matrix of the camera matrix of the target viewpoint v and the reference viewpoint e; S13. Determine a measure of consistency between views by using an energy function in a multi-view correlation calculation module.

3. The method for designing a dental crown based on a diffusion model based on multi-view consistency and shape prior according to claim 2, characterized in that: The S11 comprises the following steps: S111, randomly selected 3D scanned dental plaster models from dental hospital patients to form a crown design dataset, where each plaster model contains 1416 meshes; S112. The dental crown design dataset includes a training set, a test set, and a validation set. The training set, the test set, and the validation set contain 1150, 133, and 133 samples, respectively. Each sample is subjected to instance segmentation to classify the teeth, and a tooth is selected. In each sample, 60 images of the selected tooth are cropped, and 40 images contain the selected tooth. S113. Train the 3D tooth nerve radiation field model according to the input Shining3D crown design data set to obtain the initial nerve radiation field parameter θ0, and render 40 RGB images x0 of a given viewing angle on each sample through the rendering function g(θ0,c).

4. The method for designing a dental crown based on a diffusion model based on multi-view consistency and shape prior according to claim 2, characterized in that: In S13, the energy function is used to model the consistency between different views, and the L2 reconstruction loss is used to calculate the consistency measure between the reference viewpoint image and the target viewpoint image. The specific algorithm formula is: in, is the latent space representation of the set of rendered images of the 3D dental nerve radiation field at the i-th iteration The latent space representation of the reference view selected in , is the latent space representation of the target view, and t represents the time step of the diffusion model.

5. The method for designing a dental crown based on a diffusion model using multi-view consistency and shape prior according to claim 2, characterized in that: The S2 comprises the following steps: S21, modeling the joint image distribution of multiple views, and introducing the consistency measure between views obtained by the multi-view correlation calculation module into the joint probability distribution modeling. Specifically, the joint image distribution of multiple views is modeled and expressed as The algorithm formula is: in, represents the latent space representation of the image set rendered by the neural radiance field at the i-th iteration, represents the corresponding camera matrix set, c I and c T They represent reference images and text prompts used to control image editing, respectively, and V represents a viewpoint set; S22, use a pre-trained 2D diffusion model to represent the latent space z obtained from S1 i,t Noise prediction network Predict the noise at time t, and then minimize the image distribution sampled from the 2D pre-trained diffusion model And from the rendering function g(θ i ,c) The image distribution of the rendered image obtained in the forward diffusion process at time t The KL divergence between is used to optimize the parameter θ of the neural radiation field. The algorithm formula is: S23. Edit the multi-view image by a multi-view score distillation sampling method, wherein the multi-view score distillation function formula is:

6. The method for designing a dental crown based on a diffusion model using multi-view consistency and shape prior according to claim 5, characterized in that: The multi-view score distillation sampling method is improved from the original score distillation sampling method: Among them, the original fractional distillation sampling defines the initial neural radiation field represented by the parameter θ0, and in the i-th iteration, it is transformed by the rendering function g(θ i ,c) Get RGB image x i,0 ; Where c is the camera parameter, the RGB image x i,0 After being processed by the encoder ε(.), the corresponding latent space representation z is obtained i,0 =ε(x i,0 ), using a pre-trained 2D diffusion model to represent the input latent space z i,t Through its noise prediction network Predict the noise at time t, where c i,I and c T Respectively represent reference images and textual prompts used to control image editing; Finally, we minimize the image distribution p sampled from the 2D pre-trained diffusion model. i,t (z i,t c I ,c T ) and from the rendering function g(θ i ,c) The image distribution of the rendered image obtained in the forward diffusion process at time t The KL divergence between them is used to optimize θ, and the formula is as follows: The model is trained using the following score distillation function formula: Among them, θ i is the neural radiation field model parameter at the i-th iteration, ω(i,t) is the hyperparameter controlling the loss weight, c is the camera parameter of the current view, is the noise predicted by the noise prediction network, ε i is the real noise.

7. The method for designing a dental crown based on a diffusion model based on multi-view consistency and shape prior according to claim 1, characterized in that: The tooth shape prior knowledge is obtained by learning a denoising diffusion implicit model, which includes the following steps: The tooth model in the dental design dataset is voxelized into cubes, where the value x0 of each cube indicates that the cube is a free space x0=-1 or an occupied space x0=1; Based on the training steps of the denoising diffusion model, real noise ε is added to the cube in the forward process, and the noise is predicted using the U-Net noise prediction network with parameter π in the reverse process. The real noise ε and the predicted noise The mean square error of is used as the loss function of the 3D voxel denoising diffusion model to update the noise prediction network. The algorithm formula is: The three-dimensional tooth nerve radiation field is fine-tuned using a diffusion model that contains prior information about the tooth shape to ensure that the generated tooth model is aesthetically pleasing and conflict-free. During the fine-tuning process, the density of the cube b in the three-dimensional tooth nerve radiation field at the i-th iteration is defined as σ i,b , and binarize it into density d by comparing it with the threshold ρ i,b,t = ±1; D i,b,t Input the pre-trained 3D voxel denoising diffusion model to get the denoised cubemap density When the predicted density is inconsistent with the density generated by the neural radiation field, the neural radiation field is optimized by shape prior loss, and the algorithm formula is: in, is the binarization result of the denoised prediction value of the cube at the i-th iteration, and ω is a hyperparameter, which indicates the upper limit of the occupied space density; Fine-tuning of the 3D dental nerve radiation field via appearance loss L app , LPIPS loss L LPIPS and shape prior loss L prior The specific formula is: L all =L app +λ1L LPIPS +λ2L prior ; Among them, the appearance loss L app The LPIPS loss L is calculated by the pixel-wise L1 norm between the local region in the rendered image and the corresponding patched region. LPIPS It is used to measure the perceptual similarity between the rendering result and the target, and λ1 and λ2 are coefficients used to balance the various losses.

8. The method for designing a dental crown based on a diffusion model using multi-view consistency and shape prior according to claim 1, characterized in that: The S4 comprises the following steps: S41, for the constructed three-dimensional tooth model, perform a crown usability assessment based on the preset age groups, and perform a stress analysis on the currently constructed three-dimensional tooth model in combination with the crown material; S42. Based on the stress distribution results, the fatigue life of the crown in long-term use is predicted by Miner linear cumulative damage theory. The algorithm formula is: Among them, n d represents the dth stress cycle number, N f,d Represents the fatigue life under the corresponding stress level.

9. The method for designing a dental crown based on a diffusion model using multi-view consistency and shape prior according to claim 8, characterized in that: The S41 comprises the following steps: S411, dividing into different age groups and setting average masticatory force standard values ​​based on clinical statistical data, wherein the age groups include 12-18 years old, 19-60 years old, and over 60 years old; S412. By using the finite element analysis method and combining the patient's age, the average chewing force standard value of the corresponding age group is applied as a load to the three-dimensional tooth model. Based on Hooke's law in elastic mechanics, within the elastic range, stress is proportional to strain. Based on the corresponding elastic modulus, Poisson's ratio, and allowable stress defined by the crown material, the stress distribution and deformation of the crown under the chewing force are calculated by using finite element software. S413. Based on the maximum principal stress obtained by finite element analysis, the three-dimensional tooth model is evaluated in combination with the current allowable stress. When the maximum principal stress is less than or equal to the allowable stress, it means that the crown strength of the current three-dimensional tooth model is qualified. When the maximum principal stress is greater than the allowable stress, it means that the crown strength of the current three-dimensional tooth model is insufficient. It is necessary to further adjust the crown geometry through the topology optimization tool and re-perform finite element analysis until the maximum principal stress is less than or equal to the allowable stress.

Citation Information

Cited By

  • Video monitoring system for safe production

    CN121147606A

  • A safe production video monitoring system

    CN121147606B