Method and device for training dental mold deformation model

By extracting the feature tensor of the initial tooth model and inputting a preset network model, the dental mold deformation model is solved, and the problem of dental mold deformation relies on manual operation in the prior art is realized, and the automation deformation of the dental model is improved, and efficiency and quality are improved.

CN112884885BActive Publication Date: 2025-05-23SHINING 3D TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110287715.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-17
Publication Date
2025-05-23
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

In the prior art, 3D dental mold deformation mainly relies on manual operation, resulting in low efficiency, high cost, unreliable quality, and lack of automated solutions.

Method used

By obtaining the sample data of the initial tooth model and the corresponding target deformation model, extracting feature tensors and inputting a preset network model, obtaining the predicted deformation model, and then automatically deforming the tooth model by optimizing the network model to obtain the tooth model deformation model.

Benefits of technology

The automatic conversion of the initial dental model into a dental model that meets specific product requirements has improved efficiency and quality and reduced costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112884885B_ABST
    Figure CN112884885B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method and device for training a dental mold deformation model, which relates to the field of three-dimensional deformation technology. The method includes: obtaining sample data, the sample data includes multiple initial tooth models obtained by scanning the oral cavity and corresponding target deformation models obtained by manually processing each initial tooth model; obtaining a feature tensor corresponding to each initial tooth model, each element of the feature tensor corresponding to each initial tooth model is the TSDF value of each voxel in the cubic space where each initial tooth model is located; inputting the feature tensor corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model; according to the target deformation model and the predicted deformation model corresponding to each initial tooth model, the preset network model is optimized to obtain a dental mold deformation model. The embodiment of the present invention is used to obtain a dental mold deformation model that can automatically convert an initial tooth model into a tooth model that meets specific product requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional deformation technology, and in particular to a method and device for training a dental mold deformation model. Background Art

[0002] Dental digitalization technology aims to perform 3D modeling of teeth to obtain digital tooth models, thereby enabling subsequent processing and personalized customization.

[0003] In general, the tooth model actually used in the end is not the initial tooth model obtained by scanning the oral cavity and performing 3D reconstruction, but the tooth model that meets the specific product requirements is obtained by further processing the initial tooth model based on specific product requirements. Among them, the process of processing the initial tooth model based on specific product requirements to obtain a tooth model that meets the specific product requirements is called 3D tooth model deformation. At present, 3D tooth model deformation is generally done manually. That is, people manually process the initial tooth model based on specific product requirements so that the initial tooth model meets the specific product requirements. However, manual completion of 3D tooth model deformation has many disadvantages such as low efficiency, high cost, and unreliable quality. Therefore, how to automatically convert the initial tooth model into a tooth model that meets the specific product requirements has become a problem that needs to be solved urgently in this field. Summary of the invention

[0004] In view of this, the present invention provides a method and device for training a tooth mold deformation model, which is used to automatically convert an initial tooth model into a tooth model that meets specific product requirements.

[0005] In order to achieve the above purpose, the embodiment of the present invention provides the following technical solutions:

[0006] In a first aspect, an embodiment of the present invention provides a method for training a dental mold deformation model, comprising:

[0007] Acquire sample data, wherein the sample data includes a plurality of initial tooth models acquired by scanning the oral cavity and a target deformation model corresponding to each initial tooth model obtained by manually processing each initial tooth model;

[0008] Acquire a feature tensor corresponding to each initial tooth model, wherein each element of the feature tensor corresponding to each initial tooth model is a truncated signed distance function TSDF value of each voxel in the cubic space where each initial tooth model is located;

[0009] Inputting the feature tensors corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model;

[0010] According to the target deformation model and the predicted deformation model corresponding to each initial tooth model, the preset network model is optimized to obtain the tooth mold deformation model.

[0011] As an optional implementation of an embodiment of the present invention, the preset network model includes: an encoder component composed of multiple encoders in series structure, a self-attention component, a feature transfer component, a multi-scale analysis component, and a decoder component composed of multiple decoders in series structure; the input of the encoder component is the input of the preset network model, and the output of the encoder component is the input of the self-attention component; the output of the self-attention component is the input of the feature transfer component; the output of the feature transfer component is the input of the multi-scale analysis component, the output of the multi-scale analysis component is the input of the decoder component, and the output of the decoder component is the output of the preset network model;

[0012] Among them, the self-attention component is used to extract non-local information from the feature tensor output by the encoder component to obtain the environmental feature tensor; the feature transfer component processes the output of the self-attention component and transfers the processing result to the multi-scale analysis component; the multi-scale analysis component is used to extract the feature tensor output by the feature transfer component at multiple scales.

[0013] As an optional implementation of an embodiment of the present invention, the encoder component includes three encoders in a serial structure, each encoder includes a residual unit and a downsampling unit; the residual unit of each encoder is used to perform a convolution operation on the input of the residual unit through three convolution layers in a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the downsampling unit of each encoder is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor;

[0014] The input of each residual unit is the input of the encoder to which it belongs, the input of each downsampling unit is the residual unit output of the encoder to which it belongs, the output of each downsampling unit is the output of the encoder to which it belongs, the input of the first encoder is the input of the encoder component, the output of the third encoder is the output of the encoder component, and the inputs of the second and third encoders are the outputs of the first and second encoders, respectively.

[0015] As an optional implementation of an embodiment of the present invention, the self-attention component includes: a residual unit, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a first dot product unit, a second dot product unit, a first addition unit and a second addition unit; the residual unit is used to perform a convolution operation on the input of the residual unit through three convolutional layers of a serial structure and perform a summation operation on the convolution result of the convolution operation and the input of the residual unit, the first dot product unit and the second dot product unit are used to perform a dot product operation on the input feature tensor, and the first summation unit and the second summation unit are used to perform a summation operation on the input feature tensor;

[0016] The input of the residual unit is the output of the encoder component, and the output of the residual unit is the input of the first convolutional layer; the output of the first convolutional layer is the input of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer; the input of the first dot convolution unit is the output of the second convolutional layer and the output of the third convolutional layer; the input of the second dot convolution unit is the output of the first dot convolution unit and the output of the fourth convolutional layer; the input of the fifth convolutional layer is the output of the second dot convolution unit; the input of the first summation unit is the output of the fifth convolutional layer and the output of the first convolutional layer; the input of the sixth convolutional layer is the output of the first summation unit; the input of the second summation unit is the output of the sixth convolutional layer and the output of the residual unit, and the output of the second summation unit is the output of the self-attention component.

[0017] As an optional implementation of the embodiment of the present invention, the feature transfer component includes a downsampling unit and a residual unit;

[0018] The downsampling unit is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor, and the residual unit is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit;

[0019] The input of the downsampling unit is the output of the self-attention component, the output of the downsampling unit is the input of the residual unit, and the output of the residual unit is the output of the feature transfer component.

[0020] As an optional implementation manner of an embodiment of the present invention, the multi-scale analysis component includes: a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, and a splicing unit; the dilation rates of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer are all different; the splicing unit is configured to perform a splicing operation on the input feature tensor;

[0021] The inputs of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer are all the output of the feature transfer component. The inputs of the splicing unit are the output of the feature transfer component, the output of the seventh convolutional layer, the output of the eighth convolutional layer, the output of the ninth convolutional layer, and the output of the tenth convolutional layer. The input of the eleventh convolutional layer is the output of the splicing unit, and the input of the twelfth convolutional layer is the output of the eleventh convolutional layer; the output of the twelfth convolutional layer is the output of the multi-scale analysis component.

[0022] As an optional implementation manner of an embodiment of the present invention, the decoder component includes four decoders in series; each decoder includes: an upsampling unit, a fusion unit, and a residual unit; the residual unit of each decoder is configured to perform a convolution operation on the input of the residual unit through three convolutional layers in series and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit. The fusion unit of each decoder is configured to perform a fusion operation on the input feature tensor; the upsampling unit of each decoder is configured to upsample the input feature tensor into an output feature tensor with half the number of channels of the input feature tensor and twice the length, width, and height of the input feature tensor.

[0023] The input of the upsampling unit of the first decoder of the decoder component is the output of the multi-scale analysis component. The input of the fusion unit of the first decoder is the output of the upsampling unit of the first decoder and the output of the self-attention component; the input of the residual unit of the first decoder is the output of the fusion unit of the first decoder and the output of the self-attention component; the inputs of the upsampling units of the second decoder, the third decoder, and the fourth decoder of the decoder component are all the output of the previous decoder. The inputs of the fusion units of the second decoder, the third decoder, and the fourth decoder of the decoder component are respectively the output of the residual unit of the corresponding encoder and the output of the upsampling unit of the decoder to which they belong. The inputs of the residual units of the second decoder, the third decoder, and the fourth decoder of the decoder component are respectively the output of the residual unit of the corresponding encoder and the output of the fusion unit of the decoder to which they belong.

[0024] As an optional implementation of the embodiment of the present invention, the fusion unit of each decoder includes: a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, a third addition unit, a fourth addition unit, a third dot product unit, and a fourth dot product unit; the third addition unit and the fourth addition unit are used to perform an addition operation on the input, and the third dot product unit and the fourth dot product unit are used to perform a dot product operation on the input;

[0025] The inputs of the thirteenth convolutional layer and the fourteenth convolutional layer of the fusion unit of the first decoder of the decoder component are the output of the upsampling unit of the first decoder and the output of the self-attention component respectively. The inputs of the fusion units of the second decoder, the third decoder and the fourth decoder of the decoder component are the output of the upsampling unit of the decoder to which they belong and the output of the residual unit of the corresponding encoder. The input of the third addition unit is the output of the thirteenth convolutional layer and the output of the fourteenth convolutional layer. The input of the fifteenth convolutional layer is the output of the third addition unit. The input of the third dot product unit is the output of the thirteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth dot product unit is the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth addition unit is the output of the third dot product unit and the output of the fourth dot product unit. The output of the fourth addition unit is the output of the fusion unit to which it belongs.

[0026] As an optional implementation of the embodiment of the present invention, the method of optimizing the preset network model to obtain the tooth mold deformation model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model includes:

[0027] Constructing a loss function, and optimizing the preset network model to obtain a tooth mold deformation model according to the loss function, the target deformation model corresponding to each initial tooth model, and the predicted deformation model;

[0028] Wherein, the loss function includes:

[0029]

[0030] Among them, alpha is a constant, out i The data is obtained by processing the output of the multi-scale analysis component and the output of each decoder of the decoder component in sequence, seg is the intermediate supervision signal, and mean() is the averaging function.

[0031] In a second aspect, an embodiment of the present invention provides a device for establishing a deformable dental model, comprising:

[0032] A sample acquisition unit for acquiring sample data, where the sample data includes a plurality of initial tooth models obtained by scanning the oral cavity and target deformation models corresponding to the respective initial tooth models obtained by artificially processing each initial tooth model;

[0033] A preprocessing unit for obtaining a feature tensor corresponding to each initial tooth model, where each element of the feature tensor corresponding to each initial tooth model is the truncated signed distance function value (TSDF value) of each voxel in the cubic space where the initial tooth model is located;

[0034] A prediction unit for inputting the feature tensor corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model;

[0035] An optimization unit for optimizing the preset network model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model to obtain a dental model deformation model.

[0036] As an optional implementation manner of an embodiment of the present invention, the preset network model includes: an encoder component composed of a plurality of serially connected encoders, a self-attention component, a feature transfer component, a multi-scale analysis component, and a decoder component composed of a plurality of serially connected decoders; the input of the encoder component is the input of the preset network model, and the output of the encoder component is the input of the self-attention component; the output of the self-attention component is the input of the feature transfer component; the output of the feature transfer component is the input of the multi-scale analysis component, the output of the multi-scale analysis component is the input of the decoder component, and the output of the decoder component is the output of the preset network model;

[0037] Among them, the self-attention component is used to perform non-local information extraction on the feature tensor output by the encoder component to obtain an environmental feature tensor; the feature transfer component processes the output of the self-attention component and transfers the processing result to the multi-scale analysis component; the multi-scale analysis component is used to extract feature tensors of the feature tensor output by the feature transfer component at multiple scales.

[0038] As an optional implementation manner of an embodiment of the present invention, the encoder component includes three serially connected encoders, and each encoder includes a residual unit and a downsampling unit; the residual unit of each encoder is used to perform a convolution operation on the input of the residual unit through three serially connected convolutional layers and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the downsampling unit of each encoder is used to downsample the input feature tensor into an output feature tensor with a number of channels twice that of the input feature tensor and a length, width, and height that are half of the length, width, and height of the input feature tensor;

[0039] The input of each residual unit is the input of the encoder to which it belongs, the input of each downsampling unit is the residual unit output of the encoder to which it belongs, the output of each downsampling unit is the output of the encoder to which it belongs, the input of the first encoder is the input of the encoder component, the output of the third encoder is the output of the encoder component, and the inputs of the second and third encoders are the outputs of the first and second encoders, respectively.

[0040] As an optional implementation of an embodiment of the present invention, the self-attention component includes: a residual unit, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a first dot product unit, a second dot product unit, a first addition unit and a second addition unit; the residual unit is used to perform a convolution operation on the input of the residual unit through three convolutional layers of a serial structure and perform a summation operation on the convolution result of the convolution operation and the input of the residual unit, the first dot product unit and the second dot product unit are used to perform a dot product operation on the input feature tensor, and the first summation unit and the second summation unit are used to perform a summation operation on the input feature tensor;

[0041] The input of the residual unit is the output of the encoder component, and the output of the residual unit is the input of the first convolutional layer; the output of the first convolutional layer is the input of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer; the input of the first dot convolution unit is the output of the second convolutional layer and the output of the third convolutional layer; the input of the second dot convolution unit is the output of the first dot convolution unit and the output of the fourth convolutional layer; the input of the fifth convolutional layer is the output of the second dot convolution unit; the input of the first summation unit is the output of the fifth convolutional layer and the output of the first convolutional layer; the input of the sixth convolutional layer is the output of the first summation unit; the input of the second summation unit is the output of the sixth convolutional layer and the output of the residual unit, and the output of the second summation unit is the output of the self-attention component.

[0042] As an optional implementation of the embodiment of the present invention, the feature transfer component includes a downsampling unit and a residual unit;

[0043] The downsampling unit is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor, and the residual unit is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit;

[0044] The input of the downsampling unit is the output of the self-attention component, the output of the downsampling unit is the input of the residual unit, and the output of the residual unit is the output of the feature transfer component.

[0045] As an optional implementation of the embodiment of the present invention, the multi-scale analysis component includes: a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer and a splicing unit; the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer have different expansion rates; the splicing unit is used to perform a splicing operation on the input feature tensor;

[0046] The inputs of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer are all the outputs of the feature transfer component, the input of the splicing unit is the output of the feature transfer component, the output of the seventh convolutional layer, the output of the eighth convolutional layer, the output of the ninth convolutional layer, and the output of the tenth convolutional layer, the input of the eleventh convolutional layer is the output of the splicing unit, the input of the twelfth convolutional layer is the output of the eleventh convolutional layer; the output of the twelfth convolutional layer is the output of the multi-scale analysis component.

[0047] As an optional implementation of the embodiment of the present invention, the decoder component includes four decoders of a serial structure; each decoder includes: an upsampling unit, a fusion unit and a residual unit; the residual unit of each decoder is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the fusion unit of each decoder is used to perform a fusion operation on the input feature tensor; the upsampling unit of each decoder is used to upsample the input feature tensor to an output feature tensor whose number of channels is half of the number of channels of the input feature tensor and whose length, width and height are twice the length, width and height of the input feature tensor;

[0048] The input of the upsampling unit of the first decoder of the decoder component is the output of the multi-scale analysis component, the input of the fusion unit of the first decoder is the output of the upsampling unit of the first decoder and the output of the self-attention component; the input of the residual unit of the first decoder is the output of the fusion unit of the first decoder and the output of the self-attention component; the input of the upsampling unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are all the output of the previous decoder, the input of the fusion unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are the output of the residual unit of the corresponding encoder and the output of the upsampling unit of the decoder to which they belong, and the input of the residual unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are the output of the residual unit of the corresponding encoder and the output of the fusion unit of the decoder to which they belong.

[0049] As an optional implementation of the embodiment of the present invention, the fusion unit of each decoder includes: a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, a third addition unit, a fourth addition unit, a third dot product unit, and a fourth dot product unit; the third addition unit and the fourth addition unit are used to perform an addition operation on the input, and the third dot product unit and the fourth dot product unit are used to perform a dot product operation on the input;

[0050] The inputs of the thirteenth convolutional layer and the fourteenth convolutional layer of the fusion unit of the first decoder of the decoder component are the output of the upsampling unit of the first decoder and the output of the self-attention component respectively. The inputs of the fusion units of the second decoder, the third decoder and the fourth decoder of the decoder component are the output of the upsampling unit of the decoder to which they belong and the output of the residual unit of the corresponding encoder. The input of the third addition unit is the output of the thirteenth convolutional layer and the output of the fourteenth convolutional layer. The input of the fifteenth convolutional layer is the output of the third addition unit. The input of the third dot product unit is the output of the thirteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth dot product unit is the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth addition unit is the output of the third dot product unit and the output of the fourth dot product unit. The output of the fourth addition unit is the output of the fusion unit to which it belongs.

[0051] As an optional implementation of the embodiment of the present invention, the optimization unit is specifically used to construct a loss function, and optimize the preset network model to obtain a tooth mold deformation model according to the loss function, the target deformation model corresponding to each initial tooth model, and the predicted deformation model;

[0052] Wherein, the loss function includes:

[0053]

[0054] Among them, alpha is a constant, out i The data is obtained by processing the output of the multi-scale analysis component and the output of each decoder of the decoder component in sequence, seg is the intermediate supervision signal, and mean() is the averaging function.

[0055] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory and a processor, the memory being used to store a computer program; the processor being used to execute the method for training a dental mold deformation model described in the first aspect or any optional embodiment of the first aspect when calling the computer program.

[0056] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for training a dental mold deformation model described in the first aspect or any optional implementation manner of the first aspect is implemented.

[0057] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the method for training a dental mold deformation model as described in the first aspect or any optional implementation manner of the first aspect.

[0058] The method for training a dental mold deformation model provided by an embodiment of the present invention first obtains sample data including multiple initial tooth models and target deformation models corresponding to each initial tooth model, then obtains the TSDF value of each voxel in the cubic space where each initial tooth model is located as an element of the feature tensor corresponding to each initial tooth model, and then inputs the feature tensor corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model, and finally optimizes the preset network model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model to obtain a dental mold deformation model. Since the method for training a dental mold deformation model provided by an embodiment of the present invention can obtain a dental mold deformation model, and obtain a dental model that meets the requirements of a specific product based on the dental mold deformation model for the initial model that needs to be deformed, the dental mold deformation model obtained by the method for training a dental mold deformation model provided by an embodiment of the present invention can automatically convert the initial tooth model into a tooth model that meets the requirements of a specific product. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0061] Figure 1 A flow chart of a method for training a dental mold deformation model provided by an embodiment of the present invention;

[0062] Figure 2 An architectural diagram of a preset network model provided in an embodiment of the present invention;

[0063] Figure 3 A schematic diagram of the structure of an encoder assembly provided by an embodiment of the present invention;

[0064] Figure 4 A schematic diagram of the structure of a self-attention component provided in an embodiment of the present invention;

[0065] Figure 5 A schematic diagram of the structure of a feature transfer component provided in an embodiment of the present invention;

[0066] Figure 6 A schematic diagram of the structure of a multi-scale analysis component provided by an embodiment of the present invention;

[0067] Figure 7 A schematic diagram of the structure of a decoder component provided by an embodiment of the present invention;

[0068] Figure 8 A schematic diagram of the structure of a fusion unit provided in an embodiment of the present invention;

[0069] Fig. 9 A structural diagram of a device for training a dental mold deformation model provided by an embodiment of the present invention;

[0070] Fig.10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0071] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0072] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.

[0073] In the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way. In addition, in the description of the embodiments of the present invention, unless otherwise specified, the meaning of "multiple" refers to two or more.

[0074] The execution subject of the method for training a dental mold deformation model provided in the embodiment of the present invention may be a device for establishing a dental mold deformation model. The device for establishing a dental mold deformation model may be a terminal device such as a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart watch, a smart bracelet, or the terminal device may also be other types of terminal devices, and the embodiment of the present invention does not limit the type of the terminal device.

[0075] The embodiment of the present invention provides a method for training a dental mold deformation model, referring to Figure 1 As shown, the method for training a dental mold deformation model includes the following steps S11 to S14:

[0076] S11. Obtain sample data.

[0077] The sample data includes a plurality of initial tooth models obtained by scanning the oral cavity and a target deformation model corresponding to each initial tooth model obtained by manually processing each initial tooth model.

[0078] Specifically, oral scans of multiple users can be performed and 3D reconstruction can be performed to obtain an initial tooth model for each user. Then, part of the gingival area in the tooth model of each room can be manually removed based on specific needs, and the model can be re-scanned to obtain the target deformation model corresponding to each initial tooth model.

[0079] S12. Obtain the feature tensor corresponding to each initial tooth model.

[0080] The feature tensor corresponding to each initial tooth model is a Truncated Signed Distance Function (TSDF) value of each voxel in the cubic space where each initial tooth model is located.

[0081] Specifically, obtaining the feature tensor corresponding to each initial tooth model may include: firstly establishing a square outer bounding box of the tooth model as the cubic space where each initial tooth model is located, then voxelizing the cubic space where each initial tooth model is located, and finally using the truncated signed distance function (TSDF) to calculate the distance from each voxelization to the surface of the initial tooth model as the TSDF value of each voxel. i ,y i ,z i )=0 means that the voxel is on the surface of the tooth model, TSDF(x i ,y i ,z i )>0 means that the voxel is outside the tooth model, TSDF(x i ,y i ,z i )<0 indicates that the voxel is inside the tooth model.

[0082] S13, inputting the feature tensor corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model.

[0083] That is, a preset network model for generating a dental mold deformation model is pre-established, and the feature tensors corresponding to each initial tooth model in the sample data are input into the preset network model, and the corresponding output is obtained as the predicted deformation model corresponding to the initial tooth model.

[0084] S14. According to the target deformation model and the predicted deformation model corresponding to each initial tooth model, the preset network model is optimized to obtain a tooth mold deformation model.

[0085] The method for training a dental mold deformation model provided in an embodiment of the present invention first obtains sample data including multiple initial tooth models and target deformation models corresponding to each initial tooth model, then obtains the TSDF value of each voxel in the cubic space where each initial tooth model is located as an element of the feature tensor corresponding to each initial tooth model, and then inputs the feature tensor corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model, and finally optimizes the preset network model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model to obtain a dental mold deformation model. Since the method for training a dental mold deformation model provided in an embodiment of the present invention can obtain a dental mold deformation model, and obtain a dental model that meets the requirements of a specific product based on the initial model that needs to be deformed according to the dental mold deformation model, the dental mold deformation model obtained by the method for training a dental mold deformation model provided in an embodiment of the present invention can automatically convert the initial tooth model into a tooth model that meets the requirements of a specific product. The preset network model in the above embodiment is described in detail below.

[0086] Reference Figure 2 As shown, the preset network model in the embodiment of the present invention includes:

[0087] An encoder component 21 consisting of multiple encoders in series, a self-attention component 22, a feature transfer component 23, a multi-scale analysis component 24, and a decoder component 25 consisting of multiple decoders in series.

[0088] The input of the encoder component 21 is the input of the preset network model, and the output of the encoder component 21 is the input of the self-attention component 22; the output of the self-attention component 22 is the input of the feature transfer component 23; the output of the feature transfer component 23 is the input of the multi-scale analysis component 24, and the output of the multi-scale analysis component 24 is the input of the decoder component 25, and the output of the decoder component 25 is the output of the preset network model.

[0089] Among them, the self-attention component 22 is used to extract non-local information from the feature tensor output by the encoder component 21 to obtain the environmental feature tensor; the feature transfer component processes the output of the self-attention component and transfers the processing result to the multi-scale analysis component; the multi-scale analysis component 24 is used to extract the feature tensor output by the feature transfer component 23 at multiple scales.

[0090] Since the self-attention component can perform non-local information extraction on the feature tensor output by the encoder component and obtain the dependency between non-local features, and the multi-scale analysis component can extract the feature tensors output by the feature transfer component at multiple scales, thereby mining the correlation between feature tensors at different scales and obtaining contextual information containing multi-scale analysis results, the dental mold deformation model obtained by the method for training a dental mold deformation model provided by an embodiment of the present invention can more accurately deform the dental model and more accurately obtain a dental model that meets specific requirements.

[0091] Further, see Figure 3 As shown, the encoder component 21 includes three encoders (encoder 211, encoder 212, encoder 213) in a serial structure, and each encoder (encoder 211, encoder 212, encoder 213) includes a residual unit (residual unit E1, residual unit E2, residual unit E3) and a downsampling unit (downsampling unit Do1, downsampling unit Do2, downsampling unit Do3); the residual unit (residual unit E1, residual unit E2, residual unit E3) of each encoder is used to perform a convolution operation on the input of the residual unit through three convolution layers in a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the downsampling unit (downsampling unit Do1, downsampling unit Do2, downsampling unit Do3) of each encoder is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor.

[0092] The input of each residual unit is the input of the encoder to which it belongs (the input of residual unit E1 is the input of encoder 211, the input of residual unit E2 is the input of encoder 212, and the input of residual unit E3 is the input of encoder 213). The input of each downsampling unit is the output of the residual unit of the encoder to which it belongs (the input of downsampling unit Do1 is the output of residual unit E1 of encoder 211, the input of downsampling unit Do2 is the output of residual unit E2 of encoder 212, and the input of downsampling unit Do3 is the output of residual unit E3 of encoder 213), and the output of each downsampling unit is the output of the encoder to which it belongs (the output of downsampling unit Do1 is the output of encoder 211, the output of downsampling unit Do2 is the output of encoder 212, and the output of downsampling unit Do3 is the output of encoder 213). The input of the first encoder 211 is the input of encoder component 21, the output of the third encoder 213 is the output of encoder component 21, and the inputs of the second encoder 212 and the third encoder 213 are the outputs of the first encoder 211 and the second encoder 212, respectively.

[0093] That is, the input of the encoder component 21, the input of the first encoder 211, and the input of the residual unit E1 are the same input, and the output of the residual unit E1 is the input of the downsampling unit Do1. The output of the downsampling unit Do1 is the output of the first encoder 211. The input of the second encoder 212 and the input of the residual unit E2 are the same input, both of which are the output of the first encoder 211. The output of the residual unit E2 is the input of the downsampling unit Do2. The output of the downsampling unit Do2 is the output of the second encoder 212. The input of the third encoder 213 and the input of the residual unit E3 are the same input, both of which are the output of the second encoder 212. The output of the residual unit E3 is the input of the downsampling unit Do3. The output of the downsampling unit Do3 is the output of the third encoder 212, the output of the encoder component 21.

[0094] Optionally, the convolution kernels of the three convolution layers of the residual units of each decoder are all 3×3×3, the length, width, and height of the feature tensors output by the three convolution layers of the residual units of each decoder are the same as the length, width, and height of the input feature tensors, the number of channels of the feature tensor output by the residual unit of the first encoder is 16 times the number of channels of the input feature tensor, and the number of channels of the feature tensors output by the residual units of the second encoder and the residual units of the third encoder are the same as the number of channels of the input feature tensor. Each downsampling unit is a convolution layer with a stride of 2 and a convolution kernel of 2×2×2.

[0095] Further, see Figure 4 As shown, the self-attention component 22 includes: a residual unit E4, a first convolutional layer Co1, a second convolutional layer Co2, a third convolutional layer Co3, a fourth convolutional layer Co4, a fifth convolutional layer Co5, a sixth convolutional layer Co6, a first dot product unit Pro1, a second dot product unit Pro2, a first addition unit Add1 and a second addition unit Add2.

[0096] The residual unit E4 is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and to perform a sum operation on the convolution result of the convolution operation and the input of the residual unit. The first dot product unit Pro1 and the second dot product unit Pro2 are used to perform a dot product operation on the input feature tensor. The first addition unit Add1 and the second addition unit Add2 perform a sum operation on the input feature tensor.

[0097] The input of the residual unit E4 is the output of the encoder component 21 (the output of the downsampling unit Do3 of the third encoder 213 of the encoder component 21), and the output of the residual unit E4 is the input of the first convolution layer Co1. The output of the first convolution layer Co1 is the input of the second convolution layer Co2, the third convolution layer Co3 and the fourth convolution layer Co4. The input of the first dot product unit Pro1 is the output of the second convolution layer Co2 and the output of the third convolution layer Co3. The input of the second dot product unit Pro2 is the output of the first dot product unit Pro1 and the output of the fourth convolution layer Co4. The input of the fifth convolution layer Co5 is the output of the second dot product unit Pro2. The input of the first addition unit Add1 is the output of the fifth convolution layer Co5 and the output of the first convolution layer Co1. The input of the sixth convolution layer Co6 is the output of the first addition unit Add1. The input of the second addition unit Add2 is the output of the sixth convolution layer Co6 and the output of the residual unit E4. The output of the second adding unit Add2 is the output of the self-attention component 22.

[0098] Optionally, the convolution kernels of the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, the fifth convolution layer and the sixth convolution layer are all 1×1×1. The length, width and height of the output feature tensor of the first convolution layer are the same as the length, width and height of the input feature tensor. The number of channels of the output feature tensor of the first convolution layer is one eighth of the number of channels of the input feature tensor. The number of channels of the output feature tensor of the second convolution layer, the third convolution layer and the fourth convolution layer is one half of the number of channels of the feature tensor of the input feature tensor. The number of channels of the output feature tensor of the fifth convolution layer is twice the number of channels of the feature tensor of the input feature tensor. The number of channels of the output feature tensor of the sixth convolution layer is eight times the number of channels of the feature tensor of the input feature tensor.

[0099] Assume: The feature tensor X∈R of the output of the residual unit E4 C×H×W×L , where C is the number of channels of the feature tensor output by the residual unit E4, H, W, and L are the length, width, and height of the feature tensor output by the residual unit E4, respectively. The feature tensor X output by the first convolutional layer Co1 is 1 ∈R C1×H×W×L , C1 = C / 8, the feature tensor X output by the second convolutional layer Co2 2 ∈R C2×H×W×L , C2=C1 / 2, the feature tensor X output by the third convolutional layer Co3 3 ∈R C3×H×W×L , C3=C1 / 2.

[0100] Through xi ∈R 1 and x j ∈R 1 Represents X 2 The i-th and X 3 The label of the j-th voxel in 2 (x i )∈R C2 Represents X 2 No. x i The eigenvector of a voxel, X 3 (x j )∈R C3 Represents X 3 No. x j The feature vector of a voxel, then the attention distribution is:

[0101]

[0102] The feature tensor X output by the fourth convolutional layer Co4 4 ∈R C4×H×W×L , C4=C1 / 2. 4 By performing a dot product operation with S, we can obtain the environmental features that describe the non-local dependencies:

[0103]

[0104] Then we can get the environmental characteristics Con2∈R C5×H×W×L ,C5=C1 / 2.

[0105] The feature tensor X output by the fifth convolutional layer Co5 5 , the first addition unit Add1 is for X 5 and X 1 Perform the addition operation, the feature tensor res1∈R output by the first addition unit Add1 C1×H×W×L , the feature tensor X output by the sixth convolutional layer Co6 6 ∈R C ×H×W×L , the second addition unit Add2 adds X and X 6 Perform the sum operation to obtain the final environment feature tensor res2∈R C×W×H×L .

[0106] Further, see Figure 5 As shown, the feature transfer component 23 includes a downsampling unit Do4 and a residual unit E5;

[0107] The downsampling unit Do4 is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor. The residual unit E5 is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform a summation operation on the convolution result of the convolution operation and the input of the residual unit.

[0108] The input of the downsampling unit Do4 is the output of the self-attention component 22 , the output of the downsampling unit Do4 is the input of the residual unit E5 , and the output of the residual unit E5 is the output of the feature transfer component 23 .

[0109] Optionally, the downsampling unit Do4 is a convolution layer with a stride of 2 and a convolution kernel of 2×2×2. The convolution kernels of the three convolution layers of the residual unit E5 are all 3×3×3, the length, width, and height of the output feature tensors of the three convolution layers of each residual unit are the same as the length, width, and height of the input feature tensor, and the number of channels of the output feature tensor of the residual unit E5 is the same as the number of channels of the input feature tensor.

[0110] Further, see Figure 6 As shown, the multi-scale analysis component 24 includes: a seventh convolution layer Co7, an eighth convolution layer Co8, a ninth convolution layer Co9, a tenth convolution layer Co10, an eleventh convolution layer Co11, a twelfth convolution layer Co12 and a splicing unit MON.

[0111] Among them, the expansion rates of the seventh convolutional layer Co7, the eighth convolutional layer Co8, the ninth convolutional layer Co9, and the tenth convolutional layer Co10 are all different; the splicing unit MON is used to perform a splicing operation on the input feature tensor.

[0112] The inputs of the seventh convolutional layer Co7, the eighth convolutional layer Co8, the ninth convolutional layer Co9 and the tenth convolutional layer Co10 are all the output of the feature transfer component 23, the input of the splicing unit MON is the output of the feature transfer component 23, the output of the seventh convolutional layer Co7, the output of the eighth convolutional layer Co8, the output of the ninth convolutional layer Co9 and the output of the tenth convolutional layer Co10, the input of the eleventh convolutional layer Co11 is the output of the splicing unit MON, the input of the twelfth convolutional layer Co12 is the output of the eleventh convolutional layer Co11; the output of the twelfth convolutional layer Co12 is the output of the multi-scale analysis component 24.

[0113] Optionally, the convolution kernel of the seventh convolution layer is 1×1×1, the number of channels of the output feature tensor is the same as the number of channels of the input feature tensor, and the expansion rate is 1. The convolution kernels of the eighth convolution layer, the ninth convolution layer, and the tenth convolution layer are all 3×3×3, the number of channels of the output feature tensor is the same as the number of channels of the input feature tensor, and the expansion rates are 2, 3, and 4, respectively. The convolution kernel of the eleventh convolution layer is 3×3×3, and the number of channels of the output feature tensor is one-fifth of the number of channels of the input feature tensor; the convolution kernel of the twelfth convolution layer is 3×3×3, and the number of channels of the output feature tensor is the same as the number of channels of the input feature tensor.

[0114] Assume: The feature tensor A∈R output by the feature transfer component 23 C×H×W×L , then the feature tensor A1∈R output by the seventh convolutional layer Co7 C×H×W×C , the feature tensor A2∈R output by the eighth convolutional layer Co8 C×H×W×L , the feature tensor A3∈R output by the ninth convolutional layer Co9 C×H×W×L , the feature tensor A4∈R output by the tenth convolutional layer Co10 C×H×W×L , the feature tensor Cat∈R output by the splicing module MON 5×C×H×W×L , the feature tensor Cat1∈R output by the eleventh convolutional layer Co11 C×H×W×L , the feature tensor Cat2∈R output by the twelfth convolutional layer Co12 C×H×W×L .

[0115] Further, see Figure 7 As shown, the decoder component 25 includes four decoders in a serial structure (decoder 251, decoder 252, decoder 253, decoder 254); each decoder includes: an upsampling unit (upsampling unit Up1 of decoder 251, upsampling unit Up2 of decoder 252, upsampling unit Up3 of decoder 253, upsampling unit Up4 of decoder 254), a fusion unit (fusion unit F1 of decoder 251, fusion unit F2 of decoder 252, fusion unit F3 of decoder 253, fusion unit F4 of decoder 254) and a residual unit (residual unit E6 of decoder 251, residual unit E7 of decoder 252, residual unit E8 of decoder 253, residual unit E9 of decoder 254).

[0116] Among them, the residual unit of each decoder is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit. The fusion unit of each decoder is used to perform a fusion operation on the input feature tensor; the upsampling unit of each decoder is used to upsample the input feature tensor into an output feature tensor whose number of channels is half of the number of channels of the input feature tensor and whose length, width and height are twice the length, width and height of the input feature tensor.

[0117] The input of the upsampling unit Up1 of the first decoder 251 of the decoder component is the output of the multi-scale analysis component 24, the input of the fusion unit F1 of the first decoder 251 is the output of the upsampling unit Up1 of the first decoder 251 and the output of the self-attention component 22; the input of the residual unit E6 of the first decoder 251 is the output of the fusion unit F1 of the first decoder 251 and the output of the self-attention component 22; the input of the upsampling units (upsampling unit Up2, upsampling unit Up3, upsampling unit Up4) of the second decoder 252, the third decoder 253, and the fourth decoder 254 of the decoder component are all the output of the previous decoder, and the second decoder 252, The inputs of the fusion units (fusion unit F2, fusion unit F3, fusion unit F4) of the third decoder 253 and the fourth decoder 254 are the outputs of the residual units (residual unit E3, residual unit E2, residual unit E1) of the corresponding encoders (encoder 213, encoder 212, encoder 211) and the outputs of the upsampling units of the corresponding decoders. The inputs of the residual units (residual unit E7, residual unit E8, residual unit E9) of the second decoder 252, the third decoder 253, and the fourth decoder 254 of the decoder component are the outputs of the residual units (residual unit E3, residual unit E2, residual unit E1) of the corresponding encoders (encoder 213, encoder 212, encoder 211) and the outputs of the fusion units of the corresponding decoders.

[0118] Optionally, the convolution kernels of the three convolution layers of the residual unit of each encoder are all 3×3×3, the length, width, and height of the feature tensor output by the three convolution layers of the residual unit of each encoder are the same as the length, width, and height of the input feature tensor, and the number of channels of the feature tensor output by the residual unit of each decoder is the same as the number of channels of the input feature tensor. Each upsampling unit is a deconvolution layer with a stride of 2 and a convolution kernel of 2×2×2.

[0119] Further, see Figure 8As shown, the fusion units of each decoder include: a thirteenth convolutional layer Co13, a fourteenth convolutional layer Co14, a fifteenth convolutional layer Co15, a third addition unit Add3, a fourth addition unit Add4, a third dot product unit Pro3 and a fourth dot product unit Pro4.

[0120] The third addition unit Add3 and the fourth addition unit Add4 are used to perform addition operations on the inputs, and the third dot product unit Pro3 and the fourth dot product unit Pro4 are used to perform dot product operations on the inputs.

[0121] The thirteenth convolutional layer Co13 and the fourteenth convolutional layer Co14 inputs of the fusion unit F1 of the first decoder 251 of the decoder component 25 are respectively the output of the upsampling unit Up1 of the first decoder 251 and the output of the self-attention component 22, and the inputs of the fusion units (fusion unit F2, fusion unit F3, fusion unit F4) of the second decoder 252, the third decoder 253, and the fourth decoder 254 of the decoder component are the output of the upsampling unit of the decoder and the output of the corresponding residual unit of the encoder (the input of the fusion unit F2 is the output of the residual unit E3 of the encoder 213 and the output of the upsampling unit Up2 of the decoder 252, the input of the fusion unit F3 is the output of the residual unit E2 of the encoder 212 and the output of the upsampling unit Up3 of the decoder 253, and the input of the fusion unit F4 is The input of the third addition unit Add3 is the output of the thirteenth convolutional layer Co13 and the output of the fourteenth convolutional layer Co14, the input of the fifteenth convolutional layer Co15 is the output of the third addition unit Add3, the input of the third dot product unit Pro3 is the output of the thirteenth convolutional layer Co13 and the output of the fifteenth convolutional layer Co15, the input of the fourth dot product unit Pro4 is the output of the fourteenth convolutional layer Co14 and the output of the fifteenth convolutional layer Co15, the input of the fourth addition unit Add4 is the output of the third dot product unit Pro3 and the output of the fourth dot product unit Pro4, and the output of the fourth addition unit Add4 is the output of the fusion unit to which it belongs.

[0122] That is, Figure 8As shown, the fusion unit performs convolution operations on the two input feature tensors Ai and Bi through the thirteenth convolution layer Co13 and the fourteenth convolution layer Co14 respectively to obtain the feature tensors Ci and Di after dimensionality reduction; then the third addition unit Add3 is used to perform a sum operation on Ci and Di to fuse Ci and Di, and then the fusion result of the third addition unit Add3 is sent to the fifteenth convolution layer Co15 to obtain the encoder weight coefficient tensor Ei and the decoder weight coefficient tensor Fi respectively; then the weight coefficient tensors Ci and Ei are dot-multiplied by the third dot product unit Pro3 to obtain the result Gi, and the result Hi is obtained by the fourth dot product unit Pro4 to dot-multiply Di and Fi; finally, the fourth addition unit Add4 performs a sum operation on Gi and Hi to obtain a fused feature Zi containing the encoder output feature and the decoder output feature as the input feature tensor of the i-th module decoder residual unit.

[0123] Optionally, the convolution kernels of the thirteenth convolution layer, the fourteenth convolution layer, and the fifteenth convolution layer are all 1×1×1, the number of channels of the output feature tensors of the third convolution layer and the fourth convolution layer is half the number of channels of the input feature tensor, and the number of channels of the output feature tensor of the fifteenth convolution layer is 1.

[0124] Assume: Ai∈R C×W×H×L and Bi∈R C×W×H×L , then Ci∈R 1 / 2C×W×H×L and Di∈R 1 / 2C×W×H×L , Ei∈R 1×W×H×L Fi∈R 1×W×H×L .

[0125] As an optional embodiment of the present invention, the above step S104 (optimizing the preset network model to obtain the dental mold deformation model according to the target deformation model and the predicted deformation model corresponding to each initial dental model) includes:

[0126] Constructing a loss function, and optimizing the preset network model to obtain a tooth mold deformation model according to the loss function, the target deformation model corresponding to each initial tooth model, and the predicted deformation model;

[0127] Wherein, the loss function includes:

[0128]

[0129] Among them, alpha is a constant, out i The data is obtained by processing the output of the multi-scale analysis component and the output of each decoder of the decoder component in sequence, and seg is an intermediate supervision signal.

[0130] Optional, out 1The output of the multi-scale analysis component is convolved through a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1. The length, width, and height of the feature tensor obtained by the convolution operation are expanded by 16 times through trilinear interpolation (Trilinear), and the result is obtained by performing a Sigmoid operation on the trilinear interpolation result.

[0131] out 2 The output of the first decoder of the decoder component is convolved by a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1, and the length, width, and height of the feature tensor obtained by the convolution operation are expanded by 8 times by trilinear interpolation, and a sigmoid operation is performed on the trilinear interpolation result to obtain the result;

[0132] out 3 The output of the second decoder of the decoder component is convolved by a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1, and the length, width, and height of the feature tensor obtained by the convolution operation are expanded by 4 times by trilinear interpolation (Trilinear), and a sigmoid operation is performed on the trilinear interpolation result to obtain the result;

[0133] out 4 Perform a convolution operation on the output of the third decoder of the decoder component through a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1, and then expand the length, width, and height of the feature tensor obtained by the convolution operation by 2 times through trilinear interpolation (Trilinear), and then perform a Sigmoid operation on the trilinear interpolation result to obtain the result;

[0134] out 5 A convolution operation is performed on the output of the fourth decoder of the decoder component through a convolution layer with a convolution kernel of 1×1×1 and a channel number of 1 of the output feature tensor, and then a Sigmoid operation is performed on the feature tensor obtained by the convolution operation to obtain a result.

[0135] Optional, alpha is 0.25.

[0136] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present invention further provides a device for establishing a dental mold deformation model. The device embodiment corresponds to the aforementioned method embodiment. For ease of reading, the present device embodiment will no longer repeat the details of the aforementioned method embodiment one by one, but it should be clear that the device for establishing a dental mold deformation model in this embodiment can correspond to and implement all the contents of the aforementioned method embodiment.

[0137] Fig. 9 A schematic diagram of the structure of a device for establishing a deformable dental model according to an embodiment of the present invention, as shown in FIG. Fig. 9 As shown, the device 900 for establishing a deformable dental model provided in this embodiment includes:

[0138] A sample acquisition unit 91 is used to acquire sample data, wherein the sample data includes a plurality of initial tooth models acquired by scanning the oral cavity and a target deformation model corresponding to each initial tooth model obtained by manually processing each initial tooth model;

[0139] The preprocessing unit 92 is used to obtain the feature tensor corresponding to each initial tooth model, each element of the feature tensor corresponding to each initial tooth model is the truncated signed distance function value TSDF value of each voxel in the cubic space where each initial tooth model is located;

[0140] A prediction unit 93, used for inputting the feature tensor corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model;

[0141] The optimization unit 94 is used to optimize the preset network model to obtain the tooth mold deformation model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model.

[0142] As an optional implementation of an embodiment of the present invention, the preset network model includes: an encoder component composed of multiple encoders in series structure, a self-attention component, a feature transfer component, a multi-scale analysis component, and a decoder component composed of multiple decoders in series structure; the input of the encoder component is the input of the preset network model, and the output of the encoder component is the input of the self-attention component; the output of the self-attention component is the input of the feature transfer component; the output of the feature transfer component is the input of the multi-scale analysis component, the output of the multi-scale analysis component is the input of the decoder component, and the output of the decoder component is the output of the preset network model;

[0143] Among them, the self-attention component is used to extract non-local information from the feature tensor output by the encoder component to obtain the environmental feature tensor; the feature transfer component processes the output of the self-attention component and transfers the processing result to the multi-scale analysis component; the multi-scale analysis component is used to extract the feature tensor output by the feature transfer component at multiple scales.

[0144] As an optional implementation of an embodiment of the present invention, the encoder component includes three encoders in a serial structure, each encoder includes a residual unit and a downsampling unit; the residual unit of each encoder is used to perform a convolution operation on the input of the residual unit through three convolution layers in a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the downsampling unit of each encoder is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor;

[0145] The input of each residual unit is the input of the encoder to which it belongs, the input of each downsampling unit is the residual unit output of the encoder to which it belongs, the output of each downsampling unit is the output of the encoder to which it belongs, the input of the first encoder is the input of the encoder component, the output of the third encoder is the output of the encoder component, and the inputs of the second and third encoders are the outputs of the first and second encoders, respectively.

[0146] As an optional implementation of an embodiment of the present invention, the self-attention component includes: a residual unit, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a first dot product unit, a second dot product unit, a first addition unit and a second addition unit; the residual unit is used to perform a convolution operation on the input of the residual unit through three convolutional layers of a serial structure and perform a summation operation on the convolution result of the convolution operation and the input of the residual unit, the first dot product unit and the second dot product unit are used to perform a dot product operation on the input feature tensor, and the first summation unit and the second summation unit are used to perform a summation operation on the input feature tensor;

[0147] The input of the residual unit is the output of the encoder component, and the output of the residual unit is the input of the first convolutional layer; the output of the first convolutional layer is the input of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer; the input of the first dot convolution unit is the output of the second convolutional layer and the output of the third convolutional layer; the input of the second dot convolution unit is the output of the first dot convolution unit and the output of the fourth convolutional layer; the input of the fifth convolutional layer is the output of the second dot convolution unit; the input of the first summation unit is the output of the fifth convolutional layer and the output of the first convolutional layer; the input of the sixth convolutional layer is the output of the first summation unit; the input of the second summation unit is the output of the sixth convolutional layer and the output of the residual unit, and the output of the second summation unit is the output of the self-attention component.

[0148] As an optional implementation of the embodiment of the present invention, the feature transfer component includes a downsampling unit and a residual unit;

[0149] The downsampling unit is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor, and the residual unit is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit;

[0150] The input of the downsampling unit is the output of the self-attention component, the output of the downsampling unit is the input of the residual unit, and the output of the residual unit is the output of the feature transfer component.

[0151] As an optional implementation of the embodiment of the present invention, the multi-scale analysis component includes: a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer and a splicing unit; the expansion rates of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer are all different; the splicing unit is used to perform a splicing operation on the input feature tensor;

[0152] The inputs of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer are all the outputs of the feature transfer component, the input of the splicing unit is the output of the feature transfer component, the output of the seventh convolutional layer, the output of the eighth convolutional layer, the output of the ninth convolutional layer, and the output of the tenth convolutional layer, the input of the eleventh convolutional layer is the output of the splicing unit, the input of the twelfth convolutional layer is the output of the eleventh convolutional layer; the output of the twelfth convolutional layer is the output of the multi-scale analysis component.

[0153] As an optional implementation of the embodiment of the present invention, the decoder component includes four decoders of a serial structure; each decoder includes: an upsampling unit, a fusion unit and a residual unit; the residual unit of each decoder is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the fusion unit of each decoder is used to perform a fusion operation on the input feature tensor; the upsampling unit of each decoder is used to upsample the input feature tensor to an output feature tensor whose number of channels is half of the number of channels of the input feature tensor and whose length, width and height are twice the length, width and height of the input feature tensor;

[0154] The input of the upsampling unit of the first decoder of the decoder component is the output of the multi-scale analysis component, the input of the fusion unit of the first decoder is the output of the upsampling unit of the first decoder and the output of the self-attention component; the input of the residual unit of the first decoder is the output of the fusion unit of the first decoder and the output of the self-attention component; the input of the upsampling unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are all the output of the previous decoder, the input of the fusion unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are the output of the residual unit of the corresponding encoder and the output of the upsampling unit of the decoder to which they belong, and the input of the residual unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are the output of the residual unit of the corresponding encoder and the output of the fusion unit of the decoder to which they belong.

[0155] As an optional implementation of the embodiment of the present invention, the fusion unit of each decoder includes: a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, a third addition unit, a fourth addition unit, a third dot product unit, and a fourth dot product unit; the third addition unit and the fourth addition unit are used to perform an addition operation on the input, and the third dot product unit and the fourth dot product unit are used to perform a dot product operation on the input;

[0156] The inputs of the thirteenth convolutional layer and the fourteenth convolutional layer of the fusion unit of the first decoder of the decoder component are the output of the upsampling unit of the first decoder and the output of the self-attention component respectively. The inputs of the fusion units of the second decoder, the third decoder and the fourth decoder of the decoder component are the output of the upsampling unit of the decoder to which they belong and the output of the residual unit of the corresponding encoder. The input of the third addition unit is the output of the thirteenth convolutional layer and the output of the fourteenth convolutional layer. The input of the fifteenth convolutional layer is the output of the third addition unit. The input of the third dot product unit is the output of the thirteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth dot product unit is the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth addition unit is the output of the third dot product unit and the output of the fourth dot product unit. The output of the fourth addition unit is the output of the fusion unit to which it belongs.

[0157] As an optional implementation of the embodiment of the present invention, the optimization unit 94 is specifically used to construct a loss function, and optimize the preset network model to obtain a tooth mold deformation model according to the loss function, the target deformation model corresponding to each initial tooth model, and the predicted deformation model;

[0158] Wherein, the loss function includes:

[0159]

[0160] Among them, alpha is a constant, out j The data is obtained by processing the output of the multi-scale analysis component and the output of each decoder of the decoder component in sequence, seg is the intermediate supervision signal, and mean() is the averaging function.

[0161] The device for establishing a dental mold deformation model provided in this embodiment can execute the method for training a dental mold deformation model provided in the above method embodiment. Its implementation principle and technical effect are similar and will not be repeated here.

[0162] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device. Fig.10 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Fig.10 As shown, the electronic device provided in this embodiment includes: a memory 101 and a processor 102, the memory 101 is used to store computer programs; the processor 102 is used to execute each step in the method for training a dental mold deformation model provided in the above method embodiment when calling the computer program.

[0163] Specifically, the memory 101 can be used to store software programs and various data. The memory 101 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 101 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0164] The processor 102 is the control center of the electronic device, which uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 101, and calling data stored in the memory 101, thereby monitoring the electronic device as a whole. The processor 102 may include one or more processing units.

[0165] In addition, it should be understood that the electronic device provided in the embodiment of the present invention may also include: a radio frequency unit, a network module, an audio output unit, a sensor, a signal receiving unit, a display, a user receiving unit, an interface unit, and a power supply. Those skilled in the art will understand that the structure of the electronic device described above does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components, or combine certain components, or arrange components differently. In the embodiment of the present invention, the electronic device includes but is not limited to a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted terminal, a wearable device, and a pedometer.

[0166] The RF unit can be used to receive and send signals during information transmission or calls. Specifically, after receiving downlink data from the base station, it is sent to the processor 102 for processing; in addition, uplink data is sent to the base station. Generally, the RF unit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the RF unit can also communicate with the network and other devices through a wireless communication system.

[0167] Electronic devices provide users with wireless broadband Internet access through network modules, such as helping users to send and receive emails, browse web pages, and access streaming media.

[0168] The audio output unit can convert the audio data received by the radio frequency unit or the network module or stored in the memory 101 into an audio signal and output it as sound. Moreover, the audio output unit can also provide audio output related to a specific function performed by the electronic device (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit includes a speaker, a buzzer, a receiver, etc.

[0169] The signal receiving unit is used to receive audio or video signals. The receiving unit may include a graphics processing unit (GPU) and a microphone, and the graphics processor convolves the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on a display unit. The image frame processed by the graphics processor can be stored in a memory (or other storage medium) or sent via a radio frequency unit or a network module. The microphone can receive sound and can process such sound into audio data. The processed audio data can be converted into a format output that can be sent to a mobile communication base station via a radio frequency unit in the case of a telephone call mode.

[0170] The electronic device also includes at least one sensor, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light, and the proximity sensor can turn off the display panel and / or backlight when the electronic device is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.

[0171] The display unit is used to display information input by the user or information provided to the user. The display unit may include a display panel, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0172] The user receiving unit can be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the electronic device. Specifically, the user receiving unit includes a touch panel and other input devices. The touch panel, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on or near the touch panel using any suitable object or accessory such as a finger, stylus, etc.). The touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the contact point coordinates, and then sends it to the processor 102, receives the command sent by the processor 102 and executes it. In addition, the touch panel can be implemented using multiple types such as resistive, capacitive, infrared and surface acoustic waves. In addition to the touch panel, the user receiving unit may also include other input devices. Specifically, other input devices may include but are not limited to physical keyboards, function keys (such as volume control keys, switch keys, etc.), trackballs, mice, joysticks, which will not be repeated here.

[0173] Furthermore, the touch panel may be covered on the display panel, and when the touch panel detects a touch operation on or near it, it is transmitted to the processor 102 to determine the type of touch event, and then the processor 102 provides a corresponding visual output on the display panel according to the type of touch event. Generally, the touch panel and the display panel are used as two independent components to implement the input and output functions of the electronic device, but in some embodiments, the touch panel and the display panel can be integrated to implement the input and output functions of the electronic device, which is not limited here.

[0174] The interface unit is an interface for connecting an external device to an electronic device. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements in the electronic device or may be used to transmit data between the electronic device and the external device.

[0175] The electronic device may also include a power source (such as a battery) for supplying power to each component. Optionally, the power source may be logically connected to the processor 102 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system.

[0176] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for training a dental mold deformation model provided in the above method embodiment is implemented.

[0177] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0178] Computer readable media include permanent and non-permanent, removable and non-removable storage media. Storage media can be implemented by any method or technology to store information, and the information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0179] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0180] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments described herein, but should conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training a dental mold deformation model, It is characterized in that include: Acquire sample data, wherein the sample data includes a plurality of initial tooth models acquired by scanning the oral cavity and a target deformation model corresponding to each initial tooth model obtained by manually processing each initial tooth model; Acquire a feature tensor corresponding to each initial tooth model, wherein each element of the feature tensor corresponding to each initial tooth model is a truncated signed distance function TSDF value of each voxel in the cubic space where each initial tooth model is located; Inputting the feature tensors corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model; According to the target deformation model and the predicted deformation model corresponding to each initial tooth model, the preset network model is optimized to obtain the tooth mold deformation model; The preset network model includes: an encoder component composed of multiple encoders in series structure, a self-attention component, a feature transfer component, a multi-scale analysis component, and a decoder component composed of multiple decoders in series structure; the input of the encoder component is the input of the preset network model, the output of the encoder component is the input of the self-attention component; the output of the self-attention component is the input of the feature transfer component; the output of the feature transfer component is the input of the multi-scale analysis component, the output of the multi-scale analysis component is the input of the decoder component, and the output of the decoder component is the output of the preset network model; Among them, the self-attention component is used to extract non-local information from the feature tensor output by the encoder component to obtain the environmental feature tensor; the feature transfer component processes the output of the self-attention component and transfers the processing result to the multi-scale analysis component; the multi-scale analysis component is used to extract the feature tensor output by the feature transfer component at multiple scales.

2. The method according to claim 1, It is characterized in that The encoder component includes three encoders in a serial structure, each encoder including a residual unit and a downsampling unit; the residual unit of each encoder is used to perform a convolution operation on the input of the residual unit through three convolution layers in a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the downsampling unit of each encoder is used to downsample the input feature tensor into an output feature tensor having a channel number that is twice the number of channels of the input feature tensor and a length, width and height that are half the length, width and height of the input feature tensor; The input of each residual unit is the input of the encoder to which it belongs, the input of each downsampling unit is the residual unit output of the encoder to which it belongs, the output of each downsampling unit is the output of the encoder to which it belongs, the input of the first encoder is the input of the encoder component, the output of the third encoder is the output of the encoder component, and the inputs of the second and third encoders are the outputs of the first and second encoders, respectively.

3. The method according to claim 1, It is characterized in that The self-attention component includes: a residual unit, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a first dot product unit, a second dot product unit, a first summing unit and a second summing unit; the residual unit is used to perform a convolution operation on the input of the residual unit through three convolutional layers of a serial structure and perform a summing operation on the convolution result of the convolution operation and the input of the residual unit, the first dot product unit and the second dot product unit are used to perform a dot product operation on the input feature tensor, and the first summing unit and the second summing unit are used to perform a summing operation on the input feature tensor; The input of the residual unit is the output of the encoder component, and the output of the residual unit is the input of the first convolutional layer; the output of the first convolutional layer is the input of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer; the input of the first dot convolution unit is the output of the second convolutional layer and the output of the third convolutional layer; the input of the second dot convolution unit is the output of the first dot convolution unit and the output of the fourth convolutional layer; the input of the fifth convolutional layer is the output of the second dot convolution unit; the input of the first summation unit is the output of the fifth convolutional layer and the output of the first convolutional layer; the input of the sixth convolutional layer is the output of the first summation unit; the input of the second summation unit is the output of the sixth convolutional layer and the output of the residual unit, and the output of the second summation unit is the output of the self-attention component.

4. The method according to claim 1, It is characterized in that The feature transfer component includes a downsampling unit and a residual unit; The downsampling unit is used to downsample the input feature tensor into an output feature tensor whose number of channels is twice the number of channels of the input feature tensor and whose length, width and height are half the length, width and height of the input feature tensor, and the residual unit is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit; The input of the downsampling unit is the output of the self-attention component, the output of the downsampling unit is the input of the residual unit, and the output of the residual unit is the output of the feature transfer component.

5. The method according to claim 1, It is characterized in that The multi-scale analysis component includes: a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer and a splicing unit; the expansion rates of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer and the tenth convolutional layer are all different; the splicing unit is used to perform a splicing operation on the input feature tensor; The inputs of the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, and the tenth convolutional layer are all the outputs of the feature transfer component, the input of the splicing unit is the output of the feature transfer component, the output of the seventh convolutional layer, the output of the eighth convolutional layer, the output of the ninth convolutional layer, and the output of the tenth convolutional layer, the input of the eleventh convolutional layer is the output of the splicing unit, the input of the twelfth convolutional layer is the output of the eleventh convolutional layer; the output of the twelfth convolutional layer is the output of the multi-scale analysis component.

6. The method according to claim 1, It is characterized in that The decoder component includes four decoders of a serial structure; each decoder includes: an upsampling unit, a fusion unit and a residual unit; the residual unit of each decoder is used to perform a convolution operation on the input of the residual unit through three convolution layers of a serial structure and perform an addition operation on the convolution result of the convolution operation and the input of the residual unit, and the fusion unit of each decoder is used to perform a fusion operation on the input feature tensor; the upsampling unit of each decoder is used to upsample the input feature tensor to an output feature tensor whose number of channels is half of the number of channels of the input feature tensor and whose length, width and height are twice the length, width and height of the input feature tensor; The input of the upsampling unit of the first decoder of the decoder component is the output of the multi-scale analysis component, the input of the fusion unit of the first decoder is the output of the upsampling unit of the first decoder and the output of the self-attention component; the input of the residual unit of the first decoder is the output of the fusion unit of the first decoder and the output of the self-attention component; the input of the upsampling unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are all the output of the previous decoder, the input of the fusion unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are the output of the residual unit of the corresponding encoder and the output of the upsampling unit of the decoder to which they belong, and the input of the residual unit of the second decoder, the third decoder, and the fourth decoder of the decoder component are the output of the residual unit of the corresponding encoder and the output of the fusion unit of the decoder to which they belong.

7. The method according to claim 6, It is characterized in that The fusion unit of each decoder includes: a thirteenth convolution layer, a fourteenth convolution layer, a fifteenth convolution layer, a third addition unit, a fourth addition unit, a third dot product unit, and a fourth dot product unit; the third addition unit and the fourth addition unit are used to perform an addition operation on the input, and the third dot product unit and the fourth dot product unit are used to perform a dot product operation on the input; The inputs of the thirteenth convolutional layer and the fourteenth convolutional layer of the fusion unit of the first decoder of the decoder component are the output of the upsampling unit of the first decoder and the output of the self-attention component respectively. The inputs of the fusion units of the second decoder, the third decoder and the fourth decoder of the decoder component are the output of the upsampling unit of the decoder to which they belong and the output of the residual unit of the corresponding encoder. The input of the third addition unit is the output of the thirteenth convolutional layer and the output of the fourteenth convolutional layer. The input of the fifteenth convolutional layer is the output of the third addition unit. The input of the third dot product unit is the output of the thirteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth dot product unit is the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer. The input of the fourth addition unit is the output of the third dot product unit and the output of the fourth dot product unit. The output of the fourth addition unit is the output of the fusion unit to which it belongs.

8. The method according to claim 6, It is characterized in that The step of optimizing the preset network model to obtain the tooth mold deformation model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model includes: Constructing a loss function, and optimizing the preset network model to obtain a tooth mold deformation model according to the loss function, the target deformation model corresponding to each initial tooth model, and the predicted deformation model; Wherein, the loss function includes: Among them, alpha is a constant, out 1 The output of the multi-scale analysis component is convolved through a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1. The length, width, and height of the feature tensor obtained by the convolution operation are expanded by 16 times through trilinear interpolation (Trilinear), and the sigmoid operation is performed on the trilinear interpolation result to obtain the result; out 2 The output of the first decoder of the decoder component is convolved by a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1, and then the length, width, and height of the feature tensor obtained by the convolution operation are expanded by 8 times by trilinear interpolation (Trilinear), and then the sigmoid operation is performed on the trilinear interpolation result to obtain the result; out 3 The output of the second decoder of the decoder component is convolved by a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1, and then the length, width, and height of the feature tensor obtained by the convolution operation are expanded by 4 times through trilinear interpolation (Trilinear), and then the sigmoid operation is performed on the trilinear interpolation result to obtain the result; out 4 The output of the third decoder of the decoder component is convolved by a convolution layer with a convolution kernel of 1×1×1 and an output feature tensor with a channel number of 1, and then the length, width, and height of the feature tensor obtained by the convolution operation are expanded by 2 times by trilinear interpolation (Trilinear), and then the sigmoid operation is performed on the trilinear interpolation result to obtain the result; out 5 A convolution operation is performed on the output of the fourth decoder of the decoder component through a convolution layer with a convolution kernel of 1×1×1 and a channel number of 1 of the output feature tensor, and then a Sigmoid operation is performed on the feature tensor obtained by the convolution operation to obtain a result; seg is a binary tensor representation of the voxel occupancy information of the target deformation model, and mean() is an averaging function.

9. A device for establishing a deformable dental model, It is characterized in that include: A sample acquisition unit, used to acquire sample data, wherein the sample data includes a plurality of initial tooth models acquired by scanning the oral cavity and a target deformation model corresponding to each initial tooth model obtained by manually processing each initial tooth model; A preprocessing unit is used to obtain a feature tensor corresponding to each initial tooth model, each element of the feature tensor corresponding to each initial tooth model is a truncated signed distance function value TSDF value of each voxel in the cubic space where each initial tooth model is located; A prediction unit, used for inputting the feature tensors corresponding to each initial tooth model into a preset network model to obtain a predicted deformation model corresponding to each initial tooth model; An optimization unit, configured to optimize the preset network model to obtain a tooth mold deformation model according to the target deformation model and the predicted deformation model corresponding to each initial tooth model; The preset network model includes: an encoder component composed of multiple encoders in series structure, a self-attention component, a feature transfer component, a multi-scale analysis component, and a decoder component composed of multiple decoders in series structure; the input of the encoder component is the input of the preset network model, the output of the encoder component is the input of the self-attention component; the output of the self-attention component is the input of the feature transfer component; the output of the feature transfer component is the input of the multi-scale analysis component, the output of the multi-scale analysis component is the input of the decoder component, and the output of the decoder component is the output of the preset network model; Among them, the self-attention component is used to extract non-local information from the feature tensor output by the encoder component to obtain the environmental feature tensor; the feature transfer component processes the output of the self-attention component and transfers the processing result to the multi-scale analysis component; the multi-scale analysis component is used to extract the feature tensor output by the feature transfer component at multiple scales.

Citation Information

Patent Citations

  • Dental orthodontic process prediction method

    CN111265317A