A single tooth segmentation method, device, electronic device and storage medium
Through the combination of two-stage three-dimensional convolutional neural network and watershed segmentation algorithm, the automation and accuracy of single tooth segmentation in oral CBCT images are solved, and efficient single tooth segmentation is achieved, improving the auxiliary clinical treatment effect of three-dimensional visualization of teeth.
Patent Information
- Application Number
- CN202310664805.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-06
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-06-06
AI Technical Summary
The single tooth segmentation method in the existing oral CBCT images has problems such as low automation, low efficiency, high labor cost and poor accuracy and repeatability, especially due to the similar grayscale of the teeth to the alveolar bone and jaw and blurred boundaries of adjacent teeth.
Using a two-stage three-dimensional convolutional neural network segmentation method, the tooth mask is first obtained through the first segmentation model, and then the area of interest is extracted using the tooth mask, and the heat map of a single tooth is generated through the second segmentation model and the center of mass is extracted. The single tooth is further segmented in combination with the watershed segmentation algorithm.
It realizes efficient and accurate single tooth segmentation, avoids misalignment of adjacent teeth, and improves segmentation effect, especially in complex anatomical shapes and broken or damaged teeth, supporting fast and accurate three-dimensional visualization of teeth to assist clinical treatment.
Smart Images

Figure CN116912487B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a single tooth segmentation method, device, electronic device and storage medium. Background Art
[0002] For the treatment of oral diseases, pre-treatment examination, evaluation, treatment planning, and post-treatment review and evaluation are crucial to effective treatment. The emergence and application of CBCT has marked a significant advancement in oral imaging. CBCT can generate high-resolution, three-dimensional images that accurately reflect the three-dimensional structure of oral tissues. By segmenting teeth in oral CBCT images, dentists can visually observe the patient's teeth, better conduct preoperative assessments, and develop scientific treatment plans, thereby improving the effectiveness of oral disease treatment. However, oral CBCT images contain not only teeth but also bone tissue such as the alveolar bone and jawbone. Furthermore, due to the low-dose nature of CBCT, CBCT images are subject to significant noise, which hinders the doctor's ability to extract dental information. Therefore, dentists generally do not use raw CBCT images directly. Instead, they segment teeth from CBCT images to create three-dimensional models, enabling efficient and accurate diagnosis and treatment.
[0003] However, due to the proximity of teeth to the alveolar bone and jawbone, their similar grayscale values, the blurred boundaries between adjacent teeth, and the occlusal state of the upper and lower teeth, fully automated segmentation of individual teeth is extremely difficult. Interactive segmentation is currently the mainstream single-tooth segmentation method, but this semi-automatic segmentation method is not only labor-intensive, inefficient, and labor-intensive, but also relies on the doctor's subjective experience, resulting in poor accuracy and repeatability. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a single tooth segmentation method, device, electronic device, and storage medium, which can efficiently and accurately perform single tooth segmentation.
[0005] In one aspect, an embodiment of the present invention provides a single tooth segmentation method, comprising:
[0006] Acquiring a CBCT image to be identified and preprocessing the CBCT image;
[0007] Analyzing the preprocessed CBCT image using a first segmentation model to obtain a tooth mask; wherein the first segmentation model is trained using a three-dimensional convolutional neural network using a set of CBCT images labeled with tooth masks;
[0008] The three-dimensional convolutional neural network includes a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module include multiple submodules with corresponding structures; the submodules of the feature extraction module and the feature fusion module are sequentially connected to form a U-shaped network, and the submodules with corresponding structures in the feature extraction module and the feature fusion module are jump-connected; the submodules of the feature extraction module are used for downsampling, and the submodules of the feature fusion module are used for upsampling;
[0009] extracting a tooth region of interest from the CBCT image based on the tooth mask;
[0010] Analyzing the tooth region of interest using a second segmentation model to obtain thermal maps of different individual teeth; wherein the second segmentation model is trained based on the three-dimensional convolutional neural network using a set of CBCT images with individually labeled individual teeth;
[0011] Extracting the centroid coordinates of each of the single teeth according to the thermal map;
[0012] The centroid coordinates are used as markers, and a watershed segmentation process based on the markers is performed on the tooth mask to obtain a segmented image of each of the single teeth.
[0013] Optionally, preprocessing the CBCT image includes:
[0014] Based on a preset percentile range, limiting the HU value of the CBCT image;
[0015] The CBCT image after the restriction processing is standardized according to the mean and standard deviation of the HU value after the restriction processing.
[0016] Optionally, the method further comprises:
[0017] Determine a first training sample based on a CBCT image set with labeled tooth masks;
[0018] Based on the first training sample and in combination with the constraints of a first loss function, the three-dimensional convolutional neural network is trained for mask segmentation, and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the first segmentation model; wherein the first loss function includes a cross entropy loss function and a Dice loss function.
[0019] Optionally, the method further comprises:
[0020] Determining a second training sample based on a set of CBCT images of individually labeled individual teeth, wherein each of the individual teeth is labeled with a different color and a different label value greater than 0, and the label value of oral tissues other than teeth is 0;
[0021] Based on the second training sample and in combination with the constraints of the second loss function, the three-dimensional convolutional neural network is trained for tooth segmentation, and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the second segmentation model; wherein, the second loss function includes a mean square error loss function.
[0022] Optionally, the data processing step of the three-dimensional convolutional neural network includes:
[0023] Get input data;
[0024] The input data is subjected to multi-scale downsampling by the feature extraction module to obtain downsampled data; wherein the input of each submodule of the feature extraction module is the output of the previous submodule;
[0025] The downsampled data is multi-scale upsampled by the feature fusion module to obtain output data; wherein the input of each submodule of the feature fusion module is the output of the previous submodule and the output of the submodule corresponding to the jump connection structure in the feature extraction module.
[0026] Optionally, extracting the centroid coordinates of each of the single teeth according to the heat map includes:
[0027] performing interpolation processing on the thermal map, and scaling the thermal map to match the size of the tooth region of interest;
[0028] Based on a preset threshold, performing non-maximum suppression processing on the interpolated heat map to obtain separation results of different single teeth;
[0029] Based on the separation result, the centroid coordinates of each of the single teeth are extracted.
[0030] Optionally, using the centroid coordinates as markers, performing watershed segmentation processing on the tooth mask based on the markers to obtain segmented images of each of the single teeth, including:
[0031] performing distance change processing on the tooth mask to obtain a distance transformation map;
[0032] Constructing a foreground label map with an initial voxel value of 0 and the same specifications as the distance transform map, and setting the voxel at the position of the centroid coordinate of the foreground label map to 1;
[0033] Negating the distance transformation map, using the negated distance transformation map as a topological landform, and using each connected domain in the foreground marker map as a basin bottom of the topological landform; and determining the watershed of the basin based on a watering operation on the basin bottom;
[0034] Each single tooth in the CBCT image is segmented according to the watershed to obtain a segmented image of each single tooth.
[0035] In another aspect, an embodiment of the present invention provides a single tooth segmentation device, comprising:
[0036] The first module is used to obtain a CBCT image to be identified and preprocess the CBCT image;
[0037] A second module is configured to analyze the preprocessed CBCT image using a first segmentation model to obtain a tooth mask; wherein the first segmentation model is trained using a three-dimensional convolutional neural network using a set of CBCT images labeled with tooth masks;
[0038] The three-dimensional convolutional neural network includes a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module include multiple submodules with corresponding structures; the submodules of the feature extraction module and the feature fusion module are sequentially connected to form a U-shaped network, and the submodules with corresponding structures in the feature extraction module and the feature fusion module are jump-connected; the submodules of the feature extraction module are used for downsampling, and the submodules of the feature fusion module are used for upsampling;
[0039] A third module is configured to extract a tooth region of interest from the CBCT image based on the tooth mask;
[0040] a fourth module for analyzing the tooth region of interest using a second segmentation model to obtain thermal maps of different individual teeth; wherein the second segmentation model is trained based on the three-dimensional convolutional neural network using a set of CBCT images with individually labeled individual teeth;
[0041] A fifth module is used to extract the centroid coordinates of each of the single teeth according to the thermal map;
[0042] The sixth module is configured to use the centroid coordinates as markers to perform watershed segmentation processing on the tooth mask based on the markers to obtain segmented images of each of the single teeth.
[0043] Optionally, the system further includes:
[0044] The seventh module is used to determine a first training sample based on a set of CBCT images with labeled tooth masks; based on the first training sample and in combination with the constraints of a first loss function, the three-dimensional convolutional neural network is trained for mask segmentation; and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the first segmentation model; wherein the first loss function includes a cross entropy loss function and a Dice loss function.
[0045] Optionally, the system further includes:
[0046] The eighth module is used to determine a second training sample based on a set of CBCT images of single teeth that have been individually labeled; wherein each of the single teeth is labeled with a different color and a different label value greater than 0, and the label value of other oral tissues other than teeth is 0; based on the second training sample, and in combination with the constraints of the second loss function, the three-dimensional convolutional neural network is trained for tooth segmentation, and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the second segmentation model; wherein the second loss function includes a mean square error loss function.
[0047] Optionally, the system further includes:
[0048] The ninth module is used to store the first segmentation model and / or the second segmentation model and perform data processing of the three-dimensional convolutional neural network, including:
[0049] Get input data;
[0050] The input data is subjected to multi-scale downsampling by the feature extraction module to obtain downsampled data; wherein the input of each submodule of the feature extraction module is the output of the previous submodule;
[0051] The downsampled data is multi-scale upsampled by the feature fusion module to obtain output data; wherein the input of each submodule of the feature fusion module is the output of the previous submodule and the output of the submodule corresponding to the jump connection structure in the feature extraction module.
[0052] In the first segmentation model, the input data corresponds to the preprocessed CBCT image, and the output data corresponds to the tooth mask; in the second segmentation model, the input data corresponds to the relevant image of the tooth region of interest, and the output data corresponds to the thermal map of different single teeth.
[0053] In another aspect, an embodiment of the present invention provides an electronic device including a processor and a memory;
[0054] The memory is used to store programs;
[0055] The processor executes the program to implement the above method.
[0056] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the above method.
[0057] The present invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the above method.
[0058] The embodiment of the present invention first obtains a CBCT image to be identified and preprocesses the CBCT image; uses a first segmentation model to analyze the preprocessed CBCT image to obtain a tooth mask; wherein the first segmentation model is obtained by training a CBCT image set with annotated tooth masks according to a three-dimensional convolutional neural network; wherein the three-dimensional convolutional neural network includes a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module include a plurality of submodules corresponding in structure; each of the submodules of the feature extraction module and the feature fusion module is sequentially connected to form a U-shaped network, and the submodules corresponding in structure in the feature extraction module and the feature fusion module are connected to form a U-shaped network The modules are jump-connected; each submodule of the feature extraction module is used for downsampling, and each submodule of the feature fusion module is used for upsampling; based on the tooth mask, the tooth region of interest is extracted from the CBCT image; the second segmentation model is used to analyze the tooth region of interest to obtain heat maps of different single teeth; wherein, the second segmentation model is obtained by training the 3D convolutional neural network through a set of CBCT images with individually labeled single teeth; according to the heat map, the centroid coordinates of each single tooth are extracted; using the centroid coordinates as labels, the tooth mask is subjected to watershed segmentation processing based on the labels to obtain segmented images of each single tooth. The embodiment of the present invention implements a two-stage segmentation method based on two 3D convolutional neural networks with different training biases. On the basis of the tooth mask automatically segmented by the first segmentation model, the tooth heat map is automatically obtained by the second segmentation model, and then the tooth centroid is detected and the single tooth is further segmented by the labeled watershed algorithm. Compared with directly segmenting the single tooth instance, it effectively avoids the misclassification between adjacent teeth, and the segmentation effect of the single tooth is better. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 A schematic flow chart of a single tooth segmentation method provided by an embodiment of the present invention;
[0061] Figure 2 A schematic diagram of the flow of a marker-based watershed algorithm provided in an embodiment of the present invention;
[0062] Figure 3 A schematic diagram of the distance transformation process provided by an embodiment of the present invention;
[0063] Figure 4 A schematic diagram of the overall framework process of single tooth segmentation provided by an embodiment of the present invention;
[0064] Figure 5 A schematic structural diagram of a single tooth segmentation device provided by an embodiment of the present invention;
[0065] Figure 6 A schematic diagram of the framework of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0067] To facilitate understanding of the technical principles of the present invention, relevant technical terms that may appear in the embodiments of the present invention are explained.
[0068] CBCT: Cone Beam Computerized Tomography
[0069] HU value: reflects the degree of tissue absorption of X-rays (Hounsfield Unit)
[0070] ReLu: An activation function of a neural network. Input values less than 0 are set to 0, and input values greater than 0 remain unchanged.
[0071] Region of Interest: In machine vision and image processing, a region of interest (ROI) is an image that is outlined in the image being processed using a box, circle, ellipse, or irregular polygon. In this embodiment, this refers to the image of the smallest rectangular region that completely encloses the segmented upper and lower tooth masks.
[0072] On the one hand, if Figure 1 As shown, an embodiment of the present invention provides a single tooth segmentation method, comprising:
[0073] S100, acquiring a CBCT image to be identified and preprocessing the CBCT image;
[0074] It should be noted that, in some embodiments, preprocessing the CBCT image may include: limiting the HU value of the CBCT image based on a preset percentile range; and standardizing the CBCT image after the limiting processing according to the mean and standard deviation of the HU value after the limiting processing.
[0075] In some specific embodiments, to eliminate the effects of abnormal extreme points in CBCT images caused by metal artifacts, the 0.5% and 99.5% percentile HU values of the CBCT images can be calculated and obtained, limiting the image HU values to the 0.5% to 99.5% percentile HU value range. The HU value range of the processed CBCT data is approximately 0 to 4000. Furthermore, z-score standard normalization can be used to narrow the HU value range of the data preprocessed in Step 2. The z-score standard normalization formula is: where preCT is the voxel HU value before z-score standard normalization, mean is the voxel mean HU value, std is the voxel standard deviation HU value, and postCT is the voxel HU value after z-score standard normalization.
[0076] S200, using a first segmentation model, analyzing the preprocessed CBCT image to obtain a tooth mask;
[0077] It should be noted that the first segmentation model is obtained by training a three-dimensional convolutional neural network using a set of CBCT images with labeled tooth masks.
[0078] In some embodiments, the method may further include: determining a first training sample based on a set of CBCT images with labeled tooth masks; performing mask segmentation training on the three-dimensional convolutional neural network based on the first training sample and in combination with the constraints of a first loss function, and adjusting the three-dimensional convolutional neural network based on the training results to obtain the first segmentation model; wherein the first loss function includes a cross entropy loss function and a Dice loss function.
[0079] In some specific embodiments, the three-dimensional convolutional neural network model for segmenting the upper and lower teeth masks is trained using a cross entropy loss function and a Dice loss function for constraints.
[0080] Cross entropy loss function: Calculate the loss of each voxel one by one, and then calculate the average loss of all voxels to get the loss of the entire image. The cross entropy loss formula for a single voxel is: Where M represents the number of categories that need to be segmented in the image; y c represents the true label value of the voxel; P c represents the probability that the predicted voxel belongs to category c; the cross entropy loss formula for the entire image is: Where N represents the number of voxels in the image, CE_Loss i represents the cross entropy loss of the i-th voxel;
[0081] Dice loss function: The formula is Where N represents the number of voxels in the image; t i represents the true label value of the i-th voxel; y i represents the predicted probability value of the i-th voxel;
[0082] The total loss function is: Total_Loss = Total_CE_Loss + Dice_Loss.
[0083] In some specific embodiments, the preprocessed CBCT image is input into a built and trained three-dimensional convolutional neural network model for segmentation, and the masks of the upper teeth and / or lower teeth (tooth masks) can be segmented.
[0084] S300, extracting a tooth region of interest from the CBCT image based on the tooth mask;
[0085] In some specific embodiments, the upper and / or lower teeth mask (tooth mask) is segmented according to step S200, and the upper and / or lower teeth regions of interest are extracted from the CBCT image. In machine vision and image processing, the area to be processed is outlined in the processed image using a box, circle, ellipse, irregular polygon, etc., which is called a region of interest. Because the upper and lower teeth account for a very small proportion of the entire oral CBCT image, compared to directly feeding the entire CBCT image into the model training, the efficiency and accuracy of the model training can be significantly improved by extracting the upper and lower teeth regions of interest and then feeding the regions of interest into the model for training.
[0086] In some specific embodiments, before inputting the tooth ROI into the second segmentation model, the extracted upper and lower tooth ROIs can be scaled to a preset size (e.g., 96×128×128) by interpolation, so as to conform to the data processing format of the model and achieve efficient and convenient processing.
[0087] S400, using a second segmentation model, analyzing the tooth region of interest to obtain heat maps of different single teeth;
[0088] It should be noted that the second segmentation model is obtained by training the three-dimensional convolutional neural network using a set of CBCT images of individually labeled single teeth;
[0089] In some embodiments, the method may further include: determining a second training sample based on a set of CBCT images in which single teeth have been individually labeled; wherein each of the single teeth is labeled with a different color and a different label value greater than 0, and the label value of other oral tissues other than teeth is 0; based on the second training sample, and in combination with the constraints of a second loss function, the three-dimensional convolutional neural network is trained for tooth segmentation, and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the second segmentation model; wherein the second loss function includes a mean square error loss function.
[0090] In some embodiments, the 3D convolutional neural network model for generating a single tooth heat map is trained using a mean squared loss function, where:
[0091] The formula for the mean square loss function is: Where N represents the number of voxels in the image; t i represents the true label value of the i-th voxel; y i represents the predicted probability value of the i-th voxel.
[0092] In some specific embodiments, the upper and lower teeth regions of interest scaled in the aforementioned step may be input into the constructed second segmentation model to generate a heat map of a single tooth.
[0093] It should be noted that the three-dimensional convolution application network used in steps S200 and S400 may include a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module include multiple sub-modules corresponding to the structure; the sub-modules of the feature extraction module and the feature fusion module are connected in sequence to form a U-shaped network, and the sub-modules corresponding to the structure in the feature extraction module and the feature fusion module are jump-connected; the sub-modules of the feature extraction module are used for downsampling, and the sub-modules of the feature fusion module are used for upsampling.
[0094] In some embodiments, the data processing steps of the three-dimensional convolutional neural network may include: obtaining input data; performing multi-scale downsampling on the input data through the feature extraction module to obtain downsampled data; wherein the input of each submodule of the feature extraction module is the output of the previous submodule; performing multi-scale upsampling on the downsampled data through the feature fusion module to obtain output data; wherein the input of each submodule of the feature fusion module is the output of the previous submodule and the output of the submodule corresponding to the jump connection structure in the feature extraction module.
[0095] It is easy to understand that in the first segmentation model, the input data corresponds to the preprocessed CBCT image, and the output data corresponds to the tooth mask; while in the second segmentation model, the input data corresponds to the relevant image of the tooth region of interest, and the output data corresponds to the thermal map of different single teeth.
[0096] To facilitate understanding of the technical principles, the three-dimensional convolutional neural network used in the embodiments of the present invention is described in detail below.
[0097] First, it should be noted that the CBCT image set used for model training in this embodiment of the present invention was prepared by personnel with a background in dentistry who individually annotated individual teeth in the CBCT images. Mask images of the upper and lower teeth were generated and used as label data for model training. Furthermore, different teeth were labeled with different colors and values greater than 0, while the jawbone and other oral tissues were labeled with 0 to distinguish between individual teeth and non-tooth tissue. CBCT images with different annotation emphases were selected based on different training priorities.
[0098] The model description of the 3D convolutional neural network is as follows:
[0099] The model structure consists of three parts: feature extraction (downsampling), feature fusion (upsampling) and skip connection:
[0100] Feature extraction part: It includes four downsampling modules with the same network structure. Each downsampling module consists of three submodules in the forward propagation direction: the first submodule includes a 3D convolution layer with a convolution kernel size of and a stride of 1, a 3D batch normalization layer, and a ReLu nonlinear activation layer; the second submodule is the same as the first submodule; the third submodule is a 3D maximum pooling layer with a convolution kernel size of and a stride of 2;
[0101] The feature map output by the fourth downsampling module of the feature extraction part is used as the input of the feature fusion part;
[0102] Feature fusion part: It includes four upsampling modules with the same network structure. Each downsampling module includes three sub-modules in the direction of forward propagation: the first sub-module includes a three-dimensional convolution layer with a convolution kernel size of and a stride of 1, a three-dimensional batch normalization layer, and a ReLu nonlinear activation layer; the second sub-module is the same as the first sub-module; the third sub-module is a three-dimensional upsampling layer with a convolution kernel size of and a stride of 2.
[0103] Jump connection part: jump connection is performed between the output of each module in the feature extraction part and the input of each module in the corresponding feature fusion part;
[0104] The feature map output by the fourth upsampling module in the feature fusion section passes through three submodules: the first submodule includes a 3D convolutional layer with a kernel size of and a stride of 1, a 3D batch normalization layer, and a Reinforced Lu (ReLu) nonlinear activation layer; the second submodule is identical to the first; and the third submodule is a 3D convolutional layer with a kernel size of and a stride of 1. The final segmentation result is obtained after passing through these three submodules.
[0105] Then, based on the above-mentioned three-dimensional convolutional neural network, combined with training samples of different labels and different loss functions, a first segmentation model and a second segmentation model can be trained respectively, thereby achieving the segmentation of tooth masks and heat maps of different individual teeth.
[0106] S500, extracting the centroid coordinates of each of the single teeth according to the thermal map;
[0107] It should be noted that, in some embodiments, step S500 may include: interpolating the heat map to scale the heat map to match the size of the tooth region of interest; performing non-maximum suppression processing on the interpolated heat map based on a preset threshold to obtain separation results of different single teeth; and extracting the center of mass coordinates of each single tooth based on the separation results.
[0108] In some embodiments, the upper and lower teeth regions of interest may be filled to the size of the original oral CBCT image using a value of 0, and then the centroid coordinates of each single tooth may be extracted through the separation results based on the filled image.
[0109] In some specific embodiments, the extraction of centroid coordinates can be achieved through the following steps:
[0110] Step 1: Use interpolation to scale the heat map extracted in step 7 back to the size of the original upper and lower teeth sensory interest regions;
[0111] Step 2: Set a reasonable threshold and use non-maximum suppression to retain the local maximum in the heat map and suppress the non-maximum to separate different single teeth.
[0112] Step 3: Use 0 value to fill the upper and lower teeth ROIs to the size of the original oral CBCT image;
[0113] Step 4: Extract the centroid coordinates of a single tooth area as the key point of the tooth.
[0114] S600, using the centroid coordinates as markers, performing watershed segmentation processing on the tooth mask based on the markers to obtain segmented images of each of the single teeth;
[0115] It should be noted that, in some embodiments, step S600 may include: performing distance change processing on the tooth mask to obtain a distance transformation map; constructing a foreground marker map with an initial voxel value of 0 and the same specifications as the distance transformation map, and setting the voxel at the position of the centroid coordinate of the foreground marker map to 1; negating the distance transformation map, using the negative distance transformation map as the topological landform, and using each connected domain in the foreground marker map as the bottom of the basin in the topological landform; determining the watershed of the basin based on the watering operation on the bottom of the basin; and segmenting each single tooth in the CBCT image based on the watershed to obtain a segmented image of each single tooth.
[0116] In some specific embodiments, such as Figure 2 As shown in the figure, the segmentation of a single tooth using the marker-based watershed algorithm can be achieved by the following steps:
[0117] Step 1: Perform distance transformation on the upper teeth (or lower teeth) mask image (i.e., tooth mask) obtained by segmentation in step S200 to obtain a distance transformation map; Figure 3 As shown, in the upper (or lower) tooth image, the foreground value is 1 and the background value is 0; the distance transform calculates the distance from the non-zero voxel in the image to the 0-value voxel (background). The farther away from the background, the larger the calculated value.
[0118] Step 2: Create a foreground marker map of the same shape and size as the distance transform map, with all voxel initial values set to 0; then set the voxels in the foreground marker map corresponding to the key point coordinates (center of mass coordinates) of the single tooth obtained in step S500 to 1;
[0119] Step 3: Search for 4-connected domains in the matrix foreground marker map, and assign different values to each 4-connected domain in sequence; for example, for the third 4-connected domain found, set the value of the voxel in the connected domain to 3;
[0120] Step 4: Negate the distance transform map and treat it as a topological landform in geodesy. The grayscale value of each point represents the elevation of that point. Treat each connected domain in the foreground marker map as the bottom of a basin. Water is poured from the bottom of the basin until the basins attributed to different connected domains meet on the watershed line.
[0121] Step 5: Use the watershed obtained in Step 4 to segment each single tooth in the upper (or lower) tooth image.
[0122] It should be noted that, in some specific embodiments, the embodiments of the present invention implement a two-stage segmentation method through a two-stage segmentation model (a first-stage model and a second-stage model), wherein the first-stage model is used to analyze the preprocessed CBCT image to obtain the tooth mask. The first-stage model is trained according to a three-dimensional convolutional neural network through a set of CBCT images with labeled tooth masks (the same as the first segmentation model of the aforementioned embodiment), that is, the first-stage model integration implements the step principle implemented by S200 and its related specific embodiments; and the second-stage model can be implemented by integrating the aforementioned second segmentation model (used to analyze the tooth region of interest to obtain thermal maps of different single teeth; according to a three-dimensional convolutional neural network, it is trained through a set of CBCT images with individually labeled single teeth) and a marker-based watershed algorithm, that is, the second-stage model integration implements the step principle implemented by S400 to S600 and its related specific embodiments.
[0123] Specifically, in order to fully explain the technical principles of the embodiments of the present invention, the above-mentioned overall process steps are further explained below in combination with some specific embodiments. It should be understood that the following is an explanation of the present invention and cannot be regarded as a limitation of the present invention.
[0124] like Figure 4 As shown, Figure 4 In [1], Input represents the oral CBCT image to be segmented; 3D CNN represents a three-dimensional convolutional neural network; Tooth Mask represents the segmented upper or lower tooth mask (i.e., tooth mask); Tooth ROI represents the region of interest of the upper or lower teeth extracted from the CBCT image (i.e., tooth region of interest); Heatmap represents the heat map of a single tooth; MWT represents a marker-controlled watershed transform; and Output represents the image of the segmented single tooth. The generation method proposed in the embodiment of the present invention includes the following steps:
[0125] Step 1: Preprocessing of oral CBCT images; specifically including:
[0126] Step 1.1: Obtain the 0.5% and 99.5% Hounsfield Unit (HU) values of the oral CBCT image and limit the HU value of the image to the range of 0.5% to 99.5%;
[0127] Step 1.2: Perform z-score standard normalization on the image obtained in step 1.1.
[0128] Step 2: Use the 3D convolutional neural network model to segment the CBCT image preprocessed in step 1.2, segment the upper and lower teeth separately, and then extract the upper and lower teeth regions of interest based on the segmentation results;
[0129] Step 3: The upper and lower tooth regions of interest extracted in step 2 are input into the 3D convolutional neural network model to generate a heat map of the centroid of a single tooth;
[0130] Step 4: Extract the centroid coordinates of a single tooth from the heat map generated in step 3 as the key point of the tooth;
[0131] Step 5: Use the key points of the teeth as markers and use the marker-based watershed segmentation algorithm to segment individual teeth from the upper and lower tooth masks obtained in step 2.
[0132] In summary, the method of the present invention utilizes a two-stage segmentation approach. Based on automatically segmented upper and lower tooth masks, tooth centroids are automatically detected and further segmented using the labeled watershed algorithm. Compared to direct segmentation of individual teeth, this method effectively avoids misclassification between adjacent teeth, resulting in superior single-tooth segmentation. Furthermore, compared to the best-performing single-tooth segmentation methods currently available, the model of the present invention achieves superior segmentation results for cases with complex dental anatomies, better segmenting broken or missing teeth. This facilitates three-dimensional visualization of teeth within the patient's mouth and better assists in clinical oral treatment and surgery. Furthermore, the present invention utilizes the following evaluation metrics: Dice similarity coefficient, IoU (Intersection over Union), ASSD (Average Symmetric Surface Distance), Precision, Recall, and HD (Hausdorff Distance), achieving excellent results on experimental datasets. The present invention utilizes a three-dimensional convolutional neural network combined with the labeled watershed algorithm to implement a fully automated single-tooth segmentation process for oral CBCT images, enabling rapid, accurate, and efficient automatic segmentation of individual teeth.
[0133] On the other hand, Figure 5As shown, an embodiment of the present invention provides a single tooth segmentation device 700, comprising: a first module 710, for acquiring a CBCT image to be identified and preprocessing the CBCT image; a second module 720, for analyzing the preprocessed CBCT image using a first segmentation model to obtain a tooth mask; wherein the first segmentation model is obtained by training a CBCT image set with annotated tooth masks according to a three-dimensional convolutional neural network; wherein the three-dimensional convolutional neural network comprises a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module comprise a plurality of submodules corresponding in structure; the submodules of the feature extraction module and the feature fusion module are sequentially connected to form a U-shaped network, and the submodules corresponding in structure in the feature extraction module and the feature fusion module are sequentially connected to form a U-shaped network, and ... The submodules are jump-connected; each submodule of the feature extraction module is used for downsampling, and each submodule of the feature fusion module is used for upsampling; a third module 730 is used to extract the tooth region of interest from the CBCT image based on the tooth mask; a fourth module 740 is used to use a second segmentation model to analyze the tooth region of interest to obtain thermal maps of different single teeth; wherein, the second segmentation model is obtained by training a set of CBCT images of individually labeled single teeth based on the three-dimensional convolutional neural network; a fifth module 750 is used to extract the centroid coordinates of each single tooth based on the thermal map; a sixth module 760 is used to use the centroid coordinates as a mark to perform watershed segmentation processing on the tooth mask based on the mark to obtain a segmented image of each single tooth.
[0134] The contents of the method embodiments of the present invention are all applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0135] like Figure 6 As shown, another aspect of the embodiment of the present invention further provides an electronic device 800, including a processor 810 and a memory 820;
[0136] The memory 820 is used to store programs;
[0137] The processor 810 executes the program to implement the above method.
[0138] The contents of the method embodiments of the present invention are all applicable to the electronic device embodiments. The functions specifically implemented by the electronic device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0139] Another aspect of an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the above method.
[0140] The contents of the method embodiments of the present invention are all applicable to the computer-readable storage medium embodiments. The functions specifically implemented by the computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0141] The present invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the above method.
[0142] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0143] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It will also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, a person skilled in the art will be able to implement the present invention as set forth in the claims using ordinary skill without undue experimentation. It will also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0144] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0145] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution apparatus, device, or apparatus (e.g., a computer-based apparatus, a device including a processor, or other apparatus that can fetch instructions from and execute instructions on an instruction execution apparatus, device, or apparatus). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution apparatus, device, or apparatus.
[0146] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0147] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0148] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0149] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
[0150] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A single tooth segmentation method, characterized in that: include: Acquiring a CBCT image to be identified and preprocessing the CBCT image; Analyzing the preprocessed CBCT images using a first segmentation model to obtain a tooth mask; wherein the first segmentation model is trained using a three-dimensional convolutional neural network using a set of CBCT images labeled with tooth masks; The three-dimensional convolutional neural network includes a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module include multiple submodules with corresponding structures; the submodules of the feature extraction module and the feature fusion module are sequentially connected to form a U-shaped network, and the submodules with corresponding structures in the feature extraction module and the feature fusion module are jump-connected; the submodules of the feature extraction module are used for downsampling, and the submodules of the feature fusion module are used for upsampling; extracting a tooth region of interest from the CBCT image based on the tooth mask; Analyzing the tooth region of interest using a second segmentation model to obtain thermal maps of different individual teeth; wherein the second segmentation model is trained based on the three-dimensional convolutional neural network using a set of CBCT images with individually labeled individual teeth; Extracting the centroid coordinates of each of the single teeth according to the thermal map; The step of extracting the centroid coordinates of each of the individual teeth according to the heat map includes: performing interpolation processing on the thermal map, and scaling the thermal map to match the size of the tooth region of interest; Based on a preset threshold, performing non-maximum suppression processing on the interpolated heat map to obtain separation results of different single teeth; Based on the separation results, extracting the centroid coordinates of each of the single teeth; The centroid coordinates are used as markers, and a watershed segmentation process based on the markers is performed on the tooth mask to obtain a segmented image of each of the single teeth.
2. A single tooth segmentation method according to claim 1, characterized in that: The preprocessing of the CBCT image includes: Based on a preset percentile range, limiting the HU value of the CBCT image; The CBCT image after the restriction processing is standardized according to the mean and standard deviation of the HU value after the restriction processing.
3. A single tooth segmentation method according to claim 1, characterized in that: The method further comprises: Determine a first training sample based on a set of CBCT images with labeled tooth masks; Based on the first training sample and in combination with the constraints of a first loss function, the three-dimensional convolutional neural network is trained for mask segmentation, and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the first segmentation model; wherein the first loss function includes a cross entropy loss function and a Dice loss function.
4. A single tooth segmentation method according to claim 1, characterized in that: The method further comprises: Determining a second training sample based on a set of CBCT images of individually labeled individual teeth, wherein each of the individual teeth is labeled with a different color and a different label value greater than 0, and the label value of oral tissues other than teeth is 0; Based on the second training sample and in combination with the constraints of the second loss function, the three-dimensional convolutional neural network is trained for tooth segmentation, and based on the training results, the three-dimensional convolutional neural network is adjusted to obtain the second segmentation model; wherein, the second loss function includes a mean square error loss function.
5. A single tooth segmentation method according to any one of claims 3 or 4, characterized in that: The data processing steps of the three-dimensional convolutional neural network include: Get input data; The input data is subjected to multi-scale downsampling by the feature extraction module to obtain downsampled data; wherein the input of each submodule of the feature extraction module is the output of the previous submodule; The downsampled data is multi-scale upsampled by the feature fusion module to obtain output data; wherein the input of each submodule of the feature fusion module is the output of the previous submodule and the output of the submodule corresponding to the jump connection structure in the feature extraction module.
6. A single tooth segmentation method according to claim 1, characterized in that: The method of using the centroid coordinates as markers and performing watershed segmentation on the tooth mask based on the markers to obtain segmented images of each of the single teeth includes: performing distance change processing on the tooth mask to obtain a distance transformation map; Constructing a foreground label map with an initial voxel value of 0 and the same specifications as the distance transform map, and setting the voxel at the position of the centroid coordinate of the foreground label map to 1; Negating the distance transformation map, using the negated distance transformation map as a topological landform, and using each connected domain in the foreground marker map as a basin bottom of the topological landform; and determining the watershed of the basin based on a watering operation on the basin bottom; Each single tooth in the CBCT image is segmented according to the watershed to obtain a segmented image of each single tooth.
7. A single tooth segmentation device, characterized in that: include: The first module is used to obtain a CBCT image to be identified and preprocess the CBCT image; A second module is configured to analyze the preprocessed CBCT image using a first segmentation model to obtain a tooth mask; wherein the first segmentation model is trained using a three-dimensional convolutional neural network using a set of CBCT images labeled with tooth masks; The three-dimensional convolutional neural network includes a feature extraction module and a feature fusion module, and the feature extraction module and the feature fusion module include multiple submodules with corresponding structures; the submodules of the feature extraction module and the feature fusion module are sequentially connected to form a U-shaped network, and the submodules with corresponding structures in the feature extraction module and the feature fusion module are jump-connected; the submodules of the feature extraction module are used for downsampling, and the submodules of the feature fusion module are used for upsampling; A third module is configured to extract a tooth region of interest from the CBCT image based on the tooth mask; a fourth module for analyzing the tooth region of interest using a second segmentation model to obtain thermal maps of different individual teeth; wherein the second segmentation model is trained based on the three-dimensional convolutional neural network using a set of CBCT images with individually labeled individual teeth; A fifth module is used to extract the centroid coordinates of each of the single teeth according to the thermal map; The step of extracting the centroid coordinates of each of the individual teeth according to the heat map includes: performing interpolation processing on the thermal map, and scaling the thermal map to match the size of the tooth region of interest; Based on a preset threshold, performing non-maximum suppression processing on the interpolated heat map to obtain separation results of different single teeth; Based on the separation results, extracting the centroid coordinates of each of the single teeth; The sixth module is configured to use the centroid coordinates as markers to perform watershed segmentation processing on the tooth mask based on the markers to obtain segmented images of each of the single teeth.
8. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional CBCT tooth image segmentation method based on feature transformation
CN113744275A
CBCT image-based tooth and alveolar bone segmentation method and system
CN115661141A