Two-stage CBCT tooth segmentation method based on multi-direction multi-scale feature fusion network

By adopting a two-stage tooth segmentation method with a multi-directional multi-scale feature fusion network and a central focus mechanism in CBCT images, the problem of inaccurate tooth segmentation is solved, and higher precision tooth segmentation and clearer boundary segmentation effects are achieved.

CN120107594AActive Publication Date: 2025-06-06TAIYUAN UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510270379.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and segment teeth in CBCT images, especially in the case of dense tooth arrangement, crowded or metal artifacts, resulting in inaccurate segmentation results.

Method used

The two-stage CBCT tooth segmentation method based on multi-directional multi-scale feature fusion network is adopted. By constructing a multi-directional multi-scale feature fusion module M2FF-Block and a spatial focus module SFM, combined with the central focus mechanism and post-processing algorithm, the fine segmentation of teeth is achieved.

Benefits of technology

It improves the accuracy of tooth positioning and segmentation, enhances the segmentation effect of the central tooth boundary, reduces adhesion of adjacent teeth, and significantly improves the overall segmentation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107594A_ABST
    Figure CN120107594A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of tooth image segmentation, and particularly relates to a two-stage CBCT tooth segmentation method based on a multi-direction multi-scale feature fusion network, and the method comprises the following steps: obtaining an oral cavity CBCT data set, and carrying out the preprocessing; a multi-direction and multi-scale feature fusion module M2FF-Block is constructed; based on the M2FF-Block, respectively constructing a multi-direction multi-scale feature fusion network M2FF-UNet of the two stages; tooth recognition is carried out in a global perspective to obtain the position of the center of mass of each tooth; cutting the original CBCT image according to the centroid position to obtain tooth block data corresponding to each tooth, and performing refined segmentation on a single tooth block by combining a center focusing mechanism; and constructing a post-processing algorithm suitable for a tooth segmentation problem, and correcting a segmentation result to obtain a final segmentation result. According to the method, the inherent multi-direction and multi-scale characteristics of the teeth can be effectively utilized, and the balance between the context information capture and the calculation efficiency is also realized, so that the positioning and segmentation of the teeth are more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tooth image segmentation, and in particular relates to a two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network. Background Art

[0002] As an advanced three-dimensional imaging technology, cone beam computed tomography (CBCT) technology has been widely used in stomatology. Compared with traditional two-dimensional images, CBCT provides higher-resolution three-dimensional images, helping doctors to more accurately evaluate teeth, bones and surrounding tissues, and providing more comprehensive information support for clinical diagnosis and treatment planning. Therefore, high-precision tooth recognition and segmentation technology for CBCT images has become the basis for accurate diagnosis and personalized treatment. However, the natural arrangement characteristics of human teeth may cause dense crowding or overlap between adjacent teeth, and the tooth boundaries become blurred, making it difficult to identify the outline of each tooth. In addition, in actual applications, due to the presence of metal objects such as metal restorations and orthodontic brackets in the mouths of some patients, these metal materials will produce serious artifacts in CBCT images, resulting in a significant decrease in image quality, making it difficult for doctors to accurately identify and evaluate the tooth structure of the affected area, increasing the difficulty of diagnosis and treatment planning.

[0003] At present, research on tooth segmentation in CBCT images has made some progress. Traditional image processing methods mainly rely on edge detection, threshold segmentation and other technologies, but it is often difficult to obtain good segmentation results when processing images with complex structures or large noise. In recent years, with the rapid development of deep learning technology, especially the widespread application of convolutional neural networks (CNN), through its powerful feature learning ability, it can automatically learn multi-level features in the image, better understand the image content, more accurately identify complex structures, and thus achieve high-precision segmentation.

[0004] However, the current deep learning-based methods still have the following shortcomings: existing methods are difficult to capture the inherent directional characteristics of teeth, that is, the morphology and structure of teeth vary in different spatial directions, which makes it difficult to accurately identify and segment teeth. At the same time, the current two-stage instance segmentation method fails to pay enough attention to the segmentation of the central teeth when performing fine segmentation of tooth blocks, resulting in the adhesion of adjacent teeth in the segmentation results. Existing technical solutions are difficult to achieve high-quality segmentation of densely arranged, crowded, missing, misplaced teeth, and metal artifact areas. Summary of the invention

[0005] In view of the technical problems existing in the above-mentioned traditional image processing methods, the present invention provides a two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] The two-stage CBCT tooth segmentation method based on multi-directional and multi-scale feature fusion network includes the following steps:

[0008] S1, obtain oral CBCT data set and preprocess it;

[0009] S2, construct the multi-directional and multi-scale feature fusion module M2FF-Block;

[0010] S3, based on M2FF-Block, construct two stages of multi-directional and multi-scale feature fusion network M2FF-UNet;

[0011] S4, performing tooth recognition from a global perspective to obtain the position of the center of mass of each tooth;

[0012] S5, cropping the original CBCT image according to the centroid position to obtain the tooth block data corresponding to each tooth, and performing fine segmentation of the single tooth block in combination with the center focusing mechanism;

[0013] S6. Construct a post-processing algorithm suitable for the tooth segmentation problem, and correct the segmentation result to obtain the final segmentation result.

[0014] The method for constructing the multi-directional and multi-scale feature fusion module M2FF-Block in S2 is:

[0015] First, a multi-directional and multi-scale feature extraction module (M2FCM) is constructed to systematically extract multi-scale features in different directions. Then, a spatial focusing module (SFM) is constructed to adaptively emphasize the key spatial positions. Then, the input and the output of the two modules are multiplied and added to obtain enhanced multi-directional and multi-scale features.

[0016] The construction method of M2FCM in S3 is:

[0017] The input feature vector is divided into 5 branches. The first branch retains the original feature information without any processing. The second branch uses 3×3×3 deep convolution to extract local spatial features. The remaining three branches use large-core deep convolution blocks of different sizes of 1×1×k, 1×k×1, and k×1×1 to replace k×k×k convolution blocks to extract multi-directional and multi-scale features along different directions. Large-core convolution is used to perceive contextual information, and deep convolution is used to improve computational efficiency. Then, 1×1×1 convolution is used to aggregate these multi-directional and multi-scale features. Finally, all outputs are spliced ​​in the channel dimension to achieve feature fusion.

[0018] The structure of the M2FCM includes:

[0019] Branch 1: No processing;

[0020] Branch 2: convolution layer br2_conv1, the convolution kernel size is 3×3×3, the step size is 1, and the number of groups is the number of input channels;

[0021] Branch three: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer is brb3_conv1, the convolution kernel size is 1×1×7, the step size is 1, the padding is (0, 0, 3), and the number of groups is the number of input channels; the second convolutional layer is br3_conv2, the convolution kernel size is 1×1×9, the step size is 1, the padding is (0, 0, 4), and the number of groups is the number of input channels; the third convolutional layer is br3_conv3, the convolution kernel size is 1×1×11, the step size is 1, the padding is (0, 0, 5), and the number of groups is the number of input channels;

[0022] Branch 4: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer br4_conv1 has a convolution kernel size of 1×7×1, a step size of 1, a padding of (0, 3, 0), and the number of groups is the number of input channels; the second convolutional layer br4_conv2 has a convolution kernel size of 1×9×1, a step size of 1, a padding of (0, 4, 0), and the number of groups is the number of input channels; the third convolutional layer br4_conv3 has a convolution kernel size of 1×11×1, a step size of 1, a padding of (0, 5, 0), and the number of groups is the number of input channels;

[0023] Branch 5: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer br5_conv1 has a convolution kernel size of 7×1×1, a step size of 1, a padding of (3, 0, 0), and the number of groups is the number of input channels; the second convolutional layer br5_conv2 has a convolution kernel size of 9×1×1, a step size of 1, a padding of (4, 0, 0), and the number of groups is the number of input channels; the third convolutional layer br5_conv3 has a convolution kernel size of 11×1×1, a step size of 1, a padding of (5, 0, 0), and the number of groups is the number of input channels;

[0024] Multi-directional and multi-scale feature fusion: The outputs of the third, fourth, and fifth branches are aggregated through a 1×1×1 convolution layer with a convolution kernel of 1 and a stride of 1. The aggregated features are concatenated with the outputs of the first and second branches in the channel dimension.

[0025] The steps of constructing the spatial focusing module SFM are as follows:

[0026] First, a 3D average pooling layer is used to compress spatial information, and a 1×1×1 convolution is used to refine the interaction and integration between channels. In order to capture cross-directional spatial information, 1×1×k, 1×k×1, and k×1×1 convolutions are applied along the height, width, and depth directions, respectively, to focus on feature changes specific to each spatial direction. Finally, an attention map is generated through 1×1×1 convolution and sigmoid activation to guide the focus on spatially important areas in the multi-directional and multi-scale features extracted by M2FCM.

[0027] The structure of the spatial focusing module SFM includes:

[0028] The first pooling layer, pool1, uses average pooling with a window size of 7, a step size of 1, and a padding of 3;

[0029] The first convolutional layer is conv1, the convolution kernel size is 1×1×1, and the stride is 1;

[0030] The second convolutional layer conv2 has a convolution kernel size of 1×1×11, a stride of 1, a padding of (0, 0, 5), and the number of groups is the number of input channels;

[0031] The third convolutional layer conv3 has a convolution kernel size of 1×11×1, a stride of 1, a padding of (0, 5, 0), and the number of groups is the number of input channels;

[0032] The fourth convolutional layer conv4 has a convolution kernel size of 11×1×1, a stride of 1, a padding of (5, 0, 0), and the number of groups is the number of input channels;

[0033] The fifth convolution layer conv5 has a convolution kernel size of 1×1×1 and a stride of 1;

[0034] Finally, the sigmoid function outputs an attention map from 0 to 1.

[0035] The method of the two-stage multi-directional and multi-scale feature fusion network M2FF-Une in S3 is as follows: first, the convolution blocks of the four encoder layers in the UNet architecture are replaced with different numbers of M2FF-Blocks; for the tooth recognition M2FF-UNet in the tooth recognition stage, two decoder branches are constructed, one branch is used to segment the global tooth area, and the other branch is used to regress the offset of each voxel relative to its center of mass; for the tooth segmentation M2FF-UNet in the tooth segmentation stage, two decoder branches are constructed, one branch is used to segment the tooth area within the tooth block, and the other branch is used to segment the boundary area within the tooth block.

[0036] The method for obtaining the position of the centroid of each tooth in S4 is: first obtain the global binary segmentation map and the global voxel centroid offset map output by the tooth recognition subnetwork, then map all predicted foreground tooth voxels to new coordinates according to the offset map, count the pointing frequency of each new coordinate to obtain the density map P, then calculate the minimum distance from each voxel to the voxel with a larger density value, construct the distance map D, and finally select points with density values ​​greater than 60 and distance values ​​greater than 10 as the tooth centroids.

[0037] The center-focusing mechanism in S5 includes remapping the label category of the tooth block in the tooth segmentation stage and then inputting it into the tooth segmentation M2FF-UNet, setting the tooth category located at the center of the tooth block as the central tooth, and modifying the corresponding label value to 1, setting all teeth except the central tooth in the tooth block as adjacent teeth, and modifying the corresponding label value to 2, and the label value corresponding to the background remains unchanged at 0. The tooth block segmentation task is a three-category segmentation task, which will ensure that the model's attention is focused on the segmentation of the central tooth.

[0038] The construction method of the post-processing algorithm suitable for the tooth segmentation problem in S6 is: traverse each centroid point as the target centroid, find the coordinates of the 6 adjacent centroid points closest to the target centroid through the kd tree algorithm, and then calculate the Euclidean distances of all voxels in the tooth block to these 7 points to obtain the centroid index corresponding to the minimum distance, and put all voxels corresponding to non-zero indexes into set 1, then find all voxels in the tooth block that have the same predicted category as the target centroid and put them into set 2, delete the voxels in set 1 and set 2 in common in set 1, and finally modify all accelerated prediction categories in set 1 to background.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The multi-directional and multi-scale feature fusion module M2FF-Block proposed in the present invention can effectively utilize the inherent multi-directional and multi-scale features of teeth, and also achieve a balance between context information capture and computational efficiency, thereby making the positioning and segmentation of teeth more accurate.

[0041] 2. The center focusing mechanism proposed in the present invention can enable the model to focus on the segmentation of the central teeth, improve the segmentation effect on the boundaries of the central teeth, make the boundaries between adjacent teeth clearer, improve the refined segmentation accuracy of the tooth block, and thus improve the overall segmentation performance.

[0042] 3. The post-processing algorithm proposed in the present invention starts from the spatial relationship between the centroid and the nearest neighbor centroid, and adjusts the predicted category of the foreground voxels in the predicted fuzzy area according to the distance from the foreground voxels to the target centroid, further ensuring the accurate segmentation between adjacent teeth, thereby effectively solving the adhesion problem between adjacent teeth. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0044] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0045] Figure 1 is a flow chart of the method of the present invention;

[0046] Figure 2 This is a schematic diagram of the structure of the M2FF-Block proposed in the present invention;

[0047] Figure 3 This is the structural diagram of the M2FF-UNet proposed in the present invention;

[0048] Figure 4 It is a flow chart of the post-processing method proposed by the present invention;

[0049] Figure 5 This is the segmentation result diagram of the present invention. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than to limit the claims of the present invention. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0051] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0052] like Figures 1 to 5As shown, the present invention proposes a two-stage CBCT tooth instance segmentation method based on a multi-directional multi-scale feature fusion network (M2FF-UNet), comprising the following steps:

[0053] Step 1: Dataset preparation

[0054] 1.1 Data Collection

[0055] A total of 129 CBCT datasets were collected from public databases, and 305 CBCT datasets were collected from the oral radiology department of a certain hospital’s oral center with the patients’ consent.

[0056] 1.2 Data Preprocessing

[0057] The CBCT dataset collected from the hospital was marked as gold standard by dentists, and then the physical resolution of all CBCT images was standardized to 0.4 mm, and the intensity value was limited to 500 to 2500 before normalization.

[0058] Step 2: Design of multi-directional and multi-scale feature fusion module M2FF-Block

[0059] 2.1 Design of multi-directional and multi-scale feature extraction module M2FCM

[0060] Branch 1: No processing;

[0061] Branch 2: convolution layer br2_conv1, the convolution kernel size is 3×3×3, the step size is 1, and the number of groups is the number of input channels;

[0062] Branch three: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer is brb3_conv1, the convolution kernel size is 1×1×7, the step size is 1, the padding is (0, 0, 3), and the number of groups is the number of input channels; the second convolutional layer is br3_conv2, the convolution kernel size is 1×1×9, the step size is 1, the padding is (0, 0, 4), and the number of groups is the number of input channels; the third convolutional layer is br3_conv3, the convolution kernel size is 1×1×11, the step size is 1, the padding is (0, 0, 5), and the number of groups is the number of input channels;

[0063] Branch 4: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer br4_conv1 has a convolution kernel size of 1×7×1, a step size of 1, a padding of (0, 3, 0), and the number of groups is the number of input channels; the second convolutional layer br4_conv2 has a convolution kernel size of 1×9×1, a step size of 1, a padding of (0, 4, 0), and the number of groups is the number of input channels; the third convolutional layer br4_conv3 has a convolution kernel size of 1×11×1, a step size of 1, a padding of (0, 5, 0), and the number of groups is the number of input channels;

[0064] Branch 5: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer br5_conv1 has a convolution kernel size of 7×1×1, a step size of 1, a padding of (3, 0, 0), and the number of groups is the number of input channels; the second convolutional layer br5_conv2 has a convolution kernel size of 9×1×1, a step size of 1, a padding of (4, 0, 0), and the number of groups is the number of input channels; the third convolutional layer br5_conv3 has a convolution kernel size of 11×1×1, a step size of 1, a padding of (5, 0, 0), and the number of groups is the number of input channels;

[0065] Multi-directional and multi-scale feature fusion: The outputs of the third, fourth, and fifth branches are aggregated through a 1×1×1 convolution layer with a convolution kernel of 1 and a stride of 1. The aggregated features are concatenated with the outputs of the first and second branches in the channel dimension.

[0066] 2.2 Spatial Focusing Module (SFM) Design

[0067] The first pooling layer, pool1, uses average pooling with a window size of 7, a step size of 1, and a padding of 3;

[0068] The first convolutional layer is conv1, the convolution kernel size is 1×1×1, and the stride is 1;

[0069] The second convolutional layer conv2 has a convolution kernel size of 1×1×11, a stride of 1, a padding of (0, 0, 5), and the number of groups is the number of input channels;

[0070] The third convolutional layer conv3 has a convolution kernel size of 1×11×1, a stride of 1, a padding of (0, 5, 0), and the number of groups is the number of input channels;

[0071] The fourth convolutional layer conv4 has a convolution kernel size of 11×1×1, a stride of 1, a padding of (5, 0, 0), and the number of groups is the number of input channels;

[0072] The fifth convolution layer conv5 has a convolution kernel size of 1×1×1 and a stride of 1;

[0073] Finally, the sigmoid function outputs an attention map from 0 to 1.

[0074] 2.3 Overall structural design

[0075] The input feature vector is input to two branches at the same time. The first branch systematically extracts multi-scale features in different directions through the multi-directional multi-scale feature extraction module (M2FCM). The second branch adaptively emphasizes the key spatial positions through the spatial focusing module (SFM). Then, the enhanced features obtained by multiplying the input features with the output results of the two branches are added to obtain the enhanced multi-directional multi-scale features.

[0076] Step 3. Structural design of tooth recognition M2FF-UNet and tooth segmentation M2FF-UNet

[0077] The convolution blocks in the four encoder layers of UNet are replaced with the multi-directional and multi-scale feature integration modules M2FF-Block. The first encoder layer contains 3 M2FF-Blocks, the second encoder layer contains 3 M2FF-Blocks, the third encoder layer contains 15 M2FF-Blocks, and the fourth encoder layer contains 3 M2FF-Blocks.

[0078] For tooth recognition M2FF-UNet, two different decoder branches are constructed. The first branch is used to segment the global tooth area, and the other branch is used to regress the offset of each voxel relative to its center of mass. The loss function of tooth recognition M2FF-UNet during training is:

[0079] Lid=αLbseg+βLoff

[0080] Lbseg=0.8×Ldice+0.2×Lce

[0081] Among them, L bseg is the global binary segmentation loss, L dice is the dice loss, L ce is the cross entropy loss,

[0082] L off For the regression offset loss, use L 1 loss, α is 0.65, and β is 0.35.

[0083] For tooth segmentation M2FF-UNet, two different decoder branches are constructed. The first branch is used to segment the tooth area within the tooth block, and the other branch is used to segment the boundary area within the tooth block. The loss function of tooth segmentation M2FF-UNet during training is:

[0084] Lseg=αLtooth+βLbd

[0085] Ltooth=0.6×Ldice+0.4×Lce

[0086] Lbd=0.6×Ldice+0.4×Lce

[0087] Among them, L tooth is the segmentation loss of the tooth region in the tooth block, L dice is the dice loss, L ce is the cross entropy loss, L bdis the boundary area segmentation loss in the tooth block, α is 0.65, and β is 0.35.

[0088] Step 4: Identify teeth from a global perspective

[0089] The specific steps for tooth recognition from a global perspective to obtain the center of mass of each tooth are:

[0090] 1. First, obtain the global binary segmentation map output by the tooth recognition sub-network and the centroid offset map of all voxels;

[0091] 2. Accumulate all predicted foreground tooth voxels according to the offset map to obtain the mapped new coordinates;

[0092] 3. Count the pointing frequencies of each new coordinate to obtain the density map P;

[0093] 4. Retain voxels with density values ​​greater than 20 as potential centroids;

[0094] 5. Calculate the minimum distance from each voxel in the potential centroid to the voxel with a larger density value, construct a distance map D, and filter points with a density value greater than 60 and a distance value greater than 10 as the tooth centroid.

[0095] Step 5: Fine segmentation of individual tooth blocks

[0096] According to the position of the tooth centroid obtained in the tooth recognition stage, the original image is centrally cropped to obtain a tooth block of 80×80×128 centered on each tooth. The present invention takes into account that the number of teeth contained in each tooth block is not fixed, and we are only concerned with the segmentation of the central teeth of the tooth block, while the existing method directly regards it as a two-category segmentation task, and then relies on some morphological algorithms or connected domain analysis to separate the central teeth, which will cause the boundary between the central teeth and the adjacent teeth to be unclear, resulting in segmentation adhesion. Therefore, the label category of the tooth block should be remapped, and the category of the tooth located at the center of the tooth block is set to "central tooth", and the corresponding label value is modified to 1. All tooth categories except the central tooth in the tooth block are set to "adjacent teeth", and the corresponding label value is 2. The label value corresponding to "background" is 0, and the tooth block segmentation task becomes a three-category segmentation task, ensuring that the model focuses on the segmentation of the central teeth.

[0097] Step 6: Construct a post-processing algorithm suitable for tooth segmentation and correct the segmentation results. The specific steps of the post-processing algorithm are:

[0098] 1. Traverse each centroid point as the target centroid, and find the 6 adjacent centroid points closest to the target centroid through the kd tree algorithm;

[0099] 2. Calculate the Euclidean distance from all voxels in the tooth block to these 7 points, and get the centroid index corresponding to the minimum distance. If the index value corresponding to a voxel is 0, it means that the centroid closest to the voxel is the target centroid, and the voxel belongs to the center tooth and needs to be retained. Otherwise, the voxel does not belong to the center tooth, and the predicted label value should be modified to 0;

[0100] 3. Store all voxels corresponding to non-zero indexes into set 1;

[0101] 4. Find all voxels in the tooth block that have the same category as the target centroid prediction and store them in set 2;

[0102] 5. Delete the voxels in set 1 that are common to set 1 and set 2. This step is to maintain the integrity of the output prediction segmentation results of the segmentation network;

[0103] Change the predicted class of all voxels in set 1 to background.

[0104] Step 7: Model Evaluation

[0105] The present invention uses four indicators to evaluate the segmentation performance: Dice similarity coefficient, intersection over union (IoU), average symmetric surface distance (ASSD) and 95% Hausdorff distance (HD). Specifically, the Dice coefficient calculates the proportion of voxel-level segmentation overlap, while IoU quantifies the overlap between the predicted area and the ground truth area. HD measures the maximum distance between the closest points on two surfaces, and ASSD is mainly used to calculate the average distance between the ground truth and the segmented surface. The specific formula is as follows:

[0106]

[0107] Among them, V seg Represents the predicted segmentation result, V gt represents the true label, S seg represents the predicted surface point set, S gt represents the surface point set of the true label, and p is the predicted surface point set S seg In the point, g is a point in the real label surface point set. In addition, d(p,S gt ) is the distance from point p to set S gt The distance between the nearest points, d(g,S seg ) is the point g to the set S seg TP, TN, FP and FN are True Positive, True Negative, False Positive and False Negative respectively.

[0108] Six SOTA methods and the method of the present invention were verified on public CBCT datasets and CBCT datasets collected from hospitals. The effects of different methods on the segmentation results were compared, and the Dice coefficient, IoU, 95% HD and ASSD of different methods were obtained, as shown in Tables 1 and 2.

[0109] Table 1 Comparison of segmentation results of different methods on the CBCT dataset collected from the hospital

[0110]

[0111]

[0112] Table 2 Comparison of segmentation results of different methods on public CBCT datasets

[0113] method Dice IoU 95% HD ASSD 3DUNet 0.918 0.908 1.940 0.703 U-NetR 0.923 0.899 1.957 0.669 Swin U-NetR 0.931 0.911 1.836 0.612 3D UX-Net 0.944 0.919 1.688 0.472 HMG-Net 0.942 0.928 1.531 0.423 HFF-Net 0.953 0.939 1.307 0.349 The present invention 0.968 0.947 1.241 0.301

[0114] As shown in Table 1 and Table 2, the two-stage CBCT tooth instance segmentation method based on multi-directional multi-scale feature fusion M2FF-UNet proposed in the present invention has achieved the best performance in all indicators on the data set collected from the hospital, with a Dice coefficient of 0.977, an IoU of 0.954, a 95% HD of 1.013mm, and an ASSD of 0.233. It is worth noting that the method proposed in the present invention improves Dice by 3.4% and reduces 95% HD by 0.398 mm, respectively. In addition, on the public data set, our method also outperforms all other methods, with a Dice coefficient of 0.965, an IoU of 0.947, a 95% HD of 1.249mm, and an ASSD of 0.301mm. This is sufficient to prove that the method proposed in the present invention fully improves the accuracy of tooth instance segmentation on the CBCT data set, significantly outperforming other methods. And from the segmentation result diagram, it can be seen that the method proposed in the present invention effectively solves the problem of segmentation adhesion in complex areas such as closely arranged teeth, and achieves the best segmentation effect.

[0115] The present invention proposes a two-stage CBCT tooth instance segmentation method based on a multi-directional and multi-scale feature fusion network M2FF-UNet. The M2FF-UNet proposed in the present invention plays a key role in extracting multi-directional and multi-scale features, enabling the model to better adapt to the morphological size and spatial characteristics of tooth changes. The center-focusing mechanism proposed in the present invention achieves refined segmentation by giving priority to the central teeth within each tooth block and enhancing the segmentation boundary segmentation, especially in areas with dense or overlapping teeth. In addition, the present invention proposes a tooth segmentation post-processing algorithm to further optimize the segmentation results by redistributing voxels in areas with blurred predicted boundaries. A large number of experiments have demonstrated the effectiveness and robustness of the present invention, highlighting its potential to enhance diagnostic accuracy and optimize treatment strategies in clinical practice.

[0116] Only the preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention, and various changes should be included in the protection scope of the present invention.

Claims

1. A two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network, characterized in that: The following steps are involved: S1, obtain oral CBCT data set and preprocess it; S2, construct the multi-directional and multi-scale feature fusion module M2FF-Block; S3, based on M2FF-Block, construct two stages of multi-directional and multi-scale feature fusion network M2FF-UNet; S4, performing tooth recognition from a global perspective to obtain the position of the center of mass of each tooth; S5, cropping the original CBCT image according to the centroid position to obtain the tooth block data corresponding to each tooth, and performing fine segmentation of the single tooth block in combination with the center focusing mechanism; S6. Construct a post-processing algorithm suitable for the tooth segmentation problem, and correct the segmentation result to obtain the final segmentation result.

2. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 1 is characterized in that: The method for constructing the multi-directional and multi-scale feature fusion module M2FF-Block in S2 is: First, a multi-directional and multi-scale feature extraction module (M2FCM) is constructed to systematically extract multi-scale features in different directions. Then, a spatial focusing module (SFM) is constructed to adaptively emphasize the key spatial positions. Then, the input and the output of the two modules are multiplied and added to obtain enhanced multi-directional and multi-scale features.

3. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 1, characterized in that: The construction method of M2FCM in S3 is: The input feature vector is divided into 5 branches. The first branch retains the original feature information without any processing. The second branch uses 3×3×3 deep convolution to extract local spatial features. The remaining three branches use large-core deep convolution blocks of different sizes of 1×1×k, 1×k×1, and k×1×1 to replace k×k×k convolution blocks to extract multi-directional and multi-scale features along different directions. Large-core convolution is used to perceive contextual information, and deep convolution is used to improve computational efficiency. Then, 1×1×1 convolution is used to aggregate these multi-directional and multi-scale features. Finally, all outputs are spliced ​​in the channel dimension to achieve feature fusion.

4. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 3 is characterized in that: The structure of the M2FCM includes: Branch 1: No processing; Branch 2: convolution layer br2_conv1, the convolution kernel size is 3×3×3, the step size is 1, and the number of groups is the number of input channels; Branch three: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer is brb3_conv1, the convolution kernel size is 1×1×7, the step size is 1, the padding is (0, 0, 3), and the number of groups is the number of input channels; the second convolutional layer is br3_conv2, the convolution kernel size is 1×1×9, the step size is 1, the padding is (0, 0, 4), and the number of groups is the number of input channels; the third convolutional layer is br3_conv3, the convolution kernel size is 1×1×11, the step size is 1, the padding is (0, 0, 5), and the number of groups is the number of input channels; Branch 4: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer br4_conv1 has a convolution kernel size of 1×7×1, a step size of 1, a padding of (0, 3, 0), and the number of groups is the number of input channels; the second convolutional layer br4_conv2 has a convolution kernel size of 1×9×1, a step size of 1, a padding of (0, 4, 0), and the number of groups is the number of input channels; the third convolutional layer br4_conv3 has a convolution kernel size of 1×11×1, a step size of 1, a padding of (0, 5, 0), and the number of groups is the number of input channels; Branch 5: After the normalization layer, there are three parallel convolutional layers. The first convolutional layer br5_conv1 has a convolution kernel size of 7×1×1, a step size of 1, a padding of (3, 0, 0), and the number of groups is the number of input channels; the second convolutional layer br5_conv2 has a convolution kernel size of 9×1×1, a step size of 1, a padding of (4, 0, 0), and the number of groups is the number of input channels; the third convolutional layer br5_conv3 has a convolution kernel size of 11×1×1, a step size of 1, a padding of (5, 0, 0), and the number of groups is the number of input channels; Multi-directional and multi-scale feature fusion: The outputs of the third, fourth, and fifth branches are aggregated through a 1×1×1 convolution layer with a convolution kernel of 1 and a stride of 1. The aggregated features are concatenated with the outputs of the first and second branches in the channel dimension.

5. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 2, characterized in that: The steps of constructing the spatial focusing module SFM are as follows: First, a 3D average pooling layer is used to compress spatial information, and a 1×1×1 convolution is used to refine the interaction and integration between channels. In order to capture cross-directional spatial information, 1×1×k, 1×k×1, and k×1×1 convolutions are applied along the height, width, and depth directions, respectively, to focus on feature changes specific to each spatial direction. Finally, an attention map is generated through 1×1×1 convolution and sigmoid activation to guide the focus on spatially important areas in the multi-directional and multi-scale features extracted by M2FCM.

6. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 5, characterized in that: The structure of the spatial focusing module SFM includes: The first pooling layer, pool1, uses average pooling with a window size of 7, a step size of 1, and a padding of 3; The first convolutional layer is conv1, the convolution kernel size is 1×1×1, and the stride is 1; The second convolutional layer conv2 has a convolution kernel size of 1×1×11, a stride of 1, a padding of (0, 0, 5), and the number of groups is the number of input channels; The third convolutional layer conv3 has a convolution kernel size of 1×11×1, a stride of 1, a padding of (0, 5, 0), and the number of groups is the number of input channels; The fourth convolutional layer conv4 has a convolution kernel size of 11×1×1, a stride of 1, a padding of (5, 0, 0), and the number of groups is the number of input channels; The fifth convolution layer conv5 has a convolution kernel size of 1×1×1 and a stride of 1; Finally, the sigmoid function outputs an attention map from 0 to 1.

7. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 1, characterized in that: The method of the two-stage multi-directional and multi-scale feature fusion network M2FF-Une in S3 is as follows: first, the convolution blocks of the four encoder layers in the UNet architecture are replaced with different numbers of M2FF-Blocks; for the tooth recognition M2FF-UNet in the tooth recognition stage, two decoder branches are constructed, one branch is used to segment the global tooth area, and the other branch is used to regress the offset of each voxel relative to its center of mass; for the tooth segmentation M2FF-UNet in the tooth segmentation stage, two decoder branches are constructed, one branch is used to segment the tooth area within the tooth block, and the other branch is used to segment the boundary area within the tooth block.

8. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 1, characterized in that: The method for obtaining the position of the centroid of each tooth in S4 is: first obtain the global binary segmentation map and the global voxel centroid offset map output by the tooth recognition subnetwork, then map all predicted foreground tooth voxels to new coordinates according to the offset map, count the pointing frequency of each new coordinate to obtain the density map P, then calculate the minimum distance from each voxel to the voxel with a larger density value, construct the distance map D, and finally select points with density values ​​greater than 60 and distance values ​​greater than 10 as the tooth centroids.

9. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 1, characterized in that: The center-focusing mechanism in S5 includes remapping the label category of the tooth block in the tooth segmentation stage and then inputting it into the tooth segmentation M2FF-UNet, setting the tooth category located at the center of the tooth block as the central tooth, and modifying the corresponding label value to 1, setting all teeth except the central tooth in the tooth block as adjacent teeth, and modifying the corresponding label value to 2, and the label value corresponding to the background remains unchanged at 0. The tooth block segmentation task is a three-category segmentation task, which will ensure that the model's attention is focused on the segmentation of the central tooth.

10. The two-stage CBCT tooth segmentation method based on a multi-directional and multi-scale feature fusion network according to claim 1, characterized in that: The construction method of the post-processing algorithm suitable for the tooth segmentation problem in S6 is: traverse each centroid point as the target centroid, find the coordinates of the 6 adjacent centroid points closest to the target centroid through the kd tree algorithm, and then calculate the Euclidean distances of all voxels in the tooth block to these 7 points to obtain the centroid index corresponding to the minimum distance, and put all voxels corresponding to non-zero indexes into set 1, then find all voxels in the tooth block that have the same predicted category as the target centroid and put them into set 2, delete the voxels in set 1 and set 2 in common in set 1, and finally modify all accelerated prediction categories in set 1 to background.

Citation Information

Patent Citations

  • CBCT tooth instance segmentation method combining global attention and scale perception

    CN116188513A

  • Tooth instance segmentation method and system based on self-attention and receptive field adjustment

    CN116485809A

  • Three-stage tooth intelligent accurate segmentation method based on shape distance field

    CN119477938A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1

Cited By

  • Recognition processing method and device based on multi-angle anterior tooth photo

    CN120953209A