2D panoramic image tooth segmentation network based on boundary perception
By introducing boundary prediction network and feature fusion module into the tooth segmentation network, the problems of boundary errors and insufficient feature learning ability in tooth image segmentation are solved, and higher segmentation accuracy and boundary prediction accuracy are achieved.
Patent Information
- Application Number
- CN202410100099.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing tooth image segmentation method has weak discrimination ability in the face of the pixel feature representation between the teeth and the background tissue, resulting in boundary errors, and it is difficult for a single-task deep learning network to learn enough feature complementarity, affecting the segmentation accuracy.
A 2D panoramic image tooth segmentation network based on boundary perception is adopted. By introducing a boundary prediction network as an auxiliary task, combining the gated fusion module and the channel space attention module, the tooth main features and edge features are fused, and deep supervision is carried out in the decoder, and the feature fusion module is used to improve segmentation accuracy.
The accuracy of tooth segmentation and the accuracy of boundary prediction are significantly improved, and the feature learning ability and segmentation performance of the model are enhanced by learning additional boundary features.
Smart Images

Figure CN120339307A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of dental image segmentation models, and particularly relates to a 2D panoramic image dental segmentation network based on boundary perception. Background Art
[0002] With the development of society, people's lifestyles and eating habits have changed greatly. Due to unreasonable oral cleaning methods, the oral health situation has become increasingly severe, and the incidence of oral diseases has been rising continuously. Oral diseases account for a relatively large proportion among global diseases and there are many types of diseases. Among them, dental diseases are an important factor in oral diseases. The diagnosis and treatment of dental diseases require professional dentists to observe and analyze the health status of dental organs to make further treatment plans. Therefore, medical imaging is particularly important for improving the preoperative diagnosis efficiency. In the field of oral image analysis, dental image segmentation is a key technology for disease detection and recognition, and it is also one of the most difficult tasks in the image processing process. Dental image segmentation is to accurately segment and locate the dental area or region of interest in the panoramic oral X-ray image, so that doctors can better evaluate the condition of the teeth and formulate corresponding treatment plans. Figure 8 A 2D panoramic dental image and the corresponding ground truth segmentation mask are shown. Since the elderly generally have oral problems such as gum atrophy and tooth loss, and due to the similarity of human tissue density, the existence of permanent artifacts (such as dental fillings and dental implants) and temporary artifacts (such as orthodontic brackets, impacted teeth, tooth crowding, tooth spacing), there is a lack of clear boundaries between regions of interest, making dental segmentation very challenging.
[0003] Some early dental image segmentation methods mainly used traditional image processing algorithms based on thresholds, regions, and edges. However, medical images are different from natural images in terms of good clarity. They usually have problems such as noise, low contrast, and uneven exposure. Relying on manual feature extraction often fails to achieve accurate segmentation of teeth.
[0004] In recent years, with the development of deep learning, deep convolutional neural networks have been widely used in the field of image segmentation. Compared with image segmentation algorithms based on handcrafted features, data-driven deep learning networks can extract richer feature structures, and the models have good generalization ability and more accurate prediction results, which are extremely effective for diverse segmentation problems. In the field of panoramic tooth image segmentation based on deep learning, although many segmentation methods have been proposed to address the challenges faced in tooth segmentation, these methods often focus on using a single task to improve segmentation performance. This is limited in the following aspects: (1) The feature learning ability of the model: In deep learning, the performance of the model depends on its feature learning ability for key problems. The single-task segmentation method cannot learn the ability to perform feature complementarity; (2) Boundary errors: In tooth images, the discriminative ability of pixel features between the tooth region and the background tissue is weak. Therefore, a large number of missegmentations are concentrated on the tooth boundaries. To address the problem of the model's feature learning ability, Ronneberger et al. introduced the U-Net network into the tooth segmentation task and successfully achieved automatic segmentation of dental X-ray images. Subsequent methods have made various improvements to U-Net. For example, Tao et al. proposed a U-Net method with attention. However, these methods based on improved U-Net greatly increase the complexity of the model and cannot learn additional discriminative features.
[0005] Early tooth image segmentation mainly relied on segmentation methods based on handcrafted features, including the gray threshold, region, and edge of teeth. Due to the limited expressive ability of handcrafted features, traditional tooth segmentation models often could not achieve satisfactory segmentation performance. Summary of the Invention
[0006] To solve the technical problems existing in the prior art, the present invention provides a simple and effective network structure, the BS-Net multi-task segmentation model, for 2D panoramic tooth image segmentation. By using tooth boundary prediction as an auxiliary task to add additional discriminative features of tooth information, this model can accurately predict the tooth boundary, thereby improving the segmentation accuracy. To further improve the performance of the tooth segmentation model, the present invention introduces a deep supervision strategy and applies supervision signals to the output features of different levels of the decoder during the training process. The present invention proposes a spatial-channel attention module to highlight the important features of the encoder-decoder bottleneck layer. The present invention obtains a fine tooth segmentation result by designing a feature fusion module.
[0007] To achieve the above object, the technical solution adopted by the present invention is: A 2D panoramic image tooth segmentation network based on boundary awareness, including a backbone network for tooth segmentation and a boundary prediction network for boundary prediction.
[0008] The first-layer features and second-layer features in the backbone network encoder are fused through a gated fusion module, and the output of the gated fusion module is used as the input of the boundary prediction network. The boundary prediction network extracts tooth edge features, and feature addition is used to fuse features in similar spaces and scales. The output features of the 4th decoding blocks of the backbone network and the boundary prediction network are added together, and two decoding features, namely the main body features of the teeth and the edge features of the teeth, are fused through a feature fusion module. The output of the feature fusion module is used as the final network to predict the tooth segmentation map.
[0009] This segmentation network also includes a single-input dual-output structure composed of a pair of encoder-decoders. Max-pooling is used between blocks of the encoder to reduce the feature size to extract features at different scales. The bilinear interpolation algorithm is used in the decoder to gradually restore the image resolution layer by layer. The output features of each layer in the encoder are cascaded into the corresponding decoder blocks to supplement the shallow features lost due to multiple convolutions. The output features of each layer in the decoder are used with output convolutions to predict the segmentation results for deep supervision.
[0010] In the edge prediction network, the Canny edge detection algorithm (a commonly used image processing algorithm for detecting edges in images) is applied to the label image of the tooth panoramic image dataset to obtain the edge map of the teeth, and the obtained tooth edge map is used as the label of the boundary prediction network for loss calculation.
[0011] In the backbone network, the same attention module as in the boundary prediction network is used to highlight the features of the bottleneck layer in the backbone network. The output features of the last decoding block in the decoder of the backbone network (as shown in Figure 1 ) and the output features of the decoding block from the boundary prediction network are fed into a feature fusion module (the specific structure is as shown in Figure 4 ) to supplement the fine tooth edge features.
[0012] Each image in the tooth panoramic image dataset has a corresponding tooth segmentation binary mask map. The tooth panoramic image is a pixel-level binary classification task, that is, two categories: the tooth region and the background. The combination of the binary cross-entropy loss function and the soft dice loss function is used as the loss function for optimizing the network. The binary cross-entropy loss function is defined as in Equation (1), and the soft dice loss is defined as in Equation (4):
[0013]
[0014]
[0015]
[0016] L softdice = 1 - (Dice + Dice b ) / 2 (4)
[0017] L BCE is the binary cross-entropy loss function, where N is the total number of pixels, i represents the i-th pixel, and g i ∈ {0, 1} represents two category labels, and p i ∈ {0, 1} represents the predicted category label. The parameter ε is an extremely small number to ensure numerical stability and prevent the denominator from being zero. L softdice is the soft Dice loss function.
[0018] Let L represent the combined loss function of tooth segmentation and boundary prediction, which is defined as follows:
[0019] L = L BCE + L softdice (5)
[0020] Upsample the four outputs of the backbone network and the boundary prediction network decoder to have the same size as the label, and use the same Ground truth (the true segmentation mask corresponding to the tooth image) for supervision. Use i, m ∈ (1, 2, 3, 4) respectively represent the prediction results of the backbone network and the boundary prediction network decoder. G represents the tooth segmentation Ground truth, and Y represents the true tooth edge segmentation mask corresponding to the tooth edge image. Therefore, the total loss function is expressed as:
[0021]
[0022] In Equation (6) defining the total loss function, L is a combination of the soft Dice loss and the binary cross-entropy loss function. 2i and 2m are the loss function weight coefficients of different decoding layers, and the weight coefficient of each layer increases in multiples of 2. The weight coefficients of each decoding layer are (2, 4, 6, 8). α and β are the loss function weight coefficients of the two networks.
[0023] Compared with the prior art, the specific beneficial effects of the present invention are as follows: By introducing a boundary prediction network as an auxiliary task, the present invention proposes a simple and effective tooth segmentation model. The model can learn additional fine boundary features through the boundary prediction network, which helps to improve the segmentation accuracy. The gated fusion module fuses the shallow features containing rich boundary information as the input of the boundary prediction network; the channel-spatial attention module highlights the significant features of the bottleneck layer in the form of feature compression; deep supervision is used in both networks. In addition, introducing an additional boundary learning task in the model can significantly improve the segmentation performance. Therefore, the method proposed by the present invention can be easily applied to other models. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is the overall block diagram of the present invention.
[0025] Figure 2 This is the structural diagram of the gating fusion module in the present invention.
[0026] Figure 3 This is the structural diagram of the channel spatial attention module in the present invention.
[0027] Figure 4 This is the structural diagram of the feature fusion module in the present invention
[0028] Figure 5 These are the structural diagrams of several convolutional blocks in the overall block diagram
[0029] Figure 6 This is the Dice box plot of the present method and other methods on the test set
[0030] Figure 7 This is the comparison chart of the qualitative test results of various methods.
[0031] Figure 8 This is a panoramic tooth image. Detailed implementation manners
[0032] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] As Figure 1-2 shown, a 2D panoramic image tooth segmentation network based on boundary perception Figure 1Shows the architecture of the proposed boundary-aware network (BS-Net) for panoramic dental image segmentation based on boundary awareness, which consists of a backbone network for dental segmentation and a boundary prediction network for boundary prediction. Specifically, first, the panoramic dental image is fed into the backbone network encoder (ResNet-34 as the backbone) to extract multi-level features. Since the low-level features contain rich boundary information, the present invention proposes a gated fusion module (GFM module) to fuse the first-layer features and the second-layer features in the backbone network encoder. The output of the feature fusion module is used as the input of the boundary prediction network. The boundary prediction network can extract fine dental edge features. Feature addition is usually used to fuse features in similar spaces and scales, enabling the network to learn more rich and complex feature representations. Therefore, the present invention adds the output features of the 4th decoding blocks of the backbone network and the boundary prediction network and fuses the two decoded features, including the main body features of the teeth and the edge features of the teeth, through a feature fusion module (FFM). The output of the feature fusion module is used as the final network prediction of the dental segmentation map. The proposed channel-spatial attention CSAM module is placed after the encoders of the backbone network and the boundary prediction network to highlight the significant features in the feature blocks. A deep supervision strategy is adopted in the decoder to predict the output features of each decoding block using output convolution to improve the segmentation performance of the network.
[0034] The boundary-aware network BS-Net proposed by the present invention is a single-input dual-output structure composed of a pair of encoder-decoders. Max-pooling is used between blocks of the encoder to reduce the feature size and extract features at different scales. Bilinear interpolation algorithm is used in the decoder to gradually restore the image resolution layer by layer. The present invention cascades the output features of each layer in the encoder into the corresponding decoder block to supplement the shallow features lost due to multiple convolutions. The present invention uses output convolution to predict the segmentation result for the output features of each layer in the decoder for deep supervision. As Figure 5 shown, the output convolution block consists of a 1x1 convolution layer and bilinear interpolation. The number of feature channels is reduced by 1x1 convolution, then the image resolution is restored to the same size as the input image through bilinear interpolation, and finally, Sigmoid is used to predict the segmentation probability map.
[0035] The tooth edge information presented in the panoramic tooth image dataset is blurred, and it is difficult for a single backbone segmentation network to segment the precise tooth edge. Therefore, training a dedicated boundary prediction network for tooth edge information can improve tooth segmentation performance. To this end, the present invention proposes a boundary prediction network based on an encoder-decoder structure to generate a boundary map, and combines the fine tooth boundary features learned by the boundary prediction network into the backbone network to improve the segmentation performance. Existing research has shown that shallow features in convolutional neural networks retain rich boundary information. Therefore, the present invention uses the output features of the first two convolutional blocks in the backbone network to be fused through a gated fusion module GFM. The output features of the gated fusion module are further sent to three convolutional blocks for feature extraction. As Figure 5 shown, the encoding block has a similar internal structure to the decoding block in the boundary-aware network BS-Net. The difference is that the shallow features from the encoding block and the input features of the previous level need to be cascaded before the decoding block. Since the attention mechanism is widely used in convolutional neural networks and has been proven to improve the prediction performance of the network, the present invention proposes a channel-spatial attention CSAM module to highlight the important features in the feature block as the bottleneck layer between the encoder and the decoder. The specific structure of the CSAM module is as Figure 3 shown, where GMP and GAP respectively represent global max pooling and global average pooling in the spatial dimension of the feature block, and Max pool and Avg pool respectively represent max pooling and average pooling in the channel dimension of the feature block. The present invention performs Canny edge detection on the label image of the panoramic tooth image dataset to obtain the edge map of the teeth, and uses the obtained tooth edge map as the label of the boundary prediction network for loss calculation.
[0036] The input panoramic tooth image first undergoes feature extraction through an input convolutional block with 64 7x7 convolutional kernels, so as to obtain a larger receptive field, which is beneficial for the network to capture the global features of the panoramic tooth image. Max pooling operation is used for downsampling after the input convolutional block. These features are encoded through 4 levels in ResNet34, and each level is composed of multiple 3x3 convolutional kernels with different numbers. The height and width of the feature map are halved and the number of channels is doubled between adjacent levels. The present invention uses the same attention module CSAM module as in the boundary prediction network to highlight the important features of the bottleneck layer in the backbone network. The output features of the last decoding block in the decoder and the output features of the decoding block from the boundary prediction network are sent into a feature fusion module FFM to supplement the fine tooth edge features. As Figure 4 shown in the FFM in
[0037] Each image in the dental panoramic image dataset has a corresponding dental segmentation binary mask map. The dental panoramic image is a pixel-level binary classification task, namely two categories: the dental area and the background. The present invention uses a combination of the binary cross-entropy loss function (BCE) and the soft dice loss function as the loss function for optimizing the network. The binary cross-entropy loss function is defined as in Equation (1), and the soft dice loss is defined as in Equation (4).
[0038]
[0039]
[0040]
[0041] L softdice =1 - (Dice + Dice b ) / 2 (4)
[0042] L BCE is the binary cross-entropy loss function, where N is the number of pixels, i represents the i-th pixel, g i ∈{0, 1} represents the two category labels, p i ∈{0, 1} represents the predicted category label, and the parameter ε∈R is a very small value to ensure numerical stability and prevent the denominator from being zero.
[0043] Let L represent the combination of the loss functions for dental segmentation and boundary prediction, which is defined as follows:
[0044] L = L BCE + L softdice (5)
[0045] As Figure 1 shown, the four outputs of the backbone network and the decoder of the boundary prediction network are upsampled to have the same size as the label, and the same Ground truth (the true segmentation mask map corresponding to the dental image) is used for supervision. In the present invention, i, m ∈ (1, 2, 3, 4) respectively represent the prediction results of the backbone network and the decoder of the boundary prediction network, G represents the dental segmentation Ground truth, Y represents the dental edge Ground truth, so the total loss function can be expressed as:
[0046]
[0047] In Equation (6) defining the total loss function, 2i and 2m are the loss function weight coefficients of different decoding layers, and the weight coefficient of each layer increases in multiples of 2. The weight coefficients of each decoding layer are (2, 4, 6, 8). α and β are the loss function weight coefficients of the two networks.
[0048] To verify the effectiveness of the segmentation model, the training dataset of the present invention was randomly divided into 70% for training, 10% for validation, and 20% for testing. Finally, the training set contained 1400 images, the validation set contained 200 images, and the testing set contained 400 images. In the experiment of the present invention, the original data size of 320x640 pixels was used as the input of the network.
[0049] The network proposed by the present invention was implemented using an NVIDIA GPU (GeForce RTX 3090) in the PyTorch library. The Adam optimizer was used to train the model, with a batch size of 6 and the initial learning rate set to 1x10 -3 。
[0050] The weight decay was set to 0.001. Data augmentation was used during the training phase, including random horizontal and vertical flipping, rotation, and random cropping. The present invention trained a total of 150 epochs and used the best model on the validation set as the final test model to prevent overfitting.
[0051] The present invention uses five performance metrics commonly used in medical image segmentation to evaluate the performance of the model, including the Dice coefficient (DC), accuracy (ACC), Hausdorff distance (HD95), sensitivity (SE), and specificity (SP), as shown specifically below:
[0052]
[0053]
[0054]
[0055]
[0056]
[0057] Where TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative respectively, h represents the Hausdorff distance, A and B are the sets of points on the boundaries of the Ground truth and the predicted segmentation mask, a and b are the points in sets A and B respectively, and d represents the Euclidean distance.
[0058] To study the effectiveness of different key components in the model, including the Gated Fusion Module (GFM), Channel-Spatial Attention Module (CSAM), Boundary Prediction Network, and Feature Fusion Module (FFM), the present invention conducts ablation studies by removing or replacing them from the complete model. First, the Edge Prediction Network and the Channel-Spatial Attention Module (CSAM) proposed in the present invention are removed, and only the pre-trained ResNet34 is used as the encoder and a decoder, denoted as Model 1, which serves as the baseline model. In the baseline model, the Channel-Spatial Attention Module (CSAM) proposed in the present invention is introduced, denoted as Model 2. In Model 3, the present invention introduces the Boundary Prediction Network. First, the Gated Fusion Module (GFM) in the Boundary Prediction Network is removed, and a concatenation operation is used instead of the Gated Fusion Module, and the output features of the fourth decoding block of the backbone network are concatenated with the output features of the fourth decoding block of the Boundary Prediction Network. The concatenated features are output through the output convolution block to verify the effectiveness of the Boundary Prediction Network. On the basis of Model 3, the Gated Fusion Module (GFM) is introduced, denoted as Model 4. To verify the impact of the Feature Fusion Module (FFM) proposed in the present invention on the performance of the network, the present invention replaces the operation of concatenating the output features of the decoding blocks in Model 4 with the Feature Fusion Module (FFM) for experiments, denoted as BS-Net. The ablation experiment results of each module and component are listed in Table 1.
[0059] Table 1 Ablation study of different components
[0060]
[0061] Ablation experiment of the loss function: The total loss function is a combination of the backbone segmentation network and the boundary prediction network. The present invention experiments with the combined weight coefficients α and β of the loss functions of the two networks. The experimental results are listed in Table 2. Through experiments, the present invention finds that the combined weight coefficients α and β of the loss functions of the two networks achieve the best performance indicators when taking 1. Therefore, in the total loss function, α = 1 and β = 1.
[0062] Table 2 Ablation experiment of the weight coefficients α and β of the loss functions of the boundary prediction network and the backbone segmentation network
[0063]
[0064] Quantitative results: The BSNet of the present invention was compared with several state-of-the-art methods, including UNet, AttUNet, CENet, CPFNet, MALUNet, LDNet, CANet, MSCANet. Experiments were conducted on a 2D panoramic tooth segmentation image dataset. The present invention used the same batch_size and number of training epochs as described in the experimental settings and followed the same evaluation process. As shown in Table 3, the BS-Net of the present invention achieved better performance than other methods in most metrics. The proposed BS-Net obtained the best Dice, IOU, ACC, and SE scores, outperforming other methods. Figure 6 The box plot of Dice on the test set for the method of the present invention and several advanced methods is shown, where the dots represent outliers. It can be seen from the figure that the method of the present invention achieved a higher median Dice score compared to UNet, AttUNet, LDNet, and MALUNet, and the outliers were concentrated above 0.8.
[0065] Table 3 Quantitative comparison with advanced methods
[0066]
[0067] Some examples of the segmentation results of different models on the 2D panoramic tooth segmentation image dataset of the present invention are visualized, such as Figure 7 shown. The first row is the input image, and the following rows are the corresponding ground truth segmentation masks and the segmentation results of different models respectively. The present invention selected test samples including complete tooth images, tooth defect images, and tooth boundary-blurred images to analyze the segmentation results of different models on challenging samples. It can be seen from the figure that the BS-Net of the present invention can effectively segment difficult samples with blurred boundaries because the present invention learns boundary features by adding an additional segmentation task. Therefore, the method of the present invention performs well on these boundary-difficult samples.
[0068] The present invention proposed a simple and effective tooth segmentation model by introducing a boundary prediction network as an auxiliary task. The model can learn additional fine boundary features through the boundary prediction network, which helps to improve the segmentation accuracy. The gated fusion module (GFM) fuses shallow features containing rich boundary information as the input of the boundary prediction network; the channel-spatial attention module (CSAM) highlights the significant features of the bottleneck layer in the form of feature compression; deep supervision is used in both networks. In addition, introducing an additional boundary learning task in the model can significantly improve the segmentation performance. Therefore, the method proposed by the present invention can be easily applied to other models. The results of a large number of experiments and ablation studies verified the effectiveness of the proposed BS-Net of the present invention.
[0069] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of the present invention.
Claims
1. A 2D panoramic image tooth segmentation network based on boundary awareness, characterized in that, It includes a backbone network for segmenting teeth and a boundary prediction network for predicting tooth boundaries; The gated fusion module fuses the features of the first layer and the second layer in the backbone network encoder. The output of the gated fusion module is used as the input of the boundary prediction network. The boundary prediction network extracts tooth edge features. Feature addition is used to fuse features in similar spaces and scales. The output features of the 4th decoding blocks of the backbone network and the boundary prediction network are added together, and the tooth body features and tooth edge features are fused through a feature fusion module. The output of the feature fusion module is used as the predicted tooth segmentation map of the final network.
2. The 2D panoramic image tooth segmentation network based on boundary perception according to claim 1, wherein It includes a single-input double-output structure composed of a pair of encoders and decoders. Max pooling is used between blocks of the encoder to reduce the feature size to extract features at different scales. The bilinear interpolation algorithm is used in the decoder to gradually restore the image resolution layer by layer. The output features of each layer in the encoder are concatenated into the corresponding decoder blocks to supplement the shallow features lost due to multiple convolutions. The output features of each layer of the decoder are used with output convolutions to predict the segmentation results for deep supervision.
3. The 2D panoramic image tooth segmentation network based on boundary perception according to claim 2, characterized in that, In the edge prediction network, an edge detection algorithm is applied to the label image of the panoramic tooth image dataset to obtain the edge map of the teeth. The obtained tooth edge map is used as the label of the boundary prediction network for loss calculation; In the backbone network, the same attention module as in the boundary prediction network is used to highlight the features of the bottleneck layer in the backbone network. In the attention module of the backbone network, the output features of the last decoding block in the decoder and the output features of the decoding block from the boundary prediction network are fed into a feature fusion module to supplement the fine tooth edge features.
4. The 2D panoramic image tooth segmentation network based on boundary awareness according to claim 3, wherein, Each image in the panoramic tooth image dataset has a corresponding tooth segmentation binary mask map. The panoramic tooth image is a pixel-level binary classification task. A combination of the binary cross-entropy loss function and the soft dice loss function is used as the loss function for optimizing the network. The binary cross-entropy loss function is defined as shown in Equation (1), and the soft dice loss is defined as shown in Equation (4): L softdice = 1 - (Dice + Dice b ) / 2 (4) Among them, L BCE is the binary cross-entropy loss function, N is the total number of pixels, i is the i-th pixel, and g i ∈{0,1} are two category labels, p i ∈{0,1} is the predicted category label, the parameter ε is an extremely small number, and L softdice is the soft dice loss function; Let L represent the combination of the loss functions for tooth segmentation and boundary prediction, which is defined as follows: L = L BCE + L softdice (5) Upsample the four outputs of the backbone network and the decoder of the boundary prediction network so that the sampling results have the same size as the label, and use the ground truth segmentation mask map corresponding to the same tooth image for supervision. Use respectively represent the prediction results of the backbone network and the decoder of the boundary prediction network, G is the ground truth segmentation mask map corresponding to the tooth segmentation image, Y is the ground truth segmentation mask map corresponding to the tooth edge image, and the total loss function is expressed as: L total is the loss function. In Equation (6) defining the total loss function, 2i and 2m are the loss function weight coefficients of different decoding layers. The weight coefficient of each layer increases in multiples of 2, and the weight coefficients of each decoding layer are (2, 4, 6, 8). α and β are the loss function weight coefficients of the two networks.