Multi-level and scale feature fused tooth image segmentation method and system

By adopting a multi-level and scale feature fusion method in tooth image segmentation, combined with Focal Loss and AdamW technology, the problems of class imbalance and high feature similarity in tooth image segmentation are solved, and the generalization ability and segmentation accuracy of the model are significantly improved.

CN120182302APending Publication Date: 2025-06-20FUJIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201561.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

There are problems of class imbalance in tooth image segmentation and high similarity between diseased teeth and normal teeth, resulting in insufficient model generalization ability and segmentation accuracy.

Method used

Using a multi-level and scale feature fusion method, a tooth image segmentation model is created through the Unet++ module and the output module, the sample weight is adjusted using the Focal Loss function, and combined with the AdamW optimizer to prevent overfitting, improving the generalization ability and segmentation accuracy of the model.

Benefits of technology

It significantly improves the generalization ability and accuracy of tooth image segmentation, solves the problems of class imbalance and high feature similarity, and improves the model's ability to segment the diseased teeth area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182302A_ABST
    Figure CN120182302A_ABST
Patent Text Reader

Abstract

The invention provides a multi-level and scale feature fusion tooth image segmentation method and system in the technical field of tooth image segmentation, and the method comprises the steps: S1, obtaining a large number of historical tooth images, carrying out the tooth filling, diseased tooth, hyperplastic tooth, background marking and pixel value normalization preprocessing of each historical tooth image, and obtaining a tooth filling image; constructing a data set based on the preprocessed historical tooth images; s2, creating a tooth image segmentation model based on a Unit + + module and an output module, setting a loss function of the tooth image segmentation model as a Foca Loss function, and setting an optimizer of the tooth image segmentation model as AdamW; s3, training a tooth image segmentation model through the data set, and deploying the trained tooth image segmentation model; and S4, performing tooth image segmentation through the deployed tooth image segmentation model. The method has the advantage that the generalization ability and accuracy of tooth image segmentation are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dental image segmentation, and particularly to a dental image segmentation method and system for fusing multi-level and multi-scale features. Background Art

[0002] Dental images (X-ray images) are used to obtain the two-dimensional anatomical structures of teeth using X-ray imaging technology. These dental images are important means for detecting, observing, and diagnosing different types of diseased teeth in clinical diagnosis, and can provide information about dental diseases and abnormalities for clinical diagnosis, thereby supporting the diagnosis, treatment, and monitoring of the progression of bad teeth.

[0003] During the process of diagnosing through dental images, it is necessary to locate and label the damaged tissues in the dental images. Precise location and labeling usually require experienced doctors to invest a large amount of time and effort, which is both time-consuming and laborious. With the rapid development of hardware computing power, deep learning technology has gradually been applied to the segmentation task of dental images. With its characteristics such as automation, standardization, and strong adaptability, it has greatly saved manpower and time.

[0004] However, due to the distribution characteristics of human dental images and diseased teeth themselves, there are still the following problems in the process of combining them with deep learning technology: 1. Class imbalance, the number of pixels of diseased teeth is much less than that of non-diseased teeth, resulting in the model over-learning the features of non-diseased teeth, that is, paying too much attention to background pixels and lacking learning of the foreground, thus affecting the generalization ability of the model; 2. The feature similarity between diseased teeth and normal teeth in the original dental images is high, making the boundaries blurred and the gray values similar. Precise segmentation of diseased teeth requires extremely high boundary positioning, and this similar characteristic will affect the segmentation accuracy of the diseased tooth area.

[0005] Therefore, how to provide a dental image segmentation method and system for fusing multi-level and multi-scale features to improve the generalization ability and accuracy of dental image segmentation has become an urgent technical problem to be solved. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a dental image segmentation method and system for fusing multi-level and multi-scale features to improve the generalization ability and accuracy of dental image segmentation.

[0007] In a first aspect, the present invention provides a dental image segmentation method for fusing multi-level and multi-scale features, including the following steps:

[0008] Step S1: Obtain a large number of historical dental images, perform preprocessing on each of the historical dental images, including labeling of filled teeth, diseased teeth, hyperplastic teeth, and background, and pixel value normalization, and construct a data set based on each of the preprocessed historical dental images;

[0009] Step S2: Create a tooth image segmentation model based on the Unet++ module and the output module. Set the loss function of the tooth image segmentation model as the Focal Loss function, and set the optimizer of the tooth image segmentation model as AdamW;

[0010] Step S3: Train the tooth image segmentation model with the dataset, and deploy the trained tooth image segmentation model;

[0011] Step S4: Segment the tooth image with the deployed tooth image segmentation model.

[0012] Further, in step S2, the Unet++ module is a U-shaped network structure, constructed by five layers of encoders and four layers of decoders; a skip connection is used to bridge between each encoder and decoder;

[0013] The encoder is used to gradually reduce the spatial dimension of the historical tooth image through max-pooling operations, while increasing the number of feature channels, in order to capture the hierarchical features of the historical tooth image;

[0014] The decoder is used to gradually restore the spatial dimension of the historical tooth image through upsampling, and combine each historical tooth image with the restored spatial dimension with the corresponding and deeper-level historical tooth images in the encoder, in order to capture the scale features of the historical tooth image;

[0015] The output module is constructed based on a fully connected layer, and is used to fuse each hierarchical feature and scale feature, and output the tooth image segmentation result.

[0016] Further, in step S2, the Focal Loss function is used to adjust the weights of easy-to-classify samples and hard-to-classify samples, so that the tooth image segmentation model focuses more on the learning of hard-to-classify samples.

[0017] Further, in step S2, AdamW is used to prevent overfitting during the training of the tooth image segmentation model.

[0018] Further, step S3 is specifically as follows:

[0019] Divide the dataset into a training set and a test set according to a splitting ratio of 4:1. Train the tooth image segmentation model with the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the optimizer, including at least hyperparameters such as the update step size, the decay rate of the first moment estimate, the decay rate of the second moment estimate, the L2 regularization strength coefficient, the division-by-zero error suppression coefficient, and the learning rate;

[0020] Test the trained dental image segmentation model using the test set, and deploy the dental image segmentation model that passes the test.

[0021] In a second aspect, the present invention provides a dental image segmentation system for multi-level and scale feature fusion, including the following modules:

[0022] A dataset construction module for obtaining a large number of historical dental images, performing preprocessing on each of the historical dental images, including filling teeth, diseased teeth, hyperplastic teeth, background annotation, and pixel value normalization, and constructing a dataset based on the preprocessed historical dental images;

[0023] A dental image segmentation model creation module for creating a dental image segmentation model based on the Unet++ module and the output module, setting the loss function of the dental image segmentation model as the Focal Loss function, and setting the optimizer of the dental image segmentation model as AdamW;

[0024] A dental image segmentation model training module for training the dental image segmentation model using the dataset and deploying the trained dental image segmentation model;

[0025] A dental image segmentation module for segmenting dental images using the deployed dental image segmentation model.

[0026] Further, in the dental image segmentation model creation module, the Unet++ module has a U-shaped network structure and is constructed by five layers of encoders and four layers of decoders; each encoder and decoder are bridged through skip connections;

[0027] The encoder is used to gradually reduce the spatial dimension of the historical dental image through max-pooling operations while increasing the number of feature channels to capture the hierarchical features of the historical dental image;

[0028] The decoder is used to gradually restore the spatial dimension of the historical dental image through upsampling, and combine each of the historical dental images with the restored spatial dimension with the corresponding level and deeper level historical dental images in the encoder to capture the scale features of the historical dental image;

[0029] The output module is constructed based on a fully connected layer and is used to fuse each of the hierarchical features and scale features and output the dental image segmentation result.

[0030] Further, in the dental image segmentation model creation module, the Focal Loss function is used to adjust the weights of easy-to-classify samples and difficult-to-classify samples, so that the dental image segmentation model focuses more on the learning of difficult-to-classify samples.

[0031] Further, in the tooth image segmentation model creation module, AdamW is used to prevent overfitting during the training of the tooth image segmentation model.

[0032] Further, the tooth image segmentation model training module is specifically used for:

[0033] Divide the dataset into a training set and a test set according to a segmentation ratio of 4:1, train the tooth image segmentation model through the training set until the loss value of the loss function is less than a preset loss threshold, and continuously optimize the optimizer during the training process, including at least hyperparameters such as the update step size, the decay rate of the first-order moment estimate, the decay rate of the second-order moment estimate, the L2 regularization strength coefficient, the division-by-zero error suppression coefficient, and the learning rate;

[0034] Test the trained tooth image segmentation model through the test set, and deploy the tooth image segmentation model that passes the test.

[0035] The advantages of the present invention are as follows:

[0036] By obtaining a large number of historical tooth images, preprocessing each historical tooth image by annotating fillings, diseased teeth, hyperplastic teeth, and backgrounds and normalizing pixel values to construct a dataset; then creating a tooth image segmentation model based on the Unet++ module and the output module, setting the loss function of the tooth image segmentation model as the Focal Loss function and the optimizer as AdamW; then training the tooth image segmentation model through the dataset, and deploying the trained tooth image segmentation model; finally, segmenting the tooth image through the deployed tooth image segmentation model; the Unet++ module is a U-shaped network structure, constructed by five layers of encoders and four layers of decoders; each encoder and decoder are bridged through skip connections; the encoder is used to gradually reduce the spatial dimension of the historical tooth image through max-pooling operations while increasing the number of feature channels to capture the hierarchical features of the historical tooth image; the decoder is used to gradually restore the spatial dimension of the historical tooth image through upsampling, and combine each historical tooth image with the corresponding and deeper-level historical tooth images in the encoder to capture the scale features of the historical tooth image; that is, the Unet++ module fully extracts the local features of the tooth image and captures long-distance dependence information through a multi-scale feature fusion mechanism, filters out the key information in the channels through skip connections, makes the decoding process of each layer more refined through a layer-by-layer super-resolution strategy (upsampling gradually restores the spatial dimension), adjusts the weights of easy-to-classify samples and hard-to-classify samples through the Focal Loss function to solve the class imbalance problem, enhances the attention to diseased tooth pixels (hard-to-classify samples), and combines the AdamW optimizer to prevent overfitting, ultimately greatly improving the generalization ability and accuracy of tooth image segmentation. Description of the Drawings

[0037] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0038] Figure 1 It is a flowchart of a method for segmenting tooth images by fusing multi-level and multi-scale features of the present invention.

[0039] Figure 2 It is a schematic structural diagram of a tooth image segmentation system by fusing multi-level and multi-scale features of the present invention.

[0040] Figure 3 It is a schematic diagram of a tooth image segmentation model of the present invention.

[0041] Figure 4 It is a schematic diagram of the VGGBlock module of the present invention. Specific embodiments

[0042] The technical solutions in the embodiments of the present application generally have the following idea: segment the tooth images through a tooth image segmentation model created by the Unet++ module and the output module, and set the loss function of the tooth image segmentation model as the FocalLoss function and the optimizer as AdamW; the Unet++ module fully extracts the local features of the tooth images through the multi-scale feature fusion mechanism, captures the long-distance dependence information, filters out the key information in the channels through the skip connection, makes the decoding process of each layer more refined through the Layer-wise Super-Resolution Strategy, solves the class imbalance problem through the Focal Loss function, enhances the attention to the pixels of diseased teeth, and combines the AdamW optimizer to prevent overfitting, so as to improve the generalization ability and accuracy of tooth image segmentation.

[0043] Please refer to Figures 1 to 4 As shown, a preferred embodiment of a method for segmenting tooth images by fusing multi-level and multi-scale features of the present invention includes the following steps:

[0044] Step S1: Obtain a large number of historical tooth images, perform preprocessing on each of the historical tooth images, including labeling the filling teeth, diseased teeth, hyperplastic teeth, and background, and normalizing the pixel values, and construct a dataset based on the preprocessed historical tooth images; that is, the labeled tags are filling teeth, diseased teeth, hyperplastic teeth, or background; the pixel value normalization is to divide the pixel value by 255.0 to ensure that the range of each pixel value is between [0,1], so as to eliminate the influence brought by different pixel value ranges, make the tooth image segmentation model learn more stably, rather than making the tooth image segmentation model overly sensitive to certain features, that is, pixel value normalization helps to improve the performance of the tooth image segmentation model on different data distributions, improve the training stability, and accelerate the training process;

[0045] Step S2: Create a dental image segmentation model based on the Unet++ module and the output module. Set the loss function of the dental image segmentation model as the Focal Loss function, and set the optimizer of the dental image segmentation model as AdamW;

[0046] Step S3: Train the dental image segmentation model with the dataset, and deploy the trained dental image segmentation model;

[0047] Step S4: Segment the dental image with the deployed dental image segmentation model. That is, preprocess the obtained dental image and input it into the dental image segmentation model to output the dental image segmentation result After preprocessing, input it into the dental image segmentation model to output the dental image segmentation result where, represents a real number, H represents the height of the dental image, W represents the width of the dental image, and both C and C’ represent the number of segmentation categories, that is, the number of labels, taking the value of 4 (filling teeth, diseased teeth, hyperplastic teeth, background).

[0048] The dental image segmentation model of the present invention learns from the dataset to train a mapping function, which is the weight file of the dental image segmentation model. It can classify each pixel point in the dental image and assign each pixel point to four predefined labels (filling teeth, diseased teeth, hyperplastic teeth, background), and finally generate an accurate dental image segmentation result (segmentation mask).

[0049] In the step S2, the Unet++ module is a U-shaped network structure with an asymmetric number of layers, constructed by five encoders and four decoders; each encoder and decoder are bridged through skip connections;

[0050] By setting the encoder to five layers and the decoder to four layers, the extra layer of the encoder can capture deeper features, which contain more abstract representations and help the dental image segmentation model understand the complex structures and patterns of dental images. Since the number of encoder layers exceeds that of the decoder, the decoder can combine with multiple encoder layers through skip connections during the upsampling process to achieve multi-scale feature fusion, thereby enhancing the ability of the dental image segmentation model to capture details.

[0051] The encoder is used to gradually reduce the spatial dimension of the historical dental image through max-pooling operations, while increasing the number of feature channels to capture the hierarchical features of the historical dental image;

[0052] The decoder is used to gradually restore the spatial dimension (spatial resolution) of historical tooth images through upsampling, and combine each of the historical tooth images with restored spatial dimensions with the corresponding and deeper-level historical tooth images in the encoder to capture the scale features of the historical tooth images;

[0053] That is, through a layer-by-layer super-resolution strategy in each layer of the model, the resolution is gradually increased through upsampling, so that the decoding process of each layer can more finely restore the details of the image to enhance the model's ability to restore image details.

[0054] The output module is constructed based on a fully connected layer and is used to fuse each of the hierarchical features and scale features to output a tooth image segmentation result, that is, to output a tooth image segmentation result that matches the size of the original historical tooth image.

[0055] Each layer of the encoder and decoder, as well as the nested nodes (skip connections / upsampling paths connecting the decoder and encoder), contains a VGGBlock module for fusing feature maps (tooth images) at different levels input to the node (nested node, encoder node, and decoder node).

[0056] Assume that the size of the input tooth image in the first layer of the encoder is [H, W, C], and the size of the output feature map after passing through the first layer of the encoder is [H / 2, W / 2, 32]. During the entire downsampling process of the encoder, each layer of the encoder performs two convolutions on the input feature map. Except for the fifth layer, each layer of the encoder performs one downsampling on the input feature map, and the size of the feature map is gradually reduced to {1 / 2, 1 / 4, 1 / 8, 1 / 16} of the original image, and the number of feature channels gradually increases to {32, 64, 128, 256, 512}. At the same time, the encoder nodes (each layer of the encoder is regarded as an encoder node, and similarly, each layer of the decoder is regarded as a decoder node) input four scales of feature maps to the decoder and nested nodes through upsampling and skip connections, and the up and down convolutions layer by layer realize feature fusion. The nested structure and the decoder fuse feature maps of different scales through upsampling and skip connections to encoder nodes or other nested nodes. At the same time, the number of channels of the feature map in the decoder gradually decreases to {256, 128, 64, 32} through upsampling, and the size is gradually restored to the size of the original tooth image, and finally a tooth image segmentation result is output through a 1×1 convolution.

[0057] In each node of the nested structure, all nodes receive the feature maps output by upsampling from other nodes and one or more skip connections. The output node can be an encoder node or other nested nodes. Multiple input feature maps are concatenated in the channel dimension, and then the concatenated feature maps are subjected to feature fusion through a VGGBlock module. The VGGBlock module processes the input feature maps through two consecutive 3×3 convolutional layers. Both of these convolutional layers use convolutional kernels of the same size to enhance the ability to learn local features, that is, the consecutive 3×3 convolutional layers can focus on smaller regions during each convolution operation, which helps to capture fine-grained features in dental images, such as edges, textures, etc. And by stacking two convolutional layers, the model can learn more complex features through deeper non-linear transformations. Compared with other larger-sized convolutional kernels (such as 5×5 or 7×7), it has fewer parameters and higher computational efficiency. Since the weights of each convolutional layer are learned independently, different feature information can be captured. After the convolutional layer, BatchNorm is used for feature normalization, and the ReLU activation function is used to filter important features to enhance the model's ability to express features. A connection layer is set before each convolutional layer, and through dense connections (multiple dense skip connections), the encoder feature map is made closer to the feature map waiting in the decoder.

[0058] Let F i_j be the output of node X i_j , where i represents the downsampling layer of the encoder, j represents the convolutional layers of each node (encoder node, decoder node, nested node) in the path of the skip connection, and the calculation formula of F i_j is:

[0059]

[0060] Among them, VGG() represents the operation of the VGGBlock module; U() represents the upsampling operation; when j = 0, it represents the starting layer of the skip connection.

[0061] To further improve the segmentation accuracy, the present invention can optionally enable deep supervision in the output module:

[0062]

[0063] Among them, F i represents multiple feature maps finally generated by the model; represents the convolutional operation in the i-th stage, that is, converting F i into the prediction of the number of categories (labels); Output iDenote the output feature map of the $i$-th stage, with the shape of (batch_size, num_classes, height, width); deep_supervision represents the flag for enabling deep supervision. Whether to enable deep supervision is determined by the complexity of the dataset and the difficulty of optimizing the training of the deep neural network. The meaning of the above formula is: when deep supervision is not enabled, only the prediction results of the 4th stage are output; when deep supervision is enabled, the prediction results of multiple stages are output simultaneously, and an additional supervision loss is introduced to assist training.

[0064] In the step S2, the Focal Loss function is used to adjust the weights of easy-to-classify samples and hard-to-classify samples, making the tooth image segmentation model more focused on the learning of hard-to-classify samples, that is, used to solve the class imbalance problem, accelerate the convergence speed of the tooth image segmentation model, and improve the performance and generalization ability of the tooth image segmentation model.

[0065] The formula of the Focal Loss function is:

[0066] FL = -α t (1 - p t ) γ log(p t );

[0067] where $p$ t represents the predicted probability of the model for class $t$; $\alpha$ t represents the weight coefficient for balancing the importance of different classes, $\alpha$ t = [0.25, 0.25, 0.25, 0.1]; $\gamma$ represents the focusing parameter for adjusting the weights of easy-to-classify samples and hard-to-classify samples, $\gamma = 3$.

[0068] The calculation formula of $p$ t is:

[0069]

[0070] where $pred=(pred_1, pred_2,...., pred$ C ) represents the output of the model prediction; $C$ represents the total number of classes, taking the value of 4; represents the exponential function of $pred$ i ; represents each element in the probability distribution after Softmax calculation, and the sum of all is 1, that is The shape is (batch_size, num_classes, height, width); lables represent the true labels, with the shape of (batch_size, 1, height, width); the gather() function selects the probability p of the corresponding class for each pixel according to the labels t , p t contains the probability of the corresponding class for each pixel, with the shape of (batch_size, height, width, 1).

[0071] Select the weight α from the α array according to the class C t , multiply the Focal Loss of each sample by the corresponding weight α t , use the formula of FL to calculate the average or sum:

[0072]

[0073] Among them, FL i represents the Focal Loss of each sample; N represents the total number of samples; size_average is used to determine whether to calculate the average or sum.

[0074] In the step S2, the AdamW is used to prevent overfitting during the training of the dental image segmentation model.

[0075] AdamW is a variant of Adam. Compared with Adam, it directly integrates weight decay (L2 regularization) into the optimizer instead of being a separate part of the parameter update, which can more effectively control overfitting. Even when using a relatively large learning rate during the training of the dental image segmentation model, it can well maintain stability and improve the generalization ability of the dental image segmentation model.

[0076] The step S3 is specifically as follows:

[0077] The data set is divided into a training set and a test set at a split ratio of 4:1, and the tooth image segmentation model is trained by the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the optimizer is continuously optimized, including at least the update step size (Learning Rate), the decay rate of the first-order moment estimate (β1), the decay rate of the second-order moment estimate (β2), the L2 regularization strength coefficient (Weight Decay), the zero division error suppression coefficient (Epsilon), and the learning rate (Learning Rate) hyperparameters; that is, the hyperparameter combination that optimizes the performance of the tooth image segmentation model is found; the zero division error suppression coefficient is a very small constant to prevent zero division errors when calculating gradients; in specific implementation, the default values ​​of Learning Rate, β1, β2, and Weight Decay can be kept unchanged, and the Learning Rate is set to 2e-3, and AdamW automatically adjusts the Learning Rate as needed during the training process;

[0078] The trained tooth image segmentation model is tested using the test set, and the tooth image segmentation model that passes the test is deployed.

[0079] A preferred embodiment of a tooth image segmentation system integrating multi-level and scale features of the present invention includes the following modules:

[0080] A data set construction module is used to obtain a large number of historical tooth images, annotate each of the historical tooth images with filled teeth, diseased teeth, hyperplastic teeth, and background, and perform pixel value normalization preprocessing, and construct a data set based on each of the preprocessed historical tooth images; that is, the annotated labels are filled teeth, diseased teeth, hyperplastic teeth, or background; the pixel value normalization is to divide the pixel value by 255.0 to ensure that the range of each pixel value is between [0,1] to eliminate the influence of different pixel value ranges, so that the tooth image segmentation model can learn more stably, rather than making the tooth image segmentation model overly sensitive to certain features, that is, pixel value normalization helps to improve the performance of the tooth image segmentation model on different data distributions, improve training stability, and accelerate the training process;

[0081] A tooth image segmentation model creation module, used to create a tooth image segmentation model based on the Unet++ module and the output module, set the loss function of the tooth image segmentation model to the Focal Loss function, and set the optimizer of the tooth image segmentation model to AdamW;

[0082] A tooth image segmentation model training module, used to train the tooth image segmentation model using the data set and deploy the trained tooth image segmentation model;

[0083] A tooth image segmentation module for segmenting tooth images by means of the deployed tooth image segmentation model. That is, for the acquired tooth images After preprocessing, they are input into the tooth image segmentation model, and the tooth image segmentation results are output Among them, represents a real number, H represents the height of the tooth image, W represents the width of the tooth image, and both C and C’ represent the number of segmentation categories, that is, the number of labels, with a value of 4 (filling, diseased tooth, hyperplastic tooth, background).

[0084] The tooth image segmentation model of the present invention learns from a data set to train a mapping function, which is the weight file of the tooth image segmentation model and can classify each pixel point in the tooth image, assign each pixel point to four predefined labels (filling, diseased tooth, hyperplastic tooth, background), and finally generate an accurate tooth image segmentation result (segmentation mask).

[0085] In the tooth image segmentation model creation module, the Unet++ module is an asymmetric-layer U-shaped network structure constructed by five encoders and four decoders; a skip connection is used to bridge between each of the encoders and decoders;

[0086] By setting the encoder to five layers and the decoder to four layers, the extra layer of the encoder can capture deeper features, which contain more abstract representations and help the tooth image segmentation model understand the complex structures and patterns of tooth images. Since the number of encoder layers exceeds that of the decoder, the decoder can be combined with multiple encoder layers through skip connections during the upsampling process to achieve multi-scale feature fusion, thereby enhancing the ability of the tooth image segmentation model to capture details.

[0087] The encoder is used to gradually reduce the spatial dimension of historical tooth images through max-pooling operations while increasing the number of feature channels to capture the hierarchical features of the historical tooth images;

[0088] The decoder is used to gradually restore the spatial dimension (spatial resolution) of historical tooth images through upsampling, and combine each of the historical tooth images with restored spatial dimensions with the corresponding levels and deeper levels of historical tooth images in the encoder to capture the scale features of the historical tooth images;

[0089] That is, through a layer-by-layer super-resolution strategy in each layer of the model, the resolution is gradually increased through upsampling, so that the decoding process of each layer can more finely restore the details of the image to enhance the model's ability to restore image details.

[0090] The output module is constructed based on a fully connected layer and is used to fuse each of the hierarchical features and scale features, and output a tooth image segmentation result, that is, output a tooth image segmentation result that matches the size of the original historical tooth image.

[0091] Each layer of the encoder and decoder, as well as the nested nodes (skip connections / upsampling paths connecting the decoder and encoder), contains a VGGBlock module, which is used to perform feature fusion on feature maps (tooth images) at different levels input to the node (nested node, encoder node, and decoder node).

[0092] Assume that the size of the input tooth image of the first layer of the encoder is [H, W, C]. After passing through the first layer of the encoder, the size of the output feature map is [H / 2, W / 2, 32]. During the entire encoder downsampling process, each layer of the encoder performs two convolutions on the input feature map. Except for the fifth layer, each layer of the encoder performs one downsampling on the input feature map, and the size of the feature map is gradually reduced to {1 / 2, 1 / 4, 1 / 8, 1 / 16} of the original image, and the number of feature channels gradually increases to {32, 64, 128, 256, 512}. At the same time, encoder nodes (each layer of the encoder is regarded as an encoder node, and similarly, each layer of the decoder is regarded as a decoder node) input four scales of feature maps to the decoder and nested nodes through upsampling and skip connections, and feature fusion is achieved through convolution layer by layer upward and downward. The nested structure and the decoder both fuse feature maps of different scales through upsampling and skip connections to encoder nodes or other nested nodes. At the same time, the decoder reduces the number of channels of the feature map layer by layer to {256, 128, 64, 32} through upsampling, and the size is gradually restored to the size of the original tooth image. Finally, a 1×1 convolution is used to output the tooth image segmentation result.

[0093] In each node of the nested structure, all nodes receive the feature maps output by upsampling from other nodes and one or more skip connections. The output node can be an encoder node or other nested nodes. Multiple input feature maps are concatenated in the channel dimension, and then the concatenated feature maps are subjected to feature fusion through a VGGBlock module. The VGGBlock module processes the input feature maps through two consecutive 3×3 convolutional layers. Both of these convolutional layers use the same-sized convolutional kernels to enhance the ability to learn local features. That is, the consecutive 3×3 convolutional layers can focus on smaller regions during each convolution operation, which helps to capture fine-grained features in the tooth image, such as edges and textures. Moreover, by stacking two convolutional layers, the model can learn more complex features through deeper non-linear transformations. Compared with other larger-sized convolutional kernels (such as 5×5 or 7×7), it has fewer parameters and higher computational efficiency. Since the weights of each convolutional layer are learned independently, different feature information can be captured. After the convolutional layer, BatchNorm is used for feature normalization, and the ReLU activation function is used to filter important features to enhance the model's ability to express features. A connection layer is set before each convolutional layer, and through dense connections (multiple dense skip connections), the encoder feature maps are made closer to the feature maps waiting in the decoder.

[0094] Let F i_j be the output of node X i_j , where i represents the downsampling layer of the encoder, j represents the convolutional layer of each node (encoder node, decoder node, nested node) in the path of the skip connection, and F i_j is calculated as follows:

[0095]

[0096] Among them, VGG() represents the operation of the VGGBlock module; U() represents the upsampling operation; when j = 0, it represents the starting layer of the skip connection.

[0097] To further improve the segmentation accuracy, the present invention can choose whether to enable deep supervision in the output module:

[0098]

[0099] where F i represents multiple feature maps finally generated by the model; represents the convolutional operation in the i-th stage, that is, converting F i into the prediction of the number of categories (labels); Output iDenote the output feature map of the $i$-th stage, with the shape of (batch_size, num_classes, height, width); deep_supervision represents the flag for enabling deep supervision. Whether to enable deep supervision is determined by the complexity of the dataset and the optimization difficulty of training the deep neural network. The meaning of the above formula is: when deep supervision is not enabled, only the prediction results of the 4th stage are output; when deep supervision is enabled, the prediction results of multiple stages are output simultaneously, and additional supervision losses are introduced to assist training.

[0100] In the tooth image segmentation model creation module, the Focal Loss function is used to adjust the weights of easy-to-classify samples and hard-to-classify samples, making the tooth image segmentation model more focused on the learning of hard-to-classify samples, that is, used to solve the class imbalance problem, accelerate the convergence speed of the tooth image segmentation model, and improve the performance and generalization ability of the tooth image segmentation model.

[0101] The formula of the Focal Loss function is:

[0102] FL = -α t (1 - p t ) γ log(p t );

[0103] where $p$ t represents the prediction probability of the model for class $t$; $\alpha$ t represents the weight coefficient for balancing the importance of different classes, $\alpha$ t = [0.25, 0.25, 0.25, 0.1]; $\gamma$ represents the focusing parameter for adjusting the weights of easy-to-classify samples and hard-to-classify samples, $\gamma = 3$.

[0104] The calculation formula of $p$ t is:

[0105]

[0106] where $pred=(pred_1, pred_2,...., pred$ C ) represents the output of the model prediction; $C$ represents the total number of classes, taking the value of 4; represents the exponential function of $pred$ i ; represents each element in the probability distribution after Softmax calculation, and the sum of all is 1, that is The shape is (batch_size, num_classes, height, width); lables represent the true labels, with the shape of (batch_size, 1, height, width); the gather() function selects the probability p of the corresponding class for each pixel according to the labels t , p t contains the probability of the corresponding class for each pixel, with the shape of (batch_size, height, width, 1).

[0107] Select the weight α from the α array according to the class C t , multiply the Focal Loss of each sample by the corresponding weight α t , and use the formula of FL to calculate the average or sum:

[0108]

[0109] Among them, FL i represents the Focal Loss of each sample; N represents the total number of samples; size_average is used to determine whether to calculate the average or sum.

[0110] In the tooth image segmentation model creation module, the AdamW is used to prevent overfitting during the training of the tooth image segmentation model.

[0111] AdamW is a variant of Adam. Compared with Adam, it integrates weight decay (L2 regularization) directly into the optimizer instead of as a separate part of parameter update, which can more effectively control overfitting. Even when the tooth image segmentation model uses a large learning rate during training, it can maintain good stability and improve the generalization ability of the tooth image segmentation model.

[0112] The tooth image segmentation model training module is specifically used for:

[0113] Divide the dataset into a training set and a test set according to a splitting ratio of 4:1, and train the dental image segmentation model with the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the optimizer, which at least includes hyperparameters such as the update step size (Learning Rate), the decay rate of the first-order moment estimation (β1), the decay rate of the second-order moment estimation (β2), the L2 regularization strength coefficient (Weight Decay), the zero-division error suppression coefficient (Epsilon), and the learning rate (Learning Rate); that is, find the combination of hyperparameters that makes the performance of the dental image segmentation model the best. The zero-division error suppression coefficient is a very small constant to prevent division-by-zero errors when calculating gradients. Specifically, in implementation, the default values of Learning Rate, β1, β2, and Weight Decay can be kept unchanged, while the Learning Rate is set to 2e-3, and AdamW automatically adjusts the Learning Rate as needed during the training process;

[0114] Test the trained dental image segmentation model with the test set, and deploy the dental image segmentation model that passes the test.

[0115] In summary, the advantages of the present invention are as follows:

[0116] By obtaining a large number of historical dental images, preprocessing each historical dental image, including filling teeth, diseased teeth, hyperplastic teeth, and background annotation as well as pixel value normalization, a dataset is constructed. Then, based on the Unet++ module and the output module, a dental image segmentation model is created. The loss function of the dental image segmentation model is set as the Focal Loss function, and the optimizer is AdamW. Then, the dental image segmentation model is trained with the dataset, and the trained dental image segmentation model is deployed. Finally, the dental image segmentation is performed through the deployed dental image segmentation model. The Unet++ module has a U-shaped network structure, which is constructed by five layers of encoders and four layers of decoders. A skip connection is used to bridge between each encoder and decoder. The encoder is used to gradually reduce the spatial dimension of the historical dental image through max-pooling operations while increasing the number of feature channels to capture the hierarchical features of the historical dental image. The decoder is used to gradually restore the spatial dimension of the historical dental image through upsampling and combine each historical dental image with the corresponding and deeper-level historical dental images in the encoder to capture the scale features of the historical dental image. That is, the Unet++ module fully extracts the local features of the dental image and captures long-range dependency information through a multi-scale feature fusion mechanism, filters out the key information in the channels through skip connections, makes the decoding process of each layer more refined through a layer-by-layer super-resolution strategy (upsampling gradually restores the spatial dimension), adjusts the weights of easy-to-classify samples and difficult-to-classify samples through the Focal Loss function to solve the class imbalance problem, enhances the attention to diseased tooth pixels (difficult-to-classify samples), and combines with the AdamW optimizer to prevent overfitting, ultimately greatly improving the generalization ability and accuracy of dental image segmentation.

[0117] Although the specific implementation manners of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.

Claims

1. A tooth image segmentation method integrating multi-level and scale features, characterized in that: The steps include: Step S1, obtaining a large number of historical tooth images, annotating each of the historical tooth images with filled teeth, diseased teeth, hyperplastic teeth, and background, and preprocessing pixel value normalization, and constructing a data set based on each of the preprocessed historical tooth images; Step S2, creating a tooth image segmentation model based on the Unet++ module and the output module, setting the loss function of the tooth image segmentation model to the Focal Loss function, and setting the optimizer of the tooth image segmentation model to AdamW; Step S3, training the tooth image segmentation model using the data set, and deploying the trained tooth image segmentation model; Step S4: segment the tooth image using the deployed tooth image segmentation model.

2. The tooth image segmentation method of claim 1, characterized in that: In step S2, the Unet++ module is a U-shaped network structure, which is constructed by a five-layer encoder and a four-layer decoder; each encoder and decoder is bridged by a jump connection; The encoder is used to gradually reduce the spatial dimension of the historical dental image through a maximum pooling operation, while increasing the number of feature channels to capture the hierarchical features of the historical dental image; The decoder is used to gradually restore the spatial dimensions of the historical dental images by upsampling, and combine each of the historical dental images with the restored spatial dimensions with the historical dental images of the corresponding level and deeper levels in the encoder to capture the scale characteristics of the historical dental images; The output module is constructed based on a fully connected layer, and is used to fuse the hierarchical features and scale features, and output a tooth image segmentation result.

3. The tooth image segmentation method of claim 1, characterized in that: In step S2, the Focal Loss function is used to adjust the weights of easy-to-classify samples and difficult-to-classify samples, so that the tooth image segmentation model is more focused on learning difficult-to-classify samples.

4. The tooth image segmentation method of claim 1, characterized in that: In step S2, the AdamW is used to prevent overfitting during the training of the tooth image segmentation model.

5. The tooth image segmentation method of claim 1, characterized in that: The step S3 is specifically as follows: The data set is divided into a training set and a test set at a segmentation ratio of 4:1, and the tooth image segmentation model is trained by the training set until the loss value of the loss function is less than a preset loss threshold, and the optimizer is continuously optimized during the training process, including at least the update step size, the decay rate of the first-order moment estimate, the decay rate of the second-order moment estimate, the L2 regularization strength coefficient, the zero division error suppression coefficient, and the hyperparameters of the learning rate; The trained tooth image segmentation model is tested using the test set, and the tooth image segmentation model that passes the test is deployed.

6. A tooth image segmentation system integrating multi-level and scale features, characterized by: Includes the following modules: A data set construction module is used to obtain a large number of historical tooth images, perform preprocessing on each of the historical tooth images to mark the filled teeth, diseased teeth, hyperplastic teeth, and background, and normalize the pixel values, and construct a data set based on each of the preprocessed historical tooth images; A tooth image segmentation model creation module, used to create a tooth image segmentation model based on the Unet++ module and the output module, set the loss function of the tooth image segmentation model to the Focal Loss function, and set the optimizer of the tooth image segmentation model to AdamW; A tooth image segmentation model training module, used to train the tooth image segmentation model using the data set and deploy the trained tooth image segmentation model; The tooth image segmentation module is used to segment the tooth image by using the deployed tooth image segmentation model.

7. The tooth image segmentation system with multi-level and scale feature fusion as claimed in claim 6, characterized in that: In the tooth image segmentation model creation module, the Unet++ module is a U-shaped network structure, which is constructed by a five-layer encoder and a four-layer decoder; each encoder and decoder is bridged by a jump connection; The encoder is used to gradually reduce the spatial dimension of the historical dental image through a maximum pooling operation, while increasing the number of feature channels to capture the hierarchical features of the historical dental image; The decoder is used to gradually restore the spatial dimensions of the historical dental images by upsampling, and combine each of the historical dental images with the restored spatial dimensions with the historical dental images of the corresponding level and deeper levels in the encoder to capture the scale characteristics of the historical dental images; The output module is constructed based on a fully connected layer, and is used to fuse the hierarchical features and scale features, and output a tooth image segmentation result.

8. The tooth image segmentation system with multi-level and scale feature fusion as claimed in claim 6, characterized in that: In the tooth image segmentation model creation module, the Focal Loss function is used to adjust the weights of easy-to-classify samples and difficult-to-classify samples, so that the tooth image segmentation model is more focused on learning difficult-to-classify samples.

9. The tooth image segmentation system with multi-level and scale feature fusion as claimed in claim 6, characterized in that: In the tooth image segmentation model creation module, the AdamW is used to prevent overfitting during tooth image segmentation model training.

10. The tooth image segmentation system with multi-level and scale feature fusion as claimed in claim 6, characterized in that: The tooth image segmentation model training module is specifically used for: The data set is divided into a training set and a test set at a segmentation ratio of 4:1, and the tooth image segmentation model is trained by the training set until the loss value of the loss function is less than a preset loss threshold, and the optimizer is continuously optimized during the training process, including at least the update step size, the decay rate of the first-order moment estimate, the decay rate of the second-order moment estimate, the L2 regularization strength coefficient, the zero division error suppression coefficient, and the hyperparameters of the learning rate; The trained tooth image segmentation model is tested using the test set, and the tooth image segmentation model that passes the test is deployed.