Oral cavity panoramic X-ray image tooth segmentation method based on cavity convolution segmentation network

By integrating dilated convolutions and data augmentation techniques, the method addresses the limitations of deep learning in dental segmentation, enhancing generalization and accuracy in panoramic X-ray image analysis.

CN120318506APending Publication Date: 2025-07-15HUNAN FIRST NORMAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251350.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing oral panoramic X-ray image tooth segmentation method based on deep learning has problems such as poor generalization ability and large demand for training samples. Conventional convolution leads to loss of image details and severe overfitting of the model, making it difficult to effectively segment teeth in practical applications.

Method used

The hollow convolution segmentation network is used to amplify the training data sets with stochastic translation rotation transformation, image distortion deformation, grayscale value transformation and noise injection techniques to improve the diversity and richness of the model, and to extract image features using hollow convolution in the downsampling stage, and optimize model performance with Focal loss function and attention mechanism.

Benefits of technology

It significantly improves the generalization ability and segmentation performance of the model, improves the accuracy and robustness of tooth segmentation, and enhances the applicability and accuracy of the model under different conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318506A_ABST
    Figure CN120318506A_ABST
Patent Text Reader

Abstract

The invention discloses an oral cavity panoramic X-ray image tooth segmentation method based on a cavity convolution segmentation network. The method comprises the following steps: acquiring an oral cavity panoramic X-ray image and preprocessing the oral cavity panoramic X-ray image; a tooth area in the oral cavity panoramic X-ray image is marked in a manual mode, and a third person independently checks a marking result; performing geometric transformation, gray value transformation, noise injection and distortion on the oral panoramic X-ray image; an improved cavity convolution SegNet semantic segmentation network is adopted to train an oral cavity panoramic X-ray image tooth segmentation model; and comprehensively evaluating the performance of the training model based on three indexes of pixel accuracy, intersection-to-union ratio and contour matching degree, and converting the tooth segmentation model meeting the requirement into an ONNX format and exporting the tooth segmentation model. According to the invention, on the basis of a semantic segmentation network model, oral panoramic X-ray image training data is augmented through various image transformation so as to improve the generalization ability of the model; and the convolution kernel receptive field is improved by adopting cavity convolution so as to improve the segmentation performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing, and particularly relates to a method for segmenting teeth in panoramic dental X-ray images based on a dilated convolutional segmentation network. Background Art

[0002] With the continuous improvement of living standards and medical technology, oral health has received increasing attention. In oral medical cases, dental diagnosis and treatment play a dominant role. Dental diagnosis and treatment mainly involve the diagnosis and treatment of the internal and external morphology, lesions, and distribution of hard dental tissues. Among them, the misidentification of dental morphological characteristics is the main factor leading to the failure of dental treatment. Therefore, accurately identifying the dental morphology of patients is of great value and significance for dental diagnosis and treatment. Moreover, based on panoramic dental X-ray images, tooth segmentation is the basis for realizing dental morphology recognition. In addition, with the rapid development of computer technology and big data technology, digital oral medicine and computer-aided diagnosis have become the future development trend. Accurately segmenting the tooth region in X-ray scanned oral images and converting it into a direct or indirect reference basis for clinical diagnosis is the basis for realizing the above-mentioned cutting-edge technologies.

[0003] Some scholars have conducted certain research on tooth segmentation in panoramic dental X-ray images. Their technical routes can be mainly divided into threshold-based segmentation methods, edge-based segmentation methods, energy function-based segmentation methods, and deep learning-based segmentation methods. The threshold-based segmentation method mainly uses the statistical characteristics of the gray values of hard dental tissues to segment teeth from the image background. This method is greatly affected by noise, scanning parameters, and patient individual characteristics, and the segmentation results often require manual correction. The edge-based segmentation method segments teeth according to the characteristics of the sudden change in gray values at the boundary between teeth and the background. Noise, boundary clarity, interfering objects, etc. have a great impact on the segmentation results of this method, making it difficult to apply in clinical practice. The energy function-based segmentation method obtains the tooth contour by iteratively solving the position of the curve when the minimum value of the energy functional is defined according to the defined energy function. This method is greatly affected by the number of iterations and the initial position of the curve, and has poor timeliness.

[0004] The tooth segmentation method based on deep learning utilizes machine learning and related technologies to achieve tooth segmentation by learning oral panoramic X-ray image training samples, demonstrating good segmentation effects and application prospects. However, there are still the following problems: (1) When using conventional convolution for image feature extraction, pooling processing is often required to extract high-dimensional features. However, the pooling process will lose the detailed information of the image and cannot be restored during the upsampling process, thus affecting the tooth segmentation results; (2) A large number of images with labels and strong diversity are required to train a relatively general and reliable tooth segmentation model. However, it is difficult to obtain a sufficient number of oral panoramic X-ray images with strong diversity in practice, and the model trained based on a small number of images is severely overfitted and often does not have generalization ability. Therefore, the tooth segmentation method based on deep learning is still significantly limited in practical applications.

[0005] Based on the semantic segmentation model of deep learning, dilated convolution is introduced to replace conventional convolution to increase the receptive field of the convolution kernel. While retaining the detailed information of the image, downsampling of the image is performed, which can improve the segmentation performance of the model; Techniques such as random translation and rotation transformation, image warping, contrast adjustment, and noise injection are used to augment the training oral panoramic X-ray images and labels, improving the diversity and richness of the training dataset, which can improve the generalization ability of the model to a certain extent. For this reason, the applicant has conducted beneficial exploration and attempts and found solutions to the above problems. The invention solution to be introduced below was generated under such a background. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned prior art and to solve the problems of poor generalization ability and large demand for training samples in the tooth segmentation of oral panoramic X-ray images based on deep learning, the present invention provides a tooth segmentation method for oral panoramic X-ray images based on a dilated convolution segmentation network. By introducing dilated convolution, the segmentation performance of the model is improved, and techniques such as random translation and rotation transformation, image warping, gray value transformation, and noise injection are used to augment the diversity and richness of the training dataset, improving the universality and generalization ability of the model.

[0007] The object of the present invention is achieved through the following technical solutions:

[0008] A tooth segmentation method for oral panoramic X-ray images based on a dilated convolution segmentation network includes the following steps:

[0009] S1: Obtain an oral panoramic X-ray image dataset and perform preprocessing;

[0010] S2: Annotate the oral panoramic X-ray images;

[0011] S3: Augment the data of the oral panoramic X-ray images;

[0012] S4: Construct and train an oral panoramic X-ray image tooth segmentation model based on a dilated convolutional semantic segmentation network;

[0013] S5: Test and validate the oral panoramic X-ray image tooth segmentation model;

[0014] S6: Perform format conversion and export the oral panoramic X-ray image tooth segmentation model;

[0015] Furthermore, step S1 is specifically as follows:

[0016] The oral panoramic X-ray image dataset used contains open-source data from Alibaba Cloud Tianchi and case data from relevant hospitals. To facilitate the training and testing of the deep learning model, all X-ray images are converted into 8-bit PNG format. After conversion, the pixel gray value of the oral panoramic X-ray image is:

[0017]

[0018] where round is the rounding function, S png is the pixel gray value of the converted image, S org is the pixel gray value of the original X-ray image, S min and S max are the minimum and maximum values of the pixel gray value of the original X-ray image respectively. After the oral panoramic X-ray image is converted into PNG format, the resolution of all images is adjusted to 512×256 pixels, and the bilinear interpolation method is used to calculate the gray value of the image after resolution adjustment. The specific formula is:

[0019]

[0020] where Ψ is the bilinear interpolation sampling operator, I png is the PNG format image matrix after format conversion, I R is the image matrix after size conversion, W org and H org are the pixel width and pixel height of the PNG format image after size conversion respectively. The oral panoramic X-ray image after size conversion is as shown in Figure 2 shown.

[0021] Furthermore, step S2 is specifically as follows:

[0022] Mark the tooth area by manual annotation, and use software such as LabelMe or ImageLabel for annotation. To ensure the accuracy of the annotation, the annotation results are independently checked by a third party.

[0023] Furthermore, step S3 is specifically as follows:

[0024] S31: Geometric transformation

[0025] The geometric transformation includes randomly flipping, translating, rotating, and scaling the panoramic dental X-ray image and its label. In the geometric transformation, the random translation scale in the height direction is controlled within [-0.1H, 0.1H], where H is the height of the panoramic dental X-ray image after size conversion (256 pixels), and the random translation scale in the width direction is controlled within [-0.1W, 0.1W], where W is the width of the panoramic dental X-ray image after size conversion (512 pixels); the random rotation angle is controlled within [-10°, 10°], and the rotation center is the center of the image; the scaling factor of the random scaling is controlled within [0.9, 1.1]. When the scaling factor is greater than 1, the enlarged image is cropped to 512×256 pixels with the center as the reference.

[0026] S32: Gray value transformation

[0027] During the gray value transformation, the image label remains unchanged. Random brightness adjustment and random contrast adjustment are performed on the panoramic dental X-ray image after size conversion. The pixel gray value of the image after random brightness adjustment is:

[0028] S A = min(αS O + β, 255)

[0029] where α is the brightness gain coefficient, β is the brightness offset, S O is the pixel gray value of the panoramic dental X-ray image after size conversion, and S A is the pixel gray value of the panoramic dental X-ray image after brightness adjustment. In the random contrast adjustment, histogram equalization, logarithmic transformation, gamma transformation, and local histogram equalization are performed on the panoramic dental X-ray image, so as to change the contrast of the panoramic dental X-ray image with multiple strategies.

[0030] S33: Noise injection

[0031] In the noise injection, the image label remains unchanged. Gaussian noise, salt-and-pepper noise, uniform noise, and Poisson noise are randomly added to the panoramic dental X-ray image after size conversion. The mean value of the added Gaussian noise is set to 0, and the variance is not greater than 0.02; the noise density of the added salt-and-pepper noise is not greater than 0.05; the variance of the added uniform noise is not greater than 0.05; the Poisson noise is automatically calculated based on the input image without manual parameter setting.

[0032] S34: Distortion

[0033] During the image distortion process, the image and its labels change synchronously. Distorting the image and its labels to a certain extent and changing the shapes of teeth and the background can significantly improve the diversity and richness of samples, which plays an important role in enhancing the generalization ability and universality of the training model. The technology adopted for image distortion is block-based random non-linear transformation. This method first divides the image into multiple grid regions, and then applies random non-linear offset amounts to different regions of the image to achieve local distortion of the image, which can effectively increase the diversity of the image while maintaining the overall structure of the image. For block-based random non-linear transformation, it is first necessary to calculate the pixel coordinates of the deformed image, and then calculate the pixel gray values of the deformed image. The displacement field is used to calculate the pixel coordinates of the deformed image, and the specific formula is:

[0034]

[0035] where Τ(x,y) is the displacement field function, representing the new coordinates (x',y') after the deformation of the input coordinates (x,y). x and y are the original pixel coordinates of the image before deformation, and x' and y' are the pixel coordinates of the image after deformation. K and L are the number of horizontal blocks and the number of vertical blocks of the image respectively, and φ kl (x,y) is the interpolation weight function affected by the block vertices k and l, and Δx kl 、Δy kl are the random displacement amounts of the block vertices in the x and y directions respectively.

[0036] The bilinear interpolation method is used to calculate the pixel gray values of the deformed image, and the specific formula is:

[0037]

[0038] where I def (x',y') is the pixel gray value of the deformed image, I R (i,j) is the pixel gray value of the image after size conversion, ω i (x) is the horizontal interpolation weight, ω j (y) is the vertical interpolation weight, x f 、x c are the values of x rounded down and rounded up respectively, and y f 、y c are the values of y rounded down and rounded up respectively.

[0039] Furthermore, step S4 is specifically as follows:

[0040] The basic semantic segmentation network adopted by the tooth segmentation model is the SegNet network. During the network downsampling process, dilated convolutions with different dilation rates are used instead of conventional convolutions to extract image features, so as to improve the segmentation performance of the model. The formula for dilated convolution is:

[0041]

[0042] Among them, Y is the output feature map, X is the input feature map, ω is the convolution kernel, and r is the dilation rate. In the upsampling stage of the network, decoding is performed using the recorded pooling position information. Specifically, a sparse feature map is first generated according to the position information, and then restored to a dense feature map using convolution operations.

[0043] The downsampling stage of the semantic segmentation model consists of 4 groups of nodes. Each group of nodes contains multiple dilated convolution layers, batch normalization layers, activation layers, and pooling layers, where the dilation rate of each dilated convolution layer is different. Every time a group of nodes is passed through in the downsampling stage, the width and height of the feature map are halved. The upsampling stage of the semantic segmentation model also consists of 4 groups of nodes. During upsampling, the pooling index of the downsampling process is used to perform non-linear upsampling, and then convolution operations are used to generate a dense feature map. In the upsampling stage, the width and height of the feature map are doubled every time a group of nodes is passed through.

[0044] The semantic segmentation model uses the Focal loss function during training to solve the problem of uneven pixel ratio between teeth and the image background in oral panoramic X-ray images. The Focal loss function reduces the loss contribution to well-classified samples by introducing a modulation factor, thereby focusing the training on difficult-to-classify samples. The specific expression is:

[0045] Loss = -α t (1 - p t ) γ log(p t )

[0046] Among them, Loss is the calculated loss value, p t is the probability that the model predicts as the positive class; α t is the weight coefficient for balancing teeth and the background, taking a smaller value for positive samples and a larger value for negative samples; γ is the focusing parameter for modulating the weights of easy-to-classify samples, used to adjust the loss contribution of easy-to-classify samples. Hyperparameters during the training process are also the key to affecting the segmentation performance of oral panoramic X-ray images. The semantic segmentation network uses a gradually decreasing learning rate during training. Specifically, the initial learning rate is 0.01, and the learning rate is reduced by 20% every 10 training epochs; the optimization function is selected as sdgm.

[0047] An attention mechanism can be added to the semantic segmentation network. By calculating the similarity between the query vector and the key vector, the attention weights are determined, allowing the model to dynamically adjust its attention weights when processing input data, thereby highlighting important information and ignoring unimportant information. Additionally, an extra prediction head can be added to the intermediate layer of the semantic segmentation network to increase the depth supervision mechanism, promoting the intermediate layer to learn more meaningful features, thus alleviating the vanishing gradient problem and improving the overall performance of the model.

[0048] Further, step S5 is specifically as follows:

[0049] After completing the training of the oral panoramic X-ray image tooth segmentation model, it is necessary to comprehensively test and verify the performance of the model to confirm whether it meets the requirements. The metrics used in verifying the model performance include pixel accuracy, intersection over union (IoU), and contour matching degree. Pixel accuracy is used to evaluate the proportion of correctly segmented pixels in the oral panoramic X-ray image. The expression of this metric is:

[0050]

[0051] where T P is the number of true positive pixels, and F N is the number of false negative pixels. Intersection over union (IoU) is used to evaluate the overlap degree between the segmentation result of the oral panoramic X-ray image and the ground truth label. The expression of this metric is:

[0052]

[0053] where T P is the number of true positive pixels, F N is the number of false negative pixels, and F P is the number of false positive pixels. Contour matching degree is used to evaluate the contour similarity between the segmentation result and the ground truth label. Based on the comprehensive evaluation of the above three metrics, the performance of the trained model is evaluated. If the performance meets the requirements, the model is exported; if the performance does not meet the requirements, the improvement method is determined by analyzing the image features with large errors, and improvements are made in the network model and data augmentation stage, and then retraining and verification are carried out.

[0054] Further, step S6 is specifically as follows:

[0055] When the performance of the oral panoramic X-ray image tooth segmentation model meets the requirements, first convert the network model to the ONNX format, then test whether the performance of the converted model is consistent with the performance of the original network; finally, export the model so that it can be used for oral panoramic X-ray image tooth segmentation on multiple platforms and in multiple languages.

[0056] Compared with the prior art, the beneficial effects of the present application are as follows: The present application proposes a method for segmenting teeth in oral panoramic X-ray images based on a dilated convolutional segmentation network. This method is based on the semantic segmentation network SegNet model. By performing random geometric transformation, grayscale value transformation, noise injection, and distortion deformation to augment the training data of oral panoramic X-ray images, the richness and diversity of the training data are greatly increased, and the generalization ability and universality of the model are improved. By using dilated convolution in the downsampling stage of the semantic segmentation network to extract image features and using more global information for inference, the segmentation performance of the model is improved. Generally speaking, the method for segmenting teeth in oral panoramic X-ray images provided by the present invention has the advantages of strong generalization ability and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0058] Figure 1 It is a flowchart of the method for segmenting teeth in oral panoramic X-ray images based on a dilated convolutional segmentation network.

[0059] Figure 2 It is the preprocessed oral panoramic X-ray image.

[0060] Figure 3 It is the annotation result of the oral panoramic X-ray image. (a) is the preprocessed oral panoramic X-ray image, and (b) is the image label.

[0061] Figure 4 It is the oral panoramic X-ray image and its augmented images. (a) is the preprocessed oral panoramic X-ray image and label, (b) is the image and label after noise injection, (c) is the image and label after contrast adjustment, and (d) is the image and label after distortion deformation.

[0062] Figure 5 It is the network architecture of the semantic segmentation model for oral panoramic X-ray images in the solution of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] In order to make the technical means, creative features, achieved purposes, and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0064] Figure 1The flowchart of a tooth segmentation method for panoramic dental X-ray images based on a dilated convolution segmentation network provided by the present invention is shown. By introducing dilated convolution, the segmentation performance of the model is improved. Image augmentation techniques such as random translation and rotation transformation, image warping, gray value transformation, and noise injection are used to increase the diversity and richness of the training dataset, and improve the universality and generalization ability of the model. The implementation method shown in this embodiment includes the following steps:

[0065] S1: Obtain a panoramic dental X-ray image dataset and perform preprocessing

[0066] The dataset is the basis for training a deep learning model. The panoramic dental X-ray image dataset used in the present invention includes open-source data from Alibaba Cloud Tianchi and case data from relevant hospitals. To facilitate the training and testing of the deep learning model, all X-ray images are converted into 8-bit PNG format. After conversion, the pixel gray value of the panoramic dental X-ray image is:

[0067]

[0068] where round is the rounding function, S png is the pixel gray value of the converted image, S org is the pixel gray value of the original X-ray image, S min and S max are the minimum and maximum values of the pixel gray value of the original X-ray image respectively. After the panoramic dental X-ray image is converted into PNG format, the resolution of all images is adjusted to 512×256 pixels, and the bilinear interpolation method is used to calculate the gray value of the image after resolution adjustment. The specific formula is:

[0069]

[0070] where Ψ is the bilinear interpolation sampling operator, I png is the PNG format image matrix after format conversion, I R is the image matrix after size conversion, W org and H org are the pixel width and pixel height of the PNG format image after size conversion respectively. The panoramic dental X-ray image after size conversion is as Figure 2 shown.

[0071] S2: Annotate the panoramic dental X-ray image

[0072] Oral panoramic X-ray image annotation means marking the tooth area from the size-converted oral panoramic X-ray image, so as to distinguish the teeth from the image background for the training and verification of deep learning models. Since the clarity and contrast of oral panoramic X-ray images are poor, this solution uses manual annotation to mark the tooth area, and software such as LabelMe or ImageLabel is used for annotation. To ensure the accuracy of the annotation, the annotation results are independently checked by a third person. The oral panoramic X-ray image annotation results are as Figure 3 shown, where Figure 3 (a) is the preprocessed image, Figure 3 (b) is the annotation result.

[0073] S3: Data augmentation for oral panoramic X-ray images

[0074] In order to improve the generalization ability and universality of the semantic segmentation model under the premise of limited oral panoramic X-ray image data, the technical solution of the present invention uses image transformation technology to generate more training samples based on limited original images to expand the richness and diversity of training data. The data augmentation methods adopted in the technical solution of the present invention include: geometric transformation, gray value transformation, noise injection, and distortion.

[0075] S31: Geometric transformation

[0076] The geometric transformation includes randomly flipping, randomly translating, randomly rotating, and randomly scaling the oral panoramic X-ray image and its label. In the geometric transformation, the random translation scale in the height direction is controlled between [-0.1H, 0.1H], where H is the height (256 pixels) of the size-converted oral panoramic X-ray image, and the random translation scale in the width direction is controlled in the interval [-0.1W, 0.1W], where W is the width (512 pixels) of the size-converted oral panoramic X-ray image; the random rotation angle is controlled in the interval [-10°, 10°], and the rotation center is the center of the image; the scaling factor of random scaling is controlled in the interval [0.9, 1.1]. When the scaling factor is greater than 1, the enlarged image is cropped to 512×256 pixels based on the center.

[0077] S32: Gray value transformation

[0078] During the gray value transformation process, the image label remains unchanged, and the randomly adjusted brightness and randomly adjusted contrast are performed on the size-converted oral panoramic X-ray image. The pixel gray value of the image after random brightness adjustment is:

[0079] S A =min(αS O +β,255)

[0080] where α is the brightness gain coefficient, β is the brightness offset, and SO is the pixel gray value of the panoramic dental X-ray image after size conversion, S A is the pixel gray value of the panoramic dental X-ray image after brightness adjustment. In random contrast adjustment, histogram equalization, logarithmic transformation, gamma transformation, and local histogram equalization are performed on the panoramic dental X-ray image, so as to change the contrast of the panoramic dental X-ray image with multiple strategies.

[0081] S33: Noise injection

[0082] In noise injection, the image label remains unchanged, and Gaussian noise, salt-and-pepper noise, uniform noise, and Poisson noise are randomly added to the panoramic dental X-ray image after size conversion. The mean value of the added Gaussian noise is set to 0, and the variance is not greater than 0.02; the noise density of the added salt-and-pepper noise is not greater than 0.05; the variance of the added uniform noise is not greater than 0.05; Poisson noise is automatically calculated based on the input image without manual parameter setting.

[0083] S34: Distortion

[0084] During the image distortion process, the image and its label change synchronously. Distorting the image and its label to a certain extent and changing the shape of teeth and background can significantly improve the sample diversity and richness, and play an important role in improving the generalization ability and universality of the training model. The technology adopted for image distortion is block random non-linear transformation. This method first divides the image into multiple grid regions, and then applies random non-linear offset amounts to different regions of the image to achieve local distortion of the image, which can effectively increase the diversity of the image and at the same time maintain the overall structure of the image. Block random non-linear transformation first needs to calculate the pixel coordinates of the deformed image, and then calculate the pixel gray value of the deformed image. Among them, the displacement field is used to calculate the pixel coordinates of the deformed image, and the specific formula is:

[0085]

[0086] where Τ(x,y) is the displacement field function, representing the new coordinates (x',y') after the input coordinates (x,y) are deformed. x and y are the original pixel coordinates of the image before deformation, x' and y' are the pixel coordinates of the image after deformation, K and L are the number of horizontal blocks and vertical blocks of the image respectively, and φ kl (x,y) is the interpolation weight function affected by the block vertices k and l, and Δx kl and Δy kl are the random displacement amounts of the block vertices in the x and y directions respectively.

[0087] In the solution of the present invention, the bilinear interpolation method is used to calculate the pixel gray value of the deformed image, and the specific formula is:

[0088]

[0089] Among them, I def (x', y') is the pixel gray value of the deformed image, and I R (i, j) is the pixel gray value of the image after size conversion, and ω i (x) is the horizontal interpolation weight, and ω j (y) is the vertical interpolation weight, and x f and x c are the floor and ceiling values of x respectively, and y f and y c are the floor and ceiling values of y respectively. Figure 4 is the panoramic dental X-ray image and its augmented image, among which Figure 4 (a) is the panoramic dental X-ray image and label after size conversion, Figure 4 (b) is the image and label after noise injection, Figure 4 (c) is the image and label after contrast adjustment, Figure 4 (d) is the image and label after distortion.

[0090] S4: Construct and train a panoramic dental X-ray image tooth segmentation model based on the dilated convolutional semantic segmentation network

[0091] In the solution of the present invention, the basic semantic segmentation network adopted by the tooth segmentation model is the SegNet network. During the downsampling process of the network, dilated convolutions with different dilation rates are used instead of conventional convolutions to extract image features to improve the segmentation performance of the model. The formula for dilated convolution is:

[0092]

[0093] Among them, Y is the output feature map, X is the input feature map, ω is the convolution kernel, and r is the dilation rate. In the upsampling stage of the network, the recorded pooling position information is used for decoding. Specifically, a sparse feature map is first generated according to the position information, and then a dense feature map is restored by using convolution operations.

[0094] The downsampling stage of the semantic segmentation model consists of 4 groups of nodes. Each group of nodes contains multiple dilated convolutional layers, batch normalization layers, activation layers, and pooling layers, and the dilation rate of each layer of dilated convolution is different. After passing through each group of nodes in the downsampling stage, the width and height of the feature map are halved. The upsampling stage of the semantic segmentation model also consists of 4 groups of nodes. During upsampling, the pooling index of the downsampling process is used to perform non-linear upsampling, and then convolution operations are used to generate a dense feature map. In the upsampling stage, the width and height of the feature map are doubled after passing through each group of nodes. The network architecture of the semantic segmentation model is as Figure 5 shown.

[0095] The semantic segmentation model uses the Focal loss function during training to address the problem of uneven pixel ratios between teeth and the image background in panoramic dental X-ray images. The Focal loss function reduces the loss contribution to well-classified samples by introducing a modulation factor, thereby focusing the training on difficult-to-classify samples. The specific expression is:

[0096] Loss = -α t (1 - p t ) γ log(p t )

[0097] where Loss is the calculated loss value, and p t is the probability that the model predicts as the positive class; α t is the weight coefficient for balancing teeth and the background, taking a smaller value for positive samples and a larger value for negative samples; γ is the focusing parameter for modulating the weights of easy-to-classify samples, used to adjust the loss contribution of easy-to-classify samples. Hyperparameters during the training process are also crucial for affecting the segmentation performance of panoramic dental X-ray images. The semantic segmentation network uses a gradually decreasing learning rate during training, specifically with an initial learning rate of 0.01, and the learning rate is reduced by 20% every 10 training epochs; the optimization function is selected as sdgm.

[0098] An attention mechanism can be added to the semantic segmentation network to determine the attention weights by calculating the similarity between the query vector and the key vector, allowing the model to dynamically adjust its attention weights when processing input data, thereby highlighting important information and ignoring unimportant information; an additional prediction head can also be added to the intermediate layer of the semantic segmentation network to increase the depth supervision mechanism, promoting the intermediate layer to learn more meaningful features, thereby alleviating the problem of vanishing gradients and improving the overall performance of the model.

[0099] S5: Test and validate the panoramic dental X-ray image tooth segmentation model

[0100] After completing the training of the panoramic dental X-ray image tooth segmentation model, it is necessary to comprehensively test and validate the performance of the model to confirm whether it meets the requirements. The metrics used to validate the model performance include pixel accuracy, intersection over union (IoU), and contour matching degree. Pixel accuracy is used to evaluate the proportion of correctly segmented pixels in panoramic dental X-ray images. The expression for this metric is:

[0101]

[0102] where T P is the number of pixels of true positives, and F N is the number of pixels of false negatives. Intersection over union (IoU) is used to evaluate the overlap degree between the segmentation result of panoramic dental X-ray images and the ground truth label. The expression for this metric is:

[0103]

[0104] where T P is the number of pixels of true positives, F N is the number of pixels of false negatives, and F P is the number of pixels of false positives. The contour matching degree is used to evaluate the contour similarity between the segmentation result and the ground truth label. Based on the above three indicators, the performance of the trained model is comprehensively evaluated. If the performance meets the requirements, the model is exported; if the performance does not meet the requirements, the improvement method is determined by analyzing the image features with large errors, and improvements are made in the network model and data augmentation stage, and then retraining and validation are carried out again.

[0105] S6: Perform format conversion and export on the dental segmentation model for panoramic dental X-ray images

[0106] When the performance of the dental segmentation model for panoramic dental X-ray images meets the requirements, first convert the network model into the ONNX format, then test whether the performance of the converted model is consistent with the performance of the original network; finally, export the model so that it can be used for dental segmentation of panoramic dental X-ray images on multiple platforms and in multiple languages.

[0107] The above describes the basic principles and main features of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements. The scope of protection required by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for segmenting teeth in panoramic dental X-ray images based on a dilated convolutional segmentation network, characterized in that It includes the following steps: S1: Obtain an oral panoramic X-ray image dataset and perform preprocessing; S2: Annotate the oral panoramic X-ray images; S3: Augment the data of the oral panoramic X-ray images; S4: Construct and train a tooth segmentation model for oral panoramic X-ray images based on a dilated convolutional semantic segmentation network; S5: Test and validate the tooth segmentation model for oral panoramic X-ray images; S6: Convert the format of the tooth segmentation model for oral panoramic X-ray images and export it.

2. A method for segmenting teeth in oral panoramic X-ray images based on a dilated convolutional segmentation network according to claim 1, wherein Step S3 is specifically as follows: S31: Geometric transformation; S32: Gray value transformation; S33: Noise injection; S34: Distortion.