Two-stage multi-task oral cavity CBCT image segmentation method

By employing a two-stage, multi-task oral CBCT image segmentation method, combining nnUNet and W-UNETR networks, the feature extraction capability is enhanced, solving the problems of tooth segmentation accuracy and interference from complex structures. This method supports user interaction and achieves efficient and accurate tooth segmentation and quadrant classification.

CN120953304APending Publication Date: 2025-11-14CHINA UNIV OF MINING & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510991253.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing oral CBCT image segmentation methods struggle to accurately segment teeth, maxilla, mandible, maxillary sinus, and mandibular canal, especially under complex structures and metal artifact interference, resulting in insufficient accuracy and a lack of user interaction functions, leading to inaccurate and time-consuming segmentation results.

Method used

A two-stage multi-task oral CBCT image segmentation method is adopted. First, nnUNet is used for preliminary segmentation to remove metal artifacts. Then, the dual-output multi-task learning network W-UNETR is used for tooth segmentation and quadrant classification. The parallel spatial-channel attention module PSCA-SE is used to enhance feature extraction. The training is optimized by combining the DiceCELoss loss function and the cosine annealing learning rate strategy.

Benefits of technology

It improves the accuracy and precision of tooth segmentation, reduces interference from metal artifacts, supports user-interactive ROI selection, alleviates classification imbalance problems, and achieves efficient and accurate oral CBCT image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953304A_ABST
    Figure CN120953304A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage multi-task oral cavity CBCT image segmentation method, and belongs to the technical field of oral medical image segmentation, and the method comprises the steps: obtaining an in-oral cavity CBCT image, and carrying out the preprocessing; a medical image segmentation model nnUNet is adopted as a basic network to carry out first-stage multi-structure segmentation training on a CBCT image, and metal artifacts are removed; designing a dual-output multi-task learning network W-UNETR to carry out a second stage of tooth segmentation and quadrant classification; designing a parallel space-channel attention module PSCA-SE for the 3D medical image data for feature enhancement; constructing a multi-task loss function and a learning rate strategy until the loss function converges; and evaluating the performance of the model by adopting Dice similarity coefficients DSC and HD95. According to the method, through a two-stage multi-task segmentation strategy, the feature extraction capability is enhanced in combination with PSCA-SE, the segmentation precision is improved in cooperation with a DiceCELoss multi-task loss function, user interaction selection of ROI is supported to obtain a complete single tooth, and efficient and accurate oral cavity CBCT image segmentation is realized through DSC and HD95 accurate evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oral medical image segmentation technology, and in particular to a two-stage multi-task oral CBCT image segmentation method. Background Technology

[0002] Automatic segmentation of oral medical images has significant application value in dental image analysis. With the rapid development of medical imaging technology, especially the widespread application of cone-beam computed tomography (CBCT), the acquisition of dental imaging data has become increasingly common, and digital technology has gradually become the foundation of modern dental diagnosis and treatment. Automated oral segmentation methods can greatly improve the processing efficiency of dental medical images and reduce the workload of dentists. Traditional tooth segmentation methods rely heavily on manual intervention, especially when processing complex or noisy images, resulting in significant subjectivity and error. Deep learning-based automatic segmentation methods can reduce the influence of human factors. Simultaneously, automated segmentation technology supports personalized treatment for patients.

[0003] For a long time, manual segmentation has been the gold standard for dividing anatomical structures and pathological regions, but this process is time-consuming, labor-intensive, and often requires highly specialized knowledge. Semi-automatic or fully automatic segmentation methods can significantly reduce the required time and manpower, improve consistency, and enable the analysis of large-scale datasets. However, due to the complexity of the intraoral structure and the variability of tooth morphology, the task of automatically segmenting teeth, the maxilla, mandible, maxillary sinus, and mandibular canal remains extremely challenging, facing a series of technical difficulties. Specifically, the main challenges in tooth segmentation are as follows: First, the contrast between teeth and surrounding tissues is low, especially the indistinct boundary between the tooth root and alveolar bone, which poses significant difficulties for traditional segmentation methods. Second, there are significant individual differences in tooth morphology and position, especially between different patients, where occlusal relationships, arrangement, and growth patterns can vary significantly, requiring automated segmentation methods to have strong generalization capabilities. Furthermore, imaging technologies such as CBCT are susceptible to metal artifacts when acquiring oral structures, which affects image quality and consequently, the accuracy of subsequent segmentation. Therefore, existing segmentation methods struggle to achieve accurate segmentation results, and automatic segmentation of the oral cavity remains a major challenge in the field of medical imaging.

[0004] Therefore, deep learning technology, especially deep neural networks such as convolutional neural networks (CNNs), has become an important tool for addressing this problem. Deep learning can not only automatically learn features from data but also process large amounts of complex image data, opening up new possibilities for solving the tooth segmentation problem. However, these methods still face the following challenges: 1) Teeth are symmetrically distributed vertically and horizontally, and models can easily misclassify a certain type of tooth as its symmetrical position, introducing noise into the segmentation. 2) The accuracy of existing deep learning models cannot meet practical needs, so model improvement remains an unresolved problem. 3) Metal artifacts significantly interfere with the segmentation results. 4) The lack of platforms supporting user interaction and personalized services also makes oral diagnosis time-consuming and laborious for doctors. Summary of the Invention

[0005] The purpose of this invention is to provide a two-stage, multi-task oral CBCT image segmentation method that can finely segment the oral cavity, especially the teeth, giving it the advantages of high efficiency and accuracy while reducing the interference of metal artifacts and supporting user interaction functions.

[0006] To achieve the above objectives, the present invention provides a two-stage multi-task oral CBCT image segmentation method, comprising the following steps:

[0007] S1. Acquire cone-beam computed tomography (CBCT) images of the oral cavity and perform preprocessing;

[0008] S2. The medical image segmentation model nnUNet is used as the basic network to perform the first stage of multi-structure segmentation training on CBCT images to remove metal artifacts.

[0009] S3. Design a dual-output multi-task learning network W-UNETR to carry out the second stage of tooth segmentation and quadrant classification;

[0010] S4. A parallel spatial-channel attention module (PSCA-SE) for 3D medical image data is designed and added to the encoder side of the backbone network W-UNETR for feature enhancement.

[0011] S5. Construct a multi-task loss function and learning rate strategy to reduce the classification imbalance problem until the loss function converges to the set threshold.

[0012] S6. Establish an evaluation system and use the Dice similarity coefficient (DSC) and HD95 to evaluate the performance of the model.

[0013] Preferably, in S1, CBCT image preprocessing is divided into two stages, with the specific steps as follows:

[0014] S11, CBCT images are resampled with a voxel spacing of 0.3 mm, and the intensity in the range of [500, 3500] HU to [0.1, 1.1] is normalized;

[0015] S12. Resample the CBCT image again to a voxel spacing of 0.3 mm, normalize the intensity in the range of [500, 3500] HU to [0.1, 1.1], randomly crop four ROI images with a size of 128×96×96 and a foreground-to-background ratio of 3:1, and perform random intensity transformation and relabeling.

[0016] Preferably, S2 specifically includes the following steps:

[0017] S21, 3D CBCT data X∈ H×W×D Using nnUNet as the base network, the mandible, maxilla, dental nerves, maxillary sinus, and teeth were segmented, resulting in segmentation M1∈{0,1,2,3,4}. H×W×D ;

[0018] S22. Based on the voxel range of non-zero labels in M1, M2∈{1,2,3,4} H×W×D Element-wise multiplication with the original image X:

[0019] X f =X⊙M2 (1);

[0020] Obtain the image X after filtering the background. f ;

[0021] S23. Perform dental arch region cropping to obtain a ROI image containing only teeth:

[0022] X c =T γ (X f (2);

[0023] Where T γ (·) indicates that the coordinates are based on a preset anatomical coordinate system. Including the upper / lower tooth boundary parameters, an affine transformation is performed to obtain the input image X for the W-UNETR multi-task learning network adapted for the second-stage dual-output model. c ∈ H′×W′×D′ Where H'×W'×D' is the resolution of the input image, specifically:

[0024] X c =T γ (X⊙M2) (3);

[0025] S24. The user selects the ROI image that needs further cropping, and performs binary classification segmentation to obtain individual teeth, avoiding the model's inability to completely segment the same tooth.

[0026] Preferably, in S3, the dual-output multi-task learning network W-UNETR consists of two sub-networks: an encoder-decoder network for segmentation and a network for quadrant classification. W-UNETR is based on UNETR and uses the features extracted by the encoder to serve both tasks simultaneously, specifically including the following steps:

[0027] S31, X-ray the 3D CBCT data c ∈ B×C×H′×W′×D′ Input W-UNETR, the input image is divided into non-overlapping 3D blocks. The resolution is 128×96×96. Each block is flattened and mapped to a C-dimensional embedding vector through a learnable linear layer to form a sequence input, where C takes the value of 768.

[0028] S32. Add learnable 3D positional encoding to the embedding vector.

[0029] Preferably, the encoder consists of 12 stacked standard Transformer modules, each layer including multi-head self-attention (MSA), a feedforward network (FFN), layer normalization and residual connections, and multi-scale feature extraction. The MSA, with a value of 12, is used to capture diverse contextual features in parallel. The FFN is a two-layer fully connected network with ReLU activation. Layer normalization and residual connections are used before and after the MSA and FFN to stabilize training. Multi-scale feature extraction extracts intermediate features from layers 3, 6, 9, and 12, which are then reshaped into 3D feature maps. s i The downsampling rate is denoted as 1. The decoder consists of multiple upsampling modules that gradually restore spatial resolution. Each level receives the features of the corresponding encoder layer and the upsampling results of the previous level, and splices them together through skip connections. The convolutional block contains 3×3×3 convolution, instance normalization IN, and ReLU activation, which are repeated twice to refine the features.

[0030] Preferably, in S4, PSCA-SE consists of two parallel parts: shared multi-semantic space attention (SMSA) and progressive channel self-attention (PCSA-SE) with SE. The intermediate features extracted from layers 3, 6, 9, and 12 are simultaneously fed into both SMSA and PCSA-SE. SMSA decomposes the input features into multiple sub-features with different semantic levels and applies depthwise separable convolutions to capture multi-scale spatial priors. The PCSA-SE module mitigates semantic differences and preserves discriminative features through a channel self-attention mechanism. The specific process of feature enhancement is as follows:

[0031] S41. In the SMSA module, for a given input X∈ B×C×H×W×DDecompose along height, width, and depth, then perform global average pooling to obtain three 1D sequences: X H ∈ B×C×H X W ∈ B×C×W and X D ∈ B×C×D ;

[0032] S42. Divide the feature set obtained from each dimension into K sub-features of equal size. and The number of channels for each sub-feature is When K is 4, the decomposition process is as follows:

[0033]

[0034] X i It is the i-th sub-feature, i∈[1,K];

[0035] S43. Extracting multi-semantic spatial information, the principle of which is as follows:

[0036]

[0037] k represents the spatial structure information of the i-th sub-feature obtained after a lightweight convolution operation. i This represents the convolution kernel applied to the i-th sub-feature;

[0038] S44. Concatenate different semantic sub-features and normalize them using K-group group normalization (GN) to construct a spatial attention map. Use the Sigmoid activation function to generate spatial attention. The calculation process of the output features is as follows:

[0039]

[0040] SMSA(X) = X s =Attn H ×Attn W ×Attn D ×X (13);

[0041] σ(·) represents Sigmoid normalization, while and These represent GNs normalized to K groups along the H, W, and D dimensions, respectively.

[0042] S45. After global average pooling, the intermediate features of the input PCSA-SE are compressed by a factor of four. After passing through a ReLU layer, they are expanded back to C. Finally, the output is normalized by the Sigmoid function. The specific implementation is as follows:

[0043]

[0044] Among them, X attn This indicates the features processed in the first half of PCSA. This represents a pooling operation with a kernel size of K×K. This is a channel compression operation, where δ(·) is the ReLU activation function. It is a channel expansion;

[0045] S46. Multiply the outputs of SMSA and PCSA-SE to obtain the final processing result.

[0046] Preferably, in S5, the DiceCELoss function, which integrates the dice loss and cross-entropy loss, is used, and its formula is as follows:

[0047]

[0048] Where i is the number of voxels; j is the number of classes; Y i,j and G i,j Let w represent the probability output of class j at voxel i and the ground truth value of the one-hot encoding, respectively. t The coefficient representing the primary task of tooth segmentation is 0.7; w q The coefficient used to assist in task quadrant classification is set to 0.3. This is the total loss; the model parameters are continuously updated through gradient backpropagation until the loss function converges to the set threshold; the learning rate adopts a cosine annealing strategy with an initial value of 0.00001.

[0049] Preferably, in S6, the overlap between the DSC-measured volume segmentation prediction and the true voxel is defined as follows:

[0050]

[0051] Where Y and P represent the true value and output probability of all voxels, respectively;

[0052] HD95 effectively evaluates the boundary accuracy of segmentation results by calculating the 95th percentile value of the distance between the predicted segmentation boundary and the true segmentation boundary. Its mathematical definition is as follows:

[0053] HD 95 (Y,P)=max{d 95 (Y,P),d 95 (P,Y)} (18);

[0054] Where, d 95(Y,P) is the maximum 95th percentile distance between the true voxel and the predicted voxel, d 95 (P,Y) is the maximum 95th percentile distance between the predicted voxel and the true voxel.

[0055] Therefore, the present invention employs the above-mentioned two-stage multi-task oral CBCT image segmentation method, which has the following beneficial effects:

[0056] 1) Improve segmentation accuracy and precision: A two-stage segmentation strategy is adopted. In the first stage, five oral structures, including the mandible and maxilla, are segmented based on nnUNet. The background is filtered and the ROI region of the teeth is cropped, laying the foundation for the fine segmentation in the second stage. In the second stage, the dual-output multi-task network W-UNETR is used to complete tooth segmentation and quadrant classification at the same time. The spatial location information provided by quadrant classification is used to assist tooth segmentation, effectively solving the problem of misclassification of symmetrical tooth positions and improving the accuracy of single tooth segmentation.

[0057] 2) Enhanced feature representation and anti-interference ability: The designed parallel spatial-channel attention module PSCA-SE captures multi-scale spatial priors through shared multi-semantic spatial attention SMSA, and retains discriminative features by combining progressive channel self-attention PCSA-SE with SE, thereby enhancing the model's ability to extract features from complex oral structures and reducing the interference of metal artifacts on the segmentation results.

[0058] 3) Mitigating the class imbalance problem: The DiceCELoss multi-task loss function is adopted, and the training is optimized through the cosine annealing learning rate strategy, which effectively reduces the impact of class imbalance on the model and improves the stability and consistency of the segmentation results.

[0059] 4) Support for user interaction and personalization: Allows users to select ROI regions for further trimming, obtain individual teeth through binary classification segmentation, make up for the incomplete segmentation problem that may exist in the model, provide doctors with personalized services, and improve diagnostic efficiency.

[0060] 5) Accurately evaluate segmentation performance: An evaluation system is constructed based on the Dice similarity coefficient (DSC) and HD95 index to quantify the segmentation effect from two aspects: voxel overlap and boundary accuracy, ensuring the reliability and clinical applicability of the model output.

[0061] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0062] Figure 1 This is an overall flowchart of an embodiment of a two-stage multi-task oral CBCT image segmentation method of the present invention;

[0063] Figure 2This is a schematic diagram of metal artifact processing in an embodiment of a two-stage multi-task oral CBCT image segmentation method of the present invention;

[0064] Figure 3 This is a W-UNETR model structure diagram of an embodiment of a two-stage multi-task oral CBCT image segmentation method of the present invention;

[0065] Figure 4 This is a structural diagram of the PSCA-SE module in an embodiment of a two-stage multi-task oral CBCT image segmentation method of the present invention;

[0066] Figure 5 This is a plug-in interface diagram of an embodiment of a two-stage multi-task oral CBCT image segmentation method of the present invention;

[0067] Figure 6 This is a schematic diagram of the operation flow of an embodiment of a two-stage multi-task oral CBCT image segmentation method of the present invention. Detailed Implementation

[0068] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0069] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0070] Example 1

[0071] This invention provides a two-stage, multi-task oral CBCT image segmentation method, the overall process of which is as follows: Figure 1 As shown, the operation process is as follows: Figure 6 As shown, the developed plugin interface is as follows: Figure 5As shown. The data in this embodiment includes 446 CBCT images, which are used for training, validation, and testing in a 7:1:2 split ratio. In the second stage, data augmentation is performed on both tasks simultaneously. Specifically, this includes: image resampling with a voxel spacing of 0.3 mm, normalization of intensity in the range of [500, 3500] HU to [0.1, 1.1], random cropping of four ROI images with a foreground-to-background ratio of 3:1 and a size of 128×96×96, and random intensity transformation, with a batch size of 4.

[0072] The tooth segmentation task takes single-channel data as input and outputs 35 categories (background, alveolar bone, implants, and 32 teeth). The quadrant classification task shares an encoder with the former and outputs segmentation results for 7 categories (background, alveolar bone, implants, and teeth in 4 quadrants).

[0073] Specifically, the following steps are included:

[0074] S1. Acquire intraoral cone-beam computed tomography (CBCT) images and perform preprocessing; CBCT image preprocessing consists of two stages, specifically including the following steps:

[0075] S11, CBCT images are resampled with a voxel spacing of 0.3 mm, and the intensity in the range of [500, 3500] HU to [0.1, 1.1] is normalized;

[0076] S12. Resample the CBCT image again to a voxel spacing of 0.3 mm, normalize the intensity in the range of [500, 3500] HU to [0.1, 1.1], randomly crop four ROI images with a size of 128×96×96 and a foreground-to-background ratio of 3:1, and perform random intensity transformation and relabeling.

[0077] S2. The nnUNet medical image segmentation model is used as the base network to perform the first stage of multi-structure segmentation training on CBCT images to remove metal artifacts. The metal artifact processing is illustrated in the diagram below. Figure 2As shown, nnUNet, as a benchmark framework in the field of medical image segmentation, has its core advantages in the organic combination of standardized processing procedures and adaptive network architecture. This framework innovatively establishes a fully automated process including data preprocessing, network configuration, and post-processing: 1) It autonomously determines the optimal interpolation strategy and normalization scheme by intelligently analyzing the spatial features of the input data, such as voxel spacing and image size; 2) It dynamically adjusts the network architecture parameters, including key hyperparameters such as network depth and initial number of channels, based on the statistical characteristics of the dataset; 3) It adopts a cascaded training strategy, firstly locating the target region on the low-resolution image, and secondly performing fine segmentation on the cropped high-resolution ROI region. In this embodiment, nnUNet is used as the base network for multi-structure segmentation training. After obtaining accurate segmentation results of the mandible, maxilla, dental nerve, maxillary sinus, and teeth through an automated processing flow, the network performs element-wise multiplication and cropping on the original image based on the spatial location information of the dental arch region to generate ROI images containing only teeth, providing a data foundation for the fine segmentation in the second stage. Finally, the plugin allows users to select the ROI to be cropped and perform binary classification segmentation to obtain a single tooth, making up for the model's inability to completely segment the same tooth.

[0078] The specific steps are as follows:

[0079] S21, 3D CBCT data X∈ H×W×D Using nnUNet as the base network, the mandible, maxilla, dental nerves, maxillary sinus, and teeth were segmented, resulting in segmentation M1∈{0,1,2,3,4}. H×W×D ;

[0080] S22. Based on the voxel range of non-zero labels in M1, M2∈{1,2,3,4} H×W×D Element-wise multiplication with the original image X:

[0081] X f =X⊙M2 (1);

[0082] Obtain the image X after filtering the background. f ;

[0083] S23. Perform dental arch region clipping. Dental arch region clipping also reduces the computational burden of segmenting individual teeth in the second stage, obtaining a ROI image containing only the teeth:

[0084] X c =T γ (X f (2);

[0085] Where T γ (·) indicates that the coordinates are based on a preset anatomical coordinate system. Including the upper / lower tooth boundary parameters, an affine transformation is performed to obtain the input image X for the W-UNETR multi-task learning network adapted for the second-stage dual-output model. c ∈ H′×W′×D′ Where H'×W'×D' is the resolution of the input image, specifically:

[0086] X c =T γ (X⊙M2) (3);

[0087] S24. The user selects the ROI image that needs further cropping, and performs binary classification segmentation to obtain individual teeth, avoiding the model's inability to completely segment the same tooth.

[0088] S3. Design the dual-output multi-task learning network W-UNETR to perform the second stage of tooth segmentation and quadrant classification. The W-UNETR model structure is as follows: Figure 3 As shown, the dual-output multi-task learning network W-UNETR consists of two sub-networks: an encoder-decoder network for segmentation and a network for quadrant classification. W-UNETR is based on UNETR, simultaneously applying the features extracted by the encoder to both tasks. UNETR is a deep learning architecture specifically designed for 3D medical image segmentation. Compared to traditional convolutional neural network-based segmentation methods, UNETR significantly improves long-range dependency capture capabilities by fusing the global context modeling capabilities of Transformer with the symmetric encoder-decoder structure of U-Net. Specifically, it includes the following steps:

[0089] S31, X-ray the 3D CBCT data c ∈ B×C×H′×W′×D′ Input W-UNETR, the input image is divided into non-overlapping 3D blocks. The resolution is 128×96×96. Each block is flattened and mapped to a C-dimensional embedding vector through a learnable linear layer to form a sequence input, where C takes the value of 768.

[0090] S32. Add learnable 3D positional encoding to the embedding vector.

[0091] The encoder consists of 12 stacked standard Transformer modules. Each layer includes Multi-Head Self-Attention (MSA), a Feedforward Network (FFN), layer normalization and residual connections, and multi-scale feature extraction. The MSA, with a value of 12, is used to capture diverse contextual features in parallel. The FFN is a two-layer fully connected network with ReLU activation. Layer normalization and residual connections are used before and after the MSA and FFN to stabilize training. Multi-scale feature extraction extracts intermediate features from layers 3, 6, 9, and 12, which are then reconstructed into 3D feature maps. s iThe downsampling rate is denoted as 1. The decoder consists of multiple upsampling modules that gradually restore spatial resolution. Each level receives the features of the corresponding encoder layer and the upsampling results of the previous level, and splices them together through skip connections. The convolutional block contains 3×3×3 convolution, instance normalization IN, and ReLU activation, which are repeated twice to refine the features.

[0092] Multi-task learning has achieved success in many areas of machine learning. Its principle is to improve the generalization ability of a model by utilizing domain-specific information contained in the training signals of related tasks. In this embodiment, the dual-output multi-task learning network W-UNETR simultaneously performs tooth segmentation and quadrant classification tasks. The two tasks share a single encoder, but use different decoders for upsampling. The auxiliary task is used to predict the quadrant in which the tooth is located, providing rich spatial location information for tooth segmentation.

[0093] S4. A parallel spatial-channel attention module (PSCA-SE) with SE is designed for 3D medical image data and added to the encoder side of the backbone network W-UNETR for feature enhancement; the structure of the PSCA-SE module is as follows: Figure 4 As shown, the system consists of two parallel parts: Shared Multi-Semantic Spatial Attention (SMSA) and Progressive Channel Self-Attention (PCSA-SE) with Semantic Context (SE). The intermediate features extracted from layers 3, 6, 9, and 12 are simultaneously fed into both SMSA and PCSA-SE. SMSA decomposes the input features into multiple sub-features with different semantic levels and applies depthwise separable convolutions to capture multi-scale spatial priors. The PCSA-SE module mitigates semantic differences and preserves discriminative features through a channel self-attention mechanism. The specific process of feature enhancement is as follows:

[0094] S41. In the SMSA module, for a given input X∈ B×C×H×W×D Decompose along height, width, and depth, then perform global average pooling to obtain three 1D sequences: X H ∈ B×C×H X W ∈ B×C×W and X D ∈ B×C×D ;

[0095] S42. Divide the feature set obtained from each dimension into K sub-features of equal size. and The number of channels for each sub-feature is When K is 4, the decomposition process is as follows:

[0096]

[0097] X iIt is the i-th sub-feature, i∈[1,K];

[0098] S43. To capture different semantic information within each sub-feature, SMSA employs one-dimensional convolutions with kernel sizes of 3, 5, 7, and 9, and applies lightweight shared convolutions to enable the network to learn consistent features across the three dimensions to implicitly model their dependencies and extract multi-semantic spatial information. The principle is as follows:

[0099]

[0100] k represents the spatial structure information of the i-th sub-feature obtained after a lightweight convolution operation. i This represents the convolution kernel applied to the i-th sub-feature;

[0101] S44. Different semantic sub-features are concatenated and normalized using K-group group normalization (GN) to construct a spatial attention map. GN performs better in terms of semantic differences between sub-features and does not introduce batch statistical noise. The spatial attention is generated using the Sigmoid activation function. The calculation process of the output features is as follows:

[0102]

[0103] SMSA(X) = X s =Attn H ×Attn W ×Attn D ×X (13);

[0104] σ(·) represents Sigmoid normalization, while and These represent GNs normalized to K groups along the H, W, and D dimensions, respectively.

[0105] S45, a lightweight SE module is introduced at the tail of PCSA to form PCSA-SE to reduce computational cost. PCSA-SE adopts a progressive compression method based on average pooling, which has stronger input sensing ability. The intermediate features input to PCSA-SE are compressed by a factor of four after global average pooling. After passing through a ReLU layer, they are expanded back to C, and finally normalized by the Sigmoid function. The specific implementation is as follows:

[0106]

[0107] Among them, X attn This indicates the features processed in the first half of PCSA. This represents a pooling operation with a kernel size of K×K. This is a channel compression operation, where δ(·) is the ReLU activation function. It is a channel expansion;

[0108] S46. Multiply the outputs of SMSA and PCSA-SE to obtain the final processing result. Compared with SCSA, the parallel spatial-channel attention mechanism PSCA-SE proposed in this embodiment can achieve parallel acceleration, enhance the robustness of the model, and the parallel structure allows spatial and channel information to interact in the shallow network, avoiding the problem that high-level features may lose low-level details in the serial structure. In addition, the SE module is introduced into the tail of the progressive channel self-attention PCSA, which increases nonlinearity while reducing the amount of computation.

[0109] S5. Construct a multi-task loss function and learning rate strategy to reduce classification imbalance until the loss function converges to a set threshold; use the DiceCELoss function, which integrates the dice loss and cross-entropy loss, with the following formula:

[0110]

[0111] Where i is the number of voxels; j is the number of classes; Y i,j and G i,j Let w represent the probability output of class j at voxel i and the ground truth value of the one-hot encoding, respectively. t The coefficient representing the primary task of tooth segmentation is 0.7; w q The coefficient used to assist in task quadrant classification is set to 0.3. This is the total loss; the model parameters are continuously updated through gradient backpropagation until the loss function converges to the set threshold; the learning rate adopts a cosine annealing strategy with an initial value of 0.00001.

[0112] S6. Establish an evaluation system and use the Dice similarity coefficient (DSC) and HD95 to evaluate the performance of the model.

[0113] The overlap between the DSC-measured volume segmentation prediction and the true voxel is defined as follows:

[0114]

[0115] Where Y and P represent the true value and output probability of all voxels, respectively;

[0116] HD95 effectively evaluates the boundary accuracy of segmentation results by calculating the 95th percentile value of the distance between the predicted segmentation boundary and the true segmentation boundary. Its mathematical definition is as follows:

[0117] HD 95(Y,P)=max{d 95 (Y,P),d 95 (P,Y)} (18);

[0118] Where, d 95 (Y,P) is the maximum 95th percentile distance between the true voxel and the predicted voxel, d 95 (P,Y) is the maximum 95th percentile distance between the predicted voxel and the true voxel.

[0119] Therefore, this invention employs a two-stage, multi-task oral CBCT image segmentation method. Through this two-stage multi-task segmentation strategy, the first stage utilizes nnUNet to segment the oral cavity into five parts and remove metal artifacts. The second stage uses a dual-output network W-UNETR to simultaneously achieve tooth segmentation and quadrant classification. Combined with the parallel spatial-channel attention module PSCA-SE to enhance feature extraction capabilities, and the DiceCELoss multi-task loss function and cosine annealing learning rate strategy to alleviate classification imbalance, this method not only improves segmentation accuracy and robustness to complex structures but also supports user-interactive selection of ROIs to obtain complete single teeth. Finally, through precise evaluation using DSC and HD95, efficient and accurate oral CBCT image segmentation is achieved, providing strong support for oral medicine diagnosis.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A two-stage, multi-task oral CBCT image segmentation method, characterized in that, Includes the following steps: S1. Acquire cone-beam computed tomography (CBCT) images of the oral cavity and perform preprocessing; S2. The medical image segmentation model nnUNet is used as the basic network to perform the first stage of multi-structure segmentation training on CBCT images to remove metal artifacts. S3. Design a dual-output multi-task learning network W-UNETR to carry out the second stage of tooth segmentation and quadrant classification; S4. A parallel spatial-channel attention module (PSCA-SE) for 3D medical image data is designed and added to the encoder side of the backbone network W-UNETR for feature enhancement. S5. Construct a multi-task loss function and learning rate strategy to reduce the classification imbalance problem until the loss function converges to the set threshold. S6. Establish an evaluation system and use the Dice similarity coefficient (DSC) and HD95 to evaluate the performance of the model.

2. The two-stage multi-task oral CBCT image segmentation method according to claim 1, characterized in that: In S1, CBCT image preprocessing is divided into two stages, with the specific steps as follows: S11, CBCT images are resampled with a voxel spacing of 0.3 mm, and the intensity in the range of [500, 3500] HU to [0.1, 1.1] is normalized; S12. Resample the CBCT image again to a voxel spacing of 0.3 mm, normalize the intensity in the range of [500, 3500] HU to [0.1, 1.1], randomly crop four ROI images with a size of 128×96×96 and a foreground-to-background ratio of 3:1, and perform random intensity transformation and relabeling.

3. The two-stage multi-task oral CBCT image segmentation method according to claim 2, characterized in that: S2 specifically includes the following steps: S21, 3D CBCT data X∈ H×W×D Using nnUNet as the base network, the mandible, maxilla, dental nerves, maxillary sinus, and teeth were segmented, resulting in segmentation M1∈{0,1,2,3,4}. H×W×D ; S22. Based on the voxel range of non-zero labels in M1, M2∈{1,2,3,4} H×W×D Element-wise multiplication with the original image X: X f =X⊙M2 (1); Obtain the image X after filtering the background. f ; S23. Perform dental arch region cropping to obtain a ROI image containing only teeth: X c =T γ (X f ) (2); Where T γ (·) indicates that the coordinates are based on a preset anatomical coordinate system. Including the upper / lower tooth boundary parameters, an affine transformation is performed to obtain the input image X for the W-UNETR multi-task learning network adapted for the second-stage dual-output model. c ∈ H’×W’×D’ Where H'×W'×D' is the resolution of the input image, specifically: X c =T γ (X⊙M2) (3); S24. The user selects the ROI image that needs further cropping, and performs binary classification segmentation to obtain individual teeth, avoiding the model's inability to completely segment the same tooth.

4. The two-stage multi-task oral CBCT image segmentation method according to claim 3, characterized in that: In S3, the dual-output multi-task learning network W-UNETR consists of two sub-networks: an encoder-decoder network for segmentation and a network for quadrant classification. W-UNETR is based on UNETR and uses the features extracted by the encoder to serve both tasks simultaneously. Specifically, it includes the following steps: S31, X-ray the 3D CBCT data c ∈ B×C×H’×W’×D’ Input W-UNETR, the input image is divided into non-overlapping 3D blocks. The resolution is 128×96×96. Each block is flattened and mapped to a C-dimensional embedding vector through a learnable linear layer to form a sequence input, where C takes the value of 768. S32. Add learnable 3D positional encoding to the embedding vector.

5. The two-stage multi-task oral CBCT image segmentation method according to claim 4, characterized in that: The encoder consists of 12 stacked standard Transformer modules. Each layer includes Multi-Head Self-Attention (MSA), a Feedforward Network (FFN), layer normalization and residual connections, and multi-scale feature extraction. The MSA, with a value of 12, is used to capture diverse contextual features in parallel. The FFN is a two-layer fully connected network with ReLU activation. Layer normalization and residual connections are used before and after the MSA and FFN to stabilize training. Multi-scale feature extraction extracts intermediate features from layers 3, 6, 9, and 12, which are then reconstructed into 3D feature maps. s i The downsampling rate is denoted as 1. The decoder consists of multiple upsampling modules that gradually restore spatial resolution. Each level receives the features of the corresponding encoder layer and the upsampling results of the previous level, and splices them together through skip connections. The convolutional block contains 3×3×3 convolution, instance normalization IN, and ReLU activation, which are repeated twice to refine the features.

6. The two-stage multi-task oral CBCT image segmentation method according to claim 5, characterized in that: In S4, PSCA-SE is composed of two parallel parts: shared multi-semantic space attention (SMSA) and progressive channel self-attention (PCSA-SE) with SE. The intermediate features extracted from layers 3, 6, 9, and 12 are simultaneously fed into both SMSA and PCSA-SE. SMSA decomposes the input features into multiple sub-features with different semantic levels and applies depthwise separable convolution to capture multi-scale spatial priors. The PCSA-SE module alleviates semantic differences and preserves discriminative features through a channel self-attention mechanism. The specific process of feature enhancement is as follows: S41. In the SMSA module, for a given input X∈ B×C×H×W×D Decompose along height, width, and depth, then perform global average pooling to obtain three 1D sequences: X H ∈ B×C×H X W ∈ B×C×W and X D ∈ B×C×D ; S42. Divide the feature set obtained from each dimension into K sub-features of equal size. and The number of channels for each sub-feature is When K is 4, the decomposition process is as follows: X i It is the i-th sub-feature, i∈[1,K]; S43. Extracting multi-semantic spatial information, the principle of which is as follows: k represents the spatial structure information of the i-th sub-feature obtained after a lightweight convolution operation. i This represents the convolution kernel applied to the i-th sub-feature; S44. Concatenate different semantic sub-features and normalize them using K-group group normalization (GN) to construct a spatial attention map. Use the Sigmoid activation function to generate spatial attention. The calculation process of the output features is as follows: SMSA(X)=X s =Attn H ×Attn W ×Attn D ×X (13); σ(·) represents Sigmoid normalization, while and These represent GNs normalized to K groups along the H, W, and D dimensions, respectively. S45. After global average pooling, the intermediate features of the input PCSA-SE are compressed by a factor of four. After passing through a ReLU layer, they are expanded back to C. Finally, the output is normalized by the Sigmoid function. The specific implementation is as follows: Among them, X attn This indicates the features processed in the first half of PCSA. This represents a pooling operation with a kernel size of K×K. This is a channel compression operation, where δ(·) is the ReLU activation function. It is a channel expansion; S46. Multiply the outputs of SMSA and PCSA-SE to obtain the final processing result.

7. The two-stage multi-task oral CBCT image segmentation method according to claim 6, characterized in that: In S5, the DiceCELoss function, which integrates the dice loss and cross-entropy loss, is used. The formula is as follows: Where i is the number of voxels; j is the number of classes; Y i,j and G i,j Let w represent the probability output of class j at voxel i and the ground truth value of the one-hot encoding, respectively. t The coefficient representing the primary task of tooth segmentation is 0.7; w q The coefficient used to assist in task quadrant classification is set to 0.

3. This is the total loss; the model parameters are continuously updated through gradient backpropagation until the loss function converges to the set threshold; the learning rate adopts a cosine annealing strategy with an initial value of 0.00001.

8. The two-stage multi-task oral CBCT image segmentation method according to claim 7, characterized in that: In S6, the overlap between the DSC-measured volume segmentation prediction and the true voxel is defined as follows: Where Y and P represent the true value and output probability of all voxels, respectively; HD95 effectively evaluates the boundary accuracy of segmentation results by calculating the 95th percentile value of the distance between the predicted segmentation boundary and the true segmentation boundary. Its mathematical definition is as follows: HD 95 (Y,P)=max{d 95 (Y,P),d 95 (P,Y)} (18); Where, d 95 (Y,P) is the maximum 95th percentile distance between the true voxel and the predicted voxel, d 95 (P,Y) is the maximum 95th percentile distance between the predicted voxel and the true voxel.

Citation Information

Patent Citations

  • Oral cavity CBCT image segmentation method based on superpixel statistical characteristic and graph attention network

    CN113470045A

  • Oral cavity CBCT image tooth and soft tissue segmentation model method based on improved U-Net model

    CN117115132A

  • Tooth root angle three-dimensional intelligent measurement method based on oral cavity CBCT image

    CN119887630A

  • Method, device, storage medium, and medical system for generating a restoration dental model

    US20250143852A1