A method for segmentation of impaction teeth and neural tube based on media region
By combining deep learning models and medium probability maps, and introducing a perceptual attention mechanism and curvature-aware edge supervision, the accuracy and robustness issues of segmenting impacted teeth and neural tubes in dental CT images were solved, achieving precise semantic segmentation of the medium region.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NINGBO DENTAL HOSPITAL CO LTD
- Filing Date
- 2025-06-24
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to accurately segment impacted teeth and nerve canals in dental CT images, especially when the medial region is blurred or the boundaries are discontinuous, posing a challenge to the segmentation task.
We construct a deep learning model that combines channel attention and spatial attention with a medium probability graph and a perceptual attention mechanism. We also introduce a curvature-aware edge supervision mechanism to optimize the curvature differences and semantic feature separation of the medium region, thereby improving the accuracy and robustness of the segmentation model.
It significantly improves the segmentation accuracy and stability of impacted teeth and nerve canals in dental CBCT images, and solves the problem of accurate semantic segmentation of regions with blurred media and discontinuous boundaries.
Smart Images

Figure CN120411524B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a method for segmenting impacted teeth and neural tubes based on a media region. Background Technology
[0002] In medical imaging, particularly in dental cone-beam CT (CBCT) images or craniomaxillary CT images, anatomical structures are often not demarcated by clear, sharp edges, but rather by transitional regions of soft tissue type or different grayscale characteristics. These transitional regions typically appear on clinical imaging as being between the target structure (such as teeth, neural tubes, and bone tissue) and the background tissue, and possess the following notable characteristics:
[0003] 1) Gray-level blurring transition characteristics: Compared to the target structure (such as teeth, nerve canals, and bone tissue), the gray-level value of the medium is usually between that of the teeth and the nerve canals, forming an intermediate transition region connecting the teeth and the nerve canals, exhibiting a transitional gray-level distribution. This blurring makes it difficult for traditional image segmentation methods based on gray-level or edge gradients to accurately delineate target boundaries, especially in multi-tissue critical regions (such as between the root apex and the mandibular nerve canal);
[0004] 2) The distribution pattern is complex and variable, and the morphology of the media region is greatly affected by individual anatomical differences and pathological conditions. For example, in areas with tightly packed dentition or in the presence of clinical phenomena such as overlapping teeth, impacted teeth, dental cysts, or bone resorption, the spatial morphology of the media may be abnormally compressed, locally fused, or fractured, resulting in a non-standardized and irregular spatial distribution. This complexity places higher demands on the generalization ability and boundary adaptation ability of image segmentation models;
[0005] 3) Spatial non-uniformity and boundary asymmetry: In images, the medium is often not uniformly or symmetrically distributed. For example, there are significant differences in the sharpness and grayscale transition speed of the boundaries on both sides of the mandible, the upper and lower edges of the alveolar bone, or the palatal region. Some areas may have extremely sharp edges, while others may exhibit diffused and blurred boundaries, resulting in high heterogeneity of the boundary features of the same anatomical structure in different regions. This asymmetric blurring characteristic greatly weakens the adaptability of traditional image processing algorithms;
[0006] 4) Increased interference with structural boundaries: Since the medium region usually surrounds the target structure, its signals or features may be highly coupled with the target structural region during image convolution processing, causing model misjudgment or structural "fusion" phenomena, especially in semantically similar and gray-level transition zones such as the pulp cavity, periodontal ligament, and bone. This poses a challenge to the edge separation ability and semantic feature decoupling ability of deep segmentation networks.
[0007] In clinical settings such as oral and maxillofacial surgery planning, implant design, and orthodontic treatment, accurate identification and segmentation of impacted teeth and the inferior alveolar nerve canal are crucial. However, due to the abnormal location and complex growth direction of impacted teeth, they are often adjacent to or even overlap with the nerve canal pathway, making the segmentation task extremely challenging.
[0008] Existing segmentation methods mainly include:
[0009] 1. Image processing methods based on thresholding or edge detection perform binarization or edge enhancement based on differences in image grayscale values, such as using the Otsu method, adaptive thresholding, and Canny edge detection. These methods are significantly affected by boundary blurring and media transitions. They are suitable for images with strong structural contrast and clear boundaries, but for cases where the grayscale contrast between the neural tube and alveolar bone is not obvious, and because the neural tube and impacted teeth are often separated by soft tissue media, this area exhibits grayscale blurring and unclear boundaries. This method struggles to accurately distinguish the transition area from the true structural boundary, easily leading to "structural fusion" or "boundary drift," resulting in missed detections or unclear boundaries.
[0010] 2. Methods based on morphological models or statistical priors. These methods utilize the elongated tubular features of the neural tube in CT scans, or the near-ellipsoidal geometry of teeth, to establish morphological constraint models for fitting. Examples include active contour models, graph-cut algorithms, or random walk algorithms. These methods are typically highly dependent on image quality and anatomical consistency, have poor adaptability to ectopic or deformed teeth and neural structures, and are often insufficiently robust to impacted tooth regions with significant anatomical variations and dental deformities. Impacted teeth often exhibit complex conditions such as horizontal embedment, inclined growth, and cross-obstruction, leading to significant deviations from standard morphology, thus limiting the adaptability of this method.
[0011] 3. Traditional machine learning methods, such as SVM and random forest, utilize features like local texture and grayscale statistics to identify target regions through classifiers. These methods typically require extensive manual feature design and are difficult to generalize. They lack the ability to model semantic dependencies between structures, perform poorly in highly deformable structures or areas with strong media interference, and cannot effectively model the spatial, semantic, and structural dependencies between teeth and nerve canals. They are also more prone to misjudgment, especially when there is spatial contact or compression.
[0012] 4. Shallow Convolutional Neural Network (CNN) Methods. While shallower U-Net or 2D CNN models can be used to segment impacted teeth or neural tubes, initially addressing some structural ambiguity, they lack the ability to model three-dimensional spatial relationships and complex boundaries. Furthermore, they fail to consider media interference and structural semantic context, exhibiting weak anti-interference capabilities and sensitivity to artifacts and low-quality images. Medical images are often accompanied by motion artifacts, image noise, and reconstruction artifacts; this method lacks robustness strategies and is prone to failure on low-quality images.
[0013] In summary, while existing methods for segmenting impacted teeth and nerve tubes have a certain application basis, they still have significant limitations when faced with complex anatomical structures, interference from ambiguous media, and non-standard morphological structures.
[0014] Therefore, those skilled in the art are dedicated to developing a method for segmenting impacted teeth and nerve tubes based on a media region. Summary of the Invention
[0015] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is how to achieve accurate semantic segmentation of regions with structural overlap, media blurring and boundary discontinuity in dental CT images.
[0016] In 3D oral medical imaging, the positional relationship between impacted teeth (especially wisdom teeth) and the mandibular nerve canal is a key factor in surgical planning. This application constructs a segmentation model for impacted teeth and the nerve canal based on deep learning. The transition region between the impacted tooth and the nerve canal is defined as the media region. A media probability map is established, and a perceptual strategy for the media region is constructed. The structure of the blurred transition region between the impacted tooth and the nerve canal, such as adipose tissue and bone tissue, is modeled. During the segmentation process, interference caused by fewer boundary judgments is actively identified, improving the ability to distinguish adjacent structural boundaries. Then, a perceptual attention mechanism is introduced, combining channel attention and spatial attention to jointly guide feature responses. When dealing with complex structural compression, multiple adjacent teeth, or overlapping root canals, it can better focus on the structural body rather than the background or mixed regions, enhancing the model's expressive power and local recognition ability. Based on a curvature-aware edge supervision mechanism, edge curvature information is introduced to construct a refined supervision term, enhancing the perception and learning ability of edge contour changes and improving the performance of sharp tooth junctions and curved edge details.
[0017] In one embodiment of the present invention, a method for separating impacted teeth and nerve tubes based on a media region is provided, comprising the following steps:
[0018] S1000. Preparation: Collect CBCT training set, and initially construct a segmentation model of impacted teeth and neural tube based on deep learning. Initialize the number of training rounds to 0.
[0019] S2000, Model Structure Optimization: Input the CBCT training set into the impacted tooth and nerve tube segmentation model, divide the media region into high media region and low media region, optimize the curvature difference of the high media region, enhance the semantic feature separation of the low media region, predict the contour of the media region and perform geometric correction, and construct the overall loss function and the impacted tooth and nerve tube segmentation loss function.
[0020] S3000: Model training. The AdamW optimizer is used for training. The number of training epochs is incremented by one. When the number of training epochs is less than the maximum number of training epochs and the performance threshold is met, the training ends, and the trained impacted tooth and neural tube segmentation model is obtained. Step S5000 is then executed. When the number of training epochs is less than the maximum number of training epochs and the performance threshold is not met, the process returns to step S2000. When the number of training epochs is greater than or equal to the maximum number of training epochs, step S4000 is executed.
[0021] S4000: The temporary model continues to be trained. When the number of training rounds is greater than or equal to the maximum number of training rounds, a temporary impacted tooth and neural tube segmentation model is obtained. The weights of each loss function in the overall loss function are modified according to the loss function adjustment strategy. The temporary impacted tooth and neural tube segmentation model is used to continue training until the performance threshold is met. Training ends, and a well-trained impacted tooth and neural tube segmentation model is obtained.
[0022] S5000, Impacted Tooth and Neural Tube Segmentation Prediction: Using a trained impacted tooth and neural tube segmentation model, inference and segmentation prediction are performed on CBCT, and the output image includes the mask of the impacted tooth region and the mask of the neural tube region.
[0023] Optionally, in the method for segmenting impacted teeth and neural tubes based on the medium region in the above embodiments, the impacted tooth and neural tube segmentation model adopts a general deep learning architecture.
[0024] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2000 includes:
[0025] S2100, Media Region Division: Input the CBCT training set into the impacted tooth and neural tube segmentation model, and divide the media region into high media region and low media region according to the media filling degree characteristics. Construct a gradient regularization loss function to generate a media probability map.
[0026] S2200, high medium region curvature difference optimization, define curvature edge supervision auxiliary loss function, calculate curvature map, define curvature edge map, fuse curvature edge map with medium probability map, perform boundary enhancement, and generate enhanced boundary prediction map;
[0027] S2300, low-medium region semantic feature separation enhancement, introduces a perceptual attention guidance mechanism in the low-medium region to extract semantic features, performs global average pooling to extract channel descriptions of semantic features, calculates channel weighted feature maps, performs channel compression on semantic features, calculates spatially enhanced semantic feature maps, and obtains perceptual features by connecting the channel weighted feature maps and spatially enhanced semantic feature maps through residuals, and constructs a perceptual supervised loss function;
[0028] S2400, Medium Region Contour Prediction: The enhanced boundary prediction map, channel-weighted feature map, and spatially enhanced semantic feature map are used to extract the boundary to obtain the predicted contour of the medium region.
[0029] S2500, geometric correction, uses CML (Connecting Middle Line) as the segmentation guide line to perform geometric correction on the predicted contour of the medium region, defines the geometric correction loss function, and constructs the overall loss function and the segmentation loss function for impacted teeth and nerve canals.
[0030] Furthermore, in the method for segmenting impacted teeth and nerve tubes based on the medium region in the above embodiments, the medium filling characteristics include gray-level gradient and structural similarity.
[0031] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the media region in the above embodiments, step S2100 includes:
[0032] S2110. Input image processing: Input the CBCT training set. CBCT is three-dimensional volume data. Two-dimensional cross-sectional images are extracted through sagittal and coronal slices. , means as follows:
[0033] ;
[0034] in, R represents The set of real numbers is the set of all real numbers. H Indicates the height of the input cross-sectional image. W Indicates the width of the input cross-sectional image;
[0035] S2120, intermediate structure construction, using intermediate structure features to describe the geometric and topological probability characteristics of the medium region, defining medium filling degree features, and constructing a filling scoring function;
[0036] S2130, Dielectric region division, defining high dielectric region and low dielectric region;
[0037] S2140. Construct a gradient regularization loss function, and set supervision labels for the medium region, as shown in the following formula:
[0038] ;
[0039] Since tissue structures exhibit spatial continuity in images, a gradient regularization loss function is constructed to increase the spatial smoothness of the media region and reduce excessive fragmentation, as shown in the following formula:
[0040] ;
[0041] S2150. Medium probability map generation: Perform classification prediction for each pixel, generate and output the medium probability map. :
[0042] ;
[0043] in, Represents pixels The probability of being a medium.
[0044] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the media region in the above embodiments, the intermediate structural features include:
[0045] CML (Connecting Middle Line) is the central axis connecting the impacted tooth and the neural canal, defining the central path of the medial region.
[0046] MBO (Middle Boundary Offset) is the boundary offset distance from CML outward along the normal projection direction, used to reconstruct the rough geometric profile of the medium region.
[0047] CMO (Central Motion Orientation) is the orientation (unit vector) of CML in three-dimensional space, which enhances the spatial consistency of region modeling by maintaining its continuity.
[0048] MCC (Middle Connection Confidence) is the probability that the current pixel is in the medium region. It is output by regression through the Sigmoid function and is used to determine the existence of the blurred region.
[0049] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the media region in the above embodiments, step S2120 includes:
[0050] S2121, Intermediate Structure Feature Extraction: This section describes the geometric and topological probabilistic characteristics of the medium region using intermediate structure features. It defines the intermediate structure feature loss function, including:
[0051] CML Loss Function The definition is as follows:
[0052] ;
[0053] in, Indicates the predicted coordinates of the central axis point. Represents the true coordinates of the central axis point, indicating its position. The coordinates of points on the real CML in the cross-sectional image are used to supervise the model to learn and generate the correct central axis position; the points on the CML refer to the sampling points of the neural tube or impacted tooth on the skeleton line in the cross-sectional image, which are simplified representations of the morphology and orientation in a local area;
[0054] MBO loss function The definition is as follows:
[0055] ;
[0056] in, This represents the predicted boundary offset. This represents the distance from the CML to the true boundary along the central axis at the current point, i.e., the offset, which is used to constrain the width or thickness of the region generated by the neural network model.
[0057] CMO Loss Function The definition is as follows:
[0058] ;
[0059] in, Represents the predicted direction vector. The true vector representing the direction of the central axis is used to constrain the predicted direction field, ensuring that its spatial distribution has directional continuity.
[0060] MCC Loss Function The definition is as follows:
[0061] ;
[0062] The MCC loss function is a binary classification cross-entropy loss function used to determine whether a pixel is a medium region. It controls the recognition accuracy of pixel-level medium regions. The real label, i.e. the supervision label, represents the probability that the pixel is a medium region; To predict the probability that the pixel belongs to the medium;
[0063] S2122. Definition of Medium Filling Degree Characteristics: The medium filling degree characteristics include structural boundary strength and structural regularity, as defined below:
[0064] Structural boundary strength measures the sharpness of a boundary in a cross-sectional image, corresponding to the spatial rate of change of the structural boundary. The definition is as follows:
[0065] ;
[0066] in, This is the first-order partial derivative of the cross-sectional image, i.e., the cross-sectional image in the horizontal direction ( x The gradient of the axis. The first-order partial derivative of a two-dimensional cross-sectional image, i.e., the cross-sectional image in the vertical direction ( y The gradient of the axis;
[0067] Structural regularity is measured by the consistency between a local area of a cross-sectional image and a standard structure. Local structural similarity index, degree of structural regularity The definition is as follows:
[0068] ;
[0069] in, I For cross-sectional images, The mean map of the blurred cross-sectional image provides a structural reference for the cross-sectional image. In order to be in Calculate local structural similarity;
[0070] S2123. Construction of the Filling Scoring Function: To standardize the measurement of structural ambiguity in the medium region, a filling scoring function is defined. as follows:
[0071] ;
[0072] in, These are weighting coefficients. Preferred , , This indicates that the eigenvalues are normalized to... Interval, fill scoring function The larger the value of , the more blurred the medium region is, the higher the structural uncertainty, and the more likely the medium is to fill.
[0073] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the media region in the above embodiments, step S2130 includes:
[0074] S2131. Define the high-dielectric region and set the high-dielectric empirical threshold. The region where the medium probability is higher than the high medium empirical threshold is defined as the high medium region, as shown in the following formula:
[0075] ;
[0076] in, For medium probability diagram; For high medium empirical threshold, Preferred high medium empirical threshold It is 0.75;
[0077] S2132. Define the low-medium region and set the low-medium empirical threshold. The region where the medium probability is lower than the empirical threshold for high medium is defined as the low medium region, as shown in the following formula:
[0078] ;
[0079] in, For medium probability diagram; The empirical threshold for low-medium media is [value], and its range is [value]. Preferred low medium empirical threshold It is 0.25.
[0080] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2200 includes:
[0081] S2210. Define the curvature edge-supervised auxiliary loss function. To improve the response to boundaries and enhance edge recognition capabilities, a curvature edge-supervised auxiliary loss function is defined. as follows:
[0082] ;
[0083] in, The supervisory signal in the curvature edge supervision auxiliary loss function is Gaussian smoothed on the target boundary in the high medium region to remove noise, and a logical AND is performed with the real boundary region to retain the reliable region and generate the supervisory signal.
[0084] S2220. Calculate the first and second derivatives of the cross-sectional image. For a two-dimensional cross-sectional image... The partial derivatives are calculated using the following formula:
[0085] , , , , ;
[0086] in: , It is a two-dimensional cross-sectional image. The first-order partial derivative; , , It is a cross-sectional image. The second-order partial derivative;
[0087] S2230. Calculate the curvature diagram, using a discrete algorithm for discrete derivation, calculate local curvature, measure the degree of boundary curvature, and generate the curvature diagram. The formula is:
[0088] ;
[0089] in, It is a very small positive number. Preferred To avoid a denominator of 0; Ensure the normalization of the curvature scale;
[0090] S2240, Curvature Edge Map Definition, Setting Curvature Intensity Threshold From the curvature diagram Extract edge information and define a curvature edge map. The formula is as follows:
[0091] ;
[0092] in, Curvature intensity threshold, range The preferred value is 0.1. This represents a local maximum of curvature;
[0093] S2250, Generate an enhanced boundary prediction map, and combine it with the medium probability map. With curvature edge map Fusion generates an enhanced boundary prediction map. The formula is as follows:
[0094] ;
[0095] in, It is the fusion weighting coefficient. .
[0096] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the medium region in the above embodiments, the discrete algorithm includes Sobel, Scharr, or central difference.
[0097] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, the perceptual attention includes:
[0098] Channel attention learns which channels are more important from the channel dimension of the feature map, highlighting the role of channels in semantic recognition, including feature channels of impacted teeth, neural tubes, and mediator structures;
[0099] Spatial attention learns from the spatial dimensions of the feature map which parts of the space are more important, including the boundary between impacted teeth and the neural canal, structural abrupt changes in curved and concave areas, and areas with significant changes in curvature.
[0100] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2300 includes:
[0101] S2310. Semantic feature extraction: Extracting semantic features. F , means as follows:
[0102] ;
[0103] in: The number of channels, i.e., the number of semantic categories. H Indicates the height of the input cross-sectional image. W Indicates the width of the input cross-sectional image;
[0104] S2320, Channel-weighted feature map calculation, for semantic features F Global average pooling is used to extract channel descriptions, predict channel attention weights, and calculate channel weighted feature maps. ;
[0105] S2330, Spatial Enhancement Semantic Feature Map Calculation, for semantic features F Channel compression is performed, followed by average pooling and max pooling along the channel dimension to obtain the average pooled feature map. and max pooling feature map splicing average pooling feature maps and max pooling feature map Then perform convolution to generate a spatial attention map. Spatial attention map Broadcast back channel dimension, and semantic features F Dot product yields a spatially enhanced semantic feature map. ;
[0106] S2340, Perceptual feature calculation, weighted feature map through residual connection channel. and spatially enhanced semantic feature maps To obtain perceptual features The formula is as follows:
[0107] ;
[0108] S2350, Perceptual Feature Re-prediction, for perceptual features Perform further segmentation and prediction using the following formula:
[0109] ;
[0110] in, , This indicates a lightweight prediction head that enables fast prediction in low-medium regions without the need for large-scale upsampling and complex skip connections.
[0111] S2360. Construction of the perceptual supervised loss function: Calculate the auxiliary loss of the attention mechanism in low-media regions, i.e., the perceptual supervised loss, to improve the perceptual ability of difficult-to-identify regions. The perceptual supervised loss function... The definition is as follows:
[0112] ;
[0113] in, It is a low dielectric region. It is a perceptual feature. It is the result of re-predicting based on perceptual features. These are the corresponding real tags. This is the standard segmentation loss function.
[0114] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, the standard segmentation loss is cross-entropy loss (CE loss) + Dice loss (Dice loss).
[0115] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2320 includes:
[0116] S2321, Channel description vector calculation, for semantic features F Global average pooling is used to extract channel descriptions. For each channel, global average pooling (GAP) is performed to obtain the average description vector for a single channel. The formula is as follows:
[0117] ;
[0118] in, , C The total number of channels. The position on the c-th channel is The feature map values of the pixels, and These represent the height and width of the feature map, respectively. i , j These are the row and column indices of the pixels being traversed, respectively. For the first c A single-channel average description vector;
[0119] The set of all individual channel description vectors is the channel description vector, as shown in the following formula:
[0120] ;
[0121] in, R To represent the set of real numbers, that is, the set of all real numbers;
[0122] S2322, Channel attention weight prediction: Channel attention weights are predicted using a multilayer perceptron (MLP) in the impacted tooth and neural tube segmentation model. The channel description vectors are then used to predict these weights. z The input consists of two MLP layers, which are essentially two fully connected layers. The first layer performs dimensionality reduction, and the second layer performs dimensionality increase and outputs channel weights. It contains one hidden layer to generate channel attention weights. The attention weight of the c-th channel The formula is as follows:
[0123] ;
[0124] in, , These are the weight matrices of two fully connected layers in a channel attention MLP. It is the ReLU activation function; It's the Sigmoid function, which normalizes the weights to... ;
[0125] S2323, Channel-weighted feature map calculation, using broadcast channel attention weights. and semantic features F Dot product yields the channel-weighted feature map. :
[0126] .
[0127] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2330 includes:
[0128] S2331, Channel compression, for semantic features F Channel compression is performed, followed by average pooling and max pooling along the channel dimension, resulting in two feature maps: the average pooling feature map and the max pooling feature map. and max pooling feature map The formula is as follows:
[0129] ;
[0130] ;
[0131] S2332. Concatenate pooling features and perform convolution, then concatenate the average pooling feature maps. and max pooling feature map , obtain the size splicing pooling feature map and through a Convolution generates spatial attention maps , The formula is as follows:
[0132] ;
[0133] S2333, Spatial Enhancement Semantic Feature Map Calculation, Including Spatial Attention Map Broadcast back channel dimension, and semantic features F Dot product yields a spatially enhanced semantic feature map. The formula is as follows:
[0134] .
[0135] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2400 includes:
[0136] S2410. Feature Fusion: The enhanced boundary prediction map, channel-weighted feature map, and spatially enhanced semantic feature map are fused using the following formula:
[0137] ;
[0138] in, Indicates feature fusion;
[0139] S2420, Practical threshold binarization, will fuse features Practical threshold Binarization is performed using the following formula:
[0140] ;
[0141] in, Indicates an indicator function that satisfies the condition. It is 1 if it is true, otherwise it is 0; The optimal selection was determined through three-fold cross-validation. ; Represents a binary mask;
[0142] S2430, Boundary Extraction: Boundaries are obtained using the Canny edge detection algorithm.
[0143] ;
[0144] in, The initial contour obtained is the edge detection function. That is, the predicted profile of the medium region.
[0145] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S2500 includes:
[0146] S2510. Homogeneous representation of the initial contour: The initial contour (predicted contour of the medium region) is represented homogeneously using the following formula:
[0147] ;
[0148] S2520, Corrected Profile Definition: Geometric correction is performed on the initial profile (predicted profile for the medium region) to obtain the corrected profile. The corrected profile is defined as follows:
[0149] ;
[0150] in, As the initial outline, This is the geometric correction transformation matrix;
[0151] S2530. Definition of Geometric Correction Loss Function: The geometric correction loss function is defined as follows:
[0152] ;
[0153] in, Represents coordinate mapping, Indicates the first i The coordinates of the boundary points in a homogeneous coordinate system. This represents the coordinates of the boundary points in the homogeneous coordinate system after geometric correction. Represents the coordinates of the true boundary points in a homogeneous coordinate system. This represents the prediction confidence level of the i-th boundary point. A value close to 1 indicates that the boundary point is located in a high-medium region, which is preferred. =Above 0.85 is the high dielectric region;
[0154] S2540, Definition of the overall loss function, overall loss function The formula is as follows:
[0155] ;
[0156] in, These are the gradient regularization terms. Weighting coefficients, CML connection axis loss function Weighting coefficients, MBO boundary offset loss function The weighting coefficients of the CMO spatial orientation loss function Weighting coefficients of the MCC medium probability graph loss function Curvature edge supervision auxiliary loss function Weighting coefficients, perceptual supervision loss function Weighting coefficients, geometric correction loss function Weighting coefficients;
[0157] S2550, the segmentation loss function for impacted teeth and neural tubes is defined, integrating high-medial and low-medial regions, and performing unified feature fusion with geometric correction. The segmentation results for impacted teeth and neural tubes are output through iterative optimization. The segmentation loss function for impacted teeth and neural tubes is expressed as:
[0158] ;
[0159] in, These are the weights of the overall loss function and the geometric correction loss function, respectively.
[0160] Optionally, in the method for segmenting impacted teeth and nerve canals based on the medium region in any of the above embodiments, geometric correction includes translation, rotation, scaling, and local adjustment or correction of the shape of the contour.
[0161] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, the maximum number of training rounds is... =1000.
[0162] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, the performance threshold includes a Dice coefficient greater than 0.95 and an average segmentation value of impacted teeth and neural tubes. IoU Greater than 0.9.
[0163] Optionally, in the method for segmenting impacted teeth and neural tubes based on the media region in any of the above embodiments, step S4000 includes:
[0164] S4100, Continue training begins. When the number of training rounds is greater than or equal to the maximum number of training rounds, a temporary impacted tooth and neural tube segmentation model is obtained. The weights of each loss function in the overall loss function are modified according to the loss function adjustment strategy. The number of training rounds is initialized to 0.
[0165] S4200, continue training, continue training the temporary impacted tooth and neural tube segmentation model, continue training round number one;
[0166] S4300: Determine the number of training epochs and performance threshold. If the number of training epochs is greater than or equal to the maximum number of training epochs, modify the weights of each loss function in the overall loss function according to the loss function adjustment strategy, set the number of training epochs to 0, and return to step S4200. If the number of training epochs is less than the maximum number of training epochs and meets the performance threshold, proceed to step S4400. If the number of training epochs is less than the maximum number of training epochs and does not meet the performance threshold, proceed to step S4200.
[0167] S4400, Continue training to complete, and obtain the trained impacted tooth and neural tube segmentation model.
[0168] Optionally, in the method for segmenting impacted teeth and neural tubes based on the medium region in any of the above embodiments, the loss function adjustment strategy is: preferentially increasing the curvature edge supervision auxiliary loss function according to a predetermined optimal weight step size. Weighting coefficients and perceptual supervision loss function The weight coefficients are adjusted until the maximum value of the optimal weight change is reached; then, the gradient regularization term is reduced according to a predetermined suboptimal weight step size. Weighting coefficients and geometric correction loss function The weighting coefficients are adjusted until the suboptimal maximum weight change is reached; finally, the CML connection axis loss function is reduced by a predetermined normal weight step size. Weighting coefficients, MBO boundary offset loss function The weighting coefficients of the CMO spatial orientation loss function Weighting coefficients of the MCC medium probability graph loss function This continues until the maximum value of the ordinary weight change is reached.
[0169] Furthermore, in the method for segmenting impacted teeth and nerve tubes based on the medium region in the above embodiments, the predetermined optimal weight step size is 0.05, and the maximum optimal weight change value is 0.3.
[0170] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the medium region in the above embodiments, the predetermined suboptimal weight step size is 0.03, and the maximum suboptimal weight change value is 0.2.
[0171] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the medium region in the above embodiments, the predetermined ordinary weight step size is 0.01, and the maximum ordinary weight change value is 0.1.
[0172] Furthermore, in the method for segmenting impacted teeth and neural tubes based on the medium region in the above embodiments, the maximum number of training rounds is 100.
[0173] This invention spatially perceives and divides the target region by constructing a medium probability distribution map. Based on the degree of medium filling (e.g., grayscale gradient, structural similarity), it divides the region into high-medium and low-medium areas. In high-medium areas with clear boundaries, a curvature edge supervision mechanism is used to calculate the curvature difference between the predicted and true boundaries for boundary optimization. In low-medium areas with blurred boundaries and overlapping structures, a combination of channel attention and spatial attention is used to enhance semantic feature separation capabilities. This invention improves the accuracy and fit of boundary positions, significantly enhances the segmentation accuracy and stability of low-medium areas, and solves the technical problem of accurate semantic segmentation of structurally overlapping, blurred, and discontinuous boundary regions in dental CBCT images.
[0174] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description
[0175] Figure 1 This is a flowchart of an exemplary embodiment of a method for segmenting impacted teeth and nerve tubes based on a media region;
[0176] Figure 2 This is a flowchart illustrating the continued training of a temporary model of a method for segmenting impacted teeth and neural tubes based on a media region, as described in an exemplary embodiment. Detailed Implementation
[0177] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0178] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of components is schematically exaggerated in some places in the drawings.
[0179] This application provides a method for segmenting impacted teeth and nerve canals based on a media region, such as... Figure 1 As shown, it includes the following steps:
[0180] S1000. Preparation: Collect CBCT training set, initially construct a segmentation model for impacted teeth and neural tube based on deep learning, adopt a general deep learning architecture, and initialize the number of training rounds to 0.
[0181] S2000, Model Structure Optimization: Input the CBCT training set into the impacted tooth and neural tube segmentation model, divide the media region into high-media and low-media regions, optimize the curvature difference in the high-media region, enhance semantic feature separation in the low-media region, predict the contour of the media region and perform geometric correction, and construct the overall loss function and the impacted tooth and neural tube segmentation loss function; specifically including:
[0182] S2100, Media Region Segmentation: Input the CBCT training set into the impacted tooth and neural tube segmentation model. Based on media filling characteristics, including gray-level gradient and structural similarity, divide the media region into high-media and low-media regions. Construct a gradient regularization loss function to generate a media probability map; specifically including:
[0183] S2110. Input image processing: Input the CBCT training set. CBCT is three-dimensional volume data. Two-dimensional cross-sectional images are extracted through sagittal and coronal slices. , means as follows:
[0184] ;
[0185] in, R represents The set of real numbers is the set of all real numbers. H Indicates the height of the input cross-sectional image. W Indicates the width of the input cross-sectional image;
[0186] S2120. Intermediate structure construction: Intermediate structure features are used to describe the geometric and topological probabilistic characteristics of the medium region, defining medium fill degree features and constructing a fill scoring function; intermediate structure features include:
[0187] CML (Connecting Middle Line) is the central axis connecting the impacted tooth and the neural canal, defining the central path of the medial region.
[0188] MBO (Middle Boundary Offset) is the boundary offset distance from CML outward along the normal projection direction, used to reconstruct the rough geometric profile of the medium region.
[0189] CMO (Central Motion Orientation) is the orientation (unit vector) of CML in three-dimensional space, which enhances the spatial consistency of region modeling by maintaining its continuity.
[0190] MCC (Middle Connection Confidence) is the probability that the current pixel is in the medium region. It is output by regression through the Sigmoid function and is used to determine the existence of the blurred region.
[0191] Step S2120 specifically includes:
[0192] S2121, Intermediate Structure Feature Extraction: This section describes the geometric and topological probabilistic characteristics of the medium region using intermediate structure features. It defines the intermediate structure feature loss function, including:
[0193] CML Loss Function The definition is as follows:
[0194] ;
[0195] in, Indicates the predicted coordinates of the central axis point. Represents the true coordinates of the central axis point, indicating its position. The coordinates of points on the real CML in the cross-sectional image are used to supervise the model to learn and generate the correct central axis position; the points on the CML refer to the sampling points of the neural tube or impacted tooth on the skeleton line in the cross-sectional image, which are simplified representations of the morphology and orientation in a local area;
[0196] MBO loss function The definition is as follows:
[0197] ;
[0198] in, This represents the predicted boundary offset. This represents the distance from the CML to the true boundary along the central axis at the current point, i.e., the offset, which is used to constrain the width or thickness of the region generated by the neural network model.
[0199] CMO Loss Function The definition is as follows:
[0200] ;
[0201] in, Represents the predicted direction vector. The true vector representing the direction of the central axis is used to constrain the predicted direction field, ensuring that its spatial distribution has directional continuity.
[0202] MCC Loss Function The definition is as follows:
[0203] ;
[0204] The MCC loss function is a binary classification cross-entropy loss function used to determine whether a pixel is a medium region. It controls the recognition accuracy of pixel-level medium regions. The real label, i.e. the supervision label, represents the probability that the pixel is a medium region; To predict the probability that the pixel belongs to the medium;
[0205] S2122. Definition of Medium Filling Degree Characteristics: The medium filling degree characteristics include structural boundary strength and structural regularity, as defined below:
[0206] Structural boundary strength measures the sharpness of a boundary in a cross-sectional image, corresponding to the spatial rate of change of the structural boundary. The definition is as follows:
[0207] ;
[0208] in, This is the first-order partial derivative of the cross-sectional image, i.e., the cross-sectional image in the horizontal direction ( x The gradient of the axis. The first-order partial derivative of a two-dimensional cross-sectional image, i.e., the cross-sectional image in the vertical direction ( y The gradient of the axis;
[0209] Structural regularity is measured by the consistency between a local area of a cross-sectional image and a standard structure. Local structural similarity index, degree of structural regularity The definition is as follows:
[0210] ;
[0211] in, I For cross-sectional images, The mean map of the blurred cross-sectional image provides a structural reference for the cross-sectional image. In order to be in Calculate local structural similarity;
[0212] S2123. Construction of the Filling Scoring Function: To standardize the measurement of structural ambiguity in the medium region, a filling scoring function is defined. as follows:
[0213] ;
[0214] in, These are weighting coefficients. Preferred , , This indicates that the eigenvalues are normalized to... Interval, fill scoring function The larger the value of , the more blurred the medium region is, the higher the structural uncertainty, and the more likely the medium is to fill.
[0215] S2130, Media Zone Division, defining high-media zones and low-media zones; specifically including:
[0216] S2131. Define the high-dielectric region and set the high-dielectric empirical threshold. The region where the medium probability is higher than the high medium empirical threshold is defined as the high medium region, as shown in the following formula:
[0217] ;
[0218] in, For medium probability diagram; For high medium empirical threshold, Preferred high medium empirical threshold It is 0.75;
[0219] S2132. Define the low-medium region and set the low-medium empirical threshold. The region where the medium probability is lower than the empirical threshold for high medium is defined as the low medium region, as shown in the following formula:
[0220] ;
[0221] in, For medium probability diagram; The empirical threshold for low-medium media is [value], and its range is [value]. Preferred low medium empirical threshold It is 0.25;
[0222] S2140. Construct a gradient regularization loss function, and set supervision labels for the medium region, as shown in the following formula:
[0223] ;
[0224] Since tissue structures exhibit spatial continuity in images, a gradient regularization loss function is constructed to increase the spatial smoothness of the media region and reduce excessive fragmentation, as shown in the following formula:
[0225] ;
[0226] S2150. Medium probability map generation: Perform classification prediction for each pixel, generate and output the medium probability map. :
[0227] ;
[0228] in, Represents pixels The probability of being a medium.
[0229] S2200, high-medium region curvature difference optimization, defines a curvature edge supervised auxiliary loss function, calculates a curvature map, defines a curvature edge map, fuses the curvature edge map with the medium probability map, performs boundary enhancement, and generates an enhanced boundary prediction map; specifically including:
[0230] S2210. Define the curvature edge-supervised auxiliary loss function. To improve the response to boundaries and enhance edge recognition capabilities, a curvature edge-supervised auxiliary loss function is defined. as follows:
[0231] ;
[0232] in, The supervisory signal in the curvature edge supervision auxiliary loss function is Gaussian smoothed on the target boundary in the high medium region to remove noise, and a logical AND is performed with the real boundary region to retain the reliable region and generate the supervisory signal.
[0233] S2220. Calculate the first and second derivatives of the cross-sectional image. For a two-dimensional cross-sectional image... The partial derivatives are calculated using the following formula:
[0234] , , , , ;
[0235] in: , It is a two-dimensional cross-sectional image. The first-order partial derivative; , , It is a cross-sectional image. The second-order partial derivative;
[0236] S2230. Calculate the curvature diagram using the Sobel discrete algorithm for discrete derivation, calculate local curvature, measure the degree of boundary curvature, and generate the curvature diagram. The formula is:
[0237] ;
[0238] in, It is a very small positive number. Preferred To avoid a denominator of 0; Ensure the normalization of the curvature scale;
[0239] S2240, Curvature Edge Map Definition, Setting Curvature Intensity Threshold From the curvature diagram Extract edge information and define a curvature edge map. The formula is as follows:
[0240] ;
[0241] in, Curvature intensity threshold, range The preferred value is 0.1. This represents a local maximum of curvature;
[0242] S2250, Generate an enhanced boundary prediction map, and combine it with the medium probability map. With curvature edge map Fusion generates an enhanced boundary prediction map. The formula is as follows:
[0243] ;
[0244] in, It is the fusion weighting coefficient. .
[0245] S2300, semantic feature separation enhancement in low-medium regions: A perceptual attention-guided mechanism is introduced in low-medium regions to extract semantic features. Global average pooling is used to extract channel descriptions from the semantic features, and channel-weighted feature maps are calculated. Channel compression is applied to the semantic features, and spatially enhanced semantic feature maps are calculated. Perceptual features are obtained by concatenating the channel-weighted feature maps and the spatially enhanced semantic feature maps through residual connections. A perceptual supervised loss function is constructed. Perceptual attention includes:
[0246] Channel attention learns which channels are more important from the channel dimension of the feature map, highlighting the role of channels in semantic recognition, including feature channels of impacted teeth, neural tubes, and mediator structures;
[0247] Spatial attention learns from the spatial dimension of the feature map which part of the space is more important, including the boundary between impacted teeth and the nerve canal, structural abrupt points in curved and concave areas, and areas with significant changes in curvature.
[0248] Step S2300 specifically includes:
[0249] S2310. Semantic feature extraction: Extracting semantic features. F , means as follows:
[0250] ;
[0251] in: The number of channels, i.e., the number of semantic categories. H Indicates the height of the input cross-sectional image. W Indicates the width of the input cross-sectional image;
[0252] S2320, Channel-weighted feature map calculation, for semantic features F Global average pooling is used to extract channel descriptions, predict channel attention weights, and calculate channel weighted feature maps. Specifically, this includes:
[0253] S2321, Channel description vector calculation, for semantic features F Global average pooling is used to extract channel descriptions. For each channel, global average pooling (GAP) is performed to obtain the average description vector for a single channel. The formula is as follows:
[0254] ;
[0255] in, , C The total number of channels. The position on the c-th channel is The feature map values of the pixels, and These represent the height and width of the feature map, respectively. i , j These are the row and column indices of the pixels being traversed, respectively. For the first c A single-channel average description vector;
[0256] The set of all individual channel description vectors is the channel description vector, as shown in the following formula:
[0257] ;
[0258] in, R To represent the set of real numbers, that is, the set of all real numbers;
[0259] S2322, Channel attention weight prediction: Channel attention weights are predicted using a multilayer perceptron (MLP) in the impacted tooth and neural tube segmentation model. The channel description vectors are then used to predict these weights. z The input consists of two MLP layers, which are essentially two fully connected layers. The first layer performs dimensionality reduction, and the second layer performs dimensionality increase and outputs channel weights. It contains one hidden layer to generate channel attention weights. The attention weight of the c-th channel The formula is as follows:
[0260] ;
[0261] in, , These are the weight matrices of two fully connected layers in a channel attention MLP. It is the ReLU activation function; It's the Sigmoid function, which normalizes the weights to... ;
[0262] S2323, Channel-weighted feature map calculation, using broadcast channel attention weights. and semantic features F Dot product yields the channel-weighted feature map. :
[0263] ;
[0264] S2330, Spatial Enhancement Semantic Feature Map Calculation, for semantic features F Channel compression is performed, followed by average pooling and max pooling along the channel dimension to obtain the average pooled feature map. and max pooling feature map splicing average pooling feature maps and max pooling feature map Then perform convolution to generate a spatial attention map. Spatial attention map Broadcast back channel dimension, and semantic features F Dot product yields a spatially enhanced semantic feature map. Specifically, this includes:
[0265] S2331, Channel compression, for semantic features F Channel compression is performed, followed by average pooling and max pooling along the channel dimension, resulting in two feature maps: the average pooling feature map and the max pooling feature map. and max pooling feature map The formula is as follows:
[0266] ;
[0267] ;
[0268] S2332. Concatenate pooling features and perform convolution, then concatenate the average pooling feature maps. and max pooling feature map , obtain the size splicing pooling feature map and through a Convolution generates spatial attention maps , The formula is as follows:
[0269] ;
[0270] S2333, Spatial Enhancement Semantic Feature Map Calculation, Including Spatial Attention Map Broadcast back channel dimension, and semantic features F Dot product yields a spatially enhanced semantic feature map. The formula is as follows:
[0271] .
[0272] S2340, Perceptual feature calculation, weighted feature map through residual connection channel. and spatially enhanced semantic feature maps To obtain perceptual features The formula is as follows:
[0273] ;
[0274] S2350, Perceptual Feature Re-prediction, for perceptual features Perform further segmentation and prediction using the following formula:
[0275] ;
[0276] in, , This indicates a lightweight prediction head that enables fast prediction in low-medium regions without the need for large-scale upsampling and complex skip connections.
[0277] S2360. Construction of the perceptual supervised loss function: Calculate the auxiliary loss of the attention mechanism in low-media regions, i.e., the perceptual supervised loss, to improve the perceptual ability of difficult-to-identify regions. The perceptual supervised loss function... The definition is as follows:
[0278] ;
[0279] in, It is a low dielectric region. It is a perceptual feature. It is the result of re-predicting based on perceptual features. These are the corresponding real tags. The standard segmentation loss function is the sum of cross-entropy loss (CE loss) and Dice loss (Dice loss).
[0280] S2400, Medium Region Contour Prediction: This involves extracting boundaries from the enhanced boundary prediction map, channel-weighted feature map, and spatially enhanced semantic feature map to obtain the predicted medium region contour; specifically including:
[0281] S2410. Feature Fusion: The enhanced boundary prediction map, channel-weighted feature map, and spatially enhanced semantic feature map are fused using the following formula:
[0282] ;
[0283] in, Indicates feature fusion;
[0284] S2420, Practical threshold binarization, will fuse features Practical threshold Binarization is performed using the following formula:
[0285] ;
[0286] in, Indicates an indicator function that satisfies the condition. It is 1 if it is true, otherwise it is 0; The optimal selection was determined through three-fold cross-validation. ; Represents a binary mask;
[0287] S2430, Boundary Extraction: Boundaries are obtained using the Canny edge detection algorithm.
[0288] ;
[0289] in, The initial contour obtained is the edge detection function. That is, the predicted profile of the medium region.
[0290] S2500, geometric correction, uses CML (Connecting Middle Line) as the segmentation guide line to perform geometric correction on the predicted contour of the medium region. It defines a geometric correction loss function and constructs an overall loss function and a segmentation loss function for impacted teeth and neural tubes, specifically including:
[0291] S2510. Homogeneous representation of the initial contour: The initial contour (predicted contour of the medium region) is represented homogeneously using the following formula:
[0292] ;
[0293] S2520, Corrected Profile Definition: Geometric correction is performed on the initial profile (predicted profile for the medium region), including translation, rotation, scaling, and local adjustments or corrections to the shape of the profile, to obtain the corrected profile. The corrected profile is defined as follows:
[0294] ;
[0295] in, This is the initial profile (predicted profile for the medium region). This is the geometric correction transformation matrix;
[0296] S2530. Definition of Geometric Correction Loss Function: The geometric correction loss function is defined as follows:
[0297] ;
[0298] in, Represents coordinate mapping, Indicates the first i The coordinates of the boundary points in a homogeneous coordinate system. This represents the coordinates of the boundary points in the homogeneous coordinate system after geometric correction. Represents the coordinates of the true boundary points in a homogeneous coordinate system. This represents the prediction confidence level of the i-th boundary point. A value close to 1 indicates that the boundary point is located in a high-medium region, which is preferred. =Above 0.85 is the high dielectric region;
[0299] S2540, Definition of the overall loss function, overall loss function The formula is as follows:
[0300] ;
[0301] in, These are the gradient regularization terms. Weighting coefficients, CML connection axis loss function Weighting coefficients, MBO boundary offset loss function The weighting coefficients of the CMO spatial orientation loss function Weighting coefficients of the MCC medium probability graph loss function Curvature edge supervision auxiliary loss function Weighting coefficients, perceptual supervision loss function Weighting coefficients, geometric correction loss function Weighting coefficients;
[0302] S2550, the segmentation loss function for impacted teeth and neural tubes is defined, integrating high-medial and low-medial regions, and performing unified feature fusion with geometric correction. The segmentation results for impacted teeth and neural tubes are output through iterative optimization. The segmentation loss function for impacted teeth and neural tubes is expressed as:
[0303] ;
[0304] in, These are the weights of the overall loss function and the geometric correction loss function, respectively.
[0305] S3000 model training uses the AdamW optimizer. The training epoch count is incremented by one. When the training epoch count is less than the maximum training epoch count... =1000 and meets the performance threshold, which includes a Dice coefficient greater than 0.95 and average segmentation of impacted teeth and neural tubes. IoU If the value is greater than 0.9, the training ends and a well-trained segmentation model of impacted teeth and neural tubes is obtained. Step S5000 is executed. If the number of training rounds is less than the maximum number of training rounds and the performance threshold is not met, the process returns to step S2000.
[0306] S4000: The temporary model continues training. When the number of training epochs is greater than or equal to the maximum number of training epochs, a temporary impacted tooth and neural tube segmentation model is obtained. The weights of each loss function in the overall loss function are modified according to the loss function adjustment strategy. The temporary impacted tooth and neural tube segmentation model continues training until the performance threshold is met. Training then ends, resulting in a well-trained impacted tooth and neural tube segmentation model. Figure 2 As shown, it specifically includes:
[0307] S4100, Continue training begins. When the number of training rounds is greater than or equal to the maximum number of training rounds, a temporary impacted tooth and neural tube segmentation model is obtained. The weights of each loss function in the overall loss function are modified according to the loss function adjustment strategy. The number of training rounds is initialized to 0.
[0308] S4200, continue training, continue training the temporary impacted tooth and neural tube segmentation model, continue training round number one;
[0309] S4300: Determine the number of training epochs and performance threshold. When the number of training epochs is greater than or equal to the maximum number of training epochs (100), modify the weights of each loss function in the overall loss function according to the loss function adjustment strategy, set the number of training epochs to 0, and return to step S4200. The loss function adjustment strategy is: prioritize adding a curvature edge supervision auxiliary loss function with a predetermined optimal weight step size of 0.05. Weighting coefficients and perceptual supervision loss function The weight coefficients are adjusted until the optimal weight change maximum value of 0.3 is reached; then the gradient regularization term is reduced according to a predetermined suboptimal weight step size of 0.03. Weighting coefficients and geometric correction loss function The weighting coefficients are adjusted until the suboptimal maximum weight change of 0.2 is reached; finally, the CML connection axis loss function is reduced by a predetermined normal weight step size of 0.01. Weighting coefficients, MBO boundary offset loss function The weighting coefficients of the CMO spatial orientation loss function Weighting coefficients of the MCC medium probability graph loss function Continue until the maximum value of the normal weight change is reached (0.1); when the number of training rounds is less than the maximum number of training rounds (100) and the performance threshold is met, execute step S4400; when the number of training rounds is less than the maximum number of training rounds and the performance threshold is not met, execute step S4200.
[0310] S4400, Continue training to complete, and obtain the trained impacted tooth and neural tube segmentation model.
[0311] S5000, Impacted Tooth and Neural Tube Segmentation Prediction: Using a trained impacted tooth and neural tube segmentation model, inference and segmentation prediction are performed on CBCT, and the output image includes the mask of the impacted tooth region and the mask of the neural tube region.
[0312] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for segmenting impacted teeth and nerve tubes based on a media region, characterized in that, Includes the following steps: S1000. Preparation: Collect CBCT training set, and initially construct a segmentation model of impacted teeth and neural tube based on deep learning. Initialize the number of training rounds to 0. S2000, Model Structure Optimization: Input the CBCT training set into the impacted tooth and neural tube segmentation model, divide the media region into high-media region and low-media region, optimize the curvature difference of the high-media region, enhance the semantic feature separation of the low-media region, predict the contour of the media region and perform geometric correction, and construct the overall loss function and the impacted tooth and neural tube segmentation loss function; specifically including: S2100, Media region segmentation: Input the CBCT training set into the impacted tooth and neural tube segmentation model, and divide the media region into the high media region and the low media region according to the media filling degree characteristics. Construct a gradient regularization loss function and generate a media probability map. S2200, High-medium region curvature difference optimization, define curvature edge supervision auxiliary loss function, calculate curvature map, define curvature edge map, fuse the curvature edge map with the medium probability map, perform boundary enhancement, and generate enhanced boundary prediction map; S2300, Semantic feature separation enhancement in low-medium regions: A perceptual attention guidance mechanism is introduced in the low-medium region to extract semantic features. The semantic features are then subjected to global average pooling to extract channel descriptions. Channel-weighted feature maps are calculated. The semantic features are then subjected to channel compression. Spatially enhanced semantic feature maps are calculated. The channel-weighted feature maps and spatially enhanced semantic feature maps are then connected by residuals to obtain perceptual features. A perceptual supervised loss function is then constructed. S2400, Medium region contour prediction: Boundary extraction is performed on the enhanced boundary prediction map, the channel weighted feature map, and the spatially enhanced semantic feature map to obtain the predicted contour of the medium region. S2500, geometric correction, using CML as a segmentation guide line, performs geometric correction on the predicted contour of the medium region, defines the geometric correction loss function, and constructs the overall loss function and the segmentation loss function for impacted teeth and nerve canals; S3000: Model training is performed using the AdamW optimizer. The number of training epochs is incremented by one. When the number of training epochs is less than the maximum number of training epochs and the performance threshold is met, training ends, and a trained impacted tooth and neural tube segmentation model is obtained. Step S5000 is then executed. When the number of training epochs is less than the maximum number of training epochs and the performance threshold is not met, the process returns to step S2000. When the number of training epochs is greater than or equal to the maximum number of training epochs, step S4000 is executed. S4000. The temporary model continues to train. When the number of training rounds is greater than or equal to the maximum number of training rounds, a temporary impacted tooth and neural tube segmentation model is obtained. The weights of each loss function in the overall loss function are modified according to the loss function adjustment strategy. The temporary impacted tooth and neural tube segmentation model is used to continue training until the iteration meets the performance threshold. The training ends, and the trained impacted tooth and neural tube segmentation model is obtained. S5000, Impacted Tooth and Neural Tube Segmentation Prediction: Using the trained impacted tooth and neural tube segmentation model, inference and segmentation prediction are performed on CBCT, and an image containing the impacted tooth region mask and the neural tube region mask is output.
2. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 1, characterized in that, Step S2100 includes: S2110. Input image processing: Input the CBCT training set and extract two-dimensional cross-sectional images through sagittal and coronal slices. , means as follows: ; Where R represents the set of real numbers, that is, the set of all real numbers, H represents the height of the input cross-sectional image, and W represents the width of the input cross-sectional image; S2120. Intermediate structure construction: Use intermediate structure features to describe the geometric and topological probability characteristics of the medium region, define the medium filling degree features, and construct a filling scoring function. S2130, Dielectric region division, defining the high dielectric region and the low dielectric region; S2140. Construct a gradient regularization loss function, and set a supervision label for the medium region, as shown in the following formula: ; S2150. Medium probability map generation: Perform classification prediction for each pixel, generate and output the medium probability map. : 。 3. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 2, characterized in that, Step S2120 includes: S2121, Intermediate Structure Feature Extraction: This section describes the geometric and topological probabilistic characteristics of the medium region using intermediate structure features. It defines the intermediate structure feature loss function, including: CML Loss Function The definition is as follows: ; in, Indicates the predicted coordinates of the central axis point. Represents the true coordinates of the central axis point, indicating its position. The coordinates of points on the real CML in the cross-sectional image are used to supervise the model to learn and generate the correct central axis position; the points on the CML refer to the sampling points of the neural tube or impacted teeth on the skeleton line in the cross-sectional image, which are simplified representations of the morphology and orientation in a local area; MBO loss function The definition is as follows: ; in, This represents the predicted boundary offset. This represents the distance from the CML to the true boundary along the central axis at the current point, i.e., the offset, which is used to constrain the width or thickness of the region generated by the neural network model. CMO Loss Function The definition is as follows: ; in, Represents the predicted direction vector. The true vector representing the direction of the central axis is used to constrain the predicted direction field, ensuring that its spatial distribution has directional continuity. MCC Loss Function The definition is as follows: ; The MCC loss function is a binary classification cross-entropy loss function used to determine whether a pixel is a medium region. It controls the recognition accuracy of pixel-level medium regions. The real label, i.e. the supervision label, represents the probability that the pixel is a medium region; To predict the probability that the pixel belongs to the medium; S2122. Definition of Medium Filling Degree Characteristics: The medium filling degree characteristics include structural boundary strength and structural regularity, as defined below: Structural boundary strength measures the sharpness of a boundary in a cross-sectional image, corresponding to the spatial rate of change of the structural boundary. The definition is as follows: ; in, This represents the first-order partial derivative of the cross-sectional image, i.e., the gradient of the cross-sectional image in the horizontal direction. It is the first-order partial derivative of the two-dimensional cross-sectional image, that is, the gradient of the cross-sectional image in the vertical direction; Structural regularity is measured by the consistency between a local area of a cross-sectional image and a standard structure. Local structural similarity index, degree of structural regularity The definition is as follows: ; in, I For cross-sectional images, The mean map of the blurred cross-sectional image provides a structural reference for the cross-sectional image. In order to be in Calculate local structural similarity; S2123. Construction of the Filling Scoring Function: To standardize the measurement of structural ambiguity in the medium region, a filling scoring function is defined. as follows: ; in, These are weighting coefficients. Preferred , , This indicates that the eigenvalues are normalized to... Interval, fill scoring function The larger the value of , the more blurred the medium region is, the higher the structural uncertainty, and the greater the potential medium filling probability.
4. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 3, characterized in that, Step S2200 includes: S2210. Define the curvature edge supervision auxiliary loss function. as follows: ; in, The supervisory signal in the curvature edge supervision auxiliary loss function is Gaussian smoothed on the target boundary in the high medium region to remove noise, and a logical AND is performed with the real boundary region to retain the reliable region and generate the supervisory signal. S2220. Calculate the first and second derivatives of the cross-sectional image. The partial derivative is calculated using the following formula: 、 、 、 、 ; in: , It is a two-dimensional cross-sectional image. The first-order partial derivative; , , It is a cross-sectional image. The second-order partial derivative; S2230. Calculate the curvature diagram, using a discrete algorithm for discrete derivation, and calculate the local curvature. The formula is: ; in, It is a very small positive number. Preferred To avoid a denominator of 0; Ensure the normalization of the curvature scale; S2240, Curvature Edge Map Definition, Setting Curvature Intensity Threshold From the curvature diagram Extract edge information and define the curvature edge map. The formula is as follows: ; in, Curvature intensity threshold, range The preferred value is 0.
1. This represents a local maximum of curvature; S2250, Generate an enhanced boundary prediction map, and combine the medium probability map... With the curvature edge map The enhanced boundary prediction map is generated by fusion. The formula is as follows: ; in, It is the fusion weighting coefficient. .
5. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 4, characterized in that, Step S2300 includes: S2310. Semantic feature extraction: Extracting semantic features. F , means as follows: , in, This refers to the number of channels, i.e., the number of semantic categories. S2320, Channel-weighted feature map calculation, for the semantic features F Global average pooling is used to extract channel descriptions, channel attention weights are predicted, and the channel weighted feature maps are calculated. ; S2330, Spatial augmentation semantic feature map calculation, for the semantic features F Channel compression is performed, followed by average pooling and max pooling along the channel dimension to obtain the average pooled feature map. and max pooling feature map splicing the average pooling feature map and the max pooling feature map Then perform convolution to generate a spatial attention map. The spatial attention map Broadcast back channel dimension, and with the semantic features F Dot product yields the spatially enhanced semantic feature map. ; S2340, Perceptual feature calculation, connecting the channel weighted feature map through residuals. and the spatially enhanced semantic feature map The perceived features are obtained. The formula is as follows: ; S2350, Re-predict the perceived features. Perform further segmentation and prediction using the following formula: ; in, , This indicates a lightweight prediction head that enables fast prediction in low-medium regions without the need for large-scale upsampling and complex skip connections. S2360. Construct the perceptual supervised loss function, calculate the auxiliary loss of the attention mechanism in the low-media region, i.e., the perceptual supervised loss, the perceptual supervised loss function. The definition is as follows: ; in, It is a low dielectric region. It is a perceptual feature. It is the result of re-predicting based on perceptual features. These are the corresponding real tags. This is the standard segmentation loss function.
6. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 5, characterized in that, Step S2320 includes: S2321. Channel description vector calculation, for the semantic features F Global average pooling is used to extract channel descriptions. For each channel, global average pooling (GAP) is performed to obtain the average description vector for a single channel. The formula is as follows: ; in, , C The total number of channels. The position on the c-th channel is The feature map values of the pixels, and These represent the height and width of the feature map, respectively. i , j These are the row and column indices of the pixels being traversed, respectively. For the first c A single-channel average description vector; The set of all individual channel description vectors is the channel description vector, as shown in the following formula: ; S2322, Channel attention weight prediction: The channel attention weights are predicted using a multilayer perceptron (MLP) in the impacted tooth and neural tube segmentation model, and the channel description vectors are... z Input two layers of MLP to generate the channel attention weights. The attention weight of the c-th channel The formula is as follows: , in, , These are the weight matrices of two fully connected layers in a channel attention MLP. It is the ReLU activation function; It's the Sigmoid function, which normalizes the weights to... ; S2323, Channel weighted feature map calculation, by broadcasting the channel attention weights. and with the semantic features F Dot product yields the channel-weighted feature map. : 。 7. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 6, characterized in that, Step S2330 includes: S2331, Channel compression, for the semantic features F Channel compression is performed, followed by average pooling and max pooling along the channel dimension, resulting in two... The feature maps are respectively the average pooling feature maps. and the max pooling feature map The formula is as follows: ; ; S2332. Concatenate the pooling features and perform convolution, then concatenate the average pooling feature map. and the max pooling feature map , obtain the size splicing pooling feature map and through a Convolution generates the spatial attention map. , The formula is as follows: ; S2333, Spatial augmentation semantic feature map calculation, and the spatial attention map Broadcast back channel dimension, and with the semantic features F Dot product yields the spatially enhanced semantic feature map. The formula is as follows: 。 8. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 7, characterized in that, Step S2400 includes: S2410. Feature fusion: The enhanced boundary prediction map, the channel-weighted feature map, and the spatially enhanced semantic feature map are fused using the following formula: ; in, Indicates feature fusion; S2420, Practical threshold binarization, will fuse features Practical threshold Binarization is performed using the following formula: ; in, Indicates an indicator function that satisfies the condition. It is 1 if it is true, otherwise it is 0; The optimal selection was determined through three-fold cross-validation. ; Represents a binary mask; S2430, Boundary Extraction: Boundaries are obtained using the Canny edge detection algorithm. ; in, The initial contour obtained is the edge detection function. That is, the predicted profile of the medium region.
9. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 7, characterized in that, Step S2500 includes: S2510. Initial contour homogeneous representation: The initial contour is represented homogeneously using the following formula: ; in, For edge detection functions, Represents a binary mask; S2520. Correct the contour definition: Perform geometric correction on the initial contour to obtain the corrected contour, wherein the corrected contour is defined as follows: ; in, As the initial outline, This is the geometric correction transformation matrix; S2530. Definition of geometric correction loss function, wherein the geometric correction loss function is defined as follows: ; in, Represents coordinate mapping, Indicates the first i The coordinates of the boundary points in a homogeneous coordinate system. This represents the coordinates of the boundary points in the homogeneous coordinate system after geometric correction. Represents the coordinates of the true boundary points in a homogeneous coordinate system. This represents the prediction confidence level of the i-th boundary point. A value close to 1 indicates that the boundary point is located in a high-medium region, which is preferred. =Above 0.85 is the high dielectric region; S2540, Definition of the overall loss function, the overall loss function The formula is as follows: ; in, These are the gradient regularization terms. Weighting coefficients, CML connection axis loss function Weighting coefficients, MBO boundary offset loss function The weighting coefficients of the CMO spatial orientation loss function Weighting coefficients of the MCC medium probability graph loss function Curvature edge supervision auxiliary loss function Weighting coefficients, perceptual supervision loss function Weighting coefficients, geometric correction loss function Weighting coefficients; S2550, the loss function for impacted tooth and neural tube segmentation is defined, integrating the high-medium region and the low-medium region, and performing unified feature fusion with the geometric correction. The segmentation result of impacted tooth and neural tube is output through iterative optimization. The loss function for impacted tooth and neural tube segmentation is expressed as: ; in, These are the weights of the overall loss function and the geometric correction loss function, respectively.
10. The method for segmenting impacted teeth and nerve tubes based on a media region as described in claim 9, characterized in that, Step S4000 includes: S4100. Continue training begins. When the number of training rounds is greater than or equal to the maximum number of training rounds, the temporary impacted tooth and neural tube segmentation model is obtained. The weights of each loss function in the overall loss function are modified according to the loss function adjustment strategy. The number of training rounds is initialized to 0. S4200, Continue training, continue training the temporary impacted tooth and neural tube segmentation model, and increase the number of training rounds by one; S4300: Determine the number of training epochs and performance threshold. When the number of training epochs is greater than or equal to the maximum number of training epochs, modify the weights of each loss function in the overall loss function according to the loss function adjustment strategy, set the number of training epochs to 0, and return to step S4200. When the number of training epochs is less than the maximum number of training epochs and meets the performance threshold, execute step S4400. When the number of training epochs is less than the maximum number of training epochs and does not meet the performance threshold, execute step S4200. S4400, Continue training to complete, and obtain the trained impacted tooth and neural tube segmentation model.
Citation Information
Patent Citations
Oral cavity CBCT image tooth and soft tissue segmentation model method based on improved U-Net model
CN117115132A
Method for segmenting liver and focus thereof in medical image
WO2021184817A1