A panoramic interactive tooth segmentation method based on feature enhancement
The feature-enhanced panoramic interactive dental segmentation method addresses user-dependent and morphology-inconsistent issues by using foreground and background seeds and multi-level evaluation, enhancing interaction robustness and accuracy.
Patent Information
- Application Number
- CN202411512295.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing interactive tooth segmentation method requires repeated observation perspective on three-dimensional complex objects. The user's subjectivity has a great influence, the interaction lacks robustness, the unified range of weakening optimization effects, the foreground seed point information is attenuated, the self-evaluation module is not accurate enough, and the background seed point characteristics are not fully utilized, resulting in segmentation errors.
The panoramic interactive tooth segmentation method based on feature enhancement is adopted to interactively mark seed points through two-dimensional panoramic maps, and the foreground and background seed points characteristics are pre-enhanced in advance, combined with the Transformer model and 3DUNet for segmentation, multi-level evaluation and optimization, and adaptively adjust the seed points position and range.
It improves the robustness and accuracy of interaction, reduces user perspective switching, makes full use of foreground and background features, enhances the effectiveness of segmentation results and the accuracy of evaluation, and improves the overall effect of tooth segmentation.
Smart Images

Figure CN119445110B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image analysis and processing in computer vision, and particularly relates to a panoramic interactive tooth segmentation method based on feature enhancement. Background Art
[0002] The tooth segmentation task refers to segmenting all teeth based on a given one or more teeth using a segmentation algorithm, which belongs to the instance segmentation task in computer vision. The interactive tooth segmentation method based on cone beam computed tomography (CBCT) images effectively integrates the individual priors of the target object by designing convenient interaction strategies, depends less on the scale of training data than fully automatic segmentation, has higher segmentation efficiency and more accurate results, can extract the precise tooth morphology from the three-dimensional CT raw data, and the results can be used for quantitative analysis of clinical key indicators and visualization of the actual invisible internal morphology.
[0003] Currently, the interactive segmentation field mainly attempts to develop from two aspects: interaction strategies and effective interaction propagation. In the prior art, Lin et al. [1] proposed the first click attention network, which combines the guiding information of the first click with the obtained features to reduce the number of interactions and improve the interaction efficiency. Song et al. [2] proposed a bidirectional seed attention network. By adding a bidirectional seed attention module to the backbone segmentation network to enhance the semantic information of the seeds, the weakening of seed information is avoided to obtain better segmentation results. Zhang et al. [3] proposed a method of training patch samples by embedding a gated memory propagation unit ConvRNN network and capturing the spatial relationship between adjacent patches to achieve the fusion of seed point information and the image. Sakinis [4] et al. proposed a semi-automatic segmentation method of convolutional neural network, which is trained on a relatively limited training set and can produce fast and accurate segmentation outputs on CT images. Luo et al. [5] proposed a segmented MIDeepSeg framework based on minimum interaction deep learning, which encodes user interactions using exponential geodesic distance transformation without additional parameters and thresholds. Zhou et al. [6] proposed a new volume memory network called VMN, which realizes interaction on two-dimensional slices and can adjust the three-dimensional segmentation results, and proposed an interaction self-evaluation method to assist multi-round interactions.
[0004] Lin Z, Zhang Z, Chen L Z, et al. Interactive image segmentation with first click attention[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020: 13339-13348.
[0005] Song G, Lee K M. Bi-directional seed attention network for interactive image segmentation[J]. IEEE Signal Processing Letters, 2020, 27: 1540-1544.
[0006] Zhang J, Shi Y, Sun J, et al. Interactive medical image segmentation via a point-based interaction[J]. Artificial Intelligence in Medicine, 2021, 111: 101998.
[0007] Sakinis T, Milletari F, Roth H, et al. Interactive segmentation of medical images through fully convolutional neural networks[J]. arxiv preprint arxiv:1903.08205, 2019.
[0008] Luo X, Wang G, Song T, et al. MIDeepSeg: Minimally interactive segmentation of unseen objects from medical images using deep learning[J]. Medical image analysis, 2021, 72: 102102.
[0009] Zhou T, Li L, Bredell G, et al. Volumetric memory network for interactive medical image segmentation[J]. Medical Image Analysis, 2023, 83: 102599.
[0010] Currently, the following disadvantages and problems exist in interactive segmentation methods:
[0011] (1) In the existing technologies [1, 2, 3], the interaction only adopts the click strategy. For three-dimensional complex objects, it is necessary to repeatedly observe three standard viewpoints to determine the position of the seed points, which is greatly affected by user subjectivity and lacks robustness in interaction.
[0012] (2) In the existing technologies [3, 4], by training global or uniformly sized square patches, the scope of adjustment is restricted. However, the overall shape of the teeth is not uniformly symmetric, so applying a unified scope will weaken the optimization effect after interaction to a certain extent.
[0013] (3) In the existing technology [5], only foreground seed points are focused on during the initial interaction, while the characteristics of background seed points are ignored, and the semantic association between multiple types of seed points is not fully utilized.
[0014] (4) In the existing technology [6], the information of interactive seed points is directly used for segmentation. As local priors, the seed points inevitably decay through network learning, resulting in low guidance efficiency. The self-evaluation module has a single input and completely relies on the features of the segmentation network itself, which will lead to inaccurate evaluation results, and the biased interaction feedback will ultimately lead to segmentation errors.
[0015] Due to the problems faced by the interactive segmentation method in the existing technology, it is necessary to propose a panoramic interactive tooth segmentation method based on feature enhancement to solve the above problems. It can intelligently identify and correct certain deviation interactions, quantitatively evaluate the interaction quality, and improve the effectiveness of the interaction results in deep model learning. So as to be better applicable to application scenarios that require precise interaction, such as medical image analysis, human-computer interaction, and augmented reality. Summary of the Invention
[0016] The present invention designs a panoramic interactive tooth segmentation method based on feature enhancement, which can automatically identify and correct user interactions, provide convenient interaction viewpoints, and improve the effectiveness of interactive prior propagation by pre-enhancing features before segmentation. This method takes foreground seeds and background seeds as two independent priors, learns in parallel, and pre-enhances before the segmentation network to achieve accurate interactive segmentation results for teeth.
[0017] A panoramic interactive tooth segmentation method based on feature enhancement includes:
[0018] Step 1: On the basis of the three-dimensional image of the original cone-beam CT image, reconstruct a two-dimensional panoramic image, and mark the two-dimensional foreground seed point a, the two-dimensional background seed point b, and the bounding box c containing a single tooth according to user interaction; calculate the three-dimensional initial seed point and the single-tooth region of interest, perform semantic discrimination on the three-dimensional initial seed point, and determine whether to adjust the position of the three-dimensional initial seed point according to the discrimination result to generate the final three-dimensional foreground seed point a2 and three-dimensional background seed point b2;
[0019] Step 2: Combine the three-dimensional foreground seed point a2 and the three-dimensional background seed point b2 with the single-tooth cone-beam CT image T respectively according to the exponential geodesic distance method to generate the corresponding foreground distance map P F and background distance map P B ; The foreground distance map P F and the background distance map P B are respectively input into the Transformer model for encoding together with the single-tooth cone-beam CT image T, and the in-layer feature Fp and the inter-layer feature F d are calculated and generated. The in-layer feature Fp and the inter-layer feature F d are fused to generate the foreground enhanced distance map F(P F ) and the background enhanced distance map F(P B ), and 3DUNet is used for image segmentation to output the initial single-tooth segmentation result R0;
[0020] Step 3: Multilevel evaluation and optimization, determine the input of the evaluation network and construct the evaluation network, evaluate the initial segmentation result, and generate the evaluation score for each slice; screen the slices with scores lower than the threshold, and the user performs secondary interaction on the slices with scores lower than the threshold to generate secondary seed points; finally, optimize the initial segmentation result according to the secondary seed points.
[0021] Further, the reconstruction and generation of the two-dimensional panoramic image in Step 1 is processed according to formula (1)
[0022] where X is the tooth target image to be reconstructed, s is a point on the projected dental arch curve, Z is the Z-th slice in the CT image, P(S,Z) represents the result of the Z-th slice projected according to the point S on the dental arch curve, r(s) is the coordinate of point S, n(s) is the normal vector of point S, (α, -α) is the projection range along the normal vector, and t is the projection variable.
[0023] Further, the calculation of the three-dimensional initial seed points and the single-tooth region of interest in step 1 includes inverse-projecting the two-dimensional foreground seed points a, the two-dimensional background seed points b, and the bounding box c according to formula (1) to generate the three-dimensional foreground initial seed points a1, the three-dimensional background initial seed points b1, and the three-dimensional bounding box c1; calculating the circumscribed cuboid c2 of the three-dimensional bounding box c1 to form a three-dimensional single-tooth region of interest, and cropping the original cone-beam CT image according to the c2 region to generate a single-tooth cone-beam CT image T composed of multiple slices of the same size.
[0024] Further, the semantic discrimination of the three-dimensional initial seed points in step 1 includes discriminating whether the three-dimensional initial seed points are located in the crown region or the root region of the tooth; then identifying the seed semantics; the identification of the seed semantics further includes that if the three-dimensional initial seed points are in the root, as long as the seed position is within the tooth foreground range, it is identified as a foreground point, otherwise it is a background point; if the three-dimensional initial seed points are in the crown, the position coordinates of the three-dimensional initial seed points are vertically projected, and it is discriminated as a foreground point or a background point according to the projection position.
[0025] Further, the discrimination of whether the three-dimensional initial seeds are located in the crown region or the root region of the tooth further includes discriminating whether the seed points are in the root or crown region through the variance statistics of the tooth foreground areas of consecutive slices; the discrimination of whether it is a foreground point or a background point according to the projection position further includes that if the projection position is within the foreground with the largest area in the slice, it is a foreground point, otherwise it is a background point.
[0026] Further, the adjustment of the positions of the three-dimensional initial seed points in step 1 includes adjusting according to the semantics specified during user interaction. If the user marks it as a foreground point and the semantic discrimination is a background point, the seed point is moved to the centroid of the foreground; if the user marks it as a background point and the semantic discrimination is a foreground point, the seed point is adjusted to the background area and there is a certain distance from the foreground boundary.
[0027] Further, the calculation of the in-layer feature Fp and the inter-layer feature F in step 2 d further includes, according to the foreground distance map P F the encoded K feature and V feature, the encoded Q feature of the image T, and the background distance map P B the encoded K feature and V feature, and the encoded Q feature of the image T, calculating the in-layer pixel relationship to generate the in-layer feature Fp, and calculating the inter-layer pixel relationship to generate the inter-layer feature F d ; the further use of 3DUNet for image segmentation includes using the foreground enhanced distance map F(P F ), the background enhanced distance map F(P B ), and the single-tooth CBCT image T as the input of 3DUNet for image segmentation, and the output result is the initial segmentation result R0.
[0028] Further, the in-layer feature Fp in the calculation layer further includes the foreground distance map P F The encoded K feature and the encoded Q feature of the image T are subjected to an inner product operation. After softmax normalization, they are weighted and summed with the V feature to obtain the foreground distance map P F of the in-layer feature Fp; the background distance map P B The encoded K feature and the encoded Q feature of the image T are subjected to an inner product operation. After softmax normalization, they are weighted and summed with the encoded V feature of the background distance map P B to obtain the in-layer feature of the background distance map P B .
[0029] The calculation layer-to-layer pixel relationship generation layer-to-layer feature F d further includes the foreground distance map P F The encoded K feature, the encoded Q feature of the image T, and the background distance map P B After the encoded K feature is subjected to matrix permutation, the K features generated by the foreground distance map and the background distance map after matrix conversion are respectively subjected to inner product operation and softmax normalization with the Q feature, and then are respectively weighted and summed with the V features generated by the foreground distance map and the background distance map to obtain the layer-to-layer feature F d .
[0030] Further, in step 3, the determination of the evaluation network input and the construction of the evaluation network further include unifying the sizes of the initial segmentation result R0, the foreground enhanced distance map F(PF), the background enhanced distance map F(PB), and the deepest layer feature F in the 3DUNet model and connecting them by channels as the evaluation input; constructing the evaluation network through a convolutional layer, the fourth layer structure of ResNet50, an average pooling layer, and a fully connected layer. s
[0031] Further, in step 2, the fusion of the in-layer feature Fp and the layer-to-layer feature F d further includes adding the in-layer feature Fp and the layer-to-layer feature F d to obtain F0, upsampling F0 to obtain F1 with a scale of 1×D×H×W, and finally connecting the foreground distance map P F and the background distance map P B to F1 by channels respectively to generate the foreground enhanced distance map F(P F ), and the background enhanced distance map F(P B ), where D is the depth of the original image, H is the height of the original image, and W is the width of the original image.
[0032] The effects of the present invention are as follows:
[0033] 1. The interaction of this method is implemented on a two-dimensional panoramic view with a unified perspective. The panoramic view is a reconstructed two-dimensional map, which serves as the object directly operated by the user and is limited to the user interaction stage. However, the object of feature enhancement is three-dimensional single-tooth data, which is obtained by intelligently identifying and acting on the original data based on the user interaction result. This unified perspective can comprehensively reflect the spatial position of three-dimensional objects without the need to switch perspectives. At the same time, the seed point generation method based on morphological continuity proposed in this method can identify and correct possibly incorrect user interaction points, improve the effectiveness of the interaction result, and thus improve the robustness of user interaction.
[0034] 2. This method proposes an adaptive optimization method, which sets different optimization ranges according to the different tooth positions where the seeds are located, replacing the traditional setting with a unified optimization range.
[0035] 3. The interaction information input in this method is richer, including both foreground seed points and background seed points. The foreground enhanced distance map and the background enhanced distance map are used as two independent sources and input into the segmentation network for processing. Compared with the traditional segmentation method that only takes the foreground seeds as the interaction source, this method can ensure the semantic prior of the learning seeds themselves while fully learning the difference features between the background and the foreground.
[0036] 4. Feature pre-enhancement processing is added before segmentation in this method. Taking the foreground distance map and the background distance map as objects, enhanced features are extracted to form a more significant enhanced distance map for interaction (including foreground and background). Pre-enhancing the traditional distance map before segmentation has a more global guiding role than directly using the distance map for segmentation. At the same time, the foreground enhanced distance map and the background enhanced distance map will be reused as the input of the subsequent evaluation network. Compared with the traditional method that only uses the deep features of the segmentation network, because the added enhanced distance map directly comes from user specification, it has better supervision, so the evaluation is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the logic flowchart of the interactive segmentation method of the present invention;
[0038] Figure 2 is the schematic diagram of self-identifying interactive processing of the panoramic view;
[0039] Figure 3 is the single-tooth segmentation process based on feature pre-enhancement;
[0040] Figure 4 is the flowchart of feature pre-enhancement processing;
[0041] Figure 5 is the flowchart of multi-level evaluation and optimization. DETAILED DESCRIPTION OF THE INVENTION
[0042] The technical solution of the present invention will be specifically described below in conjunction with embodiments.
[0043] This solution consists of three parts: 1) The interactive processing step of panoramic image self-identification, which identifies and generates effective single-tooth interest regions and seed points; 2) The single-tooth segmentation step with feature pre-enhancement to obtain an initial segmentation result; 3) The multi-level evaluation and optimization step. The segmentation result will be evaluated slice by slice to obtain an evaluation coefficient, and the pictures with coefficients lower than the threshold will be selected and interacted with again. After interaction, the seed propagation range is partitioned by tooth, and different ranges are applied. Multiple rounds of interaction can be performed until a satisfactory result is obtained, which is used as the final segmentation result. The three parts are executed in sequence, and the overall process is as Figure 1 shown.
[0044] Step 1: The interactive processing step of panoramic image self-identification. Based on the three-dimensional image of the original cone beam CT image (CBCT), a two-dimensional panoramic image is reconstructed. According to user interaction, the two-dimensional foreground seed point a, the two-dimensional background seed point b, and the bounding box c containing a single tooth are marked; the three-dimensional initial seed point and the single-tooth interest region are calculated, and semantic discrimination is performed on the three-dimensional initial seed point, and whether to adjust the position of the initial seed point is determined according to the discrimination result.
[0045] A simpler and more intuitive two-dimensional interaction method is adopted, using an interaction mode combining points and bounding boxes. To ensure the objectivity of the interaction, in this solution, the interaction is not performed on the original image slices, but on the reconstructed image in the panoramic view, that is, the panoramic image. This interaction separates the original data into single-tooth interest regions and obtains several foreground or background seed points. There are two key parts in this processing process, namely 1) Reconstructing a two-dimensional enhanced panoramic image (corresponding to step 101); 2) The seed generation method based on morphological continuity (corresponding to steps 104-105). The specific interaction processing of step 1 is as Figure 2 shown, and specifically includes the following steps:
[0046] Step 101: Take the three-dimensional image of a group of original cone beam CT images (CBCT) as the input. A group of images represents the complete jaw data of a patient, which usually consists of 300-400 slices. After threshold segmentation of this image to remove the background region, a two-dimensional panoramic image is reconstructed, and the reconstruction formula is as shown in formula (1). Where X is the tooth target image to be reconstructed, s is a certain point on the projected dental arch curve, Z is the Z-th slice in the CT image, P(S,Z) represents the result of the Z-th slice projected according to the point S on the dental arch curve, r(s) is the coordinate of point S, n(s) is the normal vector of point S, (α, -α) is the projection range along the normal vector, and t is the projection variable.
[0047]
[0048] In step 101 of the present invention, a unified panoramic view is provided for the user to interact, enabling the user to determine whether the interaction position follows their intention without switching among three standard perspectives, thereby making the interaction more convenient.
[0049] Step 102: The user interacts with the panoramic image P generated in step 101, marking the two-dimensional foreground seed point a of the tooth, the two-dimensional background seed point b, and the bounding box c containing a single tooth, as Figure 2 shown.
[0050] Step 103: Calculate the three-dimensional initial seed points and the single-tooth region of interest. The two-dimensional foreground seed point a, the two-dimensional background seed point b, and the bounding box c are back-projected (using a general projection method) according to formula (1) to form the three-dimensional foreground initial seed point a1, the three-dimensional background initial seed point b1, and the three-dimensional bounding box c1; then calculate the circumscribed cuboid c2 of c1 to form the three-dimensional single-tooth region of interest, that is, the region of interest that minimally covers a single tooth. According to the c2 region, the original CBCT image is cropped to obtain the single-tooth CBCT image T, which consists of multiple slices of the same size.
[0051] Step 104: Semantic discrimination of the three-dimensional initial seed points. The semantic discrimination is divided into two steps. First, determine whether the three-dimensional initial seeds (including a1 and b1) are located in the crown region or the root region of the tooth; then identify the seed semantics. For the discrimination of the seed position region, first, taking the root as the starting point and the crown as the ending point direction as the direction of the slice access order in the single-tooth CBCT image T, calculate the area variance of the tooth blocks in each slice. If the variance is greater than 50 and the slice number is greater than two-thirds of the total number of slices in the single-tooth region of interest, it is determined that the three-dimensional initial seed is located in the crown, otherwise it is determined to be in the root. Then, identify the semantics of the three-dimensional initial seed points. If the initial seed point is in the root, as long as the seed position is within the tooth foreground range, it is identified as a foreground point, otherwise it is a background point; if the initial seed point is in the crown, taking the slice where the initial seed is located as the center, symmetrically take 5 adjacent slices, a total of 11 slices, and project the position coordinates of the three-dimensional seed points vertically onto the 11 slices. If more than 6 slices have the projection position within the largest foreground area in the slice, it is a foreground point, otherwise it is a background point. Through the above operations, the three-dimensional initial seeds (including a1 and b1) obtain the corresponding semantic discrimination results regarding background or foreground meaning.
[0052] Step 105: Semantic adjustment of 3D initial seed points. If the semantic discrimination result of the 3D initial seed points (including a1 and b1) is consistent with the specified semantics during user interaction, no adjustment is made; if not, the position of the initial seed points is adjusted. The adjustment methods are divided into the following situations: 1) If the semantic discrimination is a background point and the user designates the initial seed as a foreground point for adjustment. The initial seed position is reset to the centroid position of the tooth foreground, and the adjusted seed semantics are foreground points. 2) If the recognition is a foreground point and the user designates the seed as a background point for adjustment. The seed position is reset to a position within the background range that is at least 20 pixels away from the foreground edge, and the adjusted seed semantics are background points. The final foreground seed points a2 and background seed points b2 are obtained through the above steps.
[0053] During the processing of Step 1, the user interacts on the panoramic image without having to confirm the interaction position by switching among the three perspective cross-sections, namely the transverse plane, the coronal plane, and the sagittal plane, on the original slice. This ensures that the user operation view is unified and comprehensive, eliminating the need for additional perspective switching and greatly improving the convenience and efficiency of interaction. At the same time, to ensure a clearer panorama, this solution does not perform reconstruction on the original image but rather on the tooth target image after threshold segmentation of the original data and removal of the background area.
[0054] The present invention proposes a method for generating seeds with continuous morphology in Steps 104 - 105 to confirm whether the recognized semantics of the seed points are consistent with the user-marked semantics. If not, the adjustment is made according to the user semantics based on the centroid of the foreground area. This reduces the accuracy of user interaction. Even if the user interaction is incorrect, it can be recognized and corrected, thereby improving the robustness of user interaction.
[0055] The seed point generation method based on morphological continuity is mainly used to identify and generate effective seed points. In this application, the interactive two-dimensional interaction points need to be converted into three-dimensional seed points. Through reverse projection using formula (1), a two-dimensional interaction point corresponds to the center point of a three-dimensional region. This type of direct mapping may lead to semantic errors in seed points due to spatial information compression and inaccurate user interaction. For example, for a foreground point marked by the user, the initial seed point after direct mapping may be a background point, and vice versa. Therefore, a seed point generation method based on morphological continuity is proposed to identify the semantic attribution of the initial seed point and adjust it if the semantic attribution is incorrect. This method combines the differences between the tooth root and the tooth crown for adaptive recognition, specifically including: 1) By statistically analyzing the variance of the foreground area of continuous tooth slices, it is determined whether the seed point is in the tooth root or tooth crown area. If the variance is less than 50, it is the tooth root; if the variance is greater than 50, it is the tooth crown. 2) If the tooth root seed point is within the foreground range, it is a foreground point; otherwise, it is a background point. If the tooth crown seed point is in the largest foreground area, it is a foreground point; otherwise, it is a background point. 3) According to the semantic adjustment specified during user interaction, if the user marks a foreground point but the semantic discrimination is a background point, the seed is moved to the centroid of the foreground. If the user marks a background point but the semantic discrimination is a foreground point, the seed point is adjusted to the background area and is at a certain distance from the foreground boundary.
[0056] Step 2: The single-tooth segmentation step with feature pre-enhancement, using the foreground points and background points as two independent inputs to strengthen the difference between the foreground and the background. This method mainly consists of three parts: calculating the foreground indexed geodesic distance and the background indexed geodesic distance based on the foreground points and background points respectively; then, using the feature pre-enhancement module to extract enhanced features; finally, implementing single-tooth segmentation with the enhanced features as the input. In this way, global image enhancement is pre-implemented in the early stage of segmentation to prevent the problem that seed points are prone to attenuation in the network. An enhanced distance map is generated before the segmentation network, and the background enhanced distance map and the foreground enhanced distance map are output as two independent streams of the feature pre-enhancement module. The specific process is as follows Figure 3 shown, and the specific implementation is as follows.
[0057] Step 201: Calculate the distance map to obtain the foreground distance map P F and the background distance map P B . According to the indexed geodesic distance method (classical method), the foreground seed point a2 and the background seed point b2 are respectively combined with the single-tooth CBCT image T to generate the corresponding foreground distance map P F and the background distance map P B , and their specific calculation formulas are shown in formulas (2) and (3). Among them, S s represents the final seed point set (including a2 and b2), i is a single voxel in the single-tooth CBCT image T, and Ρ i,jDenote the set of all paths between pixels i and j. p is a feasible path, which is determined by a parameter n ∈ [0, 1]. u(n) = p(n)'p(n)' is the unit vector tangent to the path direction. EGD(i, Ss, T) represents the exponential geodesic distance result from point i to the seed point S in image T s in the image T, and D geo (i, j, T) represents the geodesic distance from point i to point j in image T.
[0058]
[0059] Step 202: Feature pre-enhancement processing, including two parallel processes, namely obtaining the foreground enhancement feature F F and the background enhancement feature F B . The two processing model structures are exactly the same, and the only difference is the input source. Specifically, to obtain the foreground enhancement feature F F , the foreground distance map P F and the single-tooth CBCT image T are used as inputs; while to obtain the background enhancement feature F B , the background distance map P B and the single-tooth CBCT image T are used as inputs. Each parallel process is specifically as Figure 4 shown. First, encode the input source with an encoder (Step 2021), then calculate the intra-layer features and inter-layer features respectively (Steps 2022 and 2023), and finally fuse the intra-layer features and inter-layer features (Step 2024). The interactive prior information is diffused from local points to global data through intra-layer and inter-layer.
[0060] Step 2021: Encode the input source. Taking the extraction of the foreground enhancement feature F(P F ) as an example, its encoding details are as follows. Input the foreground distance map P F and the single-tooth CBCT image T, both with a scale of 1×D×H×W, and use the Transformer model to encode them respectively. After encoding the image T, the Q feature with a size of is obtained; after encoding the foreground distance map P F , the K feature with a size of and the V feature with a size of are obtained; the Q feature, K feature, and V feature are all standard parameter names in the Transformer model. The K feature and V feature have the same scale but different values. Where C is the number of image channels, D is the image depth, H is the image height, W is the image width, are 1 / 8 of the original image sizes D, H, and W respectively. Taking the single-tooth CBCT image as the center, search for similarities voxel by voxel in the foreground or background distance to learn the weights as Figure 4 shown.
[0061] Step 2022: Calculate the in-layer pixel relationship, and convert the K feature and the Q feature into the in-layer feature F p First, perform an inner product operation on the K feature and the Q feature as shown in formula (4), then use the softmax normalization process as shown in formula (5), and finally perform weighted summation as shown in formula (6) to obtain the in-layer feature F p :
[0062]
[0063] where M(i, j) represents the foreground distance map P F the relationship between the position i on the foreground distance map P and the position j on the original image feature map. Next, perform Softmax normalization to unify the values between 0 and 1, which is expressed as:
[0064]
[0065] Perform weighted summation with the V feature through the weight value W(i, j), which is expressed as:
[0066]
[0067] Finally, form the in-layer feature Fp of the foreground distance map P F For the calculation of the in-layer feature Fp of the background distance map P B it is completed by repeating Step 2021 to Step 2022. When repeating, only need to replace the input foreground distance map P F with the background distance map P B that's it.
[0068] Step 2023: Calculate the inter-layer pixel relationship, that is, convert the K feature and the Q feature into the inter-layer feature
[0069] F d First, the K feature and the Q feature are resized from the original size (C, are the number of channels, depth, height, and width of K or Q respectively), adjusted to This operation can be obtained through simple matrix permutation. After that, through inner product operation, Softmax normalization, and weighted summation as shown in formulas (4), (5), and (6), the inter-layer feature F of the foreground distance map P F is obtained d For the background distance map P B the inter-layer feature F d is extracted by repeating Step 2021 and Step 2023. When repeating, only need to replace the input foreground distance map P F with the background distance map P B that's it.
[0070] Step 2024: Fuse the intra-layer features and inter-layer features. First, add the intra-layer feature F p and the inter-layer feature F d to obtain F0, then upsample F0 to obtain F1 with a scale of 1×D×H×W. Finally, combine the foreground distance map P F and the background distance map P B , and concatenate them with F1 channel-wise to obtain the foreground enhanced distance map F(P F ), and the background enhanced distance map F(P B ).
[0071] In step 202 of the present invention, feature pre-enhancement processing is used to achieve interactive feature enhancement. Specifically, by obtaining the highly similar part between the original CBCT image and the distance map, and using the addition method to strengthen the feature values on the distance map, more significant information is provided for the segmentation input. At the same time, enhancing before segmentation can prevent the problem that the interactive information represented by the distance map is likely to be weakened in the deep network.
[0072] Step 203: Single tooth segmentation. The foreground enhanced distance map F(P F ), the background enhanced distance map F(P B ), and the single tooth CBCT image T are used as the inputs of the 3D UNet for image segmentation, and the output result is the initial segmentation result R0.
[0073] Step 3: Multi-level evaluation and optimization step. This step consists of three parts. First, evaluate the score of the initial segmentation result (step 301); then screen the slices with scores lower than the threshold, and the user performs secondary interaction on such slices to generate secondary seed points (step 302); finally, optimize the initial segmentation result according to the secondary seed points (step 303). The specific process is shown in Figure 5 as follows.
[0074] Step 301: Evaluating the initial segmentation result includes two aspects: determining the input of the evaluation network and constructing the evaluation network. First, determine the input source of the evaluation network, and its input composition includes four parts: the initial segmentation result R0, the foreground enhanced distance map F(P F ), the background enhanced distance map F(P B ), and the deepest layer feature F s in the 3D UNet model; then unify the sizes of the four parts and concatenate them channel-wise to form the input of the evaluation network A. The output result of this network is the evaluation score Score for each slice, as shown in formula (7) specifically.
[0075] Score = A(R0, F(P F ), F(P F ), F s ) (7)
[0076] Construct the specific structure Q of the evaluation network. First, there is a 3×3 convolutional layer, followed by the fourth layer structure of ResNet50, which is composed of two 1×1 convolutional layers and a 3×3 convolutional layer repeated three times. Then, there is an average pooling layer, and finally, it consists of two fully connected layers. This network adopts a lightweight model design, uses fewer network layers, and realizes the reduction of the model size.
[0077] In step 301 of the present invention, the input source of the evaluation network is enriched. In addition to the deep features of the segmentation network, the deep features in the feature pre-enhancement module are added. This type of information is generated from the seed prior, which is the direct information after user and semantic recognition, more in line with the real result, more supervised, and thus improves the accuracy of the evaluation.
[0078] Step 302: User secondary interaction. The value of the score for each slice is Score(i)∈[0,1], where i is the slice number, N is the total number of slices in the initial segmentation result R0, and i∈[1,N]. The higher the score, the higher the quality of the slice segmentation, and vice versa. Set the threshold to 0.4. The slices with an evaluation score lower than 0.4 are used as secondary interaction slices, forming a set of secondary interaction slices. The user browses the set of secondary interaction slices and performs secondary interaction to generate several secondary seeds X, which include background seeds and foreground seeds.
[0079] Step 303: Adaptive optimization. This paper uses the IF-GC method proposed by MIDeepSeg. However, the IF-GC method calculates the exponential geodesic distance for seed points, and the calculation range of this geodesic distance is the entire single-tooth region of interest. This unified propagation range is obviously not suitable for objects such as teeth with large differences in different parts, such as the crown being thicker and the root being thinner. Therefore, the present invention improves the seed propagation range of IF-GC to obtain an adaptive optimization method.
[0080] The adaptive optimization is divided into three parts. First, determine the input sources of the IF-GC method: secondary seeds X and the initial segmentation result R0. Then, determine whether the secondary seeds X are in the crown or root part (step 104). If the secondary seeds are in the crown area, with the seed point as the center, the distance calculation range is limited to 0.3 times the length, width, and height of the image T. If the secondary seeds are in the root area, with the seed point as the center, the distance calculation range is set to 0.1 times the length, width, and height of the image T. Finally, use the IF-GC method to process the three input sources to obtain the secondary optimization result R1.
[0081] If the result after optimization is still not satisfactory, steps 301-303 can be repeated multiple times until the segmentation result is adjusted to the best effect.
[0082] The adaptive optimization of the present invention centers around the seed point and creates a cuboid region. If the seed is in the crown part of the tooth, it is set to 0.3 times the length, width, and height of the original image, while if the seed is in the root part, it is 0.1 times the length, width, and height of the original image. Such a setting ensures the propagation range exerted by the seeds in different parts and prevents adverse effects on the correct segmentation regions far from the seed point.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A panoramic interactive tooth segmentation method based on feature enhancement, comprising: Step 1: On the basis of the three-dimensional image of the original cone-beam CT image, reconstruct and generate a two-dimensional panoramic image, and mark the two-dimensional foreground seed point a, the two-dimensional background seed point b, and the bounding box c containing a single tooth according to user interaction; Calculate the three-dimensional initial seed points and the single-tooth region of interest, perform semantic discrimination on the three-dimensional initial seed points, and determine whether to adjust the positions of the three-dimensional initial seed points according to the discrimination results, and generate the final three-dimensional foreground seed point a2 and three-dimensional background seed point b2; Step 2: Combine the three-dimensional foreground seed point a2 and the three-dimensional background seed point b2 with the single-tooth cone beam CT image T respectively according to the exponential geodesic distance method to generate the corresponding foreground distance map P F and the background distance map P B ; The foreground distance map P F and the background distance map P B are simultaneously input into the Transformer model for encoding, and the in-layer feature Fp and the inter-layer feature F d are calculated and generated. The in-layer feature Fp and the inter-layer feature F d are fused to generate the foreground enhanced distance map F(P F ) and the background enhanced distance map F(P B ). 3DUNet is used for image segmentation to output the initial single-tooth segmentation result R0; The in-layer feature Fp and the inter-layer feature F generated by the calculation described in step 2 d further includes, according to the foreground distance map P F the encoded K feature and V feature, the encoded Q feature of the image T, and the background distance map P B calculate the in-layer pixel relationship to generate the in-layer feature Fp based on the encoded K feature and V feature and the encoded Q feature of the image T, and calculate the inter-layer pixel relationship to generate the inter-layer feature F d ; the use of 3D UNet for image segmentation further includes using the foreground enhanced distance map F(P F ), the background enhanced distance map F(P B ), and the single-tooth CBCT image T as the input of the 3D UNet for image segmentation, and the output result is the initial segmentation result R0; the fusion of the in-layer feature Fp and the inter-layer feature F d further includes adding the in-layer feature Fp and the inter-layer feature F d to obtain F0, upsampling F0 to obtain F1 with a scale of 1×D×H×W, and finally using the foreground distance map P F and the background distance map P B , respectively connecting with F1 by channels to generate the foreground enhanced distance map F(P F ), the background enhanced distance map F(P B ), where D is the depth of the original image, H is the height of the original image, and W is the width of the original image; Step 3: Multi-level evaluation and optimization, determine the input of the evaluation network and construct the evaluation network, evaluate the initial segmentation result, and generate an evaluation score for each slice; screen the slices with scores lower than the threshold, and the user performs secondary interaction on the slices with scores lower than the threshold to generate secondary seed points; finally, optimize the initial segmentation result according to the secondary seed points.
2. The panoramic interactive tooth segmentation method based on feature enhancement according to claim 1, wherein the reconstruction and generation of the two-dimensional panoramic image in step 1 is based on formula (1) Processing, where X is the target image of the tooth to be reconstructed, s is a point on the projected dental arch curve, Z is the Z-th slice in the CT image, P(S, Z) represents the result of projecting the Z-th slice according to the point S on the dental arch curve, r(s) is the coordinate of point S, n(s) is the normal vector of point S, (α, -α) is the projection range along the normal vector, and t is the projection variable.
3. The panoramic interactive tooth segmentation method based on feature enhancement according to claim 2, wherein the calculation of the three-dimensional initial seed points and the single-tooth region of interest in step 1 includes inverse-projecting the two-dimensional foreground seed point a, the two-dimensional background seed point b, and the bounding box c according to formula (1) to generate the three-dimensional foreground initial seed point a1, the three-dimensional background initial seed point b1, and the three-dimensional bounding box c1; calculate the circumscribed cuboid c2 of the three-dimensional bounding box c1 to form a three-dimensional single-tooth region of interest, and crop the original cone-beam CT image according to the c2 region to generate a single-tooth cone-beam CT image T composed of multiple slices of the same size.
4. The panoramic interactive tooth segmentation method based on feature enhancement as described in claim 1, wherein the semantic discrimination of the three-dimensional initial seed points in step 1 includes discriminating whether the three-dimensional initial seed points are located in the crown area or the root area of the tooth; Then identify the seed semantics; The identification of the seed semantics further includes that if the three-dimensional initial seed point is at the root of the tooth, as long as the seed position is within the tooth foreground range, it is identified as a foreground point, otherwise it is a background point; If the three-dimensional initial seed point is at the crown of the tooth, project the position coordinates of the three-dimensional initial seed point vertically, and determine whether it is a foreground point or a background point according to the projection position.
5. The panoramic interactive tooth segmentation method based on feature enhancement according to claim 4, wherein the discrimination of whether the three-dimensional initial seed is located in the crown area or the root area further includes discriminating whether the seed point is in the root or crown area by statistical variance of the tooth foreground area of consecutive slices; the determination of whether it is a foreground point or a background point according to the projection position further includes that if the projection position is within the foreground with the largest area in the slice, it is a foreground point, otherwise it is a background point.
6. The panoramic interactive tooth segmentation method based on feature enhancement according to claim 1, wherein the adjustment of the position of the three-dimensional initial seed point in step 1 includes adjusting according to the semantics specified during user interaction. If the user marks it as a foreground point and the semantic discrimination is a background point, the seed point is moved to the centroid of the foreground; if the user marks it as a background point and the semantic discrimination is a foreground point, the seed point is adjusted to the background area and there is a certain distance from the foreground boundary.
7. The panoramic interactive tooth segmentation method based on feature enhancement according to claim 1, The in-layer feature Fp in the pixel relationship generation layer in the computing layer further includes the foreground distance map P F The encoded K feature and the encoded Q feature of the image T are subjected to an inner product operation, and after softmax normalization, they are weighted and summed with the V feature to obtain the foreground distance map P F of the in-layer feature Fp; the background distance map P B The encoded K feature and the encoded Q feature of the image T are subjected to an inner product operation, and after softmax normalization, they are weighted and summed with the B encoded V feature of the background distance map P to obtain the background distance map P B of the in-layer feature; The calculation of the inter-layer pixel relationship generates the inter-layer feature F d Further included is the foreground distance map P F The encoded K feature, the encoded Q feature of the image T, and the background distance map P B After the encoded K feature undergoes matrix permutation, the K features respectively generated from the foreground distance map and the background distance map that have undergone matrix transformation are respectively subjected to inner product operations with the Q feature, softmax normalization, and then weighted summation is performed respectively with the V features respectively generated from the foreground distance map and the background distance map to obtain the inter-layer feature F d .
8. The panoramic interactive tooth segmentation method based on feature enhancement according to claim 1, wherein the determination of the evaluation network input and the construction of the evaluation network in step 3 further include connecting the initial segmentation result R0, the foreground enhancement distance map F(PF), the background enhancement distance map F(PB), and the deepest layer features F in the 3D UNet model after unifying the sizes according to the channels as the evaluation input; and constructing an evaluation network through a convolutional layer, the fourth layer structure of ResNet50, an average pooling layer, and a fully connected layer. s Connect them after unifying the sizes according to the channels as the evaluation input; construct an evaluation network through a convolutional layer, the fourth layer structure of ResNet50, an average pooling layer, and a fully connected layer.
Citation Information
Patent Citations
Interactive mode image partitioning method based on geodesic distance
CN101710418A
Video character separation method and device
CN103119625A