Mandibular neural tube automatic segmentation method based on multi-view guidance
By introducing multi-view guidance and shape feature information in the automatic segmentation of mandible nerve tubes, the problem of inaccurate segmentation in the prior art is solved, and more efficient and accurate segmentation of mandible nerve tubes is achieved.
Patent Information
- Application Number
- CN202510149096.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-11
AI Technical Summary
In the process of automatic segmentation of mandibular nerve tubes, the prior art relies on the knowledge learned by the network itself and lacks prior knowledge guidance on shapes, resulting in inaccurate segmentation results.
Using a multi-view guidance method, a neural prediction network and a key point detection network are established by collecting and pre-processing CBCT image data, and combining shape features and texture direction features to accurately segment the three-dimensional mandibular neural tube.
It significantly improves the accuracy and efficiency of mandibular nerve tube segmentation, reduces computational complexity, and minimizes the missing spatial information in 2-dimensional plane detection.
Smart Images

Figure CN120088475A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oral medicine, and specifically to an automatic segmentation method for the mandibular nerve canal based on multi-view guidance. Background Art
[0002] According to the "Global Oral Health Status Report" released by the World Health Organization, approximately 3.5 billion people worldwide suffer from oral diseases. Although oral diseases can be largely prevented, they impose a huge health burden on many countries and affect people's entire lives, leading to pain, discomfort, disfigurement, and even death. Existing research has shown that there is a close relationship between oral diseases and some systemic diseases. For example, there are relationships between oral diseases and cardiovascular diseases as well as diabetes. Therefore, oral problems have attracted increasing attention.
[0003] With the rapid development of "artificial intelligence + medicine", digital technology is changing the way of treating and diagnosing oral diseases. Among them, the mandibular nerve canal is the most concerned tissue in maxillofacial surgery. Whether it is during the extraction of the third molar or the process of dental implantation, damage to the mandibular nerve canal needs to be avoided. The complete and accurate segmentation of the mandibular nerve canal not only helps in the measurement of various intraoperative quantitative indicators, but also greatly improves the efficiency and accuracy of the surgery, achieving "zero damage" to the mandibular nerve canal.
[0004] Traditional segmentation of the mandibular nerve canal generally relies on prior knowledge of shape models. The mandible is segmented using a statistical shape model (SSM), and then a fast matching or tracking algorithm is used to segment the mandibular canal. However, these methods rely on prior knowledge and a robust mandible segmentation model.
[0005] For existing digital methods for automatic segmentation of the mandibular nerve canal, most of them adopt a coarse-to-fine segmentation method, using the cascade of multiple segmentation networks to achieve the best segmentation result. However, this cascade segmentation method has a large mutual influence, and only relies on the knowledge learned by the network itself during the segmentation process, without introducing prior knowledge of shape to guide the segmentation, resulting in the final segmentation result not meeting the expected effect.
[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] The purpose of the present invention is to provide an automatic segmentation method for the mandibular nerve canal based on multi-view guidance to solve the problems raised in the above background art.
[0008] To achieve the above purpose, the present invention provides the following technical solutions:
[0009] A method for automatic segmentation of the mandibular nerve canal based on multi-view guidance, the specific steps include:
[0010] Collect a number of original oral CBCT image data with known mandibular nerve canal feature information, perform image preprocessing on the collected original oral CBCT image data to obtain sample CBCT image data, and based on the obtained sample CBCT image data, use the method of manual marking to label the mandibular bone region in the sample CBCT image to obtain training CBCT image data. The mandibular nerve canal feature information includes the position information of the key points of the mandibular nerve canal and the corresponding heat map.
[0011] Based on the obtained training CBCT image data, establish a nerve prediction network, use the training CBCT image as the input of the nerve prediction network, and use the annotation of the mandibular bone region in the image as the label to train the nerve prediction network to obtain a mandibular bone segmentation network.
[0012] Based on the mandibular bone segmentation network, segment the training CBCT image to obtain a mandibular bone tissue image, perform contrast enhancement on the obtained mandibular bone tissue image to obtain a target mandibular bone tissue image, and project the target mandibular bone tissue image along three orthogonal planes: the coronal plane, the sagittal plane, and the cross-sectional plane, respectively, to obtain the projection plan views of the three orthogonal planes.
[0013] According to the known mandibular nerve canal feature information of the training CBCT image data, determine the key points of the mandibular nerve canal in each projection plan view of the three orthogonal planes, label the key points in each projection plan view, extract the position information, and correspond it to the known heat map of the key points.
[0014] Based on the position information of the key points and the corresponding heat maps of the key points in each projection plan view, establish a key point detection network, use the projection plan views of the three orthogonal planes as the input, and use the position information of the key points and the corresponding heat maps of the key points in each projection plan view as the label to train the key point detection network.
[0015] After passing the oral CBCT image to be segmented through the mandibular bone segmentation network, contrast enhancement, projection, and key point detection network, extract the shape feature information of the heat map output by the key point detection network through an encoder, and combine the shape feature information of the mandibular nerve canal with the texture direction feature as the guidance information and input it into the 3D CBCT mandibular nerve canal segmentation network to guide the network to perform the final accurate 3D segmentation of the mandibular nerve canal.
[0016] Furthermore, perform image preprocessing on the collected original oral CBCT image data, and the preprocessing includes signal denoising processing and enhancement processing.
[0017] Further, train the neural prediction network to obtain a mandible segmentation network, where the mandible segmentation network is established based on the 3D-Unet segmentation network. Feed the preprocessed CBCT data into the mandible segmentation network to remove redundant regions from the data and segment the mandible as the region of interest. The specific steps include: using the training CBCT image data as input and the labeled mandible region as the output label to train the mandible segmentation network, and obtaining the optimal parameters of the segmentation network by minimizing the loss during the training process; separately segmenting the mandible tissue through the trained mandible segmentation network to obtain the mandible tissue image. The specific steps for obtaining the optimal parameters of the segmentation network include: cropping the training CBCT image data and the corresponding labeled data into slices of size 128×128×128.
[0018] Perform data augmentation operations such as randomly rotating, translating, scaling, and flipping the data; feed the data after the data augmentation operations into the 3D-Unet segmentation network. The encoder gradually extracts image features and reduces its spatial dimension through a combination of multiple convolutional layers and pooling layers; extract the finest image features at the bottleneck layer; in the decoding stage, use transposed convolutional layers to upsample the feature maps, restore the spatial resolution, and use skip connections to splice the feature maps of the corresponding layers of the encoder with the decoder feature maps to restore the spatial information.
[0019] Train the network parameters using a weighted combination of cross-entropy loss and Dice loss to obtain the optimal parameters of the mandible segmentation network.
[0020] Further, the specific steps for obtaining the optimal parameters of the segmentation network include: defining the cross-entropy loss as defining the Dice loss as defining the weighted combination of cross-entropy loss and Dice loss as that is, the loss of the mandible segmentation network is specifically defined as:
[0021]
[0022] In the formula, λ represents the weighting coefficient, and the value range is λ∈[0,1]; the cross-entropy loss and the Dice loss are defined as:
[0023]
[0024] In the formula, N represents the total number of voxels in the training CBCT image data; y i represents the true label value of the i-th voxel, and the value range is y i ∈[0,1], 0 represents the background region, and 1 represents the mandible region. is the predicted value of the segmentation model for the i-th voxel, where i is the index of the voxel in the training CBCT image data, and i ∈ [1, N].
[0025] Further, the obtained mandibular tissue image is subjected to contrast enhancement to obtain the target mandibular tissue image. The contrast enhancement technique used is adaptive histogram equalization. The specific steps include: dividing the mandibular tissue image into 16×16×16 patches, calculating the pixel histogram for each small patch; setting a contrast limit threshold within each small patch; equalizing the histogram of each small patch, that is, mapping the cumulative distribution function of each pixel histogram to the [0, 255] pixel range; merging each patch, and using bilinear interpolation method at the boundary part to smooth the transition between each small patch to obtain the final target mandibular tissue image.
[0026] Further, using the perspective projection method, each target mandibular tissue image is projected along the sagittal plane, coronal plane and cross-sectional plane to obtain their respective projection sections;
[0027] Taking the projection plan views of the three orthogonal sections as the input, and using the position information of the key points in each projection plan view and the heat map of the corresponding key points as the labels, the key point detection network is trained. Among them, the specific steps of training the key point detection network include:
[0028] Pair the sagittal plane, coronal plane and cross-sectional projection images corresponding to each target mandibular tissue image, and use the nearest neighbor interpolation algorithm to unify the sizes of the three projection images;
[0029] Input the projection images into the key point detection network HR-Net network, and extract the initial feature map through 3×3 convolution + Batch Normalization + ReLU activation function containing 2 layers;
[0030] Introduce the category information of the projection images, that is, the sagittal plane, coronal plane and cross-sectional projection images have their own exclusive categories. The category information obtains the projection image category features through the information embedding module, and fuses the category features of the projection images into the initial feature map to obtain the advanced image feature map containing the projection image category information;
[0031] The advanced feature map extracts feature maps with different resolutions through parallel multi-branches. The low-resolution branch captures global context information, the high-resolution branch captures detail information, and the high-resolution feature map is fused with the low-resolution feature map to capture multi-scale information;
[0032] All the low-resolution feature maps are upsampled and fused with the high-resolution feature map to output the predicted key point heat map;
[0033] Based on the predicted key-point heatmap and the known true key-point heatmap, a weighted combination of mean squared error loss, heatmap loss, and focal loss is used to constrain the training of the key-point detection network.
[0034] Furthermore, the loss function for training the key-point detection network is defined as:
[0035]
[0036] where α, β, γ belong to the weight hyperparameters, where γ > β > α and α + β + γ = 1, is the mean squared error loss, is the heatmap loss, is the focal loss, where the mean squared error loss and the heatmap loss as well as the focal loss are defined as:
[0037]
[0038] where Q is the total number of key points, k s is the true position of the s-th key point, is the predicted position of the s-th key point; H s represents the true heatmap of the s-th key point, the predicted heatmap of the s-th key point; M is the total number of pixel points in each heatmap, p sj is the probability that the j-th pixel on the s-th heatmap is predicted as foreground or background, α sj is the weight factor for adjusting class imbalance, δ is the adjustment factor of the focal loss, where s is the index of the key point, s ∈ [1, Q], and j is the index of the pixel on the heatmap, j ∈ [1, M].
[0039] Furthermore, the heatmap output by the key-point detection network is used to extract shape feature information through an encoder, and the shape feature information of the mandibular nerve canal and the texture direction feature are combined as guiding information and input into the 3D CBCT mandibular nerve canal segmentation network, where the formula for calculating the texture direction feature is:
[0040]
[0041] where C is the texture direction feature vector, W is the set voxel window, |W| is the number of voxels in the window, is the gradient vector of the voxel, is the transpose of the gradient vector, (x, b, z) represents the position coordinates of the voxel, where the gradient vector of the voxel is calculated according to the formula:
[0042]
[0043] In the formula, represents the change rate of the voxel value in the X-axis direction, represents the change rate of the voxel value in the Y-axis direction, represents the change rate of the voxel value in the Z-axis direction.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0045] The present invention uses a segmentation network to first segment the mandible to obtain the region of interest, resulting in less redundant information in the entire CBCT image. Compared with locating the inferior alveolar nerve canal from the complete CBCT image, this method greatly saves the computational complexity of the subsequent key point detection network and the 3D segmentation network, and also lays a foundation for more accurate segmentation of the inferior alveolar nerve canal subsequently.
[0046] The key point detection network is used to detect key points on the three orthogonal projection planes obtained from the sagittal plane, coronal plane, and cross-sectional plane. The network captures the basic shape information contained in each plane by extracting primary features, advanced features, and high-level features, and finally combines the information of the three projection planes to obtain the basic shape information of the entire inferior alveolar nerve canal. This method not only reduces the computational complexity of the model but also minimizes the spatial information missing in key point detection on the 2D plane.
[0047] During the 3D segmentation process of the inferior alveolar nerve canal, the constraint of the basic shape information is added, enabling the segmentation network to pay more attention to extracting shape-related features of the inferior alveolar nerve canal during the feature extraction process, thereby guiding the decoder to finally generate a more complete and accurate 3D inferior alveolar nerve canal. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic diagram of the overall method flow of the present invention;
[0049] Figure 2 is a structural diagram of the mandible segmentation network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.
[0051] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0052] Embodiment:
[0053] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution:
[0054] An automatic segmentation method of the mandibular nerve canal based on multi-view guidance, the specific steps include:
[0055] Step 1: Collect a number of original oral CBCT image data with known mandibular nerve canal feature information, perform image preprocessing on the collected number of original oral CBCT image data to obtain sample CBCT image data, and based on the obtained sample CBCT image data, use the method of manual marking to label the mandibular bone region in the sample CBCT image to obtain training CBCT image data. The mandibular nerve canal feature information includes the position information of the key points of the mandibular nerve canal and the corresponding heat maps.
[0056] Perform image preprocessing on the collected number of original oral CBCT image data, and the preprocessing includes signal denoising processing and enhancement processing.
[0057] Among them, the specific methods of signal denoising processing and enhancement processing are: use the denoising method of wavelet transform to perform denoising processing on the infrared spectral image. The specific steps of the wavelet transform denoising method include: decompose the distortion-corrected image through wavelet transform to obtain wavelet coefficients of the image at different scales and directions; perform threshold processing on the wavelet coefficients, set the wavelet coefficients with low amplitudes to zero, and retain the wavelet coefficients with high amplitudes; perform inverse transform on the wavelet coefficients after threshold processing, and reconstruct the processed coefficients into an image to complete the image denoising processing; use bilateral filtering to enhance the details of the original oral CBCT image.
[0058] Among them, label the mandibular bone region by the method of manual marking through the Label Img tool.
[0059] Step 2: Based on the obtained training CBCT image data, establish a neural prediction network. Use the training CBCT images as the input of the neural prediction network, and use the labeled mandibular region in the images as the labels to train the neural prediction network, obtaining a mandibular segmentation network.
[0060] Train the neural prediction network to obtain a mandibular segmentation network. The mandibular segmentation network is established based on the 3D-Unet segmentation network. Feed the preprocessed CBCT data into the mandibular segmentation network to remove redundant regions and segment the mandible as the region of interest. The specific steps are as follows: Use the training CBCT image data as the input and the labeled mandibular region as the output label to train the mandibular segmentation network. During the training process, minimize the loss to obtain the optimal parameters of the segmentation network. Separate and process the mandibular tissue through the trained mandibular segmentation network to obtain the mandibular tissue image. The specific steps to obtain the optimal parameters of the segmentation network are as follows: Crop the training CBCT image data and the corresponding labeled data into patches of size 128×128×128.
[0061] Perform data augmentation operations such as random rotation, translation, scaling, and flipping on the data. Feed the data after the data augmentation operations into the 3D-Unet segmentation network. The encoder gradually extracts image features and reduces its spatial dimension through a combination of multiple convolutional layers and pooling layers. Extract the finest image features at the bottleneck layer. In the decoding stage, use transposed convolutional layers to upsample the feature maps, restore the spatial resolution, and use skip connections to stitch the feature maps of the corresponding layers in the encoder with the decoder feature maps to restore the spatial information.
[0062] Train the network parameters using a weighted combination of cross-entropy loss and Dice loss to obtain the optimal parameters of the mandibular segmentation network.
[0063] The specific steps to obtain the optimal parameters of the segmentation network include: Define the cross-entropy loss as Define the Dice loss as Define the weighted combination of cross-entropy loss and Dice loss as That is, the loss of the mandibular segmentation network Is specifically defined as:
[0064]
[0065] In the formula, λ represents the weighting coefficient, and the value range is λ∈[0,1]; the cross-entropy loss And the Dice loss Are defined as:
[0066]
[0067] In the formula, N represents the total number of voxels in the training CBCT image data; y i represents the true label value of the i-th voxel, and the value range is y i ∈[0, 1], where 0 represents the background area and 1 represents the mandibular region; is the predicted value of the segmentation model for the i-th voxel. i is the index of the voxel in the training CBCT image data, and i ∈ [1, N].
[0068] Step 3: Segment the training CBCT image based on the mandibular segmentation network to obtain a mandibular tissue image. Enhance the contrast of the obtained mandibular tissue image to obtain a target mandibular tissue image. Project the target mandibular tissue image along three orthogonal planes, namely the coronal plane, sagittal plane, and cross-sectional plane, to obtain projection plane graphs of the three orthogonal planes respectively;
[0069] Using the perspective projection method, project each target mandibular tissue image along the sagittal plane, coronal plane, and cross-sectional plane respectively to obtain their respective projection sections; where the direction along the sagittal plane is the X-axis coordinate direction, the direction along the coronal plane corresponds to the Y-axis coordinate direction, and the direction along the cross-sectional plane corresponds to the Z-axis coordinate direction.
[0070] Enhance the contrast of the obtained mandibular tissue image to obtain a target mandibular tissue image. The contrast enhancement technique used is adaptive histogram equalization. The specific steps include: dividing the mandibular tissue image into 16×16×16 blocks, calculating the pixel histogram for each small block; setting a contrast limit threshold within each small block; equalizing the histogram of each small block, that is, mapping the cumulative distribution function of each pixel histogram to the [0, 255] pixel range; merging each block and using bilinear interpolation at the boundary to smooth the transition between each small block to obtain the final target mandibular tissue image.
[0071] Step 4: According to the known mandibular canal feature information in the training CBCT image data, determine the key points of the mandibular canal in the projection plane graphs of the three orthogonal planes, mark the key points in each projection plane graph, extract the position information, and correspond it to the known key point heat map.
[0072] Among them, the key points of the mandibular canal usually include: the starting point and ending point of the mandibular canal. The starting point is usually located at the mandibular foramen, which is the entrance where the inferior alveolar nerve enters the mandible from the base of the skull, and the ending point is usually located at the mental foramen, which is the exit where the inferior alveolar nerve exits the mandible to reach the face.
[0073] Path feature points of the nerve canal. The mandibular nerve canal follows a curved course along the mandible, and there may be significant bending points or inflection points in the path. These feature points reflect the spatial orientation of the nerve canal. For example, the middle inflection point (middle path point) of the nerve canal is used to describe the three-dimensional curve shape of the entire nerve canal, and the structural change points, such as the places where the inner diameter of the nerve canal suddenly narrows or widens.
[0074] Corresponding points of specific anatomical landmarks. The mandibular nerve canal is associated with other important anatomical structures (such as tooth roots, mental region of the mandible, etc.) near some specific anatomical positions. For example, the adjacent point of the tooth root: the point where the mandibular nerve canal is near the apical region of the posterior mandibular teeth (such as molars or premolars).
[0075] Step 5: Based on the position information of the key points in each projected planar graph and the heat maps corresponding to the key points, establish a key point detection network. Using the projected planar graphs of three orthogonal sections as the input and the position information of the key points in each projected planar graph and the heat maps corresponding to the key points as the labels, train the key point detection network.
[0076] Using the projected planar graphs of three orthogonal sections as the input and the position information of the key points in each projected planar graph and the heat maps corresponding to the key points as the labels, train the key point detection network. Among them, the specific steps for training the key point detection network include:
[0077] Pair the sagittal, coronal, and cross-sectional projection images corresponding to each target mandibular tissue image, and use the nearest neighbor interpolation algorithm to unify the sizes of the three projection images;
[0078] Input the projection images into the key point detection network HR-Net network, and extract the initial feature map through a 3×3 convolution + Batch Normalization + ReLU activation function containing 2 layers;
[0079] Introduce the category information of the projection images, that is, the sagittal, coronal, and cross-sectional projection images have their own exclusive categories. The category information obtains the projection image category features through the information embedding module, and fuses the projection image category features into the initial feature map to obtain an advanced image feature map containing the projection image category information;
[0080] The advanced feature map extracts feature maps of different resolutions through parallel multi-branches. The low-resolution branch captures global context information, the high-resolution branch captures detail information, and the high-resolution feature map is fused with the low-resolution feature map to capture multi-scale information;
[0081] Upsample all the low-resolution feature maps and fuse them with the high-resolution feature map to output the predicted key point heat map;
[0082] Based on the predicted key-point heatmap and the known ground-truth key-point heatmap, a weighted combination of mean squared error loss, heatmap loss, and focal loss is used to constrain the training of the key-point detection network.
[0083] The loss function for training the key-point detection network is defined as:
[0084]
[0085] where α, β, γ are weight hyperparameters, is the mean squared error loss, is the heatmap loss, is the focal loss, where the mean squared error loss and the heatmap loss and the focal loss are defined as:
[0086]
[0087] where Q is the total number of key points, k s is the ground-truth position of the s-th key point, is the predicted position of the s-th key point; H s represents the ground-truth heatmap of the s-th key point, the predicted heatmap of the s-th key point; M is the total number of pixels in each heatmap, p sj is the probability that the j-th pixel on the s-th heatmap is predicted as foreground or background, α sj is the weight factor for adjusting class imbalance, δ is the adjustment factor of the focal loss, where s is the index of the key point, s ∈ [1, Q], j is the index of the pixel on the heatmap, j ∈ [1, M], and since the main role of the focal loss is to enhance the attention to the key-point region while suppressing the interference of the background region. This is crucial for the key-point detection task because the key-point region only accounts for a very small part of the image, while the background region is usually the main interference item, so its weight is the largest, and the heatmap loss functions to guide the model to generate a probability heatmap consistent with the true distribution, thereby enhancing the local detection ability of the key points. In contrast, the mean squared error loss is a direct constraint on the final key-point coordinates. Since the distribution information of the heatmap is more important for the training of the network, the weight of the heatmap loss is usually higher than that of the mean squared error loss, so γ > β > α and α + β + γ = 1.
[0088] Step 6: After passing the oral CBCT image to be segmented through the mandibular bone segmentation network, contrast enhancement, projection, and key point detection network, the heat map output by the key point detection network is used to extract shape feature information through an encoder. The shape feature information of the mandibular nerve canal is combined with the texture direction feature as guiding information and input into the 3D CBCT mandibular nerve canal segmentation network to guide the network to perform a final accurate 3D segmentation of the mandibular nerve canal.
[0089] The heat map output by the key point detection network is used to extract shape feature information through an encoder. The shape feature information of the mandibular nerve canal is combined with the texture direction feature as guiding information and input into the 3D CBCT mandibular nerve canal segmentation network, where the formula for calculating the texture direction feature is:
[0090]
[0091] In the formula, C is the texture direction feature vector, W is the set voxel window, |W| is the number of voxels in the window, is the gradient vector of the voxel, is the transpose of the gradient vector, (x, b, z) represents the position coordinates of the voxel, where the gradient vector of the voxel is calculated according to the formula:
[0092]
[0093] In the formula, represents the rate of change of the voxel value in the X-axis direction, represents the rate of change of the voxel value in the Y-axis direction, represents the rate of change of the voxel value in the Z-axis direction.
[0094] Among them, the specific steps for guiding the network to perform a final accurate 3D segmentation of the mandibular nerve canal include:
[0095] The sagittal, coronal, and cross-sectional heat maps with key point position information output by the key point detection network are respectively input into three image encoders to extract shape information;
[0096] The shape information and texture direction feature information extracted from the projection plane are sent to the feature fusion module to inject the shape information into the 3D CBCT image encoder;
[0097] The corresponding 3D training CBCT image data is sent into the image encoder based on Swin Transformer. The image encoder combines the injected basic shape information and the powerful global modeling ability of Transformer to extract the overall structural features focused on the mandibular nerve canal;
[0098] In the decoder part, following the traditional U-Net decoding structure, the features extracted by the encoder are connected to the decoder at various scales through skip connections, and finally an accurately segmented 3D mandibular nerve canal is obtained.
[0099] Among them, the loss function for training the 3D CBCT mandibular nerve canal segmentation network is defined as Its expression is as follows:
[0100]
[0101] In the formula, μ is the weighting factor, and its value range is μ ∈ [0, 1], is the Dice loss, is the clDice loss, which can simultaneously consider the topological structure and connectivity of the mandibular nerve canal. Its definition is:
[0102]
[0103] In the formula, V P is the result predicted by the network, V L is the ground truth, S P and S L are obtained through the soft-skeleton method, and T pre (S P , V L ) and T sen (S L , V P ) are defined as:
[0104]
[0105] In the formula, in the mandibular nerve canal segmentation, S P represents the central axis of the nerve canal extracted from the result V P predicted by the network, and S L is the skeleton information extracted from the ground truth V L , representing the central axis of the true position of the nerve canal. T pre (S P , V L ) is the topological accuracy, which is used to calculate the proportion of the overlapping part between S P and the ground truth V L , and T sem (S L , V P ) is the topological sensitivity, which is used to represent the proportion of the overlapping part between S L and the result V P predicted by the network.
[0106] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0107] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0108] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0109] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application.
Claims
1. A method for automatic segmentation of the mandibular nerve canal based on multi-view guidance, characterized in that: The specific steps include: Collecting a number of original oral CBCT image data with known mandibular nerve canal feature information, performing image preprocessing on the collected number of original oral CBCT image data to obtain sample CBCT image data, and annotating the mandibular bone area in the sample CBCT image by manual marking based on the obtained sample CBCT image data to obtain training CBCT image data, wherein the mandibular nerve canal feature information includes the position information of key points of the mandibular nerve canal and the corresponding thermal map; Based on the obtained training CBCT image data, a neural prediction network is established, the training CBCT image is used as the input of the neural prediction network, and the mandibular region annotation in the image is used as a label to train the neural prediction network, thereby obtaining a mandibular segmentation network; The training CBCT image is segmented based on the mandibular segmentation network to obtain a mandibular tissue image, the obtained mandibular tissue image is contrast enhanced to obtain a target mandibular tissue image, and the target mandibular tissue image is projected along three orthogonal planes, namely, the coronal plane, the sagittal plane and the cross-sectional plane, to obtain projection plane images of the three orthogonal planes respectively; According to the known feature information of the mandibular nerve canal in the training CBCT image data, the key points of the mandibular nerve canal in each projection plane are determined in the projection plane of three orthogonal sections, the key points are marked in each projection plane, the position information is extracted, and the key points are matched with the known key point heat map; Based on the position information of key points in each projection plane and the heat map of the corresponding key points, a key point detection network is established. The projection planes of three orthogonal sections are used as inputs, and the position information of key points in each projection plane and the heat map of the corresponding key points are used as labels to train the key point detection network. After the oral CBCT image to be segmented passes through the mandibular segmentation network, contrast enhancement, projection and key point detection network, the heat map output by the key point detection network is used to extract the shape feature information through the encoder, and the shape feature information of the mandibular nerve canal combined with the texture direction feature is input into the three-dimensional CBCT mandibular nerve canal segmentation network as guidance information, guiding the network to perform the final accurate 3D segmentation of the mandibular nerve canal.
2. The method for automatic segmentation of the mandibular nerve canal based on multi-view guidance according to claim 1, characterized in that: The collected original oral CBCT image data are subjected to image preprocessing, wherein the preprocessing includes signal denoising processing and enhancement processing.
3. The method for automatic segmentation of the mandibular nerve canal based on multi-view guidance according to claim 2, characterized in that: The neural prediction network is trained to obtain a mandibular segmentation network, wherein the mandibular segmentation network is established based on a 3D-Unet segmentation network, the preprocessed CBCT data is sent to the mandibular segmentation network, redundant regions of the data are removed, and the mandibular bone is segmented as a region of interest, wherein the specific steps include: taking the training CBCT image data as input and the annotated mandibular region as output label, training the mandibular segmentation network, and obtaining the optimal parameters of the segmentation network by minimizing the loss during the training process; the mandibular tissue is segmented separately by completing the trained mandibular segmentation network to obtain a mandibular tissue image, wherein the optimal parameters of the segmentation network are obtained, wherein the specific steps include: cropping the training CBCT image data and the corresponding annotated data into 128×128×128 size blocks; Perform data augmentation operations such as random rotation, translation, scaling, and flipping on the data; feed the data augmented with data into the 3D-Unet segmentation network, where the encoder gradually extracts image features and reduces its spatial dimension through a combination of multiple convolutional layers and pooling layers; extracts the finest image features at the bottleneck layer; in the decoding stage, use the transposed convolutional layer to upsample the feature map to restore the spatial resolution, and use jump connections to concatenate the feature map of the encoder's corresponding layer with the decoder's feature map to restore the spatial information; The network parameters are trained using a weighted combination of cross entropy loss and Dice loss to obtain the optimal parameters of the mandibular segmentation network.
4. The method for automatic segmentation of the mandibular canal based on multi-view guidance according to claim 3, characterized in that: The specific steps of obtaining the optimal parameters of the segmentation network include: defining the cross entropy loss as The Dice loss is defined as The weighted combination of cross entropy loss and Dice loss is defined as The loss of the mandibular segmentation network The specific definition is: In the formula, λ represents the weighting coefficient, and its value range is λ∈[0,1]; cross entropy loss and Dice loss is defined as: Where N represents the total number of voxels in the training CBCT image data; y i Represents the true label value of the i-th voxel, ranging from y i ∈[0,1], 0 represents the background area, 1 represents the mandibular area; is the predicted value of the segmentation model for the i-th voxel, i is the index of the voxel in the training CBCT image data, i∈[1,N].
5. The method for automatic segmentation of the mandibular nerve canal based on multi-view guidance according to claim 1, characterized in that: The obtained mandibular tissue image is contrast enhanced to obtain a target mandibular tissue image, wherein the contrast enhancement technique used is adaptive histogram equalization, and the specific steps include: dividing the mandibular tissue image into 16×16×16 blocks, and calculating the pixel histogram of each small block; setting a contrast limit threshold in each small block; equalizing the histogram of each small block, that is, mapping the cumulative distribution function of each pixel histogram to the [0,255] pixel range; merging the blocks, and using the bilinear interpolation method at the boundary to smooth the transition between the small blocks, to obtain the final target mandibular tissue image.
6. The method for automatic segmentation of the mandibular canal based on multi-view guidance according to claim 5, characterized in that: Using the perspective projection method, each target mandibular tissue image is projected along the sagittal plane, coronal plane, and cross-sectional plane to obtain respective projection sections; The key point detection network is trained with the projection plane images of three orthogonal sections as input, and the position information of the key points in each projection plane image and the heat map of the corresponding key points as labels. The specific steps of training the key point detection network include: The sagittal, coronal, and cross-sectional projection images corresponding to each target mandibular tissue image were paired, and the sizes of the three projection images were unified using the nearest neighbor interpolation algorithm; The projected image is input into the key point detection network HR-Net network, and the initial feature map is extracted through a 2-layer 3×3 convolution + BatchNormalization + ReLU activation function; The category information of the projection image is introduced, that is, the sagittal, coronal and cross-sectional projection images have their own exclusive categories. The category information is used to obtain the category features of the projection image through the information embedding module. The category features of the projection image are fused into the initial feature map to obtain an advanced image feature map containing the category information of the projection image. The advanced feature map extracts feature maps of different resolutions through multiple parallel branches. The low-resolution branch captures global context information, the high-resolution branch captures detail information, and the high-resolution feature map is fused with the low-resolution feature map to capture multi-scale information. All low-resolution feature maps are upsampled and fused with high-resolution feature maps to output predicted key point heat maps; Based on the predicted keypoint heatmap and the known true keypoint heatmap, a weighted combination of mean squared error loss, heatmap loss and focal loss is used to constrain the training of the keypoint detection network.
7. The method for automatic segmentation of the mandibular canal based on multi-view guidance according to claim 6, characterized in that: Loss function for training keypoint detection network Defined as: In the formula, α, β, and γ are weight hyperparameters, where γ>β>α and α+β+γ=1. is the mean square error loss, is the heat map loss, is the focal loss, where the mean square error loss and heatmap loss And focal loss is defined as: Where Q is the total number of key points, k s is the true position of the sth key point, is the predicted position of the sth key point; H s represents the true heat map of the sth key point, The predicted heat map of the sth key point; M is the total number of pixels in each heat map, p sj is the probability that the jth pixel on the sth heat map is predicted to be foreground or background, α sj is the weight factor for adjusting category imbalance, δ is the adjustment factor for focal loss, where s is the index of the key point, s∈[1,Q], and j is the index of the pixel on the heat map, j∈[1,M].
8. The method for automatic segmentation of the mandibular canal based on multi-view guidance according to claim 7, characterized in that: The heat map output by the key point detection network is used to extract shape feature information through the encoder, and the shape feature information of the mandibular nerve canal is combined with the texture direction feature as guidance information and input into the 3D CBCT mandibular nerve canal segmentation network, wherein the formula for calculating the texture direction feature is: Where C is the texture direction feature vector, W is the set voxel window, |W| is the number of voxels in the window, is the gradient vector of the voxel, is the transpose of the gradient vector, (x, b, z) represents the position coordinates of the voxel, where the gradient vector of the voxel The calculation is based on the formula: In the formula, Indicates the rate of change of the voxel value in the X-axis direction, Indicates the rate of change of the voxel value in the Y-axis direction, Indicates the rate of change of the voxel value in the Z-axis direction.
Citation Information
Patent Citations
Deep convolutional neural networks for tumor segmentation with positron emission tomography
CN113711271A
Mandibular neural tube segmentation method, device, electronic equipment and storage medium
CN114037665A
Method for extracting central line of mandibular neural tube and calculating radius of mandibular neural tube
CN114677374A
CBCT mandibular neural tube refined segmentation method based on multi-view feature fusion
CN119399117A
Deep convolutional neural networks for tumor segmentation with positron emission tomography
US20210401392A1
Cited By
Method and system for detecting symmetry of mandible
CN120451158A
Oral disease image report generation method and device, equipment and medium
CN120452657A
Mandibular image segmentation method and device based on deep learning, equipment and medium
CN122089733A