Training Method for Mandibular Nerve Canal CBCT Panoramic Image Segmentation Network Based on Mamba Architecture

By adopting a training method based on Mamba architecture in the medical image segmentation network, using technical means such as gated spatial convolution module and tubular convolution block, the difficulty of medical image segmentation in tubular structures in the existing technology is solved, and a more efficient mandibular nerve tube segmentation effect is achieved.

CN119152212BActive Publication Date: 2025-06-17GANYUE MEDICAL TECH (CHENGDU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411630401.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-06-17
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

The existing medical image segmentation network has difficulties in processing tubular structure medical images and has failed to effectively solve the segmentation problem of this type of image.

Method used

The CBCT panoramic image segmentation network training method based on the Mamba architecture is adopted. By constructing a training data set and training framework, including the stem layer, the encoder part, the decoder part and the jump connection part, the gated spatial convolution module, the tubular convolution block and the three-way Mamba module are used to extract and segment image features.

Benefits of technology

The ability to segment the mandibular nerve tube is improved, providing an effective data segmentation scheme for tubular structures, especially in the field of medical image segmentation where data sources are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152212B_ABST
    Figure CN119152212B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method for a mandibular nerve canal CBCT panoramic image segmentation network based on the Mamba architecture, which relates to the image processing technology in the field of computer vision. The segmentation network training method adopts an encoder-decoder architecture. Its feature extraction operation based on tubular convolutional kernels extracts tubular structure features by expanding the receptive field and learning deformation, and then a three-vision feature fusion strategy is further used to guide the model to supplement the attention to key features from different angles; the encoder part models global and multi-scale features through the Mamba architecture and maintains high efficiency during the training and inference processes, ultimately enabling the network to have a more complete segmentation effect on CBCT images of tubular structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the image processing technology in the field of computer vision. Specifically, it relates to a method for training a panoramic CBCT image segmentation network of the mandibular nerve canal based on the Mamba architecture. Background Art

[0002] The research on the segmentation of medical images of tubular structures has always been an important research content. Medical images of tubular structures have the characteristics of long, slender and tortuous structures, and many key research parts of human organs are tubular structures, such as cardiac blood vessels, arterial blood vessels, mandibular nerve canals, etc. However, the vast majority of medical image segmentation networks do not specifically study medical images of such structures, which causes difficulties in the segmentation and processing of such medical images. Summary of the Invention

[0003] To solve the problem of the segmentation and processing of medical images of tubular structures in the prior art, the present invention proposes a method for training a panoramic CBCT image segmentation network of the mandibular nerve canal based on the Mamba architecture, including the following steps:

[0004] S1. Construct a training dataset: Collect CBCT images of the mandibular nerve canal, and then use manual annotation means to process the collected CBCT images to generate panoramic CBCT images. Label the parts to be segmented in the panoramic CBCT images and convert them into label maps. Pixel points representing the parts to be segmented are assigned a value of 1, and pixel points representing the background are assigned a value of 0. The label map and the corresponding CBCT image form a set of training samples and are sent into the network for training;

[0005] S2. Construct a training framework: The training framework includes a stem layer, an encoder part, a decoder part, and a skip connection part. The encoder part consists of four encoders, and each encoder includes a Mamba module, a GSC module, and a TUC module. The decoder part includes a convolutional layer, an upsampling layer, a segmentation head, skip connections, and residual connection layers. The process of one iteration is as follows:

[0006] S2-1. The panoramic CBCT image to be trained is first sent into the stem layer. The stem layer uses depth convolution, and the image extracts the first feature scale after passing through the stem layer ; then is sent into the Mamba module of the encoder part and the corresponding downsampling layer;

[0007] S2-2. In the first Mamba module, first extract the image spatial features from the stem layer through the gated spatial convolution module GSC. In each subsequent Mamba module, first extract the image spatial features from the previous Mamba module through the gated spatial convolution module GSC; the m-th input spatial feature is first fed into two TUC modules to obtain and ; then, these two features are multiplied pixel by pixel to control the information transmission in the way of a gating mechanism; finally, use another TUC module to further fuse these two features into , and send the feature into the decoder through the residual connection layer in the decoder part;

[0008] S2-3. After obtaining the feature through GSC and three TUC modules, then model the dependency relationship between the features in multiple directions through the ToM module; the feature is unfolded into three sequences along three different directions, and then perform the corresponding feature interaction operations, and finally obtain the integrated feature map ;

[0009] S2-4. Repeat the steps of S2-1, S2-2, and S2-3 to obtain four feature maps output by the encoder , , , and input them into the skip connection layer to obtain the results of each skip connection layer , , , ;

[0010] S2-5. Input the feature map obtained by the last encoder into the decoder for upsampling operation to obtain the feature map . The decoder part also receives , and splices it with the feature map in the channel dimension to obtain ; then perform convolutional upsampling to obtain ; Splice with the feature map in the channel dimension to obtain ; then perform convolutional upsampling to obtain ; Splice with the feature map in the channel dimension to obtain ; then perform convolutional upsampling to obtain ; For After performing another convolutional upsampling, a feature map with the same resolution as the input feature map is obtained. ;

[0011] S2-6. Output the obtained feature map through a segmentation head, calculate the loss function between the output result and the label map, and perform backpropagation based on the calculated result. Thus, one round of iteration is completed.

[0012] Optionally, in step S2-1, the panoramic CBCT image data is uniformly represented as:

[0013]

[0014] where C represents the number of input channels, , H, and W respectively represent the values of the length, width, and height of the panoramic CBCT image; for the stem layer, the kernel size is , the padding size , the stride , and the first-scale feature of the panoramic CBCT image extracted by the stem is represented as:

[0015] .

[0016] Optionally, in step S2-2, the GSC calculation method is represented by the following formula:

[0017] ,

[0018] where, represents the m-th spatial feature of the input, and TUC represents the tubular convolution block.

[0019] Optionally, in step S2-2, in the TUC, first give a standard 3D convolution coordinate with the center coordinate ; a 3×3×3 convolution kernel K is represented as:

[0020] ,

[0021] where the change in the x-axis direction is represented as:

[0022]

[0023] The change in the y-axis direction is:

[0024]

[0025] The change in the z-axis direction is:

[0026]

[0027] Among them, i, j, and k respectively represent the changes in the x, y, and z axis directions, and ∆x, ∆y, and ∆z respectively represent the offsets in the corresponding directions. Since the offsets are decimals while the coordinates are in integer form, bilinear interpolation is adopted, which is expressed as:

[0028]

[0029] Among them, represents the decimal positions of the equation , the equation and the equation . K’ enumerates all integer space positions, and B is the bilinear interpolation kernel, which is decomposed into three one-dimensional kernels, namely:

[0030] ,

[0031] Among them, b represents the bilinear interpolation kernel in each direction, , , respectively represent the convolution kernels along the three directions; for each K, three feature maps from the l-th layer are extracted from the x-axis, y-axis, and z-axis , , , which is expressed as:

[0032] ,

[0033] Among them, , , represent the weights at the corresponding positions, and the features extracted by the convolution kernel K of the l-th layer are calculated using the cumulative method.

[0034] Optionally, in step S2-3, the calculation method of the ToM module can be expressed by the following formula:

[0035] ,

[0036] Among them, Mamba is the Mamba layer used to model global information in the sequence, represents the features in the forward direction, represents the features in the reverse direction, represents the features in the inter-slice direction;

[0037] For the m-th Mamba layer, its calculation process can be defined as:

[0038] ,

[0039] ,

[0040] ,

[0041] Among them, GSC and ToM respectively represent the gated spatial convolution module and the tri-directional Mamba module in the network, , LN represents layer normalization, and MLP represents a multi-layer perceptron layer, which is used to enrich the feature representation, represents the feature from the th layer.

[0042] Optionally, in step S2-4, is used as the input, and then after convolution and residual connection, and finally after convolution for feature fusion and size adjustment, which is expressed as:

[0043] ,

[0044] where represents the final output result of the skip connection of the i-th Mamba module, represents performing a convolution operation.

[0045] Optionally, in step S2-5, the feature map obtained after each upsampling in the decoder is concatenated with the output by the encoder of the same layer in the channel dimension to obtain a concatenated feature map, and then through a convolutional layer, and finally a feature map with the same size as the original image is output.

[0046] Optionally, in step S2-6, finally, the final output result of the network is obtained through the segmentation head, and the cross-entropy loss is calculated between the output result and the label map. The calculation formula is as follows:

[0047] ,

[0048] where N is the number of samples, C is the number of categories, is the indicator function that the pixel i in the true label belongs to the category c, is the probability that the model predicts that the pixel i belongs to the category c.

[0049] The present invention also proposes a method for processing panoramic CBCT images of the mandibular nerve canal, which uses a segmentation network obtained by the aforementioned method for training a panoramic CBCT image segmentation network of the mandibular nerve canal based on the Mamba architecture.

[0050] The beneficial effect of the present invention is that: through the method for training a panoramic CBCT image segmentation network of the mandibular nerve canal based on the Mamba architecture of the present invention, the segmentation ability of the mandibular canal and the mandibular nerve canal is improved under the condition of limited data sources in the field of medical image segmentation, providing an effective segmentation scheme for tubular structure data in intelligent medical image processing. Brief Description of the Drawings

[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0052] Figure 1 is the network framework diagram of the present invention;

[0053] Figure 2 is the schematic diagram of the calculation process of the downsampling layer of the Mamba module of the present invention;

[0054] Figure 3 is the schematic diagram of the calculation process of the GSC module including TUC of the present invention;

[0055] Figure 4 is the schematic diagram of the fusion of three visual features in the present invention;

[0056] Figure 5 is the comparison diagram of partial output results between the present invention and the existing method. Detailed Embodiment

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments.

[0058] As Figures 1-4 shown, the method for training a mandibular nerve canal CBCT panoramic image segmentation network based on the Mamba architecture disclosed in the present invention includes the following steps:

[0059] S1. Construct a training dataset: Collect mandibular nerve canal CBCT images, and then use manual annotation means to process the collected CBCT images to generate panoramic CBCT images. Label the parts to be segmented in the panoramic CBCT images and convert them into label maps. Pixel points representing the parts to be segmented are assigned a value of 1, and pixel points representing the background are assigned a value of 0. The label map and the corresponding CBCT image form a set of training samples and are fed into the network for training;

[0060] S2. Construct the training framework: The training framework consists of three parts: the stem layer, the encoder part, the decoder part, and the skip connection part. The encoder part is composed of four encoders, each encoder includes a Mamba module, a GSC module, and a TUC module. The decoder part includes a convolutional layer, an upsampling layer, a segmentation head, skip connections, and a residual connection layer. The process of one iteration is as follows:

[0061] S2-1. The panoramic CBCT image to be trained is first fed into the stem layer. The stem layer uses depth convolution, and the image extracts the first feature scale after passing through the stem layer ; then it is fed into the Mamba module and the corresponding downsampling layer in the encoder part;

[0062] S2-2. In the first Mamba module, the gated spatial convolution module GSC is first used to extract the image spatial features from the stem layer. In each subsequent Mamba module, the gated spatial convolution module GSC is used to extract the image spatial features from the previous Mamba module. The input m-th spatial feature is first fed into two TUC modules to obtain and ; then, these two features are multiplied pixel by pixel to control the information transmission in the form of a gating mechanism; finally, another TUC module is used to further fuse these two features into , and the feature is fed into the decoder through the residual connection layer in the decoder part;

[0063] S2-3. After obtaining the feature through the GSC and three TUC modules, the ToM module is used to model the dependency relationship between the features in multiple directions; the feature is unfolded into three sequences along three different directions, and then the corresponding feature interaction operations are performed to finally obtain the integrated feature map ;

[0064] S2-4. Repeat the steps of S2-1, S2-2, and S2-3 to obtain the feature maps , , , output by the four encoders and input them into the skip connection layer to obtain the results , , , of each skip connection layer;

[0065] S2-5. Input the feature map obtained from the last encoder into the decoder for upsampling operation to obtain a feature map , and the decoder part simultaneously receives , which is concatenated with the feature map in the channel dimension to obtain ; then perform convolutional upsampling to obtain ; concatenate with the feature map in the channel dimension to obtain ; then perform convolutional upsampling to obtain ; concatenate with the feature map in the channel dimension to obtain ; then perform convolutional upsampling to obtain ; perform one more convolutional upsampling on to obtain a feature map with the same resolution as the input feature map ;

[0066] S2-6. Pass the obtained feature map through a segmentation head for output, calculate the loss function between the output result and the label map, and perform backpropagation according to the calculated result. Thus, one round of iteration is completed.

[0067] Optionally, in step S2-1, the panoramic CBCT image data is uniformly represented as:

[0068]

[0069] where C represents the number of input channels, , and H, W respectively represent the length, width, and height values of the panoramic CBCT image; for the stem layer, use a kernel size of , a padding size , and a stride . The first-scale feature of the panoramic CBCT image extracted by the stem is represented as:

[0070] .

[0071] Optionally, in step S2-2, the GSC calculation method is expressed by the following formula:

[0072] ,

[0073] where, represents the m-th spatial feature of the input, and TUC represents the tubular convolution block.

[0074] Optionally, in step S2-2, in the TUC, first give a standard 3D convolution coordinate with the center coordinate being ; a 3×3×3 convolution kernel K is expressed as:

[0075] ,

[0076] where the change in the x-axis direction is expressed as:

[0077]

[0078] The change in the y-axis direction is:

[0079]

[0080] The change in the z-axis direction is:

[0081]

[0082] where i, j, k respectively represent the changes in the x, y, z-axis directions, and ∆x, ∆y, ∆z respectively represent the offsets in the corresponding directions. Since the offsets are decimals while the coordinates are in integer form, bilinear interpolation is adopted, expressed as:

[0083]

[0084] where represents the decimal positions of the equation , the equation and the equation . K’ lists all integer space positions, and B is the bilinear interpolation kernel, decomposed into three one-dimensional kernels, namely:

[0085] ,

[0086] where b represents the bilinear interpolation kernel in each direction, , , respectively represent the convolution kernels along the three directions; for each K, three feature maps , , from the l-th layer are extracted from the x-axis, y-axis, and z-axis, expressed as:

[0087] ,

[0088] where , , represent the weights at the corresponding positions, and the features extracted by the l-th layer convolution kernel K are calculated using the cumulative method.

[0089] Optionally, in step S2-3, the calculation method of the ToM module can be expressed by the following formula:

[0090] ,

[0091] where Mamba is the Mamba layer used to model global information in the sequence, represents the features in the forward direction, represents the features in the reverse direction, represents the features in the inter-slice direction;

[0092] For the m-th Mamba layer, its calculation process can be defined as:

[0093] ,

[0094] ,

[0095] ,

[0096] where GSC and ToM represent the gated spatial convolution module and the three-way Mamba module in the network respectively, , LN represents layer normalization, and MLP represents the multi-layer perceptron layer, which is used to enrich the feature representation, represents the features from the -th layer.

[0097] Optionally, in step S2-4, is used as the input, and then after convolution and residual connection, and finally after convolution for feature fusion and size adjustment, which is expressed as:

[0098] ,

[0099] where represents the final output result of the skip connection of the i-th Mamba module, represents the convolution operation.

[0100] Optionally, in step S2-5, the feature map obtained after each upsampling in the decoder is concatenated with the output by the encoder of the same layer in the channel dimension to obtain the concatenated feature map, and then through the convolutional layer, and finally output a feature map with the same size as the original image.

[0101] Optionally, in step S2-6, finally, the network's final output result is obtained through the segmentation head, and the cross-entropy loss is calculated between the output result and the label map. The calculation formula is as follows:

[0102] ,

[0103] where N is the number of samples, C is the number of classes, is the indicator function that pixel i in the true label belongs to class c, is the probability that the model predicts pixel i belongs to class c.

[0104] The present invention also proposes a method for processing panoramic CBCT images of the mandibular nerve canal, using the segmentation network obtained by the aforementioned method for training a panoramic CBCT image segmentation network of the mandibular nerve canal based on the Mamba architecture.

[0105] Figure 5 This is a comparison chart of the segmentation effects of the method of the present invention and four other efficient segmentation networks on the CBCT of the mandibular nerve canal of the mandibular canal from three different visual angles. As shown in the figure, compared with other segmentation networks, the method proposed in this application can achieve more complete segmentation of the mandibular nerve canal of the mandibular canal and is stronger in dealing with fine crack details than other segmentation methods. At the same time, it is better in dealing with the overall structure of the panoramic surface and the integrity of the edge area.

[0106] Embodiment

[0107] Construct a training dataset according to step S1. The data is the image of the mandibular CBCT containing the region of interest obtained by scanning and the labeled map containing the region of interest manually marked. Each CBCT image forms a pair of samples with the corresponding labeled map.

[0108] According to step S2-1, the CBCT image to be trained is first sent to the stem layer for convolution operation to extract the first-scale feature, and then further features are extracted through four encoders, and five downsampling operations are performed. At the same time, the feature maps obtained by downsampling each encoder are used as the input of the skip connection layer.

[0109] According to steps S2-2 and S2-3, the first-scale feature extracted by the stem layer is sent to the Mamba module for feature extraction. First, it passes through the GSC module, performs convolution and three-vision feature fusion in the TUC, then layer normalization, and then passes through the ToM module to model the dependencies between features, extract features, and finally perform a downsampling operation and send it to the next Mamba module.

[0110] According to steps S2-4 and S2-5, the feature maps obtained by each layer of skip connection are concatenated with the upsampled feature maps, and then an upsampling operation is performed. After five upsampling operations, a feature map with the same size as the input feature map is obtained.

[0111] Output the feature map obtained from the last upsampling through a segmentation head according to step S2-6. Compare the output result of the network with the label map, calculate the loss function, and perform backpropagation based on the loss function. Thus, a round of training process is completed. Repeat the above process and save the parameters of the final network model.

[0112] Testing: Select samples that did not participate in the training for testing. Compare the output result of the network with the label and calculate the Dice evaluation index to judge the segmentation effect of the model.

[0113] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A network training method for mandibular nerve canal CBCT panoramic image segmentation based on Mamba architecture, characterized in that: The following steps are involved: S1. Construct a training data set: collect CBCT images of the mandibular nerve canal, and then use manual annotation to process the collected CBCT images to generate panoramic CBCT images. Annotate the parts that need to be segmented in the panoramic CBCT images and convert them into label maps. The pixel points representing the parts that need to be segmented are assigned a value of 1, and the pixel points representing the background are assigned a value of 0. The label map and the corresponding CBCT images constitute a set of training samples and are sent to the network for training; S2. Constructing the training framework: The training framework consists of three parts: stem layer, encoder part, decoder part, and jump connection part. The encoder part consists of four encoders, each of which includes Mamba module, GSC module and TUC module. The decoder part includes convolution layer, upsampling layer, segmentation head, jump connection and residual connection layer. The process of one iteration is as follows: S2-1. Panoramic CBCT images to be trained First send it to the stem layer, the stem layer uses deep convolution, the image Extract the first feature scale through the stem layer ; then Feed into the Mamba module of the encoder part and the corresponding lower adoption layer; S2-2. In the first Mamba module, the image spatial features from the stem layer are first extracted through the gated spatial convolution module GSC. In each subsequent Mamba module, the image spatial features from the previous Mamba module are first extracted through the gated spatial convolution module GSC. The mth spatial feature of the input First, it is sent to two TUC modules to obtain and ; Then, the two features are multiplied pixel by pixel to control information transmission in a gating mechanism; Finally, another TUC module is used to further fuse the two features into , and the features are transformed through the residual connection layer of the decoder Send to decoder; S2-3. Features obtained through GSC and three TUC modules After that, the ToM module models features from multiple directions. Dependencies between; Features Expand into three sequences along three different directions, then perform corresponding feature interaction operations, and finally obtain the integrated feature map ; S2-4. Repeat steps S2-1, S2-2, and S2-3 to obtain the feature maps output by the four encoders , , , And input to the jump connection layer to get the result of each jump connection layer , , , ; S2-5. The feature map obtained by the last encoder Input to the decoder for convolution upsampling operation to obtain the feature map , the decoder part receives , and the feature map Concatenate in the channel dimension to get ; Then perform convolution upsampling to obtain ;Will With feature map Concatenate in the channel dimension to get ; Then perform convolution upsampling to obtain ;Will With feature map Concatenate in the channel dimension to get ; Then perform convolution upsampling to obtain ;right After another convolution upsampling, a feature map with the same resolution as the input feature map is obtained. ; S2-6. The obtained feature map The output is sent through a segmentation head, and the loss function is calculated with the output result and the label map. Back propagation is performed based on the calculated result, and a round of iteration is completed.

2. The method for training a mandibular canal CBCT panoramic image segmentation network based on the Mamba architecture according to claim 1, characterized in that: In step S2-1, the panoramic CBCT image data is uniformly expressed as: Where C represents the number of input channels, , H, W represent the length, width, and height of the panoramic CBCT image respectively; for the stem layer, the kernel size is , fill size , step length , the first scale feature of the panoramic CBCT image extracted by stem is expressed as: 。 3. The method for training a mandibular canal CBCT panoramic image segmentation network based on the Mamba architecture according to claim 1, characterized in that: In step S2-2, the GSC calculation method is expressed by the following formula: , in, represents the mth spatial feature of the input, and TUC represents the tubular convolution block.

4. The method for training a mandibular canal CBCT panoramic image segmentation network based on the Mamba architecture according to claim 1, characterized in that: In step S2-2, in TUC, a standard 3D convolution coordinate is first given, and the center coordinate is ; A 3×3×3 convolution kernel K is expressed as: , The change in the x-axis direction is expressed as: The change in the y-axis direction is: The change in the z-axis direction is: Among them, i, j, k represent the changes in the x, y, and z axis directions respectively, and ∆x, ∆y, ∆z represent the offsets in the corresponding directions respectively. Since the offsets are decimals, but the coordinates are in integer form, bilinear interpolation is used, which is expressed as: in, Representation equation ,equation and equation The decimal position of , K' lists all integer spatial positions, B is the bilinear interpolation kernel, which is decomposed into three one-dimensional kernels, namely: , Where b represents the bilinear interpolation kernel in each direction, , , Represents the convolution kernels along three directions respectively; for each K, three feature maps from the lth layer are extracted from the x-axis, y-axis, and z-axis , , , expressed as: , in, , , Represents the weight at the corresponding position, and the features extracted by the convolution kernel K in the lth layer are calculated using the accumulation method.

5. The method for training a mandibular canal CBCT panoramic image segmentation network based on the Mamba architecture according to claim 1, characterized in that: In step S2-3, the ToM module calculation method is expressed by the following formula: , Among them, Mamba is the Mamba layer used to model global information in the sequence, represents the feature in the forward direction, Features that indicate the reverse direction, Features representing the orientation between slices; For the mth Mamba layer, the calculation process is defined as: , , , Among them, GSC and ToM represent the gated spatial convolution module and the three-way Mamba module in the network respectively. , LN represents layer normalization, MLP represents multi-layer perceptron layer, which is used to enrich feature representation, Indicates that from Characteristics of the layer.

6. The method for training a mandibular canal CBCT panoramic image segmentation network based on the Mamba architecture according to claim 1, characterized in that: In step S2-4, As input, then pass Convolution, residual connection, and finally Convolution performs feature fusion and size adjustment, expressed as: , in represents the final output result of the ith Mamba module jump connection, Indicates a convolution operation.

7. The method for training a mandibular canal CBCT panoramic image segmentation network based on the Mamba architecture according to claim 1, characterized in that: In step S2-5, the feature map obtained after each upsampling in the decoder is compared with the feature map output by the encoder at the same layer. The concatenation is performed in the channel dimension to obtain the concatenated feature map, which is then passed through the convolution layer and finally output as a feature map of the same size as the original image.

8. The mandibular canal CBCT panoramic image segmentation network training method based on Mamba architecture according to claim 1, characterized in that: In step S2-6, the final output of the network is obtained through the segmentation head, and the cross entropy loss is calculated between the output result and the label map. The calculation formula is as follows: , Where N is the number of samples, C is the number of categories, is the indicator function that pixel i belongs to category c in the true label, is the probability that the model predicts that pixel i belongs to category c.

9. A method for processing CBCT panoramic images of mandibular canal, characterized in that: A segmentation network is obtained by the mandibular nerve canal CBCT panoramic image segmentation network training method based on the Mamba architecture as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Three-dimensional oral hard palate image segmentation method based on multidirectional state space model

    CN118941585A

  • Few-shot semantic image segmentation using dynamic convolution

    US20230154007A1