Pancreas Segmentation Network Based on Multi-Span Complementary Information Capture and Fusion
By designing a pancreatic segmentation network with multi-span complementary information capture and fusion, the problem of insufficient multi-scale feature sensing in the prior art is solved, and a more accurate pancreatic segmentation effect is achieved.
Patent Information
- Application Number
- CN202211108345.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-09-13
AI Technical Summary
The existing pancreatic medical image segmentation technology cannot complete multi-scale feature sensing, resulting in the gradual loss of some detailed features during the feature extraction process, and the segmentation effect is not ideal.
A pancreatic segmentation network based on multi-span complementary information capture and fusion is designed, including encoders and decoders with dimensional reduction and dimensional increase grid residual convolution, multi-scale feature capture module, spatiotemporal attention module and pyramid multi-scale fusion post-processing module. Through these modules, the capture, fusion and detail recall of multi-scale features are achieved.
It effectively solves the problem of easily losing small target information and edge information in the prior art, and improves the accuracy and effect of pancreatic segmentation.
Smart Images

Figure CN115423831B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly relates to a pancreatic segmentation network based on multi-span complementary information capture and fusion. Background Art
[0002] Achieving accurate and automatic segmentation of the pancreas and its cysts is an urgent need in clinical practice. At present, deep neural networks have been widely used in the field of automatic segmentation of abdominal organs. However, existing segmentation methods are only applicable to the segmentation of organs with a large abdominal cavity occupancy and clear morphology, such as the liver and lungs. For organs like the pancreas with variable morphology, large individual differences, unclear edges, and a small target area, the segmentation results are not ideal.
[0003] Cai et al. first pre-trained a 2DCNN using pancreatic CT slices to complete a relatively rough segmentation of a single layer of pancreatic slices. Subsequently, the segmentation results of multiple adjacent single-layer pancreatic CT slices were input into the next-layer RNN model to obtain a more accurate and smooth pancreatic segmentation result. Although this end-to-end segmentation model with multiple networks stacked achieved a relatively high accuracy in pancreatic segmentation, it requires pre-training of the network. Zhou et al. first performed a rough segmentation of pancreatic CT slices using a 2DFCN to remove irrelevant background pixels. Then, the cropped images were input into the next segmentation network to obtain a more refined semantic segmentation result. Although it reduced the interference of invalid segmentation information, it lost the global information of the slices, and the final segmentation result was poor. Yu et al. combined two 2DFCNs in a cascaded manner to form a segmentation network, obtained a probability map related to the pancreas, multiplied the map by the pancreatic CT slices and then performed regional cropping, and then input it into the next 2DFCN for fine segmentation. Finally, the energy function was jointly optimized to obtain a more accurate network segmentation model. The above solutions all achieved image segmentation of a single pancreatic organ by stacking two or more networks. This multi-network segmentation scheme has a relatively large overall number of parameters, high requirements for experimental equipment, and such methods only reduce the interference of redundant information and have not achieved high segmentation accuracy.
[0004] The above research believes that the task cannot complete multi-scale feature sensing, and with each step of feature convolution, some detailed features are gradually lost, and it is impossible to complete detail complementation during the feature extraction process. Summary of the Invention
[0005] In order to overcome the defects and deficiencies existing in the prior art, the present invention proposes a pancreatic segmentation network based on multi-span complementary information capture and fusion, which is used to solve the problems existing in the existing pancreatic medical image segmentation technology, namely, the inability to complete multi-scale feature sensing and the gradual loss of some detailed features with each step of feature convolution.
[0006] It utilizes the characteristics of pancreatic CT images and designs a pancreatic segmentation network with the ability to capture and fuse multi-scale features, including: an encoder branch with dimensionality reduction grid residual convolution, a decoder branch with dimensionality increase grid residual convolution, a multi-scale feature capture module MFCM, a spatio-temporal attention module CSAM, and a post-processing module PMFM with pyramid multi-scale fusion; the multi-scale feature capture module MFCM is arranged after the encoder branch and is used to extract the multi-scale detailed information of the pancreas in the feature map of the encoder branch; the spatio-temporal attention module CSAM is arranged between the encoder branch and the decoder branch and uses the output features of each layer of the decoder branch to mine the complementary features of the previous layer. The post-processing module PMFM with multi-scale fusion is arranged after the decoder and is used to further recall the multi-scale details of the pancreatic features. It solves the problems of easy loss of small target information and edge information and unsatisfactory segmentation effect of the pancreas in the existing solutions.
[0007] To achieve the above object, the present invention specifically adopts the following technical solutions:
[0008] A pancreatic segmentation network based on multi-span complementary information capture and fusion, characterized by including: an encoder branch with dimensionality reduction grid residual convolution, a decoder branch with dimensionality increase grid residual convolution, a multi-scale feature capture module MFCM, a spatio-temporal attention module CSAM, and a post-processing module PMFM with pyramid multi-scale fusion;
[0009] The encoder-decoder is composed of dimensionality reduction grid residual convolution and dimensionality increase grid residual convolution;
[0010] The multi-scale feature capture module MFCM is arranged after the encoder branch and is used to extract the multi-scale detailed information of the pancreas in the feature map of the encoder branch;
[0011] The spatio-temporal attention module CSAM is arranged between the encoder branch and the decoder branch and uses the output features of each layer of the decoder branch to mine the complementary spatio-temporal features of the previous layer;
[0012] The post-processing module PMFM with multi-scale fusion is arranged after the decoder and is used to further recall the multi-scale details of the pancreatic features.
[0013] Further, the encoder is divided into four layers, each layer is composed of a dimensionality reduction grid residual convolution, the decoder is divided into four layers, each layer is composed of a dimensionality increase grid residual convolution, and each layer of the encoder branch is connected to the corresponding layer of the decoder branch through a skip connection.
[0014] Further, the multi-scale feature capture module MFCM is arranged after the single-layer branch of the encoder and is used to extract the multi-scale detailed information of the pancreas in the feature map of the encoder branch.
[0015] Furthermore, the spatio-temporal attention module CSAM is arranged between each layer of the multi-scale feature capture module MFCM and the convolutional block of each layer of the decoder branch, and uses the output features of each layer of the decoder branch to mine the complementary spatio-temporal features of the previous layer;
[0016] The decoder branch is followed by a post-processing module PMFM for pyramid multi-scale fusion. After PMFM, a 1×1 convolution and a softmax activation function are connected to obtain a prediction probability map.
[0017] Furthermore, in the dimensionality reduction grid residual convolution, each grid residual convolution module is composed of a total of 7 branches, including 3 vertical branches and 4 horizontal branches:
[0018] The vertical branch 1 is composed of a 1×1 convolution block with a stride of 1 for 1 time, a 3×3 convolution block with a stride of 1 for 2 times, and a 1×1 convolution block with a stride of 2 for 1 time. Each convolution block is composed of convolution, batch normalization, and relu activation function. The original feature map passes through branch 1 to obtain a feature Figure 1 ;
[0019] The vertical branch 2 is composed of a 1×1 convolution block with a stride of 1 for 1 time, a 5×5 convolution block with a stride of 1 for 2 times, and a 1×1 convolution block with a stride of 2 for 1 time. Each convolution block is composed of convolution, batch normalization, and relu activation function. The original feature map passes through branch 2 to obtain a feature Figure 2 ;
[0020] The vertical branch 3 is composed of a 1×1 convolution block with a stride of 1 for 1 time, a 7×7 convolution block with a stride of 1 for 2 times, and a 1×1 convolution block with a stride of 2 for 1 time. Each convolution block is composed of convolution, batch normalization, and relu activation function. The original feature map passes through branch 3 to obtain a feature Figure 3 ;
[0021] The horizontal branch 1 adds the three feature maps of the 1×1 convolution block with a stride of 1 in the first layer element by element to obtain a feature Figure 4 ;
[0022] The horizontal branch 2 adds the three feature maps of the 3×3 convolution block with a stride of 1, the 5×5 convolution block with a stride of 1, and the 7×7 convolution block with a stride of 1 in the second layer element by element to obtain a feature Figure 5 ;
[0023] The horizontal branch 3 adds the three feature maps of the 3×3 convolution block with a stride of 1, the 5×5 convolution block with a stride of 1, and the 7×7 convolution block with a stride of 1 in the third layer element by element to obtain a feature Figure 6 ;
[0024] The horizontal branch 4 adds the three feature maps of the 1×1 convolution block with a stride of 2 in the fourth layer element by element to obtain feature map 7;
[0025] Add the original feature map element-wise with the feature Figure 4 , 5, 6, perform a 1×1 convolution with a stride of 2 once, batch normalization, and relu activation function to complete feature dimensionality reduction to obtain feature map 8; subsequently, add feature map 8 element-wise with the feature Figure 1 , 2, 3 to obtain a new feature map. Multiple grid branches simultaneously capture the horizontal and vertical multi-scale information of the pancreas, and use grid residual convolution to retain more original detailed features.
[0026] Furthermore, in the upsampling grid residual convolution, each grid residual convolution module is composed of a total of 7 branches, including 3 vertical branches and 4 horizontal branches:
[0027] Vertical branch 1 is composed of a 1×1 transposed convolution block with a stride of 2 once, 2 3×3 convolution blocks with a stride of 1, and a 1×1 convolution block with a stride of 2. Each convolution block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 1 to obtain the feature Figure 1 ;
[0028] Vertical branch 2 is composed of a 1×1 transposed convolution block with a stride of 2 once, 2 5×5 convolution blocks with a stride of 1, and a 1×1 convolution block with a stride of 2. Each convolution block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 2 to obtain the feature Figure 2 ;
[0029] Vertical branch 3 is composed of a 1×1 transposed convolution block with a stride of 2 once, 2 7×7 convolution blocks with a stride of 1, and a 1×1 convolution block with a stride of 2. Each convolution block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 3 to obtain the feature Figure 3 ;
[0030] Horizontal branch 1 adds the three feature maps of the first-layer 1×1 transposed convolution block with a stride of 2 element-wise to obtain the feature Figure 4 ;
[0031] Horizontal branch 2 adds the three feature maps of the second-layer 3×3 convolution block with a stride of 1, 5×5 convolution block with a stride of 1, and 7×7 convolution block with a stride of 1 element-wise to obtain the feature Figure 5 ;
[0032] Horizontal branch 3 adds the three feature maps of the third-layer 3×3 convolution block with a stride of 1, 5×5 convolution block with a stride of 1, and 7×7 convolution block with a stride of 1 element-wise to obtain the feature Figure 6 ;
[0033] Horizontal branch 4 adds the three feature maps of the fourth-layer 1×1 convolution block with a stride of 2 element-wise to obtain feature map 7;
[0034] The original feature map is subjected to 1x1 transposed convolution with a stride of 2, batch normalization, and relu activation function to complete feature upsampling; subsequently, this feature map is added element-wise to the features Figure 4 , 5, 6 to obtain feature map 8, and then feature map 8 is added element-wise to the features Figure 1 , 2, 3 to obtain a new feature map. Multiple grid branches simultaneously capture the horizontal and vertical multi-scale information of the pancreas, and use grid residual convolution to retain more original detailed features.
[0035] Furthermore, in the multi-scale feature capture module MFCM, MFCM consists of 3 branches:
[0036] Branch 1 consists of a 3x3 convolution block with a stride of 1 and a 1x1 convolution block with a stride of 1. Each convolution block consists of convolution, batch normalization, and relu activation function. After the original feature map passes through a 1x1 convolution block with a stride of 1 to obtain a pre-classification feature map, it is added element-wise to the feature map obtained from branch 1 to obtain the feature Figure 1 ;
[0037] Branch 2 consists of a 5x5 convolution block with a stride of 1 and a 1x1 convolution block with a stride of 1. Each convolution block consists of convolution, batch normalization, and relu activation function. After the original feature map passes through a 1x1 convolution block with a stride of 1 to obtain a pre-classification feature map, it is added element-wise to the feature map obtained from branch 2 to obtain the feature Figure 2 ;
[0038] Branch 3 consists of a 7x7 convolution block with a stride of 1 and a 1x1 convolution block with a stride of 1. Each convolution block consists of convolution, batch normalization, and relu activation function. After the original feature map passes through a 1x1 convolution block with a stride of 1 to obtain a pre-classification feature map, it is added element-wise to the feature map obtained from branch 3 to obtain the feature Figure 3 ;
[0039] Subsequently, the features Figure 1 , 2, 3 are concatenated along the channel dimension, and after passing through a 1x1 convolution block with a stride of 1, the number of channels is restored to the size of the original feature map and added element-wise to the original feature map.
[0040] Furthermore, in the spatio-temporal attention module CSAM, CSAM consists of two branches:
[0041] In the spatial branch, the original feature map passes through a 3×3 convolutional block with a stride of 1 to obtain Mask1, passes through a 5×5 convolutional block with a stride of 1 to obtain Mask2, and passes through a 7×7 convolutional block with a stride of 1 to obtain Mask3. Then, Mask1, 2, and 3 are added element-wise to obtain the spatial Mask. The original feature map is multiplied element-wise with the spatial Mask to obtain the spatial detail recall feature map;
[0042] In the temporal branch, the original feature map passes through GlobalAveragePooling to obtain Mask4, passes through GlobalMaxPooling to obtain Mask5. Then, Mask4 and 5 are added element-wise to obtain the temporal Mask. The original feature map is multiplied element-wise with the temporal Mask to obtain the temporal detail recall feature map;
[0043] Finally, the original feature map, the spatial detail recall feature map, and the temporal detail recall feature map are added element-wise to obtain the spatio-temporal detail completion feature map.
[0044] Furthermore, in the post-processing module PMFM of the pyramid multi-scale fusion, PMFM is composed of 3 branches:
[0045] Branch 1 is composed of a 3×3 dilated convolutional block with a dilation rate of 3 for 1 time. After the original feature map is dimension-reduced by the dilated convolutional block, it is added element-wise to the input image to obtain the feature Figure 1 ;
[0046] Branch 2 is composed of a 3×3 dilated convolutional block with a dilation rate of 6 for 1 time. After the original feature map is dimension-reduced by the dilated convolutional block, it is added element-wise to the input image to obtain the feature Figure 2 ;
[0047] Branch 3 is composed of a 3×3 dilated convolutional block with a dilation rate of 9 for 1 time. After the original feature map is dimension-reduced by the dilated convolutional block, it is added element-wise to the input image to obtain the feature Figure 3 ;
[0048] Finally, the feature Figure 1 , 2, and 3 are concatenated with the original input feature map along the channel dimension to obtain a new feature map.
[0049] Furthermore, the loss function is a weighted edge detail recall loss function, and the weighted edge detail recall loss function is composed of a weighted cross-entropy loss function and an MIOU loss function:
[0050] Among them, the weight calculation formula is:
[0051]
[0052] Among them, Num c is the number of pixels of class c, Numall is the total number of pixels in a single slice, and σ is a hyperparameter;
[0053] The weight coefficient adaptively performs class weighting on the cross-entropy loss function according to the pixel ratio, increasing the relevant loss ratio of the pancreas and enabling the network to learn more target region features;
[0054] Among them, the calculation formula of the weighted loss function is:
[0055]
[0056] Among them, M is the total number of categories in the application scenario; p c is the probability that the sample belongs to category c. w c is the weight value of category c, and the value range is [1, 70]; y c is an indicator variable, which is 1 if the category of the sample is the same as its corresponding category, otherwise it is 0;
[0057] Among them, the calculation formula of the MIOU loss function is:
[0058]
[0059] LMIoU = 1 - w c CMIoU
[0060] In the formula, k is the total number of categories in the application scenario, and y ij is the true marked value processed manually at the j-th pixel position in the i-th category, is the predicted marked value of the network model at the j-th pixel position in the i-th category;
[0061] The calculation formula of the weighted edge detail recall loss function is:
[0062] WERL = α * WCE + β * (1 - CMIoU)
[0063] Among them, α is the weighted cross-entropy coefficient, and β is the MIoU loss coefficient, set as α = 0.3 and β = 0.7.
[0064] Compared with the prior art, the present invention and its preferred solutions utilize the characteristics of pancreatic CT images and design a pancreatic segmentation network with the ability to capture and fuse multi-scale features, including: an encoder branch with a dimensionality reduction grid residual convolution, a decoder branch with a dimensionality increase grid residual convolution, a multi-scale feature capture module MFCM, a spatio-temporal attention module CSAM, and a post-processing module PMFM with pyramid multi-scale fusion; the multi-scale feature capture module MFCM is arranged after the encoder branch and is used to extract the multi-scale detailed information of the pancreas in the feature map of the encoder branch; the spatio-temporal attention module CSAM is arranged between the encoder branch and the decoder branch and uses the output features of each layer of the decoder branch to mine the complementary features of the previous layer. The post-processing module PMFM with multi-scale fusion is arranged after the decoder and is used to further recall the multi-scale details of the pancreatic features. It solves the problems in the existing solutions that small target information and edge information are easily lost and the segmentation effect of the pancreas is not ideal. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 FIG. is a structural diagram of a pancreatic segmentation network based on multi-span complementary information capture and fusion provided by an embodiment of the present invention;
[0066] Figure 2 FIG. is a schematic diagram of a dimensionality reduction grid residual convolution structure provided by an embodiment of the present invention;
[0067] Figure 3 FIG. is a schematic diagram of a dimensionality increase grid residual convolution structure provided by an embodiment of the present invention;
[0068] Figure 4 FIG. is a schematic diagram of the structure of the multi-scale feature capture module MFCM provided by an embodiment of the present invention;
[0069] Figure 5 FIG. is a schematic diagram of the structure of the spatio-temporal attention module CSAM provided by an embodiment of the present invention;
[0070] Figure 6 FIG. is a schematic diagram of the structure of the post-processing module PMFM with pyramid multi-scale fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0072] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0073] AsFigures 1 - 6 As shown, the pancreatic segmentation network based on multi-span complementary information capture and fusion provided in this embodiment includes: an encoder branch with dimensionality reduction grid residual convolution, a decoder branch with dimensionality increase grid residual convolution, a multiscale feature capture module (Multiscale Feature Capture Module, MFCM), a channel spatial attention module (Channel Spatial Attention Module, CSAM), and a post-processing module with pyramid multiscale fusion (Pyramid Multiscale Fusion Module, PMFM).
[0074] The encoder-decoder consists of a reduced-dimensional grid residual convolution and an increased-dimensional grid residual convolution;
[0075] The multi-scale feature capture module MFCM is set after the encoder branch and is used to extract the multi-scale detail information of the pancreas in the encoder branch feature map;
[0076] The spatiotemporal attention module CSAM is set between the encoder branch and the decoder branch, and uses the output features of each layer of the decoder branch to mine the complementary spatiotemporal features of the previous layer;
[0077] The multi-scale fusion post-processing module PMFM is set after the decoder to further recall the multi-scale details of the pancreatic features.
[0078] This embodiment uses UNet as the basic framework of the pancreatic segmentation network. The encoder is divided into four layers, each layer consists of a dimension reduction grid residual convolution, and the decoder is divided into four layers, each layer consists of a dimension increase grid residual convolution. Each layer of the encoder branch is connected to the corresponding layer of the decoder branch through a jump connection.
[0079] The multi-scale feature capture module MFCM is set after the encoder single-layer branch to extract the multi-scale detail information of the pancreas in the encoder branch feature map;
[0080] The spatiotemporal attention module CSAM is set after each layer of multi-scale feature capture module MFCM and between the convolution blocks of each layer of decoder branch, and uses the output features of each layer of the decoder branch to mine the complementary spatiotemporal features of the previous layer.
[0081] The decoder branch is followed by a pyramid multi-scale fusion post-processing module PMFM, which is followed by a 1×1 convolution and a Softmax activation function to obtain a predicted probability map.
[0082] Each grid residual convolution module consists of 7 branches, including 3 vertical branches and 4 horizontal branches:
[0083] The vertical branch 1 consists of one 1×1 convolution block (stride 1), two 3×3 convolution blocks (stride 1), and one 1×1 convolution block (stride 2). Each convolution block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 1 to obtain features Figure 1 ;
[0084] The vertical branch 2 consists of one 1×1 convolution block (stride 1), two 5×5 convolution blocks (stride 1), and one 1×1 convolution block (stride 2). Each convolution block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 2 to obtain features Figure 2 ;
[0085] The vertical branch 3 consists of one 1×1 convolution block (stride 1), two 7×7 convolution blocks (stride 1), and one 1×1 convolution block (stride 2). Each convolution block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 3 to obtain features Figure 3 ;
[0086] The horizontal branch 1 adds the three feature maps of the first-layer 1×1 convolution block (stride 1) element-wise to obtain features Figure 4 ;
[0087] The horizontal branch 2 adds the three feature maps of the second-layer 3×3 convolution block (stride 1), 5×5 convolution block (stride 1), and 7×7 convolution block (stride 1) element-wise to obtain features Figure 5 ;
[0088] The horizontal branch 3 adds the three feature maps of the third-layer 3×3 convolution block (stride 1), 5×5 convolution block (stride 1), and 7×7 convolution block (stride 1) element-wise to obtain features Figure 6 ;
[0089] The horizontal branch 4 adds the three feature maps of the fourth-layer 1×1 convolution block (stride 2) element-wise to obtain feature map 7;
[0090] The original feature map is added element-wise with the features Figure 4 , 5, 6, and after one 1×1 convolution (stride 2), batch normalization, and relu activation function, feature dimension reduction is completed to obtain feature map 8. Subsequently, feature map 8 is added element-wise with the features Figure 1 , 2, 3 to obtain a new feature map. Multiple grid branches simultaneously capture the horizontal and vertical multi-scale information of the pancreas and use grid residual convolution to retain more original detailed features.
[0091] Each grid residual convolution module consists of a total of seven branches, including three vertical branches and four horizontal branches:
[0092] The vertical branch 1 consists of a 1×1 transposed convolution block (stride 2) once, 3×3 convolution blocks (stride 1) twice, and a 1×1 convolution block (stride 2) once. Each convolution block consists of convolution, batch normalization, and a relu activation function. The original feature map passes through branch 1 to obtain features Figure 1 ;
[0093] The vertical branch 2 consists of a 1×1 transposed convolution block (stride 2) once, 5×5 convolution blocks (stride 1) twice, and a 1×1 convolution block (stride 2) once. Each convolution block consists of convolution, batch normalization, and a relu activation function. The original feature map passes through branch 2 to obtain features Figure 2 ;
[0094] The vertical branch 3 consists of a 1×1 transposed convolution block (stride 2) once, 7×7 convolution blocks (stride 1) twice, and a 1×1 convolution block (stride 2) once. Each convolution block consists of convolution, batch normalization, and a relu activation function. The original feature map passes through branch 3 to obtain features Figure 3 ;
[0095] The horizontal branch 1 adds the three feature maps of the first 1×1 transposed convolution block (stride 2) element-wise to obtain features Figure 4 ;
[0096] The horizontal branch 2 adds the three feature maps of the second 3×3 convolution block (stride 1), 5×5 convolution block (stride 1), and 7×7 convolution block (stride 1) element-wise to obtain features Figure 5 ;
[0097] The horizontal branch 3 adds the three feature maps of the third 3×3 convolution block (stride 1), 5×5 convolution block (stride 1), and 7×7 convolution block (stride 1) element-wise to obtain features Figure 6 ;
[0098] The horizontal branch 4 adds the three feature maps of the fourth 1×1 convolution block (stride 2) element-wise to obtain feature map 7;
[0099] The original feature map undergoes 1×1 transposed convolution (stride 2), batch normalization, and a relu activation function to complete feature upsampling. Subsequently, this feature map is added element-wise to the features Figure 4 , 5, 6 to obtain feature map 8. Subsequently, feature map 8 is added element-wise to the features Figure 1 , 2, 3 to obtain a new feature map. Multiple grid branches simultaneously capture the horizontal and vertical multi-scale information of the pancreas and use grid residual convolution to retain more original detailed features.
[0100] The multi-scale feature capture module MFCM consists of 3 branches:
[0101] Branch 1 consists of a 3×3 convolutional block (stride 1) and a 1×1 convolutional block (stride 1). Each convolutional block consists of convolution, batch normalization, and a relu activation function. After the original feature map passes through a 1×1 convolutional block (stride 1) to obtain a pre-classification feature map, it is added element-wise to the feature map obtained from Branch 1 to get the feature. Figure 1 ;
[0102] Branch 2 consists of a 5×5 convolutional block (stride 1) and a 1×1 convolutional block (stride 1). Each convolutional block consists of convolution, batch normalization, and a relu activation function. After the original feature map passes through a 1×1 convolutional block (stride 1) to obtain a pre-classification feature map, it is added element-wise to the feature map obtained from Branch 2 to get the feature. Figure 2 ;
[0103] Branch 3 consists of a 7×7 convolutional block (stride 1) and a 1×1 convolutional block (stride 1). Each convolutional block consists of convolution, batch normalization, and a relu activation function. After the original feature map passes through a 1×1 convolutional block (stride 1) to obtain a pre-classification feature map, it is added element-wise to the feature map obtained from Branch 3 to get the feature. Figure 3 ;
[0104] Subsequently, the features Figure 1 , 2, 3 are concatenated along the channel dimension, and after passing through a 1×1 convolutional block (stride 1), the number of channels is restored to the size of the original feature map and added element-wise to the original feature map.
[0105] The spatio-temporal attention module CSAM consists of two branches:
[0106] In the spatial branch, the original feature map passes through a 3×3 convolutional block (stride 1) to get Mask1, through a 5×5 convolutional block (stride 1) to get Mask2, and through a 7×7 convolutional block (stride 1) to get Mask3. Then Mask1, 2, 3 are added element-wise to get the spatial Mask, and the original feature map is multiplied element-wise by the spatial Mask to get the spatial detail recall feature map;
[0107] In the temporal branch, the original feature map passes through GlobalAveragePooling to get Mask4, through GlobalMaxPooling to get Mask5. Then Mask4, 5 are added element-wise to get the temporal Mask, and the original feature map is multiplied element-wise by the temporal Mask to get the temporal detail recall feature map.
[0108] Finally, the original feature map, the spatial detail recall feature map, and the temporal detail recall feature map are added element-wise to get the spatio-temporal detail completion feature map.
[0109] The post - processing module PMFM for pyramid multi - scale fusion consists of 3 branches:
[0110] Branch 1 consists of a single 3×3 dilated convolutional block (dilation rate = 3). After the original feature map is dimension - reduced by the dilated convolutional block, it is added element - by - element to the input image to obtain a feature Figure 1 ;
[0111] Branch 2 consists of a single 3×3 dilated convolutional block (dilation rate = 6). After the original feature map is dimension - reduced by the dilated convolutional block, it is added element - by - element to the input image to obtain a feature Figure 2 ;
[0112] Branch 3 consists of a single 3×3 dilated convolutional block (dilation rate = 9). After the original feature map is dimension - reduced by the dilated convolutional block, it is added element - by - element to the input image to obtain a feature Figure 3 ;
[0113] Finally, the features Figure 1 , 2, 3 are concatenated with the original input feature map along the channel dimension to obtain a new feature map.
[0114] The loss function of the pancreas segmentation network is the weighted edge - detail recall loss function, which is composed of the weighted cross - entropy loss function and the MIOU loss function:
[0115] Among them, the weight calculation formula is:
[0116]
[0117] Among them, Num c is the number of pixels of class c, Num all is the total number of pixels in a single slice, and σ is a hyper - parameter, which is generally set to 1.02 through experimental verification. This weight coefficient can adaptively perform class weighting on the cross - entropy loss function according to the pixel ratio, and the relevant loss ratio of the pancreas will be appropriately increased, enabling the network to learn more target - area features.
[0118] Among them, the weighted loss function calculation formula is:
[0119]
[0120] Among them, M is the total number of classes in the application scenario; p c is the probability that the sample belongs to class c. w c is the weight value of class c, and the value range is restricted to [1, 70]; y c is an indicator variable, which is 1 if the class of the sample is the same as its corresponding class, otherwise it is 0;
[0121] Among them, the MIOU loss function calculation formula is:
[0122]
[0123] LMIoU = 1 - CMIoU
[0124] In the formula, k is the total number of categories in the application scenario, and y ij is the true marked value processed manually at the j-th pixel position in the i-th category, and is the predicted marked value of the network model at the j-th pixel position in the i-th category.
[0125] The calculation formula of the weighted edge detail recall loss function is as follows:
[0126] WERL = α * WCE + β * (1 - CMIoU)
[0127] Among them, α is the weighted cross-entropy coefficient, and β is the MIoU loss coefficient, which are set as α = 0.3 and β = 0.7.
[0128] Based on the above model design, the experiment designed in this embodiment uses the publicly available dataset of the Medical Image Decathlon for training. The Task07_Pancreas publicly available dataset consists of 420 3D CT scan sequences, and the number of examples with true marks is 281. The spatial resolution is equal to 512×512 pixels, and the number of slices is between 181 and 466. 281 examples of data are randomly divided into 266 for training, 5 for validation, and 10 for testing.
[0129] The slice CT value is intercepted to [-100, 240] and normalized to [0, 1] through HU constraint, so as to roughly remove irrelevant tissues. The pre-trained unet network is used to obtain the rough segmentation mask of the pancreas. The mask information of multiple slices in the same batch is used to extract the central point position of the pancreas in the whole batch, and the whole batch is cropped to a size of 224*224. The same cropping strategy is used in the verification process.
[0130] The network is trained on an NVIDIA GeForce RTX 3070 GPU with 8GB. The epoch is set to 100, the batch_size is set to 32, and the Adam algorithm is used to optimize the network parameters. The initial learning rate is 0.001. The model with the best performance on the validation set is selected as the final version.
[0131] As described above, it is only the preferred embodiment of the present invention, and is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
[0132] This patent is not limited to the above best implementation manner. Anyone inspired by this patent can obtain various other forms of pancreatic segmentation networks based on multi-span complementary information capture and fusion. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. A pancreatic segmentation network based on multi-span complementary information capture and fusion, characterized in that, it includes: an encoder branch with a dimensionality reduction grid residual convolution, a decoder branch with a dimensionality increase grid residual convolution, a multi-scale feature capture module MFCM, a spatio-temporal attention module CSAM, and a post-processing module PMFM with pyramid multi-scale fusion; The encoder-decoder structure is composed of a dimensionality reduction grid residual convolution and a dimensionality increase grid residual convolution; The multi-scale feature capture module MFCM is arranged after the encoder branch and is used to extract the multi-scale detailed information of the pancreas in the feature map of the encoder branch; The spatio-temporal attention module CSAM is arranged between the encoder branch and the decoder branch, and uses the output features of each layer of the decoder branch to mine the complementary spatio-temporal features of the previous layer; The post-processing module PMFM with multi-scale fusion is arranged after the decoder and is used to further recall the multi-scale details of the pancreatic features; In the multi-scale feature capture module MFCM, MFCM is composed of 3 branches: Branch 1 is composed of a 3×3 convolution block with a stride of 1 and a 1×1 convolution block with a stride of 1. Each convolution block is composed of convolution, batch normalization, and relu activation function. After the original feature map obtains a pre-classification feature map through a 1×1 convolution block with a stride of 1, it is added element-wise to the feature map obtained by Branch 1 to obtain Feature Map 1; Branch 2 is composed of a 5×5 convolution block with a stride of 1 and a 1×1 convolution block with a stride of 1. Each convolution block is composed of convolution, batch normalization, and relu activation function. After the original feature map obtains a pre-classification feature map through a 1×1 convolution block with a stride of 1, it is added element-wise to the feature map obtained by Branch 2 to obtain Feature Map 2; Branch 3 is composed of a 7×7 convolution block with a stride of 1 and a 1×1 convolution block with a stride of 1. Each convolution block is composed of convolution, batch normalization, and relu activation function. After the original feature map obtains a pre-classification feature map through a 1×1 convolution block with a stride of 1, it is added element-wise to the feature map obtained by Branch 3 to obtain Feature Map 3; Subsequently, Feature Maps 1, 2, and 3 are concatenated along the channel dimension, and after passing through a 1×1 convolution block with a stride of 1, the number of channels is restored to the size of the original feature map and added element-wise to the original feature map; In the spatio-temporal attention module CSAM, CSAM is composed of two branches: In the spatial branch, the original feature map passes through a 3×3 convolution block with a stride of 1 to obtain Mask1, passes through a 5×5 convolution block with a stride of 1 to obtain Mask2, passes through a 7×7 convolution block with a stride of 1 to obtain Mask3. Then, Mask1, 2, and 3 are added element-wise to obtain a spatial Mask, and the original feature map is multiplied element-wise by the spatial Mask to obtain a spatial detail recall feature map; In the temporal branch, the original feature map passes through GlobalAveragePooling to obtain Mask4, and through GlobalMaxPooling to obtain Mask5. Then, Mask4 and Mask5 are added element-wise to obtain the temporal Mask. The original feature map is multiplied element-wise with the temporal Mask to obtain the temporally detailed recalled feature map; Finally, the original feature map, the spatially detailed recalled feature map, and the temporally detailed recalled feature map are added element-wise to obtain the spatio-temporal detailed complemented feature map; In the post-processing module PMFM of pyramid multi-scale fusion, PMFM consists of three branches: Branch 1 consists of a 3×3 dilated convolutional block with a dilation rate of 3 for 1 time. After the original feature map is dimension-reduced by the dilated convolutional block, it is added element-wise to the input image to obtain Feature Map 1; Branch 2 consists of a 3×3 dilated convolutional block with a dilation rate of 6 for 1 time. After the original feature map is dimension-reduced by the dilated convolutional block, it is added element-wise to the input image to obtain Feature Map 2; Branch 3 consists of a 3×3 dilated convolutional block with a dilation rate of 9 for 1 time. After the original feature map is dimension-reduced by the dilated convolutional block, it is added element-wise to the input image to obtain Feature Map 3; Finally, Feature Maps 1, 2, and 3 and the original input feature map are concatenated along the channel dimension to obtain a new feature map.
2. The pancreatic segmentation network based on multi-span complementary information capture and fusion according to claim 1, characterized in that: The encoder is divided into four layers, each layer consists of a dimension-reducing grid residual convolution, and the decoder is divided into four layers, each layer consists of a dimension-increasing grid residual convolution. Each layer of the encoder branch is connected to the corresponding layer of the decoder branch through a skip connection.
3. The pancreatic segmentation network based on multi-span complementary information capture and fusion according to claim 2, characterized in that: The multi-scale feature capture module MFCM is arranged after the single-layer branch of the encoder and is used to extract the multi-scale detailed information of the pancreas in the feature map of the encoder branch.
4. The pancreatic segmentation network based on multi-span complementary information capture and fusion according to claim 2, characterized in that: The spatio-temporal attention module CSAM is arranged between each layer of the multi-scale feature capture module MFCM and the convolutional block of each layer of the decoder branch, and uses the output feature of each layer of the decoder branch to mine the complementary spatio-temporal features of the previous layer; The decoder branch is followed by a post-processing module PMFM of pyramid multi-scale fusion. After PMFM, a 1×1 convolution and a softmax activation function are connected to obtain a prediction probability map.
5. The pancreatic segmentation network based on multi-span complementary information capture and fusion according to claim 2, characterized in that: In the dimension-reducing grid residual convolution, each grid residual convolution module consists of a total of 7 branches, 3 vertical branches and 4 horizontal branches: The vertical branch 1 consists of a 1×1 convolutional block with a stride of 1 for 1 time, a 3×3 convolutional block with a stride of 1 for 2 times, and a 1×1 convolutional block with a stride of 2 for 1 time. Each convolutional block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 1 to obtain Feature Map 1; The vertical branch 2 is composed of a 1×1 convolutional block with a stride of 1 for 1 time, a 5×5 convolutional block with a stride of 1 for 2 times, and a 1×1 convolutional block with a stride of 2 for 1 time. Each convolutional block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 2 to obtain feature map 2; The vertical branch 3 is composed of a 1×1 convolutional block with a stride of 1 for 1 time, a 7×7 convolutional block with a stride of 1 for 2 times, and a 1×1 convolutional block with a stride of 2 for 1 time. Each convolutional block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 3 to obtain feature map 3; The horizontal branch 1 adds the three feature maps of the first-layer 1×1 convolutional block with a stride of 1 element-wise to obtain feature map 4; The horizontal branch 2 adds the three feature maps of the second-layer 3×3 convolutional block with a stride of 1, the 5×5 convolutional block with a stride of 1, and the 7×7 convolutional block with a stride of 1 element-wise to obtain feature map 5; The horizontal branch 3 adds the three feature maps of the third-layer 3×3 convolutional block with a stride of 1, the 5×5 convolutional block with a stride of 1, and the 7×7 convolutional block with a stride of 1 element-wise to obtain feature map 6; The horizontal branch 4 adds the three feature maps of the fourth-layer 1×1 convolutional block with a stride of 2 element-wise to obtain feature map 7; The original feature map is added to feature maps 4, 5, and 6 element-wise, and after 1×1 convolution with a stride of 2, batch normalization, and relu activation function, feature dimension reduction is completed to obtain feature map 8; Subsequently, feature map 8 is added to feature maps 1, 2, and 3 element-wise to obtain a new feature map. Multiple grid branches simultaneously capture the horizontal and vertical multi-scale information of the pancreas and use grid residual convolution to retain more original detailed features.
6. The pancreas segmentation network based on multi-span complementary information capture and fusion according to claim 1, characterized in that: In the upsampling grid residual convolution, each grid residual convolution module is composed of a total of 7 branches, including 3 vertical branches and 4 horizontal branches: The vertical branch 1 is composed of a 1×1 transposed convolutional block with a stride of 2 for 1 time, a 3×3 convolutional block with a stride of 1 for 2 times, and a 1×1 convolutional block with a stride of 2 for 1 time. Each convolutional block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 1 to obtain feature map 1; The vertical branch 2 is composed of a 1×1 transposed convolutional block with a stride of 2 for 1 time, a 5×5 convolutional block with a stride of 1 for 2 times, and a 1×1 convolutional block with a stride of 2 for 1 time. Each convolutional block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 2 to obtain feature map 2; The vertical branch 3 is composed of a 1×1 transposed convolutional block with a stride of 2 for 1 time, a 7×7 convolutional block with a stride of 1 for 2 times, and a 1×1 convolutional block with a stride of 2 for 1 time. Each convolutional block consists of convolution, batch normalization, and relu activation function. The original feature map passes through branch 3 to obtain feature map 3; The horizontal branch 1 adds the three feature maps of the first-layer 1×1 transposed convolutional block with a stride of 2 element-wise to obtain feature map 4; The horizontal branch 2 adds the three feature maps of the second-layer 3×3 convolutional block with a stride of 1, the 5×5 convolutional block with a stride of 1, and the 7×7 convolutional block with a stride of 1 element-wise to obtain feature map 5; The horizontal branch 3 adds the feature maps of the 3×3 convolution block with a stride of 1, the 5×5 convolution block with a stride of 1, and the 7×7 convolution block with a stride of 1 in the third layer element-wise to obtain the feature map 6; The horizontal branch 4 adds the feature maps of the three feature maps of the 1×1 convolution block with a stride of 2 in the fourth layer element-wise to obtain the feature map 7; The original feature map is upsampled through 1 time of 1×1 transposed convolution with a stride of 2, followed by batch normalization and the relu activation function to complete feature upscaling; Subsequently, this feature map is added to the feature maps 4, 5, and 6 element-wise to obtain the feature map 8. Then, the feature map 8 is added to the feature maps 1, 2, and 3 element-wise to obtain a new feature map. Multiple grid branches simultaneously capture the horizontal and vertical multi-scale information of the pancreas and use grid residual convolution to retain more original detailed features.
7. The pancreas segmentation network based on multi-span complementary information capture and fusion according to claim 4, characterized in that: The loss function is a weighted edge detail recall loss function, and the weighted edge detail recall loss function is composed of a weighted cross-entropy loss function and a MIOU loss function: where the weight calculation formula is: where Num c is the number of pixels of class c, and Num all is the total number of pixels in a single slice, and σ is a hyperparameter; The weight adaptively performs class weighting on the cross-entropy loss function according to the pixel ratio, increasing the relevant loss ratio of the pancreas, so that the network learns more target region features; where the weighted loss function calculation formula is: where M is the total number of categories in the application scenario; p c is the probability that the sample belongs to category c; w c is the weight value of category c, and the value range is [1, 70]; y c is an indicator variable, which is 1 if the category of the sample is the same as its corresponding category, otherwise it is 0; where the MIOU loss function calculation formula is: LMIoU = 1 - w c CMIoU where k is the total number of categories in the application scenario, and y ij is the true marked value processed manually at the j-th pixel position in the i-th category, and is the predicted marked value of the network model at the j-th pixel position in the i-th category; The weighted edge detail recall loss function calculation formula is: WERL = α * WCE + β * (1 - CMIoU) where α is the weighted cross-entropy coefficient and β is the MIoU loss coefficient, set as α = 0.3 and β = 0.7.
Citation Information
Patent Citations
Image segmentation method for kidney tumor
CN112085743A
Pancreas segmentation network in CT image based on improved U-shaped network
CN114119448A