Pancreas segmentation method based on WEUnet
By adopting a WEUnet network-based method in pancreatic segmentation, using the wavelet convolution submodule and edge feature attention module, the problems of unclear edges and imbalance in pancreatic segmentation are solved, and a high-precision pancreatic segmentation effect is achieved.
Patent Information
- Application Number
- CN202510211478.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art has problems of unclear edges, large deformations and unbalanced categories in pancreatic segmentation, resulting in low segmentation accuracy and inability to apply to clinical practice.
The pancreatic segmentation method based on WEUnet network is adopted, and the wavelet convolution submodule is used to fuse low-frequency and high-frequency information, which enhances the network's attention to the pancreas edge, and introduces a marginal feature attention module into the jump connection to improve boundary recognition capabilities.
The problems of unclear edges and imbalance in pancreatic segmentation were effectively solved, segmentation indicators were improved, and an average DSC score of 87.61±2.34% was achieved, which performed well on the NIH dataset.
Smart Images

Figure CN120031898A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of pancreas segmentation, and in particular relates to a pancreas segmentation method based on a WEUnet network. Background Art
[0002] In recent years, with the rapid development of deep learning and neural networks, abdominal organs such as the liver, lungs, and kidneys can be effectively segmented by artificial intelligence-based systems. However, the accuracy of pancreas segmentation is still not high enough to be applied in clinical practice. The main difficulties come from the following aspects:
[0003] 1) The pancreas occupies only a small part of the entire CT image. The size of the target and background is seriously unbalanced, which makes it easy for the network to focus on non-target background areas, resulting in overfitting problems in the background area. 2) The shape of the pancreas varies greatly. The shape of the pancreas is irregular and easily deformed. The shape, size, and position of the pancreas in the abdomen of different patients vary greatly. 3) The boundary of the pancreas is not clear. Since the pancreas has a similar density to the surrounding tissues (duodenum, small intestine, etc.), the contrast between the pancreas and the surrounding tissues in the CT image is weak, and there is a problem of boundary disturbance, which makes the boundary not well defined.
[0004] In view of these problems, it is difficult for traditional machine learning image segmentation algorithms to accurately segment the pancreatic border with blurred edges. At present, the main technology for medical image segmentation is based on data-driven feature learning. One of the commonly used neural network architectures is the convolutional neural network (CNN) U-Net. It consists of three parts: encoder, decoder and jump connection. The encoder is used to extract high-level semantic features of the image while gradually reducing the spatial resolution. The decoder is used to gradually restore the spatial resolution of the feature map while retaining the semantic information extracted by the encoder. The role of the jump connection is to directly pass the features of the corresponding layer in the encoder to the decoder, fuse the low-level and high-level features, and make up for the information loss.
[0005] Although skip connections combined with upsampling work well, they still fall short in achieving optimal model performance. To address this problem, researchers have used a variety of strategies. Non-local neural networks by Hu and Wang et al. deeply studied the spatial attention mechanism, enabling the neural network to grasp long-range dependencies and complex spatial structures in images. Li et al. outlined a multi-scale attention dense residual network, which enhances the model's ability to accurately locate the nuances of the pancreas by combining dense residual blocks and multi-scale convolutional kernels. Yan et al. proposed an architecture that combines 2D and 3D convolutional layers, which reduces computational requirements while preserving spatial information. Chen et al. introduced fuzzy logic in the skip connection of U-Net to improve the model's sensitivity to small variable structures. Wang et al. added an attention-guided control mechanism to the skip connection to improve the model's ability to recognize small structures. Although the attention mechanism has advantages in handling multi-scale objects, it is still limited by the kernel size and faces challenges in capturing the characteristics of a slender organ such as the pancreas. Summary of the invention
[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a pancreas segmentation method based on the WEUnet network to solve the problem of unclear edges in pancreas segmentation.
[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a pancreas segmentation method based on WEUnet network, comprising the following steps:
[0008] S1, obtain abdominal CT images and input the abdominal CT images into the WEUnet network;
[0009] S2, performing convolution encoding processing on the abdominal CT image five times in a row through the encoding module to obtain the encoded visual feature map output by each layer of the encoding module;
[0010] S3, inputting the lowest level encoded visual feature map into a decoding module to obtain the lowest level decoded prediction feature map;
[0011] S4, the decoding module performs four consecutive convolution decoding processes on the decoding prediction feature map of the bottom layer to obtain the decoding prediction feature map of the top layer, and in each decoding process, the edge feature attention module connected by the jump connection calculates the intermediate feature map according to the encoded visual feature map of the current layer, the decoding prediction feature map of the next layer and the high-frequency feature map of the current layer, and inputs the intermediate feature map into the decoding module, and the generated result is spliced on the generated decoding prediction feature map of the current layer;
[0012] S5. Obtain a pancreatic image based on the decoded prediction feature map of the top layer.
[0013] Further: in said S2, the encoding module includes a wavelet convolution submodule, a depth convolution submodule and a point-by-point convolution submodule connected in sequence;
[0014] Among them, the output end of the wavelet convolution submodule is also connected to the input end of the first convolution submodule and the input end of the first normalization layer respectively, the output end of the point-by-point convolution submodule, the output end of the first convolution submodule and the output end of the first normalization layer are all connected to the input end of the first element addition, the output end of the first element addition is respectively connected to the input end of the second convolution submodule and the input end of the second normalization layer through the first ReLu activation layer, the output end of the second convolution submodule and the output end of the second normalization layer are both connected to the input end of the second element addition, and the output end of the second element addition is connected to the input end of the second ReLu activation layer;
[0015] The wavelet convolution submodule includes a wavelet convolution, a third normalization layer, and a third ReLu activation layer connected in sequence;
[0016] The structure of the depthwise convolution submodule is the same as that of the pointwise convolution submodule, both of which include the first 3×3 convolution and the fourth normalization layer connected to each other;
[0017] The first convolution submodule and the second convolution submodule have the same structure, both including a first 1×1 convolution and a fifth normalization layer connected to each other.
[0018] Further: In S2, the method of encoding the visual feature map output by the encoding module is specifically:
[0019] S21, obtaining a first input feature map of an input encoding module, and inputting the first input feature map into a wavelet convolution submodule to obtain a first feature map;
[0020] S22, inputting the first feature map into the depth convolution submodule and the point-by-point convolution submodule in sequence, and using the residual branch to input the first feature map into the first convolution submodule and the first normalization layer respectively, and adding the outputs of the point-by-point convolution submodule, the first convolution submodule and the first normalization layer by the first element to obtain a second feature map;
[0021] S23, passing the second feature map through the first ReLu activation layer and the second convolution submodule in sequence, and inputting the second feature map into the second normalization layer using the residual branch, and adding the outputs of the second convolution submodule and the second normalization layer through the second element to obtain a third feature map;
[0022] S24. Input the third feature map into the second ReLu activation layer to obtain an encoded visual feature map.
[0023] Further: S21 includes the following sub-steps:
[0024] S211, generating a first low-frequency component and a first high-frequency component by wavelet transforming the first input feature map;
[0025] S212, generating a second low-frequency component and a second high-frequency component by wavelet transforming the first low-frequency component;
[0026] S213, performing convolution and transposition convolution processing on the second low-frequency component and the second high-frequency component in sequence, and taking all processed components as the convolution result of the second wavelet transform;
[0027] S214, adding the convolution result of the second wavelet transform to the first low-frequency component and the first high-frequency component, and sequentially performing convolution and transposed convolution processing on the result of the addition, and using all the processed components as the convolution result of the first wavelet transform;
[0028] S215: Add the convolution result of the first wavelet transform to the first input feature map, and perform convolution processing on the result of the addition to obtain the first feature map.
[0029] Further: In S3, the decoding module includes a third convolution submodule, a fourth ReLu activation layer, and a fourth convolution submodule connected in sequence, the input end of the third convolution submodule is also connected to the input end of the fifth convolution submodule, the output end of the fourth convolution submodule and the output end of the fifth convolution submodule are both connected to the input end of the third element addition, and the output end of the third element addition is sequentially connected to the fifth ReLu activation layer and the upsampling operation layer;
[0030] The third convolution submodule and the fourth convolution submodule have the same structure, both including a second 3×3 convolution and a sixth normalization layer connected to each other;
[0031] The fifth convolution submodule includes a second 1×1 convolution and a seventh normalization layer connected to each other.
[0032] Further: In S3, the method of processing the second input feature map input by the decoding module is specifically:
[0033] S31, obtaining a second input feature map of the input decoding module, and inputting the second input feature map into the third convolution submodule, the fourth ReLu activation layer and the fourth convolution submodule;
[0034] S32, using the residual branch to input the second input feature map into the fifth convolution submodule;
[0035] S33. The outputs of the fourth convolution submodule and the fifth convolution submodule are added together through the third element to obtain a fourth feature map, and the fourth feature map is sequentially input into the fifth ReLu activation layer and the upsampling operation layer to obtain the result of the second input feature map processed by the decoding module.
[0036] Further: in said S4, the edge feature attention module is provided with a convolutional attention submodule, and the convolutional attention submodule includes a spatial attention submodule interconnected with a channel attention submodule;
[0037] Among them, the channel attention submodule is provided with a first convolution of a 1×1×C convolution kernel, and the spatial attention submodule is provided with a second convolution of a H×W×1 convolution kernel.
[0038] Further: in said S4, the high frequency feature map is obtained by DWT processing of the abdominal CT image;
[0039] The method for calculating the intermediate feature map during each decoding process includes the following steps:
[0040] S41, obtaining a coded visual feature map of the current layer, a high-frequency feature map of the current layer, and a decoded prediction feature map of the next layer, multiplying the high-frequency feature map of the current layer and the decoded prediction feature map of the next layer by the coded visual feature map of the current layer element by element, concatenating the generated feature map with the coded visual feature map of the current layer to generate a fused feature map;
[0041] S42, inputting the fused feature map into the channel attention submodule, using the first branch to perform maximum pooling and multi-layer perceptron processing on the fused feature map in sequence, and optimizing it through the first convolution after the maximum pooling and multi-layer perceptron processing to generate a first attention feature map;
[0042] S43, using the second branch to perform average pooling and multi-layer perceptron processing on the fusion feature map in sequence, and optimizing it through the first convolution after the average pooling and multi-layer perceptron processing to generate a second attention feature map;
[0043] S44, adding the first attention feature map and the second attention feature map, subjecting the addition result to Sigmoid activation function processing and first convolution optimization in sequence, multiplying the optimization result with the fusion feature map to generate a third attention feature map;
[0044] S45, input the third attention feature map into the spatial attention submodule, use the third branch to perform maximum pooling and second convolution optimization on the third attention feature map, and generate a fourth attention feature map; use the fourth branch to perform average pooling and second convolution optimization on the third attention feature map, and generate a fifth attention feature map;
[0045] S46. Fuse the fourth attention feature map and the fifth attention feature map. Perform the second convolution optimization, Sigmoid activation function processing and the second convolution optimization on the fusion result in turn. Multiply the optimization result with the third attention feature map to generate an intermediate feature map.
[0046] The beneficial effects of the present invention are:
[0047] (1) The present invention provides a pancreas segmentation method based on the WEUnet network. When dealing with complex pancreas segmentation tasks, the network effectively solves the problems of pancreas deformation, unclear edges, and class imbalance, and further improves the segmentation index. The present invention proposes a wavelet convolution submodule using Dobesi wavelet transform, which makes the network pay more attention to the edge of the pancreas by fusing low-frequency information and high-frequency information, and can well fuse local features and global features, making the network expression clearer.
[0048] (2) The edge feature attention module proposed in the present invention makes good use of the additional high-frequency information introduced, and solves the problem of blurred boundaries to a certain extent. In addition, the adopted transfer learning strategy reduces the training cost and improves the generalization ability. The method of the present invention is evaluated on the public NIH dataset and MSD dataset. The experimental results show that the segmentation results of the proposed WEUnet network on both datasets are better than those of other methods, proving that the proposed network is effective.
[0049] (3) The present invention proposes a novel neural network, WEUNet, which achieves comparable performance to mainstream methods while reducing the number of parameters and training costs, achieving an average DSC score of 87.61±2.34% on the NIH dataset and an average DSC score of 87.29±2.52% on the MSD dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flow chart of a pancreas segmentation method based on the WEUnet network of the present invention.
[0051] Figure 2 This is the overall structure diagram of the WEUnet network.
[0052] Figure 3 A schematic diagram of the encoding module.
[0053] Figure 4 Schematic diagram of the wavelet convolution submodule.
[0054] Figure 5 A schematic diagram of the decoding module.
[0055] Figure 6 Schematic diagram of the edge feature attention module.
[0056] Figure 7 Box plot representation of four-fold cross validation on the NIH dataset.
[0057] Figure 8 Segmentation results on the NIH dataset. DETAILED DESCRIPTION
[0058] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0059] like Figure 1 and Figure 2 As shown, in one embodiment of the present invention, a pancreas segmentation method based on the WEUnet network includes the following steps:
[0060] S1, obtain abdominal CT images and input the abdominal CT images into the WEUnet network;
[0061] S2, performing convolution encoding processing on the abdominal CT image five times in a row through the encoding module to obtain the encoded visual feature map output by each layer of the encoding module;
[0062] S3, inputting the lowest level encoded visual feature map into a decoding module to obtain the lowest level decoded prediction feature map;
[0063] S4, the decoding module performs four consecutive convolution decoding processes on the decoding prediction feature map of the bottom layer to obtain the decoding prediction feature map of the top layer, and in each decoding process, the edge feature attention module connected by the jump connection calculates the intermediate feature map according to the encoded visual feature map of the current layer, the decoding prediction feature map of the next layer and the high-frequency feature map of the current layer, and inputs the intermediate feature map into the decoding module, and the generated result is spliced on the generated decoding prediction feature map of the current layer;
[0064] S5. Obtain a pancreatic image based on the decoded prediction feature map of the top layer.
[0065] In this embodiment, the overall structure diagram is as follows: Figure 2 As shown, it adopts a U-shaped jump connection structure, and five encoding modules and decoding modules are connected in series in the encoding path and decoding path respectively.
[0066] The encoding module is based on the MobileOne structure, combining its lightweight and efficient features, and introduces a wavelet convolution submodule to enhance the ability to extract specific features. Specifically, in the main path, the input feature map first passes through the wavelet convolution WTconv, followed by normalization Norm and ReLu activation. The feature map then passes through depthwise convolution and pointwise convolution in turn. The specific structure is consistent with MobileOne, so that by replacing the pooling layer with a depthwise separable convolution, richer features can be learned during downsampling.
[0067] In the decoding module, the feature map undergoes two consecutive 3×3 convolutions, followed by normalization and ReLu activation. In order to prevent network degradation, a residual branch is introduced to further improve the feature restoration and fusion capabilities.
[0068] Traditional CNN structures usually lose key information when downsampling to deep layers, which is very important in medical segmentation. In addition, in order to enhance the learning ability of small objects and edge features. An edge feature attention module (SEA) with residual properties is designed between the same-level codec blocks of the jump connection. This module runs between the upsampling path and the downsampling path at each resolution level, and receives three inputs: the output of each layer in the downsampling path, the output of the lower layer in the upsampling path, and the additional extracted edge features. Through these inputs, the edge feature attention module builds a coherent bridge between the contraction path and the expansion path. This design can not only effectively retain and emphasize edge features, but also strengthen the representation ability of small objects, thereby significantly improving the network's ability to capture details and global features in medical segmentation tasks.
[0069] like Figure 3 As shown, in S2, the encoding module includes a wavelet convolution submodule, a depth convolution submodule and a point-by-point convolution submodule connected in sequence;
[0070] Among them, the output end of the wavelet convolution submodule is also connected to the input end of the first convolution submodule and the input end of the first normalization layer respectively, the output end of the point-by-point convolution submodule, the output end of the first convolution submodule and the output end of the first normalization layer are all connected to the input end of the first element addition, the output end of the first element addition is respectively connected to the input end of the second convolution submodule and the input end of the second normalization layer through the first ReLu activation layer, the output end of the second convolution submodule and the output end of the second normalization layer are both connected to the input end of the second element addition, and the output end of the second element addition is connected to the input end of the second ReLu activation layer;
[0071] like Figure 4 As shown, the wavelet convolution submodule includes a wavelet convolution, a third normalization layer, and a third ReLu activation layer connected in sequence;
[0072] The structure of the depthwise convolution submodule is the same as that of the pointwise convolution submodule, both of which include the first 3×3 convolution and the fourth normalization layer connected to each other;
[0073] The first convolution submodule and the second convolution submodule have the same structure, both including a first 1×1 convolution and a fifth normalization layer connected to each other.
[0074] In this embodiment, due to the blurred boundary between the pancreas and adjacent organs, the traditional convolutional network faces great challenges in capturing its unique features. This is mainly because the anatomical position of the pancreas is complex and its size in the image is small, which is easily interfered by background noise and adjacent organs. In order to solve these problems, the present invention introduces the wavelet convolution technology and improves the characteristics of small target organs in the abdomen to improve the network's learning ability of the pancreas features. DWT (discrete wavelet transform) is a multi-resolution analysis tool that can decompose an input signal or image into low-frequency and high-frequency information. In image processing, DWT provides a way to decompose an image into different frequency components while retaining spatial information. This decomposition generates a set of sub-images with different details and resolutions. Specifically, the original image is decomposed into low-frequency components and high-frequency components. The low-frequency component LL represents the overall structural information of the image and contains lower frequency content, which is equivalent to presenting the image at a lower spatial resolution. The high-frequency components include LH, HL, and HH: describing the details and textures in the image, where LH represents horizontal details, HL represents vertical details, and HH represents details in the diagonal direction.
[0075] In S2, the method of encoding the visual feature map output by the encoding module is specifically as follows:
[0076] S21, obtaining a first input feature map of an input encoding module, and inputting the first input feature map into a wavelet convolution submodule to obtain a first feature map;
[0077] S22, inputting the first feature map into the depth convolution submodule and the point-by-point convolution submodule in sequence, and using the residual branch to input the first feature map into the first convolution submodule and the first normalization layer respectively, and adding the outputs of the point-by-point convolution submodule, the first convolution submodule and the first normalization layer by the first element to obtain a second feature map;
[0078] S23, passing the second feature map through the first ReLu activation layer and the second convolution submodule in sequence, and inputting the second feature map into the second normalization layer using the residual branch, and adding the outputs of the second convolution submodule and the second normalization layer through the second element to obtain a third feature map;
[0079] S24. Input the third feature map into the second ReLu activation layer to obtain an encoded visual feature map.
[0080] The S21 comprises the following sub-steps:
[0081] S211, generating a first low-frequency component and a first high-frequency component by wavelet transforming the first input feature map;
[0082] S212, generating a second low-frequency component and a second high-frequency component by wavelet transforming the first low-frequency component;
[0083] S213, performing convolution and transposition convolution processing on the second low-frequency component and the second high-frequency component in sequence, and taking all processed components as the convolution result of the second wavelet transform;
[0084] S214, adding the convolution result of the second wavelet transform to the first low-frequency component and the first high-frequency component, and sequentially performing convolution and transposed convolution processing on the result of the addition, and using all the processed components as the convolution result of the first wavelet transform;
[0085] S215: Add the convolution result of the first wavelet transform to the first input feature map, and perform convolution processing on the result of the addition to obtain the first feature map.
[0086] In this embodiment, the first input feature map is input into the wavelet convolution submodule to obtain the first feature map. This method combines the global information of the low-frequency component and the detailed features of the high-frequency component, and improves the model's ability to recognize fuzzy boundaries. According to the first input feature map T, four sets of filters are used:
[0087]
[0088] In the formula, f LL is a low-pass filter, f LH , f HL , f HH is a high pass filter.
[0089] Then use the above filter to convolve each channel and output low-frequency and high-frequency components.
[0090] [Z LL ,Z LH ,Z HL ,Z HH ]=Conv([f LL ,f LH ,f HL ,f HH ],T)
[0091] In the formula, Z LL is the low-frequency component of T, Z LH ,Z HL ,Z HH are the high-frequency components in the horizontal, vertical and diagonal directions respectively, and Conv(·) is the convolution processing;
[0092] Finally, the processed components are restored through transposed convolution to generate the convolution result Z of the wavelet transform.
[0093] Z=Convtransposed([f LL ,f LH ,fHL ,f HH ],Z c )
[0094] Z c =Concat(Z LL ,Z LH ,Z HL ,Z HH )
[0095] In the formula, Convtransposed(·) is the transposed convolution processing, Z c It is the result of connecting the low-frequency component and the high-frequency component, and Concat is the connection operation.
[0096] like Figure 5 As shown, in S3, the decoding module includes a third convolution submodule, a fourth ReLu activation layer, and a fourth convolution submodule connected in sequence, the input end of the third convolution submodule is also connected to the input end of the fifth convolution submodule, the output end of the fourth convolution submodule and the output end of the fifth convolution submodule are both connected to the input end of the third element addition, and the output end of the third element addition is sequentially connected to the fifth ReLu activation layer and the upsampling operation layer;
[0097] The third convolution submodule and the fourth convolution submodule have the same structure, both including a second 3×3 convolution and a sixth normalization layer connected to each other;
[0098] The fifth convolution submodule includes a second 1×1 convolution and a seventh normalization layer connected to each other.
[0099] In S3, the method for processing the second input feature map input by the decoding module is specifically:
[0100] S31, obtaining a second input feature map of the input decoding module, and inputting the second input feature map into the third convolution submodule, the fourth ReLu activation layer and the fourth convolution submodule;
[0101] S32, using the residual branch to input the second input feature map into the fifth convolution submodule;
[0102] S33. The outputs of the fourth convolution submodule and the fifth convolution submodule are added together through the third element to obtain a fourth feature map, and the fourth feature map is sequentially input into the fifth ReLu activation layer and the upsampling operation layer to obtain the result of the second input feature map processed by the decoding module.
[0103] like Figure 6 As shown, in S4, the edge feature attention module is provided with a convolutional attention submodule, and the convolutional attention submodule includes a spatial attention submodule interconnected with a channel attention submodule;
[0104] Among them, the channel attention submodule is provided with a first convolution of a 1×1×C convolution kernel, and the spatial attention submodule is provided with a second convolution of a H×W×1 convolution kernel.
[0105] In this embodiment, the main goal of the edge feature attention module is to stably retain high-frequency information at multiple scales, thereby effectively solving the boundary fuzziness problem in pancreas segmentation. In addition, the edge feature attention module plays a key role in bridging the semantic gap between the low-level features extracted by the encoder and the high-level features produced by the decoder before their fusion.
[0106] In S4, the high-frequency feature map is obtained by DWT processing of the abdominal CT image;
[0107] The method for calculating the intermediate feature map during each decoding process includes the following steps:
[0108] S41, obtain the coded visual feature map of the current layer, the high-frequency feature map of the current layer and the decoded prediction feature map of the next layer, multiply the high-frequency feature map of the current layer and the decoded prediction feature map of the next layer by the coded visual feature map of the current layer element by element, concatenate the generated feature map with the coded visual feature map of the current layer, and generate a fused feature map f i a ;
[0109]
[0110] In the formula, f i E To obtain the encoded visual feature map of the current layer, f i DWT is the high-frequency feature map of the current layer, is the decoding prediction feature map of the next layer, i is the current layer number, is element multiplication, [] is feature concatenation;
[0111] After DWT processing, multiple high-frequency components are obtained. In this embodiment, vertical high-frequency components are used. The high-frequency components obtained in the i-th layer need to be filtered and then downsampled by 2 times the features of the i-1 layer. However, this process may reduce the intensity of the high-frequency information in the i-th layer. In order to retain the high-frequency information f in each layer i DWT The present invention is from Derive the high-frequency features of each layer to most effectively preserve the high-frequency information.
[0112] f i DWT =(down(f l )) i ifi≥1
[0113] where down is twice downsampling, (down(f l )) i is f l after i times of downsampling operations, f l is the high-frequency feature map, and the high-frequency feature map is obtained by processing the abdominal CT image through DWT.
[0114] S42. Input the fused feature map into the channel attention sub-module. Use the first branch to perform max-pooling and multi-layer perceptron processing on the fused feature map in sequence, and optimize it through the first convolution after max-pooling and multi-layer perceptron processing to generate the first attention feature map;
[0115] S43. Use the second branch to perform average-pooling and multi-layer perceptron processing on the fused feature map in sequence, and optimize it through the first convolution after average-pooling and multi-layer perceptron processing to generate the second attention feature map;
[0116] S44. Add the first attention feature map and the second attention feature map, process the addition result through the Sigmoid activation function and the first convolution optimization in sequence, and multiply the optimization result by the fused feature map to generate the third attention feature map;
[0117] S45. Input the third attention feature map into the spatial attention sub-module. Use the third branch to perform max-pooling and the second convolution optimization on the third attention feature map to generate the fourth attention feature map; use the fourth branch to perform average-pooling and the second convolution optimization on the third attention feature map to generate the fifth attention feature map;
[0118] S46. Fuse the fourth attention feature map and the fifth attention feature map. The fusion result is optimized through the second convolution, processed by the Sigmoid activation function, and optimized through the second convolution in sequence. The optimization result is multiplied by the third attention feature map to generate the intermediate feature map.
[0119] In this embodiment, in order to capture the feature correlation between the boundary region and the background region, the present invention inputs the fused feature map into the convolutional block attention module (CBAM) for recalibration. CBAM contains two consecutive modules. The channel attention module focuses on the channel dimension and optimizes the attention feature map using a 1×1×C convolutional kernel. The spatial attention module focuses on the spatial dimension and optimizes the attention feature map using an H×W×1 convolutional kernel. Through this process, the optimized decoded features are obtained.
[0120] In order to evaluate the advantages of the model of the present invention from different perspectives, we first compare it with the network that performs well on the NIH dataset. After that, we visualize the segmentation results of the proposed WEUnet network.
[0121] Table 1 shows the four-fold cross validation comparison results of the network of the present invention and the current network on the NIH dataset. The comparison results show that the proposed WEUnet network is superior to other methods, achieving an optimal DSC of 87.61%. The standard deviation of 2.34 indicates that the method of the present invention has high robustness in different cases of the NIH dataset. In addition, it can be seen from the recall rate of 89.92% that the model of the present invention exhibits excellent pixel-level pancreas recognition capabilities. And compared with other models, the model of the present invention has fewer parameters and is more lightweight while maintaining the leading DSC.
[0122] Table 1. Pancreas segmentation results of NIH dataset
[0123]
[0124]
[0125] Table 2 provides a comparison of model parameters, showing the model size in more detail. The model of the present invention excels due to its simple structure and smaller number of parameters, which is only 6.1M. While maintaining the leading performance of DSC, the number of parameters of the present invention is about four times less than the best model. This substantial improvement is attributed to transfer learning and fine-tuning techniques, which have consistently proven their effectiveness in the field of medical image segmentation.
[0126] Table 2 Comparison of model parameters (NIH dataset)
[0127]
[0128] from Figure 7 It can be seen that the network of the present invention achieves excellent performance on the NIH dataset, and the distribution of DSC values in different sample sets is very close, which shows that the proposed network has high robustness and can mitigate the impact of sample changes.
[0129] Figure 8 The segmentation results of the WEUnet network on the NIH dataset are shown. The visualization of the segmentation results and the presentation of the DSC index show that the result prediction of the present invention is very close to the actual situation, which means that the network of the present invention can effectively capture the differences in pancreatic shape and size between individuals.
[0130] Table 3 shows the comparison results of the proposed WEUnet network with the current methods that perform well on the MSD dataset. In terms of the Dice similarity coefficient, WEUnet reaches 87.29±2.52, which is in the leading position in segmentation accuracy. In addition, in terms of Precision and Recall, WEUnet reaches 86.37±2.15 and 88.94±1.67 respectively, showing high segmentation accuracy and low missed detection rate. Overall, WEUnet maintains high accuracy while taking into account strong robustness, which fully demonstrates its effectiveness and reliability in segmentation tasks.
[0131] Table 3 Pancreas segmentation results of MSD dataset
[0132]
[0133] The beneficial effects of the present invention are as follows: the present invention provides a pancreas segmentation method based on the WEUnet network. When processing complex pancreas segmentation tasks, the network effectively solves the problems of pancreas deformation, unclear edges, and imbalanced categories, and further improves the segmentation index. The present invention proposes a wavelet convolution submodule using Dobesi wavelet transform, which makes the network pay more attention to the edge of the pancreas by fusing low-frequency information and high-frequency information, and can well fuse local features and global features, making the network expression clearer.
[0134] The edge feature attention module proposed in the present invention makes good use of the additional high-frequency information introduced, and solves the problem of blurred boundaries to a certain extent. In addition, the adopted transfer learning strategy reduces the training cost and improves the generalization ability. The method of the present invention is evaluated on the public NIH dataset and MSD dataset. The experimental results show that the segmentation results of the proposed WEUnet network on both datasets are better than those of other methods, proving that the proposed network is effective.
[0135] The present invention proposes a novel neural network, WEUNet, which achieves comparable performance to mainstream methods while reducing the number of parameters and training costs, achieving an average DSC score of 87.61±2.34% on the NIH dataset and an average DSC score of 87.29±2.52% on the MSD dataset.
[0136] In the description of the present invention, it is necessary to understand that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used only for descriptive purposes, and cannot be understood as indicating or implying the relative importance or the number of implicitly specified technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of the features.
Claims
1. A pancreas segmentation method based on WEUnet network, characterized in that: The following steps are involved: S1, obtain abdominal CT images and input the abdominal CT images into the WEUnet network; S2, performing convolution encoding processing on the abdominal CT image five times in a row through the encoding module to obtain the encoded visual feature map output by each layer of the encoding module; S3, inputting the lowest level encoded visual feature map into a decoding module to obtain the lowest level decoded prediction feature map; S4, the decoding module performs four consecutive convolution decoding processes on the decoding prediction feature map of the bottom layer to obtain the decoding prediction feature map of the top layer, and in each decoding process, the edge feature attention module connected by the jump connection calculates the intermediate feature map according to the encoded visual feature map of the current layer, the decoding prediction feature map of the next layer and the high-frequency feature map of the current layer, and inputs the intermediate feature map into the decoding module, and the generated result is spliced on the generated decoding prediction feature map of the current layer; S5. Obtain a pancreatic image based on the decoded prediction feature map of the top layer.
2. The pancreas segmentation method based on WEUnet network according to claim 1, characterized in that: In S2, the encoding module includes a wavelet convolution submodule, a depth convolution submodule and a point-by-point convolution submodule connected in sequence; Among them, the output end of the wavelet convolution submodule is also connected to the input end of the first convolution submodule and the input end of the first normalization layer respectively, the output end of the point-by-point convolution submodule, the output end of the first convolution submodule and the output end of the first normalization layer are all connected to the input end of the first element addition, the output end of the first element addition is respectively connected to the input end of the second convolution submodule and the input end of the second normalization layer through the first ReLu activation layer, the output end of the second convolution submodule and the output end of the second normalization layer are both connected to the input end of the second element addition, and the output end of the second element addition is connected to the input end of the second ReLu activation layer; The wavelet convolution submodule includes a wavelet convolution, a third normalization layer, and a third ReLu activation layer connected in sequence; The structure of the depthwise convolution submodule is the same as that of the pointwise convolution submodule, both of which include the first 3×3 convolution and the fourth normalization layer connected to each other; The first convolution submodule and the second convolution submodule have the same structure, both including a first 1×1 convolution and a fifth normalization layer connected to each other.
3. The pancreas segmentation method based on WEUnet network according to claim 2, characterized in that: In S2, the method of encoding the visual feature map output by the encoding module is specifically as follows: S21, obtaining a first input feature map of an input encoding module, and inputting the first input feature map into a wavelet convolution submodule to obtain a first feature map; S22, inputting the first feature map into the depth convolution submodule and the point-by-point convolution submodule in sequence, and using the residual branch to input the first feature map into the first convolution submodule and the first normalization layer respectively, and adding the outputs of the point-by-point convolution submodule, the first convolution submodule and the first normalization layer by the first element to obtain a second feature map; S23, passing the second feature map through the first ReLu activation layer and the second convolution submodule in sequence, and inputting the second feature map into the second normalization layer using the residual branch, and adding the outputs of the second convolution submodule and the second normalization layer through the second element to obtain a third feature map; S24. Input the third feature map into the second ReLu activation layer to obtain an encoded visual feature map.
4. The pancreas segmentation method based on WEUnet network according to claim 3, characterized in that: The S21 comprises the following sub-steps: S211, generating a first low-frequency component and a first high-frequency component by wavelet transforming the first input feature map; S212, generating a second low-frequency component and a second high-frequency component by wavelet transforming the first low-frequency component; S213, performing convolution and transposition convolution processing on the second low-frequency component and the second high-frequency component in sequence, and taking all processed components as the convolution result of the second wavelet transform; S214, adding the convolution result of the second wavelet transform to the first low-frequency component and the first high-frequency component, and sequentially performing convolution and transposed convolution processing on the result of the addition, and using all the processed components as the convolution result of the first wavelet transform; S215: Add the convolution result of the first wavelet transform to the first input feature map, and perform convolution processing on the result of the addition to obtain the first feature map.
5. The pancreas segmentation method based on WEUnet network according to claim 1, characterized in that: In the S3, the decoding module includes a third convolution submodule, a fourth ReLu activation layer, and a fourth convolution submodule connected in sequence, the input end of the third convolution submodule is also connected to the input end of the fifth convolution submodule, the output end of the fourth convolution submodule and the output end of the fifth convolution submodule are both connected to the input end of the third element addition, and the output end of the third element addition is sequentially connected to the fifth ReLu activation layer and the upsampling operation layer; The third convolution submodule and the fourth convolution submodule have the same structure, both including a second 3×3 convolution and a sixth normalization layer connected to each other; The fifth convolution submodule includes a second 1×1 convolution and a seventh normalization layer connected to each other.
6. The pancreas segmentation method based on WEUnet network according to claim 5, characterized in that: In S3, the method for processing the second input feature map input by the decoding module is specifically: S31, obtaining a second input feature map of the input decoding module, and inputting the second input feature map into the third convolution submodule, the fourth ReLu activation layer and the fourth convolution submodule; S32, using the residual branch to input the second input feature map into the fifth convolution submodule; S33. The outputs of the fourth convolution submodule and the fifth convolution submodule are added together through the third element to obtain a fourth feature map, and the fourth feature map is sequentially input into the fifth ReLu activation layer and the upsampling operation layer to obtain the result of the second input feature map processed by the decoding module.
7. The pancreas segmentation method based on WEUnet network according to claim 1, characterized in that: In S4, the edge feature attention module is provided with a convolutional attention submodule, and the convolutional attention submodule includes a spatial attention submodule interconnected with a channel attention submodule; Among them, the channel attention submodule is provided with a first convolution of a 1×1×C convolution kernel, and the spatial attention submodule is provided with a second convolution of a H×W×1 convolution kernel.
8. The pancreas segmentation method based on WEUnet network according to claim 7, characterized in that: In S4, the high-frequency feature map is obtained by DWT processing of the abdominal CT image; The method for calculating the intermediate feature map during each decoding process includes the following steps: S41, obtaining a coded visual feature map of the current layer, a high-frequency feature map of the current layer, and a decoded prediction feature map of the next layer, multiplying the high-frequency feature map of the current layer and the decoded prediction feature map of the next layer by the coded visual feature map of the current layer element by element, concatenating the generated feature map with the coded visual feature map of the current layer to generate a fused feature map; S42, inputting the fused feature map into the channel attention submodule, using the first branch to perform maximum pooling and multi-layer perceptron processing on the fused feature map in sequence, and optimizing it through the first convolution after the maximum pooling and multi-layer perceptron processing to generate a first attention feature map; S43, using the second branch to perform average pooling and multi-layer perceptron processing on the fusion feature map in sequence, and optimizing it through the first convolution after the average pooling and multi-layer perceptron processing to generate a second attention feature map; S44, adding the first attention feature map and the second attention feature map, subjecting the addition result to Sigmoid activation function processing and first convolution optimization in sequence, multiplying the optimization result with the fusion feature map to generate a third attention feature map; S45, input the third attention feature map into the spatial attention submodule, use the third branch to perform maximum pooling and second convolution optimization on the third attention feature map, and generate a fourth attention feature map; use the fourth branch to perform average pooling and second convolution optimization on the third attention feature map, and generate a fifth attention feature map; S46. Fuse the fourth attention feature map and the fifth attention feature map. Perform the second convolution optimization, Sigmoid activation function processing and the second convolution optimization on the fusion result in turn. Multiply the optimization result with the third attention feature map to generate an intermediate feature map.
Citation Information
Patent Citations
Pancreatic tumor segmentation method based on easily confused region information mining
CN116823843A
Intestinal ultrasound image segmentation method, system and equipment based on diffusion network and depth metric learning, and medium
CN118762040A
Ischemic stroke image segmentation method based on wavelet pooling neural network
CN119445116A
Apparatus for estimating under monocular infrared thermal imaging vision pose of object grasped by manipulator, and method thereof
WO2024148645A1