Post - decoder 3D axial decoupling enhancement method for 3D medical image segmentation
By introducing the post-decoder three-dimensional axial decoupling enhancement method (PaR) in three-dimensional medical image segmentation, the attention mechanism is used to enhance the expression of information feature in the axial time dimension, which solves the problem of existing methods ignoring the difference in axial characteristics, and improves the accuracy and robustness of three-dimensional medical image segmentation.
Patent Information
- Application Number
- CN202510211250.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing three-dimensional medical image segmentation methods ignore the differences in physical and scanning characteristics of three-dimensional medical imaging data in different axial directions, resulting in poor performance in three-dimensional medical image segmentation tasks.
A post-decoder three-dimensional axial decoupling enhancement method (PaR) is proposed. By analyzing the differences in each axial feature of three-dimensional medical image data, an attention mechanism is introduced to independently enhance the expression of axial time dimension information.
It improves the accuracy and robustness of three-dimensional medical image segmentation, and is suitable for a variety of medical image segmentation tasks such as abdominal organ segmentation, fundus retinal tissue segmentation and cardiac system segmentation, providing high-precision automated segmentation tools for clinical diagnosis.
Smart Images

Figure CN119693398B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional medical image segmentation, and particularly to a post-decoder three-dimensional axial decoupling enhancement method for three-dimensional medical image segmentation, PaR (Post-Axial Refiner). For three-dimensional medical image segmentation of fundus retina tissue segmentation of three-dimensional medical images, PaR analyzes the spatio-temporal feature differences of different axial directions of three-dimensional medical image data, introduces an attention mechanism to enhance the model's ability to express features on the time axis, thereby more accurately mapping the structural differences and detail changes on the image time axis, and improving the accuracy and robustness of three-dimensional medical image segmentation. This method is applicable to various medical image segmentation tasks such as abdominal organ segmentation, fundus retina tissue segmentation, and cardiac system segmentation, providing an efficient automated segmentation tool for clinical diagnosis to assist diagnosis. Background Art
[0002] Three-dimensional medical imaging data based on medical imaging technologies such as MRI and CT are extremely crucial for accurate clinical diagnosis. These data contain the precise spatial structures of various tissues and organs in the patient's body and their dynamic changes over time, such as blood flow or organ movement. The three-dimensional nature of these data enables doctors to deeply analyze pathological states and physiological processes. In order to extract accurate biological structure information from these complex data sets for effective disease diagnosis and subsequent treatment analysis, precise three-dimensional medical image segmentation is of vital importance. Traditional two-dimensional image segmentation techniques ignore the hierarchical relationship between time-axis slices of three-dimensional medical imaging data and cannot effectively express the spatio-temporal correlation of three-dimensional data, resulting in a lack of three-dimensional cohesion and accuracy in the segmentation results. Three-dimensional convolution analyzes and processes information by simultaneously expanding the convolution kernel in three dimensions, directly extracts features in three-dimensional space, can utilize information in each dimension of three-dimensional medical image data, ensures that the segmentation result can truly reflect the integrity of the three-dimensional structure, retains the spatial continuity, and enhances the local correlation of information.
[0003] However, the symmetric structure design of the three-dimensional convolution kernel makes it process the time-axis information (z-axis) in the same way as the slice-plane information (XY plane), equally collecting information from all three dimensions. This processing strategy ignores the differences in the physical properties and information densities of three-dimensional medical image data between the z-axis and the xy plane, which will lead to inefficient utilization and misunderstanding of the axial information by the deep learning model. In three-dimensional medical image data, the xy-plane image information is directly collected by the sensor array, and the brightness changes of the pixels map the spatial distribution of tissue density to reflect the exact positions of tissues with different densities in physical space. In contrast, the z-axis is a scanning sequence along the time axis, and these sequences are obtained by layer-by-layer scanning, and the interval between each layer is usually much larger than the pixel resolution in the xy plane. Therefore, the characteristics of the z-axis reflect a complex of time process and spatial structure, and the spatio-temporal characteristics of its information are fundamentally different from the pure spatial characteristics of the xy axis. However, the structure of the three-dimensional convolution is the same in all directions and cannot reflect this information difference. This axial information asymmetry makes it impossible for simple three-dimensional convolution to effectively distinguish and utilize the feature information of each axis, especially when processing z-axis information. Since the information density of the z-axis is much lower than that of the xy plane, traditional three-dimensional convolution may introduce spatial deviations and errors when uniformly extracting features on the z-axis and the xy plane, and applying the xy-plane features reflecting specific spatial relationships to the z-axis features will ultimately affect the segmentation accuracy.
[0004] In the field of medical image analysis and processing, high-precision three-dimensional (3D) medical imaging data segmentation is crucial for accurate and efficient clinical diagnosis and subsequent treatment analysis, such as fundus retina tissue segmentation. However, most existing deep learning methods use 3D convolutions to process the information in the three dimensions of 3D medical image data. This method has certain limitations as it ignores the physical and scanning characteristic differences of 3D image data in different axial directions, resulting in poor performance in 3D medical image segmentation tasks. To address this problem, the present invention proposes a method for 3D medical image segmentation, namely the Post-Axial Refiner (PaR), which optimizes the coarse mask of the decoder and enhances 3D axial decoupling. PaR splits the coarse mask obtained after the decoder of the deep learning model along the time axis and uses the attention mechanism to enhance feature representation, achieving 3D axial decoupling to enhance the deep learning model's understanding of the spatio-temporal information correlation in 3D image data. By analyzing the feature differences in each axial direction of 3D medical image data, PaR introduces the attention mechanism to independently enhance the ability to express z-axis feature information, thereby more accurately mapping the structural differences and detail changes on the time axis, and improving the accuracy and robustness of 3D medical image segmentation. This module is applicable to various medical image segmentation tasks such as abdominal organ segmentation, fundus retina tissue segmentation, and cardiac system segmentation, providing a high-precision automated segmentation tool for clinical diagnosis to assist in diagnosis. The proposed PaR is a plug-and-play module that can be easily integrated into any neural network architecture, not limited to 3D medical image segmentation models. Summary of the Invention
[0005] The objective of the present invention is to provide a post - decoder three - dimensional axial decoupling enhancement method for three - dimensional medical image segmentation in view of the deficiencies of the prior art. By analyzing the spatio - temporal characteristics of different axial information in three - dimensional medical image data, an attention mechanism is introduced to independently enhance the feature expression of the axial time - dimension information, thereby more accurately mapping the structural differences and detailed changes on the time axis (z - axis), and improving the accuracy and robustness of three - dimensional medical image segmentation. This method is applicable to various medical image segmentation tasks such as abdominal organ segmentation, fundus retina tissue segmentation, and cardiac system segmentation, providing a high - precision automated segmentation tool for clinical diagnosis to assist diagnosis. Based on the idea of three - dimensional axial decoupling, this method enhances the feature expression ability of the model by specifically processing the time - axial information of three - dimensional medical image data to improve the performance of the model in three - dimensional medical image segmentation tasks. This method mainly includes two main pipelines: the Permutative Dimensional Intensification (PDI) pipeline and the Sequential Axial Intensification (SAI) pipeline. PDI learns and understands different axial features from the perspective of global spatial information and enhances the attention of z - axis features, while SAI directly groups and processes the z - axis information on the time axis to enhance the model's feature expression ability. Through the independent work and collaborative processing of the PDI pipeline and the SAI pipeline, the present invention can perform axial decoupling enhancement on the rough masks of three - dimensional medical image segmentation with high complexity and difficulty, achieving more accurate and robust segmentation results, and being able to adapt to any three - dimensional image segmentation deep - learning model, demonstrating its wide application potential in the field of three - dimensional medical image segmentation.
[0006] The technical solutions included in the present invention to solve its technical problems are as follows:
[0007] S1: Collect three - dimensional medical image segmentation datasets corresponding to various organ segmentation tasks from the network.
[0008] S2: Divide the collected three - dimensional medical image data into a training set and a test set according to a certain proportion;
[0009] S3: Perform data pre - processing and data augmentation on the three - dimensional medical image data in step S2;
[0010] S4: Collect advanced deep - learning models widely used for three - dimensional medical image segmentation from the network as benchmark models;
[0011] S5: Design a post - decoder three - dimensional axial decoupling enhancement method (PaR) that can be used for rough mask optimization in various three - dimensional medical image segmentation tasks;
[0012] S6: Integrate the PaR method into the collected benchmark models.
[0013] S7: On the preprocessed and data-augmented dataset, use the benchmark model and the benchmark model integrated with the PaR body respectively for model training, and save the best model weights of each according to the evaluation results on the validation set;
[0014] S8: Validate the saved best model weights on the test set to evaluate the model performance, and compare the performance differences between different benchmark models and verify the effectiveness of the proposed PaR method according to evaluation metrics including Dice Similarity Coefficient (DSC), Mean Intersection over Union (mIoU), and 95% Hausdorff Distance (HD95).
[0015] The beneficial effects of the present invention are as follows:
[0016] Based on the deep learning segmentation architecture, the present invention takes into account the problem of low utilization efficiency of the feature information on the time axis (z-axis) caused by the difference between the isotropic convolution kernel inherent in the current 3D convolution used for 3D medical image segmentation and the anisotropic characteristics of 3D medical image data, resulting in poor model segmentation accuracy and affecting the quality of clinical diagnosis and subsequent treatment analysis. The present invention proposes a Post-Axial Refiner (PaR) method for optimizing the coarse mask of 3D medical image segmentation. By analyzing the spatio-temporal characteristics of different axial information in 3D medical image data, PaR introduces an attention mechanism to independently enhance the feature expression of the axial time dimension information, thereby more accurately mapping the structural differences and detailed changes on the spatio-temporal axis (z-axis), and improving the accuracy and robustness of 3D medical image segmentation. This module is applicable to various medical image segmentation tasks such as abdominal organ segmentation, fundus retina tissue segmentation, and cardiac system segmentation, providing a high-precision automated segmentation tool for clinical diagnosis to assist diagnosis. As a plug-and-play deep learning module, the invention can be conveniently integrated into any deep learning model to improve the model performance. The application of the invention can improve the accuracy of 3D medical image segmentation by deep learning models, thereby enhancing the efficiency and effect of clinical diagnosis and subsequent treatment analysis, bringing important technological progress to the field of 3D medical image segmentation, and ultimately improving the diagnosis efficiency of doctors and the quality of treatment plans. Description of the Drawings
[0017] Figure 1 It is the overall structure diagram of the PaR method in the embodiment of the present invention.
[0018] Figure 2It is the detailed structure diagram of the Permutation Dimension Enhancement Pipeline (PDI) in the embodiments of the present invention.
[0019] Figure 3 It is the detailed structure diagram of the Sequence Axial Enhancement Pipeline (SAI) in the embodiments of the present invention. Detailed implementation manners
[0020] The present invention will be further explained and illustrated below in conjunction with embodiments.
[0021] As Figure 1 shown, the present invention proposes a Post-Axial Refiner (PaR) method for three-dimensional medical image segmentation. The aim is to enhance and optimize the coarse mask of three-dimensional medical image segmentation obtained by a deep learning model, which specifically includes the following steps:
[0022] Step 1: Collect publicly available three-dimensional medical image datasets corresponding to various organ segmentation tasks from the network, including but not limited to high-quality datasets such as the abdominal segmentation dataset FLARE2021, the fundus retina segmentation dataset OIMHS, and the heart segmentation dataset SegTHOR. Each dataset contains the original images and annotations, which can be used for the training, validation, and testing of subsequent deep learning models.
[0023] Step 2: Divide the collected three-dimensional medical image datasets of multiple organs into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0024] Step 3: Perform data preprocessing on the three-dimensional medical image data in Step 2 to ensure the consistency and quality of the input data to obtain better training effects. This includes reinterpolating these data, interpolating sequences with a length less than 108 to 108; performing image intensity clipping on the image pixels and applying maximum-minimum normalization. For the threshold of image intensity clipping, it is set to [-125, 275] for the FLARE2021 dataset, [0, 400] for the OIMHS dataset, and [-1000, 300] for the SegTHOR dataset, and randomly cropping 96x96x96 for training, and using the slide_windows method for validation, with the overlap between images set to 0.5 during validation. Perform data augmentation on the experimental data, including but not limited to random flipping, random rotation, random scaling, and random image intensity offset.
[0025] Step 4: Collect advanced deep learning models widely used for 3D medical image segmentation from the network as benchmark models, including but not limited to 3D U-Net, V-Net, RAUNet, ResUNet, SegResNet, MultiResUNet, UNETR, SwinUNETR, TransBTS, nnFormer, and 3D UX-Net.
[0026] Step 5: To improve the performance of the benchmark deep learning model in 3D medical image segmentation, this study proposes a post-decoder 3D axial decoupling enhancement (PaR) method. First, define the decoder output of the general deep learning model for segmentation as the coarse mask , where, represents the set of real numbers, represents the number of segmentation classes, , and represent the width, height, and temporal axis depth of the 3D medical image data, respectively. Compared with 2D planar images, 3D medical images have richer spatio-temporal and detailed texture information. Therefore, the PaR method uses 3D axial decoupling to learn the feature differences of 3D medical image data on the temporal axial plane and the slice plane, thereby optimizing the effect of the coarse mask. To effectively capture the temporal and spatial feature differences, it adopts two parallel 3D decoupling enhancement pipelines: the permutation dimension enhancement (PDI) pipeline and the sequential axial enhancement (SAI) pipeline, which extract key spatio-temporal features from different spatial perspectives and perform 3D axial decoupling in the temporal axis direction, thus comprehensively refining the segmentation coarse mask. The structure and function of the PaR method, as well as the PDI pipeline and the SAI pipeline, are introduced in detail below, and the relevant mathematical formulas are given.
[0027] As Figure 1 shown, the outputs of the PDI pipeline and the SAI pipeline are combined by addition, then compressed to the range (0, 1) by the sigmoid function, and then multiplied by the coarse mask and added to the coarse mask through residual connection to obtain the final optimized and enhanced segmentation mask. The mathematical formula is as follows:
[0028]
[0029] Where, is the final optimized and enhanced segmentation mask. represents the refined segmentation mask from the global perspective of the PDI pipeline output; represents the refined segmentation mask from the temporal axis perspective of the SAI pipeline output.
[0030] As Figure 2As shown, the PDI pipeline enhances the accuracy and robustness of 3D image segmentation by analyzing and processing key timeline feature information under different spatial distributions in the coarse mask. For the coarse mask output by the decoder of a general deep learning model,
[0031] the PDI pipeline reconstructs the feature spatial distribution through permutation operations. This operation rearranges each dimension of the coarse mask to generate multiple feature maps that are presented under different spatial configurations, such as , , and etc. Among them, axes are all rearranged to the last dimension to separately enhance the subsequent attention mechanism's focus on the timeline (z-axis) information. These rearranged feature maps are merged from a five-dimensional tensor into a three-dimensional tensor , , making their format suitable for input into the multi-head self-attention mechanism to extract and analyze the spatial features of the z-axis information. The formula for multi-head self-attention is expressed as:
[0032]
[0033] where, is the attention output, represents multi-head self-attention.
[0034] Next, the attention output is readjusted to a five-dimensional tensor through a reshaping operation, and then dimensionally aligned with the coarse mask through permutation operations, denoted as the attention-weighted segmentation mask . The last step of the PDI pipeline is to accumulate and sum these attention-weighted segmentation masks to generate an optimized segmentation mask , thereby optimizing the coarse mask by extracting and processing the z-axis information from a global perspective.
[0035] The inherent shortness of the sequence length in 3D medical images is a fundamental defect, usually alleviated by interpolation in the data preprocessing stage, but this also leads to significant changes in the timeline information. Since most segmentation networks usually identify segmentation targets based on the slice plane, they have difficulty understanding the rich spatio-temporal information across sequences. As Figure 3 shown, the described SAI pipeline extracts spatio-temporal information from the coarse mask and learns unified features across different timeline planes, enhancing the benchmark model's understanding of the rich spatio-temporal information across sequences, thereby for the coarse mask Optimize to improve the accuracy and robustness of 3D image segmentation.
[0036] For the coarse mask , the SAI pipeline first passes through an initial feature extraction layer, which includes 3D convolution, group normalization (GN), ReLU activation function, and residual connection. This initial feature extraction layer performs preliminary feature extraction to generate an initial feature map . To capture temporal axis features through subsequent attention mechanisms, the initial feature map is split along the axis into multiple subsets . The subsets are further divided into two groups: the channel group and the spatial group .
[0037] For the channel group, temporal axis channel information is extracted through adaptive max pooling and adaptive average pooling , then merged by addition and input into an adaptive weighting function to generate a feature map :
[0038] where, and are both trainable hyperparameters, and the initial hyperparameter of is set to 1, and the initial hyperparameter of is set to 0.
[0039] Then, the values are restricted to the interval (0,1) through the sigmoid function and fused with element-wise multiplication to extract the key channel features in .
[0040] For the spatial group, passes through a 3D convolutional layer to capture the spatial features within the local neighborhood of the feature map, reflecting the importance of different spatial regions in . These features are then passed through an adaptive weighting function to generate a spatial weighted feature map :
[0041]
[0042] where, represents the 3D convolutional layer, and are also trainable hyperparameters, and their initial hyperparameters are set to 1 and 0 respectively.
[0043] Similarly, After being activated by the sigmoid function, ensure that its value is in the interval (0, 1), and fuse through the element-wise multiplication fusion method to extract the key spatial features in
[0044] Finally, the outputs of the channel group and the spatial group are concatenated and reshaped to the dimension of the coarse mask to obtain the time-axis decoupled feature map .
[0045] To further extract and utilize the spatial information from the time-axis decoupled feature map, the SAI pipeline introduces a spatial redistribution method to capture shared features from different spatial states. First, reshape the shape of the time-axis decoupled feature map to , then perform a dimension swap between the dimension and the dimension to obtain . Then, reshape the transformed tensor back to and add it to to finally generate the output of the SAI pipeline, the optimized segmentation mask . Through the supplementation of spatial information, the spatial redistribution enhances the ability of the baseline model to understand spatio-temporal axial dynamics and helps capture the complex spatio-temporal relationships inherent in 3D medical imaging.
[0046] Step 6: Integrate the implemented PaR method into the collected baseline models to obtain the segmentation networks 3D U-Net+PaR, V-Net+PaR, RAUNet+PaR, ResUNet+PaR, SegResNet+PaR, MultiResUNet+PaR, UNETR+PaR, SwinUNETR+PaR, TransBTS+PaR, nnFormer+PaR, and 3D UX-Net+PaR.
[0047] Step 7: Input the multi-organ 3D medical image data that has undergone data preprocessing and data augmentation in Step 3 into different segmentation networks for full-supervised learning. For the training on all datasets, we use As the global loss function, the AdamW optimizer with an initial learning rate of 0.0001 is used, and the training is carried out for 80,000 iterations. The batch size is set to 2, and the sliding window (slide_windows) is used for validation. The image overlap is set to 0.5, and the model is evaluated on the validation set every 1000 iterations to save the best model weights. All experiments are completed under the same software and hardware conditions. All models are trained and validated on multiple computer servers equipped with 2 NVIDIA GeForce RTX 4090 GPUs and 128G of memory. Our framework is implemented using Python 3.9, PyTorch 2.0.0, and monai 0.9.0, and the DDP distributed training framework is used for training and evaluation.
[0048] Step 8: Validate the saved best model weights on the test set to evaluate the model performance. Compare the performance differences between different baseline models and verify the effectiveness of the proposed axial decoupling module according to evaluation metrics including the Dice Similarity Coefficient (DSC), Mean Intersection over Union (mIoU), and 95% Hausdorff Distance (HD95). Among them, the Mean Intersection over Union can effectively reflect the overall accuracy of the multi-class segmentation task. The Dice Similarity Coefficient performs better for small target regions in medical image segmentation. The 95% Hausdorff Distance is used to measure the accuracy of the segmentation boundary and is more sensitive to boundary detail errors.
[0049] The present invention has carried out extensive and sufficient comparative experiments of baseline models on publicly available high-quality three-dimensional medical image segmentation datasets FLARE2021, OIMHS, and SegTHOR. The experimental results are shown in Table 1, Table 2, and Table 3. It can be seen from the index results in the tables that the PaR method proposed by the present invention can help all 11 advanced baseline segmentation models collected to improve the segmentation performance. Thus, it can be verified that PaR, as a plug-and-play module, has wide application value in the three-dimensional medical image segmentation task.
[0050] Table 1 Benchmark experiment results on the FLARE2021 dataset
[0051] Method IOU Dice HD95 3D U-Net 87.92 93.08 16.31 +PaR 89.20+1.28 93.89+0.81 2.49-13.82 RAUNet 87.93 93.08 27.37 +PaR 88.32+0.39 93.28+0.20 26.71-0.66 ResUNet 87.38 92.56 30.15 +PaR 87.83+0.45 92.94+0.38 11.80-18.35 SegResNet 86.24 91.81 3.22 +PaR 88.02+1.78 93.13+1.32 2.78-0.44 V-Net 84.13 89.89 12.93 +PaR 85.99+1.86 91.49+1.60 5.98-6.95 UNETR 84.74 90.70 4.63 +PaR 85.87+1.13 91.48+0.78 3.64-0.99 Swin UNETR 88.28 93.23 3.25 +PaR 89.47+1.19 94.04+0.81 2.61-0.64 nnFormer 85.50 91.43 5.41 +PaR 88.83+3.33 93.69+2.26 2.35-3.06 TransBTS 87.63 92.84 3.54 +PaR 88.36+0.73 93.27+0.43 2.90-0.64 MultiResUNet 85.92 91.35 9.04 +PaR 86.76+0.84 91.85+0.50 3.68-5.36 3D UX-NET 88.40 93.31 8.85 +PaR 89.24+0.84 93.84+0.53 2.43-6.42
[0052] The present invention tests the performance of 11 benchmark models after integrating PaR on the FLARE dataset, and the results are shown in Table 1. As can be seen from Table 1, the performance of all basic models has been improved after integrating PaR to decouple the axial information of the coarse mask and enhance it separately. Performance indicators including average intersection-over-union ratio and Dice similarity coefficient show that the PaR method improves the accuracy of the overall segmentation performance of the model, while the 95% Hausdorff distance shows that the PaR method improves the reliability of the model boundary performance. After PaR, the average intersection-over-union ratio and Dice similarity coefficient indicators of all benchmark models have increased by 0.39% to 3.33% and 0.2% to 2.26% respectively, which reflects that the multi-head attention of the PDI pipeline to global information enhances the model's ability to analyze the overall spatial features, and can extract the most important information from complex spatial features, which improves the overall segmentation performance of the model. The integrated PaR also significantly improves the boundary segmentation performance of the baseline model. The 95% Hausdorff distance index of U-Net drops from 16.31 to 2.49, which shows that the SAI pipeline not only learns and understands the axial spatial information, but also corrects the obvious errors of independent segmentation.
[0053] Table 2. Benchmark experimental results of OIMHS dataset
[0054] Method IOU Dice HD95 3D U-Net 86.60 92.49 3.40 +PaR 87.55+0.95 93.08+0.59 2.89-0.51 RAUNet 84.52 91.14 13.61 +PaR 86.20+1.68 92.25+1.11 5.37-8.24 ResUNet 84.06 90.84 3.92 +PaR 87.07+3.01 92.79+1.95 3.33-0.59 SegResNet 83.59 90.52 12.05 +PaR 84.82+1.23 91.35+0.83 5.07-6.98 V-Net 81.00 88.53 18.13 +PaR 83.21+2.21 90.26+1.73 13.52-4.61 UNETR 81.52 89.05 29.15 +PaR 83.73+2.21 90.55+1.50 7.24-21.91 Swin UNETR 87.11 92.82 5.21 +PaR 87.92+0.81 93.27+0.45 2.92-2.29 nnFormer 80.54 88.29 25.32 +PaR 85.50+4.96 91.80+3.51 7.36-17.96 TransBTS 79.40 87.39 33.52 +PaR 83.68+4.28 90.55+3.16 21.00-12.52 MultiResUNet 86.53 92.44 3.23 +PaR 88.22+1.69 93.49+1.05 2.85-0.38 3D UX-NET 87.45 93.01 4.61 +PaR 88.51+1.06 93.66+0.65 2.65-1.96
[0055] The performance on the OIMHS dataset is shown in Table 2. As can be seen from the table, the performance of all the benchmark models has been steadily improved after integrating PaR to decouple and re-enhance the axial features. After integrating PaR, the average intersection-over-union ratio and Dice similarity coefficient of the benchmark models have increased by 0.81% to 4.96% and 0.45% to 3.51%, respectively. After integrating PaR, the 95% Hausdorff distance index of UNETR has dropped from 29.15 to 7.24, reflecting the boundary correction function that comes with PaR's centralized decoupling of axial information. The PaR method still has a good improvement effect on the latest SOAT model. When the original segmentation mask of 3DUXNET was relatively high-quality (Dice similarity coefficient index of 93.01), it still achieved an average intersection-over-union ratio improvement of 1.06% and a Dice similarity coefficient improvement of 0.65% after integrating PaR. This shows that PaR's way of learning and extracting features from the segmentation map from the perspective of axial space is different from the idea of analyzing and extracting features of conventional three-dimensional image segmentation models. The ability to improve performance can surpass the boundaries of traditional three-dimensional convolution, so it can still achieve certain improvements on the field SOTA model.
[0056] Table 3. Benchmark experimental results on the SegTHOR dataset
[0057] Method IOU Dice HD95 3D U-Net 78.69 87.59 4.32 +PaR 82.08+3.39 89.82+2.23 2.98-1.34 RAUNet 79.34 88.13 14.58 +PaR 81.04+1.70 89.08+0.95 2.99-11.59 ResUNet 79.61 88.26 3.22 +PaR 80.23+0.62 88.71+0.45 3.14-0.08 SegResNet 77.78 86.99 3.38 +PaR 81.77+3.99 89.60+2.61 2.87-0.51 V-Net 75.17 85.12 16.10 +PaR 76.23+1.06 85.81+0.69 11.49-4.61 UNETR 73.76 84.03 4.71 +PaR 74.13+0.37 84.18+0.15 4.65-0.06 Swin UNETR 78.19 87.26 3.87 +PaR 78.51+0.32 87.42+0.16 3.62-0.25 nnFormer 77.27 86.65 5.11 +PaR 79.00+1.73 87.69+1.04 3.51-1.60 TransBTS 77.70 86.88 3.84 +PaR 81.02+3.32 89.13+2.25 3.75-0.09 MultiResUNet 79.87 88.53 26.75 +PaR 80.81+0.94 89.10+0.57 11.06-15.69 3D UX-NET 78.30 87.34 4.69 +PaR 79.01+0.71 87.77+0.43 3.61-1.08
[0058] The effects of PaR for axial feature decoupling enhancement on the SegThOR dataset are shown in Table III, and the performance of the basic model is steadily improved after integrating PaR. Among all the benchmark models, in terms of the overall segmentation performance, the improvement ranges of the mean intersection over union and Dice similarity coefficient indicators after integrating PaR are 0.32% - 3.99% and 0.15% - 2.61%. The boundary correction effect of three-dimensional axial decoupling enhancement can also be observed on the SegThOR dataset. After integrating PaR into RAUNet, the 95% Hausdorff distance indicator decreased from 14.58 to 2.99.
[0059] The experimental results on the three public datasets of FLARE, OIMHS, and SegTHOR show that connecting PaR after the decoder of the benchmark model for axial information decoupling and independent enhancement can achieve stable performance improvement, can perform result self-correction on models with poor effects, and this improvement can break through the upper limit of traditional convolutional structure design and still provide performance gains in the SOTA three-dimensional segmentation methods, proving the effectiveness and reliability of three-dimensional axial decoupling for three-dimensional medical image data analysis.
[0060] Table IV Ablation Experiment I
[0061] Method mIoU Dice HD95 UNETR 81.52 89.05 29.15 SAI 81.88 89.21 13.35 PDI 82.88 89.98 22.59 PDR-cat 82.61 89.77 16.87 PDR-mult 82.07 89.42 27.31 PDR-add 83.73 90.55 7.24
[0062] Table V Ablation Experiment II
[0063] Method mIoU Dice HD95 UNet 87.92 93.08 16.31 SAI 88.45 93.49 20.33 PDI 88.92 93.73 10.72 PDR-cat 88.67 93.56 8.24 PDR-mult 88.58 93.48 13.37 PDR-add 89.20 93.89 2.49
[0064] To verify the performance differences of the proposed method on traditional convolutional structures and Transformer structures, we respectively used U-Net and UNETR as the basic models to conduct ablation experiments on the key pipeline components and pipeline fusion methods of PaR on the FLARE2021 dataset and the OIMHS dataset, as shown in Table IV and Table V. When only integrating the PDI pipeline or the SAI pipeline, the model performance was improved. For the fusion method selection of the PDI pipeline and the SAI pipeline in the post-PaR method, we experimented with three fusion methods: addition, element-wise matrix multiplication, and concentrate. Among them, the best effect was achieved when using addition, and all indicators were improved, indicating the structural complementarity of the two key pipelines. The improvement degrees of the PDI pipeline or the SAI pipeline on different datasets and different models are not the same, but after using PaR that combines the two methods, the model exceeds all basic models in terms of segmentation accuracy and boundary fineness, which proves that decoupling the axial information of three-dimensional image data and using feature learning enhancement strategies such as attention mechanisms and information redistribution can produce certain performance gains.
[0065] In summary, the PaR method proposed by the present invention decouples and analyzes the feature information under different axis space state distributions from a global perspective, performs spatial decoupling on the segmentation coarse mask of the deep learning model from the axial depth perspective, and learns and enhances the information distribution characteristics unique to the axial space, significantly improving the segmentation accuracy of the three-dimensional medical image segmentation benchmark model and having broad application prospects.
[0066] The above content is only the preferred implementation manner of the present invention and does not limit the protection scope of the present invention. It should be emphasized that any minor adjustments and improvements made by those skilled in the art without departing from the basic principles of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A post-decoder 3D axial decoupling enhancement method for 3D medical image segmentation, characterized in that: The method comprises the following steps: S1: Collect 3D medical image segmentation datasets corresponding to various organ segmentation tasks from the Internet; S2: Divide the collected 3D medical image data into a training set and a test set in proportion; S3: performing data preprocessing and data enhancement on the three-dimensional medical image data in step S2; S4: Collect advanced deep learning models widely used for 3D medical image segmentation from the Internet as benchmark models; S5: Design a post-decoder coarse mask volume axial decoupling enhancement module that can be used for coarse mask optimization in various 3D medical image segmentation tasks; the designed module includes two main pipelines: a permutation dimension enhancement pipeline and a sequence axial enhancement pipeline; The post-decoder coarse mask volume axial decoupling enhancement module adopts two parallel 3D decoupling enhancement pipelines: the permutation dimension enhancement pipeline PDI and the sequence axial enhancement pipeline SAI, which extract key spatiotemporal features from different spatial perspectives and perform 3D axial decoupling from the temporal axis direction, thereby comprehensively refining the segmentation coarse mask; the permutation dimension enhancement pipeline learns and understands different axial features from the perspective of global information, while the sequence axial enhancement pipeline directly analyzes and enhances the temporal axial features from the perspective of axial spatial information; The outputs of the permutation dimension enhancement pipeline and the sequence axis enhancement pipeline are combined by addition, and then the values are compressed to the range of (0,1) by the sigmoid function, and then combined with the coarse mask y o After multiplication, it is added to the coarse mask y through the residual connection o The final optimized and enhanced segmentation mask is obtained, and the formula is as follows: where y r ∈R C×W×H×D is the final optimized and enhanced segmentation map. Represents the refined segmentation mask from the global perspective output by the PDI pipeline; Represents the refined segmentation mask from the timeline perspective output by the SAI pipeline; S6: Integrate the post-decoder 3D axial decoupling enhanced PaR module into the collected benchmark models; S7: Use the baseline model and the baseline model integrated with PaR on the preprocessed and data augmented datasets to perform model training and save the respective best model weights according to the evaluation results on the validation set; S8: Validate the saved best model weights on the test set to evaluate model performance.
2. The post-decoder 3D axial decoupling enhancement method for 3D medical image segmentation according to claim 1, characterized in that: First, define the decoder output of the deep learning model for segmentation as a coarse mask y o ∈R C×W×H×D , where R represents a real number set, C represents the number of segmentation categories, W, H and D represent the width, height and time axis depth of the three-dimensional medical image data respectively; the PDI pipeline analyzes and processes the coarse mask y o The key time axis feature information under different spatial distributions is as follows: For the decoder of the deep learning model, the coarse mask y is output o ∈R C×W×H×D , the PDI pipeline reconstructs the feature space distribution through a permutation operation; this operation rearranges the coarse mask y o ∈R C×W×H×D Each dimension generates multiple feature maps The W axis is rearranged to the last dimension to enhance the attention of the subsequent attention mechanism on the time axis information. The rearranged feature maps are merged from a five-dimensional tensor into a three-dimensional tensor N = H × D × W, making its format suitable for input into the multi-head self-attention mechanism to extract and analyze the information space features of the time axis z axis. The formula of multi-head self-attention is expressed as: in, is the attention output, SelfAttention() represents multi-head self-attention; Attention Output Reshape into a 5-dimensional tensor through the reshape operation Then, the coarse mask y is transformed into o The dimension of is aligned and recorded as the segmentation mask after attention weighting Finally, the PDI pipeline converts the attention-weighted segmentation mask Perform cumulative summation to generate an optimized segmentation mask Thus, the coarse mask y is optimized by extracting and processing the z-axis information from a global perspective. o .
3. The post-decoder three-dimensional axial decoupling enhancement method for three-dimensional medical image segmentation according to claim 2, characterized in that: The SAI pipeline starts from the coarse mask y o Extract spatiotemporal information from the coarse mask y and learn unified features across different time axis planes, enhancing the baseline model to understand the rich spatiotemporal information across sequences. o Optimize as follows: For the coarse mask y o , the SAI pipeline first passes through an initial feature extraction layer, which contains three-dimensional convolution, group normalization, ReLU activation function and residual connection; the initial feature extraction layer performs preliminary feature extraction and generates the initial feature map y θ ∈R C ×W×H×D , the initial feature map y θ Divided into multiple subsets along the W axis The subsets are further divided into two groups: channel groups and space group Extract key channel features from the channel group and key spatial features from the spatial group; The key channel features and key spatial features are concatenated and reshaped into a coarse mask y o The dimension of time axis decoupling feature map is obtained 4. The post-decoder three-dimensional axial decoupling enhancement method for three-dimensional medical image segmentation according to claim 3, characterized in that: The key channel features of the channel group are obtained as follows: First, the channel group Through adaptive maximum pooling f ω and adaptive mean pooling f η Extract the time axis channel information, then merge it by addition and input it into the adaptive weighting function In the above example, we generate feature maps Among them, W c ∈R 1×1×1×1 and b c ∈R C×1×1×1 are all trainable hyperparameters, W c The initial hyperparameters are set to 1, b c The initial hyperparameters of are set to 0; Then, the feature map The sigmoid function is used to limit the value to the interval (0, 1) and multiply the value by the channel group element by element. Fusion to extract Key channel features in .
5. The post-decoder 3D axial decoupling enhancement method for 3D medical image segmentation according to claim 4, characterized in that: The key spatial features of the space group are obtained as follows: First, the space group After a 3D convolutional layer to capture the spatial features in the local neighborhood of the feature map, it reflects the importance of different spatial regions in the These features are then weighted by an adaptive weighting function Generating spatially weighted feature maps Among them, Conv3D represents the dimensional convolutional layer, W s ∈R 1×1×1×1 and They are also trainable hyperparameters, and their initial hyperparameters are set to 1 and 0 respectively; Then, the spatially weighted feature map After activation by the sigmoid function, its value is ensured to be in the range (0,1) and The fusion method is fused by element-wise multiplication to extract Key spatial features in .
6. The post-decoder three-dimensional axial decoupling enhancement method for three-dimensional medical image segmentation according to claim 5, characterized in that: The SAI pipeline introduces a spatial redistribution method to capture shared features from different spatial states. The specific implementation is as follows: First, decouple the time axis from the feature map The shape is reshaped to Then in the G dimension and Dimensions are exchanged between dimensions to obtain Then resize the swapped tensor back to And add to the time axis decoupling feature map The final output of the SAI pipeline is the optimized segmentation mask.
7. The post-decoder 3D axial decoupling enhancement method for 3D medical image segmentation according to claim 1, characterized in that: The benchmark models include 3D U-Net, V-Net, RAUNet, ResUNet, SegResNet, MultiResUNet, UNETR, SwinUNETR, TransBTS, nnFormer and 3D UX-Net.
8. The post-decoder 3D axial decoupling enhancement method for 3D medical image segmentation according to claim 1, characterized in that: Data preprocessing for step 2 includes re-interpolation, interpolating sequence lengths less than 108 to 108; performing image intensity cropping on image pixels and applying maximum-minimum normalization. The threshold for image intensity cropping is set to [-125, 275] for the FLARE2021 dataset, [0, 400] for the OIMHS dataset, and [-1000, 300] for the SegTHOR dataset. 96x96x96 is randomly cropped for training and verified using the sliding window method. The overlap between images is set to 0.5 during verification. Data augmentation is performed on the experimental data, including random flipping, random rotation, random scaling, and random image intensity shift.
9. The post-decoder 3D axial decoupling enhancement method for 3D medical image segmentation according to claim 7, characterized in that: Step 6 integrates the implemented PaR into the collected benchmark models to obtain the segmentation networks 3D U-Net+PaR, V-Net+PaR, RAUNet+PaR, ResUNet+PaR, SegResNet+PaR, MultiResUNet+PaR, UNETR+PaR, SwinUNETR+PaR, TransBTS+PaR, nnFormer+PaR and 3D UX-Net+PaR.
Citation Information
Patent Citations
CT image lung lobe image segmentation system based on attention mechanism
CN113936011A
3D medical image segmentation model establishment method based on mask modeling and application thereof
CN116664588A