Lymphovessel segmentation method for sensing 3DU-Net based on dynamic integrated structure

By dynamically integrating multiple 3D-UNet models, combining structure perception modules and multi-scale feature fusion strategies, the problem of insufficient robustness of existing lymphatic segmentation models is solved, and high-precision and stable lymphatic segmentation is achieved, supporting clinical diagnosis and treatment.

CN120388035AActive Publication Date: 2025-07-29AFFILIATED HOSPITAL OF JIANGNAN UNIV +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510511613.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing deep learning-based lymphatic segmentation model has insufficient robustness caused by model singularity, making it difficult to achieve efficient and accurate lymphatic segmentation in medical imaging.

Method used

Using a method based on dynamic integrated structure-aware 3DU-Net, we use multiple 3D-UNet models with different initialization and hyperparameter settings, combining structure-aware modules and multi-scale feature fusion strategies to dynamically integrate the output of multiple models, and perform morphological post-processing to optimize the segmentation results.

Benefits of technology

It significantly improves the accuracy and robustness of lymphatic tract segmentation, provides more reliable imaging data support, helping to clinical diagnosis and treatment, especially in lymphedema and cancer diagnosis, providing accurate lymphatic tract segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388035A_ABST
    Figure CN120388035A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, in particular to a lymphatic vessel segmentation method for sensing 3DU-Net based on a dynamic integrated structure, and the method comprises a training stage and a use stage. The training stage comprises construction, training and optimization of a plurality of 3D-UNet models, and each 3D-UNet model adopts different initialization and hyper-parameter settings to ensure that the models have diversity and difference. The use stage comprises the steps of integrating prediction results of a plurality of 3D-UNet models, and improving segmentation precision and robustness through a dynamic weight integration strategy. Finally, an integration result is optimized by adopting a post-processing technology, and the accuracy and stability of a segmentation result are ensured. According to the method, a plurality of 3D-UNet models are dynamically integrated, and a structure sensing module and a multi-scale feature fusion strategy are combined, so that the accuracy and robustness of lymphatic vessel segmentation are remarkably improved, and more reliable image data support is provided for clinical diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a lymphatic vessel segmentation method based on a dynamic integrated structure-aware 3D U-Net. Background Art

[0002] Lymphatic vessels are part of the human lymphatic system, mainly responsible for the transportation of lymph fluid and participation in immune responses. They play a crucial role in maintaining the body's fluid balance, transporting fats and metabolic wastes. Lymphatic vessels are distributed throughout the body and are particularly closely connected to other immune tissues through lymph nodes, forming a complex network. The health and function of lymphatic vessels directly affect the body's immune system and play an important role in the occurrence and development of various diseases.

[0003] The accurate segmentation of lymphatic vessels is of great significance in medical imaging, especially in the diagnosis and treatment of cancer, lymphedema, and other diseases related to the lymphatic system. Traditional lymphatic vessel segmentation methods usually rely on doctors' experience, with low efficiency and prone to errors. In recent years, with the development of medical imaging technology, especially the application of deep learning and artificial intelligence, the automatic segmentation of lymphatic vessels has become possible.

[0004] With the rapid development of computer vision technology, especially the application of deep learning algorithms, the accuracy of lymphatic vessel segmentation has been continuously improved. Lymphatic vessel segmentation methods based on deep learning models such as 3D-UNet can efficiently and automatically process a large amount of medical image data and provide reliable segmentation results. However, although existing deep learning technologies (such as 3D-UNet) can improve the segmentation accuracy, there are still problems of insufficient robustness caused by model singularity.

[0005] Therefore, there is an urgent need for a new technical solution to solve the above technical problems. Summary of the Invention

[0006] The purpose of the present invention is to overcome the problems of the above existing technologies, and provide a lymphatic vessel segmentation method based on a dynamic integrated structure-aware 3D U-Net to solve the technical problem of insufficient robustness caused by model singularity in the existing deep learning-based lymphatic vessel segmentation models.

[0007] The above purpose is achieved through the following technical solutions:

[0008] The lymphatic vessel segmentation method based on a dynamic integrated structure-aware 3D U-Net includes a training stage and a usage stage, wherein:

[0009] Training Stage

[0010] Construct multiple pre-training 3D-UNet models with different initializations and hyperparameter settings, and each of the pre-training 3D-UNet models adopts a differentiated initialization strategy;

[0011] Preprocess the MRI image data to generate a training dataset;

[0012] Based on the training dataset, and using a structure-aware module and a multi-scale feature fusion strategy to train each of the pre-training 3D-UNet models to obtain post-training 3D-UNet models;

[0013] Fuse the outputs of multiple post-3D-UNet models through a dynamic ensemble learning strategy to obtain an integrated segmentation result; Usage phase

[0014] Preprocess the MRI image data to be segmented to generate a dataset of images to be processed;

[0015] Use multiple post-training 3D-UNet models to respectively predict the dataset of images to be processed and obtain corresponding prediction results;

[0016] Adopt a dynamic weight integration strategy to fuse each of the prediction results to obtain a final segmentation result;

[0017] Perform morphological post-processing on the final segmentation result to optimize the accuracy and stability of the segmentation result.

[0018] Further, the initialization strategy of the pre-training 3D-UNet model includes at least one of the following combinations:

[0019] Convolution layer initialization: Xavier normal distribution, Kaiming orthogonal initialization, orthogonal initialization, sparse initialization, or prediction-level weight transfer;

[0020] Normalization layer initialization: He uniform distribution, zero-mean Gaussian distribution, fixed scaling factor, dynamic range scaling, or freezing the first 3 layers;

[0021] Learning rate strategy: cosine annealing, stepwise descent, adaptive AdamW, cyclic learning rate, or linear warmup.

[0022] Further, there are 5 pre-training 3D-UNet models, including:

[0023] Pre-training 3D-UNet model 1:

[0024] The convolution layer is initialized with the Xavier normal distribution;

[0025] The normalization layer is initialized with the He uniform distribution;

[0026] The learning rate adopts the cosine annealing strategy;

[0027] Pre-training 3D-UNet Model 2:

[0028] The convolutional layer implements Kaiming orthogonal initialization;

[0029] The parameters of the normalization layer are initialized to a Gaussian distribution with (σ = 0.1);

[0030] The learning rate decreases in a stepwise manner;

[0031] Pre-training 3D-UNet Model 3:

[0032] The convolutional layer uses strict orthogonal initialization to satisfy the constraint condition of W T W = I;

[0033] The scaling factor of the normalization layer is fixed to a constant initialization of γ = 0.8;

[0034] The optimizer uses adaptive AdamW;

[0035] Pre-training 3D-UNet Model 4:

[0036] The convolutional layer implements sparse initialization, setting 50% of the weights to 0;

[0037] The normalization layer uses dynamic range scaling initialization;

[0038] The learning rate changes cyclically according to a triangular period;

[0039] Pre-training 3D-UNet Model 5:

[0040] The convolutional layer uses predictive-level weight transfer, inheriting the convolutional kernel parameters from a pre-trained liver segmentation model;

[0041] The normalization layer freezes the γ and β parameters of the first three layers;

[0042] The learning rate has a linear warmup: (the first 5 epochs);

[0043] Furthermore, in the preprocessing of the MRI image data and the preprocessing of the MRI image data to be segmented, the preprocessing methods are the same, including:

[0044] Perform rigid registration on the MRI image, and use 6-degree-of-freedom affine transformation to align;

[0045] Perform isotropic resampling based on B-spline interpolation on MRI images of different sizes, unify them to the network input size of 512×512×32, and the voxel spacing is 1mm×1mm×3mm;

[0046] Adopt a dynamic data augmentation strategy, including adaptive adjustment of window width and window level, spatial transformation, and noise injection.

[0047] Furthermore, the structure perception module includes an edge-enhanced Dice loss and a topology-preserving loss, and the loss function is:

[0048]

[0049] where is the edge-enhanced Dice loss, is the topology-preserving loss function, and its functional expression is as follows:

[0050]

[0051] where y edge is extracted through the Canny operator (threshold is 0.2, σ is 1.0);

[0052]

[0053] where χ() is the Euler number feature, used to constrain the tubular topology of lymphatic vessels; P = 8-neighborhood connectivity.

[0054] Furthermore, the dynamic ensemble learning strategy includes:

[0055] Calculate the weight of each model based on the Dice score of the validation set, and fuse the prediction results of multiple models using weighted averaging. The formula is as follows:

[0056]

[0057] where the weight calculation is:

[0058]

[0059] where Dice m is the Dice score of the m-th model on the validation set.

[0060] Furthermore, the morphological post-processing includes:

[0061] Perform anisotropic closing using an oblate structural element, with the parameter I post = I·S 3×3×1 ;

[0062] Eliminate fragments and retain the tubular structure according to the connected component screening condition. The retention condition is:

[0063] V region > 8 voxels and

[0064] Furthermore, it is characterized in that it includes:

[0065] A data preprocessing module, which is used to preprocess the input MRI image data and generate a standardized three-dimensional matrix;

[0066] A data augmentation module, which is used to improve the robustness of the model to differences in scanning parameters;

[0067] A multi-model dynamic integration inference module, including multiple 3D-UNet models, which is used to predict the MRI image data and output corresponding prediction results;

[0068] A dynamic integration module, which is used to perform weighted fusion on the outputs of the 3D-UNet models to obtain the final segmentation result;

[0069] A post-processing module, which is used to perform morphological post-processing on the final segmentation result to optimize the accuracy and stability of the segmentation result.

[0070] The lymphatic vessel segmentation method based on dynamic integration structure-aware 3DU-Net provided by the present invention significantly improves the accuracy and robustness of lymphatic vessel segmentation by dynamically integrating multiple 3D-UNet models, combining a structure-aware module and a multi-scale feature fusion strategy, and provides more reliable image data support for clinical diagnosis and treatment. Description of the Drawings

[0071] Figure 1 It is a flowchart of the lymphatic vessel segmentation method based on dynamic integration structure-aware 3DU-Net according to the present invention;

[0072] Figure 2 It is a block diagram of the lymphatic vessel segmentation system based on dynamic integration structure-aware 3DU-Net according to the present invention;

[0073] Figure 3 It is a structural diagram of the 3D-UNet model in the lymphatic vessel segmentation method based on dynamic integration structure-aware 3DU-Net according to the present invention;

[0074] Figure 4 It is a training curve diagram of the 3D-UNet dynamic integration model in the lymphatic vessel segmentation method based on dynamic integration structure-aware 3DU-Net according to the present invention. Detailed Embodiments

[0075] The present invention will be further described in detail below with reference to the drawings and embodiments. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0076] As Figure 1As shown in the figure, this solution provides a lymphatic vessel segmentation method based on a dynamically integrated structure-aware 3D U-Net, including a training stage and a usage stage. The training stage includes the construction, training, and optimization of multiple 3D U-Net models, where each 3D U-Net model adopts different initializations and hyperparameter settings to ensure the diversity and difference of the models; the usage stage includes integrating the prediction results of multiple 3D U-Net models and improving the segmentation accuracy and robustness through a dynamic weight integration strategy; finally, post-processing techniques are used to optimize the integration results to ensure the accuracy and stability of the segmentation results. Specifically:

[0077] During the training stage

[0078] Construct multiple pre-training 3D U-Net models with different initializations and hyperparameter settings, and each of the pre-training 3D U-Net models adopts a differentiated initialization strategy;

[0079] Preprocess the MRI image data to generate a training dataset;

[0080] Based on the training dataset, and using a structure-aware module and a multi-scale feature fusion strategy to train each of the pre-training 3D U-Net models to obtain post-training 3D U-Net models; specifically, in this step, the structure-aware module uses an edge-enhanced Dice loss and a topology-preserving loss (L Topo ) to ensure that the model not only focuses on the overall segmentation accuracy during training, but also particularly strengthens the learning of lymphatic vessel boundaries (such as fine branches) and tubular topologies (such as connectivity); while multi-scale feature fusion uses the multi-level convolutional structure of 3D U-Net to simultaneously extract local details (shallow features) and global context (deep features) to solve the problem of large differences in the sizes of lymphatic vessels in MRI images (such as the main trunk and fine branches). Through the training of this step, each post-training 3D U-Net model can independently capture different features of lymphatic vessels (such as model 3 being sensitive to small-scale structures and model 4 being robust to noise), and due to initialization differences, different models may produce different error patterns for the same image after training (such as model 1 having more accurate boundary segmentation and model 5 being more robust to low-contrast regions), providing a diversity basis for subsequent integration.

[0081] Fuse the outputs of multiple post-3D U-Net models through a dynamic ensemble learning strategy to obtain an integrated segmentation result, providing an optimization basis and performance verification basis for the dynamic weight integration in the subsequent usage stage, so as to ensure that the dynamic weight integration and post-processing in the usage stage can stably output segmentation results that meet clinical requirements; this process reflects the closed-loop design idea of "training verification - feedback optimization". The specific functions include:

[0082] Reduce the bias of a single model: Through weighted averaging Fuse the prediction results of multiple models to reduce errors caused by model initialization or training randomness.

[0083] Adaptive weight allocation: According to the Dice score of the validation set Dynamically adjust the model weights so that models with better performance contribute more to the final result (as shown in Table 3, the integrated Dice score -0.67156 -0.67156 is better than a single model).

[0084] The functions of the segmentation results include:

[0085] Clinical assistant diagnosis: Provide accurate three-dimensional segmentation results of lymphatic vessels to help doctors locate lesions (such as lymphedema or tumor metastasis pathways).

[0086] Quantitative analysis: Used to measure the morphological parameters of lymphatic vessels (such as diameter, number of branches), supporting disease progression assessment or surgical planning.

[0087] Research tool: Provide reliable data support for the mechanism research of lymphatic system-related diseases (such as cancer lymphatic metastasis).

[0088] In brain MRI, the integrated model can avoid a single model misjudging vascular artifacts as lymphatic vessels (constraining the topology through the structure perception module), while retaining small lymphatic vessels (through multi-scale feature fusion), and finally output can be used to evaluate the severity of the condition of patients with lymphatic reflux disorders.

[0089] In this embodiment, the integrated segmentation result generated through dynamic integration in the training stage essentially establishes an "expert decision-making system", specifically:

[0090] In the training stage, the system learns the weight allocation rules of each sub-model (3D-UNet1-5) through the validation set data (such as the private brain MRI dataset of the Affiliated Hospital of Jiangnan University);

[0091] In the usage stage, the system directly applies the weight knowledge learned in the training stage to new data without recalculating the weights. As shown in Table 3, the integrated Dice score (-0.67156) is better than any sub-model, and this optimization ability is completely transferred to the application scenario.

[0092] In addition, the integration results in the training stage (as shown in Table 3) and the output in the usage stage share the same set of fusion algorithms, ensuring:

[0093] Preprocessing consistency: The B-spline interpolation and window width and window level adjustment in the training stage are followed in the usage stage;

[0094] Postprocessing consistency: Anisotropic closing operations are both used for morphological optimization.

[0095] This design enables the accuracy (such as Dice score) verified during the training phase to be directly extended to clinical applications.

[0096] When processing pediatric brain MRIs (lymphatic vessels are finer), the known model 3 (fixed scaling factor γ = 0.8) during the training phase is sensitive to fine structures, and the system will automatically increase its weight; when processing images of elderly patients (with more artifacts), the noise suppression characteristics of model 4 (sparse initialization) will obtain a higher weight. This dynamic adjustment ability is completely inherited from the integration rules established during the training phase.

[0097] During the usage phase

[0098] Preprocess the MRI image data to be segmented to generate a dataset of images to be processed;

[0099] Use multiple of the trained 3D-UNet models to respectively predict the dataset of images to be processed and obtain corresponding prediction results;

[0100] Adopt a dynamic weight integration strategy to fuse the prediction results to obtain a final segmentation result;

[0101] Perform morphological post-processing on the final segmentation result to optimize the accuracy and stability of the segmentation result.

[0102] This embodiment includes during the training phase:

[0103] Initial 3D-UNet model construction phase: Using medical imaging principles and deep learning techniques, construct multiple 3D-UNet models with different initializations and hyperparameter settings by preprocessing and enhancing MRI data.

[0104] Specifically, adopt a differential model initialization strategy, and the 5 sub-models adopt the initialization combinations shown in Table 1. Each model processes different lymphatic vessel image features respectively to ensure that the network has good learning ability for data diversity, thereby constructing the initial 3D-UNet model;

[0105] Integrated model training phase: Train multiple 3D-UNet models with different initializations and hyperparameters, and use metrics such as structure-aware loss to optimize the model parameters to ensure that each model has good performance in the lymphatic vessel segmentation task and effectively improve the segmentation accuracy;

[0106] Mean integration strategy phase: Use the prediction results of multiple trained 3D-UNet models to fuse the prediction results of each model through mean integration to obtain the final integrated segmentation result.

[0107] As shown in Table 1, it is the model initialization strategy table. In this embodiment, the initialization strategy of the pre-training 3D-UNet model includes at least one of the following combinations:

[0108] Convolutional layer initialization: Xavier normal distribution, Kaiming orthogonal initialization, orthogonal initialization, sparse initialization, or prediction-level weight transfer;

[0109] Normalization layer initialization: He uniform distribution, zero-mean Gaussian distribution, fixed scaling factor, dynamic range scaling, or freezing the first 3 layers;

[0110] Learning rate strategy: cosine annealing, stepwise descent, adaptive AdamW, cyclic learning rate, or linear warmup.

[0111] Table 1 Model Initialization Strategy

[0112] Model Number Convolution Layer Initialization Normalization Layer Initialization Learning Rate Policy 1 Xavier Normal He Uniform Distribution Cosine Annealing 2 Kaiming Orthogonal Zero-Mean Gaussian (σ = 0.1) Step Decay 3 Orthogonal Initialization Fixed Scaling Factor (γ = 0.8) Adaptive AdamW 4 Sparse Initialization Dynamic Range Scaling Cyclic Learning Rate 5 Pre-trained Weight Transfer Freeze the First 3 Layers Linear Warmup

[0113] As shown in Table 1, there are 5 pre-training 3D-UNet models, including:

[0114] Pre-training 3D-UNet Model 1:

[0115] The convolutional layer is initialized with the Xavier normal distribution, and the mathematical expression is where n in and n out are the number of input and output channels respectively;

[0116] The normalization layer is initialized with the He uniform distribution and follows distribution;

[0117] The learning rate adopts the cosine annealing strategy: where η max = 1e -3 , η min = 1e -5 ;

[0118] Technical effect: Maintaining the variance stability of the activation values of each layer, suitable for processing standard lymph nodes;

[0119] Pre-training 3D-UNet Model 2:

[0120] The convolutional layer implements Kaiming orthogonal initialization, and an orthogonal matrix W = QR is generated through QR decomposition, where Q is the orthogonal basis;

[0121] The parameters of the normalization layer are initialized to a Gaussian distribution with (σ = 0.1);

[0122] The learning rate decreases stepwise: multiplying by a decay factor of 0.5 every 20 epochs;

[0123] Technical effect: The orthogonality constraint makes the gradient update direction more stable, which is suitable for long-range lymphatic vessel tracking;

[0124] 3D-UNet model 3 before training:

[0125] The convolutional layer adopts strict orthogonal initialization, satisfying the constraint condition of W T W = I;

[0126] The scaling factor of the normalization layer is initialized with a fixed constant of γ = 0.8;

[0127] The optimizer uses adaptive AdamW, satisfying AdamW(β1 = 0.9, β2 = 0.99);

[0128] Technical effect: The forced low activation intensity is suitable for fine lymphatic vessel segmentation;

[0129] 3D-UNet model 4 before training:

[0130] The convolutional layer implements sparse initialization, setting 50% of the weights to 0;

[0131] The normalization layer adopts dynamic range scaling initialization;

[0132] The learning rate changes cyclically according to the triangular period: The period T = 100 steps;

[0133] Technical effect: Simulating the sparse connection characteristics of biological neurons enhances the recognition of broken lymphatic vessels;

[0134] 3D-UNet model 5 before training:

[0135] The convolutional layer adopts predictive-level weight transfer, inheriting the convolutional kernel parameters W from the pre-trained liver segmentation model tpretrained ;

[0136] The normalization layer freezes the γ and β parameters of the first three layers;

[0137] The learning rate has a linear warmup: (for the first 5 epochs);

[0138] Technical effect: Transfer learning significantly improves the generalization ability under small samples.

[0139] This design enables the integrated system to have:

[0140] Feature diversity: Each model focuses on lymphatic vessel features at different scales;

[0141] Error decorrelation: The missegmentation of a single model can be corrected by other models;

[0142] Scene adaptability: Different initialization combinations can handle various imaging quality problems.

[0143] Experimental data (Table 2) shows that the Dice scores of the 5 models on the test set are distributed between -0.6570 and -0.6769, verifying the effective differences in the initialization strategy. After dynamic weight integration (Table 3), the final Dice score is improved to -0.67156, proving that this technical solution achieves an integration effect of 1 + 1 > 2.

[0144] This embodiment in the usage stage includes:

[0145] Obtain the MRI image data to be segmented, and generate the initial features of the image to be processed through the same preprocessing technology ① (such as denoising, standardization, etc.) as in the training stage;

[0146] Then, use multiple trained 3D-UNet models to predict the MRI image respectively, generating the segmentation results of multiple models;

[0147] Then, use the mean integration strategy to fuse the prediction results of each model to obtain the final lymphatic vessel segmentation result;

[0148] Finally, adopt morphological post-processing technology to further optimize the segmentation result, remove artifacts and enhance the boundary accuracy of the lymphatic vessels, ensuring the accuracy and stability of the finally output segmentation result.

[0149] Among them, the MRI image is a digital image generated by scanning human tissues, containing a large amount of pixel information. The deep convolutional neural network cannot directly process the original image data, so it is necessary to preprocess the MRI image and convert it into a numerical form that the network can accept.

[0150] In this embodiment, the preprocessing of the MRI image data and the preprocessing of the MRI image data to be segmented have the same preprocessing method, including:

[0151] Perform rigid registration on the MRI image and use 6-degree-of-freedom affine transformation to align;

[0152] Perform isotropic resampling based on B-spline interpolation on MRI images of different sizes, unify them into the network input size of 512×512×32, and the voxel spacing is 1mm×1mm×3mm;

[0153] Adopt a dynamic data augmentation strategy, including adaptive adjustment of window width and window level, spatial transformation (such as Z-axis sampling, random flipping) and noise injection (such as Gaussian noise).

[0154] Specifically, the original DCM data are first synthesized into a patient-level MRI dataset, and the MRI images are rigidly registered (using a 6-DOF affine transformation) with the data center as the unit to ensure that the data sets of all centers have the same direction.

[0155] Next, isotropic resampling based on B-spline interpolation is performed on MRI images of different sizes to unify them into a network input size of 512×512×32 (voxel spacing 1mm×1mm×3mm). Specifically, a third-order B-spline interpolation is performed on the z-axis:

[0156]

[0157] in, The image resolution error after interpolation is controlled within 0.5mm.

[0158] To improve model generalization, a dynamic data augmentation strategy was employed. The window width and window position were adaptively adjusted: W = μ + 3σ, L = μ - σ. μ is the image mean and σ is the standard deviation, covering 99.7% of the soft tissue signal range. Spatial transformations included z-axis resampling (±15% pitch jitter), random rotation (range ±25°), and flipping (50% probability). Noise injection was performed: Gaussian noise (σ∈[0,0.1]) and Rayleigh noise (intensity ratio 1:3) were added to simulate MRI scanning artifacts.

[0159] While the aforementioned MRI image preprocessing methods can help extract some basic features, due to the vast amount of information contained in images, a single preprocessing method cannot fully exploit all valuable features. Therefore, employing more complex feature extraction methods can further enhance the model's expressiveness. Compared to traditional two-dimensional features, three-dimensional features can more comprehensively capture the spatial information of the image, thereby improving segmentation performance.

[0160] MRI image data contains rich spatial structural information, and preserving this spatial information is particularly important when performing lymphatic vessel segmentation. To this end, this method uses a 3D convolutional neural network (3D-UNet) to extract features from MRI images, fully utilizing the three-dimensional structure of the images. Unlike conventional two-dimensional convolutional networks, 3D convolution can directly process the depth, width, and height of the image, thereby more accurately capturing the spatial relationships in the image. Through the 3D convolutional layer, the network can extract deep spatial features, which are more expressive for the lymphatic vessel segmentation task.

[0161] Due to the differences in the resolution of MRI images and the size of the target area, single-scale feature extraction may not cover all details. Therefore, this method introduces a multi-scale feature fusion strategy in 3D-UNet. Through convolutional operations at different scales, the network can extract detailed information in the image at different levels and fuse this information together to further improve the segmentation accuracy. Multi-scale feature fusion can effectively capture lymphatic vessels of different sizes and shapes, avoiding segmentation errors caused by inconsistent scales.

[0162] In addition to spatial information and multi-scale features, the morphological structure of lymphatic vessels is also important. Therefore, this method also post-processes the output of the 3D convolutional network through morphological operations (such as dilation, erosion, etc.) to further enhance the morphological information of lymphatic vessels. These operations help to remove artifacts, strengthen the boundaries of lymphatic vessels, and ensure more accurate segmentation results.

[0163] The morphological post-processing includes:

[0164] Performing anisotropic closing operation using an oblate structural element with parameter I post = I·S 3×3×1 ;

[0165] Filtering out fragments and retaining tubular structures through connected component filtering conditions, and the retention conditions are: retaining tubular structures with a volume greater than 8 pixels and a surface-to-volume ratio value less than 2.7:

[0166] V region > 8 voxels and

[0167] Specifically, the parameter of the anisotropic closing operation is I post = I·S 3×3×1 , because the axial continuity is stronger than the coronal plane, so an oblate structural element S 3×3×1 is used; at the same time, connected component filtering conditions are used, and the retention conditions are: V region > 8 voxels and

[0168] Thus, fragments are removed and tubular structures are retained (volume threshold 8 voxels, surface-to-volume ratio threshold 2.7)

[0169] By combining the advantages of 3D convolution and multi-scale feature extraction, this method can effectively improve the segmentation accuracy of lymphatic vessels in MRI images and provide more accurate image data for subsequent medical analysis.

[0170] As a specific embodiment of this solution, it includes:

[0171] Step 1: Use the preprocessing matrix of the original MRI image data as the initial features, and perform operations such as denoising and normalization so that the network can accept and effectively extract image features.

[0172] Step 2: Utilize medical imaging principles and deep learning techniques to input the processed MRI image data into multiple 3D-UNet models, and extract lymphatic vessel features in the images through convolutional neural networks; each model is trained under different initializations and hyperparameter settings to ensure the performance of the model under different image features.

[0173] Step 3: Adopt a dynamic ensemble learning strategy to fuse the outputs of multiple trained 3D-UNet models, and obtain the final segmentation result through the mean ensemble strategy. The advantage of ensemble learning is that it can reduce the bias of a single model and improve the overall segmentation accuracy and robustness.

[0174] As Figure 3 shown, it is the structure diagram of one of the five ensemble 3D-UNet models. The input and output of the 3D MRI image are represented by light blue and light red respectively.

[0175] Assume the input shape is (1, 32, 512, 512), which means the input received by the model is a 3D tensor, where 1 is the batch size, 32 is the number of input channels, and 512x512 is the spatial dimension.

[0176] The input is processed through DoubleConv3D (consisting of two layers of convolution, batch normalization, and ReLU activation). This layer transforms the 32 input channels into 64 channels, and the output shape is (1, 64, 512, 512).

[0177] Downsampling stage:

[0178] The first downsampling: Use Down3D for downsampling. First, MaxPool3D reduces the dimension of the input tensor to get (1, 64, 256, 256), and then it is processed through DoubleConv3D (increasing the number of channels from 64 to 128), and the output shape is (1, 128, 256, 256).

[0179] The second downsampling: Similar to the first downsampling, the number of channels is increased from 128 to 256, and the output shape is (1, 256, 128, 128).

[0180] The third downsampling: Similarly, the number of channels is increased from 256 to 512, and the output shape is (1, 512, 64, 64).

[0181] Fourth downsampling: Continuing to reduce the dimension, the number of channels increases from 512 to 1024, and the output shape is (1, 1024, 32, 32).

[0182] Upsampling stage:

[0183] First upsampling: Using Up3D to reduce the number of channels from 1024 to 512. First, perform upsampling, and the shape becomes (1, 512, 64, 64). Then, concatenate it with the output of the previous layer (shape (1, 512, 64, 64)), and the resulting shape is (1, 1024, 64, 64). Then, process it through DoubleConv3D, and the number of channels is reduced to 512, and the output shape is (1, 512, 64, 64).

[0184] Second upsampling: The number of channels is reduced from 512 to 256, the shape becomes (1, 256, 128, 128), and after concatenation and processing through DoubleConv3D, the output is (1, 256, 128, 128).

[0185] Third upsampling: The number of channels is reduced from 256 to 128, the shape becomes (1, 128, 256, 256), and after concatenation and processing through DoubleConv3D, the output is (1, 128, 256, 256).

[0186] Fourth upsampling: The number of channels is reduced from 128 to 64, the shape becomes (1, 64, 512, 512), and after concatenation and processing through DoubleConv3D, the output is (1, 64, 512, 512).

[0187] Finally, through the OutConv3D layer, the output channels are reduced from 64 to the required number of classes (for example, for a binary classification task, the number of classes is 1), and the output shape is (1, n_classes, 512, 512).

[0188] After the convolution operations mentioned above, batch normalization will be immediately performed, and its functional expression is as follows:

[0189]

[0190] Subsequently, the ReLU function is used as the activation function, and the ReLU function expression is as follows:

[0191] ReLU(x) = max(0, x)

[0192] This scheme adopts cross-entropy and uses structure-aware weighting as the loss function. The loss function is defined as follows:

[0193]

[0194] Where is the edge-enhanced Dice loss, is the topology-preserving loss function, and its functional expression is as follows:

[0195]

[0196] where y edge is extracted by the Canny operator (threshold is 0.2, σ is 1.0);

[0197]

[0198] where χ() is the Euler number feature, which is used to constrain the tubular topology of the lymphatic vessels; P = 8-neighborhood connectivity.

[0199] Finally, the Dice loss function is used as the evaluation index of the segmentation network, and this function is defined as follows:

[0200]

[0201] where y represents the true lymphatic vessel segmentation annotation, which is a three-dimensional binary matrix, where: 1 (or 255) represents the lymphatic vessel voxel, 0 represents the background area, and it is obtained by manual annotation by the operator on the MRI image; represents the segmentation result predicted by the model, which is a three-dimensional probability matrix. After sigmoid activation, the value ∈[0,1]. The closer it is to 1, the higher the probability that the voxel belongs to the lymphatic vessel. Usually, it is binarized with a threshold of 0.5 (if >0.5, it is regarded as positive); represents the number of overlapping voxels calculated between the true annotation and the prediction result. The implementation method is to sum after element-wise multiplication of y and reflects the lymphatic vessel area correctly recognized by the model; |y| represents the total number of lymphatic vessel voxels in the true annotation, represents the total number of voxels predicted by the model to be lymphatic vessels, as the denominator is essentially the sum of the voxel numbers of the true and predicted regions.

[0202] This evaluation index is sensitive to class-imbalanced data (such as the lymphatic vessels only accounting for 1-5% of the image), and its value range is [0,1]. The larger the value, the higher the segmentation accuracy. If it takes 1, it is a perfect match. If it takes 0, it is a complete mismatch.

[0203] In this scheme, 5 3D-UNets are trained respectively, using different hyperparameters. The main differences are the learning rate change curve, batch size, and loss function. Finally, the prediction results of the 5 models are voted and decided using the dynamic integration method, that is, the dynamic integration learning strategy includes:

[0204] Calculate the weights of each model based on the Dice score of the validation set, and use weighted averaging to fuse the prediction results of multiple models. The formula is as follows:

[0205]

[0206] Among them, the weight is calculated as:

[0207]

[0208] Among them, Dice m is the Dice score of the m-th model on the validation set.

[0209] Specifically, the weight calculation is based on the Dice score of the validation set; the dynamic weight allocation function realizes the weighted fusion of the prediction results of multiple models based on the variant of the Softmax function; the adaptive weighting in model integration is realized through exponential amplification and normalization. Its core advantage is to adjust the weight distribution through the temperature parameter, making the integration result more inclined to the high-performance model. In practical applications, attention should be paid to the selection of the temperature parameter and the reliability verification of the Dice value, as Figure 4 shown.

[0210] As Figure 2 shown, this solution also provides a lymphatic vessel segmentation system based on the dynamic integrated structure-aware 3D U-Net for implementing the lymphatic vessel segmentation method based on the dynamic integrated structure-aware 3D U-Net, including:

[0211] A data preprocessing module for preprocessing the input MRI image data to generate a standardized three-dimensional matrix;

[0212] A data augmentation module for enhancing the robustness of the model to scanning parameter differences;

[0213] A multi-model dynamic integration inference module, including multiple 3D U-Net models, for predicting the MRI image data and outputting the corresponding prediction results;

[0214] A dynamic integration module for weighted fusion of the outputs of the 3D U-Net models to obtain the final segmentation result;

[0215] A post-processing module for morphological post-processing of the final segmentation result to optimize the accuracy and stability of the segmentation result.

[0216] Data flow during work: Original MRI image data → Orientation alignment → Z-axis resampling → Window width adjustment → [Gaussian noise / Random flipping] → 3D U-Net1-5 → Probability fusion → Binarization → Morphological optimization → Lymphatic vessel MASK.

[0217] Adopt a dynamic weight mechanism to adjust the fusion weights in real time according to the Dice performance of each model on the validation set (Table 3).

[0218] Adopt structure-aware training, through and the loss function to constrain the tubular morphology;

[0219] Adopt multi-scale collaboration. Model 1 and Model 3 focus on the main trunk segmentation, Model 2 and Model 4 handle noise / fractures, and Model 5 ensures the stability of small samples.

[0220] The input size of 512×512×32 adapts to the resolution of common MRI devices; noise injection simulates the motion artifacts of cancer patients, and rotation augmentation covers the morphological variations of lymphedema.

[0221] Through the technical path of differential initialization → collaborative training → dynamic integration of this system, an optimization effect with a Dice score of -0.67156 on the test set can be achieved (Table 3), which is about 12% higher than the traditional single-model method.

[0222] As a specific embodiment of this solution, according to the implementation method in the training stage, an embodiment is completed for the private brain MRI dataset of the Affiliated Hospital of Jiangnan University. The final training results are shown in Table 2 (the following are the scores on the test set).

[0223] Table 2 Training Results before Voting

[0224] Model ID Cross Entropy Loss Dice Score 1 0.7289 -0.6617 2 0.7056 -0.6769 3 0.7326 -0.6713 4 0.7302 -0.6759 5 0.7068 -0.6570

[0225] As another specific embodiment of this solution, use the mean integration algorithm to vote on the above 5 3D-UNets to obtain the final model training results in Table 3 (the following are the scores on the validation set).

[0226] Table 3 Final Prediction Results

[0227] Cross Entropy Loss Dice Score 0.72082 -0.67156

[0228] The above is only to illustrate the implementation manner of the present invention and is not used to limit the present invention. For those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A lymphatic vessel segmentation method based on a dynamically integrated structure-aware 3D U-Net, characterized in that, It includes a training phase and a usage phase, where: Training phase Construct multiple pre-training 3D-UNet models with different initializations and hyperparameter settings, and each of the pre-training 3D-UNet models adopts a differentiated initialization strategy; Preprocess the MRI image data to generate a training dataset; Based on the training dataset, and using a structure-aware module and a multi-scale feature fusion strategy to train each of the pre-training 3D-UNet models to obtain post-training 3D-UNet models; Fuse the outputs of multiple post-3D-UNet models through a dynamic ensemble learning strategy to obtain an integrated segmentation result; Usage phase Preprocess the MRI image data to be segmented to generate a dataset of images to be processed; Use multiple post-training 3D-UNet models to predict the dataset of images to be processed respectively and obtain corresponding prediction results; Adopt a dynamic weight integration strategy to fuse each of the prediction results to obtain a final segmentation result; Perform morphological post-processing on the final segmentation result to optimize the accuracy and stability of the segmentation result.

2. The lymphatic vessel segmentation method based on the dynamic integrated structure-aware 3D U-Net according to claim 1, wherein The initialization strategy of the pre-training 3D-UNet model includes at least one of the following combinations: Convolutional layer initialization: Xavier normal distribution, Kaiming orthogonal initialization, orthogonal initialization, sparse initialization, or prediction-level weight transfer; Normalization layer initialization: He uniform distribution, zero-mean Gaussian distribution, fixed scaling factor, dynamic range scaling, or freezing the first three layers; Learning rate strategy: cosine annealing, stepwise descent, adaptive AdamW, cyclic learning rate, or linear warmup.

3. The lymphatic vessel segmentation method based on the dynamically integrated structure-aware 3D U-Net according to claim 2, wherein There are 5 pre-training 3D-UNet models, including: Pre-training 3D-UNet model 1: The convolutional layer is initialized with the Xavier normal distribution; The normalization layer is initialized with the He uniform distribution; The learning rate adopts the cosine annealing strategy; Pre-training 3D-UNet model 2: The convolutional layer implements Kaiming orthogonal initialization; The parameters of the normalization layer are initialized to a Gaussian distribution with (σ = 0.1); The learning rate decreases step by step; Pre-training 3D-UNet model 3: The convolutional layer adopts strict orthogonal initialization, satisfying the constraint condition of W T W = I; The scaling factor of the normalization layer is initialized to a constant of γ = 0.8; The optimizer uses adaptive AdamW; Pre-training 3D-UNet model 4: The convolutional layer implements sparse initialization, setting 50% of the weights to 0; The normalization layer is initialized with dynamic range scaling; The learning rate changes cyclically according to a triangular period; Pre-training 3D-UNet model 5: The convolutional layer adopts prediction-level weight transfer and inherits the convolutional kernel parameters from a pre-trained liver segmentation model; The γ and β parameters of the first three layers of the normalization layer are frozen; Learning rate linear warmup: (for the first 5 epochs).

4. The lymphatic vessel segmentation method based on the dynamically integrated structure-aware 3D U-Net according to claim 1, wherein In the preprocessing of the MRI image data and the preprocessing of the MRI image data to be segmented, the preprocessing methods are the same, including: Perform rigid registration on the MRI image and align it using a 6-degree-of-freedom affine transformation; Perform isotropic resampling based on B-spline interpolation on MRI images of different sizes to unify them into the network input size of 512×512×32, with a voxel spacing of 1mm×1mm×3mm; Adopt a dynamic data augmentation strategy, including adaptive adjustment of window width and window level, spatial transformation, and noise injection.

5. The lymphatic vessel segmentation method based on the dynamically integrated structure-aware 3D U-Net according to claim 1, wherein The structure perception module includes edge enhancement Dice loss and topology preservation loss, and the loss function is: Among them is the edge-enhanced Dice loss is the topology-preserving loss function, and its functional expression is as follows: where y edge extracted by Canny operator (threshold is 0.2, σ is 1.0); where χ() is the Euler number feature, used to constrain the tubular topology of lymphatic vessels; P = 8-neighborhood connectivity.

6. The lymphatic vessel segmentation method based on the dynamically integrated structure-aware 3D U-Net according to claim 5, wherein, The dynamic ensemble learning strategy includes: Calculate the weight of each model based on the Dice score of the validation set, and fuse the prediction results of multiple models using weighted average. The formula is as follows: where the weight calculation is: Among them, Dice m is the Dice score of the m-th model on the validation set.

7. The lymphatic vessel segmentation method based on the dynamically integrated structure-aware 3D U-Net according to claim 6, wherein The morphological post-processing includes: Perform anisotropic closing operation using an oblate structural element, with parameter I post = I · S 3×3×1 ; Filter out fragments and retain the tubular structure according to the connected component screening conditions. The retention conditions are: V region > 8 voxels and 8. The lymphatic vessel segmentation method based on the dynamically integrated structure-aware 3D U-Net according to claim 1, wherein, Include a lymphatic vessel segmentation system based on a dynamic ensemble structure-aware 3D U-Net, including: A data preprocessing module for preprocessing the input MRI image data to generate a standardized three-dimensional matrix; A data augmentation module for enhancing the robustness of the model to differences in scanning parameters; A multi-model dynamic ensemble inference module, including multiple 3D U-Net models, for predicting MRI image data and outputting corresponding prediction results; A dynamic ensemble module for weighted fusion of the outputs of the 3D U-Net models to obtain the final segmentation result; A post-processing module for morphological post-processing of the final segmentation result to optimize the accuracy and stability of the segmentation result.

Citation Information

Patent Citations

  • Intracranial aneurysm detection method based on multi-dimensional feature fusion

    CN113240654A

  • Brain tumor segmentation method based on brain three-dimensional MRI image design

    CN115496771A

  • Integrated MP-Unet segmentation method based on multi-modal MRI

    CN116229071A

  • Multi-stage segmentation method based on multi-modal MRI heart image

    CN116797618A

  • 3D medical image segmentation method based on double U-Net convolutional neural network

    CN117408962A