Cerebrovascular segmentation method based on multi-scale tubular structure enhanced coding convolutional neural network

The multi-scale tubular structure-enhanced CNN improves brain blood vessel segmentation by integrating Frangi filters and selective attention, addressing discontinuities and topology errors in nnU-Net, achieving precise and robust segmentation results.

CN120318173APending Publication Date: 2025-07-15BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385283.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing nnU-Net has problems such as discontinuous segmentation results, incomplete segmentation of small vessels, and errors in the topological structure of vascular segmentation results in the cerebrovascular segmentation task, making it difficult to adapt to structures such as cerebrovascular with large scale changes and complex branches.

Method used

A multi-scale tubular structure is used to enhance the encoded convolutional neural network, combined with the Frangi vascular enhancement filter and the feature enhancement module of the selective attention mechanism, improve the encoder of nnU-Net, enhance the vascular signal strength and suppress noise, and adaptively adjust the weight of structural features and spatial features.

Benefits of technology

The comprehensive capture of cerebrovascular vessels at different scales is achieved, the accuracy and robustness of vascular segmentation are improved, and the accuracy and topological retention of segmentation results are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318173A_ABST
    Figure CN120318173A_ABST
Patent Text Reader

Abstract

The invention provides a cerebrovascular segmentation method based on a multi-scale tubular structure enhanced coding convolutional neural network. According to the method, four TOF-MRA public data sets including IXI-45, Brains, TubeTK-42 and EDEN are adopted for verification, firstly, data are randomly divided into a training set and a test set, and then a deep convolutional neural network model with an enhanced multi-scale tubular structure is constructed. The method has the advantages that the multi-scale tubular structure is designed to enhance the encoder, end-to-end optimization is achieved, and the feature extraction capacity of different-scale blood vessels is improved; and designing a self-adaptive blood vessel feature enhancement module to dynamically optimize the weight, suppress background noise and enhance the blood vessel contrast. Experimental results show that according to the method, the average Dice coefficients on four data sets respectively reach 88.81%, 82.22%, 79.54% and 92.45%, the topology maintenance Dice coefficients are 88.71%, 78.32%, 75.06% and 93.02%, the average surface distances are 0.1594 mm, 0.3119 mm, 0.6246 mm and 0.2222 mm, and all indexes are superior to those of an existing optimal method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a cerebrovascular segmentation method based on a multi-scale tubular structure enhanced convolutional neural network. Background Art

[0002] In modern medicine, the diagnosis and treatment of brain diseases face many challenges. Cerebrovascular diseases, such as stroke, cerebral angioma, and cerebrovascular malformation, are the main causes of death and disability worldwide. In addition, other neurological diseases such as epilepsy also pose a serious threat to the health of patients. The treatment of these diseases, especially when invasive procedures (such as electrode implantation) are required, requires doctors to have an in-depth understanding of the cerebrovascular anatomy of the patient in order to plan accurate surgical paths and ensure the safety and effectiveness of the surgery. However, the complexity and individual differences of the cerebrovascular system make it a challenging task to accurately identify and locate blood vessel structures before surgery.

[0003] In this context, accurate cerebrovascular segmentation technology is particularly important. As a core component of computer-aided diagnosis and treatment systems, medical image segmentation technology plays a key role in the diagnosis and treatment of brain diseases. By processing medical images such as time-of-flight magnetic resonance angiography (TOF-MRA), cerebrovascular segmentation technology can construct a three-dimensional model of the patient's cerebrovasculature, helping doctors accurately locate the lesion area and develop personalized treatment plans.

[0004] Before deep learning technology was widely applied in the medical field, traditional unsupervised algorithms such as threshold-based, region-growing-based, Hessian matrix-based, and filtering-enhanced were often used for cerebrovascular segmentation. These methods are relatively mature and interpretable, but the parameter selection is complex and vulnerable to image artifacts, resulting in a cumbersome cerebrovascular segmentation workflow, poor blood vessel segmentation accuracy, and unsatisfactory recognition effects.

[0005] In recent years, with the development of artificial intelligence technology, deep learning methods have been applied to medical image segmentation tasks. Compared with traditional methods, deep learning methods can automatically learn task-related features from training data, making the segmentation results closer to expert annotations while providing a more standardized and consistent segmentation scheme for clinical diagnosis. In the cerebrovascular segmentation task, the segmentation effect of deep learning methods is far better than that of traditional methods. Among them, the convolutional neural network (CNN) has been widely used in the field of image processing, and its calculation method of small window sliding convolution makes it particularly suitable for image segmentation tasks. In the field of medical image segmentation, nnU-Net (based on the U-shaped convolutional neural network) proposed in recent years has further improved the segmentation performance. nnU-Net adopts an encoder-decoder architecture and ensures that the context information of the image can be effectively transmitted to the decoder through multi-layer skip connections, so as to more accurately restore the details of the target structure. As an adaptive, end-to-end and general medical image segmentation framework without manual parameter tuning, nnU-Net has achieved excellent performance in the retinal vessel segmentation task since its proposal, significantly improving the segmentation accuracy and robustness, and is widely used in the field of medical image analysis. However, nnU-Net still has defects: nnU-Net is designed for organs or tissues such as the aorta, heart, and liver, and it is difficult to adapt to cerebrovascular structures with large scale changes and complex branches, resulting in discontinuous blood vessel segmentation, incomplete segmentation of small blood vessels, and errors in the topological structure of the blood vessel segmentation results. Summary of the Invention

[0006] To solve the problems of discontinuous segmentation results, incomplete segmentation of small blood vessels, and errors in the topological structure of the blood vessel segmentation results shown by nnU-Net in the cerebrovascular segmentation task, the present invention proposes a cerebrovascular segmentation method based on a multi-scale tubular structure enhanced encoding convolutional neural network. Through four steps of data collection, model construction, preprocessing and data augmentation, and model training and model testing, high-accuracy segmentation of cerebrovascular vessels is achieved. Validation of effectiveness was carried out on four internationally public time-of-flight magnetic resonance angiography (TOF-MRA) datasets, namely IXI-45, Brains, TubeTK-42, and EDEN. The Dice similarity coefficients reached 88.81%, 82.22%, 79.54%, and 92.45% respectively, the topologically preserved Dice coefficient (clDice) reached 88.71%, 78.32%, 75.06%, and 93.02% respectively, and the average surface distance (ASD) reached 0.1594mm, 0.3118mm, 0.6246mm, and 0.2222mm respectively, all of which are better than other existing optimal methods.

[0007] To solve the above problems in the prior art, the present invention provides a cerebral vascular segmentation method based on a multi-scale tubular structure enhanced coding convolutional neural network. The inventors of the present invention verified the effectiveness of this method using the internationally public time-of-flight magnetic resonance angiography datasets IXI-45, Brains, TubeTK-42, and EDEN. 40 images were selected from the IXI-45 dataset as the training set and 5 images as the test set; 44 images were selected from the Brains dataset as the training set and 12 images as the test set; 33 images were selected from the TubeTK-42 dataset as the training set and 9 images as the test set; 12 images were selected from the EDEN dataset as the test set and 3 images as the training set. The training set data was input into the model for preprocessing, data augmentation, and model training. The trained model was saved, and the test set data was input into the model for preprocessing and inference to obtain the cerebral vascular segmentation result, and the performance of the proposed method was evaluated.

[0008] A cerebral vascular segmentation method based on a multi-scale tubular structure feature enhanced coding convolutional neural network according to an embodiment of the present invention includes the following steps:

[0009] Step 1: Obtain the internationally public time-of-flight magnetic resonance angiography dataset and divide it into a training set and a test set;

[0010] Step 2: Use the Pytorch deep learning framework to construct a multi-scale tubular structure enhanced coding convolutional neural network;

[0011] Step 3: Input the training set data established in Step 1 into the model of the multi-scale tubular structure enhanced coding convolutional neural network constructed in Step 2 for preprocessing and data augmentation, and perform training on the cerebral vascular segmentation model. After training is completed, save the model parameters;

[0012] Step 4: Load the model parameters saved in Step 3 to obtain the trained cerebral vascular segmentation model. Input the test set selected in Step 1 into the model of the multi-scale tubular structure enhanced coding convolutional neural network for preprocessing and inference to obtain the cerebral vascular segmentation result.

[0013] In step 2, the present invention uses nnU-Net as the baseline model, improves the original network encoder into a multi-scale tubular structure enhanced encoder, and designs a feature enhancement module based on the Frangi vascular enhancement filter and the selective attention mechanism. The multi-scale tubular structure enhanced encoder can enhance the foreground pixel intensities belonging to different diameters from coarse to fine. The feature enhancement module based on the Frangi vascular enhancement filter and the selective attention mechanism can be trained together with the network, automatically optimize the Frangi filter parameters, enhance the vascular signal intensity, and effectively suppress noise; at the same time, it can adaptively allocate weights between the structural features and the spatial features under the condition of preserving the spatial information of the original feature map.

[0014] In step 3, the data preprocessing includes performing Z-score normalization on the original time-of-flight magnetic resonance cerebral angiography, generating data blocks, random rotation and scaling, random rotation and scaling, random image attribute transformation, and random mirror flipping.

[0015] The advantages of the cerebral vascular segmentation method based on the multi-scale tubular structure enhanced coding convolutional neural network according to the present invention include:

[0016] 1. The multi-scale tubular structure enhanced encoder designed by the present invention can enhance the vascular pixel intensities of different diameters from coarse to fine, achieve a comprehensive capture of cerebral blood vessels of different scales, and improve the accuracy of vascular segmentation;

[0017] 2. The feature enhancement module designed by the present invention can be trained together with the network, automatically optimize the Frangi filter parameters, enhance the vascular signal intensity while effectively suppressing noise, and adaptively adjust the weights of the structural features and the spatial features on the basis of preserving the spatial information of the original feature map, improving the robustness and accuracy of the segmentation result. Brief Description of the Drawings

[0018] Figure 1 It is a flowchart of a cerebral vascular segmentation method based on a multi-scale tubular structure enhanced coding convolutional neural network according to an embodiment of the present invention.

[0019] Figure 2 It is an overall structure diagram of a multi-scale tubular structure enhanced coding convolutional neural network according to an embodiment of the present invention.

[0020] Figure 3 It is a schematic diagram of the feature enhancement module.

[0021] Figure 4 It shows a comparison between partial cerebral vascular segmentation results of a cerebral vascular segmentation method based on a multi-scale tubular structure enhanced coding convolutional neural network according to an embodiment of the present invention on a test data set and the gold standard.

[0022] Figure 5Shows the comparison of partial cerebral vascular segmentation results of a cerebral vascular segmentation method based on a multi-scale tubular structure enhanced encoding convolutional neural network according to an embodiment of the present invention with the local details of the segmentation results of cerebral vascular segmentation models in recent years on a test dataset. Detailed implementation manners

[0023] In a cerebral vascular segmentation method based on a multi-scale tubular structure enhanced encoding convolutional neural network according to an embodiment of the present invention, as Figure 3 shown, a feature enhancement module based on the Frangi vessel enhancement filter and the selective attention mechanism is designed. At the same time, on the basis of the nnU-Net architecture, the encoder part is improved into a multi-scale tubular structure enhanced encoder, and finally the cerebral vascular segmentation result is obtained.

[0024] The following combines the accompanying drawings and specific embodiments to detail a cerebral vascular segmentation method based on a multi-scale tubular structure enhanced encoding convolutional neural network proposed by the present invention. The overall process is shown in Figure 1 .

[0025] A cerebral vascular segmentation method based on a multi-scale tubular structure enhanced encoding convolutional neural network according to an embodiment of the present invention includes:

[0026] Step S1: Obtain the publicly available time-of-flight magnetic resonance angiography datasets IXI-45, Brains, TubeTK-42, and EDEN;

[0027] Step S2: Use the function library of the Pytorch deep learning framework in Python to establish a multi-scale tubular structure enhanced encoding convolutional neural network based on the nnU-Net architecture. The overall structure of the network is shown in Figure 2 , which includes a multi-scale tubular structure enhanced encoder and a decoder. Among them, the feature enhancement module based on the Frangi vessel enhancement filter and the selective attention mechanism is inserted into the first three stages of the multi-scale tubular structure enhanced encoder, specifically including:

[0028] A) Construct a multi-scale tubular structure enhanced encoder, including:

[0029] A1) Construct the first stage, including:

[0030] Input the preprocessed time-of-flight magnetic resonance cerebral vascular imaging data block through the input layer. The size of the data block is D×H×W;

[0031] Set the first feature enhancement module, the first three-dimensional convolution module, and the second three-dimensional convolution module, where:

[0032] As Figure 3As shown, the first feature enhancement module consists of a local spatial feature extraction module, a local structural feature extraction module, and a feature fusion module, where:

[0033] The local spatial feature extraction module includes a first convolutional layer with a convolutional kernel size of 3×3×3, a convolutional stride of 1, and 1 convolutional kernel, a first instance normalization layer using instance normalization, and a first Leaky ReLU activation layer using the Leaky ReLU activation function, which is used to output spatial features;

[0034] The local structural feature extraction module is used to first perform Gaussian smoothing on the input feature map, and then perform convolution operations along the x, y, and z axes respectively using one-dimensional Gaussian second derivative convolutional kernels of size 9 to obtain the Hessian matrix H(x, y, z) at each voxel position (x, y, z). Subsequently, the eigenvalues of this matrix are solved and sorted according to the absolute value size, satisfying |λ1| ≤ |λ2| ≤ |λ3|. The plate-like structure response value P, spherical structure response value B, and foreground response value F at this position are calculated using the eigenvalues. The calculation formulas are as follows:

[0035]

[0036]

[0037] The calculation formula for the vesselness v of the voxel (x, y, z) is:

[0038] v = P × B × F (4)

[0039] Among them, exp(·) represents the exponential function, and α, β, and c are all trainable parameters. The parameter α controls the sensitivity of the vesselness v to the plate-like structure response value P, the parameter β controls the sensitivity of the vesselness v to the spherical structure response value B, and the parameter c controls the sensitivity of the vesselness v to the foreground response value. The initial values of α, β, and c are 0.5, 0.5, and 1 respectively. The vesselness v(x, y, z) at each position (x, y, z) constitutes a vessel enhancement feature map. Subsequently, the vessel enhancement feature map is normalized by Z-score. The second convolutional layer with a convolutional kernel size of 1 and a convolutional stride of 1 is used to perform adaptive weight assignment on the vessel feature enhancement map. After passing through the second Leaky ReLU activation layer to obtain a non-linear mapping, a residual connection is made with the input of the local structural feature extraction module, and then the output structural features are obtained after passing through the second instance normalization layer;

[0040] The feature fusion module first adds the spatial feature and the structural feature to obtain the feature map U, and then generates the spatial feature weight W s% and the structural feature weight W st , and the calculation formulas are:

[0041]

[0042] Among them, exp(·) represents the exponential function, and f c (·) is a fully connected layer, GAP(·) represents global average pooling, and f sp (·) and f st (·) respectively represent the linear mappings for generating spatial feature weights and structural feature weights.

[0043] Finally, the spatial features and structural features are weighted and fused by these two weights to achieve the adaptive integration of features, thereby obtaining the final output feature map.

[0044] The first three-dimensional convolution module includes a third convolution layer with 32 convolution kernels and a convolution stride of 1×1×1, a third instance normalization layer, and a third Leaky ReLU activation layer.

[0045] The second three-dimensional convolution module includes a fourth convolution layer with a convolution kernel size of 3×3×3, 32 convolution kernels, and a convolution stride of 1×1×1, a fourth instance normalization layer, and a fourth Leaky ReLU activation layer.

[0046] The output of stage one is x1.

[0047] A2) Construct stage two, including:

[0048] Input the output x1 of stage one.

[0049] Set the second feature enhancement module, the third three-dimensional convolution module, and the fourth three-dimensional convolution module, where:

[0050] The second feature enhancement module is the same as the first feature enhancement module in stage one.

[0051] The third three-dimensional convolution module includes a fifth convolution layer with 64 convolution kernels, a fifth instance normalization layer, and a fifth Leaky ReLU activation layer.

[0052] The fourth convolution module includes a sixth convolution layer with a convolution kernel size of 3×3×3, 64 convolution kernels, and a convolution stride of 1×1×1, a sixth instance normalization layer, and a sixth Leaky ReLU activation layer.

[0053] The output of stage two is x2.

[0054] A3) Construct stage three, including:

[0055] Input the output x2 of stage two.

[0056] Set the third feature enhancement module, the fifth three-dimensional convolution module, and the sixth three-dimensional convolution module, where:

[0057] The third feature enhancement module is the same as the feature enhancement module in Phase 1;

[0058] The fifth 3D convolutional module includes a seventh convolutional layer with 128 convolutional kernels, a seventh instance normalization layer, and a seventh Leaky ReLU activation layer;

[0059] The sixth 3D convolutional module includes an eighth convolutional layer with a convolutional kernel size of 3×3×3, 128 convolutional kernels, and a convolutional stride of 1×1×1, an eighth instance normalization layer, and an eighth Leaky ReLU activation layer;

[0060] The output of Phase 3 is x3;

[0061] A4) Construct Phase 4, including:

[0062] Input the output x3 of Phase 3;

[0063] Set the seventh 3D convolutional module and the eighth 3D convolutional module, where:

[0064] The seventh 3D convolutional module includes a ninth convolutional layer with 256 convolutional kernels, a ninth instance normalization layer, and a ninth Leaky ReLU activation layer;

[0065] The eighth 3D convolutional module includes a tenth convolutional layer with a convolutional kernel size of 3×3×3, 256 convolutional kernels, and a convolutional stride of 1×1×1, a tenth instance normalization layer, and a tenth Leaky ReLU activation layer;

[0066] The output of Phase 4 is x ( ;

[0067] A5) Construct Phase 5, including:

[0068] Input the output x of Phase 4 through the fourth hidden layer ( ;

[0069] Set the ninth and tenth 3D convolutional modules, where:

[0070] The ninth 3D convolutional module includes an eleventh convolutional layer with 320 convolutional kernels, an eleventh instance normalization layer, and an eleventh Leaky ReLU activation layer;

[0071] The tenth 3D convolutional module includes a twelfth convolutional layer with a convolutional kernel size of 3×3×3, 320 convolutional kernels, and a convolutional stride of 1×1×1, a twelfth instance normalization layer, and a twelfth Leaky ReLU activation layer;

[0072] The output of Phase 5 is x5;

[0073] A6) Construct Phase 6, including:

[0074] The output x5 of the input stage five;

[0075] Set the eleventh and twelfth 3D convolutional modules, where:

[0076] The eleventh 3D convolutional module includes the thirteenth convolutional layer with 320 convolutional kernels, the thirteenth instance normalization layer, and the thirteenth Leaky ReLU activation layer;

[0077] The twelfth 3D convolutional module includes the fourteenth convolutional layer with a convolutional kernel size of 3×3×3, 320 convolutional kernels, and a convolutional stride of 1×1×1, the fourteenth instance normalization layer, and the fourteenth Leaky ReLU activation layer;

[0078] The output of stage six is x6;

[0079] B) Construct the decoder: including constructing five stages from stage seven to stage eleven from bottom to top, including:

[0080] B1) Construct stage seven, including:

[0081] Use the first transposed convolutional layer with a convolutional kernel size of 2×2×2 and 320 convolutional kernels to perform the first upsampling on the output x6 of stage six;

[0082] Subsequently, perform the first channel-level concatenation on the result of the first upsampling and the output x5 of stage five;

[0083] Let the result of the first concatenation pass through two identical thirteenth 3D convolutional modules and fourteenth 3D convolutional modules; where the thirteenth 3D convolutional module includes the fifteenth convolutional layer with a convolutional kernel size of 3×3×3, 320 convolutional kernels, and a convolutional stride of 1×1×1, the fifteenth instance normalization layer, and the fifteenth Leaky ReLU activation layer;

[0084] The output of stage seven is d1;

[0085] B2) Construct stage eight, including:

[0086] Use the second transposed convolutional layer with a convolutional kernel size of 2×2×2 and 256 convolutional kernels to perform the second upsampling on the output d1 of stage seven;

[0087] Subsequently, perform the second channel-level concatenation on the result of the second upsampling and the output x ( ;

[0088] Let the result of the second concatenation pass through two identical fifteenth 3D convolutional modules and sixteenth 3D convolutional modules; the fifteenth 3D convolutional module includes the sixteenth convolutional layer with a convolutional kernel size of 3×3×3, 256 convolutional kernels, and a convolutional stride of 1×1×1, the sixteenth instance normalization layer, and the sixteenth Leaky ReLU activation layer;

[0089] The output of stage eight is d2;

[0090] B3) Construct stage nine, including:

[0091] Perform third upsampling on the output d2 using a third transposed convolution layer with a convolution kernel size of 2×2×2 and 128 convolution kernels;

[0092] Subsequently, perform third concatenation at the channel level on the result of the third upsampling and the output x3;

[0093] Let the result of the third concatenation pass through two identical seventeenth and eighteenth three-dimensional convolution modules; The seventeenth three-dimensional convolution module includes a seventeenth convolution layer with a convolution kernel size of 3×3×3, 128 convolution kernels, and a convolution stride of 1×1×1, a seventeenth instance normalization layer, and a seventeenth Leaky ReLU activation layer;

[0094] The output of stage nine is d3;

[0095] B4) Construct stage ten, including:

[0096] Perform fourth upsampling on the output d3 using a fourth transposed convolution layer with a convolution kernel size of 2×2×2 and 64 convolution kernels;

[0097] Subsequently, perform fourth concatenation at the channel level on the result of the fourth upsampling and the output x2;

[0098] Let the result of the fourth concatenation pass through two identical nineteenth and twenty-third three-dimensional convolution modules; The nineteenth three-dimensional convolution module includes an eighteenth convolution layer with a convolution kernel size of 3×3×3, 64 convolution kernels, and a convolution stride of 1×1×1, an eighteenth instance normalization layer, and an eighteenth Leaky ReLU activation layer;

[0099] The output of stage ten is d ( ;

[0100] B5) Construct stage eleven, including:

[0101] Through the output layer, perform fifth upsampling on the output d ( using a fifth transposed convolution layer with a convolution kernel size of 2×2×2 and 32 convolution kernels;

[0102] Subsequently, perform fifth concatenation at the channel level on the result of the fifth upsampling and the output x1;

[0103] Let the result of the fifth concatenation pass through the twenty-first, twenty-second, and twenty-third three-dimensional convolution modules; The twenty-first and twenty-second three-dimensional convolution modules are the same;

[0104] The twenty - first 3D convolution module includes the nineteenth convolutional layer with a convolutional kernel size of 3×3×3, 32 convolutional kernels, and a convolutional stride of 1×1×1, the nineteenth instance normalization layer, and the nineteenth Leaky ReLU activation layer;

[0105] The twenty - third convolution module includes a convolutional layer with a convolutional kernel size of 1×1×1 and a Softmax function layer, and outputs a binary image of the cerebral vascular segmentation result;

[0106] For the IXI - 45, Brains, TubeTK - 42, and EDEN datasets, the convolutional kernel sizes and convolutional strides used in the first 3D convolution module of each stage of the multi - scale tubular structure enhancement encoder and the transposed convolutions of each stage of the decoder are shown in the following table:

[0107]

[0108]

[0109] Step S3.1: Input the training set selected in Step S1 into the cerebrovascular segmentation model. The model first performs Z-score normalization on the original data, and then uses the preprocessing rules of the nnU-Net framework to generate data chunks for the original data: for the IXI-45 dataset, the data chunk size is 28×320×224; for the Brains dataset, the data chunk size is 96×160×160; for the TubeTK-42 dataset, the data chunk size is 64×192×192; for the EDEN dataset, the data chunk size is 64×192×192; According to the data augmentation rules of the nnU-Net framework, the processing of data chunks is performed in the following order: First, rotation and scaling are applied simultaneously to improve the calculation speed. Rotation and scaling are applied with a probability of 0.2 respectively. Therefore, the probability of only scaling is 0.16, the probability of only rotation is 0.16, and the probability of triggering both at the same time is 0.08; for isotropic 3D data chunks, the rotation angle (unit: degree) is sampled from U(-30, 30) on the x, y, and z axes respectively; for non-isotropic 3D data chunks, the rotation angle is sampled from U(-180, 180); Scaling is achieved by multiplying the scaling factor in the voxel grid, and the scaling factor is sampled from U(0.7, 1.4); Secondly, Gaussian noise is independently added to each voxel of each sample with a probability of 0.15. The variance of the noise is sampled from U(0, 0.1), and after all samples are intensity-normalized, their voxel intensities are close to zero mean and unit variance; In addition, Gaussian blur is applied to the samples with a probability of 0.2, and the width of the Gaussian kernel (unit: voxel) is independently sampled from U(0.5, 1.5); Brightness enhancement is applied with a probability of 0.15 by multiplying the voxel intensity by an x value sampled from U(0.7, 1.3), and contrast enhancement is also applied with a probability of 0.15, and the voxel intensity is multiplied by an x value sampled from U(0.65, 1.5); Low-resolution simulation is applied to the samples with a probability of 0.25, downsampled by a factor of U(1, 2) using nearest-neighbor interpolation, and then resampled to the original size using cubic interpolation. This enhancement is only applied within the 2D plane and does not change the axes outside the plane; Gamma transformation is applied with a probability of 0.15. First, the voxel intensity is normalized to between [0, 1], and then each voxel performs a non-linear intensity transformation. The Gamma correction coefficient is sampled from U(0.7, 1.5), and the voxel intensity after transformation is scaled back to the original range; In addition, with a probability of 0.15, the voxel intensity is inverted before the transformation; Finally, the data chunk is mirrored and flipped along all axes with a probability of 0.5; where U(a, b) represents randomly and uniformly sampling a value from the interval [a, b];

[0110] Step S3.2: Adjust the training parameters by observing the convergence of the loss function during training. Set the training batch size to 2, adopt the Poly scheduling strategy for the learning rate, with the initial value set to 0.01 for the initial network learning rate. Use Stochastic Gradient Descent (SGD) with Nesterov momentum (μ = 0.99) as the optimizer, and use the average of the cross-entropy loss function and the Dice loss function as the loss function. Save the model parameters after 200 iterations of training;

[0111] Step S4: Load the model parameters saved in Step S3 to obtain the trained cerebrovascular segmentation model. Input the test sets of the four datasets IXI-45, Brains, TubeTK-42, and EDEN selected in Step S1 into the model. The model generates data blocks for the original data and completes preprocessing according to the method in Step S3.1. Then the model performs inference to obtain the cerebrovascular segmentation result.

[0112] To illustrate the advantages of the method proposed in the present invention, compare the performance metrics of the latest cerebrovascular segmentation models in recent years on the same test sets.

[0113] The segmentation results of some images of the cerebrovascular model based on the multi-scale tubular structure enhanced encoder convolutional neural network on the test sets of the datasets IXI-45, Brains, TubeTK-42, and EDEN are visually compared with the gold standard in Figure 4 , and the performance metrics of all images on the four datasets are shown in Tables 1, 2, 3, and 4 respectively. The bold data indicates the optimal results in that column.

[0114] Table 1 Performance Metrics of the Multi-scale Tubular Structure Enhanced Encoding Convolutional Neural Network on the IXI-45 Test Set

[0115]

[0116] Table 2 Performance Metrics of the Multi-scale Tubular Structure Enhanced Encoding Convolutional Neural Network on the Brains Test Set

[0117]

[0118] Table 3 Performance Metrics of the Multi-scale Tubular Structure Enhanced Encoding Convolutional Neural Network on the TubeTK-42 Test Set

[0119]

[0120]

[0121] Table 4 Performance Metrics of the Multi-scale Tubular Structure Enhanced Encoding Convolutional Neural Network on the EDEN Test Set

[0122]

[0123] As can be seen from Table 1, Table 2, Table 3 and Table 4, for the cerebral vascular segmentation method based on the multi-scale tubular structure enhanced convolutional neural network proposed by the present invention, the average Dice similarity coefficient (Dice) on the public dataset IXI-45 is 88.81%, the topologically-preserved Dice coefficient (clDice) is 88.71%, and the average surface distance (ASD) is 0.1594 mm; on the public dataset Brains, the average Dice similarity coefficient is 82.22%, the topologically-preserved Dice coefficient is 78.32%, and the average surface distance is 0.3119 mm; on the public dataset TubeTK-42, the average Dice similarity coefficient is 79.54%, the topologically-preserved Dice coefficient is 75.06%, and the average surface distance is 0.6246 mm; on the public dataset EDEN, the average Dice similarity coefficient is 92.45%, the topologically-preserved Dice coefficient is 93.02%, and the average surface distance is 0.2222 mm. For multiple image segmentation results on the four datasets, the Dice similarity coefficient reaches over 80%, the topologically-preserved Dice coefficient reaches over 80%, and the average surface distance is below 0.5 mm. This shows that the cerebral segmentation result of the method of the present invention has high accuracy, can accurately segment fine blood vessels, and maintain the topological structure of the blood vessels.

[0124] For the cerebral vascular segmentation model based on the multi-scale tubular structure enhanced convolutional neural network designed by the present invention, the performance indicators on the test set and the visual comparison of some segmentation results with the cerebral vascular segmentation models in recent years are shown in Figure 5 , and the performance comparison of the segmentation results is shown in Table 5, Table 6, Table 7 and Table 8, where the bold data represents the optimal result in the column.

[0125] Table 5 Average performance indicators of cerebral vascular segmentation models on the test set of the dataset IXI-45 in recent years

[0126]

[0127] Table 6 Average performance indicators of cerebral vascular segmentation models on the test set of the dataset Brains in recent years

[0128]

[0129]

[0130] Table 7 Average performance indicators of cerebral vascular segmentation models on the test set of the dataset TubeTK-42 in recent years

[0131]

[0132] Table 8 Average performance metrics of cerebrovascular segmentation models on the EDEN test set of the dataset in recent years

[0133]

[0134] As can be seen from Table 5, Table 6, Table 7 and Table 8, the cerebrovascular segmentation method based on the multi-scale tubular structure enhanced encoding convolutional neural network proposed by the present invention is superior to the existing advanced methods in terms of the three metrics of Dice similarity coefficient, topology-preserving Dice coefficient and average surface distance. Moreover, the segmentation results obtained by this model have better visual perception and have important clinical application value in the fields of clinical disease assisted diagnosis and so on.

[0135] The above has described in detail the multi-scale tubular structure enhanced encoding convolutional neural network cerebrovascular segmentation method provided by the present invention, but obviously the scope of the present invention is not limited thereto. Without departing from the scope of protection defined by the appended claims, various changes to the above embodiments are within the scope of the present invention.

Claims

1. A construction method of a multi-scale tubular structure enhanced coding convolutional neural network based on the nnU-Net framework for a cerebrovascular segmentation method based on a multi-scale tubular structure enhanced coding convolutional neural network, characterized in that Including: A) Constructing a multi-scale tubular structure enhanced encoder, including: A1) The construction of stage one, including: Inputting the preprocessed time-of-flight magnetic resonance cerebral angiography data block through the input layer, and the size of the data block is D×H×W; Setting a first feature enhancement module, a first three-dimensional convolution module, and a second three-dimensional convolution module, where: The first feature enhancement module consists of a local spatial feature extraction module, a local structural feature extraction module, and a feature fusion module, where: The local spatial feature extraction module includes a first convolutional layer with a convolution kernel size of 3×3×3, a convolution stride of 1, and a convolution kernel number of 1, a first instance normalization layer using instance normalization, and a first LeakyReLU activation layer using the Leaky ReLU activation function, for outputting spatial features; The local structural feature extraction module first performs Gaussian smoothing on the input feature map, and then performs convolution operations along the x, y, and z axes respectively using one-dimensional Gaussian second derivative convolution kernels of size 9 to obtain the Hessian matrix H(x,y,z) at each voxel position (x,y,z). Subsequently, the eigenvalues of this matrix are solved and sorted according to the absolute value size, satisfying |λ1|≤|λ2|≤|λ3|. The plate-like structure response value P, spherical structure response value B, and foreground response value F at this position are calculated using the eigenvalues. The calculation formulas are: The calculation formula for the vesselness v of the voxel (x,y,z) is: v = P×B×F (4) where exp(·) represents the exponential function, and α, β, and c are all trainable parameters. The parameter α controls the sensitivity of the vesselness v to the plate-like structure response value P, the parameter β controls the sensitivity of the vesselness v to the spherical structure response value B, and the parameter c controls the sensitivity of the vesselness v to the foreground response value F; the initial values of α, β, and c are 0.5, 0.5, and 1 respectively. The vesselness v(x,y,z) at each position (x,y,z) constitutes a vessel-enhanced feature map. Subsequently, the vessel-enhanced feature map is normalized by Z-score. A second convolutional layer with a convolution kernel size of 1 and a convolution stride of 1 is used to perform adaptive weight allocation on the vessel feature enhancement map. After passing through the second LeakyReLU activation layer, a non-linear mapping is obtained, and then a residual connection is made with the input of the local structural feature extraction module. After passing through the second instance normalization layer, the structural features are output; The feature fusion module first adds the spatial feature and the structural feature to obtain the feature map U, and then generates the spatial feature weight W s% and the structural feature weight W st , and the calculation formula is as follows: Among them, exp(·) represents the exponential function, and f c (·) is a fully connected layer, GAP(·) represents global average pooling, and f sp (·) and f st (·) respectively represent the linear mappings for generating spatial feature weights and structural feature weights. Through the spatial feature weight W s% and the structural feature weight W st weighted fusion of spatial features and structural features is performed to achieve adaptive integration of features, thereby obtaining the final output feature map; The first three-dimensional convolution module includes a third convolutional layer with a convolution kernel number of 32 and a convolution stride of 1×1×1, a third instance normalization layer, and a third Leaky ReLU activation layer; The second three-dimensional convolution module includes a fourth convolutional layer with a convolution kernel size of 3×3×3, a convolution kernel number of 32, and a convolution stride of 1×1×1, a fourth instance normalization layer, and a fourth Leaky ReLU activation layer; The output of stage one is x1; A2) The construction of stage two, including: Inputting the output x1 of stage one; Setting a second feature enhancement module, a third three-dimensional convolution module, and a fourth three-dimensional convolution module, where: The second feature enhancement module is the same as the first feature enhancement module in stage one; The third 3D convolution module includes a fifth convolutional layer with 64 convolutional kernels, a fifth instance normalization layer, and a fifth Leaky ReLU activation layer; The fourth convolution module includes a sixth convolutional layer with a convolutional kernel size of 3×3×3, 64 convolutional kernels, and a convolutional stride of 1×1×1, a sixth instance normalization layer, and a sixth Leaky ReLU activation layer; The output of stage two is x2; A3) Construct stage three, including: Input the output x2 of stage two; Set a third feature enhancement module, a fifth 3D convolution module, and a sixth 3D convolution module, where: The third feature enhancement module is the same as the feature enhancement module in stage one; The fifth 3D convolution module includes a seventh convolutional layer with 128 convolutional kernels, a seventh instance normalization layer, and a seventh Leaky ReLU activation layer; The sixth 3D convolution module includes an eighth convolutional layer with a convolutional kernel size of 3×3×3, 128 convolutional kernels, and a convolutional stride of 1×1×1, an eighth instance normalization layer, and an eighth Leaky ReLU activation layer; The output of stage three is x3; A4) Construct stage four, including: Input the output x3 of stage three; Set a seventh 3D convolution module and an eighth 3D convolution module, where: The seventh 3D convolution module includes a ninth convolutional layer with 256 convolutional kernels, a ninth instance normalization layer, and a ninth Leaky ReLU activation layer; The eighth 3D convolution module includes a tenth convolutional layer with a convolutional kernel size of 3×3×3, 256 convolutional kernels, and a convolutional stride of 1×1×1, a tenth instance normalization layer, and a tenth Leaky ReLU activation layer; The output of stage four is x ( ; A5) Construct stage five, including: Through the fourth hidden layer, the output x of input stage four ( ; Set a ninth and a tenth 3D convolution module, where: The ninth 3D convolution module includes an eleventh convolutional layer with 320 convolutional kernels, an eleventh instance normalization layer, and an eleventh Leaky ReLU activation layer; The tenth 3D convolution module includes a twelfth convolutional layer with a convolutional kernel size of 3×3×3, 320 convolutional kernels, and a convolutional stride of 1×1×1, a twelfth instance normalization layer, and a twelfth Leaky ReLU activation layer; The output of stage five is x5; A6) Construct stage six, including: Input the output x5 of stage five; Set an eleventh and a twelfth 3D convolution module, where: The eleventh 3D convolution module includes a thirteenth convolutional layer with 320 convolutional kernels, a thirteenth instance normalization layer, and a thirteenth Leaky ReLU activation layer; The twelfth 3D convolution module includes a fourteenth convolutional layer with a convolutional kernel size of 3×3×3, 320 convolutional kernels, and a convolutional stride of 1×1×1, a fourteenth instance normalization layer, and a fourteenth Leaky ReLU activation layer; The output of stage six is x6; B) Construct the decoder: including constructing five stages from stage seven to stage eleven from bottom to top, including: B1) Construct stage seven, including: Use a first transposed convolutional layer with a convolutional kernel size of 2×2×2 and 320 convolutional kernels to perform a first upsampling on the output x6 of stage six; Subsequently, perform a first concatenation at the channel level on the result of the first upsampling and the output x of stage five ) ; Subject the result of the first concatenation to two identical thirteenth 3D convolutional modules and a fourteenth 3D convolutional module; the thirteenth 3D convolutional module includes a fifteenth convolutional layer with a convolution kernel size of 3×3×3, 320 convolution kernels, and a convolution stride of 1×1×1, a fifteenth instance normalization layer, and a fifteenth Leaky ReLU activation layer; The output of stage seven is d1; B2) Construct stage eight, including: Perform a second upsampling on the output d1 of stage seven using a second transposed convolutional layer with a convolution kernel size of 2×2×2 and 256 convolution kernels; Subsequently, perform a second concatenation at the channel level on the second upsampling result and the output x ( ; Subject the result of the second concatenation to two identical fifteenth 3D convolutional modules and a sixteenth 3D convolutional module; the fifteenth 3D convolutional module includes a sixteenth convolutional layer with a convolution kernel size of 3×3×3, 256 convolution kernels, and a convolution stride of 1×1×1, a sixteenth instance normalization layer, and a sixteenth Leaky ReLU activation layer; The output of stage eight is d2; B3) Construct stage nine, including: Perform a third upsampling on the output d2 using a third transposed convolutional layer with a convolution kernel size of 2×2×2 and 128 convolution kernels; Subsequently, perform a third concatenation at the channel level on the result of the third upsampling and the output x3; Subject the result of the third concatenation to two identical seventeenth and eighteenth 3D convolutional modules; the seventeenth 3D convolutional module includes a seventeenth convolutional layer with a convolution kernel size of 3×3×3, 128 convolution kernels, and a convolution stride of 1×1×1, a seventeenth instance normalization layer, and a seventeenth Leaky ReLU activation layer; The output of stage nine is d3; B4) Construct stage ten, including: Perform a fourth upsampling on the output d3 using a fourth transposed convolutional layer with a convolution kernel size of 2×2×2 and 64 convolution kernels; Subsequently, perform a fourth concatenation at the channel level on the result of the fourth upsampling and the output x2; Subject the result of the fourth concatenation to two identical nineteenth and twenty-third 3D convolutional modules; the nineteenth 3D convolutional module includes an eighteenth convolutional layer with a convolution kernel size of 3×3×3, 64 convolution kernels, and a convolution stride of 1×1×1, an eighteenth instance normalization layer, and an eighteenth Leaky ReLU activation layer; The output of stage ten is d ( ; B5) Construct stage eleven, including: Perform fifth upsampling on the output d using a fifth transposed convolutional layer with a convolutional kernel size of 2×2×2 and 32 convolutional kernels. ( Perform fifth upsampling; Subsequently, perform a fifth concatenation at the channel level on the result of the fifth upsampling and the output x1; Subject the result of the fifth concatenation to twenty-first, twenty-second, and twenty-third 3D convolutional modules; the twenty-first and twenty-second 3D convolutional modules are identical; The twenty-first 3D convolutional module includes a nineteenth convolutional layer with a convolution kernel size of 3×3×3, 32 convolution kernels, and a convolution stride of 1×1×1, a nineteenth instance normalization layer, and a nineteenth Leaky ReLU activation layer; The twenty-third convolutional module includes a convolutional layer with a convolution kernel size of 1×1×1 and a Softmax function layer, and outputs a binary image of the cerebrovascular segmentation result.

2. The construction method of the multi-scale tubular structure enhanced convolutional neural network according to claim 1, wherein Further include: C) Training step, input the training set data into the multi-scale tubular structure enhanced encoding convolutional neural network for data preprocessing and data augmentation and complete the training to generate a trained multi-scale tubular structure enhanced encoding convolutional neural network.

3. The construction method of the multi-scale tubular structure enhanced convolutional neural network according to claim 2, characterized in that: Select images for training the model from the publicly available time-of-flight magnetic resonance angiography datasets IXI-45, Brains, TubeTK-42, and EDEN obtained in advance, and establish the training set data.

4. The construction method of the multi-scale tubular structure enhanced convolutional neural network according to claim 3, characterized in that In the step C: Data preprocessing includes performing Z-score normalization on the original time-of-flight magnetic resonance cerebral angiography, generating data blocks, random rotation and scaling, random rotation and scaling, random image attribute transformation, and random mirror flipping; Train the multi-scale tubular structure enhanced convolutional neural network: adjust the training parameters by observing the convergence of the loss function during the training process, set the training batch size to 2, adopt the Poly scheduling strategy for the learning rate, set the initial value of the network learning rate to 0.01, use the stochastic gradient descent with Nesterov momentum as the optimizer, and use the average of the cross-entropy loss function and the Dice loss function as the loss function. Save the model parameters after 200 iterations of training.

5. The construction method of the multi-scale tubular structure enhanced convolutional neural network according to claim 2, characterized in that In step C: Generate data blocks for the original data using the preprocessing rules of the nnU-Net framework; For the IXI-45 dataset, the data block size is 28×320×224; For the Brains dataset, the data block size is 96×160×160; For the TubeTK-42 dataset, the data block size is 64×192×192; For the EDEN dataset, the data block size is 64×192×192, The processing of the data blocks is performed in the following order: First, rotation and scaling are applied simultaneously, and rotation and scaling are applied with a probability of 0.2 respectively; for isotropic 3D data blocks, the rotation angles (unit: degrees) are sampled from U(-30, 30) on the x, y, and z axes respectively; for non-isotropic 3D data blocks, the rotation angles are sampled from U(-180, 180); Scaling is achieved by multiplying the scaling factor in the voxel grid, and the scaling factor is sampled from U(0.7, 1.4); Secondly, Gaussian noise is independently added to each voxel of each sample with a probability of 0.15, and the variance of the noise is sampled from U(0, 0.1). After all samples are intensity-normalized, their voxel intensities are close to zero mean and unit variance; Apply Gaussian blur to the sample with a probability of 0.

2. The width of the Gaussian kernel (unit: voxel) is independently sampled from U(0.5, 1.5); brightness enhancement is applied with a probability of 0.15 by multiplying the voxel intensity by an x value sampled from U(0.7, 1.3), while contrast enhancement is also applied with a probability of 0.15, and the voxel intensity is multiplied by an x value sampled from U(0.65, 1.5); low-resolution simulation is applied to the sample with a probability of 0.25, downsampled by a factor of U(1, 2) using nearest-neighbor interpolation, and then resampled back to the original size using cubic interpolation. This enhancement is only applied within the 2D plane and does not change the out-of-plane axes; Gamma transformation is applied with a probability of 0.

15. First, the voxel intensity is normalized between [0, 1], and then a nonlinear intensity transformation is performed on each voxel. The Gamma correction coefficient is sampled from U(0.7, 1.5), and the transformed voxel intensity is scaled back to the original range; With a probability of 0.15, invert the voxel intensity before transformation; finally, mirror-flip the data block along all axes with a probability of 0.5; where U(a, b) represents randomly and uniformly sampling a value from the interval [a, b].

6. The method for constructing a multi-scale tubular structure enhanced encoding convolutional neural network according to claim 1, characterized in that: Use the functions of the Pytorch deep learning framework in Python and the nnU-Net framework to construct a multi-scale tubular structure enhanced encoding convolutional neural network.

7. The method for constructing a multi-scale tubular structure enhanced encoding convolutional neural network according to claim 3, characterized in that: In step S1, for the IXI-45, Brains, TubeTK-42, and EDEN datasets, the convolution kernel sizes and convolution strides used in the third, fifth, seventh, ninth, eleventh, and thirteenth convolutional layers of the multi-scale tubular structure enhanced encoder and the transposed convolutions in each stage of the decoder are as follows:

8. A cerebrovascular segmentation method based on a multi-scale tubular structure enhanced encoding convolutional neural network, characterized in that Including: S1) Execute the method for constructing a multi-scale tubular structure enhanced encoding convolutional neural network according to any one of claims 1 - 7, S2) Input the test set cerebrovascular imaging data into the finally trained multi-scale tubular structure enhanced encoding convolutional neural network for preprocessing and inference, obtain the cerebrovascular segmentation result, and calculate the metrics; Wherein: Select the images for testing from the publicly available time-of-flight magnetic resonance angiography datasets IXI-45, Brains, TubeTK-42, and EDEN obtained in advance to establish the test set data.

9. The cerebrovascular segmentation method according to claim 8, characterized in that: The step S2 includes: using the Dice similarity coefficient, the topology-preserving Dice coefficient, and the average surface distance to evaluate the cerebrovascular segmentation effect of the model.

10. A computer-readable storage medium storing a computer program, and the computer program enables a processor to execute the method according to any one of claims 1 - 9.