Cerebrovascular segmentation method and device based on dynamic snakelike convolution and Transform

By adopting the U-shaped neural network model TRDSCN-UNet based on dynamic serpentine convolution and Transformer in the cerebrovascular segmentation task, the problem of oversegment or missegment in the cerebrovascular segmentation task in the prior art is solved, and higher segmentation accuracy and continuity are achieved.

CN120032125APending Publication Date: 2025-05-23ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510092917.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems of oversegment or missegment in cerebrovascular segmentation tasks, especially in segmenting small and continuous blood vessels, and due to the small source of cerebrovascular data, it affects the learning and representation ability of neural networks in this field.

Method used

The U-shaped neural network model TRDSCN-UNet based on dynamic serpentine convolution and Transformer is adopted to design encoders and decoders by preprocessing cerebrovascular image data, and utilize Transformer's long-distance dependency modeling capabilities and local tubular structure feature extraction capabilities of dynamic serpentine convolution. Combined with the spatial channel attention module and the multi-view feature extraction module, the model's ability to segment small and continuous blood vessels is improved.

Benefits of technology

It improves the accuracy and continuity of cerebrovascular segmentation, effectively improves the model's performance on segmenting small and continuous blood vessels, and enhances the utilization rate of context information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032125A_ABST
    Figure CN120032125A_ABST
Patent Text Reader

Abstract

The invention discloses a cerebrovascular segmentation method and device based on dynamic snakelike convolution and Transform. The method comprises the following steps: (1) preprocessing cerebrovascular image data; (2) designing a U-shaped neural network model TRDSCN-UNet based on the dynamic snakelike convolution and the Transform, and constructing a U-shaped neural network model TRDSCN-UNet based on the dynamic snakelike convolution and the Transform; (3) training the neural network model TRDSCN-UNet proposed in the step (2); and (4) inputting test data into the model weight which is best in performance on the verification set to execute reasoning, completing voxel-level prediction classification, and performing three-dimensional reconstruction on a result obtained by reasoning. And the capability of segmenting fine and continuous blood vessels by the model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of 3D medical image segmentation and relates to a cerebral blood vessel segmentation method and device based on dynamic snake convolution and Transformer. Background Art

[0002] Cerebrovascular disease has become one of the major diseases threatening human health and safety. Common cerebrovascular diseases such as stroke, also known as infarction, are characterized by lesions in the cerebral arteries, including vascular stenosis, embolism, and atherosclerosis. Cerebrovascular disease is a key biomarker that is crucial in the diagnosis process and can be used to predict the possibility of a stroke or provide key information for clinical surgery, such as the radius and direction of cerebral blood vessels.

[0003] TOF-MRA (Time of Flight Magnetic Resonance Angiography) is an angiographic imaging technique with the advantages of being non-invasive, non-ionizing radiation-free, and not requiring the injection of vascular contrast agents. However, it is very challenging to accurately segment cerebral blood vessels from TOF-MRA images. During the imaging process, changes in blood flow velocity in cerebral blood vessels can cause changes in the contrast of TOF-MRA images, thereby generating additional noise. As for the structure of cerebral blood vessels, it is a very complex geometric topological structure, and the proportion of voxels in the vascular part to the entire TOF-MRA image is very low. In addition, the differences in cerebral vascular structure between different individuals are very large. If manual segmentation is used, it will take a lot of time and there will be some problems in accuracy. Therefore, automatic segmentation and reconstruction of cerebral blood vessels are of great significance for the treatment and research of brain diseases.

[0004] With the rapid development of deep learning, fully convolutional neural networks such as UNet have almost become the established standard for modern 3D medical image segmentation. They learn contextual representations in global and local features, but due to the inductive bias and weight sharing in the convolutional layer of the fully convolutional neural network, its receptive field is limited, which in turn restricts the ability of the neural network to learn global contextual representations. In the task of cerebrovascular segmentation, it mainly manifests as over-segmentation or wrong segmentation, especially in the segmentation of small and continuous blood vessels. In addition, there are very few sources of cerebrovascular data, which restricts the ability of many existing cerebrovascular segmentation neural networks to learn and represent cerebrovascular features. Summary of the invention

[0005] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a cerebral blood vessel segmentation method and device based on dynamic snake convolution and Transformer to improve the accuracy and continuity of cerebral blood vessel segmentation.

[0006] To achieve the above objectives, the first aspect of the present invention relates to a cerebrovascular segmentation method based on dynamic snake-shaped convolution and Transformer, comprising the following steps:

[0007] (1) Preprocess the cerebrovascular image data;

[0008] (2) Design a U-shaped neural network model TRDSCN-UNet based on dynamic snake-shaped convolution and Transformer;

[0009] (3) Train the neural network model TRDSCN-UNet proposed in step (2);

[0010] (4) Input the test data into the model weights that perform best on the validation set for inference, complete the voxel-level prediction classification, and then perform three-dimensional reconstruction on the inference results.

[0011] Preferably, in step (1), the cerebrovascular image data is preprocessed, the data is converted into the NIfTI format required for neural network input, and the Bet in the FSL software library is used for brain tissue extraction to complete the denoising of the data, where the Fractional intensity threshold parameter of Bet is set to 0.5.

[0012] Preferably, in step (2), a U-shaped neural network model TRDSCN-UNet based on dynamic snake-shaped convolution and Transformer is designed. To improve the model's ability to segment thin and continuous blood vessels, a TRDSCN-UNet model is proposed. The overall architecture is similar to that of 3D UNet, consisting of an encoder and a decoder, and skip connections are used to transfer features at the horizontal positions of the encoder and decoder; where the encoder is completely composed of Transformer blocks, encoding the input into a series of Patches and using self-attention to learn the weighted sum of the hidden layer values, improving the model's long-range dependence modeling ability; in the skip connection, the features from the encoder are processed by the spatial channel attention module (SCA) to extract significant features in the spatial and channel dimensions respectively, and perform multi-scale fusion with the features in the decoder; in the decoder, all upsampled features are extracted by the multi-view feature extraction module MVFE. The dynamic snake-shaped convolution from three different views and a CONV_Block are used to extract features from the same input respectively, and the outputs of the four convolutions are fused in the channel dimension and processed by the CONV_Block again to make the obtained features more conform to the thin and curved tubular structure of blood vessels; finally, through a 1×1×1 convolution operation, the prediction classification of cerebrovascular image voxels is completed; including:

[0013] (2.1) Design the encoder;

[0014] (2.2) Design skip connections;

[0015] (2.3) Design a decoder.

[0016] Further, step (2.1) specifically includes: the encoder first divides the input 3D data into several small patches, and adds the position encoding to the patch embedding to obtain the final representation of each patch; the number of patches is calculated as follows:

[0017]

[0018] Where H, W, D are the height, width and depth of the input image, P is the size of the small patch, and N patches is the number of patches, the specific value of P is 16;

[0019] Since the Transformer itself does not have the ability to perceive spatial structure, it is necessary to add position encoding, which is a learnable parameter with a dimension of 1×N patches × embed_dim; add spatial position information to each patch; add the position encoding to the patch embedding to get the final representation of each patch, where embed_dim is set to 768;

[0020] The data then passes through 12 layers of Transformer blocks, each of which consists of a multi-head self-attention MSA and a multi-layer perceptron MLP; the output of each layer is the input of the next layer; MSA divides the input query Q, key K, and value V into multiple heads, each of which calculates attention independently, and then concatenates the outputs of multiple heads to finally obtain a more expressive attention representation, which consists of n parallel self-attention (SA) heads, whose goal is to calculate the similarity between each position in the input sequence, and generate a new representation based on the weighted sum of this similarity; the number of self-attention heads n is 12; for each self-attention head h, Q, K, and V are generated through different linear transformations, and the calculation formula is as follows:

[0021]

[0022] in is the parameter matrix of each self-attention head;

[0023] For each head, calculate the similarity between Q and K to get the attention score, which is as follows:

[0024]

[0025] where d kThe dimension of each attention head, used to maintain the numerical stability of the model;

[0026] Then the attention score is normalized by Softmax to obtain the attention weight, and the obtained attention weight is used to calculate the value V h Perform weighted summation and finally concatenate all the attention head context vectors to get the output. The formula is as follows:

[0027]

[0028] MSA output =Concat(Context 1 ,Context 2 ,…,Context h )W o (6)

[0029] Among them, Context i is the context vector of the ith self-attention head, W o Multi-head attention trainable weights;

[0030] After self-attention, the representation of each position will pass through an MLP; residual connection and layer normalization are performed after each self-attention layer and MLP, which helps to speed up training and prevent gradient disappearance; during the feature extraction process, the outputs of the 3rd, 6th, 9th, and 12th layers are saved as features transferred by skip connections for subsequent multi-scale feature fusion with the decoder; this process regards feature extraction as sequence to sequence, which can more effectively model long-distance dependencies;

[0031] Step (2.2) specifically includes: the output of the third layer from the Transformer encoder is first subjected to two 3×3×3 convolutions, and the outputs of the 6th, 9th, and 12th layers are respectively subjected to 3, 2, and 1 DecBlocks, and then the data of each layer is weighted by the spatial channel attention module SCA; wherein DecBlock is composed of 2×2×2 deconvolution, 3×3×3 convolution, BatchNorm, and ReLU in series; the spatial attention part of the SCA module is processed by 1×1×c convolution, activated by sigmoid, and multiplied by the original input data element by element; the channel attention part is first subjected to Ada ptiveAvgPool3d operation, and then reduce the feature dimension from channels to channels / reduction through the Linear layer. After ReLU activation, another Linear operation is performed to restore the feature dimension to channels. The reduction parameter is set to 16. The processed data is multiplied element-wise with the original data after Sigmoid activation. Finally, the data processed by spatial attention and channel attention are added to the original data to obtain the fused features, so that the neural network can focus on significant features from both spatial and channel dimensions. The SCA calculation formula is as follows:

[0032] F out =F sp +F ch +F (7)

[0033]

[0034] F sp =mul(F,σ(conv 1×1×c (F))) (9)

[0035] Where F is the input feature, F ch is the feature extracted on the channel, F sp is the feature extracted in space, mul is the element-by-element multiplication, σ is the sigmoid operation, ch ext To perform a Linear(channels, channels / reduction) operation on the input, the output is subjected to a ReLU operation and then a Linear(channels / reduction, channels) operation. 1×1×c It is a convolution operation, and its convolution kernel size is 1×1×c, avg pool For AdaptiveAvgPool3d operation;

[0036] Step (2.3) specifically includes: each step in the decoder is to fuse the features transmitted by the jump connection with the upsampled features and process them through the multi-view feature extraction module MVFE; the data in MVFE are processed by three dynamic snake convolutions DS_CONV and a single 3D convolution module CONV_Block respectively and feature fusion is performed in the channel dimension, and the fused features are extracted again through CONV_Block, thereby expanding the receptive field of the model and improving the fitting effect of the vascular structure; among them, the morph parameters of the three dynamic snake convolutions are set to 0, 1, and 2 respectively, kernel_size is 3, and extend_scope is 1; CONV_Block is composed of 3×3×3 convolution, GroupNorm and ReLU in series.

[0037] Preferably, the step (3) specifically comprises: inputting the preprocessed cerebrovascular image in the form of a data-label Patch pair into TRDSCN-UNet for training, the image and label size is 448×448×128, divided into Patch pairs of 96×96×96; training uses the Adam optimizer, the initial learning rate is 0.002, a total of 2000 epochs are trained, and the learning rate decays to 80% of the original every 20 epochs; during training, the batch-size is 1, and the loss function is the weighted sum of CrossEntropy and Dice losses (1:1); saving the model every 50 rounds, and updating the weight of the model with the best performance on the validation set, and the loss function formula is as follows:

[0038] Loss = L ce +L dice (10)

[0039]

[0040] Among them, L ce is the CrossEntropy loss, L dice is the dice loss, N represents the number of voxels, C is the number of categories, the value of C is 2, y i is the true label of sample i, is the predicted probability of sample i, and ∈ is a small constant used to prevent division by zero.

[0041] Preferably, the step (4) inputs the test data into the model weights that perform best on the validation set to perform inference, complete voxel-level prediction classification, and perform three-dimensional reconstruction of the inference results.

[0042] The second aspect of the present invention relates to a cerebral blood vessel segmentation device based on dynamic snake convolution and Transformer, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer of the present invention.

[0043] The third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer of the present invention.

[0044] The advantages of the present invention are: data preprocessing reduces the impact of noise on data training; adding an SCA module to the jump connection enhances the model's ability to perceive significant features in space and channels; combining the long-distance dependency modeling capability of the Transformer in the encoder and the local tubular structure feature extraction capability of the dynamic snake convolution in the decoder improves the utilization rate of the neural network model for contextual information; and can effectively improve the model's ability to segment small and continuous blood vessels. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Flow chart of the method of the present invention.

[0046] Figure 2 This is the architecture diagram of the TRDSCN-UNet model of the present invention.

[0047] Figure 3 This is a structural diagram of the spatial channel attention module SCA of the present invention.

[0048] Figure 4 It is a structural diagram of the multi-view feature extraction module MVFE of the present invention. DETAILED DESCRIPTION

[0049] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0050] Example 1

[0051] The implementation process of the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer of the present invention is as follows: Figure 1 As shown. It mainly includes the following steps:

[0052] 1. Preprocess cerebrovascular imaging data;

[0053] The data were converted into the NIfTI format required by the neural network input, and Bet of the FSL software library was used to extract brain tissue and complete the data denoising, where the Fractional intensity threshold parameter of Bet was set to 0.5.

[0054] 2. Design a U-shaped neural network model TRDSCN-UNet based on dynamic snake convolution and Transformer;

[0055] In order to improve the model's ability to segment small and continuous blood vessels, a TRDSCN-UNet model is proposed. Figure 2 As shown. The overall architecture is similar to 3D UNet, consisting of an encoder and a decoder, with jump connections used to transfer features at the horizontal position of the encoder and decoder; the encoder is composed entirely of Transformer blocks, which encodes the input into a series of patches and uses self-attention to learn the weighted sum of hidden layer values, improving the model's long-distance dependency modeling ability; in the jump connection, the features from the encoder are processed by the spatial channel attention module (SCA), and significant features are extracted in space and channels respectively, and multi-scale fusion is performed with the features in the decoder; in the decoder, all upsampled features are extracted by the multi-view feature extraction module MVFE, and three dynamic snake convolutions from different perspectives and a CONV_Block are used to extract features from the same input respectively, and the outputs of the four convolutions are fused in the channel dimension and processed again by CONV_Block, so that the obtained features can be more consistent with the small and curved tubular structure of blood vessels; finally, a 1×1×1 convolution operation is used to complete the prediction and classification of cerebrovascular image voxels.

[0056] The specific design plan is as follows:

[0057] 2.1 Designing the Encoder

[0058] The encoder first embeds the input 3D data and divides the input 3D image into several small patches. The size of each small patch is P×P×P. The formula is as follows:

[0059]

[0060] Where H, W, D are the height, width and depth of the input image, P is the size of the small patch, and N patches is the number of patches, and the specific value of P is 16.

[0061] Since the Transformer itself does not have the ability to perceive spatial structure, it is necessary to add position encoding, which is a learnable parameter with a dimension of 1×Npatches × embed_dim. Add spatial position information to each patch. Add the position encoding to the patch embedding to get the final representation of each patch. In this paper, embed_dim is set to 768.

[0062] The data then passes through 12 layers of Transformer blocks, each of which consists of multi-head self-attention (MSA) and multi-layer perceptron (MLP). The output of each layer is the input of the next layer. The core idea of ​​MSA is to divide the input query (Q), key (K), and value (V) into multiple heads, each of which calculates attention independently, and then concatenates the outputs of multiple heads to finally obtain a more expressive attention representation, which consists of n parallel self-attention (SA) heads. Its goal is to calculate the similarity between each position in the input sequence, and generate a new representation based on the weighted sum of this similarity. The number of self-attention heads n is 12. For each self-attention head h, Q, K, and V are generated through different linear transformations, and the calculation formula is as follows:

[0063]

[0064] in is the parameter matrix of each head.

[0065] For each head, calculate the similarity between Q and K to get the attention score, which is as follows:

[0066]

[0067] where d k The dimension of each attention head, used to maintain the numerical stability of the model.

[0068] Then the attention score is normalized by Softmax to obtain the attention weight, and the obtained attention weight is used to calculate the value V h Perform weighted summation and finally concatenate all the attention head context vectors to get the output. The formula is as follows:

[0069]

[0070] MSA output =Concat(Context 1 ,Context 2 ,…,Context h )W o (6)

[0071] Among them, Context i is the context vector of the i-th attention head, W oTrainable weights for multi-head attention.

[0072] After self-attention, the representation of each position will pass through an MLP; residual connection and layer normalization after each self-attention layer and MLP help speed up training and prevent gradient disappearance; during feature extraction, the outputs of the 3rd, 6th, 9th, and 12th layers are saved as features passed by skip connections for subsequent multi-scale feature fusion with the decoder; this process treats feature extraction as sequence-to-sequence, which can more effectively model long-distance dependencies;

[0073] 2.2 Design of skip connections

[0074] The output of the third layer from the Transformer encoder first passes through two 3×3×3 convolutions, and the outputs of the 6th, 9th, and 12th layers pass through 3, 2, and 1 DecBlocks respectively. Then, the data of each layer is weighted by the spatial channel attention module SCA. DecBlock is composed of 2×2×2 deconvolution, 3×3×3 convolution, BatchNorm, and ReLU in series. The structure of the SCA module is as follows: Figure 3 As shown in the figure, the spatial attention part of the SCA module is processed by 1×1×c convolution, activated by sigmoid, and multiplied by the original input data element by element; the channel attention part first performs AdaptiveAvgPool3d operation, and then reduces the feature dimension from channels to channels / reduction through the Linear layer. After ReLU activation, a Linear operation is performed to restore the feature dimension to channels. The reduction parameter is set to 16. The processed data is activated by Sigmoid and multiplied by the original data element by element; finally, the data processed by spatial attention and channel attention are added to the original data to obtain the fused features, so that the neural network can focus on significant features from the two dimensions of space and channel; the SCA module in this step is the fusion of the features processed by spatial attention, the features processed by channel attention, and the original data. The SCA calculation formula is as follows:

[0075] F out =F sp +F ch +F (7)

[0076]

[0077] F sp =mul(F,σ(conv 1×1×c (F))) (9)

[0078] Where F is the input feature, F ch is the feature extracted on the channel, Fsp is the feature extracted in space, mul is the element-by-element multiplication, σ is the sigmoid operation, ch ext To perform a Linear(channels, channels / reduction) operation on the input, the output is subjected to a ReLU operation and then a Linear(channels / reduction, channels) operation. 1×1×c It is a convolution operation, and its convolution kernel size is 1×1×c, avg pool For AdaptiveAvgPool3d operation;

[0079] 2.3 Designing the decoder

[0080] Each step in the decoder is to fuse the features transmitted by the jump connection with the upsampled features and process them through the multi-view feature extraction module MVFE; the structure of the MVFE module is as follows Figure 4 As shown in the figure, the data in MVFE are processed by three dynamic snake convolutions DS_CONV and a single 3D convolution module CONV_Block, and feature fusion is performed in the channel dimension. The fused features are extracted again by CONV_Block, thereby expanding the receptive field of the model and improving the fitting effect of the vascular structure. Among them, the morph parameters of the three dynamic snake convolutions are set to 0, 1, and 2, kernel_size is 3, and extend_scope is 1; CONV_Block is composed of 3×3×3 convolution, GroupNorm and ReLU in series.

[0081] 3. Train the neural network model TRDSCN-UNet proposed in step (2);

[0082] The preprocessed cerebrovascular images were input into TRDSCN-UNet in the form of data-label Patch pairs for training. The image and label size was 448×448×128, divided into Patch pairs of 96×96×96. The Adam optimizer was used for training, with an initial learning rate of 0.002. A total of 2000 epochs were trained, and the learning rate decayed to 80% of the original every 20 epochs. The batch-size during training was 1, and the loss function was the weighted sum of CrossEntropy and Dice losses (1:1). The model was saved every 50 rounds, and the weight of the model with the best performance on the validation set was updated. The loss function formula is as follows:

[0083] Loss = L ce +L dice (10)

[0084]

[0085] Among them, L ce is the CrossEntropy loss, L dice is the dice loss, N represents the number of voxels, C is the number of categories, the value of C is 2, y i is the true label of sample i, is the predicted probability of sample i, and ∈ is a small constant used to prevent division by zero.

[0086] 4. Input the test data into the model weights that performed best on the validation set to perform inference, complete voxel-level prediction classification, and reconstruct the inference results in three dimensions.

[0087] Example 2

[0088] The present embodiment relates to a cerebral blood vessel segmentation device based on dynamic snake convolution and Transformer, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer of Example 1.

[0089] Example 3

[0090] This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer of Embodiment 1 is implemented.

[0091] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A cerebral vascular segmentation method based on dynamic snake convolution and Transformer, comprising the following steps: (1) Preprocessing of cerebrovascular imaging data; (2) Design a U-shaped neural network model TRDSCN-UNet based on dynamic snake convolution and Transformer; (3) Training the neural network model TRDSCN-UNet proposed in step (2); (4) The test data is input into the model weights that perform best on the validation set to perform inference, complete voxel-level prediction and classification, and then the inference results are reconstructed in three dimensions.

2. The cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer as claimed in claim 1, characterized in that: Step (1) preprocess the cerebrovascular imaging data, convert the data into the NIfTI format required by the neural network input, and use Bet of the FSL software library to extract brain tissue and complete the data denoising, where the Fractional intensity threshold parameter of Bet is set to 0.

5.

3. The cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer as claimed in claim 1, characterized in that: Step (2) Design a U-shaped neural network model TRDSCN-UNet based on dynamic snake convolution and Transformer. In order to improve the model's ability to segment small and continuous blood vessels, a TRDSCN-UNet model is proposed. The overall architecture is similar to 3D UNet, consisting of an encoder and a decoder. Skip connections are used at the horizontal position of the encoder and decoder to transfer features. The encoder is composed entirely of Transformer blocks, which encode the input into a series of patches and use self-attention to learn the weighted sum of hidden layer values, improving the model's ability to model long-distance dependencies. In the jump connection, the features from the encoder are processed by the spatial channel attention module (SCA), and significant features are extracted in space and channels respectively, and then multi-scale fused with the features in the decoder; in the decoder, all upsampled features are extracted by the multi-view feature extraction module MVFE, and three dynamic snake convolutions of different views and a CONV_Block are used to extract features from the same input respectively, and the outputs of the four convolutions are fused in the channel dimension and processed again by CONV_Block, so that the obtained features can better conform to the small and curved tubular structure of blood vessels; finally, a 1×1×1 convolution operation is used to complete the prediction and classification of cerebrovascular image voxels; include: (2.1) Design the encoder; (2.2) Design skip connections; (2.3) Design a decoder.

4. According to claim 3, a cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer is characterized in that: Step (2.1) specifically includes: the encoder first divides the input 3D data into several small patches, and adds the position encoding to the patch embedding to obtain the final representation of each patch; the number of patches is calculated as follows: Where H, W, D are the height, width and depth of the input image, P is the size of the small patch, and N patches is the number of patches, the specific value of P is 16; Since the Transformer itself does not have the ability to perceive spatial structure, it is necessary to add position encoding, which is a learnable parameter with a dimension of 1×N patches × embed_dim; add spatial position information to each patch; add the position encoding to the patch embedding to get the final representation of each patch, and embed_dim is set to 768; The data then passes through 12 layers of Transformer blocks, each of which consists of a multi-head self-attention MSA and a multi-layer perceptron MLP; the output of each layer is the input of the next layer; MSA divides the input query Q, key K, and value V into multiple heads, each of which calculates attention independently, and then concatenates the outputs of multiple heads to finally obtain a more expressive attention representation, which consists of n parallel self-attention SA heads. Its goal is to calculate the similarity between each position in the input sequence, and generate a new representation based on the weighted sum of this similarity; the number of self-attention heads n is 12; for each self-attention head h, Q, K, and V are generated through different linear transformations, and the calculation formula is as follows: in is the parameter matrix of each self-attention head; For each head, calculate the similarity between Q and K to get the attention score, which is as follows: where d k The dimension of each attention head, used to maintain the numerical stability of the model; Then the attention score is normalized by Softmax to obtain the attention weight, and the obtained attention weight is used to calculate the value V h Perform weighted summation and finally concatenate all the attention head context vectors to get the output. The formula is as follows: MSA output =Concat(Context1,Context2,…,Context h )W o (6) Among them, Context i is the context vector of the ith self-attention head, W o Multi-head attention trainable weights; After self-attention, the representation of each position will pass through an MLP; residual connection and layer normalization are performed after each self-attention layer and MLP, which helps to speed up training and prevent gradient disappearance; during the feature extraction process, the outputs of the 3rd, 6th, 9th, and 12th layers are saved as features transferred by skip connections for subsequent multi-scale feature fusion with the decoder; this process regards feature extraction as sequence to sequence, which can more effectively model long-distance dependencies; Step (2.2) specifically includes: the output of the third layer from the Transformer encoder is first subjected to two 3×3×3 convolutions, and the outputs of the 6th, 9th, and 12th layers are respectively subjected to 3, 2, and 1 DecBlocks, and then the data of each layer is weighted by the spatial channel attention module SCA; wherein DecBlock is composed of 2×2×2 deconvolution, 3×3×3 convolution, BatchNorm, and ReLU in series; the spatial attention part of the SCA module is processed by 1×1×c convolution, activated by sigmoid, and multiplied by the original input data element by element; the channel attention part is first subjected to Ada ptiveAvgPool3d operation, and then reduce the feature dimension from channels to channels / reduction through the Linear layer. After ReLU activation, another Linear operation is performed to restore the feature dimension to channels. The reduction parameter is set to 16. The processed data is multiplied element-wise with the original data after Sigmoid activation. Finally, the data processed by spatial attention and channel attention are added to the original data to obtain the fused features, so that the neural network can focus on significant features from both spatial and channel dimensions. The SCA calculation formula is as follows: F out =F sp +F ch +F (7) F sp =mul(F,σ(conv 1×1×c (F))) (9) Where F is the input feature, F ch is the feature extracted on the channel, F sp is the feature extracted in space, mul is the element-by-element multiplication, σ is the sigmoid operation, ch ext To perform a Linear(channels, channels / reduction) operation on the input, the output is subjected to a ReLU operation and then a Linear(channels / reduction, channels) operation. 1×1×c It is a convolution operation, and its convolution kernel size is 1×1×c, avg pool For AdaptiveAvgPool3d operation; Step (2.3) specifically includes: each step in the decoder is to fuse the features transmitted by the jump connection with the upsampled features and process them through the multi-view feature extraction module MVFE; the data in MVFE are processed by three dynamic snake convolutions DS_CONV and a single 3D convolution module CONV_Block respectively and feature fusion is performed in the channel dimension, and the fused features are extracted again through CONV_Block, thereby expanding the receptive field of the model and improving the fitting effect of the vascular structure; among them, the morph parameters of the three dynamic snake convolutions are set to 0, 1, and 2 respectively, kernel_size is 3, and extend_scope is 1; CONV_Block is composed of 3×3×3 convolution, GroupNorm and ReLU in series.

5. The cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer according to claim 1, characterized in that: The step (3) specifically includes: inputting the preprocessed cerebrovascular image in the form of a data-label patch pair into TRDSCN-UNet for training, the image and label size is 448×448×128, divided into 96×96×96 patch pairs; using the Adam optimizer for training, the initial learning rate is 0.002, a total of 2000 epochs are trained, and the learning rate decays to 80% of the original every 20 epochs; during training, the batch-size is 1, and the loss function is the weighted sum of CrossEntropy and Dice losses (1:1); saving the model every 50 rounds, and updating the weight of the model with the best performance on the validation set, and the loss function formula is as follows: Loss=L ce +L dice (10) Among them, L ce is the CrossEntropy loss, L dice is the dice loss, N represents the number of voxels, C is the number of categories, the value of C is 2, y i is the true label of sample i, is the predicted probability of sample i, and ∈ is a small constant used to prevent division by zero.

6. The cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer according to claim 1, characterized in that: The step (4) inputs the test data into the model weights that perform best on the validation set to perform reasoning, complete voxel-level prediction classification, and perform three-dimensional reconstruction of the reasoning results.

7. A cerebral vascular segmentation device based on dynamic snake convolution and Transformer, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the cerebral blood vessel segmentation method based on dynamic snake convolution and Transformer described in any one of claims 1 to 6 is implemented.