A vascular image segmentation method and system based on a three-dimensional deep network
By integrating attention-guided feature fusion, scale-aware feature enhancement and multi-scale feature aggregation modules in the U-shaped encoder-decoder network, the problem of coronary segmentation method in the prior art is solved, and a higher precision coronary image segmentation is achieved.
Patent Information
- Application Number
- CN202211026435.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-08-25
AI Technical Summary
Existing coronary segmentation methods are difficult to effectively extract multi-scale context information when dealing with complex coronary anatomy and morphology, and traditional jump connections introduce irrelevant background noise, making it difficult to distinguish coronary arteries from surrounding tissues.
Using a three-dimensional deep network based on U-shaped encoder-decoder, an attention-guided feature fusion module, a scale-aware feature enhancement module and a multi-scale feature aggregation module are integrated. The attention-guided feature fusion module replaces the jump connection between the traditional encoder and the decoder. Combined with the scale-aware feature enhancement module and a multi-scale feature aggregation module, multi-scale feature information is adaptively extracted and aggregated, suppressed background noise and separated coronary arteries and veins.
It improves the accuracy and robustness of coronary artery segmentation, can handle complex anatomical structures with large scale changes, effectively separate coronary artery from noise, enhances the network's feature representation ability, and achieves more accurate coronary artery image segmentation.
Smart Images

Figure CN115546570B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image recognition, and in particular relates to a vascular image segmentation method and system based on a three-dimensional deep network. Background Art
[0002] Coronary computed tomography angiography (CCTA) is a widely used, non-invasive, and highly sensitive imaging method for routine clinical diagnosis of cardiovascular disease. It utilizes the injection of iodinated contrast agents to visualize the coronary artery anatomy. In clinical trials, coronary artery segmentation is a crucial step for a range of tasks, including plaque assessment, stenosis detection, and centerline extraction. However, with the increasing number of CCTA examination requests and the complex coronary artery anatomy and morphology, manual processing is time-consuming and technically demanding, overwhelming the post-processing workforce. Furthermore, this process can be subject to inter-observer inconsistencies. Therefore, accurate, automated coronary artery segmentation is crucial and in increasing demand.
[0003] In recent years, a vast literature has been published on vessel segmentation methods, encompassing both classic machine learning and modern deep learning approaches, to assist in this challenging task. The former primarily relies on rule-based approaches, such as tracking-based algorithms (Moccia et al., 2018), active contour models (Lesage et al., 2016), or image filtering and enhancement algorithms (Frangi et al., 1998), which utilize various vascular image features to segment vessels. Compared to successful approaches based on deep learning strategies, these traditional machine learning methods can achieve satisfactory performance with limited training data. However, all traditional methods are based on some prior knowledge (starting points, seed points, or coarse segmentation points), which exhibits limitations when dealing with challenging conditions such as image noise, brightness variations, and blurred boundaries. Consequently, these methods generally suffer from a heavy reliance on handcrafted features and learning methods, making it difficult to achieve the desired level of robustness for segmenting entire coronary arteries.
[0004] Deep learning based on convolutional neural networks (CNNs) has attracted widespread attention in the field of medical image computing and has achieved milestones in several medical image analysis applications. Unlike traditional manual feature extraction methods, this data-driven approach can automatically extract complex image features, facilitating various vascular segmentation tasks. Currently, the most effective vascular segmentation work relies on CNN-based architectures, particularly U-Net and its variants, as well as V-Net and its variants. Early vascular segmentation methods, such as retinal vessel segmentation, were based on 2D methods. Later, as the segmentation target shifted to 3D images, 3D methods became mainstream. Due to the spatial continuity of 3D images along the z-axis, most 2D methods that ignore 3D spatial context cannot be directly converted to 3D images. Therefore, the current state-of-the-art coronary artery segmentation solutions focus primarily on 2D multipath (2.5D) and 3D methods.
[0005] While these U-shaped architectures achieve impressive performance, they are still insufficient to address the challenges of coronary artery segmentation. First, the networks typically lack the ability to effectively extract multi-scale contextual information. Second, traditional skip connections at each stage typically directly incorporate local information, introducing excessive irrelevant background noise, making it challenging to distinguish coronary arteries from surrounding tissues (such as pulmonary vessels and coronary veins). Summary of the Invention
[0006] To address the problems existing in the prior art, the present invention provides a vascular image segmentation method based on a three-dimensional deep network. This method effectively extracts multi-scale contextual information, adaptively combines semantic and spatial information while suppressing irrelevant background noise, and separates coronary arteries from veins and noise, achieving excellent performance in coronary artery image segmentation tasks.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a vascular image segmentation method based on a three-dimensional deep network, based on a multi-attention, multi-scale three-dimensional deep network, wherein the three-dimensional deep network is based on a U-shaped encoder-decoder, and an attention-guided feature fusion module, a scale-aware feature enhancement module, and a multi-scale feature aggregation module are integrated into the three-dimensional deep network of the U-shaped encoder-decoder. The attention-guided feature fusion module is used to replace the traditional skip connection between the encoding and decoding stages, the scale-aware feature enhancement module is embedded in the bottom of the network, and the multi-scale feature aggregation module is integrated into the decoding path; the method comprises the following steps:
[0008] The encoder extracts low-level features F from the vascular image l (i) , the decoder extracts high-level features The feature map generated in the last stage of the encoder is split into four parallel feature groups. Four dilated convolutions and soft attention mechanisms with four different dilation rates are deployed in parallel branches. The four branch features are spliced using channels to obtain hierarchical features F. The hierarchical features F are input into three convolutional layers respectively to generate three feature maps Q, K and V. The matrix transposition of the feature map Q is performed to obtain Q. T , Q T After matrix multiplication with the feature map K, it is input into the softmax activation layer to obtain the encoding of the feature relationship between the sagittal and coronal positions. Finally, the encoding result is matrix multiplied with the feature map V and then connected with the residual of the hierarchical feature F to obtain the pixel-level attention enhancement feature;
[0009] Input the pixel-level attention-enhanced features and the features generated by the fourth stage E4 in the encoder into the attention-guided feature fusion module;
[0010] The attention-guided feature fusion module uses channel splicing to combine low-level features F l (i) and advanced features Combined, the low-level features are refined by calculating the weight vector α, and the refined low-level features are added to the high-level features to obtain the fusion feature F (i) ; Fusion feature F (i) As the input of the i-th stage in the decoder, i = 1, 2, 3, 4;
[0011] The convolutional layer is used to compress the features generated by the four stages in the decoder. The features of the middle and high stages of adjacent scales are then upsampled and concatenated with the features of the low stage. The concatenated features are then subjected to a soft attention guidance operation. The obtained soft attention-guided features are input into the convolutional layers of the two channels. Finally, the sigmoid activation function is used to calculate the aggregated features between the features of adjacent scales, i.e., the vascular image segmentation results.
[0012] The low-level features F are concatenated using channel concatenation. l (i) and advanced features The specific combination is as follows:
[0013]
[0014] in is the splicing operation; f u Represents an upsampling operation.
[0015] Use average pooling information to stimulate feature channel information, and use maximum pooling features to retain information; then use a shared multi-layer perception layer to perform two pooling operations, input the obtained results into the Sigmoid function, and obtain the weight vector α:
[0016]
[0017] where f mlp represents the MLP operator; f gap represents the global average pooling; f gmp represents the global maximum pool; f σ Represents the sigmoid activation function.
[0018] By gradually aggregating adjacent scale features in the decoding path to obtain more global semantic representations of blood vessels, the vascular map can be refined. Specifically, the high-level semantic information of small-scale features is transferred to large-scale features:
[0019]
[0020]
[0021]
[0022] in For splicing operation; represents matrix multiplication; f u represents the upsampling operation; f σ Represents the sigmoid activation function, f c″ It is a convolution operation with a convolution kernel of 3×3×3.
[0023] Weighted cross entropy loss L WCE As the loss function, the dice similarity coefficient loss L is introduced into it DSC ,
[0024]
[0025] Where α is L WCE and L DSC The weight balance parameter between is determined based on experience.
[0026]
[0027]
[0028] ω is the weight of the vascular structure and can be used to estimate the probability p of all pixels. i get:
[0029]
[0030] N represents the number of pixels, p i ∈[0, 1] and g i ∈0, 1 represent the predicted probability and background true value i respectively thThe vascular structure of the pixel, the Laplace smoothing factor ∈ is used to avoid numerical instability problems and accelerate the convergence of the training process.
[0031] In the multi-scale feature aggregation module, a 1×1×1 convolution is first used to uniformly compress the number of channels of the D4, D3, D2 and D1 feature maps to 8 to reduce the number of parameters for subsequent operations. The features after channel compression in the four stages are recorded as F4, F3, F2 and F1 respectively; then the feature aggregation operation is performed on adjacent scales; the feature aggregation operation on adjacent scales is specifically as follows: first, the high-stage feature F4 is upsampled and channel-joined with the low-stage feature F3, and then a 1×1×1 convolution and a sigmoid activation function are used to obtain the attention map of the adjacent scale features and perform matrix multiplication with the low-stage feature F3, and then a 3×3×3 convolution is used to generate information-rich and detailed features. right After upsampling, it is passed to the next scale feature F2, and the feature aggregation operation is performed with the next scale feature F2 to obtain After upsampling, feature aggregation operation is performed with feature F1 to obtain Finally, a 1×1×1 convolution with a channel number of 2 and a sigmoid activation function are used to obtain the output of the multi-scale feature aggregation module.
[0032] Based on the concept of the method described in the present invention, a vascular image segmentation system based on a three-dimensional deep network is also provided. The system is based on a multi-attention, multi-scale three-dimensional deep network. The three-dimensional deep network is based on a U-shaped encoder-decoder and includes an attention-guided feature fusion module, a scale-aware feature enhancement module, and a multi-scale feature aggregation module. The attention-guided feature fusion module, the scale-aware feature enhancement module, and the multi-scale feature aggregation module are integrated into the three-dimensional deep network of the U-shaped encoder-decoder. The attention-guided feature fusion module replaces the traditional skip connection between the encoding and decoding stages. The scale-aware feature enhancement module is embedded at the bottom of the network, and the multi-scale feature aggregation module is integrated into the decoding path.
[0033] The encoder extracts low-level features F from the vascular image l (i) , the decoder extracts high-level features
[0034] The scale-aware feature enhancement module is used to split the feature map generated in the last stage of the encoder into four parallel feature groups, deploy four dilated convolutions and soft attention mechanisms with four different dilation rates in parallel branches, and use channels to splice the four branch features to obtain hierarchical features F; the hierarchical features F are input into three convolutional layers respectively to generate three feature maps Q, K and V, and matrix transposition is performed on the feature map Q to obtain QT , Q T After matrix multiplication with the feature map K, it is input into the softmax activation layer to obtain the encoding of the feature relationship between the sagittal and coronal positions. Finally, the encoding result is matrix multiplied with the feature map V and then connected with the residual of the hierarchical feature F to obtain the pixel-level attention enhancement feature;
[0035] The attention-guided feature fusion module is used to combine low-level features F l (i) and advanced features Combined, the low-level features are refined by calculating the weight vector α, and the refined low-level features are added to the high-level features to obtain the fusion feature F (i) ; Fusion feature F (i) As the input of the i-th stage in the decoder, i = 1, 2, 3, 4;
[0036] The multi-scale feature aggregation module uses convolutional layers to perform channel compression on the features generated at each stage in the decoder. It then upsamples the features of the middle and high stages at adjacent scales and concatenates them with the features of the low stages. Soft attention is then performed on the concatenated features, and the obtained soft attention-guided features are input into the convolutional layers of the two channels. Finally, the sigmoid activation function is used to calculate the aggregated features between the features of adjacent scales, i.e., the vascular image segmentation results.
[0037] In addition, a computer device is provided, comprising a processor and a memory, wherein an executable program is stored in the memory, and when the processor executes the executable program, the blood vessel image segmentation method based on the three-dimensional deep network described in the present invention can be executed.
[0038] A computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is executed by a processor, the blood vessel image segmentation method based on a three-dimensional deep network of the present invention can be implemented.
[0039] Compared with the prior art, the present invention has at least the following beneficial effects:
[0040] The method described in the present invention is based on a novel multi-attention, multi-scale three-dimensional deep network (CAS-Net). The proposed method can handle large-scale variations in coronary arteries and extract representative features from their complex anatomical structure and morphology. The present invention designs an attention-guided feature fusion (AGFF) module to adaptively select useful semantic and spatial information, which can not only suppress low-level irrelevant background noise but also retain more detailed local semantic information, thereby separating coronary arteries from veins and noise. The scale-aware feature enhancement (SAFE) module is used to effectively extract hidden multi-scale contextual information and aggregate multi-scale features. This module can enhance the proposed CAS-Net's ability to handle complex situations, such as the large variation in size and shape of coronary artery image features, and even many of them are intertwined. The multi-scale feature aggregation (MSFA) module is further applied to learn more global semantic representations in the decoding path, improving segmentation accuracy and thus refining the vascular map. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Schematic diagram of the structure of the model of the present invention;
[0042] Figure 2 Schematic diagram of the attention-guided feature fusion module structure;
[0043] Figure 3 This is a schematic diagram of the structure of the scale-aware feature enhancement module;
[0044] Figure 4 This is a schematic diagram of the structure of the multi-scale feature aggregation module;
[0045] Figure 5 To compare the results of different methods; DETAILED DESCRIPTION
[0046] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention, with reference to the accompanying drawings. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, but not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0047] It has been shown that high-level features generally capture global semantic information (i.e., vessel shape), while low-level features contain spatial structural details (i.e., vessel contours). Therefore, the present invention fully combines features from different layers in the network to facilitate effective fusion of features from adjacent layers.
[0048] This paper proposes a multi-attention, multi-scale 3D deep network, CAS-Net, to comprehensively address the challenge of coronary artery segmentation. Based on the classic encoder-decoder architecture, it comprises three core modules. First, the paper replaces traditional skip connections with an attention-guided feature fusion (AGFF) module. The AGFF module adaptively combines semantic and spatial information while suppressing irrelevant background noise. The AGFF module effectively guides the fusion of features from adjacent layers in the encoding and decoding stages, extracting more distinguishable semantic information and thus separating coronary arteries from veins and noise. Second, the paper proposes a scale-aware feature enhancement (SAFE) module, added at the bottom of the network, to effectively extract hidden multi-scale contextual information and dynamically aggregate multi-scale features, thereby enhancing the network's feature representation capabilities. Specifically, the paper first designs a parallel architecture, applying convolutional layers with different dilation rates to parallel branches to obtain richer feature maps with different receptive fields. Then, by calculating the spatial importance of different scales, these hierarchical features are dynamically adjusted to accommodate vessels of varying scales. Furthermore, the present invention designs a Multi-Scale Feature Aggregation (MSFA) module to learn more global semantic representations in the decoding path, improving segmentation accuracy and thus refining the vascular map. Experiments on a CCTA image dataset collected by the present invention demonstrate that the proposed method achieves good performance, outperforming other state-of-the-art methods.
[0049] Network structure
[0050] In order to achieve the accuracy and robustness of blood vessel segmentation, this paper proposes a multi-attention, multi-scale three-dimensional deep network (CAS-Net), such as Figure 1 As shown in Figure 2, the three modules proposed in this paper, AGFF, SAFE, and MSFA, are seamlessly integrated into a U-shaped encoder-decoder architecture. The encoder consists of five stages, E1, E2, E3, E4, and E5, which gradually extract vascular features at various levels and then gradually extract detailed features into higher-order semantic features. The decoder consists of four stages, D1, D2, D3, and D4. The AGFF module replaces the traditional skip connections between the encoding and decoding stages, the SAFE module is embedded at the bottom of the network, and the MSFA module is finally integrated into the decoding path.
[0051] Attention guided feature fusion module AGFF, such as Figure 2 As shown, where F l (i) Low-level features from the encoding stage, Represents the high-level features from the decoding stage. First, After upsampling and F l (i)Perform channel splicing and mark the result as Then They enter the global maximum pooling (GMP) and global average pooling (GAP) branches respectively, which are called branch 1 and branch 2 from left to right. Then the pooling results of branch 1 and branch 2 are input into the multi-layer perceptron (MLP) of the shared structure respectively, and the results are added to obtain more global context information, which is marked as Will Input into a sigmoid activation function to obtain the weight vector α; in addition, in order to refine the low-level features, the weight vector α is combined with the low-level features F l (i) Multiply to obtain reweighted low-level features Finally, the reweighted low-level features With advanced features Add together to produce the final result F (i) To suppress the interference of irrelevant background noise and retain more important semantic context information. Figure 1 As shown in the figure, between the first stage E1 and D2, the redesigned attention-guided feature fusion module AGFF is used to replace the traditional jump connection, and between the second stage E2 and D3, between the third stage E3 and D4, and between the fourth stage E4 and D5 (SAFE module), the attention-guided feature fusion module AGFF is used to replace the traditional jump connection.
[0052] Scale-aware feature enhancement module SAFE Figure 3 As shown in the figure, conv(x×y×z) indicates that the convolution has different convolution kernel sizes in the three dimensions of x, y, and z. Dilated Conv(rate=x) indicates the dilation rate of the dilated convolution. The input of the SAFE module is the feature map generated by the fifth stage E5 in the encoder, see Figure 1 .like Figure 3 As shown, the feature map generated by the fifth stage E5 in the encoder is first split into four parallel feature groups, marked as f i , i∈{1, 2, 3, 4}; then the dilated convolution and soft attention mechanisms with different dilation rates are deployed in four parallel branches to obtain rich features with multiple receptive fields, marked as f i 'i∈{1, 2, 3, 4}, in the present invention, the expansion rate of the dilated convolution is set to 1, 2, 3, 4 in the four parallel branches, and the soft attention mechanism is implemented by the sigmoid activation function; then the four branch features f'1, f'2, f'3, f'4 are concatenated by channels to obtain the hierarchical feature F; the hierarchical feature F is input into the three convolution layers respectively to generate three feature maps Q, K and V. Matrix transposition is performed on the feature map Q to obtain Q T , QT After matrix multiplication with the feature map K, it is input into the softmax activation layer to obtain the encoding of the feature relationship between the sagittal and coronal planes. Finally, the encoding result is matrix multiplied with the feature map V and then residually connected with the hierarchical feature F to obtain the pixel-level attention enhancement feature.
[0053] Multi-scale feature aggregation module MSFA, such as Figure 4 As shown in , where in_c represents the number of channels of the convolution input, out_c represents the number of channels of the convolution output, kernel_size represents the size of the convolution kernel, and ratio represents the upsampling ratio. The input of the MSFA module is the feature map generated by the four stages D4, D3, D2 and D1 in the decoder, see Figure 1 . As shown in the figure, a 1×1×1 convolution is first used to uniformly compress the number of channels of the D4, D3, D2 and D1 feature maps to 8 to reduce the number of parameters for subsequent operations. The features after channel compression in the four stages are recorded as F4, F3, F2 and F1 respectively; then feature aggregation operations are performed on adjacent scales. The specific operation is to take F4 and F3 as an example. First, the high-stage feature F4 is upsampled and then channel-joined with the low-stage feature F3. Then, a 1×1×1 convolution and a sigmoid activation function are used to obtain the attention map of the adjacent scale features and perform matrix multiplication with the low-stage feature F3. Then, a 3×3×3 convolution is used to generate information-rich and detailed features. The results are marked as After upsampling, it is passed to the next scale feature F2, and the feature aggregation operation is performed with the next scale feature F2 to obtain After upsampling, feature aggregation operation is performed with feature F1 to obtain Finally, a 1×1×1 convolution with a channel number of 2 and a sigmoid activation function are used to obtain the output of the MSFA module.
[0054] In order to adaptively select useful semantic information and spatial information, this paper proposes an attention-guided feature fusion (AGFF) module, which can not only suppress low-level irrelevant background noise but also retain more detailed local semantic information to separate coronary arteries from veins and noise. Figure 2 In order to further enhance feature learning and network training, a scale-aware feature enhancement module (SAFE) is proposed to effectively extract hidden multi-scale contextual information and aggregate multi-scale features. Figure 3 The present invention proposes a multi-scale feature aggregation module (MSFA) module that can refine the vascular map by gradually aggregating adjacent scale features to obtain a more effective vascular semantic representation. Figure 4 .
[0055] Attention-guided feature fusion module, in coronary artery segmentation, due to the complex anatomy and morphology of coronary arteries, high-level features from the decoder and low-level features from the encoder (ref. Figure 1 ). Semantically rich high-level features and spatially rich low-level features are inherently complementary. However, most existing U-Net-based methods directly connect features of different scales. This simple connection is not sufficient to consider the complementarity between high-level and low-level features.
[0056] This paper proposes an attention guided feature fusion (AGFF) module that automatically utilizes the most useful features between two adjacent layers, i.e., high-level features and low-level features. Figure 2 As shown in the figure, by calculating the weight vector α to refine the low-level features and suppress the interference of irrelevant background noise, more important semantic context information can be retained to achieve more accurate positioning. The high-level features of the present invention are The low-level feature is F l (i) , i∈{1, 2, 3, 4}, such as Figure 1 As shown, the high-level features and low-level features F l (i) Send it to the attention-guided feature fusion module to obtain the fusion feature F (i) Specifically, we first use channel concatenation to transform the low-level features F l (i) and advanced features Combined together, they are shown in Equation 1.
[0057]
[0058] in is the splicing operation; f u Represents an upsampling operation.
[0059] Then, channel attention is introduced to obtain the weight vector α, which enables automatic refinement of low-level features. This method not only uses average pooling to stimulate feature information, but also uses maximum pooling to retain more feature information. A shared multi-layer perception layer (MLP) is then used to perform two pooling operations. The results are aggregated and input into a sigmoid function to obtain the weight vector α, as shown in Equation 2.
[0060]
[0061] where f mlp represents the MLP operator; f gap represents the global average pooling; f gmp represents the global maximum pool; f σ Represents the sigmoid activation function.
[0062] Finally, the refined low-level features are added to the high-level features to produce the final result F (i) As shown in formula 3.
[0063]
[0064] in Represents element-level multiplication; in the network framework of the present invention, each layer has a corresponding AGFF module, as shown in Figure 1.
[0065] Scale-aware feature enhancement module: Due to the sparse distribution of coronary arteries, vascular fragmentation in the results is an inevitable problem in the coronary artery segmentation task. Therefore, it is crucial to improve the feature representation capability of coronary artery segmentation. The present invention proposes a scale-aware feature enhancement (SAFE) module to improve its representation capability. The SAFE module can automatically select an appropriate receiving area for feature mapping. The SAFE module consists of two parts: a multi-scale feature extraction unit and an adaptive feature aggregation unit. The configuration of the SAFE module is as follows: Figure 3 As shown; for the multi-scale feature extraction part, the feature map generated by the fifth stage E5 in the encoder is first divided into 4 parallel feature groups, see Figure 1 and Figure 3. Then, dilated convolutional layers with different dilation rates are applied to the parallel branches to obtain richer feature maps with different receptive fields.
[0066] The multi-scale feature extraction unit derives richer feature maps to handle scale variations between different image instances. Specifically, first, the feature map generated by the fifth stage E5 in the encoder (see Figure 1 ) is evenly divided into 4 parallel feature groups f i , i∈{1, 2, 3, 4}, each group f i The feature size is the same as the input feature, but the number of channels is one-fourth of the original feature. The visual perception process of the parallel feature of the present invention is described as follows:
[0067]
[0068] Among them, f i ′ is the result of visual perception of one of the four parallel feature groups, Indicates r i The dilation rate expands the convolution layer to f i feature maps, thereby obtaining richer feature maps with multiple receptive fields; f σ represents the soft attention implemented by the sigmoid activation function. Finally, the hierarchical features F are concatenated and delivered to the adaptive feature aggregation unit:
[0069]
[0070] in For splicing operation.
[0071] In order to make full use of the useful features of the coronary artery along the sagittal view (x-axis), coronal view (y-axis) and axial view (z-axis), first, the hierarchical features F are fed into three convolutional layers through batch normalization and ReLU function, where the convolution kernel sizes are 3×1×1, 1×3×1 and 1×1×3 respectively, generating three feature maps. (where C, H, W, and D represent the channel, height, width, and depth of the input feature F, respectively.) Then, the present invention performs matrix transposition on the feature map Q to obtain Q T , Q T After matrix multiplication with the feature map K, it is input into the softmax activation layer to obtain the encoding of the feature relationship between the sagittal and coronal planes. Finally, the encoding result is matrix multiplied with the feature map V and then residually connected with the hierarchical feature F to obtain the pixel-level attention enhancement feature. The output of the SAFE module is expressed as:
[0072]
[0073] in Represents matrix multiplication. The feature representation capability of the network is enhanced by adding the SAFE module after the fifth encoder stage E5.
[0074] Aggregating multi-scale features is an effective way to improve performance. In order to obtain more semantic details to refine the segmentation results, this paper further designs a multi-scale feature aggregation (MSFA) module. Figure 4 The MSFA module obtains more effective coronary artery prediction by gradually aggregating adjacent scale features in the decoding path. In other words, for each scale feature aggregation, the features of the previous scale and the current scale are adaptively fused. Specifically, D i The resulting input feature channels are compressed to 8, as shown in Equation 7. Features at adjacent scales are then concatenated channel-wise from higher to lower levels. An attention mechanism is then used to refine features at different scales, making the feature representation more discriminative.
[0075] F i =f c (D i ), i∈{1, 2, 3, 4} (7)
[0076] f c It is a convolution operation with a convolution kernel of 1×1×1, and then the feature F between two adjacent scales is i, i∈{1, 2, 3, 4} is used to perform channel splicing in the direction from high stage to low stage, and the soft attention map is used to guide the effective expression of vascular features at different scales; the high-level semantic information of small-scale features is further transferred to large-scale features, see formulas (8), (9), and (10). The "large scale" and "small scale" mentioned in this invention are two relative scales, and do not represent specific large scales and small scales.
[0077]
[0078]
[0079]
[0080] in For splicing operation; represents matrix multiplication; f u represents the upsampling operation; f σ represents the sigmoid activation function, obtaining the attention map of the current scale; f c″ The convolution operation with a kernel of 3×3×3 is used to generate rich and detailed features, which are passed to the next attention block until the final prediction of the category. Finally, the output of the MSFA module is obtained by passing it through a convolution layer with two channels and then through a Sigmoid activation function, as shown in Equation 11.
[0081]
[0082] By utilizing the MSFA module to gradually guide the aggregation features between adjacent scale features, the proposed CAS-Net can predict coronary arteries more effectively.
[0083] The labels of the coronary artery region are sparse; the present invention selects the weighted cross entropy (WCE) loss L WCE As a loss function, the loss function can evaluate the learning deviation between blood vessels and background during training. In addition, the Dice Similarity Coefficient (DSC) loss L is introduced DSC To ensure the segmentation of small blood vessels. Finally, the three-dimensional optimization loss function of the proposed CAS-Net training is defined as:
[0084]
[0085] Where α is L WCE and L DSC The weight balance parameter between them is empirically set to α = 0.6. For the binary segmentation task, WCE loss and DSC loss are defined:
[0086]
[0087]
[0088] ω is the weight of the vascular structure and can be used to estimate the probability p of all pixels. i get:
[0089]
[0090] N represents the number of pixels, p i ∈[0, 1] and g i ∈0, 1 represent the predicted probability and background true value i respectively th The parameter ∈ Laplace smoothing factor is used to avoid numerical instability and accelerate the convergence of the training process (∈=1.0 in the present invention).
[0091] In practice, the present invention first quantitatively evaluates coronary artery segmentation results using three classic performance metrics: the dice similarity coefficient (DSC), recall, and precision. The DSC score represents the overlap between actual vessels and identified vessels; recall (also known as sensitivity) measures the proportion of actual vessels that are correctly identified as vessels; and precision measures the proportion of actual vessels that are correctly identified as vessels. These three values range from 0 to 1. Higher values indicate better segmentation accuracy, calculated using the following formulas:
[0092]
[0093]
[0094]
[0095] TP and FP stand for true positive and false positive, respectively, representing the number of vessel pixels correctly segmented by the model and the number of background pixels incorrectly segmented. In addition, FN stands for false negative, which represents vessel pixels that are incorrectly labeled as background pixels.
[0096] Then, in order to prove that the model described in the present invention can learn more vascular features from sparsely labeled annotations and has better recognition ability for non-vascular patterns, the following indicators are followed and adopted: over-segmentation rate (OR) and under-segmentation rate (UR) to evaluate the model. The smaller the OR and UR values, the better the performance of the method. The dice similarity coefficient (DSC) is a very important evaluation criterion in semantic segmentation. Therefore, when quantitatively evaluating the performance of the network, the present invention gives DSC more weight, followed by recall rate, precision, over-segmentation rate (OR) and under-segmentation rate (UR). The implementation results are the average and standard deviation on the test set.
[0097] In addition, the p-value of the metric (DSC) of the proposed method and the compared method on each dataset was calculated for statistical analysis. The smaller the p-value, the more significant the difference between the test methods, and p < 0.05 is considered statistically significant.
[0098] To reduce irrelevant details and increase image contrast, we normalized the training and test data. First, we truncated the intensities of all pixels to a specific range. For example, for the CCTA-119 dataset, the range was [-250, 450]. We performed five cross-validations on all datasets to obtain fair and reliable performance across different methods. The average performance across all evaluation criteria is reported.
[0099] The CAS-Net proposed in this paper is implemented in the PyTorch framework using an NVIDIA 3090Ti GPU (24GB). The optimizer used in all comparative experiments is Adam, with a poly learning strategy and an initial learning rate of 1×10 -4 , weight decay is 5×10 -4 In addition, set the batch size to 2, when the learning rate is lower than 10 -8 Training is stopped when the training exceeds 600 epochs. For model training, subcubes of size 128×256×256 are randomly cropped from the training input data. Online data augmentation is used to enlarge the training dataset (e.g., rotation, flipping, Gaussian blurring). Comparison is made with state-of-the-art methods.
[0100] The proposed method is compared with other methods, such as U-Net3D ( et al., 2016), V-Net (Milletari et al., 2016), HighResNet3D (Li et al., 2017), DenseVoxelNet (Yu et al., 2017), CS 2 Net (Mou et al., 2021), and Di-Vnet (Dong et al., 2021). Most of the existing segmentation methods are based on U-Net, which is the first network for biomedical image segmentation. V-Net is designed to solve the problem of 3D volume segmentation and is applied to many CT image data segmentation tasks. In addition, DenseVoxelNet, CS 2 Net and Di-Vnet are specially designed for blood vessel segmentation.
[0101] Table 1 quantitatively demonstrates the performance of CAS-Net and six comparison methods on the CCTA-119 dataset, including three general-purpose segmentation networks and three networks specifically designed for vessel segmentation. All evaluation criteria in Table 1 were averaged using five-fold cross-validation. As can be seen, the proposed CAS-Net achieves the best performance, achieving 91.76%, 97.26%, and 92.66% in DSC, recall, and precision, respectively, outperforming other state-of-the-art methods. Compared to the traditional U-Net3D, DSC improves from 87.75% to 91.76%. Recall and precision also improve from 95.49% / 91.48% to 97.26% / 92.66%, respectively, validating the effectiveness of the proposed CAS-Net. Furthermore, the OR metric in Table 1 indicates that the Dense Voxel Net and the proposed method have the highest similarity to the ground truth of the predicted CA, with over-segmentation rates of only 1.76% and 1.72%, respectively. However, the larger UR achieved by DenseVoxelNet indicates that more labeled coronary arteries were not segmented into vessels. Overall, this method has better reliability for coronary artery segmentation in terms of DSC, Recall, Precision, OR, and UR. In addition, the p-value of the DSC metric for our method is less than 0.001 compared to the selected methods, indicating that our method significantly outperforms the other methods in segmentation performance.
[0102] The qualitative experimental results of this method and other methods on typical cases are as follows Figure 5 As shown in Figure 2, after 3D morphological closing and maximum connected region preservation, the blood vessel surface becomes smooth and some noise blocks are removed. The results are qualitatively compared using the 3D Slicer toolbox and enlarged patches. The 3D visualization of the segmentation results is shown in Figure 2. Figure 5 The enlarged segments marked in the results are shown in the figure below. As can be seen, most methods accurately generate the ascending aorta of the coronary arteries. However, for coronary artery branches, other methods often produce segments that do not belong to the coronary arteries, particularly in the areas represented by solid circles in the second and third rows of the ground truth. Alternatively, they lose detailed information about the coronary arteries, particularly in the areas represented by dashed circles in the second and third rows of the ground truth. The above qualitative analysis demonstrates that these modules guided by multi-attention mechanisms help refine and fuse complementary information between multi-scale features, thereby achieving better coronary artery image segmentation.
[0103] Table 1
[0104]
[0105]
[0106] Table 1 quantitatively compares the method described in this paper with other methods on the CCTA-119 dataset (mean ± standard deviation). The best results are shown in bold. "↑" and "↓" indicate that larger values have better performance, while "↓" indicates that smaller values have better performance. Figure 5 Comparison of the results of different methods. The dashed and solid circles in the second and third rows of the ground truth values indicate that the compared methods may have under-segmentation and over-segmentation problems, respectively.
[0107] In addition, the present invention can also provide a computer device, including a processor and a memory, the memory is used to store a computer executable program, the processor reads part or all of the computer executable program from the memory and executes it, and when the processor executes part or all of the computer executable program, it can implement the vascular image segmentation method based on the three-dimensional deep network described in the present invention.
[0108] The computer device may be a laptop computer, a desktop computer or a workstation.
[0109] The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or an off-the-shelf field programmable gate array (FPGA).
[0110] A computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is executed by a processor, the blood vessel image segmentation method based on a three-dimensional deep network of the present invention can be implemented.
[0111] The memory of the present invention may be an internal storage unit of a laptop computer, desktop computer or workstation, such as a memory or a hard disk; or an external storage unit, such as a mobile hard disk or a flash memory card.
[0112] Computer-readable storage media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSD) or optical disks, etc. Among them, random access memory may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).
[0113] In summary, this paper proposes a multi-attention-guided deep network capable of generating scale-aware attention features. This network efficiently selects useful information from multi-scale features through a multi-attention mechanism, fully promoting the fusion of features at different levels and obtaining more semantic representations. Three core modules are proposed: the AGFF module, the SAFE module, and the MSFA module. The AGFF module is designed to fuse features from adjacent layers while adaptively suppressing background noise. The SAFE module, placed at the bottom of the network, effectively extracts multi-scale contextual information by implicitly and dynamically adjusting the receptive field of the feature map. Furthermore, the MSFA module is used to learn more semantic representations to improve vascular segmentation images. Compared to other state-of-the-art methods, the proposed method achieves superior performance on coronary artery image segmentation using a collected CCTA dataset. Extensive experimental results demonstrate the promising potential of this network in solving a range of competitive organ vascular image feature segmentation tasks.
[0114] The above description is only a preferred specific implementation method of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, can make equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A vascular image segmentation method based on a three-dimensional deep network, characterized in that: A multi-attention, multi-scale three-dimensional deep network is based on a U-shaped encoder-decoder. The three-dimensional deep network integrates an attention-guided feature fusion module, a scale-aware feature enhancement module, and a multi-scale feature aggregation module into the three-dimensional deep network of the U-shaped encoder-decoder. The attention-guided feature fusion module replaces the traditional skip connection between the encoding and decoding stages, the scale-aware feature enhancement module is embedded in the bottom of the network, and the multi-scale feature aggregation module is integrated into the decoding path. The method comprises the following steps: The encoder extracts low-level features from the vascular image The decoder extracts high-level features The feature map generated in the last stage of the encoder is split into four parallel feature groups. Four dilated convolutions and soft attention mechanisms with four different dilation rates are deployed in parallel branches. The four branch features are spliced using channels to obtain hierarchical features F. The hierarchical features F are input into three convolutional layers respectively to generate three feature maps Q, K and V. The matrix transposition of the feature map Q is performed to obtain Q. T , Q T After matrix multiplication with the feature map K, it is input into the softmax activation layer to obtain the encoding of the feature relationship between the sagittal and coronal positions. Finally, the encoding result is matrix multiplied with the feature map V and then connected with the residual of the hierarchical feature F to obtain the pixel-level attention enhancement feature; Input the pixel-level attention-enhanced features and the features generated by the fourth stage E4 in the encoder into the attention-guided feature fusion module; The attention-guided feature fusion module uses channel splicing to combine low-level features and advanced features Combined, the low-level features are refined by calculating the weight vector α, and the refined low-level features are added to the high-level features to obtain the fusion feature F (i) ; Fusion feature F (i) As the input of the i-th stage in the decoder, i = 1, 2, 3, 4; The convolutional layer is used to compress the features generated by the four stages in the decoder. Then, the features of the middle and high stages of adjacent scales are upsampled and concatenated with the features of the low stage. The concatenated features are then subjected to a soft attention guidance operation. The soft attention-guided features are input into the convolutional layers of the two channels. Finally, the sigmoid activation function is used to calculate the aggregation features between the features of adjacent scales, i.e., the vascular image segmentation results.
2. The vascular image segmentation method based on a three-dimensional deep network according to claim 1, characterized in that: The low-level features are stitched together using channels and advanced features The specific combination is as follows: in is the splicing operation; f u Represents an upsampling operation.
3. The vascular image segmentation method based on a three-dimensional deep network according to claim 1, characterized in that: Use average pooling information to stimulate feature channel information, and use maximum pooling features to retain information; then use a shared multi-layer perception layer to perform two pooling operations, input the obtained results into the Sigmoid function, and obtain the weight vector α: where f mlp represents the MLP operator; f gap represents global average pooling; f gmp represents the global maximum pool; f σ Represents the sigmoid activation function.
4. The vascular image segmentation method based on a three-dimensional deep network according to claim 1, characterized in that: By gradually aggregating adjacent scale features in the decoding path to obtain more global semantic representations of blood vessels, the vascular map can be refined. Specifically, the high-level semantic information of small-scale features is transferred to large-scale features: Among them For splicing operation; represents matrix multiplication; f u represents the upsampling operation; f σ Represents the sigmoid activation function, f c″ It is a convolution operation with a convolution kernel of 3×3×3.
5. The vascular image segmentation method based on a three-dimensional deep network according to claim 1, characterized in that: Weighted cross entropy loss L WCE As the loss function, the dice similarity coefficient loss L is introduced into it DSC , Where α is L WCE and L DSC The weight balance parameter between is determined based on experience. ω is the weight of the vascular structure and can be used to estimate the probability p of all pixels. i get: N represents the number of pixels, p i ∈[0, 1] and g i ∈0, 1 represent the predicted probability and background true value i respectively th The vascular structure of the pixel, the Laplace smoothing factor ∈ is used to avoid numerical instability problems and accelerate the convergence of the training process.
6. The blood vessel image segmentation method based on a three-dimensional deep network according to claim 1, characterized in that: In the multi-scale feature aggregation module, a 1×1×1 convolution is first used to uniformly compress the number of channels of the D4, D3, D2, and D1 feature maps to 8 to reduce the number of parameters in subsequent operations. The features after channel compression in the four stages are recorded as F4, F3, F2, and F1 respectively; then, feature aggregation operations are performed on adjacent scales. The feature aggregation operation performed on adjacent scales is as follows: first, the high-level feature F4 is upsampled and then channel-joined with the low-level feature F3. Then, a 1×1×1 convolution and a sigmoid activation function are used to obtain the attention map of the adjacent scale features and perform matrix multiplication with the low-level feature F3. Then, a 3×3×3 convolution is used to generate information-rich and detailed features. right After upsampling, it is passed to the next scale feature F2, and the feature aggregation operation is performed with the next scale feature F2 to obtain After upsampling, feature aggregation operation is performed with feature F1 to obtain Finally, a 1×1×1 convolution with a channel number of 2 and a sigmoid activation function are used to obtain the output of the multi-scale feature aggregation module.
7. A vascular image segmentation system based on a three-dimensional deep network, characterized in that: A multi-attention, multi-scale 3D deep network based on a U-shaped encoder-decoder, including an attention-guided feature fusion module, a scale-aware feature enhancement module, and a multi-scale feature aggregation module; The attention-guided feature fusion module, the scale-aware feature enhancement module, and the multi-scale feature aggregation module are integrated into a three-dimensional deep network of a U-shaped encoder-decoder. The attention-guided feature fusion module replaces the traditional skip connections between the encoding and decoding stages. The scale-aware feature enhancement module is embedded at the bottom of the network, and the multi-scale feature aggregation module is integrated into the decoding path. The encoder extracts low-level features from the vascular image The decoder extracts high-level features The scale-aware feature enhancement module is used to split the feature map generated in the last stage of the encoder into four parallel feature groups, deploy four dilated convolutions and soft attention mechanisms with four different dilation rates in parallel branches, and use channels to splice the four branch features to obtain hierarchical features F; the hierarchical features F are input into three convolutional layers respectively to generate three feature maps Q, K and V, and matrix transposition is performed on the feature map Q to obtain Q T , Q T After matrix multiplication with the feature map K, it is input into the softmax activation layer to obtain the encoding of the feature relationship between the sagittal and coronal positions. Finally, the encoding result is matrix multiplied with the feature map V and then connected with the residual of the hierarchical feature F to obtain the pixel-level attention enhancement feature; The attention-guided feature fusion module is used to combine low-level features into and advanced features Combined, the low-level features are refined by calculating the weight vector α, and the refined low-level features are added to the high-level features to obtain the fusion feature F (i) ; Fusion feature F (i) As the input of the i-th stage in the decoder, i = 1, 2, 3, 4; The multi-scale feature aggregation module uses convolutional layers to perform channel compression on the features generated at each stage in the decoder. It then upsamples the features of the middle and high stages at adjacent scales and concatenates them with the features of the low stages. Soft attention is then performed on the concatenated features, and the obtained soft attention-guided features are input into the convolutional layers of the two channels. Finally, the sigmoid activation function is used to calculate the aggregated features between the features of adjacent scales, i.e., the vascular image segmentation results.
8. A computer device, characterized in that: The invention comprises a processor and a memory, wherein an executable program is stored in the memory, and when the processor executes the executable program, the blood vessel image segmentation method based on a three-dimensional deep network according to any one of claims 1 to 6 can be executed.
9. A computer-readable storage medium, characterized in that A computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the blood vessel image segmentation method based on the three-dimensional deep network according to any one of claims 1 to 6 can be implemented.
Citation Information
Patent Citations
Polyp segmentation method combining attention U-shaped network and multi-scale feature fusion
CN114820635A
Computer Vision Systems and Methods for Detecting and Aligning Land Property Boundaries on Aerial Imagery
US20220156493A1