Double-cascade-stage retinal blood vessel image segmentation method
Through the retinal vascular image segmentation method in the bicascade stage, the deformable convolution and self-attention mechanism are used to improve the segmentation ability of the tiny blood vessels, solve the segmentation problem in the existing technology, and achieve more efficient and accurate retinal vascular segmentation.
Patent Information
- Application Number
- CN202510524540.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing retinal vascular segmentation methods are difficult to effectively segment small blood vessels, are easily disturbed by pathological artifacts, and are difficult to adapt to the multi-scale characteristics of blood vessels.
The retinal vascular image segmentation method in the bicascade stage is adopted, and the weight of the convolution kernel is adjusted through deformable convolution, combined with the self-attention mechanism and Haar wavelet transformation, the ability to extract the features of the tiny blood vessels is improved, and segmentation accuracy and operation efficiency are enhanced.
It improves the segmentation ability of tiny blood vessels, enhances the accuracy and operation efficiency of retinal blood vessel segmentation, reduces information loss and noise interference, and improves the perception and decision-making capabilities of the model.
Smart Images

Figure CN120495311A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning medical image segmentation, and specifically relates to a dual-cascade stage retinal vascular image segmentation method. Background Art
[0002] Retinal blood vessels are the only directly observable blood vessels in the human body, and their morphological changes aid in the assessment of ocular and cardiovascular diseases. The structure of retinal blood vessels is complex and varied, consisting not only of main vessels of varying thickness but also of a dense network of branches. The contrast between retinal blood vessels and surrounding background tissue is low, especially for small vessels, which are often difficult to distinguish clearly from their surroundings, posing a significant challenge for accurate segmentation.
[0003] Traditional retinal vessel segmentation methods primarily rely on manually designed features, including filtering-based and statistical techniques. These methods typically highlight vascular features through grayscale conversion, contrast enhancement, and filtering, and then combine threshold segmentation to extract the vascular network. While these methods offer advantages such as high computational efficiency and strong interpretability, they are less effective at detecting small vessels, are susceptible to interference from pathological artifacts, and struggle to adapt to the multi-scale characteristics of vessels. Summary of the Invention
[0004] The technical problem of the present invention is: by replacing the traditional convolution layer in U-Net with deformable convolution, the network can adaptively adjust the weight of the convolution kernel during the learning process, improve the ability to extract small blood vessel features, enhance the accuracy and operating efficiency of retinal blood vessel segmentation, and improve the segmentation ability of small blood vessels.
[0005] The purpose of the present invention is to solve the above problems and propose a dual-cascade stage retinal vessel image segmentation method, comprising the following steps: S1: Construct a retinal vascular image dataset and perform preprocessing; S2: Constructing a dual-cascade stage retinal segmentation model; S3: Input the test set in the dataset into the retinal segmentation model for training; S4: Input the retinal blood vessel image into the network to segment the blood vessels.
[0006] Furthermore, in step S1, the image dataset includes images collected by a high-resolution fundus camera and annotated with vascular regions after manual annotation and expert review.
[0007] Preferably, in step S1, the data set is preprocessed, including converting the original RGB original image into a grayscale image, and extracting image blocks from the processed grayscale image using a sliding window with a step size of 6 and a size of 48×48.
[0008] Furthermore, in step S2, the retinal segmentation model of the dual-cascade stage includes a convolution module, a first cascade module V1, a second cascade module V2, an upsampling module and a skip connection module.
[0009] Preferably, the retinal vascular image is subjected to feature extraction through a convolution module to generate a preliminary feature map, and sequentially passes through the first cascade module V1 and the second cascade module V2 to obtain feature map P1 and feature map P2 respectively, and then generates the final vascular segmentation result through an upsampling module, a jump connection module and a convolution operation.
[0010] Furthermore, the first cascade module V1 and the second cascade module V2 adopt the same structure, including a convolution module, a feature refinement module and a wavelet downsampling module; the feature map is processed through the convolution module, including the following steps: Conv3×3→Bn→LeakyRelu→Conv3×3→Bn→LeakyRelu; Conv3×3 represents a 3×3 convolution operation, Bn represents a batch normalization operation, and LeakyRelu represents an activation function with a negative slope of 0.1.
[0011] Furthermore, the feature refinement module includes the following steps: 1) Capture global context information through the self-attention mechanism to obtain the feature map C1; 2) Divide the feature map C1 into four equal parts in the channel dimension, perform feature extraction on three of the feature maps through 3×3 convolution to capture local context information, and fuse global and local information with the remaining feature maps through 1×1 convolution; 3) The fused feature map is divided into two parts in the channel dimension. The first part is enhanced with local correlation through 3×3 convolution, and the second part is transposed. The first and second parts are then multiplied pixel by pixel and integrated using 1×1 convolution. 4) Perform residual fusion on the feature map C1 and the integrated feature map, and perform normalization and ReLU activation operations on the fusion result.
[0012] Preferably, the wavelet downsampling module uses Haar wavelet transform to decompose image information, and the calculation formula is: ; ; ; Where, represents the low-frequency component, represents the high-frequency component, x represents the two-dimensional image signal, , j represents the scale, and k represents the translation parameter.
[0013] Preferably, the wavelet downsampling processing step includes the following steps: 1) Perform one-dimensional discrete wavelet transform on the horizontal direction of the two-dimensional image signal to obtain high-frequency and low-frequency components; 2) Perform one-dimensional discrete wavelet transform on the vertical direction of the two-dimensional image signal and decompose it into horizontal high-frequency variables, vertical high-frequency components, diagonal high-frequency components and low-frequency components.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1) The dual-cascade stage retinal vessel image segmentation method proposed in this invention improves the ability to extract small vessel features, enhances the accuracy and efficiency of retinal vessel segmentation, and improves the ability to segment small vessels.
[0015] 2) The feature refinement module proposed in this paper refines the original features and solves information loss, noise interference or feature level mismatch during the extraction process, thereby enhancing the model's perception and decision-making capabilities for the target task, while improving the discriminative power and robustness of the features.
[0016] 3) The downsampling of wavelet decomposition proposed in the present invention encodes the information of the spatial dimension into the channel dimension through wavelet transform, thereby avoiding the loss of spatial dimension information during the downsampling process and effectively preserving the subtle but important local details in the blood vessels. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention will be further described below with reference to the accompanying drawings and examples.
[0018] Figure 1 Flowchart of a dual-cascade retinal vessel image segmentation method according to an embodiment of the present invention.
[0019] Figure 2 This is a diagram of a dual-cascade stage network structure of an embodiment of the present invention.
[0020] Figure 3 This is a structural diagram of the feature refinement module of an embodiment of the present invention.
[0021] Figure 4 2 is a comparison diagram of the segmentation effects of an embodiment of the present invention. DETAILED DESCRIPTION
[0022] like Figure 1 As shown, a dual-cascade stage retinal vessel image segmentation method includes the following steps: S1: Construct a retinal vascular image dataset and perform preprocessing.
[0023] In step S1 , the image dataset includes images collected by a high-resolution fundus camera and annotated with vascular regions after manual annotation and expert review.
[0024] In step S1, the dataset is preprocessed, including converting the original RGB original image into a grayscale image, and extracting image blocks from the processed grayscale image using a sliding window with a step size of 6 and a size of 48×48.
[0025] S2: Construct a dual-cascade stage retinal segmentation model.
[0026] In step S2, the retinal segmentation model of the dual-cascade stage includes a convolution module, a first cascade module V1, a second cascade module V2, an upsampling module and a skip connection module.
[0027] The retinal vascular image is subjected to feature extraction through the convolution module to generate a preliminary feature map, which is then passed through the first cascade module V1 and the second cascade module V2 in sequence to obtain feature maps P1 and P2 respectively. The final vascular segmentation result is then generated through the upsampling module, the skip connection module and the convolution operation.
[0028] like Figure 2 As shown, the retinal segmentation model in the cascade stage processes the retinal blood vessel image in the following steps: 1) Preprocessed retinal vascular image → convolution module → obtain feature map A1; 2) Feature map A1 → wavelet downsampling → first cascade stage convolution module extracts the first layer feature map P11 → wavelet downsampling → convolution module → obtain the second layer feature map P12 → wavelet downsampling → convolution module → feature refinement module → obtain the third layer feature map P13; 3) Perform upsampling operation on feature map P13 and skip connection with feature map P12 → extract convolution module → obtain feature map P14; 4) Perform upsampling operation on feature map P14 and skip connection with feature map P11 → extract through convolution module → obtain feature map P15; 5) Feature map P15 → the second cascade structure convolution module extracts the first layer feature map P21 → wavelet downsampling → convolution module → obtain the second layer feature map P22 → wavelet downsampling → convolution module → feature refinement module → obtain the third layer feature map P23; 6) Perform upsampling operation on feature map P23 and skip connection with feature map P22 → extract convolution module → obtain feature map P24; 7) Perform upsampling operation on feature map P24 and skip connection with feature map P21 → extract through convolution module → obtain feature map P25; 8) Upsample the feature map P25 and skip-connect it with the feature map A1 → extract it through the convolution module → obtain the feature map A2; 9) Feature map A2 → 1×1 convolution → obtain segmentation result; The first cascade module V1 and the second cascade module V2 have the same structure, including a convolution module, a feature refinement module, and a wavelet downsampling module. The feature map is processed through the convolution module, which includes the following steps: Conv3×3→Bn→LeakyRelu→Conv3×3→Bn→LeakyRelu; Conv3×3 represents a 3×3 convolution operation, Bn represents a batch normalization operation, and LeakyRelu represents an activation function with a negative slope of 0.1.
[0029] like Figure 3 As shown, the feature refinement module includes the following steps: 1) Capture global context information through the self-attention mechanism to obtain the feature map C1; 2) Divide the feature map C1 into four equal parts in the channel dimension, perform feature extraction on three of the feature maps through 3×3 convolution to capture local context information, and fuse the global and local information with the remaining feature maps through 1×1 convolution to obtain feature map C2; 3) The fused feature map C2 is divided into two parts in the channel dimension. The first part is enhanced with local correlation through 3×3 convolution, and the second part is transposed. The first and second parts are then multiplied pixel by pixel and integrated using 1×1 convolution to obtain feature map C3. 4) Perform residual fusion on the feature map C1 and the integrated feature map C3, and perform normalization and ReLU activation operations on the fusion result.
[0030] The wavelet downsampling module uses Haar wavelet transform to decompose image information. The calculation formula is: ; ; ; Where, represents the low-frequency component, represents the high-frequency component, x represents the two-dimensional image signal, , j represents the scale, and k represents the translation parameter.
[0031] The wavelet downsampling,processing steps include the following steps: 1) Perform one-dimensional discrete wavelet transform on the horizontal direction of the two-dimensional image signal to obtain high-frequency and low-frequency components; 2) Perform one-dimensional discrete wavelet transform on the vertical direction of the two-dimensional image signal and decompose it into horizontal high-frequency variables, vertical high-frequency components, diagonal high-frequency components and low-frequency components.
[0032] S3: Input the test set in the dataset into the retinal segmentation model for training.
[0033] S4: Input the retinal blood vessel image into the network to segment the blood vessels.
[0034] like Figure 4 As shown, in order to verify the segmentation ability of the present invention, the model of the present invention was compared with the U-Net, DUNet, SN-Unet, and EDEA network models. The five evaluation indicators, AUC, accuracy (Acc), sensitivity (Se), specificity (Sp), and F1 score, were used to measure the segmentation effect. The model of the present invention was applied to the public datasets DRIVE and CHASE_DB1. The DRIVE dataset contains 40 fundus images of 564×584 pixels, of which 33 showed signs of diabetic retinopathy and 7 showed mild diabetic retinopathy. The first 20 images were used as training sets and the last 20 images were used as test sets. The CHASE_DB1 dataset contains 28 fundus images of 999×960 pixels, collected from the left and right eyes of 14 school children, and manually annotated and segmented by two experts. The first 20 images were used as training sets and the last 8 images were used as test sets.
[0035] The present invention is applied to the public dataset DRIVE to segment the image. The segmentation results are shown in Table 1: Table 1
[0036] As can be seen from the results in the table, on the DRIVE dataset, the proposed network achieved 0.9891, 0.8321, 0.9707, 0.8341, and 0.9840 in the five indicators of AUC, F1 score, Acc, Se, and Sp, respectively. Among them, the two key indicators of AUC and Acc achieved the best results. Compared with U-Net, AUC increased by 0.24%, Acc increased by 0.2%, and Se increased by 1.73%. Compared with DUnet and SN-Unet networks, although the number of parameters of this network increased slightly, the two indicators of AUC and F1 increased by 0.1%, 0.19%, and 0.37%, and 1.21%, respectively.
[0037] The present invention is applied to the public dataset CHASE_DB1 to segment images. The segmentation results are shown in Table 2: Table 2
[0038] On the CHASE_DB1 dataset, the proposed network also performed well, achieving optimal performance in three key metrics: AUC, F1, and Acc. Compared to the recently popular EDEA network, the proposed network not only significantly reduced the number of parameters but also improved AUC, F1, and Acc by 0.2%, 2.65%, and 0.53%, respectively. Compared to the DUnet and SN-Unet networks, although the proposed network had slightly more parameters, it improved the F1 metric by 0.8% and 2.14%, respectively.
[0039] In summary, in comparative experiments on two datasets, under the same experimental environment, our method demonstrated significant performance improvements over the commonly used U-Net, DUNet, SN-Unet, and EDEA models. The improvements in AUC and F1 scores, in particular, demonstrate the excellent performance of our method for segmenting small blood vessels and segmenting broken vessels. The experimental results of our network on both datasets demonstrate its superior generalization capabilities.
[0040] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions recited in the claims, including equivalent alternatives to the technical features of the technical solutions recited in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.
Claims
1. A dual-cascade stage retinal vessel image segmentation method, characterized in that: The following steps are involved: S1: Construct a retinal vascular image dataset and perform preprocessing; S2: Constructing a dual-cascade stage retinal segmentation model; S3: Input the test set in the dataset into the retinal segmentation model for training; S4: Input the retinal blood vessel image into the network to segment the blood vessels.
2. The dual-cascade stage retinal vessel image segmentation method according to claim 1, characterized in that: In step S1, the image dataset includes images collected by a high-resolution fundus camera and annotated with vascular regions after manual annotation and expert review.
3. The dual-cascade stage retinal vessel image segmentation method according to claim 2, characterized in that: In step S1, the data set is preprocessed, including converting the original RGB original image into a grayscale image, and extracting image blocks from the processed grayscale image using a sliding window with a step size of 6 and a size of 48×48.
4. The dual-cascade stage retinal vessel image segmentation method according to claim 1, characterized in that: In step S2, the dual-cascade stage retinal segmentation model includes a convolution module, a first cascade module V1, a second cascade module V2, an upsampling module and a skip connection module.
5. The dual-cascade stage retinal vessel image segmentation method according to claim 4, characterized in that: The retinal vascular image is subjected to feature extraction through a convolution module to generate a preliminary feature map, and is sequentially passed through the first cascade module V1 and the second cascade module V2 to obtain feature maps P1 and P2, respectively. The final vascular segmentation result is then generated through an upsampling module, a skip connection module, and a convolution operation.
6. The dual-cascade stage retinal vessel image segmentation method according to claim 5, characterized in that: The first cascade module V1 and the second cascade module V2 have the same structure, including a convolution module, a feature refinement module and a wavelet downsampling module.
7. The dual-cascade stage retinal vessel image segmentation method according to claim 6, characterized in that: The feature map passes through the convolution module, which includes the following steps: Conv3×3→Bn→LeakyRelu→Conv3×3→Bn→LeakyRelu; Conv3×3 represents a 3×3 convolution operation, Bn represents a batch normalization operation, and LeakyRelu represents an activation function with a negative slope of 0.
1.
8. The dual-cascade stage retinal vessel image segmentation method according to claim 7, characterized in that: The feature refinement module includes the following steps: 1) Capture global context information through the self-attention mechanism to obtain the feature map C1; 2) Divide the feature map C1 into four equal parts in the channel dimension, perform feature extraction on three of the feature maps through 3×3 convolution to capture local context information, and fuse global and local information with the remaining feature maps through 1×1 convolution; 3) The fused feature map is divided into two parts in the channel dimension. The first part is enhanced with local correlation through 3×3 convolution, and the second part is transposed. The first and second parts are then multiplied pixel by pixel and integrated using 1×1 convolution. 4) Perform residual fusion on the feature map C1 and the integrated feature map, and perform normalization and ReLU activation operations on the fusion result.
9. The dual-cascade stage retinal vessel image segmentation method according to claim 8, characterized in that: The wavelet downsampling module uses Haar wavelet transform to decompose image information. The calculation formula is: ; ; ; Where, represents the low-frequency component, represents the high-frequency component, x represents the two-dimensional image signal, , j represents the scale, and k represents the translation parameter.
10. The dual-cascade stage retinal vessel image segmentation method according to claim 9, characterized in that: The wavelet downsampling processing step includes the following steps: 1) Perform one-dimensional discrete wavelet transform on the horizontal direction of the two-dimensional image signal to obtain high-frequency and low-frequency components; 2) Perform one-dimensional discrete wavelet transform on the vertical direction of the two-dimensional image signal and decompose it into horizontal high-frequency variables, vertical high-frequency components, diagonal high-frequency components and low-frequency components.