A lightweight retinal vessel segmentation method based on the attention mechanism
By optimizing the U-Net network structure, introducing attention mechanism and cascade design, the problems of high model complexity and overfitting in the existing technology are solved, and efficient retinal vascular segmentation is achieved, and segmentation performance and model robustness are improved.
Patent Information
- Application Number
- CN202211390709.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-11-07
AI Technical Summary
While improving segmentation accuracy, the existing retinal vascular segmentation algorithm has high model complexity and large parameters, resulting in high computational complexity, high hardware requirements, and prone to overfitting problems.
A lightweight retinal vascular segmentation method based on attention mechanism is designed. By optimizing the U-Net network structure, introducing attention mechanism, reducing the number of pooled layers, reducing the number of channels of codec blocks, and using cascade design and improved convolution blocks to reduce network complexity.
While improving segmentation performance, it significantly reduces the network complexity and parameter quantity, improves the robustness and segmentation performance of the model, and avoids overfitting problems.
Smart Images

Figure CN115760872B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image segmentation, and designs a lightweight retinal vessel segmentation method based on an attention mechanism, constructing a lightweight retinal vessel segmentation network while ensuring the segmentation quality. Background Art
[0002] In the retinal vessel segmentation algorithm based on convolutional neural network, most adopt the fully convolutional network (FCN) and U-Net as the benchmark networks. FCN fuses shallow semantic features and deep semantic features to achieve accurate segmentation tasks. U-Net consists of a contracting path that captures context and a symmetric expanding path that supports precise localization. Inspired by the above two classic networks, many improved networks have emerged and been applied to the field of retinal vessel segmentation.
[0003] Wu et al. proposed a cascaded structure network to improve the connectivity of segmented vessels. The forward network converts the input into a coarse vessel segmentation map, and the subsequent network adjusts the pixels with classification errors in the coarse vessel segmentation map, thereby performing a secondary segmentation on the semantic feature map to correct the misclassified pixels and re-optimize the spatial structure of the vessels.
[0004] Lian et al. designed a WUN (Weighted U-Net) module for coarse segmentation and a WRUN (Weighted Res-UNet) module for fine segmentation. By using the image slices of the globally enhanced fundus map as the input of WUN to generate the coarse segmentation vessel map, the locally enhanced image slices, the corresponding ground truth image slices, and the coarse segmentation vessel map of the previous network are used as the combined input of WRUN to train the network. This model can well segment small and tortuous vessels and maintain the geometric connection of retinal vessels.
[0005] The deep learning model proposed by Yan et al. is divided into three stages: thick vessel segmentation, thin vessel segmentation, and vessel feature fusion. Separating and segmenting thick and thin vessels can obtain better discriminative features, minimizing the negative impact brought by the imbalance of the proportion of thick and thin vessels. The vessel fusion stage improves the overall thickness consistency of vessels by further identifying non-vessel pixels and refines the results.
[0006] The above method focuses on improving the segmentation accuracy of retinal blood vessels, but does not consider the complexity of the model. There are certain limitations in the method, making it difficult to be deployed in actual application scenarios. The segmentation performance of some networks is excellent, but they have high computational complexity, a large number of parameters, and high hardware requirements, which hinder their routine clinical applications. In addition, due to the small sample data volume, using a network structure with too many parameters in the retinal blood vessel segmentation task is prone to overfitting problems. At the same time, the original U-Net network has too many layers and the convolutional blocks are simply stacked, which dilutes the semantic information excessively, and the network cannot capture the blood vessel features well, resulting in unsatisfactory segmentation performance. Summary of the Invention
[0007] In view of the above-mentioned defects of the prior art, the present invention proposes a lightweight retinal blood vessel segmentation method based on an attention mechanism. The technical solution designed by the present invention includes the following steps, specifically including:
[0008] S1: Preprocess the original retinal blood vessel image to obtain an enhanced retinal blood vessel image;
[0009] S2: Construct a retinal blood vessel segmentation model based on U-Net and a fully convolutional network. The retinal blood vessel segmentation model includes a front-stage network, a rear-stage network, a spatial grouping enhancement module, and an encoding-decoding block;
[0010] S3: Input the enhanced retinal blood vessel image into the front-stage network to obtain an initial feature map;
[0011] S4: Perform two downsamplings on the initial feature map to obtain two downsampled feature maps, input the two downsampled feature maps into the spatial grouping enhancement module, and output two enhanced feature maps;
[0012] S5: Perform two transposed convolution upsamplings on the downsampled feature map obtained by the second downsampling to obtain an upsampled feature map, perform an element-wise addition operation on the upsampled feature map and the enhanced feature map to obtain a complete feature map;
[0013] S6: Perform a 1×1 convolution operation on the complete feature map to obtain a single-channel feature map, perform a channel splicing operation on the single-channel feature map and the enhanced retinal blood vessel image to obtain a new initial feature map;
[0014] S7: Input the new initial feature map into the rear-stage network to obtain a retinal blood vessel segmentation map;
[0015] Among them, the front-stage network is the U-Net in the front, the rear-stage network is the U-Net in the back, and the encoding-decoding block is the activation function PReLU and the convolutional block.
[0016] Further, the preprocessing of the original retinal blood vessel image in S1 specifically includes:
[0017] S2001: Perform contrast - limited adaptive histogram equalization on the original retinal vessel image to obtain an enhanced retinal vessel image;
[0018] S2002: Crop, rotate, and flip the enhanced retinal vessel image;
[0019] S2003: Divide it into a training set and a test set according to the ratio of 7:3.
[0020] Furthermore, the retinal vessel segmentation model adopts a lightweight design:
[0021] Simplify the U - Net with the initial 5 - layer pooling layer to a U - Net with 3 - layer pooling layer, and reduce the number of channels in each stage of the encoding - decoding block.
[0022] Furthermore, the retinal vessel segmentation model adopts a cascaded design:
[0023] Cascade two U - Nets. Concatenate the initial feature map output by the previous - stage network with the enhanced retinal vessel image in the channel dimension and input them together into the subsequent - stage network. The subsequent - stage network performs secondary segmentation on each pixel based on the initial feature map output by the previous - stage network. Under the auxiliary network and the main supervision network, perform end - to - end two - time learning on the labels of each pixel point of the initial feature map. The loss function is expressed as:
[0024] Loss aux = BCE(PM, GT)
[0025] Loss main = BCE(rPM, GT)
[0026] Loss = Loss aux + Loss main
[0027] In the formula, Loss aux is the auxiliary loss function, Loss main is the main loss function, Loss is the total loss function, BCE is the binary cross - entropy loss function, PM is the initial feature map, rPM is the initial feature map after learning, and GT is the label.
[0028] Furthermore, the retinal vessel segmentation model adopts an improved convolutional block:
[0029] S5001: Use residual learning for the encoding - decoding block, add the input feature map to the convolutional feature map to form an output feature map;
[0030] S5002: Adopt the parametric rectified linear unit PReLU as the activation function of the encoding - decoding block;
[0031] S5003: Add a 1×1 convolutional layer to the skip connection.
[0032] Furthermore, the spatial grouping enhancement module includes:
[0033] A complete feature is composed of many sub-features, and the sub-features are distributed in groups in the features of each layer. A spatial grouping enhancement module is introduced in the skip connection part between the encoding and decoding structures to highlight the semantic features of the blood vessel area. The enhancement mechanism is as follows:
[0034] (1) Group the features according to the channel dimension, and each group is expressed as:
[0035]
[0036] where X i is the sub-feature vector, C is the total number of channels, G is the number of selected groups for grouping, H is the height of the feature map, and W is the width of the feature map;
[0037] (2) Use the global average pooling function to shrink the spatial dimension. After grouping the channels, the approximate network learns the semantic vector representing the blood vessel semantic features in this channel grouping according to the input label to obtain the global semantic feature g:
[0038]
[0039] where F(X) is the global average pooling function; x i is the grouped sub-feature vector;
[0040] (3) Generate corresponding importance coefficients for the local features:
[0041] c i = g · x i
[0042] The importance coefficient is also expressed as:
[0043] ||g||||x i ||cos(θ i )
[0044] where θ i is the angle between the i-th semantic sub-feature and the global semantic feature g, and c i is the importance coefficient;
[0045] (4) To prevent the importance coefficients of different samples from deviating too much, normalize the importance coefficients:
[0046]
[0047]
[0048]
[0049] In the formula, μ c is the mean value, σ c is the variance, ε is a constant, is c i is the normalized importance coefficient;
[0050] (5) Perform an affine transformation and Sigmoid activation on the normalized importance coefficient, and then multiply the transformed importance coefficient to each local feature, so that the relevant features are amplified and the irrelevant features are reduced:
[0051]
[0052]
[0053] In the formula, γ is the scale coefficient, β is the offset coefficient, σ(·) is the Sigmoid function, a′ i is the importance coefficient after the affine transformation, is the enhanced sub-feature vector;
[0054] (6) Grouping after enhancement:
[0055]
[0056] In the formula, is the enhanced feature.
[0057] Beneficial effects: By optimizing the U-Net network structure and introducing the attention mechanism, the present invention improves the segmentation performance while reducing the network complexity, and focuses on solving the problem that it is impossible to effectively balance the network complexity and the segmentation accuracy in the field of retinal vessel segmentation. At the same time, experiments are carried out on the public dataset to verify the robustness of the improved network. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic diagram of the segmentation process of a preferred embodiment of the present invention;
[0059] Figure 2 is a schematic diagram of the original U-Net structure of a preferred embodiment of the present invention;
[0060] Figure 3 is a schematic diagram of the lightweight U-Net structure of a preferred embodiment of the present invention;
[0061] Figure 4 is a schematic diagram of the cascade structure of a preferred embodiment of the present invention;
[0062] Figure 5Schematic diagram of the improved convolutional block ICB according to a preferred embodiment of the present invention;
[0063] Figure 6 Schematic diagram of the spatial grouping enhancement module according to a preferred embodiment of the present invention;
[0064] Figure 7 Schematic diagram of the final segmentation network structure according to a preferred embodiment of the present invention. Detailed implementation manners
[0065] The embodiments of the present invention will be described in detail below. The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0066] The present invention designs a lightweight retinal vessel segmentation method based on the attention mechanism. The segmentation process is as Figure 1 shown, which is mainly divided into two parts: data preprocessing and constructing a retinal vessel segmentation model. The technical solution includes the following steps, specifically including:
[0067] S1: Preprocess the original retinal vessel image to obtain an enhanced retinal vessel image;
[0068] S2: Construct a retinal vessel segmentation model based on U-Net and the fully convolutional network. The retinal vessel segmentation model includes a front-stage network, a rear-stage network, a spatial grouping enhancement module, and an encoder-decoder block;
[0069] S3: Input the enhanced retinal vessel image into the front-stage network to obtain an initial feature map;
[0070] S4: Perform two downsamplings on the initial feature map to obtain two downsampled feature maps. Input the two downsampled feature maps into the spatial grouping enhancement module, and output two enhanced feature maps;
[0071] S5: Perform two deconvolution upsamplings on the downsampled feature map obtained by the second downsampling to obtain an upsampled feature map. Perform a corresponding pixel addition operation on the upsampled feature map and the enhanced feature map to obtain a complete feature map;
[0072] S6: Perform a 1×1 convolution operation on the complete feature map to obtain a single-channel feature map. Perform a channel splicing operation on the single-channel feature map and the enhanced retinal vessel image to obtain a new initial feature map;
[0073] S7: Input the new initial feature map into the rear-stage network to obtain a retinal vessel segmentation map;
[0074] Among them, the front-stage network is the U-Net in the front, the back-stage network is the U-Net in the back, and the encoding and decoding block is the activation function PReLU and the convolutional block.
[0075] Furthermore, the preprocessing of the original retinal vessel image in S1 specifically includes:
[0076] S2001: Since the original retinal vessel image is affected by insufficient illumination or overexposure during the acquisition process, the boundaries of the retinal vessels are often unclear. Preprocessing can solve these interferences to a certain extent. In the present invention, the original retinal vessel image is subjected to contrast-limited adaptive histogram equalization, with clipLimit (clipping threshold) set to 2 and tileGridSize (grid size for pixel equalization) set to 8×8. The equalization processing reduces noise and improves the overall contrast between the vessels and the background, making the morphology, orientation, and quantity of the vessels clearer;
[0077] S2002: The datasets used in the present invention mainly include the low-resolution dataset DRIVE and the high-resolution dataset HRF. The images of the low-resolution dataset and the high-resolution dataset are respectively cropped to 512×512 and 1024×1024, and then rotated and flipped to alleviate the overfitting phenomenon;
[0078] S2003: The processed datasets are respectively divided into a training set and a test set according to a ratio of 7:3 to form a dataset for the retinal vessel segmentation method.
[0079] Furthermore, the retinal vessel segmentation model adopts a lightweight design:
[0080] The schematic diagram of the initial U-Net structure is as Figure 2 shown, and the optimized structure is as Figure 3 shown. The numbers in the boxes represent the number of channels of each layer of the structure.
[0081] First, since the ratio of positive sample pixels to negative sample pixels in the original retinal vessel image is usually 1:9, which is extremely unbalanced, and the vessels are less than 10 pixels in the image, especially the microvessels are only 12 pixels. And the pooling layer will dilute the detailed information, resulting in some tiny capillary structures being difficult to capture. Therefore, it is necessary to reduce the pooling layer to pay more attention to the detailed information. Therefore, the present invention simplifies the original 5-layer U-Net structure to 3 layers, reduces the number of pooling layers, and uses as few pooling layers as possible to retain the boundary information and more context information.
[0082] Secondly, the retinal vessel dataset has a small amount of data, between 20 - 40 images. Therefore, the overall number of features is small. If a deep network is used, overfitting is likely to occur. Therefore, in each stage of the encoding and decoding blocks of the present invention, the number of channels is reduced to better fit the segmentation task and prevent overfitting.
[0083] The above design simplifies the five - layer structure [64 -> 128 -> 256 -> 512 -> 1024] of the initial U - Net to a three - layer structure [8 -> 16 -> 32], significantly reducing the number of parameters from 31.03M to 0.05M. It optimizes the network structure, retains more vascular semantic information, and improves the AUC, F1, and SE metrics by 0.22%, 0.41%, and 6.6% respectively, greatly enhancing the comprehensive performance of the segmentation network.
[0084] Furthermore, the retinal vessel segmentation model adopts a cascaded design:
[0085] Pixels predicted as vessels or background by the network may be misclassified because the prediction probability is less than the threshold. To solve this problem, the model of the present invention adopts a cascaded design. By cascading two U - Nets, the initial feature map output by the previous - stage network is concatenated with the enhanced retinal vessel image in the channel dimension and then input into the subsequent - stage network. The subsequent - stage network can perform secondary segmentation on each pixel based on the initial feature map provided by the previous - stage network. The cascaded design is based on the principle of processing from coarse to fine and gradual abstraction, allowing end - to - end two - time learning of the labels of each pixel point under the auxiliary network and the main supervised network. It is expressed by the loss function as:
[0086] Loss aux = BCE(PM, GT)
[0087] Loss main = BCE(rPM, GT)
[0088] Loss = Loss aux + Loss main
[0089] In the formula, Loss aux is the auxiliary loss function, Loss main is the main loss function, Loss is the total loss function, BCE is the binary cross - entropy loss function, PM is the initial feature map, rPM is the initial feature map after learning, and GT is the label. The subsequent - stage network can inherit the learning experience of the previous - stage network, thus accelerating the training process and effectively solving the problem of data imbalance. It also corrects some misclassified pixels (especially pixels with probability values close to the threshold), improving the segmentation performance. The cascaded network structure is as Figure 4 shown.
[0090] Furthermore, the retinal vessel segmentation model adopts an improved convolutional block:
[0091] In a convolutional neural network, simply stacking convolutional blocks will result in the loss of feature information, causing problems such as vanishing gradients and exploding gradients. To solve these problems, the present invention uses residual learning in the encoder-decoder block, adding the input feature map to the convolutional feature map to form the output feature map. Through this design, the network enhances the feature transmission ability and the ability to extract retinal vessel features, and more completely preserves the vascular semantic information.
[0092] Due to the small number of sample images in the retinal segmentation task, the network is prone to overfitting. An appropriate activation function should be selected to alleviate this problem. The traditional U-Net uses the rectified linear unit (ReLU) as the activation function, which can effectively alleviate the problems of overfitting and vanishing gradients. However, when the input is negative, the output of ReLU is zero, which means that these neurons will become inactivated. To solve the above problems, this paper uses the parametric rectified linear unit (PReLU) as the activation function of the encoder-decoder block. PReLU adds a negative response to ReLU, avoiding the disadvantage of neuron inactivation when the input is negative. At the same time, the activation function can learn adaptive rectification parameters, improving the learning efficiency of the network in the retinal segmentation task.
[0093] Through the above improvements, the encoder-decoder convolutional block in the original U-Net is improved to Figure 5 the improved convolutional block (ICB) shown in the figure. The ICB contains two standard 3×3 convolutional layers, two batch normalization layers, and two PReLU layers. To match the feature dimensions of the input feature layer and the output feature layer, a 1×1 convolutional layer is added in the skip connection.
[0094] Furthermore, the spatial group enhancement module includes:
[0095] A complete feature is composed of many sub-features, and the sub-features are distributed in the features of each layer in groups. To make each sub-feature spatially robust and well-distributed, and to improve the utilization rate of the feature maps in the middle layer of the network, the present invention introduces a group spatial enhancement module in the skip connection part between the encoder-decoder structures, as Figure 6 shown in the figure. It uses the similarity between the global statistical features and the local position features to establish a spatial enhancement mechanism within each feature group to highlight the semantic features of the vascular region. The specific enhancement mechanism is as follows:
[0096] (1) Group the features according to the channel dimension, and each group is represented as:
[0097]
[0098] Wherein, X i is the sub-feature vector, C is the total number of channels, G is the number of selected groups, H is the height of the feature map, and W is the width of the feature map;
[0099] (2) The global average pooling function is used to shrink the spatial dimension. After grouping the channels, the approximate network obtains the global semantic feature g from the semantic vector representing the vascular semantic features learned by the approximate network in this channel group according to the input label:
[0100]
[0101] Wherein, F(X) is the global average pooling function; x i is the grouped sub-feature vector;
[0102] (3) The local features generate corresponding importance coefficients:
[0103] c i = g · x i
[0104] The importance coefficient is also expressed as:
[0105] ||g||||x i ||cos(θ i )
[0106] Wherein, θ i is the included angle between the i-th semantic sub-feature and the global semantic feature g, and c i is the importance coefficient;
[0107] (4) To prevent the importance coefficients of different samples from deviating too much, the importance coefficients are normalized:
[0108]
[0109]
[0110]
[0111] Wherein, μ c is the mean value, σ c is the variance, ε is a constant, is c i after the normalized importance coefficient;
[0112] (5) The normalized importance coefficients are subjected to affine transformation and Sigmoid activation, and then the transformed importance coefficients are multiplied by each local feature, so that the relevant features are amplified and the irrelevant features are reduced:
[0113]
[0114]
[0115] In the formula, γ is the scale coefficient, β is the offset coefficient, σ(·) is the Sigmoid function, and a′ i is the importance coefficient after affine transformation, is the enhanced sub-feature vector;
[0116] (6) Enhanced grouping:
[0117]
[0118] In the formula, is the enhanced feature.
[0119] After the above improvements, the final segmentation network structure is obtained as Figure 7 shown, with a total number of parameters of 0.1M.
[0120] Furthermore, the specific evaluation metrics and results used in the present invention are as follows:
[0121] (1) Randomly select color retinal images from the test set and input them into the trained retinal vessel segmentation model to obtain the test set vessel segmentation result map and the parameter values of the segmentation evaluation metrics. The evaluation metrics include accuracy (Acc), sensitivity (Sen), F1-Score, and the area AUC value under the ROC curve. At the same time, the model complexity of this method is also compared with mainstream algorithms, and the main metrics are the number of parameters (Params) and the number of floating-point operations (FLOPs). The performance of the entire algorithm is evaluated on the DRIVE and HRF datasets used for testing. The method of the present invention obtains competitive results and has higher sensitivity and better segmentation performance compared with current mainstream algorithms.
[0122] (2) The experimental results and comparisons are shown in the following table
[0123] Table 1 Comparison of segmentation performance with different methods on the DRIVE dataset
[0124]
[0125] Table 2 Comparison of segmentation performance with different methods on the HRF dataset
[0126]
[0127] Table 3 Comparison of model complexity with different methods on the DRIVE dataset
[0128]
[0129] As can be seen from the above table, the performance of the entire algorithm of the present invention was evaluated on the DRIVE and HRF datasets during testing. Compared with the current mainstream algorithms, the model has the lowest complexity, higher sensitivity, and better segmentation performance, obtaining competitive results.
[0130] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A lightweight retinal vessel segmentation method based on the attention mechanism, characterized in that Including: S1: Preprocess the original retinal vessel image to obtain an enhanced retinal vessel image; S2: Construct a retinal vessel segmentation model based on U-Net and a fully convolutional network. The retinal vessel segmentation model includes a front-stage network, a rear-stage network, a spatial group enhancement module, and an encoder-decoder block; S3: Input the enhanced retinal vessel image into the front-stage network to obtain an initial feature map; S4: Perform two downsamplings on the initial feature map to obtain two downsampled feature maps. Input the two downsampled feature maps into the spatial group enhancement module, and output to obtain two enhanced feature maps; S5: Perform two transposed convolutional upsamplings on the downsampled feature map obtained from the second downsampling to obtain an upsampled feature map. Perform a corresponding pixel addition operation on the upsampled feature map and the enhanced feature map to obtain a complete feature map; S6: Perform a 1×1 convolution operation on the complete feature map to obtain a single-channel feature map. Perform a channel concatenation operation on the single-channel feature map and the enhanced retinal vessel image to obtain a new initial feature map; S7: Input the new initial feature map into the rear-stage network to obtain a retinal vessel segmentation map; Among them, the front-stage network is the U-Net in the front, the rear-stage network is the U-Net in the back, and the encoder-decoder block is the activation function PReLU and the convolutional block.
2. The lightweight retinal vessel segmentation method based on the attention mechanism according to claim 1, wherein The preprocessing of the original retinal vessel image includes: S2001: Perform contrast-limited adaptive histogram equalization processing on the original retinal vessel image to obtain an enhanced retinal vessel image; S2002: Perform cropping, rotation, and flipping processing on the enhanced retinal vessel image; S2003: Divide it into a training set and a test set according to a ratio of 7:
3.
3. The lightweight retinal vessel segmentation method based on the attention mechanism according to claim 1, characterized in that The retinal vessel segmentation model adopts a lightweight design: Simplify the U-Net with an initial 5-layer pooling layer to a U-Net with 3-layer pooling layers, and reduce the number of channels in each stage of the encoder-decoder block.
4. A lightweight retinal vessel segmentation method based on an attention mechanism according to claim 1, characterized in that, The retinal vessel segmentation model adopts a cascaded design: Cascade two U-Nets. Perform a channel dimension concatenation on the initial feature map output by the front-stage network and the enhanced retinal vessel image, and input them together into the rear-stage network. The rear-stage network performs a secondary segmentation on each pixel based on the initial feature map output by the front-stage network, and performs end-to-end two-learnings on the labels of each pixel point of the initial feature map under the auxiliary network and the main supervision network. The loss function is expressed as: Loss aux = BCE(PM, GT) Loss main = BCE(rPM, GT) Loss=Loss aux +Loss main where Loss aux is the auxiliary loss function, Loss main is the main loss function, Loss is the total loss function, BCE is the binary cross-entropy loss function, PM is the initial feature map, rPM is the learned initial feature map, and GT is the label.
5. A lightweight retinal vessel segmentation method based on the attention mechanism according to claim 1, characterized in that The retinal vessel segmentation model adopts an improved convolutional block: S5001: Use residual learning for the encoder-decoder block, add the input feature map to the convolutional feature map to form an output feature map; S5002: Adopt the parametric rectified linear unit PReLU as the activation function of the encoder-decoder block; S5003: Add a 1×1 convolutional layer in the skip connection.
6. The lightweight retinal vessel segmentation method based on the attention mechanism according to claim 1, characterized in that, The spatial group enhancement module includes: A complete feature is composed of many sub-features. The sub-features are distributed in the features of each layer in groups. Introduce a spatial group enhancement module in the skip connection part between the encoder-decoder structures to highlight the semantic features of the vascular region. The enhancement mechanism is as follows: (1) Group the features according to the channel dimension, and each group is expressed as: where X i is the sub-feature vector, C is the total number of channels, G is the number of selected groups, H is the height of the feature map, and W is the width of the feature map; (2) Use the global average pooling function to shrink the spatial dimension. After channel grouping, the approximate network obtains the global semantic feature g from the semantic vector representing the vascular semantic features learned by the approximate network according to the input label in this channel grouping: where F(X) is the global average pooling function; x i is the grouped sub-feature vector; (3) The local features generate corresponding importance coefficients: c i = g·x i The importance coefficients are also expressed as: ||g||||x i ||cos(θ i ) where θ i is the included angle between the i-th semantic sub-feature and the global semantic feature g, and c i is the importance coefficient; (4) To prevent the importance coefficients of different samples from deviating too much, normalize the importance coefficients: where μ c is the mean, σ c is the variance, ε is a constant, is c i is the normalized importance coefficient; (5) Perform an affine transformation and a Sigmoid activation on the normalized importance coefficients, and then multiply the transformed importance coefficients by each local feature. The relevant features are amplified and the irrelevant features are reduced: where γ is the scale coefficient, β is the offset coefficient, σ(·) is the Sigmoid function, and a′ i is the importance coefficient after the affine transformation, and is the enhanced sub-feature vector; (6) The enhanced grouping: In the formula, is the enhanced feature.
Citation Information
Patent Citations
U-Net + +-based retinal vessel segmentation method and device
CN113205534A
Retinal blood vessel segmentation method and device
CN113793348A