A method, apparatus, and computer device for segmenting retinal vessel images
Through the improved U-Net network, combined with residual pyramid convolution and attention mechanism, the problem of low accuracy in retinal vascular image segmentation is solved, and a higher quality vascular segmentation effect is achieved.
Patent Information
- Application Number
- CN202111490173.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The prior art is difficult to effectively segment the retinal blood vessel images of complex morphology, resulting in low segmentation accuracy and loss of blood vessel contour information.
The improved U-Net network is adopted, combining residual pyramid convolution and attention mechanism, and the different levels of retinal blood vessel images are extracted through the encoder, and feature splicing is used to stitch the features of retinal blood vessels to output the segmentation results of retinal blood vessels.
It improves the segmentation accuracy and integrity of retinal blood vessel images, reduces the interference of background noise, and enhances the segmentation performance of tiny blood vessels.
Smart Images

Figure CN114283158B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and particularly relates to a method, apparatus and computer device for segmenting retinal vessel images. Background Art
[0002] Retinal vessels have various morphological structures with different lengths, widths and angles. Since the blood vessels are complexly distributed, uneven in size and the contrast between the target vessels and the background image is low in fundus images, it is time-consuming and laborious to manually segment retinal vessels, and there is a large subjectivity. Therefore, using computer technology to achieve accurate and efficient segmentation of retinal vessels is of great significance for assisting diagnosis and treatment.
[0003] Many domestic and foreign scholars have studied the segmentation of retinal vessels. Delibasis used a model-based automatic vessel tracking algorithm and introduced a multi-scale filter during the initialization of seed pixels to achieve the segmentation of retinal vessels. Alhussein used morphological filtering to denoise the image, extracted thick vessels and thin vessels enhanced images using Hessian matrices of different scales respectively, and finally used different threshold segmentation methods to segment thin vessels and thick vessels to obtain the final retinal vessel segmentation image. Li extracted the vessel network using the connected tube MPP model, and then applied the pipeline segmentation algorithm to the extended pipeline target area for vessel segmentation.
[0004] With the development of artificial intelligence technology, convolutional neural networks have been widely used in the field of medical image processing. Francia proposed a new method of linking two convolutional neural networks. The second CNN adopted a residual network block design and was added to the information flow from the first module to achieve accurate segmentation of blood vessels. Jin proposed a deformable vessel segmentation network, which segmented blood vessels in an end-to-end manner using the local features of retinal vessels. Zhou used CNN to extract vessel features and used a set of filters to enhance thin blood vessels, reducing the intensity difference between thin blood vessels and thick blood vessels, and finally used a dense CRF to segment blood vessels. U-Net combined with DenseNet network fully utilized the feature information of the output layer and incorporated dilated convolution into the network to enhance the receptive field of the network, and could segment more blood vessels. Li proposed a method for segmenting retinal vessels based on a U-shaped network, which segmented retinal vessels using the advantages of deformable convolution and dual attention modules. Guo proposed using dense blocks to replace the skip connections in the traditional U-shaped network to achieve feature fusion. During the training stage, a generative adversarial network was adopted. The dense U network based on the initial module was used as the generator of the GAN, and a multi-layer neural network was established as the discriminator of the GAN to achieve the segmentation of retinal vessels.
[0005] Although some of the above methods generally perform well in retinal vessel segmentation, due to the complex structure information, large shape differences of retinal vessels, and the influence of lesions, etc., the loss of some vascular contour information will occur, making it difficult to segment the complete retinal vessels. Therefore, there is an urgent need for an image segmentation model suitable for retinal vessel images for image segmentation processing. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention proposes a method, device and computer device for retinal vessel image segmentation, which realizes the complete segmentation of retinal vessel images by improving the U-Net network, that is, by obtaining the retinal vessel image to be segmented, preprocessing and data augmentation of the obtained retinal vessel image; inputting the processed image into the trained improved U-Net network for image recognition and segmentation to obtain the segmented vessel image; the improved U-Net network includes an encoder, a decoder, a residual pyramid convolution and a skip connection part.
[0007] In the first aspect of the present invention, the present invention provides a method for retinal vessel image segmentation, the method comprising:
[0008] Obtain a retinal vessel image and preprocess the retinal vessel image;
[0009] Input the preprocessed retinal vessel image into the trained U-Net network;
[0010] Use each residual pyramid convolution layer and the corresponding pooling layer in the encoder of the U-Net network to extract the convolution features and pooling features of the retinal vessel image at different levels;
[0011] Transfer the convolution features of each layer to the corresponding attention mechanism layer through skip connection, and select the attention features of the target area of the retinal vessel image from the convolution features of each layer;
[0012] Input the pooling features of the last layer into the first residual pyramid convolution layer of the decoder in the U-Net network, and use the upsampling layer to output the sampling features;
[0013] Use each residual pyramid convolution layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampling features of the upsampling layer and the corresponding attention features, and transfer them to the last residual pyramid convolution layer of the decoder to output the obtained feature map;
[0014] Perform a 1×1 convolution on the output obtained feature map, and finally obtain the segmentation result of the retinal vessels.
[0015] In a second aspect of the present invention, the present invention further provides a retinal vascular image segmentation device, which includes:
[0016] An image acquisition module for acquiring retinal vascular images;
[0017] An image processing module for preprocessing the acquired retinal vascular images;
[0018] An image input module for inputting the preprocessed retinal vascular images into a trained U-Net network;
[0019] An encoder module that uses each residual pyramid convolutional layer and the corresponding pooling layer in the encoder of the U-Net network to extract convolutional features and pooling features of the retinal vascular images at different levels;
[0020] A skip connection module for transmitting the convolutional features of each layer to the corresponding attention mechanism layer in a skip connection manner;
[0021] An attention mechanism module for selecting attention features of the target region of the retinal vascular images from the convolutional features of each layer;
[0022] A feature connection module for inputting the pooling features of the last layer into the first residual pyramid convolutional layer of the decoder in the U-Net network and outputting sampled features using an upsampling layer;
[0023] A decoder module that uses each residual pyramid convolutional layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampled features of the upsampling layer and the corresponding attention features and transmit them to the last residual pyramid convolutional layer of the decoder;
[0024] An image output module for obtaining the segmentation result of the retinal blood vessels by performing a 1×1 convolution on the feature map of the last residual pyramid convolutional layer.
[0025] In a third aspect of the present invention, the present invention further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect of the present invention are implemented.
[0026] Advantages of the present invention:
[0027] In view of the problems such as low segmentation accuracy caused by the complex and variable scale information and morphological structure of retinal blood vessels, the present invention proposes a method, device and computer equipment for segmenting retinal blood vessel images. Based on the U-Net network, in the encoding stage, the residual pyramid module RPC is used to extract features from the retinal blood vessel images using convolutional kernels of different sizes and depths, so as to capture retinal blood vessel information of different scales; an attention mechanism is introduced in the skip connection to focus on the feature information of the target area and reduce interference; finally, the final segmentation result is obtained through the SoftMax activation function. Description of the Drawings
[0028] Figure 1 It is a flowchart of a method for segmenting retinal blood vessel images in an embodiment of the present invention;
[0029] Figure 2 It is a flowchart of a method for segmenting retinal blood vessel images in a preferred embodiment of the present invention;
[0030] Figure 3 It is a schematic diagram of the improved U-Net network structure of the present invention;
[0031] Figure 4 It is a diagram of the pyramid convolution module of the present invention;
[0032] Figure 5 It is a diagram of the residual pyramid convolution layer of the present invention;
[0033] Figure 6 It is a diagram of the residual pyramid convolution module of the present invention;
[0034] Figure 7 It is a diagram of the improved attention mechanism module of the present invention;
[0035] Figure 8 It is a structural diagram of a device for segmenting retinal blood vessel images in an embodiment of the present invention;
[0036] Figure 9 It is an enhanced image of the dataset of the present invention;
[0037] Figure 10 It is a diagram of the blood vessel segmentation result of the present invention and the prior art;
[0038] Figure 11 It is a diagram of the detailed comparison result between the present invention and the prior art. Detailed Embodiment
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0040] A method for segmenting retinal blood vessel images provided by this application can be applied to a server application environment. Specifically, the server acquires a retinal blood vessel image and preprocesses the retinal blood vessel image; the server inputs the preprocessed retinal blood vessel image into a trained U-Net network; the server uses each residual pyramid convolution layer and the corresponding pooling layer in the encoder of the U-Net network to extract convolution features and sampling features of the retinal blood vessel image at different levels; the server transfers the convolution features of each layer to the corresponding attention mechanism layer through skip connections, and selects the attention features of the target area of the retinal blood vessel image from the convolution features of each layer; the server inputs the sampling features of the last layer into the first residual pyramid convolution layer of the decoder in the U-Net network, and outputs sampling features using the upsampling layer; the server uses each residual pyramid convolution layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampling features of the upsampling layer and the corresponding attention features, and transfers them to the last residual pyramid convolution layer of the decoder to obtain the segmentation result of the retinal blood vessels.
[0041] Those skilled in the art of this technology can understand that the "server" used here can be implemented by an independent server or a server cluster composed of multiple servers.
[0042] Figure 1 It is a flowchart of a method for segmenting retinal blood vessel images in an embodiment of the present invention. As Figure 1 shown, the method includes inputting a retinal blood vessel image to be segmented, and performing image denoising enhancement and data augmentation processing on the retinal blood vessel image to be segmented; using the processed retinal blood vessel image to train a U-Net network model until the training is completed and meets the standard of qualified training, then storing the network model, and using the network model to output the segmentation result of the retinal blood vessel image to be segmented.
[0043] Figure 2 It is a flowchart of a method for segmenting retinal blood vessel images in a preferred embodiment of the present invention. As Figure 2 shown, the method includes:
[0044] 101. Acquire a retinal blood vessel image and preprocess the retinal blood vessel image;
[0045] In the embodiments of the present invention, in the actual field of medical images, there are a large number of retinal blood vessel images. These images are generated by medical staff (detecting personnel) using an instrument system during medical activities. These image information have various morphological structures. The purpose of the present invention is to segment these retinal blood vessel images with different morphological structures. Therefore, an instrument system can be accessed to obtain actual retinal blood vessel images.
[0046] In the embodiments of the present invention, the preprocessing of the retinal blood vessel images mainly includes denoising and enhancing the contrast of the obtained retinal blood vessel images, and processing the retinal blood vessel images using different color channels to obtain retinal blood vessel images with the highest contrast between blood vessels and the background.
[0047] In the preferred embodiments of the present invention, the preprocessing of the retinal blood vessel images further includes performing a data augmentation operation on the obtained retinal blood vessel images, randomly cropping and combining each retinal blood vessel image to the same size, so as to obtain augmented retinal blood vessel images.
[0048] 102. Input the preprocessed retinal blood vessel images into the trained U-Net network;
[0049] In the embodiments of the present invention, the preprocessed retinal blood vessel images to be segmented can be directly input into the trained U-Net network for subsequent segmentation and recognition. The U-Net network is a network improved by the present invention. As Figure 3 shown, the improved U-Net network of the present invention also includes an encoder, a decoder, and a skip connection module. The encoder is used to extract shallow features, deep features, and small blood vessel features of the retinal blood vessel images, that is, to extract convolutional features and sampling features at different levels of the retinal blood vessel images using each residual pyramid convolutional layer and the corresponding pooling layer in the encoder of the U-Net network; the decoder is composed of transposed convolutional layers and is used to restore the size of the feature map output by the encoder, that is, to use each residual pyramid convolutional layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampling features of the upsampling layer and the corresponding attention features and transfer them to the last residual pyramid convolutional layer of the decoder; the skip connection module is used to transfer the convolutional features of each layer to the corresponding attention mechanism layer through skip connection.
[0050] The core improvement of the present invention lies in replacing the convolutional layers in the encoder and decoder with residual pyramid convolutional layers. Specifically, the residual pyramid convolutional layer (RPC) includes two cascaded pyramid convolutional modules, and the output of the first pyramid convolutional module is connected to the output of the second pyramid convolutional module through a residual connection layer; each pyramid convolutional module includes two first units and one second unit, and the two first units are connected through the second unit. The first unit includes a batch normalization layer (BN), a convolutional layer (Conv), and a rectified linear unit layer (Relu); the second unit includes a batch normalization layer (BN), a pyramid convolutional layer (PyConv), and a rectified linear unit layer (Relu); each of the pyramid convolutional layers includes convolutional kernels of various different sizes.
[0051] In the embodiment of the present invention, pyramid convolution (Pyramid Ponvolution, PyConv) can process the input at multiple filter scales. PyConv contains a kernel pyramid, where each level contains different types of filters with different sizes and depths, capable of capturing different levels of details in the scene. In addition to these improved recognition capabilities, PyConv is also very efficient, and compared with standard convolution, it does not increase the computational cost and parameters. Moreover, it is very flexible and scalable, providing a huge potential network architecture space for different applications. As Figure 4 shown, the kernel size increases from the bottom of the pyramid (level 1 of PyConv) to the top (level n of PyConv). At the same time, as the spatial size increases, the depth of the kernel decreases from level 1 to level n. Therefore, this results in two interconnected pyramids in opposite directions. One pyramid has a bottom at the bottom (evolving to the top by reducing the kernel depth), and the other pyramid has a bottom at the top, where the spatial size of the convolutional kernel is the largest (evolving to the bottom by reducing the spatial size of the kernel). For the input feature map FM i , each layer {1, 2, 3,..., n} of the pyramid convolution corresponds to convolutional kernels of different sizes {K1 2 , K2 2 , K3 2 ,......, K n 2}, and each layer of convolutional kernels has different depths and can output different numbers of output feature maps. Therefore, the parameter number and computational cost formulas for pyramid convolution are:
[0052] Parameter number formula: Computational cost formula:
[0053] where H represents the height of the feature matrix, W represents the width of the feature matrix; FM on+...+FM o3 +FM o2 +FM o1 =FM o Each line in these equations represents the number of parameters and computational cost at a certain level in the pyramid convolution. If each layer in the pyramid convolution outputs an equal number of feature maps, then the number of parameters and computational cost in the pyramid convolution will be evenly distributed along each pyramid layer.
[0054] Inspired by residual learning and pyramid convolution, this paper proposes a residual pyramid module, whose structure is as Figure 5 shown. For the input feature Figure X , it is processed by two PyConvBlocks, and then connected to the original input feature map through a residual connection. After passing through the Relu activation function, the final feature map is output. The residual connection further strengthens the feature propagation and improves the performance of the network. The structure of the PyConvBlock is as Figure 6 shown, including normalization, convolution, and Relu operations. In the second convolution in the PyConvBlock, this paper uses PyConv with four different sizes of convolutional kernels: 3×3, 5×5, 7×7, and 9×9. The smaller convolutional kernel has a smaller receptive field and can obtain small target and local detail information. The larger convolutional kernel has a larger receptive field and can obtain large target and global semantic information. In the RPC module as Figure 5 shown, PyConv with different kernels is used to capture fine blood vessel branches, extract features at different levels in the retinal vessel image for combination, and improve the integrity and accuracy of segmentation. Finally, a residual structure is used for output, which solves the degradation problem caused by too many network cascade layers and speeds up the network convergence speed.
[0055] In the embodiment of the present invention, the encoder can be composed of 4 groups of residual pyramid convolution layer RPC modules and 4 groups of pooling layers. A pooling layer follows each RPC module. After the retinal vessel image is processed, it enters the RPC module. Each RPC module contains two pyramid convolution layers PyConv, and each PyConv contains four convolutional kernels Conv with different sizes, which are used to extract information of the retinal vessels at different scales for fusion, and then connect to a convolutional kernel Cat with a size of 1×1 to reduce the number of mapped input information. A max-pooling layer with a size of 2×2 is connected after each RPC module, which realizes the function of reducing the size of the retinal vessel feature map to half of the size of the previous layer's feature map. Therefore, the encoder is composed of 4 groups of residual pyramid convolution modules RPC and 4 groups of 2×2 max-pooling layers. Each group of RPC modules contains a convolutional layer with a size of 1×1 at the front and back, and a pyramid convolution with convolutional kernel sizes of 3×3, 5×5, 7×7, and 9×9 is included between these two ordinary convolutions.
[0056] In an embodiment of the present invention, the decoder may be composed of an upsampling layer (Upsampling layer) and an RPC module. The upsampling layer is a transposed convolutional layer with a kernel size of 2×2, which upsamples the output feature map to restore it to the original size, and then uses the SoftMax activation function to classify the blood vessel and background images, and outputs the segmentation result. An attention mechanism is introduced in the skip connection module to fuse the ratio of the background image and the blood vessels, so as to reduce the influence of the background on blood vessel segmentation. Therefore, the decoder is composed of 4 layers of 2×2 transposed convolutions and 1 layer of 1×1 ordinary convolution.
[0057] Based on the above analysis, the training process of the U-Net network may include:
[0058] Obtain the original retinal blood vessel image, and preprocess the original retinal blood vessel image to obtain a training data set. Among them, each retinal blood vessel image has a corresponding segmentation label, that is, a label image; input the image data in the training data set into the improved U-Net network for processing; the RPC module of the encoder performs shallow feature extraction on the input data to obtain the shallow features of the image; use residual connections in the RPC module to avoid the network degradation phenomenon that occurs when the number of network layers is too large; the skip connection module transmits the extracted shallow features to the attention mechanism module; use the attention mechanism to select the target region features, and transmit the selected features to the output layer of the encoder; the transposed convolutional layer of the decoder restores the feature map size of the deep features obtained after multiple convolutions and downsamplings by the encoder; splice the features upsampled by the encoder and the features output by the attention mechanism, and transmit the spliced feature map to the last convolutional layer to obtain the final feature map; compare the final feature map with the label image pixel by pixel to obtain the error; calculate the loss function of the model according to the error result, calculate the gradient of the target loss function by backpropagation, and use the stochastic descent algorithm to determine the minimum value of the target loss function. When the loss function is the smallest, the training of the model is completed.
[0059] 103. Use each residual pyramid convolutional layer and the corresponding pooling layer in the encoder of the U-Net network to extract the convolutional features and pooling features of the retinal blood vessel image at different levels;
[0060] In an embodiment of the present invention, in the U-Net network, the encoder includes a plurality of residual pyramid convolutional layers and a plurality of pooling layers. In this embodiment, these residual pyramid convolutional layers and pooling layers are connected in an interleaved manner. Therefore, the output result of the previous residual pyramid convolutional layer is input into the next pooling layer, and the output result of the previous pooling layer is input into the next residual pyramid convolutional layer. In this connection manner, each residual pyramid convolutional layer can output a layer of convolutional features, and each pooling layer can output a layer of sampling features. Therefore, convolutional features and sampling features from different levels can be extracted; these convolutional features and sampling features at different levels can reflect the shallow features, deep features, and small blood vessel features of the retinal blood vessel image.
[0061] 104. Transfer the convolutional features of each layer to the corresponding attention mechanism layer through skip connections, and select the attention features of the target region of the retinal blood vessel image from the convolutional features of each layer;
[0062] In an embodiment of the present invention, during the image segmentation process, considering that the retinal image may have background and other noise information that affects the segmentation result and reduces the segmentation accuracy, the present invention introduces an attention mechanism. Considering that the blood vessels in the retinal image are relatively scattered, the original attention mechanism may have errors in the weight distribution process for blood vessel regions and background regions, etc. Therefore, the attention mechanism is improved. On the basis of the original input, an additional input is added, and then the outputs of the two inputs are added and then subsequent operations are performed, which can more prominently highlight the target region and enable the model to be more focused on the feature learning of the target region in the retinal blood vessel image. As Figure 7 shown, the attention coefficient of the attention mechanism model is used to identify significant retinal blood vessel regions and prune the feature responses, only retaining the activations related to blood vessel information. The output of the attention mechanism model is the element-wise multiplication of the input feature map and the attention coefficient. The formula is:
[0063]
[0064] where the input feature map is Output feature c is the number of network layers input, which can correspond to the attention mechanism layer and the pooling layer, l is the size of the channel, and i is the size of the pixel space. During image segmentation, there are multiple semantic categories, and multi-dimensional attention coefficients can be used. Each AG will focus on its target segmentation situation. Compared with multiplicative attention, additive attention is used to obtain the gating coefficient to obtain better segmentation results. Additive attention formula:
[0065]
[0066]
[0067] Among them is the attention mechanism, is the weight parameter vector of the 1×1 convolution, and F int is the length of the pixel feature vector, is the attention coefficient, corresponds to the Sigmoid activation function, and Ω1 is the Relu activation function. By including the linear transformation parameter Θ att the features of AG can be obtained. This parameter includes the linear transformation coefficient matrix and and Among them, represents the first input transposed matrix of the first linear transformation coefficient matrix, represents the second input transposed matrix of the first linear transformation coefficient matrix; represents the first input transposed matrix of the second linear transformation coefficient matrix, represents the second input transposed matrix of the second linear transformation coefficient matrix; b g represents the first bias term; b Ψ represents the second bias term, and b Ψ ∈R, x i and g i are the input feature map and the gating signal respectively.
[0068] By analyzing the input features and the gating signal, the attention mechanism model can obtain the gating correlation coefficient. So that when segmenting an image, the attention mechanism model can focus on the feature information of the retinal blood vessels and eliminate noise information such as the background. It can be seen from the improved U-Net network structure that the attention mechanism model in the encoding part skips the pooling layer and the convolutional layer through the skip connection and is directly cascaded to the decoding part, fuses complementary information and uses a 1×1 convolutional layer for linear transformation, which helps to further reduce the breakage or gap phenomenon that occurs during the segmentation of the small blood vessels in the retinal image due to insufficient recovery. Improve the accuracy and integrity of blood vessel segmentation.
[0069] 105. Input the pooling features of the last layer into the first residual pyramid convolutional layer of the decoder in the U-Net network, and use the upsampling layer to output the sampled features;
[0070] In the embodiment of the present invention, the sampled features of the last layer in the encoder are used as the input of the first residual pyramid convolutional layer in the decoder to realize the normal connection between the encoder and the decoder. Of course, the complete connection between the encoder and the decoder also needs to be completed with the help of the skip connection.
[0071] 106. Use each residual pyramid convolutional layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampling features of the upsampling layer and the corresponding attention features, and transmit them to the last residual pyramid convolutional layer of the decoder to output the obtained feature map.
[0072] 107. Pass the output feature map through a 1×1 convolution to finally obtain the segmentation result of the retinal blood vessels.
[0073] In the embodiment of the present invention, since the encoder and decoder in the U-Net network have an approximately symmetric structure, in the embodiment of the present invention, the decoder includes multiple residual pyramid convolutional layers and multiple upsampling layers. In this embodiment, these residual pyramid convolutional layers and upsampling layers are connected in an interleaved manner. Therefore, the output result of the next-layer residual pyramid convolutional layer is input into the previous-layer upsampling layer, and the output result of the next-layer upsampling layer is input into the previous-layer residual pyramid convolutional layer. In this connection mode, each residual pyramid convolutional layer can output a layer of convolutional features until it is transmitted to the last residual pyramid convolutional layer. Through this last residual pyramid convolutional layer combined with the softmax layer, the segmentation result of the retinal blood vessels can be obtained.
[0074] Figure 8 It is a structural diagram of a retinal blood vessel image segmentation device in an embodiment of the present invention. As Figure 8 shown, the image segmentation device 200 includes:
[0075] An image acquisition module 201, configured to acquire a retinal blood vessel image;
[0076] An image processing module 202, configured to preprocess the acquired retinal blood vessel image;
[0077] An image input module 203, configured to input the preprocessed retinal blood vessel image into the trained U-Net network;
[0078] An encoder module 204, which uses each residual pyramid convolutional layer and the corresponding pooling layer in the encoder of the U-Net network to extract the convolutional features and pooling features of the retinal blood vessel image at different levels;
[0079] A skip connection module 205, configured to transmit the convolutional features of each layer to the corresponding attention mechanism layer in a skip connection manner;
[0080] An attention mechanism module 206, configured to select the attention features of the target area of the retinal blood vessel image from the convolutional features of each layer;
[0081] A feature connection module 207, configured to input the pooling features of the last layer into the first residual pyramid convolution layer of the decoder in the U-Net network, and output sampled features by using an upsampling layer;
[0082] A decoder module 208, configured to use each residual pyramid convolution layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampled features of the upsampling layer and the corresponding attention features, and transmit them to the last residual pyramid convolution layer of the decoder;
[0083] An image output module 209, configured to output the convolution features of the last residual pyramid convolution layer to obtain the segmentation result of the retinal blood vessels.
[0084] In one embodiment, a computer device is provided. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining a retinal blood vessel image, and preprocessing the retinal blood vessel image; inputting the preprocessed retinal blood vessel image into a trained U-Net network; using each residual pyramid convolution layer and the corresponding pooling layer in the encoder of the U-Net network to extract the convolution features and sampled features of the retinal blood vessel image at different levels; transmitting the convolution features of each layer to the corresponding attention mechanism layer in a skip connection manner, and selecting the attention features of the target area of the retinal blood vessel image from the convolution features of each layer; inputting the sampled features of the last layer into the first residual pyramid convolution layer of the decoder in the U-Net network, and outputting sampled features by using an upsampling layer; using each residual pyramid convolution layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampled features of the upsampling layer and the corresponding attention features, and transmitting them to the last residual pyramid convolution layer of the decoder to obtain the segmentation result of the retinal blood vessels.
[0085] The present invention will be described below with reference to the effect diagrams. In the embodiments of the present invention, during the training process, the selected databases are all RGB images, and the pixels are composed of a mixture of red, green, and blue colors. Since the contrast between the retinal blood vessels and the background is relatively low, in order to obtain the features of the fine blood vessels, it is necessary to perform fundus image enhancement to highlight the contrast between the blood vessels and the background. Figure 9It can be seen that in the green channel, the contrast between blood vessels and the background is high, and the noise interference is low, while in the red and blue channels, the contrast between blood vessels and the background is low, and the noise interference is high. Although the fundus image as a whole appears red, after calculating and analyzing the RGB channels after image grayscale conversion, it is found that the pixel value difference between blood vessels and the background in the red channel is small, which is not conducive to blood vessel segmentation. In addition, the high brightness display in the optic disc area results in the loss of some blood vessel information. In contrast, the pixel value difference between the retinal blood vessels and the background in the green channel is large, which is more conducive to blood vessel segmentation. Therefore, the green channel of the fundus image is used for processing.
[0086] As Figure 10 shown, the results of retinal blood vessel image segmentation using the network model proposed in the present invention are as follows. The (a) column is the original image, the (b) column is the standard segmentation result image, the (c) column is the segmentation image of U-Net, and the (d) column is the experimental result image obtained according to the network model proposed in the present invention. It can be seen from the segmentation result image of the U-Net algorithm that there are phenomena of blood vessel breakage and incomplete blood vessel segmentation, and the segmentation performance for the ends of small blood vessels is relatively poor. In addition, the segmentation result contains more noise information. The network model proposed in the present invention has a great improvement compared with the U-Net algorithm, can effectively suppress the influence of noise information, accurately segment more detailed blood vessel information, and improve the segmentation performance of small blood vessels.
[0087] As Figure 11 shown, (a) is the original image, and (b)-(e) are the local detail images of the original image, standard image, U-Net algorithm, and the present invention respectively. It can be intuitively seen that when approaching the blood vessel area, the segmentation result of the U-Net algorithm for small blood vessels is relatively blurred, and there are phenomena of blood vessel loss and breakage in the blood vessel crossing area. In addition, interference information will also appear. In contrast, by introducing the PRC module and attention mechanism into the classic U-Net in this paper, the morphological structure information of blood vessels at different levels can be captured, and the phenomenon of blood vessel breakage is effectively avoided in the segmentation of blood vessel crossing areas, and the interference of noise information is reduced. The experimental results show that the method of introducing the RPC module and attention mechanism to improve the original U-Net segmentation result can segment different blood vessel areas and obtain better segmentation results.
[0088] Finally, the result comparison on the DRIVE dataset is shown in Table 1:
[0089]
[0090] According to the comparison results in Table 1 above, it can be seen that the segmentation results of the present invention have a greater improvement in accuracy and sensitivity than those of U-Net, Recurrent U-Net, Residual U-Ne, RCBAM-Net, R2U-Net, and LadderNet. The overall segmentation effect is better and the results are more accurate.
[0091] It should be understood that although the steps in the flowchart of the drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0092] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "coaxial", "bottom", "one end", "top", "middle", "the other end", "upper", "one side", "top", "inner", "outer", "front", "center", "both ends", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0093] In the present invention, unless otherwise clearly specified and defined, the terms "installation", "setting", "connection", "fixation", "rotation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0094] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for segmenting retinal blood vessel images, characterized in that, The method includes: Obtain a retinal vascular image and preprocess the retinal vascular image; Input the preprocessed retinal vascular image into the trained U-Net network; Use each residual pyramid convolution layer and the corresponding pooling layer in the encoder of the U-Net network to extract the convolution features and pooling features of the retinal vascular image at different levels; Transfer the convolution features of each layer to the corresponding attention mechanism layer in a skip connection manner, and select the attention features of the target area of the retinal vascular image from the convolution features of each layer; The improved attention mechanism formula adopted by the attention mechanism layer is expressed as: Among them, represents the attention feature of the i-pixel space output by the c-th attention mechanism layer in the l-th channel, represents the pooling feature of the i-pixel space output by the c-th pooling layer in the l-th channel; is the attention coefficient of the c-th attention mechanism layer, Ω2 represents the Sigmoid activation function, represents the attention mechanism of the l-th channel, represents the input feature map of the i-pixel space output in the l-th channel; g i represents the gating signal in the i-pixel space; b g represents the first bias term; b Ψ represents the second bias term; Ψ T represents the weight parameter vector of the 1×1 convolution; Ω1 is the Relu activation function; the attention feature is obtained by a set of parameters Θ including linear transformation att obtained, and the parameter includes and Among them, represents the first input transposed matrix of the first linear transformation coefficient matrix, represents the second input transposed matrix of the first linear transformation coefficient matrix; represents the first input transposed matrix of the second linear transformation coefficient matrix, represents the second input transposed matrix of the second linear transformation coefficient matrix; Input the pooling feature of the last layer into the first residual pyramid convolution layer of the decoder in the U-Net network, and use the upsampling layer to output the sampling feature; Use each residual pyramid convolution layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampling feature of the upsampling layer and the corresponding attention feature, and transfer it to the last residual pyramid convolution layer of the decoder to output the obtained feature map; Perform a 1×1 convolution on the output feature map to finally obtain the segmentation result of the retinal blood vessels; Among them, the residual pyramid convolution layer includes two cascaded pyramid convolution modules, and the output of the first pyramid convolution module is connected to the output of the second pyramid convolution module through a residual connection layer; Each pyramid convolution module includes two first units and one second unit, and the two first units are connected through the second unit. The first unit includes a batch normalization layer, a convolution layer, and an activation function layer; The second unit includes a batch normalization layer, a pyramid convolution layer, and an activation function layer; Each of the pyramid convolution layers includes convolution kernels of various different sizes.
2. The retinal blood vessel image segmentation method according to claim 1, wherein Preprocessing the retinal vascular image includes denoising and enhancing the contrast of the obtained retinal vascular image, and processing the retinal vascular image using different color channels to obtain the retinal vascular image with the highest contrast between blood vessels and the background.
3. A method for segmenting retinal blood vessel images according to claim 1, characterized in that, Preprocessing the retinal vascular image also includes performing a data augmentation operation on the obtained retinal vascular image, randomly cropping and combining each retinal vascular image to the same size, so as to obtain the augmented retinal vascular image.
4. A method for segmenting retinal vessel images according to claim 1, characterized in that, The training process of the U-Net network includes performing a pixel-by-pixel comparison between the segmentation result of the retinal vascular image and its corresponding label image to obtain an error image; calculating the target loss function of the U-Net network according to the error image, calculating the gradient value by backpropagation of the target loss function, and using the stochastic descent algorithm to determine the minimum value of the target loss function. When the target loss function is the smallest, the training of the U-Net network model is completed.
5. A method for segmenting retinal vessel images according to claim 1, characterized in that, Each convolution kernel in the pyramid convolution layer has a different depth.
6. A retinal blood vessel image segmentation device, characterized in that, The device includes: An image acquisition module for acquiring a retinal vascular image; An image processing module for preprocessing the acquired retinal vascular image; An image input module for inputting the preprocessed retinal vascular image into the trained U-Net network; An encoder module that extracts convolutional features and pooling features at different levels of a retinal blood vessel image by using each residual pyramid convolutional layer and the corresponding pooling layer in the encoder of the U-Net network; A skip connection module for transmitting the convolutional features of each layer to the corresponding attention mechanism layer in a skip connection manner; An attention mechanism module for selecting attention features of the target region of the retinal blood vessel image from the convolutional features of each layer; the improved attention mechanism formula adopted by the attention mechanism layer is expressed as: Among them, represents the attention feature of the i-pixel space output by the c-th attention mechanism layer in the l-th channel, represents the pooling feature of the i-pixel space output by the c-th pooling layer in the l-th channel; is the attention coefficient of the c-th attention mechanism layer, Ω2 represents the Sigmoid activation function, represents the attention mechanism of the l-th channel, represents the input feature map of the i-pixel space output in the l-th channel; g i represents the gating signal in the i-pixel space; b g represents the first bias term; b Ψ represents the second bias term; Ψ T represents the weight parameter vector of the 1×1 convolution; Ω1 is the Relu activation function; the attention feature is obtained by a set of parameters Θ including linear transformation att obtained, and the parameters include and Among them, represents the first input transpose matrix of the first linear transformation coefficient matrix, represents the second input transpose matrix of the first linear transformation coefficient matrix; represents the first input transpose matrix of the second linear transformation coefficient matrix, represents the second input transpose matrix of the second linear transformation coefficient matrix; A feature connection module for inputting the pooling features of the last layer into the first residual pyramid convolutional layer of the decoder in the U-Net network and outputting sampled features by using an upsampling layer; A decoder module for using each residual pyramid convolutional layer and the corresponding upsampling layer in the decoder of the U-Net network to splice the sampled features of the upsampling layer and the corresponding attention features and transmit them to the last residual pyramid convolutional layer of the decoder; An image output module for obtaining the segmentation result of the retinal blood vessels by performing a 1×1 convolution on the feature map of the last residual pyramid convolutional layer; Among them, the residual pyramid convolutional layer includes two cascaded pyramid convolutional modules, and the output of the first pyramid convolutional module is connected to the output of the second pyramid convolutional module through a residual connection layer; each pyramid convolutional module includes two first units and one second unit, and the two first units are connected through the second unit. The first unit includes a batch normalization layer, a convolutional layer, and an activation function layer; the second unit includes a batch normalization layer, a pyramid convolutional layer, and an activation function layer; each pyramid convolutional layer includes convolutional kernels of various different sizes.
7. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Retinal vessel segmentation method in fundus image and computer readable storage medium
CN112233135A
Deep convolutional neural network suitable for corneal ulcer segmentation of fluorescent staining slit lamp image
CN112767406A