A fundus image blood vessel segmentation method based on line filtering and deep learning
By using Hessian matrix linear filtering and deep learning technology in fundus image processing, VSegNet network was established, which solved the problem of low blood vessel segmentation performance in fundus image in the prior art, and achieved more efficient and accurate blood vessel segmentation.
Patent Information
- Application Number
- CN202210546837.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The prior art has low performance in fundus image vascular segmentation, is susceptible to noise and requires manual selection of thresholds.
The linear filtering algorithm based on Hessian matrix is used to enhance the vascular region, and combined with MobileNetV3 as the basic model, a segmentation network VSegNet is established, and a recursive module is added for downsampling and decoder is added for upsampling, and vascular segmentation is performed through supervised deep learning.
It improves the adaptability and performance of vascular segmentation, can more accurately segment the blood vessel area in the fundus image, and enhances the ability to extract feature information.
Smart Images

Figure CN115205308B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and relates to a method for segmenting blood vessels in fundus images based on line filtering and deep learning. Background Art
[0002] Fundus images are widely used in disease detection and diagnosis. During the diagnosis process, the analysis of blood vessels in fundus images is an important way to identify systemic diseases such as arteriosclerosis, hypertension, and diabetic retinopathy. As an automated and intelligent technology, image processing has developed rapidly and achieved great success. In image processing, feature-based techniques can encode features of interest or individual features to complete automatic image analysis tasks. They have been successfully applied in various fields. For example, by exploring the global context information of SIFT features, a global context verification scheme for image copy detection was proposed; by using global features, local features, and a coarse-to-fine clustering method of the PageRank algorithm, an approximate duplicate elimination method was introduced for visual sensor networks; by developing the CCMSL model, a common feature space was learned from a heterogeneous aging database and the CRL model was constructed for age estimation across heterogeneous databases; an activity detection method based on fingerprint features, multi-scale local phasors, and principal component analysis; and a set of features based on quaternion wavelet transform that provides valuable information for distinguishing photographic images from computer-generated images. Using image processing techniques based on blood vessel features for automatic blood vessel analysis of fundus images can assist or even improve disease diagnosis, especially when analyzing a large amount of image data of patient groups.
[0003] Vascular segmentation is an important first step in fundus image vascular analysis. Many image segmentation methods have been used to study vascular segmentation of fundus images, including vascular tracing, filtering, mathematical morphology, deformable models, and machine learning; using a multi-scale line tracing program for vascular segmentation, selecting a small group of initial pixels and performing post-processing; filtering the initially acquired image using the Hessian matrix and using entropy thresholding to extract blood vessels from the background, and also applying connectivity constraints to reduce speckle noise; a vascular segmentation algorithm using mathematical morphology and curvature evaluation, which performs cross-curvature evaluation to distinguish blood vessels from similar background patterns after morphological operations; using snake models and specific domain knowledge including vascular topological properties for vascular segmentation; developing a vascular segmentation method using the extreme learning machine method, which is based on pixel classification using a 7-dimensional feature vector obtained from preprocessed retinal images. Among these methods, the preprocessing method based on the Hessian matrix and the introduction of the new network structure of VSegNet provide an effective supervised segmentation algorithm to obtain a more accurate network model for vascular segmentation, which is suitable for vascular segmentation of retinal images. However, the performance of most Hessian-based linear region segmentation methods is largely related to manually selected thresholds and is vulnerable to noise.
[0004] Therefore, there is an urgent need for a new method to achieve adaptive and high-performance fundus image segmentation. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a fundus image vascular segmentation method based on line filtering and deep learning, which can perform adaptive, high-performance, and supervised segmentation deep learning on images. The present invention strengthens the feature information extraction ability, thereby improving the model segmentation performance.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A fundus image vascular segmentation method based on line filtering and deep learning, comprising the following steps:
[0008] S1: Input a fundus image, and use a line filtering algorithm based on the Hessian matrix to enhance the vascular region;
[0009] S2: Use MobileNetV3 as the basic model of the vascular segmentation model, establish the segmentation network VSegNet, and then perform downsampling by adding an encoder based on a recursive module to the segmentation network VSegNet to highlight local key features;
[0010] S3: Add a decoder to the segmentation network VSegNet to perform upsampling and aggregation on the feature map output by the encoder;
[0011] S4: When training the segmentation network VsegNet, the L1 norm between the segmentation prediction result and the ground truth image is used to calculate the loss value of the segmentation result.
[0012] Furthermore, in step S1, the Hessian matrix utilizes the second-order structure of the local gray-scale variation of each pixel point in the fundus image, i.e., the characteristic represented by the second derivative.
[0013] Furthermore, in step S2, the first layer of the encoder is a standard 3×3 convolution with a stride of 2, and the first layer has 64 output channels.
[0014] The segmentation network VSegNet iteratively uses a recursive module to generate multi-scale feature maps; the recursive module consists of 5 inverted residual blocks. Among them, the stride of the middle block is 2, and the stride of the remaining blocks is 1. Therefore, each iteration of the recursive module will halve the size of the feature values. By setting the number of input and output channels of the recursive module to 64, reusability is achieved. Under the influence of the VSegNet series, the ratio of the feature resolution of the input image to the final output resolution of the encoder part is the output stride set to 32 in the encoder part. Therefore, the halved feature map from the first layer will iteratively pass through the recursive module four times, and each time the segmentation feature size will be halved. At the same time, in the recursive model, by setting the expansion ratio of the first three inverted residual blocks to 2 and the expansion ratio of the last two blocks to 4, the damage of ReLU to the feature map is reduced. To improve performance, without increasing the parameters of the vascular segmentation model, a squeeze and excitation (SE) block is used to regularize the feature map of the recursive module. The SE block is placed between the depth convolution inside the inverted residual block and the last point convolution, and the reduction ratio of the fully connected layer in the SE module is set to 16.
[0015] Furthermore, in step S2, the inverted residual block consists of a 1×1 convolution with ReLU6, ReLU6, a depth convolution with a stride of 1 or 2, an SE block, and a convolution without any non-linear activation.
[0016] Furthermore, in step S2, the SE block consists of global pooling, two fully connected (FC) layers, a REU non-linearity, a Sigmoid operation, and channel multiplication.
[0017] Further, in step S3, upsampling is implemented using a lightweight upsampling block, where the lightweight upsampling block consists of three residual DSconv blocks with shortcut connections between the input and output, namely upsampling, concatenation, and Sigmoid operations; after the output segmentation ground truth of the first convolutional layer and the recurrent module is split and jump-connected to the corresponding upsampling block in a concatenated manner, under the concatenation operation, all residual DSconv blocks except the second one have the input channel number; using the third residual DSconv block, and then performing Sigmoid operation to obtain the multi-scale segmentation map P t ; To improve the prediction accuracy, bilinear interpolation is performed on the multi-scale segmentation map to make it have the same feature resolution as the preprocessed black-and-white image.
[0018] Further, in step S3, the multi-scale segmentation map is expressed as follows:
[0019] D t = 1 / (aP t + b)
[0020] where the constants a and b are set to 10 and 0.01 respectively to always constrain the predicted segmentation map D t to be positive within the valid range, and finally enabling the segmentation network VsegNet to achieve a balance between high-depth prediction accuracy and low model parameters.
[0021] Further, in step S4, training the segmentation network VsegNet specifically includes: iterating through the preprocessed input image four times through the recursive module of the segmentation network VsegNet, and then using a new efficient upsampling block in the generated new network structure VSegNet to upsample and aggregate the feature map output by the encoder, thereby training to obtain an optimal fundus image blood vessel segmentation model; and using this model to map all inputs to corresponding outputs, and performing absolute deviation loss analysis on the output segmentation ground truth image, so that it has the ability to segment and predict fundus blood vessel images.
[0022] Further, in step S4, when training the segmentation network VsegNet, the stochastic gradient descent method with the backpropagation learning rule is used to calculate the minimized loss; by minimizing the loss L a between the predicted segmentation ground truth image and the corresponding segmentation prediction result, learning the mapping function between the segmentation ground truth image and the corresponding prediction result;
[0023] L a = ||I s - I d′ ||1
[0024] where I s represents the segmentation ground truth image of the synthetic image, Id′ Represents the segmentation prediction result of the synthetic image, ||||1 represents the L1 norm. It strengthens the feature information extraction ability, thereby improving the overall performance of the model.
[0025] Furthermore, the image preprocessing of the present invention uses a line filter algorithm, the core content of which is the combination of the line region enhancement of the Hessian matrix and the segmentation network, so that the blood vessel part in the generated fundus image obtains a visual enhancement effect, and can characteristically complete the sampling of the blood vessel image, which is an important step in completing the subsequent overall image segmentation.
[0026] In the subsequent image segmentation algorithm, a neural network based on supervised segmentation deep learning is adopted. Deep neural networks currently have a very broad application prospect and have relatively mature applications in image processing related aspects such as semantic segmentation, object detection, and image classification. In addition, recent research has shown that deep learning can recover pixel-level depth maps from a single image in an end-to-end manner. In the current general processing method of monocular depth estimation, many neural network models have proven their effectiveness, such as recurrent neural networks, variational autoencoders, convolutional neural networks, and adversarial neural networks. The VSegNet network model adopted by the present invention has been optimized in both the encoder part and the decoder part. For example, the introduction of the recursive module can make the data acquisition in the encoder part more complete, and the spatial feature size of the data can be iteratively processed to become more detailed and easy to analyze. And a supervised learning segmentation algorithm is adopted in the upsampling module, and the L1 norm is calculated between the feature value of the segmentation result and the feature value of the true segmentation result, so as to correct the model through the loss value of the segmentation result. The deep network performs supervised learning deep inference, which is also one of the core elements of the present invention.
[0027] The present invention realizes the segmentation of fundus images. Therefore, first of all, it is necessary to ensure the quality of the segmented fundus images, and the Hessian matrix and the supervised segmentation deep learning network VSegNet are adopted for each segmentation stage to improve the fundus image segmentation ability.
[0028] The beneficial effects of the present invention are as follows: The present invention can enhance the blood vessel region in the fundus image, combine the efficient encoder and decoder network design to find stable information in data changes, enhance the feature information extraction ability, thereby improving the performance of the network model and obtaining a more accurate blood vessel segmentation image.
[0029] Other advantages, objects, and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art upon examination of the following, or may be learned from the practice of the present invention. The objects and other advantages of the present invention may be realized and attained by the means of the instrumentalities and combinations particularly pointed out hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to make the objectives, technical solutions, and advantages of the present invention more clear, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0031] Figure 1 Schematic structural diagram of the encoder part of the deep network;
[0032] Figure 2 Schematic structural diagram of the inverted residual block and the SE block;
[0033] Figure 3 Schematic diagram of the network generator model architecture;
[0034] Figure 4 Schematic structural diagram of the upsampling module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically, and the following embodiments and the features in the embodiments can be combined with each other without conflict.
[0036] Vascular segmentation is the key and difficulty in fundus image processing and is a prerequisite and necessary first step for further vascular measurement and diagnosis. Aiming at the problem of vascular segmentation in fundus images, a new hybrid vascular automatic segmentation method is proposed. The method includes two main steps: linear filtering based on the Hessian matrix and vascular segmentation using the supervised segmentation deep learning network VSegNet. First, the vascular region is enhanced by adopting linear filtering based on the Hessian matrix; after filtering, the image is input into the VSegNet network model for processing; then, downsampling is performed through the encoder part composed of recursive modules to improve the ability to capture feature information; finally, the output feature map of the encoder is processed by the decoder of the more efficient upsampling module. Compared with the traditional network segmentation model, the method of the present invention applies a supervised algorithm model and can automatically obtain more accurate and complete vascular segmentation results.
[0037] Please refer to Figures 1 to 4 , the algorithm model used in the present invention mainly includes the following steps:
[0038] Step 1: Implement blood vessel segmentation by combining the line filter algorithm and the segmentation network. The core content is to first enhance the blood vessel area in the fundus image through the Hessian matrix and then combine the segmentation method to achieve blood vessel segmentation;
[0039] Step 2: Based on MobileNetV3 as the basic model, establish a new network structure VSegNet. A recursive module is added to the encoder structure of this network to form the new network structure VSegNet;
[0040] Step 3: In the generated new network structure VSegNet, a new efficient upsampling block is used to upsample and aggregate the feature maps output by the encoder;
[0041] Step 4: When training the network, use the L1 norm of the segmentation result and the true segmentation result to calculate the loss value of the segmentation result.
[0042] In Step 1, linear feature filtering based on the Hessian. For a two-dimensional image I, the Hessian matrix is used to describe the second-order derivatives in all directions of each pixel point, expressed as
[0043]
[0044] where, I xx , I xy , I yx , I yy represents the second-order derivative of the two-dimensional image I, and Q Hessian is a symmetric matrix when I xy is equal to I yx . According to the linear scale theory, the second-order derivative can be obtained by convolution with the Gaussian derivative. For example, at a pixel point (x, y)
[0045]
[0046] where, G is the Gaussian function with scale σ.
[0047]
[0048] Two eigenvalues λ1 and λ2 (|λ1|≥|λ2|) can be decomposed from the Hessian matrix in Equation (1). These eigenvalues are important parameters of the evidence, as well as the eigenvectors corresponding to the two eigenvalues. The maximum absolute value |λ1| of the eigenvalue λ1 represents the local maximum of the curvature, and each pixel has a principal direction, which we define as the direction of the eigenvector corresponding to its maximum absolute eigenvalue. As Figure 3 shown, in the context of a two-dimensional image, the curvature of pixels in a linear region is usually small along the principal direction of the linear structure and large in the perpendicular direction. Therefore, it can be used to represent |λ1|>>|λ2|≈0. The linear filter is represented as
[0049]
[0050] R in Equation (4) B =(λ2 / λ1) 2 to distinguish between bubble-like and linear structures and to distinguish between objects and the background. α and β are two parameters, where α usually refers to the linear filter and β usually refers to the sensitivity. In this work, α and β are fixed at 0.5 and 15 respectively. The linear structures in the image usually vary in size, so the scale factor σ of the linear structure is extended to multiple scales, as shown in Equation (5)
[0051]
[0052] When the scale factor σ matches the half-width of the linear structure, ν(x,y;σ) is maximized. Since the Hessian-based feature filter is widely used in image processing, including linear filtering in two-dimensional images, tubular filtering in three-dimensional images, and planar filtering. Also, in two-dimensional space, blood vessels are linear structures. Therefore, we decided to use two-dimensional linear filtering in our study.
[0053] In the second step, a new network structure VSegNet composed of a recursive module is added. The output feature maps from the first convolutional layer and the recursive module will be jump-connected to the corresponding upsampling blocks in a cascaded manner. The deep network iteratively uses the recursive module to generate multi-scale feature maps, where i and T represent the iteration time and the total number of iterations respectively. S represents the stride number of the convolutional layer. The multi-scale segmentation prediction will be bilinearly upsampled to the same feature resolution as the preprocessed black-and-white image. As Figure 1As shown, the encoder part of the deep network consists of standard convolutional layers and a recursive module. Similar to MobileNetV3, the first layer of the encoder part proposed in this invention is a standard 3×3 convolution with a stride of 2, and the first layer has 64 output channels, and then ReLU is activated. The recursive module in this paper consists of 5 inverted residual blocks, where the middle block has a stride of 2 and the other blocks have a stride of 1. Therefore, each iteration of the recursive module will halve the size of the feature values. To achieve reusability, the input and output channels of the recursive module have the same design, that is, both are 64. Under the influence of the VSegNet series, the ratio of the input image feature resolution to the final output resolution of the encoder part will be used as the output stride set to 32 in the encoder part. Therefore, the halved feature map from the first layer will iterate through the recursive module four times, and each time the spatial feature size will be halved.
[0054] The recursive module is based on the inverted residual block of MobileNetV3, which has an inverted residual and a linear bottleneck to mitigate the damage of ReLU to the feature map. As shown in Figure 2 (a), the inverted residual block consists of a 1×1 convolution of ReLU6, ReLU6 and a depth convolution with a stride of 1 or 2, a squeeze and excitation (SE) block, and a convolution without any non-linear activation. If there is a depth convolution block with a stride of 1, then the input and output ends will be connected in a direct connection manner. The expansion ratio of the inverted residual block is set to 2 or 4, that is, the ratio of the output channels of the first pointwise convolution to the input channels. To achieve a compromise between model parameters and depth prediction accuracy, in the recursive model of this paper, the expansion ratios of the first three inverted residual blocks are set to 2, and the expansion ratios of the last two blocks are set to 4. To improve performance, without increasing model parameters, the SE block is used to regularize the feature map of the recursive module. As shown in Figure 2 (b), the SE block consists of global pooling, two fully connected (FC) layers, REU non-linearity, Sigmoid operation, and channel multiplication. Just as done in MobileNetV3, the SE block is placed between the depth convolution and the last pointwise convolution inside the inverted residual block. In the SE module, the reduction rate of the fully connected layer is set to 16.
[0055] As shown in Figure 3As shown. The encoder structure is organically composed of multi-scale residual blocks and downsampled by four convolutions with a stride of 2. The bottleneck includes six residual dense blocks (RDB). On the other hand, the decoder includes a concatenation of transposed convolutions for upsampling and skip connection layers, followed by a convolution. The output of the last decoder block summarizes the feature maps from the refinement module, followed by a convolution operation to achieve the same size as the input image. To improve the ability of the generator to produce images with better details and edge information, the downsampled images obtained from different unsharp masks are connected. The generator architecture consists of multi-scale residual blocks in each encoder layer, the bottleneck consists of six residual dense blocks, and the decoder layer consists of transposed convolutions and a concatenation with skip connection layers, a refinement module, and a few convolutional layers to generate an image of the input dimension. Generator G DS The generator G will perform unsharp masking on the input feature map.
[0056] Generator G DS Helps to generate synthetic fundus vascular images, and the encoder layer of the network does not contain the connection of the image obtained from the unsharp mask and the input feature map. The sharpened image obtained by unsharp masking is as follows:
[0057] g(x,y) = f smooth (x, y) (6)
[0058] f sharp (x, y) = f(x, y) + k × g(x, y) (7)
[0059] where f(x,y) is the input image, f smooth (x,y) is the smoothed image obtained by convolution, and g(x,y) is the image with high-frequency information. Multiply this high-frequency information by the quantity k and add it to the original image, thus producing the expected clearer image f sharp (x,y) with more details and better edge information. Utilizing the function of the sharpened mask image to improve the high-frequency image elements and the overall contrast of the image, the low-contrast and slowly changing non-vascular parts of the fundus are suppressed. Each layer of the encoder has been downsampled relative to the previous layer, and the size of the sharpened image is consistent with the size of a specific encoder layer, enabling channel-wise concatenation. The kernel sizes for generating the sharpened images are 24, 12, 6, and 3, with different downsampled images having 1, 1 / 2, 1 / 4, 1 / 8 dimensional values relative to the original image.
[0060] The encoder used in the present invention records the position of the maximum value during the maximum pooling operation in the operation process, and then realizes non-linear upsampling through the corresponding pooling index operation during decoding. The sparse feature map is generated through the upsampling operation, and then the dense feature map is obtained. Ordinary convolution is used in this step, and then the upsampling operation is repeated several times. Finally, the activation function is used to generate the one-hot encoding classification result. VSegNet mainly compares with FCN. We know that FCN has different characteristics during decoding. The unique operation during FCN decoding is the transposed convolution operation. Through this operation, the feature map can be obtained, and then it is coupled with the corresponding encoded feature map to finally obtain the output. In contrast, VSegNet has two major advantages: one is that it does not need to completely save the feature map of the encoding part, and only needs to save the result of the pooling index, which can greatly save the internal storage space; the other is that it does not use the transposed convolution operation, and only needs convolution learning after the upsampling is completed, and no learning is required during the upsampling stage.
[0061] In step three, in order to meet the requirements of high precision and real-time performance, a new efficient upsampling block is adopted in the generated new network structure VSegNet to upsample and aggregate the feature map output by the encoder. As Figure 4 shown, the output feature maps from the first convolutional layer and the recurrent module will be connected to the corresponding upsampling block through the cascaded skip connection. Different from PYD-Net, the lightweight upsampling block adopted in the present invention consists of three residual DSconv blocks, namely upsampling, cascading, and Sigmoid operations. PYD-Net uses a heavy decoding block with one transposed convolution and four standard convolutions. These residual DSConv blocks are plug-in replacements for standard convolutions. The standard convolution consists of two parts: depth and point convolution (i.e., depthwise separable convolution), and has a shortcut connection between the input and the output. Due to the cascading operation, all residual DSconv blocks except the second DSconv block with 2c have the input channel number C. Using the third residual DSconv block, and then performing the Sigmoid operation, the multi-scale segmentation map Pt is obtained. In order to improve the prediction accuracy, the multi-scale segmentation map is bilinearly interpolated to make its feature resolution the same as that of the preprocessed black and white map. The multi-scale segmentation ground truth map can be expressed as follows:
[0062] D t = 1 / (aP t + b) (8)
[0063] where the constants a and b are set to 10 and 0.01 to always constrain the predicted segmentation ground truth D t to be positive within the valid range.
[0064] After the decoder completes the upsampling operation and the convolution operation, each pixel is sent to the softmax classifier. During upsampling, the max-pooling indices at the corresponding encoder layers are called to perform upsampling multiple times. Finally, the class of each pixel needs to be predicted, and a K-class softmax classifier is used for this process. The advantage of the upsampling module is that while restoring the size of the feature map, it effectively improves the segmentation accuracy and reduces the computational complexity. The possible reasons are as follows: 1) The new upsampling model can greatly improve the image reconstruction ability; 2) The upsampling operation of the decoder can flexibly utilize the combination of features from different layers of any CNN decoder. By using upsampling, the increase in computational cost and memory occupancy caused by reducing the stride of the decoder is avoided.
[0065] In step four, when training the network, the L1 norm of the segmentation result and the true segmentation result is used to calculate the loss value of the segmentation result. Through the supervised segmentation algorithm, deep learning can be achieved, and the network model can be optimized to make the segmentation network model more accurate. In the supervised method, the true data of the segmented image is used as the supervision signal, so that the estimation of the image segmentation data is regarded as a regression problem. The deep convolutional neural network uses the difference between the predicted depth information and the true depth information to supervise the network for training, and can effectively perform depth estimation on a single image. Generally speaking, the process of supervised learning is actually a constraint process on the following function, that is, a minimization process of the objective function.
[0066] w * = rgmin w ∑ i L(y i ,(x i ,))+Ω(w) (9)
[0067] Among them, the first term L(y i ,f(x i ; w)) is used to describe the error generated between the predicted value f(x i ; w) of the i-th sample and the true label y i in the classification problem or regression problem and is used to measure it. In order to make the model better fit the training requirements of the training samples, we require the value of this term to be minimized, that is, we need the model we adopt to have a high fitting degree with the training data. For this reason, we need to use the regularization function Ω(w) to constrain the parameter w.
[0068] In the present invention, the L1 norm is adopted as the regularization term. As shown in the following formula, it is the absolute deviation loss between the true segmentation image and the segmentation prediction result. By minimizing the loss between the predicted true segmentation image and the corresponding segmentation prediction result, the mapping function between the true segmentation image and the corresponding prediction map is learned.
[0069] L a = ||I s - I d′ ||1(10)
[0070] where I s represents the ground truth image of the segmentation of the synthetic image, and I d′ represents the segmentation prediction result of the synthetic image, and || ||1 represents the L1 norm.
[0071] Combined with the supervised learning segmentation algorithm, an image segmentation process with automatic optimization and automatic precise feedback can be achieved.
[0072] The decoder of the present invention is very efficient and simple, and is used for the semantic segmentation module of processing images. The decoder mainly uses the prediction results of upsampling pixel by pixel. The adverse effects generated by the CNN when calculating high-resolution feature maps, such as the problem of low computational efficiency, will be greatly reduced, and the coupling effect between the features to be fused and the final output is released, enabling more flexible selection of the features to be fused. Moreover, the method proposed by the present invention does not apply the upsampling operation to the deep features of low resolution, greatly reducing the computational amount of the decoding module. A linear filtering algorithm based on the Hessian matrix is added at the sampling end of the source image to further enhance the blood vessel area for preprocessing. Then, the new segmentation network structure VSegNet is used to process the image. Finally, through the supervision network model and L1 norm optimization, the fundus image is transformed from the most abstract model into an accurate, effective, and medical-demand-satisfying image. It can solve many complex medical problems and generate relatively accurate images under many uncertain situations. At the same time, it can shield interference to a certain extent, reduce noise, self-regulate, and reduce the changes caused by data variation or uneven data distribution.
[0073] In future development, both the quantification and diagnosis of blood vessels should be emphasized, which is very important from the perspective of clinical applications. The basic principle of the present invention is to first study whether the segmentation of blood vessels in fundus images is feasible and then perform other tasks. Another limitation of this work is the connection of linear structures and the segmentation problem at the ends, which is an inherent problem of the Hessian-based method. The nodes and endpoints of linear structures are usually not ideal linear structures, resulting in small filtering results. Therefore, the connection and ends of linear structures in the results may not be detected as blood vessels. It may have an adverse impact on more complex cases. Many researchers have also proposed methods for node detection and enhancement in Hessian-based linear filtering. We will further study this point in the future and improve the functions and uses of the present invention.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A fundus image blood vessel segmentation method based on linear filtering and deep learning, characterized in that, The method comprises the following steps: S1: Input a fundus image, and enhance the blood vessel region using a linear filtering algorithm based on the Hessian matrix; S2: Use MobileNetV3 as the base model of the blood vessel segmentation model to establish a segmentation network VSegNet, and then perform downsampling by adding an encoder based on a recursive module to the segmentation network VSegNet; The first layer of the encoder is a standard 3×3 convolution with a stride of 2, and the first layer has 64 output channels; The segmentation network VSegNet iteratively uses the recursive module to generate multi-scale feature maps; the recursive module consists of 5 inverted residual blocks, where the middle block has a stride of 2 and the remaining blocks have a stride of 1, so each iteration of the recursive module will halve the size of the feature values; the ratio of the feature resolution of the input image to the final output resolution of the encoder part is the output stride in the encoder part; without increasing the parameters of the blood vessel segmentation model, a squeeze-and-excitation (SE) block is used to regularize the feature maps of the recursive module, and the SE block is placed between the depth convolution inside the inverted residual block and the last point convolution; The inverted residual block consists of a 1×1 convolution of ReLU6, ReLU6 and a depth convolution with a stride of 1 or 2, an SE block, and a convolution without any non-linear activation; S3: Add a decoder to the segmentation network VSegNet to upsample and aggregate the feature maps output by the encoder; The upsampling is implemented using lightweight upsampling blocks. Among them, the lightweight upsampling block consists of three residual DSconv blocks with shortcut connections between the input and output, namely, upsampling, concatenation, and Sigmoid operations. After the output of the first convolutional layer and the recursive module is split into true values, they will be jump-connected to the corresponding upsampling blocks in a concatenated manner. Under the concatenation operation, all residual DSconv blocks except the second one have the input channel number. The third residual DSconv block is used, and then the Sigmoid operation is performed to obtain the multi-scale segmentation map. P t Bilinear interpolation is performed on the multi-scale segmentation map to make it have the same feature resolution as the preprocessed black-and-white image. S4: When training the segmentation network VsegNet, calculate the loss value of the segmentation result using the L1 norm between the segmentation prediction result and the segmentation ground truth image.
2. The fundus image blood vessel segmentation method according to claim 1, characterized in that, In step S1, the Hessian matrix utilizes the second-order structure of the local gray change of each pixel point in the fundus image, i.e., the characteristic represented by the second-order derivative.
3. The fundus image blood vessel segmentation method according to claim 1, characterized in that, In step S2, the SE block consists of global pooling, two fully connected layers, a REU non-linearity, a Sigmoid operation, and channel multiplication.
4. The fundus image blood vessel segmentation method according to claim 1, characterized in that, In step S3, the multi-scale segmentation map is represented as follows: Among them, the constants a and b are set to 10 and 0.01 respectively to always constrain the predicted segmentation map to be positive within the valid range, ultimately enabling the segmentation network VsegNet to achieve a balance between high-depth prediction accuracy and low model parameters.
5. The fundus image blood vessel segmentation method according to claim 1, characterized in that, In step S4, the specific training of the segmentation network VsegNet includes: iterating the preprocessed input image through the recursive module four times through the segmentation network VSegNet, and then using a new efficient upsampling block from the generated new network structure VSegNet to upsample and aggregate the feature maps output by the encoder, so as to train an optimal blood vessel segmentation model for fundus images; and using this model to map all inputs to corresponding outputs, and performing absolute deviation loss analysis on the output segmentation ground truth image, so that it has the ability to segment and predict fundus blood vessel images.
6. The fundus image blood vessel segmentation method according to claim 5, characterized in that, In step S4, when training the segmentation network VsegNet, the stochastic gradient descent method with the backpropagation learning rule is used to calculate the minimized loss; by minimizing the loss between the predicted segmentation ground truth image and the corresponding segmentation prediction result, the mapping function between the segmentation ground truth image and the corresponding prediction result is learned. , the mapping function between the segmentation ground truth image and the corresponding prediction result is learned. Among them, I s represents the segmentation ground truth image of the synthetic image, represents the segmentation prediction result of the synthetic image, and || ||1 represents L the L1 norm.
Citation Information
Patent Citations
Fundus image retinal vessel segmentation method and system based on deep learning
CN106408562A
Retinal vessel segmentation method and device based on guided filtering
CN113658207A