Retinal vessel extraction method based on strong feature dynamic fusion gating network

By adopting a strong feature dynamic fusion gating network in retinal vascular segmentation technology, combined with adaptive gating residual blocks, parallel information focus modules and feature navigation hubs, the shortcomings of the existing technology in vascular fine-grained feature extraction, dynamic adaptability processing and feature balance are solved, and higher segmentation accuracy and robustness are achieved.

CN120198673AInactive Publication Date: 2025-06-24CHANGCHUN UNIV OF SCI & TECH +1

Patent Information

Application Number
CN202510668002.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing retinal vascular segmentation technology has shortcomings in vascular fine-grained and complex feature extraction, dynamic adaptive processing, and balance of global context information and local detailed feature.

Method used

The retinal vascular extraction method based on a strong feature dynamic fusion gating network is adopted, and the dynamic fusion of multi-scale features and the application of attention mechanisms through core modules such as adaptive gating residual blocks, parallel information focus modules and feature navigation hubs.

Benefits of technology

It significantly improves the accuracy and robustness of retinal vascular segmentation, especially in terms of small blood vessel sensitivity, lesion area robustness and overall segmentation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198673A_ABST
    Figure CN120198673A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a retinal vessel extraction method based on a strong feature dynamic fusion gating network, which adopts strong feature dynamic fusion gating U-Net, designs a core module as a center node by introducing a centrality thought in a graph theory, dynamically aggregates key information of a multi-layer encoder, and improves the extraction efficiency of the retinal vessel. Therefore, effective fusion of global and local features is realized. The feature fusion hub module balances the relationship between global context information and local detail features by using the long-range dependency relationship of multilayer features, and the adaptive gating residual block screens key features and inhibits redundant information through a dynamic gating mechanism, so that the segmentation robustness in complex vascular morphology and lesion regions is enhanced, and the segmentation efficiency is improved. The parallel information focusing module combines a channel and a space attention mechanism, global position information and long-range dependence are effectively modeled, fine-grained and complex features of blood vessels are captured, and the segmentation performance is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a method for extracting retinal blood vessels based on a strong feature dynamic fusion gating network. Background Art

[0002] Retinal blood vessel segmentation technology has irreplaceable clinical value in modern medicine. As an important part of the human microcirculation system, the morphological characteristics of retinal blood vessels are of great significance for the diagnosis of various eye diseases and systemic diseases. The severity of hypertensive retinopathy is also closely related to the risk of cardiovascular events. Evaluating vascular changes through retinal images can provide a scientific basis for the prediction of events such as stroke and myocardial infarction. The results of image analysis can also reflect the treatment effect of patients and provide real-time reference for hypertension management. Accurately analyzing these vascular morphological characteristics can assist doctors in more effectively identifying diseases and formulating more personalized treatment plans.

[0003] Although current research has made certain progress in improving the performance of retinal blood vessel segmentation, there are still the following deficiencies: existing methods still need to be improved in extracting fine-grained and complex features of blood vessels; for the characteristics of complex blood vessel morphologies (such as thick blood vessels and tiny blood vessels) and lesion areas (such as exudates and bleeding points), there is a lack of an effective dynamic adaptation processing mechanism; in terms of balancing global context information and local detail features, the capabilities of existing methods are still insufficient. Summary of the Invention

[0004] To achieve the above object, this application provides the following technical solutions: According to the first aspect of the present invention, the present invention claims protection for a method for extracting retinal blood vessels based on a strong feature dynamic fusion gating network, including: Collecting an original image of the retinal blood vessels to be extracted, and inputting the original image into an encoder to obtain a plurality of intermediate encoded features and a target encoded feature; Inputting the plurality of intermediate encoded features through skip connections into a feature navigation hub to obtain corresponding multiple fused encoded features, and inputting the target encoded feature into a parallel information focusing module to obtain parallel information focusing features; Inputting the plurality of fused encoded features after skip connections and the parallel information focusing features into a decoder together to obtain integrated blood vessel features; After each layer of decoder processing, enhancing the feature fusion effect and gradually restoring the image details to finally obtain a binary output image of the retinal blood vessels.

[0005] Further, the collecting an original image of the retinal blood vessels to be extracted and inputting the original image into an encoder further includes: An adaptive gating residual block is adopted, which combines residual connection and adaptive gating mechanism, and introduces a spatial dropout module to alleviate overfitting; In the encoder, the adaptive gating residual block includes two convolutional layers and their corresponding batch normalization and activation layers.

[0006] Furthermore, obtaining multiple intermediate encoded features and target encoded features further includes: The first feature of the original image is processed through the first convolutional layer, and after batch normalization and ReLU activation function, the first intermediate encoded feature is obtained; The first intermediate encoded feature is passed through the second convolutional layer and batch normalization again to obtain the second intermediate encoded feature; The spatial dropout module processes the second intermediate encoded feature to reduce noise in the feature map and enhance the regularization effect; The adaptive gating mechanism of the adaptive gating residual block extracts global context information through global average pooling, and generates a gating coefficient through two fully connected layers and Sigmoid activation function, which is used to adjust the weight of the residual term; The residual connection part performs channel matching on the first feature to ensure that the number of channels is consistent with the main branch feature, then performs weighted addition with the second intermediate encoded feature, and obtains the corresponding target encoded feature through ReLU activation.

[0007] Furthermore, inputting the target encoded feature into the parallel information focusing module to obtain the parallel information focusing feature further includes: A channel attention module is used to aggregate the global features of each channel, generate channel importance weights, and adjust the channel expression ability of the input feature map; Input the target encoded feature, extract the global descriptor of each channel through global average pooling operation, and generate channel weights through two-layer fully connected network to achieve feature compression and expansion; The generated channel weights are multiplied element-wise with the input feature map to complete the enhancement in the channel dimension; The spatial attention module captures global spatial information through maximum pooling and average pooling of the feature map in the spatial dimension, and generates spatial attention weights; Perform maximum pooling and average pooling operations on the first feature in the channel dimension to generate two spatial feature maps; The two spatial feature maps are concatenated in the channel dimension, and the result passes through a convolutional layer to generate spatial attention weights; The spatial weights are multiplied element-wise with the input feature map to complete the enhancement in the spatial dimension; After completing the attention optimization in the channel and space, the final parallel information focusing feature is generated through a fusion operation.

[0008] Further, the step of inputting the multiple intermediate encoded features into the feature navigation hub through skip connections to obtain corresponding multiple fused encoded features further includes: Performing convolution, batch normalization, and ReLU activation operations on each intermediate encoded feature using the skip connection features of a preset encoder to obtain uniformly processed features; In each layer of decoding, adjusting the target resolution of the intermediate encoded features through bilinear interpolation; Using the attention mechanism of the Transformer to capture the dependencies between multi-scale features, stacking all the adjusted features along a new dimension, and reshaping them into a sequence representation; Reshaping the features into a tensor of the target shape and further permuting the dimensions to meet the input requirements of the Transformer encoder; After passing through the Transformer encoder, global context information can be integrated to output multiple fused encoded features. Further, the step of inputting the multiple fused encoded features into the decoder through skip connections together with the parallel information focusing feature to obtain integrated vascular features further includes:

[0009] Adopting an adaptive gated residual upsampling block that includes upsampling and skip connection operations; Upsampling the input parallel information focusing feature to half the size of the original image through bilinear interpolation, and then splicing it with the feature map of the corresponding layer in the feature navigation hub to integrate the semantic information in the encoding stage; The spliced feature map is further processed by the adaptive gated residual block to enhance the feature fusion effect and restore image details to obtain integrated vascular features.

[0010] The present application relates to the technical field of image processing, and in particular to a retinal vessel extraction method based on a strong feature dynamic fusion gating network. By adopting a strong feature dynamic fusion gating U-Net, the centrality idea in graph theory is introduced, and the core module is designed as the central node to dynamically aggregate the key information of multiple layers of encoders, thereby realizing the effective fusion of global and local features. The feature fusion hub module balances the relationship between global context information and local detail features by utilizing the long-range dependencies of multi-layer features. The adaptive gated residual block screens key features through a dynamic gating mechanism and suppresses redundant information, thereby enhancing the segmentation robustness in complex vessel morphologies and lesion regions. The parallel information focusing module combines channel and spatial attention mechanisms, effectively models global position information and long-range dependencies, captures fine-grained and complex features of blood vessels, and further improves the segmentation performance. Description of the Drawings

[0011] Figure 1 The flowchart of a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 2 The overall model diagram of EFDG-UNet for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 3 The structural diagram of an adaptive gating residual block for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 4 The structural diagram of a parallel information focusing module for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 5 The structural diagram of a feature navigation hub for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 6 The structural diagram of a Transformer for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 7 The structural diagram of an adaptive gating residual upsampling block for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application; Figure 8 The schematic diagram of a retinal dataset for a retinal vessel extraction method based on a strong feature dynamic fusion gating network requested to be protected by the embodiments of the present application. Detailed implementation manners

[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0013] The terms "first", "second", and "third" in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of this application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units not listed, or optionally also includes other steps or units inherent to these processes, methods, products, or devices.

[0014] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0015] Integrity and distribution anomalies can directly serve as early warning signals for cardiovascular and cerebrovascular diseases, providing key evidence for the risk assessment of diseases such as hypertension, coronary heart disease, and stroke. The impact of hypertension on retinal blood vessels is particularly significant. Under long-term hypertension, phenomena such as arterial stenosis, sclerosis, and arteriovenous crossing compression frequently occur. These features are called "copper wire-like" or "silver wire-like" blood vessels and are important indicators of the course of hypertension and potential damage to the systemic vascular system.

[0016] Traditional retinal vessel segmentation methods rely on professional physicians to manually annotate images, which is time-consuming and laborious. The annotation results are affected by the physicians' experience and judgment. Especially in low-contrast regions and complex vascular structures, the segmentation results may be discontinuous or missed detections. To address this limitation, Ronneberger et al. proposed the U-Net model in 2015, which is the first deep learning architecture dedicated to biomedical image segmentation. U-Net achieves multi-scale feature extraction and precise pixel-level localization through a symmetric encoder-decoder structure and skip connections. Nevertheless, U-Net still has deficiencies in processing vessels with complex morphologies and low-contrast regions. Its segmentation accuracy for small vessels is relatively low, and gradient explosion or gradient vanishing problems may occur as the network depth increases. To address these deficiencies, subsequent studies have optimized the U-Net architecture. Gegundez-Arias et al. added batch normalization and residual modules to U-Net and introduced a new loss function to better capture the separability between pixels and the vascular tree. Although these improvements have enhanced the overall performance, the ability to capture small vessels is still limited. To solve this problem, DUNet replaces the standard convolutional layer with deformable convolutional blocks, enhancing the encoder's ability to capture context information. At the same time, by fusing low-level and high-level feature maps, the decoder achieves more precise localization, significantly improving the segmentation effect for complex vascular morphologies. Mou et al. added a dual self-attention mechanism between the encoder and decoder of U-Net, including spatial attention and channel attention modules, to adaptively integrate local features and global dependencies, enhancing the model's sensitivity to small vessels. These methods usually only add additional modules at the lowest scale stage, neglecting the processing ability of high-resolution details, resulting in insufficient precision in segmenting tiny vessels. To further improve the segmentation performance of small vessels, Kamran et al. proposed RV-GAN, a new multi-scale generative architecture that uses two generators and two multi-scale autoencoder discriminators to better localize and segment tiny vessels. Although RV-GAN demonstrates powerful vessel segmentation capabilities in the literature, its relatively low sensitivity indicates deficiencies in extracting small vessels. Zhou et al. designed a noisy label synthesis process and proposed a learning group learning scheme to improve the performance of models trained with imperfect labels, but these strategies mainly focus on improving label quality and fail to perform dynamic adaptive processing for the characteristics of different vascular morphologies and lesion regions.

[0017] According to the first embodiment of the present invention, the present invention claims a retinal vessel extraction method based on a strong feature dynamic fusion gating network, referring to Figure 1 and 2 , including: Collect the original image of the retina blood vessels to be extracted, and input the original image into the encoder to obtain multiple intermediate encoded features and a target encoded feature; Input the multiple intermediate encoded features obtained by skip connection into the feature navigation hub to obtain corresponding multiple fused encoded features, and input the target encoded feature into the parallel information focusing module to obtain a parallel information focusing feature; After skip connection of the multiple fused encoded features, input them together with the parallel information focusing feature into the decoder to obtain the integrated blood vessel feature; After each layer of decoder processing, enhance the feature fusion effect and gradually restore the image details to finally obtain the binary output image of the retina blood vessels.

[0018] The main objective of the present invention is to overcome the deficiencies of the prior art in retina blood vessel segmentation through multiple innovative technologies, especially for the problems of small blood vessel sensitivity, multi-scale adaptability, and lesion area interference. Through the three core modules of the adaptive gating residual block, parallel information focusing module, and feature navigation hub, the model has been comprehensively optimized and improved in data feature extraction, attention mechanism application, and multi-scale feature fusion, thus significantly improving the accuracy and robustness of retina blood vessel segmentation.

[0019] Aiming at the problems of lesion area interference and poor multi-scale blood vessel adaptability in retina blood vessel segmentation, the present invention proposes an adaptive gating residual block (AGRB). This module combines residual connection and adaptive gating mechanism, significantly improving the performance of the model in multi-scale blood vessel feature extraction and enhancement, thus enhancing the sensitivity to blood vessels of different scales. To further avoid overfitting, AGRB also introduces a spatial dropout module to enhance the generalization ability and robustness of the model. In the encoder part, AGRB is responsible for layer-by-layer extraction and compression of multi-scale features, while in the decoder, the adjusted AGRB is used to restore the image resolution, ensuring accurate segmentation of blood vessels of different scales and fine restoration of details.

[0020] In terms of feature fusion and global modeling, the present invention combines the feature navigation hub (FN-Hub) with a Transformer-based global attention mechanism to dynamically integrate multi-layer features of different scales in the encoder. This module can effectively capture the long-range dependence relationship between multi-layer features, promote the efficient fusion of global and local features, and greatly improve the model's ability to process the global context information of thick and thin blood vessels.

[0021] To address the problem of extracting fine-grained and complex features of blood vessels, the present invention designs a Parallel Information Focusing Module (PFAM). By fusing channel attention and spatial attention mechanisms, this module can enhance useful features in both the channel and spatial dimensions, thereby improving the feature expression ability and the ability to capture fine-grained features. Through this parallel optimization, PFAM effectively enhances the ability to extract fine-grained and complex features of blood vessels, further improving the accuracy and robustness of segmentation.

[0022] The synergistic effect of these modules significantly improves the performance of the model in complex blood vessel morphologies and low-contrast scenarios, making EFDG-UNet perform excellently in terms of small blood vessel sensitivity, lesion area robustness, and overall segmentation performance, especially demonstrating efficient and robust segmentation capabilities in complex blood vessel regions.

[0023] Furthermore, the step of collecting the original image of the retinal blood vessels to be extracted and inputting the original image into the encoder further includes: Adaptive Gated Residual Blocks are adopted, which combine residual connections and adaptive gating mechanisms, and introduce a spatial dropout module to mitigate overfitting; In the encoder, the Adaptive Gated Residual Block contains two convolutional layers and their corresponding batch normalization and activation layers.

[0024] Furthermore, the step of obtaining multiple intermediate encoded features and target encoded features further includes: The first feature of the original image is processed by the first convolutional layer, and through batch normalization and the ReLU activation function, the first intermediate encoded feature is obtained; The first intermediate encoded feature is processed again through the second convolutional layer and batch normalization to obtain the second intermediate encoded feature; The spatial dropout module processes the second intermediate encoded feature to reduce noise in the feature map and enhance the regularization effect; The adaptive gating mechanism of the Adaptive Gated Residual Block extracts global context information through global average pooling, and through two fully connected layers and the Sigmoid activation function, generates a gating coefficient for adjusting the weight of the residual term; The residual connection part performs channel matching on the first feature to ensure that the number of channels is consistent with the main branch feature, then performs weighted addition with the second intermediate encoded feature, and through ReLU activation, obtains the corresponding target encoded feature.

[0025] Among them, referring to Figure 3 , in this embodiment, in the encoder, the core structure of AGRB contains two convolutional layers and their corresponding batch normalization and activation layers. First, the input feature map Processed by the first convolutional layer, and through batch normalization and ReLU activation function to obtain intermediate features , and the process is shown in Equation (1).

[0026] (1); Next, the intermediate features are processed again through the second convolutional layer and batch normalization to obtain features , and the process is shown in Equation (2).

[0027] (2); The spatial dropout module processes to reduce the noise in the feature map and enhance the regularization effect. The adaptive gating mechanism of AGRB extracts global context information through global average pooling, and generates gating coefficients through two fully connected layers and Sigmoid activation function, as shown in Equation (3).

[0028] (3); Among them, ranges between and is used to adjust the weight of the residual term. The residual connection part performs channel matching on to ensure that its number of channels is consistent with the main branch features, and then performs weighted addition with and obtains the output through ReLU activation, as shown in Equation (4).

[0029] (4).

[0030] Furthermore, the step of inputting the target encoded features into the parallel information focusing module to obtain parallel information focusing features further includes: Using a channel attention module to aggregate the global features of each channel, generate channel importance weights, and adjust the channel expression ability of the input feature map; Input the target encoded features, extract the global descriptors of each channel through global average pooling operation, and generate channel weights through two fully connected networks to achieve feature compression and expansion; The generated channel weights are multiplied element-wise with the input feature map to complete the enhancement in the channel dimension; The spatial attention module captures global spatial information by performing max pooling and average pooling on the feature map in the spatial dimension, and generates spatial attention weights; Perform max pooling and average pooling operations on the first feature in the channel dimension to generate two spatial feature maps; Concatenate the two spatial feature maps in the channel dimension, and the result passes through a The convolutional layer generates spatial attention weights; The spatial weights are multiplied element-wise with the input feature map to complete the enhancement in the spatial dimension; After completing the attention optimization in the channel and spatial dimensions, the final parallel information focusing features are generated through a fusion operation.

[0031] Among them, referring to Figure 4 , in this embodiment, first, the input feature map extracts the global descriptor of each channel through global average pooling operation, denoted as . Subsequently, these descriptors go through a two-layer fully connected network to achieve feature compression and expansion, generating the channel weights , as shown in Equation 5.

[0032] (5); Among them, and are the weight matrices of the fully connected layers, is the channel compression rate, represents the ReLU activation function, is the Sigmoid activation function. The generated channel weights are multiplied element-wise with the input feature map to complete the enhancement in the channel dimension, as shown in Equation 6.

[0033] (6); The spatial attention module captures the global spatial information by performing max-pooling and average-pooling on the feature map in the spatial dimension, thereby generating the spatial attention weights. As Figure 6 shown, first, max-pooling and average-pooling operations are performed on the input feature map in the channel dimension to generate two spatial feature maps, as shown in Equation 7.

[0034] (7); Among them, . Then, the two are concatenated in the channel dimension, and the result passes through a convolutional layer to generate the spatial attention weights, as shown in Equation 8.

[0035] (8); Among them, represents the concatenation in the channel dimension, is a convolution operation, is the Sigmoid activation function. Finally, the spatial weights are multiplied element-wise with the input feature map to complete the enhancement in the spatial dimension, as shown in Equation 9.

[0036] (9); After completing the attention optimization of channels and space, the module generates the final output feature map through a fusion operation. The results of channel attention and spatial attention are first concatenated in the channel dimension, as shown in Equation (10).

[0037] (10); where . Then, the fused features pass through a convolutional layer for channel compression and are further optimized through the BatchNormalization and ReLU activation functions, as shown in Equation (11).

[0038] (11); where represents the ReLU activation function, and represents BatchNormalization. The final feature map is used for subsequent task processing.

[0039] The PFAM module independently optimizes channel attention and spatial attention to capture the channel global dependencies and spatial context information of the feature map respectively. The fusion operation further integrates the advantages of these two attention mechanisms to provide high-quality feature representations for subsequent tasks.

[0040] Furthermore, the step of inputting the multiple intermediate encoded features with skip connections into the feature navigation hub to obtain corresponding multiple fused encoded features further includes:[[]] Presetting the skip connection features of the encoder, and performing convolution, batch normalization, and ReLU activation operations on each intermediate encoded feature to obtain the uniformly processed features; In each layer of decoding, the intermediate encoded features are adjusted to the target resolution through bilinear interpolation; Using the attention mechanism of Transformer to capture the dependencies between multi-scale features, stacking all the adjusted features along a new dimension, and reshaping them into a sequence representation; Reshaping the features into a tensor of the target shape and further permuting the dimensions to meet the input requirements of the Transformer encoder; After passing through the Transformer encoder, the global context information can be integrated to output multiple fused encoded features.

[0041] Among them, referring to Figure 5 , in this embodiment, first, it is assumed that the skip connection features of the encoder are denoted as , where represents the number of encoding layers, and To unify the number of channels, we perform convolution, batch normalization, and ReLU activation operations on each feature map to obtain the feature after unified processing, as shown in Equation (12).

[0042] (12); Among them, is the processed feature, is the unified number of channels, , and respectively represent convolution, batch normalization, and the ReLU activation function.

[0043] In each layer of the decoder , to match the spatial dimensions of the decoder, we resize the feature map to the target resolution . This process is completed through bilinear interpolation, and its mathematical expression is shown in Equation (13).

[0044] (13); Among them, is the pixel value at position under the target resolution, are the values of the four nearest pixels to in the original feature map, is the interpolation weight, as shown in Equation (14).

[0045] (14); Next, we use the attention mechanism of the Transformer to capture the dependencies between these multi-scale features. First, stack all the resized features along a new dimension and reshape them into a sequence representation, as shown in Equation (15).

[0046] (15); Among them, is the stacked multi-scale feature, is the number of feature layers, is the unified number of channels, is the target resolution.

[0047] We reshape the feature into a tensor with the shape of and further permute the dimensions to to meet the input requirements of the Transformer encoder.

[0048] After passing through the Transformer encoder, the features are integrated with global context information, and the output representation is shown in Equation (16).

[0049] (16); where, is the output feature after encoding, containing multi-scale global context information.

[0050] We restore the output feature to the original spatial dimension , and perform an average operation on the encoding layer dimension to obtain the fused feature, as shown in Equation (17).

[0051] (17); where, is the fused feature, represents the layer encoding feature fusion result.

[0052] The fused feature is processed through convolution, batch normalization, and ReLU activation to match the number of channels required by the decoder , generating the final feature map , as shown in Equation (18).

[0053] (18); where, is the generated decoder feature, is the target number of channels.

[0054] In the Transformer module, we adopt the multi-head self-attention mechanism to model the complex dependencies between features. The structure diagram is shown in Figure 6 . Let the input feature be represented as , where , . The core calculation formula of multi-head self-attention is shown in Equation (19).

[0055] (19); where, , , , , , are the linear transformation weight matrices of query, key, and value respectively, is the dimension of the key vector. Through this process, the decoder can more accurately utilize multi-scale information to generate a segmentation result with rich details and clear semantics.

[0056] Further, the step of inputting the fused encoded features after skip connection and the parallel information focused feature into the decoder to obtain the integrated vascular features further includes: Using an adaptive gated residual upsampling block that includes upsampling and skip connection operations; The input parallel information focused feature is upsampled to half the size of the original image through bilinear interpolation, and then concatenated with the feature map of the corresponding layer in the feature navigation hub to integrate the semantic information in the encoding stage; The concatenated feature map is further processed by an adaptive gated residual block to enhance the feature fusion effect and restore the image details to obtain the integrated vascular features.

[0057] Among them, referring to Figure 7 , in this embodiment, first, the input feature map is upsampled to half the size of the original image through bilinear interpolation, and then concatenated with the feature map of the corresponding layer in the FN-Hub to integrate the semantic information in the encoding stage. The concatenated feature map is further processed by the AGRB module to enhance the feature fusion effect and restore the image details. The adaptive gating mechanism in the UAGRB dynamically adjusts the weights of the residual connections, thereby optimizing the importance of different features in the upsampling process and enabling the network to better retain fine structures.

[0058] For the embodiments of the present invention, after extracting the retinal blood vessels, an evaluation of the extraction algorithm will be performed; In the image segmentation task, in order to more comprehensively evaluate the segmentation effect, the following performance metrics are used: intersection over union (IoU), accuracy (ACC), sensitivity (Se), specificity (Sp), F1-score (F1), and area under the curve (AUC), and the specific definitions are shown in Formulas 20-21.

[0059] (20); (21); (22); (23); Among them, TP, TN, FP, and FN respectively represent true positive, true negative, false positive, and false negative.

[0060] The meanings of each index are as follows: ACC represents the proportion of pixels correctly classified. Se (sensitivity) reflects the ability of the model to identify positive class pixels, while Sp (specificity) represents the correct recognition rate of the model for negative class pixels. The F1-score combines precision and sensitivity to evaluate the comprehensive classification performance of the model. AUC measures the overall performance of the model at different discrimination thresholds.

[0061] Referring toFigure 8 In the experiment, the DRIVE and CHASE_DB1 datasets were used for training and testing. These two datasets are public standard datasets for retinal vessel segmentation.

[0062] The DRIVE dataset contains 40 RGB color fundus images, each with a resolution of 565x584. We divided the training set and validation set from the official training set according to a ratio of 8:2, and the remaining 20 test sets remained unchanged. The CHASE_DB1 dataset contains 28 RGB color fundus images, each with a resolution of 999x960. We divided the data into a training set, a validation set, and a test set according to a ratio of 7:2:1, with the first 19 as the training set and 6 as the validation set, and randomly selected 3 as the test set.

[0063] In medical image segmentation tasks, appropriate preprocessing can significantly improve the performance of the model. The dataset was successively subjected to green channel extraction, contrast-limited adaptive histogram equalization (CLAHE), and gamma correction to optimize the quality of color retinal images.

[0064] To reduce the color differences caused by different acquisition devices or lighting conditions, we first extracted the green channel from the RGB images. The green channel usually has higher contrast, which is particularly beneficial for the extraction of retinal blood vessels. Subsequently, we converted the RGB image to a grayscale image through the formula to further reduce the interference of color information on segmentation: (24); where represents the pixel intensity of the grayscale image, , and respectively represent the pixel intensities of the red, green, and blue channels.

[0065] Next, to enhance the contrast between blood vessels and the background in the image, we applied the CLAHE operation. CLAHE can suppress noise while enhancing contrast. This method sets a threshold, evenly distributes the pixel values exceeding the threshold to other gray levels to form a new histogram, and performs adaptive equalization on it. We set the threshold to 2 and selected a grid size of 8×8, and divided the image into non-overlapping regions and processed them separately to ensure the stability and consistency of the results.

[0066] To further highlight the brightness difference between blood vessels and surrounding tissues, we used gamma correction technology. Gamma correction changes the brightness distribution by adjusting the gamma curve of the image, thereby enhancing the contrast in low-brightness and high-brightness regions. The mathematical expression of gamma correction is: (25); Among them, and represent the output and input pixel intensities respectively, is the gamma value. In this study, we set the gamma value to 0.8 to enhance the contrast in low-brightness regions, thereby highlighting the details of retinal blood vessels. After gamma correction, the brightness difference between the blood vessel region and the background increases significantly, especially the structure of small blood vessels becomes more clearly visible.

[0067] After a series of preprocessing operations, the originally low-contrast and color-blurred retinal images have been significantly improved. The processed images not only have more prominent blood vessel details, but also the contrast between blood vessels and the surrounding non-blood vessel regions is significantly enhanced. To increase the diversity of training data, we crop the preprocessed images into non-overlapping small blocks of 48×48. This slicing method can increase the number of training samples, alleviate the overfitting problem of the deep learning model, and at the same time provide more diverse feature samples for the model, thereby improving the performance of the retinal blood vessel segmentation task.

[0068] In terms of accuracy (Acc), EFDG-UNet reaches 0.9736, which is competitive compared with current high-performance models (such as UNet3+ and SDDC-Net), and is significantly higher than traditional U-Net (0.9691) and LadderNet (0.9561), indicating its precision in the overall segmentation task. The specificity (Sp) is 0.9856, also in a leading position, second only to a few models such as ResUNet++ (0.9837), reflecting its efficient recognition ability for background regions.

[0069] In terms of sensitivity (Se), the performance of EFDG-UNet (0.8438) is significantly better than classical methods (such as 0.7948 of U-Net) and some modern models (such as 0.7856 of ResUNet++), indicating its greater advantage in capturing small blood vessels. Although it is slightly lower than SDDC-Net (0.8603) and DCA-CNN (0.8745) in sensitivity, its high balance of specificity and accuracy makes its overall performance more balanced.

[0070] The AUC (0.9886) and F1 value (0.8412) of EFDG-UNet are both among the top. Among them, AUC reflects the robustness of the segmentation model at different thresholds, while the F1 value reflects the coordination ability of precision and recall. Compared with other recently proposed models (such as the F1 value of UNet3+ is 0.8254), the segmentation performance of EFDG-UNet shows better comprehensiveness.

[0071] EFDG-UNet performs excellently in multiple key metrics, especially in terms of accuracy (Acc) and comprehensive segmentation performance, taking the lead. Its relatively high sensitivity (Se) and specificity (Sp) values indicate that EFDG-UNet can well balance the segmentation capabilities for small blood vessels and background regions. The excellent performance of its AUC and F1 metrics further validates the robustness and reliability of the model, making it one of the leading methods for retinal vessel segmentation tasks.

[0072] EFDG-UNet shows significant advantages in segmenting complex vascular regions. In the segmentation of thick blood vessels, EFDG-UNet demonstrates higher coherence and integrity. Compared with other models, EFDG-UNet can accurately maintain the overall shape of blood vessels. In the segmentation of tiny blood vessels, EFDG-UNet can not only accurately identify blood vessel branches and intersections but also significantly reduce misidentifications, avoiding mis-segmenting background regions as blood vessels. Generally speaking, EFDG-UNet has stronger robustness and higher segmentation accuracy in the fine-grained blood vessel segmentation task.

[0073] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0074] In addition, the functional units in each embodiment of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. The above is only the implementation mode of this application, and does not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of this application.

[0075] The specific embodiments of the invention have been described in detail above, but they are only examples, and the present application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of the present application. Therefore, equivalent transformations, modifications, improvements, etc. made without departing from the spirit and principles of the present application should all be covered within the scope of the present application.

Claims

1. A retinal vessel extraction method based on a strong feature dynamic fusion gating network, characterized in that, Including: Collect the original image of the retina blood vessels to be extracted, input the original image into the encoder, and obtain multiple intermediate encoded features and a target encoded feature; Input the multiple intermediate encoded features obtained by skip connection into the feature navigation hub to obtain corresponding multiple fused encoded features, and input the target encoded feature into the parallel information focusing module to obtain parallel information focusing features; Input the multiple fused encoded features obtained by skip connection and the parallel information focusing features into the decoder together to obtain the integrated blood vessel features; After each layer of decoder processing, enhance the feature fusion effect and gradually restore the image details to finally obtain the binary output image of the retina blood vessels.

2. The retinal vessel extraction method based on a strong feature dynamic fusion gating network according to claim 1, wherein The step of collecting the original image of the retina blood vessels to be extracted and inputting the original image into the encoder further includes: Adopt an adaptive gated residual block, combine the residual connection and the adaptive gating mechanism, and introduce a spatial dropout module to reduce overfitting; In the encoder, the adaptive gated residual block includes two convolutional layers and their corresponding batch normalization and activation layers.

3. The retinal blood vessel extraction method based on a strong feature dynamic fusion gating network according to claim 2, wherein The step of obtaining multiple intermediate encoded features and a target encoded feature further includes: The first feature of the original image is processed by the first convolutional layer, and after batch normalization and ReLU activation function, the first intermediate encoded feature is obtained; The first intermediate encoded feature is processed again by the second convolutional layer and batch normalization to obtain the second intermediate encoded feature; The spatial dropout module processes the second intermediate encoded feature to reduce the noise in the feature map and enhance the regularization effect; The adaptive gating mechanism of the adaptive gated residual block extracts global context information through global average pooling, and generates a gating coefficient through two fully connected layers and a Sigmoid activation function, which is used to adjust the weight of the residual term; The residual connection part performs channel matching on the first feature to ensure that the number of channels is consistent with the main branch feature, then performs weighted addition with the second intermediate encoded feature, and obtains the corresponding target encoded feature through ReLU activation.

4. A method for extracting retinal blood vessels based on a strong feature dynamic fusion gating network according to claim 3, wherein The step of inputting the target encoded feature into the parallel information focusing module to obtain parallel information focusing features further includes: Adopt a channel attention module to aggregate the global features of each channel, generate channel importance weights, and adjust the channel expression ability of the input feature map; Input the target encoded feature, extract the global descriptor of each channel through global average pooling operation, and perform feature compression and expansion through two fully connected networks to generate channel weights; The generated channel weights are multiplied element-wise with the input feature map to complete the enhancement in the channel dimension; The spatial attention module captures global spatial information through max-pooling and average-pooling of the feature map in the spatial dimension, and generates spatial attention weights; Perform max-pooling and average-pooling operations on the first feature in the channel dimension to generate two spatial feature maps; Concatenate two spatial feature maps in the channel dimension, and the result passes through a convolutional layer to generate spatial attention weights; The spatial weights are multiplied element-wise with the input feature map to complete the enhancement in the spatial dimension; After completing the attention optimization in the channel and spatial dimensions, generate the final parallel information focusing features through a fusion operation.

5. A method for extracting retinal blood vessels based on a strong feature dynamic fusion gated network according to claim 4, characterized in that The step of jump-connecting the multiple intermediate encoded features into the input feature navigation hub to obtain corresponding multiple fused encoded features further includes: The skip connection features of the preset encoder are used to perform convolution, batch normalization, and ReLU activation operations on each intermediate encoded feature to obtain the uniformly processed features; In each layer of decoding, bilinear interpolation is used to adjust the intermediate encoded features to the target resolution; The attention mechanism of Transformer is used to capture the dependencies between multi-scale features. All the adjusted features are stacked along a new dimension and reshaped into a sequence representation; The features are reshaped into a tensor of the target shape and the dimensions are further permuted to meet the input requirements of the Transformer encoder; After passing through the Transformer encoder, global context information is integrated to output multiple fused encoded features.

6. The retinal blood vessel extraction method based on a strong feature dynamic fusion gating network according to claim 5, wherein The step of jump-connecting the multiple fused encoded features and inputting them together with the parallel information focusing features into the decoder to obtain the integrated vascular features further includes: An adaptive gated residual upsampling block is adopted, which includes upsampling and jump-connecting operations; The input parallel information focusing features are upsampled to half the size of the original image by bilinear interpolation, and then concatenated with the feature maps of the corresponding layers in the feature navigation hub to integrate the semantic information in the encoding stage; The concatenated feature maps are further processed by the adaptive gated residual block to enhance the feature fusion effect and restore the image details to obtain the integrated vascular features.

Citation Information

Patent Citations

  • Eye fundus blood vessel image segmentation method based on visual attention fusion network

    CN117523202A

  • Retinal blood vessel image segmentation method based on multi-stage feature analysis

    CN117635642A

  • Retinal vessel segmentation method and system based on multiple attention mechanisms

    CN118840550A

  • Ironmaking product quality evaluation method and device based on multi-source data

    CN119624970A

Cited By

  • Dense pedestrian detection method and system based on improved D-FINE-N network

    CN122493497A