Eye fundus image blood vessel segmentation method and system
Through the dual-branch feature extraction architecture and topological structure constraint mechanism, the problems of insufficient accuracy of vascular segmentation and insufficient maintenance of topological structure in the prior art are solved, and the vascular segmentation effect with high accuracy and high reliability are achieved.
Patent Information
- Application Number
- CN202510161066.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has insufficient accuracy when dealing with complex vascular structures, lacks topological maintenance, and poor adaptability to vascular systems at different scales.
A two-branch feature extraction architecture is adopted, and macroscopic details are processed by global branches and local branches respectively, combined with multi-head self-attention mechanism and position coding, a topological structure constraint mechanism based on persistence homomodulation theory is introduced to achieve cross-scale feature fusion and structural optimization.
It significantly improves the accuracy and reliability of vascular segmentation, effectively maintains the structural integrity and topological characteristics of the vascular network, and reduces the fracture and missegment in the segmentation results.
Smart Images

Figure CN120071409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and medical image analysis, and particularly relates to a method and system for segmenting fundus image blood vessels. Background Art
[0002] Fundus vessel segmentation is of great significance for the early diagnosis and prevention of eye diseases. In recent years, with the development of deep learning technology, significant progress has been made in the research of fundus vessel segmentation. The existing fundus vessel segmentation methods can be mainly divided into traditional methods and deep learning-based methods. Among these existing methods, Ronneberger O et al. first proposed the U-Net network structure, which is a milestone medical image segmentation architecture. The biggest innovation of this method lies in the design of a symmetric encoder-decoder structure and the effective fusion of multi-scale features through the skip connection mechanism. In the network structure design, the encoding path adopts the classic convolutional neural network architecture. Through successive convolutional operations and max-pooling operations, the input image is gradually transformed into a feature map with rich semantic information, while gradually reducing the spatial resolution of the feature map. The decoding path uses upsampling and transposed convolution operations to reconstruct the segmentation result by gradually restoring the spatial resolution of the feature map. Its innovative skip connection design enables the network to utilize both the fine spatial information of the lower layer and the abstract semantic information of the higher layer by directly connecting the low-level features in the encoding path to the corresponding layers in the decoding path, significantly improving the accuracy of medical image segmentation. Based on the pioneering work of U-Net, Guo C et al. proposed SA-UNet, which effectively enhanced the feature representation of the vascular region by introducing a spatial attention mechanism. The core innovation of this method lies in the design of a dedicated spatial attention module, which can automatically learn to generate an attention weight map to adaptively adjust the importance weights of different spatial positions in the feature map. Specifically, the spatial attention module first extracts spatial attention features through a series of convolutional operations, then generates a normalized attention weight map through the softmax function, and finally multiplies this weight map with the original feature map to achieve the adaptive enhancement of the vascular region features. At the same time, this method also innovatively optimizes the design of the skip connection by introducing a channel attention mechanism to achieve the adaptive selection and recombination of feature channels. The design of this dual attention mechanism significantly improves the network's ability to identify small blood vessels. TopoNet proposed by T Li et al. is a pioneering end-to-end deep learning framework that first applies graph neural network technology to the topological structure inference task in complex scenarios. The core innovation points of this method are mainly reflected in three aspects: first, an innovative feature embedding module is designed, which can map input information of different modalities to a unified feature space to achieve the effective fusion of multi-source information; second, a scene knowledge graph is constructed based on prior knowledge, which guides the propagation and update of features by encoding the topological relationships between objects in the scene; finally, an adaptive information flow mechanism is proposed, which can dynamically adjust the way and intensity of feature propagation according to the complexity of the scene.These innovative designs enable TopoNet to more accurately maintain and infer topological relationships in complex scenarios, demonstrating significant performance advantages in challenging tasks such as autonomous driving scene understanding. The DeepVessel method proposed by Fu H et al. innovatively combines deep learning with conditional random fields (CRF), presenting a novel hybrid model architecture. An important innovation of this method is the design of a multi-scale deep supervision network, which realizes effective learning and supervision of vascular features at different scales by strategically adding auxiliary loss functions at different levels of the network. In terms of the model architecture, this method first uses CRF as a post-processing module of the deep network, optimizing the initial segmentation results by modeling the spatial relationships between pixels, effectively improving the spatial continuity and accuracy of segmentation. In addition, this method also innovatively proposes a new weighted cross-entropy loss function, which effectively solves the common class imbalance problem in the vascular segmentation task by adaptively adjusting the weights of different category samples, significantly enhancing the model's detection ability for small blood vessels.
[0003] However, the existing methods still have the following problems: 1) Most of the existing technologies focus on using spatial attention mechanisms to improve the accuracy of vascular segmentation, which is based on the consideration that the vascular structure has local spatial continuity characteristics. However, due to the complexity of the fundus vascular network structure, methods relying only on local features are prone to missegmentation in vascular bifurcation and crossing regions. Especially when dealing with fine blood vessels, local features often cannot provide sufficient discriminative information, resulting in broken or missed detections in the segmentation results. This situation is particularly obvious in fundus images with poor image quality or severe lesions.
[0004] 2) Existing technologies usually need to extract multi-scale features through deep convolutional neural networks to handle vascular segmentation tasks of different thicknesses. However, when the vascular network presents a highly complex hierarchical structure, the performance of these methods is often significantly affected. A common improvement scheme is to fuse features at different levels through skip connections, but this simple feature fusion method is difficult to effectively handle the topological structure relationships of blood vessels. Although SA-UNet improves the feature representation ability by introducing a spatial attention mechanism, its attention module only focuses on feature enhancement in local regions and fails to fully utilize the overall topological information of the vascular network. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for fundus image vascular segmentation, which are beneficial to improving the accuracy and reliability of vascular segmentation of fundus images.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is: A method for fundus image vascular segmentation, comprising: Construct a fundus image vascular segmentation model, where the fundus image vascular segmentation model includes a dual-branch feature extraction module, a feature fusion and topological constraint module, a multi-task joint optimization module, and a post-processing optimization module; the dual-branch feature extraction module includes a global feature extraction module and a local feature extraction module. For the input fundus image, the global feature extraction module performs initial feature extraction through a deep convolutional network, and combines position encoding enhancement and a multi-head self-attention mechanism to capture the overall topological structure of the vascular network. The local feature extraction module performs local feature extraction based on an improved U-Net architecture, and maintains the fine structural features of blood vessels through multi-scale feature extraction and skip connection mechanisms; the feature fusion and topological constraint module realizes the adaptive fusion of global features and local features at different scales through a cross-scale feature fusion strategy, and at the same time introduces a topological constraint mechanism based on persistent homology theory, and maintains the structural integrity of the vascular network through topological graph construction and connectivity analysis, and finally outputs an initial segmentation result; the multi-task joint optimization module comprehensively considers topological loss, segmentation loss, and contrast loss, and balances different optimization objectives through a dynamic weight adjustment mechanism; among them, the topological loss ensures the structural integrity of the vascular network, the segmentation loss optimizes the classification accuracy at the pixel level, and the contrast loss enhances the discriminative ability of feature representation; the post-processing optimization module converts the initial segmentation result into a graph structure representation, then performs structural optimization through a multi-layer graph convolutional network, and finally uses conditional random fields to achieve fine processing of the boundary and outputs an optimized segmentation result; Train the constructed fundus image vascular segmentation model to obtain a trained fundus image vascular segmentation model; Use the trained fundus image vascular segmentation model to perform vascular segmentation on the input fundus image to obtain corresponding vascular segmentation results.
[0007] Furthermore, the implementation method of the global feature extraction module is as follows: In the initial feature extraction stage, perform feature mapping on the input fundus image through a deep convolutional neural network to extract the initial feature map of the image; at the same time, introduce a position encoding mechanism to inject position encoding information into the feature map to better perceive the spatial distribution features of blood vessels: Among them, F init is the extracted initial feature map, F g ' is the feature representation after position encoding enhancement, and ξ( ) is the position encoding; In the feature enhancement stage, construct a feature modeling module based on the multi-head self-attention mechanism. The feature modeling module captures the global association between features through h parallel attention heads, and each attention head H i independently learns the interaction relationship between features: Among them, ΛQ i , ΛK i and ΛV i are the linear transformation matrices of query, key, and value respectively, and σ( ) is the softmax normalization function, which prevents the vanishing gradient problem through the scaling factor ; finally, the outputs of all attention heads are concatenated and passed through the linear transformation Ω to obtain the fused global features: Among them, H h is the h-th attention head; To achieve dynamic optimization of features, a feature optimization feedback path is constructed inside the global feature extraction module. Through the backpropagation of the optimization signal, the feature extraction and attention calculation processes are dynamically adjusted to further improve the discriminative ability of feature representation.
[0008] Furthermore, the implementation method of the local feature extraction module is as follows: The local feature extraction module effectively captures different-scale features of fundus blood vessels through a multi-level encoder-decoder structure and a multi-scale feature extraction module; In the encoder part, a three-layer encoder structure is constructed, corresponding to feature extraction at scales of 1 / 1, 1 / 2, and 1 / 4 respectively; each layer of the encoder is implemented through a combination of depthwise separable convolution and residual connection, ensuring effective feature extraction while reducing computational complexity; In the decoder part, a three-layer decoder structure is correspondingly adopted to restore the spatial resolution of the feature map through step-by-step upsampling operations; in each layer of the decoding process, the feature extraction process is expressed as: Among them, f i represents the output feature of the i-th layer decoder; η i is a feature conversion function based on depthwise separable convolution, which is used to process the input features, f i-1 is the output feature of the (i - 1)-th layer decoder, and F i skip is the skip connection feature from the i-th layer encoder; The multi-scale feature extraction module outputs local feature sequences F l ={f 1 , f 2 , f 3} corresponding to different scales at three encoder layers respectively, where f 1 , f 2 , and f 3 correspond to features at scales of 1 / 1, 1 / 2, and 1 / 4 respectively; Each element in the local feature sequence undergoes feature concatenation and feature refinement processing; feature concatenation adopts a skip connection mechanism to achieve effective fusion of high-level and low-level features: Among them, δ( ) represents the feature fusion function, which is used to adaptively integrate feature information at different scales and effectively fuse the features at three scales of 1 / 1, 1 / 2, and 1 / 4; Feature refinement introduces a feature guidance mechanism to dynamically adjust the feature extraction process by optimizing the backpropagation of signals; this feature guidance mechanism not only acts between the layers of the encoder but also further optimizes the extracted features through the feature refinement unit; the feature refinement process is expressed as: Among them, i = 1, 2, 3, γ( ) represents the feature refinement function, which is used to perform a non-linear transformation on the input features to extract more discriminative representations; β is a learnable scaling parameter used to adaptively adjust the importance of features; Finally, the local feature sequence F output by the local feature extraction module is obtained l ={f 1 ', f 2 ', f 3 '}.
[0009] Furthermore, the implementation method of the feature fusion and topological constraint module is as follows: The global feature F output by the global feature extraction module g and the local feature sequence F output by the local feature extraction module l are input into the feature fusion and topological constraint module, which includes two parts: cross-scale feature fusion and topological structure constraint; In the cross-scale feature fusion stage, first, the global feature and the local feature are preliminarily integrated through the feature fusion module, and then an adaptive weight mechanism based on channel attention is introduced for feature enhancement; the fusion process is expressed as: Among them, F g is the global feature, f i ' ∈ F l is the i-th feature in the local feature sequence F l , and β( ) and γ( ) respectively represent the channel attention functions for the global feature and the local feature, and ⊙ represents the element-wise product operation; the channel attention weights are dynamically learned through the attention weight calculation module: Among them, Wα and W α are learnable weight matrices, b α and b α are the corresponding bias terms, and σ( ) is the sigmoid activation function, which is used to normalize the attention weights to the interval [0,1]; in order to make full use of the statistical information of the features, the information of global average pooling and max pooling is considered simultaneously in the pooling operation Pool(): where x represents the input feature map, AvgPool( ) represents average pooling, MaxPool( ) represents max pooling, and concat() represents feature concatenation; In the topological structure constraint stage, a topological constraint mechanism based on persistent homology theory is introduced. First, the feature representation is converted into a topological space through the topological graph construction unit, and then the topological graph is made to maintain the topological characteristics of the blood vessel network through connectivity analysis and structure optimization; the construction process of the topological graph is expressed as: where (b k , d k ) represents the lifespan of the k-th topological feature, including its birth time b k and death time d k ; based on the constructed topological graph, a topological loss function is constructed to constrain the structural integrity of the segmentation result: where S represents the segmentation result predicted by the model, Y represents the true label, PH 0 ( ) represents the operator for calculating topological features, and D( ) represents the function for measuring the distance between two sets of topological features; To achieve the collaborative optimization of feature fusion and topological constraint, a structural feedback mechanism is constructed inside the feature fusion and topological constraint module. This structural feedback mechanism feeds back the optimization signal of topological constraint to the feature fusion process and dynamically adjusts the feature fusion strategy through topological guidance.
[0010] Furthermore, the implementation method of the multi-task joint optimization module is as follows: The multi-task joint optimization module realizes the overall optimization of the model performance through the collaborative action of three parts: topological loss, segmentation loss, and contrast loss; the overall loss function is expressed as: where L seg represents the segmentation loss, which is used to optimize the classification accuracy at the pixel level; L topo represents the topological loss, which is used to maintain the structural integrity of the blood vessel network; L contrDenotes the contrast loss, which is used to enhance the discriminative ability of feature representations; λ 1 , λ 2 and λ 3 are the corresponding weight coefficients, which are used to balance the contributions of various losses; For the topological loss, two key constraints are designed based on the persistent homology theory: structural consistency and connectivity constraints; the topological loss function is defined as: where S represents the segmentation result predicted by the model, Y represents the ground truth label, and PH 0 ( ) represents the operator for calculating topological features, and D( ) represents the function for measuring the distance between two sets of topological features; b i and b i ' respectively represent the birth times of the i-th topological feature in the prediction result and the ground truth label, and μ is the balance coefficient, which is used to adjust the relative importance of the structural consistency constraint and the connectivity constraint; The topological loss function not only considers the topological differences between the prediction result and the ground truth label, but also enhances the stability of topological features through the life cycle regularization term; For the segmentation loss, cross-entropy loss and Dice loss are adopted to achieve: where L CE represents the cross-entropy loss, which is used to optimize the classification accuracy at the pixel level; L Dice represents the Dice coefficient loss, which is used to measure the overlap degree between the predicted segmentation result and the ground truth label; α is the weight coefficient, which is used to balance the relative importance of the two losses, 0 ≤ α ≤ 1, and the best balance between the pixel-level accuracy and the regional overlap degree is achieved by adjusting α; Among them, the cross-entropy loss is improved through the Focal loss mechanism to enhance the learning ability for difficult-to-classify samples: where y i represents the ground truth label of the i-th pixel, p i represents the probability that the model predicts that this pixel belongs to a blood vessel, and ∑ represents the summation over all pixel positions; For the contrast loss, by constructing positive and negative sample pairs, the model is promoted to learn more discriminative feature representations: where s + represents the similarity score between positive sample pairs, s - represents the similarity score between negative sample pairs, τ represents the temperature coefficient, which is used to adjust the smoothness of the feature distribution; ∑ represents the summation over the similarity scores of all negative sample pairs; To achieve the dynamic balance among loss functions, an adaptive weight adjustment mechanism based on gradient statistics is constructed; the weight coefficient is updated in the following way: where λ i represents the weight coefficient of the i-th loss function, θ i is the parameter value corresponding to the i-th loss function, and ∑exp(θ j ) represents the normalization of the parameters of all loss functions; where the update of the parameter θ i considers the gradient correlation between loss functions: where η is the learning rate, and represent the gradients of the i-th and j-th loss functions respectively, ρ( ) represents the function for calculating gradient correlation, and σ( ) is the activation function used to map the correlation to an appropriate range; To ensure the stability of the training process, a smoothing mechanism based on exponential moving average is introduced: where represents the smoothed weight of the i-th loss function at the t-th moment, is the weight value at the previous moment, and β is the smoothing coefficient used to control the influence degree of the historical weight value, 0 ≤ β ≤ 1.
[0011] Furthermore, the implementation method of the post-processing optimization module is as follows: The post-processing optimization module performs refined processing on the initial segmentation result through optimization in three stages: graph structure construction, graph convolution optimization, and conditional random field refinement; In the graph structure construction stage, an initial graph structure is first constructed through node feature extraction and edge feature calculation; for each superpixel region p i extracted from the segmentation result, the construction process of the node feature is expressed as: where p i represents the i-th superpixel region, f i is the corresponding depth feature, v i is the constructed node feature vector, is the feature fusion function used to fuse the position information and depth feature, and P is the set of all superpixel regions; the weight of the edge is calculated by comprehensively considering the spatial relationship and feature similarity between nodes: where e ijdenotes the weight of the edge between node i and node j, v i and v j are the feature vectors of two nodes respectively, σ is the Gaussian kernel parameter that controls the influence degree of feature similarity, d i,j represents the spatial distance between two nodes, r is the preset distance threshold, δ( ) is the indicator function, which takes the value of 1 when the condition holds, otherwise 0; In the graph convolution optimization stage, a multi-layer graph convolution network is constructed for structure optimization; the feature update rule for each layer is: where, H l+1 and H l represent the node feature matrices of the (l + 1)-th layer and the l-th layer respectively, σ( ) is the non-linear activation function, is the adjacency matrix with self-connection added, is the corresponding degree matrix, W l is the learnable weight matrix of the l-th layer; where this term realizes the symmetric normalization of the adjacency matrix; To enhance the effectiveness of feature propagation, an attention mechanism is introduced to dynamically modulate message passing: where, h i and h j represent the feature vectors of node i and node j respectively, h i ' is the updated feature representation of node i, W α is the attention weight matrix, W is the feature transformation matrix, g( ) is the attention scoring function used to calculate feature similarity, softmax( ) is used to normalize the attention scores to the [0, 1] interval, σ( ) is the non-linear activation function; N i represents the set of neighbor nodes of node i, α ij represents the attention weight from node j to node i; In the conditional random field refinement stage, an improved energy function is constructed: where, E(y) is the energy function, y represents the segmentation label configuration, is the unary potential function, representing the local confidence of the y i label of node i, is the pairwise potential function used to model the label dependency between adjacent nodes i and j; the pairwise potential function comprehensively considers spatial position, color features and edge information: Among them, μ(y i , y j ) is a label compatibility function used to measure the compatibility degree between adjacent node labels y i and y j ; the spatial position coordinates p i and p j represent the position information of nodes i and j in the image, I i and I j represent the color feature vectors of nodes i and j respectively; w 1 and w 2 represent weight coefficients; the Gaussian kernel parameters σ α and σ β respectively control the influence ranges of spatial distance and color similarity, and are used to calculate the distance metric between nodes.
[0012] The present invention also provides a fundus image vascular segmentation system, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above-mentioned method can be implemented.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a fundus image vascular segmentation method and system to solve the problems of insufficient accuracy, lack of topological structure preservation, and poor adaptability to blood vessels of different scales in the prior art when dealing with complex vascular structures. The present invention proposes a vascular segmentation architecture based on dual-branch feature extraction, which processes macroscopic structures and microscopic details through a global branch and a local branch respectively; the global branch introduces a multi-head self-attention mechanism and position encoding to enhance the ability to understand the overall topological structure of blood vessels; the local branch adopts an improved U-Net structure to accurately extract the detailed features of blood vessels. This dual-branch structure can effectively balance global semantic information and local geometric features, significantly improving the accuracy of vascular segmentation. At the same time, in order to ensure that the segmentation result meets the physiological characteristics of vascular connectivity, the present invention innovatively introduces a topological structure constraint mechanism based on persistent homology theory. This mechanism directly realizes end-to-end topological structure constraint in the deep learning framework by constructing a filter complex and calculating persistent homology features. By integrating the topological loss function into the network training process, the connectivity and hierarchical relationship of the vascular network are effectively maintained, and the cases of breakage and mis-segmentation in the segmentation result are significantly reduced. In addition, the present invention proposes a cross-scale feature fusion strategy to dynamically adjust the importance of features at different scales through an adaptive weight mechanism; this strategy not only realizes the effective fusion of features of the global branch and the local branch, but also enhances the recognition ability of blood vessels of different thicknesses through the attention mechanism. Combining the post-processing optimization of the graph convolutional network and the conditional random field further improves the accuracy and reliability of the segmentation result. Brief Description of the Drawings
[0014] Figure 1 It is a schematic diagram of the overall architecture of the fundus image vascular segmentation model in the embodiment of the present invention; Figure 2 It is a schematic diagram of the structure of the global feature extraction module in the embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of the local feature extraction module in the embodiment of the present invention; Figure 4 It is a schematic diagram of the structure of the feature fusion and topological constraint module in the embodiment of the present invention; Figure 5 It is a schematic diagram of the structure of the multi-task joint optimization module in the embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of the post-processing optimization module in the embodiment of the present invention. Detailed implementation manners
[0015] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0016] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0017] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0018] This embodiment provides a fundus image vascular segmentation method, which realizes high-precision segmentation of blood vessels in fundus images by innovatively combining technologies such as dual-branch feature extraction, topological structure constraint, and multi-task joint optimization. The method specifically includes: S1. Construct a fundus image vascular segmentation model.
[0019] As Figure 1 shown, the fundus image vascular segmentation model includes a dual-branch feature extraction module, a feature fusion and topological constraint module, a multi-task joint optimization module, and a post-processing optimization module.
[0020] The dual-branch feature extraction module includes a global feature extraction module and a local feature extraction module. For the input fundus image, the global feature extraction module performs initial feature extraction through a deep convolutional network, and combines positional encoding enhancement and multi-head self-attention mechanism to capture the overall topological structure of the vascular network. The local feature extraction module performs local feature extraction based on an improved U-Net architecture, and maintains the fine structural features of blood vessels through multi-scale feature extraction and skip connection mechanism. This dual-branch design can take into account both the macroscopic structure and microscopic details of the vascular network, laying a solid foundation for subsequent feature fusion and optimization.
[0021] The feature fusion and topological constraint module realizes the adaptive fusion of global features and local features at different scales through a cross-scale feature fusion strategy. At the same time, it introduces a topological constraint mechanism based on persistent homology theory, and maintains the structural integrity of the vascular network through topological graph construction and connectivity analysis, effectively avoiding the phenomena of breaks and mis-segmentation in the segmentation results, and finally outputs the initial segmentation result.
[0022] To achieve the overall optimization of the model performance, the multi-task joint optimization module comprehensively considers the topological loss, segmentation loss and contrast loss, and balances different optimization objectives through a dynamic weight adjustment mechanism; among them, the topological loss ensures the structural integrity of the vascular network, the segmentation loss optimizes the classification accuracy at the pixel level, and the contrast loss enhances the discriminative ability of feature representation. This multi-task collaborative optimization strategy significantly improves the generalization ability and robustness of the model.
[0023] The post-processing optimization module converts the initial segmentation result into a graph structure representation, then performs structural optimization through a multi-layer graph convolutional network, and finally uses conditional random fields to achieve refined boundary processing, and outputs the optimized segmentation result. The entire post-processing process forms an iterative optimization mechanism, continuously improving the accuracy of the segmentation result.
[0024] These modules form a closed-loop optimization system through feature flow and gradient feedback. The modules cooperate closely and promote each other, jointly improving the performance of fundus vessel segmentation. This modular system design not only improves the interpretability of the algorithm, but also provides a flexible expansion space for subsequent optimization and improvement.
[0025] S2. Train the constructed fundus image vascular segmentation model to obtain a trained fundus image vascular segmentation model.
[0026] S3. Use the trained fundus image vascular segmentation model to perform vascular segmentation on the input fundus image to obtain the corresponding vascular segmentation result.
[0027] Global feature extraction module The present invention first realizes the feature extraction and enhancement of the input fundus image through the global feature extraction branch (G-Branch). As Figure 2 shown, the global feature extraction module adopts a novel feature extraction and enhancement architecture, mainly including key components such as deep convolutional feature extraction, position encoding enhancement, and multi-head self-attention mechanism. Each component forms a closed loop through feature mapping and optimization, realizing the effective modeling of the overall structure of the vascular network.
[0028] In the initial feature extraction stage, the input fundus image x is feature-mapped through the deep convolutional neural network φ(x) to extract the initial feature map of the image; at the same time, a position encoding mechanism is introduced to inject the position encoding information into the feature map to better perceive the spatial distribution characteristics of blood vessels: where, F init is the extracted initial feature map, F g ' is the feature representation after position encoding enhancement, and ξ( ) is the position encoding.
[0029] In the feature enhancement stage, a feature modeling module based on the multi-head self-attention mechanism is constructed. The feature modeling module captures the global correlation between features through h parallel attention heads, and each attention head H i independently learns the interaction relationship between features: where, ΛQ i , ΛK i and ΛV i are the linear transformation matrices of query, key, and value respectively, σ( ) is the softmax normalization function, and the gradient vanishing problem is prevented through the scaling factor .
[0030] Finally, the outputs of all attention heads are concatenated and passed through the linear transformation Ω to obtain the fused global feature: where, H h is the h-th attention head.
[0031] To realize the dynamic optimization of features, a feature optimization feedback path is constructed inside the global feature extraction module. Through the backpropagation of the optimization signal, the feature extraction and attention calculation processes are dynamically adjusted to further improve the discriminative ability of the feature representation. This closed-loop design of feature extraction and optimization ensures that the model can accurately grasp the overall topological structure of the vascular network, laying a foundation for subsequent accurate segmentation.
[0032] The design of the above global feature extraction module not only enhances the model's understanding ability of the overall structure of the vascular network, but also effectively captures long-range dependencies through the multi-head self-attention mechanism, overcoming the limitation of traditional methods that only focus on local features. At the same time, the introduction of the position encoding enhancement mechanism and the feature optimization feedback mechanism further improves the integrity and accuracy of feature representation, significantly enhancing the segmentation performance of the model.
[0033] Local Feature Extraction Module In the present invention, an improved U-Net architecture is adopted in the local branch (L-Branch) for local feature extraction. As Figure 3 shown, the local feature extraction module effectively captures different-scale features of fundus blood vessels through a multi-level encoder-decoder structure and a multi-scale feature extraction module.
[0034] In the encoder part, a three-layer encoder structure is constructed, corresponding to feature extraction at scales of 1 / 1, 1 / 2, and 1 / 4 respectively; each layer of the encoder is implemented through a combination of depthwise separable convolution and residual connection, ensuring effective feature extraction while reducing computational complexity; among them, depthwise separable convolution is used to perform specific convolution operations on the input features to extract key feature information, and the residual connection helps to retain part of the information of the original input features, avoiding information loss during the convolution process, and the two work together to ensure the effect of feature extraction.
[0035] In the decoder part, a three-layer decoder structure is correspondingly adopted, and the spatial resolution of the feature map is restored through step-by-step upsampling operations; during each layer of decoding process, the feature extraction process is expressed as: where f i represents the feature mapping result of the i-th layer decoder; η i is a feature transformation function based on depthwise separable convolution, used to process the input features, f i-1 is the output feature of the (i - 1)-th layer decoder, and F i skip is the skip connection feature from the i-th layer encoder. This design follows the basic architecture idea of U-Net, and through the skip connection, the detailed feature information in the encoder can be effectively fused with the semantic information in the decoder.
[0036] The multi-scale feature extraction module outputs local feature sequences F l ={f 1 , f 2 , f 3} corresponding to different scales respectively at three encoder layers, where f 1 、f 2 、f3 Features corresponding to scales of 1 / 1, 1 / 2, and 1 / 4 respectively.
[0037] Each element in the local feature sequence undergoes feature concatenation and feature refinement processing; in the feature concatenation stage, a skip connection mechanism is adopted to achieve effective fusion of high-level and low-level features: Among them, δ( ) represents a feature fusion function, which is used to adaptively integrate feature information of different scales and effectively fuse the features of the three scales of 1 / 1, 1 / 2, and 1 / 4.
[0038] In terms of feature refinement, a feature guidance mechanism is introduced to dynamically adjust the feature extraction process by optimizing the backpropagation of signals; this feature guidance mechanism not only acts between the layers of the encoder, but also further optimizes the extracted features through a feature refinement unit; the feature refinement process is expressed as: Among them, i = 1, 2, 3, γ( ) represents a feature refinement function, which is used to perform a non-linear transformation on the input features to extract more discriminative representations; β is a learnable scaling parameter, which is used to adaptively adjust the importance of features. Through the combination of feature refinement and adaptive scaling, effective optimization of features is achieved.
[0039] Finally, the local feature sequence F output by the local feature extraction module is obtained l ={f 1 ', f 2 ', f 3 '}.
[0040] This improved design of the local feature extraction module has the following advantages: First, the multi-level encoder-decoder structure can effectively capture vascular features of different scales; second, the feature concatenation and refinement mechanisms ensure the integrity and accuracy of feature representations; finally, the feedback mechanism for optimizing signals provides the ability to dynamically adjust the feature extraction process. The organic combination of these innovative designs significantly improves the model's ability to identify the fine structure of blood vessels and lays a solid foundation for subsequent precise segmentation.
[0041] Through the design of the above local feature extraction module, the present invention successfully solves the adaptability problem of traditional methods in dealing with vascular structures of different scales, especially showing excellent performance in dealing with fine blood vessels. The synergistic effect of this module and the global feature extraction module further improves the overall segmentation performance of the model.
[0042] Feature Fusion and Topological Constraint Module The present invention innovatively designs a feature fusion and topological constraint module, which fuses the global feature F output by the global feature extraction module gand the local feature sequence F output by the local feature extraction module l Input Feature Fusion and Topological Constraint Module. As Figure 4 shown, the feature fusion and topological constraint module includes two core parts: cross-scale feature fusion and topological structure constraint. Through the organic combination of adaptive feature fusion and topological structure optimization, an accurate modeling of the vascular network structure is achieved.
[0043] In the cross-scale feature fusion stage, first, the global feature and the local feature are preliminarily integrated by the feature fusion module, and then an adaptive weight mechanism based on channel attention is introduced for feature enhancement; the fusion process is expressed as: where, F g is the global feature, f i ' ∈ F l is the i-th feature in the local feature sequence F l , β( ) and γ( ) respectively represent the channel attention functions for the global feature and the local feature, and ⊙ represents the element-wise product operation; the channel attention weight is dynamically learned through the attention weight calculation module: where, W α and W α are learnable weight matrices, b α and b α are the corresponding bias terms, and σ( ) is the sigmoid activation function, which is used to normalize the attention weight to the [0,1] interval.
[0044] To make full use of the statistical information of the features, the information of both global average pooling and max pooling is considered simultaneously in the pooling operation Pool( ): where, x represents the input feature map (which can be a local feature or a global feature), AvgPool( ) represents average pooling, MaxPool( ) represents max pooling, and concat() represents feature concatenation. This combination of pooled features can more comprehensively express the statistical characteristics of the features, contribute to the accurate calculation of the attention weight, and thus improve the effect of feature fusion.
[0045] In the topological structure constraint stage, a topological constraint mechanism based on persistent homology theory is introduced. First, the feature representation is converted into a topological space through the topological graph construction unit, and then through connectivity analysis and structure optimization, the topological graph can effectively maintain the topological characteristics of the vascular network; the construction process of the topological graph is expressed as: Among them, (b k , d k ) represents the life cycle of the k-th topological feature, including its birth time b k and death time d k . Based on the constructed topological graph, a topological loss function is constructed to constrain the structural integrity of the segmentation result: Among them, S represents the segmentation result predicted by the model, Y represents the true label, and PH 0 ( ) represents the operator for calculating topological features, and D( ) represents the function for measuring the distance between two sets of topological features.
[0046] To achieve the collaborative optimization of feature fusion and topological constraint, a structural feedback mechanism is constructed inside the feature fusion and topological constraint module. This structural feedback mechanism feeds back the optimization signal of topological constraint to the feature fusion process, and dynamically adjusts the feature fusion strategy through topological guidance. This design not only ensures the accuracy of feature fusion, but also effectively maintains the topological structure of the vascular network.
[0047] Through the above design of the feature fusion and topological constraint module, the present invention successfully solves the deficiencies of traditional methods in feature fusion and structure preservation. This module effectively integrates global semantic information and local detail features through cross-scale feature fusion, and at the same time ensures the structural integrity of the segmentation result through topological structure constraint, significantly improving the performance of fundus vascular segmentation.
[0048] , Multi-task joint optimization module The present invention innovatively designs a multi-task joint optimization module. As Figure 5 shown, the multi-task joint optimization module realizes the overall optimization of the model performance through the synergistic effect of three parts: topological loss, segmentation loss, and contrast loss.
[0049] Through the dynamic weight adjustment mechanism and the backpropagation optimization strategy, an organic optimization closed-loop is formed among the loss functions. The overall loss function is expressed as: Among them, L seg represents the segmentation loss, which is used to optimize the classification accuracy at the pixel level; L topo represents the topological loss, which is used to maintain the structural integrity of the vascular network; L contr represents the contrast loss, which is used to enhance the discriminative ability of feature representation; λ 1 , λ 2 and λ 3 are the corresponding weight coefficients, which are used to balance the contributions of various losses.
[0050] For the topological loss, two key constraints are designed based on the persistent homology theory: structural consistency and connectivity constraint; the topological loss function is defined as: where S represents the segmentation result predicted by the model, Y represents the ground truth label, PH 0 ( ) represents the operator for calculating topological features, D( ) represents the function for measuring the distance between two sets of topological features; b i and b i ' respectively represent the birth times of the i-th topological feature in the prediction result and the ground truth label, and μ is the balance coefficient used to adjust the relative importance of the structural consistency constraint and the connectivity constraint. The first term measures the topological difference between the prediction result and the ground truth label, and the second term enhances the preservation of connectivity by constraining the birth times of topological features.
[0051] The topological loss function not only considers the topological difference between the prediction result and the ground truth label, but also enhances the stability of topological features through the life cycle regularization term.
[0052] For the segmentation loss, cross-entropy loss and Dice loss are adopted to achieve: where L CE represents the cross-entropy loss, which is used to optimize the classification accuracy at the pixel level; L Dice represents the Dice coefficient loss, which is used to measure the overlap degree between the predicted segmentation result and the ground truth label; α is the weight coefficient used to balance the relative importance of the two losses, 0 ≤ α ≤ 1, and the best balance between the pixel-level accuracy and the region overlap degree can be achieved by adjusting α.
[0053] Among them, the cross-entropy loss is improved through the Focal loss mechanism to enhance the learning ability for difficult-to-classify samples: where y i represents the ground truth label (taking 0 or 1) of the i-th pixel, p i represents the probability that the model predicts that this pixel belongs to the blood vessel, and ∑ represents the summation over all pixel positions. By accumulating the cross-entropy loss at each pixel position, the overall classification accuracy can be evaluated. The Focal loss mechanism makes the model pay more attention to these challenging regions during training by giving higher weights to difficult-to-classify samples.
[0054] For the contrastive loss, by constructing positive and negative sample pairs, the model is promoted to learn more discriminative feature representations: Among them, s + represents the similarity score between positive sample pairs, and s - represents the similarity score between negative sample pairs. τ represents the temperature coefficient, which is used to adjust the smoothness of the feature distribution. The design objective of this loss function is to make the feature representations of positive sample pairs more similar (s + is larger), while the feature representations of negative sample pairs are more different (s - is smaller), thereby enhancing the discriminative ability of the model features. ∑ represents the summation of the similarity scores of all negative sample pairs.
[0055] To achieve the dynamic balance between loss functions, an adaptive weight adjustment mechanism based on gradient statistics is constructed; the weight coefficient is updated as follows: Among them, λ i represents the weight coefficient of the i-th loss function, and θ i is the parameter value corresponding to the i-th loss function. ∑exp(θ j ) represents the normalization of the parameters of all loss functions.
[0056] Among them, the update of the parameter θ i considers the gradient correlation between loss functions: Among them, η is the learning rate, and represent the gradients of the i-th and j-th loss functions respectively. ρ( ) represents the function for calculating the gradient correlation, and σ( ) is the activation function, which is used to map the correlation to an appropriate range.
[0057] To ensure the stability of the training process, a smoothing mechanism based on exponential moving average is introduced: Among them, represents the smoothed weight of the i-th loss function at the t-th moment, is the weight value at the previous moment, and β is the smoothing coefficient, which is used to control the influence degree of the historical weight value, and 0 ≤ β ≤ 1.
[0058] This multi-task joint optimization design has the following advantages: First, the three loss functions optimize the model performance from three dimensions: topological structure, pixel classification, and feature representation; second, the dynamic weight adjustment mechanism can adaptively balance each optimization objective; finally, the backpropagation optimization strategy ensures that each sub-module can co-evolve and continuously improve the overall performance of the model. This design not only improves the generalization ability of the model but also enhances its robustness in practical applications.
[0059] , and a post - processing optimization module The present invention innovatively designs a post - processing optimization module. As Figure 6 shown, the post - processing optimization module realizes the refined processing of the initial segmentation result through three - stage optimization: graph structure construction, graph convolution optimization, and conditional random field refinement.
[0060] In the graph structure construction stage, an initial graph structure is first constructed through node feature extraction and edge feature calculation. For each super - pixel region p i extracted from the segmentation result, the construction process of the node feature is expressed as: where p i represents the i - th super - pixel region, f i is the corresponding depth feature, v i is the constructed node feature vector, is a feature fusion function for fusing position information and depth features, P is the set of all super - pixel regions; the edge weight is calculated by comprehensively considering the spatial relationship and feature similarity between nodes: where e ij represents the weight of the edge between node i and node j, v i and v j are the feature vectors of the two nodes respectively, σ is the Gaussian kernel parameter controlling the influence degree of feature similarity, d i,j represents the spatial distance between the two nodes, r is a preset distance threshold, and δ( ) is an indicator function that takes the value of 1 when the condition holds and 0 otherwise. This design of edge weight takes into account both the similarity of node features (exp term) and spatial proximity (δ term).
[0061] In the graph convolution optimization stage, a multi - layer graph convolution network is constructed for structure optimization; the feature update rule for each layer is: where H l+1 and H l represent the node feature matrices of the (l + 1)-th layer and the l - th layer respectively, σ( ) is a non - linear activation function (such as ReLU), is the adjacency matrix with self - connection added, is the corresponding degree matrix (the diagonal element of which is the sum of the elements in each row of ), W l is the learnable weight matrix of the l - th layer; where This implementation achieves symmetric normalization of the adjacency matrix, which helps prevent numerical instability problems. This update rule updates node features by aggregating the neighborhood information of each node, thereby achieving structural optimization.
[0062] To enhance the effectiveness of feature propagation, an attention mechanism is introduced to dynamically modulate message passing: where h i and h j represent the feature vectors of nodes i and j respectively, h i ' is the updated feature representation of node i, W α is the attention weight matrix, W is the feature transformation matrix, g( ) is the attention scoring function for calculating feature similarity, softmax( ) is used to normalize the attention scores to the [0,1] interval, and σ( ) is the non-linear activation function; N i represents the set of neighborhood nodes of node i, and α ij represents the attention weight from node j to node i.
[0063] The first formula calculates the attention coefficient between node pairs, determining the importance weight of information transmission by evaluating the correlation between the transformed feature vectors. The second formula performs weighted aggregation of the features of neighborhood nodes based on the calculated attention weights, thus achieving adaptive feature update. This attention-based message passing mechanism can dynamically adjust the information flow according to the correlation of node features, effectively improving the quality of feature propagation.
[0064] In the conditional random field refinement stage, an improved energy function is constructed: where E(y) is the energy function, y represents the segmentation label configuration, is the unary potential function, representing the local confidence of the y i label of node i, is the pairwise potential function for modeling the label dependence between adjacent nodes i and j.
[0065] The pairwise potential function comprehensively considers spatial location, color features, and edge information: where μ(y i , y j ) is the label compatibility function, used to measure the compatibility degree between the labels y i and y j of adjacent nodes; the spatial location coordinates p i and pj represents the position information of nodes i and j in the image, I i and I j represent the respective color feature vectors of nodes i and j; w 1 and w 2 represent weight coefficients; in order to reasonably balance the influence of the two factors of spatial distance and color similarity, the weight coefficients w 1 and w 2 are introduced for weighting. The Gaussian kernel parameter σ α and σ β respectively control the influence ranges of spatial distance and color similarity, which is used to calculate the distance metric between nodes.
[0066] These three optimization stages form a closed-loop optimization system, and iterative optimization is achieved through a structural feedback mechanism: wherein, represents the graph structure update function, E t is the value of the energy function after the t-th round of optimization.
[0067] This multi-stage post-processing optimization design has the following advantages: First, the graph structure representation provides a high-level semantic understanding of the segmentation result; Second, the graph convolution optimization realizes structure-based feature optimization; Finally, the conditional random field refinement ensures the precise processing of boundaries. The organic combination of the three stages and the dynamic feedback mechanism form a complete optimization closed-loop, continuously improving the quality of the segmentation result.
[0068] Through the design of the above post-processing optimization module, the present invention successfully solves the deficiencies of traditional methods in boundary refinement and structure preservation, and significantly improves the accuracy and reliability of fundus vascular segmentation. The cooperation of this module with the aforementioned feature extraction and optimization module further improves the segmentation performance of the entire system.
[0069] In this embodiment, the method is experimentally verified on multiple public datasets. The experimental results show that compared with the prior art, the present invention shows significant advantages in aspects such as vascular segmentation accuracy, topological structure preservation, and processing of blood vessels of different scales. Especially when dealing with complex lesion images and fine blood vessels, the present invention demonstrates stronger robustness and higher segmentation accuracy, providing more reliable technical support for the clinical diagnosis of fundus diseases.
[0070] This embodiment also provides a fundus image vascular segmentation system, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above method can be implemented.
[0071] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0072] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0075] As mentioned above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for segmenting blood vessels in fundus images, characterized in that: include: A fundus image vascular segmentation model is constructed, which includes a dual-branch feature extraction module, a feature fusion and topology constraint module, a multi-task joint optimization module and a post-processing optimization module; the dual-branch feature extraction module includes a global feature extraction module and a local feature extraction module. For the input fundus image, the global feature extraction module performs initial feature extraction through a deep convolutional network, and combines position encoding enhancement and multi-head self-attention mechanism to capture the overall topological structure of the vascular network; the local feature extraction module performs local feature extraction based on an improved U-Net architecture, and maintains the fine structural features of the blood vessels through multi-scale feature extraction and jump connection mechanism; the feature fusion and topology constraint module realizes full-scale feature fusion through a cross-scale feature fusion strategy. The local features are adaptively integrated with local features of different scales, and a topological constraint mechanism based on persistent homology theory is introduced. The structural integrity of the vascular network is maintained through topological graph construction and connectivity analysis, and the initial segmentation result is finally output; the multi-task joint optimization module comprehensively considers topological loss, segmentation loss and contrast loss, and balances different optimization goals through a dynamic weight adjustment mechanism; among them, the topological loss ensures the structural integrity of the vascular network, the segmentation loss optimizes the classification accuracy at the pixel level, and the contrast loss enhances the discriminative ability of feature representation; the post-processing optimization module converts the initial segmentation result into a graph structure representation, and then performs structural optimization through a multi-layer graph convolutional network, and finally uses conditional random fields to achieve refined processing of the boundary, and outputs the optimized segmentation result; Training the constructed fundus image blood vessel segmentation model to obtain a trained fundus image blood vessel segmentation model; The trained fundus image vascular segmentation model is used to perform vascular segmentation on the input fundus image to obtain the corresponding vascular segmentation result.
2. The method for segmenting blood vessels in fundus images according to claim 1, characterized in that: The implementation method of the global feature extraction module is: In the initial feature extraction stage, the input fundus image is feature mapped through a deep convolutional neural network to extract the initial feature map of the image; at the same time, a position encoding mechanism is introduced to inject the position encoding information into the feature map to better perceive the spatial distribution characteristics of the blood vessels: Among them, F init is the extracted initial feature map, F g ' is the feature representation after position coding enhancement, ξ( ) is the position coding; In the feature enhancement stage, a feature modeling module based on a multi-head self-attention mechanism is constructed. The feature modeling module captures the global correlation between features through h parallel attention heads. Each attention head H i Learn interactions between features independently: Among them, ΛQ i , ΛK i and ΛV i are the linear transformation matrices of query, key, and value respectively, σ( ) is the softmax normalization function, and To prevent the gradient vanishing problem; finally, the outputs of all attention heads are concatenated and linearly transformed Ω to obtain the fused global features: Among them, H h is the hth attention head; In order to achieve dynamic optimization of features, a feature optimization feedback path is constructed inside the global feature extraction module. By optimizing the back propagation of signals, the feature extraction and attention calculation processes are dynamically adjusted to further improve the discriminative ability of feature representation.
3. The method for segmenting blood vessels in fundus images according to claim 1, characterized in that: The implementation method of the local feature extraction module is: The local feature extraction module effectively captures the features of fundus blood vessels at different scales through a multi-level encoder-decoder structure and a multi-scale feature extraction module; In the encoder part, a three-layer encoder structure is constructed, corresponding to the feature extraction of 1 / 1, 1 / 2 and 1 / 4 scales respectively; Each layer of the encoder is implemented through a combination of depth-wise separable convolution and residual connection, which reduces the computational complexity while ensuring effective feature extraction. In the decoder part, a three-layer decoder structure is used to restore the spatial resolution of the feature map through step-by-step upsampling operations; in each layer of decoding, the feature extraction process is expressed as: Among them, f i represents the output feature of the i-th layer decoder; η i is a feature conversion function based on depthwise separable convolution, which is used to process the input features. i-1 is the output feature of the i-1th layer decoder, F i skip is the skip connection feature from the i-th layer encoder; The multi-scale feature extraction module outputs the local feature sequence F of the corresponding scale at three encoder layers through the scale feature generation unit l ={f1, f2, f3}, where f1, f2, f3 correspond to the features of 1 / 1, 1 / 2 and 1 / 4 scales respectively; Each element in the local feature sequence undergoes feature concatenation and feature refinement. The feature concatenation uses a skip connection mechanism to achieve effective fusion of high-level and low-level features: Among them, δ( ) represents the feature fusion function, which is used to adaptively integrate feature information of different scales and effectively fuse the features of the three scales of 1 / 1, 1 / 2 and 1 / 4; Feature refinement introduces a feature guidance mechanism to dynamically adjust the feature extraction process by optimizing the back propagation of the signal; this feature guidance mechanism not only acts between the layers of the encoder, but also further optimizes the extracted features through the feature refinement unit; the feature refinement process is expressed as: Where i=1,2,3, γ( ) represents the feature refinement function, which is used to perform nonlinear transformation on the input features to extract more discriminative representations; β is a learnable scaling parameter, which is used to adaptively adjust the importance of features; Finally, the local feature sequence F output by the local feature extraction module is obtained. l ={f1', f2', f3'}.
4. The method for segmenting blood vessels in fundus images according to claim 1, characterized in that: The implementation method of the feature fusion and topology constraint module is: The global feature F output by the global feature extraction module g And the local feature sequence F output by the local feature extraction module l An input feature fusion and topology constraint module is provided, wherein the feature fusion and topology constraint module includes two parts: cross-scale feature fusion and topology structure constraint; In the cross-scale feature fusion stage, the global features and local features are first integrated through the feature fusion module, and then the adaptive weight mechanism based on channel attention is introduced to enhance the features; the fusion process is expressed as: Among them, F g is the global feature, f i '∈F l is the local feature sequence F l The i-th feature in , β( ) and γ( ) represent the channel attention functions for global features and local features respectively, ⊙ represents the element-by-element product operation; the channel attention weight is dynamically learned by the attention weight calculation module: Among them, W α and W α is the learnable weight matrix, b α and b α is the corresponding bias term, σ( ) is the sigmoid activation function, which is used to normalize the attention weight to the interval [0,1]. In order to make full use of the statistical information of the features, the global average pooling and maximum pooling information are considered in the pooling operation Pool( ): Where x represents the input feature map, AvgPool() represents average pooling, MaxPool() represents maximum pooling, and concat() represents feature concatenation; In the topological structure constraint stage, a topological constraint mechanism based on persistent homology theory is introduced. First, the feature representation is converted into a topological space through the topological map construction unit, and then the topological map maintains the topological characteristics of the vascular network through connectivity analysis and structural optimization. The construction process of the topological map is expressed as: Among them, (b k , d k ) represents the life cycle of the kth topological feature, including its birth time b k and extinction time d k ; Based on the constructed topological map, a topological loss function is constructed to constrain the structural integrity of the segmentation result: Among them, S represents the segmentation result predicted by the model, Y represents the true label, PH0( ) represents the operator for calculating the topological feature, and D( ) represents the function for measuring the distance between two topological feature sets; In order to achieve the coordinated optimization of feature fusion and topological constraints, a structural feedback mechanism is constructed inside the feature fusion and topological constraint module, which feeds back the optimization signal of the topological constraint to the feature fusion process and dynamically adjusts the feature fusion strategy through topological guidance.
5. The method for segmenting blood vessels in fundus images according to claim 1, characterized in that: The implementation method of the multi-task joint optimization module is: The multi-task joint optimization module achieves overall optimization of model performance through the synergy of topological loss, segmentation loss and contrast loss; the overall loss function is expressed as: Among them, L seg represents the segmentation loss, which is used to optimize the classification accuracy at the pixel level; L topo represents the topological loss, which is used to maintain the structural integrity of the vascular network; L contr represents contrast loss, which is used to enhance the discriminative ability of feature representation; λ1, λ2 and λ3 are corresponding weight coefficients, which are used to balance the contribution of each loss; For the topological loss, two key constraints are designed based on the persistent homology theory: structural consistency and connectivity constraints; the topological loss function is defined as: Where S represents the segmentation result predicted by the model, Y represents the true label, PH0( ) represents the operator for calculating the topological feature, and D( ) represents the function for measuring the distance between two topological feature sets; i and b i 'represents the birth time of the i-th topological feature in the prediction result and the true label respectively, μ is the balance coefficient, which is used to adjust the relative importance of the structural consistency constraint and the connectivity constraint; The topological loss function not only considers the topological difference between the predicted result and the true label, but also enhances the stability of the topological feature through the life cycle regularization term; For segmentation loss, cross entropy loss and Dice loss are used to implement it: Among them, L CE represents the cross entropy loss, which is used to optimize the classification accuracy at the pixel level; L Dice represents the Dice coefficient loss, which is used to measure the overlap between the predicted segmentation result and the true label; α is the weight coefficient, which is used to balance the relative importance of the two losses, 0≤α≤1, and the best balance between pixel-level accuracy and regional overlap is achieved by adjusting α; Among them, the cross entropy loss is improved through the Focal loss mechanism to enhance the learning ability of difficult-to-classify samples: Among them, y i represents the true label of the i-th pixel, p i It represents the probability that the model predicts that the pixel belongs to a blood vessel, and ∑ represents the sum of all pixel positions; For contrast loss, by constructing positive and negative sample pairs, the model is encouraged to learn more discriminative feature representations: Among them, s + Represents the similarity score between positive sample pairs, s - Represents the similarity score between negative sample pairs, τ represents the temperature coefficient, which is used to adjust the smoothness of feature distribution; ∑ represents the sum of the similarity scores of all negative sample pairs; In order to achieve a dynamic balance between loss functions, an adaptive weight adjustment mechanism based on gradient statistics is constructed; the weight coefficients are updated in the following way: Among them, λ i represents the weight coefficient of the i-th loss function, θ i is the parameter value corresponding to the i-th loss function, ∑exp(θ j ) means normalizing the parameters of all loss functions; Among them, the parameter θ i The update of takes into account the gradient correlation between the loss functions: Where η is the learning rate, and They represent the gradients of the i-th and j-th loss functions, ρ( ) represents the function for calculating the gradient correlation, and σ( ) is the activation function used to map the correlation to an appropriate range; In order to ensure the stability of the training process, a smoothing mechanism based on exponential moving average is introduced: in, represents the smoothed weight of the i-th loss function at the t-th time, is the weight value at the previous moment, β is the smoothing coefficient, which is used to control the influence of the historical weight value, 0≤β≤1.
6. The method for segmenting blood vessels in fundus images according to claim 1, characterized in that: The implementation method of the post-processing optimization module is: The post-processing optimization module refines the initial segmentation results through three stages of optimization: graph structure construction, graph convolution optimization, and conditional random field refinement. In the graph structure construction stage, the initial graph structure is first constructed by extracting node features and calculating edge features; for each superpixel region p extracted from the segmentation result, i , the construction process of node features is expressed as: Among them, p i represents the i-th superpixel region, f i is the corresponding depth feature, v i is the constructed node feature vector, is a feature fusion function used to fuse position information and depth features, P is the set of all superpixel regions; the edge weight is calculated by comprehensively considering the spatial relationship and feature similarity between nodes: Among them, e ij represents the weight of the edge between node i and node j, v i and v j are the feature vectors of the two nodes, σ is the Gaussian kernel parameter that controls the influence of feature similarity, and d i,j represents the spatial distance between two nodes, r is the preset distance threshold, δ( ) is an indicative function, which takes the value of 1 when the condition is met, otherwise it takes the value of 0; In the graph convolution optimization stage, a multi-layer graph convolution network is constructed for structural optimization; the feature update rule of each layer is: Among them, H l+1 and H l They represent the node feature matrices of the l+1th layer and the lth layer respectively, σ( ) is a nonlinear activation function, To add the self-connected adjacency matrix, for The corresponding degree matrix, W l is the learnable weight matrix of the lth layer; This term realizes the symmetric normalization of the adjacency matrix; In order to enhance the effectiveness of feature propagation, an attention mechanism is introduced to dynamically modulate message passing: Among them, h i and h j denote the feature vectors of node i and node j respectively, h i ' is the updated feature representation of node i, W α is the attention weight matrix, W is the feature transformation matrix, g( ) is the attention score function used to calculate feature similarity, softmax( ) is used to normalize the attention score to the [0,1] interval, σ( ) is the nonlinear activation function; N i represents the set of neighboring nodes of node i, α ij represents the attention weight from node j to node i; In the conditional random field refinement stage, an improved energy function is constructed: Among them, E(y) is the energy function, y represents the segmentation label configuration, is a single-point potential function, representing the y of node i i The local confidence of the label, The pairwise potential function is used to model the label dependency relationship between adjacent nodes i and j; the pairwise potential function comprehensively considers spatial position, color features and edge information: Among them, μ(y i , y j ) is a label compatibility function used to measure the adjacent node labels y i and j The degree of compatibility between them; spatial position coordinates p i and p j Represents the location information of nodes i and j in the image, I i and I j represents the color feature vector of nodes i and j respectively; w1 and w2 represent weight coefficients; Gaussian kernel parameter σ α and σ β Control the influence range of spatial distance and color similarity respectively, A distance metric used to calculate the distance between nodes.
7. A fundus image blood vessel segmentation system, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 6 can be implemented.
Citation Information
Cited By
Medical image field generalization segmentation method and system
CN120355929A
Cell microtubule array uncertainty quantitative segmentation method fused with large vision model
CN120807557A
Airway tree segmentation method and device, computer equipment and storage medium
CN120852445A
ROP lesion detection and partition positioning method based on fundus color photo
CN120976155A
Method for detecting and partitioning retinopathy of prematurity (ROP) lesions based on fundus color photograph
CN120976155B