Retinal blood vessel image segmentation method based on dynamic weight distribution
The retinal vascular image segmentation method optimized by dynamically adjusting the weight of loss function and index trend, combined with the hybrid encoder of EfficientNet-B4 and ResNet-50, solves the problem of insufficient synergy between global information capture and loss function, and achieves higher segmentation accuracy and effective segmentation of small blood vessels.
Patent Information
- Application Number
- CN202510308071.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-15
- Publication Date
- 2025-07-04
AI Technical Summary
The existing retinal vascular image segmentation method has insufficient global information capture ability and loss function optimization synergy, resulting in limited segmentation accuracy, especially in complex and small blood vessel areas.
The retinal vascular image segmentation method based on dynamic weight allocation is adopted, and the mixed encoder of EfficientNet-B4 and ResNet-50 is combined with the LM-TransUNet model, and the loss function weight is dynamically adjusted using the LT-DWA module, and the weight is optimized according to the segmentation index trend through the MT-GO module to realize multi-scale feature fusion and global feature extraction.
The segmentation accuracy of retinal vascular images is significantly improved, especially in complex and small blood vessel areas, which significantly improves the segmentation effect, improving the flexibility and segmentation performance of the model.
Smart Images

Figure CN120260087A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image segmentation, and relates to a method for segmenting retinal blood vessel images based on dynamic weight allocation. Background Art
[0002] Currently, the encoding and decoding architecture has become the mainstream in the field of medical image segmentation. Among them, U-Net is widely used in various segmentation tasks due to its efficient U-shaped structure and multi-scale feature fusion ability. However, the encoder of the traditional U-Net is not comprehensive enough in feature extraction. The main reason is the limitation of convolution, resulting in insufficient ability to capture global information. It can be improved by modifying the encoding and decoding structure or optimizing the convolution structure. Huang et al. used full-scale skip connections to combine high-level and low-level semantic information to enhance feature extraction. Zhang et al. used Swin Transformer as the encoder and proposed cross-layer feature enhancement to recover low-level features for accurate semantic segmentation. Garbaz et al. used the proposed multi-scale information attention to replace the stacked module of UNet, thereby expanding the receptive field to improve the segmentation performance. Tang et al. used the attention mechanism and convolutional weighted fusion to enhance the shallow features, and then combined the enhanced shallow features with the deep features to achieve multi-scale feature extraction. However, the existing methods have limitations in the static setting of model-related weights, and further research is needed on dynamic weight adjustment to improve the performance and flexibility of the model. Yuan et al. extracted semantic information at different stages and assigned different weights to it to improve the feature learning ability. Han et al. used dynamic weights to constrain strong interference at different scales to solve the problem of microaneurysm feature confusion. Wang et al. improved the multi-scale context extraction ability by dynamically adjusting the receptive field. You et al. introduced a dynamic sparse attention mechanism to selectively focus on the image lesion area by precisely controlling the sparse ratio. The multi-level long-range relationship modeling network proposed by Pan et al. can adaptively and dynamically adjust different attention maps, assign different weights to different features, and effectively distinguish the contributions of different information to the model, thereby improving the segmentation ability of fundus blood vessels. The above methods all perform dynamic adjustment from the perspective of module parameter weights, but ignore the coordination of multiple loss functions.
[0003] In order to improve the global feature extraction ability of the encoding and decoding architecture neural network and the coordination of loss function optimization, the present invention proposes a method for segmenting retinal blood vessel images based on dynamic weight allocation. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for segmenting retinal blood vessel images based on dynamic weight allocation, which strengthens and improves the global feature extraction ability of the encoding and decoding architecture neural network and the coordination of loss function optimization to improve the segmentation accuracy of retinal blood vessel images.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A retinal vessel image segmentation method based on dynamic weight allocation, comprising:
[0007] Obtain the retinal vessel image to be segmented;
[0008] Input the retinal vessel image to be segmented into the LM-TransUNet model to obtain the visualization result of image segmentation, wherein the LM-TransUNet model is trained based on a training set, and the training set includes retinal vessel images and corresponding ground truths;
[0009] During the process of training the LM-TransUNet model based on the training set, it includes: predicting the retinal vessel image through the initial LM-TransUNet model, calculating multiple loss function values according to the prediction result of the retinal vessel image and the ground truth; inputting the loss function values into the dynamic weight allocation module, and dynamically adjusting the weights by combining the historical loss gradient change and the index trend information; using the adjusted weighted total loss for backpropagation iteration until convergence to obtain the LM-TransUNet model.
[0010] Optionally, the LM-TransUNet model includes a hybrid encoder, a decoder, an LT-DWA module, and an MT-GO module, wherein the hybrid encoder includes an EfficientNet-B4 global encoder and a ResNet-50 multi-scale skip connection module, and the dynamic weight allocation includes the LT-DWA module and the MT-GO module;
[0011] The LT-DWA module is used to dynamically adjust the loss weights according to the historical gradient change and the current gradient change rate of the loss function;
[0012] The MT-GO module is used to perform secondary optimization on the weights according to the change trend of the segmentation index;
[0013] The decoder is used to fuse the global encoding features and the multi-scale skip connection features and output the new prediction result.
[0014] The multi-scale feature fusion method of the decoder is: embedding the global features output by EfficientNet-B4 and successively splicing the low, medium, and high-resolution skip connection feature maps of ResNet-50; gradually restoring the spatial resolution through deconvolution operations, and adding a 3×3 convolutional layer and a ReLU activation function at each upsampling stage for feature enhancement.
[0015] The weight adjustment method of the LT-DWA module includes:
[0016] Optionally, the method for calculating the gradient change rate of each loss function at the current moment and the previous moment is as follows:
[0017]
[0018] Optionally, the method for integrating historical gradient changes and normalizing them is as follows:
[0019]
[0020] Optionally, the method for calculating the gradient change rate of the loss function is as follows:
[0021]
[0022] Among them, α and β are the weight coefficients of the current gradient and the historical gradient, and ε is a small positive number used for stable calculation.
[0023] Optionally, the method for updating the weights based on the normalization result is as follows:
[0024]
[0025] Among them, η is the learning rate.
[0026] Normalize the adjusted comprehensive weights and calculate the weighted total loss for backpropagation.
[0027] The dynamic change values of the index trend information, including sensitivity (SE), specificity (SP), accuracy (AC), and Dice coefficient (DI), are statistically analyzed for their historical trends through the sliding window method and input into the MT-GO module.
[0028] The weight optimization method of the MT-GO module includes:
[0029] Optionally, calculate the overall index change trend:
[0030]
[0031] Optionally, adjust the weights according to the index change direction:
[0032]
[0033] Among them, λ is a dynamic adjustment factor used to amplify the weights that are significantly improved for the index.
[0034] The multiple loss functions include BoundaryLoss, FocalLoss, SkeletonRecallLoss, and CrossEntropyLoss.
[0035] Optionally, the weighted total loss calculation formula is as follows:
[0036] L = w bl ·L bl + w fl ·L fl + w srl ·L srl + w ce ·L ce
[0037] Wherein, w bl , w fl , w srl , w ce are the weight coefficients after dynamic adjustment.
[0038] The beneficial effects of the present invention are as follows: The method of the present invention uses EfficientNet-B4 as the encoder and combines the skip connection structure of ResNet-50 to enhance the multi-scale feature fusion and global feature extraction capabilities of the model. Secondly, by comprehensively analyzing the change trend of the segmentation index, the dynamic adjustment of the weights of each loss function is guided; on this basis, the weights are optimized and adjusted by combining the loss gradient and the index fluctuation, so as to improve the optimization efficiency of the model for different loss terms. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Figure 1 It is the overall flowchart of a retinal vessel image segmentation method improved based on dynamic weight allocation according to an embodiment of the present invention.
[0040] Figure 2 It is the flowchart of the preprocessing of the retinal image dataset according to an embodiment of the present invention.
[0041] Figure 3 It is the flowchart of the retinal vessel image segmentation method based on dynamic weight allocation according to an embodiment of the present invention.
[0042] Figure 4 It is the architecture diagram of LM-TransUNet according to an embodiment of the present invention.
[0043] Figure 5 It is the comparison diagram of the visualization results of the segmentation results of different methods on different datasets according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] Inspired by the dynamic parameter adjustment mechanism, this paper proposes a medical image segmentation method based on dynamic weight allocation on the basis of TransUNet, including: a loss-trend guided dynamic weight adjustment module (Loss-Trend Guided Dynamic Weight Adjustment, LT-DWA) and a metric-trend guided optimization module (Metric-Trend Guided Optimization, MT-GO), named LM-TransUNet. As Figure 1 shown, from step 101 to step 104. This embodiment provides a retinal vessel image segmentation method based on dynamic weight allocation, including: as Figure 2 shown, obtain the retinal vessel image to be segmented, from step 201 to step 204. Input the retinal vessel image to be segmented into the LM-TransUNet model to obtain a visual result of image segmentation, where the LM-TransUNet model is obtained by training based on a training set, and the training set includes retinal vessel images and corresponding ground truths.
[0047] The process of training the LM-TransUNet model based on the training set includes: predicting the retinal vessel image through the initial LM-TransUNet model, calculating multiple loss function values according to the prediction result of the retinal vessel image and the ground truth; inputting the loss function values into the dynamic weight allocation module, and dynamically adjusting the weights in combination with the historical loss gradient change and metric trend information; using the adjusted weighted total loss for backpropagation iteration until convergence to obtain the LM-TransUNet model.
[0048] Furthermore, the LM-TransUNet model includes a hybrid encoder, a decoder, an LT-DWA module, and an MT-GO module, where the hybrid encoder includes an EfficientNet-B4 global encoder and a ResNet-50 multi-scale skip connection module. The LT-DWA module is used to dynamically adjust each loss weight according to the historical gradient change and the current gradient change rate of the loss function. The MT-GO module is used to perform secondary optimization on the weights according to the change trend of the segmentation metric. The decoder is used to fuse the global encoding features and the multi-scale skip connection features and output the new prediction result.
[0049] As Figure 4 shown, in this embodiment, based on the proposed LT-DWA and MT-GO modules, an LM-TransUNet model is constructed on the basis of TransUNet. The LM-TransUNet model consists of two parts: an encoder and a decoder. The encoder is mainly composed of a CNN hybrid module and an attention module, where the CNN hybrid module is composed of an EfficientNet-B4 and a ResNet-50 multi-scale skip connection module, and the attention module contains 12 Transformer sub-modules. The decoder has the same structure as the decoder of TransUNet and is mainly composed of a 1x1 convolution, an upsampling layer, and a convolution layer. After the decoder outputs the prediction result, the LT-DWA and MT-GO modules are used for dynamic weight adjustment.
[0050] The overall process of the method in this embodiment is as Figure 3As shown below. First, replace ResNet50 with Efficientnet-b4 for subsequent encoder operations, and use three feature maps of the original ResNet50 for the skip connections of the decoder. Next, obtain the prediction results through the forward propagation of LM-TransUNet, and calculate various loss values according to the prediction results and GT, including BoundaryLoss, FocalLoss, SkeletonRecallLoss, and CrossEntropyLoss. Input the loss values and the historical loss trends into the LT-DWA module, dynamically adjust the weights of each loss according to the current gradient and the gradient at the previous moment, and further optimize the weight update strategy by combining the loss adjustment coefficient. Then, input the adjusted weights of each loss and the index trend information into the MT-GO module, update the weights through comprehensive information such as index changes, and perform normalization and weighted summation. Finally, perform backpropagation on the dynamically allocated weights, and repeat this process until the algorithm converges. Specifically, in this embodiment, LM-TransUNet extracts global feature embeddings through EfficientNet-B4, and at the same time uses ResNet-50 to extract skip connection features of low, medium, and high resolutions. After these features are fused, the final prediction results are obtained through the decoder. The prediction results participate in the calculation of four loss functions, and the loss values and index values of each training are recorded. The index value calculates the index change gradient in the MT-GO module, and the weight update is based on the LT-DWA module, comprehensively considering the historical changes of the loss values, index change trends, and other weight values. Finally, perform backpropagation through the weighted total loss to optimize the network parameters. Repeat the above process in subsequent training until the maximum number of training epochs is reached.
[0051] Furthermore, the multi-scale feature fusion method in model training is as follows: Gradually splice the global feature embeddings output by EfficientNet-B4 with the low, medium, and high resolution skip connection feature maps of ResNet-50; gradually restore the spatial resolution through deconvolution operations, and add 3×3 convolutional layers and ReLU activation functions at each upsampling stage for feature enhancement.
[0052] Perform dynamic weight allocation for multiple loss functions to obtain the current optimal weight value.
[0053] Specifically, in this embodiment, in order to establish an internal connection between the combination and optimization of multiple loss functions, a multi-loss combination strategy with dynamic weight allocation is proposed.
[0054] First, calculate the gradient change rate of the i-th loss function at the current moment t and the previous moment t-1:
[0055]
[0056] where, Li (t) represents the value of the i-th loss at time t, representing the gradient change of the current loss.
[0057] Next, calculate the historical gradient change of each loss function, that is, the difference between the previous two gradient changes:
[0058]
[0059] Then, combine the current gradient change and the historical gradient change, and normalize the combined gradient change of all losses:
[0060]
[0061] where, represents the combined gradient change of all losses at time t, α and β are the weight coefficients of the current gradient and the historical gradient respectively, and after normalization, it can ensure the stability of the influence of the gradient change on the weight adjustment.
[0062] After that, the gradient change rate of the loss function can be calculated:
[0063]
[0064] where, represents the gradient change rate of the i-th loss at time t, and ε is a small positive number used for stable calculation.
[0065] Finally, combine the change trend of the loss function to update the weight of each loss. The update formula of the loss weight is:
[0066]
[0067] where, represents the weight of the i-th loss at time t, and η is the learning rate, which is used to control the influence degree of the change rate on the weight adjustment.
[0068] Take the change trend of the index as the optimization strategy to enhance the dynamic weight allocation.
[0069] Input the index of the previous round of training and the current index into the MT-GO module to calculate the index change gradient. Specifically, in this embodiment, the index includes the degree of training. In order to make full use of this information to guide the weight allocation, a weight optimization method based on the index change is proposed. The MT-GO module includes two parts: the historical index value and the current index value. First, save the index value during model training, and calculate the index change trend at the beginning of the second round of training:
[0070]
[0071] Among them, λ is a dynamic adjustment factor used to amplify the weights that significantly improve the indicators. represents the overall change trend of the indicators. When it indicates that the overall performance has improved, the dynamic adjustment factor λ is used to promote the weight increase of the larger loss term to accelerate the optimization effect. Otherwise, it remains unchanged to ensure the stability of training.
[0072] Normalize the adjusted comprehensive weights and calculate the weighted total loss for backpropagation.
[0073] Compared with the traditional segmentation model, the method in this paper optimizes the global feature extraction using EfficientNet-B4 while retaining the skip connections of ResNet to enhance the multi-scale feature fusion ability. At the same time, precise weighting of multiple loss functions is achieved through the change of training indicators and the dynamic loss weight strategy allocation. The main effects are as follows:
[0074] A hybrid encoder structure based on the combination of EfficientNet-B4 and ResNet is designed. EfficientNet-B4 is used for global feature extraction, and the ResNet skip connections retain local detail information. The combination of the two significantly improves the feature expression ability and segmentation performance.
[0075] A dynamic loss weight allocation strategy is proposed. During the training process, the dynamic changes of the loss function and related indicators are monitored in real time, and the weight allocation of each loss function is adjusted according to the change trends of the loss and training indicators to achieve the collaborative optimization of multiple loss functions, further improving the flexibility and segmentation performance of the model.
[0076] This example is conducted on three retinal vessel segmentation datasets, namely DRIVE, CHASE_DB1, and LES-AV. Compared with other segmentation methods such as DPL-GFT-EFA and DA-TransUNet, LM-TransUNet has higher AC values, approximately 98.81%, 98.75%, and 98.40% respectively. The experiment is implemented using PyTorch 1.13 and runs on a computer with a 3.5GHz Intel i5-13600KF processor, 32GB of 3600MHz DDR4 RAM, and a 12GB NVIDIA RTX4070Ti, with the system environment being Windows 11. The experiment is carried out on three public retinal vessel datasets, DRIVE, CHASE_DB1, and LES-AV. First, a comparison of segmentation accuracy is made with SOTA methods such as DA-TransUNet, DPL-GFT-EFA, and G2Vit. Second, the effectiveness of EfficientNet-B4, LT-DWA, and MT-GO is verified through ablation experiments. Finally, a visual analysis and discussion of the segmentation results are conducted. The relevant parameter settings for the experiment are as follows: The optimizer used is SGD, with a weight decay of 0.0001, a momentum of 0.9, an initial learning rate of 0.01, and the learning rate is updated by exponential decay. Max_Epoch = 200, batch size = 1.
[0077] The relevant parameters of the datasets are shown in Table 1. The DRIVE dataset contains 40 retinal vessel images and their corresponding GTs, with a resolution of 565×584. It is pre-divided into 20 samples for training and 20 samples for testing. The CHASE_DB1 dataset contains 28 retinal vessel images and their corresponding GTs, with a resolution of 990×960. Among them, the training subset contains 20 samples, and the remaining are used as the test subset. The LES-AV dataset contains 22 retinal vessel images and their corresponding GTs, with an image resolution of 1620×1444. Its training set contains 16 images, and the remaining images are used as the test set. To alleviate the overfitting problem, a 512×512 region of the original image is randomly intercepted as the training sample during training.
[0078] Table 1
[0079]
[0080] The segmentation results are compared with the manually segmented GTs. By calculating sensitivity (SE), specificity (SP), accuracy (AC), Jaccard coefficient (JA), and Dice coefficient (DI, also known as F-measure), the degree to which each pixel is correctly classified as background or vessel is measured. The formulas for these metrics are as follows:
[0081]
[0082] Among them, TP represents the number of true positive samples, FP represents the number of false positive samples, TN represents the number of true negative samples, and FN represents the number of false negative samples.
[0083] First, it was compared with twelve methods such as IterMiU-Net, HRD-Net, IMFF-Net, MILU-Net, DA-TransUNet, TransUNext, DPL-GFT-EFA, and G2Vit on the DRIVE dataset. As shown in Table 2, the SE, AC, DI, and JA values of LM-TransUNet all ranked first. The AC value of DA-TransUNet ranked second. It is also an improvement based on the TransUNet structure and integrates spatial and channel attention. G2ViT is a comprehensive model constructed by combining GNN, CNN, and VIT and adding methods such as feature fusion, and its AC value ranked third. DPL-GFT-EFA improves accuracy by image denoising and fusing the gate mechanism and attention mechanism, and HRD-Net enhances features by fusing the outputs of different resolution branches. The AC and other index values of these two methods are relatively low. MILU-Net enhances features by fusing multi-scale detailed features, and TransUNext optimizes the computational complexity of U-Net inspired by ConvNeXt and introduces self-attention and multi-scale global feature fusion to enhance the segmentation performance. The AC and other index values of these two methods are also relatively low. IMFF-Net adopts a four-layer encoder-decoder framework based on attention pool feature fusion, and IterMiU-Net increases samples based on the patch method and reduces parameters based on the internet integrated encoder-decoder structure. The AC and other index values of these two methods are low.
[0084] Then, it was compared with eleven methods including IterMiU-Net, HRD-Net, IMFF-Net, DA-TransUNet, TransUNext, DPL-GFT-EFA, and G2Vit on the CHASE_DB1 dataset. As shown in Table 3, LM-TransUNet has the highest SE (94.71%), AC (98.75%), DI (89.15%), and JA (80.68%). The AC and other index values of DA-TransUNet rank second. It uses a dual attention mechanism in the skip connection to capture finer features, but it is prone to model overfitting. The AC and other index values of G2Vit rank third. It is a comprehensive model that integrates three architectures and reduces information loss through methods such as dilated convolution and edge detection. The AC values of HRD-Net and DPL-GFT-EFA are similar. The former fuses the outputs of different resolution branches through global aggregation operations and uses feature enhancement cascades to integrate features. The diffusion probability learning method proposed by the latter highlights the feature representation. TransUNext is based on the U-shaped architecture and combines the advantages of Transformer and ConvNeXt to improve the segmentation performance. The AC and other index values of TransUNext, IterMiU-Net, and IMFF-Net are all low. Finally, it was compared with eight methods including DA-TransUNet, FSE-Net, and RRWNet on the LES-AV dataset. As shown in Table 4, LM-TransUNet also has the highest AC (98.26%), DI (85.35%), and JA (75.08%). The AC value of TransUNet ranks second. It combines the self-attention mechanism of Transformer and the encoder-decoder structure of U-Net. DA-TransUNet adopts a skip connection with a dual attention mechanism, and its AC value ranks third. The AC value of FSE-Net ranks fourth. It improves the segmentation performance by expanding the network receptive field and eliminating the upsampling operation. RRWNet is jointly trained through module stacking and recursive refinement, and divides the segmentation map into multiple categories to improve the topological consistency of segmentation, and its AC value is low.
[0085] Table 2
[0086]
[0087] Table 3
[0088]
[0089] Table 4
[0090]
[0091] Figure 5 shows the visualization results of image segmentation of LM-TransUNet, TransUnet, and DA-TransUNet on three datasets. As Figure 5 can be seen, each method can effectively segment blood vessels in most regions, but in low-contrast regions, especially the segmentation effect of fine blood vessels is different. From Figure 5 the local enlarged view, it can be seen that TransUNet only relies on the global and local features of the image for segmentation, and its segmentation effect is relatively limited. Especially in the region of fine blood vessels, many tiny blood vessels are misclassified as the background. The segmentation result of DA-TransUnet is slightly better than that of TransUnet. On this basis, it enhances the features by introducing a bidirectional attention mechanism, and its segmentation effect is improved. It can identify some relatively obvious fine blood vessels, but still fails to completely segment all small blood vessels, and some blood vessels show discontinuous phenomena. LM-TransUNet shows the best segmentation effect and can correctly segment most of the fine blood vessels. This is due to the fact that it enhances the sensitivity of the model to low-contrast fine blood vessels through a dynamic integration loss mechanism and effectively fuses multi-scale features, thus significantly improving the segmentation accuracy. It can be seen that LM-TransUNet shows obvious advantages in the segmentation task of retinal blood vessel images, especially in the segmentation of low-contrast fine blood vessels.
[0092] The method of this embodiment uses EfficientNet-B4 and ResNet-50 to construct a hybrid encoder, and fuses its output with the multi-scale feature maps in the skip connection of ResNet-50. At the same time, the LT-DWA and MT-GO modules guide the adjustment of dynamic weights by analyzing the change trends of loss gradients and metrics, and effectively balance the contributions of each loss function. Its effectiveness is proved by experimental comparison on 3 retinal blood vessel datasets. The proposed hybrid encoder is relatively independent of the LT-DWA and MT-GO modules and can be applied to different semantic segmentation models.
[0093] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for segmenting retinal vessel images by improving the TrasnUNet network based on dynamic weight allocation, characterized in that Including: Obtain the retinal vascular images to be segmented from the Kaggle official website and the LES_AV official website; Preprocess the obtained retinal vascular dataset to obtain the processed retinal vascular dataset; Input the retinal vascular images to be segmented into the LM-TransUNet model to obtain the visualization results of image segmentation. Among them, the LM-TransUNet model is obtained by training based on a training set, and the training set includes retinal vascular images and corresponding ground truths; During the process of training the LM-TransUNet model based on the training set, it includes: predicting the retinal vascular images through the initial LM-TransUNet model, calculating multiple loss function values according to the prediction results of the retinal vascular images and the ground truths; inputting the loss function values into the dynamic weight allocation module, and dynamically adjusting the weights by combining the historical loss gradient changes and the index trend information; using the adjusted weighted total loss for backpropagation iteration until convergence to obtain the LM-TransUNet model.
2. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 1, wherein The LM-TransUNet model includes a hybrid encoder, a decoder, an LT-DWA module, and an MT-GO module. Among them, the hybrid encoder includes an EfficientNet-B4 global encoder and a ResNet-50 multi-scale skip connection module, and the dynamic weight allocation includes the LT-DWA module and the MT-GO module; The LT-DWA module is used to dynamically adjust the loss weights according to the historical gradient changes and the current gradient change rate of the loss function; The MT-GO module is used to perform secondary optimization on the weights according to the change trend of the segmentation metrics; The decoder is used to fuse the global encoding features and the multi-scale skip connection features and output the new prediction results.
3. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 2, wherein The implementation method of the hybrid encoder is: Use EfficientNet-B4 as the backbone network to extract global feature embeddings; Retain the three different resolution skip connection feature maps of ResNet-50, corresponding to low, medium, and high spatial resolutions respectively; Gradually fuse the global feature embeddings and the multi-scale feature maps in the decoder.
4. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 3, wherein The multi-scale feature fusion method of the decoder is: Gradually splice the global feature embeddings output by EfficientNet-B4 with the low, medium, and high resolution skip connection feature maps of ResNet-50; Gradually restore the spatial resolution through deconvolution operations, and add a 3×3 convolutional layer and a ReLU activation function at each upsampling stage for feature enhancement.
5. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 2, wherein The weight adjustment method of the LT-DWA module includes: Calculate the gradient change rate of each loss function at the current moment and the previous moment: ; Integrate the historical gradient changes and perform normalization: ; Calculate the gradient change rate of the loss function: ; where α and β are the weight coefficients of the current gradient and the historical gradient, is a small positive number for stable calculation; Update the weights based on the normalization results: ; where η is the learning rate; Normalize the adjusted comprehensive weights and calculate the weighted total loss for backpropagation.
6. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 5, characterized in that, The said index trend information includes the dynamic change values of sensitivity (SE), specificity (SP), accuracy (AC), and Dice coefficient (DI). The historical trend is statistically analyzed by the sliding window method and input into the MT-GO module.
7. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 6, wherein The weight optimization method of the said MT-GO module includes: Calculating the overall index change trend: ; Adjusting the weight according to the index change direction: ; Among them, λ is a dynamic adjustment factor used to amplify the weight that significantly improves the index.
8. The method for segmenting retinal vessel images based on dynamic weight allocation according to claim 2, wherein, The said multiple loss functions include Boundary Loss, Focal Loss, Skeleton Recall Loss, and Cross Entropy Loss. The calculation formula for their weighted total loss is: ; Among them, w bl , w fl , w srl , w ce are the dynamically adjusted weight coefficients.
Citation Information
Cited By
Model training method and device and drivable area detection method and device
CN121789185A
Model training method and device, drivable area detection method and device
CN121789185B