EdgeAttenNet Glomerular Image Precise Segmentation System and Method Based on Camouflaged Target Detection
EdgeAttenNet addresses the challenges of kidney glomeruli segmentation in pathological images by employing PVTv2 with AGBM, EFEM, and AFCM modules to enhance feature extraction and boundary detection, resulting in improved accuracy and robustness in complex scenarios.
Patent Information
- Application Number
- CN202411255905.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-09
AI Technical Summary
The existing glomerulus segmentation methods are difficult to achieve high-precision segmentation when facing complex backgrounds and blurred boundaries, especially the detection ability of small-sized glomerulus is limited, and the existing deep learning models have limitations in boundary recognition and feature extraction.
The EdgeAttenNet model is adopted, combined with PVTv2 backbone network and innovative modules such as AGBM, EFEM, and AFCM, and the accuracy and robustness of glomerular segmentation are improved through multi-scale feature extraction, boundary detection and context fusion.
It significantly improves the accuracy and robustness of glomerular segmentation, can better handle complex morphology and blurred boundaries, enhances the detection ability of small targets, and provides more efficient medical image analysis support.
Smart Images

Figure CN119206218B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image analysis, and particularly to an EdgeAttenNet glomerular image precise segmentation system and method based on camouflaged target detection. Background Art
[0002] In recent years, with the rapid development of deep learning and computer vision technologies, significant progress has been made in the field of medical image analysis. The precise segmentation of glomeruli is crucial for the diagnosis, condition assessment, and treatment plan formulation of chronic kidney diseases. However, due to the complex morphology, irregular boundaries, and high similarity to surrounding tissues of glomeruli in pathological images, traditional image segmentation methods are difficult to achieve ideal results, and more advanced technical means are urgently needed.
[0003] Deep learning models, especially those based on convolutional neural networks (CNNs), have demonstrated excellent performance in image segmentation tasks. Classic network architectures such as U-Net, FCN, and DeepLab3+ have also made significant progress in medical image segmentation. However, when faced with glomerular pathological images, these models still face many challenges, such as blurred boundaries, detail loss, and difficulty in identifying small targets (such as small-sized glomeruli). In addition, how to effectively utilize multi-scale features and simultaneously process target objects of different sizes remains the focus and difficulty of current research.
[0004] Existing glomerular segmentation methods mainly include traditional image processing techniques, machine learning methods, and deep learning methods. Traditional image processing techniques, such as threshold segmentation, edge detection, and region growing, have limited effects in complex backgrounds and are difficult to handle the morphological diversity and structural variability of glomeruli. Machine learning methods such as support vector machines (SVMs) and random forests, although improved, still have limitations in feature extraction and generalization ability.
[0005] In contrast, deep learning-based segmentation methods, especially models based on architectures such as U-Net, perform excellently in medical image segmentation. However, these models still have some deficiencies in the segmentation of the complex structure of glomeruli, such as inaccurate boundary recognition and limited detection ability for small target glomeruli. These challenges mainly stem from the high similarity between glomeruli and surrounding tissues, as well as the morphological diversity and variability of glomeruli in pathological images. Although existing deep learning methods perform excellently in global feature extraction, they still have limitations in the processing of boundary details and small-sized targets. Especially when the glomerular boundary is blurred or fused with the background, the model is prone to misclassification or missed detection.
[0006] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention
[0007] In view of the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide an EdgeAttenNet glomerular image precise segmentation system and method based on camouflaged target detection, and to propose an optimization solution for the complex challenges in glomerular segmentation. This system uses the EdgeAttenNet model, combined with PVTv2 as the backbone network, which can efficiently capture multi-scale features, and through its hierarchical design, enhance the global perception ability and maintain efficient feature extraction performance. The model also introduces an ASPP (Atrous Spatial Pyramid Pooling) module, which uses different dilation rates to expand the receptive field and accurately capture spatial information from details to the global level, especially suitable for the glomerular segmentation task with complex morphology and blurred boundaries.
[0008] The innovative modules in EdgeAttenNet include the AGBM (Attention-Guided Boundary Module) and the EFEM (Edge-Guided and Feature-Enhancement Module), which significantly improve the boundary detection ability and feature expression. The AFCM (Adaptive Fusion of Context Module) further improves the model's accuracy in segmenting glomeruli of different sizes and complex morphologies through the fusion of multi-scale features. The synergistic effect of these modules effectively solves the limitations of traditional models in dealing with small targets and complex boundaries, and significantly improves the segmentation performance and robustness of the model.
[0009] To achieve the above objectives, the present invention adopts the following technical solutions:
[0010] In a first aspect, an EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection includes the following modules:
[0011] A feature extraction module based on multi-level convolution, including five consecutive feature extraction layers (f1, f2, f3, f4, f5), for extracting image features from low-level to high-level layer by layer;
[0012] An attention-guided boundary (AGBM) module, responsible for precise boundary detection, which improves the recognition accuracy of glomerular boundaries in complex scenarios by combining low-level and high-level features (such as f2 and f5);
[0013] An edge-guided feature enhancement (EFEM) module, connected to each feature extraction layer (f2, f3, f4, f5), for enhancing the edge information in the features of each layer to ensure that multi-scale edge features can be accurately captured and processed;
[0014] An adaptive fusion of context (AFCM) module, which concatenates multiple features from the EFEM and gradually fuses context information at different levels to enhance the model's global understanding ability of complex structures.
[0015] These modules work together to form an end-to-end segmentation network structure, designed specifically for high-precision glomerulus recognition and segmentation. By integrating multiple innovative modules such as multi-level convolutional feature extraction, boundary detection, edge enhancement, and context information fusion, the system successfully addresses the limitations of traditional methods in dealing with complex backgrounds, blurred boundaries, and small object recognition, significantly improving the accuracy and robustness of segmentation.
[0016] Furthermore, the feature extraction module adopts the PVTv2 (Pyramid Vision Transformer v2) network architecture, which can extract multi-scale feature information (such as f1 to fn) from the input glomerulus pathological images. These features have different spatial resolutions and semantic levels, corresponding to features of {C1, C2,... Cn} dimensions respectively. Through multi-level spatial downsampling and local attention mechanisms, PVTv2 reduces the computational complexity while maintaining high-efficient feature extraction capabilities, ensuring that the system can process high-resolution medical images. The multi-scale feature extraction strategy of this module enables the system to capture both local details and understand the global structure, thus better handling glomeruli of different sizes and shapes.
[0017] Furthermore, the attention-guided boundary module adopts spatial attention mechanism and channel attention mechanism, and captures multi-scale context information through Atrous Spatial Pyramid Pooling (ASPP). The core process of the attention-guided boundary module includes:
[0018] feat_1 = self.conv1(feat_1) feat_1 = self.spatial_attention(feat_1) *feat_1 feat_4 = self.aspp(feat_4) feat_4 =self.channel_attention(feat_4) *feat_4
[0019] xa = torch.cat((feat_4, feat_1), 1)
[0020] xa = self.block(xa)
[0021] Among them, feat_1 represents low-level features. After being reduced in dimension by a 1x1 convolution and weighted by a spatial attention mechanism, it highlights important spatial location information and captures local details. feat_4 represents high-level features. After extracting multi-scale context information through the ASPP module and expanding the receptive field, it is then weighted by a channel attention mechanism to highlight key channels. feat_1 and feat_4 are concatenated in the channel dimension (torch.cat((feat_4, feat_1), 1)) and passed to the block for processing. The block module fuses multi-scale information through multiple convolution operations, further refining the boundary features to ensure the accuracy of segmentation;
[0022] With the overall design of this module, the model performs excellently in processing glomerular boundaries with complex shapes. The spatial attention mechanism effectively focuses on important spatial location information in the feature map, ensuring the accurate capture of local details; the channel attention mechanism highlights important feature channels related to the segmentation task, improving the accuracy of overall feature expression; and the ASPP module enables the model to handle glomerular structures with different shapes and sizes through the capture of multi-scale context information and the expansion of the receptive field. The overall design ensures the precise detection and segmentation of boundary features by the model in complex backgrounds, significantly enhancing the robustness and performance of the segmentation task.
[0023] Furthermore, the edge-guided feature enhancement module realizes feature enhancement through the following steps:
[0024] Combined with the boundary prediction information, the input feature c is weighted, and the boundary weight att is used to highlight the edge region. The core operation is:
[0025] x = c * att + c
[0026] where c is the input feature and att is the boundary prediction weight. This operation highlights the features related to the glomerular edge by multiplying the channel attention weight att with the original feature map c. The residual connection preserves the original feature information, ensuring that the enhanced features can emphasize important edge information;
[0027] The adaptive pooling adopts a combination of average pooling and max pooling to converge global features at different scales; subsequently, one-dimensional convolution is used to capture the dependencies between channels, generating the fused attention weight:
[0028] wei = self.avg_pool(x)
[0029] wei = self.conv1d(wei.squeeze(-1).transpose(-1, -2)).transpose(-1, -2).unsqueeze(-1)
[0030] wei2 = self.max_pool(x)
[0031] wei2 = self.conv1d2(wei2.squeeze(-1).transpose(-1, -2)).transpose(-1,-2).unsqueeze(-1)
[0032] Among them, self.avg_pool(x) uses adaptive average pooling to aggregate the input feature map x into global features, and the generated result is used for subsequent attention weight calculation; self.max_pool(x) uses adaptive max pooling to downsample the input feature map x into a one-dimensional vector, providing the maximum value representation of the global features; wei.squeeze(-1) removes the redundant dimension, and the transpose operation is used to adjust the tensor dimension order to adapt to the input format of the one-dimensional convolution. self.conv1d represents processing the global features through one-dimensional convolution to generate the attention weights between channels.
[0033] Subsequently, the channel weights generated by average pooling and max pooling are fused to ensure that the model is smoother and more robust when integrating different global features:
[0034] wei = 0.5 * wei + 0.5 * wei2
[0035] Among them, wei and wei2 are the attention weights obtained through average pooling and max pooling respectively. In order to make full use of the advantages of the two different pooling strategies, the two are weighted and fused with a weight of 0.5 to ensure the comprehensiveness and robustness of the global features.
[0036] Next, the fused weights are normalized through the Sigmoid activation function:
[0037] wei = self.sigmoid(wei)
[0038] Among them, the Sigmoid function scales the attention weights to the range of [0,1], ensuring the stability and flexibility of the attention weights, thereby effectively adjusting the weight distribution between channels, highlighting the key edge features and suppressing the irrelevant background information.
[0039] Finally, further convolution processing is performed on the feature map x after attention weighting, and the enhancement is finally completed through the following operations:
[0040] x = x * wei + x
[0041] This step performs a weighting operation on the feature map x. This method not only enhances the edge features but also maintains the integrity of the input features through residual connections, ensuring that the model has stronger robustness and adaptability for detecting complex boundaries.
[0042] This module effectively captures the dependencies between channels through adaptive pooling, dynamically calculated convolution kernel sizes, and channel attention mechanisms, significantly improving the detection accuracy of the model for complex glomerular boundaries. This module introduces edge prediction information, further strengthening the expression of edge features through attention weights related to the edges, and achieving precise enhancement of edge semantics. This edge guidance mechanism can highlight the regions highly relevant to the glomerular boundaries in the feature map and suppress irrelevant background information. At the same time, the original feature information is retained through the residual connection mechanism, ensuring that the enhanced features can better capture key edge details without losing global information.
[0043] Furthermore, the adaptive fusion context module realizes the aggregation, feature recombination, and enhancement of context information through the following steps, and introduces new context information to improve the feature expression ability of the model and the understanding ability of complex structures:
[0044] First, the low-level feature (lf) and the high-level feature (hf) are concatenated along the channel dimension (i.e., dim = 1) through the torch.cat operation, and then a 1x1 convolution (conv1_1) is used to fuse the concatenated features:
[0045] x = self.conv1_1(torch.cat((lf, hf), dim=1))
[0046] Among them, lf is the low-level feature, hf is the high-level feature, torch.cat concatenates them along the channel dimension, and conv1_1 is a 1x1 convolution operation for the preliminary fusion and dimension unification of features;
[0047] The fused feature map is divided into four sub-feature maps with equal channels in the channel dimension using torch.chunk:
[0048] xc = torch.chunk(x, 4, dim=1)
[0049] Among them, the torch.chunk function divides x into 4 sub-feature blocks along the channel dimension for subsequent multi-scale feature fusion operations.
[0050] Perform convolution operations on the first and second sub-feature blocks, and use the AttentionFuoin2 module to weight and fuse the two feature maps through the attention mechanism:
[0051] x0 = self.conv3_1(xc[0], xc[1])
[0052] Among them, conv3_1 is the AttentionFuoin2 module, which is used to generate weights through Softmax to weight xc[0] and xc[1], ensuring that important local details are retained during the fusion process.
[0053] Perform multi-scale convolution operations on the second, third, and the feature block x0 after the first step of fusion, and use the AttentionFuoin3 module to handle the fusion of the three feature blocks:
[0054] x1 = self.dconv5_1(xc[1], x0,xc[2])
[0055] Among them, dconv5_1 is the AttentionFuoin3 module, which fuses xc[1], x0, and xc[2] in a weighted manner, ensuring the interaction between features and strengthening the capture of multi-scale context information.
[0056] Similarly, perform convolution operations on the third, fourth, and the feature blocks after fusion in the previous steps, and use the AttentionFuoin3 to continue handling the fusion of the three feature blocks:
[0057] x2 = self.dconv7_1(xc[2], x1,xc[3])
[0058] Among them, dconv7_1 is also the AttentionFuoin3 module, which is used to further fuse xc[2], x1, and xc[3], strengthening the context dependence between different scales.
[0059] Finally, perform convolution operations on the fourth feature block and the feature block x2 after the third step of fusion to complete the weighted fusion of the last two features:
[0060] x3 = self.dconv9_1(xc[3], x2)
[0061] Among them, dconv9_1 is the AttentionFuoin2 module, which is used to fuse xc[3] and x2, and retain key features through residual connections.
[0062] Re - splice the four feature blocks after being fused through multiple convolution operations into a complete feature map, and further compress the channels through convolution operations to complete the final feature fusion:
[0063] xx = self.conv1_2(torch.cat((x0, x1, x2, x3), dim = 1))
[0064] Among them, torch.cat re - splices the four feature blocks in the channel dimension, and conv1_2 uses 1x1 convolution to compress the channels, further fusing multi - scale feature information to ensure the compactness of feature representation.
[0065] Finally, through further convolution processing, refine the fused feature map to enhance the feature expression ability:
[0066] x = self.conv3_3(xx, x)
[0067] Among them, conv3_3 is a convolution operation used to refine the fused features, ensuring that the model can make full use of high - level and low - level features to improve the segmentation accuracy in complex scenarios.
[0068] The core logics of AttentionFuoin2 and AttentionFuoin3 are as follows:
[0069] AttentionFuoin2 generates dynamic weights through Softmax, performs weighted fusion on the two input features, and retains the original features through residual connections:
[0070] w = self.act(x1 + x2)
[0071] x1 = w[:, 0,...].unsqueeze(1) * x1+ x1
[0072] x2 = w[:, 1,...].unsqueeze(1) * x2+ x2
[0073] Among them, the weights generated by the act function are normalized through Softmax and weighted on x1 and x2. The residual connection ensures that the information of the original features is not lost during the fusion process.
[0074] AttentionFuoin3 processes three input feature maps in a similar way, generates three-dimensional weights, weights each feature map, and retains key features through residual connections:
[0075] w = self.act(x1 + x2 + x3)
[0076] x1 = w[:, 0, ...].unsqueeze(1) * x1+ x1
[0077] x2 = w[:, 1, ...].unsqueeze(1) * x2+ x2
[0078] x3 = w[:, 2, ...].unsqueeze(1) * x3+ x3
[0079] Among them, Softmax generates a three-dimensional weight matrix for weighted fusion of the three features, ensuring the integration of multi-scale information. The residual connection also retains the original features, enhancing the robustness of the model.
[0080] The adaptive fusion context module can effectively capture multi-scale information in complex glomerular structures through the adaptive fusion of multi-scale features and the dynamic weight generation mechanism combining AttentionFuoin2 and AttentionFuoin3. The design of the residual connection ensures that key information is not lost during the feature fusion process, improving the segmentation ability of the model in complex backgrounds.
[0081] Second, a segmentation method for an EdgeAttenNet glomerular image precise segmentation system based on camouflaged object detection, the method comprising the following steps:
[0082] Step 1: Use PVTv2 (Pyramid Vision Transformer v2) as the feature extraction module to perform multi-scale feature extraction on the input glomerular pathological image. PVTv2 extracts multi-scale features layer by layer through its hierarchical Transformer structure, which can not only effectively capture local details (such as the texture and edge information of glomeruli), but also provide global context information to adapt to the feature requirements of glomeruli of different sizes and shapes. The multi-scale feature extraction ability of PVTv2 helps to improve the adaptability of the model to complex scenarios while maintaining a low computational complexity;
[0083] Step 2: Use the Attention-Guided Boundary Module (AGBM module) to perform boundary enhancement processing on multi-scale features. This module combines spatial attention mechanism, channel attention mechanism and ASPP (Atrous Spatial Pyramid Pooling) mechanism, captures multi-scale context information through different dilation rates, and expands the receptive field of the model. The spatial attention mechanism is used to strengthen the local details related to the glomerular edge in the low-level features, while the channel attention mechanism is used to enhance the global information expression in the high-level features, thereby improving the detection accuracy of edge features and ensuring more accurate segmentation of the glomerular boundary;
[0084] Step 3: Through the Edge-Guided Feature Enhancement Module (EFEM module), further enhance the features related to the glomerular edge. The EFEM module combines boundary prediction information to generate channel-level attention weights, and dynamically adjusts feature expressions through the edge-guided mechanism. EFEM uses a dual-pooling strategy (average pooling and max pooling) to comprehensively capture global and local information, effectively highlighting key features, ensuring that the edge details of the glomerulus are strengthened, and further improving the segmentation accuracy;
[0085] Step 4: Use the Adaptive Fusion Context Module (AFCM module) to adaptively fuse features at different levels. The AFCM captures multi-scale context information by combining convolutional operations with different dilation rates, and dynamically weights and fuses features through the attention mechanism. During the feature fusion process, the AFCM can focus on important context information, suppress irrelevant background noise, and thus improve the segmentation performance of the model in complex glomerular structures;
[0086] Step 5: Finally, decode to generate the glomerular segmentation result and apply morphological operations to optimize the segmentation result.
[0087] This method realizes the accurate segmentation of the glomerular structure by gradually refining and enhancing features. Each step specifically addresses specific challenges in glomerular segmentation; this progressive feature processing and fusion strategy not only improves the segmentation accuracy but also enhances the adaptability of the model to different types and degrees of diseased glomeruli.
[0088] In a third aspect, a training method for an EdgeAttenNet glomerular image accurate segmentation system based on camouflage object detection includes:
[0089] Step 1: Standardize the input glomerular pathological image III to ensure that the eigenvalue distributions of the input images are consistent, and improve the training effect and stability of the model. The standardization formula is as follows:
[0090]
[0091] Among them, μ is the mean of the image, and σ is the standard deviation. Through this step, the brightness and contrast differences between different images are eliminated, enabling the model to process different images on the same scale and ensuring the consistency of feature extraction;
[0092] Step 2: Load the EdgeAttenNet model and initialize the feature extraction module with the pre-trained PVTv2 (Pyramid Vision Transformer v2) weights. PVTv2 is pre-trained on a large-scale image dataset. Through its multi-scale feature extraction ability, it provides rich context information for glomerular images, especially having significant advantages in capturing complex structures and details. The initialization of the model helps to accelerate the training convergence speed and improve the initial segmentation performance;
[0093] Step 3: Adopt a weighted combination of Cross-Entropy Loss and Dice Loss to optimize the objective function of the segmentation task. Cross-Entropy Loss focuses on pixel-level classification accuracy, while Dice Loss pays more attention to the overlap degree and matching accuracy of the target region. The loss function is defined as:
[0094]
[0095] Among them, LCE is the Cross-Entropy Loss, LDice is the Dice Loss, and α is the balance factor. By adjusting α, a good balance can be achieved between global classification accuracy and local region matching. Especially in the case of class imbalance, Dice Loss can effectively prevent the omission of small targets (such as glomeruli);
[0096] Step 4: Use the Adam optimizer to update the parameters. The Adam optimizer combines the advantages of momentum and adaptive learning rate adjustment and has good robustness in dealing with complex non-convex optimization problems. Set the initial learning rate:
[0097]
[0098] The adaptive learning rate mechanism of Adam helps to improve the training convergence speed and reduce the need for manual adjustment of the learning rate;
[0099] Step 5: Adopt the Polynomial Decay strategy to dynamically adjust the learning rate to ensure that the model can reasonably control the learning rate at different stages of training. The learning rate update formula is:
[0100]
[0101] Among them, lr0 is the initial learning rate, t is the current iteration number, T is the total iteration number, and β is a hyperparameter that controls the decay rate. This strategy can maintain a low learning rate in the later stage of training, avoid the model from falling into oscillation or overfitting, and at the same time ensure fine-grained feature adjustment;
[0102] Step 6: In each training iteration, first perform forward propagation to calculate the loss function value of the current batch. Then, update the model parameters according to the gradient information of the loss function through backpropagation. After each update, the Adam optimizer adaptively adjusts the learning rate of the parameters to ensure that the model can gradually converge;
[0103] Step 7: Regularly evaluate the model performance on the validation set. By calculating metrics such as the mean intersection over union (mIoU) and the mean Dice coefficient (mDice), quantify the segmentation performance of the model. mIoU measures the segmentation consistency of the model in each category, while mDice focuses on the overlap degree of the target area. These evaluation metrics can accurately reflect the performance of the model in complex scenarios. After the validation evaluation, save the model weights with the best current performance to ensure that the optimal parameter configuration can be selected during the training process. Through this process, overfitting of the model is avoided, and the robustness of the model in practical applications is ensured.
[0104] This method ensures that the EdgeAttenNet model can achieve excellent performance in the glomerular image segmentation task through a multi-step optimization strategy and an adaptive mechanism. Each step is closely combined. Through data preprocessing, model initialization, loss function optimization, dynamic learning rate adjustment, etc., the model can still maintain high precision and stability in complex backgrounds.
[0105] The technical solution adopted by the present invention has the following beneficial effects:
[0106] The present invention proposes the EdgeAttenNet network for the glomerular segmentation problem. Through innovative module design, it effectively improves the segmentation accuracy, especially performs excellently in dealing with complex backgrounds and fuzzy boundaries. EdgeAttenNet has been optimized in terms of the width and depth of the network, and can achieve better detection results.
[0107] The attention-guided boundary module (AGBM) proposed by the present invention combines spatial attention, channel attention, and ASPP, significantly enhancing the ability to perceive the glomerular boundary. The design of this multiple attention mechanism enables the model to more accurately locate the glomerular boundary and overcomes the deficiencies of traditional methods in dealing with fuzzy boundaries.
[0108] The edge-guided feature enhancement module (EFEM) proposed in the present invention realizes fine enhancement of edge regions by combining boundary prediction information and input features, effectively improving the accuracy of segmentation. The design of EFEM not only enhances edge features but also preserves the original feature information through residual connections, avoiding the problem of overemphasizing edges while ignoring internal structures.
[0109] The adaptive fusion context module (AFCM) proposed in the present invention realizes effective capture and dynamic fusion of multi-scale context information, improving the adaptability of the model to glomeruli of different sizes. The design of AFCM allows the model to adaptively adjust the importance of different features according to the input, enabling the system to better handle changing glomerular morphologies and complex background information.
[0110] The overall network architecture of the present invention, through the synergistic effect of multiple innovative modules, effectively solves problems such as detail loss and boundary blur while maintaining high-precision segmentation, providing reliable technical support for glomerular pathological analysis. EdgeAttenNet not only improves the accuracy of segmentation but also enhances the interpretability and scalability of the model, laying a foundation for future applications in other medical image segmentation tasks.
[0111] The multi-objective loss function design adopted in the present invention combines cross-entropy loss and Dice loss, which can better balance pixel-level classification accuracy and region overlap, thus achieving more accurate glomerular boundary localization. This loss function design is particularly suitable for the requirements of medical image segmentation tasks. Brief Description of the Drawings
[0112] Figure 1 It is a schematic structural diagram of the EdgeAttenNet glomerular segmentation system based on the camouflage target detection technology of the present invention;
[0113] Figure 2 It is a schematic structural diagram of the attention-guided boundary module (AGBM) of the present invention;
[0114] Figure 3 It is a schematic structural diagram of the edge-guided feature enhancement module (EFEM) of the present invention;
[0115] Figure 4 It is a schematic structural diagram of the adaptive fusion context module (AFCM) of the present invention;
[0116] Figure 5 It is a comparison diagram of the segmentation effects between the present invention and other methods. Detailed Embodiments
[0117] To make the objectives, technical solutions, and effects of the present invention clearer and more explicit, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0118] The present invention takes glomeruli as a type of camouflage target in pathological images and, for the first time, innovatively applies a deep learning-based camouflage target detection algorithm to glomerular segmentation, proposing an EdgeAttenNet algorithm that adopts a brand-new attention mechanism and network structure to achieve high-precision glomerular segmentation. This innovative application not only improves the accuracy of glomerular segmentation but also opens up a new research direction for the field of medical image analysis.
[0119] Figure 1 Figure 1 is a schematic structural diagram of the EdgeAttenNet glomerular segmentation system based on the camouflage target detection technology of the present invention. As shown in the figure, the EdgeAttenNet glomerular segmentation system includes a feature extraction module, an attention-guided boundary module (AGBM), an edge-guided feature enhancement module (EFEM), and an adaptive fusion context module (AFCM). These modules work together to form an end-to-end segmentation network structure for achieving high-precision glomerular recognition and segmentation.
[0120] The feature extraction module is used to extract multi-scale feature information from the input glomerular pathological images. Specifically, in the feature extraction module, an improved Pyramid Vision Transformer v2 (PVTv2) is adopted to complete the feature extraction work. As an advanced vision Transformer architecture, while maintaining the powerful feature extraction ability of the original Transformer, PVTv2 significantly reduces the computational complexity by introducing spatial downsampling and local attention mechanisms. This enables the model to efficiently process high-resolution medical images.
[0121] The feature extraction module is represented by PVTv2(i), where i = 1, 2, 3, 4, 5 is the index of the extraction block. The size of the input glomerular pathological image is 224×224×3. PVTv2 generates a series of feature maps {f1, f2,..., f5} by gradually reducing the spatial resolution and increasing the number of channels, where f1 represents the lowest-level feature and f5 represents the highest-level feature. This multi-scale feature extraction strategy enables the model to simultaneously focus on local details and global structures, thus better dealing with glomeruli of different sizes and morphologies. The feature dimensions output by the feature extraction module are respectively: f1: 128×128×64 f2: 64×64×128 f3: 32×32×320 f4: 16×16×512 f5: 16×16×512
[0122] Reference Figure 1 and Figure 2 , the innovation of the Attention-Guided Boundary Module (AGBM) lies in combining the Spatial Attention mechanism (SA) and the Channel Attention mechanism (CA), and introducing Atrous Spatial Pyramid Pooling (ASPP), which enhances the perception ability of glomerular boundaries. The SA mechanism highlights important spatial positions in the feature map, and the CA mechanism weights the channels to further optimize the feature expressions related to boundaries. In addition, the ASPP module expands the receptive field through multi-scale convolutions, effectively dealing with glomeruli of different sizes and morphologies. Through the fusion and refinement of multi-scale features, this module achieves accurate segmentation of boundaries in complex backgrounds.
[0123] Reference Figure 1 and 3 , the innovation of the Edge-Guided Feature Enhancement Module (EFEM) is to achieve fine enhancement of the edge region by combining boundary prediction information and input features. EFEM adopts a dual pooling strategy, capturing significant features through max pooling while obtaining the overall feature distribution through average pooling, thus enhancing the edges while retaining global context information. This design effectively improves the model's ability to recognize glomerular edges, avoiding overemphasizing the edges while ignoring the internal structure. In addition, the module retains the original feature information through residual connections, ensuring the integrity of the input features while enhancing the edges, and enhancing the robustness and adaptability of the model in complex boundary detection.
[0124] Reference Figure 1 and Figure 4 , the innovation of the Adaptive Fusion Context Module (AFCM) is to capture multi-scale context information through convolutional operations with different dilation rates and achieve dynamic weighted fusion of features. This design allows the system to adaptively adjust the fusion weights according to the different scales and context information of the input features, thus enhancing the model's ability to understand complex glomerular structures. Especially when dealing with glomeruli with large variations in size and morphology, the AFCM module shows superior adaptability. In addition, the introduction of the Attentional Feature Fusion (AFF) technique further enhances the accuracy of feature fusion, ensuring that important features can be dynamically highlighted during multi-scale information fusion and improving the segmentation performance of the model.
[0125] To verify the glomerular segmentation performance of the system and method of the present invention, we conducted a comprehensive experimental evaluation on the KPIS2024 dataset. The KPIS2024 dataset contains 60 high-resolution whole-slide images of rat kidneys from 20 rodents, covering three chronic kidney disease (CKD) models and normal kidney tissues. These images were stained with PAS and captured at 40x magnification using a Leica SCN400 slide scanner. This dataset represents different CKD conditions and stages, providing diverse samples for kidney pathological image segmentation research.
[0126] We compared the EdgeAttenNet method proposed in the present invention with a variety of cutting-edge methods, including UNet, Segformer, DeepLabV3+, SINetV2, PFNet, and BGNet, etc. To ensure the rigor and fairness of the experiment, the prediction results of all benchmark methods were either obtained directly from the original authors or re-implemented by using publicly available source codes. This approach ensured the reproducibility of our comparative analysis and reduced potential biases in model performance evaluation.
[0127] We adopted multiple metrics widely used in medical image segmentation to evaluate the segmentation results, including mean pixel accuracy (mPA), mean intersection over union (mIoU), and mean Dice coefficient (mDice). These metrics comprehensively evaluated the performance of the model from the perspectives of pixel-level classification accuracy, region overlap degree, and region similarity.
[0128] The experimental results are shown in Table 1 below. EdgeAttenNet achieved the best results in all metrics, demonstrating significant performance advantages. Specifically, without using pre-trained weights, EdgeAttenNet achieved an mIoU of 83.10% on the KPIS2024 dataset, which is approximately 5% higher than the 78.03% of the traditional UNet model. This result not only highlights the superiority of the EdgeAttenNet architecture but also proves its excellent generalization ability and efficiency in processing complex visual datasets.
[0129] Table 1 Quantitative comparison and evaluation of method performance - without using pre-trained weights
[0130]
[0131] More notably, when using pre-trained weights as shown in Table 2 below, the performance of EdgeAttenNet is further improved. In this case, EdgeAttenNet achieved an mIoU of 91.23%, which is 3.5% higher than 87.58% of UNet. This significant performance improvement not only demonstrates the robust performance of the EdgeAttenNet model structure but also showcases its ability to effectively utilize pre-trained knowledge.
[0132] Table 2 Quantitative comparison and evaluation of method performance - using pre-trained weights
[0133]
[0134] In terms of qualitative analysis, we conducted a visual comparison of representative samples in the KPIS2024 dataset. The results showed that traditional segmentation methods often have difficulties in accurately depicting the boundaries of glomeruli, frequently leading to over-segmentation or under-segmentation problems. This is particularly evident in areas where glomeruli have similar visual characteristics to the surrounding tissues. In contrast, models specifically designed for camouflaged object detection demonstrated stronger capabilities to distinguish these subtle differences, thus improving segmentation accuracy.
[0135] Our EdgeAttenNet method effectively addresses these limitations by leveraging advanced attention mechanisms and boundary-aware features. As Figure 5 shown, EdgeAttenNet consistently produced segmentation results highly consistent with the ground truth annotations. It can accurately capture the complex morphological changes of glomeruli, precisely depict the boundaries even in challenging situations with tissue overlap or similarity, and retain the fine structural details crucial for accurate pathological analysis. This performance highlights the effectiveness and robustness of our model architecture.
[0136] This qualitative comparison not only confirms the previously presented quantitative results but also emphasizes the practical significance of the improvement of our method. The enhanced segmentation quality provided by EdgeAttenNet may significantly assist pathologists in their analysis, potentially leading to more accurate diagnoses and a better understanding of kidney diseases at the glomerular level. Additionally, this demonstration of excellent performance indicates that EdgeAttenNet has significant advantages in handling complex medical image segmentation tasks, especially when dealing with objects with camouflaged features.
[0137] To further validate the effectiveness of the proposed model, we also conducted comprehensive ablation experiments. These experiments aimed to evaluate the effects of each module we proposed and how they work together to improve the overall performance. All ablation experiments were conducted without using pre-trained weights to ensure an unbiased assessment of the intrinsic learning ability and architectural merits of each model.
[0138] The ablation experiment results are shown in Table 3. Removing the AFCM module causes the mIoU to decrease by 0.22%, removing the AGBM module causes the mIoU to decrease by 0.31%, and removing the EFEM module causes the mIoU to decrease by 0.37%. These results indicate that each module plays a crucial role in improving the segmentation performance. In particular, the combination of AGBM + EFEM + AFCM, as the core innovation of this study, achieves the highest mIoU (83.1%) among all tested model combinations. This result further emphasizes the complementarity of the AGBM, EFEM, and AFCM modules, which jointly enhance the model's ability to handle the segmentation of camouflaged targets in complex backgrounds.
[0139] Table 3 Ablation Experiment Results
[0140]
[0141] In addition, we also tested the performance of the model on input images with different resolutions. The results show that EdgeAttenNet can still maintain a high segmentation accuracy when processing high-resolution images (such as 2048×2048), while the computing time only increases slightly, demonstrating the efficiency and scalability of the model. This feature is particularly important for processing large-scale pathological image datasets.
[0142] In summary, the EdgeAttenNet glomerular segmentation system based on camouflaged target detection technology proposed in this invention effectively solves the challenges faced by existing methods in processing complex medical images by innovatively applying camouflaged target detection technology to medical image analysis. This system not only improves the accuracy and efficiency of glomerular segmentation but also provides new ideas and methods for other medical image analysis tasks, opening up a new research direction for applying camouflaged target detection technology to medical image analysis.
[0143] After considering the specification and practicing the disclosed solutions herein, those skilled in the art will readily conceive of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in this technical field not disclosed in this disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.
Claims
1. An EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection, characterized in that, It includes the following modules: The feature extraction module contains five consecutive convolutional feature extraction layers f1, f2, f3, f4, and f5. By extracting feature information at different depths layer by layer, the extracted features gradually capture the details, local features, and high-level semantic information in the image from shallow to deep; the features at different levels are used for processing and fusion by subsequent different modules; The attention-guided boundary module obtains information from the low-level feature f2 and the high-level feature f5 in the feature extraction module; the low-level feature f2 usually contains more edge and local texture information, while the high-level feature f5 contains rich context and semantic information; by fusing the features of the low-level feature f2 and the high-level feature f5, the glomerular boundary in a complex background is recognized; The attention-guided boundary module adopts a spatial attention mechanism and a channel attention mechanism, and captures multi-scale context information through the atrous spatial pyramid pooling ASPP; feat_1 represents the low-level feature, which is dimensionally reduced by a 1x1 convolution and then weighted by the spatial attention mechanism to highlight important spatial position information and capture local details; feat_4 represents the high-level feature, which extracts multi-scale context information through the ASPP module to expand the receptive field, and then is weighted by the channel attention mechanism to highlight the key channels; The edge-guided feature enhancement module is connected to the four feature extraction layers f2, f3, f4, and f5 in the feature extraction module; these edge-guided feature enhancement modules process the feature maps from each layer, especially using edge information for feature enhancement, ensuring that the model can extract key edge-related information at multiple scales. Through this multi-scale fusion, the edge-guided feature enhancement module can effectively capture the edge details and structural features of the target area, ensuring that the model has the ability to process multi-scale edge features in a complex scene, thereby improving the edge recognition accuracy in the segmentation task; The adaptive fusion context module receives the enhanced features f2, f3, f4, and f5 at different levels processed by the edge-guided feature enhancement module and performs multi-scale context information fusion on these features; the core role of this module is to capture context information from different levels through standard convolution and dilated convolution operations at different scales, ensuring that the model can effectively obtain global semantic information under a larger receptive field; The adaptive fusion context module fuses features from different scales through an adaptive mechanism. This fusion mechanism can dynamically adjust the weights according to the importance of the features, thereby improving the feature expression ability and the utilization efficiency of context information; by combining context information at different levels and different scales, the adaptive fusion context module significantly improves the feature expression ability and segmentation accuracy of the model in a complex background.
2. The EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection according to claim 1, wherein The feature extraction module adopts the PVTv2 network architecture, which can extract multi-scale feature information f1, f2, …, fn from the input glomerular pathological images. These features have different spatial resolutions and semantic levels, where f1 represents the lowest-level feature and fn represents the highest-level feature, and the feature dimensions are C1, C2, …, Cn respectively. The PVTv2 network architecture significantly reduces the computational complexity while ensuring a strong feature extraction ability through layer-by-layer spatial downsampling and local attention mechanism, enabling it to efficiently process high-resolution medical images.
3. The EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection according to claim 1, characterized in that, The attention-guided boundary module includes: The spatial attention mechanism is used to identify and emphasize important spatial positions in the feature map, improving the precise localization of the target area. The channel attention mechanism highlights the key information related to boundary detection by weighted selection of feature channels, ensuring the precise capture of the target boundary. Multi-scale convolution operations are adopted to capture context information at different scales, enhancing the multi-scale representation ability of the feature map, thereby improving the boundary segmentation effect in complex scenarios. The attention-guided boundary module is partially represented as: xa = torch.cat((feat_4, feat_1), 1) xa = self.block(xa) Among them, xa is the feature map after fusion processing. The torch.cat() function concatenates and fuses the deep feature map feat_4 and the shallow feature map feat_1 in the channel dimension. feat_4 represents the high-level feature map, providing rich context information, while feat_1 is the low-level feature map, retaining the detailed information of the target. self.block is a predefined feature processing module used to further enhance the feature expression ability after fusion and achieve the precise segmentation of the target boundary.
4. The EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection according to claim 1, characterized in that, The edge-guided feature enhancement module realizes feature enhancement through the following steps: Element-wise multiplication of the input feature c and the boundary prediction weight att is performed to combine the boundary information and enhance the feature expression related to the edge: x = c * att + c Among them, c represents the input feature map, and att is the weight obtained through boundary prediction, used to highlight the important features at the boundary. The input feature map is processed using adaptive average pooling and adaptive max pooling operations to extract global and local information, and channel-level attention weights are generated through one-dimensional convolution. This step enhances the feature expression by adaptively adjusting the weights of different channels.
5. The EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection according to claim 1, characterized in that, The adaptive fusion context module realizes feature fusion through the following steps: Initial fusion of the low-level feature lf and the high-level feature hf is performed through 1x1 convolution operations. First, lf and hf are concatenated in the channel dimension, and the number of channels is compressed through 1x1 convolution operations to obtain the initial fused feature: x = self.conv1_1(torch.cat((lf, hf), dim = 1)) Among them, lf and hf are the low-level and high-level features respectively, and conv1_1 is the 1x1 convolution operation used to reduce the number of channels of the feature map and achieve initial fusion. The fused feature map x is divided along the channel dimension to obtain four feature blocks, and each block will perform convolution operations with different dilation rates to capture context information at different scales: xc = torch.chunk(x, 4, dim=1) Among them, torch.chunk is an operation in PyTorch for splitting a tensor into multiple chunks along a specified dimension; Convolution operations with different dilation rates are respectively applied to the chunked features to capture context information at different scales; the features generated by different dilated convolutions are weighted through an attention mechanism to dynamically adjust the weights of each feature block, thereby achieving adaptive fusion of multi-scale information.
6. A segmentation method for the EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection as described in claim 1, characterized in that, Including: Step 1: Use the feature extraction module to perform multi-scale feature extraction on the input image based on PVTv2; PVTv2 extracts multi-scale features layer by layer through its hierarchical structure and has good global context capture ability; PVTv2 can provide efficient global feature representation while maintaining low computational complexity, adapting to the multi-scale feature requirements in complex scenes; Step 2: Use the attention-guided boundary module to perform boundary enhancement processing on the multi-scale features; the attention-guided boundary module combines spatial attention, channel attention, and the ASPP mechanism to capture multi-scale context information and highlights the features related to the glomerular edge through the attention mechanism; this module strengthens the low-level features through spatial attention and processes the high-level features using the ASPP module to enhance the perception ability of edge features, thereby improving the accurate segmentation of the glomerular boundary; Step 3: Through the edge-guided feature enhancement module, perform edge-guided enhancement processing on the input features; this module uses the boundary prediction information and the adaptive pooling mechanism to generate channel-level attention weights to further strengthen the feature expression related to the edge and improve the segmentation accuracy; Step 4: Use the adaptive fusion context module to perform adaptive fusion on features at different levels; by combining convolution operations with different dilation rates, capture multi-scale context information, and achieve dynamic weighted fusion through the attention mechanism to further enhance the feature expression ability and improve the accurate segmentation of the glomerular structure by the model; Step 5: Finally, decode the fused features to generate an accurate glomerular segmentation result.
7. The segmentation method of the EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection according to claim 6, characterized in that, The edge-guided feature enhancement module adopts a channel attention mechanism that combines adaptive pooling and one-dimensional convolution, where the kernel size k of the one-dimensional convolution is determined by the number of feature channels C, and the calculation formula is: C represents the number of feature channels of the image, and the log function calculates the logarithm of these channel numbers; the formula calculates t by adding 1 to the logarithmic result, dividing by 2, and taking the integer part, and finally doubles the t value and adds 1 to determine the convolution kernel size k, ensuring that k is always an odd number, which helps to maintain the spatial symmetry of the convolution operation.
8. The segmentation method of the EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection according to claim 6, characterized in that, The adaptive fusion context module uses the AttentionFusion2 and AttentionFusion3 components to achieve dynamic weighted fusion of features, and its core operation is: AttentionFusion2(x1, x2) = Conv(w1·x1 + w2·x2 + x1 + x2) AttentionFusion3(x1,x2,x3) = Conv(w1·x1+w2·x2+w3·x3+x1+x2+x3) where w = Softmax(Conv1x1(x1+x2+x3)) is the attention weight. This adaptive fusion strategy can dynamically adjust the fusion weight according to the importance of the input features.
9. A training method for an EdgeAttenNet glomerular image precise segmentation system based on camouflaged target detection as claimed in claim 1, characterized in that, Including: Step 1: Data preprocessing: Standardize the input glomerular pathological image I: where μ is the mean of the image and σ is the standard deviation. This step ensures that the input feature values of different images are on the same scale, improving the training effect of the model. Step 2: Model initialization: Load the EdgeAttenNet model and initialize the feature extraction module with the pre-trained PVTv2 weights. After being pre-trained on a large-scale image dataset, PVTv2 can provide richer context information for multi-scale feature extraction of glomerular images, contributing to the rapid convergence and performance improvement of the model. Step 3: Loss function definition: Adopt a weighted combination of cross-entropy loss and Dice loss to optimize the objective function of the segmentation task, ensuring that the model can still achieve good segmentation results in the case of class imbalance. The loss function is defined as: L = α·L CE +(1 - α)·L Dice where LCE is the cross-entropy loss, LDice is the Dice loss, and α is the balance factor. This combination can take into account both the accuracy of global classification and the precision of local region matching. Step 4: Optimizer selection: Use the Adam optimizer. The Adam optimizer has good robustness and fast convergence when dealing with complex optimization problems, and the initial learning rate lr = 1e-4. Step 5: Learning rate scheduling: Adopt a polynomial decay strategy to dynamically adjust the learning rate, gradually reducing the learning rate as the training progresses to prevent overfitting or oscillation in the later stage of training. The learning rate update formula is: where t is the current iteration number, T is the total iteration number, lr_0 is the initial learning rate, and β is the hyperparameter controlling the decay rate. This scheduling strategy can ensure that the model is more stable for fine-tuning in the later stage of training. Step 6: Iterative training: In each training iteration, perform forward propagation to calculate the loss, and then update the model parameters through backpropagation. Step 7: Validation and evaluation: Regularly evaluate the model performance on the validation set, calculate metrics such as the mean intersection over union and the mean Dice coefficient to measure the segmentation performance of the model. After each round of evaluation, save the model parameters with the best performance according to the validation results. This process ensures that the model can be continuously optimized during training and avoid overfitting.
Citation Information
Patent Citations
Camouflage target image segmentation method and system based on multilevel feature fusion
CN116703950A
Arophic gastritis area segmentation method based on multi-scale boundary refinement and fusion
CN118115490A