Multi-scale self-attention mechanism-based glomerular segmentation method
By improving the U-Net network through the multi-scale self-attention mechanism and edge-aware feature fusion module, the problem of low efficiency of glomerular segmentation in traditional methods is solved, high-precision segmentation of diseased glomeruli is achieved, and the edge recognition ability and the meticulousness of the segmentation results are enhanced.
Patent Information
- Application Number
- CN202510556674.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional methods are inefficient in glomerular lesion segmentation and have difficulty achieving accurate segmentation, especially in the face of significant heterogeneity and blurred edges of diseased glomeruli.
A glomerulus segmentation method based on the multi-scale self-attention mechanism is adopted. The multi-scale self-attention feature extraction module and the edge-aware feature fusion module are combined to improve the U-Net network. The segmentation accuracy is improved through multi-scale feature extraction and edge deep supervision training.
It effectively addresses lesion heterogeneity, improves the segmentation accuracy of complex lesions and edge segmentation accuracy, reduces the loss of lesion features during upsampling, and improves the meticulousness and robustness of segmentation results.
Smart Images

Figure CN120655665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image technology, and in particular to a glomerulus segmentation method based on a multi-scale self-attention mechanism. Background Art
[0002] Chronic kidney disease (CKD) is a common and growing disease worldwide, with a prevalence of approximately 13.4%. It is one of the fastest-growing diseases in terms of morbidity and mortality worldwide and is projected to become the fifth leading cause of death by 2040. The glomerulus is a key structure in the kidney that filters blood and has important physiological functions, such as maintaining fluid balance, regulating blood pressure, and acid-base balance. Accurate identification and segmentation of glomerular lesions are crucial in renal pathology. However, traditional pathology image scanning and manual observation methods are time-consuming, labor-intensive, and inefficient.
[0003] With the continuous advancement of image processing technology and computing power, the application of deep learning in computer-aided diagnosis has attracted widespread attention. In recent years, deep learning-based segmentation algorithms have been applied to the segmentation of diseased glomeruli and have made significant progress in this field. Compared with traditional manual interpretation methods, deep learning methods have shown great potential in medical image analysis. They can not only effectively reduce the workload of physicians and the influence of human factors, but also improve diagnostic efficiency and the reliability of screening results. However, we still need to recognize the limitations of deep learning in cellular medical image analysis, which require further optimization and breakthroughs. In recent years, with the maturity and widespread application of deep learning technology, some progress has been made in addressing these issues. Currently, most glomerular segmentation datasets are based on U-Net designs, which have high segmentation accuracy. However, due to the significant heterogeneity and blurred edges of diseased glomeruli, traditional networks still have difficulty achieving accurate segmentation.
[0004] To this end, we designed a glomerulus segmentation method based on a multi-scale self-attention mechanism to provide another technical solution to the above technical problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a glomerulus segmentation method based on a multi-scale self-attention mechanism to solve the technical problems raised in the background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a glomerulus segmentation method based on a multi-scale self-attention mechanism, comprising at least the following steps:
[0007] S1: Acquire a diseased glomerulus dataset, including but not limited to a training set and a test set. The test set and the training set are used to construct and train a model. After generating the optimal model, actual operation is performed, and a dataset of images of glomerular cells to be segmented is also acquired;
[0008] S2: Preprocess the glomerular pathology images in the training dataset;
[0009] S3: Construct a multi-scale self-attention feature extraction module and an edge-aware feature fusion module, and build an improved U-Net network based on the multi-scale self-attention module and the edge-aware feature fusion module;
[0010] S4: Input the preprocessed training set into the improved U-Net network, output the training results, and save the trained model;
[0011] S5: Use the optimal improved U-Net network model to perform glomerulus segmentation and retain the segmentation result image.
[0012] Furthermore, the step S2 at least includes the following steps:
[0013] Unify the size of each image in the dataset so that all images are adjusted to the same preset size;
[0014] Then the images in the dataset are randomly flipped horizontally, randomly flipped vertically, and rotated at random angles;
[0015] Then add image noise, adjust color, and add occlusion;
[0016] Finally, the Canny operator is used to extract multi-scale image edges for edge depth supervision training.
[0017] Furthermore, the multi-scale self-attention feature extraction module constructed in S3 combines the ideas of multi-scale feature information and global spatial context modeling, applies the self-attention mechanism based on the extracted multi-scale lesion features, dynamically allocates feature weights, and improves feature representation, and includes at least the following steps:
[0018] Assume that the input feature map X∈R C×H×W , where C is the number of channels, H and W are the height and width;
[0019] First, use 1×1 convolution to perform channel transformation to generate the initial feature map, and use the GeLU activation function to introduce nonlinear processing, see the following formula:
[0020] X′=σ(W in *X)
[0021] Where Win is a 1×1 convolution kernel, * represents the convolution operation, and σ represents the GeLU activation function;
[0022] Then, 5×5 depth convolution is used to extract the input features of the initial multi-scale module, as shown in the following formula:
[0023] X5×5 =W 5×5 *X′
[0024] Subsequently, three parallel asymmetric convolutions are used to extract multi-scale features. These branches use different sizes of convolution kernels to extract features of different receptive fields, as shown in the following formula:
[0025] X3=W (1,3) *X 5×5 , X3=W (3,1) *X3
[0026] X7=W (1,7 )*X 5×5 , X=W (7,1 )*X7
[0027] X 11 =W (1,11) *X 5×5 , X 11 =W (11,1) *X 11
[0028] Therefore, feature fusion is performed by element-by-element addition, and the input residual information is added to retain the original features. The multi-scale feature expression is:
[0029] X ms =X+X3+X7+X 11
[0030] Based on the fused multi-scale features, the query Q, key K, and value V are extracted through 1×1 convolution respectively, as shown in the following formula:
[0031] Q=W q *X ms ,K=W k *X ms ,V=W v *X ms
[0032] Among them, Wq, Wk and Wv are all 1×1 convolution kernels;
[0033] Then calculate the self-attention weight as follows:
[0034]
[0035] Among them, d k is the dimension of the key, used for scaling to stabilize the gradient;
[0036] The self-attention output is adjusted through an additional 1×1 convolution to obtain an output feature map that combines local multi-scale information and global context information. The final output is:
[0037] X out =W out *(X att +X ms )
[0038] Among them, W out Represents a 1X1 convolution kernel.
[0039] Furthermore, the edge-aware feature fusion module is used to reduce the information loss of edge features during upsampling and improve the edge segmentation accuracy. Constructing the edge-aware feature fusion module includes at least the following steps:
[0040] Assume that the input feature map (X, Y) ∈ R C×H×W , and W1, W2 and W3 represent 3×3 convolution kernels, * represents convolution operation, and σ represents ReLU activation function;
[0041] First, the addition branch is calculated. The addition branch strengthens the preservation of edge details through feature accumulation, enabling the model to more accurately identify the contours of the lesion area. Multi-layer 3x3 convolution is used to extract input features, and skipping is added to the last layer. See the following formula:
[0042] X1=σ(W1*X),Y1=σ(W1*Y)
[0043] X2=σ(W2*X1),Y2=σ(W2*Y1)
[0044] X3=σ(W3*(X2+X1)), Y3=σ(W3*(Y2+Y1))
[0045] The features of all levels are then summed and the dimension is reduced by 1×1 convolution, as shown in the following formula:
[0046] X add =W out *(X1+X2+X3+Y1+Y2+Y3)
[0047] Then the subtraction branch is calculated. The subtraction branch highlights the local significant change area through feature difference calculation, enhances the sensitivity to the morphology of the lesion area, and helps detect structural abnormalities and minor lesions. In addition, the dense connection of the x and y branches allows information to be gradually accumulated, avoiding feature loss and enhancing the feature expression ability of the model. The subtraction branch is similar to the addition branch, using multi-layer 3x3 convolution to extract input features, but the subtraction branch uses subtraction fusion features, see the following formula:
[0048] X diff =W out *(X1+X2+X3)-(Y1+Y2+Y3)
[0049] Feature enhancement: The outputs of the addition and subtraction branches are instance normalized respectively, and the nonlinear features are enhanced through the ReLU activation function, where W4 and W5 represent 1×1 convolution kernels, as shown in the following formula:
[0050] X add′ =o(InstanceNorm(W4*X add ))
[0051] X diff′ =σ(InstanceNorm(W4*X diff ))
[0052] Among them, Instance Normalization means normalization;
[0053] The results of the addition branch and the subtraction branch are synthesized to obtain the final output feature map:
[0054] X out =X add′ +X diff′
[0055] The fused feature map is reasonably distributed after dimensionality reduction. One part is used for edge deep supervision to strengthen the constraint of boundary information and improve edge recognition accuracy. The other part is passed to the upper layer to ensure the integrity of the global features. See the following formula:
[0056] y=W5*X out
[0057]
[0058] Among them, Bce is the binary cross entropy, which is used to calculate the single-scale boundary loss to ensure that the edge information is strengthened, e is the edge map, The actual edge annotation.
[0059] Furthermore, constructing the improved U-Net network in S3 at least includes the following steps:
[0060] A multi-scale self-attention module is introduced in each jump link of U-Net, and the outputs of these multi-scale self-attention modules are summed layer by layer as the bottleneck feature;
[0061] The feature map generated by each multi-scale self-attention module is not only used as a bottleneck feature, but also passed as input to the edge-aware feature fusion module. The edge-aware feature fusion module combines the model's upsampled feature map with the multi-scale self-attention feature map to extract richer feature information and generate new fused features.
[0062] Through the edge-aware feature fusion module, the generated fusion feature map will be upsampled and gradually restored to the original image resolution, ultimately obtaining a high-quality segmentation output result.
[0063] Furthermore, the S4 at least includes the following steps:
[0064] The training set is divided into multiple batches using the batch training method, where the training validation batch is set to 8;
[0065] Traversing all the images in the training set once is considered an iteration;
[0066] The total loss of the model is the edge loss of each scale plus the segmentation head loss. The loss function can be expressed as:
[0067]
[0068] Where y represents gt, Represents the prediction result, e i Represents gt edge images of different scales, Indicates the predicted glomerular margin.
[0069] An improved U-Net is used as the basic network for training, and the basic network includes a backbone network, a multi-scale self-attention mechanism and an edge-aware feature fusion module.
[0070] Furthermore, the S5 at least includes the following steps:
[0071] Input the test set to be segmented, segment the glomerular cell image using the trained improved U-Net segmentation model, and obtain the output mask map;
[0072] The root output mask map draws the glomerulus target edge on the original image and saves the final result.
[0073] Compared with the prior art, the present invention has the following beneficial effects:
[0074] The present invention effectively addresses lesion heterogeneity and improves the segmentation accuracy of complex lesions by combining receptive fields of different scales to extract macroscopic structures and microscopic details. The edge-aware feature fusion module adopts dual-path dense feature fusion and deep edge supervision, which are gradually introduced in the decoding stage to reduce the loss of lesion features during upsampling and improve the accuracy and meticulousness of edge segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0076] Figure 1 is a flowchart of an embodiment of the present invention;
[0077] Figure 2 Flowchart of the glomerular cell image segmentation method of the present invention;
[0078] Figure 3 This is an example diagram of glomerular diseased cells of the present invention;
[0079] Figure 4 This is the structural diagram of the multi-scale self-attention module of the present invention;
[0080] Figure 5 This is a structural diagram of the edge-aware feature fusion module of the present invention;
[0081] Figure 6 This is a diagram of the improved U-Net model structure of the present invention;
[0082] Figure 7 This is a diagram showing the improved U-Net network effect using the present invention. DETAILED DESCRIPTION
[0083] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0084] See also Figure 1-Figure 7 , a glomerulus segmentation method based on a multi-scale self-attention mechanism, comprising at least the following steps:
[0085] S1: Obtain a dataset of diseased glomeruli. The dataset includes but is not limited to a training set and a test set. The training set and the test set are used to build and train the model. After the optimal model is generated, actual operation is performed. A dataset of images of glomerular cells to be segmented is also obtained.
[0086] S2: Preprocess the glomerular pathology images in the training dataset;
[0087] S2 includes at least the following steps:
[0088] Unify the size of each image in the dataset so that all images are adjusted to the same preset size;
[0089] Then the images in the dataset are randomly flipped horizontally, randomly flipped vertically, and rotated at random angles;
[0090] Then add image noise, adjust color, and add occlusion;
[0091] Finally, the Canny operator is used to extract multi-scale image edges for edge depth supervision training.
[0092] In order to compensate for the impact of data set sample imbalance on model recognition performance and avoid network overfitting, the present invention performs enhancement processing on the sample data before training. The enhancement method used in this embodiment is as follows, and different enhancement methods and data may exist in different embodiments.
[0093] 1) Adjust the uniform image size: adjust the dataset to a uniform image size of 512×512; 2) Random horizontal flip: flip the image horizontally and vertically with a probability of 50%; 3) Random translation, scaling and rotation: translate the image (up to 6.25% displacement), scale (up to 20%) and rotate (up to 45) with a probability of 50%; 4) Color adjustment: adjust the hue, saturation and lightness, brightness and contrast of the image with a random probability of 50%; 5) Randomly select one of Gaussian blur, motion blur and median blur with a probability of 20%; 6) Random occlusion: randomly occlude a certain number of small areas on the image with a probability of 20% and fill them with 0 values to simulate the situation of local information loss; 7) Use the Canny operator to extract multi-scale image edges for edge deep supervision training.
[0094] S3: Construct a multi-scale self-attention feature extraction module and an edge-aware feature fusion module, and build an improved U-Net network based on the multi-scale self-attention module and the edge-aware feature fusion module;
[0095] The multi-scale self-attention feature extraction module constructed in S3 combines the ideas of multi-scale feature information and global spatial context modeling. It applies the self-attention mechanism based on the extracted multi-scale lesion features, dynamically assigns feature weights, and improves feature representation. It includes at least the following steps:
[0096] Assume that the input feature map X∈R C×H×W , where C is the number of channels, H and W are the height and width;
[0097] First, use 1×1 convolution to perform channel transformation to generate the initial feature map, and use the GeLU activation function to introduce nonlinear processing, see the following formula:
[0098] X′=σ(W in *X)
[0099] Where Win is a 1×1 convolution kernel, * represents the convolution operation, and σ represents the GeLU activation function;
[0100] Then, 5×5 depth convolution is used to extract the input features of the initial multi-scale module, as shown in the following formula:
[0101] X 5×5 =W 5×5 *X′
[0102] Subsequently, three parallel asymmetric convolutions are used to extract multi-scale features. These branches use different sizes of convolution kernels to extract features of different receptive fields, as shown in the following formula:
[0103] X3=W (1,3) *X 5×5 ,X3=W (3,1) *X3
[0104] X7=W (1,7) *X 5×5 , X7=W (7,1) *X7
[0105] X 11 =W (1,11) *X 5×5 , X 11 =W (11,1) *X 11
[0106] Therefore, feature fusion is performed by element-by-element addition, and the input residual information is added to retain the original features. The multi-scale feature expression is:
[0107] X ms =X+X3+X7+X 11
[0108] Based on the fused multi-scale features, the query Q, key K, and value V are extracted through 1×1 convolution respectively, as shown in the following formula:
[0109] Q=W q *X ms ,K=W k *X ms ,V=W v *X ms
[0110] Among them, Wq, Wk and Wv are all 1×1 convolution kernels;
[0111] Then calculate the self-attention weight as follows:
[0112]
[0113] Among them, d kis the dimension of the key, used for scaling to stabilize the gradient;
[0114] The self-attention output is adjusted through an additional 1×1 convolution to obtain an output feature map that combines local multi-scale information and global context information. The final output is:
[0115] X out =W out *(X att +X ms )
[0116] Among them, W out Represents a 1X1 convolution kernel.
[0117] In Table 1 below, the present invention designed a set of comparative experiments to test the performance of the multi-scale self-attention module in feature extraction, specifically:
[0118] Table 1 Results of feature extraction with multi-scale self-attention module
[0119]
[0120] In the experimental results in Table 1, the use of the multi-scale self-attention module for feature extraction improves DSC by 1.92% compared with the original model.
[0121] Based on the above performance, the process of constructing a multi-scale self-attention module can be expressed as:
[0122] The design of the multi-scale self-attention module has a significant impact on the model's performance. First, channel compression reduces computational costs through 1×1 convolution, making subsequent calculations more efficient, while nonlinear transformations (GeLU activation) enhance feature expression. Subsequently, the introduction of parallel asymmetric convolutions not only reduces model parameters but also enhances the ability to capture vertical and horizontal edge and texture features, making it more adaptable to the morphological heterogeneity caused by glomerular lesions. In addition, depthwise separable convolutions further reduce computational complexity and improve the model's applicability to high-resolution glomerular images.
[0123] After feature fusion, the self-attention mechanism dynamically adjusts feature expression through query-key-value (QKV) calculations, enabling the model to adaptively focus on key areas and improve feature differentiation. Softmax normalization also ensures the stability of attention weights, preventing interference from irrelevant information. To alleviate the vanishing gradient problem and improve information flow efficiency, the module introduces shortcuts during multi-scale feature extraction and self-attention calculations to ensure that key features are not weakened, making training more stable. Overall, the construction of this module not only enhances the model's ability to identify lesions but also achieves a good balance between computational efficiency and feature expression.
[0124] After the multi-scale self-attention module is built, the edge-aware feature fusion module is then built;
[0125] Constructing an edge-aware feature fusion module is used to reduce the information loss of edge features during upsampling and improve the edge segmentation accuracy. Constructing an edge-aware feature fusion module includes at least the following steps:
[0126] Assume that the input feature map (X, Y) ∈ R C×H×W , and W1, W2 and W3 represent 3×3 convolution kernels, * represents convolution operation, and σ represents ReLU activation function;
[0127] First, the addition branch is calculated. The addition branch strengthens the preservation of edge details through feature accumulation, enabling the model to more accurately identify the contours of the lesion area. Multi-layer 3x3 convolution is used to extract input features, and skipping is added to the last layer. See the following formula:
[0128] X1=σ(W1*X),Y1=σ(W1*Y)
[0129] X2=σ(W2*X1),Y2=σ(W2*Y1)
[0130] X3=σ(W3*(X2+X1)), Y3=σ(W3*(Y2+Y1))
[0131] The features of all levels are then summed and the dimension is reduced by 1×1 convolution, as shown in the following formula:
[0132] X add =W out *(X1+X2+X3+Y1+Y2+Y3)
[0133] Then the subtraction branch is calculated. The subtraction branch highlights the local significant change area through feature difference calculation, enhances the sensitivity to the morphology of the lesion area, and helps detect structural abnormalities and minor lesions. In addition, the dense connection of the x and y branches allows information to be gradually accumulated, avoiding feature loss and enhancing the feature expression ability of the model. The subtraction branch is similar to the addition branch, using multi-layer 3x3 convolution to extract input features, but the subtraction branch uses subtraction fusion features, see the following formula:
[0134] X diff =W out *|(X1+X2+X3)-(Y1+Y2+Y3)|
[0135] Feature enhancement: The outputs of the addition and subtraction branches are instance normalized respectively, and the nonlinear features are enhanced through the ReLU activation function, where W4 and W5 represent 1×1 convolution kernels, as shown in the following formula:
[0136] Xadd′ =σ(InstamceNorm(W4*X add ))
[0137] X diff′ =σ(InstanceNorm(W4*X diff ))
[0138] Among them, Instance Normalization means normalization;
[0139] The results of the addition branch and the subtraction branch are synthesized to obtain the final output feature map:
[0140] X out =X add′ +X diff′
[0141] The fused feature map is reasonably distributed after dimensionality reduction. One part is used for edge deep supervision to strengthen the constraint of boundary information and improve edge recognition accuracy. The other part is passed to the upper layer to ensure the integrity of the global features. See the following formula:
[0142] y=W5*X out
[0143]
[0144] Among them, Bce is the binary cross entropy, which is used to calculate the boundary loss to ensure that the edge information is strengthened, e is the edge map, The actual edge annotation.
[0145] Overall, the design of the edge-aware feature fusion module strikes a balance between preserving edge information and enhancing lesion area detection, enabling the model to not only accurately identify lesion contours but also effectively capture the morphological changes of local lesions, thereby significantly improving the accuracy and robustness of medical image segmentation. In Table 2 below, the present invention designed a set of comparative experiments to test the performance of the edge-aware feature fusion module, specifically:
[0146] Table 2 Results of edge-aware feature fusion module
[0147]
[0148] In the experimental results in Table 2, the use of the edge-aware feature fusion module improves DSC by 1.98% and reduces HD95 by 3.8% compared to the original model. The effect of adding the edge-aware feature fusion module to DSC is better because the edge-aware feature fusion module enhances the ability to constrain the fuzzy boundaries of the lesion, thereby significantly improving the accuracy of boundary segmentation.
[0149] In Table 3, a set of ablation experiments were designed to evaluate the impact of the size of the convolution kernel on the model's feature extraction. Convolution kernels of different sizes vary greatly in capturing spatial features and texture information. For example, medium-to-large-sized convolution kernels such as 11 have an expanded receptive field and can capture more contextual information, which is beneficial for extracting higher-level semantic feature information of diseased glomeruli and making the segmentation results more structurally consistent. Small-sized convolution kernels such as 3, 5, and 7 are better at capturing fine-grained texture features, fully retaining local spatial information, and facilitating the segmentation of lesion boundaries. However, if the convolution kernel scale is too large (such as kernelsize of 11), too many parameters will be added, leading to model overfitting and reduced generalization ability. The multi-scale approach takes advantage of the advantages of convolution kernels of different scales. Large convolution kernels extract structural information, while small convolution kernels extract texture details. This synergy enhances the performance of the model in segmenting diseased glomeruli.
[0150] Table 3 Effects of different convolution kernel sizes on multi-scale feature extraction
[0151]
[0152]
[0153] Building an improved U-Net network in S3 includes at least the following steps:
[0154] A multi-scale self-attention module is introduced in each jump link of U-Net, and the outputs of these multi-scale self-attention modules are summed layer by layer as the bottleneck feature;
[0155] The feature map generated by each multi-scale self-attention module is not only used as a bottleneck feature, but also passed as input to the edge-aware feature fusion module. The edge-aware feature fusion module combines the model's upsampled feature map with the multi-scale self-attention feature map to extract richer feature information and generate new fused features.
[0156] Through the edge-aware feature fusion module, the generated fusion feature map will be upsampled and gradually restored to the original image resolution, ultimately obtaining a high-quality segmentation output result.
[0157] S4: Input the preprocessed training set into the improved U-Net network, output the training results, and save the trained model;
[0158] S4 includes at least the following steps:
[0159] The training set is divided into multiple batches using the batch training method, where the training validation batch is set to 8;
[0160] Traversing all the images in the training set once is considered an iteration;
[0161] The total loss of the model is the edge loss of each scale plus the segmentation head loss. The loss function can be expressed as:
[0162]
[0163] Where y represents gt, Represents the prediction result, e i Represents gt edge images of different scales, Indicates the predicted glomerular margin.
[0164] The improved U-Net is used as the basic network for training. The basic network includes a backbone network, a multi-scale self-attention mechanism and an edge-aware feature fusion module.
[0165] S5: Use the optimal improved U-Net network model to perform glomerulus segmentation and retain the segmentation result image.
[0166] S5 includes at least the following steps:
[0167] Input the test set to be segmented, segment the glomerular cell image using the trained improved U-Net segmentation model, and obtain the output mask map;
[0168] The root output mask map draws the glomerulus target edge on the original image and saves the final result.
[0169] The effect diagram generated by the detection method of the present invention is as follows Figure 7 shown.
[0170] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A glomerulus segmentation method based on a multi-scale self-attention mechanism, characterized by: At least the following steps are included: S1: Acquire a diseased glomerulus dataset, including but not limited to a training set and a test set. The test set and the training set are used for model construction and training. After the optimal model is generated, actual operation is performed, and a dataset of images of glomerular cells to be segmented is also acquired; S2: Preprocess the glomerular pathology images in the training dataset; S3: Construct a multi-scale self-attention feature extraction module and an edge-aware feature fusion module, and build an improved U-Net network based on the multi-scale self-attention module and the edge-aware feature fusion module; S4: Input the preprocessed training set into the improved U-Net network, output the training results, and save the trained model; S5: Use the optimal improved U-Net network model to perform glomerulus segmentation and retain the segmentation result image.
2. The glomerulus segmentation method based on a multi-scale self-attention mechanism according to claim 1, characterized in that: Said S2 at least comprises the following steps: Unify the size of each image in the dataset so that all images are adjusted to the same preset size; Then the images in the dataset are randomly flipped horizontally, randomly flipped vertically, and rotated at random angles; Then add image noise, adjust color, and add occlusion; Finally, the Canny operator is used to extract multi-scale image edges for edge depth supervision training.
3. The glomerulus segmentation method based on a multi-scale self-attention mechanism according to claim 2, characterized in that: The multi-scale self-attention feature extraction module constructed in S3 combines the ideas of multi-scale feature information and global spatial context modeling. It applies the self-attention mechanism based on the extracted multi-scale lesion features, dynamically allocates feature weights, and improves feature representation. It includes at least the following steps: Assume that the input feature map X∈R C×H×W , where C is the number of channels, H and W are the height and width; First, use 1×1 convolution to perform channel transformation to generate the initial feature map, and use the GeLU activation function to introduce nonlinear processing, see the following formula: X′=σ(W in *X) Where Win is a 1×1 convolution kernel, * represents the convolution operation, and σ represents the GeLU activation function; Then, 5×5 depth convolution is used to extract the input features of the initial multi-scale module, as shown in the following formula: X 5X5 =W 5X5 *X′ Subsequently, three parallel asymmetric convolutions are used to extract multi-scale features. These branches use different sizes of convolution kernels to extract features of different receptive fields, as shown in the following formula: X3=W (1,3 )*X 5×5 ,X3=W (3,1) *X3 X7=W (1,7) *X 5×5 ,X7=W (7,1) *X7 X 11 =W (1,11) *X 5×5 ,X 11 =W (11,1 )*X 11 Therefore, feature fusion is performed by element-by-element addition, and the input residual information is added to retain the original features. The multi-scale feature expression is: X ms =X+X3+X7+X 11 Based on the fused multi-scale features, the query Q, key K, and value V are extracted through 1×1 convolution respectively, as shown in the following formula: Q=W q *X ms ,K=W k *X ms ,V=W v *X ms Among them, Wq, W k and Wv are both 1×1 convolution kernels; Then calculate the self-attention weight as follows: Among them, d k is the dimension of the key, used for scaling to stabilize the gradient; The self-attention output is adjusted through an additional 1×1 convolution to obtain an output feature map that combines local multi-scale information and global context information. The final output is: X out =W out *(X att +X ms ) Among them, W out Represents a 1X1 convolution kernel.
4. The glomerulus segmentation method based on a multi-scale self-attention mechanism according to claim 3, characterized in that: The edge-aware feature fusion module is used to reduce the information loss of edge features during upsampling and improve the edge segmentation accuracy. Constructing the edge-aware feature fusion module includes at least the following steps: Assume that the input feature map (X, Y) ∈ R C×H×W , and W1, W2 and W3 represent 3×3 convolution kernels, * represents convolution operation, and σ represents ReLU activation function; First, the addition branch is calculated. The addition branch strengthens the preservation of edge details through feature accumulation, enabling the model to more accurately identify the contours of the lesion area. Multi-layer 3x3 convolution is used to extract input features, and skipping is added to the last layer. See the following formula: X1=σ(W1*X),Y1=σ(W1*Y) X2=σ(W2*X1),Y2=σ(W2*Y1) X3=σ(W3*(X2+X1)), Y3=σ(W3*(Y2+Y1)) The features of all levels are then summed and the dimension is reduced by 1×1 convolution, as shown in the following formula: X add =W out *(X1+X2+X3+Y1+Y2+Y3) Then the subtraction branch is calculated. The subtraction branch highlights the local significant change area through feature difference calculation, enhances the sensitivity to the morphology of the lesion area, and helps detect structural abnormalities and minor lesions. In addition, the dense connection of the x and y branches allows information to be gradually accumulated, avoiding feature loss and enhancing the feature expression ability of the model. The subtraction branch is similar to the addition branch, using multi-layer 3x3 convolution to extract input features, but the subtraction branch uses subtraction fusion features, see the following formula: X diff =W out *|(X1+X2+X3)-(Y1+Y2+Y3)| Feature enhancement: The outputs of the addition and subtraction branches are instance normalized respectively, and the nonlinear features are enhanced through the ReLU activation function. W4 and W5 represent 1×1 convolution kernels, see the following formula: X add′ =σ(InstanceNorm(W4*X add )) X diff′ =σ(InstanceNorm(W4*X diff )) Among them, Instance Normalization means normalization; The results of the addition branch and the subtraction branch are synthesized to obtain the final output feature map: X out =X add′ +X diff′ The fused feature map is reasonably distributed after dimensionality reduction. One part is used for edge deep supervision to strengthen the constraint of boundary information and improve edge recognition accuracy. The other part is passed to the upper layer to ensure the integrity of the global features. See the following formula: y=W5*X out Among them, Bce is the binary cross entropy, which is used to calculate the boundary loss to ensure that the edge information is strengthened, e is the edge map, The actual edge annotation.
5. The glomerulus segmentation method based on a multi-scale self-attention mechanism according to claim 4, characterized in that: Constructing the improved U-Net network in S3 includes at least the following steps: A multi-scale self-attention module is introduced in each jump link of U-Net, and the outputs of these multi-scale self-attention modules are summed layer by layer as the bottleneck feature; The feature map generated by each multi-scale self-attention module is not only used as a bottleneck feature, but also passed as input to the edge-aware feature fusion module. The edge-aware feature fusion module combines the model's upsampled feature map with the multi-scale self-attention feature map to extract richer feature information and generate new fused features. Through the edge-aware feature fusion module, the generated fusion feature map will be upsampled and gradually restored to the original image resolution, ultimately obtaining a high-quality segmentation output result.
6. The glomerulus segmentation method based on a multi-scale self-attention mechanism according to claim 1, characterized in that: The S4 at least includes the following steps: The training set is divided into multiple batches using the batch training method, where the training validation batch is set to 8; Traversing all the images in the training set once is considered an iteration; The total loss of the model is the edge loss of each scale plus the segmentation head loss. The loss function is expressed as: Where y represents gt, Represents the prediction result, e i Represents gt edge images of different scales, Indicates the predicted glomerular margin. An improved U-Net is used as the basic network for training, and the basic network includes a backbone network, a multi-scale self-attention mechanism and an edge-aware feature fusion module.
7. The glomerulus segmentation method based on a multi-scale self-attention mechanism according to claim 1, characterized in that: The S5 at least includes the following steps: Input the test set to be segmented, segment the glomerular cell image using the trained improved U-Net segmentation model, and obtain the output mask map; The root output mask map draws the glomerulus target edge on the original image and saves the final result.