Medical image segmentation method and system based on multi-scale residual feature aggregation network

The Multi-Scale Residual Feature Aggregation Network (MS-ResDuck) addresses the shortcomings of existing medical feature map segmentation networks in feature fusion and computational efficiency, achieving efficient and accurate medical image segmentation and improving segmentation accuracy and computational efficiency.

CN120852281APending Publication Date: 2025-10-28ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510758966.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-10-28

Smart Images

  • Figure CN120852281A_ABST
    Figure CN120852281A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, and discloses a medical image segmentation method based on a multi-scale residual feature aggregation network, and the method comprises the following steps: constructing a multi-scale residual feature aggregation network (MS-ResDuck) architecture; inputting and importing a medical feature map; and outputting a final segmentation structure. The medical image segmentation system based on the multi-scale residual feature aggregation network is used for executing the segmentation method and comprises an input module, an encoder, a decoder, a multi-stage feature aggregation module, a double-branch collaborative fusion module and an output module. The input module is used for importing a medical feature map. According to the method, the MS-ResD feature block is designed, seven convolution paths are integrated, and the feature expression ability is enhanced; wherein the SBA unit adopts a two-way attention mechanism to adaptively adjust the feature weight so as to realize dynamic gating fusion of high-layer and low-layer features; and the double-branch collaborative fusion module is used for adding and fusing the features output by the multi-stage feature aggregation branch and the features output by the decoding branch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to a medical image segmentation method and system based on a multi-scale residual feature aggregation network. Background Technology

[0002] Deep learning has made significant progress in the field of medical feature map segmentation, especially with network architectures based on U-Net and its variants becoming the mainstream approach. U-Net, through its unique encoder-decoder structure and skip connections, effectively combines low-level detail features with high-level semantic features, demonstrating excellent performance in medical feature map segmentation tasks. However, existing methods still have significant shortcomings in feature fusion mechanisms and computational efficiency.

[0003] Traditional U-Net networks employ a simple feature concatenation method for multi-scale feature fusion, which suffers from two main drawbacks: first, it lacks differentiation in the importance of features at different levels, leading to the introduction of irrelevant background noise into the segmentation results; second, its fixed receptive field design struggles to adapt to the large variations in target size in medical feature maps. While subsequent research has proposed various improvements, such as enhancing feature selection through attention mechanisms, these methods often come at the cost of significantly increased computational complexity, resulting in poor performance in maintaining model lightweightness.

[0004] In existing technologies, the prominent problems faced by medical feature map segmentation networks include: a simplistic and crude feature fusion mechanism that fails to achieve adaptive selection of hierarchical features; a limited receptive field design that makes it difficult to simultaneously capture local details and global contextual information; and high computational complexity, which hinders deployment on resource-constrained medical devices. These problems severely restrict the practical application of deep learning in medical image analysis.

[0005] To address the aforementioned issues, there is an urgent need to develop a novel medical feature map segmentation network that can achieve more efficient feature fusion and lower computational costs while maintaining high segmentation accuracy. An ideal solution should possess the following characteristics: an intelligent multi-level feature selection mechanism, an adaptively adjustable receptive field design, and a lightweight network structure. Such innovations will significantly improve the accuracy and practicality of medical feature map segmentation. Summary of the Invention

[0006] To address the technical problems mentioned in the background section, this invention provides a medical image segmentation method and system based on a multi-scale residual feature aggregation network.

[0007] This invention employs the following technical solution: a medical image segmentation method based on a multi-scale residual feature aggregation network, comprising the following steps:

[0008] Construct a multi-scale residual feature aggregation network (MS-ResDuck) architecture;

[0009] Import medical feature maps;

[0010] Output the final segmentation structure.

[0011] This invention proposes a medical image segmentation system based on a multi-scale residual feature aggregation network, which performs the aforementioned segmentation method, including...

[0012] Input module, encoder, decoder, multi-level feature aggregation module, dual-branch collaborative fusion module, output module;

[0013] The input module is used to import medical feature maps;

[0014] The encoder uses multi-scale feature extraction blocks (MS-ResD) to extract multi-level features from the input medical feature map;

[0015] The decoder part gradually restores the feature map size through bilinear interpolation upsampling operation, and combines the skip connection mechanism to perform channel splicing of the feature maps of each stage of the encoder with the feature maps of the corresponding level of the decoder, so as to achieve effective fusion of shallow spatial information and deep semantic information.

[0016] The multi-level feature aggregation module fuses feature maps from different stages and embeds a symmetric bidirectional attention module (SBA) to enhance feature representation capabilities.

[0017] The dual-branch collaborative fusion module is used to add and fuse the features output by the multi-level feature aggregation branch with the features output by the decoding branch;

[0018] The output module processes the output structure of the dual-branch collaborative fusion module by embedding multi-scale feature extraction blocks (MS-ResD) convolution, and finally outputs the segmentation result.

[0019] Furthermore, the multi-scale feature extraction (MS-ResD) module includes parallel multi-level atrous blocks, midscope blocks, separated blocks, and residual blocks for extracting features from different receptive fields.

[0020] The Multi-Scale Feature Extraction (MS-ResD) module achieves multi-scale feature capture by combining convolutional kernels of different scales;

[0021] During feature extraction, the encoder downsamples the feature map using a convolutional layer with a stride of 2, gradually compressing the spatial dimension while increasing the channel dimension; and fuses features through skip links.

[0022] Furthermore, both the encoder and decoder include five-layer paths: the encoding path includes a downsampling layer and a multi-scale feature extraction block (MS-ResD), and the decoding path includes an upsampling layer and a multi-scale feature extraction block (MS-ResD), and the features of the corresponding layers are fused through skip connections.

[0023] Furthermore, the multi-level feature aggregation module includes high-level and low-level feature processing paths and an SBA fusion unit, which dynamically fuses multi-scale features through a gating mechanism.

[0024] Furthermore, the low-level feature processing path in the multi-level feature aggregation module is used to perform 3×3 convolution processing on the mid-scale features output by the fourth layer of the encoder; the high-level feature processing path in the multi-level feature aggregation module is used to perform cross-scale concatenation of the features of the fifth layer of the encoder with the features of the fourth layer.

[0025] Furthermore, the dual-branch collaborative fusion module includes a dual-branch structure for parallel processing, wherein:

[0026] Decoding branch: Inherits the U-shaped structure of the encoder-decoder module, gradually restores spatial resolution through five levels of upsampling, and finally outputs decoded features and initial convolutional features;

[0027] The multi-level feature aggregation branch consists of two paths: a low-level feature processing path that performs 3×3 convolution on the mid-scale features output from the fourth layer of the encoder; a high-level feature processing path that concatenates the features from the fifth layer of the encoder with those from the fourth layer across scales; and a symmetric bidirectional attention module (SBA) that fuses the low-level and high-level features through a gating mechanism. This module uses an adaptive weight allocation mechanism to enable bidirectional interaction between low-resolution and high-resolution feature maps, significantly enhancing the network's ability to represent features in lesion regions.

[0028] Furthermore, the operations of the Symmetric Bidirectional Attention Module (SBA) include feature preprocessing, attention weight generation, cross-resolution feature interaction, feature integration and output, specifically;

[0029] The feature preprocessing part performs 1×1 convolution operations on the input low-resolution and high-resolution features respectively, adjusts the number of channels to a uniform dimension, and removes bias terms to reduce the number of parameters. The attention weight generation part performs Sigmoid activation on the preprocessed low-resolution and high-resolution features respectively to generate a dynamic attention weight map, which is used to control the fusion ratio of cross-resolution features. The cross-resolution feature interaction part upsamples the high-resolution features to the size of the low-resolution features through transposed convolution and performs weighted fusion with the low-resolution features. It downsamples the low-resolution features to the size of the high-resolution features through average pooling and performs weighted fusion with the high-resolution features. During fusion, the contribution of the feature itself and the low-resolution supplementary features is dynamically balanced by attention weights. The feature integration and output part concatenates the enhanced low-resolution features and high-resolution features after spatial alignment and further fuses multi-scale information through 3×3 convolution. Finally, the number of channels is compressed through 1×1 convolution to generate the final output feature map.

[0030] The present invention also proposes a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the medical image segmentation method based on a multi-scale residual feature aggregation network as described in claim 1.

[0031] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the medical image segmentation method based on a multi-scale residual feature aggregation network as described in claim 1.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] While maintaining the characteristics of the U-Net network, this invention significantly improves the performance of medical feature map segmentation through the following innovations.

[0034] 1. Design a multi-scale feature extraction block (MS-ResD) that integrates seven convolutional paths to enhance feature representation capabilities;

[0035] 2. Design a feature aggregation module, in which the Symmetric Bidirectional Attention (SBA) module adopts a bidirectional attention mechanism to adaptively adjust feature weights, thereby achieving dynamic gating fusion of high- and low-level features;

[0036] 3. Design a dual-branch collaborative fusion module to add and fuse the features output by the multi-level feature aggregation branch with the features output by the decoding branch;

[0037] Experiments show that the Dice coefficient of this invention reaches 0.8996 on the Kvasir public dataset. Attached Figure Description

[0038] Figure 1 The flowchart of the Multi-Scale Residual Feature Aggregation Network (MS-ResDuck) method is shown below.

[0039] Figure 2 This is an architecture diagram of the Multi-Scale Feature Extraction Block (MS-ResD).

[0040] Figure 3 Component detail diagram for Multi-Scale Feature Extraction Block (MS-ResD);

[0041] Figure 4 The flowchart for the Symmetric Bidirectional Attention Module (SBA) is shown below.

[0042] Figure 5 This is a visual comparison diagram of the relevant networks on the Kvasir-SEG dataset. Detailed Implementation

[0043] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0044] Example 1:

[0045] Reference Figure 1-Figure 5 This embodiment proposes a medical image segmentation system based on a multi-scale residual feature aggregation network. The specific implementation of the present invention will be described in detail below.

[0046] The system includes an input module, an encoder, a decoder, a multi-level feature aggregation module, a dual-branch collaborative fusion module, and an output module.

[0047] The input module works as follows: First, it loads a 352×352 resolution medical feature map dataset, divides the dataset into training, validation, and test sets, and fixes the random seed to ensure experimental reproducibility; it implements dynamic data augmentation strategies, including spatial transformation and color adjustment; the preprocessed data is input into the MS-ResDuck model in float32 format.

[0048] The encoder consists of five layers. Except for the fifth layer, each of the remaining layers contains two paths: a main path and a multi-scale feature extraction path. The main path uses a 2×2 convolutional kernel with a stride of 2 for downsampling. The multi-scale feature extraction path first adds and fuses the input of this layer with the feature map processed by the upper layer, and then performs multi-scale feature extraction on the feature map.

[0049] The input layer takes an RGB feature map with a resolution of 352×352 as input. The main path convolves the feature map with a 2*2 convolution kernel with a stride of 2, and outputs 34 channels with equal-length padding. The multi-scale feature extraction path extracts features from the feature map using a multi-scale feature extraction block (MS-ResD), and outputs a feature map with a resolution of 352*352 and 17 channels.

[0050] The first layer inputs a feature map with a resolution of 176*176 and 34 channels. The main path convolves the feature map with a 2*2 convolution kernel with a stride of 2, resulting in an output feature map with a resolution of 88*88 and 64 channels, padded with equal length. The multi-scale feature extraction path adds and merges the first layer input feature map with the feature map processed by the input layer, resulting in an output feature map with a resolution of 176*176 and 34 channels. Then, multi-scale feature extraction (MS-ResD) is performed on the feature map, resulting in an output feature map with a resolution of 176*176 and 34 channels.

[0051] The second layer takes a feature map with a resolution of 88*88 and 68 channels as input. The main path convolves the feature map with a 2*2 convolution kernel with a stride of 2, resulting in an output feature map with a resolution of 44*44 and 136 channels, padded with equal length. The multi-scale feature extraction path adds and merges the second-layer input feature map with the feature map processed by the first layer, resulting in an output feature map with a resolution of 88*88 and 68 channels. Then, multi-scale feature extraction (MS-ResD) is performed on the feature map, resulting in an output feature map with a resolution of 88*88 and 68 channels.

[0052] The third layer takes a feature map with a resolution of 44*44 and 136 channels as input. The main path convolves the feature map with a 2*2 convolution kernel with a stride of 2, resulting in an output feature map with a resolution of 22*22 and 272 channels, padded with equal length. The multi-scale feature extraction path adds and merges the input feature map from the third layer with the feature map processed by the second layer, resulting in an output feature map with a resolution of 44*44 and 136 channels. Then, multi-scale feature extraction (MS-ResD) is performed on the feature map, resulting in an output feature map with a resolution of 44*44 and 136 channels.

[0053] The fourth layer takes a feature map with a resolution of 22*22 and 272 channels as input. The main path convolves the feature map with a 2*2 convolution kernel with a stride of 2, resulting in an output feature map with a resolution of 11*11 and 544 channels, padded with equal length. The multi-scale feature extraction path adds and merges the fourth layer's input feature map with the feature map processed by the third layer, resulting in an output feature map with a resolution of 22*22 and 272 channels. Then, multi-scale feature extraction (MS-ResD) is performed on the feature map, resulting in an output feature map with a resolution of 22*22 and 272 channels.

[0054] The fifth layer takes a feature map with a resolution of 11*11 and 544 channels as input. The feature map input from the fifth layer is then fused with the feature map processed by the fourth layer, resulting in an output feature map with a resolution of 11*11 and 544 channels. Then, a residual convolution is performed on the feature map, resulting in an output feature map with a resolution of 11*11 and 544 channels. Finally, a second residual convolution is performed, resulting in an output feature map with a resolution of 11*11 and 272 channels.

[0055] Furthermore, the multi-scale feature extraction path is performed by a multi-scale feature extraction block (MS-ResD). The multi-scale feature extraction block (MS-ResD) contains seven parallel paths. The output feature maps of the seven branches are fused by adding them element by element. The fused result is then processed by batch normalization to form the final output of the multi-scale feature extraction block.

[0056] Furthermore, the multi-scale feature extraction block (MS-ResD) feature block includes:

[0057] Multi-level Atrous Block: Employs a cascaded three-stage dilated convolution path. The first stage uses a 3x3 kernel with a dilation rate of 1, followed by ReLU activation and batch normalization. The second stage uses a 3x3 kernel with a dilation rate of 2, again followed by ReLU activation and batch normalization. The third stage uses a 3x3 kernel with a dilation rate of 2, again followed by ReLU activation and batch normalization. All edge padding maintains consistent input and output spatial dimensions. All convolutional layer weights are initialized using a He normal distribution.

[0058] Midscope Block: Employs a cascaded two-stage dilated convolution. The first stage uses a 3x3 kernel with a dilation rate of 1, followed by ReLU activation and batch normalization. The second stage uses a 3x3 kernel with a dilation rate of 2, followed by ReLU activation and batch normalization. All edge padding maintains consistent input and output spatial dimensions. All convolutional layer weights are initialized using a He normal distribution.

[0059] Separated Block: A cascaded two-dimensional separated convolutional structure is adopted. The first stage uses a 1×3 convolutional kernel for horizontal feature extraction, followed by ReLU activation and batch normalization. The second stage uses a 3×1 convolutional kernel for vertical feature extraction, followed by ReLU activation and batch normalization. Edge padding maintains consistent input and output spatial dimensions. All convolutional layer weights are initialized using the He normal distribution method.

[0060] Residual Convolution Path: There are three residual convolution paths. The first path uses one residual convolution block, the second path uses two cascaded residual convolution blocks, and the third path uses three cascaded residual convolution blocks. Each residual convolution block includes two parallel branches. Path one uses a cascaded two-stage convolution. The first stage uses a 3x3 convolution kernel, then uses a ReLU function for activation and batch normalization. The second stage uses a 3x3 convolution kernel, then uses a ReLU function for activation and batch normalization. Path two uses a 1x1 convolution kernel, then uses a ReLU function for activation.

[0061] The decoder consists of five layers. Except for the fifth layer, each of the remaining layers contains two paths: a multi-scale feature fusion path and a backbone path. The multi-scale feature fusion path first adds and fuses the feature map with the encoder feature map through skip connections, and then performs multi-scale feature extraction (MS-ResD) on the feature map. The backbone path uses bilinear interpolation for upsampling.

[0062] The feature map of the fifth layer with a resolution of 11*11 and 272 channels is upsampled using bilinear interpolation, and the output feature map has a resolution of 22*22 and 272 channels.

[0063] The fourth-layer multi-scale feature fusion path first performs a skip connection between the upsampled feature map from the fifth layer and the feature map from the fourth-layer encoder, outputting a feature map with a resolution of 22*22 and 272 channels. Then, it performs multi-scale feature extraction (MS-ResD) on the feature map, outputting a feature map with a resolution of 22*22 and 136 channels. The fourth-layer backbone path upsamples the 22*22, 136-channel feature map from the fourth layer using bilinear interpolation, outputting a feature map with a resolution of 44*44 and 136 channels.

[0064] The third-layer multi-scale feature fusion path first performs a skip connection between the upsampled feature map from the fourth layer and the feature map from the third layer encoder, outputting a feature map with a resolution of 44*44 and 136 channels. Then, multi-scale feature extraction (MS-ResD) is performed on the feature map, outputting a feature map with a resolution of 44*44 and 68 channels. The third-layer backbone path upsamples the 44*44 and 68-channel feature map from the third layer using bilinear interpolation, outputting a feature map with a resolution of 88*88 and 68 channels.

[0065] The second-layer multi-scale feature fusion path first performs a skip connection between the upsampled feature map from the third layer and the feature map from the second-layer encoder, outputting a feature map with a resolution of 88*88 and 68 channels. Then, it performs multi-scale feature extraction (MS-ResD) on the feature map, outputting a feature map with a resolution of 88*88 and 34 channels. The second-layer backbone path upsamples the 88*88, 34-channel feature map from the second layer using bilinear interpolation, outputting a feature map with a resolution of 176*176 and 34 channels.

[0066] The first-layer multi-scale feature fusion path first performs a skip connection between the upsampled feature map from the second layer and the feature map from the first layer encoder, outputting a feature map with a resolution of 176*176 and 34 channels. Then, multi-scale feature extraction (MS-ResD) is performed on the feature map, outputting a feature map with a resolution of 176*176 and 17 channels. The first-layer backbone path upsamples the first-layer feature map with a resolution of 176*176 and 17 channels using bilinear interpolation, outputting a feature map with a resolution of 352*352 and 17 channels.

[0067] Furthermore, the multi-scale feature fusion path is performed by the multi-scale feature extraction block (MS-ResD). The multi-scale feature extraction block (MS-ResD) contains seven parallel paths. The output feature maps of the seven branches are fused by adding them element by element. The fusion result is then processed by batch normalization to form the final output of the multi-scale feature fusion block.

[0068] Furthermore, the multi-scale feature extraction block (MS-ResD) feature block includes:

[0069] Multi-level Atrous Block: Employs a cascaded three-stage dilated convolution path. The first stage uses a 3x3 kernel with a dilation rate of 1, followed by ReLU activation and batch normalization. The second stage uses a 3x3 kernel with a dilation rate of 2, again followed by ReLU activation and batch normalization. The third stage uses a 3x3 kernel with a dilation rate of 2, again followed by ReLU activation and batch normalization. All edge padding maintains consistent input and output spatial dimensions. All convolutional layer weights are initialized using a He normal distribution.

[0070] Midscope Block: Employs a cascaded two-stage dilated convolution. The first stage uses a 3x3 kernel with a dilation rate of 1, followed by ReLU activation and batch normalization. The second stage uses a 3x3 kernel with a dilation rate of 2, followed by ReLU activation and batch normalization. All edge padding maintains consistent input and output spatial dimensions. All convolutional layer weights are initialized using a He normal distribution.

[0071] Separated Block: A cascaded two-dimensional separated convolutional structure is adopted. The first stage uses a 1×3 convolutional kernel for horizontal feature extraction, followed by ReLU activation and batch normalization. The second stage uses a 3×1 convolutional kernel for vertical feature extraction, followed by ReLU activation and batch normalization. Edge padding maintains consistent input and output spatial dimensions. All convolutional layer weights are initialized using the He normal distribution method.

[0072] Residual Convolution Path: There are three residual convolution paths. The first path uses one residual convolution block, the second path uses two cascaded residual convolution blocks, and the third path uses three cascaded residual convolution blocks. Each residual convolution block includes two parallel branches. Path one uses a cascaded two-stage convolution. The first stage uses a 3x3 convolution kernel, then uses the ReLU function for activation and batch normalization. The second stage uses a 3x3 convolution kernel, then uses the ReLU function for activation and batch normalization. Path two uses a 1x1 convolution kernel, then uses the ReLU function for activation.

[0073] The feature aggregation module is divided into two paths, with path one having two branches and path two having two branches.

[0074] Branch 1 performs a 4x upsampling on the feature map of the fifth layer with a resolution of 11*11 and 272 channels, outputting a feature map with a resolution of 44*44 and 272 channels. Branch 2 performs a 2x upsampling on the feature map of the fourth layer encoder with a resolution of 22*22 and 272 channels, outputting a feature map with a resolution of 44*44 and 272 channels. The output feature maps of the two branches are then concatenated with the feature map of the third layer encoder with a resolution of 44*44 and 136 channels. The feature map is then sequentially processed by a 1x1 convolution, batch normalization, ReLU activation, another 1x1 convolution, and an 8x upsampling operation, outputting a feature map with a resolution of 44*44.

[0075] Path 2 has two branches: Branch 1 is a low-resolution feature path, and Branch 2 is a high-resolution feature path. The low-resolution feature path upsamples the 11*11 feature map with 272 channels in the fifth layer by a factor of 2, outputting a 22*22 feature map with 272 channels. This output feature map is then concatenated with the 22*22 feature map in the fourth layer encoder by a channel dimension concatenation operation. The feature map is then subjected to a 1*1 convolution, batch normalization, and ReLU activation. The high-resolution feature path uses the 88*88 feature map in the second layer encoder with 68 channels. The outputs of both the low-resolution and high-resolution feature paths are fed into a Symmetric Bidirectional Attention (SBA) module. The SBA-processed feature map is upsampled by a factor of 4, and the feature maps from Path 1 and Path 2 are then fused together. Finally, the output feature map from the decoder path and the output feature map from the feature aggregation path are fused together, resulting in a final feature map after multi-scale feature fusion.

[0076] Furthermore, the Symmetric Bidirectional Attention Module (SBA) includes a dual-path feature interaction structure, processing low-resolution and high-resolution inputs respectively. Each path performs the following operations sequentially: first, a 1×1 convolutional kernel is used to adjust the channel dimension; then, an attention weight map is generated using a sigmoid activation function. The low-resolution path uses transposed convolution for a 2x upsampling, while the high-resolution path uses average pooling for a 2x downsampling. Finally, a weighted fusion mechanism is used to achieve bidirectional feature enhancement. After upsampling and alignment, the fused dual-path features are concatenated along the channel dimension and then processed through two levels of feature processing: the first level uses a 3×3 convolutional kernel for feature extraction, combined with ReLU activation and batch normalization; the second level uses a 1×1 convolutional kernel for channel compression, finally outputting the fused result. All convolutional operations maintain the spatial dimension and use a bias-free design.

[0077] The loss function used is based on the Dice coefficient, the optimizer is RMSprop, and the initial learning rate is set to a specific value. The training process employs mini-batch gradient descent with a batch size of 4 and 600 training epochs.

[0078] The output prediction module contains a 1×1 convolutional layer and a sigmoid activation function, which converts the feature map into a probability map with the same spatial size as the input feature map. Each pixel value represents the probability of belonging to the target region, and the value ranges from 0 to 1.

[0079] During the model evaluation phase, a multi-indicator comprehensive evaluation system is adopted, including:

[0080] Dice coefficient: measures the degree of overlap between the predicted result and the actual annotation.

[0081] Intersection over Union (IoU): Calculates the ratio of the intersection to the union of the predicted and actual regions.

[0082] Precision: The proportion of samples that were predicted to be positive but were actually positive.

[0083] Recall: The proportion of samples that are actually positive that are correctly predicted.

[0084] Accuracy: The proportion of all samples that are correctly predicted.

[0085] Typical applications of this invention include, but are not limited to:

[0086] Medical feature map segmentation, remote sensing feature map segmentation.

[0087] The table below illustrates the data comparison results of the relevant networks on the Kvasir-SEG dataset.

[0088]

[0089] According to the data above, the network proposed in this invention achieved a similarity coefficient (DSC) of 0.8996, a Jaccard coefficient of 0.8175, a precision of 0.8949, a recall of 0.9044, and an accuracy of 0.9695 on the Kvasir-SEG dataset. Compared with other networks in the table, its overall performance is superior.

[0091] The multi-scale residual feature aggregation network (MS-ResDuck) proposed in this invention has the following technical advantages: it designs MS-ResD feature blocks, integrates seven convolutional paths, and enhances feature representation capabilities; it designs a feature aggregation module, in which the SBA unit adopts a bidirectional attention mechanism to adaptively adjust feature weights, realizing dynamic gating fusion of high and low layer features; and it designs a dual-branch collaborative fusion module, which adds and fuses the features output by the multi-level feature aggregation branch with the features output by the decoding branch.

[0092] Please refer to Figure 5 The comparison images show that the MS-ResDuck model proposed in this invention performs best among all models. Its segmentation results are very close to the correctly labeled images, with clear edges and good detail preservation. The HardNet-DFUS model has relatively good segmentation results, but compared with the MS-ResDuck model, its edge detail processing is slightly insufficient, and some regions exhibit oversegmentation. The HRNetV2 model's segmentation results are average, with insufficient fine edge detail processing and some undersegmentation, resulting in some regions not being fully covered. The MSRF-Net model has good segmentation results and relatively fine edge detail processing, but compared with the MS-ResDuck model, some details are still not fully captured. The U-Net model has relatively good segmentation results, but compared with the MS-ResDuck model, its edge detail processing is slightly insufficient, and some regions exhibit oversegmentation.

[0093] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A medical image segmentation method based on a multi-scale residual feature aggregation network, characterized in that, The steps include: Construct a multi-scale residual feature aggregation network architecture; Import medical feature maps; Output the final segmentation structure.

2. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, wherein the system performs the method described in claim 1, characterized in that, include Input module, encoder, decoder, multi-level feature aggregation module, dual-branch collaborative fusion module, output module; The input module is used to import medical feature maps; The encoder uses multi-scale feature extraction blocks to perform multi-level feature extraction on the input medical feature map; The decoder part gradually restores the feature map size through bilinear interpolation upsampling operation, and combines the skip connection mechanism to perform channel splicing of the feature maps of each stage of the encoder with the feature maps of the corresponding level of the decoder, so as to achieve effective fusion of shallow spatial information and deep semantic information. The multi-level feature aggregation module fuses feature maps from different stages and embeds a symmetrical bidirectional attention module to enhance feature representation capabilities; The dual-branch collaborative fusion module is used to add and fuse the features output by the multi-level feature aggregation branch with the features output by the decoding branch; The output module incorporates a multi-scale feature extraction block convolutional processing module to process the output structure of the dual-branch collaborative fusion module, ultimately outputting the segmentation result.

3. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, characterized in that, The multi-scale feature extraction module includes parallel multi-level dilated convolutional blocks, medium-dilated convolutional blocks, segregating convolutional blocks, and residual convolutions, used to extract features from different receptive fields. The multi-scale feature extraction module achieves multi-scale feature capture by combining convolutional kernels of different scales; During feature extraction, the encoder downsamples the feature map using a convolutional layer with a stride of 2, gradually compressing the spatial dimension while increasing the channel dimension; and fuses features through skip links.

4. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, characterized in that, Both the encoder and decoder consist of five layers: the encoding path includes a downsampling layer and a multi-scale feature extraction block, and the decoding path includes an upsampling layer and a multi-scale feature extraction block. Features from the corresponding layers are fused through skip connections.

5. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, characterized in that, The multi-level feature aggregation module includes high-level and low-level feature processing paths and an SBA fusion unit, which dynamically fuses multi-scale features through a gating mechanism.

6. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, characterized in that, The low-level feature processing path in the multi-level feature aggregation module is used to perform 3×3 convolution processing on the mid-scale features output by the fourth layer of the encoder; the high-level feature processing path in the multi-level feature aggregation module is used to perform cross-scale concatenation of the features of the fifth layer of the encoder with the features of the fourth layer.

7. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, characterized in that, The dual-branch collaborative fusion module includes a parallel processing dual-branch structure, wherein: Decoding branch: Inherits the U-shaped structure of the encoder-decoder module, gradually restores spatial resolution through five levels of upsampling, and finally outputs decoded features and initial convolutional features; Multi-level feature aggregation branch: Low-level feature processing path, which performs 3×3 convolution processing on the mid-scale features output by the fourth layer of the encoder; High-level feature processing path, which performs cross-scale concatenation of the features from the fifth layer of the encoder with the features from the fourth layer; Symmetrical bidirectional attention module, which fuses the low-level and high-level features through a gating mechanism.

8. The medical image segmentation system based on a multi-scale residual feature aggregation network as described in claim 1, characterized in that, The operations of the symmetric bidirectional attention module include feature preprocessing, attention weight generation, cross-resolution feature interaction, feature integration and output, specifically; The feature preprocessing part performs 1×1 convolution operations on the input low-resolution and high-resolution features respectively, adjusts the number of channels to a uniform dimension, and removes the bias term to reduce the number of parameters; the attention weight generation part performs Sigmoid activation on the preprocessed low-resolution and high-resolution features respectively to generate a dynamic attention weight map, which is used to control the fusion ratio of cross-resolution features. The cross-resolution feature interaction part upsamples high-resolution features to the size of low-resolution features through transposed convolution and then performs weighted fusion with low-resolution features. The low-resolution features are downsampled to the size of high-resolution features through average pooling and then performed weighted fusion with high-resolution features. During fusion, the contribution of the feature itself and the low-resolution supplementary features is dynamically balanced through attention weights. The feature integration and output part concatenates the enhanced low-resolution features and high-resolution features after spatial alignment and further integrates multi-scale information through 3×3 convolution. Finally, the number of channels is compressed through 1×1 convolution to generate the final output feature map.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the medical image segmentation method based on a multi-scale residual feature aggregation network as described in claim 1.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the medical image segmentation method based on a multi-scale residual feature aggregation network as described in claim 1.

Citation Information

Cited By

  • Medical image segmentation method based on multi-scale window alignment iterative fusion

    CN121504958A