Medical Image Segmentation Method and Device, Electronic Device, Readable Storage Medium
The dual-branch feature extraction and fusion method addresses CNN limitations in medical image segmentation by leveraging ConvNeXt-T and VMamba architectures to enhance feature representation and decoding, improving segmentation performance across multiple datasets.
Patent Information
- Application Number
- CN202411726294.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The existing CNN-based medical image segmentation method has limitations in capturing long-range dependencies in images, resulting in poor segmentation accuracy and high computational complexity in complex tasks.
Multi-scale local and global feature extraction is adopted, and the dual-branch feature extraction module and fusion module are combined with ConvNeXt and VMamba architectures to integrate local and global features. The channel and spatial attention mechanisms are used to perform feature fusion and decoding to generate medical image segmentation results.
It improves the accuracy and robustness of medical image segmentation, enhances the expression ability of the model, can better capture local details and global context information, and improves segmentation performance.
Smart Images

Figure CN119559197B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of image processing, and more particularly, relates to a medical image segmentation method, apparatus, electronic device, and readable storage medium. Background Art
[0002] Medical image segmentation is one of the core tasks of computer-aided diagnosis systems (CADs), and its accuracy directly affects the diagnosis, treatment, and prognosis of patients. Although clinical experts can provide high-quality segmentation results by manually annotating lesions or targets, this method is time-consuming, labor-intensive, and highly dependent on the knowledge and experience of experts, making it prone to subjective errors. Therefore, automated and accurate medical image segmentation technology is crucial for improving diagnostic efficiency, reducing the workload of doctors, and achieving standardized and objective medical image analysis. Medical image segmentation is a key task in computer-aided diagnosis and is of great significance for the early diagnosis of diseases and the formulation of treatment plans.
[0003] With the rapid development of computer hardware and deep learning technology, automatic paired medical image segmentation has shown great promise. In particular, convolutional neural networks (CNNs) have demonstrated excellent performance in medical image segmentation tasks, significantly improving the segmentation accuracy. The UNet architecture and its U-shaped derivatives have gradually become the standard paradigm for medical image segmentation, achieving state-of-the-art performance in various tasks. Inspired by U-Net, subsequent models improve the segmentation accuracy by designing unique feature enhancement modules, combining attention mechanisms, and modifying the U-Net structure. However, due to the locality of convolutional operations, CNNs have limitations in modeling global information, resulting in poor performance in some complex tasks. Traditional CNN-based segmentation methods mainly rely on local feature extraction and are difficult to capture long-range dependencies in images. Although the receptive field can be enlarged by stacking more convolutional layers, it also brings problems such as increased computational complexity and optimization difficulties. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a medical image segmentation method, apparatus, electronic device, and readable storage medium to improve the segmentation accuracy.
[0005] In a first aspect of the embodiments of the present disclosure, a medical image segmentation method is provided, including:
[0006] Performing multi-scale local feature extraction and global feature extraction on a medical image to obtain a first local feature map and a first global feature map corresponding to each scale;
[0007] Fuse the first local feature maps and the first global feature maps corresponding to each scale to obtain first fusion feature maps corresponding to each scale; fuse the first fusion feature maps corresponding to each scale to obtain a second fusion feature map;
[0008] Decode the second fusion feature map to obtain an image segmentation result.
[0009] In a second aspect of the embodiments of the present disclosure, there is provided a medical image segmentation device, including:
[0010] A dual-branch feature extraction module, configured to perform multi-scale local feature extraction on a medical image to obtain first local feature maps corresponding to each scale;
[0011] A dual-branch fusion module, configured to perform multi-scale global feature extraction on a medical image to obtain first global feature maps corresponding to each scale; fuse the first global feature maps and the first local feature maps corresponding to each scale to obtain first fusion feature maps corresponding to each scale;
[0012] A decoder, configured to decode the first fusion feature map to obtain an image segmentation result.
[0013] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, where when the processor executes the computer program, the steps of the above-mentioned medical image segmentation method are implemented.
[0014] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the steps of the above-mentioned medical image segmentation method are implemented.
[0015] The beneficial effects of the medical image segmentation method, device, electronic device, and readable storage medium provided by the embodiments of the present disclosure are as follows:
[0016] In the embodiments of the present disclosure, through multi-scale local feature extraction of a medical image, the obtained first local feature maps can capture the detailed information of local regions in the medical image; through multi-scale global feature extraction of the medical image, the obtained first global feature maps can capture more extensive context information. Fusing the first local feature maps and the first global feature maps corresponding to each scale, the obtained first fusion feature maps can comprehensively reflect local and global features, and fusing the first fusion feature maps corresponding to each scale enables full utilization of features from low-level to high-level. In summary, through the extraction and fusion of multi-scale features, the expression ability and segmentation performance of the model are enhanced. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of a medical image segmentation method provided by an embodiment of the present disclosure;
[0019] Figure 2 It is a principle block diagram of a medical image segmentation device provided by an embodiment of the present disclosure;
[0020] Figure 3 It is a principle block diagram of a dual-branch fusion module provided by an embodiment of the present disclosure;
[0021] Figure 4 It is a principle block diagram of a decoder provided by an embodiment of the present disclosure;
[0022] Figure 5 It is a summary diagram of experimental results provided by an embodiment of the present disclosure;
[0023] Figure 6 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0024] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0025] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the drawings.
[0026] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a medical image segmentation method provided by an embodiment of the present disclosure. This method is applied to a medical image segmentation device, and the method includes:
[0027] S101: Perform multi-scale local feature extraction and global feature extraction on the medical image to obtain a first local feature map and a first global feature map corresponding to each scale.
[0028] Please refer to Figure 2, in this embodiment, the dual-branch feature extraction module in the medical image segmentation device includes a spatial branch and a state space branch. The spatial branch is used to extract multi-scale local features from the medical image. The spatial branch uses ConvNeXt-T as its main feature extraction network. ConvNeXt combines the advantages of the Transformer and CNN architectures, maintaining both the efficiency and simplicity of traditional CNNs and integrating the powerful feature representation ability and flexibility of the Transformer. Compared with VGG, ResNet, DenseNet, etc., ConvNeXt uses a larger convolutional kernel (7×7) and introduces a concept similar to the attention mechanism, aiming to balance the global and local features of image processing. This method enables the network to achieve stronger feature representation ability when processing image features of different scales.
[0029] In this embodiment, the spatial branch includes four ConvNeXt Blocks, where:
[0030] The first ConvNeXt Block: processes features of size h×768;
[0031] The second ConvNeXt Block: processes features of size h×384;
[0032] The third ConvNeXt Block: processes features of size h×192;
[0033] The fourth ConvNeXt Block: processes features of size h×96;
[0034] There is a downsampling operation between each Block.
[0035] Process the input image Image ∈ Rb×3×h×w, and generate the first local feature maps of different scales through four stages:
[0036] ∈ Rb×96×h / 4×w / 4;
[0037] ∈ Rb×192×h / 8×w / 8;
[0038] ∈ Rb×384×h / 16×w / 16;
[0039] ∈ Rb×768×h / 32×w / 32.
[0040] In this embodiment, state space branches are used to extract multi-scale global features from medical images. The state space branches integrate the core function SS2D in VMamba and are responsible for capturing a wider range of context information.
[0041] Specifically, the state space branches include four VSSBlocks. The input of the VSS Block passes through an initial linear embedding layer, and the output is divided into two information streams. One stream passes through a 3×3 depth convolution layer and then enters the core SS2D module through the Silu activation function. The output of SS2D passes through a normalization layer and is then added to the output of the other information stream, which has passed through the Silu activation. Due to the causal nature of VMamba, positional embedding bias is not used. Among them:
[0042] The first VSS Block: processes features of size h×768;
[0043] The second VSS Block: processes features of size h×384;
[0044] The third VSS Block: processes features of size h×192;
[0045] The fourth VSS Block: processes features of size h×96;
[0046] There is a downsampling operation between each Block.
[0047] Similarly, the input image is processed in four stages to generate the first global feature maps of different scales:
[0048] ∈ Rb×96×h / 4×w / 4;
[0049] ∈ Rb×192×h / 8×w / 8;
[0050] ∈ Rb×384×h / 16×w / 16;
[0051] ∈ Rb×768×h / 32×w / 32;
[0052] Among them, the processing process of each stage can be expressed as:
[0053] patches = PatchPartition(images);
[0054] = VSSBlock(patches);
[0055] = VSSBlock(DownSample( ));
[0056] = VSSBlock(DownSample( ));
[0057] = VSSBlock(DownSample( ));
[0058] S102: Feature - fuse the first local feature maps and the first global feature maps corresponding to each scale to obtain the first fused feature maps corresponding to each scale; fuse the first fused feature maps corresponding to each scale to obtain the second fused feature map.
[0059] In this embodiment, local feature extraction helps to capture the detailed information of local regions in medical images. For example, in CT images, it may be the specific texture of an organ, minute lesions, etc. Global feature extraction, on the other hand, can grasp the overall features of the image, such as the general shape of the whole body part and the relative positional relationship between different organs. By feature - fusing the first local feature maps and the first global feature maps corresponding to each scale, the information in medical images can be more comprehensively mined.
[0060] In this embodiment, the feature maps of each scale can represent the image features at different resolutions, and the features at different scales have different effects on medical image segmentation. For example, features at high - resolution scales may be more conducive to detecting minute diseased tissues, while features at low - resolution scales can quickly locate the general position of organs and provide a framework for subsequent fine - grained segmentation.
[0061] By fusing the first fused feature maps corresponding to each scale again to obtain the second fused feature map, cross - scale feature integration is achieved, and the dominant features at different scales are fused. For example, the fine diseased detail features at high - resolution scales are fused with the overall organ layout features at low - resolution scales, which can provide rich and comprehensive feature bases for the final segmentation result.
[0062] S103: Decode the second fused feature map to obtain the image segmentation result.
[0063] In this embodiment, by decoding the second fused feature map, such as restoring the resolution through upsampling and converting the features into pixel - category information, etc., the second fused feature map can be transformed into the specific medical image segmentation situation, clearly dividing different tissues, lesions, etc. regions to obtain the image segmentation result.
[0064] It can be concluded from the above that in this embodiment, by performing multi-scale local feature extraction on medical images, the obtained first local feature map can capture the detailed information of local regions in medical images; by performing multi-scale global feature extraction on medical images, the obtained first global feature map can capture more extensive context information. Feature fusion is performed on the first local feature map and the first global feature map corresponding to each scale, and the obtained first fusion feature map can comprehensively reflect local and global features. The first fusion feature maps corresponding to each scale are fused, so that features from low level to high level can be fully utilized. In summary, through the extraction and fusion of multi-scale features, the expression ability and segmentation performance of the model are enhanced.
[0065] Please refer to Figure 3 , in an embodiment of the present disclosure, performing feature fusion on the first local feature map and the first global feature map corresponding to each scale to obtain the first fusion feature map corresponding to each scale includes:
[0066] Performing channel pooling on the first local feature map to aggregate information of each channel to obtain a second local feature map;
[0067] Passing the first local feature map through an activation function to output a third local feature map;
[0068] Obtaining a fourth local feature map based on the second local feature map and the third local feature map;
[0069] Performing element-wise multiplication on the first local feature map and the first global feature map to obtain a third fusion feature map;
[0070] Processing the first global feature map based on a channel attention module and a spatial attention module to obtain a second global feature map;
[0071] Fusing the second global feature map, the third fusion feature map, and the fourth local feature map to obtain a first fusion feature map.
[0072] In this embodiment, for the first local feature generated by the spatial branch , first perform channel pooling, and then aggregate information of different channels through learning ( Figure 3 's CBR module, convolution + normalization + ReLU) to enhance the robustness of the model. Next, this embodiment applies the Sigmoid function to the input feature and multiplies the results, so that important information is retained while compressing the feature representation. This method helps to retain features beneficial to the medical segmentation task and suppress noise that may affect the segmentation performance of the model. This process can be expressed as:
[0073] = Sigmoid( ) × Conv(Channelpool( ))
[0074] Meanwhile, directly apply the Hadamard product to perform pointwise multiplication of features Mf and Cf, and then perform a convolution operation to achieve feature learning. The Hadamard product emphasizes the features considered important in the two branches, pointing to relevant information. This process can be expressed as:
[0075] = Conv( ⊙ )
[0076] For the first global feature generated by the state space branch , process the first global feature map based on the channel attention module and the spatial attention module to obtain the second global feature map .
[0077] Finally, fuse the three obtained features , and by concatenation and input them into the residual module to achieve automatic feature integration:
[0078] = Residual(Concat( , , ))
[0079] It can be concluded from the above that in this embodiment, the global features are further refined through two attention strategies, and then channel-level feature fusion and residual connection are performed. Multiple feature fusion methods can effectively extract and integrate feature information at different levels.
[0080] In an embodiment of the present disclosure, processing the first global feature map based on the channel attention module and the spatial attention module to obtain the second global feature map includes:
[0081] Input the first global feature map into the channel attention module to obtain the third global feature map;
[0082] Obtain the fourth global feature map based on the third global feature map and the first global feature map;
[0083] Input the fourth global feature map into the spatial attention module to obtain the fifth global feature map;
[0084] Obtain the second global feature map based on the fourth global feature map and the fifth global feature map.
[0085] In this embodiment, the channel attention module includes:
[0086] (1)Feature pooling: including max pooling operation and average pooling operation;
[0087] (2)Feature processing: including shared MLP processing, addition operation and Sigmoid activation function;
[0088] (3)Output: generating a channel attention feature map.
[0089] The spatial attention module includes:
[0090] (1)Input: channel redefined feature;
[0091] (2)Processing flow: [max pooling, average pooling] operation, convolutional layer processing and Sigmoid activation function;
[0092] (3)Output: generating a spatial attention feature map.
[0093] It can be concluded from the above that this embodiment adopts the channel and spatial attention mechanism (CBAM) to more effectively select the medical image features extracted by the state space branch from both the channel and spatial dimensions , discarding irrelevant information.
[0094] Please refer to Figure 4 , in an embodiment of the present disclosure, fusing the first fusion feature maps corresponding to each scale to obtain a second fusion feature map, including:
[0095] Dividing the first fusion feature maps corresponding to each scale into a first fine-grained feature map and a first coarse-grained feature map; wherein, the scale of the first fine-grained feature map is larger than the size of the first coarse-grained feature map;
[0096] Sequentially passing the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing to obtain a second coarse-grained feature map;
[0097] Performing feature fusion on the first fine-grained feature maps and the second coarse-grained feature maps corresponding to each scale to obtain a second fusion feature map.
[0098] In this embodiment, in the first fusion feature maps corresponding to each scale, ( , , ) are regarded as fine-grained (low-level) features, providing structural contours and context information, such as localization, while ( ) are coarse-grained (high-level) features, with richer and more accurate semantic details, which are crucial for generating accurate segmentation predictions.
[0099] When fusing the first fusion feature maps corresponding to each scale, first, the first coarse-grained feature map ( )The second coarsely grained feature map is obtained through the PPU unit (Pyramid Pooling Unit), and then the second coarsely grained feature map is fused with the first finely grained feature maps corresponding to each scale to obtain the second fused feature map.
[0100] Specifically, in the PPU unit, it is processed through three adaptive max pooling layers to adjust it to the specified size, and then learning is performed through convolution. This method captures the detailed texture of the coarsely grained features at different scales, enhances the semantic information at each level, and integrates the features at different scales to improve the accuracy and robustness of segmentation.
[0101] The specific operation details of the PPU are as follows:
[0102] (1) Feature extraction is performed using pooling kernels of different sizes;
[0103] (2) Align the features at each scale through upsampling;
[0104] (3) Concatenate all the features and perform convolution fusion;
[0105] A1 = Conv(AMPamp=1( )) ;
[0106] A2 = Conv(AMPamp=3( )) ;
[0107] A3 = Conv(AMPamp=6( )) ;
[0108] = Concat( , up(A1), up(A2), up(A3);
[0109] Among them, AMP represents the adaptive max pooling operation, amp represents the specified size of the feature, and up represents upsampling the feature generated by AMP to the size of ( ) for splicing. Subsequently, the features generated by the PPU unit are aggregated with the low-level features from each layer ( , , ). Finally, the features at each scale are fused to generate the prediction map for the segmentation task.
[0110] It can be concluded from the above that in this embodiment, through multi-scale feature extraction and progressive pooling, the coarsely grained and finely grained feature information is effectively combined, which can enhance the feature extraction ability.
[0111] In one embodiment of the present disclosure, before fusing the first fine-grained feature maps and the second coarse-grained feature maps corresponding to each scale, the medical image segmentation method further includes:
[0112] Performing convolution processing on each first fine-grained feature map through multiple paths to obtain multiple second fine-grained feature maps corresponding to each first fine-grained feature map;
[0113] Fusing the multiple second fine-grained feature maps to obtain an updated first fine-grained feature map.
[0114] In this embodiment, before fusing the first fine-grained feature maps and the second coarse-grained feature maps corresponding to each scale, the first fine-grained feature maps corresponding to each scale are simultaneously input into a Separable Dilated Unit (SDU). The first fine-grained feature maps are processed through four different paths in the SDU: the first path uses a (1x1) convolution, while the second, third, and fourth paths combine standard convolution, strip convolution, and dilated convolution with dilation rates of 3 respectively, and the kernel sizes of the convolution in each path are 3, 5, and 7 respectively.
[0115] Specifically, the structure of the Separable Dilated Unit includes:
[0116] (1) Input processing: Initial Conv1×1 convolution;
[0117] (2) Three parallel branches:
[0118] The first branch: Conv3×3 (d = 3), Conv1×1;
[0119] The second branch: Conv1×5, Conv5×5 (d = 3);
[0120] The third branch: Conv1×7, Conv7×7 (d = 3);
[0121] (3) Output processing: The results of all branches are merged through Concat and output .
[0122] = Conv1×1( ) ;
[0123] = D3Conv3×3(Conv3×1(Conv1×3(Conv3×3( )))) ;
[0124] = D3Conv5×5(Conv5×1(Conv1×5(Conv5×5( )))) ;
[0125] = D3Conv7×7(Conv7×1(Conv1×7(Conv7×7( )))) ;
[0126] = Concat( , , , , )。
[0127] It can be concluded from the above that the decoder in this embodiment integrates a Separable Dilated Unit (SDU) and a Pyramid Pooling Unit (PPU) to more accurately identify the detailed textures and structural features in images of different sizes, thereby improving the accuracy of the prediction map.
[0128] In one embodiment of the present disclosure, before sequentially passing the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing, the medical image segmentation method further includes:
[0129] Passing the first coarse-grained feature map through convolutional processing along multiple paths to obtain multiple third fine-grained feature maps corresponding to the first coarse-grained feature map;
[0130] Fusing the multiple third fine-grained feature maps to obtain an updated first coarse-grained feature map.
[0131] In this embodiment, using the same processing method as the first fine-grained feature map, before sequentially passing the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing, passing the first coarse-grained feature map through convolutional processing along multiple paths, and by using split hole units to expand the receptive field, the feature extraction ability can be further enhanced.
[0132] In one embodiment of the present disclosure, the convolutional processing along multiple paths includes: 1x1 convolution, standard convolution, strip convolution, and dilated convolution.
[0133] In summary, the processing flow of the decoder for mixing coarse and fine features in this embodiment is as follows:
[0134] (1) Top branch ( ) : The SDU module processes the input, passes through the CBR module, addition operation, then passes through the CBR module again and outputs ;
[0135] (2) Middle branch (D f2 ) : A processing flow similar to the top branch, SDU → CBR → addition → CBR, outputs .
[0136] (3) The third branch (D f3 ): The same processing flow, SDU → CBR → addition → CBR,
[0137] Output .
[0138] (4) The bottom branch ( and PPU):
[0139] Processed by the SDU module;
[0140] Enter the PPU (Progressive Pooling Unit) module;
[0141] It contains three different AMP (AdaptiveMaxPool) layers: AMP = 1, AMP = 3, and AMP = 6;
[0142] Feature size: B×C×H×W.
[0143] (5) Feature fusion:
[0144] All branches are connected through the Concat operation; the final output is the prediction map.
[0145] Experiments were conducted using the method of this embodiment and the existing method on four medical image datasets, BUSI, DDTI, TN3K, and ISIC2016. The experimental results are as Figure 5 shown. According to the experimental results, it can be seen that the method of the present invention has significant advantages compared with the existing method:
[0146] In the BUSI dataset, YMamba proposed in this embodiment achieved state-of-the-art (SOTA) results in both metrics. It is 0.01 higher than PolypPVT in the Dice metric and reaches 0.768 in the IoU, which is 1.1% and 1.8% higher than EH-Former and NPDNet respectively. Although CMUNeXt performs well in global context modeling, YMamba improves by an average of 6.75% in both metrics, with the IoU increasing by 7.4%.
[0147] In the DDTI dataset, EH-Former also achieved SOTA on two evaluation metrics. However, YMamba in this embodiment further surpassed or tied with EH-Former on these two metrics, with the Dice increased by 1.2% and the IoU increased by 0.8%. It is worth noting that compared with VMUNet which also uses VMamba as the backbone network, YMamba in this embodiment exceeded by 0.11 and 0.19 on the Dice and IoU metrics respectively.
[0148] For TN3K, YMamba performed excellently on both Dice and IoU, with scores of 0.859 and 0.778. Compared with M2SNet, the method in this embodiment increased by 0.019 on Dice.
[0149] In the ISIC2016 dataset, YMamba performed best on IoU, only 0.006 lower than the leading VMUNet. It is worth mentioning that compared with the recently proposed CMUNet, the average performance of the method in this embodiment increased by 11.5%, and particularly increased by 15% on IoU. Compared with PraNet, which also uses a method combining rough and detailed segmentation, the method in this embodiment increased by 3.8% on Dice and 5.7% on IoU. These results indicate that YMamba has made significant progress in medical image segmentation tasks, demonstrating its wide applicability and superior performance on multiple datasets.
[0150] Corresponding to the medical image segmentation method in the above embodiment, Figure 2 is a structural block diagram of a medical image segmentation device provided by an embodiment of the present disclosure. For ease of description, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The medical image segmentation device includes: a dual-branch feature extraction module, a dual-branch fusion module, and a decoder.
[0151] Among them, the dual-branch feature extraction module is used to perform multi-scale local feature extraction and global feature extraction on medical images to obtain the first local feature map and the first global feature map corresponding to each scale;
[0152] The dual-branch fusion module is used to fuse the first local feature map and the first global feature map corresponding to each scale to obtain the first fusion feature map corresponding to each scale; fuse the first fusion feature maps corresponding to each scale to obtain the second fusion feature map;
[0153] The decoder is used to decode the second fusion feature map to obtain the image segmentation result.
[0154] In one embodiment of the present disclosure, the dual-branch fusion module is specifically configured to:
[0155] Perform channel pooling on the first local feature map to aggregate the information of each channel and obtain a second local feature map;
[0156] Pass the first local feature map through an activation function to output a third local feature map;
[0157] Obtain a fourth local feature map based on the second local feature map and the third local feature map;
[0158] Perform element-wise multiplication on the first local feature map and the first global feature map to obtain a third fusion feature map;
[0159] Process the first global feature map based on a channel attention module and a spatial attention module to obtain a second global feature map;
[0160] Fuse the second global feature map, the third fusion feature map, and the fourth local feature map to obtain a first fusion feature map.
[0161] In one embodiment of the present disclosure, the dual-branch fusion module is specifically further configured to:
[0162] Input the first global feature map into the channel attention module to obtain a third global feature map;
[0163] Obtain a fourth global feature map based on the third global feature map and the first global feature map;
[0164] Input the fourth global feature map into the spatial attention module to obtain a fifth global feature map;
[0165] Obtain a second global feature map based on the fourth global feature map and the fifth global feature map.
[0166] In one embodiment of the present disclosure, the dual-branch fusion module is specifically configured to:
[0167] Divide the first fusion feature map corresponding to each scale into a first fine-grained feature map and a first coarse-grained feature map; wherein, the scale of the first fine-grained feature map is larger than the size of the first coarse-grained feature map;
[0168] Pass the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing in sequence to obtain a second coarse-grained feature map;
[0169] Fuse the first fine-grained feature map and the second coarse-grained feature map corresponding to each scale to obtain a second fusion feature map.
[0170] In one embodiment of the present disclosure, the dual-branch fusion module is specifically further configured to:
[0171] Each first fine-grained feature map is respectively subjected to convolutional processing through multiple paths to obtain multiple second fine-grained feature maps corresponding to each first fine-grained feature map;
[0172] The multiple second fine-grained feature maps are fused to obtain an updated first fine-grained feature map.
[0173] In an embodiment of the present disclosure, the dual-branch fusion module is specifically further configured to:
[0174] The first coarse-grained feature map is respectively subjected to convolutional processing through multiple paths to obtain multiple third fine-grained feature maps corresponding to the first coarse-grained feature map;
[0175] The multiple third fine-grained feature maps are fused to obtain an updated first coarse-grained feature map.
[0176] In an embodiment of the present disclosure, the convolutional processing through multiple paths includes: 1x1 convolution, standard convolution, strip convolution, and dilated convolution.
[0177] See Figure 6 , Figure 6 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 6 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments.
[0178] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (Central Processing Unit, CPU), and this processor may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field programmable gate arrays (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0179] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0180] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may further store information about the device type.
[0181] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the medical image segmentation method provided in the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.
[0182] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the above-described method embodiments are implemented. It may also be completed by instructing related hardware through the computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-described method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0183] A computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.
[0184] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0185] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0186] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection may be an indirect coupling or communication connection through some interfaces or units, and may also be an electrical, mechanical, or other form of connection.
[0187] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.
[0188] In addition, in various embodiments of the present disclosure, each functional unit may be integrated into a processing unit, may exist individually as a physical unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0189] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A medical image segmentation method, characterized in that, Including: Performing multi-scale local feature extraction and global feature extraction on a medical image to obtain a first local feature map and a first global feature map corresponding to each scale; Performing feature fusion on the first local feature map and the first global feature map corresponding to each scale to obtain a first fusion feature map corresponding to each scale; Fusing the first fusion feature maps corresponding to each scale to obtain a second fusion feature map; Decoding the second fusion feature map to obtain an image segmentation result; The step of performing feature fusion on the first local feature map and the first global feature map corresponding to each scale to obtain a first fusion feature map corresponding to each scale includes: Performing channel pooling on the first local feature map to aggregate information of each channel to obtain a second local feature map; Passing the first local feature map through an activation function to output a third local feature map; Obtaining a fourth local feature map based on the second local feature map and the third local feature map; Performing element-wise multiplication on the first local feature map and the first global feature map to obtain a third fusion feature map; Processing the first global feature map based on a channel attention module and a spatial attention module to obtain a second global feature map; Fusing the second global feature map, the third fusion feature map, and the fourth local feature map to obtain the first fusion feature map; The step of processing the first global feature map based on a channel attention module and a spatial attention module to obtain a second global feature map includes: Inputting the first global feature map into the channel attention module to obtain a third global feature map; Obtaining a fourth global feature map based on the third global feature map and the first global feature map; Inputting the fourth global feature map into the spatial attention module to obtain a fifth global feature map; Obtaining the second global feature map based on the fourth global feature map and the fifth global feature map; The step of fusing the first fusion feature maps corresponding to each scale to obtain a second fusion feature map includes: Dividing the first fusion feature map corresponding to each scale into a first fine-grained feature map and a first coarse-grained feature map; wherein, the scale of the first fine-grained feature map is larger than the size of the first coarse-grained feature map; Sequentially passing the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing to obtain a second coarse-grained feature map; Performing feature fusion on the first fine-grained feature map and the second coarse-grained feature map corresponding to each scale to obtain a second fusion feature map; Before performing feature fusion on the first fine-grained feature map and the second coarse-grained feature map corresponding to each scale, the medical image segmentation method further includes: Passing each first fine-grained feature map through convolutional processing of multiple paths respectively to obtain multiple second fine-grained feature maps corresponding to each first fine-grained feature map; Fusing the multiple second fine-grained feature maps to obtain an updated first fine-grained feature map.
2. The medical image segmentation method according to claim 1, wherein Before sequentially passing the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing, the medical image segmentation method further includes: Passing the first coarse-grained feature map through convolutional processing of multiple paths respectively to obtain multiple third fine-grained feature maps corresponding to the first coarse-grained feature map; Fuse multiple third fine-grained feature maps to obtain an updated first coarse-grained feature map.
3. The medical image segmentation method according to claim 1 or 2, characterized in that, The convolutional processing of the multiple paths includes: 1x1 convolution, standard convolution, strip convolution, and dilated convolution.
4. A medical image segmentation device, characterized in that, It includes: A dual-branch feature extraction module for performing multi-scale local feature extraction on a medical image to obtain a first local feature map corresponding to each scale; A dual-branch fusion module for performing multi-scale global feature extraction on a medical image to obtain a first global feature map corresponding to each scale; fusing the first global feature map and the first local feature map corresponding to each scale to obtain a first fusion feature map corresponding to each scale; A decoder for decoding the first fusion feature map to obtain an image segmentation result; Specifically, the dual-branch fusion module is used for: Performing channel pooling on the first local feature map to aggregate information of each channel to obtain a second local feature map; Passing the first local feature map through an activation function to output a third local feature map; Obtaining a fourth local feature map based on the second local feature map and the third local feature map; Performing element-wise multiplication on the first local feature map and the first global feature map to obtain a third fusion feature map; Processing the first global feature map based on a channel attention module and a spatial attention module to obtain a second global feature map; Fusing the second global feature map, the third fusion feature map, and the fourth local feature map to obtain a first fusion feature map; Specifically, the dual-branch fusion module is further used for: Inputting the first global feature map into a channel attention module to obtain a third global feature map; Obtaining a fourth global feature map based on the third global feature map and the first global feature map; Inputting the fourth global feature map into a spatial attention module to obtain a fifth global feature map; Obtaining a second global feature map based on the fourth global feature map and the fifth global feature map; Specifically, the dual-branch fusion module is used for: Dividing the first fusion feature map corresponding to each scale into a first fine-grained feature map and a first coarse-grained feature map; wherein, the scale of the first fine-grained feature map is larger than the size of the first coarse-grained feature map; Sequentially passing the first coarse-grained feature map through an adaptive max pooling layer and convolutional processing to obtain a second coarse-grained feature map; Fusing the first fine-grained feature map and the second coarse-grained feature map corresponding to each scale to obtain a second fusion feature map; Specifically, the dual-branch fusion module is further used for: Passing each first fine-grained feature map through convolutional processing of multiple paths to obtain multiple second fine-grained feature maps corresponding to each first fine-grained feature map; Fusing the multiple second fine-grained feature maps to obtain an updated first fine-grained feature map.
5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Lightweight pedestrian re-identification method combined with multi-level features
CN115841683A
Global and local feature reconstruction network-based medical image segmentation method
US20230274531A1