Method for enhancing medical image segmentation through multi-scale and multi-view frequency fusion
Through the multi-scale and multi-view frequency fusion method, the characteristics of different encoder levels are integrated and the global frequency domain signals and local spatial information are captured, which solves the problem of difficulty in capturing long-distance correlation and ignoring multi-level global information in the prior art, and achieves higher medical image segmentation accuracy.
Patent Information
- Application Number
- CN202411663539.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-02
AI Technical Summary
Existing medical image segmentation technology is difficult to effectively capture long-distance correlations, and ignores the potential of multi-level and multi-scale global information, resulting in limited segmentation accuracy.
Using a multi-scale and multi-view frequency fusion method, the characteristics of different encoder levels are integrated through the MSFA module, and the MVFE module is used to capture global frequency domain signals and local spatial information to enrich feature representation.
The accuracy of multi-scale feature extraction and processing is improved, more detailed information is retained, and the accuracy of medical image segmentation is significantly improved.
Smart Images

Figure CN119919425A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image segmentation, and in particular relates to a method for enhancing medical image segmentation through multi-scale and multi-view frequency fusion. Background Art
[0002] Medical image segmentation is an essential component of intelligent healthcare systems, enabling precise quantitative assessment of health information and supporting healthcare providers in making accurate diagnoses of various medical conditions. This task is inherently complex due to the diversity of medical images, which often contain various structures such as organs, blood vessels, and bones, which vary in shape, size, and location. These inherent characteristics pose significant challenges. Specifically, medical images exhibit a high degree of irregularity, with significant differences in the morphology, scale, and location of structures. For example, when taking images of skin lesions, various external factors such as hair and blood vessels can affect the results. In addition, the irregular color variations and rough contours of lesions pose a challenge to developing accurate segmentation models, making them more difficult than traditional methods. In addition, the quality of medical images also plays a key role. Low-resolution images often have difficulty depicting sharp edges and fine details, while high-resolution images may overwhelm the segmentation process due to their complexity, resulting in inconsistent results.
[0003] The problems with the above technologies are: 1. The U-Net framework in the existing technology has made breakthrough progress in the field of medical image segmentation. The framework stands out for its balanced encoder-decoder configuration. It uses a unique U-shaped design to cleverly identify and enhance features, but they also encounter difficulties in accurately capturing long-distance correlations. This limitation stems from the fact that convolution operations mainly focus on local features and spatial interactions within a limited receptive field, thereby limiting their ability to capture broader contextual information. Solving this limitation is an important obstacle faced by CNN-based models in medical images;
[0004] 2. In the existing technology, with the continuous progress of the Transformer architecture in image recognition, its application exploration in various downstream tasks has also made significant progress, especially in medical image segmentation. TransUnet is an important progress in this field. It is an innovative hybrid framework that integrates the advantages of CNN and Transformer. Despite these advances, many existing methods mainly focus on transferring the features of the encoder-related layers to the decoder. This approach often ignores the potential of utilizing the large amount of multi-level and multi-scale global information accumulated at each stage of the hierarchical encoder block, thereby limiting the model's ability to fully utilize these rich features. Summary of the invention
[0005] In view of the problems existing in the prior art, the present invention provides a method for enhancing medical image segmentation through multi-scale and multi-view frequency fusion, which can overcome the above problems or at least partially solve the above problems.
[0006] The present invention is implemented as follows: a method for enhancing medical image segmentation by multi-scale and multi-view frequency fusion, named MSFM-UNet, the MSFM-UNet includes a VSS module, a layer for merging patches, a layer for expanding patches, and a module dedicated to multi-scale feature aggregation (MSFA);
[0007] S1. First, the patch embedding layer takes the input image The image is split into different patches, each of which is 4×4 pixels in size and does not overlap. Each patch is then converted into a vector of dimension C, which is usually set to 96. This conversion produces an embedded image.
[0008] S2. Before feature extraction in the encoder, X′ is layer normalized. The encoder is divided into four independent stages. In the initial three stages, a patch merging operation is performed after each stage. This operation can reduce the spatial dimension of the features while increasing the number of channels. The VSS modules are distributed in the four stages in the following configuration: {2, 2, 2, 2}, and the channel dimension is set to {C, 2C, 4C, 8C};
[0009] S3, the decoder and encoder have four stages, where the first three stages use patch expansion operations to reverse channel reduction and restore spatial dimensions. In the decoder, the structure of the VSS module is {2, 2, 2, 1}, and the number of channels in each stage is set to {8C, 4C, 2C, C}. After the decoding stage, the final projection layer is used to align the size of the feature map with the segmentation target. This process involves four times upsampling, where the patch expansion layer restores the spatial size of the feature map;
[0010] S4. Subsequently, different projection layers convert the feature map back to its initial number of channels. The MSFA module enhances the skip connection by combining the multi-scale features with the features of the advanced decoder, while the MVFE module utilizes the frequency domain information extracted in the MSFA module to enrich the multi-scale feature representation extracted by MSFA. This method can extract richer and more detailed image features, thereby obtaining more accurate segmentation results.
[0011] Compared with the prior art, the present invention has the following beneficial effects:
[0012] The present invention achieves accurate segmentation in medical imaging by setting up multi-scale feature aggregation (MSFA) and multi-view frequency enhancement (MVFE). The MSFM-UNet model, which contains the core VSS module, uses MSFA to integrate features from different encoder levels, and applies MVFE to enrich feature representation by simultaneously capturing global frequency domain signals and local spatial information. The synergy of these methods aims to improve the accuracy of multi-scale feature extraction and processing, and retain more detailed information. Comprehensive evaluation on skin lesion and multi-organ segmentation datasets shows that our proposed model exhibits strong performance in the field of medical image segmentation tasks. These results emphasize the robustness of the model and indicate great potential for further improvement and exploration in subsequent research work, solving the problem of difficulties and limitations in capturing long-range correlations. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of a model of MSFM-UNet provided by an embodiment of the present invention;
[0014] Figure 2 The embodiment of the present invention provides a schematic diagram of multi-scale feature aggregation;
[0015] Figure 3 It is a schematic diagram of a multi-view frequency enhancement module provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to further understand the content, features and effects of the present invention, the following embodiments are given as examples and described in detail with reference to the accompanying drawings.
[0017] The structure of the present invention is described in detail below in conjunction with the accompanying drawings.
[0018] like Figures 1 to 3 As shown, an embodiment of the present invention provides a method for enhancing medical image segmentation by multi-scale and multi-view frequency fusion, which is named MSFM-UNet. MSFM-UNet includes a VSS module, a layer for merging patches, a layer for expanding patches, and a module dedicated to multi-scale feature aggregation (MSFA);
[0019] S1. First, the patch embedding layer takes the input image The image is split into different patches, each of which is 4×4 pixels in size and does not overlap. Each patch is then converted into a vector of dimension C, which is usually set to 96. This conversion produces an embedded image.
[0020] S2. Before feature extraction in the encoder, X′ is layer normalized. The encoder is divided into four independent stages. In the initial three stages, a patch merging operation is performed after each stage. This operation can reduce the spatial dimension of the features while increasing the number of channels. The VSS modules are distributed in the four stages in the following configuration: {2, 2, 2, 2}, and the channel dimension is set to {C, 2C, 4C, 8C};
[0021] S3, the decoder and encoder have four stages, where the first three stages use patch expansion operations to reverse channel reduction and restore spatial dimensions. In the decoder, the structure of the VSS module is {2, 2, 2, 1}, and the number of channels in each stage is set to {8C, 4C, 2C, C}. After the decoding stage, the final projection layer is used to align the size of the feature map with the segmentation target. This process involves four times upsampling, where the patch expansion layer restores the spatial size of the feature map;
[0022] S4. Subsequently, different projection layers convert the feature map back to its initial number of channels. The MSFA module enhances the skip connection by combining the multi-scale features with the features of the advanced decoder, while the MVFE module utilizes the frequency domain information extracted in the MSFA module to enrich the multi-scale feature representation extracted by MSFA. This method can extract richer and more detailed image features, thereby obtaining more accurate segmentation results.
[0023] VSS Module:
[0024] The VSS module is a key component of MSFM-UNet, which is based on the VMamba architecture. The structure of the VSS module consists of several key steps. First, the input is layer normalized and then divided into two different processing paths. In the first path, the input is initially transformed through a linear layer, and then an activation function is applied to introduce nonlinearity. At the same time, the second path applies a series of processing including linear transformation, depthwise separable convolution, and additional activation functions, aiming to capture complex features and enhance the representation ability of the model. The 2D Selective Sweep (SS2D) module is designed for complex feature extraction and is used to receive processed data for further analysis. After preliminary processing, layer normalization is reapplied and the features are element-wise multiplied with the output of the first path, effectively merging the two information streams. These merged features are then passed through a final linear layer, and the final output is integrated with the original input through a residual connection, completing the function of the VSS module. In this module, the SiLU activation function is used as the main mechanism to enhance nonlinear transformations.
[0025] Multi-Scale Feature Aggregation:
[0026] A new module called Multi-Scale Feature Aggregation (MSFA) that performs feature fusion at each stage of the VSS encoder instead of using traditional skip connections. This module combines multi-scale features from different encoder stages with contextual information from lower-level decoders. This strategy compensates for the loss of spatial detail information caused by downsampling and effectively improves the model's ability to retain and utilize complex spatial details during segmentation. By fusing these multi-scale features with rich contextual information, the decoder's ability to make more accurate and informed decisions is significantly improved, thereby improving segmentation accuracy.
[0027] This approach maximizes the effective utilization of the feature maps generated by the VSS encoder modules. The vertically arranged outputs of multiple encoder modules are denoted as X1_nput, X2_nput, …, Xi_input, where each To handle the different resolutions between feature maps at different encoder stages, downsampling is performed using a patch merging layer, while upsampling is performed through a patch expansion layer. In addition, when dealing with feature maps with different dimensions, it is important to prevent the model from disproportionately prioritizing feature maps with larger dimensions, which may lead to output bias. To ensure consistency, linear projection is applied on all feature maps to standardize their dimensions. In addition, the introduction of the MBConv layer is crucial for the combination and extraction of features, which greatly improves the efficiency of the module, just as relative position embedding improves the performance of the entire network. Therefore, in the jth MSFA module from top to bottom, the multi-scale results of the VSS encoder block are first processed, resulting in more refined feature integration.
[0028]
[0029] Among them, X i Represents the i-th input X i_input The corresponding optimization results. In addition, an efficient MBConv layer is used to further enhance the feature extraction capability.
[0030] Subsequently, the processed outputs are combined using a concatenation operation to obtain the final feature representation, denoted as X.
[0031] X=Concat(X1,X2,...,X n ). (2)
[0032] Before reducing the dimensionality of this concatenated feature, we integrate multi-view frequency enhancement MVFE (see Section 3.4 for details) to capture both global and local information. The integration of frequency domain feature filtering enhances the model’s ability to comprehensively analyze image texture and edge features.
[0033] X encoder =Linear Proj(MVFE(X)), (3)
[0034] The MVFE operation represents multi-view frequency enhancement, and the LinearProj operation represents a fully connected layer. Here, X encoder Represents the output of multi-scale features processed by different encoders.
[0035] At the same time, the feature map obtained from the decoder output is carefully adjusted to match the processed result (X encoder ) are synchronized to obtain a unified feature map, namely X decoder .
[0036] X decoder =LinearProj(X d ), (4)
[0037] in, Represents the feature map generated by the lower layer decoder.
[0038] Finally, the feature maps generated by the encoder are integrated with the feature maps generated by the decoder through a cross-attention mechanism and then processed by the feed-forward network (FFN). The calculation process of the cross-attention mechanism is as follows:
[0039] z=softmax((W Q X encoder )(W K X deco d eer ) T )W V X decoder , (5)
[0040] Among them, WQ, WK and WV are randomly initialized and fine-tuned during the training process. decoder represents the features from the lower layer of the decoder, X encoder represents the processed output of the encoder and z represents the output of the cross attention.
[0041] Multi-View Frequency Enhancement
[0042] Recent developments in medical image segmentation have mainly focused on enhancing the capture of spatial information, while often ignoring the critical role of frequency domain analysis. In the spatial domain, there is considerable difficulty in clearly distinguishing the boundary between the target object and its background. In contrast, in the frequency domain, objects exhibit different frequencies, which facilitates a clearer distinction. In addition, frequency domain processing helps to enhance texture and boundary information in medical images. Figure 3As shown in Figure 1, we extracted a sample from the ISIC17 dataset and obtained the high-frequency information of the image by fast Fourier transform (FFT). Obviously, the edges and textures (called high frequencies) present a brighter appearance (white-gray), while the background and the gradually changing areas in the lesion (called low frequencies) appear darker. High-frequency edge information can well distinguish the lesion area and the background in medical images. In this study, we proposed MVFE using 2D discrete Fourier transform (DFT) from different perspectives. Our goal is to expand the local and global information collection in the frequency domain.
[0043] Consider the input feature map Where H, W, and C represent the height, width, and number of channels, respectively. First, X is divided into four parts along the channel axis. Each part is then processed by a different branch. The mathematical representation of the MVFE operation is shown in Equation 6 to Equation 9.
[0044] x1, x2, x3, x4=Split(X), (6)
[0045]
[0046] Y=Concat(x′1,x′2,x′3,x′4)+X, (9)
[0047] We use different indexes i=1, 2, 3 to represent different views:
[0048] i = 1 corresponds to a height-width view, where p and q represent the height and width dimensions, respectively;
[0049] i = 2 corresponds to the channel-width view, where p and q represent the channel and width dimensions, respectively;
[0050] i=3 corresponds to the channel-height view, where p and q represent the channel and height dimensions, respectively.
[0051] In our module, W(p,q) represents the trainable external weights, represents the 2D discrete Fourier transform (DFT) of each view, and Denotes a two-dimensional inverse DFT operation. DW stands for depthwise convolution, and the symbol ⊙ stands for element-wise multiplication. In addition, operations such as Split and Concat represent splitting along the channel dimension of the feature map and then concatenating the parts together.
[0052] In the first branch, a 2D discrete Fourier transform (DFT) is applied on the feature map along the height-width axis, effectively transforming it into the frequency domain. After this transformation, the feature map is element-wise multiplied with the trainable weights, and then a 2D inverse DFT is applied to transform the feature map back into the spatial domain. The procedure in the second and third branches repeats this operation, focusing on the channel-width and channel-height dimensions. Multi-view methods can enhance the capture of global information. In addition, capturing local details is crucial for medical image segmentation. Therefore, the fourth branch uses deep convolution to extract local features, which helps maintain computational efficiency and reduce the number of parameters. The outputs of the four branches are aggregated along the channel dimension to restore the original spatial resolution of the feature map. Afterwards, a residual connection is incorporated into the combined output, which is merged with the initial input feature map to generate the final result. Frequency domain information is crucial in medical image segmentation as it greatly helps distinguish the lesion area from the adjacent background.
[0053] The purpose of implementing MSFM-UNet is to evaluate the performance of MSFA and MVFE components in the field of medical image segmentation. In order to solve both binary segmentation and multi-class segmentation problems, we use two main loss functions: binary cross entropy combined with Dice loss (BceDice loss) for binary segmentation and cross entropy combined with Dice loss (CeDice loss) for multi-class segmentation. These loss functions are defined by Equation 10 and Equation 11, respectively.
[0054] L BceDice =λ1L Bce +λ2L Dice , (10)
[0055] L CeDice =λ1L Ce +λ2L Dice , (11)
[0056]
[0057] Among them, N represents the total number of samples and C represents the total number of categories.
[0058] For each sample m, y m and Represent the true label and model prediction respectively. If sample m is classified into category c, the binary variable y m,c The value is 1; otherwise, the value is 0. The estimated probability that sample m belongs to category c is expressed as The symbol |X| represents the actual value, while the symbol |Y| represents the value predicted by the model. In addition, λ1 and λ2 are coefficients that measure the corresponding loss function, and their initial values are both 1.
[0059] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0060] The above description is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this patent will not depart from the scope of the technical solution of the present invention.
Claims
1. A method for enhancing medical image segmentation via multi-scale and multi-view frequency fusion, MSFM-UNet includes a VSS module, a layer for merging patches, a layer for expanding patches, and a module dedicated to multi-scale feature aggregation (MSFA); S1. First, the patch embedding layer takes the input image The image is split into different patches, each of which is 4×4 pixels in size and does not overlap. Each patch is then converted into a vector of dimension C, which is usually set to 96. This conversion produces an embedded image. S2. Before feature extraction in the encoder, X′ is layer normalized. The encoder is divided into four independent stages. In the initial three stages, a patch merging operation is performed after each stage. This operation can reduce the spatial dimension of the features while increasing the number of channels. The VSS modules are distributed in the four stages in the following configuration: {2,2,2,2}, and the channel dimension is set to {C,2C,4C,8C}; S3, the decoder and encoder have four stages, where the first three stages use patch expansion operations to reverse channel reduction and restore spatial dimensions. In the decoder, the structure of the VSS module is {2,2,2,1}, and the number of channels in each stage is set to {8C,4C,2C,C}. After the decoding stage, the final projection layer is used to align the size of the feature map with the segmentation target. This process involves four times upsampling, where the patch expansion layer restores the spatial size of the feature map; S4. Subsequently, different projection layers convert the feature map back to its initial number of channels. The MSFA module enhances the skip connection by combining the multi-scale features with the features of the advanced decoder, while the MVFE module utilizes the frequency domain information extracted in the MSFA module to enrich the multi-scale feature representation extracted by MSFA. This method can extract richer and more detailed image features, thereby obtaining more accurate segmentation results.
Citation Information
Cited By
Diabetic retinopathy image classification system and method based on deep learning
CN120298846A
Diabetic retinopathy image classification system and method based on deep learning
CN120298846B