Medical image segmentation method and device based on spatial perception and frequency domain information
Through the FDMUNet network model combined with the multi-head perceived visual state space module and the focus context attention module, the problem of insufficient fusion of global and local features in medical image segmentation is solved, and efficient multi-scale feature extraction and noise filtering is realized, which significantly improves segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202510779465.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing medical image segmentation technology has problems such as insufficient fusion of global and local features, insufficient multi-scale feature extraction and low computational efficiency when processing complex images, resulting in loss of edge details and omission of lesion areas.
Using a medical image segmentation method based on spatial perception and frequency domain information, feature extraction is performed through the FDMUNet network model, combining the context feature extraction module focusing on low-frequency information, the multi-headed perceived visual state space module and the context attention module focusing on multi-scale feature fusion and noise filtering are performed to generate high-resolution segmentation results.
It significantly improves the accuracy and robustness of medical image segmentation, can effectively capture features of different scales, enhances the ability to identify lesion areas, and improves the accuracy and reliability of segmentation.
Smart Images

Figure CN120298441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and pattern recognition, and particularly relates to a medical image segmentation method and device based on spatial perception and frequency domain information. Background Art
[0002] In the field of medical image processing, image segmentation technology is one of the key links to achieve precision medicine. It precisely divides different tissues, organs or lesion areas in medical images, providing important basis for disease diagnosis, treatment planning and efficacy evaluation. In recent years, with the rise of deep learning technology, segmentation methods based on convolutional neural networks (CNNs) have made significant progress. Among them, the U-shaped network architecture (such as UNet), with its unique encoder-decoder structure and skip connection mechanism, has become the mainstream choice in medical image segmentation tasks.
[0003] However, the existing technologies still face many challenges when dealing with complex medical images. On the one hand, traditional CNN encoders perform well in extracting local features, but their fixed convolutional kernels are insufficient in capturing global semantic information and long-range spatial dependency relationships. Although some improvement methods (such as dilated spatial pyramid pooling and pyramid pooling modules) alleviate the problem of information loss through multi-scale feature fusion, they often neglect the complementarity between different scale features, resulting in insufficient learning of global and detailed features. On the other hand, the low-frequency information in medical images (such as the macroscopic structure of organs) is highly robust to noise, but existing methods fail to fully utilize frequency domain features, and low-frequency information is easily interfered by irrelevant backgrounds, thus affecting the segmentation accuracy.
[0004] In addition, when the bottleneck layer of the U-shaped network connects the encoder and the decoder, a balance needs to be achieved between global context modeling and computational efficiency. Although the bottleneck layer based on Transformer enhances the global perception ability through the self-attention mechanism, its high computational complexity limits its application in high-resolution medical images. And some lightweight models (such as Mamba combined with the state space model) reduce the computational overhead, but their feature extraction ability for multi-scale lesions (especially tiny lesions) is insufficient and difficult to adapt to the complex structure of medical images.
[0005] Generally speaking, the existing medical image segmentation technologies still have deficiencies in global and local feature fusion, multi-scale feature extraction and computational efficiency. Especially when dealing with complex medical images, problems such as edge detail loss and lesion area omission are likely to occur. Therefore, developing a medical image segmentation method that can effectively fuse frequency domain information, enhance multi-scale perception ability and improve computational efficiency is of great significance for improving the accuracy and reliability of medical image segmentation.
[0006] In view of this, the present application is proposed. Summary of the Invention
[0007] The present invention provides a medical image segmentation method and device based on spatial perception and frequency domain information, which can at least partially improve the above problems.
[0008] To achieve the above object, the present invention adopts the following technical solutions: A medical image segmentation method based on spatial perception and frequency domain information, which includes: Obtain the medical image to be processed, perform data preprocessing on the medical image, and perform quality assessment on the preprocessed medical image, calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluate the contrast and clarity of the image, detect the artifacts and noise level in the image, if the signal-to-noise ratio is lower than the preset threshold, perform adaptive histogram equalization, and selectively apply gamma correction to perform local enhancement processing on the detected artifact regions according to the image contrast situation; Call the pre-trained FDMUNet network model, and use the context feature extraction module that focuses on low-frequency information to perform feature extraction processing on the preprocessed medical image to obtain an output feature map; Adopt the multi-head perception visual state space module in the FDMUNet network model to perform extraction processing on the output feature map, and extract the global information features of the output feature maps at different scales; Use the focused context attention module in the FDMUNet network model to perform boundary enhancement processing on the feature map to obtain enhanced features; Use the decoder to perform recovery preprocessing on the global information features, and fuse the processed global information features with the enhanced features to obtain a segmentation result, repeat this operation until the preset number of times is reached, and generate a high-resolution feature map.
[0009] The present invention also provides a medical image segmentation device based on spatial perception and frequency domain information, which includes: A preprocessing unit, configured to obtain the medical image to be processed, perform data preprocessing on the medical image, and perform quality assessment on the preprocessed medical image, calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluate the contrast and clarity of the image, detect the artifacts and noise level in the image, if the signal-to-noise ratio is lower than the preset threshold, perform adaptive histogram equalization, and selectively apply gamma correction to perform local enhancement processing on the detected artifact regions according to the image contrast situation; A first extraction unit, configured to call the pre-trained FDMUNet network model, and use the context feature extraction module that focuses on low-frequency information to perform feature extraction processing on the preprocessed medical image to obtain an output feature map; A second extraction unit, which is used to perform extraction processing on the output feature map by using the multi-head perception visual state space module in the FDMUNet network model, and extract the global information features of the output feature map at different scales; An enhancement unit, which is used to perform boundary enhancement processing on the feature map by using the focus context attention module in the FDMUNet network model to obtain enhanced features; A fusion unit, which is used to perform recovery preprocessing on the global information features by using a decoder, and fuse the processed global information features with the enhanced features to obtain a segmentation result, and repeat this operation until a preset number of times is reached to generate a high-resolution feature map.
[0010] In summary, the medical image segmentation method based on spatial perception and frequency domain information significantly improves the segmentation accuracy and robustness by fusing frequency domain information and multi-head state space perception technology. The core of this method is to decompose the image into different frequency components through frequency domain transformation technology, focus on extracting low-frequency features containing global information, and effectively suppress background noise through a learnable noise filtering mechanism, enhancing the recognition ability of lesion areas. In terms of the model architecture, a multi-head perception visual state space module is introduced, which can extract image features from multiple scales, effectively capture multi-scale information from micro-lesions to the overall anatomical structure, and make up for the deficiencies of traditional methods in dealing with complex medical images. In addition, a context focus attention mechanism is added to the skip connection of this method to further optimize the fusion of global and local information and improve the adaptability of the model to complex scenarios. Finally, a high-resolution segmentation result is restored through a decoder to ensure the accuracy of the segmentation. This method is not only innovative in theory, but also shows significant performance improvement in practical applications, providing a new technical solution for the field of medical image segmentation and having important clinical application value.
[0011] Compared with the prior art, the medical image segmentation method based on spatial perception and frequency domain information has the following beneficial effects: (1) Through the multi-scale decomposition ability of frequency domain transformation and a learnable frequency domain filtering module for low-frequency components, it efficiently captures features of different frequencies in medical images and effectively suppresses noise in low-frequency information. (2) Through the multi-head perception visual state space module, the information extracted from different scales is used to enhance the understanding ability of the model, which can effectively capture fine-grained (such as micro-lesions) and coarse-grained (such as the overall anatomical structure) features, ensuring that the network has sufficient feature expression ability at different resolutions. (3) It effectively shows the important impact of detail features on the segmentation performance of the model, and detail features also have an important impact on the performance of semantic segmentation tasks. Description of the Drawings
[0012] Figure 1It is a schematic flowchart of a medical image segmentation method based on spatial perception and frequency domain information provided by the first embodiment of the present invention; Figure 2 It is an overall model structure diagram of a medical image segmentation method based on spatial perception and frequency domain information provided by an embodiment of the present invention; Figure 3 It is a structure diagram of the FLICEB module provided by an embodiment of the present invention; Figure 4 It is a structure diagram of the MPVSS module provided by an embodiment of the present invention; Figure 5 It is a structure diagram of the CFA module provided by an embodiment of the present invention; Figure 6 It is a module schematic diagram of a medical image segmentation device based on spatial perception and frequency domain information provided by the second embodiment of the present invention. Detailed implementation manners
[0013] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0014] Refer to Figure 1 、 Figure 2 As shown, the first embodiment of the present invention discloses a medical image segmentation method based on spatial perception and frequency domain information, which can be executed by a medical image segmentation device based on spatial perception and frequency domain information (hereinafter referred to as the segmentation device), and particularly, by one or more processors in the segmentation device to implement the following method: S1. Obtain the medical image to be processed, perform data preprocessing on the medical image, and perform quality evaluation on the preprocessed medical image, calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluate the contrast and clarity of the image, detect artifacts and noise levels in the image. If the signal-to-noise ratio is lower than the preset threshold, perform adaptive histogram equalization, and selectively apply gamma correction to perform local enhancement processing on the detected artifact regions according to the image contrast situation; Specifically, step S1 includes: calculating the signal-to-noise ratio and peak signal-to-noise ratio of the image to quantitatively evaluate the image quality. Among them, the signal-to-noise ratio measures the image quality by calculating the ratio of the average power of the image signal to the average power of the noise, and the peak signal-to-noise ratio evaluates the image fidelity through the logarithmic relationship between the possible maximum pixel value of the image and the mean square error; Adopt the root mean square contrast method to evaluate the global contrast of the image, and use the Laplacian operator to calculate the second derivative of the image to evaluate the clarity; A noise detection mechanism based on high-frequency component analysis is designed to address artifacts and noise in images. The mechanism combines Histogram of Oriented Gradients (HOG) features and Radon transform to identify stripe artifacts and motion artifacts present; When the signal-to-noise ratio (SNR) is detected to be lower than a preset threshold, contrast-limited adaptive histogram equalization (CLAHE) is used for enhancement, where the clip limit parameter is dynamically adjusted according to the SNR value; For regions with insufficient contrast, adaptive gamma correction based on the statistical characteristics of the image is introduced, and the gamma value is automatically adjusted by analyzing the average brightness of the image; For the detected artifact regions, local enhancement processing based on guided filtering is adopted, and Poisson fusion is used to ensure natural transition between the processed region and the surrounding image.
[0015] In this embodiment, an image quality assessment and adaptive enhancement mechanism is performed on the preprocessed image. This mechanism first quantitatively assesses the image quality by calculating the signal-to-noise ratio (SNR) and peak signal-to-noise ratio (PSNR) of the image. Among them, SNR measures the image quality by calculating the ratio of the average power of the image signal to the average power of the noise, and PSNR evaluates the image fidelity through the logarithmic relationship between the maximum possible pixel value of the image and the mean square error. At the same time, the RMS (root mean square) contrast method is used to evaluate the global contrast of the image, and the Laplacian operator is used to calculate the second derivative of the image to evaluate the sharpness. For possible artifacts and noise in the image, this method designs a noise detection mechanism based on high-frequency component analysis, which combines Histogram of Oriented Gradients (HOG) features and Radon transform to identify possible stripe artifacts and motion artifacts. Based on these evaluation metrics, this method further implements an adaptive image enhancement strategy: when the detected SNR is lower than the preset threshold (usually set to 15 dB), contrast-limited adaptive histogram equalization (CLAHE) is used for enhancement, where the clip limit parameter is dynamically adjusted according to the SNR value; for regions with insufficient contrast, adaptive gamma correction based on the statistical characteristics of the image is introduced, and the gamma value is automatically adjusted by analyzing the average brightness of the image; for the detected artifact regions, local enhancement processing based on guided filtering is adopted, and Poisson fusion is used to ensure natural transition between the processed region and the surrounding image.
[0016] S2: Call the pre-trained FDMUNet network model, and use the context feature extraction module that focuses on low-frequency information to perform feature extraction processing on the preprocessed medical image to obtain an output feature map; Specifically, step S2 includes: for the input preprocessed medical feature map , H, W, and C respectively represent the height, width, and number of channels of the image. The context feature extraction module of the FDMUNet network model that focuses on low-frequency information performs feature decomposition processing on the medical feature map X at preset different frequency scales to capture different frequency features in the medical image. The formula is as follows: , , T , , where is the frequency domain transformation operation, is the low-frequency component of the medical image, are all high-frequency components of the medical image, is the output feature map obtained through the noise filtering mechanism, is the sigmod function, is the max pooling operation, is the average pooling operation, represents concatenation by the number of channels, is the 1x1 convolution operation, is the batch normalization operation, is the ReLU activation function, is the feature information after the high-frequency information and the low-frequency information are concatenated by channels, is the extracted feature map; After three consecutive feature extractions, the output feature map of is obtained. The feature maps after three downsampling feature extractions are represented as , where , , ; , , .
[0017] Preferably, the data preprocessing includes: size scaling, random cropping, horizontal mirroring, and random cropping.
[0018] In this embodiment, first, the medical images to be processed are obtained. These images can be the original images from CT scans, MRI imaging, or other medical imaging devices in the hospital. To improve the generalization ability and segmentation accuracy of the model, data preprocessing is performed on these medical images. The specific steps of data preprocessing include resizing the images to meet the input requirements of the model; performing random cropping to increase data diversity; implementing horizontal mirroring operations to further enrich the dataset; and performing random cropping again to ensure that the model can learn different parts of the images. These preprocessing steps not only enhance the robustness of the model but also improve its adaptability to different image sizes and orientations.
[0019] Since most existing segmentation schemes use CNN encoders, which ignore the complementary information between different-scale features and thus hinder the learning of global information and detailed features, based on this, considering the complementarity between frequency-domain features and spatial features, after preprocessing, the pre-trained FDMUNet network model is called to perform feature extraction on the preprocessed medical images. The core of the FDMUNet network model lies in its context feature extraction module for focusing on low-frequency information (Frequency-domain Low-frequency Information Context Extraction Block, FLICEB). This module can be used to extract key information from the input data, thus helping to improve the performance of the model. It utilizes the multi-scale analysis ability of frequency-domain transformation to perform feature decomposition on medical images at different preset frequency scales. Specifically, this module uses frequency-domain transformation operations to decompose the input feature map into low-frequency components and high-frequency components. The low-frequency components mainly contain the global information of the image, while the high-frequency components contain detailed information. Through this decomposition, the model can capture various features in medical images more comprehensively. Based on the processing process of the context feature extraction module for focusing on low-frequency information, not only can the key features in medical images be effectively extracted, but also the attention of the model to the lesion area can be improved through the noise filtering mechanism.
[0020] During the feature decomposition process, special attention is paid to the processing of low-frequency components. Since frequency-domain transformation may introduce background noise and low-frequency components of regions irrelevant to the target, to solve this problem, a learnable noise filtering mechanism is introduced. This mechanism can automatically identify and suppress noise or irrelevant background information in the low-frequency part, enabling the model to focus more on regions with higher semantic significance, such as lesion regions. This process not only effectively suppresses noise and irrelevant background information in low-frequency information, improves the segmentation ability of the model for key regions, but also enhances the accuracy and reliability of the segmentation results, enabling the model to focus on regions with higher semantic significance (such as lesion regions), thereby enhancing the segmentation ability of the model in key regions.
[0021] Among them, the max-pooling and average-pooling operations involved in the noise filtering mechanism formula are used to extract the spatial features of low-frequency information. The noise filtering mechanism will weight each spatial position of the low-frequency information to emphasize important regions and suppress irrelevant regions.
[0022] S3. Use the multi-head perceptual visual state space module in the FDMUNet network model to extract and process the output feature map, and extract the global information features of the output feature map at different scales; Specifically, step S3 includes: using the multi-head perceptual visual state space module in the bottleneck layer of the FDMUNet network model to process the output feature map Extract and process to obtain the global information features of the output feature maps at different scales. The formula for the multi-head perception visual state space module is: , where is a 3x3 depthwise separable convolution operation, is the VSS module, is layer normalization, is the multi-scale adaptive feature fusion module, are the extracted fine-grained features and coarse-grained features; Among them, the multi-scale adaptive feature fusion module divides the output feature map into three branches at different scales for processing, and fuses the outputs of the three branches. The formula for the multi-scale adaptive feature fusion module is: , , , , where ConCat is concatenation by channel to perform feature fusion, is the output of the first branch, is the output of the second branch, is the output of the third branch, is 1x1 average adaptive pooling, is 3x3 average adaptive pooling, is 5x5 average adaptive pooling, is the result of feature fusion; After four consecutive multi-head perception visual state space modules extract and process the output feature map , obtain the feature map of, where the output feature maps of the four bottleneck layers are expressed as , , , . Among them, , , .
[0023] In this embodiment, in order to effectively process multi-scale features in dermatological images, a multi-head perception visual state space module (MPVSS) is proposed in the bottleneck layer. This module uses a multi-scale feature fusion mechanism to enhance the model's ability to understand small-sized lesions and complex structures. The module extracts global information of the image at different scales, effectively captures the multi-scale features of dermatological diseases through a multi-scale adaptive feature fusion module, and finally uses the visual state space module to effectively model the global context information of the image, thereby enhancing the model's perception ability of different-scale lesion regions during the segmentation process.
[0024] Specifically, the multi-head perception visual state space module is integrated into the bottleneck layer of the FDMUNet network model. The core of this module lies in its ability to extract global information features of the output feature map at different scales, which is crucial for understanding complex structures in medical images. For example, in dermatological images, lesion areas may exhibit different sizes, from tiny lesion points to larger lesion regions. Through the multi-head perception visual state space module, the model can simultaneously capture these features at different scales, thus more comprehensively understanding the image content. Simply put, the model enhances its understanding ability through information extracted from different scales. This process ensures the stability of features through layer normalization (LN), models the visual state space using the VSS module, extracts local features through 3x3 depthwise separable convolution operations, and finally fuses features at different scales through a multi-scale adaptive feature fusion module (MAF) to obtain a feature representation containing global information. It can effectively capture fine-grained (such as tiny lesions) and coarse-grained (such as overall anatomical structures) features, ensuring that the network has sufficient feature representation at different resolutions. This multi-scale adaptive feature fusion is particularly suitable for complex scenarios in medical imaging.
[0025] In addition, the multi-scale adaptive feature fusion module is an important part of the multi-head perceptual vision state space module. This module divides the input feature map into three branches for processing at different scales, and each branch is responsible for extracting feature information at a specific scale. The first branch expands the number of channels through a 1x1 convolutional layer, then performs normalization operations using a batch normalization layer and a ReLU activation layer, and finally captures global information through 1x1 adaptive average pooling while reducing the computational amount. The second branch processes the input features through 3x3 adaptive pooling, reduces the spatial dimension and extracts local context information, then passes through batch normalization and the ReLU activation function, and finally obtains the output features through a 1x1 convolution. The third branch processes the feature map through a 5x5 adaptive pooling window, extracts context information in a larger range, then passes through batch normalization and the ReLU activation function, and finally obtains the output features through a 1x1 convolution. The outputs of the three branches are respectively denoted as , and , and they obtain the final output feature map through feature fusion, which can effectively extract features from different scales, enhancing the multi-scale perception ability and information integration ability of the network when processing complex medical images (such as skin disease images).
[0026] Through this multi-scale adaptive feature fusion mechanism, the model can effectively capture feature information from different scales, enhancing the multi-scale perception ability and information integration ability of the network when processing complex medical images. For example, when processing skin disease images containing tiny lesions, the model can simultaneously focus on the detailed features of the lesion area and the macroscopic features of the overall anatomical structure, thereby performing segmentation more accurately. This multi-scale feature fusion mechanism not only improves the segmentation accuracy of the model for lesions at different scales but also ensures that the network has sufficient feature expression ability at different resolutions, which is very important for medical image segmentation tasks. In addition, the design of the multi-head perceptual vision state space module also considers computational efficiency. By combining operations such as depthwise separable convolution and adaptive pooling, this module can effectively extract global information features while maintaining efficient computation.
[0027] S4. Use the focused context attention module in the FDMUNet network model to perform boundary enhancement processing on the feature map to obtain enhanced features; Specifically, step S4 includes: during the skip connection process, perform boundary enhancement processing on the feature map , and respectively according to the focused context attention module in the FDMUNet network model, and adopt the method of combining average pooling and 1x1 convolution to obtain local region features; The features are expanded using two depthwise separable large kernel convolutions, and a 1x1 convolution and a sigmoid activation function are used to perform convolution activation on the processed features to obtain enhanced features. The formula is as follows: , where is the enhanced feature, is a 11x1 depthwise separable convolution operation, is a 1x11 depthwise separable convolution operation; Similarly, the features enhanced by the three-layer focused context attention module are respectively expressed as, , , , where , .
[0028] In this embodiment, in the traditional UNet, the skip connection simply cascades the corresponding feature maps of the encoder into the decoder feature maps, lacking the modeling of the context dependence between distant pixels. In medical images, the distribution of diseases is often local, but the context information of the lesion area (such as the entire organ structure) is crucial for segmentation. The traditional skip connection is often insufficient to capture this distant context information. To overcome the above defects, this method introduces a Contextual Focused Attention (CFA) module in the skip connection. While transmitting the low-level features, the CFA module can also, guided by the context information, strengthen the network's attention to the global information and improve the performance of the model in the segmentation task. This module is a simple and effective attention mechanism scheme for mining multi-granularity information around the edges at each stage in the skip connection.
[0029] The focused context attention module is integrated into the skip connection process of the FDMUNet network model. The skip connection is a key part of the U-shaped network architecture, which directly transmits the low-level features in the encoder to the decoder to retain the detailed information of the image. In this method, a focused context attention module is introduced to strengthen the network's attention to the global information, guided by the context information, and improve the performance of the model in the segmentation task.
[0030] Specifically, in this embodiment, during the boundary enhancement process, the local region features are first obtained by combining average pooling and 1x1 convolution. The average pooling operation can downsample the feature map and extract the local information in the feature map, while the 1x1 convolution is used to adjust the number of channels of the feature map to make it more suitable for subsequent processing. This combination method can not only effectively extract the features of the local region but also reduce the computational amount and improve the efficiency of the model.
[0031] Subsequently, the features are expanded using two depthwise separable large kernel convolutions. Depthwise separable convolution is an efficient convolution operation that decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution, thus greatly reducing the computational complexity. In this method, 11x1 depthwise separable convolution and 1x11 depthwise separable convolution are adopted. These two operations expand the feature map in the horizontal and vertical directions respectively, thereby enlarging the receptive field of the model and enabling the model to better capture the context information in the feature map.
[0032] Finally, a 1x1 convolution and a sigmoid activation function are used to perform convolution activation processing on the processed features to obtain enhanced features. The 1x1 convolution is used to adjust the number of channels of the feature map to match the subsequent network structure, while the sigmoid activation function maps each element in the feature map to the range between 0 and 1, thereby obtaining a probability map representing the probability that each pixel belongs to the lesion area. In this way, the model can more accurately identify the boundary of the lesion area, thereby improving the accuracy of segmentation.
[0033] Through the boundary enhancement processing of the focus context attention module, this method can effectively enhance the boundary information in the feature map, enabling the model to more accurately identify the edge of the lesion area. This improvement not only improves the accuracy of segmentation but also enhances the adaptability of the model to complex medical images. For example, when processing dermatological images containing small lesions, the model can more clearly identify the boundary of the lesion area and thus perform more accurate segmentation. In addition, by adopting depthwise separable convolution and sigmoid activation function, this module can effectively extract and enhance boundary features while maintaining efficient computation, further improving the performance of the model.
[0034] S5. Use the decoder to perform recovery preprocessing on the global information features, and fuse the processed global information features with the enhanced features to obtain the segmentation result. Repeat this operation until the preset number of times is reached to generate a high-resolution feature map.
[0035] Specifically, step S5 includes: using the decoder to perform channel-wise fusion on the global information output by the multi-head perception visual state space module and the enhanced features obtained by the focus context attention module, and performing bilinear interpolation upsampling processing to obtain , so as to restore the control size of the image; Adopt convolution operation to perform feature dimensionality reduction and batch normalization processing on the processed global information features, and retain the feature information through the ReLU activation function; Perform addition fusion processing on the enhanced features obtained from the skip connection and the global information features after retention processing to obtain the features restored by the decoder. The formula is: , where, is an upsampling operation, is to obtain the features recovered by the decoder; After three consecutive feature fusions and upsampling processes, the final high-resolution feature map G is obtained , where the features recovered by the decoder are respectively , , where, , .
[0036] In this embodiment, the decoder is first used to perform bilinear interpolation upsampling on the global information features. Bilinear interpolation upsampling is a commonly used image magnification technique that generates new pixel values by calculating the weighted average of adjacent pixels, thereby restoring the spatial resolution of the image. In this embodiment, this operation is used to restore the low-resolution global information features processed by the encoder to a resolution close to that of the original image for fusion with the enhanced features. This process can not only effectively restore the spatial structure of the image but also provide a basis for subsequent feature fusion. Subsequently, convolution operations are used to perform feature dimensionality reduction and batch normalization on the processed global information features. Feature dimensionality reduction reduces the number of channels of the feature map through convolution operations, thereby reducing the computational complexity and extracting a more compact feature representation. Batch normalization is a commonly used regularization technique that standardizes the activation values of each feature map, making the network training more stable and accelerating convergence. Through convolution operations and batch normalization processing, the expression of the global information features can be further optimized to make them more suitable for fusion with the enhanced features.
[0037] To retain feature information, a non-linear transformation is performed on the processed global information features through the ReLU activation function. The ReLU activation function can set negative values in the feature map to zero while retaining positive values, thereby enhancing the sparsity of the features and improving the model's ability to distinguish features. This operation can not only effectively retain the important information in the global information features but also provide a clearer feature representation for subsequent feature fusion. Finally, addition fusion processing is performed on the enhanced features and the global information features after retention processing to obtain the final segmentation result. Addition fusion is a simple and effective feature fusion method that combines the boundary information in the enhanced features with the background information in the global information features by adding the two feature maps pixel by pixel. This process can not only make full use of the boundary details in the enhanced features but also combine the context information in the global information features, thereby obtaining a more accurate segmentation result. After 4 consecutive upsampling operations, the model finally generates a high-resolution feature map. This process ensures that the model can combine global background information and local details, thereby outputting a more accurate segmentation result. Through continuous fusion and learning, the decoder can effectively restore the spatial structure of the image and complete complex medical image segmentation tasks.
[0038] Through this step, the method can effectively fuse the global information features and the enhanced features, ensuring the accuracy and integrity of the segmentation result. The bilinear interpolation upsampling process can restore the spatial resolution of the image, the convolution operation and the batch normalization process can optimize the feature representation, the ReLU activation function can retain important feature information, and the addition fusion process can make full use of the boundary details in the enhanced features and the context information in the global information features. The combination of these steps not only improves the accuracy of segmentation, but also enhances the adaptability of the model to complex medical images. For example, when processing dermatological images containing tiny lesions, the model can more clearly identify the boundaries of the lesion areas and perform accurate segmentation in combination with the global background information. This fusion mechanism of global and local information significantly improves the model's ability to identify lesion areas, thus providing a new technical solution for the field of medical image segmentation and having important clinical application value.
[0039] Preferably, before calling the pre-trained FDMUNet network model to perform feature extraction processing on the pre-processed medical image, it further includes: Obtaining training data, performing data pre-processing on the training data to obtain a medical segmentation task data set, where the medical segmentation task data set includes training set images and training segmentation ground truth images corresponding to the training set images; The data pre-processing includes: size scaling, random cropping, horizontal mirroring, random cropping; Training the FDMUNet network model according to the medical segmentation task data set to obtain a trained FDMUNet network model.
[0040] Specifically, in this embodiment, first, medical image data for training is obtained. These data usually come from medical image libraries in hospitals or research institutions and include various types of medical images, such as CT scans, MRI imaging, etc. These images and their corresponding segmentation ground truth images constitute the medical segmentation task data set, where the training set images are used for training the model, and the training segmentation ground truth images are used as supervision signals to guide the model to learn the correct segmentation method.
[0041] Before model training, data preprocessing of the training data is a crucial step. The specific operations of data preprocessing include resizing the images to ensure that all images have a unified size for easy model processing; performing random cropping to increase data diversity so that the model can learn different parts of the images; implementing horizontal mirroring to further enrich the dataset and improve the generalization ability of the model; and finally performing random cropping again to enhance the model's adaptability to different image sizes and orientations. These preprocessing steps can not only improve the robustness of the model but also effectively prevent overfitting, enabling the model to maintain good performance when facing new and unseen images.
[0042] After completing data preprocessing, the preprocessed medical segmentation task dataset is used to train the FDMUNet network model. During the training process, the model gradually adjusts its own parameters by learning the mapping relationship between the training set images and the corresponding training segmentation ground truth images to minimize the difference between the predicted segmentation result and the true segmentation result. This training process usually involves a large number of iterative optimizations, as well as continuous evaluation and adjustment of the model performance. Through this carefully designed training strategy, the FDMUNet network model can learn rich feature representations and thus perform well in subsequent feature extraction and segmentation tasks. For example, when dealing with medical images containing complex lesions, the pre-trained model can more accurately identify the boundaries of the lesion areas and perform precise segmentation by combining global background information. This mechanism of fusing global and local information significantly improves the model's ability to identify lesion areas, thus providing a new technical solution for the field of medical image segmentation and having important clinical application value.
[0043] In summary, the medical image segmentation method based on spatial perception and frequency domain information aims to solve the key problems existing in medical image segmentation in the prior art, such as loss of edge details, omission of lesion areas, and insufficient extraction of multi-scale lesion features. This method significantly improves the accuracy and robustness of medical image segmentation and is applicable to a variety of medical image segmentation tasks.
[0044] Specifically, during the data preprocessing stage, operations such as resizing, random cropping, and horizontal mirroring are performed on medical images to enhance the generalization ability of the model and its adaptability to different image sizes. These preprocessing steps not only improve the robustness of the model but also effectively prevent overfitting, enabling the model to maintain good performance when faced with new and unseen images. In the feature extraction stage, an improved context feature extraction module for focusing on low-frequency information (FLICEB) is introduced. Using the multi-scale analysis ability of frequency domain transformation, the input image is decomposed into low-frequency and high-frequency components. The low-frequency components contain the global information of the image, while the high-frequency components contain the detailed information. Through a learnable noise filtering mechanism, this module can automatically identify and suppress noise or irrelevant background information in the low-frequency part, enabling the model to focus more precisely on the lesion area. This innovative design significantly enhances the model's ability to capture global information and detailed features, improving the accuracy of segmentation.
[0045] Furthermore, a multi-head perceptual visual state space module (MPVSS) is designed in the bottleneck layer. Through a multi-scale adaptive feature fusion mechanism, global information features of the image are extracted from different scales. This module combines the local perception advantages of CNN and the efficient global modeling ability of the state space model, and can effectively capture multi-scale features from tiny lesions to the overall anatomical structure. This multi-scale feature fusion mechanism not only improves the segmentation accuracy of the model for lesions of different scales but also ensures that the network has sufficient feature expression ability at different resolutions, especially performing well when dealing with small-sized lesions and complex structures. In the skip connection, a focused context attention module (CFA) is introduced to perform boundary enhancement processing on the feature map. Local region features are obtained by combining average pooling and 1x1 convolution, then the receptive field of the model is expanded using depthwise separable large kernels convolution, and finally enhanced features are obtained through 1x1 convolution and sigmoid activation function. The design of this module not only strengthens the model's attention to global information but also significantly improves the ability to recognize the edges of the lesion area, further enhancing the accuracy of segmentation.
[0046] Finally, the global information features are upsampled by bilinear interpolation through the decoder to restore the spatial resolution of the image, and added to the enhanced features to obtain the final segmentation result. After continuous upsampling operations, the model finally generates a high-resolution feature map. This process not only combines the global background information and local details but also ensures the accuracy and integrity of the segmentation result through feature fusion.
[0047] Generally speaking, the medical image segmentation method based on spatial perception and frequency domain information significantly improves the accuracy and robustness of medical image segmentation through the complementary fusion of frequency domain information and spatial information, the adaptive optimization of multi-scale features, and the effective utilization of context information. This method is not only innovative in theory but also shows significant performance improvement in practical applications, providing a new technical solution for the field of medical image segmentation and having important clinical application value.
[0048] Please refer to Figure 6 , a second embodiment of the present invention provides a medical image segmentation device based on spatial perception and frequency domain information, which includes: A preprocessing unit 101, configured to obtain a medical image to be processed, perform data preprocessing on the medical image, and perform quality assessment on the preprocessed medical image, calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluate the contrast and clarity of the image, detect artifacts and noise levels in the image, and if the signal-to-noise ratio is lower than a preset threshold, perform adaptive histogram equalization, and selectively apply gamma correction according to the image contrast situation to perform local enhancement processing on the detected artifact regions; A first extraction unit 102, configured to call a pre-trained FDMUNet network model and use a context feature extraction module that focuses on low-frequency information to perform feature extraction processing on the preprocessed medical image to obtain an output feature map; A second extraction unit 103, configured to use a multi-head perception visual state space module in the FDMUNet network model to perform extraction processing on the output feature map to extract global information features of the output feature map at different scales; An enhancement unit 104, configured to use a focused context attention module in the FDMUNet network model to perform boundary enhancement processing on the feature map to obtain enhanced features; A fusion unit 105, configured to use a decoder to perform restoration preprocessing on the global information features, and fuse the processed global information features with the enhanced features to obtain a segmentation result, and repeat this operation until a preset number of times is reached to generate a high-resolution feature map.
[0049] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A medical image segmentation method based on spatial perception and frequency domain information, characterized in that Including: Obtain the medical image to be processed, perform data preprocessing on the medical image, and conduct quality assessment on the preprocessed medical image. Calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluate the contrast and clarity of the image, detect artifacts and noise levels in the image. If the signal-to-noise ratio is lower than the preset threshold, perform adaptive histogram equalization. According to the image contrast situation, selectively apply gamma correction to perform local enhancement processing on the detected artifact regions; Call the pre-trained FDMUNet network model, and use the context feature extraction module that focuses on low-frequency information to perform feature extraction processing on the preprocessed medical image to obtain an output feature map; Adopt the multi-head perception visual state space module in the FDMUNet network model to perform extraction processing on the output feature map, and extract the global information features of the output feature maps at different scales; Use the focused context attention module in the FDMUNet network model to perform boundary enhancement processing on the feature map to obtain enhanced features; Use the decoder to perform restoration preprocessing on the global information features, and fuse the processed global information features with the enhanced features to obtain a segmentation result. Repeat this operation until the preset number of times is reached to generate a high-resolution feature map.
2. The medical image segmentation method based on spatial perception and frequency domain information according to claim 1, wherein The data preprocessing includes: size scaling, random cropping, horizontal mirroring, random cropping.
3. The medical image segmentation method based on spatial perception and frequency domain information according to claim 1, wherein Conduct quality assessment on the preprocessed medical image, calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluate the contrast and clarity of the image, detect artifacts and noise levels in the image. If the signal-to-noise ratio is lower than the preset threshold, perform adaptive histogram equalization. According to the image contrast situation, selectively apply gamma correction to perform local enhancement processing on the detected artifact regions. Specifically: Calculate the signal-to-noise ratio and peak signal-to-noise ratio of the image to quantitatively evaluate the image quality. Among them, the signal-to-noise ratio measures the image quality by calculating the ratio of the average power of the image signal to the average power of the noise, and the peak signal-to-noise ratio evaluates the image fidelity through the logarithmic relationship between the maximum possible pixel value of the image and the mean square error; Adopt the root mean square contrast method to evaluate the global contrast of the image, and use the Laplacian operator to calculate the second derivative of the image to evaluate the clarity; For the artifacts and noise existing in the image, design a noise detection mechanism based on high-frequency component analysis, and combine the histogram of oriented gradients feature and Radon transform to identify the existing stripe artifacts and motion artifacts; When it is detected that the signal-to-noise ratio is lower than the preset threshold, use limited contrast adaptive histogram equalization for enhancement, where the clipping limit parameter is dynamically adjusted according to the signal-to-noise ratio value; For regions with insufficient contrast, introduce adaptive gamma correction based on the statistical characteristics of the image, and automatically adjust the gamma value by analyzing the average brightness of the image; For the detected artifact regions, adopt local enhancement processing based on guided filtering, and ensure the natural transition between the processed region and the surrounding image through Poisson fusion.
4. The medical image segmentation method based on spatial perception and frequency domain information according to claim 1, wherein Call the pre-trained FDMUNet network model to perform feature extraction processing on the preprocessed medical image to obtain an output feature map. Specifically: For the preprocessed medical feature map of the input , where H, W, and C represent the height, width, and number of channels of the image respectively. The context feature extraction module of the FDMUNet network model that focuses on low-frequency information is used to perform feature decomposition on the medical feature map X at different preset frequency scales to capture different frequency features in the medical image. The formula is as follows: , , T , , where is the frequency domain transformation operation, is the low-frequency component of the medical image, are all high-frequency components of the medical image, is the output feature map obtained through the noise filtering mechanism, is the sigmod function, is the max pooling operation, is the average pooling operation, represents concatenation by the number of channels, is the 1x1 convolution operation, is the batch normalization operation, is the ReLU activation function, is the feature information after the high-frequency information and the low-frequency information are concatenated by channels, is the extracted feature map; After three consecutive feature extractions, the output feature map of is obtained. The feature maps after three downsampling feature extractions are denoted as , where , , ; , , 。 5. The medical image segmentation method based on spatial perception and frequency domain information according to claim 3, characterized in that The multi-head perceptual visual state space module in the FDMUNet network model is used to extract and process the output feature map, and the global information features of the output feature map at different scales are extracted, specifically: The multi-head perceptual visual state space module in the bottleneck layer of the FDMUNet network model is used to process the output feature map for extraction, and the global information features of the output feature map at different scales are extracted. The formula of the multi-head perceptual visual state space module is: , where is a 3x3 depthwise separable convolution operation, is the VSS module, is layer normalization, is the multi-scale adaptive feature fusion module, are the fine-grained features and coarse-grained features extracted; Among them, the multi-scale adaptive feature fusion module processes the output feature map by dividing it into three branches with different scales, and fuses the outputs of the three branches. The formula of the multi-scale adaptive feature fusion module is: , , , , where ConCat is concatenation by channel to perform feature fusion, is the output of the first branch, is the output of the second branch, is the output of the third branch, is 1x1 average adaptive pooling, is 3x3 average adaptive pooling, is 5x5 average adaptive pooling, is the result of feature fusion; After four consecutive multi-head perception visual state space modules process the output feature map for extraction, the feature map of is obtained, where the output feature maps of the four bottleneck layers are denoted as , , , . Among them, , , .
6. The medical image segmentation method based on spatial perception and frequency domain information according to claim 4, wherein The focus context attention module in the FDMUNet network model is used to perform boundary enhancement processing on the feature map to obtain enhanced features, specifically: During the skip connection process, the feature map is focused on by the focus context attention module in the FDMUNet network model. , and Perform boundary enhancement processing respectively, and use average pooling and 1x1 convolution to obtain local area features; The features are expanded using two depthwise separable large kernel convolutions, and the processed features are convolved and activated using 1x1 convolution and sigmoid activation function to obtain enhanced features. The formula is as follows: , where is the enhanced feature, is the 11x1 depthwise separable convolution operation, is the 1x11 depthwise separable convolution operation; Similarly, the three layers of features enhanced by the focused context attention module are respectively expressed as, , , , where, , .
7. The medical image segmentation method based on spatial perception and frequency domain information according to claim 5, characterized in that The decoder is used to perform restoration preprocessing on the global information features, and the processed global information features are fused with the enhanced features to obtain a segmentation result. This operation is repeated until a preset number of times is reached to generate a high-resolution feature map, specifically: Use a decoder to process the global information output by the multi-head perceptual visual state space module and the enhanced features obtained by the focused context attention module are fused channel by channel and then upsampled by bilinear interpolation to obtain , so as to restore the control size of the image; Convolution operations are used to perform feature dimensionality reduction and batch normalization on the processed global information features, and the ReLU activation function is used to retain the feature information; Perform an addition fusion process on the enhanced features obtained from the skip connection and the globally-informed features after retention processing to obtain the features restored by the decoder. The formula is as follows: , where is an upsampling operation, is to obtain the features restored by the decoder; After three consecutive feature fusions and upsampling processes, the final high-resolution feature map G is obtained , where the features recovered by the decoder are respectively , , where , .
8. The medical image segmentation method based on spatial perception and frequency domain information according to claim 1, wherein Before calling the pre-trained FDMUNet network model to perform feature extraction processing on the preprocessed medical image, it also includes: Training data is obtained, and the training data is preprocessed to obtain a medical segmentation task data set. Among them, the medical segmentation task data set includes training set images and training segmentation ground truth images corresponding to the training set images; The data preprocessing includes: size scaling, random cropping, horizontal mirroring, random cropping; The FDMUNet network model is trained according to the medical segmentation task data set to obtain a trained FDMUNet network model.
9. A medical image segmentation device based on spatial perception and frequency domain information, characterized in that, It includes: A preprocessing unit for obtaining a medical image to be processed, preprocessing the medical image, and performing quality assessment on the preprocessed medical image, calculating the signal-to-noise ratio and peak signal-to-noise ratio of the image, evaluating the contrast and clarity of the image, detecting artifacts and noise levels in the image. If the signal-to-noise ratio is lower than a preset threshold, perform adaptive histogram equalization, and selectively apply gamma correction to perform local enhancement processing on the detected artifact regions according to the image contrast situation; A first extraction unit for calling the pre-trained FDMUNet network model and using the context feature extraction module that focuses on low-frequency information to perform feature extraction processing on the preprocessed medical image to obtain an output feature map; A second extraction unit for using the multi-head perceptual visual state space module in the FDMUNet network model to extract and process the output feature map to extract the global information features of the output feature map at different scales; An enhancement unit for using the focus context attention module in the FDMUNet network model to perform boundary enhancement processing on the feature map to obtain enhanced features; A fusion unit for using the decoder to perform restoration preprocessing on the global information features, and fusing the processed global information features with the enhanced features to obtain a segmentation result. This operation is repeated until a preset number of times is reached to generate a high-resolution feature map.
Citation Information
Patent Citations
Medical image segmentation method based on boundary perception and attention mechanism
CN117078930A
Semi-supervised polyp segmentation method based on dynamic multi-scale perception
CN118552575A
Monitoring video target detection method based on multi-frame high-frequency difference enhancement
CN119540812A
Method for breast screening in fused mammography
US20190350549A1
Cited By
Skin disease image preprocessing method and system, electronic equipment and storage medium
CN120563508A
Preprocessing method, system, electronic device and storage medium for skin disease images
CN120563508B
Abdominal tuberculosis image processing method
CN120612251A
Colorectal cancer pathological image segmentation method and device based on frequency domain characteristics and readable storage medium thereof
CN120635104A
Medical image computer-aided analysis method based on deep learning
CN120807509A