Medical image fuzzy boundary segmentation system and method based on COD frequency domain guidance
By adopting a COD frequency domain guidance method in medical image segmentation and using the FEA-Net model that works in multiple modules in a collaborative manner, the problem of difficulty in capturing subtle boundary changes in medical images in the prior art is solved, and higher segmentation accuracy and boundary positioning accuracy are achieved.
Patent Information
- Application Number
- CN202510203032.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Existing medical image segmentation methods are difficult to effectively capture subtle boundary changes in medical images, especially in the gradient areas at the junction of tissues, and over-reliance on manual parameter adjustments reduces the value of clinical application.
The medical image blur boundary segmentation method based on COD frequency domain guidance is adopted, and data processing is performed through the FEA-Net model, combined with the pre-trained PVTv2 model, and the semantic edge enhancement module, frequency domain feature guidance enhancement module, adaptive edge feature module and context attention aggregation module are used to achieve accurate segmentation of medical image boundaries.
It significantly improves the segmentation accuracy and boundary positioning accuracy of the blurred boundary areas in medical images, enhances the interpretability and scalability of the model, and provides more reliable technical support for clinical diagnosis.
Smart Images

Figure CN120125599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image analysis, and particularly to a medical image fuzzy boundary segmentation system and method guided by COD frequency domain. Background Art
[0002] In recent years, with the rapid development of deep learning and computer vision technologies, significant progress has been made in the field of medical image analysis. Precise boundary segmentation in medical images is crucial for disease diagnosis, condition assessment, and treatment plan formulation. However, due to the complex morphology, irregular boundaries, and high similarity to surrounding tissues of pathological tissues in medical images, traditional image segmentation methods are difficult to achieve ideal results. This boundary blur problem is prevalent in various medical imaging modalities: in dermoscopic image analysis, the lesion area usually presents a highly variable appearance and a gradual transition to healthy tissues; in colonoscopy, complex tissue textures and weak intensity changes at the lesion boundaries pose challenges for accurate detection; in glomerular pathological images, the subtle contrast differences between the lesion tissue and healthy tissue make boundary delineation extremely difficult.
[0003] Currently, deep learning models, especially those based on convolutional neural networks (CNNs), still face many problems when addressing these challenges: traditional spatial domain feature extraction methods have limitations in dealing with fuzzy boundaries and are difficult to effectively capture the subtle boundary changes in medical images, especially in the gradual transition regions at tissue junctions; existing boundary detection techniques rely too much on manual parameter adjustment, significantly reducing their clinical application value. Recent research has attempted to solve this problem through various method innovations: Although the Simple Generic Method introduces a boundary-aware loss function, it is still limited to spatial domain optimization and lacks a dedicated boundary feature extraction structure; AEC-Net integrates the spatial domain attention mechanism and edge constraints through a dual-branch structure, but its simplified edge branch has limited performance in dealing with complex medical image boundaries; SharpContour improves performance through contour optimization and bottom-up feature fusion, but it is still insufficient in dealing with highly fuzzy medical image boundaries. At the same time, significant achievements have been made in the recognition of hidden targets in natural scenes with COD. For example, DGNet and ERRNet have shown preliminary results in medical applications such as polyp segmentation and pneumonia image segmentation by effectively fusing context and edge information. However, due to the lack of consideration of the specific frequency domain features of medical images, when these methods are directly applied to medical images with fuzzy boundaries, their effects are significantly different from those in natural scenes. Quantitative analysis shows that these technologies require specific optimizations for the medical field to achieve ideal results.
[0004] Therefore, there is an urgent need for an effective new method to improve the segmentation accuracy of blurred boundaries in medical images and provide more reliable technical support for clinical practice. Summary of the Invention
[0005] In view of the problems of blurred boundaries and low segmentation accuracy in medical image segmentation in the prior art, the present invention proposes a medical image segmentation system and method based on COD frequency domain guidance.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, a method for segmenting blurred boundaries of medical images based on COD frequency domain guidance includes:
[0008] Step 1: Data preprocessing: Standardize the input medical image I to obtain a standard image;
[0009] Step 2: Model initialization: Establish an FEA-Net image data processing model, pre-trained PVTv2, and initialize the FEA-Net model with the weights of the pre-trained PVTv2;
[0010] Step 3: Image processing: Input the image into the FEA-Net model, perform semantic edge enhancement processing on the input standard image, continue to perform preliminary localization and extraction on the boundary features of the standard image, output the processed image data, then perform frequency domain feature-guided enhancement on the processed image data to strengthen the boundary feature representation and improve the accuracy of boundary prediction, then perform adaptive edge feature processing on the enhanced image, adaptively optimize the features output by the enhanced image to highlight the key boundary information and achieve fine-grained expression of boundary features, and finally perform context attention aggregation on the image data features processed above, perform adaptive fusion of the multi-scale features obtained above under the guidance of attention, and finally output an accurate segmentation result;
[0011] Step 4: Loss function calculation: Calculate the cross-entropy loss LCE of the segmentation task, the Dice loss LDice for measuring morphological similarity, and the edge loss LEdge for enhancing boundary learning respectively, and then combine these three losses into a comprehensive loss function L = α·LCE + β·LDice + γ·LEdge through the weighting coefficients α, β, γ to optimize the segmentation accuracy and boundary accuracy of the segmentation result;
[0012] Step 5: Iterative Training: Iteratively train the FEA-Net model. In each training iteration, first perform forward propagation to calculate the multitask loss, and then update the model parameters through backpropagation. Select the optimal model parameters through performance evaluation on the validation set. After training is completed, input the test images into the optimized model to generate segmentation prediction maps. These prediction maps accurately show the boundary contours of the target regions. Finally, the model outputs the final segmentation results of the medical images, ensuring clear segmentation of the boundary regions;
[0013] Step 6: Validation and Evaluation: Regularly evaluate the model performance on the validation set, calculate relevant metrics such as mean Pixel Accuracy, mean Intersection over Union, mean Dice coefficient, and F1-score to measure the segmentation performance of the model; after each round of evaluation, save the model parameters with the best performance according to the validation results.
[0014] Furthermore, the data preprocessing in Step 1 includes:
[0015] Normalize the input medical images. Use the mean μ = [0.485, 0.456, 0.406] and standard deviation σ = [0.229, 0.224, 0.225] for normalization. The calculation formula is:
[0016]
[0017] where μ is the mean of the image and σ is the standard deviation;
[0018] At the same time, perform data augmentation operations, including random horizontal and vertical flips with a probability of 0.5, dynamic image size adjustment, and random cropping, to enhance the robustness and generalization ability of the model
[0019] Furthermore, in Step 2, use PVTv2 pre-trained on ImageNet as the backbone network. Through its hierarchical Transformer structure and local self-attention mechanism, extract multiple-scale feature maps f 1 to f 4 from the input image. This network structure effectively captures the local detail features and global context information of the image, providing rich feature representations for subsequent boundary enhancement processing.
[0020] Furthermore, in Step 3, perform semantic edge enhancement processing on the input standard images, continue to perform preliminary localization and extraction of the boundary features of the standard images, and the output processed image data specifically includes:
[0021] First, the spatial-channel self-attention mechanism is utilized to simultaneously construct the dependencies of the feature map in both the spatial and channel dimensions. Its key mathematical expression is as follows:
[0022]
[0023] It contains 8 attention heads with a window size of 8×8. Among them, the spatial self-attention is used to highlight important spatial location information, and the channel self-attention is used to enhance the interaction between information channels.
[0024] Q h ,K h ,V h respectively represent the query matrix, key matrix, and value matrix of the h-th attention head. d represents the feature dimension for scaling the attention scores, the key and value matrices. and represent the average values of the features in the horizontal and vertical directions. SA(·) represents the spatial attention operation.
[0025] Then, dilated spatial pyramid pooling is adopted for processing. Softmax represents the normalization function used to convert the attention scores into a probability distribution.
[0026] F(X) = ∑ r∈{2,3,6} α r Conv r (X) + A(X)
[0027] where r represents the dilation rate of the dilated convolution, and the dilation rate configuration of [2, 3, 6] is adopted to capture multi-scale context information through different receptive field sizes and enhance the multi-scale representation ability of the features; α r represents the weight coefficient under different dilation rates. Conv r (·) represents the convolution operation with a dilation rate of r; X represents the input feature map, and A(X) represents the output of the aforementioned spatial-channel self-attention operation.
[0028] Finally, the processing is completed through the feature fusion and enhancement steps:
[0029] Y = φ(concat[F(feat 4 ), Conv 1×1 (feat 1 )])
[0030] where feat 1 represents the low-level features after dimensionality reduction by 1×1 convolution. After spatial attention to highlight important spatial positions, feat 4 is the high-level features after spatial-channel self-attention and ASPP processing. F(·) represents the aforementioned multi-scale feature processing function. Conv 1x1(·) indicates that the 1×1 convolution operation realizes the effective combination of low-level features and high-level features, concat(·) indicates the feature concatenation in the channel dimension, and finally, by concatenating feat 4 and feat 1 , and through the non-linear transformation including convolution, batch normalization, and ReLU activation, φ(·) further fuses and refines the features to enhance the boundary enhancement effect.
[0031] Furthermore, in step 3, frequency-domain feature-guided enhancement is performed on the processed image data to strengthen the boundary feature representation and improve the accuracy of boundary prediction, specifically including:
[0032] The Token fusion mechanism of adaptive frequency-domain enhancement is adopted in the frequency-domain feature-guided enhancement, and its key steps include the following mathematical expressions:
[0033] First, the features are transformed into the frequency-domain space through the two-dimensional fast Fourier transform:
[0034]
[0035] X reshape =Reshape(X f , [B, N b ,, S b , H, W])
[0036] where represents the orthonormal two-dimensional Fourier transform, B is the batch size, N b =8 is the number of blocks, S b is the block size, and H and W are the height and width of the feature map respectively; Reshape represents the tensor reshaping operation;
[0037] Feature enhancement is achieved through the frequency-domain block diagonal matrix transformation:
[0038] X′ real =σ(W 1 ·X real -W 2 ·X imag +b 1 )
[0039] X′ imag =σ(W 1 ·X imag +W 2 ·X real +b 2 )
[0040] where X real , X imag respectively represent the real and imaginary parts of the frequency-domain features, W 1and W 2 is a learnable complex field block diagonal weight matrix, b 1 and b 2 are bias terms, and σ represents an activation function;
[0041] Next, a soft threshold function is used for denoising and selection of frequency domain features:
[0042] X λ = S λ (X) = sign(X) max(|X| - λ, 0)
[0043] where S λ represents the soft threshold function, sign(X) represents the sign function, and λ = 0.01 is the sparsity threshold parameter used to control the sparsity degree of frequency domain features;
[0044] Finally, the spatial domain features are restored through the Hadamard product and the inverse Fourier transform:
[0045]
[0046] where represents the inverse Fourier transform, ⊙ represents the Hadamard product (element-wise product), b is the final bias term, and X λ represents the feature after soft threshold processing, and X f represents the original frequency domain feature, realizing the adaptive enhancement of frequency domain features and the reconstruction of spatial domain features.
[0047] Furthermore, in step 3, the enhanced image is subjected to adaptive edge feature processing, and the features output from the enhanced image are adaptively optimized to highlight key boundary information and realize the fine-grained expression of boundary features, specifically including:
[0048] First, a fusion operation is performed on the input feature and the boundary prediction weight:
[0049] X = C ⊙ A + C
[0050] where C represents the input feature map, A represents the boundary prediction weight, ⊙ represents element-wise multiplication, and the original feature information is retained through the residual connection +C;
[0051] Then, an adaptive channel attention mechanism AFGCA is used for feature enhancement, where the convolution kernel size is obtained through adaptive calculation:
[0052]
[0053] where C represents the channel dimension, and the convolution kernel size k is adaptively determined by the number of channels;
[0054] Finally, dynamic feature weighting is achieved through the Mix module:
[0055] Y = α·F 1 +(1 - α)·F 2
[0056] α = σ(w), w = -0.80
[0057] Where F 1 and F 2 are the features to be fused, α is the mixing factor calculated by the Sigmoid function σ, and the initial weight w is set to -0.80. This module realizes the adaptive enhancement of features through residual connection and dynamic feature weighting.
[0058] Furthermore, in step 3, the context attention aggregation is performed on the image data features processed above, and the adaptive fusion of the multi-scale features obtained above under the guidance of attention includes:
[0059] First, the low-level features and high-level features are unified and concatenated:
[0060]
[0061] Where L f and H f represent the low-level and high-level features respectively, represents the bilinear interpolation operation to adjust the high-level features to the same size as the low-level features, concat represents the feature concatenation in the channel dimension, and Conv 1×1 represents the convolution operation;
[0062] Then multi-level feature fusion is performed:
[0063]
[0064] Where is the attention weight, ⊙ represents element-wise multiplication, φ is the non-linear transformation, and X i is the feature block;
[0065] Finally, global feature fusion is achieved:
[0066] Y = Conv(concat[Y 2 , Y 3 , X 4 )
[0067] Where Y 2 and Y 3 represent the results of two-feature and three-feature fusion respectively, and X 4 is the fourth feature block.
[0068] Second aspect: A system using a medical image fuzzy boundary segmentation method based on COD frequency domain guidance, including a semantic edge enhancement module, a frequency domain feature guidance enhancement module, an adaptive edge feature module, and a context attention aggregation module;
[0069] The semantic edge enhancement module receives the features extracted from the backbone network, performs preliminary localization and extraction of the boundary, and strengthens the boundary information;
[0070] The frequency domain feature guidance enhancement module receives the features from the backbone network, performs frequency domain processing, and further improves the accuracy of the boundary features;
[0071] The adaptive edge feature module receives the processed features from the semantic edge enhancement module and the frequency domain feature guidance enhancement module, performs adaptive optimization, refines the boundary information, and highlights the key boundary features;
[0072] The context attention aggregation module receives the multi-scale features from the adaptive edge feature module, performs adaptive fusion, and combines its own feature transmission to finally output an accurate segmentation result, ensuring clear and accurate boundaries.
[0073] The beneficial effects of the technical solution of the present invention are as follows:
[0074] In view of the problem of accurate segmentation of medical images with fuzzy boundaries, the present invention proposes a segmentation network based on COD frequency domain guidance, innovatively combines frequency domain feature processing with an improved camouflaged object detection (COD) algorithm, and effectively improves the accuracy of segmentation through modular design, especially performing well in dealing with complex backgrounds and fuzzy boundaries;
[0075] The semantic edge enhancement module (SEEM) proposed by the present invention combines spatial-channel self-attention mechanism and ASPP, significantly enhancing the perception ability of the boundary. The design of this multiple attention mechanism enables the model to more accurately locate the boundary, overcoming the deficiencies of traditional methods in dealing with fuzzy boundaries;
[0076] The frequency domain feature guidance enhancement module (FTEM) proposed by the present invention realizes the effective extraction of boundary information through frequency domain processing and adaptive feature selection, while maintaining the integrity of the boundary, enhancing the expression ability of the features;
[0077] The present invention innovatively combines and strengthens the two modules of SEEM and FTEM to form a progressive boundary optimization mechanism. SEEM first realizes the preliminary localization of the boundary through spatial domain processing, and then FTEM further refines these boundary features in the frequency domain. This complementary design of spatial domain - frequency domain significantly improves the description ability of the model for the boundaries of complex medical images;
[0078] The proposed Adaptive Edge Feature Module (AEFM) and Context Attention Aggregation Module (CAAM) of the present invention achieve effective fusion of multi-scale features, improving the model's adaptability to targets of different sizes;
[0079] Through the synergistic effect of multiple innovative modules, the overall network architecture of the present invention effectively solves problems such as detail loss and boundary blurring while maintaining high-precision segmentation, providing reliable technical support for medical image analysis. The present invention not only improves the accuracy of segmentation but also enhances the interpretability and scalability of the model. Brief Description of the Drawings
[0080] Figure 1 It is a schematic structural diagram of the medical image blurred boundary segmentation system and method based on COD frequency domain guidance of the present invention;
[0081] Figure 2 It is a schematic structural diagram of the Semantic Edge Enhancement Module (SEEM) of the present invention;
[0082] Figure 3 It is a schematic structural diagram of the Frequency Domain Feature Guided Enhancement Module (FTEM) of the present invention;
[0083] Figure 4 It is a schematic structural diagram of the Adaptive Edge Feature Module (AEFM) of the present invention;
[0084] Figure 5 It is a schematic diagram of the result of the Context Attention Aggregation Module (CAAM) of the present invention;
[0085] Figure 6 It is a comparison chart of boundary prediction effects between SEEM, FTEM of the present invention and other methods;
[0086] Figure 7 It is a comparison chart of qualitative segmentation effects between FEA-Net of the present invention and other methods. Detailed Embodiments
[0087] To make the objectives, technical solutions and effects of the present invention clearer and more understandable, the present invention will be elaborated in detail below with reference to the accompanying drawings and through specific examples. It should be emphasized that the examples described herein are only for helping to understand the present invention and do not limit the scope of the present invention.
[0088] The present invention selects the improved PVTv2 as the backbone network, enhances the feature extraction and global perception capabilities through its hierarchical design, and innovatively combines the frequency-domain attention guidance mechanism with the improved camouflaged object detection algorithm (COD) to achieve precise positioning and segmentation of fuzzy boundaries, and proposes an innovative medical image segmentation framework. This innovative application not only improves the accuracy of medical image segmentation, but also opens up new research directions for the field of medical image analysis.
[0089] The core of the present invention is a new medical image segmentation framework FEA-Net, which realizes progressive boundary optimization through four collaborative modules. Specifically, it includes: the semantic edge enhancement module (SEEM) generates initial edge features using the spatial-channel self-attention mechanism. This module enhances the perception ability of target-related edge features by effectively integrating low-level detail information and high-level semantic information; the frequency-domain feature guidance enhancement module (FTEM) decomposes the features into the frequency domain for processing, enhances the high-frequency signals related to the boundary through block diagonal transformation and adaptive frequency selection operations, and suppresses noise interference at the same time; the adaptive edge feature module (AEFM) realizes a fine-grained channel attention mechanism and adaptively adjusts the feature response according to the boundary features; the context attention aggregation module (CAAM) realizes the intelligent fusion of multi-scale features through the attention mechanism, and effectively integrates global information while maintaining boundary details.
[0090] Through the collaborative effect of the above modules, the present invention establishes a progressive feature optimization framework from the spatial domain to the frequency domain and from the local to the global, significantly improving the segmentation accuracy of the fuzzy boundary region in medical images. This method is particularly suitable for dealing with complex medical image segmentation tasks with low contrast and unclear boundaries, providing more reliable technical support for clinical diagnosis.
[0091] A method for segmenting fuzzy boundaries of medical images based on COD frequency-domain guidance is as follows:
[0092] Data preprocessing:
[0093] The input medical images are standardized, normalized using the mean μ = [0.485, 0.456, 0.406] and the standard deviation σ = [0.229, 0.224, 0.225], and the calculation formula is I_norm = (I - μ) / σ; at the same time, data augmentation operations are performed, including random horizontal and vertical flips with a probability of 0.5, dynamic image size adjustment, and random cropping, to enhance the robustness and generalization ability of the model.
[0094] Feature extraction:
[0095] Use PVTv2 pre-trained on ImageNet as the backbone network. Through its hierarchical Transformer structure and local self-attention mechanism, extract features f 1 to f 4 feature maps of multiple scales from the input image. This network structure can effectively capture local detail features and global context information of the image simultaneously, providing rich feature representations for subsequent boundary enhancement processing.
[0096] Semantic edge enhancement module processing:
[0097] First, use 1×1 convolution to reduce the channel dimension of the features from the shallow layer to obtain a more compact feature representation, and then use the SCSA_RB block configured with 8 attention heads and a window size of 8×8 to process the deep f 1 features to explore long-range feature relationships; then use the ASPP module with dilation rates of [2, 3, 6] to capture multi-scale context information, match the feature map sizes through upsampling operations, and finally achieve fine fusion of multi-scale features through the ConvBNR block to enhance the expression of boundary semantic information. Specifically: 4 First, utilize the spatial-channel self-attention mechanism to simultaneously construct the dependencies of the feature map in the spatial and channel dimensions. Its key mathematical expression is:
[0098] First, utilize the spatial-channel self-attention mechanism to simultaneously construct the dependencies of the feature map in the spatial and channel dimensions. Its key mathematical expression is:
[0099]
[0100] which contains 8 attention heads and a window size of 8×8; among them, spatial self-attention is used to highlight important spatial location information, and channel self-attention is used to enhance the interaction between information channels.
[0101] Q h , K h , V h represent the query matrix, key matrix, and value matrix of the h-th attention head respectively. d represents the feature dimension for scaling the attention scores, the key and value matrices. and represent the average values of the features in the horizontal and vertical directions, and SA(·) represents the spatial attention operation;
[0102] Then, perform processing using dilated spatial pyramid pooling. Softmax represents the normalization function used to convert the attention scores into a probability distribution;
[0103] F(X) = ∑ r∈{2,3,6} α r Conv r (X) + A(X)
[0104] Among them, r represents the dilation rate of the dilated convolution. The dilation rate configuration of [2, 3, 6] is adopted to capture multi-scale context information through different receptive field sizes and enhance the multi-scale representation ability of features; α r represents the weight coefficient under different dilation rates, and Conv r (·) represents the convolution operation with a dilation rate of r; X represents the input feature map, and A(X) represents the output of the aforementioned spatial-channel self-attention operation:
[0105] Finally, the processing is completed through the feature fusion and enhancement steps:
[0106] Y = φ(concat[F(feat 4 ), Conv 1×1 (feat 1 )])
[0107] Among them, feat 1 represents the low-level features after dimensionality reduction by 1×1 convolution. After spatial attention to highlight important spatial positions, feat 4 is the high-level features processed by spatial-channel self-attention and ASPP. F(·) represents the aforementioned multi-scale feature processing function, and Conv 1x1 (·) represents the 1×1 convolution operation to achieve an effective combination of low-level and high-level features. concat(·) represents feature concatenation in the channel dimension. Finally, by concatenating feat 4 and feat 1 in the channel dimension and passing through a non-linear transformation including convolution, batch normalization, and ReLU activation, φ(·) further fuses and refines the features to enhance the boundary enhancement effect.
[0108] Frequency-domain feature-guided enhancement module processing:
[0109] First, the dimensions of the f 1 and f 4 features of the backbone extraction network are reduced by 1×1 convolution, and then the features are mapped to the frequency domain space using the two-dimensional Fourier transform; in the frequency domain, the N b = 8-block complex-domain block diagonal matrix transformation is used for feature processing, and an adaptive frequency selection operation with a soft threshold of λ = 0.01 is performed to enhance the high-frequency signals related to the boundary while suppressing noise; finally, the enhanced features are restored to the spatial domain through the Hadamard product and the inverse Fourier transform, and the original boundary detail information is retained through the residual connection, specifically including:
[0110] The adaptive frequency-domain enhancement Token fusion mechanism is adopted in the frequency-domain feature-guided enhancement, and its key steps include the following mathematical expressions:
[0111] First, the features are transformed into the frequency domain space through a two-dimensional fast Fourier transform:
[0112]
[0113] X reshape = Reshape(X f , [B, N b , S b , H, W])
[0114] where, represents the orthonormal two-dimensional Fourier transform, B is the batch size, N b = 8 is the number of blocks, S b is the block size, and H and W are the height and width of the feature map respectively; Reshape represents the tensor reshaping operation;
[0115] Feature enhancement is achieved through a frequency domain block diagonal matrix transformation:
[0116] X' real = σ(W 1 ·X real - W 2 ·X imag + b 1 )
[0117] X' imag = σ(W 1 ·X imag + W 2 ·X real + b 2 )
[0118] where, X real , X imag represent the real and imaginary parts of the frequency domain features respectively, W 1 and W 2 are learnable complex domain block diagonal weight matrices, b 1 and b 2 are bias terms, and σ represents the activation function;
[0119] Next, a soft threshold function is used for denoising and selection of the frequency domain features:
[0120] X λ = S λ (X) = sign(X)mx(|X| - λ, 0)
[0121] where, S λ represents the soft threshold function, sign(X) represents the sign function, and λ = 0.01 is the sparsity threshold parameter used to control the sparsity degree of the frequency domain features;
[0122] Finally, the spatial domain features are restored through the Hadamard product and the inverse Fourier transform:
[0123]
[0124] where denotes the inverse Fourier transform, ⊙ denotes the Hadamard product (element-wise product), b is the final bias term, and X λ denotes the feature after soft thresholding, and X f denotes the original frequency domain feature, realizing the adaptive enhancement of the frequency domain feature and the reconstruction of the spatial domain feature.
[0125] Adaptive edge feature module processing:
[0126] Perform element-wise multiplication fusion on the input feature and the boundary prediction weight and add a skip connection to retain the original information; then adopt an adaptive channel attention mechanism that dynamically calculates the convolution kernel size (t = |(log2C + 1) / 2|) based on the channel dimension C, and use a channel mixing function with an initial weight of -0.80 for feature modulation; finally, capture the local spatial pattern through a 3×3 convolution and use a residual connection for feature integration to achieve fine-grained enhancement of the boundary feature, specifically including:
[0127] First, perform a fusion operation on the input feature and the boundary prediction weight:
[0128] X = C⊙A + C
[0129] where C represents the input feature map, A represents the boundary prediction weight, ⊙ represents element-wise multiplication, and the original feature information is retained through the residual connection +C;
[0130] Then, adopt the adaptive channel attention mechanism AFGCA for feature enhancement, where the convolution kernel size is obtained through adaptive calculation:
[0131]
[0132] where C represents the channel dimension, and the convolution kernel size k is adaptively determined by the number of channels;
[0133] Finally, dynamic feature weighting is achieved through the Mix module:
[0134] Y = α·F 1 +(1 - α)·F 2
[0135] α = σ(w), w = -0.80
[0136] where F 1 and F 2is the feature to be fused, α is the mixing factor calculated by the Sigmoid function σ, and the initial weight w is set to -0.80°. This module realizes the adaptive enhancement of features through residual connection and dynamic feature weighting.
[0137] Context attention aggregation module processing:
[0138] First, the dimensions of N groups of input features are unified through 1×1 convolution, and the feature combination is calculated to obtain the global view; then the attention weights are calculated through 1×1 convolution and the Softmax function
[0139] W = softmax(Conv 1×1 (F)), and based on the weights, the weighted feature fusion operation xi' = xi·Wi + xi is performed; finally, the weighted features are processed using 3×3 convolution operations with different dilation rates to effectively capture multi-scale context information, including:
[0140] First, the low-level features and high-level features are unified and concatenated:
[0141]
[0142] where L f and H f represent the low-level and high-level features respectively, represents the bilinear interpolation operation to adjust the high-level features to the same size as the low-level features, concat represents the feature concatenation in the channel dimension, and Conv 1×1 represents the convolution operation;
[0143] Then, multi-level feature fusion is performed:
[0144]
[0145] where is the attention weight, ⊙ represents element-wise multiplication, φ is the non-linear transformation, and X i is the feature block;
[0146] Finally, global feature fusion is achieved:
[0147] Y = Conv(concat[Y 2 , Y 3 , X 4 )
[0148] where Y 2 and Y 3 represent the results of two-feature and three-feature fusion respectively, and X 4 is the fourth feature block.
[0149] Loss function calculation:
[0150] Calculate the cross-entropy loss \(L_{CE}\) of the segmentation task, the Dice loss \(L_{Dice}\) for measuring morphological similarity, and the edge loss \(L_{Edge}\) for enhancing boundary learning respectively. Then, combine these three losses into a comprehensive loss function \(L=\alpha\cdot L_{CE}+\beta\cdot L_{Dice}+\gamma\cdot L_{Edge}\) through the weighting coefficients \(\alpha\), \(\beta\), and \(\gamma\) to jointly optimize the segmentation accuracy and boundary accuracy, specifically including:
[0151] Use a weighted combination of cross-entropy loss, Dice loss, and edge loss to optimize the segmentation task, ensuring that the model can achieve good segmentation results in the case of class imbalance and complex boundaries. The loss function is defined as:
[0152] \(L = \alpha*L_{CE}+\beta*L_{Dice}+\gamma*L_{Edge}\)
[0153] where \(L_{CE}\) is the cross-entropy loss, \(L_{Dice}\) is the Dice loss, \(L_{Edge}\) is the edge loss, and \(\alpha\), \(\beta\), \(\gamma\) are balance factors. This combination can simultaneously consider the accuracy of global classification, the precision of local region matching, and the accuracy of boundary localization.
[0154] Among them, the Adam optimizer is used. The Adam optimizer has good robustness and fast convergence when dealing with complex optimization problems, and the initial learning rate \(lr = 1e-4\);
[0155] Model training:
[0156] Set the training batch size to 16, use the Adam optimizer to optimize the parameters, set the initial learning rate to \(1e-4\) and use the poly strategy with power 0.9 for dynamic adjustment; train on the NVIDIA 4090 GPU for 50 epochs until the model converges. The entire training process takes about 2 - 6 hours, and select the optimal model through the performance evaluation of the validation set;
[0157] Then, perform learning rate scheduling, adopt the polynomial decay strategy to dynamically adjust the learning rate, and gradually reduce the learning rate as the training progresses to prevent overfitting or oscillation in the later stage of training. The learning rate update formula is:
[0158]
[0159] where \(t\) is the current iteration number, \(T\) is the total iteration number, \(lr\) 0 is the initial learning rate, and power is a hyperparameter controlling the decay speed, set to 0.9. This scheduling strategy can ensure that the model is more stable for fine-tuning in the later stage of training.
[0160] Output result:
[0161] Input the test image into the trained model to generate a segmentation prediction map, which clearly shows the boundary contours of the target area; finally, output the final segmentation result of the medical image.
[0162] Figure 1 It is a schematic structural diagram of the medical image boundary segmentation system based on frequency-domain edge attention of the present invention. As shown in the figure, the system includes a feature extraction module, a semantic edge enhancement module (SEEM), a frequency-domain feature-guided enhancement module (FTEM), an adaptive edge feature module (AEFM), and a context attention aggregation module (CAAM). These modules work together to form an end-to-end segmentation network structure for achieving high-precision medical image segmentation.
[0163] Reference Figure 1 And Figure 2 , the innovation of the semantic edge enhancement module (SEEM) lies in combining the spatial channel self-attention mechanism and the atrous spatial pyramid pooling (ASPP), which enhances the perception ability of boundary features. The SA mechanism highlights the important spatial positions in the feature map, and the CA mechanism weights the channels to further optimize the feature expression related to the boundary. In addition, the ASPP module expands the receptive field through multi-scale convolution, effectively dealing with targets of different sizes and shapes. This module realizes the extraction of initial boundary features through the fusion and refinement of multi-scale features.
[0164] Reference Figure 1 And Figure 3 , the innovation of the frequency-domain feature-guided enhancement module (FTEM) is to enhance boundary features through the frequency-domain processing mechanism. This module processes the features in the frequency-domain space, and realizes feature enhancement through the block diagonal matrix transformation in the complex domain and adaptive frequency selection. Finally, the spatial features are restored through the inverse transformation, effectively enhancing the feature expression of the fuzzy boundary region. Compared with the boundary processing of the SEEM module in the spatial domain, FTEM can capture subtle boundary changes more accurately through the cooperation of frequency-domain transformation and multi-scale attention.
[0165] Reference Figure 2 And Figure 3 Shows the comparison of the boundary extraction effects of SEEM and FTEM. SEEM shows better target contour description ability by introducing the SCSA attention mechanism. To further enhance boundary feature learning, FTEM expands this ability through Fourier preprocessing and adaptive frequency selection mechanism, especially showing more accurate boundary description effect when dealing with fuzzy boundaries. This progressive enhancement from SEEM to FTEM fully demonstrates the effectiveness of our boundary optimization strategy in dealing with complex medical image boundaries. For example, in the ISIC dataset, the initial boundary captured by SEEM has breaks ( Figure 6) After frequency domain enhancement of FTEM, the boundary continuity is significantly improved. Figure 6 ) This progressive optimization mechanism realizes the accurate characterization of complex boundaries through the complementarity of spatial domain localization and frequency domain enhancement.
[0166] Reference Figure 1 、 Figure 4 and Figure 5 With reference to
[0167] To verify the performance of the system of the present invention, we conducted a comprehensive experimental evaluation on three medical image datasets, namely ISIC2018, Kvasir-SEG, and KPIs2024. These datasets cover a variety of medical imaging modalities, from histopathology to dermatology and endoscopy, and each modality poses different challenges for boundary delineation.
[0168] We compared the method of the present invention with a variety of state-of-the-art methods, including traditional segmentation methods such as DeepLabV3+ and U-Net3+, and camouflaged object detection methods such as SINetV2, PFNet, and BGNet. The experimental results are shown in Table 1. The method of the present invention achieved the best results in all indicators, demonstrating significant performance advantages. Specifically, our method achieved continuous improvements compared to the state-of-the-art models. Compared with the second-best model in each indicator and dataset, the mPA of FEA-NET increased by an average of 0.60%, the mIoU increased by an average of 1.33%, and the mDice increased by an average of 0.84%. It is worth noting that compared with the third-ranked model, the improvement of FEA-NET is even greater (for example, the improvement of mIoU is +1.63%).
[0169] Table 1 Quantitative comparison and evaluation table of method performance
[0170]
[0171] In terms of boundary exploration, Figure 6 shows the visual comparison of EAM (from the baseline), SEEM, and FTEM in boundary refinement. From the visualization results, EAM utilizes the multi-scale features of the backbone network (f 2 and f 5) Edge detection is performed, but it is difficult to capture complex boundary patterns, resulting in suboptimal boundary localization in medical image segmentation. Based on EAM, SEEM enhances the boundary detection ability by introducing SCSA attention to capture spatial-channel relationships and shows higher performance in object contour delineation. To further strengthen boundary feature learning, our FTEM extends boundary feature learning with Fourier preprocessing through an adaptive frequency selection and label mixing mechanism, achieving more accurate boundary delineation, especially for blurred edges. The gradual enhancement from EAM to SEEM and FTEM demonstrates the effectiveness of our boundary refinement strategy in handling challenging medical image boundaries, as evidenced by the increasingly accurate and detailed segmentation results.
[0172] In terms of qualitative analysis, Figure 7 shows the qualitative comparison of different methods on several typical samples. As the intuitive results show, our method can provide more accurate predictions, finer details, and more complete structure preservation. Specifically, compared with other methods such as DeepLabV3+, UNet3+, and Segformer, for complex medical image cases with different tissue morphologies, our method can maintain more precise boundary delineation. The results intuitively show that our method can better handle challenging cases where traditional methods often fail, such as cases with unclear boundaries or complex anatomical structures.
[0173] To further verify the effectiveness of the proposed model, we also conducted comprehensive ablation experiments. These experiments aim to evaluate the effects of the individual modules we proposed and how they work together to improve the overall performance.
[0174] The ablation experiment results are shown in Table 2. The integration of FTEM+AEFM+CAAM brings significant improvements on all three datasets. Most notably, on the challenging KPIs2024 dataset, the model's mIoU and F1-score are increased by 3.4% and 3.8% respectively, indicating FTEM's excellent ability to capture subtle boundary features through frequency-domain processing. Considering the complex histopathological features of the dataset, this improvement is particularly important. The continuous performance improvement on ISIC2018 (+2.2% mIoU, +0.9% F1-score) and Kvasir-SEG (+1.0% mIoU, +0.8% F1-score) further validates the effectiveness of FTEM in handling various medical imaging modalities with different boundary features;
[0175] Table 2 Ablation experiment results
[0176]
[0177] The synergistic combination of SEEM, AEFM, and CAAM has robust improvements in all evaluation metrics. On the KPIs2024 dataset, this configuration achieved significant gains (+2.4% mIoU, +2.6% F1-score), while maintaining continuous improvements on ISIC2018 (+1.8% mIoU, +0.6% F1-score) and Kvasir-SEG (+0.7% mIoU, +0.6% F1-score). These results validate the effectiveness of our enhanced semantic edge learning and adaptive feature fusion mechanisms in capturing multi-scale context information;
[0178] FTEM + SEEM + AEFM + CAAM achieved the best performance in all metrics, with an average increase of 0.93% in mPA, 2.67% in mIoU, 2.20% in mDice, and 2.47% in F1-score. In challenging cases, the synergistic effect was particularly evident, with the highest gains in KPIs2024 (+3.8% mIoU), and continuous improvements in ISIC2018 (+2.7% mIoU) and Kvasir-SEG (+1.5% mIoU).
[0179] In-depth analysis of module interactions shows that the frequency domain processing of FTEM provides complementary information for the spatial domain features of SEEM. When examining the performance of complex cases in KPIs2024, we observed that regions with subtle intensity variations showed particularly significant improvements (a 3.8% increase in boundary accuracy) compared to regions with clearer boundaries (+2.1%). This indicates that frequency domain processing can effectively capture features that are less obvious in spatial domain analysis; these comprehensive improvements demonstrate that the modules we proposed have strong complementarity when dealing with challenging medical image segmentation tasks, especially in cases where tissue boundaries are unclear and anatomical structures are complex.
[0180] In summary, the medical image fuzzy boundary segmentation system and method based on COD frequency domain guidance proposed in the present invention effectively solves the challenges faced by existing methods in dealing with fuzzy boundaries in medical images by innovatively combining frequency domain processing with an improved camouflaged object detection algorithm. The system not only significantly improves the segmentation accuracy and boundary localization accuracy of fuzzy boundary regions in medical images, but also provides new technical ideas and methods for other medical image analysis tasks, opening up a new research direction for the combined application of frequency domain enhancement and camouflaged object detection algorithms in medical image analysis.
[0181] It should be understood that those skilled in the art, inspired by the technical concept of the present invention, can make various improvements or transformations based on the above description without departing from the content of the present invention, but this still falls within the protection scope of the present invention.
Claims
1. A medical image fuzzy boundary segmentation method based on COD frequency domain guidance, characterized in that: include: Step 1: Data preprocessing: Standardize the input medical image I to obtain a standard image; Step 2: Model initialization: Establish the FEA-Net image data processing model, pre-trained PVTv2, and use the pre-trained PVTv2 weights to initialize the FEA-Net model; Step 3: Image processing: Input the image into the FEA-Net model, perform semantic edge enhancement on the input standard image, continue to preliminarily locate and extract the boundary features of the standard image, output the processed image data, and then perform frequency domain feature-guided enhancement on the processed image data to strengthen the boundary feature representation and improve the accuracy of boundary prediction. Then, perform adaptive edge feature processing on the enhanced image, adaptively optimize the features of the enhanced image output, highlight the key boundary information, and achieve fine-grained expression of boundary features. Finally, perform contextual attention aggregation on the image data features processed above, and perform adaptive fusion of the multi-scale features obtained above under the guidance of attention, and finally output accurate segmentation results. Step 4: Loss function calculation: Calculate the cross entropy loss LCE of the segmentation task, the Dice loss LDice for measuring morphological similarity, and the boundary loss LEdge for enhancing boundary learning respectively, and then combine these three losses into a comprehensive loss function L = α·LCE+β·LDice+γ·LEdge through weighted coefficients α, β, and γ to optimize the segmentation accuracy and boundary accuracy of the segmentation results; Step 5: Iterative training: The FEA-Net model is iteratively trained. In each training iteration, forward propagation is performed to calculate the multi-task loss, and then the model parameters are updated through backpropagation. The optimal model parameters are selected through the performance evaluation of the validation set. The test image is then input into the trained model to generate a segmentation prediction map, which clearly shows the boundary contour of the target area; finally, the final segmentation result of the medical image is output; Step 6: Validation and evaluation: Regularly evaluate the model performance on the validation set and calculate relevant indicators such as mean Pixel Accuracy, mean Intersection over Union, mean Dice coefficient, and F1-score to measure the segmentation performance of the model. After each round of evaluation, the model parameters with the best performance will be saved based on the validation results.
2. The medical image fuzzy boundary segmentation method based on COD frequency domain guidance according to claim 1 is characterized in that: The data preprocessing in step 1 includes: The input medical image is standardized using mean μ = [0.485, 0.456, 0.406] and standard deviation σ = [0.229, 0.224, 0.225], and the calculation formula is: Among them, μ is the mean of the image and σ is the standard deviation; Data augmentation operations, including random horizontal and vertical flipping with a probability of 0.5, dynamic image resizing, and random cropping, are performed to enhance the robustness and generalization ability of the model.
3. The medical image fuzzy boundary segmentation method based on COD frequency domain guidance according to claim 1 is characterized in that: In the step 2, PVTv2 pre-trained on ImageNet is used as the backbone network. Through its hierarchical Transformer structure and local self-attention mechanism, feature maps of multiple scales from f1 to f4 are extracted layer by layer from the input image. This network structure effectively captures the local detail features and global context information of the image, providing rich feature representation for subsequent boundary enhancement processing.
4. The medical image fuzzy boundary segmentation method based on COD frequency domain guidance according to claim 1 is characterized in that: In step 3, semantic edge enhancement processing is performed on the input standard image, and the boundary features of the standard image are further preliminarily located and extracted. The output processed image data specifically includes: First, the spatial-channel self-attention mechanism is used to construct the dependency relationship of the feature map in the spatial and channel dimensions. The key mathematical expression is: It contains 8 attention heads with a window size of 8×8, where spatial self-attention is used to highlight important spatial position information, channel self-attention is used to enhance the interaction between information channels, and Q h ,K h ,V h denote the query matrix, key matrix, and value matrix of the h-th attention head, respectively. d denotes the feature dimension used for scaling the attention scores, key and value matrices. and represents the average of the features in the horizontal and vertical directions, SA(·) represents the spatial attention operation; Then, the atrous spatial pyramid pooling is used for processing, and the softmax normalization function is used to convert the attention score into a probability distribution; F(X)=∑ re{2,3,6} α r Conv r (X)+A(X) Where r represents the dilation rate of the dilated convolution, and the dilation rate configuration of [2, 3, 6] is adopted to capture multi-scale context information through different receptive field sizes and enhance the multi-scale representation ability of features; α r Represents the weight coefficient under different expansion rates, Conv r (·) represents a convolution operation with a dilation rate of r; X represents the input feature map, and A(X) represents the output of the aforementioned spatial-channel self-attention operation: Finally, the processing is completed through feature fusion and enhancement steps: Y=ϕ(concat[F(feat4),Conv 1×1 (feat1)]) Among them, feat1 represents the low-level features after 1×1 convolution and the important spatial positions are highlighted after spatial attention. feat4 is the high-level features after spatial-channel self-attention and ASPP processing. F(·) represents the aforementioned multi-scale feature processing function. Conv 1x1 (·) indicates that the 1×1 convolution operation realizes the effective combination of low-level features and high-level features. concat(·) indicates the feature concatenation in the channel dimension. Finally, by concatenating feat4 and feat1 in the channel dimension and undergoing nonlinear transformations including convolution, batch normalization, and ReLU activation, φ(·) further fuses and refines the features to improve the boundary enhancement effect.
5. The medical image fuzzy boundary segmentation method based on COD frequency domain guidance according to claim 1 is characterized in that: In step 3, the processed image data is enhanced by frequency domain feature guidance to strengthen the boundary feature representation and improve the accuracy of boundary prediction, which specifically includes: The frequency domain feature guided enhancement adopts the Token fusion mechanism of adaptive frequency domain enhancement. Its key steps include the following mathematical expressions: First, the features are converted to the frequency domain through two-dimensional fast Fourier transform: X reshape =Reshape(X f ,[B,N b ,S b ,H,W]) in, represents the orthogonal normalized two-dimensional Fourier transform, B is the batch size, N b =8 is the number of blocks, S b is the block size, H and W are the height and width of the feature map respectively; Reshape represents the tensor reshaping operation; Feature enhancement is achieved through frequency domain block diagonal matrix transformation: X′ real =σ(W1·X real -W2·X imag +b1) X′ imag =σ(W1·X imag +W2·X real +b2) Among them, X real , X imag Respectively represent the real and imaginary parts of the frequency domain features, W1 and W2 are learnable complex domain block diagonal weight matrices, b1 and b2 are bias terms, and σ represents the activation function; Next, the soft threshold function is used to denoise and select frequency domain features: X λ =S λ (X)=sign(X)max(|X|-λ,0) Among them, S λ represents the soft threshold function, sign(X) represents the sign function, and λ=0.01 is the sparsity threshold parameter, which is used to control the sparsity of frequency domain features; Finally, the spatial domain features are restored through Hadamard product and inverse Fourier transform: in, represents the inverse Fourier transform, ⊙ represents the Hadamard product, b is the final bias term, X λ represents the feature after soft threshold processing, X f Represents the original frequency domain features, realizes adaptive enhancement of frequency domain features and reconstruction of spatial domain features.
6. The medical image fuzzy boundary segmentation method based on COD frequency domain guidance according to claim 1 is characterized in that: In step 3, the enhanced image is subjected to adaptive edge feature processing, the enhanced image output features are adaptively optimized, key boundary information is highlighted, and fine-grained expression of boundary features is achieved, specifically including: First, the input features and boundary prediction weights are fused: X=C⊙A+C Among them, C represents the input feature map, A represents the boundary prediction weight, ⊙ represents element-level multiplication, and the original feature information is retained through the residual connection + C; Then, the adaptive channel attention mechanism AFGCA is used for feature enhancement, where the convolution kernel size is obtained by adaptive calculation: Among them, C represents the channel dimension, and the convolution kernel size k is adaptively determined by the number of channels; Finally, dynamic feature weighting is achieved through the Mix module: Y=α·F1+(1-α)·F2 α=σ(w),w=-0.80 Among them, F1 and F2 are the features to be fused, α is the mixing factor calculated by the Sigmoid function σ, and the initial weight w is set to -0.80°. This module realizes adaptive enhancement of features through residual connection and dynamic feature weighting.
7. The medical image fuzzy boundary segmentation method based on COD frequency domain guidance according to claim 1 is characterized in that: In step 3, context attention aggregation is performed on the image data features processed above, and adaptive fusion of the multi-scale features obtained above under attention guidance includes: First, unify and concatenate low-level features and high-level features: Among them, L f and H f Represent low-level and high-level features respectively, Indicates that the bilinear interpolation operation adjusts the high-level features to the same size as the low-level features, concat indicates feature concatenation in the channel dimension, Conv 1×1 Represents the convolution operation; Then perform multi-level feature fusion: in, is the attention weight, ⊙ represents element-by-element multiplication, φ is a nonlinear transformation, X i is a feature block; Finally, global feature fusion is achieved: Y=Conv(concat[Y2, Y3, X4]) Among them, Y2 and Y3 represent the results of two-feature and three-feature fusion respectively, and X4 is the fourth feature block.
8. A system using the COD frequency domain guided medical image fuzzy boundary segmentation method as claimed in claim 1, characterized in that: It includes semantic edge enhancement module, frequency domain feature guided enhancement module, adaptive edge feature module and contextual attention aggregation module; The semantic edge enhancement module receives the features extracted from the backbone network, performs preliminary positioning and extraction of boundaries, and strengthens boundary information; The frequency domain feature guided enhancement module receives features from the backbone network and performs frequency domain processing to improve the accuracy of boundary features; The adaptive edge feature module receives the processed features from the semantic edge enhancement module and the frequency domain feature guided enhancement module, performs adaptive optimization, refines the boundary information, and highlights the key boundary features; The contextual attention aggregation module receives multi-scale features from the adaptive edge feature module, performs adaptive fusion, and combines its own feature transfer to finally output accurate segmentation results to ensure clear and accurate boundaries.
Citation Information
Patent Citations
RGB-D image-based CLANet steel rail surface defect detection system and method
CN114170174A
Superpixel segmentation method and system based on frequency domain guidance
CN117079280A
System and method for accurately segmenting EdgeAttenNet glomerular image based on camouflage target detection
CN119206218A
Medical image segmentation method based on residual axial attention
CN119229127A
Camouflaged object segmentation method with distraction mining
US20220230324A1
Cited By
Boundary enhancement model for semi-supervised medical image segmentation
CN121095274A
Semantic segmentation-based method for extracting ocean internal waves from SAR remote sensing image
CN121527630A
Method for extracting oceanic internal waves from sar remote sensing images based on semantic segmentation
CN121527630B
Medical image segmentation model training method based on multi-mode scanning Mama
CN122049381A
Method for training a medical image segmentation model based on multi-modal scan mamba
CN122049381B