An edge-heuristic medical image segmentation method based on efficient parallel coding

By combining CNN and KAN-Transformer with a parallel encoder architecture and utilizing an edge-inspired mechanism, the problem of insufficient utilization of local details and global contextual information in medical image segmentation by existing methods is solved, thereby improving segmentation accuracy.

CN120747146BActive Publication Date: 2025-11-14WUXI LANANG INTELLECTUAL PROPERTY SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511258012.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-14
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing deep learning methods struggle to effectively utilize both local details and global contextual information simultaneously in medical image segmentation, resulting in insufficient segmentation accuracy for complex medical images, particularly in areas with blurred target boundaries.

Method used

A parallel encoder architecture is designed, which combines the local feature extraction capability of CNN with the global dependency modeling capability of Transformer. The computational complexity is reduced by using KAN-Transformer encoder and wavelet transform, and an edge heuristic mechanism is introduced to enhance the utilization of edge information and improve segmentation accuracy.

Benefits of technology

It significantly improves the segmentation accuracy of blurred boundary regions in medical image segmentation, and achieves efficient and accurate segmentation of complex medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747146B_ABST
    Figure CN120747146B_ABST
Patent Text Reader

Abstract

This invention relates to an edge-heuristic medical image segmentation method based on efficient parallel encoding. Addressing the problems existing in current medical image segmentation tasks, this method proposes an encoder-decoder architecture. In the encoding stage, a parallel encoder combining a CNN and a KAN-Transformer processes the input image. The KAN-Transformer utilizes wavelet transform to reduce the dimensionality of the KAN network, preserving the advantages of KAN's nonlinear modeling while reducing computational complexity. Multi-scale features from the two encoders are fused through a channel attention mechanism. In the decoding stage, an edge extraction module first extracts edge information from the fused encoded features; subsequently, an edge enhancement module fuses the edge information with the decoded features, enabling edge information to guide segmentation. This invention effectively improves the accuracy and efficiency of medical image target region segmentation by combining the local feature extraction capability of CNNs with the global dependency modeling capability of Transformers through parallel encoding, enhancing nonlinear representation through the KAN structure, and introducing an edge-heuristic mechanism to strengthen segmentation boundaries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image segmentation research, and discloses an edge-heuristic medical image segmentation method based on efficient parallel coding. Background Technology

[0002] Medical image segmentation is a key area of ​​interdisciplinary research in artificial intelligence and medicine. Its goal is to accurately identify and segment different anatomical structures or lesion regions in images. This technology plays a vital role in clinical diagnosis, surgical planning, and treatment evaluation.

[0003] Compared to natural scene images, medical image segmentation faces significant challenges, including: the contrast between the target region and the background is usually low, resulting in blurred target boundaries; the imaging process is susceptible to noise interference, further reducing image quality; and the small contrast differences between different tissues increase the risk of missegmentation.

[0004] Traditional medical image segmentation methods primarily rely on image processing techniques and mathematical models, such as thresholding, edge detection, region growing, and level set methods. While these methods may be effective for images with relatively simple structures, they have significant limitations when dealing with complex and varied medical images.

[0005] With the advancement of computer vision and medical imaging technologies, deep learning-based medical image segmentation methods have become an important research direction.

[0006] Current mainstream deep learning methods can be divided into segmentation methods based on convolutional neural networks (CNN) and segmentation methods based on visual Transformers (ViT).

[0007] CNN-based methods excel at capturing local features of images, but they fall short in modeling long-range dependencies. ViT-based methods, on the other hand, can effectively extract global contextual information, but they often neglect to fully explore local details, resulting in limited ability to represent fine structures in medical images.

[0008] In existing technologies, most deep learning-based medical image segmentation methods employ a single encoder structure (whether pure CNN, pure ViT, or a simple combination thereof), making it difficult to simultaneously and fully utilize both local details and global contextual information of the image. This single-path feature extraction approach fails to effectively overcome the inherent limitations imposed by the architectural characteristics of CNN or Transformer, thus restricting the potential for performance improvement in complex medical image segmentation tasks.

[0009] Some existing methods attempt to fuse CNN and Transformer structures, aiming to combine local and global information. However, such fusion approaches often fail to achieve deep and efficient collaboration between local features and global contextual information, and there is still room for improvement in fully leveraging the advantages of both to significantly enhance segmentation accuracy.

[0010] It is worth noting that the Kolmogorov-Arnold Network (KAN), as a novel neural network architecture based on the Kolmogorov-Arnold representation theorem, theoretically has the potential to efficiently model complex nonlinear relationships with fewer parameters. However, the application of KAN in image processing, especially in computationally intensive and highly accurate fields such as medical image segmentation, has not been fully explored. Its high computational complexity and memory consumption when processing image data particularly hinder its effective application in practical scenarios such as medical image segmentation.

[0011] Therefore, given the inherent challenges of medical images (such as low contrast, noise interference, and rich details) and the shortcomings of existing deep learning methods in effectively fusing local details and global contextual information, there is an urgent need to develop an efficient, accurate medical image segmentation method that can effectively utilize key information such as segmentation boundaries. Summary of the Invention

[0012] To overcome the shortcomings of existing technologies, this invention provides an edge-heuristic medical image segmentation method based on efficient parallel encoding. The core of this method lies in: designing a parallel encoder architecture that integrates the local feature extraction capabilities of CNNs with the global dependency modeling capabilities of Transformers; constructing a KAN-Transformer encoder by integrating Kolmogorov-Arnold Network (KAN) and wavelet transform into the Transformer architecture, effectively reducing computational complexity while maintaining the efficient nonlinear modeling advantages of KAN; and introducing an edge-heuristic mechanism to enhance the semantic meaning of target edges in the decoded features using extracted edge information, thereby improving the segmentation accuracy of regions with blurred boundaries.

[0013] The technical solution of this invention is achieved through the following steps: an edge-heuristic medical image segmentation network based on efficient parallel coding, such as... Figure 1As shown, the process includes an encoding stage, a decoding stage, and an edge information heuristic stage. In the encoding stage, features are extracted and encoded from the input image through two parallel encoder branches. In the decoding stage, the encoded and decoded features from the same stage are fused via skip connections, and the CNN decoder is used to obtain the decoded features for the next stage. In the edge information heuristic stage, edge information is extracted from features at different scales in the encoding stage, and this information is used to enhance the edge semantic representation of the target region in the decoding stage. Finally, the decoder outputs a predicted segmentation mask.

[0014] During the encoding phase, the parallel encoder structure comprises two branches: the first branch uses a CNN encoder to extract local detail features of the image; the second branch uses a KAN-Transformer encoder to extract global contextual information of the image and enhance the ability to model nonlinear relationships. The KAN-Transformer encoder is one of the key innovations of this invention.

[0015] This invention integrates the KAN network and wavelet transform into the Transformer architecture to form the KAN-Transformer encoder, as shown below. Figure 2 As shown. The specific implementation includes: (1) replacing the multilayer perceptron (MLP) structure in the Transformer block with a KAN layer structure; (2) replacing the linear mapping matrix of query Q, key K, and value V in the Transformer self-attention mechanism with a single-layer KAN structure. Through the above replacements, the core KAN-Transformer block is constructed, enabling it to maintain efficient nonlinear representation capabilities during feature extraction and processing.

[0016] Furthermore, to reduce computational complexity and utilize high-frequency detail information, this invention performs a KAN-Transformer operation after the wavelet transform: First, the spatial image features input to the KAN-Transformer block are subjected to a wavelet transform, decomposing them into low-frequency and high-frequency components; then, the low-frequency components are input into the KAN-Transformer block for processing, and the processed low-frequency and high-frequency components are then reintegrated into complete frequency domain features. Finally, the spatial domain features are reconstructed through an inverse wavelet transform. Since the spatial size of the wavelet subband is one-quarter the size of the original image, this process significantly reduces the computational complexity of subsequent KAN-Transformer block processing. Simultaneously, processing high-frequency information helps the model capture and enhance detail information in the image, which is beneficial for improving segmentation accuracy.

[0017] The specific processing flow of the parallel encoder is as follows: The input image is simultaneously fed into the CNN encoder branch and the KAN-Transformer encoder branch. In the CNN branch, the CNN encoder is used for feature extraction, and the image is downsampled sequentially (e.g., to 1 / 4, 1 / 8, 1 / 16 of the original size) to obtain multi-scale local feature maps, denoted as... , and In the KAN-Transformer branch, a pyramid structure is used for feature extraction and downsampling (e.g., by factors 4, 8, and 16). This captures hierarchical representations of the input image through wider receptive fields, resulting in feature maps. , and .

[0018] To effectively fuse the different features (local details and global context) extracted from the two parallel branches, this invention introduces a channel attention module to combine the local and global information obtained from the CNN and KAN-Transformer branches: .in and These are the first two branches of CNN and KAN-Transformer, respectively. Characteristics of each stage This represents a channel connection operation. Here, FcaNet channel attention is used for feature enhancement. Then, through skip connections, valuable information is passed from the encoder to the decoder, resulting in pixel-level segmentation results.

[0019] In the decoding stage, the decoded features from the same stage and the fused encoded features are connected through skip connections and then input into the CNN decoder for decoding to obtain the decoded features for the next stage.

[0020] To address the problem of blurred and discontinuous boundaries of target tissues in medical images, this invention designs an edge extraction module and an edge enhancement module in the edge information heuristic stage. This stage aims to utilize multi-scale features extracted in the encoding stage to generate target-related edge prior knowledge, thereby inspiring the decoder to more accurately segment boundary regions.

[0021] The edge extraction module (such as) Figure 3 The specific workflow of (as shown) is as follows: (1) Select feature maps of different scales in the encoding stage (usually containing low-level features) and high-level characteristics Low-level features contain rich details but are noisy, while high-level features have strong semantics but low spatial resolution. (2) To suppress high-frequency noise interference, low-level features are respectively... and high-level characteristics By applying Discrete Wavelet Transform (DWT), low-frequency information that better represents the main structure of the image can be extracted. , (3) Using the Sobel operator, edge information in the horizontal and vertical directions is extracted from each low-frequency subband to obtain the edge information of each feature: Then, these features are combined through channel concatenation operations and simple convolutional layers to obtain edge features: .

[0022] The edge enhancement module (such as) Figure 4 (As shown) is used to extract edge features The implied structural information is injected into the first Features of each decoding stage In order to enhance the semantic representation of the target edge region, this module is based on a local cross-attention mechanism: firstly, the input edge features are processed... and decoding features The system is divided into multiple local regions, and the features of each local region are projected into query (Q), key (K), and value (V) vectors, respectively. Attention scores are calculated using a dot product attention mechanism. The resulting vectors are then weighted and summed, and finally, a convolutional layer is applied to obtain the fused enhanced features. .

[0023] In summary, this invention first extracts multi-scale fused features using a parallel encoder (CNN + KAN-Transformer); secondly, it utilizes skip connections to fuse encoded features during the decoding process; and crucially, it extracts edge features from the encoded features through an edge information-inspired stage. And use the edge enhancement module to combine it with the decoded features Fusion, resulting in enhanced features Finally, the decoder outputs an accurate segmentation mask based on features that incorporate edge information. This mechanism effectively enhances the model's ability to perceive the semantics of target edges, thereby significantly improving the segmentation accuracy of medical image regions with blurred boundaries. Attached Figure Description

[0024] Figure 1 This is a framework diagram of an edge-heuristic medical image segmentation method based on efficient parallel coding.

[0025] Figure 2 This is a structural diagram of the KAN-Transformer.

[0026] Figure 3 This is a structural diagram of the edge extraction module.

[0027] Figure 4 This is a structural diagram of the edge enhancement module.

[0028] Figure 5 This is a visual comparison of the segmentation effects of the method of this invention with other methods. Detailed Implementation

[0029] To make the technical solution, objectives, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.

[0030] like Figure 1 As shown, this invention provides an edge-heuristic medical image segmentation method based on efficient parallel coding, which includes a parallel dual encoder of CNN and KAN-Transformer, a decoder, an edge extraction module, and an edge enhancement module.

[0031] like Figure 1 As shown, in the encoding part, the medical image is input into a parallel encoder consisting of a CNN encoder and a KAN-Transformer encoder to extract features and generate feature maps. For the CNN branch, we use ResNet as the backbone network to capture local details of the image. We downsample the feature maps to 1 / 4, 1 / 8, and 1 / 16 of the input image, respectively, which can be expressed by the following formula:

[0032]

[0033]

[0034]

[0035] in This indicates a downsampling operation. Indicates the first The CNN encoding operations at each stage, This represents the input image.

[0036] For the other branch using the pyramidal KAN-Transformer encoder, we downsample the feature maps in the order of factors 4, 8, and 16, thereby allowing a wider receptive field and capturing a hierarchical representation of the input image. The KAN-Transformer branch encoding can be described as follows:

[0037]

[0038]

[0039]

[0040] in Indicates the first Patch partitioning operations at each stage, Indicates the first KAN-Transformer encoding operations in each stage This represents the input image.

[0041] To effectively utilize information from different encoders during the decoder stage, this invention introduces a channel attention module that combines local and global information obtained from CNN and KAN-Transformer branches: .

[0042] in and These are the first two branches of CNN and KAN-Transformer, respectively. Characteristics of each stage This represents a channel connection operation. This approach utilizes FcaNet channel attention for feature enhancement. These features are then passed through skip connections to transfer valuable information from the encoder to the decoder, resulting in pixel-level segmentation.

[0043] like Figure 3 As shown, we use an edge extraction module to mine the edge features of the target region, and we combine low-level features ( ) and high-level characteristics ( We model the edge information related to the target region. To counteract the interference caused by high-frequency noise in the image, we first apply wavelet transform to these features to separate the low-frequency information that better represents the main structure of the image: , .in This indicates that discrete wavelet transform is performed on image features.

[0044] Next, the Sobel operator is used to extract the edge information in the horizontal and vertical directions from each low-frequency subband. The horizontal and vertical Sobel kernels are defined as follows:

[0045]

[0046] For each pixel, calculate the gradient in the horizontal and vertical directions. Combine these gradients to calculate the total gradient magnitude: .

[0047] After calculating the total gradient magnitude, the edge information for each feature is obtained: .

[0048] Then, these features are combined at different levels using simple convolutional layers to obtain the final edge features: .in, Represents the convolution operation. This represents a channel connection operation. This represents an upsampling operation.

[0049] like Figure 4 As shown, this invention designs an edge enhancement module that injects edge cues related to the segmentation boundary into the learning of the segmentation task to enhance the feature representation with object structural semantics. A Transformer block with a local cross-attention mechanism is used to adjust the features of neighboring regions, and the input edge features... and the Stage decoding features The data is divided into multiple local regions and then projected into query (Q), key (K), and value (V) vectors. Attention scores are calculated using a dot product attention mechanism. The resulting vectors are then weighted and summed, and finally passed through a convolutional layer to obtain the fused enhanced feature representation.

[0050]

[0051]

[0052] Learnable matrix , , These represent the key matrix, query matrix, and sum matrix, respectively. This represents element-wise addition. This represents normalization using the Softmax function. This represents a convolution operation. The edge information extracted by the edge extraction module is injected into the segmentation semantics through the edge enhancement module, and the predicted segmentation map is output by the last decoder.

[0053] Finally, the model's predicted mask is compared with the real mask, and the model is updated based on the constraints of the loss function. The model's total loss function is as follows:

[0054]

[0055] in, and These represent the Dice loss and cross-entropy loss, respectively. and For real segmentation labels and boundary labels, and For the predicted segmentation and boundary results, the parameters The value is 0.5.

[0056] The Adam optimizer is used to iteratively update the model parameters.

[0057] The trained segmentation model is used to perform actual medical image segmentation tasks.

[0058] To verify the effectiveness of the proposed method in medical image segmentation, the method was compared with eight existing deep learning-based medical image segmentation methods on four different modalities of medical image segmentation datasets (breast ultrasound image dataset, tissue slide dataset, dermoscopy dataset, and intestinal polyp dataset). The segmentation results are visualized as follows. Figure 5 As shown, the edge-heuristic medical image segmentation method based on efficient parallel coding proposed in this invention is significantly superior to other methods.

[0059] The method proposed in this invention is not only applicable to the field of medical image segmentation related to human diseases, but also has wide applicability in early warning and control of important diseases in aquaculture and crop cultivation. For example, by segmenting abnormal parts of fish body surfaces (such as ulcers, white spots, hemorrhages, etc.) and combining them with fish behavioral data (such as abnormal swimming, body posture and movement disorders, abnormal breathing, etc.), disease early warning and disease severity can be performed, providing a basis for disease control. By accurately segmenting diseased parts of crops, the location and severity of the disease can be analyzed, guiding the adjustment of control strategies.

Claims

1. A heuristic medical image segmentation method based on efficient parallel coding, characterized in that, Includes the following steps: Parallel encoding steps: The medical image is simultaneously input into the first encoder branch based on a convolutional neural network (CNN) and the second encoder branch based on a KAN-Transformer for feature extraction and encoding, respectively obtaining CNN-encoded features and KAN-Transformer-encoded features. The KAN-Transformer encoder replaces the multilayer perceptron (MLP) structure and the linear mapping layer used to generate query vector Q, key vector K, and value vector V in the standard Transformer block with a KAN structure. Discrete wavelet transform is performed on the input features to obtain high and low frequency components. The low frequency components are input into the KAN-Transformer block for processing and then integrated with the high frequency components to form frequency domain features. Finally, the features are reconstructed back to the spatial domain through inverse discrete wavelet transform. Encoding feature fusion step: The channel attention module is used to fuse the CNN encoding features output by the first encoder branch and the KAN-Transformer encoding features output by the second encoder branch to obtain multi-scale fused encoding features; Edge information extraction step: Use the edge extraction module to extract the edge features of the segmented target from the multi-scale fused encoded features; Feature decoding steps: After connecting the decoded features and the fused encoded features at the same stage through skip connections, the features are input into the CNN decoder for decoding to obtain multi-scale decoded features; Edge information enhancement step: Using the edge enhancement module, the edge features are fused with the multi-scale decoding features to obtain enhanced decoding features; Segmentation mask generation step: Based on the enhanced decoding features, generate the final segmentation mask image.

2. The method according to claim 1, characterized in that, The parallel encoding steps specifically include: In the first encoder branch, ResNet is used as the backbone network to extract features and downsample the input image, resulting in feature maps at least three scales. and Their resolutions are 1 / 4, 1 / 8, and 1 / 16 of the input image, respectively. In the second encoder branch, a pyramid-structured KAN-Transformer encoder extracts and downsamples the input image in the order of factors 4, 8, and 16, obtaining KAN-Transformer encoded features at at least three scales, denoted as feature maps. and 3. The method according to claim 1 or 2, characterized in that, The specific steps of the coding feature fusion are as follows: For the i-th encoding stage, the CNN encoded features output from the first encoder branch of this stage are fused using the channel attention module. KAN-Transformer encoded features from the output of the second encoder branch Obtain the fusion coding features at this stage Represented as: Here, Cat(·) represents the channel connection operation, and PCA(·) represents the channel attention operation based on FcaNet.

4. The method according to claim 1, characterized in that, The edge information extraction step specifically includes: selecting low-level features from the multi-scale fusion encoded features. and high-level characteristics To each and Perform discrete wavelet transform to extract their respective low-frequency approximate subbands. and Calculate using Sobel operators respectively In the horizontal direction G x and vertical direction G y The gradient, and In the horizontal direction G x and vertical direction G y The gradients are calculated, and their respective edge response maps are plotted. and These features are combined at different levels using simple convolutional layers to obtain edge features:

5. The method according to claim 1, characterized in that, Refine the segmentation results using boundary information: By utilizing the edge enhancement module, edge cues obtained from the edge extraction module are injected into the learning of the segmentation task to enhance the edge semantics of the decoded features, and ultimately accurately segment the target region.

Citation Information

Patent Citations

  • Medical image segmentation method based on edge-guided attention mechanism

    CN119295497A

  • Colorectal lesion typing method based on KAN network and variable attention mechanism

    CN119444700A