Colorectal lesion segmentation method based on improved UNet
By improving the UNet model and combining adaptive frequency-aware fusion, morphological attention, and global-local context integration modules, the problems of time-consuming and labor-intensive segmentation and insufficient generalization ability in colorectal lesion segmentation are solved, achieving efficient and accurate lesion segmentation.
Patent Information
- Application Number
- CN202511688277.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies for colorectal lesion segmentation suffer from problems such as time-consuming and labor-intensive manual annotation, significant subjective influence, and segmentation accuracy and speed that are difficult to meet clinical needs. Furthermore, traditional algorithms have insufficient generalization ability under different lighting conditions.
An improved UNet model is adopted, which introduces an adaptive frequency-aware fusion module, a morphological attention fusion module, and a global-local context integration module. Through frequency domain decomposition, dynamic dilated convolution, and attention mechanisms, feature extraction and lesion adaptability are enhanced, computational complexity is reduced, and segmentation accuracy and speed are improved.
It achieves high-precision segmentation of colorectal lesions under different lighting conditions, reduces the workload of doctors, improves segmentation efficiency and model generalization ability, and adapts to various lesion morphologies.
Smart Images

Figure CN121564010A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and more specifically to a method for colorectal lesion segmentation based on an improved UNet. Background Technology
[0002] Colorectal cancer is one of the most serious threats to human health. Its incidence and mortality rates rank third and second globally among cancers. Due to the complex and multifaceted nature of colorectal cancer, its specific pathogenesis is not fully understood. Initially, research on colorectal cancer focused almost entirely on treatment. Later, researchers realized that finding efficient early screening methods and improving the early diagnosis rate could significantly reduce the incidence and mortality rates of colorectal cancer through early intervention. Annotating and segmenting colorectal lesions in histopathological images is a crucial step and method for medical experts to determine the grade of colorectal cancer. This work is vital for patient treatment because accurate lesion segmentation facilitates targeted and personalized treatment design, thereby improving the cure rate. However, manually annotating and segmenting glandular cells by medical experts requires high concentration and is extremely laborious and time-consuming. Furthermore, because the judgment of colorectal lesions can be influenced by the subjective opinions of medical experts, prolonged fatigue can significantly increase the probability of life-threatening situations for patients. Therefore, clinical treatment places high demands on automated segmentation methods to improve efficiency and reliability and reduce the workload of medical experts.
[0003] In reviewing the research progress in colorectal lesion segmentation both domestically and internationally in recent years, it can be observed that despite a series of advancements, research has generally encountered several key challenges. First, traditional detection methods relying on manual feature extraction have limitations in the selection of training samples. These manually extracted features often only cover a limited scope of the detected object, restricting their representational ability and thus affecting the overall accuracy of segmentation. Second, the color and texture of lesions can change significantly under different lighting conditions, making it extremely difficult to develop a comprehensive feature extraction algorithm that can integrate numerous features of the lesion. Consequently, while some segmentation models may perform well on specific datasets, their generalization ability is significantly reduced. Furthermore, due to the complexity and time consumption of traditional algorithms, their segmentation speed and accuracy are difficult to meet the requirements of clinical applications, limiting their wider application in various scenarios. Summary of the Invention
[0004] The purpose of this invention is to provide a method for colorectal lesion segmentation based on an improved UNet, addressing the problem that doctors may overlook minor lesions during direct observation due to fatigue or lack of experience, leading to inaccurate diagnoses. Furthermore, the method enables simultaneous lesion detection and segmentation, allowing doctors to directly identify the lesion category.
[0005] This invention proposes a method for colorectal lesion segmentation based on an improved UNet, comprising the following steps: Step 1: Input endoscopic image data and introduce an adaptive frequency-aware fusion module designed specifically for medical images to build an encoder backbone network that can dynamically balance multi-scale contextual information and fine-grained boundary details. Its core idea is to enhance the diversity and specificity of features through frequency domain decomposition, reduce the computational redundancy of traditional convolution when dealing with heterogeneous lesions, and improve the model's adaptability to lesions of different sizes by dynamically adjusting the receptive field. Step 2: By adding a morphological attention fusion module, the complete morphological information in the encoder features is preserved and noise is suppressed. The morphological prior knowledge of erosion and dilation is transformed into a learnable collaborative attention mechanism, so that the spatial semantic features are evenly distributed under the guidance of morphological enhancement. The output features of the two encoding paths are further aggregated through cross-branch interaction to capture the precise boundary and internal structure of the lesion. Step 3: A global-local context integration module was added, and a reasonable global and local feature collaboration strategy was designed. This strategy models long-distance dependencies through a lightweight Transformer mechanism, while refining and preserving spatial details through local convolution. This reduces the boundary blurring problem caused by an excessively large global receptive field, and focuses more on the collaborative representation of the overall layout of the lesion and local details, thereby improving the network model's generalization ability and overall segmentation performance for complex lesions.
[0006] According to claim 1, the method for colorectal lesion segmentation based on improved UNet is characterized in that the specific details of step one are as follows: This application utilizes images captured by high-resolution endoscopic equipment to obtain information on colorectal lesions. Improvements to the U-Net model include replacing the standard convolutional blocks in all U-Net encoders with the adaptive frequency-aware fusion module, AdaFuse Block. Using AdaFuse Block as the foundation module, each module consists of frequency domain decomposition, dynamic dilation rate prediction, and multi-branch dilated convolution, further enhanced by parameter reallocation for more efficient feature tradeoffs. This foundation module reduces the limitations of fixed receptive fields and alleviates the information loss problem of traditional convolutions when processing multi-scale lesions. Simultaneously, specific convolutional operations are added to each branch for local information interaction and to help introduce inductive bias. An innovative frequency domain feature enhancement mechanism improves model performance while reducing computational complexity. The feature extraction method is optimized to better handle uneven illumination and complex textures in endoscopic images, achieving accurate feature extraction across various lesion morphologies, including micropolyps, squamous adenomas, and invasive carcinomas. Compared to traditional encoder networks, AdaFuse Block significantly improves adaptability to lesion size and morphology while maintaining high performance.
[0007] According to claim 2, a method for colorectal lesion segmentation based on an improved UNet is characterized in that the AdaFuse Block employs a novel frequency domain decomposition and fusion mechanism, which aims to improve the efficiency and effectiveness of the model in processing the diversity of lesions in endoscopic images. Unlike traditional convolutional networks, the AdaFuse Block achieves differentiated enhancement of edge details and global context by decomposing features into the frequency domain, which helps reduce noise interference. Edge details are separated by a high-pass filter, while the global structure is preserved by a low-pass filter, which is particularly important for accurately delineating lesion boundaries. The efficient operations in the AdaFuse Block mainly refer to the use of Fast Fourier Transform and dynamically dilated convolution in the model. These designs enable the AdaFuse Block to significantly reduce computational complexity while maintaining model performance. These operations include using frequency domain decomposition to replace traditional single-scale convolution and employing dynamically predicted dilation rates to adapt to lesion targets of different sizes.
[0008] According to claim 1, the method for colorectal lesion segmentation based on improved UNet is characterized in that the specific details of step two are as follows: A novel morphological attention fusion module, MAFM, without channel-level dimensionality reduction, is incorporated into the skip connections of U-Net to fuse features between the encoder and decoder. MAFM has only two core processing paths, corresponding to the erosion and dilation operations, respectively. Partial channel dimensions are reshaped into batch dimensions to avoid some form of dimensionality reduction through general convolution. In addition to constructing local cross-channel interactions in each parallel subnetwork without channel-level dimensionality reduction, the output feature maps of the erosion and dilation paths are fused through a collaborative attention learning method. For any given input feature map, MAFM binarizes the input feature map to highlight salient regions. In the erosion path, large-kernel pooling smooths the image and eliminates minor noise; in the dilation path, large-kernel pooling enhances image details and fills in subtle gaps. To achieve distinct morphological interaction features between the erosion and dilation paths, collaborative attention is computed on the activation feature maps of the two paths, strengthening feature interactions by generating a similarity matrix. On the other hand, the dilation path captures broader contextual information through large-kernel operations to expand the receptive field. By using an attention mechanism to generate morphological weights, global spatial location information is compressed into morphological descriptors. Parallel substructures help networks avoid more sequential processing and deep network requirements, effectively establishing short-term and long-term dependencies.
[0009] According to claim 1, the method for colorectal lesion segmentation based on improved UNet is characterized in that the specific details of step three are as follows: The Global-Local Context Integration (GLCIM) module is added to the bottleneck layer of the U-Net algorithm, introducing a global-local collaborative attention mechanism. This mechanism uses global context labels to evaluate the importance of local features and dynamically adjusts feature weights, transforming the original single-feature fusion into a collaborative integration of global and local features. GLCIM constructs a Transformer-based global attention mechanism and further refines it through local convolutions to add spatial details, reducing the loss of boundary details caused by global modeling. It focuses more on the collaborative representation of global context and local details, thereby improving the network model's generalization ability for large lesions and small lesions, as well as its overall segmentation performance. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of the research model architecture for the colorectal lesion segmentation method based on the improved UNet of this invention.
[0011] Figure 2 This is a structural diagram of the AdaFuse Block adaptive frequency sensing fusion module used in this invention.
[0012] Figure 3 This is a structural diagram of the morphological attention fusion module (MFAM) used in this invention.
[0013] Figure 4 This is a structural diagram of the Global-Local Context Integration Module (GLCIM) used in this invention. Detailed Implementation
[0014] For those skilled in the art, certain well-known structures and their descriptions in the accompanying drawings may be omitted. The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0015] This invention provides a method for colorectal lesion segmentation based on an improved version of UNet. This method improves upon UNet by introducing an adaptive frequency-aware fusion module, a morphological attention fusion module, and a global-local context integration module, thereby enhancing the segmentation accuracy of lesion regions in colorectal endoscopy images.
[0016] Figure 1 This invention presents the research model architecture for the colorectal lesion segmentation method based on an improved UNet. The overall structure adopts a symmetrical dual-branch design. Each branch undergoes hierarchical feature processing through multiple AdaFuse Blocks. The MAFM module facilitates cross-branch feature interaction and fusion, while the GLCIM module at the bottom performs the final integration of global and local context. Through advanced techniques such as adaptive feature fusion, multi-attention mechanisms, and context integration, efficient extraction and optimized fusion of complex features are achieved.
[0017] Figure 2 The adaptive frequency-aware fusion module AdaFuse Block used in this invention has an overall architecture with a three-branch parallel design. The input signal is decomposed into high-frequency and low-frequency components by FFT. The high-frequency path processes details through a dynamic dilation network, the low-frequency path extracts structural features through convolution, pooling and upsampling, and the third path directly performs 1×1 convolution. The three are finally fused and output to achieve collaborative optimization of frequency domain and spatial domain features.
[0018] Figure 3 The Morphological Attention Fusion (MFAM) module used in this invention divides the input signal "X" into two parallel paths: the upper path passes through Softmax, MaxPool, and Tanh before entering the Co-Attention module; the lower path passes through Softmax, subtracts the output from the upper path, enters MaxPool, adds the output of the sigmoid, and then processes through Linear and sigmoid. Finally, the output of Co-Attention is fused with the sigmoid output of the lower path and combined with the original input "X" through a multiplication operation to form the final output.
[0019] Figure 4The Global-Local Context Integration (GLCIM) module used in this invention processes input data through three paths: the left side extracts global features through global average pooling, the middle side preserves local information through patching, and the right side directly performs linear projection. The global and local features are input into the core "Global-Local Attention" module for fusion. Its output is copied and linearly transformed, then added to the linear projection result of the right branch, and finally generated again through linear projection.
[0020] The specific implementation steps are as follows: Step 1.1 Input endoscopic image data and input the image data into the encoder for feature extraction. Through the adaptive frequency-aware fusion module, frequency domain decomposition, dynamic hole rate prediction and multi-branch hole convolution processing are performed on the input feature map to generate an enhanced multi-scale feature map. Step 1.2 Decompose the input feature map into high-frequency components and low-frequency components using Fast Fourier Transform; Step 1.3 are the spatially adaptive void rate maps for the high-frequency component and the low-frequency component, respectively, wherein the void rate range for the high-frequency component prediction is lower than that for the low-frequency component prediction. Step 1.4 Using the predicted dilatation rate map, perform dilated convolution on the high-frequency components and low-frequency components respectively, while retaining a pass-through convolution branch; Step 1.5 concatenates the outputs of the high-frequency branch, low-frequency branch, and direct-through branch along the channel dimension and performs dimensionality reduction fusion to generate the enhanced multi-scale feature map. As shown in Equation 1: (1) Step 2.1: Through the morphological attention fusion module, the complete morphological information in the encoder features is preserved and noise is suppressed. The output features of the two encoding paths are further aggregated through cross-branch interaction to capture the precise boundary and internal structure of the lesion. Step 2.2 Binarize the input feature map and generate eroded and dilated feature maps by performing equivalent erosion and dilation operations, respectively, as shown in Formulas 2 and 3: (2) (3) Step 2.3 Calculate the collaborative attention weights between the erosion feature map and the dilation feature map; Step 2.4 The collaborative attention weights are used to weight the feature map that fuses the erosion features and the original features to generate a morphologically enhanced cross-branch fusion feature map; Step 3.1: Add a global-local context integration module to model long-distance dependencies through a lightweight Transformer mechanism, while refining spatial details through local convolution to reduce the boundary blurring problem caused by an excessively large global receptive field; Step 3.2 Generates a global context label by global average pooling, and at the same time divides the input feature map into non-overlapping image blocks and projects them as local feature labels; Step 3.3 Using the global context marker as the query and the local feature marker as the key and value, a context-enhanced feature vector is generated through a self-attention mechanism. Step 3.4 The feature vector is copied in the spatial dimension and fused with the original input feature map through residual connection to generate a global context integrated feature map.
Claims
1. A study on a method for segmenting colorectal lesions based on an improved UNet, comprising the following steps: Step 1: Input endoscopic image data and input the image data into the encoder for feature extraction. Through the adaptive frequency-aware fusion module, frequency domain decomposition, dynamic hole rate prediction and multi-branch hole convolution processing are performed on the input feature map to generate an enhanced multi-scale feature map. Step 2: Through the morphological attention fusion module, the complete morphological information in the encoder features is preserved and noise is suppressed. The output features of the two encoding paths are further aggregated through cross-branch interaction to capture the precise boundary and internal structure of the lesion. Step 3: Add a global-local context integration module to model long-distance dependencies through a lightweight Transformer mechanism, while refining spatial details through local convolution to reduce the boundary blurring problem caused by an excessively large global receptive field.
2. The research on the method for colorectal lesion segmentation based on improved UNet according to claim 1, wherein the specific process in Step 1 is as follows: Step 1.1 Decompose the input feature map into high-frequency components and low-frequency components using Fast Fourier Transform; Step 1.2 are the spatially adaptive void rate maps for the high-frequency component and the low-frequency component, respectively, wherein the void rate range for the high-frequency component prediction is lower than that for the low-frequency component prediction. Step 1.3 Using the predicted dilatation rate map, perform dilated convolution on the high-frequency components and low-frequency components respectively, while retaining a pass-through convolution branch; Step 1.4: The outputs of the high-frequency branch, low-frequency branch, and through branch are concatenated along the channel dimension and then fused to generate the enhanced multi-scale feature map.
3. The research on the method for colorectal lesion segmentation based on improved UNet according to claim 1, wherein the specific process in Step 2 is as follows: Step 2.1 Binarize the input feature map and generate eroded and dilated feature maps by performing equivalent erosion and dilation operations respectively. Step 2.2 Calculate the collaborative attention weights between the erosion feature map and the dilation feature map; Step 2.3 The collaborative attention weights are used to weight the feature map that fuses the erosion features and the original features to generate a morphologically enhanced cross-branch fusion feature map.
4. In the study of the method for colorectal lesion segmentation based on the improved UNet according to claim 1, the specific process in Step 3 is as follows: Step 3.1 Generate a global context label by global average pooling, and at the same time divide the input feature map into non-overlapping image blocks and project them into local feature labels; Step 3.2 Using the global context marker as the query and the local feature marker as the key and value, a context-enhanced feature vector is generated through a self-attention mechanism. Step 3.3: Copy the feature vector in the spatial dimension and fuse it with the original input feature map through residual connection to generate a global context integrated feature map.