Multi-scale polyp segmentation method and segmentation system based on context awareness

Through a polyp segmentation network based on U-shaped architecture, combined with multi-scale interaction and context-aware repair modules, the problems of low accuracy and poor generalization ability in polyp segmentation are solved, and more accurate polyp segmentation and recognition are achieved, supporting clinical diagnosis and treatment.

CN120259641AActive Publication Date: 2025-07-04SOUTHWEAT UNIV OF SCI & TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510135069.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-07-04
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing polyp segmentation method is not highly segmented when dealing with small or blurred polyps, the decoder cannot accurately locate the segmentation target, and the generalization ability is poor. It fails to fully consider the characteristics of polyp images such as different sizes and diverse morphology, resulting in poor segmentation results.

Method used

A polyp segmentation network based on the U-shaped architecture is adopted, combined with the pyramid vision transformer, multi-scale interaction module, spatial attention enhancement module and context-aware repair module, through multi-scale feature interaction, spatial attention enhancement and edge area processing, the segmentation boundaries are gradually optimized, and the weighted intersection ratio and binary cross-entropy loss function are introduced for supervision.

Benefits of technology

More precise polyp segmentation is achieved, segmentation robustness and accuracy are improved, and the polyp size and morphology are adapted to the characteristics of different polyps, and the ability to identify polyp boundaries is enhanced, and clinical diagnosis and treatment decisions are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259641A_ABST
    Figure CN120259641A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale polyp segmentation method and segmentation system based on context perception, and belongs to the technical field of medical image processing. Comprising the steps of obtaining a to-be-segmented pathological image, inputting the to-be-segmented pathological image into a pre-constructed context sensing multi-scale polyp segmentation model, and outputting a segmentation result to complete polyp segmentation. According to the method, the problem that small polyps and flat polyps are easy to lose in the segmentation process is solved, the defect of detail processing based on a pure Transform method is overcome, the defect that a decoder cannot accurately position the segmentation target position during image restoration is overcome, the challenge that the polyp boundary is difficult to accurately segment is solved, and the segmentation efficiency is improved. The problem that the generalization ability of the polyp segmentation network is poor is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and more specifically, to a multi-scale polyp segmentation method and segmentation system based on context awareness. Background Art

[0002] Currently, the method based on convolutional neural network (CNN) is represented by UNet. UNet consists of a contracting path (encoder) and a symmetric expanding path (decoder), which extracts features and captures context information through layer-by-layer convolution and pooling. The expanding path restores the image spatial resolution through upsampling and concatenation, and its skip connections combine low-level and high-level features, improving the segmentation effect. However, there are some problems in the encoder-decoder architecture of CNN. In the encoding stage, the reduction of feature scale leads to the loss of small targets and details; while in the decoding stage, it is more difficult to restore details. Although skip connections can fuse low-resolution features to supplement details, the inconsistency of feature representation causes the semantic gap between the encoder and the decoder, and at the same time introduces background noise. Therefore, CNN performs well in capturing local details, but is relatively weak in summarizing global context, making it difficult to accurately segment small polyps.

[0003] In addition, the method based on pure Transformer uses the self-attention mechanism to handle long-range dependencies for efficient feature extraction. The input image is divided into windows, and self-attention is applied between each window to capture global features. This method uses an encoder-decoder structure to gradually restore the spatial resolution and achieve image segmentation. However, pure Transformer usually relies on large-scale datasets for pre-training, and its position encoding and affinity matrix calculation require more complex processing and computing resources. Although position encoding compensates for the lack of spatial position perception ability of Transformer to a certain extent, it is still not as directly effective as convolutional operations. At the same time, although Transformer performs well in modeling global context and capturing long-range dependencies, due to the lack of local inductive bias of CNN, its performance in segmenting object boundaries and processing fine structures is relatively poor.

[0004] Furthermore, the method based on the hybrid architecture of CNN and Transformer utilizes the advantages of CNN in capturing local features and texture information, as well as the ability of Transformer in capturing global context and long-range dependencies. Although these hybrid architectures have made great progress in feature extraction of polyp images, many methods only focus on improving the encoder part of the architecture, while ignoring the skip connections and decoder part. This results in the decoder being unable to accurately locate the specific position of the segmentation target at the initial stage of the decoding stage, and the lack of recovery of detail information, making the segmentation result perform poorly when dealing with small or boundary-blurred polyps.

[0005] In summary, the current polyp segmentation methods fail to fully consider the characteristics of polyp images (such as different sizes, various shapes, blurred boundaries, etc.), resulting in low segmentation accuracy. At the same time, these methods are disconnected from the actual clinical operation process and do not integrate the whole process of clinical polyp resection into the architecture design, so the effect is not good. In addition, due to the diverse sources of polyp image datasets, the performance of some methods varies on different public datasets, and the generalization ability is poor.

[0006] Therefore, how to provide a context-aware multi-scale polyp segmentation method and segmentation system that can solve the above problems is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a context-aware multi-scale polyp segmentation method and segmentation system, which solves the problem that small polyps and flat polyps are easily lost during the segmentation process, solves the deficiency of the pure Transformer method in detail processing, solves the defect that the decoder cannot accurately locate the segmentation target position when restoring the image, solves the challenge of accurately segmenting the polyp boundary, and solves the problem of poor generalization ability of the polyp segmentation network.

[0008] To achieve the above object, the present invention provides the following technical solutions: A context-aware multi-scale polyp segmentation method, including obtaining a pathological image to be segmented, inputting the pathological image to be segmented into a pre-constructed context-aware multi-scale polyp segmentation model, and outputting a segmentation result to complete polyp segmentation; wherein, the context-aware multi-scale polyp segmentation model is based on a polyp segmentation network with a U-shaped architecture, and is composed of a pyramid vision transformer as an encoder, a multi-scale interaction module as a skip connection, a spatial attention enhancement module and a context-aware repair module as a decoder; the pyramid vision transformer extracts features from the input pathological image to be segmented; the multi-scale interaction module performs multi-scale feature interaction processing on the features extracted by the pyramid vision transformer; the spatial attention enhancement module highlights the important regions in the extracted multi-scale features by fusing parallel global attention mechanism units and local attention mechanism units; the context-aware repair module detects the edge region of the pathological image to be segmented and gradually processes the segmentation boundary.

[0009] Furthermore, the Pyramid Vision Transformer extracts features from the input pathological image to be segmented, obtaining four different levels of semantic features. Subsequently, the features obtained from each layer are respectively passed to the multi-scale interaction module of each layer, and a multi-scale convolutional module is used to achieve cross-learning between different receptive fields, resulting in: ; where, is a 3×3 convolution, is a 5×5 convolution, is a 1×1 convolution, is the multi-scale feature, and are the features obtained through the parallel operations of 3×3 convolution and 5×5 convolution respectively.

[0010] Furthermore, the multi-scale interaction module performs information interaction on the two features and multiplies them with the convolutions with receptive fields of 3 and 5 to obtain feature combination. The specific process is expressed as: ; where, is the feature obtained from the complementary information under different receptive fields, represents matrix multiplication, cat represents concatenation, CRB includes a 3×3 convolution, ReLU loss, and batch normalization operations, is to , and combine and add the original feature to obtain the final output. Furthermore, the specific process of the global attention mechanism unit in the spatial attention enhancement module is expressed as: ; where, the feature processed by the multi-scale interaction module in the 4th layer is evenly split into two sub-features and along the channel dimension, represents the input feature of the global spatial attention mechanism operation; Softmax represents the normalization operation; Flatten represents flattening the input two-dimensional feature map into a one-dimensional vector; represents applying convolution to the input feature map to obtain the context mask; GSA represents the global spatial attention module; AttG represents the global attention operation; represents matrix multiplication; the first layer of MLP converts the input into a higher-dimensional space with an expansion ratio of 2, and the second layer restores the dimension to match the input.

[0011] Furthermore, the specific process of the local attention mechanism unit in the spatial attention enhancement module is expressed as: ; In the formula, represents the input feature of the local spatial attention mechanism operation; is the Sigmoid activation function, AttL represents the local attention operation, represents element-wise multiplication, and LSA represents the local spatial attention module.

[0012] Furthermore, the specific process of the context-aware repair module is expressed as: ; In the formula, represents the generated foreground attention feature; represents the upsampling operation, represents the prediction mask of the previous layer; is the Sigmoid activation function, represents element-wise multiplication, represents the current layer feature; represents the background attention feature.

[0013] A context-aware multi-scale polyp segmentation system includes: Data acquisition module: Acquire the pathological image to be segmented; Segmentation module: Input the pathological image to be segmented into a pre-constructed context-aware multi-scale polyp segmentation model, output the segmentation result, and complete polyp segmentation; Among them, the context-aware multi-scale polyp segmentation model is based on a polyp segmentation network with a U-shaped architecture, consisting of a pyramid vision transformer as the encoder, a multi-scale interaction module as the skip connection, a spatial attention enhancement module and a context-aware repair module as the decoder; The multi-scale interaction module performs multi-scale feature interaction processing on the features extracted by the pyramid vision transformer; The spatial attention enhancement module highlights the important regions in the extracted multi-scale features by fusing parallel global spatial attention mechanism units and local spatial attention mechanism units; The context-aware repair module detects the edge region of the pathological image to be segmented and gradually processes the segmentation boundary.

[0014] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a context-aware multi-scale polyp segmentation method and segmentation system. The present invention combines the advantages of convolutional neural networks (CNNs) and Transformers to better adapt to the characteristics of polyps with different sizes, various shapes, and blurred boundaries, thereby achieving more accurate segmentation. Using a pyramid vision transformer as the encoder can effectively capture the global information of the image; the multi-scale interaction convolutional block in the skip connection realizes cross-learning between different receptive fields and enhances the feature expression ability. The decoder first performs preliminary spatial localization on the features to be restored through a spatial attention enhancement module, and then the context-aware repair module focuses on the boundary region and gradually optimizes the accuracy of boundary segmentation. The specific beneficial effects are as follows: (1) Based on the polyp segmentation network with a U-shaped architecture, the present invention introduces a multi-scale interaction module to further perform interactive learning on the features obtained by the encoder through different receptive fields, further enhancing the feature representation ability.

[0015] (2) The spatial attention enhancement module in the present invention emphasizes the important regions in the extracted features by fusing parallel global attention mechanisms and local attention mechanisms, thereby performing preliminary localization on the segmentation target.

[0016] (3) The context-aware repair module in the present invention pays attention to the edge region of the segmented image in a context-aware manner and gradually improves the segmentation boundary to solve the problem of inaccurate segmentation caused by over-segmentation and under-segmentation.

[0017] (3) The entire decoder of the present invention is combined with a weighted intersection over union (IoU) loss and a weighted binary cross-entropy (BCE) loss to supervise the segmentation results of each round, ensuring that the segmentation results are close to the true mask. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0019] Figure 1 is a schematic flowchart of the method of the present invention; Figure 2 is an overall architecture diagram of the context-aware multi-scale polyp segmentation model provided by the present invention; FIG. 3(a) is a schematic diagram of the multi-scale interaction module provided by the present invention; FIG. 3(b) is a schematic diagram of the spatial attention enhancement module provided by the present invention; Figure 3 (c) is a schematic diagram of the context-aware repair module provided by the present invention; Figure 3 (d) is a schematic diagram of the context exploration module provided by the present invention; Figure 4 is a schematic diagram of the system structure of the present invention. Specific embodiments

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] See Figure 1 , a multi-scale polyp segmentation method based on context awareness, including: Obtain the pathological image to be segmented, input the pathological image to be segmented into a pre-constructed context-aware multi-scale polyp segmentation model, output the segmentation result, and complete polyp segmentation; Among them, the context-aware multi-scale polyp segmentation model is based on a polyp segmentation network with a U-shaped architecture, and is composed of a pyramid vision transformer as the encoder, a multi-scale interaction module as the skip connection, a spatial attention enhancement module and a context-aware repair module as the decoder; The pyramid vision transformer extracts features from the input pathological image to be segmented; The multi-scale interaction module performs multi-scale feature interaction processing on the features extracted by the pyramid vision transformer; The spatial attention enhancement module highlights the important regions in the extracted multi-scale features by fusing parallel global attention mechanism units and local attention mechanism units; The context-aware repair module detects the edge region of the pathological image to be segmented and gradually processes the segmentation boundary.

[0022] In a specific embodiment, the overall architecture of the context-aware multi-scale polyp segmentation model is as Figure 2As shown in the figure, a segmentation network based on a U-shaped architecture is adopted. This network consists of a Pyramid Vision Transformer as the encoder, a multi-scale interaction convolutional block as the skip connection, a spatial attention enhancement module, and a context-aware restoration module as the decoder. In addition, the loss function includes a weighted Intersection over Union (IoU) loss and a weighted Binary Cross Entropy (BCE) loss. Specifically, the encoder is responsible for capturing the global context information of the input image; the skip connection establishes a direct feature channel between the encoder and the decoder, enabling the effective fusion of low-level detailed features and high-level semantic features; the decoder gradually restores the spatial resolution of the image and uses upsampling and feature stitching techniques to generate a segmentation mask of the same size as the input image. At the same time, the loss function is used to supervise the segmentation results of each round of the decoder.

[0023] Specifically, in the first stage (observation), this embodiment uses a Pyramid Vision Transformer as the encoder to extract the features of the polyp image. The designed multi-scale interaction module further performs multi-scale aggregation on the features through convolution to adapt to the scale and morphological changes of the target. In the second stage (localization), at the fourth layer of the decoder, this embodiment uses a position attention enhancement module to locate the specific position of the polyp from the high-level features by combining global and local spatial information. In the third stage (focusing), after obtaining the preliminary localization, in order to achieve more accurate target segmentation, this embodiment uses a context-aware restoration module in the first three layers of the decoder to perform multi-scale context exploration based on the attention features of the foreground and background to solve the problems of missing segmentation or over-segmentation, thereby obtaining more accurate segmentation results. The overall architecture of the context-aware multi-scale polyp segmentation model of the present invention is as Figure 2 shown Figure 2 a detailed diagram of the modules used in the context-aware multi-scale polyp segmentation model. The entire model adopts a segmentation network with a U-shaped architecture, including an encoder, a decoder, and a skip connection. By imitating the process of clinicians removing polyps, the present invention designs a new architecture, effectively solving the problem of low polyp segmentation accuracy in existing methods and ensuring that the segmentation results are consistent with the morphological features of actual medical images, thereby improving the robustness and accuracy of the segmentation.

[0024] Regarding the overall process from the input image to the segmentation result, as shown in FIGS. 3(a)-(b) is a detailed flowchart of the context-aware multi-scale polyp segmentation model provided by the present invention; specifically, first, the resolution of the polyp pictures in the dataset is adjusted to 352×352, and then input into the encoder. In order to better extract the global features of the pictures and prevent the loss of small target polyp features, in this embodiment, the encoder uses a pre-trained Pyramid Vision Transformer to utilize the rich visual features obtained from training on a large-scale dataset. The encoder extracts features from the input pictures and obtains four different levels of semantic features Subsequently, the features obtained for each layer are respectively passed to the multi-scale interaction modules of each layer. As shown in Fig. 3(a), since the polyp shapes are diverse, the feature information extracted relying on a single receptive field will have an adverse impact on image understanding and segmentation. To better adapt to the diverse polyp shapes, in this embodiment, a multi-scale convolution module is adopted and the idea of interaction is introduced, aiming to achieve cross-learning between different receptive fields and enhance the complementarity of features.

[0025] Specifically, in this embodiment, the features are first used with three independent 1×1 convolutions to obtain three feature representations. Then, through parallel operations of convolutions with different receptive fields, and are obtained. Then, these two features undergo information interaction, and subsequently, parallel convolution processing and element-wise multiplication are performed to fully utilize the complementary information under different receptive fields to obtain . Next, in this embodiment, the obtained , and are combined together and then input into the CRB module (including a 3×3 convolution, ReLU, and batch normalization). Finally, to retain the original information, the original features are superimposed on it to obtain the final output . The entire process can be expressed as: ; In the formula, is a 3×3 convolution, is a 5×5 convolution, is a 1×1 convolution, is a multi-scale feature, and are respectively the features obtained through parallel operations of the 3×3 convolution and the 5×5 convolution; is the feature obtained from the complementary information under different receptive fields, represents matrix multiplication, cat represents concatenation, CRB includes a 3×3 convolution, ReLU loss, and batch normalization operations, is to , and are combined and added with the original feature to obtain the final output.

[0026] In a specific embodiment, the features after skip connection processing are fed into the decoder, and a spatial attention enhancement module is used in the fourth layer of the decoder to integrate features of global and local spatial information to enhance the spatial features of the highest-level semantic information. This makes the position of the segmentation target more prominent in the feature map, contributing to the preliminary localization of the segmentation target. As shown in Fig. 3(b), the present invention designs this module to process the input in parallel through channel separation, which not only reduces the computational complexity but also allows the module to focus on different subsets of features.

[0027] In a specific embodiment, the feature map is split into two sub-feature maps along the channel dimension , and then they are respectively input into the global spatial attention module and the local spatial attention module. Finally, the features obtained by the two modules are fused together. Next, this embodiment will describe the global spatial attention module and the local spatial attention module in detail: Specifically, the global spatial attention module (GSA): The global spatial attention module enables the model to dynamically adjust the weights of different parts of the feature map, enhance important features, and make the approximate position of the segmentation target more prominent. Taking as the input, the specific process can be expressed as: ; where the feature processed by the multi-scale interaction module in the fourth layer is and averaged and split into two sub-features along the channel dimension, represents the input feature for the global spatial attention mechanism operation; Softmax represents the normalization operation; Flatten represents flattening the input two-dimensional feature map into a one-dimensional vector; represents applying a convolution to the input feature map to obtain a context mask; GSA represents the global spatial attention module; AttG represents the global attention operation; represents matrix multiplication; the first layer of the MLP converts the input into a higher-dimensional space with an expansion ratio of 2, and the second layer restores the dimension to match the input.

[0028] Specifically, the local spatial attention module (LSA): The local spatial attention module helps the model focus on the key details in the image and addresses the limitations of the global spatial attention module. Taking as the input, the specific process can be expressed as: ; Here, represents the input feature for the local spatial attention mechanism operation; is the Sigmoid activation function, AttL represents the local attention operation, denotes element-wise multiplication, and LSA represents the Local Spatial Attention module.

[0029] In a specific embodiment, polyps are often difficult to be completely separated from the surrounding normal tissues, resulting in irregular and blurred edges. This usually leads to under-segmentation or over-segmentation during the initial segmentation process. The context-aware repair module is used in the first three layers of the decoder, as shown in Fig. 3(c), aiming to gradually improve these problems through context analysis, thereby improving the segmentation accuracy. When humans evaluate the boundaries of fuzzy objects, they often repeatedly compare the uncertain regions with the clear regions and analyze from multiple perspectives to finally determine the boundaries. Inspired by this process, the present invention introduces a context exploration mechanism in the predicted foreground and background regions. By using convolutional kernels with different receptive fields to evaluate the fuzzy boundaries from multiple angles and cross-validating with the information in the surrounding regions. The context-aware repair module utilizes the high-level prediction and the current layer features to perform context exploration. By upsampling and normalizing the high-level prediction and combining it with the current layer features, the foreground attention feature and the background attention feature are generated. The specific process can be expressed as: ; In the formula, represents the generated foreground attention feature; represents the upsampling operation, represents the prediction mask of the previous layer; is the Sigmoid activation function, denotes element-wise multiplication, represents the current layer features; represents the background attention feature.

[0030] In a specific embodiment, and are input into the context exploration module (CE), which identifies and improves the under-segmentation or over-segmentation problems through multi-scale convolution.

[0031] Specifically, as shown in Fig. 3(d), CE consists of four parallel branches, and each branch uses convolutional kernels of different sizes (1, 3, 5, 7) and dilated convolutions (dilation rates of 1, 2, 4, 8) to extract multi-scale features. The output of each layer is combined with the previous layer and concatenated in the channel dimension, thereby generating a feature map with rich context information. After identifying the under-segmented or over-segmented regions, element-wise operations are used to correct the segmentation errors. For over-segmented regions, subtraction is used to suppress over-segmentation, while for under-segmented regions, addition is used to enhance under-segmentation. The process is as follows: ; wherein, BR represents batch normalization and ReLU, and are learnable scaling parameters.

[0032] In a specific embodiment, in order to supervise the prediction quality of each stage of the decoder output, a multi-stage joint loss function is utilized. The loss function of this embodiment can be written as: ; The loss function of each stage is a combined function, composed of a weighted intersection over union (IoU) loss and a weighted binary cross-entropy (BCE) loss, and can be expressed as: ; wherein, represents the predicted mask of the i-th stage of the decoder, and G is the ground truth annotation of the input training image. The present invention evaluates the segmentation performance on five publicly available polyp datasets, namely Kvasir-SEG, ClinicDB, ColonDB, Endoscene, and ETIS. The evaluation metrics include average Dice, average IoU, and mean absolute error (MAE). The average Dice coefficient is used to measure the overlap degree between the predicted segmentation region and the ground truth label, and the closer the value is to 1, the better the segmentation effect; the average IoU calculates the ratio of the intersection to the union of the predicted segmentation region and the ground truth segmentation region, and the larger the value, the higher the segmentation accuracy; while the MAE measures the average error between the prediction result and the ground truth label, and the lower the value, the more accurate the prediction. On the Kvasir-SEG dataset, the average Dice, average IoU, and mean absolute error (MAE) are 0.921, 0.875, and 0.022 respectively; on the ClinicDB dataset, the average Dice, average IoU, and mean absolute error (MAE) are 0.940, 0.894, and 0.006 respectively; on the ColonDB dataset, the average Dice, average IoU, and mean absolute error (MAE) are 0.816, 0.733, and 0.028 respectively; on the Endoscene dataset, the average Dice, average IoU, and mean absolute error (MAE) are 0.911, 0.848, and 0.006 respectively; on the ETIS dataset, the average Dice, average IoU, and mean absolute error (MAE) are 0.801, 0.722, and 0.014 respectively; On the other hand, as shown in Figure 4 this embodiment also discloses a context-aware multi-scale polyp segmentation system, including: Data acquisition module: acquiring the pathological image to be segmented; Detection and segmentation module: inputting the pathological image to be segmented into a pre-constructed multi-scale segmentation model, outputting a segmentation result, and completing polyp segmentation; Among them, the multi-scale segmentation model is based on a polyp segmentation network with a U-shaped architecture, and is composed of a pyramid vision transformer as the encoder, a multi-scale interaction module as the skip connection, a spatial attention enhancement module and a context-aware repair module as the decoder; Among them, the multi-scale interaction module performs multi-scale feature interaction processing on the features extracted by the pyramid vision transformer; The spatial attention enhancement module learns and extracts important regions in the multi-scale features by fusing parallel global attention mechanism units and local attention mechanism units; The context-aware repair module detects the edge region of the pathological image to be segmented and gradually processes the segmentation boundary.

[0033] The context-aware multi-scale polyp segmentation model of the present invention combines the advantages of convolutional neural networks and Transformers, better adapts to the characteristics of polyps with different sizes, various shapes, and blurred boundaries, and thus realizes more accurate segmentation. The pyramid vision transformer is used as the encoder to effectively capture the global information of the image; the multi-scale interaction module used in the skip connection realizes cross-learning between features with different receptive fields and enhances the feature expression ability. The decoder first performs preliminary spatial positioning on the features to be restored through the spatial attention enhancement module, and then the context-aware repair module focuses on the boundary region and gradually optimizes the accuracy of boundary segmentation.

[0034] The method of the present invention can be widely applied to polyp image segmentation and diagnosis of related diseases in endoscopy. By introducing a multi-scale interaction module, a spatial attention enhancement module and a context-aware repair module, the model significantly enhances the fine recognition ability of polyp morphology, which helps to improve the early detection and segmentation accuracy of polyps. This application not only improves the efficiency of medical image analysis, but also provides clinicians with clearer polyp position and boundary information, thus providing strong support for subsequent treatment decisions and clinical interventions, and further improving the prognosis and treatment effect of patients.

[0035] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0036] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A context-aware multi-scale polyp segmentation method, characterized in that Including: Obtain the pathological image to be segmented, input the pathological image to be segmented into a pre-constructed context-aware multi-scale polyp segmentation model, output the segmentation result, and complete polyp segmentation; Among them, the context-aware multi-scale polyp segmentation model is based on a polyp segmentation network with a U-shaped architecture, and consists of a pyramid vision transformer as the encoder, a multi-scale interaction module as the skip connection, a spatial attention enhancement module, and a context-aware repair module as the decoder; The pyramid vision transformer extracts features from the input pathological image to be segmented; The multi-scale interaction module performs multi-scale feature interaction processing on the features extracted by the pyramid vision transformer; The spatial attention enhancement module highlights the important regions in the extracted multi-scale features by fusing parallel global attention mechanism units and local attention mechanism units; The context-aware repair module detects the edge regions of the pathological image to be segmented and gradually processes the segmentation boundaries.

2. The multi-scale polyp segmentation method based on context awareness according to claim 1, wherein The pyramid vision transformer extracts features from the input pathological image to be segmented, obtains four different levels of semantic features. Subsequently, the features obtained by each layer are respectively passed to the multi-scale interaction module of each layer, and a multi-scale convolution module is used to achieve cross-learning between different receptive fields, and obtain: ; Wherein, is a 3×3 convolution, is a 5×5 convolution, is a 1×1 convolution, is a multi-scale feature, and are features obtained by parallel operations of 3×3 convolution and 5×5 convolution respectively.

3. A context-aware multi-scale polyp segmentation method according to claim 2, characterized in that, The multi-scale interaction module performs information interaction on the two features, and multiplies the convolution with a receptive field of 3 and a receptive field of 5 to obtain a feature combination. The specific process is expressed as: ; Wherein, is the feature obtained from complementary information under different receptive fields, represents matrix multiplication, cat represents concatenation, and CRB includes a 3×3 convolution, ReLU loss, and batch normalization operation, is to , and combine and add the original feature to obtain the final output.

4. A context-aware multi-scale polyp segmentation method according to claim 1, characterized in that, The specific process of the global attention mechanism unit in the spatial attention enhancement module is expressed as: ; Among them, the features after the 4th layer is processed by the multi-scale interaction module are averaged and split into two sub-features along the channel dimension and , represents the input feature of the global spatial attention mechanism operation; Softmax represents the normalization operation; Flatten represents flattening the input two-dimensional feature map into a one-dimensional vector; represents applying convolution to the input feature map to obtain a context mask; GSA represents the global spatial attention module; AttG represents the global attention operation; represents matrix multiplication; the first layer of the MLP converts the input into a higher-dimensional space with an expansion ratio of 2, and the second layer restores the dimension to match the input.

5. A context-aware multi-scale polyp segmentation method according to claim 1, characterized in that, The specific process of the local attention mechanism unit in the spatial attention enhancement module is expressed as: ; In the formula, represents the input feature of the local spatial attention mechanism operation; is the Sigmoid activation function, AttL represents the local attention operation, represents element-wise multiplication, and LSA represents the local spatial attention module.

6. A context-aware multi-scale polyp segmentation method according to claim 1, characterized in that According to the context-aware multi-scale polyp segmentation method described in claim 1, wherein the specific process of the context-aware repair module is expressed as: ; Wherein, represents the generated foreground attention feature; represents the upsampling operation, represents the prediction mask of the previous layer; is the Sigmoid activation function, represents element-wise multiplication, represents the current layer feature; represents the background attention feature.

7. A context-aware multi-scale polyp segmentation system using the context-aware multi-scale polyp segmentation method according to any one of claims 1-6, characterized in that, Including: Data acquisition module: Obtain the pathological image to be segmented; Segmentation module: Input the pathological image to be segmented into a pre-constructed context-aware multi-scale polyp segmentation model, output the segmentation result, and complete polyp segmentation; Among them, the context-aware multi-scale polyp segmentation model is based on a polyp segmentation network with a U-shaped architecture, and consists of a pyramid vision transformer as the encoder, a multi-scale interaction module as the skip connection, a spatial attention enhancement module, and a context-aware repair module as the decoder; The multi-scale interaction module performs multi-scale feature interaction processing on the features extracted by the pyramid vision transformer; The spatial attention enhancement module highlights the important regions in the extracted multi-scale features by fusing parallel global spatial attention mechanism units and local spatial attention mechanism units; The context-aware repair module detects the edge regions of the pathological image to be segmented and gradually processes the segmentation boundaries.

Citation Information

Patent Citations

  • Deep learning-based skin lesion image segmentation method and device, and storage medium

    CN114066904A

  • Multi-focus image fusion method based on multi-scale context perception

    CN116630763A

  • Bladder tumor image segmentation method and system based on detail enhancement reverse attention network

    CN117975011A

  • Image semantic segmentation method fusing space detail context and multi-scale interaction

    CN118230323A

  • Eye fundus image segmentation method and equipment based on multi-scale cross fusion attention

    CN118644679A