Polyp image segmentation method and system based on combination of SAM model and graph neural network
By combining SAM model and graph neural network, using graph attention mechanism and SAM hybrid network architecture, the problems of accuracy and robustness of the existing technology in early polyp segmentation are solved, and efficient polyp detection and segmentation are achieved, which is suitable for clinical applications.
Patent Information
- Application Number
- CN202510108848.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Existing deep learning models face the challenge of polyp segmentation with blurred boundary, complex structure or smaller size when dealing with early polyps, and have high sensitivity to complex backgrounds and computing resources.
Using a polyp image segmentation method based on the combination of SAM model and graph neural network, local and global graph structural features are extracted and segmentation results are generated through graph attention mechanism and SAM hybrid network architecture.
It significantly improves the detection accuracy, robustness and computing efficiency in early diagnosis of colorectal cancer, reduces the occurrence of misdiagnosis and misdiagnosis, and is suitable for real-time clinical applications.
Smart Images

Figure CN120070886A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a technique for precise automatic segmentation of polyps in the diagnosis of colorectal cancer (CRC). Specifically, it is a polyp image segmentation method and system based on the combination of the SAM model and graph neural network. Background Art
[0002] Colorectal cancer is the second most lethal cancer globally and usually presents as polyps in the early stage. In the field of medical image processing, especially in the early diagnosis of colorectal cancer, polyp segmentation technology is one of the key technologies. Therefore, early detection and precise segmentation of polyps play a crucial role in improving the survival rate of patients.
[0003] Currently, polyp segmentation technology mainly relies on deep learning models, especially models based on U-Net and its variants. The U-Net model is a popular convolutional neural network (CNN) architecture and has been widely used due to its excellent performance in image segmentation tasks. These models effectively identify and segment lesion areas such as polyps by capturing local features and context information in the image. In recent years, Transformer-based models have also been introduced into the polyp segmentation task due to their ability to handle long-range dependencies.
[0004] Nevertheless, existing deep learning models still face several challenges in dealing with early polyps, especially in the segmentation of polyps with blurred boundaries, complex structures, or small sizes. These challenges mainly include:
[0005] 1. Neglect of structural connections and interactions between pixel regions: Existing models, especially those based on U-Net, tend to focus on capturing local features and boundary information, while ignoring the complex structural connections between polyps and surrounding tissues and the interactions between pixel regions, which may lead to the model's inability to accurately identify and segment polyps with unclear boundaries.
[0006] 2. Insufficient detection of small and early polyps: Early polyps are usually small in size and have little difference from surrounding tissues. Existing models are difficult to accurately detect and segment these polyps, increasing the risk of missed diagnosis or misdiagnosis.
[0007] 3. Sensitivity to complex backgrounds: The complexity of the background in polyp images (such as bleeding, inflammation, etc.) may interfere with the performance of the model, resulting in inaccurate segmentation results.
[0008] 4. Requirement for computing resources: Although Transformer-based models can handle long-range dependencies, these models usually require high computing resources, which may be impractical in actual clinical applications. Summary of the Invention
[0009] To address the limitations of existing medical image processing techniques in accurately segmenting colorectal polyps, especially the challenges in dealing with early-stage, small-sized, or complex-structured polyps, the present invention provides a polyp image segmentation method and system based on the combination of the SAM model and graph neural network. By adopting a graph attention mechanism and a hybrid network architecture of SAM (Segment Anything Model), the detection accuracy, robustness, and computational efficiency of the model in the early diagnosis of colorectal cancer are significantly improved.
[0010] According to one aspect of the specification of the present invention, there is provided a polyp image segmentation method based on the combination of the SAM model and graph neural network, including:
[0011] Obtain the image to be segmented;
[0012] Input the obtained image to be segmented into the trained image segmentation model to output the image segmentation result; wherein, the training of the image segmentation model includes:
[0013] Construct a training data set;
[0014] Construct an image graph convolutional attention encoder module for local graph structure feature extraction;
[0015] Based on the pre-trained SAM model, introduce a spatial and channel dual adapter module to fine-tune the image encoder of the SAM model, and use the fine-tuned SAM model to extract global graph structure features;
[0016] Fuse the extracted local graph structure features and global graph structure features;
[0017] Generate a segmentation result based on the fused image feature representation;
[0018] Use the training data set for training and output the trained image segmentation model.
[0019] As a further technical solution, constructing a training data set includes:
[0020] Collect a diverse polyp image data set, including pathological images and related annotation masks;
[0021] Perform preprocessing on the collected images;
[0022] Perform conversion from image data to graph data based on the preprocessed images.
[0023] As a further technical solution, the constructed image graph convolutional attention encoder module includes four layers of graph convolutional networks and one layer of graph attention network. The graph convolutional network is used for feature extraction, and the graph attention network is used to strengthen the feature representation between nodes.
[0024] As a further technical solution, a spatial and channel dual adapter module is introduced to fine-tune the image encoder of the SAM model, including:
[0025] Insert the spatial and channel dual adapter module into the Transformer module of the SAM model. The channel adapter reduces the spatial dimension through average pooling and performs feature transformation through a multi-layer perceptron; the spatial adapter adjusts the dimension of the input features through convolution and deconvolution;
[0026] Fuse the channel adaptation result and the spatial adaptation result so that the fine-tuned SAM model is adapted to polyp images.
[0027] As a further technical solution, based on the fused image feature representation, a segmentation result is generated, including:
[0028] Use the prompt encoder to guide the model to focus on key regions in the image according to predefined prompts;
[0029] Use the mask decoder of the SAM model to receive the image feature representation and convert it into a segmentation mask with the same resolution as the input image.
[0030] According to one aspect of the specification of the present invention, a polyp image segmentation system based on the combination of the SAM model and a graph neural network is provided, including:
[0031] An image input module for obtaining an image to be segmented;
[0032] An image segmentation module for inputting the obtained image to be segmented into a trained image segmentation model and outputting an image segmentation result; wherein, the training of the image segmentation model includes:
[0033] Construct a training data set;
[0034] Construct an image graph convolutional attention encoder module for local graph structure feature extraction;
[0035] Based on the pre-trained SAM model, introduce a spatial and channel dual adapter module to fine-tune the image encoder of the SAM model, and use the fine-tuned SAM model to extract global graph structure features;
[0036] Fuse the extracted local graph structure features and global graph structure features;
[0037] Generate a segmentation result based on the fused image feature representation;
[0038] Use the training data set for training and output a trained image segmentation model.
[0039] According to one aspect of the specification of the present invention, there is provided a polyp image segmentation device based on the combination of the SAM model and the graph neural network, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the polyp image segmentation method based on the combination of the SAM model and the graph neural network.
[0040] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the steps of the polyp image segmentation method based on the combination of the SAM model and the graph neural network.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] The image segmentation model (Polyp-GSAM) of the present invention, through its innovative graph attention-guided hybrid SAM architecture, significantly improves the accuracy and efficiency of polyp segmentation in medical images. This model ingeniously integrates local structural features and global information, significantly enhancing the detection ability for early and structurally complex polyps, effectively reducing the occurrence of misdiagnosis and missed diagnosis. At the same time, this model optimizes the use of computing resources, improves the processing speed, and makes it more suitable for real-time clinical application scenarios.
[0043] Specifically, the innovation of the Polyp-GSAM model also includes a fine-tuned CS-Adapter module, which enhances the adaptability of the model to different medical image characteristics and further improves the generalization ability of the model. In addition, its optimized feature fusion strategy effectively reduces the risk of overfitting, ensuring the performance consistency of the model on diverse datasets.
[0044] Generally speaking, the present invention not only significantly improves the quality and efficiency of medical diagnosis, but also provides a more reliable and convenient diagnostic tool for clinicians, greatly promoting the development of early colorectal cancer diagnosis technology. Through these technological innovations, the present invention not only improves the effect of medical services, but also helps to save medical resources, which is of great significance for improving public health levels. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1Schematic flowchart of image segmentation provided by an embodiment of the present invention.
[0047] Figure 2 Schematic architecture diagram of an image segmentation model provided by an embodiment of the present invention.
[0048] Figure 3 Schematic diagram of an image graph convolutional attention encoding module provided by an embodiment of the present invention.
[0049] Figure 4 Schematic diagram of a spatial adapter module and a channel adapter module provided by an embodiment of the present invention. Detailed implementation manners
[0050] The method provided by the present invention is specifically applied to the fields of endoscopy and medical imaging, and automatically identifies and accurately segments colorectal polyps through advanced image analysis methods. The application of this technology not only improves the accuracy of polyp detection, but also greatly enhances the efficiency of diagnosis and treatment, thereby contributing to improving the survival rate and quality of life of patients.
[0051] Aiming at the limitations of existing medical image processing technologies in accurately segmenting colorectal polyps, especially the challenges in dealing with early-stage, small-sized or structurally complex polyps, the present invention proposes an innovative image segmentation method that adopts a graph attention mechanism and a SAM hybrid network architecture to significantly improve the detection accuracy, robustness and computational efficiency of the model in the early diagnosis of colorectal cancer.
[0052] The above technical innovation not only solves the problems of accuracy and robustness of polyp segmentation in the prior art, but also improves the practicality and efficiency of the operation, making the present invention have important application value in the early diagnosis of colorectal cancer.
[0053] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or a single embodiment provided by the present invention can be combined with each other arbitrarily to form a new technical solution. This combination is not restricted by the order of steps and / or the pattern of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0054] An embodiment of the present invention provides a polyp image segmentation method based on the combination of the SAM model and the graph neural network. First, an image to be segmented is obtained; then, the obtained image to be segmented is input into a trained image segmentation model (i.e., the Polyp-GSAM model), and an image segmentation result is output.
[0055] The Polyp-GSAM model provided by the embodiment of the present invention is a graph attention-guided hybrid SAM (SegmentAnything Model) architecture model, aiming to achieve accurate segmentation of polyps in medical images. This model comprehensively uses deep learning, graph theory, and attention mechanisms to improve the detection and segmentation accuracy of polyps, especially when dealing with early-stage, small-sized, or structurally complex polyps. The model mainly includes: an image graph convolutional attention encoding module, a spatial and channel dual adapter module, an image encoder module, a feature fusion module, a mask decoder module, and a prompt encoder module.
[0056] In the field of image segmentation, especially medical image segmentation, such as polyp detection and segmentation, although traditional convolutional neural networks (CNNs) and attention mechanisms have achieved certain results, they often have difficulty capturing complex structural information in images and relationships between pixels. To improve the segmentation accuracy and robustness, an embodiment of the present invention proposes a novel image graph convolutional attention encoder module.
[0057] Specifically, the image graph convolutional attention encoding (IGCA) module is responsible for converting the input medical image into graph-structured data. In this process, each pixel point is regarded as a node in the graph, and the connection relationship between nodes is established based on the neighborhood relationship between pixels. The image graph convolutional attention encoding (IGCA) module uses a graph convolutional network (GCN) and a graph attention network (GAT) to extract local features and deep structural features of the image, thereby capturing complex interaction relationships between pixels.
[0058] The goal of the IGCA module is to convert the graph-structured data into a high-dimensional feature representation, with a focus on local image features. This module consists of four layers of graph convolutional network (Graph Convolutional Network) and graph attention network (GraphAttention Network).
[0059] Among them, the graph convolutional network (GCN) captures local features by simulating the adjacency relationships between pixels in an image. The features of each node (pixel) not only contain its own information but also the information of its neighboring nodes, which helps the model understand the spatial structure in the image. In GCN, the weights of the convolutional kernels are shared across all nodes and edges, which reduces the number of model parameters and makes the model more efficient. GCN can be stacked in multiple layers, and each layer can capture features at different scales, thus achieving the fusion of multi-scale features. The working principle of GCN is based on an aggregation operation, where the new features of each node are updated by aggregating the features of itself and its neighboring nodes.
[0060] After passing through the GCN layer, the final feature representation is input into the GAT layer to further enhance the feature representation between nodes. GAT assigns different weights to different nodes through an attention mechanism, which enables the model to focus more on the important features in the image. GAT does not rely on a fixed order or distance metric of nodes, so it is more flexible and can capture more complex dependencies. The calculation of GAT can be parallelized, which makes it more efficient on large-scale graphs. GAT updates the features by calculating the attention coefficients of each node to other nodes.
[0061] By combining GCN and GAT, the image graph convolutional attention encoder module can effectively extract the local features and global structure information of the image, providing a powerful feature representation for the polyp segmentation task. This combination utilizes the local connectivity of GCN and the dynamic attention mechanism of GAT, enabling the model to more accurately identify and segment the polyp region.
[0062] With the development of deep learning technology, significant progress has been made in the field of image segmentation. As an advanced image segmentation model, the SAM model can handle various image segmentation tasks. However, there are significant differences in color, texture, and pixel intensity between medical images and natural images, and directly applying the SAM model to medical image segmentation does not yield satisfactory results. Therefore, it is necessary to fine-tune the SAM model to adapt to the characteristics of medical images.
[0063] To introduce the relevant knowledge of polyp images into the SAM model and enable the SAM encoder to better adapt to the characteristics of polyp images, the present invention introduces a spatial and channel dual adapter module to fine-tune the image encoder of SAM, allowing SAM to learn new knowledge while maintaining the basic knowledge. The spatial and channel dual adapter module is inserted into the TransformerBlocks of the SAM encoder to adjust the channel and spatial dimensions of the feature map respectively. This fine-tuning process enables the model to more accurately capture the features of polyps and improves the segmentation accuracy.
[0064] The feature fusion module is used to integrate the local graph structure features extracted by the IGCA module and the global features extracted by the SAM encoder in the feature fusion layer to obtain a more comprehensive image feature representation. This fusion strategy not only enhances the model's understanding of the complex structure of polyps but also optimizes the integration of features, thereby improving the segmentation performance and generalization ability.
[0065] The prompt encoder module is used to process additional information related to the segmentation task, such as bounding boxes or click positions, to further improve the segmentation accuracy. This module enhances the model's responsiveness to user input, making it perform better in interactive segmentation tasks.
[0066] The mask decoder module is used to further process the fused features and finally generate the polyp segmentation mask. The mask decoder is responsible for converting the integrated features into the final segmentation result, thereby achieving the precise localization and segmentation of polyps.
[0067] A polyp image segmentation method based on the combination of the SAM model and the graph neural network provided by the embodiments of the present invention introduces the graph neural network into the SAM model to solve the problem of insufficient understanding of complex medical images by SAM. The zero-shot segmentation ability of SAM can greatly improve the segmentation efficiency of the model, reduce the consumption of computing resources, and provide a reference segmentation mask for professional doctors to assist in treatment.
[0068] Please refer to Figure 1 and Figure 2 , which shows the flow chart and model architecture diagram of the graph attention-guided hybrid SAM architecture model. The embodiments of the present invention adopt a dual-network encoder structure, the SAM Image Encoder and the Image Graph Convolutional Attention Encoder module (IGCA). After the image is extracted with features by the double-layer encoder, it enters the Mask Decoder, and then the Prompt Encoder gives point and box prompts to guide the Mask Decoder to give the segmented mask. The training process of the model includes the following steps:
[0069] Step 1: Preparation and preprocessing of the dataset.
[0070] Step 1.1, Construction of the dataset: Collect a diverse polyp dataset, including pathological images and related annotated masks, to ensure the robustness of the SAM model. The dataset consists of multiple public polyp image datasets, including ETIS-Larib, CVC-ClinicDB, CVC-ColonDB, CVC-300, and Kvasir-SEG. These datasets provide different numbers of polyp images and their corresponding pixel-level annotations, and the resolutions are also different.
[0071] Step 1.2, Division of the dataset: Using a recognized dataset division method, a total of 1450 samples, including 900 Kvasir-SEG images and 550 CVC-ClinicDB images, are used as the training set; the test set includes the remaining 100 images of Kvasir-SEG, the remaining 62 images of CVC-ClinicDB, 380 samples of CVC-ColonDB, 196 samples of ETIS, and 60 images of CVC-300.
[0072] Step 1.3, Preprocessing of the dataset, including operations such as cropping, rotation, and deformation for image enhancement.
[0073] Further including:
[0074] Step 1.3.1, Resizing of the images, uniformly adjusting the size of the input images to 256×256 pixels, which helps with the input consistency of the model.
[0075] Step 1.3.2, Using an image-to-graph data conversion module to convert the input images into graph-structured data (i.e., data of nodes and edges that can be recognized by the Image Graph Convolutional Attention Encoding (IGCA) module). This involves converting each pixel point of the image into a node in the graph and establishing edges based on the neighborhood relationship of the pixel points.
[0076] Step 1.3.3, Through grayscale processing, convert the color images into grayscale images to better capture the local features and structural information of the images.
[0077] For each image I, where the shape of I is [C, H, W], convert it to a grayscale image G, and its calculation method is as follows:
[0078]
[0079] where I C represents the C-th channel of the image, C represents the number of channels, and H and W represent the height and width of the image respectively. For each pixel p in the grayscale image G xy , at the position (x, y), create a node v for each pixel, and the feature f(v) of the node is the grayscale value of the pixel. For each pixel in the image, create edges to connect it with its upper, lower, left, and right neighbor pixels. If the pixel p xy is at the position (x, y), for the existing neighbors, add an edge (v xy , v neighbor ), where v neighbor is the node corresponding to the neighbor pixel. The specific set of edges can be represented as:
[0080] E = {(v xy , v x-1,y ), (vxy ,v x+1,y ),(v xy ,v x,y-1 ),(v xy ,v x,y+1 )}
[0081] Finally, for the entire image, a graph G=(V, E) is obtained, where V is the set of all nodes defined by pixel gray values, and E is the set of all edges defined by pixel adjacency relationships.
[0082] Step 2: The image graph convolutional attention encoder module performs graph feature extraction.
[0083] Please refer to Figure 3 , the image graph convolutional attention encoder module is composed of 4 layers of graph neural networks (GCN) and one layer of GAT (graph attention mechanism).
[0084] The working principle of GCN is based on the aggregation operation, where the new features of each node are updated by aggregating the features of itself and its neighbor nodes. Specifically, for the graph convolutional network part, there are L GCN layers, where L = num_gcn_layers. For each node v in the graph, its features at the L-th layer can be updated in the following way.
[0085]
[0086] where N(v) is the set of neighbor nodes of node v, c uv is the normalization constant, and is W l the weight matrix of the L-th layer, is the feature representation of node u at the l-1-th layer, is the initial feature of node v (e.g., pixel intensity) and ReLU is the activation function used to introduce non-linearity.
[0087] The graph attention network GAT assigns different weights to different nodes through the attention mechanism, which enables the model to pay more attention to the important features in the image.
[0088] For each node v in the graph, its updated feature X IGCA at the GAT layer can be calculated in the following way:
[0089]
[0090] where, where is the attention coefficient of node u to node v, k is the number of network layers, W is the weight matrix, and σ is the softmax function used to normalize the attention weights. The attention coefficient Calculated in the following way:
[0091]
[0092] where a is a learnable weight vector used to calculate the relative importance between node pairs, || represents the concatenation operation, and σ is the softmax function used to normalize the attention weights of all neighbors of v. Each node undergoes a series of graph convolutional and graph attention transformations, and the final obtained feature X IGCA , contains comprehensive information of local neighborhoods and global structures. This enables the model to capture complex spatial relationships in image data and provides powerful feature representations for downstream tasks.
[0093] By combining GCN and GAT, the image graph convolutional attention encoder module can effectively extract local features and global structure information of images, providing powerful feature representations for the polyp segmentation task. This combination utilizes the local connectivity of GCN and the dynamic attention mechanism of GAT, enabling the model to more accurately identify and segment polyp regions.
[0094] Step 3: Fine-tuning based on the SAM model.
[0095] With the development of deep learning technology, significant progress has been made in the field of image segmentation. As an advanced image segmentation model, the SAM model can handle various image segmentation tasks. However, there are significant differences in color, texture, and pixel intensity between medical images and natural images, and directly applying the SAM model to medical image segmentation yields unsatisfactory results. Therefore, it is necessary to fine-tune the SAM model to adapt to the characteristics of medical images.
[0096] The advantages of fine-tuning the SAM model include: improved adaptability, fine-tuning the SAM model enables it to better adapt to the characteristics of medical images and improve segmentation accuracy; reduced computational resources, by fine-tuning rather than completely retraining, the consumption of computational resources is reduced; rapid deployment, the fine-tuning process can quickly adapt to new datasets and accelerate the deployment and application of the model.
[0097] Please refer to Figure 4 , the working principle of fine-tuning the SAM model is: by using the Channel Adapter and Space Adapter, which are integrated into the Transformer blocks of the SAM model to enhance the model's ability to learn medical image features without significantly changing the existing model architecture or parameters.
[0098] The detailed steps include:
[0099] Step 3.1, Pre-trained Model Selection: Select the SAM model pre-trained on a large-scale general image dataset to obtain a basic feature representation.
[0100] Step 3.2, Data Preparation: Collect and annotate a medical image dataset for the fine-tuning process.
[0101] Step 3.3 Freeze Parameters: Freeze most of the parameters in the SAM model, especially the weights of the underlying layers, to retain the features learned by the model on general images.
[0102] Step 3.4, Introduce Spatial Adapter and Channel Adapter, and only update these parts of the parameters during training.
[0103] The main purpose of the channel adapter is to adjust the channel dimension of the feature map to better adapt to downstream tasks. This process is completed through average pooling and a multi-layer perceptron (MLP).
[0104] Average Pooling: First, the input feature map X reduces the spatial dimension through average pooling, which helps the model capture global information.
[0105] Feature Transformation: The pooled feature map undergoes feature transformation through an MLP. The MLP contains two weight matrices W1 and W2, as well as ReLU activation function and Sigmoid function.
[0106] X channel =σ(W 2 (Relu(W 1 (AvgPool(X)))))
[0107] where and are the weight matrices obtained through learning, C′ is the dimension of the intermediate layer, and σ is the Sigmoid function, which is used to map the weights between 0 and 1, indicating the importance of the channels.
[0108] The spatial adapter adjusts the spatial resolution of the input features through convolution and transposed convolution operations.
[0109] X spatial =ConvTranspose2d(Relu(Conv2d(X channel )))
[0110] Conv2d and ConvTranspose2d are convolution and transposed convolution operations, which are used to adjust the spatial resolution of the features.
[0111] Convolution: The output of the channel adapter is first passed through the convolutional layer Conv2d to further extract spatial features.
[0112] Deconvolution: Then, the spatial resolution of the feature map is adjusted through the transposed convolutional layer ConvTranspose2d to match the spatial dimensions of the original input feature map.
[0113] Step 3.5, Feature Fusion: The feature fusion strategy fuses the outputs of the channel adapter and the spatial adapter with the original input feature map to enhance the model's ability to integrate local and global information.
[0114] Element-wise Multiplication: The output X of the channel adapter channel is multiplied element-wise with the original input feature map X to integrate channel attention information.
[0115] Weighted Summation: The result of the element-wise multiplication is summed with the output X of the spatial adapter spatial to obtain the final fused feature map.
[0116] Specifically, the channel attention weights are applied to the original input and fused with the result of the spatial attention transformation, as shown in the formula:
[0117] X fusion = X channel ⊙ X + X spatial
[0118] If the skip connection is enabled, X spatial is added to the input X; otherwise, X spatial is directly used, where ⊙ represents element-wise multiplication.
[0119] Step 3.6, Loss Function Customization: Design a loss function suitable for the medical image segmentation task, such as a combined loss function, to optimize the model's ability to identify lesion regions.
[0120] Step 3.7, Data Augmentation and Cross-Validation: Use data augmentation techniques to increase the diversity of medical image data and perform cross-validation to ensure that the model can perform stably well on different datasets.
[0121] Step 4, Feature Fusion.
[0122] Fuse the graph features obtained from the graph convolutional attention encoder module and the output features of the fine-tuned SAM graph encoder module, and the formula is:
[0123] X fusion = W vit × X vit + W IGCA × XIGCA
[0124] W vit and W IGCA are the weights of two features respectively. To maintain feature consistency, we make W vit and W IGCA add up to 1 constantly. X fusion represents the result after fusion. Finally, through the Neck layer including a series of convolutional and normalization operations, it is used to output the final image feature representation.
[0125] Step 4, Mask Decoder and Prompt Encoder.
[0126] Retain the mask decoder structure of SAM, update the parameters to process the final image feature representation, and generate predicted masks.
[0127] The Mask Decoder includes:
[0128] Feature representation: The mask decoder receives the high-dimensional feature representation from the image encoder and converts it into a segmentation mask with the same resolution as the input image.
[0129] Convolutional and normalization operations: Process the features through a series of convolutional layers and normalization layers to enhance the expressive ability of the features and gradually restore the spatial resolution of the image.
[0130] Upsampling and fusion: Use upsampling techniques (such as transposed convolution) to increase the spatial resolution of the feature map to the same as the input image, and fuse the low-level features through skip connections to retain more detailed information.
[0131] Prompt Encoder:
[0132] Prompt enhancement: The prompt encoder uses predefined prompts (such as bounding boxes, key points, or natural language descriptions) to guide the model to focus on the key regions in the image.
[0133] Feature fusion: Combine the prompt information and image features, and strengthen the model's ability to recognize key regions through the attention mechanism.
[0134] Multi-modal output: Provide multiple output modes, including BOX-based segmentation and Point-based segmentation, to adapt to different application scenarios and requirements.
[0135] In practical applications, the mask decoder and prompt encoder can be jointly trained to optimize the overall performance of the model. The mask decoder is responsible for generating accurate segmentation masks, while the prompt encoder provides additional context information to help the model better understand and segment the target regions in the image.
[0136] Step 6, Model Training.
[0137] Loss Function Definition: Define the loss function as a combined loss function, including weighted cross-entropy loss and Dice loss, to optimize the segmentation performance of the model.
[0138] L = α·LWCE + β·LDice
[0139] Where α and β are the weight coefficients of weighted cross-entropy loss and Dice loss respectively.
[0140] Optimizer Selection: Select the Adam optimizer for model training, as its adaptive learning rate characteristic is suitable for processing complex medical image data.
[0141] Training Epoch Setting: Set the model training epoch to 50 epochs to ensure that the model learns sufficiently and avoids overfitting.
[0142] Learning Rate Adjustment: Set the initial learning rate to 1x10 -4 And adopt a learning rate decay strategy according to the performance of the validation set during training.
[0143] Batch Size Selection: Select an appropriate batch size according to the GPU memory capacity and the dataset size to ensure the stability and efficiency of model training.
[0144] Data Augmentation Techniques: Apply data augmentation techniques such as rotation, flipping, and scaling to increase data diversity and improve the generalization ability of the model.
[0145] Cross-Validation: Adopt cross-validation to evaluate the model performance and ensure that the model can perform stably and well on different datasets.
[0146] Hyperparameter Definition: Designed multiple hyperparameters such as fusion weights, Dropout, and the number of GCN layers, which need to be adjusted in real time according to the specific experimental situation.
[0147] Step 7, Performance Evaluation.
[0148] Use Dice Coefficient and Mean Intersection over Union as evaluation metrics. Run the model on a 4080 GPU with 16GB of memory.
[0149] The Dice coefficient, also known as a variant of the F1 score, is a performance metric used to evaluate binary classification problems, especially in image segmentation tasks. It measures the similarity between the segmentation regions predicted by the model and the ground truth annotation regions. The value of the Dice coefficient ranges from 0 to 1. A value of 1 indicates perfect segmentation, i.e., the segmentation regions predicted by the model are exactly the same as the ground truth annotation regions; a value of 0 indicates no overlap and segmentation failure.
[0150] mIoU is the average of the intersection over union (IoU) for multiple classes and is used for performance evaluation in multi-class image segmentation tasks. It measures the degree of overlap between the segmentation regions predicted by the model and the ground truth annotation regions. The value of mIoU also ranges from 0 to 1. A value of 1 indicates that the segmentation of all classes is completely correct; a value of 0 indicates no overlap and segmentation failure. mIoU takes into account the segmentation performance of all classes, so it is a global performance metric.
[0151] The present invention has successfully designed and implemented a novel hybrid architecture model called Polyp-GSAM specifically for polyp segmentation tasks. The model is based on the core framework of SAM (Segment Anything Model) and integrates an innovative image graph convolutional attention encoding module to facilitate collaborative training. By fine-tuning SAM for polyp images and allowing the graph convolutional attention encoding module to focus on extracting local structural features and integrating non-local pixel relationships, the method of the present invention effectively combines these two types of features. This approach creates a comprehensive feature representation that contains long-range dependencies and short-range structural information, significantly improving the performance of the model on multiple standard polyp datasets. This success not only confirms the effectiveness of the collaborative operation between SAM and the graph convolutional attention encoding but also indicates the potential for the method to be extended to broader challenges in the field of medical image segmentation.
[0152] Experimental results show that the Polyp-GSAM model outperforms existing state-of-the-art methods on multiple benchmark datasets, marking it as a powerful tool for early colorectal cancer intervention. Additionally, the model of the present invention also achieved good performance on the CVC-ColonDB, CVC-300, and ETIS-Larib datasets that were not included in the training data, further demonstrating the robustness and generalization ability of the model. Through quantitative and qualitative evaluations, the model of the present invention performed excellently in metrics such as intersection over union and Dice coefficient, and these findings together validate the effectiveness of our model in diverse data configurations.
[0153] In summary, the Polyp-GSAM model provides a new solution in the field of medical image segmentation, especially demonstrating excellent performance when dealing with complex structures and subtle features.
[0154] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides a polyp image segmentation system based on the combination of the SAM model and the graph neural network, which is used to execute the polyp image segmentation method based on the combination of the SAM model and the graph neural network in the above method embodiments.
[0155] The system includes: an image input module for acquiring the image to be segmented; an image segmentation module for inputting the acquired image to be segmented into a trained image segmentation model and outputting an image segmentation result; wherein, the training of the image segmentation model includes: constructing a training data set; constructing an image graph convolutional attention encoder module for extracting local graph structure features; based on the pre-trained SAM model, introducing a spatial and channel dual adapter module to fine-tune the image encoder of the SAM model, and using the fine-tuned SAM model to extract global graph structure features; fusing the extracted local graph structure features and global graph structure features; generating a segmentation result based on the fused image feature representation; and training using the training data set to output a trained image segmentation model.
[0156] The polyp image segmentation system based on the combination of the SAM model and the graph neural network provided by the embodiment of the present invention faces the limitations of existing medical image processing technologies in accurately segmenting colorectal polyps, especially the challenges in dealing with early, small-sized or complex-structured polyps. By adopting the above-mentioned several modules and using a graph attention mechanism and a SAM hybrid network architecture, it significantly improves the detection accuracy, robustness and computational efficiency of the model in the early diagnosis of colorectal cancer.
[0157] It should be noted that the system embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, is also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting corresponding functional modules, and its principle is basically the same as that of the above system embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above system embodiment, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicability of the technical solutions, improve the modules in the above system embodiment to obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.
[0158] It should be understood that the implementation of the system embodiment of the present invention can be realized with reference to the foregoing method embodiment, and the present invention will not be elaborated herein.
[0159] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention further provides a polyp image segmentation device based on the combination of the SAM model and the graph neural network, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the polyp image segmentation method based on the combination of the SAM model and the graph neural network.
[0160] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the polyp image segmentation method based on the combination of the SAM model and the graph neural network.
[0161] In summary of the above embodiments, the present invention solves the problems of the prior art through the following technical innovations:
[0162] 1. Introduce an image graph convolutional encoder module: Through the graph attention mechanism, the present invention can more effectively process long-range dependencies and complex structures in images, thereby improving the recognition and segmentation capabilities for polyps with blurred boundaries or complex structures.
[0163] 2. Adopt a SAM hybrid network architecture: The SAM architecture combines the multi-scale processing ability of deep learning and the detailed analysis of specific regions, enhancing the model's detection ability for small-sized polyps and improving the segmentation accuracy in complex backgrounds.
[0164] 3. Introduce a spatial and channel dual adapter module to fine-tune SAM, and complete the learning of polyp image features by SAM with a small amount of computing resources.
[0165] 4. Optimize the computing efficiency: This architecture design also considers the computing resource limitations in actual clinical applications. By optimizing the algorithm and network structure, it reduces the dependence on high-performance computing resources, making the present technology more suitable for deployment in resource-constrained environments.
[0166] The above technical innovations not only solve the problems of the accuracy and robustness of polyp segmentation in the prior art, but also improve the practicality and efficiency of the operation, making the present invention have important application value in the early diagnosis of colorectal cancer.
[0167] The terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A polyp image segmentation method based on the combination of SAM model and graph neural network, characterized in that: include: Obtain the image to be segmented; Input the acquired image to be segmented into the trained image segmentation model, and output the image segmentation result; wherein the training of the image segmentation model includes: Build a training dataset; Construct an image graph convolutional attention encoder module to extract local graph structure features; Based on the pre-trained SAM model, the spatial and channel dual adapter module is introduced to fine-tune the image encoder of the SAM model, and the fine-tuned SAM model is used to extract global graph structure features; The extracted local graph structure features are fused with the global graph structure features to obtain image feature representation; Generate segmentation results based on the fused image feature representation; The training data set is used for training, and a trained image segmentation model is output.
2. According to claim 1, the polyp image segmentation method based on the combination of SAM model and graph neural network is characterized in that: Construct a training dataset, including: Collect a diverse dataset of polyp images, including pathological images and associated annotated masks; Perform preprocessing based on the collected images; Perform image to graph data conversion based on the preprocessed image.
3. According to claim 1, the polyp image segmentation method based on the combination of SAM model and graph neural network is characterized in that: The constructed image graph convolutional attention encoder module includes a four-layer graph convolutional network and a one-layer graph attention network. The graph convolutional network is used for feature extraction, and the graph attention network is used to strengthen the feature representation between nodes.
4. According to claim 1, the polyp image segmentation method based on the combination of SAM model and graph neural network is characterized in that: The spatial and channel dual adapter module is introduced to fine-tune the image encoder of the SAM model, including: Inserting the spatial and channel dual adapter modules into the Transformer module of the SAM model, the channel adapter reduces the spatial dimension by average pooling and performs feature conversion through a multi-layer perceptron; the spatial adapter adjusts the dimension of the input feature by convolution and deconvolution; The channel adaptation results are fused with the spatial adaptation results to make the fine-tuned SAM model adapt to the polyp image.
5. According to claim 1, the polyp image segmentation method based on the combination of SAM model and graph neural network is characterized in that: Based on the fused image feature representation, the segmentation results are generated, including: Using a hint encoder to guide the model to focus on key areas in the image based on predefined hints; The mask decoder of the SAM model receives the image feature representation and converts it into a segmentation mask with the same resolution as the input image.
6. The polyp image segmentation system based on the combination of SAM model and graph neural network is characterized by: include: An image input module is used to obtain the image to be segmented; The image segmentation module is used to input the acquired image to be segmented into the trained image segmentation model and output the image segmentation result; wherein the training of the image segmentation model includes: Build a training dataset; Construct an image graph convolutional attention encoder module to extract local graph structure features; Based on the pre-trained SAM model, the spatial and channel dual adapter module is introduced to fine-tune the image encoder of the SAM model, and the fine-tuned SAM model is used to extract global graph structure features; The extracted local graph structure features are fused with the global graph structure features to obtain image feature representation; Generate segmentation results based on the fused image feature representation; The training data set is used for training, and a trained image segmentation model is output.
7. A polyp image segmentation device based on the combination of SAM model and graph neural network, characterized in that: It includes a memory and a processor, the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the polyp image segmentation method based on the combination of the SAM model and the graph neural network as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which enable the computer to execute the steps of the polyp image segmentation method based on the combination of the SAM model and the graph neural network as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Semantic SAM large model-based three-dimensional point cloud robustness component segmentation method
CN118397282A
Self-adaptive medical image segmentation method and device based on prompt learning, and medium
CN118521595A
Specific dermatitis image segmentation method based on SAM model
CN119048526A
Device for Unsupervised Domain Adaptation in Semantic Segmentation Exploiting Inter-pixel Correlations and Driving Method Thereof
KR102437959B1
System and method for structure learning for graph neural networks
US20220101103A1
Cited By
Dynamic graph neural network trauma image segmentation method, electronic equipment and computer program product
CN120543863A