Medical image segmentation method based on SAM and prompter

By adopting a SAM and prompter-based method in medical image segmentation, and using Transformer structure and prompter mechanism to adaptively learn the best embedded prompt, the existing medical image segmentation method framework is solved, and efficient and accurate medical image segmentation effect is achieved.

CN120014262AActive Publication Date: 2025-05-16CHINA UNIV OF PETROLEUM (EAST CHINA)

Patent Information

Application Number
CN202510056567.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have problems such as weak universality, long training time and low accuracy.

Method used

Using a medical image segmentation method based on SAM and prompter, the feature aggregation and segmentation of medical images are realized through the combination of SAM image encoder, feature enhancement block, prompter module and SAM image decoder. This method utilizes the Transformer structure and prompter mechanism to adaptively learn the best embed prompts, locate objects and infer their semantic categories and instance masks.

Benefits of technology

The medical image segmentation effect with short training time, high accuracy and strong universality is achieved, and the problems of type, location and number of SAM prompts are overcome, and the accuracy and efficiency of segmentation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014262A_ABST
    Figure CN120014262A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image segmentation, and particularly discloses a medical image segmentation method based on SAM and a prompter. According to the method, firstly, an original medical image is preprocessed, enhanced features are obtained through an SAM image encoder and a feature enhancement block in sequence, then position codes and level codes are combined, the combined codes are input into a Transform encoder to generate multi-scale features of prompts, enhanced prompt features are generated through a Transform decoder, and the prompts are subjected to multi-scale feature extraction. Inputting a prompt linear projection layer and then combining the prompt linear projection layer with the multi-scale features to form sparse prompts; splicing the features with intermediate features generated by each layer of an SAM image encoder, and inputting the features into an SAM image decoder to obtain a mask image; and finally training the model and segmenting the to-be-segmented medical image to finally obtain a segmented image. According to the method, the problem that the SAM generation result is obviously influenced by the type, the position and the number of SAM prompts is solved, and the method has the advantages of short training time, high training accuracy, high universality and convenience in operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image segmentation, and in particular is a medical image segmentation method based on SAM and a prompter. Background Art

[0002] With the rapid development of computer technology, medical imaging technology has been significantly developed and widely used, such as computer tomography (CT) and magnetic resonance imaging (MRI). The medical images generated by these technologies provide important tools for doctors to diagnose diseases, and also support medical researchers to conduct pathological analysis in the field of scientific research. The rapid growth of medical images has brought about a huge amount of data, which has posed new challenges to storage, processing and analysis; the effective use of this data requires the further development of computer and artificial intelligence technologies. By applying artificial intelligence, especially machine learning and deep learning technologies, medical images can be analyzed more accurately and the deep information contained in medical image data can be mined. This can not only help doctors make more accurate diagnoses, but also play a key role in the formulation of treatment plans and disease monitoring.

[0003] In computer vision, the segmentation task refers to the process of classifying each pixel in an image into a specific category. This usually includes two types: semantic segmentation and instance segmentation. Semantic segmentation divides each pixel in the image into predefined categories, without distinguishing different entities in the same category; while instance segmentation not only distinguishes categories, but also distinguishes different instances in the same category. In the medical field, image segmentation technology is particularly important. Accurate image segmentation can help doctors identify and quantify lesion areas, such as tumors, inflammation or other abnormal structures; it can also accurately identify the boundaries of organs such as the heart, liver, and brain, and track changes in lesion areas before and after treatment; at the same time, it is also an indispensable tool in medical education and can be used to train medical students to understand complex anatomical structures and pathological changes.

[0004] The Chinese invention patent application with publication number CN108038862A discloses an interactive medical image intelligent segmentation modeling method for achieving target area segmentation, optimizing segmentation results and improving modeling efficiency. However, this method relies heavily on user input and interaction at the segmentation level, which means that the final segmentation quality and efficiency will be affected by user experience and skill level. For users who have not been fully trained, it is difficult to achieve the optimal segmentation effect. Moreover, when processing large or complex medical image data, the process of manual marking and adjusting contours is very time-consuming, and the application efficiency is low in fast-paced clinical environments. At the same time, in practical applications, there is a lack of analysis of the overall information of the initial image, and it performs poorly in high-dimensional data analysis. Summary of the invention

[0005] The purpose of the present invention is to provide a medical image segmentation method based on SAM (Segment Anything Model, SAM) and a prompter to solve the problems of weak framework universality, long training time, low accuracy, etc. in existing medical image segmentation methods.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A medical image segmentation method based on SAM and a prompter comprises the following steps:

[0008] Step 1: Obtain original medical images, preprocess the original medical images, and construct a training data set;

[0009] Step 2: Build a medical image segmentation model based on the SAM model, which includes a SAM image encoder, a feature enhancement block, a prompter module, and a SAM image decoder;

[0010] The SAM image encoder uses a multi-layer Transformer Block structure;

[0011] The prompter module includes a Transformer encoder, a Transformer decoder, and a prompt linear projection layer;

[0012] The SAM image decoder uses a multi-layer Transformer decoder structure;

[0013] Among them, the processing process of the preprocessed original medical image in the medical image segmentation model is as follows:

[0014] The preprocessed original medical image first passes through the SAM image encoder composed of multiple layers of Transformer Block to obtain the intermediate features generated by each layer, and then the feature enhancement block recursively accepts the features output by each layer of Transformer Block and performs enhancement processing;

[0015] The enhanced features output by the feature enhancement block are input to the prompter module and processed as follows: first, the enhanced features input to the prompter module are flattened after merging the position encoding and the level encoding, and then input to the Transformer encoder. The Transformer encoder extracts the multi-scale features of the high-level prompts, and then passes through the Transformer decoder, combined with the initialization parameters, to recursively generate the enhanced prompt features of each layer; then input to the prompt linear projection layer, and finally combined with the multi-scale features generated by the Transformer encoder to form a sparse prompt;

[0016] The sparse prompts output by the prompter module are concatenated with the intermediate features generated by each layer of the SAM image encoder and input into the SAM image decoder;

[0017] In the SAM image decoder, a Transformer decoder block is used to interact between the total image features generated by the encoder and the hint embedding generated by the hinter to obtain a mask image;

[0018] Step 3: Train the medical image segmentation model built in step 2 based on the training data set in step 1, and then use the trained medical image segmentation model to segment the medical image to be segmented, and finally obtain a segmented image.

[0019] Preferably, in step 1, the preprocessing process includes unifying the image size and data normalization, and the specific process is as follows:

[0020] First, an image with the largest pixel value in the original medical image is selected as the reference matrix, and the size of other original medical images is adjusted to match the reference matrix by filling zeros in the blank areas;

[0021] Then, the pixel values ​​of all original medical images are divided by 255 to ensure that the pixel values ​​of all original medical images are uniformly between [0, 1].

[0022] Preferably, the SAM image encoder comprises an embedding layer for converting the input image into a 16*16 block, and multiple layers of Transformer Block;

[0023] The input image is split into fixed-size blocks through the embedding layer, each block is flattened and mapped into a vector; position encoding is added to the embedding vector of each block to add position information to the feature vector of each block;

[0024] Each Transformer Block layer includes an attention layer with 256-dimensional input features and an MLP layer for amplifying the features by 4 times to 1024 dimensions. The attention layer uses a self-attention mechanism.

[0025] The blocks with added position information from the embedding layer will pass through each layer of Transformer Block. Through the self-attention mechanism, the model can calculate the similarity between any two blocks in the image and combine the information of each block weightedly according to the similarity;

[0026] After each layer of self-attention, residual connections are added to help avoid the vanishing gradient problem;

[0027] At the end of each Transformer Block, a layer of MLP is passed to amplify the features by 4 times to 1024 dimensions for further nonlinear mapping in order to capture more complex visual information.

[0028] Preferably, the feature enhancement block includes a feature aggregator and a feature splitter;

[0029] The feature aggregator is used to receive the intermediate feature maps output by each layer of the Transformer Block in the SAM image encoder and learn representative semantic features. The specific processing process is as follows:

[0030] The feature maps output by each layer of Transformer Block in the SAM image encoder first use a 1×1 convolution-relu block to reduce the number of channels from the original number of feature maps to 32, and then use a 3×3 convolution-relu block to increase spatial information; then down-sample to produce down-sampled features, and merge with the features from the previous layer, and then pass through a 3×3 convolution-relu block, and finally pass through a fused convolution layer consisting of two 3×3 convolution layers and a 1×1 convolution layer. The fused convolution layer is used to restore the channel dimension;

[0031] The feature splitter is used to receive features from the feature aggregator and use multiple transposed convolutional layers to perform upsampling to obtain enhanced features.

[0032] Preferably, in step 2, the Transformer encoder is composed of N layers of stacked self-attention layers and feedforward layers;

[0033] The process of the Transformer encoder extracting multi-scale features of high-level prompts is as follows:

[0034]

[0035] Among them, PE i represents the positional encoding of the i-th layer, LE i represents the level encoding of the i-th layer, Cat(·) represents the connection of the tensor along the channel dimension, and PE i LE i and the intermediate features from the i-th layer of the SAM encoder Merge into Indicates the generation of multiple layers of intermediate features, T_enc represents the Transformer encoder layer;

[0036] The multi-scale features extracted by the Transformer encoder pass through the Transformer decoder and the prompt linear projection layer, and are finally combined with the multi-scale features to form a sparse prompt as shown below:

[0037]

[0038]

[0039] in, represents the zero-initialized learnable parameters and the multi-scale features from each block of the Transformer encoder Input Transformer decoding layer T_dec recursively to obtain feature parameters mlp prompt is a two-layer MLP used to obtain the prompt linear embedding e i ; sin represents the sine function, Indicates a sparse prompt.

[0040] Preferably, the process of training the medical image segmentation model in step 3 is as follows:

[0041] First, by combining SAM and the prompter module, the loss function of the medical image segmentation model is obtained. The expression of the loss function is:

[0042]

[0043] N p Indicates the number of prompt groups; Represents the cross entropy loss calculated between the predicted category and the target; represents the binary cross entropy loss between the predicted mask and the matched ground-truth instance mask, including both the predicted coarse-grained mask and the fine-grained mask; a i Indicates that the match is confirmed to be positive; where, and The calculation formula is:

[0044]

[0045] Among them, y represents a binary label on a certain category, and the value of y is 0 or 1; Represents the probability of the corresponding category predicted by the neural network, The value of is between (0,1);

[0046] When y = 0, there is only the second term in the above formula. The closer it is to 0, the smaller the loss. When y=1, only the first term in the above formula has a value. The closer it is to 1, the smaller the loss;

[0047] Finally, the Adam optimization algorithm is used to solve the minimum value of the loss function.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] As described above, the medical image segmentation method based on SAM and prompter described in the present invention realizes the aggregation and segmentation of medical image features under the SAM framework, which fully considers the characteristics of high-dimensional data of medical images. The medical image obtains intermediate features through the SAM encoder, and uses the category-related prompter mechanism to adaptively learn the best embedding prompt, locate objects and infer their semantic categories and instance masks, which overcomes the problem that the type, position and number of SAM prompts will significantly affect the SAM generation results, and can segment medical images more accurately and quickly. The method of the present invention has the advantages of short training time, high training accuracy, strong universality and convenient operation; at the same time, it is more coherent and has more theoretical guidance than the method of segmenting after manually giving prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0051] Figure 1 Flow chart of a medical image segmentation method based on SAM and prompter in an embodiment of the present invention;

[0052] Figure 2 It is a structural schematic diagram of a medical image segmentation method based on SAM and a prompter in an embodiment of the present invention;

[0053] Figure 3 Schematic diagram of the structure of a feature enhancement block in an embodiment of the present invention;

[0054] Figure 4 Schematic diagram of the structure of the prompter in the embodiment of the present invention. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0056] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making any creative work shall fall within the scope of protection of the present invention.

[0057] Example

[0058] like Figures 1 to 4As shown, the present invention proposes a medical image segmentation method based on SAM and prompter, which can realize medical image feature aggregation and segmentation under the SAM framework, and at the same time use the category-related prompter mechanism to adaptively learn the best embedding prompt, locate objects and infer their semantic categories and instance masks. This method has strong universality, short training time and high accuracy.

[0059] The following describes the specific process of medical image segmentation in combination with lung X-ray images:

[0060] Step 1: Obtain the original medical image, preprocess the original medical image, and construct a training data set; the preprocessing process includes unifying the image size and data normalization.

[0061] Step 1.1: Save the original medical image in a certain image format on the hard disk. For example, taking a lung X-ray image as an example, the image size is 128*128*1 pixel.

[0062] Step 1.2, read the original medical image, select the image with the largest pixel value in the original medical image as the reference matrix, and adjust the size of other original medical images by filling zeros in the blank area to match the reference matrix; then normalize the pixel values ​​of the original medical images, divide the pixel values ​​of all original medical images by 255, and ensure that the pixel values ​​of all original medical images are uniformly between [0,1] or [-1,1] to ensure the numerical consistency of the model's input data, help accelerate training and improve accuracy, and finally obtain the training data set.

[0063] Step 2: Build a medical image segmentation model based on the SAM model, which includes a SAM image encoder, a feature enhancement block, a prompter module and a SAM image decoder.

[0064] The SAM image encoder consists of an embedding layer that converts the image into a 16*16 block and twelve TransformerBlocks. Each Transformer Block includes an attention layer with a 256-dimensional input feature and a multi-layer perceptron (MLP) that amplifies the feature by 4 times to 1024 dimensions.

[0065] Among them, the number of layers of Transformer Block can be reasonably changed according to the amount of data.

[0066] The SAM encoder accepts the input sequence and divides it into N learnable blocks (N is a hyperparameter) through the embedding layer and position encoding. The divided blocks are input into the attention layer. Each block is used as a query, key and value at the same time. The dependency between blocks can be captured through the attention mechanism. The multi-layer perceptron is used to amplify the features by 4 times through the feedforward layer. After residual connection and layer normalization, the training stability is ensured. The SAM encoder finally encodes the input sequence into an intermediate feature map. The specific processing process is as follows:

[0067] First, the preprocessed image is divided into fixed-size blocks through the embedding layer, and each block is flattened and mapped into a vector. Position encoding is added to the embedding vector of each block to add position information to the feature vector of each block, ensuring that the model can understand the relative positions between different blocks.

[0068] The blocks with added position information from the embedding layer will pass through each layer of Transformer Block. Through the self-attention mechanism, the model can calculate the similarity between any two blocks in the image, and combine the information of each block in a weighted manner based on this similarity; it can capture long-distance dependencies and effectively transfer information between multiple regions in the image, and the output features remain 256-dimensional. After each layer of attention, a residual connection will be added to avoid the gradient disappearance problem.

[0069] At the end of each Transformer Block, a multi-MLP layer is passed to amplify the features by 4 times to 1024 dimensions for further nonlinear mapping in order to capture more complex visual information.

[0070] The feature enhancement block recursively accepts the features output by each layer of the Transformer Block of the SAM image encoder and enhances them, such as Figure 3 shown.

[0071] The feature enhancement block includes a feature aggregator and a feature splitter. The feature aggregator aims to learn representative semantic features from the feature maps output by each layer of the Transformer Block in the SAM image encoder. First, a 1×1 convolution-relu block is used to reduce the number of channels from the original number of feature maps to 32. Then a 3×3 convolution-relu block is used to increase spatial information. After that, downsampling is performed to generate downsampled features, which are merged with the features from the previous layer through a 3×3 convolution-relu block. Finally, a fusion convolution layer is used to restore the channel dimension. The fusion convolution layer consists of two 3×3 convolution layers and one 1×1 convolution layer.

[0072] The feature splitter is used to receive features from the feature aggregator and use multiple transposed convolutional layers for upsampling to obtain enhanced features, which are then delivered to the prompter module.

[0073] like Figure 4 As shown in Figure 1, the prompter module includes a Transformer encoder, a Transformer decoder, and a prompt linear projection layer. The Transformer encoder consists of N layers of stacked self-attention layers and feedforward layers, which are used to generate multi-scale features of the prompt. The Transformer decoder consists of M layers of stacked cross-attention layers and feedforward layers, which are used to generate enhanced prompt features.

[0074] The specific processing process of the prompter module is as follows: first, the enhanced features input to the prompter module are flattened after merging the position encoding and level encoding, and then input into the Transformer encoder. The Transformer encoder extracts the multi-scale features of the high-level prompt, as shown below.

[0075]

[0076] Among them, PE i represents the positional encoding of the i-th layer, LE i represents the level encoding of the i-th layer, Cat(·) represents the connection of the tensor along the channel dimension, and PE i LE i and the intermediate features from the i-th layer of the SAM encoder Merge into Indicates the generation of multiple layers of intermediate features, and T_enc represents the Transformer encoder layer.

[0077] The Transformer encoder extracts the multi-scale features of the high-level prompt and delivers them to the Transformer decoder. Combined with the initialization parameters, it recursively generates the enhanced prompt features of each layer; then it inputs the prompt linear projection layer, and finally combines it with the multi-scale features generated by the Transformer encoder to form a sparse prompt, as shown below.

[0078]

[0079] in, represents the zero-initialized learnable parameters and the multi-scale features from each block of the Transformer encoder Input Transformer decoding layer T_dec recursively to obtain feature parameters mlpprompt is a two-layer MLP for obtaining the prompt linear embedding e i ; sin represents the sine function, Indicates a sparse prompt.

[0080] Finally, the sparse prompts output by the prompter module are concatenated with the intermediate features produced by each layer of the SAM image encoder and input into the SAM image decoder.

[0081] In the SAM image decoder, the Transformer decoder block is used to interact between the total image features generated by the encoder and the hint embedding generated by the hinter to obtain the mask image. The specific process is as follows:

[0082] The SAM decoder receives the feature map from the SAM encoder and the prompter module, and uses the cross attention mechanism to compare the similarity between the query of the SAM decoder and the key and value of the SAM encoder, strengthen the important features, and suppress the unimportant features; finally, the feature conversion is performed through the feedforward neural network. In the last step of the SAM decoder, the output passes through the linear layer and the Softmax operation to generate the class probability for each pixel, thus producing the segmentation result.

[0083] Step 3: Train the medical image segmentation model built in step 2 based on the training data set in step 1, and then use the trained medical image segmentation model to segment the medical image to be segmented, and finally obtain a segmented image.

[0084] The process of training the medical image segmentation model is as follows. First, the SAM and prompter modules are combined to obtain the loss function of the medical image segmentation model. The expression of the loss function is:

[0085]

[0086] N p Indicates the number of prompt groups; Represents the cross entropy loss calculated between the predicted category and the target; represents the binary cross entropy loss between the predicted mask and the matched ground-truth instance mask, including both the predicted coarse-grained mask and the fine-grained mask; a i Indicates that the match is confirmed to be positive. and The calculation formula is:

[0087]

[0088] Among them, y represents a binary label on a certain category, and the value of y is 0 or 1; It indicates the probability of the corresponding category predicted by the neural network. Because it has been processed by softmax, The value of is between (0,1);

[0089] When y = 0, there is only the second term in the above formula. The closer it is to 0, the smaller the loss. When y=1, only the first term in the above formula has a value. The closer it is to 1, the smaller the loss.

[0090] Taking lung X-ray images as an example, the number of prompt groups N p Set to 5, corresponding to pneumonia, tuberculosis, tumor, COPD, and pulmonary embolism, and the corresponding number of predicted categories is 5. Calculate the cross entropy loss between the probability of each class and the true class, Calculate the cross entropy loss between the probability of the corresponding category of each point in the predicted mask image under one-hot encoding and the true corresponding probability of each point as 1; calculate the average of the losses calculated for each category to get the final loss value. Taking the true mask of the lung X-ray image as an example, y is the probability of each pixel in the true mask being 1, Indicates the probability that the predicted mask is 1.

[0091] After obtaining the loss function of the medical image segmentation model, the Adam optimization algorithm can be used to solve the minimum value of the loss function.

[0092] Finally, the trained medical image segmentation model is used to segment and evaluate the lung X-ray images, and the best segmentation result is obtained through comparative calculation.

[0093] The present invention uses the Dice coefficient to evaluate the quality of the segmentation effect. The larger the Dice value, the better the segmentation effect. The calculation formula is as follows:

[0094]

[0095] Among them, TP (True Positive) is the number of pixels correctly marked as foreground (object of interest), FP (False Positive) is the number of pixels incorrectly marked as foreground, and FN (False Negative) is the number of pixels incorrectly marked as background. This indicator takes into account both precision and recall, and is a balanced method to measure the segmentation effect.

[0096] The present invention fully considers the characteristics of high-dimensional data of medical images. The medical images obtain intermediate features through the SAM encoder, and use the category-related prompt mechanism to adaptively learn the best embedding prompts, locate objects and infer their semantic categories and instance masks, overcoming the problem that the type, position and number of SAM prompts will significantly affect the results generated by SAM, and can segment medical images more accurately and quickly.

[0097] The embodiments of the present invention are only used to illustrate the technical solutions of the present invention rather than to limit the present invention. It can be understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the attached claims and their equivalents.

Claims

1. A medical image segmentation method based on SAM and prompter, characterized in that: The steps include: Step 1: Obtain original medical images, preprocess the original medical images, and construct a training data set; Step 2: Build a medical image segmentation model based on the SAM model, which includes a SAM image encoder, a feature enhancement block, a prompter module, and a SAM image decoder; The SAM image encoder uses a multi-layer Transformer Block structure; The prompter module includes a Transformer encoder, a Transformer decoder, and a prompt linear projection layer; The SAM image decoder uses a multi-layer Transformer decoder structure; Among them, the processing process of the preprocessed original medical image in the medical image segmentation model is as follows: The preprocessed original medical image first passes through the SAM image encoder composed of multiple layers of Transformer Block to obtain the intermediate features generated by each layer, and then the feature enhancement block recursively accepts the features output by each layer of Transformer Block and performs enhancement processing; The enhanced features output by the feature enhancement block are input to the prompter module and processed as follows: First, the enhanced features input to the prompt module are flattened after merging the position encoding and level encoding, and then input to the Transformer encoder. The Transformer encoder extracts the multi-scale features of the high-level prompts, and then passes through the Transformer decoder, combined with the initialization parameters, to recursively generate the enhanced prompt features of each layer; then input to the prompt linear projection layer, and finally combined with the multi-scale features generated by the Transformer encoder to form a sparse prompt; The sparse prompts output by the prompter module are concatenated with the intermediate features generated by each layer of the SAM image encoder and input into the SAM image decoder; In the SAM image decoder, a Transformer decoder block is used to interact between the total image features generated by the encoder and the hint embedding generated by the hinter to obtain a mask image; Step 3: Train the medical image segmentation model built in step 2 based on the training data set in step 1, and then use the trained medical image segmentation model to segment the medical image to be segmented, and finally obtain a segmented image.

2. A medical image segmentation method based on SAM and prompter according to claim 1, characterized in that: In step 1, the preprocessing process includes unifying the image size and data normalization, and the specific process is as follows: First, an image with the largest pixel value in the original medical image is selected as the reference matrix, and the size of other original medical images is adjusted to match the reference matrix by filling zeros in the blank areas; Then, the pixel values ​​of all original medical images are divided by 255 to ensure that the pixel values ​​of all original medical images are uniformly between [0, 1].

3. The medical image segmentation method based on SAM and prompter according to claim 1, characterized in that: The SAM image encoder includes an embedding layer for converting the input image into a 16*16 block, and multiple layers of Transformer Block; The input image is split into fixed-size blocks through the embedding layer, each block is flattened and mapped into a vector; position encoding is added to the embedding vector of each block to add position information to the feature vector of each block; Each Transformer Block layer includes an attention layer with 256-dimensional input features and an MLP layer for amplifying the features by 4 times to 1024 dimensions. The attention layer uses a self-attention mechanism. The blocks with added position information from the embedding layer will pass through each layer of Transformer Block. Through the self-attention mechanism, the model can calculate the similarity between any two blocks in the image and combine the information of each block weightedly according to the similarity; After each layer of self-attention, residual connections are added to help avoid the vanishing gradient problem; At the end of each Transformer Block, a layer of MLP is passed to amplify the features by 4 times to 1024 dimensions for further nonlinear mapping in order to capture more complex visual information.

4. The medical image segmentation method based on SAM and prompter according to claim 1, characterized in that: The feature enhancement block includes a feature aggregator and a feature splitter; The feature aggregator is used to receive the intermediate feature maps output by each layer of the Transformer Block in the SAM image encoder and learn representative semantic features. The specific processing process is as follows: The feature maps output by each layer of Transformer Block in the SAM image encoder first use a 1×1 convolution-relu block to reduce the number of channels from the original number of feature maps to 32, and then use a 3×3 convolution-relu block to increase spatial information; then down-sample to produce down-sampled features, and merge with the features from the previous layer, and then pass through a 3×3 convolution-relu block, and finally pass through a fused convolution layer consisting of two 3×3 convolution layers and a 1×1 convolution layer. The fused convolution layer is used to restore the channel dimension; The feature splitter is used to receive features from the feature aggregator and use multiple transposed convolutional layers to perform upsampling to obtain enhanced features.

5. The medical image segmentation method based on SAM and prompter according to claim 1, characterized in that: In step 2, the Transformer encoder in the prompter module is composed of N layers of stacked self-attention layers and feed-forward layers; The process of the Transformer encoder extracting multi-scale features of high-level prompts is as follows: Among them, PE i represents the positional encoding of the i-th layer, LE i represents the level encoding of the i-th layer, Cat(·) represents the connection of the tensor along the channel dimension, and PE i LE i and the intermediate features from the i-th layer of the SAM encoder Merge into Indicates the generation of multiple layers of intermediate features, T_enc represents the Transformer encoder layer; The Transformer decoder consists of M layers of stacked cross-attention layers and feed-forward layers; The multi-scale features extracted by the Transformer encoder pass through the Transformer decoder and the prompt linear projection layer, and are finally combined with the multi-scale features to form a sparse prompt as shown below: in, represents the zero-initialized learnable parameters and the multi-scale features from each block of the Transformer encoder Input Transformer decoding layer T_dec recursively to obtain feature parameters mlp prompt is a two-layer MLP used to obtain the prompt linear embedding e i ; sin represents the sine function, Indicates a sparse prompt.

6. The medical image segmentation method based on SAM and prompter according to claim 1, characterized in that: The process of training the medical image segmentation model in step 3 is as follows: First, we combine SAM and the prompter module to obtain the loss function of the medical image segmentation model. The expression of the loss function is: Among them, N p Indicates the number of prompt groups; Represents the cross entropy loss calculated between the predicted category and the target; represents the binary cross entropy loss between the predicted mask and the matched ground-truth instance mask, including both the predicted coarse-grained mask and the fine-grained mask; a i Indicates that the match is confirmed to be positive; where, and The calculation formula is: Among them, y represents a binary label on a certain category, and the value of y is 0 or 1; Represents the probability of the corresponding category predicted by the neural network, The value of is between (0,1); When y = 0, there is only the second term in the above formula. The closer it is to 0, the smaller the loss. When y=1, only the first term in the above formula has a value. The closer it is to 1, the smaller the loss; Finally, the Adam optimization algorithm is used to solve the minimum value of the loss function.

Citation Information

Patent Citations

  • Interactive medical image intelligent segmentation modeling method

    CN108038862A

  • Medical image segmentation method combining curve structure prompts and deep neural network

    CN118314121A

  • 3D medical image segmentation method and device based on multi-scale self-prompting fine tuning

    CN118334060A

  • Cell nucleus segmentation method of cascade coding segmentation network based on large model guidance

    CN118366153A

  • Medical image cross-center generalization segmentation method and system based on query guidance

    CN118448014A

Cited By

  • Photovoltaic module waste prediction method based on geographic clustering and installation prediction

    CN120258258A

  • SAM-based automatic prompt ultrasonic image segmentation method and system

    CN120689610A

  • Pork marbling intelligent grading method and system based on image pre-training model

    CN120707924A

  • Skin scar image segmentation method

    CN121280463A

  • Instance segmentation and calculation method and system based on recursion prompt, medium and equipment

    CN121366170A