Image segmentation method and device based on SAM (Segmentation All Model)

By fine-tuning the mask decoder in the SAM segmentation all model, and segmenting the image and prompt information features, the segmentation difficulty of SAM in processing small objects or dense target scenes is solved, and the accuracy and efficiency of segmentation are improved.

CN120219398APending Publication Date: 2025-06-27BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202311792957.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Segmentation All Model SAM is difficult to perform accurate segmentation when processing small objects or internal objects in the image is too dense, and performs poorly when it is necessary to perform accurate segmentation of the background.

Method used

By setting the mask decoder in SAM as learnable parameters, and fine-tuning them using a backpropagation algorithm, combining image embedding features and prompt information embedding features for segmentation, the segmentation accuracy of the model is improved.

Benefits of technology

Improve the accuracy and efficiency of image segmentation, especially when dealing with small objects or dense scenes of targets inside the image, enhance the ability to segment the background.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219398A_ABST
    Figure CN120219398A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method and device based on an SAM (Segmentation All Model), and the method comprises the steps: obtaining a to-be-segmented image and first segmentation prompt information; inputting the to-be-segmented image and the first segmentation prompt information into a segmentation cutting model SAM obtained by pre-training to output a first image segmentation result; wherein the SAM obtained through pre-training is obtained by taking a mask decoder as a learnable parameter of model training and updating the weight of the learnable parameter through a back propagation algorithm. Therefore, according to the embodiment of the invention, the mask decoder of the SAM is set as the learnable parameter and the mask decoder is finely adjusted through the back propagation algorithm, and the whole SAM does not need to be adjusted, so that the fine adjustment efficiency of the model is ensured, the accuracy of the model is improved, and the accuracy and efficiency of image segmentation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to an image segmentation method and device based on the Segment Anything Model (SAM). Background Art

[0002] Image segmentation is an important research direction in the field of computer vision. Its purpose is to segment an image into different regions or objects for more advanced semantic analysis and processing.

[0003] The Segment Anything Model (SAM) is an open-source large image segmentation model. SAM adopts the method of a prompt-based model. It is based on the Vision Transformer architecture and performs segmentation by predicting the masks of objects in the image. SAM has trained more than 1 billion masks on 11 million images, achieving powerful zero-shot generalization ability. It has learned the general concept of objects, and this understanding enables it to perform zero-shot generalization on unfamiliar objects and images without additional training. Compared with other image segmentation models, SAM has higher segmentation accuracy and a wider application range.

[0004] However, SAM also has relatively large limitations in some specific scenarios. For example, in the case of small objects or overly dense internal targets in the image, SAM may fail to segment out these foreground targets; when precise segmentation of the background is required, SAM also performs poorly. Summary of the Invention

[0005] In view of this, this application provides an image segmentation method and device based on the Segment Anything Model (SAM) to overcome the limitations of SAM in some specific scenarios and improve the accuracy and efficiency of image segmentation.

[0006] The technical solution is as follows:

[0007] In a first aspect, an embodiment of this application provides an image segmentation method based on the Segment Anything Model (SAM), and the method includes:

[0008] Obtain an image to be segmented and first segmentation prompt information;

[0009] Input the image to be segmented and the first segmentation prompt information into a pre-trained Segment Anything Model (SAM) to output a first image segmentation result; wherein, the pre-trained Segment Anything Model (SAM) uses the mask decoder as the learnable parameters of model training and updates the weights of the learnable parameters through the backpropagation algorithm.

[0010] Optionally, the Segment Anything Model (SAM) further includes: an image encoder and a prompt encoder;

[0011] Inputting the image to be segmented and the first segmentation prompt information into a preset Segment Anything Model (SAM) to output a first image segmentation result includes:

[0012] Inputting the target image to be segmented into the image encoder to obtain image embedding features;

[0013] Inputting the first segmentation prompt information into the prompt encoder to obtain prompt information embedding features;

[0014] Inputting the image embedding features and the prompt information embedding features into the mask decoder to output the first image segmentation result.

[0015] Optionally, the prompt information embedding features include: prompt point embedding features, prompt box embedding features, prompt text embedding features, and prompt mask embedding features. The method further includes:

[0016] Adding the prompt point embedding features, the prompt box embedding features, and the prompt text embedding features to obtain updated prompt information embedding features;

[0017] Adding the prompt mask embedding features and the image embedding features to obtain updated image embedding features;

[0018] Inputting the image embedding features and the prompt information embedding features into the mask decoder to output the first image segmentation result includes:

[0019] Upsampling the image embedding features based on a transposed convolutional neural network to obtain upsampled image embedding features;

[0020] Combining the updated image embedding features and the updated prompt information embedding features based on a self-attention mechanism and a cross-attention mechanism to obtain updated prompt mask embedding features;

[0021] Determining target segmentation pixels based on the dot product result of the upsampled image embedding features and the updated prompt mask embedding features, where the target segmentation pixels are used to determine the first image segmentation result.

[0022] Optionally, the training steps of the Segment Anything Model (SAM) include:

[0023] Obtaining a target scene dataset, where the target scene dataset is associated with the scene of the image to be segmented;

[0024] Divide the target scene dataset into a training dataset, a validation dataset, and a test dataset;

[0025] Determine model training constraints according to image segmentation requirements, where the model training constraints include: mean squared error (MSE) loss function or cross-entropy (CE) loss function;

[0026] Under the model training constraints, based on the backpropagation algorithm (BP), train the initial Segment Anything Model (SAM) according to the prompt point information obtained by random sampling and the training dataset until the model loss meets the preset convergence condition. Among them, the mask decoder in the initial SAM is set as the learnable parameter for model training, and the image encoder and prompt encoder in the initial SAM are set as non-learnable parameters for model training.

[0027] Optionally, before dividing the target scene dataset into a training dataset, a validation dataset, and a test dataset, the method further includes:

[0028] Preprocess the target scene dataset to obtain a processed scene dataset, where the preprocessing includes at least one of: random rotation processing, random cropping and adjustment, size adjustment, and random inversion processing;

[0029] Normalize the processed scene dataset to obtain the scene dataset to be used.

[0030] Optionally, the method further includes:

[0031] Obtain second segmentation prompt information based on the first image segmentation result;

[0032] Input the target image to be segmented and the second segmentation prompt information into the pre-trained SAM to obtain a second image segmentation result.

[0033] In a second aspect, an embodiment of the present application provides an image segmentation device based on the SAM, where the device includes:

[0034] An acquisition module, configured to acquire an image to be segmented and first segmentation prompt information;

[0035] A segmentation module, configured to input the image to be segmented and the first segmentation prompt information into the pre-trained SAM to output a first image segmentation result; where the SAM uses the mask decoder as the learnable parameter for model training and updates the weights of the learnable parameters through the backpropagation algorithm.

[0036] Optionally, the SAM also includes an image encoder and a prompt encoder;

[0037] The segmentation module includes:

[0038] A feature extraction sub-module for inputting the target image to be segmented into the image encoder to obtain image embedding features;

[0039] The feature extraction sub-module is also used to input the first segmentation prompt information into the prompt encoder to obtain prompt information embedding features;

[0040] An image segmentation sub-module for inputting the image embedding features and the prompt information embedding features into the mask decoder to output the first image segmentation result.

[0041] Optionally, the prompt information embedding features include prompt point embedding features, prompt box embedding features, prompt text embedding features, and prompt mask embedding features. The segmentation module further includes:

[0042] A feature update sub-module for adding the prompt point embedding features, the prompt box embedding features, and the prompt text embedding features to obtain updated prompt information embedding features;

[0043] The feature update sub-module is also used to add the prompt mask embedding features and the image embedding features to obtain updated image embedding features;

[0044] The image segmentation sub-module is specifically configured to: upsample the image embedding features based on a transposed convolutional neural network to obtain upsampled image embedding features; combine the updated image embedding features and the updated prompt information embedding features based on a self-attention mechanism and a cross-attention mechanism to obtain updated prompt mask embedding features; determine target segmentation pixels based on the dot product operation result of the upsampled image embedding features and the updated prompt mask embedding features, and the target segmentation pixels are used to determine the first image segmentation result.

[0045] Optionally, the device further includes a model training module, and the model training module includes:

[0046] A data acquisition sub-module for acquiring a target scene dataset, which is associated with the scene of the image to be segmented;

[0047] A data division sub-module for dividing the target scene dataset into a training dataset, a validation dataset, and a test dataset;

[0048] A constraint determination sub-module for determining model training constraints according to image segmentation requirements, where the model training constraints include: mean squared error (MSE) loss function or cross-entropy (CE) loss function;

[0049] A model training sub-module for training an initial Segment Anything Model (SAM) based on the hint point information obtained by random sampling and the training data set under the model training constraints using the Backpropagation (BP) algorithm until the model loss meets the preset convergence condition. Among them, the mask decoder in the initial SAM is set as the learnable parameter for model training, and the image encoder and hint encoder in the initial SAM are set as non-learnable parameters for model training.

[0050] The above technical solution has the following beneficial effects:

[0051] An image segmentation method based on the Segment Anything Model (SAM) provided by an embodiment of the present application, when executing the method, obtains an image to be segmented and first segmentation hint information; inputs the image to be segmented and the first segmentation hint information into the pre-trained SAM to output a first image segmentation result; where the pre-trained SAM uses the mask decoder as the learnable parameter for model training and updates the weights of the learnable parameter through the Backpropagation algorithm. It can be seen that in the embodiment of the present application, the mask decoder of SAM is set as the learnable parameter and the mask decoder is fine-tuned through the Backpropagation algorithm, without adjusting the whole of SAM, which ensures the efficiency of model fine-tuning and improves the accuracy of the model, and improves the accuracy and efficiency of image segmentation.

[0052] An embodiment of the present application also provides a device corresponding to the above method, which has the same beneficial effects as the above method. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in this embodiment or the prior art, the following will briefly introduce the drawings required for the description of the embodiment or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0054] Figure 1 It is a schematic flowchart of an image segmentation method based on the Segment Anything Model (SAM) provided by an embodiment of the present application;

[0055] Figure 2 It is a schematic architecture diagram of a Segment Anything Model (SAM) provided by an embodiment of the present application;

[0056] Figure 3The test effect diagram of an image segmentation method based on the Segment Anything Model (SAM) provided by an embodiment of this application;

[0057] Figure 4 The structural schematic diagram of an image segmentation device based on the Segment Anything Model (SAM) provided by an embodiment of this application. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0059] As described in the background art, SAM has trained more than 1 billion masks on 11 million images, achieving powerful zero-shot generalization ability. It has learned the general concept of objects, and this understanding enables it to perform zero-shot generalization on unfamiliar objects and images without additional training. However, as a pre-trained model, in actual use, SAM cannot accurately perform segmentation in scenarios where small objects or internal targets in the image are too dense or when accurately segmenting the background.

[0060] To overcome the limitations of SAM in some specific scenarios and improve the accuracy and efficiency of image segmentation, an embodiment of this application provides an image segmentation method based on the Segment Anything Model (SAM), Figure 1 The flow schematic diagram of this method is shown, and this method may include:

[0061] Step S100: Obtain the image to be segmented and the first segmentation prompt information.

[0062] Specifically, the Segment Anything Model (SAM) is a prompt-based model. When SAM segments an image, in addition to the image to be segmented, segmentation prompt information also needs to be input. The segmentation prompt information may be at least one of segmentation prompt points, segmentation prompt boxes, segmentation prompt texts, and segmentation prompt masks. The first segmentation prompt information may be input manually according to the image segmentation requirements.

[0063] Step S200: Input the image to be segmented and the first segmentation prompt information into the pre-trained Segment Anything Model (SAM) to output the first image segmentation result; among them, the pre-trained Segment Anything Model (SAM) uses the mask decoder as the learnable parameters for model training and updates the weights of the learnable parameters through the backpropagation algorithm.

[0064] Specifically, since SAM, as a pre-trained model, performs poorly in scenarios such as dealing with small objects, scenes where objects inside the image are too dense, or precise segmentation of the background, in the embodiments of the present application, the mask decoder in SAM is set as a learnable parameter and trained through the Backpropagation (BP) algorithm, thereby achieving fine-tuning of some parameters of SAM, improving the model segmentation accuracy, and ensuring the efficiency of model fine-tuning. In the embodiments of the present application, the image to be segmented and the first segmentation prompt information are input into the pre-trained SAM, and the image segmentation result of the image to be segmented can be obtained.

[0065] The following combines Figure 2 the schematic architecture diagram of a Segment Anything Model (SAM) shown to illustrate the working principle of SAM.

[0066] The Segment Anything Model (SAM) is a network architecture based on the encoder-decoder form. SAM includes: an image encoder, a prompt encoder, and a mask decoder. Among them, the image encoder is used to receive the image to be segmented and output image embedding features; the prompt encoder is used to receive artificial prompt points and output prompt embedding features; then the image embedding features and the prompt embedding features are input into the mask decoder to obtain the image segmentation result corresponding to the prompt points.

[0067] In the embodiments of the present application, since we only need to perform efficient fine-tuning on the mask decoder, in practical applications, when deploying the SAM model, only the mask decoder can be set as a learnable parameter, while the gradients of the image encoder and the prompt encoder are masked to ensure that their internal parameters are not updated by backpropagation during training, and at the same time improve the efficiency of model fine-tuning.

[0068] The following explains the working principles of the image encoder, the prompt encoder, and the mask decoder.

[0069] (1) Image encoder: Receives the image to be segmented and processes it into a complex image embedding feature map. Specifically, the image encoder E ViT consists of a Vision Transformer pre-trained with MAE self-supervision. Therefore, for the image x, its image embedding feature F x is calculated as follows:

[0070] F x = E ViT (x)

[0071] (2) Prompt encoder: Receives prompt information, including: prompt points, prompt boxes, prompt texts, prompt masks, etc., and processes them into a series of prompt information embedding feature vectors.

[0072] The prompt encoder designs different encoders for different prompt forms:

[0073] For prompt point and prompt box information, use the Fourier Neural Tangent Kernel E Fourier For the prompt point P points and the box P boxes are encoded to obtain the prompt point embedding feature and the prompt box embedding feature:

[0074] F points = E Fourier (P points ) F boxes = E Fourier (P boxes )

[0075] For prompt text information, use the text encoder E of the pre-trained CLIP model (Contrastive Language-Image Pre-Training, CLIP) CLIP For the prompt text P texts are encoded to obtain the prompt text embedding feature:

[0076] F texts = E CLIP (P texts )

[0077] For the new prompt mask, use a convolutional neural network E CNN For the prompt mask P masks are encoded to obtain the prompt mask embedding feature:

[0078] F masks = E CNN (P masks )

[0079] Before the obtained feature vectors are input to the mask decoder, the fusion of these features is required. Add the prompt mask embedding feature F masks to the image embedding feature F of the same length and width x to obtain the updated image embedding feature; add the prompt point embedding feature F points and the prompt box embedding feature F boxes and the prompt text embedding feature F texts to obtain the updated prompt information embedding feature:

[0080] F' x = F x + F masks F prompts = F points + F boxes + F texts

[0081] (3) Mask decoder: Combines the image embedding feature map and the prompt embedding feature vector, and processes to obtain the image segmentation result corresponding to the prompt input.

[0082] Specifically, on the one hand, the mask decoder uses a transposed convolutional neural network to upsample the updated image embedding feature to increase the resolution and obtain a high-resolution image embedding feature T x :

[0083] T x = D CNN (F x )

[0084] On the other hand, it uses the self-attention and cross-attention mechanisms to combine the updated image embedding F' x and the updated prompt embedding F prompts to update and obtain a series of prompt mask embedding features T masks :

[0085] T masks = D Attention (F' x , F prompts )

[0086] Finally, use the dot product method to find the pixel positions related to each prompt mask embedding feature vector in the image embedding feature, and obtain the segmentation result M corresponding to each prompt mask embedding feature vector prompts :

[0087] M prompts = T masks ·T x

[0088] After obtaining the segmentation result, according to the number of categories included in the training data, the mean square error loss function (MSE, Mean Square Error) or the cross-entropy loss function (CE, Cross Entropy) can be used to calculate the model loss as follows:

[0089] L2 = MSE(M prompts , GT prompts )

[0090] L n = CE(M prompts , GT prompts ), n > 2

[0091] Among them, when the loss function is used as the model training constraint, it is preset. In the embodiments of the present application, in the case of only segmenting the foreground and background, the mean square error loss function is adopted; when multiple categories need to be segmented, the cross-entropy loss function is adopted.

[0092] Based on the working principle of the above-mentioned SAM, in the embodiments of the present application, step S200 may specifically include the following steps.

[0093] Step S201: Input the target image to be segmented into an image encoder to obtain image embedding features.

[0094] Step S202: Input the first segmentation prompt information into the prompt encoder to obtain prompt information embedding features; wherein, the prompt information embedding features include: prompt point embedding features, prompt box embedding features, prompt text embedding features, and prompt mask embedding features.

[0095] Step S203: Add the prompt point embedding features, prompt box embedding features, and prompt text embedding features to obtain updated prompt information embedding features, and add the prompt mask embedding features and the image embedding features to obtain updated image embedding features.

[0096] Step S204: Upsample the updated image embedding features based on a transposed convolutional neural network to obtain sampled image embedding features.

[0097] Step S205: Combine the updated image embedding features and the updated prompt information embedding features based on a self-attention mechanism and a cross-attention mechanism to obtain updated prompt mask embedding features.

[0098] Step S206: Determine target segmentation pixels based on the dot product result of the sampled image embedding features and the updated prompt mask embedding features, and the target segmentation pixels are used to determine the first image segmentation result.

[0099] Specifically, steps S201 to S206 may refer to the working principle of the above-mentioned SAM, which will not be elaborated here.

[0100] In an alternative implementation, the image segmentation method provided in the embodiments of the present application further includes a training step for the segmentation model SAM, including:

[0101] Step S301: Obtain a target scene dataset, and the target scene dataset is associated with the scene of the image to be segmented.

[0102] Specifically, obtain a target scene dataset associated with the scene of the image to be segmented. For example, if the image to be segmented is an urban landscape, the target scene dataset is also an image dataset related to the urban landscape. Among them, the target scene dataset can be selected to include a real image dataset and a mask image dataset. Exemplarily, when the image to be segmented is an urban landscape, in this embodiment, the real image dataset is the Cityscapes dataset, which contains 5,000 images captured by on-vehicle cameras and fine-labeled semantic segmentation masks with 19 categories.

[0103] Step S302: Divide the target scene dataset into a training dataset, a validation dataset, and a test dataset.

[0104] Specifically, the dataset is divided according to the given criteria of each dataset or according to a preset division ratio. Exemplarily, for the Cityscapes dataset, 3,475 images are used for training, 500 images are used for validation, and 1,525 images are used for testing. Among the 19 categories, the Road category is selected as the Free Space category in the autonomous driving scene.

[0105] Step S303: Determine the model training constraints according to the image segmentation requirements. The model training constraints include: the mean square error (MSE) loss function or the cross-entropy (CE) loss function.

[0106] Specifically, if it is necessary to segment the foreground and background, the mean square error (MSE) loss function is used as the model training constraint. When multiple categories need to be segmented, the cross-entropy (CE) loss function is used as the model training constraint.

[0107] Step S304: Under the model training constraints, based on the backpropagation algorithm BP, train the initial segment-anything model SAM according to the prompt point information obtained by random sampling and the training dataset until the model loss meets the preset convergence condition. Among them, the mask decoder in the initial segment-anything model SAM is set as the learnable parameter of the model training, and the image encoder and the prompt encoder in the initial segment-anything model SAM are set as the non-learnable parameters of the model training.

[0108] Specifically, using the model training constraints determined in the above steps, adopt the backpropagation algorithm BP to iteratively update and optimize the learnable network parameters (i.e., the mask decoder) in the initial segment-anything model SAM according to the prompt point information obtained by random sampling and the training dataset until the model loss tends to converge, thereby obtaining the fine-tuned SAM, that is, the pre-trained segment-anything model SAM provided in this application.

[0109] In actual model training, both the fine-tuning training and evaluation of SAM are completed on the PyTorch platform. The model training is completed on a single NVIDIA GTX 3080TI GPU (12GB), the batch size is set to 4, and the AdamW optimizer is used to optimize our generator and discriminator. The initial learning rate is 1×10 -6 The model is efficiently fine-tuned, and then the learning rate is gradually decreased to 1×10 using a cosine schedule. -7 For the Cityscapes dataset, the model needs to be trained for 100 epochs to fit.

[0110] It should be noted that the prompt point information used in the SAM fine-tuning training is obtained based on random sampling. In practical applications, several points can be randomly sampled within the region belonging to each category according to the semantic segmentation annotation given in the dataset to simulate artificial prompt points and improve the training efficiency. Moreover, this method of random sampling rather than manual annotation helps to improve the generalization and robustness of the model.

[0111] Specifically, for the model training fine-tuning of the Cityscapes dataset, 5 points can be randomly sampled within the Road category region as positive prompt points to be provided to SAM (prompting the model to segment the region where these 5 points are located), and 10 points can be randomly sampled in other regions as negative prompt points to be provided to SAM (prompting the model not to segment the region where these 10 points are located).

[0112] In an optional implementation manner, before the above step S302, the method further includes: preprocessing the target scene dataset to obtain a processed scene dataset. The preprocessing includes at least one of random rotation processing, random cropping adjustment, size adjustment, and random inversion processing; normalizing the processed scene dataset to obtain the scene dataset to be used.

[0113] Specifically, the image enhancement operations can include random image rotation, random image cropping, size adjustment, and random inversion, etc. Exemplarily, in the embodiments of the present application, images with a size of 1024×1024 are randomly cropped from the images for training, and two image enhancement methods of randomly horizontally inverting with a 50% probability are adopted. Finally, the images are normalized according to the mean and variance of each RGB channel commonly used. The image preprocessing operations including image enhancement operations and image normalization operations can improve the image quality of the training dataset and balance the training samples, thereby improving the efficiency of model training.

[0114] In an alternative implementation, the image segmentation method provided by the embodiments of the present application may further include: obtaining second segmentation prompt information based on the first image segmentation result; inputting the target image to be segmented and the second segmentation prompt information into the pre-trained Segment Anything Model (SAM) to obtain a second image segmentation result.

[0115] Specifically, since SAM is a human-computer interaction image segmentation model, after obtaining the first image segmentation result, the segmentation result can be subjectively evaluated manually. If the segmentation result needs to be further adjusted, hint points can be added / deleted, that is, the second segmentation prompt information is further determined based on the first image segmentation result, and then the image to be segmented is segmented again through the second segmentation prompt information to obtain the second image segmentation result. And the above steps can be repeated until the obtained image segmentation result does not need to be adjusted.

[0116] For further reference Figure 3 to the test effect diagram of an image segmentation method based on the Segment Anything Model (SAM) shown in Figure 3 In the upper part of the effect diagram, the vehicle in front is well segmented and recognized, Figure 3 In the middle part of the effect diagram, the left lane line is well segmented and recognized, Figure 3 In the lower part of the effect diagram, the road surface is well segmented and recognized. Thus, the image segmentation method provided by the embodiments of the present application can effectively complete the human-computer interactive image segmentation task and performs well in the segmentation tasks of multiple categories (cars, lane lines, and road surfaces).

[0117] In summary, the embodiments of the present application provide an image segmentation method based on the Segment Anything Model (SAM), which performs efficient fine-tuning on the downstream task dataset based on SAM, makes up for the defect that the SAM model has poor background segmentation effect, and at the same time optimizes the efficiency of fine-tuning the pre-trained model through partial parameter fine-tuning, realizing the rapid adaptation of the human-computer interactive segmentation model.

[0118] Corresponding to the above method, the embodiments of the present application also provide an image segmentation device based on the Segment Anything Model (SAM). Please refer to Figure 4 , which shows the structural schematic diagram of the device. The device may include:

[0119] An acquisition module 401, configured to acquire the image to be segmented and the first segmentation prompt information;

[0120] A segmentation module 402, configured to input the image to be segmented and the first segmentation prompt information into the pre-trained Segment Anything Model (SAM) to output a first image segmentation result; wherein, the pre-trained Segment Anything Model (SAM) uses the mask decoder as the learnable parameter for model training and updates the weights of the learnable parameters through the backpropagation algorithm.

[0121] In an alternative implementation, the segmentation model SAM includes: an image encoder and a prompt encoder;

[0122] The segmentation module 402 includes:

[0123] A feature extraction sub-module for inputting the target image to be segmented into the image encoder to obtain image embedding features;

[0124] The feature extraction sub-module is also used to input the first segmentation prompt information into the prompt encoder to obtain prompt information embedding features;

[0125] An image segmentation sub-module for inputting the image embedding features and the prompt information embedding features into the mask decoder to output a first image segmentation result.

[0126] In an alternative implementation, the prompt information embedding features include: prompt point embedding features, prompt box embedding features, prompt text embedding features, and prompt mask embedding features. The segmentation module 402 further includes:

[0127] A feature update sub-module that adds the prompt point embedding features, the prompt box embedding features, and the prompt text embedding features to obtain updated prompt information embedding features;

[0128] The feature update sub-module is also used to add the prompt mask embedding features and the image embedding features to obtain updated image embedding features;

[0129] The image segmentation sub-module is specifically configured to: upsample the image embedding features based on a transposed convolutional neural network to obtain upsampled image embedding features; combine the updated image embedding features and the updated prompt information embedding features based on a self-attention mechanism and a cross-attention mechanism to obtain updated prompt mask embedding features; determine target segmentation pixels based on the dot product operation result of the upsampled image embedding features and the updated prompt mask embedding features, and the target segmentation pixels are used to determine the first image segmentation result.

[0130] In an alternative implementation, the device further includes: a model training module, and the model training module includes:

[0131] A data acquisition sub-module for acquiring a target scene data set, and the target scene data set is associated with the scene of the image to be segmented;

[0132] A data division sub-module for dividing the target scene data set into a training data set, a validation data set, and a test data set;

[0133] A constraint determination sub-module, configured to determine model training constraints according to image segmentation requirements, where the model training constraints include: mean squared error (MSE) loss function or cross-entropy (CE) loss function;

[0134] A model training sub-module, configured to train an initial Segment Anything Model (SAM) based on the hint point information obtained by random sampling and the training data set under the model training constraints using the backpropagation algorithm (BP) until the model loss meets a preset convergence condition. Among them, the mask decoder in the initial Segment Anything Model (SAM) is set as a learnable parameter for model training, and the image encoder and hint encoder in the initial Segment Anything Model (SAM) are set as non-learnable parameters for model training.

[0135] In an optional implementation manner, the device provided in the embodiments of the present application further includes: a data preprocessing module, configured to preprocess the target scene data set to obtain a processed scene data set, and the preprocessing includes at least one of random rotation processing, random cropping and adjustment, size adjustment, and random inversion processing;

[0136] A data normalization module, configured to normalize the processed scene data set to obtain a scene data set to be used.

[0137] In an optional implementation manner, the acquisition module 401 is further configured to obtain second segmentation hint information based on the first image segmentation result;

[0138] The segmentation module 402 is further configured to input the target image to be segmented and the second segmentation hint information into the pre-trained Segment Anything Model (SAM) to obtain a second image segmentation result.

[0139] It should be noted that the steps executed by each module in the image segmentation device based on the Segment Anything Model (SAM) provided in the embodiments of the present application and the related technical features correspond to the method provided in the embodiments of the application. The description of the device part can refer to the embodiments of the foregoing method part, and will not be elaborated here.

[0140] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0141] Those skilled in the art can understand that the flowchart shown in the figure is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of this flowchart.

[0142] In several embodiments provided in the present application, it should be understood that the disclosed methods, devices, and equipment can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0143] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0144] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0145] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image segmentation method based on the Segment Anything Model (SAM), characterized in that, The method includes: Obtaining an image to be segmented and first segmentation prompt information; Inputting the image to be segmented and the first segmentation prompt information into a pre-trained Segment Anything Model (SAM) to output a first image segmentation result; wherein, the pre-trained Segment Anything Model (SAM) uses a mask decoder as the learnable parameter for model training and updates the weights of the learnable parameter through the backpropagation algorithm.

2. The method according to claim 1, wherein The Segment Anything Model (SAM) includes: an image encoder and a prompt encoder; The step of inputting the image to be segmented and the first segmentation prompt information into the preset Segment Anything Model (SAM) to output a first image segmentation result includes: Inputting the target image to be segmented into the image encoder to obtain image embedding features; Inputting the first segmentation prompt information into the prompt encoder to obtain prompt information embedding features; Inputting the image embedding features and the prompt information embedding features into the mask decoder to output the first image segmentation result.

3. The method according to claim 2, wherein The prompt information embedding features include: prompt point embedding features, prompt box embedding features, prompt text embedding features, and prompt mask embedding features. The method further includes: Adding the prompt point embedding features, the prompt box embedding features, and the prompt text embedding features to obtain updated prompt information embedding features; Adding the prompt mask embedding features and the image embedding features to obtain updated image embedding features; The step of inputting the image embedding features and the prompt information embedding features into the mask decoder to output the first image segmentation result includes: Upsampling the updated image embedding features based on a transposed convolutional neural network to obtain sampled image embedding features; Combining the updated image embedding features and the updated prompt information embedding features based on a self-attention mechanism and a cross-attention mechanism to obtain updated prompt mask embedding features; Determining target segmentation pixels based on the dot product result of the sampled image embedding features and the updated prompt mask embedding features, and the target segmentation pixels are used to determine the first image segmentation result.

4. The method according to claim 1, wherein The training steps of the Segment Anything Model (SAM) include: Obtaining a target scene dataset, where the target scene dataset is associated with the scene of the image to be segmented; Dividing the target scene dataset into a training dataset, a validation dataset, and a test dataset; Determining model training constraints according to image segmentation requirements, where the model training constraints include: mean squared error (MSE) loss function or cross-entropy (CE) loss function; Under the model training constraints, training the initial Segment Anything Model (SAM) based on the backpropagation algorithm (BP) according to the randomly sampled prompt point information and the training dataset until the model loss meets a preset convergence condition, where the mask decoder in the initial Segment Anything Model (SAM) is set as the learnable parameter for model training, and the image encoder and the prompt encoder in the initial Segment Anything Model (SAM) are set as non-learnable parameters for model training.

5. The method according to claim 4, wherein Before dividing the target scene dataset into a training dataset, a validation dataset, and a test dataset, the method further includes: Preprocessing the target scene dataset to obtain a processed scene dataset, where the preprocessing includes at least one of random rotation processing, random cropping and adjustment, size adjustment, and random inversion processing; Performing normalization processing on the processed scene dataset to obtain a scene dataset to be used.

6. The method according to claim 1, characterized in that, The method further includes: Obtaining second segmentation prompt information based on the first image segmentation result; Inputting the target image to be segmented and the second segmentation prompt information into the pre-trained Segment-Anything Model SAM to obtain a second image segmentation result.

7. An image segmentation device based on the Segment Anything Model (SAM), characterized in that, The device includes: An acquisition module, configured to acquire an image to be segmented and first segmentation prompt information; A segmentation module, configured to input the image to be segmented and the first segmentation prompt information into the pre-trained Segment-Anything Model SAM to output a first image segmentation result; wherein, the Segment-Anything Model SAM uses a mask decoder as the learnable parameters of model training and updates the weights of the learnable parameters through the backpropagation algorithm.

8. The device according to claim 7, wherein, The Segment-Anything Model SAM further includes: an image encoder and a prompt encoder; The segmentation module includes: A feature extraction sub-module, configured to input the target image to be segmented into the image encoder to obtain image embedding features; The feature extraction sub-module is further configured to input the first segmentation prompt information into the prompt encoder to obtain prompt information embedding features; An image segmentation sub-module, configured to input the image embedding features and the prompt information embedding features into the mask decoder to output the first image segmentation result.

9. The device according to claim 8, characterized in that, The prompt information embedding features include: prompt point embedding features, prompt box embedding features, prompt text embedding features, and prompt mask embedding features. The segmentation module further includes: A feature update sub-module, configured to add the prompt point embedding features, the prompt box embedding features, and the prompt text embedding features to obtain updated prompt information embedding features; The feature update sub-module is further configured to add the prompt mask embedding features and the image embedding features to obtain updated image embedding features; The image segmentation sub-module is specifically configured to: perform upsampling on the image embedding features based on a transposed convolutional neural network to obtain upsampled image embedding features; combine the updated image embedding features and the updated prompt information embedding features based on a self-attention mechanism and a cross-attention mechanism to obtain updated prompt mask embedding features; determine target segmentation pixels based on the dot product operation result of the upsampled image embedding features and the updated prompt mask embedding features, and the target segmentation pixels are used to determine the first image segmentation result.

10. The device according to claim 7, characterized in that, The device further includes: a model training module, and the model training module includes: A data acquisition sub-module, configured to acquire a target scene dataset, where the target scene dataset is associated with the scene of the image to be segmented; A data division sub-module for dividing the target scenario data set into a training data set, a validation data set, and a test data set; A constraint determination sub-module for determining model training constraints according to image segmentation requirements, where the model training constraints include: mean squared error (MSE) loss function or cross-entropy (CE) loss function; A model training sub-module for training an initial Segment Anything Model (SAM) based on the backpropagation algorithm (BP) according to the prompt point information obtained by random sampling and the training data set under the model training constraints until the model loss meets the preset convergence condition. Among them, the mask decoder in the initial SAM is set as a learnable parameter for model training, and the image encoder and prompt encoder in the initial SAM are set as non-learnable parameters for model training.

Citation Information

Cited By

  • High-dimensional coupling segmentation method for target fine semantics based on SAM

    CN121861293A

  • A high-dimensional coupled segmentation method based on SAM for fine-grained target semantics

    CN121861293B

  • Trademark infringement detection method and system based on commodity pictures

    CN121962872A

  • A trademark infringement detection method and system based on commodity pictures

    CN121962872B