A dynamic decision-making image segmentation method based on self-prompt guidance
By introducing a self-prompt guidance mechanism and dynamic decision strategy in the image segmentation method, combining the task self-promptrator and the SAM model driven by spatial position information, the problem of insufficient model adaptability and dynamic decision-making capabilities in the prior art is solved, and more efficient and accurate image automatic segmentation is achieved.
Patent Information
- Application Number
- CN202510248236.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing automatic segmentation method based on SAM model has shortcomings in model adaptability, interactive dependence and dynamic decision-making capabilities. Especially when processing complex image content, the segmentation effect is poor and it is difficult to cope with large-scale data processing and dynamically changing image content.
A dynamic decision image segmentation method based on self-prompt guidance is proposed. By constructing a task self-prompt guide-guided SAM model and spatial position information driven by driving self-prompt guide-guided SAM model, and fusing it, dynamically adjusting the inference strategy to generate multiple mask areas, and achieving precision and semantic consistency through synthesis, noise removal and region screening.
It significantly improves the accuracy and efficiency of automatic image segmentation, reduces the need for manual intervention, enhances the model's processing ability of complex images, and can better adapt to the diversity and variability of image content.
Smart Images

Figure CN119762787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a dynamic decision-making image segmentation method based on self-prompt guidance. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the breakthroughs in the fields of deep learning and computer vision, image segmentation technology has gradually become one of the key technologies in multiple industries, including remote sensing images, medical images, autonomous driving, industrial inspection and other fields. Image segmentation is not only the core link of computer vision tasks, but also the basis for various practical applications. However, existing image segmentation methods still face many bottlenecks when facing the need for large-scale data annotation.
[0003] Traditional image segmentation methods rely on manual annotation or early algorithm technologies, such as template matching, region growing, etc. These methods have acceptable effects in simple scenarios, but when dealing with complex image content (such as multi-object scenarios, occlusion, blurred boundaries), they often face great challenges, with limited segmentation effects, low accuracy, and difficulty in meeting the needs of large-scale data processing. In addition, the computational efficiency of traditional methods is low, and it is difficult to achieve flexible adaptive adjustment, which limits their scalability in practical applications.
[0004] In recent years, the booming development of deep learning technology, especially the introduction of generative adversarial networks, self-attention mechanisms, and large-scale pre-trained models, has significantly improved image segmentation technology. Pre-trained models represented by the SAM model, with their powerful instance segmentation capabilities in natural images and excellent performance in the zero-shot case, have become an important breakthrough in the field of image segmentation. However, although SAM performs well in the application of natural images, when applied to fields such as remote sensing images and medical images, it still faces several problems. First, the SAM model is not optimized for the needs of specific fields, and its segmentation effect is often unsatisfactory in scenarios such as complex backgrounds, object deformations, or occlusions. In addition, as an interactive segmentation model independent of categories, the SAM model lacks the ability to distinguish object categories, and its segmentation results rely heavily on manually provided prompts (such as points, boxes, or coarse-grained masks). This artificial dependence limits the application of the model in automated segmentation tasks, especially in large-scale data processing, where the efficiency of manual prompts is low and rapid batch processing cannot be achieved.
[0005] In addition, existing segmentation methods based on the SAM model lack an effective dynamic decision-making mechanism. Many methods rely on fixed local regions for segmentation and lack flexible dynamic decision-making capabilities. When the target morphology changes or is occluded, it is difficult for existing technologies to adaptively adjust the segmentation strategy, resulting in insufficient annotation accuracy and stability and being unable to cope with the diversity and variability of image content. Finally, the insufficient automatic segmentation accuracy remains an urgent problem to be solved. Although existing technologies can provide relatively accurate segmentation results in simple scenarios, in complex image scenarios, especially remote sensing images and medical images, there are still problems such as mis-segmentation, over-segmentation, or non-segmentation, which seriously affect the quality of the segmentation results and restrict their popularization and application effects in practical applications. Summary of the Invention
[0006] Aiming at the deficiencies of existing automatic segmentation methods based on the SAM model in terms of model adaptability, interactive dependence, and dynamic decision-making ability, the present invention proposes a dynamic decision-making image segmentation method guided by self-prompting.
[0007] The object of the present invention is achieved through the following technical solutions: A dynamic decision-making image segmentation method guided by self-prompting, comprising the following steps:
[0008] Obtain the original image of the target field as a sample, take out part of the samples for manual annotation and make corresponding labels to construct a training data set, and the training data set includes a training set and a validation set;
[0009] Construct a SAM model guided by a task self-promptor and a SAM model guided by a spatial position information-driven self-promptor respectively according to the training data set;
[0010] Train and fine-tune the two SAM models;
[0011] After fusing the two fine-tuned SAM models, the fused model can dynamically adjust the inference strategy according to the specific image content and task requirements of the image to be annotated, so as to generate multiple mask regions;
[0012] Through the synthesis of multiple mask regions, noise removal, region screening, and similarity-based optimization, the accuracy and semantic consistency of the mask image are realized, the quality and visualization effect of the segmentation result are improved, and finally a semantic-level instance segmentation mask image is obtained.
[0013] Further, the construction of the training dataset includes: First, ensure that the extracted samples cover the diversity of the target objects; Then, for each original image, use an artificial segmentation annotation tool to generate a segmentation mask of the target object, while annotating the bounding box and class label of the target; Finally, based on the original image and the generated segmentation mask, bounding box and class label, construct a training dataset for training and validation, and randomly divide it into a training set and a validation set according to a ratio of 9:1.
[0014] Further, the SAM model guided by the task self-prompter includes: an image encoder module, a feature aggregator module, a feature separator module, a task-based prompter module, and a mask decoder module;
[0015] The image encoder module uses a pre-trained SAM visual encoder to extract image features;
[0016] The feature aggregator module fuses and processes the input features through convolution and normalization operations to extract and integrate information at different scales;
[0017] The feature separator module integrates features at different scales through upsampling, downsampling, convolution, and lateral connection operations to enhance the multi-scale perception ability of the model;
[0018] The task-based prompter module includes: a transformer encoder, which performs self-attention operations on four different-scale input feature maps from the feature separator module, extracts global semantic information, and generates four-scale feature maps; a transformer decoder, which interacts with the zero-initialized task-related query tokens and fuses with the four-scale feature maps; and generates a mask filter, a prompt embedding, and a semantic class for each instance through a mask projection layer, a prompt projection layer, and a class derivation layer, and finally generates a rough mask according to the largest-scale feature map;
[0019] The mask decoder module generates a high-precision image segmentation mask through image embedding, mask tokenization, 2 dual-path transformer modules, 4 multi-layer perceptron modules, and 2 transposed convolutional layers for upsampling.
[0020] Further, the SAM model guided by the self-prompter driven by spatial position information includes: an image encoder module, a feature aggregator module, a feature separator module, a prompter module driven by spatial position information, and a mask decoder module;
[0021] The image encoder module uses a pre-trained SAM visual encoder to extract image features;
[0022] The feature aggregator module fuses and processes the input features through convolution and normalization operations to extract and integrate information at different scales;
[0023] The feature separator module integrates features of different scales through upsampling, downsampling, convolution, and lateral connection operations to enhance the multi-scale perception ability of the model;
[0024] The spatial position information-driven prompter module includes: a lightweight region proposal network for generating class scores and bounding box regression values from multi-scale input feature maps, removing redundant boxes through non-maximum suppression, and outputting multiple region proposal boxes; a position encoding module for applying sine position encoding to the input feature map to enhance its spatial position information; a region of interest alignment layer for extracting features corresponding to each proposal box from multi-scale feature maps; a shared feature extractor for processing the extracted features and generating shared features; a class derivation layer and a bounding box derivation layer for generating the semantic class and bounding box of each instance respectively; a region of interest feature extraction module for extracting target features; a spatial information prompter layer for generating sparse target features and transmitting them to the mask decoder to generate the final multi-semantic segmentation mask;
[0025] The mask decoder module generates a high-precision image segmentation mask through image embedding, mask tokenization, two dual-path transformer modules, four multi-layer perceptron modules, and two upsampling transposed convolutional layers.
[0026] Furthermore, the training and fine-tuning of the two SAM models include:
[0027] Preprocess the training dataset, and the preprocessing includes image loading, normalization, data augmentation, target screening, and image size unification to ensure the consistency and effectiveness of image and mask information;
[0028] Use the preprocessed data as input and input it into the SAM model guided by the task self-prompter and the SAM model guided by the spatial position information-driven self-prompter. Freeze the parameters of the image encoder and fine-tune the parameters of the feature aggregator, feature separator, prompt encoder, and mask decoder modules;
[0029] The training parameters are fine-tuned using the AdamW optimizer. The training parameters include weight decay, initial learning rate, cosine annealing scheduler, linear warm-up strategy, batch size, and number of iterative updates. The total loss function is calculated for two SAM models respectively. Among them, for the SAM model guided by the task self-prompter, the Hungarian algorithm is used to match the predicted mask with the true instance mask, and the class matching loss, mask cross-entropy loss, and mask Dice loss are calculated to ensure that each predicted instance is paired with the most suitable true instance, and supervised learning of multi-class classification loss and mask classification loss is carried out. For the SAM model guided by the self-prompter driven by spatial position information, the total loss includes region proposal loss, classification loss, regression loss, and mask loss. Hyperparameter optimization is performed using the training set, the model performance is evaluated using the validation set, and the best model configuration is selected based on the mean average precision (mAP). Finally, the model with the best performance on the validation set is selected.
[0030] Furthermore, the fusion of the two fine-tuned SAM models, and the inference strategy of the fused model dynamically adjusts according to the specific image content and task requirements of the image to be annotated, including:
[0031] The image to be annotated is preprocessed, including loading the image, performing normalization, scaling it to the target size proportionally, and making necessary padding to ensure that the image meets the model input requirements.
[0032] The preprocessed image is sent into the fused model for inference. The fused model includes two guiding mechanisms: the SAM model guided by the task self-prompter, which is used to adaptively select the most suitable feature extraction and inference method according to the task requirements to ensure that the task characteristics are precisely processed; the SAM model driven by spatial position information, which is used to improve the model's ability to model the spatial features of the target by analyzing the distribution and layout of the target in space. A set of instance segmentation masks are generated by the fused model, representing the regions of each instance in the image. Each mask is stored together with its corresponding label, score, and other information as part of the inference result.
[0033] After generating multiple instance segmentation masks, each mask is processed by the fused model to dynamically select the optimal mask.
[0034] Furthermore, the processing of each mask by the fused model includes:
[0035] The area of each mask is calculated by the fused model, and the valid masks with larger areas are screened out.
[0036] If there are overlapping regions of multiple masks, the fused model dynamically determines which masks to retain based on the area of the overlapping regions and the score values of each mask; among them, for overlapping regions with an overlapping area smaller than the set threshold, the mask with a larger area is selected; while for larger overlapping regions, the mask with a higher score is preferentially selected;
[0037] The selected masks, labels, and scores are stored in the final result array for subsequent analysis and processing.
[0038] Furthermore, the realization of the precision and semantic consistency of the mask image through the synthesis of multi-mask regions, noise removal, region screening, and similarity-based optimization includes:
[0039] Extract multiple masks and class labels from the model fusion process; according to each class label, obtain the corresponding color from the color palette and apply the color to the corresponding mask region; by assigning different colors to each binary mask, multiple masks are drawn on the same image to generate an instance mask map containing semantic information, and the semantic category of the map is represented by the corresponding color to ensure that each region can clearly identify its semantic category;
[0040] For the synthesized mask image, perform connected component labeling on the masks of each color to identify different regions in the image; set an area threshold and filter out connected components smaller than the threshold to remove noise regions; this process ensures that only larger valid regions are retained in the image; merge the screened mask regions into the final image to optimize the visualization effect and ensure the accuracy and effectiveness of the image;
[0041] Merge the overlapping regions existing between different masks.
[0042] Furthermore, the merging of the overlapping regions existing between different masks is specifically as follows:
[0043] Define the expansion threshold of adjacent regions to determine the expansion range of the mask regions during synthesis, ensuring that adjacent regions can be merged and avoiding omission;
[0044] Perform connected component analysis on each color mask to extract the statistical information of the regions; based on the statistical information, expand the bounding boxes of each connected component and use the set expansion threshold to cover more adjacent regions;
[0045] Check the overlapping parts of each connected component with other color masks, judge whether the overlapping parts belong to adjacent regions by calculating the intersection, and at the same time apply the minimum pixel threshold to exclude too small noise regions; for adjacent regions that meet the merging conditions, merge them into one region and replace the color label of the adjacent region with the label of the current region;
[0046] Update the synthesized mask image to ensure that the image only contains important and coherent regions, removing noise and small irrelevant regions, thereby generating a concise and semantically consistent final mask image.
[0047] The present invention also provides a dynamic decision-making image segmentation device based on self-prompt guidance, including one or more processors for implementing the dynamic decision-making image segmentation method based on self-prompt guidance described above.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention proposes a fusion and dynamic decision-making method that combines the SAM model guided by a task self-prompter and the SAM model guided by a spatial position information-driven self-prompter. This method can dynamically adjust the inference strategy according to the specific image content and task requirements, thereby achieving more accurate and efficient automatic image segmentation and significantly improving the instance segmentation accuracy.
[0049] (2) Aiming at the deficiency in the prior art that prior guidance information (such as points, boxes, or rough masks) needs to be provided manually, the present invention introduces a heuristic prompt learning mechanism, enabling the model to automatically generate adaptive prompts and complete the precise segmentation of the target without manual intervention. This innovation greatly improves the segmentation efficiency, reduces the labor cost, and enhances the practicality of the method.
[0050] (3) The present invention inherits the zero-shot generalization ability of the SAM model and further introduces semantic category information on this basis. During the process of automatic image segmentation, the addition of semantic information makes the segmentation result not only more accurate in morphology but also provide stronger discrimination ability at the semantic level. This innovation improves the overall accuracy of the image segmentation result. Especially in specific fields (such as land cover classification in remote sensing images or organ recognition in medical images), it can provide more accurate and reliable support, thus broadening the application scenarios and practical value of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of the present invention;
[0052] Figure 2 is a schematic diagram of the SAM model guided by the task self-prompter of the present invention;
[0053] Figure 3 is a schematic diagram of the SAM model guided by the spatial position information-driven self-prompter of the present invention;
[0054] Figure 4 is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.
[0056] As Figure 1 shown, the embodiment of the present invention provides a dynamic decision-making image segmentation method based on self-prompt guidance, including the following steps:
[0057] Step 1, Dataset acquisition and annotation
[0058] Acquire the original image data from a specific field (such as remote sensing images, medical images, etc.), select about 1000 sample images to ensure that these images can cover the diversity of the target objects. For each image containing the target object, use an artificial segmentation annotation tool to generate a segmentation mask for the target object, and at the same time generate the bounding box and class label of the target.
[0059] Step 2, Model construction and design
[0060] 2.1, Hypothesis Denote the training data contained in the training dataset. Among them, denotes the th training data. Among them, is 3D volume data of the image, denotes the height of the volume data, denotes the width of the volume data, denotes the number of layers of the volume data; is the ground truth annotation of the corresponding training data, including the coordinates of object bounding boxes ([[]] ), its attached semantic class ([[]] ) and binary mask ([[]] ). The main goal is to train two self-prompt-guided SAM models so that they can process any image from the test set and infer its semantic class and instance mask.
[0061] 2.2, Construct a task self-prompt-guided SAM model: As Figure 2 shown, this model includes five main parts, namely an image encoder module, a feature aggregator module, a feature separator module, a task-based prompt module, and a SAM mask decoder module.
[0062] The image encoder module uses a pre-trained SAM vision encoder to extract image features. First, the input image is processed by the patch embedding module, which uses convolutional operations to divide the image into small patches and maps each patch to a high-dimensional feature space of 1280 dimensions, thereby extracting rich local information. Next, the image features are processed through 32 layers of Transformer encoder layers, each layer containing a multi-head self-attention layer with 32 heads, two layers of feed-forward neural networks, layer normalization, and residual connections, thereby capturing global context information and enhancing the expressive power of the feature representation.
[0063] The feature aggregator module fuses and processes the input features through a series of convolutional and normalization operations, aiming to extract and integrate information at different scales. First, the input 256-channel features are extended to 512 channels through a 1x1 convolution, and then layer normalization is performed to stabilize the training process. Next, the features are further processed using a 3x3 convolution, maintaining the number of channels at 512, and layer normalization is performed again. Finally, another 3x3 convolution reduces the feature dimension to 256 channels, and layer normalization is performed again to ensure the stability and consistency of the features. This series of operations enhances the feature expression ability through multiple convolutions and normalizations, helping the model to effectively utilize the fused multi-scale information in subsequent tasks.
[0064] Feature Separator Module. Through a series of refined upsampling, downsampling, convolution, and lateral connection operations, it aims to integrate features at different scales to enhance the model's multi-scale perception ability. First, the module uses two layers of upsampling operations to process low-resolution feature maps. Specifically, in the first layer, a transposed convolution is used to upsample a 256-channel feature map to 128 channels, then the Gaussian Error Linear Unit (GELU) activation function is applied, and then it is further upsampled to 64 channels through a transposed convolution; in the second layer, a transposed convolution is used to upsample the 256-channel feature map to 128 channels. At the same time, for high-resolution feature maps, the module uses max-pooling operations for downsampling to reduce its size. These operations create a hierarchical structure for feature maps at different scales, enabling features at each scale to be effectively processed and utilized. To achieve the unification of features at different scales, the module further performs channel number mapping between features at different scales through lateral connection operations. Specifically, the module uses 4 1x1 convolutions to adjust the channel number of the feature map to a unified 256 channels: first, the 64-channel feature map is mapped to 256 channels through convolution; then, the 128-channel feature map is mapped to 256 channels through convolution; for the 256-channel feature, the channel number remains unchanged through a 1x1 convolution. Next, the module uses 4 3x3 convolutional layers to further process all the features obtained through lateral connection, ensuring that the fused features are consistent in channel number and enhancing their stability and representational ability through layer normalization. The operation of each convolutional layer unifies the channel number to 256, further improving the representational ability and quality of the feature map. Specifically, during the processing of the feature map by the four 3x3 convolutional layers, the consistency of the channel number is ensured, and the feature information at different scales is smoothly fused and optimized. Through these upsampling, downsampling, convolution, and lateral connection operations, features from different scales are effectively separated and fused, providing a more stable and high-quality feature representation with multi-scale information for subsequent tasks, thereby enhancing the model's perception and understanding ability of complex image structures.
[0065] The Task-based Prompt Module combines the encoder and decoder structures of the Transformer, aiming to accurately generate multi-semantic segmentation masks through a task-related query mechanism. First, the input feature maps from four different scales of the Feature Separator are concatenated after positional encoding and hierarchical encoding, then flattened and input into a Transformer encoder composed of 3 layers of attention layers and feed-forward neural networks for self-attention operations to extract global semantic information. Next, the feature map output by the encoder is split into a set of feature maps at 4 scales , , , , and these feature maps are consistent in spatial size with the input feature maps of the original Feature Separator, where is the feature map of the largest size, used to generate a rough mask The transformer decoder then interacts with the learnable task-related query tokens initialized to zero and interacts with the 4-scale feature maps from the encoder , , , and fuses them. The decoder generates the mask filter for each instance through the mask projection layer , generates the prompt embedding for each instance through the prompt projection layer , and generates the semantic category for each instance through the category derivation layer . These operations are formulated as: , , . Among them , , , , , where is a transformer decoder composed of cross-attention and a feed-forward neural network and are linear projection layers is a prompt projection layer composed of two-layer multi-layer perceptrons represents the number of prompt groups represents the embedding number of each prompt
[0066] The SAM mask decoder module generates a high-precision image segmentation mask through image embedding, mask tokens, two dual-path transformer modules, 4 multi-layer perceptron modules, and two upsampling transposed convolutional layers. First, the high-dimensional image features from the image encoder module are processed by the image embedding module, where the dimension of the features is reduced from 1280 to 256 through pointwise convolution and standard convolution, and combined with layer normalization operations to improve the stability and usability of the features. Next, the mask token part is the rough mask generated by the task-based prompt module Start, and further generate the prompt embedding features of the mask through the prompt embedding layer. These prompt embedding features, together with the prompt features from the prompt projection layer, are processed by the dual-path transformer module. In the dual-path transformer module, the self-attention layer and the cross-modal attention layer enable the mask tokens to deeply interact with the image features, thereby capturing rich spatial and semantic information in the image. This module includes a self-attention layer, a cross-modal token-image cross-attention layer, a multi-modal perceptron layer, and two cross-modal image-token cross-attention layers. These layers extract and transform the image features layer by layer, promoting information fusion and representation. After these attention layers, the multi-layer perceptron module further transforms the features to generate the mask filter. To address the resolution issue in image segmentation, the decoder also includes two upsampling convolutional layers, which gradually restore the low-resolution feature maps to higher-resolution fine-grained feature maps, thereby ensuring the fineness and accuracy of the segmentation mask. In addition, the SAM mask decoder uses position encoding and class embedding layers to provide spatial position information and semantic class guidance, further enhancing the segmentation accuracy. Finally, after these multi-level processes, the SAM mask decoder can generate accurate and detailed image segmentation masks.
[0067] 2.3. Construct a SAM model guided by a self-prompter driven by spatial position information: As Figure 3 shown, this model contains five core modules, namely an image encoder module, a feature aggregator module, a feature separator module, a prompter module driven by spatial position information, and a mask decoder module.
[0068] Among them, the image encoder module, the feature aggregator module, the feature separator module, and the mask decoder module correspond to the task self-prompter-guided SAM model constructed in step 2.2.
[0069] The prompter module driven by spatial position information combines a lightweight region proposal network and region of interest pooling, aiming to accurately generate multi-semantic segmentation masks through the accurately generated candidate boxes and the features related to the spatial position information extracted from them. The workflow of this module can be divided into multiple stages and proceeds in the following order: First, the input feature maps of four different scales from the feature separator are input into the lightweight proposal network, which consists of 1 convolutional layer, 1 classification layer, and 1 regression layer, and is used to generate the class scores and bounding box regression values for each anchor point. By combining the prior boxes, the network transforms the classification and regression features extracted from the multi-scale input feature maps into the final bounding boxes, and removes redundant high-overlap boxes through non-maximum suppression. Finally, 1000 region proposal boxes are obtained, denoted as , . Secondly, in order to enhance the model's ability to perceive the spatial position of objects, the four input feature maps of different scales from the feature separator are applied with sinusoidal position encoding (PE). Each feature map is added to the corresponding sinusoidal position encoding to obtain a multi-scale feature map , This enhances the location features and helps improve the spatial positioning capability of the model. Next, the extracted region proposal box Then, the target region feature extraction module consisting of 4 different scale region alignment layers is used to extract the features corresponding to each region of interest from the 4 different scale multi-layer feature maps. The operation formula is , where the extracted features The size is Then each The feature map is flattened into a one-dimensional vector and extracted through a shared feature extractor consisting of two fully connected layers to obtain the shared feature Then, it passes through a fully connected layer of category information derivation layer. , generating the semantic category of each instance , and is derived from the boundary information of another fully connected layer Generate bounding boxes for each instance The formulas are ))), ))). Next, based on the features extracted by the target region feature extraction module, non-maximum suppression is applied to each type of target frame to remove overlapping low-scoring frames and obtain the corresponding features. These features are then fed into another region of interest feature extraction module, which further extracts features corresponding to each region of interest from feature maps of four different scales. , The operation formula is , extracted The size of the feature map is .Then It is fed into a spatial information hint layer consisting of 1 convolutional layer and 3 fully connected layers to generate sparse target features. These target features are further passed to the mask decoder for processing. Through the above steps, the module can accurately generate multiple semantic segmentation masks and effectively handle complex target recognition and segmentation tasks.
[0070] Step 3: Model training and fine-tuning
[0071] The present invention trains and fine-tunes two different SAM models, namely, the SAM model guided by the task self-prompter and the SAM model guided by the spatial location information-driven self-prompter. The specific steps of the training and fine-tuning process are as follows:
[0072] 3.1. Preprocess the manually annotated training set and validation set: First, load the images from the file and convert them to the floating-point format. Then, the images are normalized, using fixed mean and standard deviation for normalization to ensure data consistency. Subsequently, load the annotation information of the images, including the target bounding boxes and segmentation masks of each image, to ensure that each image corresponds to accurate instance segmentation data. To improve the generalization ability of the model and enhance data diversity, the present invention adopts a variety of data augmentation techniques. These augmentation operations include image rotation, horizontal flipping, vertical flipping, scale scaling, translation, random cropping, sharpening, smoothing, as well as adjusting the grayscale, contrast, and brightness of the images. In addition, the present invention also introduces techniques such as random occlusion and image denoising. These augmentation operations ensure that the images and the corresponding mask information remain aligned during the transformation, thus ensuring the accuracy of the masks. During the preprocessing process, the targets are also screened to filter out too small targets to ensure the effectiveness and representativeness of the targets in the training set and avoid unnecessary noise interfering with the training process. Finally, all the images and their annotation information are packed into the input format required by the model, and padding operations are used to ensure that the images and masks in the batch have a unified size, thus ensuring data consistency during the training process.
[0073] 3.2. Data input and model initialization: First, use the preprocessed image data as input and input it into the SAM model guided by the task self-prompter and the SAM model guided by the spatial location information-driven self-prompter respectively. During the fine-tuning process, the present invention freezes the parameters of the image encoder to ensure that they are not updated during the training process. To retain the features that the image encoder has learned and avoid losing effective image features during the fine-tuning process. The focus of fine-tuning is to update the parameters of the following modules: the feature aggregator, the feature separator, the prompt encoder, and the mask decoder.
[0074] 3.3. Optimizer settings and training parameters: The AdamW optimizer is used during the fine-tuning process, which has good convergence performance and stability. During the training, the present invention sets the following key parameters: weight decay, initial learning rate, cosine annealing scheduler and linear warm-up strategy, batch size, and number of iterative updates.
[0075] 3.4. Calculation of the total loss function: During the fine-tuning process, the present invention calculates the total loss function to guide model optimization and parameter update. The total loss functions of different models are slightly different. For the SAM model guided by the task self-prompter, its total loss function consists of a classification loss and a mask loss. The classification loss uses cross-entropy to calculate the difference between the predicted category and the true category, and the mask loss calculates the difference between the predicted mask and the true mask through binary cross-entropy, specifically including a rough mask loss and a fine-grained mask loss. The form of this loss function is:
[0076] ;
[0077] where is the classification loss, is the mask loss.
[0078] For the SAM model guided by the spatial location information-driven self-prompter, the loss function includes a region proposal loss, a classification loss, a regression loss, and a mask loss. The region proposal loss evaluates the performance of the region proposal network. The classification loss uses cross-entropy to calculate the difference between the predicted category and the true category. The regression loss calculates the smooth absolute error loss based on the difference between the predicted coordinate offset and the target offset. The mask loss calculates the difference between the predicted mask and the true instance mask through binary cross-entropy. The form of this loss function is:
[0079] ;
[0080] where is the region proposal loss, is the classification loss, is the regression loss, is the mask loss.
[0081] 3.5. Matching process and supervised learning: When training the SAM model guided by the task self-prompter, first, the predicted mask is matched with the true instance mask. This matching process uses the Hungarian algorithm to find the best matching relationship to ensure that each predicted instance is paired with the most suitable true instance. The matching loss consists of three parts: the class matching loss, which is used to calculate the difference between the predicted class and the true class; the mask cross-entropy loss, which is used to calculate the cross-entropy difference between the predicted mask and the true mask; and the mask Dice loss, which evaluates the similarity between the predicted mask and the true mask through the Dice coefficient. After the matching is completed, it enters the supervised learning stage. In this stage, the supervision items during the training process include the multi-class classification loss and the mask classification loss. The multi-class classification loss is used to calculate the classification error of different class targets, while the mask classification loss is used to calculate the error between the predicted mask and the matched true instance mask, thereby further optimizing the performance of the model.
[0082] 3.6. Hyperparameter Tuning and Model Selection: During the model fine-tuning process, the present invention uses the training set for parameter optimization and learning, enhances the generalization ability of the model through data augmentation, while the validation set is used to regularly evaluate the performance of the model to ensure that it does not overfit and can perform well on unseen data. The hyperparameter tuning process is based on the mean average precision (mAP) on the validation set. By monitoring the change of mAP, hyperparameters such as the learning rate, batch size, and number of training epochs are optimized to select the best model configuration. Finally, based on the comprehensive evaluation of the training set and the validation set, the present invention selects the model with the best performance on the validation set to ensure its robustness and accuracy in practical applications.
[0083] Step 4. Model Fusion and Dynamic Decision Making
[0084] In practical applications, factors such as the scale, position, and shape of the target often change. Therefore, it is crucial to have the ability to adaptively adjust according to these changes. For this purpose, the present invention proposes a fusion method that combines the SAM model guided by the task self-prompter and the SAM model guided by the self-prompter driven by spatial position information. Through model fusion and dynamic decision making, this method can dynamically adjust the inference strategy according to the specific image content and task requirements, thereby significantly improving the accuracy of instance segmentation. The specific process of this process is as follows:
[0085] 4.1. Preprocessing of Samples to be Annotated: The image to be annotated is first loaded from a file and undergoes a series of standardization processes. Specifically, the image is first normalized using the same mean and standard deviation as the training data to ensure that the distribution of the image data is consistent with the training data. Then, the image is scaled proportionally to the target size while keeping the aspect ratio unchanged to prevent image distortion or deformation. If the size of the image does not match the target input requirements, necessary padding will be performed to ensure that the image meets the size requirements of the model input. The preprocessed image and related metadata (such as image ID, path, and original size, etc.) will be packaged into the model input format for subsequent inference.
[0086] 4.2. Model Inference and Fusion: The preprocessed image will be fed into the fused model for inference. This fused model consists of two guiding mechanisms: the SAM model guided by the task self-prompter and the SAM model guided by the self-prompter driven by spatial position information. The SAM model guided by the task self-prompter can adaptively select the most suitable feature extraction and inference methods according to the task requirements to ensure that the task characteristics are precisely processed; while the SAM model driven by spatial position information improves the model's ability to model the spatial features of the target by analyzing the distribution and layout of the target in space. By combining these two guiding mechanisms, the model can generate a set of instance segmentation masks representing the regions of each instance in the image. Each mask will be stored together with its corresponding label, score, and other information as part of the inference result.
[0087] 4.3. Dynamic Selection of the Optimal Mask: After generating multiple instance masks, the model will process each mask to dynamically select the optimal mask. First, the model calculates the area of each mask and filters out the valid masks with larger areas (e.g., masks with areas smaller than the set threshold will be discarded). Then, if there are overlapping regions among multiple masks, the model will dynamically decide which masks to retain based on the area of the overlapping regions and the score values of each mask. The specific operations include: for overlapping regions with an overlapping area smaller than the set threshold, select the mask with a larger area; for larger overlapping regions, preferentially select the mask with a higher score. Finally, according to this rule, the model will store the most suitable mask, label, and score in the final result array for subsequent analysis and processing.
[0088] Step 5. Optimization and Synthesis of Multi-Mask Regions
[0089] This module realizes the precision and semantic consistency of the mask image, and improves the quality and visualization effect of the segmentation result through the synthesis of multi-mask regions, noise removal, region screening, and similarity-based optimization. The specific steps are as follows:
[0090] 5.1. Synthesis of Multi-Mask Regions and Presentation of Semantic Information: During the optimization process of multi-mask regions, multiple masks and instance labels are first extracted from the model fusion and dynamic decision-making steps. According to each label value, the corresponding color is obtained from the color palette and applied to the corresponding mask region. By assigning different colors to each binary mask, multiple masks are drawn on the same image to generate an instance mask map containing semantic information. The semantic categories in this map are represented by the corresponding colors, ensuring that each region can clearly identify its semantic category.
[0091] 5.2. Removal of Small-Area Noise and Screening of Mask Regions: For the synthesized mask image, first, the connected regions of each color mask are labeled to identify different regions in the image. Then, an area threshold is set, and the connected regions smaller than this threshold are filtered out to remove the noise regions. This process ensures that only the larger valid regions are retained in the image, avoiding the influence of small regions on the clarity of the final result. Finally, the screened mask regions are merged into the final image to optimize the visualization effect and ensure the accuracy and effectiveness of the image.
[0092] 5.3 Mask Synthesis and Optimization Based on Region Similarity and Minimum Pixel Threshold: During the superposition of multiple masks, there may be overlapping regions between different masks. To reduce color label errors, remove noise, and enhance region coherence, it is necessary to merge the overlapping regions. First, define the expansion threshold for adjacent regions to determine the expansion range of the mask region during synthesis, ensuring that adjacent regions can be merged to avoid omission. Then, perform connected component analysis on each color mask to extract statistical information of the regions, such as area and bounding box. Based on this information, expand the bounding box of each connected component and use the set expansion threshold to cover more adjacent regions. Next, check the overlapping parts of each connected component with other color masks. Determine whether they belong to adjacent regions by calculating the intersection, and at the same time apply the minimum pixel threshold to exclude too small noise regions. For adjacent regions that meet the merging conditions, merge them into one region and replace the color label of the adjacent region with the label of the current region. To improve processing efficiency, use multi-threaded parallel processing for multiple color masks to quickly complete mask synthesis and optimization. Finally, update the synthesized mask image to ensure that the image only contains important and coherent regions, remove noise and irrelevant small regions, thereby generating a concise and semantically consistent final mask image.
[0093] Through the above steps, the finally obtained image is a high-quality semantic-level instance segmentation mask image generated by the automated segmentation method.
[0094] Corresponding to the foregoing embodiment of a dynamic decision-making image segmentation method based on self-prompt guidance, the present invention also provides an embodiment of a dynamic decision-making image segmentation device based on self-prompt guidance.
[0095] As Figure 4 shown, an embodiment of a dynamic decision-making image segmentation device based on self-prompt guidance provided by an embodiment of the present invention includes one or more processors for implementing a dynamic decision-making image segmentation method based on self-prompt guidance in the foregoing embodiment.
[0096] The embodiment of the dynamic decision-making image segmentation device based on self-prompt guidance of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the dynamic decision-making image segmentation device based on self-prompt guidance of the present invention is located. Except for Figure 4In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located may generally include other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0097] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.
[0098] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
Claims
1. A dynamic decision image segmentation method based on self-prompt guidance, characterized in that: The steps include: Obtaining original images of the target area as samples, taking out some samples for manual annotation and making corresponding labels to construct a training data set, wherein the training data set includes a training set and a validation set; Constructing a SAM model guided by a task self-prompter and a SAM model guided by a spatial position information driven self-prompter according to the training data set; The SAM model guided by the task self-prompter includes: an image encoder module, a feature aggregator module, a feature separator module, a task-based prompter module and a mask decoder module; The image encoder module uses a pre-trained SAM visual encoder to extract image features; The feature aggregator module fuses and processes the input features through convolution and normalization operations to extract and integrate information of different scales; The feature separator module is used to integrate features of different scales through upsampling, downsampling, convolution and lateral connection operations to enhance the multi-scale perception ability of the model; The task-based prompter module includes: a transformer encoder for performing self-attention operations on four different scales of input feature maps from the feature separator module, extracting global semantic information and generating feature maps of four scales; a transformer decoder for interacting with zero-initialized task-related query tags and fusing with feature maps of four scales; and a mask filter, prompt embedding and semantic category for each instance are generated through a mask projection layer, a prompt projection layer and a category inference layer, and a coarse mask is generated according to the maximum size feature map; The mask decoder module generates high-precision image segmentation masks through image embedding, mask labeling, 2 dual-path transformer modules, 4 multi-layer perceptron modules, and 2 up-sampled transposed convolutional layers; The SAM model driven by spatial position information and guided by a self-prompter includes: an image encoder module, a feature aggregator module, a feature separator module, a prompter module driven by spatial position information, and a mask decoder module; The image encoder module uses a pre-trained SAM visual encoder to extract image features; The feature aggregator module fuses and processes the input features through convolution and normalization operations to extract and integrate information of different scales; The feature separator module is used to integrate features of different scales through upsampling, downsampling, convolution and lateral connection operations to enhance the multi-scale perception ability of the model; The prompter module driven by spatial position information includes: a lightweight region proposal network, which is used to generate category scores and bounding box regression values from multi-scale input feature maps, remove redundant boxes through non-maximum suppression, and output multiple region proposal boxes; a position encoding module, which is used to apply sinusoidal position encoding to the input feature map to enhance its spatial position information; a region of interest alignment layer, which is used to extract features corresponding to each proposal box from the multi-scale feature map; a shared feature extractor, which is used to process the extracted features and generate shared features; a category inference layer and a boundary inference layer, which are used to generate semantic categories and bounding boxes for each instance, respectively; a region of interest feature extraction module, which is used to extract target features; a spatial information prompting layer, which is used to generate sparse target features and pass them to a mask decoder to generate a final multi-semantic segmentation mask; The mask decoder module generates high-precision image segmentation masks through image embedding, mask labeling, 2 two-way transformer modules, 4 multi-layer perceptron modules, and 2 up-sampled transposed convolutional layers; Train and fine-tune two SAM models; After fusing the two fine-tuned SAM models, the fusion model dynamically adjusts the reasoning strategy according to the specific image content and task requirements of the image to be labeled, thereby generating multiple mask regions; Through the synthesis of multiple mask regions, noise removal, region screening and similarity-based optimization, the mask image is refined and semantically consistent, and finally a semantic-level instance segmentation mask image is obtained.
2. The method for dynamic decision-making image segmentation based on self-prompt guidance according to claim 1, characterized in that: The construction of the training data set includes: ensuring that the extracted samples cover the diversity of the target objects; for each original image, using a manual segmentation and annotation tool to generate a segmentation mask of the target object, and annotating the bounding box and category label of the target; constructing a training data set for training and verification based on the original image and the generated segmentation mask, bounding box and category label, and randomly dividing it into a training set and a verification set in a ratio of 9:
1.
3. The method for dynamic decision-making image segmentation based on self-prompt guidance according to claim 1, characterized in that: The training and fine-tuning of the two SAM models includes: Preprocessing the training data set, including image loading, standardization, data enhancement, target screening, and image size unification to ensure the consistency and validity of image and mask information; The preprocessed data is fed as input to the SAM model guided by the task self-cue and the SAM model guided by the spatial position information self-cue, freezing the parameters of the image encoder and fine-tuning the parameters of the feature aggregator, feature separator, cue encoder and mask decoder modules. The AdamW optimizer is used to fine-tune the training parameters, including weight decay, initial learning rate, cosine annealing scheduler, linear warm-up strategy, batch size, and number of iterations. The total loss function is calculated for the two SAM models respectively. For the SAM model guided by the task self-prompter, the Hungarian algorithm is used to match the predicted mask with the real instance mask, and the category matching loss, mask cross entropy loss and mask Dice loss are calculated to ensure that each predicted instance is paired with the most appropriate real instance, and supervised learning of multi-category classification loss and mask classification loss is performed; for the SAM model guided by the self-prompter driven by spatial location information, the total loss includes region proposal loss, classification loss, regression loss and mask loss; hyperparameter optimization is performed through the training set, model performance is evaluated using the validation set, the best model configuration is selected based on the average accuracy, and finally the model with the best performance on the validation set is selected.
4. The method for dynamic decision image segmentation based on self-prompt guidance according to claim 1, characterized in that: The two fine-tuned SAM models are fused, and the fused model dynamically adjusts the reasoning strategy according to the specific image content and task requirements of the image to be labeled, including: Preprocess the image to be annotated, including loading the image, normalizing it, scaling it to the target size, and performing necessary padding to ensure that the image meets the model input requirements; The preprocessed image is sent to the fused model for reasoning; the fused model includes two guiding mechanisms: the SAM model guided by the task self-prompter is used to adaptively select the most appropriate feature extraction and reasoning method according to the task requirements to ensure that the task characteristics are accurately processed; the SAM model driven by spatial position information is used to improve the model's modeling ability of the target spatial characteristics by analyzing the distribution and layout of the target in space; a set of instance segmentation masks are generated by the fused model to represent the area of each instance in the image; each mask is stored together with its corresponding label and score information as part of the reasoning result; After generating multiple instance segmentation masks, each mask is processed through the fused model to dynamically select the optimal mask.
5. The method for dynamic decision image segmentation based on self-prompt guidance according to claim 4, characterized in that: The processing of each mask by the fused model includes: The area of each mask is calculated through the fused model, and the effective masks with larger areas are screened out; If there are overlapping areas of multiple masks, the fused model dynamically decides which masks to keep based on the area of the overlapping area and the score of each mask. For overlapping areas whose overlapping area is smaller than the set threshold, the mask with a larger area is selected; for overlapping areas with a larger area, the mask with a higher score is selected first. The selected masks, labels, and scores are stored in the final result array.
6. The method for dynamic decision-making image segmentation based on self-prompt guidance according to claim 1, characterized in that: The method of achieving precision and semantic consistency of the mask image by synthesizing multiple mask regions, removing noise, selecting regions, and optimizing based on similarity includes: Extract multiple masks and category labels from the model fusion process; according to each category label, obtain a corresponding color from the color palette, and apply the color to the corresponding mask area; by assigning different colors to each mask, draw multiple masks on the same image to generate an instance mask map containing semantic information, and the semantic category of the instance mask map is represented by the corresponding color to ensure that each area can clearly identify its semantic category; For the synthesized mask image, mark the connected areas of each color mask to identify different areas in the image; set an area threshold and filter out connected areas smaller than the threshold to remove noise areas; merge the filtered mask areas into the final image to ensure the accuracy and effectiveness of the image; Merge the overlapping areas between different masks.
7. The method for dynamic decision image segmentation based on self-prompt guidance according to claim 6, characterized in that: The merging of overlapping areas between different masks is specifically as follows: Define the expansion threshold of the neighboring area to determine the expansion range of the mask area during synthesis; Performing a connected region analysis on each color mask to extract statistical information of the region; based on the statistical information, expanding the bounding box of each connected region to cover more adjacent regions using a set expansion threshold; Check the overlap between each connected region and other color masks, and determine whether the overlap belongs to an adjacent region by calculating the intersection. At the same time, apply the minimum pixel threshold to exclude the noise region that is too small. For adjacent regions that meet the merging conditions, merge them into one region and replace the color label of the adjacent region with the label of the current region. The synthesized mask image is updated to ensure that the image only contains important and coherent areas, and removes noise and irrelevant small areas to generate the final mask image.
8. A dynamic decision image segmentation device based on self-prompt guidance, characterized in that: It comprises one or more processors for implementing a dynamic decision image segmentation method based on self-prompt guidance as described in any one of claims 1-7.
Citation Information
Patent Citations
High-precision three-dimensional medical image semantic segmentation system and method based on improved SAM model
CN118864855A
Unsupervised zero-shot segmentation mask generation and semantic labeling
US20250045930A1