Click-based interactive image segmentation system and method
By employing a collaborative architecture of a global prompting module and a local correction module, the problem of insufficient utilization of interactive graph information in existing technologies is solved, achieving efficient interactive image segmentation, especially the ability to generate high-quality segmentation results in complex scenes.
Patent Information
- Application Number
- CN202510970351.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing click-based interactive image segmentation methods fail to fully utilize the global guidance information of the interactive graph and fail to effectively use the previous click results as prior information, resulting in inconsistent segmentation results and insufficient accuracy.
A collaborative architecture of Global Cueing Module (GHM) and Local Correction Module (LCM) is adopted. The GHM deeply encodes user click information into the Transformer backbone network, and the LCM refines the initial segmentation mask, thereby realizing cross-modal fusion and local optimization of interactive signals and visual features.
It significantly improves the boundary awareness capability of interactive segmentation, and can generate high-quality segmentation results with a small number of interactive clicks. In particular, it maintains efficient interaction in complex scenes while having pixel-level precision control capability.
Smart Images

Figure CN120821418A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation based on computer vision, and in particular relates to a click-based interactive image segmentation system and method. Background Art
[0002] Current deep learning still requires a large amount of data. The problem of data annotation is becoming increasingly serious, especially for pixel-level image annotation, where manual annotation consumes a significant amount of labor and time. Consequently, interactive image segmentation has attracted widespread attention. The goal of interactive image segmentation is to obtain high-quality segmentation masks with minimal user interaction. Interactive image segmentation encompasses a variety of interactive methods, such as click, scribble, bounding box, and their combination. Click-based methods are favored by most researchers due to their simplicity and ease of pointing to objects of interest. Specifically, within an image, the user indicates their interest by clicking on a target object once, and the model guides the user to obtain an initial segmentation mask. Subsequently, the user continues clicking based on the previous round of segmentation results, applying positive clicks to areas they wish to identify as foreground and negative clicks to areas they wish to identify as background, to obtain the next segmentation mask. This process iterates until a satisfactory segmentation mask is obtained. Notably, starting from the second round of interaction, the segmentation results from the previous round are incorporated into the current process as feedback. This feedback is crucial to the model, providing prior information that helps promote convergence and improve segmentation accuracy.
[0003] Recent research on click-based interactive segmentation has focused on obtaining detailed object masks. For example, SimpleClick employs a standard Visual Transformer (ViT) backbone network to capture global information for accurate segmentation. These methods typically generate a single mask based on user clicks, hence the term single-mask approach. SimpleClick simply appends the interaction map to the original image before feeding it into the network, failing to fully utilize the interaction map's guiding role. Furthermore, in subsequent interactions, SimpleClick treats the current click and previous clicks equally, limiting its ability to more effectively focus on the current click. Failure to use previous segmentation results as prior information leads to inconsistent segmentation results. In another case, while previous segmentation results are utilized, they are simply appended to the original image before feeding it into the network. While this approach incorporates prior information, the amount of information is relatively small compared to the original image. Therefore, this simple addition approach does not fully utilize the prior information provided by the interaction map. Summary of the Invention
[0004] To address the above problems, the present invention provides a click-based interactive image segmentation system in a first aspect, comprising a global prompt module, a self-attention feature extraction module, a feature pyramid sampling module, a nonlinear segmentation head, and a local correction module; The global hint module introduces a patch embedding layer symmetrical to the patch embedding layer in the backbone. It encodes user click information into a disk map and concatenates it with the segmentation mask from the previous round to form a three-channel interaction map. This map is then embedded into the feature extraction network, enabling the network to learn the object boundaries and region information indicated by the user click. The self-attention feature extraction module uses ViT as the backbone network and a single-scale feature map structure to avoid information loss caused by feature map downsampling, learns local detail information of the image, and uses Transformer to learn global context information of the image; The feature pyramid sampling part uses the last feature map to construct a multi-scale feature map based on the single-scale feature pyramid, so that the network can capture the features of objects at different scales; The nonlinear segmentation head module uses the MLP layer to perform feature transformation and fusion, converting the multi-scale feature map into a segmentation probability map, avoiding complex convolution operations; The local correction module decomposes a heavy inference on the entire image into two light predictions on small blocks, and uses historical information to focus on the refined processing of the local area of the new click area.
[0005] Preferably, the global prompt module is specifically: A patch embedding layer symmetrical to the one in the backbone is introduced, followed by element-wise feature addition; user clicks are encoded in a two-channel disk map, one for positive clicks and the other for negative clicks, with positive clicks placed on the foreground and negative clicks on the background; the segmentation mask and the two-channel click map in the visual deformer are concatenated into a three-channel map for patch embedding, and the two symmetrical embedding layers operate on the image and the concatenated three-channel map respectively. The input is patched, flattened and projected to two vector sequences of the same dimension, and then element-wise addition is performed before being input into the self-attention feature extraction module.
[0006] Preferably, the self-attention feature extraction module is specifically: The self-attention feature extraction module uses the traditional ViT as the backbone network, which maintains only a single-scale feature map throughout. The patch embedding layer divides the input image into non-overlapping fixed-size patches (e.g., 16×16 for ViT-B). Each patch is flattened and linearly projected into a fixed-length vector (e.g., 768 for ViT-B). The resulting vector sequence is fed into a queue of Transformer blocks (e.g., 12 for ViT-B) for self-attention. The backbone is pre-trained as MAEs on ImageNet-1k. During fine-tuning, the pre-trained backbone is adapted to a higher-resolution input using non-shifting windowed attention, supplemented by a number of global self-attention blocks (e.g., 2 for ViT-B). Furthermore, by incorporating global cues, user interactions are encoded as a two-channel disk map, with one channel assigned to positive clicks, representing foreground, and the other to negative clicks, representing background. The pre-existing segmentation mask and the two-channel click map are merged to create a three-channel interaction map. GHM processes this map by splitting it into non-overlapping blocks of equal size (16×16 pixels). Each block is then flattened and linearly projected into a fixed-length vector space with 768 dimensions. This projection is performed in triplicate, and the resulting set of three vectors is superimposed with the outputs of the first, sixth, and twelfth conversion layers to enhance the self-attention mechanism and address the problem of dilution of interaction graph information.
[0007] Preferably, the feature pyramid sampling uses the last feature map output by the self-attention feature extraction module, and uses four convolutional layers with different step sizes to downsample the last feature map to resolutions of 1 / 32, 1 / 16, 1 / 8, and 1 / 4, respectively, to generate a multi-scale feature map, so that the network can better capture the features of objects at different scales.
[0008] Preferably, the nonlinear segmentation head implements a lightweight segmentation head using an MLP layer. It uses a simple feature pyramid to generate a segmentation probability map at a scale of 1 / 4, which is then upsampled to restore the original resolution. Note that this segmentation head avoids computationally demanding components and accounts for only 1% of the model parameters. Crucially, using a powerful pre-trained backbone, the lightweight segmentation head is sufficient for interactive segmentation. The proposed full MLP segmentation head works in three steps. First, each feature map from the simple feature pyramid is passed through an MLP layer to convert it to the same channel dimension. Second, all feature maps are upsampled to the same resolution for concatenation. Third, the concatenated features are fused with another MLP layer to produce a single-channel feature map, which is then fused using a sigmoid function to obtain a segmentation probability map, which is then converted to a binary segmentation given a predefined threshold (i.e., 0.5).
[0009] Preferably, the local correction module decomposes a heavy inference on the entire image into two light predictions for small patches; first, an image patch around the target object is selected, resized to a small scale, and sent to Backbone to predict a coarse mask; then, a local area is selected around the click, and the local patch that needs to be refined and enlarged is fed into the local correction module; finally, progressive merging aligns the local prediction back to the full-size mask, refining only a small local area after each click, and all pixels of the final prediction are refined through calculations assigned to different rounds.
[0010] Preferably, the target crop is used to filter out background information irrelevant to the target object; first, the minimum external box of the previous mask and the newly added click is calculated, and then the ratio r is used to calculate the minimum external box of the previous mask and the newly added click. TC = 1.4 to expand it, then crop the input tensor and resize it; The focus cropping (Focus Crop) is used to locate the area that the user intends to modify; first, the difference between the original segmentation result and the previous mask is compared to obtain the differential mask Mxor, and then the maximum connected area containing the newly clicked Mxor is calculated, and the external box is generated for this maximum connected area in r FC = 1.4 ratio; crop local patches on the input image and click map; The refinement (Refiner) is used to recover the details of the coarse prediction in the focus cropping; first, a depth-wise separable convolution is used to extract low-level features from the cropping tensor; at the same time, the number of channels of the regional features is adjusted and fused with the extracted low-level features, and the detail map M is predicted using two heads. d and boundary graph M b , and update the rough prediction logarithm M l The boundary area is used to calculate the refined prediction Mr.
[0011] A second aspect of the present invention provides a click-based interactive image segmentation method, using the interactive image segmentation system as described in the first aspect, and comprising the following steps: First, the user clicks are determined. The user clicks on the image to instruct the model to focus on the object. The first click determines the center of the segmentation target, and subsequent clicks are in the center of the area with the largest error. The clicks are then encoded using disk encoding to form an interaction graph, represented as a tensor; The interaction graph and the original image are fed into the self-attention feature extraction module and the self-attention is enhanced through the global hint module to extract the global information of the image and the interaction graph; The last feature map is then processed using a nonlinear segmentation head module and upsampled to generate a coarse segmentation mask; At the same time, the local correction module is used to identify the local areas that need to be optimized, and a lightweight network is used to generate local fine segmentation masks; The entire system is iteratively optimized. Through user clicks, a segmentation mask is formed. Click again, and it is iteratively optimized to gradually improve the accuracy of the segmentation mask.
[0012] Compared with the prior art, the present invention has the following beneficial effects: Based on the SimpleClick model, this paper proposes an innovative Glclick framework. Through the collaborative architecture of the Global Hint Module (GHM) and the Local Correction Module (LCM), it significantly improves the boundary perception ability of interactive segmentation and explores the deep integration mechanism of global semantic guidance and local feature correction. The technical architecture of this paper consists of two core components: the Global Hint Module deeply encodes user click information into the Transformer backbone network, realizing cross-modal fusion of interactive signals and visual features; and the Local Correction Module uses a multi-scale feature optimization mechanism to perform boundary refinement on the initial segmentation mask, enabling the model to maintain efficient interaction while possessing pixel-level precision control capabilities.
[0013] The Glclick model of the present invention demonstrates excellent performance in complex scene applications such as remote sensing image analysis and medical image processing. Even in the face of challenges such as blurred target edges, interwoven multi-scale structures, and mixed texture and background, it can still generate high-quality segmentation results with a small number of interactive clicks. The technological breakthrough is reflected in the following: on the basis of retaining the lightweight advantage of the classic interactive segmentation framework, through a global-local dual-path optimization mechanism, without significantly increasing the computational load, it not only maintains the cross-domain generalization ability of the model, but also realizes the multi-level feature fusion of interactive information - the GHM module strengthens the global representation of the target through the attention reweighting mechanism, and the LCM module uses a dynamic convolution kernel group to adaptively correct the boundary area, ultimately forming an accurate depiction of the complex target structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, what is described below is only one embodiment of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 This is the overall structural block diagram of the interactive image segmentation model of the present invention.
[0016] Figure 2 This is the structural diagram of the local correction module.
[0017] Figure 3This is the visualization experiment result of the embodiment of the present invention; Figure 4 Convergence analysis of the model trained on the SBD dataset in the embodiment of the present invention Figure 5 This is a convergence analysis of the model trained on the DAVIS dataset according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The invention will be further described below with reference to specific embodiments.
[0019] The goal of click-based interactive image segmentation is to obtain pixel-level segmentation masks with only a few manual clicks. This approach simplifies the pixel-level annotation and image editing process. Many studies have focused on this area. In particular, SimpleClick
[15] implemented a simple and straightforward design using visual transformers and demonstrated its effectiveness in interactive image segmentation. However, two issues remain: First, this simple design does not fully utilize the global guidance provided by the interactive map. Second, by treating all clicks equally, it fails to maximize the guiding role of new clicks in each iteration.
[0020] This paper proposes a network model GlClick that integrates global prompt information and corrects local details. Figure 1 This invention addresses the inherent limitations of click-based interaction image segmentation using traditional visual deformers. Initially, GlClick includes a Global Hints Module (GHM), which enhances the model's ability to more effectively utilize guidance cues from the interaction graph. Subsequently, after obtaining a segmentation mask, the system carefully evaluates the differences between the generated mask and previously generated masks and then performs targeted local optimization to minimize these differences.
[0021] like Figure 1 As shown in Figure 1, first, users can suggest targets of interest to the segmentation model by clicking. In the first interaction, the model obtains the center point of the segmentation target based on the ground truth, thereby guiding the model to perform the first segmentation; in subsequent interactions, the user compares the ground truth with the previous round of segmentation results and clicks at the center of the area with the largest error. Secondly, disk encoding
[27] is used to convert the labeled clicks into a click map S∈R H×W×2 , representing each click as a disk with a smaller radius. Finally, concatenate S with the segmentation result P of the previous round at the channel level to obtain the interaction graph M∈R H×W×3 . At the first interaction, initialize P to an all-zero matrix.
[0022] The present invention follows the SimpleClick network architecture and uses the standard ViT-B as the backbone network. Two patch embedding layers segment the input image I and interaction map M into non-overlapping patches of fixed size (16×16). Each image patch is projected into a vector of fixed dimension (768) through a linear transformation. The vector sequence generated from the image and interaction map is element-wise added to obtain the final vector sequence. The sequence is then input into a series of Transformer blocks for self-attention processing. In the self-attention stage, GlClick superimposes the vector sequence enhanced with global information extracted by the Global Hints Module into the self-attention mechanism. Since the last feature map has been processed by all attention blocks, it is considered to have the most comprehensive representation ability. Therefore, the present invention only uses this last feature map to construct a simple multi-scale feature pyramid. The pyramid is then optimized by the All-MLP Segmentation Head and a segmentation mask is generated by upsampling. Finally, the Local Correction Module uses the image, interaction map, segmentation mask and previous segmentation mask to perform local enhancement and achieve fine adjustment.
[0023] About the Global Hint Module GHM: To fully leverage the global contextual information in the interaction graph and seamlessly integrate it with the traditional backbone network architecture, we developed a Global Hints Module (GHM). User interactions are encoded as a two-channel disk map, with one channel for positive clicks (representing the foreground) and the other for negative clicks (representing the background). The existing segmentation mask is merged with the two-channel click map to generate a three-channel interaction map. The GHM processes the interaction map by segmenting it into non-overlapping blocks of uniform size (16×16 pixels). Each block is flattened and mapped into a 768-dimensional fixed-length vector space via linear projection. This projection process is repeated three times, and the three resulting vectors are superimposed with the outputs of the first, sixth, and twelfth Transformer blocks in the backbone network to enhance the effect of the self-attention mechanism.
[0024] About the Local Correction Module LCM: In order to achieve better local optimization, such as Figure 2As shown, the local correction module (LCM) first determines the local region that requires refinement. First, the present invention compares the current segmentation result with the previous segmentation mask to obtain a difference mask, Mxor. Next, the maximum connected region containing the newly clicked image in Mxor is calculated and a bounding box for this region is generated. The present invention expands this region by a ratio of r = 1.4. Observing that this region is the core area for focused cropping, the present invention performs local cropping on the input image I to obtain I', and then crops the interaction map M to obtain M'.
[0025] In order to keep the network lightweight and simple, the present invention does not use a complex model to extract more features from I' and M', but uses three layers of convolutional layers for basic feature extraction to obtain feature Fs. However, in order to capture richer features, the present invention utilizes the feature pyramid output by the backbone network. By aligning the scale and upsampling, more information-rich features are obtained. These features are cropped using RoiAlign to obtain Fc. Fc is then scale-aligned with Fs and added element-by-element to finally generate a local mask. The ground truth is cropped using the same bounding box, and the loss L between the cropped ground truth and the local mask is calculated. loss . Global loss G loss and local loss L loss The binary cross entropy loss function is used.
[0026] Specific embodiments based on the above interactive image segmentation system are as follows: User clicks: The user clicks on the image to indicate the object the model should focus on. With the first click, the model determines the center of the segmented object based on the ground truth to guide the initial segmentation. In subsequent interactions, the user compares the ground truth with the previous segmentation result and makes the next click in the center of the region with the largest error.
[0027] Click encoding: Use the disk encoding method to convert the annotated clicks into a click map S, which is represented as a tensor of R^HxWx2, where each click is represented by a disk of smaller radius.
[0028] Global Hints Module (GHM): This module combines the interaction map and the previous segmentation results into a three-channel interaction map M, which is then segmented into non-overlapping blocks of 16x16 pixels. Each block is flattened and linearly projected into a 768-dimensional vector space, then interleaved with the outputs of the first, sixth, and twelfth transformation layers in the backbone network to enhance the self-attention mechanism.
[0029] ViT-B Backbone Network: Use the standard ViT-B as the backbone network to extract global information of images and interaction graphs.
[0030] All-MLP Segmentation Head: Use the All-MLP segmentation head to process the last feature map and upsample it to generate the segmentation mask.
[0031] Local Corrections Module (LCM): This module first identifies the local region requiring optimization. It compares the segmentation result with the mask from the previous pass to obtain a differential mask Mxor. It then calculates the largest connected region containing the new click, generates an outer bounding box for it, and proportionally enlarges the region. LCM then crops the local patch from the input image I and the interaction map M and uses three convolutional layers for basic feature extraction to obtain features Fs. To capture richer features, LCM also utilizes the feature pyramid output by the backbone network and crops it using RoiAlign to obtain features Fc. Finally, Fc and Fs are element-wise summed to generate the local mask.
[0032] Iterative Optimization: Glclick gradually improves the accuracy of the segmentation mask by iteratively optimizing user clicks and the network. In each iteration, the model performs local optimization based on user clicks and the previous segmentation results to reduce errors and improve segmentation accuracy.
[0033] Experimental results show that: Table 1 Fine-grained experimental results on four common datasets
[0034] Method: method name; Berkeley, DAVIS, COCO_MVal, and SBD are all names of datasets; NoC85, NoC90, NoC95: the number of clicks required to achieve an average intersection-union ratio of 85%, 90%, and 95%; †: indicates training on the SBD dataset; ‡: indicates training on the COCO_LVIM dataset; As shown in Table 1, the number of clicks required for the proposed method to achieve 85%, 90% and 95% IoU on two training sets is lower than that of previous methods.
[0035] like Figure 3 As shown in Figure 2, the figure is divided into two columns, with simple images on the left and more refined images on the right. It can be clearly seen that for the same number of clicks, the intersection-over-union ratio of both types of images is higher than that of the Simple-Click model.
[0036] In addition, if Figure 4 and Figure 5 As shown in the figure, the convergence analysis of the model trained on the DAVIS dataset and the convergence analysis of the model trained on the SBD dataset, the horizontal axis is the number of clicks, and the vertical axis is the average intersection-over-union ratio. Figure 4 and Figure 5 The superiority and convergence of the method of the present invention are demonstrated. Since the method of the present invention and the SOTA method SimpleClick model share the same backbone and training data set, a detailed comparison of the SimpleClick model is performed to demonstrate the superiority of the method of the present invention. The solid yellow line represents the SimpleClick model and the solid blue line represents the method of the present invention. Figure 4 and Figure 5 It can be seen that the proposed method is significantly superior to the SimpleClick model in terms of convergence speed and final accuracy.
[0037] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
[0038] Although the above describes the specific implementation methods of the present invention, it does not limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A click-based interactive image segmentation system, characterized by: It includes a global prompt module, a self-attention feature extraction module, a feature pyramid sampling module, a nonlinear segmentation head, and a local correction module; The global hint module introduces a patch embedding layer symmetrical to the patch embedding layer in the backbone. It encodes user click information into a disk map and concatenates it with the segmentation mask from the previous round to form a three-channel interaction map. This map is then embedded into the feature extraction network, enabling the network to learn the object boundaries and region information indicated by the user click. The self-attention feature extraction module uses ViT as the backbone network and a single-scale feature map structure to avoid information loss caused by feature map downsampling, learns local detail information of the image, and uses Transformer to learn global context information of the image; The feature pyramid sampling part uses the last feature map to construct a multi-scale feature map based on the single-scale feature pyramid, so that the network can capture the features of objects at different scales; The nonlinear segmentation head module uses the MLP layer to perform feature transformation and fusion, converting the multi-scale feature map into a segmentation probability map, avoiding complex convolution operations; The local correction module decomposes a heavy inference on the entire image into two light predictions on small blocks, and uses historical information to focus on the refined processing of the local area of the new click area.
2. The click-based interactive image segmentation system according to claim 1, wherein: The global prompt module is specifically: A patch embedding layer symmetrical to the one in the backbone is introduced, followed by element-wise feature addition; user clicks are encoded in a two-channel disk map, one for positive clicks and the other for negative clicks, where positive clicks are placed on the foreground and negative clicks are placed on the background; The segmentation mask and two-channel click map in the visual deformer are concatenated into a three-channel map for patch embedding. The two symmetric embedding layers operate on the image and the concatenated three-channel map respectively. The input is patched, flattened and projected into two vector sequences of the same dimension, and then element-wise addition is performed before input into the self-attention feature extraction module.
3. The click-based interactive image segmentation system according to claim 1, wherein: The self-attention feature extraction module is specifically: Using the traditional ViT as the backbone network, only single-scale feature maps are maintained throughout the process; the patch embedding layer divides the input image into non-overlapping fixed-size blocks, each block is flattened and linearly projected to a fixed-length vector, and the resulting vector sequence is fed into the Transformer block queue for self-attention. The backbone is pre-trained as MAEs on ImageNet-1k, and the pre-trained backbone is adjusted to a higher resolution input during fine-tuning using some non-shifting window attention assisted by global self-attention blocks, and global hint information is added. The user's interaction is encoded as a two-channel disk map, where one channel is assigned to positive clicks, representing the foreground, and the other is used for negative clicks, representing the background. The pre-existing segmentation mask and the two-channel click map are merged to create a three-channel interaction map; the global hint module processes this map by splitting it into non-overlapping blocks of the same size, and then flattens and linearly projects each block into a fixed-length vector space with 768 dimensions. This projection is performed in triplicate, and the resulting set of three vectors is superimposed with the outputs of the first, sixth, and twelfth conversion layers to enhance the self-attention mechanism.
4. The click-based interactive image segmentation system according to claim 1, wherein: The feature pyramid sampling uses the last feature map output by the self-attention feature extraction module and uses four convolutional layers with different strides to downsample the last feature map to resolutions of 1 / 32, 1 / 16, 1 / 8 and 1 / 4 respectively, generating a multi-scale feature map so that the network can better capture the features of objects at different scales.
5. The click-based interactive image segmentation system according to claim 1, wherein: The nonlinear segmentation head uses an MLP layer as a lightweight segmentation head, adopts a simple feature pyramid, generates a segmentation probability map with a scale of 1 / 4, and then performs an upsampling operation to restore the original resolution; the MLP segmentation head works in three steps: first, each feature map from the simple feature pyramid passes through the MLP layer to convert it to the same channel dimension. Second, all feature maps are upsampled to the same resolution for concatenation; third, the concatenated features are fused with another MLP layer to produce a single-channel feature map, and then a sigmoid function is used to obtain the segmentation probability map, which is then converted to a binary segmentation given a predefined threshold.
6. The click-based interactive image segmentation system according to claim 1, wherein: The local correction module decomposes a heavy inference on the entire image into two light predictions for small patches; first, an image patch around the target object is selected, resized to a small scale, and sent to Backbone to predict a coarse mask; then, a local area is selected around the click, and the local patch that needs to be refined and enlarged is fed into the local correction module. Finally, progressive merging aligns the local prediction back to the full-size mask, refining only a small local area after each click, and all pixels of the final prediction are refined through calculations assigned to different rounds.
7. The click-based interactive image segmentation system according to claim 6, wherein: Target Crop is used to filter out background information that is irrelevant to the target object. First, the minimum external box of the previous mask and the newly added click are calculated, and then the ratio r is used to calculate the minimum external box of the previous mask and the newly added click. TC = 1.4 to expand it, then crop the input tensor and resize it; Focus Crop is used to locate the area that the user intends to modify; first, the difference between the original segmentation result and the previous mask is compared to obtain the differential mask Mxor, and then the maximum connected area containing the newly clicked Mxor is calculated, and the external box is generated for this maximum connected area in r FC = 1.4 ratio; crop local patches on the input image and click map; Refinement Refiner is used to recover the details of the coarse prediction in the focus crop. First, a depth-wise separable convolution is used to extract low-level features from the crop tensor. At the same time, the number of channels of the regional features is adjusted and fused with the extracted low-level features. The two heads are used to predict the detail map M. d and boundary graph M b , and update the rough prediction logarithm M l The boundary area is used to calculate the refined prediction Mr.
8. A click-based interactive image segmentation method, characterized in that: The interactive image segmentation system according to any one of claims 1 to 7 is used, and includes the following process: First, the user clicks are determined. The user clicks on the image to instruct the model to focus on the object. The first click determines the center of the segmentation target, and subsequent clicks are in the center of the area with the largest error. The clicks are then encoded using disk encoding to form an interaction graph, represented as a tensor; The interaction graph and the original image are fed into the self-attention feature extraction module and the self-attention is enhanced through the global hint module to extract the global information of the image and the interaction graph; The last feature map is then processed using a nonlinear segmentation head module and upsampled to generate a coarse segmentation mask; At the same time, the local correction module is used to identify the local areas that need to be optimized, and a lightweight network is used to generate local fine segmentation masks; The entire system is iteratively optimized. Through user clicks, a segmentation mask is formed. Click again, and it is iteratively optimized to gradually improve the accuracy of the segmentation mask.
Citation Information
Patent Citations
Local correction interactive medical image segmentation method and system
CN116934759A
Interactive image segmentation method and system based on fine tuning of lightweight adapter
CN117372444A
Transform-based interactive image segmentation method
CN117372701A
Unsupervised zero-shot segmentation mask generation and semantic labeling
US20250045930A1
Cited By
Interactive annotation segmentation method and device for full-slice pathological image
CN122156650A
An interactive annotation and segmentation method and device for whole-section pathological images
CN122156650B