Complex target-oriented interactive friendly segmentation method and system

Through the combination of two-stage workflow and attention mechanism, the problem of insufficient efficiency and accuracy in the complex target and high-resolution image segmentation is solved, and an efficient and accurate interactive friendly segmentation effect is achieved.

CN119991725APending Publication Date: 2025-05-13BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510048445.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing interactive segmentation technology has insufficient efficiency and accuracy when processing complex targets and high-resolution images, especially when processing small targets and complex structures, and the user's operational efficiency and experience have significantly decreased.

Method used

An interactive friendly segmentation method for complex goals is proposed, and a two-stage workflow is adopted: the first stage generates a foreground-background-uncertained area map through an explicit rough perception network, and the second stage fine-classify the uncertain areas at ultra-high resolution through a high-resolution refined network, combining grid attention and neighborhood attention mechanisms.

Benefits of technology

It significantly improves interaction efficiency and segmentation accuracy, reduces calculation complexity, is suitable for segmentation tasks of complex targets and high-resolution images, and improves user operation convenience and the quality of segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991725A_ABST
    Figure CN119991725A_ABST
Patent Text Reader

Abstract

The invention provides an interactive friendly segmentation method and system for a complex target, and belongs to the technical field of computer vision, and the method comprises the steps: obtaining a to-be-segmented image; and processing the obtained to-be-segmented image by using a pre-trained interactive friendly segmentation model facing the complex target to obtain a target segmentation result. According to the method, by introducing the noise tolerant click, the user is allowed to perform fuzzy positioning near the target without accurate click, so that the interaction difficulty and the time cost are remarkably reduced, and meanwhile, the user friendliness in a complex target segmentation task is enhanced. Two-stage workflow is adopted: in the first stage, a foreground-background-uncertain region (FBU) graph is generated, and a target region is quickly sensed; and in the second stage, the uncertain regions are accurately classified through a high-resolution refining network. And in combination with grid attention and neighborhood attention mechanisms, the detail information of the ultrahigh-resolution image is reserved, and the calculation complexity is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an interactive friendly segmentation method and system for complex targets. Background Art

[0002] Interactive image segmentation is an important topic in the field of computer vision, which aims to quickly generate high-precision target segmentation masks through simple user interactions (such as clicking or scribbling). This technology is widely used in scenarios such as data annotation, medical image analysis, and industrial inspection. In particular, it has played a key role in promoting the development of large-scale visual models (such as Segment Anything Model, SAM) in recent years. However, with the increase in the complexity and resolution of target objects, interactive segmentation technology has gradually exposed its shortcomings in efficiency and accuracy, and new methods are urgently needed to improve its practicality and robustness.

[0003] In the early stage of interactive segmentation, the research mainly relied on graph cut models (such as GrabCut, Rother et al., 2004) or Active Contour models (such as Snakes, Kass et al., 1988) to generate segmentation masks. These methods segment locally by optimizing the energy function of foreground and background, but their performance is limited when dealing with high-complexity targets. In recent years, deep learning-based interactive segmentation methods, such as DIOS (Xu et al., 2016), learn global image features through deep convolutional networks, and SimpleClick (Liu et al., 2023) uses visual Transformer as the backbone network to improve the segmentation accuracy of interactive segmentation models.

[0004] At present, click-based interaction has become the mainstream method, and users can achieve target segmentation by clicking on the foreground or background area. For example, DEXTR (Maninis et al., 2018) generates a target mask by clicking on four extreme points, further optimizing the number of user clicks and making the operation more efficient. However, when the target area is small or the boundary is complex, the need for precise positioning of clicks significantly increases the interaction cost. In this regard, graffiti-based interaction forms, such as Slim Scissors (Han et al., 2022), allow users to segment the target area by rough smearing to reduce the difficulty of operation, but the efficiency of large-scale smearing is still limited when dealing with dense and small structures. Therefore, whether it is frequent precise clicks or graffiti, there are deficiencies in interaction efficiency.

[0005] At the same time, with the enhancement of visual model capabilities, the segmentation task has an increasing demand for processing high-resolution images. However, existing methods face significant computational bottlenecks in high-resolution scenarios. Although segmentation methods with Transformers as the core architecture (such as SAM, Kirillov et al., 2023) have strong global perception capabilities, the computational complexity of their self-attention mechanism grows with the square of the resolution (Dosovitskiy et al., 2021). Therefore, the input image is usually downsampled to a fixed size (such as 1024×1024), which inevitably leads to the loss of target details, especially when dealing with fine targets (such as ThinObject5K, Liew et al., 2021) or high-resolution salient targets (such as HRSOD, Zeng et al., 2019).

[0006] Existing interactive segmentation methods mainly rely on users’ precise clicks (such as foreground clicks and background clicks) or rough scribble annotations. However, for small objects or complex structures, precise click operations are cumbersome and time-consuming, especially when the object is only a few pixels wide (such as Figure 1 The efficiency and experience of user operation are significantly reduced. At the same time, although the rough graffiti method has low requirements for precise positioning, it still requires a lot of smearing operations when processing dense target areas, and the efficiency is not ideal.

[0007] Complex objects often have higher image resolutions in order to preserve details. However, the computational complexity of mainstream visual Transformer-based segmentation models increases with the square of the image resolution. Therefore, due to hardware limitations, it is usually necessary to downsample high-resolution images to a fixed size (such as 1024×1024) to reduce computational complexity. This operation inevitably loses detail information, making it difficult for the model to accurately segment complex areas or fine edges. Although some methods attempt to improve accuracy through multi-stage refinement, their performance on ultra-high-resolution (2048×2048 and above) images is still limited due to memory and computing resources. Summary of the invention

[0008] The object of the present invention is to provide an interactive friendly segmentation method and system for complex targets to solve at least one technical problem existing in the above-mentioned background technology.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] In a first aspect, the present invention provides an interactive friendly segmentation method for complex objects, comprising:

[0011] Obtain an image to be segmented;

[0012] The acquired image to be segmented is processed by using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty region map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty region map to generate a final segmentation mask.

[0013] As a further limitation of the first aspect of the present invention, the high-resolution refinement network receives a low-resolution foreground-background-uncertain region map and upsamples it to a high resolution, and then combines it with the original image to be segmented to perform feature extraction and classification.

[0014] As a further limitation of the first aspect of the present invention, the high-resolution refinement network combines grid attention and neighborhood attention mechanisms, the grid attention constructs sparse long-distance dependencies in the image, and the neighborhood attention further improves the detail capture capability by enhancing local dependencies.

[0015] As a further limitation of the first aspect of the present invention, the explicit coarse perception network takes a low-resolution image to be segmented as input, combines click information provided by a user, including foreground clicks, background clicks, and noise-tolerant clicks, to distinguish the foreground, background, and complex areas in the image; the uncertain areas are not subclassified but marked as "uncertain".

[0016] As a further limitation of the first aspect of the present invention, the explicit coarse perception network adopts the standard ViT as the backbone network, combines the window attention and global attention modules, and a lightweight decoder to achieve three-category mask output.

[0017] As a further limitation of the first aspect of the present invention, the training of the interactive friendly segmentation model for complex targets includes: converting the fine mask provided by the training data set into the true value of the required foreground-background-uncertainty area map through a series of expansion and corrosion operations, and storing it for training the explicit coarse perception network; at the same time, the high-precision annotated data set is not processed and stored for training the high-resolution refinement network; through the true value of the foreground-background-uncertainty area map composed of three categories, the required foreground points, background points and noise tolerance points are randomly sampled in the corresponding area; in the feature extraction stage, the visual Transformer architecture is used to extract the foreground points, background points and noise tolerance points in the interactive With the assistance of points, the extraction from image to feature is completed; in the foreground-background-uncertainty area map prediction stage, the explicit coarse perception network will first predict the preliminary perception results at low resolution, including the uncertain areas of the foreground and background, and output them in the form of a foreground-background-uncertainty area map; in the high-resolution refinement stage, the high-resolution refinement network will, with the assistance of the foreground-background-uncertainty area map, perform fine perception and clear classification of each pixel; by comparing the output final segmentation result with the true value mask of the data set, the loss is calculated and the model parameters are updated with the optimizer; the above process is repeated until the model test results meet expectations or the number of training times is reached.

[0018] In a second aspect, the present invention provides an interactive friendly segmentation system for complex objects, comprising:

[0019] An acquisition module, used for acquiring an image to be segmented;

[0020] A processing module is used to process the acquired image to be segmented using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty region map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty region map to generate a final segmentation mask.

[0021] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the interactive and friendly segmentation method for complex targets as described in the first aspect is implemented.

[0022] In a fourth aspect, the present invention provides a computer device comprising a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the interactive friendly segmentation method for complex targets as described in the first aspect.

[0023] In a fifth aspect, the present invention provides an electronic device, comprising: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the interactive friendly segmentation method for complex targets as described in the first aspect.

[0024] Terminology explanation:

[0025] Interactive segmentation: Interactive segmentation is an important task in computer vision. It uses user interaction (such as clicking, drawing, etc.) to segment the foreground and background of an image. Interactive segmentation models are widely used, such as in content creation, medical image analysis, etc., to help users accurately extract target areas for further analysis or editing. Interactive segmentation based on deep learning networks has extremely high segmentation accuracy and generalization, which is of vital importance for practical applications and the exploration of basic visual models.

[0026] Fine-grained segmentation: Fine-grained segmentation is a technique that processes small objects, complex structures, or boundary areas with higher precision in image segmentation tasks. Traditional segmentation methods usually complete the preliminary division of the target area at a low resolution or coarse-grained level, which may result in loss of details or blurred boundaries. The goal of fine-grained segmentation is to accurately classify the target area pixel by pixel through higher resolution, more accurate feature extraction and classification strategies.

[0027] Beneficial effects of the present invention: By introducing noise-tolerant clicks, users are allowed to perform fuzzy positioning near the target without precise clicks, thereby significantly reducing the difficulty and time cost of interaction, while enhancing user-friendliness in complex target segmentation tasks. A two-stage workflow is adopted: the first stage generates a foreground-background-uncertainty area (FBU) map to quickly perceive the target area; the second stage accurately classifies the uncertain area through a high-resolution refinement network. Combining the grid attention and neighborhood attention mechanisms, it not only retains the detail information of the ultra-high-resolution image, but also effectively reduces the computational complexity.

[0028] Additional advantages of the present invention will be more clearly presented in the following description or learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0030] Figure 1 This is an overall framework diagram of the interactive friendly segmentation model for complex targets described in an embodiment of the present invention.

[0031] Figure 2 This is a training flowchart of the interactive friendly segmentation model for complex targets described in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below by the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.

[0033] It should be understood by those skilled in the art that unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.

[0034] It should also be understood that terms, such as those defined in commonly used dictionaries, should be understood to have a meaning consistent with that in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless as defined herein.

[0035] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements and / or groups thereof.

[0036] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. Different embodiments or examples described in this specification and features of different embodiments or examples may be combined and combined by those skilled in the art without contradiction.

[0037] To facilitate understanding of the present invention, the present invention is further explained below with reference to specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation on the embodiments of the present invention.

[0038] Those skilled in the art should understand that the drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily necessary for implementing the present invention.

[0039] This paper proposes an interactive and friendly segmentation model for complex targets, which achieves high-precision segmentation by introducing noise-tolerant clicks and a two-stage workflow. Noise-tolerant clicks allow users to perform imprecise positioning near the target, significantly improving the convenience of user operations. The two-stage workflow first predicts the FBU map (Foreground Background Uncertain) containing foreground, background and uncertain areas through an explicit rough perception network, and then uses a high-resolution refinement network to accurately classify the uncertain areas, thereby taking into account both efficiency and accuracy. Experimental results show that the model significantly outperforms existing methods on multiple high-resolution and complex target data sets, achieving an ideal balance between segmentation accuracy and interactive efficiency. Its efficient and friendly interactive mode and excellent processing capabilities for detailed targets make it have important practical application value in scenarios such as data annotation and medical image analysis.

[0040] The present invention aims to solve two key problems of fine target segmentation in interactive segmentation tasks: low interaction efficiency and limited resolution. Although the traditional click-based interaction method is flexible and efficient, when the target area is complex and small (such as only a few pixels wide), the user needs to accurately locate the click position, which significantly reduces the operation efficiency. At the same time, since the computational complexity of the self-attention mechanism in the existing method increases quadratically with the image resolution, high-resolution images are usually forced to be downsampled, which leads to the loss of target detail information and affects the segmentation quality.

[0041] Example 1

[0042] In this embodiment 1, firstly, an interactive friendly segmentation system for complex targets is provided, including: an acquisition module, used to acquire an image to be segmented. A processing module, used to process the acquired image to be segmented using a pre-trained interactive friendly segmentation model for complex targets, to obtain a target segmentation result; wherein, the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty region map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty region map to generate a final segmentation mask.

[0043] In this embodiment, the above-mentioned system is used to implement an interactive and friendly segmentation method for complex targets, including: obtaining an image to be segmented; using a pre-trained interactive and friendly segmentation model for complex targets to process the obtained image to be segmented to obtain a target segmentation result; wherein the interactive and friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty area map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty area map to generate a final segmentation mask.

[0044] like Figure 1 As shown, in this embodiment, a new interaction method called noise-tolerant click is proposed to significantly improve the efficiency and user experience of interactive image segmentation. Traditional interaction methods usually rely on users to accurately click on the target area. However, when dealing with complex or small targets, this method is not only cumbersome to operate, but may also lead to a decrease in segmentation accuracy. To solve this problem, noise-tolerant click allows users to click fuzzily near the target without accurately locating the target area. Regardless of whether the click falls on the target area, boundary or background, the model can identify the small structure near the click area through contextual information, thereby generating accurate segmentation results. Noise-tolerant clicks and foreground points are encoded together with background points into a disk map, and are input into the model after being combined with image stitching. This new form of interaction significantly reduces the difficulty of user operation, making the segmentation process more efficient and flexible, and is particularly suitable for scenes with complex targets and rich details.

[0045] ECP (Explicit Coarse Perception Network) is the first stage of the two-stage workflow of the interactive-friendly segmentation model. It is specifically designed to quickly generate a "foreground-background-uncertainty region map" (FBU map) to provide prior information for subsequent fine segmentation. ECP takes a low-resolution image as input and combines the click information provided by the user, including foreground clicks, background clicks, and noise-tolerant clicks, to distinguish the foreground, background, and complex regions in the image. Uncertain regions usually cover complex details such as edges or fine structures. These regions are not finely classified for the time being, but are marked as "uncertain". In order to train ECP efficiently, the fine annotations in the segmentation dataset are converted into coarse FBU masks through morphological processing (such as corrosion and dilation) to serve as supervision signals for the network. In terms of architecture, ECP uses the standard ViT (Vision Transformer) as the backbone network, combined with window attention and global attention modules, and a lightweight decoder, which can achieve fast and high-quality three-classification mask output at low resolution. In this way, ECP can provide strong prior support for high-resolution fine segmentation in complex scenes with low computational overhead, significantly improving the overall segmentation accuracy and efficiency.

[0046] High Resolution Refinement Network, HRR High Resolution Refinement Network (HRR) is the second stage of the model, focusing on fine segmentation of uncertain regions generated by ECP and generating the final high-quality mask. HRR receives the low-resolution FBU Map and upsamples it to high resolution, and combines it with the original RGB image for feature extraction and classification. In order to achieve efficient computation at ultra-high resolution, HRR designs a mechanism that combines grid attention and neighborhood attention. Grid attention significantly reduces the computational complexity of global attention by building sparse long-distance dependencies in the image; neighborhood attention further improves the ability to capture details by enhancing local dependencies. Compared with the standard ViT architecture that builds long-distance dependencies through global attention, HRR's way of building long-distance dependencies through grid attention also greatly reduces the spatial complexity of this link, thereby reducing the requirement for hardware video memory size. Through this combination of attention mechanisms, HRR can operate efficiently in ultra-high resolution scenarios while maintaining sensitivity to small structures and complex boundaries. Finally, the segmentation results generated by HRR have extremely high accuracy and perform well in multiple complex scenes and high-resolution tasks. HRR is trained in a completely decoupled manner, requiring only a separate loss function and optimizer to be trained directly on a high-precision labeled dataset.

[0047] At the same time, the HRR module can perform reasoning at multiple resolutions, so users can adjust the reasoning resolution size according to the device they are actually using. A larger resolution will bring greater computational overhead and higher accuracy, but even when reasoning at a relatively low resolution (such as 1024×1024), HRR has high perceptual accuracy.

[0048] like Figure 2 As shown, the training process of the interactive friendly segmentation model described in this embodiment is as follows:

[0049] S1: The fine mask provided by the training dataset is converted into the true value of the required FBUmap through a series of dilation and corrosion operations, and stored for training the ECP network. At the same time, the high-precision labeled dataset is not processed and stored for training the HRR network.

[0050] S2: Using the true value of the FBU map consisting of three categories, the required foreground points, background points, and noise tolerance points are randomly sampled in the corresponding area.

[0051] S3: In the feature extraction stage, the present invention uses the visual Transformer architecture to complete the extraction from image to feature with the assistance of interaction points.

[0052] S4: In the FBU map prediction stage, ECP will first predict the preliminary perception results at low resolution, including the uncertain areas of the foreground and background, and output them in the form of FBU map.

[0053] S5: In the high-resolution refinement stage, HRR, with the assistance of the FBU map, will perform detailed perception and clear classification of each pixel.

[0054] S6: By comparing the final segmentation result output by S5 with the true value mask of the dataset, calculate the loss and use the optimizer to update the model parameters. Repeat S1-S5 until the model test results meet expectations or the training times are reached.

[0055] Example 2

[0056] This embodiment 2 provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the interactive friendly segmentation method for complex objects as described above is implemented. The method includes:

[0057] Obtain an image to be segmented;

[0058] The acquired image to be segmented is processed by using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty region map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty region map to generate a final segmentation mask.

[0059] Example 3

[0060] This embodiment 3 provides a computer device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, the processor calls the program instructions to execute the interactive friendly segmentation method for complex objects as described above, the method comprising:

[0061] Obtain an image to be segmented;

[0062] The acquired image to be segmented is processed by using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty region map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty region map to generate a final segmentation mask.

[0063] Example 4

[0064] This embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to execute instructions for implementing the interactive friendly segmentation method for complex targets as described above, the method including:

[0065] Obtain an image to be segmented;

[0066] The acquired image to be segmented is processed by using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertainty region map to provide prior information for subsequent fine segmentation; the high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertainty region map to generate a final segmentation mask.

[0067] In summary, the present invention proposes an interactive and friendly segmentation model for complex targets to address the two major problems of low interaction efficiency and resolution limitation. First, by introducing noise-tolerant clicks, users are allowed to fuzzily click near the target without accurately locating small areas, significantly reducing the interaction cost. Compared with traditional click methods, noise-tolerant clicks are more efficient and user-friendly in complex target segmentation. Secondly, in order to break through the resolution limitation: a two-stage workflow is designed: in the first stage, a foreground-background-uncertain region map (Foreground Background Uncertain map) is generated through an explicit coarse perception network (Explicit Coarse Perception, ECP) ​​to preliminarily identify areas that may contain complex structures; in the second stage, a high-resolution refinement network (High Resolution Refinement, HRR) is used to accurately classify uncertain areas at ultra-high resolution. By introducing grid attention and neighborhood attention mechanisms, the model optimizes the construction of local and global attention mechanisms, while significantly reducing computational complexity.

[0068] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0070] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0072] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative labor on the basis of the technical solution disclosed in the present invention should be included in the scope of protection of the present invention.

Claims

1. An interactive and friendly segmentation method for complex objects, characterized in that: include: Obtain an image to be segmented; The acquired image to be segmented is processed using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertain region map to provide prior information for subsequent fine segmentation; The high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertain region map to generate a final segmentation mask.

2. The interactive friendly segmentation method for complex objects according to claim 1, characterized in that: The high-resolution refinement network receives the low-resolution foreground-background-uncertain region map and upsamples it to high resolution, and then combines it with the original image to be segmented to perform feature extraction and classification.

3. The interactive friendly segmentation method for complex objects according to claim 2, characterized in that: The high-resolution refinement network combines grid attention and neighborhood attention mechanisms. Grid attention builds sparse long-distance dependencies in the image, while neighborhood attention further improves the ability to capture details by enhancing local dependencies.

4. The interactive friendly segmentation method for complex objects according to claim 1, characterized in that: The explicit coarse perception network takes the low-resolution image to be segmented as input and combines the click information provided by the user, including foreground clicks, background clicks, and noise-tolerant clicks, to distinguish the foreground, background, and complex areas in the image. Uncertain areas are not subclassified but marked as "uncertain".

5. The interactive friendly segmentation method for complex objects according to claim 4, characterized in that: The explicit coarse perception network adopts the standard ViT as the backbone network, combines the window attention and global attention modules, and a lightweight decoder to achieve three-category mask output.

6. The interactive friendly segmentation method for complex objects according to claim 1, characterized in that: The training of the interactive friendly segmentation model for complex targets includes: converting the fine mask provided by the training data set into the true value of the required foreground-background-uncertainty region map through a series of expansion and corrosion operations, and storing it for training the explicit coarse perception network; at the same time, the high-precision labeled data set is not processed and stored for training the high-resolution refinement network; through the true value of the foreground-background-uncertainty region map composed of three categories, the required foreground points, background points and noise tolerance points are randomly sampled in the corresponding area; in the feature extraction stage, the visual Transformer architecture is used to complete the feature extraction with the assistance of the interactive points. From image to feature extraction; in the foreground-background-uncertainty region map prediction stage, the explicit coarse perception network will first predict the preliminary perception results at low resolution, including the uncertain regions of the foreground and background, and output them in the form of a foreground-background-uncertainty region map; in the high-resolution refinement stage, the high-resolution refinement network will, with the assistance of the foreground-background-uncertainty region map, perform fine perception and clear classification of each pixel; by comparing the output final segmentation result with the true value mask of the data set, calculate the loss and use the optimizer to update the model parameters; repeat the above process until the model test results meet expectations or the number of training times is reached.

7. An interactive and friendly segmentation system for complex objects, characterized in that: include: An acquisition module, used for acquiring an image to be segmented; A processing module, used to process the acquired image to be segmented using a pre-trained interactive friendly segmentation model for complex targets to obtain a target segmentation result; wherein the interactive friendly segmentation model for complex targets includes an explicit coarse perception network and a high-resolution refinement network; the explicit coarse perception network is used to generate a foreground-background-uncertain region map to provide prior information for subsequent fine segmentation; The high-resolution refinement network is used to perform fine segmentation on the generated foreground-background-uncertain region map to generate a final segmentation mask.

8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the interactive and friendly segmentation method for complex targets as described in any one of claims 1-6 is implemented.

9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the interactive friendly segmentation method for complex targets as described in any one of claims 1-6.

10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the interactive friendly segmentation method for complex targets as described in any one of claims 1-6.