Image segmentation method and device based on dynamic Prompt constraint and edge anchor point guidance of SAM model, storage medium and electronic equipment

By introducing a lightweight dynamic regulation network and edge anchor point guidance mechanism in the SAM model, the problems of deep network diffusion and semantic dispersion are solved, and stable segmentation performance and efficient boundary positioning are achieved in low-contrast medical images, improving the segmentation effect of weak edge targets.

CN120279276AActive Publication Date: 2025-07-08SICHUAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510765209.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Existing SAM models have problems with deep network diffusion and semantic representation consistency deterioration in low-contrast medical images, resulting in a decrease in target boundary positioning capabilities, especially in weak edge target segmentation tasks.

Method used

The lightweight dynamic adjustment network and edge anchor point guidance mechanism is adopted, and the intensity and spatial range of interactive prompts are controlled through the layered Prompt constraint mechanism, and combined with the edge anchor point guidance mechanism, the feature focus ability and training efficiency of the image encoder are improved.

Benefits of technology

Maintaining stable segmentation performance in low-contrast medical images improves the boundary positioning accuracy and generalization ability of the model, especially in weak edge target segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279276A_ABST
    Figure CN120279276A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method and device based on dynamic Prompt constraint and edge anchor point guidance of an SAM model, a storage medium and electronic equipment, and relates to the technical field of image segmentation, and the method comprises the steps: employing an MLP structure design and training to obtain a lightweight dynamic adjustment network; a lightweight dynamic adjustment network is utilized to control the intensity and the space range of an interactive prompt provided by a user to be introduced into an image encoder; training an image encoder based on an edge anchor point guiding mechanism; inputting the introduction result of the interactive prompt and the image into an image encoder to obtain an image embedding vector; inputting the interactive prompt into a prompt encoder to obtain a prompt embedding vector; and inputting the image embedding vector and the prompt embedding vector into a mask decoder for interaction, and obtaining a segmented mask image after multi-layer cross attention and up-sampling. According to the method, the problem that the segmentation boundary is fuzzy and inconsistent with the target under the traditional weak labeling condition is effectively solved, and extremely high medical image segmentation capability and generalization are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image segmentation technology. Specifically, it relates to an image segmentation method, device, storage medium, and electronic device based on dynamic Prompt constraint and edge anchor point guidance of the SAM model. Background Art

[0002] The typical architecture of the SAM model is as follows: The prompt information such as points and boxes provided by the user is first encoded by the Prompt Encoder, but these encoded Prompt signals only interact with the image features in the Mask prediction head of the Decoder to guide the generation of the segmentation result. At the same time, the input image extracts multi-scale features through multiple Transformer Blocks (Stage1 - Stage4) of the Image Encoder, and the features at each stage aggregate global context information through the self-attention mechanism.

[0003] Although this design can establish long-range dependencies, due to the lack of explicit constraints on the Prompt signal, it leads to two key defects: On the one hand, in difficult samples such as low-contrast medical images, the semantic representation of the deep network (Stage3 - 4) will over-diffuse, and the attention distribution shows a global diffusion characteristic. On the other hand, when dealing with hard samples, the full-level attention defocus phenomenon occurs starting from Stage2, resulting in the model almost completely losing the ability to locate the target boundary. The root cause of these defects is that the existing architecture fails to establish an effective cross-level Prompt constraint mechanism, making the Image Encoder unable to maintain continuous attention to the prompt area, and ultimately leading to a significant deterioration of the feature expression consistency as the network depth increases. In addition, since the deep network needs to learn a large number of redundant background features to identify edge details, this directly leads to a significant decline in the performance of the SAM model in the segmentation task of weak-edge targets (such as early tumors). Therefore, there is an urgent need for an image segmentation method to make up for the defects existing in the prior art. Summary of the Invention

[0004] Embodiments of this application provide an image segmentation method, device, storage medium, and electronic device based on dynamic Prompt constraint and edge anchor point guidance of the SAM model to solve the technical problems existing in the prior art.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0006] According to the first aspect of the embodiments of the present application, an image segmentation method based on dynamic Prompt constraint and edge anchor guidance of the SAM model is provided. In the SAM model, an image encoder, a prompt encoder, and a mask decoder are provided. The method includes: Design and train a lightweight dynamic adjustment network using an MLP structure; Use the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder; Train the image encoder based on the edge anchor guidance mechanism; Input the introduction result of the interactive prompt and the image provided by the user into the trained image encoder, and process to obtain an image embedding vector; Input the interactive prompt provided by the user into the prompt encoder to process and obtain a prompt embedding vector; Input the image embedding vector and the prompt embedding vector into the mask decoder for interaction, and obtain a segmented mask image after multi-layer cross-attention and upsampling.

[0007] In some embodiments of the present application, based on the foregoing solution, designing a lightweight dynamic adjustment network using an MLP structure includes: Design a lightweight dynamic adjustment network using a lightweight architecture of MLP + multi-head output.

[0008] In some embodiments of the present application, based on the foregoing solution, the process of training the lightweight dynamic adjustment network includes: Train the lightweight dynamic adjustment network using a three-stage progressive training scheme, including: In the first stage, freeze the backbone network and only train the modules related to the interactive prompt; In the second stage, unfreeze the high-level network for adjustment; In the third stage, adjust the parameters of the entire model.

[0009] In some embodiments of the present application, based on the foregoing solution, the using the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder includes: Input the interactive prompt provided by the user and the feature maps output from different layers in the image encoder into the lightweight dynamic adjustment network for processing, and generate a set of fusion weights and diffusion coefficients that adaptively change with the hierarchy.

[0010] In some embodiments of the present application, based on the foregoing solution, the training the image encoder based on the edge anchor guidance mechanism includes: Extract edge responses from the shallow features of the image encoder using depthwise separable convolutions, and cooperate with the batch normalization layer and the Gaussian error linear unit activation to generate an anchor mask; Establish a dual alignment mechanism. In the spatial dimension, use the anchor mask to shield the gradients of non-anchor regions during backpropagation, and only update the weights corresponding to the anchors. In the semantic dimension, calculate the L2 alignment loss function between the deep features and the anchor features; Adopt an alternating training strategy. In odd training epochs, only the anchor regions participate in gradient updates, and in even training epochs, the entire image resumes normal training; During the entire training process, dynamically balance the interaction between the shallow and deep features in the image encoder through learnable gating parameters.

[0011] In some embodiments of the present application, based on the foregoing solution, inputting the result of introducing the interactive prompt and the image provided by the user into the trained image encoder to process and obtain an image embedding vector includes: Encode the interactive prompt provided by the user into a Gaussian heatmap; Combine the fusion weight, the diffusion coefficient, and the Gaussian heatmap, and input the combination result and the image provided by the user into the image encoder to process and obtain an image embedding vector.

[0012] According to the second aspect of the embodiments of the present application, there is provided an image segmentation device based on a SAM model with dynamic Prompt constraints and edge anchor guidance, including: A first generation unit for designing and training a lightweight dynamic adjustment network using an MLP structure; A fusion unit for using the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder; A training unit for training the image encoder based on the edge anchor guidance mechanism; A second generation unit for inputting the result of introducing the interactive prompt and the image provided by the user into the trained image encoder to process and obtain an image embedding vector; A third generation unit for inputting the interactive prompt provided by the user into the prompt encoder to process and obtain a prompt embedding vector; A segmentation processing unit for inputting the image embedding vector and the prompt embedding vector into the mask decoder for interaction, and obtaining a segmented mask image after multi-layer cross-attention and upsampling.

[0013] According to the third aspect of the embodiments of the present application, there is provided a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions run on a computer, the computer is caused to execute the method described in the first aspect.

[0014] According to a fourth aspect of the embodiments of the present application, an electronic device is provided, including: a memory and a processor; The memory is used to store computer instructions; The processor is configured to call the computer instructions stored in the memory, so that the electronic device executes the method described in the first aspect.

[0015] The technical solution of the present application solves the problem of attention diffusion in deep networks by establishing a hierarchical interactive prompt constraint mechanism in the image encoder, enabling the model to maintain stable segmentation performance in low-contrast medical images. The edge geometric anchor point guidance mechanism slows down the problem of semantic dispersion and improves the model training efficiency. While inheriting the efficient interaction characteristics of the image segmentation method based on the SAM model, the overall technical solution effectively solves the problems of fuzzy segmentation boundaries and inconsistent targets under traditional weak annotation conditions, demonstrating strong medical image segmentation capabilities and generalization.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings: Figure 1 Shows a flowchart of an existing method for image segmentation using the SAM model according to an embodiment of the present application; Figure 2 Shows a schematic flowchart of an image segmentation method based on dynamic Prompt constraint and edge anchor point guidance of the SAM model according to an embodiment of the present application; Figure 3 Shows a block diagram of an image segmentation device based on dynamic Prompt constraint and edge anchor point guidance of the SAM model according to an embodiment of the present application; Figure 4 Shows a block diagram of an electronic device according to an embodiment of the present application; Figure 5 Shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.

[0018] DESCRIPTION OF THE REFERENCE NUMERALS 300 - Image segmentation device based on dynamic Prompt constraint and edge anchor point guidance of SAM model, 301 - First generation unit, 302 - Fusion unit, 303 - Training unit, 304 - Second generation unit, 305 - Third generation unit, 306 - Segmentation processing unit, 400 - Electronic device, 410 - Memory, 411 - Computer program 411, 420 - Processor, 500 - Computer system, 501 - Central processing unit, 502 - Read-only memory, 503 - Random access memory, 504 - Bus, 505 - Input / output interface, 506 - Output section, 507 - Output section, 508 - Storage section, 509 - Communication section, 510 - Driver, 511 - Removable medium. Detailed implementation manners

[0019] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0020] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.

[0021] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0022] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0023] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the objects used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described.

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.

[0025] The following will describe in detail some embodiments of the present application. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0026] As Figure 1 shown, the process of the existing method for image segmentation using the SAM model is as follows: 1. The most primitive SAM has two inputs: an image and an interactive prompt (including points / box / text / mask); 2. The image vector is first sent into the image encoder to obtain an image embedding vector; at the same time, the interactive prompt is also sent to the prompt encoder to obtain a prompt embedding vector; 3. Finally, the image embedding vector obtained by the image encoder and the prompt embedding vector obtained by the prompt encoder are sent into the mask decoder for interaction. After multiple layers of cross-attention and upsampling, a segmented mask image is finally obtained (by default, three masks with high confidence are generated for the user to choose from, or the user can manually set to generate only one).

[0027] Among them, mask represents a mask, points represents a point prompt, box represents a box prompt, and text represents a text prompt.

[0028] The existing technology for image segmentation using the SAM model has the following problems and defects.

[0029] (1) The problem of unbalanced hierarchical features in the image encoder of the SAM model. This problem is mainly reflected in that the shallow network can effectively capture local edge features. However, as the network depth increases, due to the role of the global attention mechanism, the interactive prompt position signal in the deep network tends to gradually decay. This decay is particularly obvious in low-contrast images. When the target and the background have similar grayscales (such as tumors and normal tissues), the deep features are easily interfered by the background, resulting in blurred target boundaries or even complete missed segmentations. In extremely difficult hard samples (target-background similarity > 90%), this problem is even more serious, and the attention distribution at each level from the shallow layer is difficult to effectively focus on the prompt area. The fundamental reason for this phenomenon is that existing methods lack a cross-hierarchical Prompt constraint mechanism. The interactive prompt information only acts on the Decoder side and is not injected into the image encoder, unable to correct the inherent attention distribution characteristics of the Transformer network. At the same time, the traditional inter-level feature transmission completely relies on the self-attention mechanism and lacks the ability to adaptively strengthen the signals in key regions.

[0030] (2) The problem of lost edge geometric information. This problem stems from the fact that the deep network needs to learn a large amount of redundant background features to identify edge details, which directly leads to a significant decline in performance in the segmentation task of weak-edge targets (such as early-stage tumors). The fundamental reason for this phenomenon is that existing methods fail to fully utilize the edge features extracted by the shallow network to explicitly constrain the distribution of deep attention. At the same time, the fixed interactive prompt injection weights cannot adapt to the differences in feature requirements at different network levels. For example, the shallow layer requires stronger edge constraints while the deep layer requires more flexible guidance. The lack of this dynamic adaptability makes it difficult for the network to complete the understanding of high-level semantics while maintaining geometric details.

[0031] Current improved research based on SAM mainly focuses on three technical directions. In terms of Decoder optimization, typical methods such as MedSAM enhance the interactive prompt interaction by introducing cross-attention mechanisms in the Mask prediction head. However, such improvements are limited to the Decoder side and are difficult to solve the problem of feature dispersion inside the Encoder. The parameter fine-tuning scheme adjusts the model parameters by introducing an adaptation layer or low-rank weight update. Although it improves the model adaptability, it significantly increases the training burden. In the direction of architecture expansion, researchers try to enhance the model capabilities through three-dimensional convolution expansion or automatic parameter adjustment mechanisms. However, these methods often face the problem of balancing computational efficiency and segmentation accuracy.

[0032] These improvement methods generally have three key limitations: First, they lack an effective cross-hierarchical interactive prompt signal transmission mechanism; second, they fail to establish a reliable cross-layer alignment path for geometric features; finally, they have insufficient adaptability to static prompts. Especially when dealing with low-contrast regions or tiny lesions in medical images, the performance of existing solutions is still unstable.

[0033] In view of the above defects and technical problems, the present application provides an image segmentation method based on dynamic Prompt constraint and edge anchor point guidance of the SAM model. The objectives of this method are as follows: 1. Establish a hierarchical Prompt constraint mechanism in the image encoder to solve the problem of attention diffusion in deep networks; 2. Through the edge geometric anchor point alignment mechanism, achieve dual technical effects: on the one hand, use shallow edge features to constrain the deep attention distribution and slow down the problem of semantic dispersion; on the other hand, improve the training efficiency through gradient control in the anchor point area.

[0034] 3. Adopt a lightweight dynamic weight design to achieve precise adjustment of the Prompt intensity at each level with only a small increase in parameters.

[0035] Refer to Figure 2 , which shows a schematic flowchart of an image segmentation method based on dynamic Prompt constraint and edge anchor point guidance of the SAM model according to an embodiment of the present application.

[0036] As Figure 2 shown, an image segmentation method based on dynamic Prompt constraint and edge anchor point guidance of the SAM model is presented, which specifically includes steps S100 to S600.

[0037] Refer to Figure 2 , in step S100, a lightweight dynamic adjustment network is designed and trained using the MLP structure.

[0038] It can be understood that in order to achieve efficient parameter utilization, a minimalist dynamic weight adjustment network is designed in this embodiment. This network adopts a lightweight structure design and adaptively generates the interactive prompt injection intensity coefficients at each level by analyzing the global information of the input features. With almost no increase in computational overhead, precise control of the hierarchical interactive prompt intensity is achieved, enabling the model to automatically adjust the constraint intensity at each level according to different input features.

[0039] In some feasible embodiments, based on the foregoing solution, a lightweight dynamic adjustment network is designed using the MLP structure, including: A lightweight dynamic adjustment network is designed using the lightweight architecture of MLP + multi-head output.

[0040] It should be noted that MLP refers to a multi-layer perceptron, which consists of an input layer, one or more hidden layers, and an output layer.

[0041] Exemplarily, the design process of the lightweight dynamic adjustment network is as follows: First, input the feature map (dimension C, with additional coordinate information, for example, 768 + n as shown in the figure) into the shared layer, and compress it to dimension D (recommended D = C / 3, such as 256; can be flexibly adjusted within the range of C / 4 to C / 2) through Linear(C→D)→GELU→LayerNorm; subsequently, configure a lightweight Head for each stage s: Linear(D→2)→GELU, map the D-dimensional feature to a binary vector [w_s, α_s], and activate them through the function Sigmoid(w_s) and the function Softplus(α_s) respectively to obtain the stage weight w_s∈(0,1) and the adjustment coefficient α_s>0. The number of parameters of the entire network mainly comes from C×D of the shared layer and 2D×S of S Heads (S is the number of stages). By selecting D≈C / 3 and restricting the number of stages, the total number of parameters can be controlled within 1.5C + 1, achieving extreme lightweighting.

[0042] Among them, Linear represents the fully connected layer, GELU represents the Gaussian Error Linear Unit, and LayerNorm represents layer normalization.

[0043] In some feasible embodiments, based on the foregoing solution, the process of training the lightweight dynamic adjustment network includes: Train the lightweight dynamic adjustment network using a three-stage progressive training scheme, including: In the first stage, freeze the backbone network and only train the modules related to interactive prompts; In the second stage, unfreeze the high-level network for adjustment; In the third stage, adjust the parameters of the entire model.

[0044] It should be noted that in the actual application process, the above training order is crucial for performance realization and cannot be adjusted, but the specific parameters of each stage can be appropriately adjusted.

[0045] For different application scenarios, the present invention provides the following optional implementation schemes: Computation-sensitive scenarios: The dynamic weight network can be omitted and fixed hierarchical weights can be used instead; Accuracy-priority scenarios: Non-maximum suppression can be added after edge detection to improve the boundary accuracy by 1.2%; Medical imaging scenarios: The heat map intensity (σ = 3) and edge detection threshold need to be specially adjusted; For the anchor-guided semantic boundary learning network, the present invention recommends the following alternating training scheme: In the 1st to Nth epochs: The anchor regions participate in backpropagation and other regions are frozen; In the (N + 1)th to 2Nth epochs: The entire map participates in backpropagation and normal training is resumed; It can be repeated for several rounds (Anchor ⇄ Full) for alternating refinement.

[0046] Continue to refer to Figure 2 , step S200, using the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder.

[0047] It can be understood that in this embodiment, the interactive prompt is innovatively injected into the image feature encoder to solve problems such as deep semantic diffusion and target blur when the SAM model processes medical images, and improve the model's attention and feature focusing ability on the prompt area.

[0048] In some feasible embodiments, based on the foregoing solution, the using the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder includes: Input the interactive prompt provided by the user and the feature maps output by different layers in the image encoder into the lightweight dynamic adjustment network for processing, and generate a set of fusion weights and diffusion coefficients that adaptively change with the layer.

[0049] Exemplarily, the specific process of the lightweight dynamic adjustment network processing the interactive prompt and the feature maps output by different layers in the image encoder is as follows: First, splice the image features representing the feature maps output by different layers with the prompt features representing the interactive prompt and input them into a shared MLP module; extract global features through a linear layer, GELU activation, and LayerNorm, and then output them to four Stage Heads respectively; each Stage Head further predicts the Prompt injection intensity coefficients (weight W and sigma coefficient σ) corresponding to the levels, where W controls the injection intensity through Sigmoid, and σ ensures non-negativity through Softplus, so as to achieve fine adjustment of the Prompt injection weights for each layer, enhancing the dynamic adaptability of the guiding ability while maintaining computational efficiency.

[0050] It should be noted that different from the traditional method that only acts on the Decoder, this solution uses a dynamic weight control mechanism to enable the shallow network to obtain stronger Prompt guidance to capture edge details, while the deep network maintains moderate constraints to retain semantic flexibility, effectively solving the problem of deep attention dispersion. This mechanism significantly improves the model's continuous attention ability to the prompt area on the premise of keeping the backbone network structure unchanged.

[0051] Continue to refer to Figure 2 , step S300, training the image encoder based on the edge anchor point guiding mechanism.

[0052] It should be noted that the edge anchor guidance mechanism means that the edge response map extracted from the shallow features extracts key geometric anchors through a lightweight depthwise separable convolution module and generates an edge anchor map; this map is sent into the deep feature map as a cross-layer guidance signal to explicitly modulate it, guiding the deep attention to focus on the key edge regions, thereby enhancing the perception ability of weak boundary targets. Through this mechanism, while maintaining the semantic expression ability, the model effectively avoids the overfitting of the deep layer to background redundancy and realizes the efficient fusion of edge perception and semantic expression.

[0053] In some feasible embodiments, based on the foregoing solution, training the image encoder based on the edge anchor guidance mechanism includes: Using depthwise separable convolution to extract edge responses from the shallow features of the image encoder, and generating an anchor mask in cooperation with a batch normalization layer and a Gaussian error linear unit activation; Establishing a dual alignment mechanism, using the anchor mask to shield the gradients of non-anchor regions in the backpropagation in the spatial dimension, and only updating the weights corresponding to the anchors, and calculating the L2 alignment loss function between the deep features and the anchor features in the semantic dimension; Adopting an alternating training strategy, only the anchor regions participate in the gradient update in odd training epochs, and the whole image resumes normal training in even training epochs; During the entire training process, the interaction between the shallow and deep features in the image encoder is dynamically balanced through learnable gating parameters.

[0054] It can be understood that for the problem of geometric information loss, the present invention proposes to extract edge anchors from shallow features to guide the deep network. Specifically, high-response geometric key points in the shallow features are identified through an edge detection algorithm, and a cross-level attention distribution constraint is established, so that the deep network can maintain sensitivity to key edge regions. This constraint method based on geometric priors not only reduces the learning requirements of the deep network for redundant background features, but also significantly improves the segmentation accuracy of weak edge targets.

[0055] It should be noted that the edge anchor guidance mechanism is only enabled in the model training stage and does not require additional calculations in the inference stage. However, because it optimizes the structural perception ability of deep semantics during training, the model can also maintain good boundary segmentation effects during inference, and is particularly suitable for the precise segmentation of weak boundary and low-contrast targets (such as lesions or tumors) in medical images.

[0056] Continue to refer to Figure 2 , step S400, input the result of the introduction of the interactive prompt and the image provided by the user into the trained image encoder, and process to obtain an image embedding vector.

[0057] In some feasible embodiments, based on the foregoing solution, inputting the introduction result of the interactive prompt and the image provided by the user into the trained image encoder to process and obtain an image embedding vector includes: Encoding the interactive prompt provided by the user into a Gaussian heatmap; Combining the fusion weight, the diffusion coefficient and the Gaussian heatmap, and inputting the combination result and the image provided by the user into the image encoder to process and obtain an image embedding vector.

[0058] Continue to refer to Figure 2 , step S500, inputting the interactive prompt provided by the user into the prompt encoder to process and obtain a prompt embedding vector.

[0059] Continue to refer to Figure 2 , step S600, inputting the image embedding vector and the prompt embedding vector into the mask decoder for interaction, and obtaining a segmented mask image after multiple layers of cross-attention and upsampling.

[0060] In summary, while inheriting the efficient interaction characteristics of the image segmentation method based on the SAM model, the technical solution of this application effectively solves the problems of fuzzy segmentation boundaries and inconsistent targets under traditional weak annotation conditions, and demonstrates extremely strong medical image segmentation capabilities and generalization.

[0061] This method shows significant technical advantages compared with the prior art. Experimental data shows that: on the ISIC skin lesion dataset, the Dice coefficient of this method reaches 0.74 when using point prompts, an 8.8% improvement compared to the baseline SAM (0.68); it reaches 0.89 when using box prompts, a 18.7% improvement compared to SAM (0.75). In the Kvasir polyp segmentation task, the Dice values under point prompts and box prompts reach 0.86 and 0.94 respectively, a 4.9% and 2.2% improvement compared to SAM. For BUSI breast ultrasound image segmentation, this method obtains Dice values of 0.85 and 0.89 respectively under point prompts and box prompts, a 6.3% and 3.5% improvement compared to SAM. In the REFUGE fundus image segmentation task, the Dice values under point prompts and box prompts reach 0.59 and 0.88 respectively, a 6.1% and 13.2% improvement compared to SAM.

[0062] These improvements stem from three core technologies: the hierarchical Prompt injection mechanism solves the problem of deep attention dispersion, enabling the model to maintain stable segmentation performance in low-contrast medical images; the edge geometric anchor constraint significantly improves the boundary localization accuracy while reducing the parameter update amount by 18% through gradient mask control; the lightweight dynamic adjustment network only increases the parameter amount by 0.5%. Especially for difficult samples with high target-background similarity, this method shows stronger robustness, providing a more reliable solution for medical image analysis.

[0063] The device embodiments of the present application are described below, which can be used to execute a dynamic Prompt constraint and edge anchor-guided image segmentation method based on the SAM model in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application above.

[0064] Refer to Figure 3 As shown, an image segmentation device 300 based on a dynamic Prompt constraint and edge anchor guidance of the SAM model according to an embodiment of the present application includes: A first generation unit 301, configured to design and train a lightweight dynamic adjustment network using an MLP structure; A fusion unit 302, configured to use the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder; A training unit 303, configured to train the image encoder based on an edge anchor guidance mechanism; A second generation unit 304, configured to input the introduction result of the interactive prompt and the image provided by the user into the trained image encoder, and process to obtain an image embedding vector; A third generation unit 305, configured to input the interactive prompt provided by the user into the prompt encoder to process and obtain a prompt embedding vector; A segmentation processing unit 306, configured to input the image embedding vector and the prompt embedding vector into the mask decoder for interaction, and obtain a segmented mask image after multiple layers of cross-attention and upsampling.

[0065] As Figure 4 shown, an embodiment of the present application further provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored on the memory 410 and executable on the processor. When the processor 420 executes the computer program 411, the steps of the above-mentioned dynamic Prompt constraint and edge anchor-guided image segmentation method based on the SAM model are implemented.

[0066] Since the electronic device introduced in this embodiment is the device adopted by an image segmentation device based on SAM model dynamic Prompt constraint and edge anchor point guidance in the embodiments of the present application, based on the methods introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the methods in the embodiments of the present application will not be described in detail here. As long as the devices adopted by those skilled in the art to implement the methods in the embodiments of the present application belong to the scope protected by the present application.

[0067] In the specific implementation process, when the computer program 411 is executed by the processor, it can implement any implementation manner in the corresponding embodiments of the first aspect.

[0068] Figure 5 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0069] It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present application.

[0070] As Figure 5 shown, the computer system 500 includes a central processing unit 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory 502 or the program loaded from the storage section 508 into the random access memory 503, such as executing the methods described in the above embodiments. In the random access memory 503, various programs and data required for system operation are also stored. The central processing unit 501, the read-only memory 502, and the random access memory 503 are connected to each other through a bus 504. The input / output interface 505 is also connected to the bus 504.

[0071] The following components are connected to the input / output interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as required. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 510 as required so that a computer program read from it can be installed into the storage section 508 as required.

[0072] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, various functions defined in the system of the present application are performed.

[0073] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the above module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0075] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the units themselves.

[0076] As another aspect, the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes a method for image segmentation based on dynamic Prompt constraint and edge anchor point guidance in the above embodiments.

[0077] As another aspect, the present application also provides a computer-readable medium. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by an electronic device, the electronic device implements a method for image segmentation based on dynamic Prompt constraint and edge anchor point guidance in the above embodiments.

[0078] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0079] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0080] Other embodiments of the present application will be readily contemplated by those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An image segmentation method based on SAM model with dynamic Prompt constraint and edge anchor point guidance, wherein, The SAM model is provided with an image encoder, a prompt encoder, and a mask decoder, and is characterized in that the method includes: Design and train a lightweight dynamic adjustment network using an MLP structure; Use the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder; Train the image encoder based on the edge anchor point guidance mechanism; Input the introduction result of the interactive prompt and the image provided by the user into the trained image encoder, and process to obtain an image embedding vector; Input the interactive prompt provided by the user into the prompt encoder to process and obtain a prompt embedding vector; Input the image embedding vector and the prompt embedding vector into the mask decoder for interaction, and obtain a segmented mask image after multi-layer cross-attention and upsampling; 2. The method according to claim 1, characterized in that, Design a lightweight dynamic adjustment network using an MLP structure, including: Design a lightweight dynamic adjustment network using a lightweight architecture of MLP + multi-head output; 3. The method according to claim 2, wherein The process of training the lightweight dynamic adjustment network includes: Train the lightweight dynamic adjustment network using a three-stage progressive training scheme, including: In the first stage, freeze the backbone network and only train the modules related to the interactive prompt; In the second stage, unfreeze the high-level network for adjustment; In the third stage, adjust the parameters of the entire model; 4. The method according to claim 1, wherein The using the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder includes: Input the interactive prompt provided by the user and the feature maps output by different layers in the image encoder into the lightweight dynamic adjustment network for processing, and generate a set of fusion weights and diffusion coefficients that adaptively change with the hierarchy; 5. The method according to claim 1, wherein The training the image encoder based on the edge anchor point guidance mechanism includes: Use depthwise separable convolution to extract edge responses from the shallow features of the image encoder, and cooperate with the batch normalization layer and the Gaussian error linear unit activation to generate an anchor mask; Establish a dual alignment mechanism. In the spatial dimension, use the anchor mask to shield the gradients of non-anchor regions in backpropagation, and only update the weights corresponding to the anchors. In the semantic dimension, calculate the L2 alignment loss function between the deep features and the anchor features; Adopt an alternating training strategy. Only the anchor regions participate in gradient update in odd training cycles, and the full image resumes normal training in even training cycles; During the entire training process, dynamically balance the interaction between the shallow and deep layer features in the image encoder through learnable gating parameters; 6. The method according to claim 4, characterized in that, The inputting the introduction result of the interactive prompt and the image provided by the user into the trained image encoder to process and obtain an image embedding vector includes: Encode the interactive prompt provided by the user into a Gaussian heatmap; Combine the fusion weights, the diffusion coefficients, and the Gaussian heatmap, and input the combination result and the image provided by the user into the image encoder to process and obtain an image embedding vector; 7. An image segmentation device based on dynamic Prompt constraint and edge anchor point guidance of the SAM model, characterized in that, Includes: A first generation unit for designing and training a lightweight dynamic adjustment network using an MLP structure; A fusion unit for using the lightweight dynamic adjustment network to control the intensity and spatial range of the interactive prompt provided by the user introduced into the image encoder; A training unit for training the image encoder based on an edge anchor point guidance mechanism; A second generation unit for inputting the introduction result of the interactive prompt and the image provided by the user into the trained image encoder to process and obtain an image embedding vector; A third generation unit for inputting the interactive prompt provided by the user into the prompt encoder to process and obtain a prompt embedding vector; A segmentation processing unit for inputting the image embedding vector and the prompt embedding vector into the mask decoder for interaction, and obtaining a segmented mask image after multiple layers of cross-attention and upsampling.

8. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1-6.

9. An electronic device, characterized in that, Comprising: A memory and a processor; The memory for storing computer instructions; The processor for calling the computer instructions stored in the memory, causing the electronic device to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • SAM-based cross-modal domain generalization medical image segmentation method

    CN117808834A

  • Target instance segmentation model establishment method based on automatic prompt learning and application thereof

    CN118279320A

  • Self-adaptive medical image segmentation method and device based on prompt learning, and medium

    CN118521595A

  • Dynamic decision image segmentation method based on self-prompt guidance

    CN119762787A