Human in-loop and visual basis model-based few-sample medical image segmentation method

Through the visual basic model and expert feedback interaction mechanism, the mask prompt is selected using data augmentation and perceived similarity, and combined with attention mechanism, medical image segmentation is solved, and the problems of insufficient generalization ability and high labeling cost in the existing technology are achieved, and efficient and accurate medical image segmentation is achieved.

CN120298698APending Publication Date: 2025-07-11EAST CHINA NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510440551.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing medical image segmentation methods are insufficient in generalization capabilities in clinical applications and cannot effectively integrate expert knowledge, which makes it difficult to meet clinical needs and is expensive to label.

Method used

The visual basic model and expert feedback interaction mechanism are adopted to generate diversified support images through data augmentation, mask prompts are selected using perceptual similarity, and preliminary segmentation is performed in combination with attention mechanism, and segmentation results are iteratively optimized through human-computer interaction correction.

Benefits of technology

It improves the accuracy and generalization performance of medical image segmentation, reduces dependence on labeled data, and improves the degree of automation and clinical applicability of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298698A_ABST
    Figure CN120298698A_ABST
Patent Text Reader

Abstract

The invention discloses a few-sample medical image segmentation method based on a human-in-loop and visual basis model. The method is characterized by comprising the following steps: firstly, carrying out diversified data enhancement on a single annotated image to construct a rich support set; then automatically selecting an optimal support image as a mask prompt for each image to be segmented through a dynamic matching algorithm, and driving the visual basic model SAM2 to realize automatic preliminary segmentation; and then, correcting an initial segmentation result by utilizing expert interaction feedback, generating a mask prompt enhancement signal, and iteratively optimizing the segmentation result of the whole sequence through a mask prompt mechanism of a visual basic model. Compared with the prior art, the method has the advantages that expert knowledge is fully utilized, efficient cooperation of automatic segmentation and expert correction is achieved, the medical image segmentation precision is remarkably improved, the problems of blurred target areas, unclear boundaries and the like in the medical images are effectively solved, the automation degree of the medical image analysis process is greatly improved, and the medical image segmentation efficiency is improved. Good clinical application prospects are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and more specifically, to a few-shot medical image segmentation method based on human-in-the-loop technology and vision foundation models. Background Art

[0002] Medical image segmentation is a key link in medical image analysis and is widely used in clinical diagnosis, surgical planning, radiotherapy planning, disease monitoring, and treatment effect evaluation. Traditional medical image segmentation methods usually train deep learning models with a large amount of manually annotated data to achieve accurate automatic segmentation. However, in actual clinical scenarios, the medical image annotation process is extremely complex, usually consuming a large amount of human resources and time costs, and requiring high professional skills of annotators, making the acquisition cost of annotated data extremely high.

[0003] In recent years, to address the problems of scarce and costly data annotation, few-shot learning has been gradually introduced into the field of medical image segmentation. This method attempts to use a very small amount of annotated data to achieve accurate segmentation of medical images. However, most existing few-shot learning methods rely on techniques such as meta-learning, which often have insufficient generalization ability when dealing with the inherent multimodality, complexity of lesion regions, and strong background interference of medical images, and the segmentation accuracy is difficult to meet clinical needs. In addition, medical images themselves usually have characteristics such as high dimensionality, complex structure, and blurred object boundaries, which make traditional few-shot methods perform poorly in practical applications.

[0004] Vision foundation models (such as the Segment Anything Model, SAM and its upgraded version SAM2) show strong generalization ability in segmentation tasks in the field of natural images. Nevertheless, when directly applying vision foundation models to medical image segmentation, challenges specific to medical images are still faced, including blurred boundaries of target regions, diverse internal features of lesion regions, high similarity between tissues and organs, and obvious differences in image features caused by different imaging devices, resulting in the segmentation accuracy relying solely on vision foundation models being difficult to meet actual clinical needs.

[0005] In summary, the medical image segmentation in the prior art has not effectively combined the rich domain knowledge of clinical experts with advanced automated algorithms, lacking an efficient human-computer interaction feedback mechanism to utilize experts' experience for result correction and model feedback optimization, resulting in the limitation of automated segmentation algorithms in actual clinical applications. There are problems such as insufficient generalization performance and inability to fully integrate expert knowledge. Therefore, there is an urgent need for a few-shot medical image segmentation method that can comprehensively utilize clinical expert feedback information and the advantages of vision-based models to achieve more accurate, robust, and practical medical image segmentation effects and improve the accuracy and efficiency of clinical diagnosis and treatment. Summary of the Invention

[0006] The object of the present invention is to provide a few-shot medical image segmentation method based on human-in-the-loop and vision-based models in view of the deficiencies of the prior art. The automatic segmentation ability of the vision-based model and the interaction mechanism of expert feedback work together to achieve accurate segmentation of medical images. This method generates a diverse support set by performing extended data augmentation on a small number of labeled support images, and uses a perceptual similarity algorithm to dynamically match the query image and the augmented support images, thereby obtaining the most suitable support image and its mask hint. Subsequently, the hint feature embedding and the query image feature are jointly input into the vision-based model, and the features are fused through an attention mechanism to achieve automatic preliminary segmentation. The expert uses the interaction interface to quickly and accurately perform local annotation correction on the incorrect or unclear regions in the preliminary segmentation result. The interaction interface automatically records the expert feedback information and generates an enhanced mask hint. The vision-based model further iteratively optimizes the segmentation result based on the enhanced mask hint through the built-in video segmentation module, realizing continuous propagation and update of the query image sequence. This method effectively reduces the dependence on a large amount of labeled data through the generalization ability of the vision-based model, and at the same time efficiently introduces expert domain knowledge by means of the "human-in-the-loop" technology, greatly improving the segmentation accuracy, being better able to adapt to the complex scenarios of clinical practice, being simple and practical, easy to implement, and having broad application prospects.

[0007] The object of the present invention is achieved as follows: A few-shot medical image segmentation method based on human-in-the-loop and vision foundation model, characterized in that the method uses data augmentation to expand the data volume of a single labeled support image, and then automatically selects the most matching augmented support image and mask based on Learned Perceptual Image Patch Similarity (LPIPS) as the mask prompt to input into the vision foundation model. Then, the vision foundation model extracts the feature embedding of the mask prompt through the image encoder, and then fuses the mask prompt feature with the query image feature through the memory attention mechanism in the built-in video segmentation module to achieve automatic preliminary segmentation. Subsequently, through human-computer interaction, experts quickly correct the local problem areas in the preliminary segmentation to automatically generate an enhanced mask prompt after feedback. Finally, the video segmentation module of the vision foundation model is used again to propagate the feedback mask prompt through the memory attention mechanism to achieve iterative optimization of the automatic segmentation result until the clinical requirements are met.

[0008] Specifically, the medical image segmentation of the present invention is carried out according to the following steps: (1) Perform data augmentation on a single labeled support image and its mask to form an extended set of support images, specifically including: 1.1: Perform affine transformation on a single support image and its corresponding labeled mask simultaneously, including rotation, scaling, translation or shearing, to form a series of spatially morphologically enhanced support images and corresponding masks; 1.2: Further perform color jitter transformation on the support image after affine transformation, and the color transformation includes random adjustment of image brightness, contrast, saturation or hue to obtain richer and more diverse image visual features.

[0009] (2) Dynamically select the support image that best matches the query image to be segmented from the enhanced set of support images based on perceptual similarity, specifically including: 2.1: Use Learned Perceptual Image Patch Similarity (LPIPS) to measure the visual perceptual similarity between the enhanced support image and the query image; 2.2: Calculate the LPIPS distance between each query image and all enhanced support images, and automatically select the support image that is most similar to the current query image as the best prompt image.

[0010] (3) Use the selected support image and its corresponding mask as the mask prompt to input into the vision foundation model to automatically obtain a preliminary segmentation result, specifically including: 3.1: Use the vision foundation model to encode the prompt composed of the selected support image and its mask to obtain the feature embedding representation of the mask prompt; 3.2: Based on the above-mentioned mask prompt feature embedding, use the video segmentation propagation mechanism built into the vision-based model to automatically perform preliminary segmentation on the query image to obtain the initial segmentation mask result.

[0011] 3.2.1: The video segmentation module of the vision-based model uses the memory attention mechanism to establish the feature correlation between the support image mask prompt and the current query image; 3.2.2: Based on the established cross-image feature correlation relationship, automatically generate the preliminary segmentation mask result.

[0012] (4) Perform local feedback correction on the preliminary segmentation result through a human-computer interaction method, and generate an enhanced mask prompt after expert feedback, specifically including: 4.1: The expert marks and corrects the local areas with errors or ambiguities in the preliminary segmentation result through the human-computer interaction interface; 4.2: The human-computer interaction interface automatically records the expert's feedback operations and generates the corresponding enhanced mask prompt for subsequent segmentation optimization steps.

[0013] (5) Use the enhanced mask prompt of the expert feedback to drive the vision-based model again, and propagate and update the segmentation result with the video segmentation mechanism, specifically including: 5.1: Input the enhanced mask prompt of the expert feedback and the current query image into the video segmentation module of the vision-based model again; 5.2: Use the video segmentation mechanism of the vision-based model to automatically propagate and update the segmentation result of the query image according to the enhanced mask prompt, and realize the iterative optimization of the segmentation result.

[0014] (6) Repeat steps (3) to (4) until the segmentation accuracy reaches the preset threshold or meets the clinical requirements.

[0015] The present invention has the following beneficial technical effects and significant technical progress compared with the prior art: 1) The present invention uses expert knowledge to efficiently correct the local segmentation result, effectively solves the problem that the existing few-shot medical image segmentation methods fail to fully utilize expert knowledge, and realizes the efficient integration of automatic segmentation and expert knowledge; 2) The dynamic support image selection method based on visual perception similarity of the present invention significantly improves the generalization performance of the vision-based model in the medical image scenario, can effectively cope with the challenges of blurred boundaries and complex structures in the lesion area of medical images, and obtains more accurate and reliable segmentation results.

[0016] 3) Through the "human-in-the-loop" expert feedback mechanism, the present invention effectively overcomes problems such as blurred target regions and unclear boundaries in medical images, makes full use of expert knowledge to achieve efficient collaboration between automatic segmentation and expert correction, significantly improves the accuracy of medical image segmentation, reduces the burden of manual annotation, and enhances the automation of the medical imaging analysis process.

[0017] 4) The method is efficient, reliable, easy to implement, and has good clinical application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the present invention; Figure 2 is a schematic diagram of the specific operation of Example 1. DETAILED DESCRIPTION OF THE INVENTION

[0019] The present invention first performs extended data augmentation on a single-labeled support image to form a rich set of support images; secondly, dynamically selects the enhanced support image and its mask that best match the query image to be segmented through a learned perceptual image patch similarity algorithm, and inputs the mask hint into the vision-based model. Then, the vision-based model generates feature embeddings of the mask hint through an image encoder, and realizes automatic propagation and preliminary segmentation of the query image sequence through the memory attention mechanism in its built-in video segmentation module. Then, through the expert human-computer interaction interface, the incorrect or blurred regions in the preliminary segmentation result are quickly locally corrected, and an enhanced mask hint after feedback is automatically generated. Finally, the enhanced mask hint is input again into the video segmentation module of the vision-based model, and the memory attention mechanism is used to further propagate and optimize the segmentation result to achieve iterative automatic segmentation optimization until the clinical requirements are met.

[0020] Refer to Figure 1 , the medical image segmentation of the present invention is carried out according to the following steps: (1) Perform data augmentation on a single labeled support image and its mask to form an extended set of support images, specifically including: 1.1: Simultaneously perform affine transformations on a single support image and its corresponding labeled mask, including rotation, scaling, translation, or shear, to form a series of spatially morphologically enhanced support images and corresponding masks; 1.2: Further perform color jitter transformation on the support image after affine transformation, and the color transformation includes random adjustment of image brightness, contrast, saturation, or hue to obtain more diverse image visual features.

[0021] (2) Based on perceptual similarity, automatically dynamically select the support image that best matches the query image to be segmented from the enhanced set of support images, specifically including: 2.1: Measure the visual perceptual similarity between the enhanced support image and the query image using Learned Perceptual Image Patch Similarity (LPIPS). 2.2: Calculate the LPIPS distance between each query image and all enhanced support images, and automatically select the support image that is most similar to the current query image as the best hint image.

[0022] (III) Use the selected support image and its corresponding mask as mask hints, input them into the visual base model, and automatically obtain the preliminary segmentation result, specifically including: 3.1: Encode the hint composed of the selected support image and its mask using the visual base model to obtain the feature embedding representation of the mask hint. 3.2: Based on the mask hint feature embedding, use the video segmentation propagation mechanism built into the visual base model to automatically perform preliminary segmentation on the query image to obtain the initial segmentation mask result.

[0023] The specific steps of 3.2 include: 3.2.1: The video segmentation module of the visual base model uses the memory attention mechanism to establish the feature correlation between the support image mask hint and the current query image. 3.2.2: Based on the established cross-image feature correlation relationship, automatically generate the preliminary segmentation mask result.

[0024] (IV) Perform local feedback correction on the preliminary segmentation result through a human-computer interaction method, and generate an enhanced mask hint after expert feedback, specifically including: 4.1: The expert marks and corrects the local areas with errors or ambiguities in the preliminary segmentation result through the human-computer interaction interface. 4.2: The human-computer interaction interface automatically records the expert's feedback operations and generates the corresponding enhanced mask hint for subsequent segmentation optimization steps.

[0025] (V) Use the enhanced mask hint of the expert feedback to drive the visual base model again, and propagate and update the segmentation result through the video segmentation mechanism, specifically including: 5.1: Input the enhanced mask hint of the expert feedback and the current query image into the video segmentation module of the visual base model again. 5.2: Use the video segmentation mechanism of the visual base model to automatically propagate and update the segmentation result of the query image according to the enhanced mask hint, and realize the iterative optimization of the segmentation result.

[0026] (VI) Repeat steps (III) - (IV) until the segmentation accuracy reaches the preset threshold or meets the clinical requirements.

[0027] The following takes a specific example of liver region segmentation in medical images to further illustrate the present invention. Embodiment

[0028] Refer to Figure 2 , this embodiment performs few-shot medical image segmentation according to the following steps: 1) Select a single annotated liver CT support image and mask, and use data augmentation methods such as affine transformation and color jitter transformation to expand the support image set.

[0029] 2) Automatically select the best support image using the LPIPS metric to match with the query CT image to be segmented.

[0030] 3) Input the matched support image and mask prompt into the vision-based model SAM2, and obtain the preliminary automatic segmentation result through the video segmentation module of the model.

[0031] 4) The expert makes feedback corrections on the boundary-blurred or incorrect regions in the initial segmentation through the human-computer interface, and automatically generates an enhanced mask prompt.

[0032] 5) Input the enhanced mask prompt and the query image sequence into the vision-based model again, and through the memory attention mechanism built in the video segmentation module, automatically propagate and optimize the segmentation result until the liver segmentation accuracy meets the clinical requirements.

[0033] The above is only a preferred embodiment to further illustrate the present invention, and is not intended to limit the patent of the present invention. All equivalent implementations of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A few-shot medical image segmentation method based on human-in-the-loop and vision foundation models, characterized in that, The method specifically includes the following steps: 1) Perform data augmentation on a single annotated support image and its mask to form an extended set of support images; 2) Dynamically select the support image that best matches the query image to be segmented from the set of support images based on perceptual similarity; 3) Use the selected support image and its corresponding mask as mask prompts and input them into the visual base model to automatically obtain a preliminary segmentation result; 4) Perform local feedback correction on the preliminary segmentation result through human-computer interaction and generate an enhanced mask prompt after expert feedback; 5) Use the expert feedback enhanced mask prompt to drive the visual base model again to propagate the updated segmentation result through the video segmentation mechanism; 6) Repeat steps 3) to 4) above until the segmentation accuracy reaches a preset threshold or meets clinical requirements.

2. The few-shot medical image segmentation method based on human-in-the-loop and vision foundation models according to claim 1, wherein Step 1) specifically includes: 1.1: Simultaneously perform affine transformations including rotation, scaling, translation, or shearing on a single support image and its corresponding annotated mask to form a series of spatially morphologically enhanced support images and corresponding masks; 1.2: Independently perform color jitter transformation on the support image after affine transformation, and the color transformation includes: randomly adjusting the image brightness, contrast, saturation, or hue to obtain more diverse image visual features.

3. The few-shot medical image segmentation method based on human-in-the-loop and vision foundation models according to claim 1, wherein Step 2) specifically includes: 2.1: Utilize the visually perceived similarity between the support image enhanced by learning to perceive image patch similarity metric and the query image; 2.2: Calculate the LPIPS distance between each query image and all enhanced support images, and automatically select the support image that is most similar to the current query image as the best prompt image.

4. The few-shot medical image segmentation method based on human-in-the-loop and vision foundation model according to claim 1, wherein Step 3) specifically includes: 3.1: Use the visual base model to encode the prompt composed of the selected support image and its mask to obtain a feature embedding representation of the mask prompt; 3.2: Based on the mask prompt feature embedding, adopt the video segmentation propagation mechanism built into the visual base model to automatically perform a preliminary segmentation on the query image to obtain an initial segmentation mask result.

5. The few-shot medical image segmentation method based on human-in-the-loop and vision foundation model according to claim 1, wherein Step 4) specifically includes: 4.1: The expert annotates and corrects the local areas with errors or ambiguities in the preliminary segmentation result through the human-computer interface; 4.2: The human-computer interface automatically records the expert's feedback operations and generates corresponding enhanced mask prompts.

6. The few-shot medical image segmentation method based on human-in-the-loop and vision foundation model according to claim 1, characterized in that Step 5) specifically includes: 5.1: Input the mask prompt enhanced by expert feedback and the current query image into the video segmentation module of the visual base model again; 5.2: Utilize the video segmentation mechanism of the visual base model to automatically propagate the enhanced mask prompt and update the segmentation result of the query image to achieve iterative optimization of the segmentation result.

7. The few-shot medical image segmentation method based on human-in-the-loop and vision foundation model according to claim 4, wherein Step 3.2 specifically includes: 3.2.1: The video segmentation module of the visual base model uses the memory attention mechanism to establish a feature association between the support image mask prompt and the current query image; 3.2.2: Based on the established cross-image feature association relationship, automatically generate a preliminary segmentation mask result.

Citation Information

Cited By

  • Interactive medical image segmentation method and system based on robust sequence prompt optimization

    CN120894387A

  • Interactive medical image segmentation method and system based on robust sequence prompt optimization

    CN120894387B