Medical image segmentation method based on SAM large model domain generalization visual cue tuning

By constructing a medical image segmentation model based on SAM large model, using divide-perception correction and detail-aware prompt learning, the stability and accuracy problems of SAM in medical image segmentation tasks are solved, and a more efficient medical image segmentation effect is achieved.

CN120339623APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510491433.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing SAM large models cannot achieve zero-sample segmentation stably and accurately on multimodal, multi-objective medical datasets, especially for objects with complex shapes, small sizes or low contrast, and perform unstable in medical image segmentation tasks.

Method used

A medical image segmentation model based on SAM large model is built, including image preprocessing module, multi-layer Transformer, boundary refinement module and SAM decoder. Through the gap-aware correction strategy and detailed-aware prompt learning, enhance feature extraction, and use pre-trained natural image datasets and knowledge of target medical scenes to correct image features, and combine the boundary refinement module to ensure the sharp and smooth edges of the segmented object.

Benefits of technology

The stability and accuracy of SAM models in medical image segmentation tasks are improved, especially in the segmentation effect of complex shapes and low-contrast objects, which improves the Dice index by more than 20%, and reduces the amount of training parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339623A_ABST
    Figure CN120339623A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image segmentation, and particularly relates to a medical image segmentation method based on SAM large model domain generalization visual cue tuning, which comprises the following steps: constructing a medical image segmentation model based on an SAM large model, and performing medical image segmentation through the medical image segmentation model; according to the method, the current image features are improved and corrected by utilizing knowledge learned from a previous pre-trained natural image data set and a target medical scene through a gap perception correction strategy, so that when a scene related to a medical image segmentation task is processed, the difference between a medical image and a pre-trained natural image can be effectively corrected; better performance is ensured; by introducing detailed task characteristic prompts, the model is helped to focus on a target area in local feature extraction, and detailed information related to tasks is deeply mined; the feature-level edge information is enhanced by introducing boundary refinement, so that the sharp and smooth edge of the segmented object is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and particularly relates to a medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model. Background Art

[0002] The Segment Anything Model (SAM), a model for segmenting anything, was proposed as an innovative foundation model for image segmentation. SAM is based on the Vision Transformer (ViT) model and is trained on a large dataset of 11 million images containing 1 billion masks. The biggest highlight of SAM is its good zero-shot segmentation performance for unknown datasets and tasks. This process is driven by different prompts, such as points and boxes, which are used to indicate the pixel-level semantics and region-level positions of the target object. It has been proven to be highly versatile and capable of handling a wide range of segmentation tasks.

[0003] However, SAM cannot stably and accurately achieve zero-shot segmentation on multi-modal and multi-object medical datasets. Secondly, different attributes of medical objects may affect the object perception ability of SAM. In particular, for objects with complex shapes / boundaries, small sizes, or low contrasts, SAM may output poor results. We believe that although SAM has the potential to become a general medical image segmentation model, its performance in medical image segmentation tasks is currently unstable. Therefore, future research should study how to effectively use a small amount of medical images to fine-tune SAM to improve its reliability. In addition, exploring the segmentation performance of SAM for three-dimensional volume data is also an interesting direction.

[0004] Currently, the most common method to adapt general large models to downstream tasks is to freeze the backbone network and add a part of the adapter to adapt to the current downstream task. Although SAM performs poorly in many medical image segmentation tasks, it has also demonstrated its potential through fine-tuning techniques. Different from directly fine-tuning a large number of parameters of the entire SAM, another flexible method is to fine-tune a part of the parameters of SAM or utilize the parameter-efficient fine-tuning (PEFT) technique to transfer the pre-trained SAM to a specific medical image segmentation task.

[0005] Segmentation is an important task in medical image processing, and its purpose is to perform quantitative analysis on organs and other substructures to make the changes in anatomical or pathological structures in the image clearer. Medical image segmentation plays a key role in intelligent medical fields such as computer-aided diagnosis, pathological analysis, and surgical planning. The success of deep learning has greatly promoted the development of medical image segmentation.

[0006] Today, the field of artificial intelligence research is undergoing a drastic transformation. The emergence of large-scale models such as SAM and SegGPT has made it possible for researchers to solve multiple types of problems within a unified framework. Deploying such large-scale models is more promising for industrial use due to their remarkable generalization ability. In the field of medical image segmentation, if large-scale CV models, such as SAM or SegGPT, can achieve high competitiveness, there is no need to deploy a single medical image segmentation model. Instead, the solution for medical image segmentation can be directly integrated into the large-scale CV model. Moreover, without the network engineering burden of medical image segmentation, the deployment and storage costs of a single medical image segmentation model can be greatly saved. Summary of the Invention

[0007] To solve the above technical problems, the present invention provides a medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model, including:

[0008] Construct a medical image segmentation model based on the SAM large model, and perform medical image segmentation through the medical image segmentation model;

[0009] The medical image segmentation model includes: an image preprocessing module, a multi-layer Transformer, a boundary refinement module, and a SAM decoder;

[0010] Performing medical image segmentation through the medical image segmentation model includes:

[0011] Step 1: The image preprocessing module normalizes the input image to obtain normalized image data;

[0012] Step 2: The normalized image data is input into the multi-layer Transformer for feature extraction. For the features extracted by each layer of the Transformer, feature enhancement is performed through a gulf perception correction strategy, and the enhanced features are introduced with detailed task-specific prompts to help the model focus more on the target area during local feature extraction and deeply mine task-related detailed information;

[0013] Step 3: The feature image extracted by the last layer of the Transformer is input into the boundary refinement module for target image segmentation masking to obtain the overall features of the image;

[0014] Step 4: The overall features of the image are input into the SAM decoder to convert the high-dimensional feature representation into a specific segmentation mask image, and the predicted segmentation map is output.

[0015] Advantages of the present invention:

[0016] The present invention uses a gap perception correction strategy to utilize the knowledge learned from a previous pre-trained natural image dataset and a target medical scenario to improve and correct the current image features, so as to effectively correct the gap between medical images and pre-trained natural images when dealing with scenarios related to medical image segmentation tasks, ensuring better performance; by introducing detailed task feature prompts, it helps the model to focus more on the target area in local feature extraction and dig deeper into task-related detailed information; by introducing boundary refinement, the feature-level edge information is strengthened, thus ensuring sharp and smooth edges of the segmented objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the overall principle block diagram of the present invention;

[0018] Figure 2 It is the structural diagram of the detailed perception prompt learning framework of the present invention;

[0019] Figure 3 It is the structural framework diagram of the boundary refinement module of the present invention;

[0020] Figure 4 It is the comparison test diagram of the present invention embodiment with other similar SAM algorithms on different datasets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] A medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model, as Figure 1 shown, includes:

[0023] Construct a medical image segmentation model based on the SAM large model, and perform medical image segmentation through the medical image segmentation model;

[0024] The medical image segmentation model includes: an image preprocessing module, a multi-layer Transformer, a boundary refinement module, and a SAM decoder;

[0025] Performing medical image segmentation through the medical image segmentation model includes:

[0026] Step 1, the image preprocessing module normalizes the input image to obtain normalized image data;

[0027] Step 2: Use the gap-aware calibration strategy for the normalized image data to more efficiently extract and utilize the shared features of cross-modal images, enabling the model to better adapt to medical tasks, such as Figure 2 as shown

[0028] Step 2.1: Input the image into the image encoder for processing and output the features of the image. After image feature extraction, the gap-aware calibration strategy utilizes the knowledge learned from the previously pre-trained natural image dataset and the target medical scenario to improve and calibrate the current image features, so as to effectively correct the gap between medical images and pre-trained natural images when dealing with scenarios related to medical image segmentation tasks, ensuring better performance.

[0029] Step 2.2: For the feature map F i generated by each layer of Transformer in the backbone network, the gap-aware calibration strategy corrects its features to enhance the feature performance. The specific implementation includes: 1. Perform feature mapping on the input image. 2. Combine the feature mapping with the gap-aware calibration strategy. 3. Update the feature map F i of each layer to enhance its expressive ability, obtaining the enhanced feature map F i ′

[0030] Step 2.3: Token serialization and Softmax processing. Transmit the calibrated feature knowledge to each Token sequence, where these Tokens represent different feature regions in the image, and adjust the feature map by calculating the similarity of each token. Then use the softmax function to further classify the relationship between each Token and the image Patch, enabling each patch to match the correct category, and helping the network adjust the allocation of each Token by calculating the similarity of each feature.

[0031] Step 2.4: Finally, the gap-aware calibration strategy further enhances the adjusted features through an MLP (Multi-Layer Perceptron), enabling the expressive ability of each layer of feature map to be enhanced, thereby obtaining the finally calibrated image features F i ′

[0032] The specific implementation of Step 2 is as follows: For the pre-processed pictures, input them into the model. After each layer of Transformer feature extraction in the model, the obtained features are passed through the gap-aware calibration strategy to obtain the features after knowledge calibration. Specifically: For the feature map generated by the i-th layer of the i-th Vision Transformer, the gap-aware calibration strategy (Gap-aware Knowledge

[0033] Rectification, GKR) generates an enhanced feature map for the next layer as follows:

[0034]

[0035] F i+1 = L i+1 (F i + Δf i ) i = 1, 2…, N - 1

[0036] F out = f N + Δf N

[0037] where F i + Δf i represents the feature map, x is the input image, which is input to the image encoder and represented by the block embedding layer, n represents the number of Blocks in ViT, and c is the dimension of F1, F2, …, F N .. Each layer L is frozen. Our focus is on learning the inter - domain gap correction between the current task and other image tasks to generate the following inter - domain knowledge correction:

[0038]

[0039] An ideal Δf i can help the vision - based model to bridge the difference gap between two types of tasks or images. First is the scene difference between the pre - training dataset and the target scene.

[0040] To effectively utilize the connection between the pre - trained backbone knowledge and the current medical image segmentation task, and break the distribution difference and the inter - domain task gap difference existing between images. GKR starts from a set of learnable Tokens , where each Token sequence Ti is randomly initialized, and m represents the sequence length of Ti. GKR freezes the image encoder of the SAM backbone network and embeds the correction knowledge learned from the fine - tuning dataset into these Tokens. Considering the essential requirements of identifying foreground and background regions or multiple instances in a single medical image in medical image segmentation, the dynamic knowledge correction of perceiving the domain gap implements an attention - inspired mechanism, enabling the basic vision model SAM to perform customized adjustments on the features of different instances or category regions, thus helping SAM to adapt to the differences between different medical image modalities. Specifically, GKR learns to use the dot - product operation to generate a similarity map Si, which captures the association between the feature vectors in fi and the tokens in T:

[0041]

[0042] Where $T_i$ represents the token sequence of the $i$-th layer, and $m$ represents the number of tokens in $T_i$. Since $S$ quantitatively evaluates the relationship between various tokens and feature vectors, the dynamic knowledge correction of the perceptual field gap can apply a softmax function to align each patch with the corresponding category:

[0043]

[0044] The similarity mapping from token to feature can be used to initially estimate $\Delta f$ using the following formula i :

[0045]

[0046] where, and represent the weights and biases of a multi-layer perceptron (MLP), respectively.

[0047] To enhance the flexibility of feature adjustment, GKR uses an MLP composed of and to generate the final inter-domain knowledge correction:

[0048]

[0049] The finally obtained knowledge-corrected feature is $F$ i ' = $F$ i + $\Delta f$ i .

[0050] Step 3. Introduce detailed task feature cues through the detail-aware cue learning module for the knowledge-corrected features obtained through the knowledge correction strategy, helping the model to focus more on the target area in local feature extraction and deeply explore task-related detailed information, as Figure 3 shown;

[0051] The specific implementation details are as follows:

[0052] Step 3.1. Extract local features through a convolutional network based on the enhanced feature map. The convolutional network includes three convolutional layers, one max pooling, and a 3×3 convolutional layer with a stride of 2;

[0053] Step 3.2. Project the feature map into an $L$-dimensional space using multiple 1×1 convolutional layers to generate similarity maps of three scales, corresponding to 1 / 2, 1 / 4, and 1 / 8 resolutions of the original image respectively. Each scale of similarity map is $L$-dimensional, denoted as $V = \{v_1, v_2, v_3\}$; where $V$ represents the set of similarity maps, and $v_1, v_2, v_3$ represent the similarity maps of 1 / 2, 1 / 4, and 1 / 8 resolutions respectively;

[0054] Step 3.3: Introduce a set of trainable vectors Q = {q1, q2, q3} for horizontal embedding to learn local detailed information. Here, Q represents the set of trainable vectors, and q1, q2, q3 represent the first, second, and third trainable vectors.

[0055] Step 3.4: Fuse the similarity map of each scale with the trainable vectors to obtain the feature map P output for each layer. i ;

[0056] Step 3.5: Use the feature map P obtained by fusing the previous layer as the query, the output feature of the previous layer as the key and value of cross-attention, and inject the pre-trained knowledge contained in i-1 into the prompt to generate the local detailed prompt information P for the current task after enhancement and optimization. as the key and value of cross-attention, and inject the pre-trained knowledge contained in into the prompt to generate the local detailed prompt information P for the current task after enhancement and optimization. q ;

[0057] Step 3.6: Use the enhanced and optimized local details prompt P q to generate the input for the next stage thereby completing the adaptive adjustment of features.

[0058] The specific details of Step 3 are as follows: Use the feature F i ' output by the previous layer of ViT for processing, and use a convolutional backbone network. The designed convolutional system includes the following parts: three convolutional layers, one max-pooling layer, and a stack of 3×3 convolutional layers with a stride of 2. This design helps to increase the number of channels of the feature map and reduce its spatial size, thereby enhancing the model's ability to process features at different scales. Finally, project the feature map into the L-dimensional space through multiple 1×1 convolutional layers. After these operations, similarity maps at three scales can be obtained, corresponding to 1 / 2, 1 / 4, and 1 / 8 resolutions of the original image respectively. The similarity map of each scale is L-dimensional, denoted as V = {q1, q2, v3}, where v1, v2, v3 represent the feature representations at different resolutions respectively. In addition, a set of trainable vectors Q = {q1, q2, q3} is introduced for horizontal embedding, aiming to learn local detailed information. These vectors are initialized using a Gaussian function to ensure that features can share some common representations during training, thereby improving the model's performance on different tasks. Finally, through these designs, the feature map P i output for each layer can be obtained through the following formula:

[0059] P i = Fusion(v i , q i )

[0060] where Fusion(,) is the feature fusion operation. Here, we first repeat V i several times so that the vector can be adjusted to the same shape as q i and then perform element-wise addition. Finally, we flatten and concatenate the prompt feature map F i ′ into local detail prompt information as the input of the auxiliary prompt information for the next layer of ViT, which is used for prompt refinement and feature adaptation. This fusion operation ensures the effective combination of each detail-enhanced embedding vector with the similarity maps of different scales, enabling each prompt to provide more accurate local information.

[0061] To gradually refine the prompt and improve the feature generalization ability through the refined prompt, we design a novel cross-attention mechanism. We use the detail-aware prompt P i-1 from the previous network stage as the query, and the output features of the previous layer as the key and value of the cross-attention, so as to inject the pre-trained knowledge contained in into the prompt to optimize the local detail prompt information for the current task and enhance the knowledge richness of the model. This process can be expressed as:

[0062]

[0063] where norm(·) represents layer normalization LayerNorm, and Attention(·) represents cross-attention operation. The attention layer uses the sparse attention mechanism to reduce the computational cost. After that, we use the improved local detail prompt P i to generate the input for the next stage to complete the adaptive adjustment of the features. This process can be expressed as:

[0064]

[0065] where is a learnable scalar used to balance the relationship between the task-specific prompt and the input features. At initialization, γ i is 0 to avoid drastic changes in the features at the initial stage. In this way, the flexibility of the local detail prompt is ensured while fully utilizing the knowledge contained in the pre-trained model.

[0066] Step 4: Pass the feature image through the boundary refinement module to obtain the boundary information for target image segmentation:

[0067] After all the above steps, that is, the gap-aware truth-seeking strategy, detail-aware cue learning and 12-layer Transformer network in the SAM image encoder, the features obtained by the last layer of the Transformer network are input into the boundary refinement module to predict the contour of the image segmentation area. Given the prediction map p output by the last layer of the backbone network structure Transformer k , the network is defined as follows:

[0068]

[0069] in and are the foreground and background attention features, represents the overall features of the last layer, conca(·) represents the feature concatenation operation, and conv3(·) represents a 3×3 convolutional network, which are expressed as:

[0070]

[0071]

[0072] Where S() and R() are sigmoid and reverse operators, i.e., element-wise subtraction of 1, and RCAB() is the residual channel attention block, which is used to emphasize those information channels and high-frequency information. In this case, edge features are explicitly incorporated to guide the segmentation process and promote edge prominence. This strengthens the feature-level edge information, thereby ensuring sharp and smooth edges of the segmented objects.

[0073] Step 5: Input the overall features of the target image into the SAM decoder and output the predicted segmentation map. After the above two stages, we use the hint encoder and mask decoder of the SAM backbone network to convert the high-dimensional feature representation into a specific segmentation mask image.

[0074] Comparative experimental analysis shows that the results are as follows Figure 4As shown: We compared our model with the SAM algorithm of similar large models and other models based on convolutional neural networks or Transformers on three intracranial hemorrhage datasets. We first observed that the performance of the model method we proposed is better than that of other similar SAM large model algorithms and convolutional neural network-based algorithms. These results confirm the effectiveness of the gap perception correction strategy, detail perception prompt learning, and boundary refinement module. For example, on the BCIHM dataset of non-contrast CT scans, it is much better than the convolutional neural network-based U-Net method and the Transformer-based method. That is, the Dice index has increased by nearly 20%. In addition, compared with other similar SAM methods, our model method has also improved the performance by nearly 10% and has fewer trainable parameters. With only 5.43M trainable parameters, it can achieve good results. These results greatly verify the effectiveness of the proposed gap perception correction strategy, detail perception prompt learning, and boundary refinement.

[0075] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model, characterized in that, include: Construct a medical image segmentation model based on the SAM large model, and perform medical image segmentation through the medical image segmentation model; The medical image segmentation model includes: an image preprocessing module, a multi-layer Transformer, a boundary refinement module and a SAM decoder; Medical image segmentation is performed through medical image segmentation models, including: Step 1: the image preprocessing module normalizes the input image to obtain normalized image data; Step 2: The normalized image data is input into a multi-layer Transformer for feature extraction. For the features extracted by each layer of Transformer, the features are enhanced through the gap-aware correction strategy, and the enhanced features are introduced into detailed task feature prompts, which helps the model focus more on the target area in local feature extraction and deeply mines the detailed information related to the task; Step 3: Input the feature image extracted by the last layer of Transformer into the boundary refinement module to perform target image segmentation masking to obtain the overall features of the image; Step 4: Input the overall features of the image into the SAM decoder for high-dimensional feature representation and convert it into a specific segmentation mask image, and output the predicted segmentation map.

2. The medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model according to claim 1, wherein Feature enhancement via gap-aware correction strategies, including: The preprocessed images are input into the medical image segmentation model. After each layer of Transformer feature extraction in the model, the gap-aware correction strategy corrects the features of the feature map generated by the i-th layer of the i-th Transformer. The gap-aware correction strategy GKR generates an enhanced feature map for the next layer. The gap-aware correction strategy GKR is used to learn the inter-domain gap correction between the current task and other image tasks to generate inter-domain knowledge correction.

3. A medical image segmentation method based on visual prompt tuning for domain generalization of the SAM large model according to claim 2, characterized in that, For the feature map generated by the i-th layer of the i-th layer Transformer, the gap-aware correction strategy GKR generates an enhanced feature map for the next layer, including: F i+1 = L i+1 (F i + Δf i ) i = 1, 2…, N - 1 F out = F N + Δf N Among them, F1 represents the feature map generated by the first-layer Transformer, L1 represents the first-layer Transformer, Embed represents the embedding process on the input image, and x represents the input image. indicates that the dimension of the feature map is n×c; F i+1 represents the feature map generated by the (i + 1)-th layer of the Transformer, F i represents the feature map generated by the i-th layer of the Transformer, Δf i represents the feature after being corrected by the i-th layer of the gap perception knowledge correction strategy, L i+1 represents the (i + 1)-th layer of the Transformer, and N represents the total number of Transformer layers. F out represents the output features after passing through the Transformer layer and the gulf perception knowledge correction, F N represents the feature map generated by the Nth layer of the Transformer, Δf N represents the features after being corrected by the Nth layer of the gulf perception knowledge correction strategy, n represents the number of image patches, and c represents the dimension of the image patches.

4. A medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model according to claim 2, characterized in that, The gap-aware correction strategy GKR is used to learn the inter-domain gap correction between the current task and other image tasks to generate inter-domain knowledge correction, including: Among them, Δf i represents the inter-domain gap correction between the image task i and other image tasks, GKR represents the gulf perception correction strategy, and f i represents the input image task i, indicates that the dimension of the image feature is n×c, n represents the number of image patches, c represents the dimension of the image patch, and N represents the total number of Transformer layers.

5. A medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model according to claim 1, characterized in that, Introducing enhanced features into detailed task feature hints helps the model focus more on the target area in local feature extraction and deeply explore task-related details, including: Step 3.1, extracting local features through a convolutional network based on the enhanced feature map, wherein the convolutional network includes three layers of convolution, a maximum pooling layer and a 3×3 convolution layer with a stride of 2; Step 3.2: Use multiple 1×1 convolutional layers to project the feature map into L-dimensional space to generate similarity maps of three scales, corresponding to 1 / 2, 1 / 4, and 1 / 8 resolutions of the original image, respectively. The similarity map of each scale is L-dimensional, denoted as V = {v1, v2, v3}; where V represents the similarity map set, v1, v2, and v3 represent similarity maps with 1 / 2, 1 / 4, and 1 / 8 resolutions, respectively; Step 3.3: Introduce a set of trainable vectors Q = {q1, q2, q3} for horizontal embedding to learn local detailed information. Here, Q represents the set of trainable vectors, and q1, q2, q3 represent the first, second, and third trainable vectors. Step 3.4: Fuse the similarity mapping at each scale with the trainable vector to obtain the feature map P output at each layer i ; Step 3.5: Take the feature map P fused in the previous layer i-1 as the query, and the output features of the previous layer as the key and value of cross-attention. Inject the pre-trained knowledge contained in into the prompt to generate the enhanced and optimized local detail prompt information P q for the current task; Step 3.6: Use the locally detailed hint P after enhancement and optimization q to generate the input for the next stage so as to complete the adaptive adjustment of features.

6. The medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model according to claim 5, wherein Fuse the similarity map at each scale with a trainable vector to obtain the feature map P output at each layer i , including: P i = Fusion(v i ,q i ) Among them, P i represents the output feature map of the i-th layer of the Transformer, Fusion represents the feature fusion operation, and v i represents the similarity mapping, and q i represents the trainable vector.

7. A medical image segmentation method based on SAM large model domain generalization visual prompt tuning according to claim 5, characterized in that, Take the feature map P obtained by the previous layer fusion i-1 as the query, and the output features of the previous layer as the key and value of the cross-attention, inject the pre-trained knowledge contained in into the prompt to generate the local detail prompt information P of the current task after enhancement and optimization q , including: Among them, P q represents the local detail hint information of the current task after enhancement and optimization, F i-1 represents the features generated by the (i - 1)-th layer Transformer, P i-1 represents the fused features of the (i - 1)-th layer, represents the corrected features output after the (i - 1)-th layer passes through the gulf perception knowledge correction strategy, Attention represents the cross-attention operation, and norm represents the layer normalization LayerNorm.

8. A medical image segmentation method based on visual prompt tuning for domain generalization of the SAM large model according to claim 5, characterized in that Using the enhanced and optimized local detail hint P q is used to generate the input for the next stage including: Among them, represents the input feature map of the next stage, represents the corrected feature generated after the i-1 layer passes through the gully perception knowledge correction strategy, represents the prompt information output by the i-1 layer through detail perception prompt learning, P q represents the local detail prompt information of the current task after enhancement and optimization, γ i represents a learnable scalar for balancing the relationship between task characteristic prompts and input features. Attention represents the cross-attention operation, and norm represents the layer normalization LayerNorm.

9. A medical image segmentation method based on domain generalization visual prompt tuning of the SAM large model according to claim 1, characterized in that, Input the feature image extracted by the last layer of the Transformer into the boundary refinement module for object image segmentation masks to obtain the overall features of the image, including: Among them, represents the overall characteristics of the image, and respectively represent the foreground and background attention features, p k represents the prediction map output by the last layer of the Transformer, conca represents the feature concatenation operation, and conv3 represents a 3×3 convolutional network.