2.5 D promptable medical image segmentation method and device based on SAM
By introducing cross-slice attention and residual learning mechanisms in the SAM model and combining LoRA fine-tuning strategy, the 2.5D can prompt medical image segmentation method is designed, which solves the problem of insufficient accuracy in medical image segmentation of SAM models and achieves a more efficient and accurate medical image segmentation effect.
Patent Information
- Application Number
- CN202510347869.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing SAM model lacks universality and adaptability in medical image segmentation, and there is a problem of insufficient accuracy when applied directly to medical image segmentation.
The 2.5D segmentation idea was introduced, and the cross-slice attention mechanism and residual learning mechanism were introduced in the SAM model, combined with the LoRA fine-tuning strategy, a 2.5D prompt medical image segmentation method based on SAM was designed, and the three-dimensional medical image data was preprocessed using the trained SAM model to generate 2.5D data blocks, and the segmentation results were optimized through the iterative point prompt and mask prompt strategy.
It improves the accuracy and consistency of medical image segmentation, especially when processing images with complex structures and blurred boundaries, significantly improves the accuracy and stability of the segmentation results and reduces the computing resource requirements.
Smart Images

Figure CN120279043A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of computer vision and medical image processing, and particularly to a 2.5D promptable medical image segmentation method and device based on SAM (SegmentAnything Model). Background Art
[0002] Medical image segmentation is a key task in medical image analysis, aiming to extract regions of interest (such as tumors, organs, or lesion areas) from medical images for subsequent diagnosis, treatment planning, or pathological analysis. Before the emergence of deep learning, medical image segmentation usually relied on traditional image processing methods, such as threshold segmentation, edge detection, region growing, active contour models, etc. These methods usually rely on manually designed features, such as gray values, textures, edges, etc., and have poor performance when dealing with complex or blurred medical images. With the development of deep learning technology, especially the successful application of convolutional neural networks in computer vision, deep learning has become one of the main technologies for medical image segmentation.
[0003] U-Net is a classic network architecture for medical image segmentation, consisting of a typical encoder-decoder architecture. The encoder gradually extracts high-dimensional features of the image through convolutional and pooling layers, and the decoder restores the spatial resolution of the image through deconvolutional layers, finally achieving fine pixel-level segmentation. With the diversification of medical image segmentation tasks, U-Net has also evolved continuously, forming many different variants to address more complex and diverse medical image segmentation problems. For example, U-Net++ adds more skip connections on the basis of U-Net and enhances feature fusion and information flow through a nested structure; V-Net adopts a similar encoder-decoder structure to U-Net, but it introduces fully convolutional operations and batch normalization in the network, which can better process 3D voxel data. In recent years, in order to further improve the performance of the network, the attention mechanism has been widely applied to medical image segmentation. By introducing attention modules, the network can automatically focus on important regions in the image and ignore irrelevant information, thereby improving the accuracy of the segmentation results.
[0004] However, traditional image segmentation methods are usually designed for specific types of images and tasks, lacking generality and being inflexible when dealing with diverse inputs or new tasks. The emergence of SAM (SegmentAnything Model) overcomes these problems. As a general image segmentation model, SAM fills the gaps in generality, adaptability, and interactivity of existing image segmentation models. With the gradual expansion of the application of SAM, its application in the field of medical image segmentation is also increasing. However, SAM was not originally designed specifically for medical image segmentation, so directly applying it to medical image segmentation has limitations and affects the accuracy of medical image segmentation results. Summary of the Invention
[0005] The purpose of this application is to provide a 2.5D promptable medical image segmentation method and device based on SAM, which can improve the accuracy of medical image segmentation results.
[0006] To achieve the above purpose, this application provides the following solutions:
[0007] In the first aspect, this application provides a 2.5D promptable medical image segmentation method based on SAM, including:
[0008] Obtain the three-dimensional medical image data to be processed;
[0009] Preprocess the three-dimensional medical image data to be processed to obtain a number of 2.5D data blocks, where each 2.5D data block includes a target slice and adjacent slices, and the adjacent slices refer to a preset number of slices adjacent to the target slice before and after.
[0010] Take each 2.5D data block as an input, and use the trained SAM model to output the corresponding segmentation result, where the SAM model includes an image encoder, a cross-slice attention mechanism module, a prompt encoder, and a mask decoder. The image encoder is used to process the 2.5D data block to obtain the context residual features of every two adjacent slices in the 2.5D data block and the image features of each slice in the 2.5D data block. The cross-slice attention mechanism module is used to process the image features by using the cross-slice attention mechanism to obtain fused image features. The prompt encoder is used to encode the initial point prompt information to obtain the initial encoded prompt information, and the initial point prompt information is any pixel point in the 2.5D data block. The mask decoder is used to obtain the segmentation result according to the context residual features, the fused image features, and the initial encoded prompt information.
[0011] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the SAM-based 2.5D promptable medical image segmentation method described in the above first aspect.
[0012] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the SAM-based 2.5D promptable medical image segmentation method described in the above first aspect.
[0013] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the SAM-based 2.5D promptable medical image segmentation method described in the above first aspect.
[0014] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0015] The present application provides a SAM-based 2.5D promptable medical image segmentation method and apparatus. The method includes: preprocessing the to-be-processed three-dimensional medical image data to obtain a plurality of 2.5D data blocks; using each 2.5D data block as an input, and outputting a corresponding segmentation result by using a trained SAM model, where the SAM model introduces a residual learning mechanism and a cross-slice attention mechanism to process the 2.5D data blocks respectively. The present application introduces the 2.5D segmentation idea into SAM, and introduces an attention mechanism at the slice feature level in the feature extraction stage, so that adjacent slices guide the segmentation of the current slice, and introduces a residual learning mechanism, so that the SAM model can better maintain the spatial consistency between slices, fully capture the correlation between each slice, thereby improving the expression ability of image features and enhancing the performance of the model on complex structure images, and thus can improve the accuracy of medical image segmentation results. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is an application environment diagram of a SAM-based 2.5D promptable medical image segmentation method in Embodiment 1 of the present application;
[0018] Figure 2Schematic flowchart of a 2.5D promptable medical image segmentation method provided in Embodiment 1 of this application;
[0019] Figure 3 Conceptual diagram of a 2.5D promptable medical image segmentation method in Embodiment 1 of this application;
[0020] Figure 4 Schematic structural diagram of a computer device provided in Embodiment 2 of this application. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0022] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in conjunction with the accompanying drawings and specific implementation manners.
[0023] Embodiment 1
[0024] Through research, it is found that since medical images are usually three-dimensional data and SAM is designed for 2D images, the existing SAM cannot fully capture the context and spatial continuity between 3D image slices, resulting in inconsistent or inaccurate segmentation results between slices. If the two-dimensional SAM is changed to a fully three-dimensional architecture, while increasing computing resources, the pre-trained model parameters cannot be used.
[0025] In response to this, this embodiment provides a 2.5D promptable medical image segmentation method based on SAM, introducing the 2.5D segmentation idea into SAM, and introducing an attention mechanism at the slice feature level in the feature extraction stage, so that adjacent slices guide the segmentation of the current slice (i.e., the target slice below). To address the problem that 2D models cannot capture the continuity between slices, a residual learning mechanism is designed to enable the model to better maintain spatial consistency between slices. During the prompt learning process, an iterative point prompt and mask prompt strategy is adopted. For point prompts, a point is randomly selected in the initial iteration, and points are reselected in the missegmented area in subsequent iterations. For mask prompts, on the one hand, during the residual learning process, the extracted residual features (i.e., context residual features), image embeddings (i.e., fused image features), and the mask features obtained in the previous iteration are input into the mask decoder to generate segmentation results. On the other hand, the segmentation result of the previous iteration is directly used as the mask prompt for the next iteration.
[0026] The 2.5D promptable medical image segmentation method based on SAM provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the three-dimensional medical image data to be processed to the server 104. After receiving the three-dimensional medical image data to be processed, the server 104 preprocesses the three-dimensional medical image data to be processed to obtain a number of 2.5D data blocks. Among them, each 2.5D data block includes a target slice and adjacent slices, and the adjacent slices refer to a preset number of slices adjacent to the target slice before and after; taking each 2.5D data block as an input, using the trained SAM model to output the corresponding segmentation result, where the SAM model includes an image encoder, a cross-slice attention mechanism module, a prompt encoder, and a mask decoder. The image encoder is used to process the 2.5D data block to obtain the context residual features of every two adjacent slices in the 2.5D data block and the image features of each slice in the 2.5D data block. The cross-slice attention mechanism module is used to process the image features by using the cross-slice attention mechanism to obtain the fused image features. The prompt encoder is used to encode the initial point prompt information to obtain the initial encoded prompt information, and the initial point prompt information is any pixel point in the 2.5D data block. The mask decoder is used to obtain the segmentation result according to the context residual features, the fused image features, and the initial encoded prompt information. The server 104 can feedback the obtained segmentation result to the terminal 102. In addition, in some embodiments, the 2.5D promptable medical image segmentation method based on SAM can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly process the three-dimensional medical image data to be processed by using the 2.5D promptable medical image segmentation method based on SAM, or the server 104 can obtain the three-dimensional medical image data to be processed from the data storage system and process it by using the 2.5D promptable medical image segmentation method based on SAM.
[0027] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablets, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0028] For example,Figure 2 As shown in Figure 2 , this embodiment provides a SAM-based 2.5D promptable medical image segmentation method, which segments three-dimensional medical images based on the SAM architecture. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of this application, taking the method applied to Figure 1 the server 104 in Figure 1 as an example for illustration, it includes the following steps 201 to 203. Among them:
[0029] Step 201, obtain the three-dimensional medical image data to be processed;
[0030] Step 202, preprocess the three-dimensional medical image data to be processed to obtain a plurality of 2.5D data blocks. Each 2.5D data block includes a target slice and adjacent slices, and the adjacent slices refer to a preset number of slices adjacent to the target slice before and after.
[0031] Step 203, use each 2.5D data block as input, and use the trained SAM model to output the corresponding segmentation result. The SAM model includes an image encoder, a cross-slice attention mechanism module, a prompt encoder, and a mask decoder. The image encoder is used to process the 2.5D data block to obtain the context residual features of every two adjacent slices in the 2.5D data block and the image features of each slice in the 2.5D data block. The cross-slice attention mechanism module is used to process the image features by using the cross-slice attention mechanism to obtain fused image features. The prompt encoder is used to encode the initial point prompt information to obtain the initial encoded prompt information. The initial point prompt information is any pixel point in the 2.5D data block. The mask decoder is used to obtain the segmentation result according to the context residual features, the fused image features, and the initial encoded prompt information.
[0032] In this embodiment, by combining slices into 2.5D data blocks and introducing a cross-slice attention mechanism and a residual learning mechanism, the SAM model can obtain richer context information, capture the dependencies between different parts of the image, maintain the spatial consistency of the image, enhance the image representation ability, improve the understanding ability of complex images, and thus improve the accuracy of medical image segmentation results.
[0033] Next, the specific execution process of the SAM-based 2.5D promptable medical image segmentation method in this embodiment will be elaborated.
[0034] This embodiment is based on SAM, and on this basis, a LoRA (Low-Rank Adaptation) fine-tuning strategy is introduced to design a 2.5D segmentation framework, enabling the SAM model to learn the dependencies between slices. The specific steps are as follows:
[0035] (1) Preprocess the three-dimensional medical image data. For each slice, select its adjacent front and back slices to form a new data sample, namely a "2.5D" data block;
[0036] (2) Design a neural network (the SAM model is a neural network model) to implement the three-dimensional medical image segmentation function. Introduce the LoRA fine-tuning strategy into the SAM model, freeze the parameters in the image encoder and mask decoder, add LoRA layers to the attention layers of the image encoder and mask decoder, and only update the model parameters of the LoRA layers;
[0037] (3) Use the image encoder in SAM to extract high-dimensional image features, and add a cross-slice attention mechanism to the image encoder, that is, learn the correlation between slices in the three-dimensional medical image, so that adjacent slices can guide the segmentation of the current slice, and fully capture the anatomical structure information between consecutive slices. In the cross-slice attention mechanism, the features of each slice interact with the features of other slices through the attention mechanism. By calculating the attention weights between slices, the model can dynamically determine the correlation of adjacent slices to the current slice, and obtain the features of each slice after weighted fusion (i.e., the fused image features);
[0038] (4) For the three-dimensional medical image, obtain the context residual information by calculating the difference between two adjacent slices, that is, taking the absolute value of the subtraction. Combine the context residual information between two adjacent slices, the initial encoded prompt information (obtained by encoding the initial point prompt information through the prompt encoder) and the fused image features, and input them into the mask decoder to enable the model to better maintain the spatial consistency between slices and achieve more accurate segmentation, obtaining the first segmentation result;
[0039] (5) During the training process, an iterative prompting learning strategy is proposed, that is, using point prompts and mask prompts to gradually guide the model to learn the correct target segmentation region. For point prompts, at the initial iteration of training, a single pixel point is selected from the foreground (the region to be segmented) as the prompt. From the second iteration onwards, the segmentation result of the previous iteration is compared with the ground truth label, and points are reselected in the missegmented regions as the point prompts for subsequent iterations. By reselecting points in the missegmented regions in this way, the model can continuously improve the previous segmentation result. For mask prompts, on the one hand, mask features are generated based on the segmentation result obtained in the previous iteration process (the segmentation result of the previous iteration passes through convolution, batch normalization, and activation functions to obtain mask features). After combining with the extracted context residual features and fused image features, they are input into the mask decoder to generate the segmentation result of the current iteration. On the other hand, the segmentation result of the previous iteration is directly used as the mask prompt for the current iteration. First, it is input into the prompt encoder, and after operations such as convolution, batch normalization, and activation functions, mask features are generated;
[0040] (6) A large number of comparative experiments are carried out on the abdominal multi-organ segmentation dataset, and the Dice metrics on each organ and the average Dice value are recorded;
[0041] (7) Ablation experiments are designed, mainly including ablation of the SegmentAnything Model (SAM), low-rank fine-tuning strategy, prompting learning mechanism, cross-slice attention mechanism, and residual learning strategy, to explore the contribution degree of these strategies to the final segmentation performance.
[0042] Aiming at the fact that the existing 2D SAM cannot learn the correlation between 3D voxel medical image slices, in this embodiment, the 2.5D segmentation idea is introduced into SAM. First, an attention mechanism for slice feature dimensions is designed to learn the dependence between adjacent slices, and a residual slice learning mechanism is introduced to capture continuous anatomical structure information. In prompting learning, an iterative learning strategy is adopted, using point prompts and mask prompts to gradually guide the model to learn the correct target segmentation region. During the overall training process, the LoRA (Low-Rank Adaptation) fine-tuning strategy is adopted to only update a small number of parameters, reducing the computational cost and storage requirements. This embodiment is applicable to various 2D and 3D medical image segmentation tasks, such as abdominal multi-organ segmentation and brain tumor segmentation, etc.
[0043] The following is described with specific examples.
[0044] Using the Synapse Abdominal Multi-Organ Segmentation Dataset, which is derived from the MICCAI 2015 Multi-Abdominal Labeling Challenge and contains a large number of abdominal CT scan images. The training set has 2,212 axial enhanced abdominal CT images, and the test set has 12 samples in 3D voxel data format, stored in HDF5 format. The images are annotated with eight organ regions (aorta, gallbladder, liver, pancreas, left kidney, right kidney, spleen, and stomach). These organs have significant differences in size, shape, and position, and the boundaries between adjacent organs are not clear. There are outliers and noise in the images, which increases the difficulty of segmentation.
[0045] As Figure 3 shown, first, preprocess the dataset. For three-dimensional voxel medical images, not only consider the slice itself, but also select its adjacent slices before and after, and combine these slices into new data. Specifically, for each target slice, several adjacent slices before and after it will be used together with the target slice as input data to form a "2.5D" data block containing local context. This way can not only retain the local details of the two-dimensional image, but also introduce cross-slice spatial information, enabling the model to more accurately understand the continuity of tissue structure and organ morphology in three-dimensional space. By this way of fusing adjacent slices, the accuracy and robustness of the segmentation result can be improved without increasing too much computational complexity.
[0046] In the segmentation algorithm, use the "vit_b" version of SAM as the backbone network, and introduce the LoRA strategy in the image encoder and mask decoder to reduce the computational cost. First, the input image passes through the image encoder to obtain image features, and then use the proposed cross-slice attention mechanism to learn the dependencies at the slice feature level, enabling the current slice to make full use of context information. At the same time, in order to maintain the spatial consistency between slices, use residual slice learning to capture continuous anatomical information. For adjacent input slices, calculate the residual information between adjacent slices (i.e., pixel residual information):
[0047] I res =|I i+1 -I i |(1);
[0048] where, I i+1 and I i represent two adjacent slices, and the superscript i represents the slice number. The residual information between slice features (i.e., feature residual information) is calculated in a similar way:
[0049] F res =|F i+1 -F i | (2);
[0050] where, F i+1 and Fi Represents the features of adjacent slices. Then, the residual information between the residual information of the original image and the slice features is combined to obtain a more representative residual feature representation (i.e., context residual feature):
[0051]
[0052] Where σ represents the ReLU activation function, BN represents the batch normalization operation, and conv represents the convolution operation. The combined residual feature representation and image features are fed into the mask decoder. At the same time, the point prompt of the first iteration is input into the mask decoder to obtain the prediction result.
[0053] In the subsequent training process, an iterative training strategy is used. For the point prompt, the segmentation result obtained from the previous iteration is compared with the ground truth annotation. In the missegmented regions, pixel points are reselected as the point prompt for the subsequent iteration. In this way, the model can gradually focus on the regions with inaccurate segmentation, continuously adjust and optimize the segmentation result. Especially for objects with complex structures or blurred boundaries, this strategy of iteratively updating the point prompt helps to improve the accuracy of segmentation. For the mask prompt, during the residual learning process, the mask prompt is combined with the extracted context residual features and the fused image features, and then input into the mask decoder to generate the subsequent segmentation result. The mask prompt plays a role in supplementing information here, helping the model better understand the structures and features in the image, so as to make more accurate segmentation predictions. At the same time, the segmentation result obtained from the previous iteration is directly used as the mask prompt for the next iteration. That is, the segmentation result is input into the prompt encoder. After operations such as convolution, batch normalization, and activation functions, a mask feature is generated. Then, this mask feature is element-wise added to the context residual features and the fused image features and used for subsequent operations. This enables the model to use the previous segmentation result to guide the next segmentation, forming an iterative optimization process and gradually improving the segmentation accuracy.
[0054] Finally, the cross-entropy loss and Dice coefficient loss are used to supervise the training of the model. In the image segmentation task, the cross-entropy loss is used to measure the difference between the predicted class and the actual class. The Dice coefficient is used to measure the similarity between two sets. The Dice loss optimizes the model by calculating the similarity between the segmented region and the ground truth region, aiming to maximize the ratio of the intersection to the union.
[0055] The 2.5D promptable medical image segmentation method based on SAM in this embodiment is based on the SAM architecture, introduces the LoRA fine-tuning strategy, updates only a small number of parameters, and reduces the computational cost. An attention mechanism at the slice feature level is introduced in the image encoder to learn the correlation between adjacent slices. Iterative point prompts and mask prompts are used to integrate the context residual features, fused image features, and mask features of adjacent slices into the mask decoder to obtain the segmentation result and use it as the guiding information for the subsequent iterative process. For point prompts, a point is randomly selected in the initial iteration, and points are reselected in the missegmented area in subsequent iterations, gradually guiding the model to learn the correct target segmentation area.
[0056] In this embodiment, by combining slices into 2.5D data blocks and using a cross-slice attention mechanism, the model can obtain richer context information, thereby improving the ability to understand complex images; the LoRA fine-tuning technology enables reducing the training resource requirements without sacrificing performance, solving the computational bottleneck in the training of large-scale image models; through cross-slice attention and residual learning, the model can capture the dependencies between different parts of the image, maintain the spatial consistency of the image, and enhance the image representation ability; the iterative prompt mechanism focuses on the error area by dynamically adjusting the prompt information, helping the model accurately adjust and correct the prediction, and significantly improving the accuracy and stability of the final prediction.
[0057] Therefore, the combination of these technical features improves the performance of the SAM model, especially in dealing with complex image structures and precise tasks, enabling more efficient and accurate image understanding. This method can be applied to real medical research and can be extended to various three-dimensional medical images, contributing a little to medical research.
[0058] Embodiment 2
[0059] This embodiment provides a computer device, which can be a server or a terminal, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data in the 2.5D promptable medical image segmentation method based on SAM. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a 2.5D promptable medical image segmentation method based on SAM in Embodiment 1.
[0060] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are realized.
[0061] Embodiment 3
[0062] This embodiment provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it realizes a 2.5D promptable medical image segmentation method based on SAM in Embodiment 1.
[0063] Embodiment 4
[0064] This embodiment provides a computer program product including a computer program, and when the computer program is executed by a processor, it realizes a 2.5D promptable medical image segmentation method based on SAM in Embodiment 1.
[0065] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0066] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0067] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0068] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0069] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A SAM-based 2.5D promptable medical image segmentation method, characterized in that, The SAM-based 2.5D promptable medical image segmentation method includes: Obtain the three-dimensional medical image data to be processed; Preprocess the three-dimensional medical image data to be processed to obtain a number of 2.5D data blocks, where each 2.5D data block includes a target slice and adjacent slices, and the adjacent slices refer to a preset number of slices adjacent to the target slice before and after; Take each 2.5D data block as input and use the trained SAM model to output the corresponding segmentation result. The SAM model includes an image encoder, a cross-slice attention mechanism module, a prompt encoder, and a mask decoder. The image encoder is used to process the 2.5D data block to obtain the context residual features of every two adjacent slices in the 2.5D data block and the image features of each slice in the 2.5D data block. The cross-slice attention mechanism module is used to process the image features using the cross-slice attention mechanism to obtain fused image features. The prompt encoder is used to encode the initial point prompt information to obtain the initial encoded prompt information, where the initial point prompt information is any pixel point in the 2.5D data block. The mask decoder is used to obtain the segmentation result based on the context residual features, the fused image features, and the initial encoded prompt information.
2. The SAM-based 2.5D promptable medical image segmentation method according to claim 1, wherein The LoRA layer is introduced into the attention layers of the image encoder and the mask decoder, and the LoRA fine-tuning strategy is used to update the model parameters of the LoRA layer.
3. The SAM-based 2.5D promptable medical image segmentation method according to claim 1, wherein The specific process of processing the 2.5D data block to obtain the context residual features of every two adjacent slices in the 2.5D data block includes: Calculate the pixel residual information between every two adjacent slices in the 2.5D data block; Calculate the feature residual information between every two adjacent slices in the 2.5D data block; Fuse the pixel residual information and the feature residual information to obtain the context residual features.
4. The SAM-based 2.5D promptable medical image segmentation method according to claim 1, wherein, The specific process of processing the image features using the cross-slice attention mechanism to obtain fused image features includes: Calculate the attention weights of each image feature; Based on the attention weights, perform weighted summation on the image features of all slices in the 2.5D data block to obtain the fused image features.
5. The SAM-based 2.5D promptable medical image segmentation method according to claim 1, wherein, The SAM model is trained using the iterative prompt learning strategy, where the iterative prompt learning strategy includes point prompts and mask prompts. The point prompt in the iterative process of the SAM model refers to based on the segmentation result of the previous iteration of the SAM model, reselecting a pixel point in the mis-segmented area as the point prompt for the current iteration. The mask prompt in the iterative process of the SAM model refers to taking the segmentation result of the previous iteration of the SAM model as the mask prompt for the current iteration, generating the mask feature of the current iteration according to the mask prompt of the current iteration, and using the mask feature of the current iteration, the context residual features, and the fused image features as the input of the mask decoder to obtain the segmentation result of the current iteration.
6. The 2.5D SAM-based medical image segmentation method with prompting according to claim 1, wherein, The SAM model is trained based on the cross-entropy loss and the Dice coefficient loss.
7. The SAM-based 2.5D promptable medical image segmentation method according to claim 3, wherein, The calculation formula of the context residual feature is as follows: Among them, represents the context residual feature, σ represents the ReLU activation function, BN represents the batch normalization operation, conv represents the convolution operation, and I res represents the pixel residual information, and F res represents the feature residual information.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the SAM-based 2.5D promptable medical image segmentation method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the SAM-based 2.5D promptable medical image segmentation method according to any one of claims 1-7.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the SAM-based 2.5D promptable medical image segmentation method according to any one of claims 1-7.
Citation Information
Cited By
Construction method and application of abdominal organ segmentation SAM model based on prompt enhancement
CN120635083A
Medical image segmentation method and segmentation system
CN120655647A
A medical image segmentation method and segmentation system
CN120655647B