Abdominal CT multi-organ segmentation method based on LoRA vertical domain MedSAM large model
By introducing LoRA vertical domain technology into the MedSAM large model, combining multi-feature information fusion and dynamic context prompts, the challenge of multi-organ segmentation of abdominal CT is solved, achieving a more efficient and accurate segmentation effect.
Patent Information
- Application Number
- CN202411973609.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
Multi-organ segmentation of abdominal CT scans has challenges, especially due to similar strength between organs, blurred boundaries, and the failure of existing deep learning algorithms to make full use of spatial relationships between image slices, resulting in poor segmentation effects.
The multi-organ segmentation method of abdominal CT based on the LoRA vertical MedSAM large model is adopted to realize the automated segmentation of abdominal CT images through multi-feature information fusion, dynamic context prompts and adaptive LoRA fine-tuning.
The accuracy and robustness of the MedSAM large model in the abdominal multi-organ CT segmentation task is improved, and the organ boundaries can be segmented more accurately, adapt to limited training data, and may achieve better results in practical applications.
Smart Images

Figure CN119992083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular to an abdominal CT multi-organ segmentation method based on a LoRA vertical domain MedSAM large model. Background Art
[0002] Segmentation of abdominal CT scans is challenging due to the complex morphology of structures, inter- and intra-subject variations, and image features such as low contrast and blurred boundaries. Segmentation is further complicated by weak inter-organ boundaries with similar intensities in the abdomen. Preprocessing and feature extraction of raw CT images can achieve accurate segmentation and identification of organs in the model. Since 3D volume data contains a large number of slices, manual and semi-automatic segmentation becomes a time-consuming and laborious task, while end-to-end automated segmentation will be more practical for clinical applications. In addition, since high-quality medical images are difficult to obtain in clinical practice, and a large amount of clinical data is thick slices with low precision and diverse data resolutions. Therefore, efficient segmentation of massive clinical low-quality data is not only a technical challenge, but also has important clinical significance. Early medical image segmentation systems relied on traditional image processing methods such as edge detection, threshold-based methods, and region growing-based methods, which focus on certain specific organ structures and are difficult to extend to unified segmentation of multi-organ models. Existing deep learning algorithms allow models to segment 2D slice sequences from CT scans, but they fail to fully utilize the spatial relationship between image slices, resulting in poor multi-organ segmentation results. Based on this, there is an urgent need to provide an abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model to solve the above-mentioned problems existing in the prior art. Summary of the invention
[0003] The purpose of the present invention is to provide an abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model, which solves the problem of achieving economical and efficient segmentation with a small amount of data through the cooperation of multi-feature information fusion, dynamic context prompts and adaptive LoRA fine-tuning.
[0004] In order to solve the above technical problems, the present invention is achieved through the following technical solutions:
[0005] The abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model comprises the following steps:
[0006] Step 1, preprocessing the abdominal CT image supervised learning dataset;
[0007] Step 2, feature selection network construction;
[0008] Step 3, construction of LoRA layer;
[0009] Step 4: Use a hybrid prompt strategy to perform inter-layer context reasoning to achieve automated abdominal CT segmentation.
[0010] Furthermore, the abdominal CT image supervised learning dataset described in step 1 is preprocessed: an abdominal CT image dataset is collected from a medical imaging database or relevant medical institutions to ensure that the images cover samples of different patients, different pathological conditions and different scanning parameters to enhance the generalization ability of the model. The FLARE2022 dataset or a similar supervised abdominal CT image dataset is used. This dataset does not require manually annotated segmentation masks. The specific steps are as follows:
[0011] Step 1.1 Development adjustment operation: by adjusting the window width and window level, according to the tissue characteristics of the abdominal organs and the grayscale distribution characteristics of the image, select appropriate window width and window level values, and adjust the image to highlight the development of the abdominal organ area;
[0012] Step 1.2 Histogram equalization: Perform histogram equalization on the image after adjusting the window width and window position. By redistributing the grayscale of the image, the grayscale histogram of the image is made more uniform, the overall contrast of the image is enhanced, the impact of factors such as uneven illumination is reduced, the detailed information of the organ is highlighted, and better image quality is provided for subsequent segmentation tasks.
[0013] Step 1.3 Image feature enhancement based on Sobel operator: Use the Sobel operator to perform edge detection on the image. The Sobel operator is a discrete differential operator that can calculate the grayscale gradient amplitude and direction of each pixel in the image. In abdominal CT images, the Sobel operator can enhance the edge information of organs and help to segment organ boundaries more accurately. In the specific operation, the horizontal and vertical gradients are calculated separately, and then the edge pixels are determined according to the gradient amplitude and direction, providing clearer edge features for subsequent model training and segmentation.
[0014] Furthermore, the feature selection network described in step 2 is constructed, and the MedSAM large model is used as the basic model architecture. The feature selection network is embedded into the MedSAM large model to integrate the image prior information and the large model processing capability, thereby improving the MedSAM large model's ability to select and extract abdominal organ features. The feature selection network extracts this information from the original CT image data and the auxiliary feature data.
[0015] Construct a feature selection network based on UNet. The encoding path extracts image features through multiple convolution and pooling layers, while the decoding path gradually restores the original image size using deconvolution and skip connections. Given an image tensor As input, through the downsampling path, feature map fusion and upsampling path, the network outputs an image tensor As a feature map;
[0016] The network is used to extract the entropy and curvature information of the image. By improving the local entropy algorithm, it can more finely reflect the degree of distribution disorder of pixel values in the image. By using the complex statistical analysis of pixel values in the local area, the differences in texture features between different organs can be more accurately captured, providing a more powerful basis for distinguishing organs. Curvature information extraction combines anatomical prior knowledge, uses the general law of organ contour curvature in anatomy, and uses a specific convolution kernel to more accurately capture the curvature characteristics of organ contours, especially in the complex convex and concave areas at the edges of abdominal organs, and more accurately depicts the boundaries of organs. The feature selection network is embedded in the MedSAM large model, and the entropy and curvature information extracted from the feature selection network are weighted fused with the feature image generated by the image encoder from the MedSAM large model to improve the MedSAM model's ability to select and extract abdominal organ features.
[0017] For curvature feature selection, IFE (Instructive Feature Enhancement for Dichotomous Medical Image Segmentation) uses simple linear convolution to obtain the 2D surface in Euclidean space R 3 The approximate mean curvature K in is as follows:
[0018] K=[K 1 K 2 K 3 ]*X......(1),
[0019] In formula (1), K 1 =[α,β,α] T , K 2 =[β,γ,β] T , K 3 =[α,β,α] T ), where * represents convolution and X represents the input image;
[0020] In order to reflect the spatial and aggregation characteristics of intensity distribution, IFE defines information entropy as follows:
[0021]
[0022] In formula (2), i n represents the grayscale value of the center pixel of the nth 3×3 sliding window, j n represents the average gray value W of the remaining pixels in the window, and formula (2) calculates the gray value of the entire image (in ,j n ) has a probability of occurrence P i,j , and E is expressed as information entropy, where each pixel corresponds to a grayscale value between 0 and 255. By selecting the corresponding proportion of channel features and combining them with the original features, IFE helps improve the performance with minor modifications to the segmentation network architecture, as shown in Equation (3):
[0023] F=|Conv(X,k,b curv )|......(3),
[0024] In formula (3), X represents the feature map of the previous convolutional block, b curv It is the bias term in the convolution operation. The spatial dimensions (height and width) of the resulting tensor F are the same as those of X, but it highlights areas with specific features. k is a specific convolution kernel used to extract curvature features or entropy features.
[0025] Then, the calculated curvature values are summed in the spatial dimension to compress the 2D feature map into a single scalar value. The result of this process is that each channel is represented as a scalar value that reflects the overall curvature strength of the channel. Through the above algorithm, the feature selection network extracts this information from the original CT image data and auxiliary feature data;
[0026] The entropy and curvature information extracted from the feature selection network are fused with the feature image generated by the image encoder in the MedSAM large model using an adaptive weighted fusion mechanism. In this fusion process, an adaptive weighted fusion mechanism is used. The adaptive weighted fusion mechanism automatically assigns weights according to the prominence of different organ regions in the image. The basis for this mechanism is the statistical analysis of organ features in a large number of abdominal CT images and the learning of feature importance by the deep learning model. When the newly extracted entropy and curvature information is fused with it, the new information combination will provide a more comprehensive and valuable basis for subsequent segmentation reasoning.
[0027] Furthermore, the construction of the LoRA layer in step 3 is based on the image encoder of the MedSAM large model, and a learnable LoRA layer is added to each Transformer layer to achieve efficient parameter training. For the mask decoder of the MedSAM large model, the LoRA layer is also added to the self-attention and cross-attention layers; the update of LoRA is as shown in equations (4) and (5):
[0028] W 0 +ΔW=W 0 +BA ...... (4),
[0029] Formula (4) represents a matrix decomposition, where B is a d×r dimensional matrix, A is a d×k dimensional matrix, and the rank r is much smaller than the minimum value of d and k. For each layer in the MedSAM large model, during the training process, the weight matrix W of the layer is o remains unchanged and does not receive gradient updates, while A and B contain trainable parameters, and ΔW is W o The low-rank update matrix of , and then, the fixed weight h imposed by this layer on the input vector x is calculated as follows (5);
[0030] h=W 0 x+ΔWx=W 0 x+BAx ...... (5),
[0031] Formula (5) indicates that LoRA only updates a small part of the parameters during fine-tuning, which reduces the computational cost and storage constraints.
[0032] Furthermore, step 3 describes LoRA fine-tuning: LoRA, as a model fine-tuning method, adds a low-rank matrix on the basis of the pre-trained model to learn the parameter update of a specific task, so as to reduce the training parameters and the amount of calculation, and realize rapid adaptation to new tasks. In the abdominal multi-organ segmentation, according to the characteristics of the abdominal CT images and the task requirements of the organ segmentation, the appropriate low-rank matrix size and parameter initialization strategy are set in LoRA to achieve the performance enhancement of the pre-trained basic model with minimal computational input. As a method for efficient parameter fine-tuning, LoRA is integrated into the image encoder and mask decoder of the MedSAM large model. For the image encoder, LoRA adapts to the specific features of the abdominal CT image data by adjusting its low-rank matrix, and supervises the abdomen. The image dataset FLARE2022 is fine-tuned using the above method. Through the stratified sampling strategy, the dataset is stratified according to factors such as organ size, position, and complexity, so that the model can learn the characteristics of different types of organs more evenly, avoiding the deviation of the model in the segmentation of these organs due to the distribution characteristics of certain organs. In terms of the mask decoder, LoRA also optimizes its parameters to better handle the mask generation in the segmentation task. Only the parameters of the low-rank matrices A and B introduced by LoRA are updated, and the rest of the pre-trained weights of the MedSAM large model are kept relatively stable. The stratified sampling strategy is adopted to improve the accuracy of the model in the abdominal CT multi-organ segmentation task. On this basis, the training process is further optimized to meet the special needs of abdominal CT segmentation.
[0033] In the image encoder, the parameters of the Transformer layer are frozen, and the LoRA layer is added next to the query and key projection layers in the attention mechanism. The LoRA layer consists of two linear layers, which operate in the image encoder as shown in Equations (6)-(9):
[0034]
[0035] Q=W q T+B q A q T......(7),
[0036] K=W k T......(8),
[0037] V=W v T+B v A v T......(9),
[0038] In formula (6) to formula (9), represents the primitives from the fully connected layer, Q, K, V represent the query, key and value layers respectively, and W q , W k and W v is a frozen projection layer from the MedSAM large model, and A q , B q , A v and B v is the trainable parameter matrix in the LoRA linear layer;
[0039] The supervised abdominal image dataset FLARE2022 is used for efficient fine-tuning. In terms of image encoder and mask decoder, LoRA also optimizes its parameters to better handle mask generation in segmentation tasks. Only the parameters of the low-rank matrices A and B introduced by LoRA are updated, and the rest of the pre-trained weights of the MedSAM large model are kept relatively stable. The accuracy of the model in the abdominal CT multi-organ segmentation task is improved by adopting a stratified sampling strategy. On this basis, the training process is further optimized to adapt to the special needs of abdominal CT segmentation.
[0040] Furthermore, the vertical domain large model reasoning optimization of the fused prompt words described in step 4: the trained model and the set reasoning strategy are integrated into an automated workflow. When a new abdominal CT image needs to be segmented, the image is first preprocessed and then input into the model. The model performs reasoning based on the contextual relationship between the mixed prompt words and adjacent layers, automatically outputs the segmentation results of each organ, and saves the segmentation results in the form of image masks to facilitate subsequent medical analysis and visualization.
[0041] Automated segmentation: Integrate the trained model and the set reasoning strategy into an automated workflow. When a new abdominal CT image needs to be segmented, the image is first preprocessed and then input into the model. The model automatically outputs the segmentation results of each organ based on mixed cues and contextual reasoning between adjacent layers. The segmentation results can be saved in the form of image masks to facilitate subsequent medical analysis and visualization:
[0042] First, select the middle layer x of the CT image sequence t Perform segmentation, which includes all target organs;
[0043] Then, for x t Apply histogram equalization to enhance its clarity and calculate x t The cumulative distribution function is used to map each pixel value to obtain a new pixel value, as shown in the following equation (10):
[0044]
[0045] In formula (10), s k is the new pixel value, L∈[0,255] represents the grayscale level, g j is the probability of the jth gray level. After denoising the image using a median filter, the region growing method is used for initial organ segmentation. For each target organ, a seed is placed within its region to start the segmentation process;
[0046] Finally, the intermediate layer x is smoothed using morphological operations t The closing operation smoothes the edge of the object and fills the small holes in the object by dilating and then corroding the edge of the segmentation result, making the result more coherent, accurate and continuous, and providing detailed guidance for the segmentation of adjacent layers, as shown in the following formula (11):
[0047]
[0048] In formula (11), represents the expansion operation, Represents the corrosion operation, A is the image, B is the structural element, and the intermediate layer x is completed. t After region growing segmentation, the result m t Used as a mask cue for adjacent layers;
[0049] A hybrid prompt is constructed by fusing the bounding box prompt with the mask prompt of the previous layer's segmentation result to complete the vertical domain large model reasoning process of the fused prompt word. The box prompt provides initial position information by drawing a rectangular box surrounding the organ on the image to guide the model to focus on a specific area. For each organ in the abdominal CT image, a corresponding rectangular box is drawn on the image according to its approximate position and size. The mask prompt provides a rough organ mask as the model's prior knowledge to help the model better understand the shape and distribution of the organ. The combination of the two can accurately locate and segment the organ. When performing organ segmentation, the contextual information between adjacent layers of the abdominal CT image is considered, because the organ structures of adjacent layers are continuous and correlated, and this contextual information is used to improve the accuracy of segmentation.
[0050] Furthermore, the automatic segmentation of the CT image in step 4 also includes: adjusting the subsequent segmentation strategy in real time according to the difference between the current segmentation result and the predefined anatomical model, correcting possible segmentation deviations, and ensuring the accuracy and completeness of the segmentation process.
[0051] The present invention has the following superior technical effects:
[0052] 1. The abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of the present invention improves the accuracy and robustness of the MedSAM large model in the abdominal multi-organ CT segmentation task through multi-feature information fusion, dynamic context prompts and adaptive LoRA fine-tuning. Multi-feature information can characterize organ characteristics more comprehensively, dynamic context prompts can more effectively utilize context information, and adaptive LoRA fine-tuning can more effectively utilize limited training data. These innovations make this method more competitive and may achieve better results in practical applications.
[0053] 2. The present invention is based on the LoRA vertical domain MedSAM large model of the abdominal CT multi-organ segmentation method. Based on the image encoder of the MedSAM large model, a learnable LoRA layer is added to each Transformer layer to achieve efficient parameter training. For the mask decoder of the MedSAM large model, the LoRA layer is also added to the self-attention and cross-attention layers. In order to make full use of the features of the CT scan image, IFE is integrated into U-Net. This structure is used as a feature selection module to capture the curvature and entropy features in the image. The feature map generated by the feature selection module will be merged with the image embedding generated by the image encoder and then input into the mask decoder. In the segmentation process, regional growth is first applied to accurately segment the body organs, starting from the middle layer of the volume data. During the inference process of the model, the model is used for automatic segmentation. The previous segmentation results are used as mask prompts and as additional primitives to input the mask decoder. The original prompt encoder remains unchanged, and the box prompts extracted and expanded from the previous segmentation mask are integrated into it. The segmentation process runs in a step-by-step manner, using the previous results to enhance the segmentation of the next layer. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the specific technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below.
[0055] Figure 1 It is a flow chart of the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of the present invention;
[0056] Figure 2 Schematic diagram of the segmentation model architecture of the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of the present invention;
[0057] Figure 3 This is a schematic diagram of information related to feature extraction of the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of the present invention;
[0058] Figure 4 This is a schematic diagram of the use of the LoRA layer in Transformer of the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model described in the present invention. DETAILED DESCRIPTION
[0059] The technical solutions in the embodiments of the present invention will be described below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0060] Example
[0061] like Figure 1As shown, the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model includes the following steps:
[0062] Step 1, preprocessing the abdominal CT image supervised learning dataset;
[0063] Step 2, feature selection network construction;
[0064] Step 3, construction of LoRA layer;
[0065] Step 4: Use a hybrid prompt strategy to perform inter-layer context reasoning to achieve automatic abdominal CT segmentation;
[0066] As an example of a specific step: Step 1 preprocesses the abdominal CT image supervised learning dataset: collect abdominal CT image datasets from medical imaging databases or relevant medical institutions to ensure that the images cover samples from different patients, different pathological conditions, and different scanning parameters to enhance the generalization ability of the model. Use the FLARE2022 dataset or a similar supervised abdominal CT image dataset, which does not require manually annotated segmentation masks. The specific steps are:
[0067] Step 1.1 Development adjustment operation: by adjusting the window width and window level, select the appropriate window width and window level values according to the tissue characteristics of the abdominal organs and the grayscale distribution characteristics of the image;
[0068] Step 1.2 Histogram equalization: Perform histogram equalization on the image after adjusting the window width and window position. By redistributing the grayscale of the image, the grayscale histogram of the image is made more uniform, the overall contrast of the image is enhanced, the impact of factors such as uneven illumination is reduced, the detailed information of the organ is highlighted, and better image quality is provided for subsequent segmentation tasks.
[0069] Step 1.3 Image feature enhancement based on Sobel operator: Use the Sobel operator to perform edge detection on the image. The Sobel operator is a discrete differential operator that can calculate the grayscale gradient amplitude and direction of each pixel in the image. In abdominal CT images, the Sobel operator can enhance the edge information of organs and help to segment organ boundaries more accurately. In the specific operation, the horizontal and vertical gradients are calculated separately, and then the edge pixels are determined according to the gradient amplitude and direction, providing clearer edge features for subsequent model training and segmentation.
[0070] As an example of a specific step: the feature selection network described in step 2 is constructed, the MedSAM large model is used as the basic model architecture, and the feature selection network is embedded into the MedSAM large model to integrate the image prior information and the large model processing capability, thereby improving the MedSAM large model's ability to select and extract abdominal organ features, such as Figure 3 As shown, the feature selection network extracts this information from the original CT image data and auxiliary feature data;
[0071] Construct a feature selection network based on UNet. The encoding path extracts image features through multiple convolution and pooling layers, while the decoding path gradually restores the original image size using deconvolution and skip connections. Given an image tensor As input, through the downsampling path, feature map fusion and upsampling path, the network outputs an image tensor As a feature map;
[0072] The network is used to extract the entropy and curvature information of the image; by improving the local entropy algorithm, it can more finely reflect the degree of distribution disorder of pixel values in the image, and use the complex statistical analysis of pixel values in the local area to more accurately capture the differences in texture features between different organs, providing a more powerful basis for distinguishing organs; curvature information extraction, combined with anatomical prior knowledge, uses the general law of organ contour curvature in anatomy, and uses a specific convolution kernel to more accurately capture the curvature characteristics of organ contours, especially in the complex convex and concave areas at the edges of abdominal organs, and more accurately depict the boundaries of organs; the feature selection network is embedded in the MedSAM large model, and the entropy and curvature information extracted from the feature selection network are weighted fused with the feature image generated by the image encoder from the MedSAM large model to improve the MedSAM model's ability to select and extract abdominal organ features;
[0073] For curvature feature selection, IFE uses simple linear convolution to obtain the 2D surface in Euclidean space R 3 The approximate mean curvature K in is as follows:
[0074] K=[K 1 K 2 K 3 ]*X......(1),
[0075] In formula (1), K 1 =[α,β,α] T , K 2 =[β,γ,β] T , K 3 =[α,β,α] T), where * represents convolution and X represents the input image;
[0076] In order to reflect the spatial and aggregation characteristics of intensity distribution, IFE defines information entropy as follows:
[0077]
[0078] In formula (2), i n represents the grayscale value of the center pixel of the nth 3×3 sliding window, j n represents the average gray value W of the remaining pixels in the window, and formula (2) calculates the gray value of the entire image (i n ,j n ) has a probability of occurrence P i,j , and E is expressed as information entropy. Each pixel corresponds to a grayscale value between 0 and 255. By selecting a certain proportion of channel features and combining them with the original features, IFE helps improve the performance with minor modifications to the segmentation network architecture, as shown in the following formula (3):
[0079] F=|Conv(X,k,b curv )|......(3),
[0080] In formula (3), X represents the feature map of the previous convolutional block, b curv It is the bias term in the convolution operation. The spatial dimensions (height and width) of the resulting tensor F are the same as those of X, but it highlights areas with specific features. k is a specific convolution kernel used to extract curvature features or entropy features.
[0081] Then the calculated curvature values are summed in the spatial dimension to compress the 2D feature map into a single scalar value. The result of this process is that each channel is represented as a scalar value that reflects the overall curvature strength of the channel. Through the above algorithm, the feature selection network extracts this information from the original CT image data and the auxiliary feature data;
[0082] The entropy and curvature information extracted from the feature selection network are fused with the feature image generated by the image encoder in the MedSAM large model using an adaptive weighted fusion mechanism. In this fusion process, an adaptive weighted fusion mechanism is used. The adaptive weighted fusion mechanism automatically assigns weights according to the prominence of different organ regions in the image. This is based on the statistical analysis of organ features in a large number of abdominal CT images and the learning of feature importance by the deep learning model. This mechanism can ensure the accurate combination of information and make the important organ features more prominent in the fused information. When the newly extracted entropy and curvature information is fused with it, the new information combination will provide a more comprehensive and valuable basis for subsequent segmentation reasoning.
[0083] As an example of a specific step: the construction of the LoRA layer in step 3 is based on the image encoder of the MedSAM large model, and a learnable LoRA layer is added to each Transformer layer to achieve efficient parameter training. For the mask decoder of the MedSAM large model, the LoRA layer is also added to the self-attention and cross-attention layers; the update of LoRA is as shown in formula (4) and formula (5):
[0084] W 0 +ΔW=W 0 +BA......(4),
[0085] Formula (4) represents a matrix decomposition, where B is a d×r dimensional matrix, A is a d×k dimensional matrix, and the rank r is much smaller than the minimum value of d and k. For each layer in the MedSAM large model, during the training process, the weight matrix W of the layer is o remains unchanged and does not receive gradient updates, while A and B contain trainable parameters, and ΔW is W o The low-rank update matrix is used to calculate the fixed weight h imposed by this layer on the input vector x, as shown in equation (5):
[0086] h=W 0 x+ΔWx=W 0 x+BAx......(5),
[0087] Formula (5) indicates that LoRA only updates a small part of the parameters during fine-tuning, which reduces the computational cost and alleviates the storage limitation.
[0088] As an example of a specific step: Figure 4As shown, step 3 describes LoRA fine-tuning: LoRA is a model fine-tuning method. By adding a low-rank matrix on the basis of the pre-trained model, it learns the parameter update of a specific task to reduce the training parameters and the amount of calculation, and realizes rapid adaptation to new tasks. In the abdominal multi-organ segmentation, according to the characteristics of the abdominal CT images and the task requirements of the organ segmentation, the appropriate low-rank matrix size and parameter initialization strategy are set in LoRA to achieve the performance enhancement of the pre-trained basic model with minimal computational input. As a method for efficient parameter fine-tuning, LoRA is integrated into the image encoder and mask decoder of the MedSAM large model. For the image encoder, LoRA adapts to the specific features of the abdominal CT image data by adjusting its low-rank matrix. By supervising the abdominal image The dataset FLARE2022 is efficiently fine-tuned using the above method. Through the stratified sampling strategy, the dataset is stratified according to factors such as organ size, position, and complexity, so that the model can learn the characteristics of different types of organs more evenly, avoiding the deviation of the model in the segmentation of these organs due to the distribution characteristics of certain organs. In terms of the mask decoder, LoRA also optimizes its parameters to better handle the mask generation in the segmentation task. Only the parameters of the low-rank matrices A and B introduced by LoRA are updated, and the rest of the pre-trained weights of the MedSAM large model are kept relatively stable. The stratified sampling strategy is adopted to improve the accuracy of the model in the abdominal CT multi-organ segmentation task. On this basis, the training process is further optimized to meet the special needs of abdominal CT segmentation.
[0089] In the image encoder, the parameters of the Transformer layer are frozen, and the LoRA layer is added next to the query and key projection layers in the attention mechanism. The LoRA layer consists of two linear layers, which operate in the image encoder as shown in Equations (6)-(9):
[0090]
[0091] Q=W q T+B q A q T......(7),
[0092] K=W k T......(8),
[0093] V=W v T+B v A v T......(9),
[0094] In formula (6) to formula (9), represents the primitives from the fully connected layer, Q, K, V represent the query, key and value layers respectively, and W q , W k and W vis a frozen projection layer from the MedSAM large model, and A q , B q , A v and B v is the trainable parameter matrix in the LoRA linear layer;
[0095] The supervised abdominal image dataset FLARE2022 is used for efficient fine-tuning. In terms of image encoder and mask decoder, LoRA also optimizes its parameters to better handle mask generation in segmentation tasks. Only the parameters of the low-rank matrices A and B introduced by LoRA are updated, and the rest of the pre-trained weights of the MedSAM large model are kept relatively stable. The accuracy of the model in the abdominal CT multi-organ segmentation task is improved by adopting a stratified sampling strategy. On this basis, the training process is further optimized to adapt to the special needs of abdominal CT segmentation.
[0096] As an example of a specific step: Step 4 is the vertical domain large model reasoning optimization of the fused prompt words: the trained model and the set reasoning strategy are integrated into an automated workflow. When a new abdominal CT image needs to be segmented, the image is first preprocessed and then input into the model. The model performs reasoning based on the contextual relationship between the mixed prompt words and adjacent layers, automatically outputs the segmentation results of each organ, and saves the segmentation results in the form of an image mask to facilitate subsequent medical analysis and visualization.
[0097] Automated segmentation: Integrate the trained model and the set reasoning strategy into an automated workflow. When a new abdominal CT image needs to be segmented, the image is first preprocessed and then input into the model. The model automatically outputs the segmentation results of each organ based on mixed cues and contextual reasoning between adjacent layers. The segmentation results can be saved in the form of image masks to facilitate subsequent medical analysis and visualization:
[0098] First, select the middle layer x of the CT image sequence t Perform segmentation, which includes all target organs;
[0099] Then, for x t Apply histogram equalization to enhance its clarity and calculate x t The cumulative distribution function is used to map each pixel value to obtain a new pixel value, as shown in the following equation (10):
[0100]
[0101] In formula (10), s k is the new pixel value, L∈[0,255] represents the grayscale level, g jis the probability of the jth gray level. After denoising the image using a median filter, the region growing method is used for initial organ segmentation. For each target organ, a seed is placed within its region to start the segmentation process;
[0102] Finally, the intermediate layer x is smoothed using morphological operations t The closing operation smoothes the edge of the object and fills the small holes in the object by dilating and then corroding the edge of the segmentation result, making the result more coherent, accurate and continuous, and providing detailed guidance for the segmentation of adjacent layers, as shown in the following formula (11):
[0103]
[0104] In formula (11), represents the expansion operation, Represents the corrosion operation, A is the image, B is the structural element, and the intermediate layer x is completed. t After region growing segmentation, the result m t Used as a mask cue for adjacent layers;
[0105] By fusing the bounding box prompt with the mask prompt of the previous layer's segmentation result, a hybrid prompt is constructed to complete the vertical domain large model reasoning process of the fused prompt word. The box prompt provides initial position information by drawing a rectangular box surrounding the organ on the image to guide the model to focus on a specific area. For each organ in the abdominal CT image, a corresponding rectangular box is drawn on the image according to its approximate position and size. The mask prompt provides a rough organ mask as the model's prior knowledge to help the model better understand the shape and distribution of the organ. The combination of the two can accurately locate and segment the organ. When performing organ segmentation, the contextual information between adjacent layers of the abdominal CT image is considered, because the organ structure of adjacent layers is continuous and correlated. This contextual information is used to improve the accuracy of segmentation. According to the difference between the current segmentation result and the predefined anatomical model, the subsequent segmentation strategy is adjusted in real time. This dynamic feedback mechanism can correct possible segmentation deviations in time to ensure that the segmentation process proceeds in the right direction. The feature relationship between adjacent layers is captured by introducing convolutional layers or recurrent neural network layers in the model.
[0106] As an example of a specific step: the automated segmentation of CT images in step 4 also includes: adjusting the subsequent segmentation strategy in real time according to the difference between the current segmentation result and the predefined anatomical model, correcting possible segmentation deviations, and ensuring the accuracy and completeness of the segmentation process.
[0107] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that technicians in the relevant technical field can better understand and utilize the present invention.
Claims
1. The abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model includes the following steps: Step 1, preprocessing the abdominal CT image supervised learning dataset; Step 2, feature selection network construction; Step 3, construction of LoRA layer; Step 4: Use a hybrid prompt strategy to perform inter-layer context reasoning to achieve automated abdominal CT segmentation.
2. According to the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of claim 1, the abdominal CT image supervised learning data set preprocessing described in step 1: collect abdominal CT image data sets from medical imaging databases or relevant medical institutions to ensure that the images cover samples of different patients, different pathological conditions and different scanning parameters to enhance the generalization ability of the model, and use the FLARE2022 data set or a similar supervised abdominal CT image data set, which does not require manually annotated segmentation masks. The specific steps are: Step 1.1 Development adjustment operation: by adjusting the window width and window level, according to the tissue characteristics of the abdominal organs and the grayscale distribution characteristics of the image, select appropriate window width and window level values, and adjust the image to highlight the development of the abdominal organ area; Step 1.2 Histogram equalization: Perform histogram equalization on the image after adjusting the window width and window position. By redistributing the grayscale of the image, the grayscale histogram of the image is made more uniform, the overall contrast of the image is enhanced, the impact of factors such as uneven illumination is reduced, the detailed information of the organ is highlighted, and better image quality is provided for subsequent segmentation tasks. Step 1.3 Image feature enhancement based on Sobel operator: Use the Sobel operator to perform edge detection on the image. The Sobel operator is a discrete differential operator that can calculate the grayscale gradient amplitude and direction of each pixel in the image. In abdominal CT images, the Sobel operator can enhance the edge information of organs and help to segment organ boundaries more accurately. In the specific operation, the horizontal and vertical gradients are calculated separately, and then the edge pixels are determined according to the gradient amplitude and direction, providing clearer edge features for subsequent model training and segmentation.
3. According to the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of claim 1, the feature selection network construction in step 2 includes the following steps: Step 3.1: The MedSAM large model is used as the basic model architecture, and the feature selection network is embedded into the MedSAM large model to integrate the image prior information and the large model processing capability, thereby improving the MedSAM large model's ability to select and extract abdominal organ features. The feature selection network extracts this information from the original CT image data and auxiliary feature data. Step 3.2 builds a feature selection network based on UNet. The encoding path extracts image features through multiple convolution and pooling layers, while the decoding path gradually restores the original image size using deconvolution and skip connections. Given an image tensor As input, through the downsampling path, feature map fusion and upsampling path, the network outputs an image tensor As a feature map; Step 3.3 uses the network to extract the entropy information and curvature information of the image. By improving the local entropy algorithm, it can better reflect the degree of distribution disorder of pixel values in the image, and use the complex statistical analysis of pixel values in the local area to capture the differences in texture features between different organs, providing a more powerful basis for distinguishing organs; curvature information extraction combines anatomical prior knowledge, uses the law of organ contour curvature in anatomy, and uses a specific convolution kernel to capture the curvature characteristics of the organ contour. In the complex convex and concave areas at the edge of the abdominal organ, it can depict the organ boundary; The feature selection network is embedded into the MedSAM large model, and the entropy and curvature information extracted from the feature selection network are weighted fused with the feature image generated by the image encoder from the MedSAM large model to improve the MedSAM model's ability to select and extract abdominal organ features. Step 3.4 For curvature feature selection, IFE uses simple linear convolution to obtain the 2D surface in Euclidean space R 3 The approximate mean curvature K in is as follows: K=[K1 K2 K3]*X......(1), In formula (1), K1 = [α, β, α] T , K2=[β,γ,β] T , K3=[α,β,α] T ), where * represents convolution and X represents the input image; In order to reflect the spatial and aggregation characteristics of intensity distribution, IFE defines information entropy as follows: In formula (2), i n represents the grayscale value of the center pixel of the nth 3×3 sliding window, j n represents the average gray value W of the remaining pixels in the window, and formula (2) calculates the gray value of the entire image (i n ,j n ) has a probability of occurrence P i,j , and E is expressed as information entropy, where each pixel corresponds to a grayscale value between 0 and 255. By selecting the corresponding proportion of channel features and combining them with the original features, IFE helps improve the performance with minor modifications to the segmentation network architecture, as shown in Equation (3): F=|Conv(X,k,b curv )|......(3), In formula (3), X represents the feature map of the previous convolutional block, b curv It is the bias term in the convolution operation. The spatial dimension of the resulting tensor F is the same as that of X, but it highlights areas with specific features. k is a specific convolution kernel used to extract curvature features or entropy features. Step 3.5 sums the calculated curvature values in the spatial dimension to compress the 2D feature map into a single scalar value. The result of this process is that each channel is represented as a scalar value that reflects the overall curvature strength of the channel. Through the above algorithm, the feature selection network extracts this information from the original CT image data and the auxiliary feature data; In step 3.6, the entropy and curvature information extracted from the feature selection network is fused with the feature image generated by the image encoder in the MedSAM large model using an adaptive weighted fusion mechanism. In this fusion process, an adaptive weighted fusion mechanism is used. The adaptive weighted fusion mechanism automatically assigns weights according to the prominence of different organ regions in the image. The basis for this mechanism is the statistical analysis of organ features in a large number of abdominal CT images and the learning of feature importance by the deep learning model. When the newly extracted entropy and curvature information is fused with it, the new information combination will provide a basis for subsequent segmentation reasoning.
4. According to the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model according to claim 1, the construction of the LoRA layer in step 3 specifically includes: Based on the image encoder of the MedSAM large model, a learnable LoRA layer is added to each Transformer layer to achieve efficient parameter training. For the mask decoder of the MedSAM large model, the LoRA layer is also added to the self-attention and cross-attention layers; the update of LoRA is as shown in equations (4) and (5): W0+ΔW=W0+BA......(4), Formula (4) represents a matrix decomposition, where B is a d×r dimensional matrix, A is a d×k dimensional matrix, and the rank r is much smaller than the minimum value of d and k. For each layer in the MedSAM large model, during the training process, the weight matrix W of the layer is o remains unchanged and does not receive gradient updates, while A and B contain trainable parameters, and ΔW is W o The low-rank update matrix of , and then, the fixed weight h imposed by this layer on the input vector x is calculated as follows (5); h=W0x+ΔWx=W0x+BAx......(5), Formula (5) indicates that LoRA only updates a small part of the parameters during fine-tuning, which reduces the computational cost and storage constraints.
5. According to the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of claim 1, the LoRA fine-tuning in step 3 specifically includes: As a model fine-tuning method, LoRA adds a low-rank matrix on the basis of the pre-trained model to learn the parameter update of specific tasks, so as to reduce the training parameters and the amount of calculation, and achieve rapid adaptation to new tasks. In the abdominal multi-organ segmentation, according to the characteristics of abdominal CT images and the task requirements of organ segmentation, the appropriate low-rank matrix size and parameter initialization strategy are set in LoRA to achieve the performance enhancement of the pre-trained basic model with minimal computational input. As a method for efficient parameter fine-tuning, LoRA is integrated into the image encoder and mask decoder of the MedSAM large model. For the image encoder, LoRA adapts to the specific characteristics of the abdominal CT image data by adjusting its low-rank matrix. By supervising the abdominal image dataset FLAR E2022 is fine-tuned using the above method. Through the stratified sampling strategy, the data set is stratified according to factors such as organ size, location, and complexity, so that the model can learn the characteristics of different types of organs more evenly and avoid the deviation of the model in the segmentation of these organs due to the distribution characteristics of certain organs. In terms of the mask decoder, LoRA also optimizes its parameters to better handle the mask generation in the segmentation task. Only the parameters of the low-rank matrices A and B introduced by LoRA are updated, and the rest of the pre-trained weights of the MedSAM large model are kept relatively stable. The accuracy of the model in the abdominal CT multi-organ segmentation task is improved by adopting a stratified sampling strategy. On this basis, the training process is further optimized to meet the special needs of abdominal CT segmentation. In the image encoder, the parameters of the Transformer layer are frozen, and the LoRA layer is added next to the query and key projection layers in the attention mechanism. The LoRA layer consists of two linear layers. Its operation in the image encoder is as shown in Equations (6)-(9): Q=W q T+B q A q T......(7), K=W k T......(8), V=W v T+B v A v T......(9), In formula (6) to formula (9), represents the primitives from the fully connected layer, Q, K, V represent the query, key and value layers respectively, and W q , W k and W v is a frozen projection layer from the MedSAM large model, and A q , B q , A v and B v is the trainable parameter matrix in the LoRA linear layer; The supervised abdominal image dataset FLARE2022 is used for efficient fine-tuning. In terms of image encoder and mask decoder, LoRA also optimizes its parameters to better handle mask generation in segmentation tasks. Only the parameters of the low-rank matrices A and B introduced by LoRA are updated, and the rest of the pre-trained weights of the MedSAM large model are kept relatively stable. The accuracy of the model in the abdominal CT multi-organ segmentation task is improved by adopting a stratified sampling strategy. On this basis, the training process is further optimized to adapt to the special needs of abdominal CT segmentation.
6. According to the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of claim 1, the vertical domain large model reasoning optimization of the fusion prompt words in step 4 includes: The trained model and the set reasoning strategy are integrated into an automated workflow. When a new abdominal CT image needs to be segmented, the image is first preprocessed and then input into the model. The model performs reasoning based on the contextual relationship between the mixed prompt words and adjacent layers, automatically outputs the segmentation results of each organ, and saves the segmentation results in the form of image masks to facilitate subsequent medical analysis and visualization. Automated segmentation includes: integrating the trained model and the set reasoning strategy into an automated workflow. When a new abdominal CT image needs to be segmented, the image is first preprocessed and then input into the model. The model automatically outputs the segmentation results of each organ based on mixed cues and contextual reasoning between adjacent layers. The segmentation results can be saved in the form of image masks to facilitate subsequent medical analysis and visualization: First, select the middle layer x of the CT image sequence t Perform segmentation, which includes all target organs; Then, for x t Apply histogram equalization to enhance its clarity and calculate x t The cumulative distribution function is used to map each pixel value to obtain a new pixel value, as shown in the following equation (10): In formula (10), s k is the new pixel value, L∈[0,255] represents the grayscale level, g j is the probability of the jth gray level. After denoising the image using a median filter, the region growing method is used for initial organ segmentation. For each target organ, a seed is placed within its region to start the segmentation process; Finally, the intermediate layer x is smoothed using morphological operations t The closing operation smoothes the edge of the object and fills the small holes in the object by dilating and then corroding, providing detailed guidance for the segmentation of adjacent layers, as shown in the following formula (11): In formula (11), represents the expansion operation, Represents the corrosion operation, A is the image, B is the structural element, and the intermediate layer x is completed. t After region growing segmentation, the result m t Used as a mask cue for adjacent layers; By fusing the bounding box prompt with the mask prompt of the previous layer's segmentation result to build a hybrid prompt, the vertical domain large model reasoning process of the fused prompt word is completed. The box prompt provides initial location information by drawing a rectangular box surrounding the organ on the image to guide the model to focus on a specific area. For each organ in the abdominal CT image, a corresponding rectangular box is drawn on the image according to its approximate position and size. The mask prompt provides a rough organ mask as the model's prior knowledge to help the model understand the shape and distribution of the organ. The combination of the two can locate and segment the organ. When performing organ segmentation, the contextual information between adjacent layers of the abdominal CT image is considered, because the organ structures of adjacent layers are continuous and relevant, and this contextual information is used to improve the accuracy of segmentation.
7. According to the abdominal CT multi-organ segmentation method based on the LoRA vertical domain MedSAM large model of claim 1, the CT image automatic segmentation in step 4 further includes: According to the difference between the current segmentation result and the predefined anatomical model, the subsequent segmentation strategy is adjusted in real time to correct possible segmentation deviations and ensure the accuracy and completeness of the segmentation process.
Citation Information
Cited By
Robust polyp segmentation method based on improved SAM-Med2D
CN120510169A
Construction method and application of abdominal organ segmentation SAM model based on prompt enhancement
CN120635083A
Construction method and application of abdominal organ segmentation sam model based on prompt enhancement
CN120635083B
Low-rank adaptation-based few-sample coronal mass ejection segmentation method and system
CN121616830A