Medical image target segmentation and three-dimensional visualization method based on DFL-MedSAM2
Through the combination of the improved MedSAM2 deep learning model and VTK three-dimensional visualization technology, the problems of low medical image segmentation accuracy and poor three-dimensional visualization effect are solved, and high-precision local organ/lesion segmentation and clear three-dimensional visualization effects are achieved.
Patent Information
- Application Number
- CN202510259906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
AI Technical Summary
Existing medical image segmentation algorithms are sensitive to image noise and are difficult to capture complex texture features, resulting in low segmentation accuracy of lesions with blurred boundaries or irregular morphology. Traditional three-dimensional visualization technology cannot effectively remove interference, affecting the clarity of local details.
The improved MedSAM2 deep learning model (DFL-MedSAM2) combined with VTK three-dimensional visualization technology is used to achieve high-precision local organ/lesion segmentation through the deep learning segmentation module, and the target area is highlighted through the three-dimensional visualization module to eliminate irrelevant interference.
High-precision segmentation and three-dimensional visualization of local organs or lesions in medical images are achieved, which meets the clinical need for attention to local details and improves segmentation accuracy and visualization effects.
Smart Images

Figure CN120182306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing and three-dimensional visualization, and in particular to a method for segmenting and visualizing local organs / lesions in medical images based on the DFL-MedSAM2 deep learning model and VTK three-dimensional visualization. Background Art
[0002] In the field of medical image analysis, accurately segmenting lesions is crucial for clinical diagnosis. Traditional segmentation algorithms rely on manually set gradient features and shape priors, and there are two main limitations: on the one hand, these methods are sensitive to image noise and prone to boundary leakage; on the other hand, they are difficult to effectively capture complex texture features in images. This makes the traditional methods have low segmentation accuracy when dealing with lesions with blurred boundaries or irregular shapes, thus affecting the accuracy of clinical decision-making.
[0003] In recent years, deep learning technology has made significant progress in the field of medical image segmentation. However, existing deep learning models usually only focus on the accuracy of segmentation results and lack support for three-dimensional visualization of segmentation results.
[0004] In addition, although traditional three-dimensional visualization technologies (such as VTK) can achieve three-dimensional model reconstruction of medical images, they rely on the original image data and cannot effectively remove the interference of irrelevant parts, resulting in the visualization results being difficult to focus and unable to meet the clinical attention requirements for local details. For example, in the three-dimensional reconstruction of lung CT images, high-density bones and blood vessels often obscure lung tumors, affecting the clarity of the tumor area and making it difficult for doctors to clearly observe the morphology and boundaries of tumors.
[0005] Therefore, there is an urgent need for a method that can combine deep learning segmentation technology and three-dimensional visualization technology to achieve accurate segmentation and local visualization of local organs or lesions in medical images. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for segmenting and visualizing local organs / lesions in medical images based on an improved MedSAM2 deep learning model and VTK three-dimensional visualization. Through the improved MedSAM2 model - DFL-MedSAM2, high-precision segmentation is achieved, and combined with VTK three-dimensional visualization technology, the target area is highlighted and the interference of irrelevant parts is excluded, so as to meet the clinical attention requirements for local details.
[0007] The technical solution is as follows:
[0008] The technical solution of the present invention includes the following modules and steps:
[0009] Data Processing Module (DPT):
[0010] Responsible for the input of original medical images (such as CT, MRI, etc.), the selection of segmentation dimensions, and format conversion.
[0011] Deep learning segmentation module (DFL-MedSAM2):
[0012] Use the DFL-MedSAM2 model to segment the target organ or lesion.
[0013] First, use the automatic segmentation model of deep learning to segment the medical image and extract specific organs or regions. Assume the input image is I(x, y, z), which passes through the segmentation model M seg Output the binary segmentation result S(x, y, z), where:
[0014] S(x, y, z) = M seg (I(x, y, z))
[0015] The segmentation result S(x, y, z) contains the pixel labels of the region to be reconstructed. Each pixel value S(x, y, z) ∈ {0, 1}, indicating the presence or absence of the target tissue, and is used for further reconstruction operations.
[0016] The DFL-MedSAM2 model framework includes:
[0017] Image Encoder: Extract the features of the input image.
[0018] Memory Attention: Construct local and global spatial information through multiple self-attention and cross-attention mechanisms.
[0019] Mask Decoder: Predict the segmentation result of the current frame based on the prompt input and previous outputs.
[0020] Memory Encoder: Convert the segmentation result of the current frame into high-dimensional features and store them in the memory bank to support the segmentation prediction of subsequent frames.
[0021] Improved loss function: For the problems of class imbalance and difficult-to-classify regions in medical images, use a combined loss function (CombinedLoss) of DiceLoss and Focal Loss instead of binary cross-entropy loss to improve the segmentation accuracy and training efficiency of the model.
[0022] Data matching and fusion module (DMF):
[0023] Perform format conversion, dimension judgment, interpolation scaling, and point-by-point multiplication operations on the segmentation results generated by DFL-MedSAM2.
[0024] Three-dimensional Visualization Module (3DVM):
[0025] By introducing a bounding box component and an Extractor class, real-time cropping and interaction of the three-dimensional model are achieved, avoiding unnecessary resampling operations and improving processing efficiency.
[0026] Adjust the inherent properties (such as transparency) of the volume rendering method to highlight the detailed features of the target area. Brief Description of the Drawings
[0027] Figure 1 It is the overall flowchart of the local organ / lesion segmentation and visualization method of the present invention;
[0028] Figure 2 It is an example of the segmentation result of the DFL-MedSAM2 model;
[0029] Figure 3 It is an example of the fusion of the segmentation result and the original data;
[0030] Figure 4 It is an example of the three-dimensional visualization result before lung segmentation;
[0031] Figure 5 It is an example of the three-dimensional visualization result after lung segmentation. Detailed Description of the Preferred Embodiments
[0032] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0033] Step 1: Data Processing Module (DPT)
[0034] The original data obtained by medical scanning devices is usually in the DICOM or Nii data format, while the segmentation model usually does not directly receive inputs in the form of DICOM or Nii files (pixel-level segmentation). Therefore, it is necessary to convert DICOM or Nii into jpg to successfully connect to the input of the segmentation model.
[0035] Taking the nii file as an example, converting each slice layer of the nii file layer by layer along the z-axis (assuming the z-axis is the top view, i.e., the horizontal plane or cross-section) into jpg is beneficial to improving the segmentation accuracy. The main reasons are as follows:
[0036] First, anatomical consistency: The cross-section (z-axis plane) is the acquisition direction of CT and MRI images. Segmentation layer by layer along the z-axis, which conforms to the anatomical cross-sectional structure, helps the segmentation model accurately identify the morphology of different organs or tissues. Secondly, resolution and spacing: In CT or MRI scans, the resolution of the x-axis and y-axis is usually high, while the interlayer spacing of the z-axis is relatively large. Due to the low information density in the z-axis direction, directly segmenting along the z-axis can avoid information loss or confusion caused by the large interlayer spacing. The model can obtain a clear cross-sectional structure in the z-axis plane and avoid being affected by insufficient interlayer information on the x or y axis. Finally, model adaptability: Deep learning models commonly use 2D segmentation methods in medical image segmentation. This method is applied layer by layer in the z-axis direction, training the model using the information of each cross-section, reducing the dependence on three-dimensional structure information. This shows good performance in the segmentation accuracy of each layer and also helps the model achieve good segmentation results in the case of insufficient training data.
[0037] Step 2: Deep learning segmentation module (DFL-MedSAM2)
[0038] To improve the performance of the MedSAM2 segmentation model in medical images, especially for the problems of class imbalance and difficult-to-classify regions, Dice Loss and Focal Loss are combined, and a new loss function CombinedLoss is used. The ratio of Dice Loss to Focal Loss is 1:20 to be applicable to multi-class problems, such as the segmentation of small objects like tumors and organs.
[0039] The formula is as follows:
[0040] L Combined = λ1·L Dice + λ2·L Focal
[0041]
[0042] L Focal = -α(1 - p t ) γ log(p t )
[0043] Among them,
[0044] λ1 and λ2 are weight coefficients.
[0045] y i is the ground truth label (0 or 1), is the predicted label.
[0046] p t is the predicted probability of the model for the correct class.
[0047] α is the balance factor that controls the weights between positive and negative samples.
[0048] γ is the focusing factor. When γ > 0, it increases the attention to difficult-to-classify samples.
[0049] Load the DFL-MedSAM2 model to segment the preprocessed medical images.
[0050] The model output is a binary mask image, where the target area is 1 and the background area is 0.
[0051] Step 3: Data Matching and Fusion Module (DMF)
[0052] Multiply the segmentation mask generated by DFL-MedSAM2 with the original medical image pixel by pixel to generate local image data containing only the target area.
[0053] After obtaining the segmentation mask S(x, y, z), we need to match it with the original medical image data V(x, y, z) to ensure that the segmented mask can correctly reflect on the original data. The key to this step is to solve the alignment problem of the segmentation result. Since the resolution or size of the segmentation mask and the original data may be different, interpolation processing is required to make the mask coincide with the original data in the same spatial dimension.
[0054] For this step, the segmentation result is spatially aligned through scale transformation and interpolation methods. Assume the size of the original data is (R x , R y , R z ), and the size of the segmentation mask is R′ x , R′ y , R′ z . Adjust S(x, y, z) through scale transformation.
[0055] Formula:
[0056] S′(x′, y′, z′) = Interpolate(S(x, y, z), (R x , R y , R z ))
[0057] where Interpolate represents the interpolation operation to make the segmentation mask match the original image data spatially.
[0058] The trilinear interpolation method is used for interpolation. After registration and interpolation operations, the segmentation result is combined with the original image to obtain the application mask M for reconstruction apply , and its expression is as follows:
[0059] M apply (x, y, z) = I(x, y, z) · S(x, y, z)
[0060] Where S(x, y, z) is the binary mask obtained by segmentation, and I(x, y, z) is the original image after registration. Through element-wise multiplication operation, it is ensured that the intensity value of each pixel is multiplied by its corresponding segmentation mask, thereby generating a new three-dimensional dataset. This step ensures that the segmented organ not only maintains the structural outline but also has internal image details.
[0061] Step 4: Three-dimensional visualization module (3DVM)
[0062] Use VTK (Visualization Toolkit) to perform three-dimensional reconstruction on the fused local image data, use the ray casting method for volume rendering, introduce a bounding box to implement the cutting of the target object, and at the same time set the rendering parameters property (such as color, transparency, etc.) to render the target area.
[0063] The reconstructed three-dimensional model is subjected to advanced visualization processing through VTK (Visualization Toolkit), aiming to improve the user's ability to understand and analyze medical images. Users can freely rotate, zoom, and view different slices and cross-sections (implemented through vtkRenderer and vtkRenderWindowInteractor) to observe the structural and functional information of the organ from multiple angles, and can emphasize the tissue characteristics corresponding to different intensity values by adjusting the transparency and coloring properties of the model.
[0064] In CT imaging, the intensity value is expressed in Hounsfield units (HU), and the HU value ranges of different tissues are obvious. For example, air is about -1000 HU, water is 0 HU, muscle is between +50 and +100 HU, and bone can be as high as +3000 HU.
[0065] The commonly used HU ranges are:
[0066] Air: about -1000 HU
[0067] Water: 0 HU
[0068] Muscle: about +50 to +100 HU
[0069] Bone: about +300 to +3000 HU
[0070] This enables users to highlight high-density tissues (such as bone) or low-density tissues (such as fat) by setting appropriate transparency to better observe the relationship between specific tissues and their adjacent structures.
[0071] Through the interactive functions of VTK, it supports users to perform operations such as rotating, scaling, adjusting transparency, and color mapping on 3D models.
[0072] Step 5: Result output and saving
[0073] Save the 3D visualization results as image or video files, and save the 3D model in.vtk or.stl format.
[0074] Save the video in MP4 format.
[0075] Save the segmentation mask and 3D model data for subsequent analysis and use.
[0076] Example
[0077] Taking lung tumor segmentation and visualization as an example:
[0078] Input chest CT images and perform preprocessing.
[0079] Use the DFL-MedSAM2 model to segment lung tumors and generate binary masks.
[0080] Fuse the segmentation mask with the original CT image to generate local image data containing only the tumor area.
[0081] Use VTK to perform 3D reconstruction on the tumor area to generate a 3D model of the tumor.
[0082] Through the visualization function of VTK, render and interact with the tumor model. Users can adjust the transparency through the slider to observe the relationship between the tumor and the surrounding tissues in real time and highlight the morphological features of the tumor.
Claims
1. A method for segmenting and visualizing local organs / lesions in medical images based on an improved MedSAM2 deep learning model and VTK 3D visualization, and a deep learning network model DFL-MedSAM2, wherein the network model is characterized by using a loss function optimization combining FocalLoss and Dice Loss, and comprises the following steps: (a) Use the DPT module to preprocess the input medical image, including: Dimension adjustment: Set the cross section as the acquisition direction of CT and MRI images, and segment layer by layer on the z-axis to conform to the anatomical cross-sectional structure. Format conversion: Convert medical formats such as DICOM / NIfTI to the input format of the DFL-MedSAM2 model. (b) Target region segmentation The DFL-MedSAM2 model is used to segment the target organs or lesions in medical images and generate binary mask images. The DFL-MedSAM2 model solves the problems of category imbalance and difficult-to-classify areas through a combined loss function (Dice Loss+Focal Loss), significantly improving the segmentation accuracy. (c) Data fusion processing Pixel-by-pixel multiplication of the mask image and the original image; Generate a local image that preserves the anatomical information of the target area; Pixel value normalization is performed to keep data distribution consistent. (d) 3D visualization reconstruction Based on VTK toolkit: Use VTK (Visualization Toolkit) to perform 3D reconstruction on the fused local image data; Real-time cropping interaction technology is used to highlight the detailed features of the target area. (e) Output and storage Visualization result export: static image (PNG / TIFF) and dynamic video (MP4); Data archive: original image, segmentation mask and 3D mesh data (STL / PLY format).
2. The method according to claim 1, characterized in that The combined loss function (CombinedLoss) consists of DiceLoss and Focal Loss, where the weight of Dice Loss is 1 / 21 and the weight of Focal Loss is 20 / 21, and is used to solve the problems of category imbalance and difficult-to-classify areas in medical image segmentation.
3. The method according to claim 1, characterized in that Apply the DFL-MedSAM2 segmentation results to VTK for visualization, which includes the following steps: Convert the segmentation results into NIfTI medical format; Create a separate mapper for the original medical data and the segmentation results; Set color and opacity properties to generate a 3D model; It supports fusion display to facilitate observation of the overall position characteristics of local organs / lesions, and supports separate display to facilitate observation of the details of local organs / lesions.
4. The method according to claim 1, characterized in that The real-time clipping interaction technology introduces a bounding box component and an Extractor class to directly clip the area of interest extracted from the three-dimensional model, thereby improving processing efficiency.
5. The method according to claim 1, characterized in that The target organs or lesions include, but are not limited to, liver tumors, lung tumors, skull tissues, heart, kidneys and brain lesions.
6. The method according to claim 1, characterized in that The three-dimensional visualization result supports user interactive operations, including settings such as rotation, scaling, transparency adjustment, and color mapping.