Deep learning-based mediastinal medical image segmentation method and system
By collecting and processing custom data sets, using the MedSAM model for transfer learning and optimization, the accuracy of mediastinal tumor segmentation is solved, and efficient mediastinal lesions area detection and medical diagnosis assistance is achieved.
Patent Information
- Application Number
- CN202510332840.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
Existing medical image segmentation methods are difficult to accurately identify mediastinal tumors, especially because mediastinal tumor data is limited and most of them are labeled training data, and traditional deep learning methods are difficult to perform correct tissue and organ segmentation and identification.
Different types of medical image data sets are collected, data cleaning, preprocessing and annotation are performed, and the pre-trained MedSAM model is used for transfer learning. The model parameters are optimized through the target loss function, and combined with Dice loss and cross entropy loss, the semantic segmentation of mediastinal lesion areas is achieved.
It improves the accuracy and robustness of mediastinal tumor segmentation, realizes accurate detection of mediastinal lesions, and provides basic medical diagnosis auxiliary functions.
Smart Images

Figure CN120339299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and relates to, but is not limited to, a mediastinal medical image segmentation method and system based on deep learning. Background Art
[0002] Medical image segmentation is an important part of clinical practice, which helps with accurate diagnosis, treatment planning, and disease monitoring. Existing segmentation methods are often tailored to specific patterns or disease types, while mediastinal tumors have a relatively low incidence and a complex variety of types, making it very difficult to diagnose solely based on images.
[0003] Due to limited mediastinal tumor data, and most of the data being unlabeled training data, it is difficult for traditional deep learning methods to correctly segment and identify tissues and organs in images. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a mediastinal medical image segmentation method and system based on deep learning, which at least solves the problem of low recognition accuracy of existing medical image segmentation methods.
[0005] The technical solution of the embodiments of the present invention is implemented as follows:
[0006] In a first aspect, embodiments of the present invention provide a mediastinal medical image segmentation method based on deep learning, the method comprising:
[0007] Collect a custom dataset of different types of medical images, and perform data cleaning, preprocessing, and annotation to obtain a sample dataset; wherein, the different types of medical images include three-dimensional CT images of patients with germ cell tumors, lymph node tumors, neuromas, teratomas, and thymomas; the sample dataset includes 2D slices generated by cutting the three-dimensional CT images and corresponding mask files; divide the sample dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1; input the training set batches into a pre-trained MedSAM model, and obtain the predicted segmentation results of the voxels in the training samples through forward calculation; based on the predicted segmentation results and the prior labels corresponding to the mask files, determine the target loss, and update the parameters of the pre-trained MedSAM model in the reverse direction until convergence to obtain a medical image semantic segmentation model; use the validation set and the test set to evaluate the performance of the medical image semantic segmentation model.
[0008] In some possible embodiments, the preprocessing includes the following operations: cutting all three-dimensional CT images perpendicular to the z-axis through the nibabel library of Python, slicing images in the mediastinal window with dimensions from 50 to 350, removing objects with a small number of pixels in the two-dimensional plane, setting the image size to 1024 to obtain a complete 2D slice containing the lesion area, and generating a 2D pixel-level segmentation label using the tumor area annotated by clinicians as a mask file.
[0009] In some possible embodiments, the pre-trained MedSAM model includes an image encoder, a prompt encoder, and a mask decoder, where: the image encoder is used to extract key features from the image using a vision transformer; the vision transformer consists of 12 transformer layers, each transformer layer includes a multi-head self-attention mechanism and a multi-layer perceptron, including layer normalization; the prompt encoder is used to map the corners of the bounding box prompt to a 256-dimensional vector embedding; the mask decoder is used to integrate the image embedding, the prompt embedding, and the output token to produce an accurate segmentation result and the corresponding confidence metric; the mask decoder includes two transformer layers for fusing the image embedding and the prompt encoding, and two transposed convolutional layers for upsampling the resolution of the embedding to 256*256; subsequently, the embedding is processed by the sigmoid function and then aligned with the original input size through bilinear interpolation.
[0010] In some possible embodiments, the target loss is the unweighted sum of the dice loss and the cross-entropy loss, calculated by the following formula:
[0011]
[0012] L = L CE + L Dice ;
[0013] where L is the target loss, L CE is the cross-entropy loss, L Dice is the dice loss, s i and g i respectively represent the predicted segmentation result and the ground truth of voxel i, and N is the number of voxels in image I.
[0014] In a second aspect, an embodiment of the present invention provides a mediastinal medical image segmentation system based on deep learning. The mediastinal medical image segmentation system includes an image processing terminal, a user system, and a parameter adjustment system, where:
[0015] The image processing terminal is deployed with a semantic segmentation model obtained by transfer learning on a custom dataset based on the MedSAM model, which is used for doctors to log in, perform basic image processing on CT images, and execute the method described in the first aspect above to segment the image lesion sites of CT images;
[0016] The user system is used for user registration and user login to query the user diagnosis results fed back by doctors at the image processing terminal;
[0017] The parameter system is used for background management login to adjust the parameters of the semantic segmentation model and conduct data tests.
[0018] Further, the image processing terminal is specifically used to perform the following operations: pre-train the FLARE22 open-source dataset using the MedSAM model to obtain a pre-trained MedSAM model; perform data preprocessing and normalization operations on the custom dataset; fine-tune the parameters of the pre-trained MedSAM model and then use the processed custom data for training and conduct comparative experiments to obtain the semantic segmentation model; feedback the training weights of the semantic segmentation model to the front end, detect a single unknown data to obtain the segmentation and system diagnosis results, feed them back to the doctor for diagnosis, and finally feed the final results back to the user system.
[0019] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0020] In the embodiments of the present invention, a custom dataset containing different types of medical images is collected and data cleaning, preprocessing, and annotation are performed; then, a series of processed custom datasets are trained based on a pre-trained model so that it can perform semantic segmentation and detection on the mediastinal lesion area. At the same time, the deep learning medical image segmentation system based on MedSAM migration provided in the embodiments of the present invention will provide basic registration, login, system feedback, and basic medical diagnosis report functions for ordinary users, CT image processing and secondary diagnosis functions for doctor users, and parameter adjustment and system testing functions for administrators to ensure the auxiliary diagnosis effect. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0022] Figure 1 It is a flowchart of the method for mediastinal medical image segmentation based on deep learning provided in the embodiments of the present invention;
[0023] Figure 2 Schematic diagram of the prediction results of the medical image semantic segmentation model provided by the embodiments of the present invention;
[0024] Figure 3 Transfer learning algorithm framework based on MedSAM provided by the embodiments of the present invention;
[0025] Figure 4 System function framework of the mediastinal medical image segmentation system provided by the embodiments of the present invention;
[0026] Figure 5 System design task flow chart of the mediastinal medical image segmentation system provided by the embodiments of the present invention;
[0027] Figure 6 Schematic diagram of the lesion location test results for random images in the test set provided by the embodiments of the present invention;
[0028] Figure 7 Schematic diagram of the comparison results of the segmentation of random images in the validation set by different models provided by the embodiments of the present invention;
[0029] Figure 8 Visualization segmentation example diagram of random test data in the test set provided by the embodiments of the present invention;
[0030] Figure 9 Login interface of the mediastinal medical image segmentation system provided by the embodiments of the present invention;
[0031] Figure 10 Function interface of the mediastinal medical image segmentation system provided by the embodiments of the present invention. Detailed implementation manners
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0033] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0034] It should be noted that the terms "first", "second", and "third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0035] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.
[0036] Weakly Supervised Domain Adaptation refers to the situation where, in the target domain with only weak labels (such as image-level labels), the performance of the target domain is improved by leveraging the label information in the source domain. It is a form of transfer learning aimed at solving the problem of difficult or expensive data annotation in the target domain.
[0037] In weakly supervised domain adaptation, there is usually a source domain and a target domain. The source domain contains labeled data, while the target domain has only weakly labeled or unlabeled data. The goal is to train a model with good generalization performance in the target domain by leveraging the label information in the source domain.
[0038] Domain Adaptation Network: A domain adaptation network is a method that achieves adaptation by minimizing the inter-domain differences between the source domain and the target domain. It usually includes a shared feature extractor and two domain-specific classifiers, one for the source domain and the other for the target domain. By minimizing the differences between the two domain classifiers, feature alignment between the source domain and the target domain can be achieved.
[0039] Due to the limited data of mediastinal tumors and most of the data being unlabeled training data, it is difficult for traditional deep learning methods to correctly segment and identify tissue organs in images. To address this problem, the embodiments of the present invention plan to adopt a semantic segmentation model based on domain adaptation, and actually use MedSAM for transfer learning on a custom tumor dataset.
[0040] Figure 1 A schematic flow diagram of a mediastinal medical image segmentation method based on deep learning provided for the embodiments of the present invention, asFigure 1 As shown, the method at least includes the following steps:
[0041] Step S110, collect a custom dataset of different types of medical images, and perform data cleaning, preprocessing, and annotation to obtain a sample dataset.
[0042] Here, the different types of medical images include three-dimensional CT images of patients with germ cell tumors, lymph node tumors, neuromas, teratomas, and thymomas; this dataset is sourced from hospital data, which contains a large amount of messy patient data. In the embodiments of the present invention, the data of various categories are first distinguished, and the above five categories are divided according to doctor's orders and outpatient results to obtain a series of required data.
[0043] The sample dataset includes 2D slices generated by cutting the three-dimensional CT images and corresponding mask files. The original dataset contains segmentation labels of tumor regions marked by clinical professional doctors. In the embodiments of the present invention, 2D pixel-level segmentation labels corresponding to each 2D slice are obtained through these annotation information and set as mask files, which are used as the final standard for verifying and testing the segmentation performance of the model. To increase the richness of the data, a total of 5271 slices are obtained after cutting the three-dimensional CT images into 2D slices, and pixel-level segmentation labels are used to verify and test the segmentation performance of the model.
[0044] It should be noted that, on the basis of performing a series of processing on the custom dataset so that it can be directly input into the model, the embodiments of the present invention add text annotation data to the image data and perform dimensional splicing for multi-classification tasks.
[0045] Step S120, divide the sample dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0046] Here, the processed 2D chest slices are divided into a training set, a validation set, and a test set according to a ratio of 80%, 10%, and 10%. To ensure the rigor of the experiment and reduce the influence of random data selection on the experimental results, random data division is used for 5 repeated experiments, and the average result of the 5 experiments is taken as the final result.
[0047] Step S130, input the training set batch by batch into the pre-trained MedSAM model, and obtain the predicted segmentation results of the voxels in the training samples through forward calculation.
[0048] Here, the pre-trained MedSAM model is used to segment the mediastinal tumor region in the CT image. The obtained 2D slices are divided into two categories: lesion images and normal images according to whether there is a tumor region in the image. Among them, the pre-trained MedSAM model is trained on the officially designated FLARE22 open-source dataset.
[0049] Step S140: Based on the predicted segmentation result and the prior label corresponding to the mask file, determine the target loss, and update the parameters of the pre-trained MedSAM model in the reverse direction until convergence to obtain a medical image semantic segmentation model.
[0050] Here, the prior label corresponding to the mask file is a 2D pixel-level segmentation label determined based on the segmentation label of the tumor region annotated by clinical professional doctors.
[0051] The target loss is the unweighted sum between the Dice loss and the cross-entropy loss, which has been proven to be robust in various segmentation tasks. This network is optimized by the AdamW optimizer (β1 = 0.9, β2 = 0.999), with an initial learning rate of le -4 , a weight decay of 0.01. The batch size is 160, data augmentation is not used, and the last checkpoint is selected as the final model.
[0052] As Figure 2 shown in the schematic diagram of the prediction result of the medical image semantic segmentation model provided by the embodiment of the present invention. From left to right, they are the original image, the image with the segmentation mask predicted by the algorithm superimposed on the original image, and the ground truth mask image. It can be seen that using the medical image semantic segmentation model provided by the embodiment of the present invention can significantly improve the performance of the segmentation of mediastinal images, thereby improving the accuracy of auxiliary medical diagnosis.
[0053] Step S150: Evaluate the performance of the medical image semantic segmentation model using the validation set and the test set.
[0054] Here, in order to verify the segmentation effect and accuracy of the medical image semantic segmentation model provided by the embodiment of the present invention, outstanding methods in the field of semantic segmentation such as DeepLabV3+ and nnU-Net are selected as experimental comparison models.
[0055] In order to evaluate the segmentation performance of the model proposed in the embodiment of the present invention and the comparison model from different perspectives, the experiment selects a variety of evaluation metrics as the final measurement criteria: The semantic segmentation model is evaluated using the segmentation accuracy (Pixel-wise Accuracy) and the Jaccard similarity coefficient (Jaccard Similarity Coefficient). The semantic segmentation accuracy on the test set can reach 99.64%. Due to the small number of cases, it may decrease in practice. The number of tests is (1000 2D slices).
[0056] The segmentation accuracy is used to measure the accuracy of all pixels in the model's segmentation results. The accuracy score is defined as shown in formula (1), and the higher the value, the better the accuracy of the model prediction.
[0057]
[0058] Among them, correct predictions are all the data with correct predictions, and total predictions are all the predicted data.
[0059] The Jaccard similarity coefficient (Jaccard Score) measures the similarity and difference between finite sample sets, and the magnitude of its value is proportional to the similarity between samples, as shown in formula (2):
[0060]
[0061] Among them, i is a variable and n is the total number of samples. is the size of the intersection between sample set A i and sample set B i indicating the number of samples shared by the two sets. is the size of the union between sample set A i and sample set B i indicating the total number of samples in the two sets.
[0062] In the embodiments of the present invention, a custom dataset containing different types of medical images is collected and data cleaning, preprocessing, and annotation are performed; then, based on a pre-trained model, a series of processed custom datasets are trained to enable semantic segmentation and detection of the mediastinal lesion area. The present invention improves the accuracy and robustness of mediastinal tumor segmentation in general medical image segmentation.
[0063] In some possible embodiments, the preprocessing in the above step S110 includes the following operations: cutting all three-dimensional CT images perpendicular to the z-axis through the nibabel library of python, slicing the images in the mediastinal window with dimensions from 50 to 350, removing objects with a small number of pixels in the two-dimensional plane, setting the image size to 1024 to obtain complete 2D slices containing the lesion area, and generating 2D pixel-level segmentation labels using the tumor areas annotated by clinicians as mask files.
[0064] Here, all 2D slices are from different cases to avoid label leakage problems caused by similar slices from the same case being separately assigned to the training set and the test set. Subsequently, Python code is used to preprocess and transform the medical image data, convert the images in NIfTI format (nii.gz) to npz files, and save the original images and the corresponding mask files.
[0065] In some possible embodiments, the pre-trained MedSAM model includes an image encoder, a prompt encoder, and a mask decoder, where: the image encoder is used to extract key features from the image using a vision transformer; the vision transformer consists of 12 transformer layers, each transformer layer includes a multi-head self-attention mechanism and a multi-layer perceptron, which includes layer normalization; the prompt encoder is used to map the corners of the bounding box prompt to a 256-dimensional vector embedding; the mask decoder is used to integrate the image embedding, the prompt embedding, and the output token to generate an accurate segmentation result and the corresponding confidence metric; the mask decoder includes two transformer layers for fusing the image embedding and the prompt encoding, and two transposed convolutional layers for upsampling the resolution of the embedding to 256*256; subsequently, the embedding is processed by a sigmoid function and then aligned with the original input size through bilinear interpolation.
[0066] Here, the transfer learning algorithm framework based on MedSAM is as Figure 3 shown, including an input layer for receiving medical image data, a feature extraction layer for using a pre-trained network to extract image features, and a transfer learning layer for adjusting the parameters of the feature extraction layer to adapt to the new dataset. The segmentation network includes an encoder, a prompt encoder, and a mask decoder for generating segmentation results. Finally, the output layer provides the final segmented image and the corresponding confidence score.
[0067] In the context of medical image semantic segmentation, MedSAM, as a base model, is adjusted and optimized through transfer learning to adapt to specific tasks and datasets. To balance segmentation performance and computational efficiency, embodiments of the present invention adopt a basic ViT model as the image encoder. Specifically, the basic ViT model consists of 12 transformer layers, and each block includes a multi-head self-attention block and a multi-layer perceptron (MLP) block, which contains layer normalization. The model is pre-trained using masked autoencoder modeling and then fully supervised trained on the SAM dataset. The input image (1024*1024*3) is formatted into a sequence of flattened two-dimensional patches of size 16*16*3. Through the processing of the image encoder, the image is compressed into an embedding with a 64×64 feature map, and the size is reduced by sixteen times. The prompt encoder maps the corners of the bounding box prompt to a 256-dimensional vector embedding. Specifically, each bounding box is represented by an embedding pair of the upper left corner and the lower right corner. With the addition of the bounding box prompt, the prior label corresponding to each mask file is also added to complete the prompt for the multi-class model. To facilitate real-time user interaction after the image embedding calculation is completed, a lightweight mask decoder framework is integrated. The framework includes two transformer layers for fusing the image embedding and the prompt encoding, and two transposed convolutional layers to enhance the embedding resolution to 256*256. Subsequently, the embedding is processed through sigmoid activation and then aligned with the original input size through bilinear interpolation.
[0068] Initializing with a pre-trained MedSAM model can encode the bounding box prompt. The bounding box prompt is simulated by randomly perturbing 0-20 pixels on the ground-truth mask.
[0069] In some possible embodiments, the objective loss is the unweighted sum of the dice loss and the cross-entropy loss, and is calculated by the following formula:
[0070]
[0071]
[0072] L = L CE + L Dice ;
[0073] where L is the objective loss, L CE is the cross-entropy loss, L Dice is the dice loss, s i and g i represent the predicted segmentation result and the ground truth of voxel i respectively, and N is the number of voxels in image I.
[0074] Here, in the embodiments of the present invention, the target loss is defined as a direct combination of the Dice coefficient loss and the cross-entropy loss, and this combination has shown strong robustness in various segmentation attempts. The network is optimized by the AdamW optimizer (β1 = 0.9, β2 = 0.999), and the initial learning rate is le -4 , and the weight decay is 0.01. The batch size is defined as 160, and data augmentation is not used. The last checkpoint is selected as the final model.
[0075] Figure 4 is the system function framework of the mediastinal medical image segmentation system provided by the embodiments of the present invention. As Figure 4 shown, the system includes an image processing terminal, a user system, and a parameter adjustment system, where:
[0076] The image processing terminal is deployed with a semantic segmentation model obtained by performing transfer learning on a custom dataset based on the MedSAM model, and is used for doctors to log in and perform basic image processing on CT images and execute any of the above mediastinal medical image segmentation methods to segment the image lesion sites of CT images;
[0077] The user system is used for user registration and user login to query the user diagnosis results fed back by doctors at the image processing terminal;
[0078] The parameter system is used for background management login to adjust the parameters of the semantic segmentation model and perform data testing.
[0079] Here, the image processing terminal logged in by doctors can perform basic image processing on CT images; segment the image lesion sites; and assist doctors in diagnosing diseases. The system will provide basic registration, login, system feedback, and basic medical diagnosis report functions for ordinary users, CT image processing and secondary diagnosis functions for doctor users, and parameter adjustment and system testing functions for administrators to ensure the auxiliary diagnosis effect.
[0080] In some possible embodiments, the image processing terminal is specifically used to perform the following operations: pre-train the FLARE22 open-source dataset using the MedSAM model to obtain a pre-trained MedSAM model; perform data preprocessing and normalization operations on the custom dataset; perform parameter fine-tuning on the pre-trained MedSAM model and then use the processed custom data for training and conduct a comparative experiment to obtain the semantic segmentation model; feedback the training weights of the semantic segmentation model to the front end, detect a single unknown data to obtain the segmentation and system diagnosis results, feedback them to the doctor for diagnosis, and finally feedback the final results to the user system.
[0081] In the embodiments of the present invention, the MedSAM model is first used to perform transfer learning on a custom dataset, and the existing mediastinal case data is trained to obtain the training weights required by the front end. Then, a lightweight visual front-end page is built based on Flask and Vue, which is linked to the training weights and can be independently developed, tested, deployed, and operated and maintained to complete the business functions required by the project.
[0082] The following describes the above-mentioned mediastinal medical image segmentation method and system based on deep learning in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better explaining the present invention and does not constitute an improper limitation of the present invention.
[0083] As Figure 5 shown is the system design task flow chart provided by the embodiments of the present invention. First, the MedSAM model is used to perform pre-training on the official FLARE22 open-source dataset, and the pre-training weights are saved. At the same time, corresponding preprocessing work is carried out on the custom dataset. After parameter fine-tuning of the MedSAM model, training is performed, and a comparative experiment is conducted. After obtaining the training weights, testing work is carried out. On this basis, the back end is built with the Flask framework, and the front-end interface is built with vue to implement a lightweight system for user interaction. Specifically, the overall design of the system provided by the embodiments of the present invention includes the following four processes:
[0084] S1. Data collection and preprocessing: Collect a dataset containing different types of medical images, and perform data cleaning, preprocessing, and annotation.
[0085] It should be noted that creating a tumor segmentation training dataset includes the following key steps: 1) Data collection: Collect data provided by volunteers or released by institutions. 2) Data preprocessing: Screen the data that meets the requirements and use algorithms to adjust and correct it. 3) Image processing: Register and standardize the images to ensure the consistency of the data format. 4) Image annotation: Experts annotate each image to provide the required labels for training the model. In the present invention, our task is to create a personalized dataset specifically for CT mediastinal images. In this work, the two steps of data collection and image annotation have been completed by professional institutions and doctors. The focus of the research of the present invention is on the stages of data preprocessing and image processing.
[0086] To verify the segmentation performance of the model proposed in the present invention and compare it with the performance of other models, a custom dataset, the CT mediastinal image dataset, was created. The dataset was sourced from hospital data, which contained a large amount of messy patient data. In the embodiments of the present invention, the data of various categories were first distinguished and divided according to doctor's orders and outpatient results to obtain the required data, which were then divided into five categories as a whole, namely 3D CT images of patients with germ cell tumors, lymph node tumors, neuromas, teratomas, and thymomas.
[0087] In the embodiments of the present invention, the following data preprocessing was performed on all 3D images: First, all 3D images were cut perpendicular to the z-axis using the third-party library nibabel in Python. The images in the 50 - 350 dimension were sliced through the mediastinal window, and objects with a small number of pixels in the 2D plane were removed. The image size was set to 1024 to obtain complete 2D slices with lesion areas and corresponding mask images.
[0088] S2. Feature extraction and encoding: This part mainly designs a feature extraction method based on a deep learning model. The Transformer structure is introduced in the encoder stage of the MedSAM model to capture the global dependencies of the images and extract global features with high-level semantic information to determine the location and category of the target lesions.
[0089] S3. Design of medical image segmentation algorithm: Based on the extracted features, a medical image segmentation algorithm is designed.
[0090] S4. Comparative experiment: To verify the segmentation effect and accuracy of the present invention, DeepLabV3+ and nnU-Net were selected as the comparative methods in the embodiments of the present invention. The evaluation metrics DSC, JA, and AC were used for comparison, as shown in Table 1. The comparison results are as Figure 6 shown.
[0091] Table 1. Various evaluation criteria for the internal validation set
[0092] Model DSC JA AC The model of the present invention 0.9912 0.9674 0.9764 Deeplabv3+ 0.8011 0.7845 0.7347 Unet 0.7899 0.7644 0.7768 SAM 0.5077 0.5103 0.5324 MedSAM 0.9363 0.9444 0.9712
[0093] In the 10% test set, the predicted DSC was measured to be 99.14%. The test results of the lesion location of a randomly selected image are as Figure 6 shown. The left figure is the original 2D slice image; the middle figure is the predicted image, where the green part is the lesion area; the right figure is the mask image corresponding to the original slice image. According to the experimental results, it can be known that the present invention has achieved good segmentation results for the lesion location segmentation of chest mediastinal tumors, and the evaluation results on the corresponding test data have reached an ideal level.
[0094] The comparison results of the segmentation of random images by different models are shown asFigure 7 As shown, it is a visual segmentation example of the internal validation set. These three examples are Unet, DeepLabV3+, and Self-Model respectively. Magenta represents the segmentation result.
[0095] The visualization segmentation example diagram of randomly selected test data in the test set is as Figure 8 shown. Two are randomly selected. The blue in the figure is the segmentation box, which is convenient for finding the segmentation result in the color image. The light yellow inside the box is the segmentation result.
[0096] S4. System implementation and interface design: Build the basic architecture of the medical image segmentation system, including the development of modules such as image input, feature extraction, and segmentation, and design a friendly user interface, such as Figure 9 and Figure 10 shown.
[0097] Here, in terms of interface construction, use PyQt5 to build a graphical user interface (GUI), including image display, buttons, and interactive annotation tools. In terms of event handling, implement mouse event handling, such as clicking, moving, and releasing, to support users in drawing rectangular boxes on the image. In terms of model loading, load a pre-trained deep learning model (MedSAM) to process image data. In terms of image processing, use PyTorch for pre-processing and post-processing of images, including resizing images, normalizing, and converting data formats. In terms of annotation conversion, convert the rectangular boxes drawn by users with the mouse into a format acceptable to the model. In terms of model inference, use the above-mentioned medical image semantic segmentation model to segment the annotated area and generate a segmentation mask.
[0098] The key processing steps include the following:
[0099] 1. Initialization and configuration: Set a random seed to ensure reproducibility of the results, and select a suitable device (CPU / GPU) according to the system configuration.
[0100] 2. Image loading and display: Use QFileDialog to load local image files. Convert the loaded images into QPixmap format and display them in QGraphicsView.
[0101] 3. Interactive annotation: Users draw rectangular boxes on the image through mouse operations. Record the positions where the mouse is pressed, moved, and released, and dynamically draw rectangular boxes.
[0102] The mediastinal medical image segmentation method and system based on deep learning provided by the embodiments of the present invention have the following characteristics and advantages:
[0103] 1) Facing a weakly supervised and small-sample data environment, use a dataset containing different types of data (such as text, images, audio, etc.), including 3D CT slices, 2D slices, and doctors' diagnostic information (as categorical annotation text) to assist in the accurate classification of mediastinal tumors in CT image data.
[0104] 2) The mediastinal medical image segmentation system supports the input of 3D and 2D CT images and performs segmentation detection.
[0105] 3) Support free annotation and segmentation of lesion areas.
[0106] 4) Use transfer learning. Utilize the similarities between data, tasks, or models with a limited dataset, and apply the model learned in the old domain to the new domain to increase the accuracy and robustness of the segmentation results.
[0107] 5) Utilize the attention mechanism to screen and express features in the CT image data of mediastinal tumors. How to design a deep learning model to model the correlation between features, mine the correlation and difference between local features and global features, generate an expression method that can visually interact spatially with data, and perform efficient semantic segmentation. Introduce the Transformer structure in the encoder stage to capture the global dependencies of the image, extract global features with high-level semantic information, and determine the location and category of the target lesion.
[0108] It should be noted that in the embodiments of the present invention, if the above-mentioned deep learning-based mediastinal medical image segmentation method is implemented in the form of software functional modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an electronic device to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs, etc., which can store program codes. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0109] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The sequence numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.
[0110] It should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
[0111] In several embodiments provided by the present invention, it should be understood that the disclosed method can be implemented in other ways. The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0112] As mentioned above, only the embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.
Claims
1. A mediastinal medical image segmentation method based on deep learning, characterized in that, Including: Collect a custom dataset of different types of medical images, and perform data cleaning, preprocessing, and annotation to obtain a sample dataset; wherein, the different types of medical images include three-dimensional CT images of patients with germ cell tumors, lymph node tumors, neuromas, teratomas, and thymomas; the sample dataset includes 2D slices generated by cutting the three-dimensional CT images and corresponding mask files; Divide the sample dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1; Input the training set batch by batch into the pre-trained MedSAM model, and obtain the predicted segmentation results of the voxels in the training samples through forward calculation; Based on the predicted segmentation results and the prior labels corresponding to the mask files, determine the target loss, and update the parameters of the pre-trained MedSAM model in reverse until convergence to obtain a medical image semantic segmentation model; Use the validation set and the test set to evaluate the performance of the medical image semantic segmentation model.
2. The method according to claim 1, wherein The preprocessing includes the following operations: Cut all three-dimensional CT images perpendicular to the z-axis through the nibabel library in python, slice the images in the mediastinal window with dimensions from 50 to 350, remove objects with a small number of pixels in the two-dimensional plane, set the image size to 1024 to obtain complete 2D slices containing the lesion area, and generate 2D pixel-level segmentation labels annotated by clinicians as mask files.
3. The method according to claim 1, wherein The pre-trained MedSAM model includes an image encoder, a prompt encoder, and a mask decoder, where: The image encoder is used to extract key features from the image using a vision transformer; the vision transformer consists of 12 transformer layers, and each transformer layer includes a multi-head self-attention mechanism and a multi-layer perceptron, which includes layer normalization; The prompt encoder is used to map the corner points of the bounding box prompt to a 256-dimensional vector embedding; The mask decoder is used to integrate the image embedding, the prompt embedding, and the output token to produce accurate segmentation results and corresponding confidence metrics; the mask decoder includes two transformer layers for fusing the image embedding and the prompt encoding, and two transposed convolutional layers for upsampling the resolution of the embedding to 256*256; subsequently, the embedding is processed by a sigmoid function and then aligned with the original input size through bilinear interpolation.
4. The method according to any one of claims 1 to 3, characterized in that, The target loss is the unweighted sum of the dice loss and the cross-entropy loss, and is calculated by the following formula: Among them, L is the target loss, and L CE is the cross-entropy loss, and L Dice is the dice loss, where s i and g i represent the predicted segmentation result and the ground truth of voxel i respectively, and N is the number of voxels in image I.
5. A mediastinal medical image segmentation system based on deep learning, characterized in that, The mediastinal medical image segmentation system includes an image processing terminal, a user system, and a parameter adjustment system, where: The image processing terminal is deployed with a semantic segmentation model obtained by performing transfer learning on a custom dataset based on the MedSAM model, and is used for doctors to log in and perform basic image processing on CT images and execute the method according to any one of claims 1 to 4 to segment the image lesion site of the CT image; The user system is used for user registration and user login to query the user diagnosis results fed back by doctors at the image processing terminal; The parameter system is used for background management login to adjust the parameters of the semantic segmentation model and conduct data testing.
6. The system according to claim 5, wherein The image processing end is specifically used to perform the following operations: Use the MedSAM model to pre-train the FLARE22 open-source dataset to obtain a pre-trained MedSAM model; Perform data preprocessing and normalization operations on the custom dataset; Fine-tune the parameters of the pre-trained MedSAM model and then use the processed custom data for training and conduct a comparative experiment to obtain the semantic segmentation model; Feed back the training weights of the semantic segmentation model to the front end, detect a single unknown data to obtain the segmentation and system diagnosis results, and feedback them to the doctor for diagnosis. Finally, feedback the final results to the user system.
Citation Information
Cited By
X-ray image quality intelligent evaluation fusion model and use method
CN120876272A
X-ray image quality intelligent evaluation fusion model and use method
CN120876272B
Transfer learning driven intracranial tumor image data full-automatic segmentation method
CN121305088A