A Weak-Supervision-Based Lung Nodule Segmentation Method, System, Device and Medium
Through the weakly supervised V-Net network method, a lung nodule segmentation model was constructed, which solved the problem that the existing technology could not evaluate the proportion of solid components of lung nodules, and achieved high-precision segmentation of ground glass and solid components of lung nodules, improving the accuracy of clinical diagnosis.
Patent Information
- Application Number
- CN202410365646.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-03-28
AI Technical Summary
The existing pulmonary nodule segmentation method can only divide the lesion area and cannot evaluate the proportion of solid components of the pulmonary nodule.
A weakly supervised lung nodule segmentation method is used to construct a lung nodule segmentation model using the V-Net network. Feature extraction and segmentation are performed through the input module, the downsampling module, the upsampling module and the output module. The output module includes two output convolution submodules to output the results of the lung nodule ground glass and solid component respectively.
The segmentation of ground glass and solid components of lung nodules can be achieved, which can accurately evaluate the proportion of solid components of lung nodules, and improve the accuracy and clinical value of lung nodules segmentation.
Smart Images

Figure CN118097151B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, relates to the segmentation of pulmonary nodules, and particularly relates to a weakly supervised pulmonary nodule segmentation method, system, device and medium. Background Technique
[0002] Lung cancer is caused by lung lesion tissues formed by uncontrollable abnormal cells that are not detected by the immune system through continuous self-replication and division. Such lung lesion tissues are called pulmonary nodules. Early-stage lung cancer exists in the form of pulmonary nodules. Pulmonary nodules are spherical lesion areas in the lungs with a diameter less than 4 cm. Usually, these nodules vary in size and location. Early detection can greatly reduce the mortality rate of lung cancer. The detection of pulmonary nodules is mainly through medical imaging techniques. For the detection of pulmonary nodules, considering factors such as easy acquisition and availability, CT scans are usually used, and the lesion areas of pulmonary nodules in CT images are segmented.
[0003] Pulmonary nodules will show four density states of ground glass, mixed ground glass (or semi-solid), solid and calcification in chest computed tomography (CT) images, which to a certain extent reflects the cell differentiation of the lesions. Evaluating the proportion of the solid component of pulmonary nodules has important clinical significance. It can not only guide doctors to formulate accurate follow-up and treatment strategies, but also predict the surgical prognosis of patients. Separately segmenting the regions of interest (ROIs) of the ground glass and solid components of pulmonary nodules can accurately evaluate the proportion of the solid component of the lesions. However, on the one hand, artificial intelligence algorithms based on deep neural network models require a large amount of accurate labeled data for model training. On the other hand, due to imaging principles and accuracy issues, the boundary positions of the lesions can only be approximately obtained and are difficult to distinguish by the naked eye. Therefore, developing an algorithm for measuring the solid component of pulmonary nodules based on deep neural networks still faces great challenges.
[0004] For the segmentation of the pulmonary nodule lesion area in CT images, most of the existing technologies are implemented through the V-Net model. The left side of the V-Net model is the compression path, that is, the encoding part, and the right side of the model is the decoding part. The encoding part of the V-Net model includes multiple different convolutional layers, and each layer corresponds to a resolution of a certain scale. As the data advances along the compression path to different layers, its resolution is reduced by a factor of two. Each layer of the V-Net model contains 1 to 3 convolutional operations, and the size of the convolutional kernel is 5x5x5. The output feature maps after 1 to 3 convolutional operations in the same layer have the same resolution. The V-Net model introduces a residual learning structure in each layer. The residual learning structure adds the output feature map obtained by the last convolutional layer in this layer to the output feature map after downsampling in the previous layer. The V-Net model performs downsampling operations at the end of each layer to obtain a larger receptive field. After downsampling, the resolution of the feature map becomes half of the original, and at the same time, the number of feature maps is twice the original. The V-Net model uses convolution to achieve downsampling. In image segmentation, downsampling can reduce the resolution of the feature map while retaining the important features of the image. In addition, due to the use of a relatively large 5x5x5 convolutional kernel and a total of 21 convolutional operations, the V-Net model has a large receptive field and strong fitting ability of the model.
[0005] The invention patent application with the application number 202210536515.8 also discloses a method for segmenting pulmonary nodule CT images based on an improved V-Net network, including the steps of: obtaining a pulmonary nodule CT image to be segmented; inputting the pulmonary nodule CT image into a trained improved V-Net network model for segmentation to obtain a segmentation result image; the improved V-Net network model includes multiple encoding layers and multiple decoding layers. The improved V-Net network model includes a first encoding layer, a second encoding layer, a third encoding layer, a fourth encoding layer, a fifth encoding layer, a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer. Among them, the first encoding layer is used to perform two convolutional operations on the pulmonary nodule CT image; the second encoding layer is used to perform two convolutional operations on the feature map output by the downsampling operation of the upper encoding layer; the third encoding layer, the fourth encoding layer, and the fifth encoding layer are used to perform three convolutional operations on the feature map output by the downsampling operation of the upper encoding layer; the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer are used to perform three convolutional operations on the feature map output by the splicing operation of this layer. This method for segmenting pulmonary nodule CT images makes full use of feature information, further alleviates the vanishing gradient, can grasp feature details, has a high detection rate for small nodules, strong anti-interference ability for interfering pixels such as blood vessels, and high segmentation accuracy for the cavity part and the adhesion area, so as to achieve high segmentation accuracy of pulmonary nodules.
[0006] Similar to the above-mentioned lung nodule segmentation method, most of the existing technologies can only segment the lesion areas of lung nodules. However, lung nodules can present four density states in chest computed tomography (CT) images, namely ground-glass, mixed ground-glass (or semi-solid), solid, and calcified. To a certain extent, the cell differentiation of the lesion can be reflected through lesion segmentation, but it is also of great clinical significance to evaluate the proportion of the solid component of the lung nodule, which cannot be achieved by the existing technologies. Summary of the Invention
[0007] The purpose of the present invention is to provide a weak-supervised lung nodule segmentation method, system, device, and medium to solve the technical problem that the existing lung nodule segmentation can only segment the lesion area and cannot evaluate the solid component of the lung nodule.
[0008] The present invention specifically adopts the following technical solutions to achieve the above purpose:
[0009] A weak-supervised lung nodule segmentation method includes the following steps:
[0010] Step S1, obtaining sample data;
[0011] Obtain a lung CT image with a lung nodule lesion and label it. The label includes a lung nodule ground-glass label and a lung nodule solid component label;
[0012] Step S2, constructing a lung nodule segmentation model;
[0013] Construct a lung nodule segmentation model. The lung nodule segmentation model adopts a V-Net network. The lung nodule segmentation model includes an input module, a downsampling module, an upsampling module, and an output module. The input module includes an input convolution sub-module. The downsampling module includes multiple downsampling convolution sub-modules. The upsampling module includes multiple upsampling convolution sub-modules. The downsampling convolution sub-modules of the downsampling module are skip-connected to the corresponding upsampling convolution sub-modules of the upsampling module. The output module includes two output convolution sub-modules. Each output convolution sub-module includes a convolutional layer and a sigmoid activation layer. The outputs of the convolutional layer and the sigmoid activation layer of each output convolution sub-module are used as the outputs of the lung nodule segmentation model;
[0014] Step S3, training the lung nodule segmentation model;
[0015] Training the lung nodule segmentation model specifically includes the following steps:
[0016] Step S3-1, pre-training;
[0017] Using the sample data, pre-train the lung nodule segmentation model with a conventional cross-entropy loss function to obtain a pre-trained model;
[0018] Step S3-2, update the model and back up the parameters;
[0019] Copy the parameters of the pre-trained model to the new pulmonary nodule segmentation model to obtain a backup model;
[0020] Step S3-3, E-step calculation;
[0021] Input the sample data into the backup model to obtain a prediction result, and combine it with the label of the sample data to calculate and obtain an estimate of the true segmentation mask ;
[0022] Step S3-4, M-step calculation;
[0023] Input the sample data into the pre-trained model to obtain a prediction result, and combine it with the estimate of the true segmentation mask and the label data to calculate the prediction error;
[0024] Step S3-5, update the model parameters;
[0025] Use the Adam algorithm to update the parameters of the backup model;
[0026] Step S3-6, repeat steps S3-3 to S3-5 until the iteration step ends;
[0027] Step S3-7, use the model after iteration as the new backup model, and use the original backup model as the pre-trained model, repeat steps S3-2 to S3-6 until the model converges or meets the termination condition;
[0028] Step S3-8, output the mature pulmonary nodule segmentation model;
[0029] Step S4, real-time segmentation;
[0030] Obtain the lung CT image to be segmented and input it into the mature pulmonary nodule segmentation model, and the pulmonary nodule segmentation model outputs the results of pulmonary nodule ground-glass and pulmonary nodule solid components.
[0031] Furthermore, in step S1, preprocess the lung CT images in the sample data, and the specific processing method is:
[0032] First, crop the pulmonary nodules from the sample data and label data;
[0033] Perform pixel spacing resampling and image normalization processing on the cropped images and annotation results;
[0034] Obtain the central position of the pulmonary nodule lesions in the lung CT images of the sample data, and crop an image block of size 48*96*96 from the original image with this central position as the center, and crop the corresponding annotated mask block in the label data;
[0035] Perform data augmentation on the sample data by means of random flipping, random rotation, random scaling, and random Gaussian noise perturbation.
[0036] Further, in step S2, the input convolution sub-module includes a convolutional layer, a batch normalization layer, and a PReLU activation layer;
[0037] The downsampling convolution sub-module and the upsampling convolution sub-module both include three consecutive convolutional layers, a batch normalization layer, and a PReLU activation layer.
[0038] Further, in step S3-3, the estimation of the true segmentation mask The calculation formula is:
[0039]
[0040] Among them, represents the sample data, represents the annotated mask in the label data, represents the true segmentation mask, represents the posterior probability of the true segmentation mask, represents the tendency probability that each pixel of the image is annotated, represents the posterior probability of the annotated mask, A B means A is proportional to B;
[0041] In step S3-4, the calculation formula of the prediction error is:
[0042]
[0043] Among them, represents the sample data, represents the annotated mask in the label data, represents the inference result of the E-step, represents the model to be optimized currently, represents the index of each pixel in the image, represents the dimension of the image, that is, the number of pixels, represents in the result of the E-step the estimation of the true class of the i th pixel; represents the i th pixel's true class, that is, whether it belongs to the ROI region; represents the model input data x after thei The probability that a pixel is predicted as a true segmentation ROI denotes the cross-entropy loss function denotes the i th pixel is labeled denotes the model input data x After i th pixel belongs to the true ROI, the probability that the pixel is labeled is predicted.
[0044] A weakly supervised lung nodule segmentation system, comprising:
[0045] A sample data acquisition module, configured to acquire a lung CT image with a lung nodule lesion and label it, and the label includes a lung nodule ground glass label and a lung nodule solid component label;
[0046] A lung nodule segmentation model construction module, configured to construct a lung nodule segmentation model, and the lung nodule segmentation model adopts a V-Net network; the lung nodule segmentation model includes an input module, a downsampling module, an upsampling module, and an output module. The input module includes an input convolutional sub-module, the downsampling module includes multiple downsampling convolutional sub-modules, the upsampling module includes multiple upsampling convolutional sub-modules, and the downsampling convolutional sub-modules of the downsampling module are skip-connected to the upsampling convolutional sub-modules of the corresponding upsampling module; the output module includes two output convolutional sub-modules, and each output convolutional sub-module includes a convolutional layer and a sigmoid activation layer. The outputs of the convolutional layer and the sigmoid activation layer of each output convolutional sub-module are used as the outputs of the lung nodule segmentation model;
[0047] A lung nodule segmentation model training module, configured to train the lung nodule segmentation model, specifically including the following steps:
[0048] Step S3-1, pre-training;
[0049] Using the sample data, pre-train the lung nodule segmentation model with a conventional cross-entropy loss function to obtain a pre-trained model;
[0050] Step S3-2, update the model and back up the parameters;
[0051] Copy the parameters of the pre-trained model to a new lung nodule segmentation model to obtain a backup model;
[0052] Step S3-3, E-step calculation;
[0053] Input the sample data into the backup model to obtain a prediction result, and combine it with the label of the sample data to calculate and obtain an estimate of the true segmentation mask ;
[0054] Step S3-4, M-step calculation;
[0055] Input the sample data into the pre-trained model to obtain the prediction results, and combine with the estimation of the true segmentation mask and the label data to calculate the prediction error;
[0056] Step S3-5, update the model parameters;
[0057] Use the Adam algorithm to update the parameters of the backup model;
[0058] Step S3-6, repeat steps S3-3 to S3-5 until the iteration step ends;
[0059] Step S3-7, use the model after iteration as the new backup model, and the original backup model as the pre-trained model, repeat steps S3-2 to S3-6 until the model converges or meets the termination conditions;
[0060] Step S3-8, output the mature pulmonary nodule segmentation model;
[0061] The real-time segmentation module is used to obtain the lung CT image to be segmented and input it into the mature pulmonary nodule segmentation model, and the pulmonary nodule segmentation model outputs the ground-glass result of the pulmonary nodule and the solid component result of the pulmonary nodule.
[0062] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the above method.
[0063] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the above method.
[0064] The beneficial effects of the present invention are as follows:
[0065] 1. In the present invention, the output module of the pulmonary nodule segmentation model includes two output convolutional sub-modules. Each output convolutional sub-module includes a convolutional layer and a sigmoid activation layer. The outputs of the convolutional layer and the sigmoid activation layer of each output convolutional sub-module are used as the outputs of the pulmonary nodule segmentation model; thus, the pulmonary nodule segmentation model can output two segmentation results of ground-glass of pulmonary nodules and solid components of pulmonary nodules, effectively solving the technical problem that the existing pulmonary nodule segmentation can only segment the lesion area and cannot evaluate the solid components of pulmonary nodules.
[0066] 2. In the present invention, a specific training method is adopted for the pulmonary nodule segmentation model. Even if the real segmentation region is not completely given in the labeled data set, the pulmonary nodule segmentation model can learn the real region of interest ROI for segmentation, improving the segmentation effect of pulmonary nodules.
[0067] 3. In the present invention, the maximization objective is decomposed into two steps, namely the E-step and the M-step, and the above loss function is constructed. The ultimate goal is to accurately identify the true segmentation ROI even without fully labeled segmentation ROIs. The constructed loss function and training algorithm enable the model to achieve this goal. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a schematic flowchart of the present invention;
[0069] Figure 2 is a schematic structural diagram of the pulmonary nodule segmentation model in the present invention;
[0070] Figure 3 is a schematic diagram when the pulmonary nodule segmentation model in the present invention is being trained. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0072] Therefore, based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0073] Embodiment 1
[0074] This embodiment provides a weakly supervised pulmonary nodule segmentation method for segmenting pulmonary nodules in pulmonary CT images. The segmentation task includes two components: ground glass of pulmonary nodules and solid components of pulmonary nodules, solving the technical problem that existing pulmonary nodule segmentation can only segment the lesion area and cannot evaluate the solid components of pulmonary nodules.
[0075] The pulmonary nodule segmentation method includes the following steps:
[0076] Step S1, obtaining sample data;
[0077] Obtain a pulmonary CT image with a pulmonary nodule lesion and label it. The labels include a ground glass label of pulmonary nodules and a solid component label of pulmonary nodules.
[0078] The sample data in this embodiment is from the data in the PACS system of West China Hospital of Sichuan University. 800 pulmonary CT images with pulmonary nodule lesions were retrieved and exported from the PACS system. Among them, 120 cases were used as the test set for model and algorithm evaluation, and the remaining 680 cases were used as the training set of the model. Then, using Slicer software, the ROI of the pulmonary nodule components in the training and test data sets was labeled respectively, including two labeling categories, namely the ROI of the pulmonary nodule lesion and the ROI of the solid component of the pulmonary nodule (the lesion area outside the solid component can be regarded as the ground-glass component). For the training data set, each case of data was labeled by a doctor with the ROI of each determined category, while for the test data set, each case of data was simultaneously outlined by four doctors with the complete ROI of each category according to their subjective judgments. While labeling each case of data, the corresponding labeling time was recorded; finally, the labeling results were exported and sorted out, the original data set was organized and summarized according to the specifications, and the labeling results of the data set were counted, including information such as the data volume, the average number of pulmonary nodules per CT, the scale distribution of nodules in each dimension, the average labeling time of each pulmonary nodule, and the labeling consistency analysis among different labelers.
[0079] Preprocess the pulmonary CT images in the sample data. The specific processing method is as follows:
[0080] The network model in this embodiment uses a three-dimensional convolutional neural network. Since the input data is the local area of the pulmonary nodule, it is necessary to first cut out the pulmonary nodules from the sample data and label data; then perform pixel spacing resampling and image normalization on the cut-out images and labeling results; when resampling, the pixel spacing of each CT image and its labeling result is resampled to 1×1×1 mm 3 in size to eliminate the influence of the pixel spacing difference of different CT images. According to the commonly used window width and window level of the lung window in clinical practice, the image pixel values are limited in the range of [-1000, 400], and then normalized; then, obtain the central position of the pulmonary nodule lesion in the pulmonary CT image of the sample data, and cut out an image block of size 48*96*96 from the original image with this central position as the center, cut out the corresponding labeled mask block in the label data, and use the voting method to determine the gold standard for the test data; finally, perform data augmentation on the sample data by means of random flipping, random rotation, random scaling, and random Gaussian noise perturbation to alleviate the risk of model overfitting.
[0081] Step S2, construct a pulmonary nodule segmentation model;
[0082] Construct a pulmonary nodule segmentation model, and the pulmonary nodule segmentation model adopts a V-Net network; the pulmonary nodule segmentation model includes an input module, a downsampling module, an upsampling module, and an output module. The input module includes an input convolutional sub-module, and the input convolutional sub-module includes a convolutional layer, a batch normalization layer, and a PReLU activation layer.
[0083] The downsampling module includes multiple downsampling convolutional sub-modules, the upsampling module includes multiple upsampling convolutional sub-modules, and there are skip connections between the downsampling convolutional sub-modules of the downsampling module and the upsampling convolutional sub-modules of the corresponding upsampling module. Each of the downsampling convolutional sub-modules and the upsampling convolutional sub-modules includes three consecutive convolutional layers, a batch normalization layer, and a PReLU activation layer. Among them, the size of the convolutional kernel of the convolutional layer is 5×5×5, the convolutional stride is 1, the number of convolutional kernels is the same as the number of channels, and padding operations are used to ensure that the size of the feature map after each convolution does not change. The output module includes two output convolutional sub-modules, and each output convolutional sub-module includes a convolutional layer and a sigmoid activation layer. The outputs of the convolutional layer and the sigmoid activation layer of each output convolutional sub-module are used as the outputs of the pulmonary nodule segmentation model; the output module outputs two channels, representing two categories of pulmonary nodule ROI and solid component ROI respectively, and the Sigmoid function is used to obtain effective probability value predictions.
[0084] Step S3, train the pulmonary nodule segmentation model;
[0085] Training the pulmonary nodule segmentation model specifically includes the following steps:
[0086] Step S3-1, pre-training;
[0087] Using sample data, pre-train the pulmonary nodule segmentation model with a conventional cross-entropy loss function to obtain a pre-trained model;
[0088] Step S3-2, update the model and back up the parameters;
[0089] Copy the parameters of the pre-trained model to a new pulmonary nodule segmentation model to obtain a backup model;
[0090] Step S3-3, E-step calculation;
[0091] Input the sample data into the backup model to obtain a prediction result, and combine it with the label of the sample data to calculate and obtain an estimate of the true segmentation mask ;
[0092] Step S3-4, M-step calculation;
[0093] Input the sample data into the pre-trained model to obtain a prediction result, and combine it with the estimate of the true segmentation mask And the label data to calculate the prediction error;
[0094] Step S3-5, update the model parameters;
[0095] Use the Adam algorithm to update the parameters of the backup model;
[0096] Step S3-6, repeat steps S3-3 to S3-5 until the iteration step ends;
[0097] Step S3-7, the model after the iteration ends is used as the new backup model, and the original backup model is used as the pre-trained model, and repeat steps S3-2 to S3-6 until the model converges or meets the termination condition;
[0098] Step S3-8, output the mature pulmonary nodule segmentation model;
[0099] Let be used to represent the input image, the ground truth segmentation mask, and the deterministic annotation mask respectively. D is the image dimension (the data dimension in this embodiment is 48×96×96). Use bold lowercase letters to represent the instances of the random variables respectively. Use regular lowercase letters and regular lowercase letters with subscripts to represent the value of a pixel in the image. For example, use to represent the value of the th pixel in the ground truth segmentation mask image. According to the problem definition of this embodiment, given a certain pulmonary nodule image data , if a certain pixel is annotated, that is , then this pixel must belong to the true ROI, that is or . If a certain pixel is not annotated, that is , then this pixel may or may not belong to the true ROI, that is or . From a statistical point of view, in this embodiment is the observable variable, while is unknown (i.e., the hidden variable). Therefore, this embodiment is based on the Expectation-Maximization (EM) algorithm to iteratively estimate the hidden variable . Specifically, first, according to the marginalization formula and the conditional probability formula, the following equation can be obtained:
[0100] (1)
[0101] According to this formula, this embodiment designs two output modules for the V-Net model to respectively convert the distributions and Parameterization means that for a given pixel, the model outputs the probability that this pixel belongs to the true ROI and the annotation tendency probability of this pixel . Maximizing the rightmost term of formula (1) is equivalent to maximizing the leftmost term of formula (1), so that the constructed model generates the given data set with the maximum likelihood probability under the dependence of the random variable , thus finally estimating the potential true segmentation mask. Formula (1) directly gives the maximization objective. However, when actually maximizing the log-likelihood , further transformation is needed, that is, the Q function:
[0102] (2)
[0103] where represent the model of the current iteration and the model of the previous EM iteration respectively. Formula (2) is the lower bound of formula (1). Then when given , any that makes formula (2) increase can also make formula (1) increase.
[0104] The above formula (2) directly derives the two stages that the model proposed in this embodiment has to go through during training, namely the E-step and the M-step.
[0105] In step S3-3, that is, the E-step. This process needs to be given the data sample and its annotation mask s, and estimate the probability of the true segmentation mask at this time. Given a certain pixel, according to Bayes' theorem, the estimate of the true segmentation mask is calculated by the formula:
[0106] (3)
[0107] where represents the sample data, represents the annotation mask in the label data, represents the true segmentation mask, represents the posterior probability of the true segmentation mask, represents the tendency probability that each pixel of the image is annotated, represents the posterior probability of the annotation mask, A B means A is proportional to B.
[0108] At this time, the estimate of the variable is jointly inferred from the output of the model of the previous EM algorithm iteration and the annotation mask . To simplify the expression, in this embodiment, the following is used to replace 。
[0109] In step S3-4, i.e., the M-step. According to the inference result of the E-step , and the given data and labeled samples , the calculation formula of the loss function (i.e., the prediction error) designed in this embodiment is:
[0110] (4)
[0111] Among them, represents the sample data, represents the annotation mask in the label data, represents the inference result of the E-step, represents the model that needs to be optimized currently, represents the index of each pixel in the image, represents the dimension of the image, i.e., the number of pixels, represents the estimation of the true class of the i th pixel in the result of the E-step; represents the i th pixel's true class, i.e., whether it belongs to the ROI region; represents the model input data x after which the probability that the i th pixel is predicted to be the true segmented ROI, represents the cross-entropy loss function, represents the i th pixel is labeled or not, represents the model input data x after which, given that the i th pixel belongs to the true ROI, the probability that this pixel is predicted to be labeled.
[0112] is the cross-entropy function.
[0113] During training, the Adam algorithm is used as the optimizer for updating the model parameters. The initial learning rate is set to 0.001. After each iteration, the initial learning rate is decreased by 10%, and during each iteration, the learning rate is exponentially decreased as the training epoch increases to prevent the learning rate from being too large and affecting the model convergence. Each time the model is trained for 40 epochs. If the optimal performance of the model has not improved after 5 iterations, the training is terminated to ensure the model converges sufficiently. During the training process, a backup of the model is always maintained, and its model parameters are the model that achieved the best result in the previous iteration.
[0114] Then, the trained model is applied to the pre-prepared test dataset to verify the effectiveness of the algorithm and the training process.
[0115] For the determination of solid components, given a test data, the proportion of the solid component can be calculated by simply counting the ratio of the number of pixels in the ROI of the pulmonary nodule lesion and the ROI of the solid component.
[0116]
[0117] Among them, and respectively represent the volumes occupied by the ROI of the solid component and the ROI of the pulmonary nodule lesion.
[0118] Step S4, real-time segmentation;
[0119] Obtain the lung CT image to be segmented and input it into a mature pulmonary nodule segmentation model, and the pulmonary nodule segmentation model outputs the results of pulmonary nodule ground-glass and the results of pulmonary nodule solid components.
[0120] In this embodiment, two output convolutional sub-modules are set in the pulmonary nodule segmentation model. The reasons are as follows: In the equation of formula (1), the term on the leftmost side of the equation can be estimated from the dataset (known), but in actual use, Y (the true segmentation result, unknown and needs to be estimated) is required. Therefore, the rightmost term in formula (1) is constructed. In this term, two probabilities need to be parameterized by the model, namely P(Y|X) and P(S|X,Y). While optimizing the loss function, the parameterized model can gradually approximate the true probability distributions P(Y|X) and P(S|X,Y). Therefore, the two output modules of the model actually model these two probability distributions. Secondly, each output convolutional sub-module outputs two channels. The reason is that the pulmonary nodule segmentation task in this embodiment includes two components: ground-glass pulmonary nodules and solid components of pulmonary nodules (since: ground-glass + solid = pulmonary nodule lesion, so when annotating, only the pulmonary nodule lesion area and the solid area need to be annotated to infer the area of the ground-glass component). Based on the above two categories, each output layer needs to model the probability distributions P(Y|X) and P(S|X,Y) for each category respectively. Therefore, the two output channels of each module represent the two categories of the segmentation task. By setting two channels, according to the annotation setting (only the determined areas are annotated, this annotation strategy can reduce the annotation burden and improve the reliability of the labels, and the annotated areas are all reliable positive samples), the dataset does not directly know the true segmentation Mask, that is, P(Y|X), and needs to be estimated through probability assumptions. Therefore, formula (1) is constructed, which means that the model needs two outputs to parameterize the probability distributions P(Y|X) and P(S|X,Y) respectively. This setting enables the true segmentation result to be estimated on the biased annotated dataset eventually, so that the model can learn to judge by itself the areas in the data that are difficult to distinguish as positive or negative samples (that is, those areas where the annotators are ambiguous about whether they are segmentation ROIs).
[0121] In this embodiment, an innovative training step is adopted. Through such a learning algorithm (including this loss function), even if the annotated dataset does not completely give the true segmentation area, the model can learn the true segmentation ROI. In addition, directly maximizing formula (1) is computationally infeasible. Therefore, formula (2) is derived and constructed. Formula (2) is the lower bound of formula (1), that is, maximizing formula (2) can also ensure the maximization of formula (1). Two steps are decomposed from the maximization objective of formula (2), namely the E-step and the M-step, and thus the above loss function and the training process here are constructed. The ultimate goal is to accurately identify the true segmentation ROI even if the segmentation ROI is not completely annotated. The loss function and training algorithm constructed above enable the model to achieve this goal.
[0122] Embodiment 2
[0123] This embodiment provides a pulmonary nodule segmentation system based on weak supervision, including:
[0124] A sample data acquisition module, configured to obtain pulmonary CT images with pulmonary nodule lesions and label them. The labels include pulmonary nodule ground-glass labels and pulmonary nodule solid component labels.
[0125] The sample data in this embodiment comes from the data in the PACS system of West China Hospital of Sichuan University. 800 pulmonary CT images with pulmonary nodule lesions were retrieved and exported from the PACS system. Among them, 120 cases were used as the test set for model and algorithm evaluation, and the remaining 680 cases were used as the training set of the model. Then, using Slicer software, the pulmonary nodule component ROIs of the training and test data sets were labeled respectively, including two labeling categories, namely the ROI of the pulmonary nodule lesion and the ROI of the solid component of the pulmonary nodule (the lesion area other than the solid component can be regarded as the ground-glass component). For the training data set, each case of data was labeled by a doctor with the ROI of each category determined by him. For the test data set, each case of data was simultaneously outlined with complete ROIs of various categories by four doctors according to their subjective judgments. When labeling each case of data, the corresponding labeling time was recorded. Finally, the labeling results were exported and sorted, the original data set was organized and summarized according to the specifications, and the labeling results of the data set were statistically analyzed, including information such as the data volume, the average number of pulmonary nodules per CT, the scale distribution of nodules in each dimension, the average labeling time of each pulmonary nodule, and the labeling consistency analysis among different labelers.
[0126] Preprocess the pulmonary CT images in the sample data. The specific processing method is as follows:
[0127] The network model in this embodiment uses a three-dimensional convolutional neural network. Since the input data is the local area of the pulmonary nodule, it is necessary to first cut out the pulmonary nodules from the sample data and label data; then perform pixel spacing resampling and image normalization processing on the cut-out images and labeling results. When resampling, the pixel spacing of each CT image and its labeling result is resampled to a size of 1×1×1 mm 3 to eliminate the influence of the pixel spacing difference of different CT images. According to the commonly used window width and window level of the lung window in clinical practice, the image pixel values are limited in the range of [-1000, 400], and then normalized; then, the central position of the pulmonary nodule lesion in the pulmonary CT image of the sample data is obtained, and an image block with a size of 48*96*96 is cut from the original image with this central position as the center. The corresponding labeled mask block in the label data is cut, and for the test data, the gold standard is determined by voting; finally, the sample data is augmented by random flipping, random rotation, random scaling, and random Gaussian noise perturbation to alleviate the risk of model overfitting.
[0128] The pulmonary nodule segmentation model construction module is used to construct a pulmonary nodule segmentation model. The pulmonary nodule segmentation model adopts the V-Net network. The pulmonary nodule segmentation model includes an input module, a downsampling module, an upsampling module, and an output module. The input module includes an input convolution sub-module, and the input convolution sub-module includes a convolutional layer, a batch normalization layer, and a PReLU activation layer.
[0129] The downsampling module includes multiple downsampling convolution sub-modules, and the upsampling module includes multiple upsampling convolution sub-modules. The downsampling convolution sub-modules of the downsampling module are skip-connected to the corresponding upsampling convolution sub-modules of the upsampling module. Both the downsampling convolution sub-module and the upsampling convolution sub-module include three consecutive convolutional layers, a batch normalization layer, and a PReLU activation layer. Among them, the size of the convolution kernel of the convolutional layer is 5×5×5, the convolution stride is 1, the number of convolution kernels is consistent with the number of channels, and through padding operations, the size of the feature map after each convolution does not change. The output module includes two output convolution sub-modules. Each output convolution sub-module includes a convolutional layer and a sigmoid activation layer. The outputs of the convolutional layer and the sigmoid activation layer of each output convolution sub-module are used as the outputs of the pulmonary nodule segmentation model. The output module outputs two channels, representing two categories of pulmonary nodule ROI and solid component ROI respectively, and uses the Sigmoid function to obtain effective probability value predictions.
[0130] The pulmonary nodule segmentation model training module is used to train the pulmonary nodule segmentation model, specifically including the following steps:
[0131] Step S3-1, pre-training;
[0132] Using the sample data, pre-train the pulmonary nodule segmentation model with a conventional cross-entropy loss function to obtain a pre-trained model;
[0133] Step S3-2, update the model and back up the parameters;
[0134] Copy the parameters of the pre-trained model to a new pulmonary nodule segmentation model to obtain a backup model;
[0135] Step S3-3, E-step calculation;
[0136] Input the sample data into the backup model to obtain a prediction result, and combine it with the label of the sample data to calculate and obtain an estimate of the true segmentation mask ;
[0137] Step S3-4, M-step calculation;
[0138] Input the sample data into the pre-trained model to obtain a prediction result, and combine it with the estimate of the true segmentation mask And the label data is used to calculate the prediction error;
[0139] Step S3-5, update the model parameters;
[0140] Use the Adam algorithm to update the parameters of the backup model;
[0141] Step S3-6, repeat steps S3-3 to S3-5 until the iteration step ends;
[0142] Step S3-7, the model after the iteration ends is used as the new backup model, and the original backup model is used as the pre-trained model. Repeat steps S3-2 to S3-6 until the model converges or meets the termination condition;
[0143] Step S3-8, output the mature pulmonary nodule segmentation model;
[0144] Let be used to represent the input image, the ground truth segmentation mask, and the deterministic annotation mask respectively. D is the image dimension (the data dimension in this embodiment is 48×96×96). Use bold lowercase letters to represent the instances of the random variables respectively. Use regular lowercase letters and regular lowercase letters with subscripts to represent the value of a pixel in the image. For example, use to represent the value of the th pixel in the ground truth segmentation mask image. According to the problem definition of this embodiment, given a certain pulmonary nodule image data , if a certain pixel is annotated, that is , then this pixel must belong to the true ROI, that is or . If a certain pixel is not annotated, that is , then this pixel may or may not belong to the true ROI, that is or . From a statistical point of view, in this embodiment is the observable variable, while is unknown (i.e., the hidden variable). Therefore, this embodiment is based on the Expectation-Maximization (EM) algorithm to iteratively estimate the hidden variable . Specifically, first, the following equations can be obtained according to the marginalization formula and the conditional probability formula:
[0145] (1)
[0146] According to this formula, this embodiment designs two output modules for the V-Net model to respectively convert the distributions and Parameterization means that for a given pixel, the model outputs the probability that this pixel belongs to the true ROI and the annotation tendency probability of this pixel . Maximizing the rightmost term of formula (1) is equivalent to maximizing the leftmost term of formula (1), such that the constructed model generates the given data set with the maximum likelihood probability under the dependence of the random variable , thereby finally estimating the potential true segmentation mask. Formula (1) directly gives the maximization objective. However, when actually maximizing the log-likelihood , further transformation is needed, that is, the Q function:
[0147] (2)
[0148] where respectively represent the model of the current iteration and the model of the previous EM iteration. Formula (2) is the lower bound of formula (1). Then, when given , any that makes formula (2) increase can also make formula (1) increase.
[0149] The above formula (2) directly derives the two stages that the model proposed in this embodiment has to go through during training, namely the E-step and the M-step.
[0150] In step S3-3, that is, the E-step. This process needs to be given the data sample and its annotation mask s, and estimate the probability of the true segmentation mask at this time. Given a certain pixel, according to Bayes' theorem, the formula for estimating the true segmentation mask is:
[0151] (3)
[0152] where represents the sample data, represents the annotation mask in the label data, represents the true segmentation mask, represents the posterior probability of the true segmentation mask, represents the tendency probability that each pixel of the image is annotated, represents the posterior probability of the annotation mask, A B means A is proportional to B.
[0153] At this time, the estimation of the variable is jointly inferred from the output of the model of the previous EM algorithm iteration and the annotation mask . For the sake of simplified expression, in this embodiment, the following is used to replace 。
[0154] In step S3-4, i.e., the M-step. According to the inference result of the E-step , and the given data and labeled samples , the calculation formula of the loss function (i.e., the prediction error) designed in this embodiment is:
[0155] (4)
[0156] Wherein, represents the sample data, represents the annotation mask in the label data, represents the inference result of the E-step, represents the model to be optimized currently, represents the index of each pixel in the image, represents the dimension of the image, i.e., the number of pixels, represents the estimation of the true class of the i th pixel in the result of the E-step; represents the i th pixel's true class, i.e., whether it belongs to the ROI region; represents the model input data x after which the probability that the i th pixel is predicted to be the true segmented ROI, represents the cross-entropy loss function, represents the i th pixel is labeled or not, represents the model input data x after which, given that the i th pixel belongs to the true ROI, the probability that this pixel is predicted to be labeled.
[0157] is the cross-entropy function.
[0158] During training, the Adam algorithm is used as the optimizer for updating the model parameters. The initial learning rate is set to 0.001. After each iteration, the initial learning rate is decreased by 10%, and during each iteration, the learning rate is exponentially decreased as the training epoch increases to prevent the learning rate from being too large and affecting the model convergence. Each iteration of the model is trained for 40 epochs. If the optimal performance of the model has not improved after 5 iterations, the training is terminated to ensure that the model converges sufficiently. During the training process, a backup of the model is always maintained, and its model parameters are the model that achieved the best result in the previous iteration.
[0159] Then, the trained model is applied to the pre-prepared test dataset to verify the effectiveness of the algorithm and the training process.
[0160] For the determination of solid components, given a test data, the proportion of the solid component can be calculated by only counting the ratio of the number of pixels of the ROI of the pulmonary nodule lesion and the ROI of the solid component.
[0161]
[0162] Among them, and respectively represent the volumes occupied by the ROI of the solid component and the ROI of the pulmonary nodule lesion.
[0163] The real-time segmentation module is used to obtain the pulmonary CT image to be segmented and input it into a mature pulmonary nodule segmentation model, and the pulmonary nodule segmentation model outputs the pulmonary nodule ground-glass result and the pulmonary nodule solid component result.
[0164] Embodiment 3
[0165] A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the pulmonary nodule segmentation method based on weak supervision.
[0166] Among them, the computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device.
[0167] The memory at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or D-interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system and various application software installed on the computer device, such as the program code of the weakly supervised lung nodule segmentation method. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0168] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the weakly supervised lung nodule segmentation method.
[0169] Embodiment 4
[0170] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the weakly supervised lung nodule segmentation method.
[0171] Wherein, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor to cause the at least one processor to execute the steps of the weakly supervised lung nodule segmentation method as described above.
[0172] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the weakly supervised lung nodule segmentation method described in the embodiments of the present application.
Claims
1. A pulmonary nodule segmentation method based on weak supervision, characterized in that: The following steps are involved: Step S1, obtaining sample data; Obtain lung CT images with lung nodule lesions and annotate them with labels, including ground glass labels for lung nodules and labels for solid components of lung nodules; Step S2, constructing a lung nodule segmentation model; A pulmonary nodule segmentation model is constructed, and the pulmonary nodule segmentation model adopts a V-Net network; the pulmonary nodule segmentation model includes an input module, a downsampling module, an upsampling module and an output module, the input module includes an input convolution submodule, the downsampling module includes multiple downsampling convolution submodules, the upsampling module includes multiple upsampling convolution submodules, and the downsampling convolution submodule of the downsampling module is jump-connected with the upsampling convolution submodule of the corresponding upsampling module; the output module includes two output convolution submodules, each of which includes a convolution layer and a sigmoid activation layer, and the outputs of the convolution layer and the sigmoid activation layer of each output convolution submodule are used as the output of the pulmonary nodule segmentation model; Step S3, training a lung nodule segmentation model; Training the lung nodule segmentation model includes the following steps: Step S3-1, pre-training; Using the sample data, a conventional cross entropy loss function is used to pre-train the lung nodule segmentation model to obtain a pre-trained model; Step S3-2, update the model and back up parameters; Copy the parameters of the pre-trained model to the new lung nodule segmentation model to obtain a backup model; Step S3-3, E-step calculation; Input the sample data into the backup model to get the prediction result, and combine it with the label of the sample data to calculate and get the estimate of the true segmentation mask ; Step S3-4, M-step calculation; Input the sample data into the pre-trained model to obtain the prediction result, and combine it with the estimate of the real segmentation mask And label data, calculate the prediction error; Step S3-5, updating model parameters; Use the Adam algorithm to update the parameters of the backup model; Step S3-6, repeating steps S3-3 to S3-5 until the iteration step ends; Step S3-7, the model after the iteration is used as the new backup model, and the original backup model is used as the pre-trained model, and steps S3-2 to S3-6 are repeated until the model converges or meets the termination condition; Step S3-8, outputting a mature lung nodule segmentation model; Step S4, real-time segmentation; The lung CT image to be segmented is obtained and input into a mature lung nodule segmentation model, which outputs the ground glass component results of the lung nodules and the solid component results of the lung nodules.
2. The method for segmenting pulmonary nodules based on weak supervision as claimed in claim 1, characterized in that: In step S1, the lung CT image in the sample data is preprocessed, and the specific processing method is as follows: First, cut out the lung nodules from the sample data and label data; Perform pixel spacing resampling and image normalization on the cropped images and annotation results; Get the center position of the lung nodule lesion in the lung CT image of the sample data, and crop a 48*96*96 image block from the original image with this center position as the center, and crop the corresponding annotated mask block in the label data; The sample data is augmented by random flipping, random rotation, random scaling, and random Gaussian noise perturbation.
3. The method for segmenting pulmonary nodules based on weak supervision as claimed in claim 1, characterized in that: In step S2, the input convolution submodule includes a convolution layer, a batch normalization layer, and a PReLU activation layer; Both the downsampling convolution submodule and the upsampling convolution submodule include three consecutive convolution layers, a batch normalization layer, and a PReLU activation layer.
4. The method for segmenting pulmonary nodules based on weak supervision as claimed in claim 1, characterized in that: In step S3-3, the estimation of the true segmentation mask The calculation formula is: in, represents sample data, represents the annotation mask in the label data, represents the true segmentation mask, represents the posterior probability of the true segmentation mask, Indicates the probability of each pixel of the image being labeled. represents the posterior probability of the labeled mask, A B means A is proportional to B; In step S3-4, the calculation formula of the prediction error is: in, represents sample data, represents the annotation mask in the label data, represents the reasoning result of E-step, Indicates the model that currently needs to be optimized. Represents the index of each pixel in the image, represents the dimension of the image, i.e. the number of pixels, Indicates the result of E-step for i An estimate of the true category of pixels; Indicates i The true category of the pixel, that is, whether it belongs to the ROI area; Represents model input data x After the first i The probability that a pixel is predicted to be a true segmentation ROI, represents the cross entropy loss function, Indicates i Whether pixels are labeled, Represents model input data x After that, given i Under the condition that the pixel belongs to the real ROI, predict the probability of the pixel being labeled.
5. A pulmonary nodule segmentation system based on weak supervision, characterized in that: include: The sample data acquisition module is used to acquire lung CT images with lung nodule lesions and annotate them with labels, including ground glass labels for lung nodules and labels for solid components of lung nodules. A pulmonary nodule segmentation model construction module is used to construct a pulmonary nodule segmentation model, and the pulmonary nodule segmentation model adopts a V-Net network; the pulmonary nodule segmentation model includes an input module, a downsampling module, an upsampling module and an output module, the input module includes an input convolution submodule, the downsampling module includes multiple downsampling convolution submodules, the upsampling module includes multiple upsampling convolution submodules, and the downsampling convolution submodule of the downsampling module is jump-connected with the upsampling convolution submodule of the corresponding upsampling module; the output module includes two output convolution submodules, each of which includes a convolution layer and a sigmoid activation layer, and the outputs of the convolution layer and the sigmoid activation layer of each output convolution submodule are used as the output of the pulmonary nodule segmentation model; The pulmonary nodule segmentation model training module is used to train the pulmonary nodule segmentation model, which specifically includes the following steps: Step S3-1, pre-training; Using the sample data, a conventional cross entropy loss function is used to pre-train the lung nodule segmentation model to obtain a pre-trained model; Step S3-2, update the model and back up parameters; Copy the parameters of the pre-trained model to the new lung nodule segmentation model to obtain a backup model; Step S3-3, E-step calculation; Input the sample data into the backup model to get the prediction result, and combine it with the label of the sample data to calculate and get the estimate of the true segmentation mask ; Step S3-4, M-step calculation; Input the sample data into the pre-trained model to obtain the prediction result, and combine it with the estimate of the real segmentation mask And label data, calculate the prediction error; Step S3-5, updating model parameters; Use the Adam algorithm to update the parameters of the backup model; Step S3-6, repeating steps S3-3 to S3-5 until the iteration step ends; Step S3-7, the model after the iteration is used as the new backup model, and the original backup model is used as the pre-trained model, and steps S3-2 to S3-6 are repeated until the model converges or meets the termination condition; Step S3-8, outputting a mature lung nodule segmentation model; The real-time segmentation module is used to obtain the lung CT image to be segmented and input a mature lung nodule segmentation model. The lung nodule segmentation model outputs the lung nodule ground glass result and the lung nodule solid component result.
6. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Pulmonary nodule CT image segmentation and training method and device based on improved V-Net network
CN114782402A
Multi-modal, multi-resolution deep learning neural networks for segmentation, outcomes prediction and longitudinal response monitoring to immunotherapy and radiotherapy
CN112771581A
Pulmonary nodule benign and malignant identification model training method, application method and system
CN116468103A
Cited By
Lung nodule risk prediction method, system, device and medium based on large model
CN120544910B