A training method for a pancreatic cancer target segmentation model
The pancreatic cancer target segmentation model uses a cyclic generative adversarial network to generate enhanced CT images from plain CT scans, integrating OAR and imageomics models to enhance segmentation accuracy and minimize radiation exposure.
Patent Information
- Application Number
- CN202411612708.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-13
AI Technical Summary
In the prior art, pancreatic cancer target segmentation methods rely on enhanced CT images to have a risk of allergic reactions, and the lack of effective constraints leads to low segmentation accuracy.
By constructing a joint training model, a reinforced CT image is generated using plain-scanned CT images, and a pancreatic cancer target segmentation is performed by combining OAR segmentation model and imaging omics model. A circular generation adversarial network is used to generate enhanced CT images, which combines tumor distribution probability and OAR segmentation information for constraints to improve segmentation accuracy.
While reducing radiation exposure to patients, it improves the accuracy and effectiveness of target segmentation of pancreatic cancer tumors, and improves the problem of target segmentation position errors caused by the lack of effective constraints in traditional methods.
Smart Images

Figure CN119579889B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target segmentation, and particularly to a method for training a pancreatic cancer target segmentation model. Background Art
[0002] Pancreatic cancer is a highly malignant tumor. Its early symptoms are not obvious, and it is often in the advanced stage when discovered, resulting in poor treatment effects. Radiotherapy is one of the main treatment methods for pancreatic cancer, and accurate gross target volume (GTV) segmentation is a key link in radiotherapy.
[0003] With the development of artificial intelligence technologies such as machine learning and deep learning, computer-aided target segmentation methods have begun to be applied clinically. In current medical practice, the automatic segmentation effect of the target area of pancreatic cancer is generally inferior to that of other types of cancers. This is mainly because the pancreas is located in a special position inside the human body, and coupled with the limitations of plain CT in imaging ability, the accurate segmentation of pancreatic cancer has become a challenge.
[0004] In the field of radiotherapy, using multi-modal images of plain CT and contrast-enhanced CT for automatic target segmentation is an important technological advancement. It can help doctors more accurately determine the location and scope of tumors, thereby improving the treatment effect. However, the contrast agent of contrast-enhanced CT may cause allergic reactions and nephrotoxicity. Multi-phase contrast-enhanced CT scans will prolong the scanning time and increase radiation exposure, which is harmful to the health of radiation-sensitive groups such as children. Therefore, it is difficult to obtain contrast-enhanced CT images. The lack of existing contrast-enhanced CT leads to a low segmentation accuracy. Summary of the Invention
[0005] In view of the above analysis, embodiments of the present invention aim to provide a method for training a pancreatic cancer target segmentation model to solve the problem of low segmentation accuracy caused by the lack of existing contrast-enhanced CT.
[0006] On the one hand, embodiments of the present invention provide a method for training a pancreatic cancer target segmentation model, including the following steps:
[0007] Obtain plain CT images of pancreatic cancer patients, and generate enhanced CT images corresponding to the plain CT based on an enhanced CT image generation model; construct samples according to the plain CT images and corresponding OAR contours, enhanced CT images and target area contours to obtain a training sample set;
[0008] Construct a joint training model, the joint training model includes an OAR segmentation model, a radiomics model, and a GTV segmentation model; the OAR segmentation model is used for OAR segmentation based on plain CT images; the radiomics model is used for tumor distribution probability prediction; the GTV segmentation model is used for pancreatic cancer GTV segmentation based on OAR segmentation and tumor distribution probability;
[0009] Jointly train the joint model based on the training sample set to obtain a trained joint model for pancreatic cancer target segmentation.
[0010] Based on a further improvement of the above method, the enhanced CT image generation model is a cyclic generative adversarial network model, and the cyclic generative adversarial network model includes: a first generator, a second generator, a first discriminator, and a second discriminator;
[0011] The first generator is used to generate enhanced CT images based on plain CT images; the first discriminator is used to discriminate the authenticity of the generated enhanced CT images;
[0012] The second generator is used to generate plain CT images from enhanced CT images; the second discriminator is used to discriminate the authenticity of the generated plain CT images;
[0013] Collect paired plain CT images and enhanced CT images of pancreatic cancer patients to construct a training set, and train the cyclic generative adversarial network model to obtain an enhanced CT image generation model.
[0014] Use the following formula to calculate the training loss of the cyclic generative adversarial network model:
[0015]
[0016] where represents the discriminator loss, represents the generator loss, represents the consistency loss.
[0017] Based on a further improvement of the above method, use the following formula to calculate the consistency loss:
[0018]
[0019] where D1(·) represents the discrimination result of the first discriminator, D2(·) represents the discrimination result of the second discriminator, G1(·) represents the generation result of the first generator, G2(·) represents the generation result of the second generator, CT1 represents the plain CT image, CT2 represents the enhanced CT image, and ||·||1 represents the 1-norm.
[0020] Based on a further improvement of the above method, jointly train the joint model based on the training sample set to obtain a trained joint model for pancreatic cancer target segmentation, including:
[0021] S21. Fix the parameters of the GTV segmentation model, train the OAR segmentation model based on the plain CT images of the training samples and the corresponding OAR contours, and train the radiomics model based on the enhanced CT images of the training samples and the target region contours;
[0022] S22. Fix the parameters of the OAR segmentation model and the radiomics model, and train the GTV segmentation model based on the enhanced CT images of the training samples, the target region contours, the OAR segmentation results output by the OAR segmentation model, and the tumor distribution probability prediction results output by the radiomics model;
[0023] S23. Alternately perform step S21 and step S22 until the GTV segmentation model converges, and end the training.
[0024] Based on a further improvement of the above method, the following formula is used to calculate the training loss of the OAR segmentation model:
[0025] L OAR = α1L OAR-contour + α2L OAR-dice + α3L OAR-CE
[0026] L OAR-contour = -(L OCdice + L OCif )
[0027]
[0028]
[0029] where L OCdice represents the boundary dice loss, L OCif represents the boundary distance loss, N represents the number of pixels of the sample, C represents the number of types of OARs, M represents the number of samples in the current training batch, represents whether the i-th pixel of the j-th sample belongs to the k-th type of OAR. If it belongs, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample predicted by the model belongs to the k-th type of OAR. represents whether the i-th pixel of the j-th sample belongs to the OAR. If it belongs, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample predicted by the model belongs to the OAR. α1, α2, and α3 represent weights.
[0030] Based on a further improvement of the above method, the following formula is used to calculate the boundary distance loss:
[0031]
[0032] where, represents the transformation matrix from the boundary of the k-th type of OAR of the j-th sample predicted by the OAR segmentation model to the true boundary of the k-th type of OAR of the j-th sample, I represents the identity matrix, ||·||F Represents the Frobenius norm of the matrix.
[0033] Based on the further improvement of the above method, the following formula is used to calculate the boundary dice loss:
[0034]
[0035] Where, Indicates whether the i-th pixel of the j-th sample is the boundary of the k-th class of OAR. If it is the boundary, it is 1; otherwise, it is 0. Represents the probability that the i-th pixel of the j-th sample learned by the model is the boundary of the k-th class of OAR.
[0036] Based on the further improvement of the above method, the GTV segmentation model includes:
[0037] The second encoding module, including a first encoding channel and a second encoding channel. The input of the first encoding channel is the enhanced image of the sample and the OAR segmentation result, and the input of the second encoding channel is the enhanced image of the sample and the tumor distribution probability prediction result; each encoding channel includes a plurality of downsampling units connected in sequence; used for extracting multi-layer semantic features from the input image.
[0038] The intermediate layer is used to fuse the semantic features extracted by the last downsampling unit of the first encoding channel and the second encoding channel, and transfer the fused semantic features to the decoding module.
[0039] The decoding module, including a plurality of upsampling units connected in sequence; used for layer-by-layer feature map restoration based on the semantic features extracted by the second encoding module to obtain the target area prediction feature map; the number of downsampling units of the second encoding module is the same as and corresponds one-to-one with the number of upsampling units of the decoding module.
[0040] The output module is used to predict the pancreatic cancer GTV based on the target area prediction feature map.
[0041] Based on the further improvement of the above method, the following formula is used to calculate the loss of the GTV segmentation model:
[0042] L GTV = β1L GTV-contour + β2L GTV-dice + β3L GTV-CE
[0043] L GTV-contour = -(L GCdice + L GCif )
[0044]
[0045]
[0046] Among them, L GCdice represents the Dice loss of the boundary of the GTV, and L GCif represents the boundary distance loss of the GTV. N represents the number of pixels of the sample, and M represents the number of samples in the current training batch. represents whether the i-th pixel of the j-th sample is in the GTV. If it is, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample predicted by the model is in the GTV. β1, β2, and β3 represent weights.
[0047] Compared with the prior art, the present invention obtains the plain CT of pancreatic cancer patients, generates the enhanced CT image corresponding to the plain CT based on the enhanced CT image generation model, and obtains multi-modal medical image data. Thus, while reducing the impact on patients, multi-modal data is utilized, that is, the automatic segmentation of the multi-modal images of the plain CT combined with the enhanced CT can improve the accuracy of the target area segmentation of pancreatic cancer tumors. By constructing a joint training model and performing joint training, information on the pancreas and organs at risk is obtained, that is, the tumor distribution probability and the OAR segmentation information. These information are used as constraint conditions, so as to more comprehensively mine and utilize the useful information in the multi-modal images, and improve the accuracy and effect of automatic segmentation. In the GTV segmentation training, two constraints are introduced simultaneously, which can effectively improve the segmentation accuracy, improve the problem of incorrect target area segmentation position caused by the lack of effective constraints in the traditional method, and improve the quality of automatic segmentation.
[0048] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description. Moreover, some advantages can be made obvious from the description, or can be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained from the content specifically pointed out in the description and the drawings. Description of the Drawings
[0049] The drawings are only for the purpose of showing specific embodiments, and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs represent the same components.
[0050] Figure 1 is a flowchart of the training method of the pancreatic cancer target area segmentation model according to the embodiment of the present invention. Detailed Embodiments
[0051] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. Among them, the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, and are not used to limit the scope of the present invention.
[0052] A specific embodiment of the present invention discloses a training method for a pancreatic cancer target segmentation model, as Figure 1 shown, which includes the following steps:
[0053] S1. Obtain the plain CT images of pancreatic cancer patients, and generate enhanced CT images corresponding to the plain CT based on an enhanced CT image generation model; construct samples according to the plain CT images and the corresponding OAR contours, and the enhanced CT images and target area contours to obtain a training sample set;
[0054] S2. Construct a joint training model, which includes an OAR segmentation model, a radiomics model, and a GTV segmentation model; the OAR segmentation model is used to segment OAR based on plain CT images; the radiomics model is used to predict the tumor distribution probability; the GTV segmentation model is used to segment pancreatic cancer GTV based on OAR segmentation and tumor distribution probability;
[0055] S3. Perform joint training on the joint model based on the training sample set to obtain a trained pancreatic cancer target segmentation joint model.
[0056] Compared with the prior art, the training method for the pancreatic cancer target segmentation model provided in this embodiment obtains the plain CT of pancreatic cancer patients, generates enhanced CT images corresponding to the plain CT based on an enhanced CT image generation model, and obtains multi-modal medical image data, thereby reducing the impact on patients while using multi-modal data, that is, automatic segmentation using the multi-modal images of plain CT combined with enhanced CT can improve the accuracy of target segmentation of pancreatic cancer tumors. By constructing a joint training model and performing joint training, information on the pancreas and organs at risk can be obtained, that is, tumor distribution probability and organ-at-risk segmentation information, and these information are used as constraint conditions, so as to more comprehensively mine and utilize the useful information in multi-modal images and improve the accuracy and effect of automatic segmentation. In the GTV segmentation training, two types of constraints are introduced simultaneously, which can effectively improve the segmentation accuracy and improve the problem of incorrect target segmentation position caused by the lack of effective constraints in the traditional method. This method not only improves the quality of target segmentation, but also expands the traditional automatic target area training method. With the development of medical technology, the demand for precision medicine is increasing day by day. As a typical precision medicine demand, the treatment of pancreatic cancer also has an increasing demand for accurate target segmentation. Therefore, this technology will have a relatively broad application prospect in the field of radiotherapy planning for pancreatic cancer tumors.
[0057] During implementation, before starting model training, the primary task is to collect medical image data of pancreatic cancer patients, that is, plain CT images, and generate enhanced CT images corresponding to the plain CT based on an enhanced CT image generation model, which provide a key information basis for subsequent analysis and training.
[0058] Specifically, the enhanced CT image generation model is a cyclic generative adversarial network model, which includes: a first generator, a second generator, a first discriminator, and a second discriminator;
[0059] The first generator is used to generate enhanced CT images based on non-enhanced CT images; the first discriminator is used to distinguish the authenticity of the generated enhanced CT images;
[0060] The second generator is used to generate non-enhanced CT images from enhanced CT images; the second discriminator is used to distinguish the authenticity of the generated non-enhanced CT images;
[0061] Collect pairs of non-enhanced CT images and enhanced CT images of pancreatic cancer patients to construct a first training set, and train the cyclic generative adversarial network model to obtain an enhanced CT image generation model.
[0062] In order to train a network model that can more accurately generate enhanced CT images based on non-enhanced CT images, the present invention adopts a cyclic generative adversarial network, that is, the trained generative adversarial network can not only generate enhanced CT images based on non-enhanced CT images, but also generate non-enhanced CT images based on enhanced CT images, thereby avoiding the situation that the enhanced CT images generated by the generator for all samples are the same.
[0063] In implementation, both the first generator and the second generator can adopt an encoder-decoder network structure, such as the UNet structure. The first discriminator and the second discriminator can adopt PatchGAN. By collecting pairs of non-enhanced CT images and enhanced CT images of patients to construct a first training set, train the cyclic generative adversarial network model to obtain an enhanced CT image generation model.
[0064] Specifically, the following formula is used to calculate the training loss of the cyclic generative adversarial network model:
[0065]
[0066] Among them, represents the discriminator loss, represents the generator loss, represents the consistency loss.
[0067] The discriminator loss is the sum of the generation loss of the first discriminator and the discrimination loss of the second discriminator. Specifically, the following formula is used to calculate the discriminator loss:
[0068]
[0069] Among them, D1(·) represents the discrimination result of the first discriminator, D2(·) represents the discrimination result of the second discriminator, G1(·) represents the generation result of the first generator, G2(·) represents the generation result of the second generator, CT1 represents the plain CT image, and CT2 represents the enhanced CT image.
[0070] The generator loss is the sum of the generation losses of the first generator and the second generator. The generation loss of the first generator is the loss between the enhanced CT image generated by the first generator based on the plain CT image and the original enhanced CT image. The generation loss of the second generator is the loss between the plain CT image generated by the second generator based on the enhanced CT image and the generation of the original plain CT image.
[0071] Specifically, the following formula is used to calculate the generator loss:
[0072]
[0073] Among them, G1(·) represents the generation result of the first generator, G2(·) represents the generation result of the second generator, CT1 represents the plain CT image, CT2 represents the enhanced CT image, represents the loss function.
[0074] In order to improve the accuracy of the generated data and avoid generating the same result for different samples, the loss of the cyclic generative adversarial network includes a consistency loss.
[0075] Specifically, the following formula is used to calculate the consistency loss:
[0076]
[0077] Among them, D1(·) represents the discrimination result of the first discriminator, D2(·) represents the discrimination result of the second discriminator, G1(·) represents the generation result of the first generator, G2(·) represents the generation result of the second generator, CT1 represents the plain CT image, CT2 represents the enhanced CT image, and ||·||1 represents the 1-norm.
[0078] During implementation, the image G1(G2(CT2)) obtained by inputting the enhanced CT image G2(CT2) generated by the second generator based on the plain scan CT image of the sample into the first generator is compared with the original plain scan CT image CT1, and the image G2(G1(CT1)) obtained by inputting the enhanced CT image G1(CT1) generated by the first generator based on the original plain scan CT image of the sample into the second generator is compared with the original enhanced CT image CT2. In addition, the first generator is used to generate enhanced CT images, and can generate corresponding enhanced CT images even when the input is an enhanced CT image. Therefore, the consistency loss includes the loss between CT2 and G1(CT2). The second generator is used to generate plain scan CT images, and can generate corresponding plain scan CT images even when the input is a plain scan CT image. Therefore, the consistency loss includes the loss between CT1 and G2(CT1).
[0079] It should be noted that the loss is the loss of one sample.
[0080] Based on the first training set, the cycle adversarial generation network model is trained, and backpropagation is performed according to the loss function to update the parameters of the cycle adversarial generation network model until the model converges, and an enhanced CT image generation model is obtained.
[0081] Among the collected plain scan CT images, doctors or radiology experts will be responsible for outlining the nearby organs at risk (OAR, Organ at Risk), including the stomach, liver, small intestine, duodenum, etc., to obtain OAR labels. Based on the enhanced CT image generation model, the enhanced CT image corresponding to the plain scan CT is generated. Similarly, in the enhanced CT image, the location of the tumor is outlined.
[0082] The plain scan CT image, the generated enhanced CT image, and the corresponding OAR labels and target area labels of pancreatic cancer patients constitute a training sample, thereby constructing a training sample set.
[0083] During implementation, the CT values of the images in the training sample set are statistically analyzed, sorted from smallest to largest, and the 5-thousandth percentile of the CT value sequence is used as the minimum value I min , and the 995-thousandth percentile is used as the maximum value I max , and the images are normalized. The formula is as follows:
[0084]
[0085] where I image is the original CT value, and CT values higher than the maximum value I max are uniformly set to I max , and CT values lower than the minimum value I min are set to Imin , which can improve the stability of model training.
[0086] Specifically, the joint model is jointly trained based on the training sample set to obtain a trained joint model for pancreatic cancer target region segmentation, including:
[0087] S21. Fix the parameters of the GTV segmentation model, train the OAR segmentation model based on the plain CT images of the training samples and the corresponding OAR contours, and train the radiomics model based on the enhanced CT images of the training samples and the target region contours;
[0088] S22. Fix the parameters of the OAR segmentation model and the radiomics model, and train the GTV segmentation model based on the enhanced CT images of the training samples and the target region contours, the OAR segmentation results output by the OAR segmentation model, and the tumor distribution probability prediction results output by the radiomics model;
[0089] S23. Alternately perform step S21 and step S22 until the GTV segmentation model converges, and end the training.
[0090] In implementation, the present invention adopts a joint training method to train a deep learning model. First, the OAR segmentation model and the radiomics model are trained based on one or more batches of samples, that is, the OAR segmentation model is trained based on the plain CT images of the training samples and the corresponding labels, and the radiomics model is trained based on the enhanced CT images of the training samples and the corresponding labels. Then, the parameters of the OAR segmentation model and the radiomics model are fixed, and the GTV segmentation model is trained based on one or more batches of samples. If the prediction network does not converge, continue to fix the parameters of the GTV segmentation model, and train the OAR segmentation model and the radiomics model based on one or more batches of samples, and so on, that is, alternately train the OAR segmentation model and the radiomics model, and the GTV segmentation model until the GTV segmentation model converges.
[0091] The OARs segmented by the OAR segmentation model will be used as OAR constraints for subsequent training, which is equivalent to a prior knowledge constraint. By using the labels in the OARs to guide the subsequent target region segmentation model, we can identify which regions are OARs and which regions are target regions, thereby further improving the accuracy and reliability of the model. The distribution probability predicted by the radiomics model describes the possibility of tumors appearing at different positions in the pancreas, and it can be used as a constraint for subsequent training to help us better predict the position and size of tumors.
[0092] Specifically, the OAR segmentation model adopts an encoder-decoder symmetric structure, specifically including:
[0093] The first encoding module, including a plurality of downsampling units connected in sequence; used for extracting multi-layer semantic features from the sample image;
[0094] The middle layer is used to compress the semantic features extracted by the last downsampling unit of the first encoding module and transfer the compressed semantic features to the decoding module;
[0095] The decoding module, including a plurality of upsampling units connected in sequence; used for layer-by-layer feature map restoration based on the semantic features extracted by the first encoding module to obtain the final feature map; the number of downsampling units of the first encoding module is the same as and corresponds one-to-one to the number of upsampling units of the decoding module; an attention module is connected between the downsampling unit and the corresponding upsampling unit, used for extracting the attention of the semantic features extracted by the downsampling unit and transferring them to the corresponding upsampling unit;
[0096] The output module includes a plurality of output channels, and each output channel is used to predict the OAR region corresponding to the output channel based on the final feature map.
[0097] In implementation, the OAR segmentation model adopts an encoder-decoder symmetric structure, which can obtain the spatial information of medical images for image semantic segmentation, improve the accuracy of medical image segmentation, and at the same time add an attention module to the skip connection structure. The convolutional attention module can weight the feature map, enhance the weight of important features, and at the same time reduce the influence of irrelevant and secondary features, improve the ability to capture important details, enable the model to pay more attention to the region to be segmented, and improve the segmentation accuracy.
[0098] In implementation, semantic features are extracted by the first encoding module. The first encoding module includes four downsampling units connected in sequence, gradually reducing the size of the feature map and increasing the number of channels of the feature map at the same time, so as to capture image features of different scales and improve the accuracy of the segmentation result.
[0099] In implementation, the downsampling unit includes a convolutional layer, an activation function, and a max pooling layer connected in sequence. The convolutional layer is used for feature extraction, and each downsampling unit usually contains two consecutive convolutional operations; an activation function is connected after each convolutional layer to increase the non-linear expression ability of the model; finally, downsampling is performed through the max pooling layer to reduce the spatial dimension of the image by half, thereby reducing the computational complexity and obtaining more abstract features.
[0100] The last downsampling unit of the first encoding module is connected to the middle layer, which is used to compress and refine the semantic features output by the downsampling unit, and at the same time maintain the semantic information of the features. In implementation, the middle layer is a convolutional layer.
[0101] The number of upsampling units of the decoding module of the OAR segmentation model is the same as and corresponds one-to-one to the number of downsampling units of the first encoding module.
[0102] During implementation, the output module is a convolutional layer for generating the final OAR segmentation result. The number of channels of the output module corresponds one-to-one with the types of OARs, that is, the first output channel is used to predict the region of the first type of OAR, the second output channel is used to predict the region of the second type of OAR, and so on.
[0103] Since segmentation errors are more likely to occur near the boundaries, to improve the segmentation accuracy of the model, the OAR segmentation model not only predicts the regions of OARs but also predicts the boundaries of OARs, considering the OAR boundary loss, thereby improving the prediction accuracy.
[0104] Specifically, the following formula is used to calculate the training loss of the OAR segmentation model:
[0105] L OAR = α1L OAR-contour + α2L OAR-dice + α3L OAR-CE
[0106]
[0107]
[0108]
[0109] where L OCdice represents the boundary dice loss, L OCif represents the boundary distance loss, N represents the number of pixels of the sample, C represents the number of types of OARs, M represents the number of samples in the current training batch, is whether the i-th pixel of the j-th sample belongs to the k-th type of OAR, if it belongs, it is 1, otherwise it is 0; represents the probability that the i-th pixel of the j-th sample predicted by the model belongs to the k-th type of OAR, and the value range is (0, 1), represents whether the i-th pixel of the j-th sample belongs to the OAR, if it belongs, it is 1, otherwise it is 0, represents the probability that the i-th pixel of the j-th sample predicted by the model belongs to the OAR, and α1, α2, and α3 represent weights. L OAR-contour represents the OAR boundary loss, L OAR-dice represents the OAR DICE loss, L OAR-CE represents the OAR cross-entropy loss.
[0110] Specifically, the following formula is used to calculate the boundary dice loss:
[0111]
[0112] where Indicates whether the i-th pixel of the j-th sample is the boundary of the k-th class of OAR. If it is the boundary, it is 1; otherwise, it is 0. Indicates the probability that the i-th pixel of the j-th sample learned by the model is the boundary of the k-th class of OAR. During implementation, are the parameters learned by the model.
[0113] Specifically, the following formula is used to calculate the boundary distance loss:
[0114]
[0115] where represents the transformation matrix from the boundary of the k-th class of OAR in the j-th sample predicted by the OAR segmentation model to the true boundary of the k-th class of OAR in the j-th sample. I represents the identity matrix, and ||·|| F represents the Frobenius norm of the matrix.
[0116] In the segmentation training, a multi-label strategy is adopted. According to the anatomical structure relationship, the OARs are labeled separately and then trained together, without the need for separate model training for each organ.
[0117] The radiomics model can adopt the vector space model (VSM) or other machine learning models.
[0118] During implementation, the input of the radiomics model is to segment the enhanced CT image into several small regions. For example, each layer of the image is segmented into square regions of 10*10 pixels. The radiomics features within each small region are extracted by inputting the small regions into the radiomics model, and the probability that the predicted region is the tumor region is obtained, thereby obtaining the tumor distribution probability map of the entire enhanced CT image. The label corresponding to the small region is based on the target mask. If there is a tumor in the small region, the corresponding label is the tumor region; otherwise, it is the non-tumor region. It can be based on the target mask. If the proportion of the pixel points within the target area in the small region is higher than the preset threshold, it is the tumor region; otherwise, it is the non-tumor region.
[0119] By predicting the probability that each small region is the tumor region through the radiomics model, the tumor distribution probability map is obtained. The probability distribution map gives a value between 0 and 1 for each pixel within each small region. The larger the value, the higher the probability of the presence of a tumor. The probability values of the pixels within the content of the small region are the same. During implementation, the cross-entropy is used as the loss function of the radiomics model. In this way, we can establish a radiomics model that can predict the tumor distribution probability. The distribution probability describes the possibility of the tumor appearing in different positions within the pancreas, and it can be used as a constraint for subsequent training to help us better predict the location and size of the tumor.
[0120] During implementation, the training losses of the OAR segmentation model and the radiomics model were calculated, and the parameters of the OAR segmentation model and the radiomics model were updated using the gradient descent method.
[0121] Then, the parameters of the OAR segmentation model and the imaging omics model are fixed, one or more batches of samples are collected, the plain CT of the samples are input into the OAR segmentation model to obtain the OAR segmentation result, and the enhanced CT image is input into the imaging omics model to obtain the tumor distribution probability prediction result. The OAR segmentation result, the tumor distribution probability prediction result, and the enhanced CT image are used as the input of the GTV segmentation model, and features are extracted for target area segmentation. Combined with the gold standard of GTV (i.e., the label), the loss is calculated, and the parameters of the GTV segmentation model are updated by gradient descent method through linear back propagation.
[0122] Specifically, the GTV segmentation model adopts an encoding-decoding symmetric structure, which includes:
[0123] The second encoding module includes a first encoding channel and a second encoding channel, the input of the first encoding channel is the enhanced image of the sample and the OAR segmentation result, and the input of the second encoding channel is the enhanced image of the sample and the tumor distribution probability prediction result; each encoding channel includes a plurality of downsampling units connected in sequence; and is used to perform multi-layer semantic feature extraction on the input image;
[0124] The middle layer is used to fuse the semantic features extracted by the last downsampling unit of the first encoding channel and the second encoding channel, and transmit the fused semantic features to the decoding module;
[0125] A decoding module, comprising a plurality of upsampling units connected in sequence; used to restore the feature map layer by layer based on the semantic features extracted by the second encoding module to obtain a target area prediction feature map; the number of downsampling units in the second encoding module is the same as the number of upsampling units in the decoding module and corresponds one to one;
[0126] An output module is used to predict the pancreatic cancer GTV based on the target area prediction feature map.
[0127] It should be noted that the OAR segmentation result input by the first coding channel is the set of segmentation results of all OAR classes output by the OAR segmentation model, and the segmentation result of each OAR class is represented by a 0, 1 mask, where 1 represents that the pixel belongs to this class of OAR, and 0 represents that the pixel does not belong to this class of OAR. The tumor distribution probability prediction result input by the second coding channel is the tumor distribution probability prediction result output by the imaging omics model.
[0128] During implementation, the structure of each encoding channel of the second encoding module is the same as that of the first encoding module. The decoding module of the GTV segmentation model has the same structure as that of the decoding module of the OAR segmentation model.
[0129] During implementation, the semantic features extracted by the last downsampling unit of the first coding channel and the second coding channel are fused in a splicing manner.
[0130] The number of upsampling units in the decoding module of the GTV segmentation model is the same as and corresponds one by one to the number of downsampling units in each channel of the second coding module.
[0131] Using the tumor distribution probability obtained from the enhanced image and radiomics, the GTV segmentation model can be informed of the areas in the pancreas where the probability of tumor presence is higher, thereby improving the accuracy of GTV semantic segmentation. At the same time, in order to further reduce the segmentation error and improve the accuracy of image semantic segmentation, the GTV segmentation model not only predicts the tumor region but also predicts the boundary of the target area, taking into account the boundary loss of the target area, thereby improving the prediction accuracy.
[0132] During implementation, the following formula is used to calculate the loss of the GTV segmentation model:
[0133] L GTV = β1L GTV-contour + β2L GTV-dice + β3L GTV-CE
[0134] L GTV-contour = -(L GCdice + L GCif )
[0135]
[0136]
[0137] Among them, L GCdice represents the boundary dice loss of GTV, L GCif represents the boundary distance loss of GTV, N represents the number of pixels of the sample, M represents the number of samples in the current training batch, represents whether the i-th pixel of the j-th sample is in GTV, if so it is 1, otherwise it is 0; represents the probability that the i-th pixel of the j-th sample predicted by the model is in GTV, and β1, β2, and β3 represent weights. Among them, L GTV-contour represents the boundary loss of GTV, L GTV-dice represents the DICE loss of GTV, L GTV-CE represents the cross-entropy loss of GTV.
[0138] Specifically, the following formula is used to calculate the boundary dice loss of GTV:
[0139]
[0140] Among them, Indicates whether the i-th pixel of the j-th sample is the boundary of the GTV. If it is the boundary, it is 1; otherwise, it is 0. Represents the probability that the i-th pixel of the j-th sample learned by the model is the boundary of the GTV. During implementation, are the parameters learned by the model.
[0141] Specifically, the following formula is used to calculate the boundary distance loss of the GTV:
[0142]
[0143] Among them, represents the transformation matrix from the boundary of the GTV of the j-th sample predicted by the target region segmentation model to the true boundary of the GTV of the j-th sample. I represents the identity matrix, and ||·|| F represents the Frobenius norm of the matrix.
[0144] Update the parameters of the GTV segmentation model according to the loss function until convergence, and end the training. Obtain the trained joint model.
[0145] For pancreatic cancer patients to be subjected to target region segmentation, input their plain CT images into the enhanced CT image generation model to generate enhanced CT images corresponding to the plain CT images. Then input the plain CT images and enhanced CT images into the trained joint model to obtain the pancreatic cancer target region segmentation result. That is, input the plain CT images into the OAR segmentation model to obtain the OAR segmentation result, and input the enhanced CT images into the radiomics model to obtain the tumor distribution probability prediction result. Then input the OAR segmentation result, the tumor distribution probability prediction result, and the enhanced CT images into the GTV segmentation model to obtain the target region segmentation result.
[0146] Those skilled in the art can understand that all or part of the processes of implementing the above embodiment methods can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0147] As mentioned above, only the preferred specific implementation manners of the present invention are described, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A training method for a pancreatic cancer target segmentation model, characterized in that, Including the following steps: Obtain the plain CT images of pancreatic cancer patients, and generate the enhanced CT images corresponding to the plain CT based on the enhanced CT image generation model; construct samples according to the plain CT images and the corresponding OAR contours, the enhanced CT images and the target area contours to obtain a training sample set; Construct a joint training model, which includes an OAR segmentation model, a radiomics model, and a GTV segmentation model; the OAR segmentation model is used to segment the OAR based on the plain CT image; the radiomics model is used to predict the tumor distribution probability; the GTV segmentation model is used to segment the pancreatic cancer GTV based on the OAR segmentation and the tumor distribution probability; Based on the training sample set, perform joint training on the joint training model to obtain a trained joint model for pancreatic cancer target segmentation, including: S21. Fix the parameters of the GTV segmentation model, train the OAR segmentation model based on the plain CT images and the corresponding OAR contours of the training samples, and train the radiomics model based on the enhanced CT images and the target area contours of the training samples; S22. Fix the parameters of the OAR segmentation model and the radiomics model, and train the GTV segmentation model based on the enhanced CT images and the target area contours of the training samples, the OAR segmentation results output by the OAR segmentation model, and the tumor distribution probability prediction results output by the radiomics model; S23. Alternately perform step S21 and step S22 until the GTV segmentation model converges, and end the training; Use the following formula to calculate the training loss of the OAR segmentation model: L OAR = α1L OAR-contour + α2L OAR-dice + α3L OAR-CE L OAR-contour = -(L OCdice + L OCif ) Among them, L OCdice represents the boundary dice loss, and L OCif represents the boundary distance loss. N represents the number of pixels in the sample, C represents the number of types of OARs, M represents the number of samples in the current training batch. indicates whether the i-th pixel of the j-th sample belongs to the k-th type of OAR. If it belongs, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample predicted by the model belongs to the k-th type of OAR. indicates whether the i-th pixel of the j-th sample belongs to the OAR. If it belongs, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample predicted by the model belongs to the OAR. α1, α2, and α3 represent weights. Use the following formula to calculate the loss of the GTV segmentation model: L GTV = β1L GTV-contour + β2L GTV-dice + β3L GTV-CE L GTV-contour = -(L GCdice + L GCif ) Among them, L GCdice represents the boundary dice loss of the GTV, and L GCif represents the boundary distance loss of the GTV. N represents the number of pixels of the sample, and M represents the number of samples in the current training batch. indicates whether the i-th pixel of the j-th sample is in the GTV. If it is in the region, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample predicted by the model is in the GTV. β1, β2, and β3 represent weights.
2. The training method of the pancreatic cancer target region segmentation model according to claim 1, wherein The enhanced CT image generation model is a cyclic generative adversarial network model, which includes: a first generator, a second generator, a first discriminator, and a second discriminator; The first generator is used to generate enhanced CT images based on plain CT images; the first discriminator is used to discriminate the authenticity of the generated enhanced CT images; The second generator is used to generate plain CT images from enhanced CT images; the second discriminator is used to discriminate the authenticity of the generated plain CT images; Collect paired plain CT images and enhanced CT images of pancreatic cancer patients to construct a first training set and obtain an enhanced CT image generation model for the cyclic generative adversarial network model.
3. The training method of the pancreatic cancer target area segmentation model according to claim 2, wherein Use the following formula to calculate the training loss of the cyclic generative adversarial network model: Among them, represents the discriminator loss, represents the generator loss, represents the consistency loss.
4. The training method of the pancreatic cancer target area segmentation model according to claim 3, wherein Use the following formula to calculate the consistency loss: Where, D1(·) represents the discrimination result of the first discriminator, D2(·) represents the discrimination result of the second discriminator, G1(·) represents the generation result of the first generator, G2(·) represents the generation result of the second generator, CT1 represents the plain CT image, CT2 represents the enhanced CT image, and ||·||1 represents the 1-norm.
5. The training method of the pancreatic cancer target region segmentation model according to claim 1, characterized in that Use the following formula to calculate the boundary distance loss: Among them, represents the transformation matrix from the boundary of the k-th type of OAR of the j-th sample predicted by the OAR segmentation model to the true boundary of the k-th type of OAR of the j-th sample. I represents the identity matrix, and ||·|| F represents the Frobenius norm of the matrix.
6. The training method of the pancreatic cancer target area segmentation model according to claim 1, wherein Use the following formula to calculate the boundary dice loss: Among them, indicates whether the i-th pixel of the j-th sample is the boundary of the k-th type of OAR. If it is the boundary, it is 1; otherwise, it is 0. represents the probability that the i-th pixel of the j-th sample learned by the model is the boundary of the k-th type of OAR.
7. The training method of the pancreatic cancer target area segmentation model according to claim 1, characterized in that The GTV segmentation model includes: The second encoding module includes a first encoding channel and a second encoding channel. The input of the first encoding channel is the enhanced image of the sample and the OAR segmentation result, and the input of the second encoding channel is the enhanced image of the sample and the tumor distribution probability prediction result. Each encoding channel includes a plurality of downsampling units connected in sequence, and is used for extracting multi-layer semantic features from the input image. The middle layer is used for fusing the semantic features extracted by the last downsampling unit of the first encoding channel and the second encoding channel, and transmitting the fused semantic features to the decoding module. The decoding module includes a plurality of upsampling units connected in sequence, and is used for gradually recovering the feature map based on the semantic features extracted by the second encoding module to obtain the target area prediction feature map. The number of downsampling units of the second encoding module is the same as that of the upsampling units of the decoding module and they correspond one by one. The output module is used for predicting the pancreatic cancer GTV based on the target area prediction feature map.
Citation Information
Patent Citations
Enhanced CT image generation method and system based on multi-task learning
CN116630463A
Training method for enhancing CT image generation model and storage medium
CN116977466A