Abdominal Multi-Organ Incremental Segmentation Method Based on Location Guidance and Consistency Learning
By introducing position guidance and consistent learning strategies into the abdominal multi-organ segmentation network, the network structure is gradually expanded, and the problem of decreasing accuracy and complex expansion process of new and old organ segmentation in the existing technology is solved, and efficient and accurate incremental segmentation of abdominal multi-organs is achieved.
Patent Information
- Application Number
- CN202310381113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-04-11
AI Technical Summary
When the existing multi-organ incremental segmentation method of abdominal multi-organ expansion is expanded, the accuracy of the old organ segmentation decreases, and the expansion process is complicated and time-consuming, which cannot meet the precise multi-organ outline of the abdominal multi-organ.
Using a method based on location guidance and consistency learning, by gradually expanding the network structure, adding decoder, position guidance module and deep supervision module, using shared encoder and consistency learning strategies, simplifying the expansion process and improving the segmentation accuracy of new and old organs.
The incremental segmentation process of multiple organs in the abdomen is simplified, the segmentation accuracy of new and old organs is improved, the old network structure is avoided, and the old organ segmentation performance is maintained.
Smart Images

Figure CN116402800B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for incremental segmentation of multiple abdominal organs, which can be used to assist doctors in automatically delineating multiple organs before abdominal disease treatment. Background Art
[0002] The abdomen refers to the part of the human body between the chest cavity and the pelvis, and its internal structure is intricate. The abdominal region contains many important organs of the human body, such as the liver, kidneys, spleen, gallbladder, stomach, pancreas, and duodenum. With the changes in the climate environment and people's living habits, more and more people suffer from abdominal-related diseases. Computed tomography (CT) imaging is widely used in the diagnosis of abdominal diseases, such as disease diagnosis, surgical navigation, and lesion delineation. Manually delineating organs in abdominal CT images is not only time-consuming but also leads to significant differences within and between delineators. Therefore, an automatic abdominal multi-organ segmentation method is crucial. In addition, the real clinical needs are complex and diverse. Many times, the segmentation model is not only required to be able to segment existing organs but also needs to be further extended to be able to segment new organs.
[0003] Many abdominal multi-organ segmentation methods based on traditional algorithms have been successively proposed. Chu et al. proposed a three-dimensional automatic segmentation method for multiple abdominal organs based on a patient-specific weighted-probabilistic atlas in the article "Multi-organ segmentation from 3D abdominal CT images using patient-specific weighted-probabilistic atlas" published in the journal Medical Imaging: Image Processing. Okada et al. introduced shape and location priors to improve the accuracy of abdominal CT organ segmentation in the article "Abdominal multi-organ segmentation from CT images using conditional shape–location and unsupervised intensity priors" published in the journal Medical Image Analysis. However, although these traditional methods have promoted the development of automatic abdominal multi-organ segmentation, their huge computational cost limits their clinical application.
[0004] In recent years, due to the strong robustness of the fully convolutional network (FCN), the end-to-end abdominal multi-organ segmentation method based on FCN has received increasing attention. In the article Automatic Multi-organ Segmentation on Abdominal CT with Dense V-networks published by Gibson et al. in the journal IEEE Transactions on Medical Imaging, a densely connected V-shaped network was proposed. Its densely connected mode and multi-scale V-shaped network structure improved the performance of abdominal multi-organ segmentation. However, this method can only segment specific abdominal organs and does not have scalability. When new organs need to be included, the network structure must be redesigned and the segmentation network must be retrained. Using new data and new organ labels to expand a trained network to include new organs is called incremental segmentation. In incremental segmentation, the problem of forgetting old organs needs to be solved. For this reason, Liu et al. used the idea of knowledge distillation in the article Incremental Learning for Multi-organ Segmentation with Partially Labeled Datasets on arXiv, that is, calculating the KL divergence loss between the prediction P' of the new network for old organs and the prediction result P of the old network for old organs, and using this loss to constrain the training of the new network, so that the new network can maintain the ability to segment old organs while segmenting new organs. However, this method has a complex expansion process, consumes time, and reduces the segmentation accuracy of old organs. In addition, in this method, the new network does not fully utilize the information of old organs that can already be segmented, so the segmentation accuracy of new organs is not high and cannot meet the needs of accurate abdominal multi-organ delineation. Summary of the Invention
[0005] The purpose of the present invention is to propose an abdominal multi-organ incremental segmentation method based on position guidance and consistency learning for the deficiencies of the above-mentioned existing technologies, so as to simplify the expansion process and improve the segmentation accuracy of new organs and old organs.
[0006] The technical idea of the present invention is as follows: First, use the consistency learning strategy to train the basic network Net1 to segment the common abdominal organs liver, kidney, spleen, gallbladder, and stomach for the first time. Subsequently, by adding a new decoder, a position guidance module, and a deep supervision module, the basic network Net1 is extended to the position-guided incremental network Net2, and the position-guided incremental network Net2 is trained to segment the liver, kidney, spleen, gallbladder, stomach, and pancreas for the second time. Similarly, the position-guided incremental network Net2 is extended to the position-guided and consistency learning incremental network Net3, and the position-guided and consistency learning incremental network Net3 is trained to segment the liver, kidney, spleen, gallbladder, stomach, pancreas, and duodenum for the third time, and so on, to complete the abdominal multi-organ incremental segmentation.
[0007] According to the above idea, the implementation of the present invention includes the following:
[0008] (1) Divide the data set:
[0009] 1.1) Obtain four publicly available abdominal CT data sets. The first three data sets are respectively labeled with the liver, kidney, spleen, gallbladder, stomach, pancreas, and duodenum, and the fourth data set has no annotation;
[0010] 1.2) Perform image preprocessing on the four data sets obtained in 1.1), and randomly divide the first three preprocessed data sets into three training sets A1, A2, A3 and three test sets B1, B2, B3, and use the fourth data set as the fourth training set C in its entirety;
[0011] (2) Based on the UNet network structure, construct a basic network Net1 composed of an encoder, a decoder, and a deep supervision module;
[0012] (3) Initialize the basic network Net1 using the He initialization method, and use the labeled first training set A1 and the unlabeled fourth training set C to alternately iterate and train it through a consistency learning strategy to obtain the trained basic network parameters W1;
[0013] (4) Load the trained basic network parameters W1 into the basic network Net1, and add a decoder, two position guidance modules, and a deep supervision module to Net1 to form a position guidance incremental network Net2;
[0014] (5) Initialize the parameters of the newly added structure of the position guidance incremental network Net2 using the He initialization method, set its loss function to the Dice loss, freeze the trained basic network parameters W1, and use the second training set A2 and the Adam optimizer to iteratively train the position guidance incremental network Net2 until the loss function value no longer decreases, obtaining the trained position guidance incremental network parameters W2;
[0015] (6) Load the trained position guidance incremental network parameters W2 into the position guidance incremental network Net2, and add a decoder, two position guidance modules, and a deep supervision module to Net2 to form a position guidance and consistency learning incremental network Net3;
[0016] (7) Initialize the parameters of the newly added structure of the position guidance and consistency learning incremental network Net3 using the He initialization method, set its loss function to the Dice loss, freeze the parameters W2 of the trained position guidance incremental network, and use the third training set A3 and the Adam optimizer to iteratively train the position guidance and consistency learning incremental network Net3 until the value of the loss function no longer decreases, obtaining the trained parameters W3 of the position guidance and consistency learning incremental network;
[0017] (8) And so on, continue to expand the position guidance and consistency learning incremental network Net3 to segment more abdominal organs and complete the incremental segmentation of multiple abdominal organs;
[0018] (9) Input the first test set B1, the second test set B2, and the third test set B3 into the position guidance and consistency learning incremental network Net3 to obtain the prediction results of the network for the liver, kidney, spleen, gallbladder, stomach, pancreas, and duodenum.
[0019] Compared with the existing technologies, the present invention has the following advantages:
[0020] 1. Simplify the process of incremental segmentation of multiple abdominal organs.
[0021] The present invention realizes the expansion from the basic network Net1 to the position guidance incremental network Net2 by using a shared encoder and adding a decoder, and further expands the position guidance incremental network Net2 to the position guidance and consistency learning incremental network Net3. And so on, the position guidance and consistency learning incremental network Net3 can be continuously expanded to segment more abdominal organs. Compared with the previous method of incremental segmentation of multiple abdominal organs based on knowledge distillation, the expansion method proposed by this method is simpler and the training time is shorter.
[0022] 2. Improve the segmentation performance of the incremental segmentation of multiple abdominal organs.
[0023] The present invention uses the idea of deep supervision, and each segmentation network includes a deep supervision module, thereby avoiding gradient disappearance and accelerating network convergence; in addition to training the basic network Net1 using the labeled training set A1, a large number of unlabeled training sets C can also be used to train the basic network Net1, and the consistency learning strategy is used during the training process, enabling the basic network Net1 to make full use of the information of unlabeled images and improving the feature expression ability and migration ability of its encoder;
[0024] The present invention proposes a position guidance module and uses the position information of old organs to guide the training of new organs in the position guidance module, making the training process of new organs pay more attention to the feature information of key positions;
[0025] The incremental expansion method proposed by the present invention does not require changing the structure of the old network, so the segmentation performance of the old organs is not affected at all, avoiding catastrophic forgetting. Description of the Drawings
[0026] Figure 1 is the implementation flowchart of the present invention;
[0027] Figure 2 is the structural schematic diagram of the basic network Net1 in the present invention;
[0028] Figure 3 is the structural schematic diagram of the position-guided incremental network Net2 in the present invention;
[0029] Figure 4 is the structural schematic diagram of the position-guided and consistency learning incremental network Net3 in the present invention;
[0030] Figure 5 is the visual comparison diagram when the present invention and the existing comparison method segment the pancreas;
[0031] Figure 6 is the visual comparison diagram when the present invention and the existing comparison method segment the duodenum. Detailed Embodiment
[0032] The embodiments and effects of the present invention will be further described in detail below with reference to the drawings.
[0033] Refer to Figure 1 , the implementation steps of this example are as follows:
[0034] Step 1. Obtain four publicly available abdominal CT datasets, preprocess them, and divide them into a training set and a test set.
[0035] (1.1) Obtain four publicly available abdominal CT datasets, including three labeled abdominal CT datasets DenseVNet, MICCAI FLARE21, and AMOS 2022, and an unlabeled abdominal CT dataset MICCAI FLARE22;
[0036] (1.2) Retain the labels of the liver, kidney, spleen, gallbladder, and stomach in the labeled DenseVNet dataset, retain the label of the pancreas in the labeled MICCAI FLARE21 dataset, retain the label of the duodenum in the labeled AMOS 2022 dataset, and set the remaining unused labels as the background;
[0037] (1.3) Crop the abdominal CT images in the four datasets to crop off the useless black edges;
[0038] (1.4) Uniformly adjust the sizes of the cropped abdominal CT images and labels to 144×144×144;
[0039] (1.5) Normalize the resized abdominal CT images to obtain the normalized images Y:
[0040]
[0041] where X is the input image, X min is the maximum grayscale value of the pixel points of the input image, and X max is the minimum grayscale value of the pixel points of the input image;
[0042] (1.6) Randomly partition the first three datasets DenseVNet, MICCAI FLARE21, and AMOS 2022 respectively to obtain three training sets A1, A2, A3 and three test sets B1, B2, B3, and use the entire fourth dataset MICCAI FLARE22 as the fourth training set C.
[0043] Step 2. Construct the basic network Net1.
[0044] As Figure 2 shown, the specific implementation of this step is as follows:
[0045] (2.1) Build an encoder consisting of five convolutional blocks and four max pooling layers. Each convolutional block consists of a 3D convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the 3D convolutional layer is set to 3×3×3, the stride is set to 1, and the padding is set to "SAME";
[0046] (2.2) Build a decoder consisting of four convolutional blocks, four upsampling layers, and a Softmax activation layer. Each convolutional block consists of a 3D convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the 3D convolutional layer is set to 3×3×3, the stride is set to 1, and the padding is set to "SAME". The output of the last convolutional block in the decoder is set to 6 channels;
[0047] (2.3) Build a deep supervision module consisting of a 3D convolutional layer and a Softmax activation layer. The convolutional kernel size of the 3D convolutional layer is 1×1×1, the stride is 1, and the number of channels is 6;
[0048] (2.4) Serially connect the encoder and the decoder, and serially connect the deep supervision module and the output of the second convolutional block in the decoder to form the basic network Net1.
[0049] Step 3. Initialize the network parameters of the basic network Net1, and use the labeled first training set A1 and the unlabeled fourth training set C to alternately iterate and train it through the consistency learning strategy to obtain the trained basic network parameters W1.
[0050] (3.1) Initialize the network parameters W1' of the basic network Net1 using the He initialization method. After initialization, the parameters W1' of the basic network Net1 follow the distribution:
[0051]
[0052] where n l is the number of neurons in the l-th layer of the basic network Net1, represents the normal distribution with an expected value of 0 and a variance of ;
[0053] (3.2) Take a CT image from the first training set A1 as the input of the basic network Net1, and obtain its prediction result P A and the deep supervision output P1, and calculate the loss function value L DC of the basic network Net1:
[0054]
[0055] where G A is the corresponding true label, and G1 is the result of reducing the true label size by 1 / 4;
[0056] (3.3) Backpropagate the calculated L DC loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters W1' of the basic network Net1;
[0057] (3.4) Repeat the process of (3.2)-(3.3) until all data in the training set A1 have been learned, and complete one supervised learning,
[0058] (3.5) Repeat (3.4) five times to obtain the parameters W1' of the basic network Net1 after multiple supervised learnings;
[0059] (3.6) Take a CT image X u from the fourth training set C as the input of the basic network Net1. At the same time, transpose X u to obtain T(X u ), and also use it as the input of the basic network Net1, and calculate the value of the consistency loss function L CL :
[0060]
[0061] where pi represents the i-th pixel of the prediction result of the basic network Net1 for T(X u ), where N represents the total number of pixels; represents the i-th pixel after transposing the prediction result of the basic network Net1 for X u , and N represents the total number of pixels;
[0062] (3.7) Backpropagate the calculated L CL loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters W1' of the basic network Net1;
[0063] (3.8) Repeat steps (3.6) - (3.7) until all data in the fourth training set C have been learned, completing one unsupervised learning;
[0064] (3.9) Repeat steps (3.2) - (3.8) a total of 10 times to obtain the parameters W1' of the basic network Net1 after alternating iterative training;
[0065] (3.10) Execute (3.5) again on the parameters W1 of the basic network Net1 to obtain the trained basic network parameters W1.
[0066] Step 4. Load the trained basic network parameters W1 into the basic network Net1. On the basis of the basic network Net1, add a new decoder branch, a position guidance module, and a deep supervision module to construct the position guidance incremental network Net2.
[0067] As Figure 3 shown, the specific implementation of this step is as follows:
[0068] (4.1) Establish a new decoder composed of three convolutional blocks, three upsampling layers, and a Softmax activation layer. Each convolutional block is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the three-dimensional convolutional layer is 3×3×3, the stride is 1, and the padding is set to "SAME". The output of the last convolutional block has 2 channels;
[0069] (4.2) Establish the first position guidance module of the position guidance incremental network. Its input is the output E1 of the first convolutional block of the encoder in the basic network Net1 and the output P1 of the deep supervision module of the basic network Net1. Calculate the background probability map BP 11 and the foreground probability map FP 11 of the position guidance incremental network in this module:
[0070] BP 11 = 1 - P10
[0071] FP11 = 1 - BP 11
[0072] where P10 is the first - dimensional channel of P1;
[0073] (4.3) Upsample the background probability map BP 11 and the foreground probability map FP 11 of the position - guiding incremental network to the same size as E1, and multiply the upsampled background probability map BP 11 and foreground probability map FP 11 of the position - guiding incremental network with E1 pixel - by - pixel respectively to obtain the background feature map E1 of the position - guiding incremental network after position attention BP and the foreground feature map E1 FP . Then, splice the background feature map E1 BP and the foreground feature map E1 FP of the position - guiding incremental network along the channel direction to obtain the position - attention feature map E1' of the position - guiding incremental network. Finally, use a channel - attention SE module and a 1×1×1 convolutional layer to perform channel - feature attention on E1';
[0074] (4.4) Establish the second position - guiding module of the position - guiding incremental network. Its input is the output E2 of the second convolutional block of the encoder in the basic network Net1 and the output P1 of the depth - supervision module of the basic network Net1. Calculate the background probability map BP 12 and the foreground probability map FP 12 in this module:
[0075] BP 12 = 1 - P10
[0076] FP 12 = 1 - BP 12
[0077] where P10 is the first - dimensional channel of P1;
[0078] (4.5) Upsample the background probability map BP 12 and the foreground probability map FP 12 of the position - guiding incremental network to the same size as E2, and multiply the upsampled background probability map BP 12 and the foreground probability map FP 12 with E2 pixel - by - pixel respectively to obtain the background feature map E2 of the position - guiding incremental network after position attention BP and the foreground feature map E2 FP . Then, along the channel direction, for the background feature map E2 BP and the foreground feature map E2 of the position - guiding incremental networkFP Perform splicing to obtain the position attention feature map E2' of the position guidance incremental network. Finally, use a channel attention SE module and a 1×1×1 convolutional layer to perform channel feature attention on E2'.
[0079] (4.6) Establish a deep supervision module composed of a three-dimensional convolutional layer and a Softmax activation layer. The convolutional kernel size of the three-dimensional convolutional layer is 1×1×1, the stride is 1, and the number of channels is 2.
[0080] (4.7) Connect the newly added decoder in series with the first convolutional block of the decoder in the basic network Net1. Add the first position guidance module of the position guidance incremental network between the first convolutional block of the encoder of the basic network Net1 and the third convolutional block of the newly added decoder. Add the second position guidance module of the position guidance incremental network between the second convolutional block of the encoder of the basic network Net1 and the second convolutional block of the newly added decoder. Connect the deep supervision module in series with the output of the first convolutional block in the newly added decoder to form the position guidance incremental network Net2.
[0081] Step 5. Use the second training set A2 and the Adam optimizer to iteratively train the position guidance incremental network Net2 to obtain the trained position guidance incremental network parameters W2.
[0082] (5.1) Initialize the parameters of the newly added structure of the position guidance incremental network Net2 using the He initialization method.
[0083] (5.2) Freeze the basic network parameters W1.
[0084] (5.3) Take a CT image from the second training set A2 as the input of the position guidance incremental network Net2 to obtain the prediction result of its newly added decoder and the deep supervision output, and calculate the current L DC loss function value.
[0085] (5.4) Backpropagate the calculated L DC loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters of the newly added decoder, two position guidance modules, and deep supervision module in the position guidance incremental network Net2.
[0086] (5.5) Repeat the process of (5.3)-(5.4) until all data in the training set A2 have been learned, completing one iteration of learning.
[0087] (5.6) Repeat the process of (5.5) until the loss function value no longer decreases to end the training and obtain the position guidance incremental network parameters W2.
[0088] Step 6. Load the trained location guidance incremental network parameter W2 into the location guidance incremental network Net2, and on the basis of the location guidance incremental network Net2, add a new decoder branch, a location guidance module, and a deep supervision module, so as to construct the location guidance and consistency learning incremental network Net3.
[0089] As Figure 4 shown, the specific implementation of this step is as follows:
[0090] (6.1) Establish a new decoder composed of three convolutional blocks, three upsampling layers, and a Softmax activation layer. Each convolutional block is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the three-dimensional convolutional layer is 3×3×3, the stride is 1, and the padding is set to "SAME". The output of the last convolutional block has 2 channels;
[0091] (6.2) Establish the first location guidance module of the location guidance and consistency learning incremental network. Its inputs are the output E1 of the first convolutional block of the encoder in the basic network Net1, the output P1 of the deep supervision module of the basic network Net1, and the output P2 of the deep supervision module of the location guidance incremental network Net2. Calculate the background probability map BP 21 and the foreground probability map FP 21 of the location guidance and consistency learning incremental network:
[0092] BP 21 = max(1 - P10, P21)
[0093] FP 21 = 1 - BP 11
[0094] where P10 is the first-dimensional channel of P1 and P21 is the second-dimensional channel of P2;
[0095] (6.3) Upsample the background probability map BP 21 and the foreground probability map FP 21 of the location guidance and consistency learning incremental network to the same size as E1, and multiply the upsampled background probability map BP 21 and foreground probability map FP 21 of the location guidance and consistency learning incremental network with E1 pixel by pixel respectively to obtain the background feature map E1 BP ' and foreground feature map E1 FP ' of the location guidance and consistency learning incremental network after location attention; then, along the channel direction, for the background feature map E1 BP ' and foreground feature map E1 FP'Concatenate to obtain the position attention feature map E1" of the position guidance and consistency learning incremental network. Finally, use a channel attention SE module and a 1×1×1 convolutional layer to perform channel feature attention on E1".
[0096] (6.4) Establish the second position guidance module of the position guidance and consistency learning incremental network. Its inputs are the output E2 of the second convolutional block of the encoder in the base network Net1, the output P1 of the depth supervision module of the base network Net1, and the output P2 of the depth supervision module of the position guidance incremental network Net2. Calculate the background probability map BP of the position guidance and consistency learning incremental network within this module 22 and the foreground probability map FP 22 :
[0097] BP 22 = max(1 - P10, P21)
[0098] FP 22 = 1 - BP 11
[0099] where P10 is the first-dimensional channel of P1 and P21 is the second-dimensional channel of P2;
[0100] (6.5) Upsample the background probability map BP 22 and the foreground probability map FP 22 of the position guidance and consistency learning incremental network to the same size as E2, and multiply the upsampled background probability map BP 22 and foreground probability map FP 22 of the position guidance and consistency learning incremental network with E2 pixel by pixel to obtain the background feature map E2' BP ' and foreground feature map E2 FP ' of the position guidance and consistency learning incremental network after position attention. Then concatenate the background feature map E2 BP ' and foreground feature map E2 FP ' of the position guidance and consistency learning incremental network along the channel direction to obtain the position attention feature map E2" of the position guidance and consistency learning incremental network. Finally, use a channel attention SE module and a 1×1×1 convolutional layer to perform channel feature attention on E2";
[0101] (6.6) Establish a depth supervision module composed of a three-dimensional convolutional layer and a Softmax activation layer. The convolutional kernel size of this three-dimensional convolutional layer is 1×1×1, the stride is 1, and the number of channels is 2;
[0102] (6.7) Connect the newly added decoder in series with the first convolutional block of the decoder in the basic network Net1. Add the first position guidance module of the position guidance and consistency learning incremental network between the first convolutional block of the encoder in the basic network Net1 and the third convolutional block of the newly added decoder. Add the second position guidance module of the position guidance and consistency learning incremental network between the second convolutional block of the encoder in the basic network Net1 and the second convolutional block of the newly added decoder. Connect the depth supervision module in series with the output of the first convolutional block in the newly added decoder to form the position guidance and consistency learning incremental network Net3.
[0103] Step 7. Use the third training set A3 and the Adam optimizer to iteratively train the position guidance and consistency learning incremental network Net3 to obtain the trained parameters W3 of the position guidance and consistency learning incremental network.
[0104] (7.1) Initialize the parameters of the newly added structure of the position guidance and consistency learning incremental network Net3 using the He initialization method;
[0105] (7.2) Freeze the parameters W2 of the position guidance incremental network;
[0106] (7.3) Take a CT image from the third training set A3 as the input of the position guidance and consistency learning incremental network Net3 to obtain the prediction result of its newly added decoder and the depth supervision output, and calculate the current L DC loss function value;
[0107] (7.4) Backpropagate the calculated L DC loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters of the newly added decoder, two position guidance modules, and depth supervision module in the position guidance and consistency learning incremental network Net3;
[0108] (7.5) Repeat the process of (7.3)-(7.4) until all data in the training set A3 have been learned, completing one iteration of learning;
[0109] (7.6) Repeat the process of (7.5) and perform multiple iterations of learning until the loss function value no longer decreases to end the training and obtain the parameters W3 of the position guidance and consistency learning incremental network.
[0110] Step 8. And so on. According to the same method as in Steps 6 and 7, the position guidance and consistency learning incremental network Net3 can be continuously extended subsequently to segment more abdominal organs.
[0111] Step 9. Input the first test set B1, the second test set B2, and the third test set B3 into the position guidance and consistency learning incremental network Net3 to obtain the prediction results of the network for the liver, kidney, spleen, gallbladder, stomach, pancreas, and duodenum.
[0112] The effects of the present invention can be further illustrated by the following simulations.
[0113] 1. Simulation conditions
[0114] The simulation platform for this experiment is a desktop computer with 32GB of memory. Its CPU model is Intel(R) Core(TM) CPU i9-10920X@3.50GHZ. The neural network model is built and trained using Python 3.7, keras 2.5.0, and tensorflow 2.5.0, and accelerated using NVIDIA GeForce RTX 3090 (24GB of video memory), Cuda 11.1, and Cudnn 8.1.0.
[0115] The segmentation performance evaluation index used in the simulation is the Dice coefficient, and its calculation formula is as follows:
[0116]
[0117] Among them, A represents the true label, and B represents the prediction result.
[0118] 2. Simulation content
[0119] In Simulation 1, using the first training set A1, the second training set A2, the third training set A3, and the fourth training set C, the incremental segmentation method proposed in the present invention and the existing six incremental segmentation methods, namely fine-tuning FT, LWF, separate segmentation, joint training 1, and joint training 2, are used to train the abdominal multi-organ incremental segmentation network respectively. Then, the test sets B1, B2, and B3 are used to test the incremental segmentation networks trained by their respective incremental segmentation strategies to obtain the test set segmentation results of each segmentation method. The Dice coefficient between the test set segmentation results of each method and the true labels of the test set is calculated, as shown in Table 1:
[0120] Table 1 Dice coefficients (%) of the test results of the present invention and existing comparison methods
[0121] Method Liver Kidney Spleen Gallbladder Stomach Pancreas Duodenum Mean FT 0 0 0 0 0 0 64.5±15.1 - LWF 91.9±2.9 80.0±16.9 89.6±8.3 71.5±21.5 64.7±17.8 56.2±17.5 63.9±13.9 74.0±14.1 Separate segmentation 95.0±1.4 91.8±14.1 93.1±7.7 79.8±16.9 83.1±10.1 78.1±9.9 64.2±14.4 83.6±10.6 Joint training 1 94.0±1.6 90.1±13.5 92.9±6.4 73.9±18.9 80.8±8.8 79.1±7.5 71.0±11.5 83.1±9.7 Joint training 2 93.3±2.1 89.9±13.7 92.0±6.0 71.7±19.4 78.1±9.3 78.9±7.9 68.6±12.1 81.8±10.1 The present invention 95.4±1.2 92.1±14.2 93.2±8.3 79.6±15.5 85.4±9.4 79.4±7.9 70.6±10.9 85.1±9.6
[0122] The FT mentioned above refers to fine-tuning both the encoder and decoder of the old network to incrementally segment new organs. As can be seen from the first row of Table 1, using FT for incremental segmentation results in a Dice coefficient of 0 for the segmentation of the liver, kidney, spleen, gallbladder, stomach, and pancreas, indicating that this method causes serious forgetting of old organs.
[0123] The LWF refers to a method of protecting old knowledge from being overwritten by new knowledge by imposing knowledge distillation constraints on the loss function of the new network. As can be seen from the second row of Table 1, the Dice coefficients of the results of incremental segmentation using LWF on each organ are lower than those of the method proposed in the present invention.
[0124] The separate segmentation refers to training three segmentation networks using the first training set A1, the second training set A2, and the third training set A3 respectively. As can be seen from the third row of Table 1, the performance of separate segmentation is inferior to that of the method of the present invention. In particular, the Dice coefficient of the duodenum is 6.4% lower than that of the method proposed in the present invention.
[0125] The joint training 1 refers to training a multi-branch network using the first training set A1, the second training set A2, and the third training set A3 simultaneously. As can be seen from the third row of Table 1, the average performance of the seven organs in joint training 1 is 2% lower than that of the method proposed in the present invention.
[0126] The joint training 2 refers to training a single-branch network using the first training set A1, the second training set A2, and the third training set A3 simultaneously. As can be seen from the third row of Table 1, the average performance of the seven organs in joint training 2 is 3.3% lower than that of the method proposed in the present invention.
[0127] Based on the analysis of Table 1 above, the average Dice coefficient of the present invention is increased by 1.5% - 11.1% compared with the existing methods, indicating that the present invention significantly improves the incremental segmentation performance.
[0128] In Simulation 2, five test cases in the test set B2 were randomly selected, and the pancreatic segmentation results of the present invention and the prior art were visualized to compare their accuracies. The results are as Figure 5 .
[0129] From Figure 5 it can be seen that the present invention has higher accuracy when expanding the segmentation of the pancreas than the existing comparative methods, with fewer missegmentations and undersegmentations.
[0130] In Simulation 3, five test cases in the test set B3 were randomly selected, and the duodenal segmentation results of the present invention and the prior art were visualized to compare their accuracies. The results are as Figure 6 .
[0131] From Figure 6 it can be seen that the present invention has higher segmentation performance when expanding the segmentation of the duodenum than the existing comparative methods, demonstrating the effectiveness and superiority of the abdominal CT multi-organ incremental segmentation method based on position guidance and consistency learning proposed in the present invention.
[0132] The above simulation results show that the present invention has higher abdominal multi-organ incremental segmentation performance than the prior art.
Claims
1. An abdominal multi-organ incremental segmentation method based on position guidance and consistency learning, characterized in that, Including: (1) Divide the dataset: 1.1) Obtain four publicly available abdominal CT datasets. The first three datasets are labeled with the liver, kidney, spleen, gallbladder, stomach, pancreas, and duodenum respectively, and the fourth dataset has no annotations; 1.2) Perform image preprocessing on the four datasets obtained in 1.1), and randomly divide the preprocessed first three datasets into three training sets A1, A2, A3 and three test sets B1, B2, B3, and use the fourth dataset as the fourth training set C in its entirety; (2) Based on the UNet network structure, construct a basic network Net1 composed of an encoder, a decoder, and a deep supervision module; (3) Initialize the basic network Net1 using the He initialization method, and use the labeled first training set A1 and the unlabeled fourth training set C to alternately iterate and train it through a consistency learning strategy to obtain the trained basic network parameters W1; (4) Load the trained basic network parameters W1 into the basic network Net1, and add a new decoder, two position guidance modules, and a deep supervision module to Net1 to form a position guidance incremental network Net2; (5) Initialize the parameters of the newly added structure of the position guidance incremental network Net2 using the He initialization method, set its loss function to the Dice loss, freeze the trained basic network parameters W1, and use the second training set A2 and the Adam optimizer to iteratively train the position guidance incremental network Net2 until the loss function value no longer decreases, obtaining the trained position guidance incremental network parameters W2; (6) Load the trained position guidance incremental network parameters W2 into the position guidance incremental network Net2, and add a new decoder, two position guidance modules, and a deep supervision module to Net2 to form a position guidance and consistency learning incremental network Net3; (7) Initialize the parameters of the newly added structure of the position guidance and consistency learning incremental network Net3 using the He initialization method, set its loss function to the Dice loss, freeze the trained position guidance incremental network parameters W2, and use the third training set A3 and the Adam optimizer to iteratively train the position guidance and consistency learning incremental network Net3 until the loss function value no longer decreases, obtaining the trained position guidance and consistency learning incremental network parameters W3; (8) And so on, the position guidance and consistency learning incremental network Net3 can be continuously expanded in the future to segment more abdominal organs; (9) Input the first test set B1, the second test set B2, and the third test set B3 into the position guidance and consistency learning incremental network Net3 to obtain the prediction results of the network for the liver, kidney, spleen, gallbladder, stomach, pancreas, and duodenum.
2. The method according to claim 1, wherein In step (1.1), the four publicly available abdominal CT datasets obtained include three labeled abdominal CT datasets DenseVNet, MICCAI FLARE21, and AMOS 2022, and an unlabeled abdominal CT dataset MICCAI FLARE22.
3. The method according to claim 1, characterized in that, In step (1.2), image preprocessing is performed on the four datasets obtained in (1.1) as follows: 1.2a) Retain the labels of the liver, kidney, spleen, gallbladder, and stomach in the labeled DenseVNet dataset, retain the label of the pancreas in the labeled MICCAI FLARE21 dataset, retain the label of the duodenum in the labeled AMOS 2022 dataset, and set the remaining unused labels as the background; 1.2b) Crop the abdominal CT images in the four datasets to crop out the useless black edges; 1.2c) Uniformly adjust the sizes of the cropped abdominal CT images and labels to 144×144×144; 1.2d) Normalize the resized abdominal CT images to obtain the normalized image Y: Among them, X is the input image, X min is the maximum gray value of the pixel points of the input image, X max is the minimum gray value of the pixel points of the input image.
4. The method according to claim 1, wherein In step (2), the basic network Net1 composed of an encoder, a decoder, and a deep supervision module is implemented as follows: 2a) Build an encoder composed of five convolutional blocks and four max-pooling layers. Each convolutional block is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the three-dimensional convolutional layer is set to 3×3×3, the stride is set to 1, and the padding is set to "SAME"; 2b) Build a decoder composed of four convolutional blocks, four upsampling layers, and a Softmax activation layer. Each convolutional block is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the three-dimensional convolutional layer is set to 3×3×3, the stride is set to 1, and the padding is set to "SAME". The output of the last convolutional block in the decoder is set to 6 channels; 2c) Build a deep supervision module composed of a three-dimensional convolutional layer and a Softmax activation layer. The size of the convolutional kernel of the three-dimensional convolutional layer is 1×1×1, the stride is 1, and the number of channels is 6; 2d) Serially connect the encoder and the decoder, and serially connect the deep supervision module and the output of the second convolutional block in the decoder to form the basic network Net1.
5. The method according to claim 1, characterized in that, In step (3), the basic network Net1 is initialized using the He initialization method, and the distribution function it follows is: Among them, represents a normal distribution with an expected value of 0 and a variance of , and n l is the number of neurons in the l-th layer of the basic network Net1.
6. The method according to claim 1, wherein In step (3), the labeled first training set A1 and the unlabeled fourth training set C are used to alternately iteratively train it through a consistency learning strategy as follows: (3a) Take a CT image from the first training set A1 as the input of the basic network Net1, and obtain its prediction result P A and the deep supervision output P1, and calculate L DC Loss function value: where G A is the corresponding true label, and G1 is the result of reducing the true label size by 1 / 4; (3b) Backpropagate the calculated L DC to obtain the gradient value, and use the Adam optimizer to update the parameter W1' of the basic network Net1; (3c) Repeat the process of (3a)-(3b) until all data in the training set A1 have been learned, completing one supervised learning, (3d) Repeat (3c) five times to obtain the parameters W1' of the basic network Net1 after multiple supervised learnings; (3e) Take a CT image X from the fourth training set C u as the input of the basic network Net1, and at the same time, for X u perform transposition to obtain T(X u ), and also use it as the input of the basic network Net1 to calculate the value of the consistency loss function L CL : where p i represents the i-th pixel of the prediction result of the base network Net1 for T(X u ), and N represents the total number of pixels; represents the i-th pixel after transposing the prediction result of the base network Net1 for X u ; N represents the total number of pixels. (3f)Backpropagate the calculated L CL loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters W1' of the basic network Net1; (3g) Repeat steps (3e)-(3f) until all data in the fourth training set C have been learned, completing one unsupervised learning; (3h) Repeat (3a)-(3g) a total of 10 times to obtain the parameters W1' of the basic network Net1 after alternating iterative training; (3i) Perform (3d) on the parameters W1 of the basic network Net1 again to obtain the parameters W1 of the basic network after training is completed.
7. The method according to claim 1, wherein In step (4), the position-guided incremental network Net2 is constructed as follows: (4a) An additional decoder is established, which consists of three convolutional blocks, three upsampling layers, and a Softmax activation layer. Each convolutional block consists of a three-dimensional convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of the three-dimensional convolutional layer is 3×3×3, the stride is 1, and the padding is set to "SAME". The output of the last convolutional block has 2 channels; (4b) Establish the first position guidance module, whose input is the output E1 of the first convolutional block of the encoder in the basic network Net1 and the output P1 of the depth supervision module of the basic network Net1, and calculate the background probability map BP within this module 11 and the foreground probability map FP 11 : BP 11 = 1 - P10 FP 11 = 1 - BP 11 where P10 is the first-dimensional channel of P1; (4c) Upsample the background probability map BP 11 and the foreground probability map FP 11 to the same size as E1, and multiply the upsampled background probability map BP 11 and the foreground probability map FP 11 pixel - by - pixel with E1 respectively to obtain the background feature map E1 after position attention BP and the foreground feature map E1 FP , concatenate the background feature map E1 BP and the foreground feature map E1 FP along the channel dimension to obtain the position - attention feature map E1'. Finally, use a channel - attention SE module and a 1×1×1 convolutional layer to perform channel - feature attention on E1'. (4d) Establish a second position guidance module, whose inputs are the output E2 of the second convolutional block of the encoder in the basic network Net1 and the output P1 of the depth supervision module of the basic network Net1, and calculate the background probability map BP within this module 12 and the foreground probability map FP 12 : BP 12 = 1 - P10 FP 12 = 1 - BP 12 where P10 is the first-dimensional channel of P1; (4e) Upsample the background probability map BP 12 and the foreground probability map FP 12 to the same size as E2, and then multiply the upsampled background probability map BP 12 and the foreground probability map FP 12 pixel by pixel with E2 respectively to obtain the background feature map E2 after position attention BP and the foreground feature map E2 FP . Concatenate the background feature map E2 BP and the foreground feature map E2 FP along the channel dimension to obtain the position attention feature map E2'. Finally, use a channel attention SE module and a 1×1×1 convolutional layer to perform channel feature attention on E2'. (4f) A deep supervision module is established, which consists of a three-dimensional convolutional layer and a Softmax activation layer. The convolutional kernel size of this three-dimensional convolutional layer is 1×1×1, the stride is 1, and the number of channels is 2; (4g) The additional decoder is concatenated with the first convolutional block of the decoder in the basic network Net1. The first position guidance module is added between the first convolutional block of the encoder in the basic network Net1 and the third convolutional block of the additional decoder. The second position guidance module is added between the second convolutional block of the encoder in the basic network Net1 and the second convolutional block of the additional decoder. The deep supervision module is concatenated with the output of the first convolutional block in the additional decoder to form the position guidance incremental network Net2.
8. The method according to claim 1, wherein In step (5), the second training set A2 and the Adam optimizer are used to iteratively train the position guidance incremental network Net2, achieving the following: (5a) Freeze the parameters W1 of the basic network; (5b) Take a CT image from the second training set A2 as the input of the position-guided incremental network Net2, obtain the prediction results of its new decoder and the deep supervision output, and calculate the current L DC loss function value; (5c)Backpropagate the calculated L DC loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters of the newly added decoder, two position guidance modules, and depth supervision module in the position guidance increment network Net2; (5d) Repeat the process of (5b)-(5c) until all data in the training set A2 have been learned, completing one iteration of learning; (5e) Repeat the process of (5d) for multiple iterations of learning until the value of the loss function no longer decreases, and then end the training to obtain the position guidance incremental network parameters W2.
9. The method according to claim 1, wherein In step (6), the position guidance and consistency learning incremental network Net3 is constructed, achieving the following: (6a) An additional decoder is established, which consists of three convolutional blocks, three upsampling layers, and a Softmax activation layer. Each convolutional block consists of a three-dimensional convolutional layer, a BN layer, and a Relu activation layer. The convolutional kernel of this three-dimensional convolutional layer is 3×3×3, the stride is 1, and the padding is set to "SAME". The output of the last convolutional block has 2 channels; (6b) Establish the first position guidance module, whose inputs are the output E1 of the first convolutional block of the encoder in the basic network Net1, the output P1 of the depth supervision module of the basic network Net1, and the output P2 of the depth supervision module of the position guidance incremental network Net2. Calculate the background probability map BP within this module 21 and the foreground probability map FP 21 : BP 21 = max(1 - P10, P21) FP 21 = 1 - BP 11 where P10 is the first-dimensional channel of P1, and P21 is the second-dimensional channel of P2; (6c) Upsample the background probability map BP 21 and the foreground probability map FP 21 to the same size as E1, and then multiply the upsampled background probability map BP 21 and the foreground probability map FP 21 pixel - by - pixel with E1 respectively to obtain the background feature map E1' after position attention BP and the foreground feature map E1 FP '. Concatenate the background feature map E1 BP ' and the foreground feature map E1 FP ' along the channel direction to obtain the position - attention feature map E1". Finally, use a channel - attention SE module and a 1×1×1 convolutional layer to perform channel - feature attention on E1". (6d) Establish a second position guidance module, whose inputs are the output E2 of the second convolutional block of the encoder in the basic network Net1, the output P1 of the depth supervision module of the basic network Net1, and the output P2 of the depth supervision module of the position guidance incremental network Net2. Calculate the background probability map BP within this module 22 and the foreground probability map FP 22 : BP 22 = max(1 - P10, P21) FP 22 = 1 - BP 11 where P10 is the first-dimensional channel of P1, and P21 is the second-dimensional channel of P2; (6e) Upsample the background probability map BP 22 and the foreground probability map FP 22 to the same size as E2, and then multiply the upsampled background probability map BP 22 and the foreground probability map FP 22 pixel by pixel with E2 respectively to obtain the background feature map E2 BP ' after position attention and the foreground feature map E2 FP '. Then concatenate the background feature map E2 BP ' and the foreground feature map E2 FP ' along the channel dimension to obtain the position attention feature map E2”. Finally, use a channel attention SE module and a 1×1×1 convolutional layer to perform channel feature attention on E2”; (6f) A deep supervision module is established, which consists of a three-dimensional convolutional layer and a Softmax activation layer. The convolutional kernel size of this three-dimensional convolutional layer is 1×1×1, the stride is 1, and the number of channels is 2; (6g) Connect the newly added decoder in series with the first convolutional block of the decoder in the basic network Net1. Add the first position guiding module between the first convolutional block of the encoder in the basic network Net1 and the third convolutional block of the newly added decoder. Add the second position guiding module between the second convolutional block of the encoder in the basic network Net1 and the second convolutional block of the newly added decoder. Connect the deep supervision module in series with the output of the first convolutional block in the newly added decoder to form the position guiding and consistency learning incremental network Net3.
10. The method according to claim 1, characterized in that In step (7), use the third training set A3 and the Adam optimizer to iteratively train the position guiding and consistency learning incremental network Net3, as follows: (7a) Freeze the parameters W2 of the position guiding incremental network; (7b) Take a CT image from the third training set A3 as the input of the position-guided and consistency learning incremental network Net3, obtain the prediction results of its new decoder and the deep supervision output, and calculate the current L DC loss function value; (7c)Backpropagate the calculated loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters of the newly added decoder, two position guidance modules, and the deep supervision module in the position guidance and consistency learning incremental network Net3; DC Backpropagate the calculated loss function value to obtain the gradient value, and use the Adam optimizer to update the parameters of the newly added decoder, two position guidance modules, and the deep supervision module in the position guidance and consistency learning incremental network Net3; (7d) Repeat the process of (7b)-(7c) until all data in the training set A3 have been learned, completing one iteration of learning; (7e) Repeat the process of (7d) for multiple iterations of learning until the value of the loss function no longer decreases, ending the training and obtaining the parameters W3 of the position guiding and consistency learning incremental network.
Citation Information
Patent Citations
Brain tumor segmentation network and segmentation method based on U-Net network
CN111192245A
Liver CT automatic segmentation method based on deep shape learning
CN113674281A