Method and device for training neural network for organ segmentation
Patent Information
- Authority / Receiving Office
- HK · HK
- Patent Type
- Patents
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2023-02-27
- Publication Date
- 2026-07-17
AI Technical Summary
Existing deep learning-based organ segmentation methods lack awareness of the anatomical shape of the target organ, resulting in inconsistent segmentation outputs that require post-processing for correction, which is especially noticeable in 3D scenes.
By training a deep 3D U-net neural network, combining a segmentation map and a signed distance map (SDM), connecting the two outputs using a differentiable Heaviside function, and jointly training with a specific loss function, we can directly predict organ segmentation with smooth surfaces and reduced noise.
It enables the direct output of organ segments with smooth surfaces and less noise without post-processing, improving the accuracy and stability of segmentation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] Cross-referencing related applications
[0002] This application claims priority to U.S. Application No. 16 / 869,012, filed May 7, 2020, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0003] This disclosure relates to computer vision (e.g., object detection (identifying objects in images and videos)) and artificial intelligence. Specifically, it relates to a computer-implemented method and apparatus for training a neural network for organ segmentation. Background Technology
[0004] List of related technologies
[0005] Non-Patent Literature 1: Scher, AL; Xu, Y.; Korf, E.; White, LR; Schelten ns, P.; Toga, AW; Thompson, PM; Hartley, S.; Witter, M.; Valentino, DJ; et al., March 12, 2007, “Hippocampal Shape Analysis in Alzheimer’s Disease: A Population-Based Study.” Neuroimage; May 15, 2007; 36(1): pp. 8-18. Electronic version March 12, 2007.
[0006] Non-patent literature 2: Moore, KL; Brame, RS; Low, DA; and Mutic, S.; 2011. “Experience-Based Quality Control of Clinical Intensity Modulated Radiotherapy Planning.” International Journal of Radiation Oncology*Biology*Physics 81(2): 545-551.
[0007] Non-patent literature 3: Kass, M.; Witkin, A.; and Terzopoulos, D. 1988. “Snakes: Active Contour Models”. IJCV 1(4): pp. 321-331.
[0008] Non-patent literature 4: Osher, S., and Sethian, JA; 1988. "Fronts Propagating with Curvature-Dependent speed: Algorithms based on Hamilton-Jacobi formulations." Journal of computational physics 79(1): pp. 12-49.
[0009] Non-patent literature 5: Cerrolaza, JJ; Summers, RM; Gonz'alez Ballester, MA'; and L1nguraru, MG; 2015 "Automatic Multi-Resolution Shape Mode Ling".
[0010] Non-Patent Literature 6: Aljabar, P.; Heckemann, RA; Hammers, A.; Hajnal, JV; and Rueckert, D; 2009; “Multi-Atlas Based Segmentation of Brain Images: Atlas Selection and Its Effect On Accuracy”; Neuroimage 46(3): 726-738.
[0011] Non-patent literature 7: Ronneberger, O.; Fischer, P.; and Brox, T.; 2015; U-Net: Convolutional Networks for Biomedical Image Segmentation; Medical Image Computing and Computer Assisted Intervention (In MICCAI, pp. 234-241; Springer).
[0012] Non-patent literature 8: ( O.; Abdulkadir, A.; Lienkamp, SS; Brox, T.; and Ronneberger, O.; 2016. “3d U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation”; (In MICCAI, pp. 424–432; Springer).
[0013] Non-Patent Literature 9: Kamnitsas, K.; Ledig, C.; Newcombe, VF; Simpson, JP; Kane, AD; Menon, DK; Rueckert, D.; and Glocker, B.; 2017; “Efficient Multi-Scale 3d CNN With Fully Connected CRF For Accurate Brain Lesion Segmentation”; MedIA 36: pp. 61-78.
[0014] Non-patent literature 10: Kohlberger, T.; Sofka, M.; Zhang, J.; Birkbeck, N.; Wetzl, J.; Kaftan, J.; Declerck, J.; and Zhou, SK; 2011; “Automatic Multi-Organ Segmentation Using Learning-Based Segmentation And Level Set Optimization”; (In MICCAI, pp. 338-345; Springer).
[0015] Non-patent literature 11: Perera, S.; Barnes, N.; He, X.; Izadi, S.; Kohli, P.; and Glocker, B.; 2015; “Motion Segmentation Of Truncated Ssigned Distance Function Based Volumetric Surfaces”; (In WACV, pp. 1046-1053. IEEE).
[0016] Non-patent literature 12: Hu, P.; Shuai, B.; Liu, J.; and Wang, G.; 2017; “Deep Level Sets for Sa1ient Object Detection”; (In CVPR, pp. 2300-2309).
[0017] Non-patent literature 13: Park, JJ; Florence, P.; Straub, J.; Newcombe, R.; and Lovegrove, S.; 2019; "Deepsdf: Learning Continuous Signed Distance Functions For Shape Representation"; arXiv preprint arXiv: 1901.05103.
[0018] Non-patent literature 14: A1 Arif, SMR; Knapp, K.; and Slabaugh, G.; 2018; “Spnet: ShapePrediction Using a Fully Convolutional Neural Network”; (In MICCAI, pp. 430-439; Springer).
[0019] Non-Patent Literature 15: Dangi, S.; Yaniv, Z.; and Linte, C.; 2019; “A Distance Map Regularized CNN For Cardiac Cine MR Image Segmentation”; arXiv preprint arXiv:1901.01238.
[0020] Non-Patent Literature 16: Navarro, F.; Shit, S.; Ezhov, I; Paetzold, J.; Gafita, A.; Peeken, JC; Combs, SE; and Menze, BH; 2019; “Shape-Aware Complementary-Task Learning For Multi-Organ Segmentation”; (In MIDL, pp. 620-627; Springer).
[0021] Non-patent literature 17: Wu, Y., and He, K.; 2018; “Group Normalization”; (In ECCV, pp. 3-19).
[0022] Description of related technologies
[0023] Organ segmentation
[0024] In medical image segmentation, organ segmentation is of great significance in disease diagnosis and surgical planning. For example, the segmented shape of organs (e.g., the hippocampus) can be used as a biomarker for neurodegenerative diseases, including Alzheimer's disease (AD). See Non-Patent Literature 1.
[0025] In radiotherapy planning, accurate segmentation of organs at risk (OARs) can help oncologists design better radiotherapy plans (e.g., appropriate beam paths) that concentrate radiation on the tumor area while minimizing the dose to surrounding healthy organs. See Non-Patent Literature 2.
[0026] Unlike general segmentation problems such as lesion segmentation, organs have relatively stable locations, shapes, and sizes. Current segmentation systems, primarily based on deep learning methods (Roth et al., 2015), often lack awareness of feasible shapes and are affected by the unevenness of training data labeled by physicians, especially in three-dimensional (3D) scenes. For example, see... Figure 5A .
[0027] For organ segmentation, traditional methods include statistical models (non-patent literature 5), atlas-based methods (non-patent literature 6), active contour models (non-patent literature 3), and level sets (non-patent literature 4).
[0028] The segmentation performance of atlas-based methods typically depends on the accuracy of registration and label fusion algorithms. During inference, the snakes and level sets need to be iteratively optimized via gradient descent. In contrast, advances in deep learning-based 2D (non-patent document 7) and 3D (non-patent document 8) segmentation methods have enabled more efficient and accurate organ segmentation.
[0029] Technical problems to be solved
[0030] Although learning-based methods offer faster reasoning speeds and higher accuracy than traditional methods, they often lack an understanding of the anatomical shape of the target organ.
[0031] Regardless of the network architecture and training loss, the segmentation output in related techniques may contain inconsistent regions and may not preserve the anatomical shape of organs.
[0032] Therefore, post-processing is required to correct errors and refine the segmentation results, such as CRF (Non-Patent Document 9) or level set (Non-Patent Document 10), to increase the smoothness of the segmented surface. Summary of the Invention
[0033] According to one aspect of this disclosure, a computer-implemented method for training a neural network for organ segmentation may include: collecting a set of digital images from a database as samples; inputting the collected set of digital images into a neural network recognition model; and training the neural network recognition model to identify a first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image.
[0034] Computer-based methods may include combining segmentation maps to predict signed distance maps (SDM).
[0035] Predictive organ segmentation can have smooth surfaces and can directly remove noisy segments without post-processing.
[0036] The method may also include connecting the segmentation map and the SDM via a differentiable approximate Heaviside function, and predicting the segmentation map and the SDM as a whole.
[0037] Training can include connecting the two outputs of a neural network recognition model via a differentiable approximate Heaviside function and training them jointly.
[0038] The method may further include acquiring a real-world captured image; inputting the captured image as input to a trained neural network recognition model; and outputting segmentation prediction data including at least one segmented organ from the trained neural network recognition model as output, wherein the trained neural network recognition model identifies the target real-world organ.
[0039] The neural network recognition model can be a deep three-dimensional (3D) U-net.
[0040] The computer-implemented method may also include modifying the 3DU-net by performing at least one of the following: (A) using downsampling in the decoder and corresponding upsampling in the decoder, (B) using group normalization instead of batch normalization, and (C) using leaked modified linear unit (ReLU) instead of ReLU as the activation function.
[0041] Modifications may include each of (A) through (C) listed above.
[0042] A graphics processing unit (GPU) can be used to perform processing of neural network recognition models.
[0043] Computer-based methods may also include: predicting the SDM of organ masks using 3D Unet.
[0044] The computer-implemented method may also include: after predicting the SDM of the organ mask in 3D Unet, using the Heaviside function to convert the SDM of the organ mask into a segmentation mask.
[0045] Training can include training a neural network by optimizing the segmentation mask together with SDF.
[0046] The regression loss used for SDM prediction can have two parts. The first part of the loss minimizes the difference between the predicted SDF and the true SDF. The second part of the loss maximizes the Dice similarity coefficient between the predicted mask and the true mask. Segmentation and distance maps can be predicted in the same branch, thus ensuring a correspondence between the segmentation and the SDM branch.
[0047] The first part of the loss can be determined by combining the common loss to be used in the regression task with a product-based regression loss, which is defined based on a formula using the true SDM and the predicted SDM.
[0048] The second part of the loss can be defined as a constant minus the Dice similarity coefficient.
[0049] According to one embodiment, the device may include: at least one memory configured to store computer program code; and at least one processor configured to access the at least one memory and operate according to the computer program code.
[0050] The computer program code may include: collection code configured to cause at least one processor to collect a set of digital sample images from a database; input code configured to cause at least one processor to input the collected set of digital images into a neural network recognition model; and training code configured to cause at least one processor to train the neural network recognition model to identify a first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image, including combining a segmentation map to predict a signed distance map (SDM).
[0051] Collection can include acquiring images captured in the real world.
[0052] The input may include feeding the captured image into a trained neural network recognition model.
[0053] The computer program code may also include output code, which is configured to cause at least one processor to output segmentation prediction data, including at least one segmented organ, from a trained neural network recognition model.
[0054] A trained neural network recognition model can identify the target's real-world organs.
[0055] The neural network recognition model can be a deep three-dimensional (3D) U-net.
[0056] Training may include modifying the 3D U-net by performing at least one of the following: (A) using downsampling in the decoder and corresponding upsampling in the decoder, (B) using group normalization instead of batch normalization, and (C) using leaked modified linear units (ReLU) instead of ReLU as the activation function.
[0057] The output may include: predicting the SDM of the organ mask using 3D Unet, and after predicting the SDM of the organ mask using 3D Unet, converting the SDM of the organ mask into a segmentation mask using the Heaviside function.
[0058] Training can include training a neural network by optimizing the segmentation mask and SDF together.
[0059] The regression loss used for SDM prediction can have two parts. The first part of the loss can minimize the difference between the predicted SDF and the true SDF, and the second part of the loss can maximize the Dice similarity coefficient between the predicted mask and the true mask. Here, the segmentation map and the distance map can be predicted in the same branch, thus ensuring the correspondence between the segmentation and the SDM branch.
[0060] According to an embodiment, a non-transitory computer-readable storage medium may be provided, which stores instructions. The instructions may cause one or more processors to perform the following operations: collect a set of digital sample images from a database; input the collected set of digital images into a neural network recognition model; and train the neural network recognition model to identify a first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image, including combining a segmentation map to predict a signed distance map (SDM).
[0061] According to embodiments of this application, the segmentation scheme is able to predict organ segmentation with smooth surfaces and less noise without any post-processing. Attached Figure Description
[0062] Other features, nature, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, wherein:
[0063] Figure 1 This is a schematic diagram of a network system architecture including an SDM learning model for organ segmentation, according to an implementation method.
[0064] Figure 2A The proposed regression loss for SDM prediction according to an implementation method is shown.
[0065] Figure 2B A graph showing the loss values according to the implementation method is provided.
[0066] Figure 3 This illustrates one aspect of what can be achieved by [the party] according to this disclosure. Figure 7 The flowchart describes the execution of a computer system, including a computer-implemented method for training a neural network for organ segmentation.
[0067] Figure 4A A formula for calculating regression loss according to an implementation method is shown.
[0068] Figure 4B A formula for calculating the Dice loss portion according to an embodiment is shown.
[0069] Figures 5A to 5C The following example hippocampal segmentation comparison is shown: Figure 5A The actual annotations are shown. Figure 5B The segmentation results from the model are shown without predicting the signed distance map; and Figure 5C The segmentation results from the model are shown when predicting a signed distance map.
[0070] Figure 6A to Figure 6E Examples of output image (organ) segmentation using GT, Dice, SDM, L1 SDM+Dice, and embodiments of this disclosure (“Ours”) are shown respectively.
[0071] Figure 7 This is a schematic diagram of a computer system according to an implementation method. Detailed Implementation
[0072] This disclosure relates to medical imaging techniques that use AI neural networks to perform organ segmentation for use in medical imaging such as computed tomography (CT) scans (which use X-ray beams aimed at parts of a patient, such as organs, to generate digital X-ray images). The generated digital X-ray images can be cross-sectional images of the body (or body organs), which may be referred to as slices.
[0073] For surgical procedures (e.g., organ transplantation), shape-aware neural networks (which combine shape knowledge of one or more organs via statistical shape models used in segmentation) can be used to perform organ segmentation.
[0074] The techniques used for organ segmentation can be implemented by one or more processors that can execute computer software with computer-readable instructions (code), which can be physically stored on one or more computer-readable media (e.g., hard disk drives). For example, as discussed in detail below... Figure 7 A computer system 700 suitable for implementing a particular embodiment of the disclosed subject matter is shown.
[0075] In traditional medical image segmentation methods, smoothness issues can be mitigated by adding physically meaningful regularization terms, as is the case in snakes (non-patent document 3) and level sets (non-patent document 4).
[0076] In order to utilize shape perception using traditional methods, the inventors proposed, according to the implementation method, to directly regress a signed distance function (SDF) from the input image using a 3D convolutional neural network.
[0077] Signed distance graph
[0078] Several works have explored the applications of signed distance maps (SDMs) or signed distance functions (SDFs) in computer vision and graphics. For example, see Non-Patent Document 11, which uses a truncated SDF to better reconstruct volumetric surfaces on RGB-D images. Non-Patent Document 12 treats a linearly shifted saliency map as an SDF and utilizes level set smoothing terms to refine the predicted saliency map across multiple training phases.
[0079] Non-patent document 13 learns continuous 3D SDF directly from point sampling through a network containing a series of fully connected layers and L1 regression loss.
[0080] The learned SDF can be used to obtain shape representations and completion results at the current state-of-the-art level. Because medical images contain richer contextual information than point sampling, more complex network architectures and training strategies need to be considered when applying SDM learning to organ segmentation tasks.
[0081] Non-patent document 14 proposes using a distance map (unsigned) as an intermediate step in a 2D organ shape prediction task. The conversion from the distance map to the shape parameter vector is performed by PCA and does not involve the segmentation map.
[0082] However, for 3D organ segmentation with a much higher dimension than the 2D case, directly applying the method of non-patent document 14 may not be effective in small organs.
[0083] Recently, non-patent literature 15 and non-patent literature 16 have used distance graph prediction as a regularizer during the training of organ segmentation.
[0084] Because non-patent literature 15 and non-patent literature 16 predict segmentation and distance maps in different branches, the correspondence between segmentation and SDM branches is not guaranteed.
[0085] In view of the problems of conventional techniques, a new deep learning scheme for segmentation and a new loss for learning organ segmentation are provided according to the implementation method. According to the implementation method, the segmentation scheme is able to predict organ segmentation with smooth surfaces and less noise without any post-processing.
[0086] like Figure 1 As shown, according to the implementation, SDM (predicted via SDF) can be used in conjunction with the segmentation map for prediction, rather than as a regularizer in the organ segmentation task.
[0087] According to the implementation method, the two outputs can be concatenated and jointly trained using a differentiable Heaviside function. According to the implementation method, a novel regression loss can be utilized, which, compared to L1 regression loss in ablation studies, leads to a larger gradient magnitude and exhibits better performance for inaccurate predictions.
[0088] Therefore, the method according to the embodiment can differ from the methods in Non-Patent Document 14 and Non-Patent Document 15. For example, according to the embodiment, the segmentation map and SDM can be connected by a differentiable Heaviside function and can be predicted as a whole.
[0089] Figure 1 A network system architecture including an SDM learning model for organ segmentation according to an embodiment is shown.
[0090] like Figure 1 As shown, according to an embodiment, an image (e.g., a 3D medical image) can be used as input to a deep 3D Unet (or U-net) neural network and can output a segmentation prediction that may include detected objects (e.g., organs).
[0091] according to Figure 1 The implementation shown allows the proposed backbone depth 3D UNet to be trained during training using a differentiable approximation of the Heaviside function, with SDM loss and segmentation loss.
[0092] Depending on the implementation method, the 3D Unet (or U-net) can be modified. For example, such as... Figure 1 As shown, the modifications may include one or more of the following: (1) using 6 downsamplings in the decoder and 6 corresponding upsamplings in the decoder; (2) using group normalization (e.g., similar to the group normalization in non-patent document 17) instead of batch normalization, since the batch size can be limited to 1 due to the limited size of the GPU memory, according to the implementation; and (3) using leaky rectified linear units (ReLU) instead of using ReLU as the activation function.
[0093] According to one implementation, 3D UNet can predict the SDM of an organ mask. According to another implementation, 3D UNet can be a model executed by a dedicated processor (e.g., a GPU) that may have limited memory.
[0094] According to the implementation, after the SDM of the organ mask is predicted by the 3D unit, the SDM can be converted into a segmentation mask using a Heviside function (e.g., similar to non-patent document 4).
[0095] Figure 3 It is possible to obtain, based on one aspect of this disclosure, from Figure 7 The flowchart of computer system execution, such as Figure 3 As shown, a computer-implemented method for training a neural network for organ segmentation may include: collecting a set of digital sample images from a database (step 301); inputting the collected set of digital images into a neural network recognition model (step 302); and training the neural network recognition model (step 303).
[0096] According to an implementation, step 303 may include training a neural network recognition model to identify the first object as a specific object based on the similarity between the first object in the first digital image and the second object in the second digital image.
[0097] Computer-based methods may include combining segmentation maps to predict signed distance maps (SDM).
[0098] Predictive organ segmentation can have smooth surfaces and can directly remove noisy segments without post-processing.
[0099] The method may also include connecting the segmentation map and the SDM via a differentiable approximate Heaviside function, and predicting the segmentation map and SDM as a whole.
[0100] Training can include connecting the two outputs of a neural network recognition model via a differentiable approximate Heaviside function and training them jointly.
[0101] The method may further include acquiring a real-world captured image; inputting the captured image as input to a trained neural network recognition model; and outputting segmentation prediction data including at least one segmented organ from the trained neural network recognition model as output, wherein the trained neural network recognition model identifies the target real-world organ.
[0102] The neural network recognition model can be a deep three-dimensional (3D) U-net.
[0103] The computer-implemented method may also include modifying the 3DU-net by performing at least one of the following: (A) using downsampling in the decoder and corresponding upsampling in the decoder, (B) using group normalization instead of batch normalization, and (C) using leaked modified linear unit (ReLU) instead of ReLU as the activation function.
[0104] Modifications may include each of (A) through (C) listed above.
[0105] A graphics processing unit (GPU) can be used to perform processing of neural network recognition models.
[0106] The computer-implemented method may also include predicting the SDM of the organ mask using 3D Unet.
[0107] The computer-implemented method may also include: after predicting the SDM of the organ mask using 3D Unet, converting the SDM of the organ mask into a segmentation mask using a Heaviside function.
[0108] Training can include training a neural network by optimizing the segmentation mask and SDF together.
[0109] The regression loss for SDM prediction can have two parts. The first part of the loss minimizes the difference between the predicted SDF and the true SDF. The second part of the loss maximizes the Dice similarity coefficient between the predicted mask and the true mask. Segmentation and distance maps can be predicted in the same branch, thus ensuring the correspondence between the segmentation and SDM branches.
[0110] The first part of the loss can be determined by combining the common loss to be used in the regression task with a product-based regression loss, which is defined based on a formula using the true SDM and the predicted SDM.
[0111] The second part of the loss can be defined as a constant minus the Dice similarity coefficient.
[0112] According to an embodiment, the device may include: at least one memory configured to store computer program code; and at least one processor configured to access the at least one memory and operate according to the computer program code.
[0113] The computer program code may include: collection code configured to cause at least one processor to collect a set of digital sample images from a database; input code configured to cause at least one processor to input the collected set of digital images into a neural network recognition model; and training code configured to cause at least one processor to train the neural network recognition model to identify a first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image, including combining a segmentation map to predict a signed distance map (SDM).
[0114] Collection can include obtaining real-world captured images.
[0115] The input may include the captured image being fed into the trained neural network recognition model.
[0116] The computer program code may also include output code, which is configured to cause at least one processor to output segmentation prediction data, including at least one segmented organ, from a trained neural network recognition model.
[0117] The trained neural network recognition model can identify the target real-world organ.
[0118] The neural network recognition model can be a deep three-dimensional (3D) U-net.
[0119] Training may include modifying the 3D U-net by performing at least one of the following: (A) using downsampling in the decoder and corresponding upsampling in the decoder, (B) using group normalization instead of batch normalization, and (C) using leaked modified linear units (ReLU) instead of ReLU as the activation function.
[0120] The output may include: predicting the SDM of the organ mask using 3D Unet, and after predicting the SDM of the organ mask using 3D Unet, converting the SDM of the organ mask into a segmentation mask using the Heaviside function.
[0121] Training can include training a neural network by optimizing the segmentation mask and SDF together.
[0122] The regression loss used for SDM prediction can have two parts: the first part minimizes the difference between the predicted SDF and the true SDF, and the second part maximizes the Dice similarity coefficient between the predicted mask and the true mask, where the segmentation map and the distance map are predicted in the same branch, thus ensuring the correspondence between the segmentation and the SDM branch.
[0123] According to an implementation, a non-transitory computer-readable storage medium may be provided to store instructions. The instructions may cause one or more processors to perform the following operations: collect a set of digital sample images from a database; input the collected digital images into a neural network recognition model; and train the neural network recognition model to identify the first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image, including combining a segmentation map to predict a signed distance map (SDM).
[0124] According to the implementation method, the neural network can be trained by optimizing the segmentation mask and SDF together.
[0125] According to the implementation, the loss can have two parts. According to the implementation, the first part of the loss can minimize the difference between the predicted SDF and the true SDF, while the second part can maximize the Dice (coefficient) between the predicted mask and the true mask.
[0126] Figure 2A The proposed regression loss for SDM prediction according to an embodiment is shown. According to the embodiment, all SDM values can be normalized.
[0127] Figure 2B A graph showing the loss value for a given true SDM value of 0:5 according to the implementation method is presented. Figure 2B In this context, line L1' can represent a combination of loss and L1 loss proposed according to an implementation of this disclosure.
[0128] According to embodiments of this disclosure, the SDM loss portion can be formulated as a regression problem. According to embodiments, the L1 loss is a common loss used in regression tasks. However, for multi-organ segmentation tasks, training using L1 loss sometimes leads to an unstable training process.
[0129] To overcome the drawbacks of L1 loss, according to one implementation, L1' can be determined by combining the L1 loss with a proposed product-based regression loss, where the product is defined by a formula. For example, according to one implementation, it can be based on... Figure 4A The formula is used to calculate the regression loss, where y t Represents a real SDM, p t This represents the predicted SDM.
[0130] According to the implementation method, the product of the prediction and the true value can be used to penalize the output SDM with incorrect symbols.
[0131] According to the implementation method, for the Dice loss component, the loss can be defined as a constant minus the Dice similarity coefficient. For example, it can be based on... Figure 4B The formula is used to calculate the Dice loss, where N is the number of classes and t represents the t-th organ class. t and p t These represent the true labels and the model predictions, respectively (ε can be a term with small values to avoid numerical problems).
[0132] While current organ segmentation systems are primarily based on deep learning methods (Roth et al., 2015), they often lack awareness of feasible shapes and are affected by the unsmoothness of the training ground truth labeled by physicians, especially in three-dimensional (3D) scenes. For example, the ground truth labeling of the hippocampus may not maintain a consistent and continuous shape because it is annotated in two-dimensional (2D) slices via contours rather than 3D surfaces. See, for example... Figure 5A .
[0133] Figures 5A to 5C The following example hippocampal segmentation comparison is shown: Figure 5A True annotations, due to inconsistencies in 2D annotations, lack smoothness in 3D views; Figure 5B Segmentation results from the model without predicting the signed distance map; and ( Figure 5C When predicting signed distance maps, the segmentation results from the model are significantly better than those from other models. Figure 5A and Figure 5B Smoother while maintaining the overall shape.
[0134] Figure 1 An exemplary process for implementing this disclosure is shown.
[0135] According to one implementation, the neural network can receive an image (e.g., a 3D medical image) as input. According to one implementation, the neural network can output an SDF prediction. According to one implementation, such as... Figure 1 As shown, a neural network may include one or more skip connections (e.g., one or more additional connections between nodes in different layers of the neural network that skip one or more non-linear processing layers).
[0136] According to one implementation, the loss may have two parts. According to another implementation, the two parts of the loss may include a first part that minimizes the difference between the predicted SDF and the true SDF, and a second part that maximizes the difference between the predicted mask and the true mask.
[0137] Figures 8A and 8B illustrate the losses according to the implementation method.
[0138] According to an implementation, the SDM loss can be formulated as a regression problem. According to an implementation, the L1 loss can be a common loss used in regression tasks. However, training with L1 loss sometimes leads to an unstable training process (e.g., when training a multi-organ segmentation task). To overcome the drawbacks of L1 loss, according to an implementation, the L1 loss is combined with a regression loss L'. According to an implementation, the regression loss L' can be based on a product... Figure 4A The formula in the text.
[0139] According to the implementation method, the product of the prediction and the true value can be used to penalize the output SDM with incorrect symbols.
[0140] Figure 6A to Figure 6E Examples of output image (organ) segmentation using GT, Dice, SDM, L1 SDM+Dice, and embodiments of this disclosure (“Ours”) are shown. Specifically, Figure 6A shows GT, Figure 6B shows Dice, Figure 6C shows SDM, and Figure 6D shows L1SDM+Dice. Figure 6E Implementations of this disclosure are shown (“Ours”).
[0141] like Figure 7 As shown, computer software can be coded using any appropriate machine code or computer language. Machine code or computer language can be processed by mechanisms such as assembly, compilation, and linking to create code that includes instructions. Instructions can be executed directly by the computer's central processing unit (CPU), graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0142] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0143] Figure 7 The components shown for computer system 700 are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any component or combination thereof shown in the exemplary embodiments of computer system 700.
[0144] Computer system 700 may include specific human-machine interface input devices. For example, the human-machine interface input device may respond to input from one or more human users via, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, clapping), visual input (such as gestures), olfactory input, etc. The human-machine interface device may also be used to capture specific media that are not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sounds), images (such as CT images, scan images, photographic images obtained from still image capture devices), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0145] The input human-machine interface device may include one or more of the following (only one is depicted): keyboard 701, mouse 702, trackpad 703, touchscreen 710, data glove 704, joystick 705, microphone 706, scanner 707, camera device 708, etc. According to one embodiment, the camera device 708 may be a CT scanner. According to another embodiment, the camera device 708 may be a medical imaging device.
[0146] The computer system 700 may also include specific human-machine interface (HMI) output devices. Such HMI output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. These HMI output devices may include tactile output devices (e.g., tactile feedback via a touchscreen 710, data glove 704, or joystick 705, but may also be tactile feedback devices not used as input devices), audio output devices (e.g., speakers 709, headphones (not depicted), visual output devices (e.g., screens 710, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability, some of which are capable of outputting two-dimensional visual output or outputting more than three-dimensional output in a manner such as stereoscopic output; virtual reality glasses, holographic displays, and smoke generators), and printers.
[0147] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVDROM / RW 720 having media 721 such as CD / DVD, thumb drives 722, removable hard disk drives or solid-state drives 723, conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0148] Those skilled in the art should also understand that the term "computer-readable medium" or "computer-readable medium" as used in connection with the subject matter disclosed herein corresponds to a non-transitory computer-readable medium and does not include transmission media, carrier waves, or other transient signals.
[0149] Computer system 700 may also include interfaces to one or more communication networks. These networks may be wireless, wired, or optical. They may also be local, wide area, metropolitan area, vehicle-mounted, industrial, real-time, latency-tolerant, etc. Examples of networks include local area networks (LANs) such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANbus, etc. Specific networks typically require external network interface adapters that connect to certain general-purpose data ports or peripheral buses (749) (e.g., the USB port of computer system 700); other systems are typically integrated into the core of computer system 700 by connecting to system buses such as Ethernet interfaces to PC computer systems or cellular network interfaces to smartphone computer systems. Using any of these networks, computer system 700 can communicate with other entities. This communication can be unidirectional (receive-only, e.g., broadcasting TV), unidirectional (transmit-only, e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Specific protocols and protocol stacks can be used on each of the networks and network interfaces mentioned above.
[0150] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be attached to the core 740 of the computer system 700.
[0151] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 743, hardware accelerators 744 for certain tasks, etc. These devices, along with internal mass storage such as read-only memory (ROM) 745, random access memory 746, and non-user-accessible hard disk drives (HDDs), SSDs, etc. 747, can be connected via a system bus 748. In some computer systems, the system bus 748 may be accessed as one or more physical connectors to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus 748 or connected via a peripheral bus 749. Peripheral bus architectures include PCI, USB, etc.
[0152] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can execute specific instructions, and combinations of these instructions can constitute the aforementioned computer code. This computer code can be stored in ROM 745 or RAM 746. Transient data can also be stored in RAM 746, while permanent data can be stored, for example, in internal mass storage 747. Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more CPUs 741, GPUs 742, mass storage devices 747, ROM 745, RAM 746, etc.
[0153] According to the implementation method, the CPU can use one or more of a GPU, FPGA, or accelerator to perform neural network processing.
[0154] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be of types known and available to those skilled in the art of computer software.
[0155] By way of example and not limitation, a computer system having architecture 700, particularly core 740, can provide functionality resulting from the execution of software contained in one or more tangible computer-readable media by a processor (including CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with: user-accessible mass storage as described above; and certain non-transitory memory of core 740, such as internal mass storage 747 or ROM 745. Software implementing various embodiments of this disclosure can be stored in such a device and executed by core 740. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause core 740, particularly the processor therein (including CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to software-defined processes. In addition, or as an alternative, the computer system may provide functionality resulting from a structure implemented in circuitry by hardwired logic or otherwise (e.g., accelerator 744), which may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. References to software, where appropriate, may include logic, and vice versa. References to computer-readable media, where appropriate, may include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry containing logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0156] advantage
[0157] 1) No post-processing is required because the direct output of the network remains smooth and without small flickers.
[0158] 2) Any existing 3D segmentation network can be easily adaptively adapted to incorporate the SDM prediction model with almost no additional overhead.
[0159] Although several exemplary embodiments have been described in this disclosure, there are variations, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.
Claims
1. A computer-implemented method for training a neural network for organ segmentation, characterized in that, The computer-implemented method includes: Collect a set of digital images from the database as samples; The collected set of digital images is input into a neural network recognition model; and The neural network recognition model is trained to identify the first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image. The computer-implemented method includes combining a segmentation map to predict a signed distance map (SDM). The neural network recognition model is a deep 3D U-net; The SDM of the organ mask is predicted by the 3D Unet; After the 3D Unet predicts the SDM of the organ mask, the Heaviside function is used to convert the SDM of the organ mask into a segmentation mask. The training includes training the neural network by optimizing the segmentation mask and the signed distance function SDF together; The regression loss of the SDM prediction has two parts: the first part of the loss minimizes the difference between the predicted SDF and the true SDF, while the second part of the loss maximizes the Dice similarity coefficient between the predicted mask and the true mask.
2. The computer-implemented method according to claim 1, characterized in that, Also includes: Predict organ segmentation with smooth surfaces and directly remove noisy segments without post-processing.
3. The computer-implemented method according to claim 1, characterized in that, It also includes connecting the segmentation map and the SDM via a differentiable approximate Heaviside function, and predicting the segmentation map and the SDM as a whole, wherein the training includes connecting the two outputs of the neural network recognition model via the differentiable approximate Heaviside function and training them jointly.
4. The computer-implemented method according to claim 1, characterized in that, Also includes: Obtain images captured from the real world; The captured images are fed into a trained neural network recognition model. as well as The output of the trained neural network recognition model includes segmentation prediction data of at least one segmented organ, wherein the trained neural network recognition model identifies the target real-world organ.
5. The computer-implemented method of claim 1, further comprising modifying the 3D U-net by performing at least one of the following: (A) using downsampling in the decoder and using corresponding upsampling in the decoder, (B) using group normalization instead of batch normalization, and (C) using leaked modified linear unit ReLU, wherein the ReLU is not used as an activation function.
6. The computer-implemented method according to claim 1, characterized in that, The graphics processing unit (GPU) is used to perform the processing of the neural network recognition model.
7. The computer-implemented method according to claim 1, characterized in that, in, The segmentation map and the distance map are predicted in the same branch, thereby ensuring the correspondence between the segmentation and the SDM branch.
8. The computer-implemented method according to claim 7, characterized in that, The first part of the loss is determined by combining the common loss used in the regression task with a product-based regression loss, which is defined based on a formula using the true SDM and the predicted SDM.
9. The computer-implemented method according to claim 7, characterized in that, The second part of the loss is defined as a constant minus the Dice similarity coefficient.
10. A device, characterized in that, include: At least one memory is configured to store computer program code; as well as At least one processor is configured to access the at least one memory and execute the method as described in any one of claims 1-9 according to the computer program code.
11. A non-transitory computer-readable storage medium for storing instructions, characterized in that, The instructions cause one or more processors to perform the method as described in any one of claims 1-9.
12. A computer-implemented apparatus for training a neural network for organ segmentation, characterized in that, The device includes: The collection module collects a set of digital images from the database as samples; The input module inputs the collected set of digital images into the neural network recognition model; and The training module trains the neural network recognition model to identify the first object as a specific object based on the similarity between a first object in a first digital image and a second object in a second digital image. The device includes a prediction module that combines the segmentation map to predict the signed distance map (SDM). The neural network recognition model is a deep 3D U-net; The SDM of the organ mask is predicted by the 3D Unet; After the 3D Unet predicts the SDM of the organ mask, the Heaviside function is used to convert the SDM of the organ mask into a segmentation mask. The training includes training the neural network by optimizing the segmentation mask and the signed distance function SDF together; The regression loss of the SDM prediction has two parts: the first part of the loss minimizes the difference between the predicted SDF and the true SDF, while the second part of the loss maximizes the Dice similarity coefficient between the predicted mask and the true mask.
13. The computer-implemented apparatus according to claim 12, characterized in that, Also includes: Predict organ segmentation with smooth surfaces and directly remove noisy segments without post-processing.
14. The computer-implemented apparatus according to claim 12, characterized in that, It also includes connecting the segmentation map and the SDM via a differentiable approximate Heaviside function, and predicting the segmentation map and the SDM as a whole, wherein the training includes connecting the two outputs of the neural network recognition model via the differentiable approximate Heaviside function and training them jointly.
15. The computer-implemented apparatus according to claim 12, characterized in that, Also includes: Obtain images captured from the real world; The captured images are fed into a trained neural network recognition model. as well as The output of the trained neural network recognition model includes segmentation prediction data of at least one segmented organ, wherein the trained neural network recognition model identifies the target real-world organ.
16. The computer-implemented apparatus of claim 12, further comprising modifying the 3D U-net by performing at least one of the following: (A) using downsampling in the decoder and using corresponding upsampling in the decoder, (B) using group normalization instead of batch normalization, and (C) using leaked modified linear unit ReLU, wherein the ReLU is not used as an activation function.
17. The computer-implemented apparatus according to claim 12, characterized in that, The graphics processing unit (GPU) is used to perform the processing of the neural network recognition model.
18. The computer-implemented apparatus according to claim 12, characterized in that, in, The segmentation map and the distance map are predicted in the same branch, thereby ensuring the correspondence between the segmentation and the SDM branch.
19. The computer-implemented apparatus according to claim 18, characterized in that, The first part of the loss is determined by combining the common loss used in the regression task with a product-based regression loss, which is defined based on a formula using the true SDM and the predicted SDM.
20. The computer-implemented apparatus according to claim 18, characterized in that, The second part of the loss is defined as a constant minus the Dice similarity coefficient.