Automatic generation of 3D anatomical models
The method addresses the complexity and time constraints of medical image segmentation by employing neural networks and fuzzy definitions for rapid, accurate 3D anatomical modeling, enhancing surgical planning with detailed nerve representation.
Patent Information
- Application Number
- JP2025535268
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-12-20
- Publication Date
- 2026-01-16
AI Technical Summary
Existing medical image segmentation methods are complex, time-consuming, and require specialized personnel, limiting the adoption of 3D technology in surgical planning, particularly for complex procedures like deformity correction and cancer resection, and fail to represent the peripheral nervous system, which is crucial for patient mobility and physiological functionality.
A method using multiple segmentation neural networks and fuzzy anatomical definitions to automatically generate 3D anatomical models, including bones, organs, and blood vessels, with specific focus on nerve detection through spatial relationship analysis, leveraging supervised machine learning and AI algorithms for rapid and accurate segmentation of pediatric medical images.
Enables rapid and accurate 3D anatomical modeling with high accuracy (>90%) and reduced processing time, facilitating preoperative planning and surgical preparation by consistently representing complex anatomical structures like the sacral plexus, especially in pediatric patients, and adapting to pathological variations.
Smart Images

Figure 2026501533000001_ABST
Abstract
Description
[Technical Field]
[0001] Various exemplary embodiments relate generally to methods / devices for automatic segmentation of medical images and for representing segmented images together with the peripheral nervous system. [Background technology]
[0002] Imaging data is nowadays essential for the construction of efficient and, if possible, non-invasive surgical strategies. This preparation is routinely performed using 2D images (slices), often in a less interactive environment, which limits the visualization of the anatomical structures in the volume. Segmentation of volumetric images (e.g., MRI, CT, computed tomography) allows the creation of a unique 3D model of the patient, providing the surgeon with a realistic view and allowing for better prediction of problems.
[0003] However, anatomical structure segmentation is a complex processing procedure requiring long production times, making it unsuitable for clinical practice. This procedure must be performed by individuals with backgrounds in radiology and computer science, significantly limiting the adoption of 3D technology for surgical patient care, particularly in complex procedures such as deformity correction or cancer resection. Therefore, automating medical image segmentation is essential for the rapid deployment of 3D solutions.
[0004] Existing solutions fail to represent the peripheral nervous system, which is important information for maintaining a patient's mobility and physiological functionality.
[0005] There is a need for efficient and automatic segmentation of medical images that is suitable for pediatric medical images with pathological cases. Summary of the Invention
[0006] The scope of protection is indicated by the independent claims. The embodiments, examples and features described herein that do not fall within the scope of protection, if any, should be interpreted as useful examples for understanding the various embodiments or examples described herein.
[0007] According to a first aspect, a method for generating a 3D anatomical model of an anatomical region of interest is disclosed, the method comprising: - obtaining an input image of said anatomical region of interest and a 3D representation of fibers within said anatomical 3D region of interest; generating an aggregated segmentation image representing an anatomical structure, the aggregated segmentation image comprising at least one bone, at least one organ, or at least one blood vessel; - detecting at least one fiber representing a nerve in a 3D representation of said fibers within an anatomical region of interest; generating the 3D anatomical model using the aggregated segmentation image and at least one fiber representing the detected nerve; generating an aggregated segmentation image; - performing a segmentation of the input image to generate a first segmented image by detecting voxels representing bones using a first segmentation neural network applied to the input image, the first segmentation neural network being a bone-specific neural network trained to detect one or more bones in the image; - performing a segmentation of the input image to generate a second segmented image by detecting voxels representing organs using a second segmentation neural network applied to the input image, the second segmentation neural network being an organ-specific neural network trained to detect one or more organs in the image; - performing a segmentation of the input image to generate a third segmented image by detecting voxels representing blood vessels using a third segmentation neural network applied to the input image, the third segmentation neural network being a blood vessel-specific neural network trained to detect one or more blood vessels in the image; - performing a segmentation of the input image to generate a fourth segmented image by detecting voxels representing bones, organs, or blood vessels using a fourth segmentation neural network applied to the input image, the fourth segmentation neural network being a multi-layered neural network trained to detect bones, blood vessels, and at least one organ in the image; combining the first, second, third, and fourth segmentation images; detecting the at least one fiber representing the nerve; obtaining a nerve definition expressed in the form of one or more spatial relationships between said nerve and one or more anatomical structures segmented in the aggregated segmentation image; - detecting that the at least one fiber satisfies a nerve definition based on at least one fuzzy 3D anatomical area corresponding to one or more spatial relationships expressed in the nerve definition.
[0008] Detecting at least one fiber satisfying a nerve definition may include searching for fibers within a 3D search space defined by the at least one fuzzy 3D anatomical area.
[0009] This method is extracting from the aggregated segmentation image each anatomical structure used in said nerve definition; obtaining a segmentation mask from each extracted anatomical structure; - Obtaining a fuzzy anatomical 3D area representing a given spatial relationship of a nerve definition by expanding a segmentation mask based on a structuring element, the structuring element encoding at least one of directional information, path information, and connectivity information of the given spatial relationship.
[0010] An expansion score may be associated with each voxel within the at least one fuzzy anatomical 3D area.
[0011] This method is - calculating a fiber score for each of at least one fiber that falls within at least one fuzzy anatomical 3D area based on the dilation scores of the voxels that represent the considered fiber; - detecting that a nerve is represented by at least one fiber when the fiber is completely or partially within at least one fuzzy anatomical area and has a fiber score higher than a threshold.
[0012] A nerve definition may include a sequence of nerve sub-definitions, each of which defines a portion of a nerve. and for a first portion of a nerve having an associated first sub-definition expressing at least one first spatial relationship between the first portion of the nerve and at least one first anatomical structure segmented in the aggregated segmentation image, the method comprising: (a) obtaining a segmentation mask of at least one first anatomical structure; (b) obtaining at least one first local fuzzy anatomical 3D area representing at least one first spatial relationship by expanding the segmentation mask based on a structuring element, the structuring element encoding at least one of direction information, path information, and connectivity information of the given spatial relationship; (c) calculating, for each of the plurality of fibers, a fiber score based on the expansion scores of voxels representing the considered fiber that fall within a first fuzzy anatomical 3D area defined by the at least one first local fuzzy 3D anatomical area; (d) for at least one of the plurality of fibers, detecting that a first portion of the nerve is represented by the considered fiber when the considered fiber is completely or partially within the first fuzzy anatomical 3D area and has a fiber score higher than a threshold.
[0013] The method may include repeating steps (a)-(d) for a second portion of the nerve and using at least one fiber detected for the first portion to narrow down the plurality of fibers for which a fiber score is calculated.
[0014] The segmentation mask of the at least one anatomical structure may be a fuzzy segmentation mask of the at least one anatomical structure, the method comprising: - obtaining an initial segmentation mask by extraction of at least one anatomical structure from the aggregated segmentation image; - obtaining an expanded segmentation mask by expanding said initial segmentation mask based on a structuring element; - extracting voxels from the dilated segmentation mask that correspond to the initial segmentation mask and normalizing said extracted voxels; - defining said fuzzy segmentation mask as said extracted voxels normalized.
[0015] If corresponding voxels in the first, second, and third segmentation images are detected as representing distinct anatomical structures, then each voxel in the aggregated segmentation image may be a corresponding voxel in the fourth segmentation image.
[0016] One or more or each of the first, second, third, and fourth segmentation neural networks may be trained using a loss function defined as a weighted sum of specific loss functions, where the weighting coefficients of the weighted sum may be specific to the segmentation neural network under consideration, and the weighting coefficients may vary dynamically during training of the segmentation neural network.
[0017] The first particular loss function can measure accuracy based on global and local information over a 3D volume in the reference segmentation image and the current segmentation image.
[0018] A second specific loss function may measure the local consistency between voxels of a reference segmentation image and voxels of a current segmentation image.
[0019] The third specific loss function is the difference between the boundary points of the reference segmentation image and the current segmentation image. The distance between the boundary points of the projection image can be measured.
[0020] One or more, or each, of the first, second, third, and fourth segmentation neural networks may be trained using image patches extracted from the input image, the centers of the image patches having coordinates determined using a sampling algorithm adapted to the type of anatomical structure to be detected by the segmentation neural network being considered.
[0021] The method may include determining, for an input image, a bounding box of an object of interest using a deep regression network, and applying one or more or each of a first, second, third, and fourth segmentation neural network to only the bounding box of the object of interest.
[0022] The method may include displaying a segmentation image generated by one of the segmentation neural networks, where the associated segmentation neural network includes one or more additional convolutional layers whose outputs are each merged with the input of a corresponding encoder block of the associated segmentation neural network; generating a segmentation mask of the anatomical structure in the segmentation image; allowing a user to select one or more voxels in the displayed segmentation image that are encoded as false positives or false negatives; and providing the segmentation mask and the selected voxels as input images to the one or more additional convolutional layers.
[0023] Generating the aggregated segmentation image may further include refining at least one of the first, second, third, and fourth segmentation images by using an iterative or automatic refinement process before combining the first, second, third, and fourth segmentation images.
[0024] At least one or each of the first, second, third, and fourth segmentation neural networks may be a 3D convolutional neural network (CNN).
[0025] One or more or each of the first, second, third, and fourth 3D CNNs is based on a U-Net architecture.
[0026] The method may include obtaining a 3D representation of fibers by applying a tractography algorithm to a diffusion-weighted MR image within a convex hull corresponding to an anatomical region of interest that envelops the anatomical structure being segmented in the aggregated segmentation image, and by using random seed points for the tractography algorithm on the convex hull.
[0027] According to another aspect, an apparatus comprises at least one processor and at least one memory containing computer program code configured to cause the apparatus, using the at least one processor, to perform one or more or all steps of the method according to the first aspect.
[0028] According to one or more examples, at least one memory and computer program code may be configured to cause an apparatus, using at least one processor, to perform the steps of a method according to the first aspect.
[0029] In general, the at least one memory and computer program code are configured to cause the device, using the at least one processor, to perform one or more or all steps of the methods for generating segmentation images representing anatomical structures disclosed herein.
[0030] One or more exemplary embodiments provide an apparatus comprising means for performing one or more or all of the steps of the methods for generating segmentation images representing anatomical structures disclosed herein.
[0031] Generally, the device comprises means for performing one or more or all of the steps of the methods for generating segmentation images representing anatomical structures disclosed herein. These means may include circuitry configured to perform one or more or all of the steps of the methods for generating segmentation images representing anatomical structures disclosed herein. These means may include at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured to cause the device, using the at least one processor, to perform one or more or all of the steps of the methods for generating segmentation images representing anatomical structures disclosed herein.
[0032] These means may comprise a Central Processing Unit (CPU) and a Graphical Processing Unit (GPU), wherein generating the segmentation image is performed in parallel by at least one processor of the CPU and at least one processor of the GPU, and detecting at least one fiber that satisfies the definition of a nerve is performed by at least one processor of the CPU.
[0033] At least one exemplary embodiment provides a non-transitory computer-readable medium storing computer-executable instructions that, when executed by at least one processor in the apparatus, cause the apparatus to perform a method according to the first aspect.
[0034] Generally, the computer-executable instructions cause the device to perform one or more or all of the steps of the methods for generating segmentation images representative of anatomical structures disclosed herein. [Brief explanation of the drawings]
[0035] Example embodiments will be more fully understood from the detailed description given herein below and the accompanying drawings, which are given by way of illustration only and therefore do not limit the present disclosure.
[0036] [Figure 1] 1 shows a block diagram of a training system for image segmentation, according to one embodiment. [Figure 2] 1 shows a schematic flowchart of a method for training a supervised machine learning algorithm, according to one embodiment. [Figure 3] 1 illustrates a statistical distribution of a training data set, according to one embodiment. [Figure 4] FIG. 1 illustrates a block diagram of a bounding box detection subsystem, according to one embodiment. [Figure 5]FIG. 1 shows a block diagram of a convolutional neural network for image segmentation, according to one embodiment. [Figure 6] FIG. 1 illustrates a block diagram of a subsystem for ensemble learning, according to one embodiment. [Figure 7] FIG. 1 illustrates a block diagram of a trained image segmentation system, according to one embodiment. [Figure 8] 1 shows a schematic flow chart of a method for generating a segmentation image according to an embodiment; [Figure 9] 1 shows a flowchart of a method for generating a 3D anatomical model. [Figure 10] 1 shows a schematic flow chart of a method for nerve detection. [Figure 11A] 1 illustrates a fuzzy geometric area for spatial definition of nerves, according to one embodiment. [Figure 11B] 1 illustrates fuzzy anatomical areas for spatial definition of nerves, according to one embodiment. [Figure 11C] 1 illustrates, by way of example, the spatial definition of a nerve. [Figure 11D] Illustrate an example of the spatial definition of a nerve. [Figure 11E] Illustrate an example of the spatial definition of a nerve. [Figure 12] 1 shows a diagram illustrating aspects of a method for neurocognition, according to one embodiment.
[0037] It should be noted that these figures illustrate general features of methods, structures, and / or materials utilized in certain exemplary embodiments and are intended to supplement the written description provided below. However, these figures are not to scale, may not accurately reflect the precise structural or performance characteristics of any given embodiment, and should not be construed as defining or limiting the range of values or properties encompassed by the exemplary embodiments. The use of similar or identical reference numbers in various figures is intended to indicate the presence of similar or identical elements or features. DETAILED DESCRIPTION OF THE INVENTION
[0038] Various exemplary embodiments will now be described more fully with reference to the accompanying drawings, in which several exemplary embodiments are shown.
[0039] Detailed exemplary embodiments are disclosed herein. However, the specific structural and functional details disclosed herein are merely representative for purposes of describing the exemplary embodiments. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as being limited to only the embodiments described herein. Thus, while the exemplary embodiments are susceptible to various modifications and alternative forms, embodiments have been shown by way of example in the drawings and will be described in detail herein. However, it should be understood that there is no intention to limit the exemplary embodiments to the particular forms disclosed.
[0040] Next, a method and corresponding system for automatic segmentation of medical images will be described in detail in the context of its application to pediatric medical images of the pelvis (e.g., MRI, magnetic resonance imaging, and / or CT, computed tomography images). The method allows for identifying major anatomical structures (including organs such as the coaxial bones, sacrum, internal and external iliac vessels, bladder, colon, digestive tract, pelvic muscles, as well as specific regions such as the sacral foramen and sacral canal) and representing peripheral nerve networks (including the sacral plexus, pudendal plexus, and hypogastric plexus).
[0041] A complete segmentation of the patient's pelvis is produced with no or very little user interaction. A refinement process for error correction and validation by the user is implemented for clinical consistency, allowing for the generation of a 3D anatomical model adapted for surgical preparation or any further use.
[0042] The method comprises two main steps: Phase A / anatomical structure segmentation (without nerves) or "Segmentation Phase A", and - Phase B / Nervous system extraction from diffusion-weighted MRI images or "Nerve detection Phase B" (nerve detection and integration in the segmentation images resulting from the anatomical structure segmentation Phase A).
[0043] This method enables the extraction of clinical information of surgical interest from medical images, such as a diffusion-weighted MRI sequence (dMRI) with the following parameters: 25 directions, b = 600, voxel size of 1.25 × 1.25 × 3.5 mm, TE = 0.0529 seconds, and TR = 4 seconds, and a T2-weighted MRI sequence (T2w) with the following parameters: TE = 0.065945 seconds and TR = 2.4 seconds (TE is the echo time, TR is the repetition time, and b is the b-value). As is known in the medical imaging arts, dMRI images are functional MRI sequences that emphasize water in tissue, and T2-weighted images are a standard diagnostic sequence.
[0044] Segmentation Phase A uses an artificial intelligence (AI) algorithm based on supervised machine learning, leveraging a database of pediatric MRI and CT scans to enable rapid and complete automatic modeling of a patient's pelvic anatomy. The AI system is trained (during Training Phase A1) on reference images manually annotated by expert surgeons and radiologists before it can be used on new images during the second phase (referred to herein as Inference Phase A2). The AI system is modular and includes subsystems for localization (determination of a bounding box defining the region of interest), semantic segmentation, and assisted refinement based on error correction. The system implements an easily usable and extremely fast 3D anatomical model creation chain adapted to clinical surgical deadlines. The 3D anatomical model can be integrated into visualization applications for more efficient preoperative planning.
[0045] Nerve detection stage B uses as input segmentation images representing bones, organs, and / or blood vessels (and / or particular regions of these types of anatomical structures) obtained as output of stage A or obtained using any other image segmentation method that produces images representing bones, organs, and / or blood vessels.
[0046] Additionally, specific regions of anatomical structures detected in the segmentation images may be identified in a 3D anatomical model that includes the segmentation images. The term "specific region" or "region of interest" is used herein to refer to regions of interest other than blood vessels, bones, nerves, and organs that may be used in nerve definition for nerve detection phase B. These regions of interest may include, for example, sacral foramina, sacral canals, intervertebral foramina, etc.
[0047] The nerve detection phase B allows for the reconstruction of the peripheral nervous system and fine nerve structures, as well as the investigation of nerves (including lumbosacral nerve roots). The reliability and clinical consistency of the reconstructed nerves is ensured by a symbolic AI system that integrates surgical anatomical knowledge expressed in the form of spatial relationships and geometric characteristics. The symbolic AI system uses fuzzy definitions to take into account imprecision in clinical definitions and to adapt to anatomical modifications resulting from various pathologies. The modeling of anatomical knowledge is performed with maximum consistency with the patient's anatomy. For example, a complete representation of the sacral plexus in children and the visualization of the autonomic nervous system in adults can be achieved.
[0048] This method has been tested on medical images of the pelvis and allows for the complete automation of the image segmentation process. An automatic segmentation algorithm of the main pelvic anatomical structures (e.g., bones, sacrum, blood vessels, bladder, colon, or their specific regions such as the sacral foramen) with an accuracy higher than 90% (measured by the Dice index) has been achieved. The algorithm always locates the pelvic region of the image during an initialization step (see below). The segmentation time required during inference stage A2 using the trained segmentation system was significantly reduced: 2 minutes instead of 2 hours for coaxial bone segmentation using the existing system.
[0049] This method consistently demonstrated excellent representation of the sacral plexus in children. Anatomical variations associated with different pathologies were fully taken into account, and users positively evaluated its contribution to preoperative planning, especially for neurogenic tumors. Presurgical nerve visualization enabled the preservation of the sciatic nerve in two cases of neuroblastoma and one case of neurofibroma. Different quantitative metrics demonstrated a correlation between the development of the pelvic nerve network and several types of anorectal malformations.
[0050] The method can be adapted to normal and pathological images, regardless of whether the patient is a child or an adult. The methods disclosed herein can be used and developed, for example, for structural and functional modeling of various body regions (e.g., the reproductive system, thoracic organs, facial nerve, etc.) for both children and adults.
[0051] Four types of anatomical structures are considered herein: blood vessels, bones, organs, and nerves. The term "organ" is used herein to refer to anatomical structures other than blood vessels, nerves, and bones, and may include the stomach, heart, muscles (piriformis, coccyx, obturator, and levator ani), bladder including ureters, female reproductive system (vagina, uterus, and ovaries), colon, rectum, digestive tract, pelvis, etc.
[0052] In the described embodiments and examples, the images are MRI images, but the system is equally applicable to other types of medical images adapted to image various anatomical structures (including blood vessels, bones, organs, and nerves, or specific regions thereof). The MRI images here are 3D (three-dimensional) images.
[0053] A system for accurate computer-aided modeling of major organs-at-risk (OAR) from preoperative MRI images of children with tumors and malformations is disclosed. The system is based on supervised learning techniques using deep neural networks with convolutional layers. The system automatically adapts functions and parameters based on the structures of interest. Additionally, a set of symbolic artificial intelligence methods is used for accurate representation of peripheral nerve networks via diffusion imaging (dMRI) processing.
[0054] A collection of over 150 training images (MRI and CT images) can be acquired. All anatomical structures of interest are manually segmented by a surgeon or radiologist to generate corresponding reference (manual) segmentation images. The segmentation algorithm is then trained using the training image / reference (manual) segmentation image pairs to learn the correspondence between voxel features and organ labels.
[0055] Figure 1 shows a block diagram of a training system 100, according to one embodiment, configured to train several machine learning algorithms that enable segmentation of medical images.
[0056] The training system 100 may include one or more of the following subsystems: a preprocessing subsystem 120 , a detection subsystem 130 , a segmentation subsystem 140 , an aggregation subsystem 150 , and a refinement subsystem 160 .
[0057] Training data 110 is provided at the input of training system 100 and may be used by different subsystems as described herein.
[0058] The training data 110 may include training input images 110a (i.e., training original images), reference segmentation images 110b (also referred to as ground truth) corresponding to each training original image, bounding box parameters 110c (i.e., reference bounding boxes 130b) defining regions of interest within the original images, and annotation data 110d.
[0059] The reference segmentation image 110b and bounding box parameters 110c can be generated manually by an expert user from the training input images 110a using, for example, some drawing software tools that allow for each anatomical structure represented in the input image to select voxels that represent this anatomical structure and provide an identification of the anatomical structure. To generate a segmentation image from an input image, the user can outline a region of interest in the input image, input a label that identifies the anatomical structure represented by this region of interest, and repeat this operation for each anatomical structure (bone, blood vessel, or organ, or specific region thereof) visible in the input image.
[0060] The annotation data 110d may be generated manually as well. The annotation data 110d may include a bounding box (e.g., defined in 3D space) of an anatomical region of interest (e.g., the pelvic region). The annotation data 110d may include labels, i.e., identifiers of anatomical structures detected in the segmentation image during manual segmentation. The annotation data 110d may further include patient identification information (gender, age) and medical information (diagnosis, pathology, surgical approach, etc.).
[0061] The preprocessing subsystem 120 is configured to perform preprocessing operations on the training input images 110a to generate preprocessed training images 130a from the training input images 110a.
[0062] The detection subsystem 130 is configured to train a bounding box detection algorithm 130c using the preprocessed training images 130a and the corresponding reference bounding boxes 130b as training data, and the bounding box detection algorithm 130c is configured to generate a bounding box 130d based on the preprocessed training images 130a.
[0063] The detection subsystem 130 is further configured to generate reduced preprocessed training images 140a from the preprocessed training images 130a generated by the preprocessing subsystem 120 and the corresponding 3D bounding boxes 130d generated by the bounding box detection algorithm 130c.
[0064] The bounding box detection algorithm 130c is a supervised machine learning algorithm adapted to detect 3D bounding boxes in 3D images that define regions of interest. The reference bounding boxes 130b of the preprocessed training images 130a can be determined from the corresponding reference segmentation images 110b and / or bounding box parameters 110c input by the surgeon.
[0065] The bounding box detection algorithm 130c may include a deep neural network. A 2D deep neural network may be used to improve speed and portability. The deep neural network may be based on a 2D Faster R-CNN architecture using a ResNet-101 backbone model trained on the ImageNet dataset. According to one example, this architecture includes a feature extraction network 420, a region of interest (ROI) pooling network, and a 2D Faster R-CNN architecture. It may include a region proposal network (RPN) 430 with modules 440, a detection network 450, and a 3D merge module 460. An example architecture is described in detail with reference to, for example, FIG.
[0066] Returning to FIG. 1 , the segmentation subsystem 140 is configured to train several segmentation algorithms 140b (i.e., segmentation neural networks 140b1, 140b2, 140b3, 140b4) using the reduced preprocessed training images 140a and the corresponding reference segmentation images 110b as training data. The segmentation algorithms 140b may be supervised machine learning algorithms adapted to perform segmentation. They may be implemented as neural networks, also referred to herein as segmentation neural networks. Each of the segmentation neural networks is implemented as a 3D convolutional neural network.
[0067] The first segmentation neural network 140b1 is a bone-specific neural network trained to detect one or more bones in an image, and is also referred to herein as a bone segmentation neural network, which is also referred to herein as a specific neural network for bones or a bone-specific neural network.
[0068] The second segmentation neural network 140b2 is a vessel-specific neural network trained to detect one or more vessels in an image, and is also referred to herein as a vessel segmentation neural network, which is also referred to herein as a specific neural network for vessels or a vessel-specific neural network.
[0069] The third segmentation neural network 140b3 is an organ-specific neural network trained to detect one or more organs in an image, and is also referred to herein as an organ segmentation neural network, which is also referred to herein as an organ-specific neural network or an organ-specific neural network.
[0070] The fourth segmentation neural network 140b4 is a multi-structure neural network trained to simultaneously detect bones, blood vessels, and organs (and / or specific regions of bones, blood vessels, and organs) in an image. The fourth segmentation neural network is also referred to herein as a multi-structure segmentation neural network, a multi-structure neural network, or a multi-structure model.
[0071] For each reduced preprocessed training image 140a received as input, four segmentation images may be generated by the fourth neural network 140b1-140b4, respectively: a first segmentation image 14c1 showing one or more detected bones is generated by the bone segmentation neural network; a second segmentation image 14c2 showing one or more detected blood vessels is generated by the blood vessel segmentation neural network; a third segmentation image 14c3 showing one or more detected organs is generated by the organ segmentation neural network; A fourth segmentation image 14c4 showing one or more detected anatomical structures (including blood vessels, bones, organs, and / or specific regions of bones, blood vessels, and organs) is generated by the multi-structure segmentation neural network.
[0072] The segmentation subsystem 140 may apply a patch extraction algorithm 141 to each training reduced preprocessed image 140a received as input to extract one or more training patches for each input image and use these training patches as inputs to train the segmentation neural networks 140b1-140b4, thereby reducing the processing time required to train the segmentation neural networks.
[0073] The aggregation subsystem 150 includes an aggregation function 150a configured to compare these three segmentation images 14c1-14c3 at the output of the unique segmentation network and detect inconsistencies between the three segmentation images 14c1-14c3 when corresponding voxels in at least two of these segmentation images represent distinct anatomical structures.
[0074] If there is a mismatch between the three segmentation images 14c1-14c3, a probability score can be calculated for each voxel to determine which label is correct for a given voxel. The probability score can be calculated using a voxel-wise softmax function applied to segmentation image 14c4 at the output of multi-structure segmentation network 140b4 as a reference, which is compared with the corresponding outputs of voxel-wise softmax functions applied to segmentation images 14b1-14b3 obtained at the outputs of the other unique segmentation networks 140b1-140b3. Aggregation function 150a assigns the mismatched voxel a label corresponding to the label with the highest probability score. In another embodiment, the label of output segmentation image 14c4 is used to replace the mismatched voxel in the combined image.
[0075] The aggregation function 150a is further configured to generate an aggregated segmentation image 150b from the four segmentation images 14c1-14c4 at the output of the unique segmentation networks 140b1-140b3 and the multi-structure segmentation neural network 140b4.
[0076] The refinement subsystem 160 is configured to implement an automated refinement process that simulates user input (e.g., generated by a random algorithm) that identifies one or more erroneous voxels encoded as false positives or false negatives, respectively. The refinement subsystem 160 is configured to receive as input a given segmentation image 14c1-14c4 and generate an interaction map that is fed to one or more additional convolutional layers of one or more associated neural networks 140b1-140b4 that generated the associated segmentation image. Examples and further aspects of the refinement subsystem are described with reference to FIG. 5.
[0077] Further aspects of the training system 100 are described with reference to FIGS.
[0078] 2 shows a schematic flowchart of a method for training a supervised machine learning algorithm, according to one or more embodiments, which may be implemented using, for example, the training system 100 described with reference to FIG.
[0079] In step 200, training data is obtained. The training data may include, for example, training input images 110a, reference segmentation images 110b corresponding to each training original image, bounding box parameters 110c defining regions of interest within the original images (i.e., reference bounding boxes 130b), and annotation data 110d. For a training input image, at least one reference segmentation image and bounding box may be Available.
[0080] Input dataset
[0081] In one or more embodiments, the input dataset may include over 200 raw abdominal MRI images of children aged 0-18 years. For each patient, one volumetric T2-weighted (T2w) MRI image of the abdominal area 110a is obtained. The MRI sequence is optimized to maximize modeling accuracy and minimize acquisition time. The MRI sequence may be used in a clinical setting. For each patient, annotation data 110d is automatically created, storing the following information: date, diagnosis, medical condition, surgical approach, organ status, and image characteristics.
[0082] group
[0083] A cohort of 131 unique unrelated patients was used with a total of 171 original MRI images, of which 32 patients underwent both pre- and postoperative imaging.
[0084] Figure 3 shows the descriptive distribution of the input dataset: (a) gender distribution, (b) age distribution by gender, and (c) pathology distribution. The population consisted of 82 females and 49 males. At the time of image acquisition, 34 patients (25.95%) were younger than 1 year old, 16 (12.21%) were between 1 and 3 years old, 20 (15.27%) were between 3 and 6 years old, 11 (8.39%) were between 6 and 9 years old, 14 (10.69%) were between 9 and 12 years old, 23 (17.56%) were between 12 and 15 years old, and 13 (9.92%) were older than 15 years old. Females had a mean age of 92.73 months (±68.11 months) and a median age of 89.50 months (25th percentile: 23.25 months, 75th percentile: 154.00 months). The men had a mean age of 59.48 months (+ / - 64.60 months) and a median age of 36.00 months (25th percentile: 5.00 months, 75th percentile: 115.00 months). The population was divided into three categories: tumor, malformation, and control. The population consisted of 64 patients with pelvic malformation, 60 patients with pelvic tumor, and 7 control patients (no visible lesions at the time of image acquisition). The tumor group had a mean age of 101.267 months (+ / - 65.70 months), the malformation group had a mean age of 50.78 months (+ / - 57.97 months), and the control group had a mean age of 170.43 months (+ / - 27.56 months).
[0085] Manual annotation
[0086] Manual segmentation of anatomical structures can be performed in various ways. Each anatomical structure can be identified by a label (e.g., a specific voxel value used to identify this anatomical structure in the segmentation image). For example, pelvic organ segmentation can be performed manually on T2w images by several expert pediatric surgeons and / or radiologists. Segmented anatomical structures can include the pelvic bones (hips and sacrum) and muscles (piriformis, coccyx, obturator, and levator ani), the bladder including the ureters, the female reproductive system (vagina, uterus, and ovaries), the colon and rectum, the iliac veins and arteries, and / or specific regions of the anatomy (sacral foramen, sacral canal, intervertebral foramina).
[0087] Due to differences in age and field of view, the location and visibility of organs and structures may vary significantly between images. For each image, a bounding box is defined that limits the region of interest. For pelvic images, the bounding box may be defined superiorly by the location of the iliac artery bifurcation and inferiorly by the lowest point of the perineum.
[0088] For each training image, manual contouring of the organ at risk is performed. Specific regions to be detected by the manual contouring are identified in the image (e.g., sacral canal, sacral foramen, intervertebral foramina, and ischial spines). The image resulting from this manual contouring is typically defined as the "ground truth" and is also referred to herein as the reference segmentation image. The ground truth is used to train the deep neural network as a reference during loss function calculation.
[0089] Each training image may be manually annotated with bounding box parameters 110c that contain the coordinates of the smallest bounding box that contains all labels (e.g., the smallest bounding box of the image that contains all anatomical structures of interest).
[0090] For the training images, a bounding box can be automatically determined using the manual segmentation image of the hip bones and sacrum (reference segmentation image 110b), and the bounding box can be used to annotate the training image to generate an annotated image showing the reference bounding box 130b. The bounding box can be defined by six values (x1, x2, y1, y2, z1, z2), where x1 and x2 are the minimum and maximum transverse coordinates of the manual segmentation, respectively, y1 and y2 are the minimum and maximum coronal coordinates of the manual segmentation, respectively, and z1 and z2 are the minimum and maximum sagittal coordinates of the manual segmentation, respectively.
[0091] k-fold training
[0092] The input dataset can be divided into a training dataset, a validation dataset, and a test dataset, for example, in a ratio of 95% (training + validation) to 5% (testing). The validation dataset is used during the training loop to check whether the associated model (i.e., convolutional neural network) is behaving as expected (decreasing loss function, increasing evaluation metric) and evaluate whether the associated model is sufficiently trained. The test dataset is used to accurately evaluate model performance independently from the data used during the training loop.
[0093] A k-fold dataset split can be performed on the input dataset to enable ensemble learning. In ensemble learning, a base model is trained for each different fold and used as a building block to design a more complex model that combines several base models. One method for combining base models is to perform random selection, according to which base models are trained in parallel and independently of each other, and then combined using a deterministic averaging process. Another method for combining base models is to perform bootstrap aggregation (bagging) selection, according to which base models are trained sequentially in an adaptive manner, and then combined using a deterministic process. An example of an ensemble learning algorithm is described, for example, with reference to FIG. 6.
[0094] Random k-fold or bootstrap aggregation can be used. Because a high correlation between the type of pathology and age and gender has been detected, the training dataset is divided into folds to stratify patients based on pathology alone (tumor vs. malformation vs. control), and thus, the pathology can be used to perform k-fold dataset division. For bagging selection, no limit is placed on the amount of replacement sampled, meaning that a sample image may be present multiple times within a fold. The training dataset is divided into, for example, five folds (e.g., all images are present at least once in the training set) to utilize all available data. Each fold is then divided into a training subset, a validation subset, and a test subset, for example, with a ratio of 80% training, 15% validation, and 5% test, respectively.
[0095] Returning to FIG. 2, steps 210-240 may be performed for each training fold or until a training criterion is met (the training criterion is checked for the total number of epochs for step 220, and for step 235 either the total number of epochs or a small validation metric improvement (1e per validation metric over the last 20 iterations). -4 (increase less than 0.001).
[0096] In step 210, the training input images are preprocessed to generate preprocessed training images. One or more of the following preprocessing operations may be performed on the preprocessed training images: cropping non-zero regions, correcting bias field artifacts, correcting image orientation, resampling images, clipping functions, and normalization.
[0097] Cropping of non-zero regions may be performed since voxels with zero values do not correspond to anatomical structures.
[0098] Correction of bias field artifacts produced by the image acquisition MRI system can be performed using, for example, non-parametric non-uniform normalization.
[0099] Since image acquisition protocols may vary depending on the acquisition system, image orientation correction may be performed to have all images in the same format (e.g., NlfTl standard) and in representation space (right-anterior-superior (RAS) space).
[0100] Resampling of some images may be performed so that all images have the same resolution: this operation may include, for example, image resampling based on trilinear interpolation. The target voxel size for resampling may be the median image resolution of the entire input / training dataset (0.88 × 0.85 × 0.88 mm).
[0101] A clipping function may be applied to voxel values to retain only intensity values between the 0.5th and 99.5th percentile of the statistical distribution of foreground voxels, while other voxels are considered noise.
[0102] Normalization (e.g., Z-score normalization) of the images I(x, y, z) can be performed so that all images have a single standard deviation.
number
number
[0103] In step 220, the preprocessed training images 130a are used as training data along with their corresponding reference bounding boxes 130b to train a bounding box detection algorithm 130c. All image voxels outside the bounding box are set as background, or all voxels outside the defined bounding box are removed from the search space during segmentation to reduce processing time and reduce the number of voxels. In order to reduce the gap, the voxels outside the bounding box are ignored. After removal of the voxels outside the bounding box, a reduced preprocessed training image is obtained. By "reduced" image, we mean that only the 3D portion of the preprocessed training image and the corresponding reference segmentation image that fits within the corresponding bounding box are used.
[0104] Training of the bounding box detection algorithm 130c may be performed using only bone labels: bones (e.g., hip joints in the pelvic region) can be detected more easily and with less error than other anatomical structures, such as blood vessels. In one or more embodiments, each training image in the training dataset is converted into a series of coronal slices, and all slices that do not contain bone labels may be discarded. The coronal slices are provided as input to a 2D deep neural network used as the bounding box detection algorithm 130c.
[0105] For training, an enlarged bounding box can be used to ensure that all anatomical structures of interest are located inside the enlarged bounding box to have a final 3D search space used for segmentation. If the model is trained only on bounding boxes detected for bones, the predicted bounding box can be automatically enlarged (e.g., by 5, 10, or 20%) to always include the iliac artery bifurcation. The hips width (HW) is calculated as the box axial width, and the coordinates of the enlarged bounding box can be defined by increasing the height of the initial bounding box, for example, as follows:
number
[0106] The detection error is determined by comparing the detected bounding box BB with a corresponding reference bounding box 130b of preprocessed training images 130a (ground truth bounding boxes), using, for example, mean average precision as a metric, and is used to update the parameters of a bounding box detection algorithm 130c. The bounding box detection algorithm may be a deep convolutional neural network that minimizes a categorical cross-entropy loss function to distinguish between foreground and background boxes. An exemplary bounding box detection subsystem 130 and related algorithms are disclosed, for example, with reference to FIG. 4.
[0107] Returning to Figure 2, in step 230, salient regions (patches) of each reduced preprocessed training image are extracted to generate training patches that are sent to the segmentation algorithm, thereby preventing exceeding the memory limitations of the graphics card. Each reduced preprocessed training image is converted into a set of training patches. The centers of the image patches have coordinates determined using a sampling algorithm adapted to the type of anatomical structure to be detected by the 3D convolutional neural network under consideration.
[0108] Several patch selection algorithms can be used: random sampling, uniform sampling, local minority sampling, or co-occurrence balancing sampling. A suitable patch selection algorithm aims to minimize class imbalance and reduce segmentation errors in the case of highly imbalanced data (e.g., organs with large differences in volume). In other words, it is designed to generate a subset of patches in which thin and large structures are equally represented. Such a patch selection algorithm prevents overfitting for large organs and underfitting for small organs. For an N-class segmentation problem, a patch selection process can be chosen that samples an equal number of voxels (1 / N) for each class. This is a challenging problem due to label co-occurrence.
[0109] For each training epoch (one epoch may correspond to the period until the backpropagation algorithm calculations are performed on all patches of images belonging to the current fold in the training dataset), a dataset composed of randomly sampled image patches is created using the algorithms best suited to different types of organs (i.e., one algorithm is selected for blood vessels, one for bones, and one for organs, respectively). Patch selection may be iteratively refined by selecting patches to minimize the Kullback-Leibler divergence between the current distribution of labels and the target distribution. As known in the art, the Kullback-Leibler divergence may be determined based on the following formula:
number
number
number
[0110] It has been determined that for training the bone segmentation neural network 140b1, the most suitable sampling algorithm is random sampling with 12 samples per image. For training the blood vessel segmentation neural network 140b2, the most suitable sampling algorithm is random sampling with 6 samples per image. For training the organ segmentation neural network 140b3, the most suitable sampling algorithm is random sampling with 12 samples per image. For training the multi-structure segmentation neural network 140b4, the most suitable sampling algorithm is co-occurrence balancing sampling.
[0111] Returning to FIG. 2, in step 235, the training patches generated by the patch extraction algorithm 141 and the corresponding reference segmentation images 110b are used to train four segmentation neural networks 140b to generate four predicted segmentation images 14c1-14c4 at the output of the segmentation subsystem 140.
[0112] Each segmentation neural network 140b1-140b3 is trained to segment a given type of anatomical structure that may be present in the anatomical region of interest. The fourth segmentation neural network 140b4 is trained using all these types of anatomical structures, i.e., bones and / or blood vessels and / or organs (and / or specific regions of such types of anatomical structures), that may be present in the anatomical region of interest.
[0113] After the last convolutional layer of the segmentation neural network, an activation function AF is applied to generate the segmented image.
[0114] For each type of anatomical structure of interest (bone, blood vessel, or organ), a segmentation neural network ("specific neural network") 140b1-140b3 is trained, for example, using a sigmoid activation function after the last convolutional layer.
number
[0115] The fourth segmentation neural network 140b4 ("multiple structure neural network") (where all structure is present in the input) is trained, for example, using a softmax activation function after the last convolutional layer.
[0116] As known in the art, a softmax activation layer is a function in a neural network used to normalize the output of the previous layer into a probability distribution. A softmax activation layer is a function that is applied to an input vector Z, where i varies between 1 and K (equal to the number of labels). i to output vector
number
number
[0117] It should be noted that for the unique neural networks and multi-layered neural networks, training for each fold in the training dataset may be performed, for example, as described with reference to Figure 6. Furthermore, for each unique neural network, a student neural network may be trained during training phase A1, and used in place of the expert neural network during inference phase A2, for example, as described with reference to Figure 6.
[0118] Loss function for training stage A1
[0119] A loss function is determined for each segmentation neural network and used to update the parameters of the associated segmentation neural network. The loss function is based on a combination of metrics based on volume, location, and distance measurements and is adapted to the type of anatomical structure (bone, vessel, or organ) being segmented.
[0120] Loss function parameters used for backpropagation (alpha, beta, gamma, L 領域 , L 局所 , L 距離 ) can be automatically adjusted based on the type of anatomical structure considered for segmentation (i.e., vessel, bone, or organ), and a learning rate scheduler with linear warm-up and cosine decay is used for all networks.
[0121] For training a segmentation neural network that is specific to blood vessels (thin structures), the following combined loss function was used: L=αL 領域 +βL 局所 +γL 距離 During the ceremony, -L 領域 is a first score (or a first specific loss function) that measures accuracy based on global and local information over the 3D volume in the reference segmentation image and the current segmentation image; -L 局所 is a second score (or a second specific loss function) that measures the local consistency between voxels of the reference segmentation image and voxels of the current segmentation image, -L 距離 is a third score (or a third specific loss function) that measures the distance between the boundary points of the reference segmentation image and the boundary points (i.e., contours) of the current segmentation image.
[0122] All these scores can be computed patch-by-patch between the predicted output and the ground truth reference for each training image batch to reduce memory usage. The coefficients a and y can be dynamically varied during the training phase A1 as a function of the training epoch ep, as illustrated by Table 1 below.
[0123] By using these three types of scores, the loss function considers three types of information: area (first score), voxel (second score), and distance (third score). Each score is weighted depending on the type of anatomical structure (bone, blood vessel, or organ) being segmented. Additionally, for each structure, the weight is dynamically modified during the training phase A1 based on the current epoch (see Table 1 below). In this way, the network can learn different information at different times during the training phase A1.
[0124] Table 1 below shows the weights used in the loss function depending on the type of anatomy and the training epoch. [Table 1]
[0125] According to Table 1, the third score is not used for bones or organs, but only for blood vessels, as a function of the training epoch (ep). The loss function focuses the model on imperfections at the edges of the labels, so dynamic adjustment of the third score as a function of the training epoch (ep) allows for increased accuracy in the final part of training. This strategy also has the advantage of not affecting training time, since the distance-based loss is only used for refinement purposes (training using a distance-based function from the first epoch would be unstable at the start, since the prediction boundaries are not clearly defined).
[0126] First score L 領域 is mostly used at the beginning of the training cycle to maximize accuracy by leveraging local and global information across the 3D volume. For training on large structures (bones or organs), the first score L 領域 is defined as the Soft Dice score. For the training of thin structures (blood vessels), the first score L 領域is defined as the Tversky score, and the parameters α and β may be set to 0.3 and 0.7, respectively. For the Soft Dice score Ds and the Tversky score Tv, the formula:
number
[0127] One possible formula for these scores Ds and Tv is provided below.
number
[0128] Second Score L 局所 is the binary cross-entropy (BCE). The second score allows measuring the local consistency between voxels. This formula follows the same notation as the Dice and Tversky scores. Using the same notation, BCE can be defined as:
number
[0129] Third Score L 距離 is the pairwise distance between the contours of the manually segmented target mask of the anatomical structure and the current prediction performed by the segmentation neural network. This is L 領域 It is useful for very thin structures where L is not robust to small errors. 距離can be computed using any distance measure (Euclidean distance, Hausdorff distance, or average distance). In practice, the average distance is more robust to local segmentation errors and can be computed using the fast Chamfer algorithm. The average distance can be defined as:
number
[0130] Due to the high memory and computational cost of distance-based losses, their direct use in the first stage of the training cycle is not feasible. Therefore, distance information can only be used in more advanced stages. Additionally, to further reduce the computational load, the computation of pairwise distances can be limited to the boundary points of the predictions. The boundary points are computed directly on the GPU using morphological erosion followed by logical operators. The erosion is computed as a MinPool operation. For multi-class problems, we use 領域 and L 局所 The generalized Dice loss and category loss entropy for .times. ...
[0131] The neural network coefficients (including the coefficients of the convolution filters) are updated to minimize the loss function using an iterative process. Various known methods can be used, such as gradient descent algorithms, e.g., methods based on stochastic gradient descent. The initial values of the neural network coefficients can be set using various strategies, e.g., using the He initialization strategy.
[0132] Learning Rate (LR)
[0133] The learning rate is a hyperparameter that controls how much the neural network weights are changed in response to the estimated error at each backpropagation step. Choosing the learning rate is difficult because too small a value can result in a long training process that leads to underfitting, while too large a value can result in an unstable training process.
[0134] Different learning rate schedulers may be implemented. A learning rate scheduler is a method that dynamically adapts the value of the learning rate depending on the current training iteration. Scheduling the learning rate This has several beneficial properties for training cycles, notably reducing overfitting and increasing training stability. Specifically, a learning rate system with an initial warm-up, stagnation, and decay can be used. The warm-up increases stability in the first epoch. Once the loss stabilizes, a fixed learning rate is used for longer, constant training. Finally, to fine-tune the neural network without incurring overfitting, the learning rate is decayed by a fixed ratio (e.g., 30% of the initial learning rate).
[0135] The following schedulers are available: -No difference between warm-up and decay strategies, linear warm-up linear decay. - Linear warm-up exponential decay with faster decay to reduce overfitting. A linear warm-up cosine annealing nonlinear decay process that delays the falloff to later epochs. This smoother decay system has been used for vascular and colonic training and has been shown to provide better accuracy results.
[0136] In one or more embodiments, λ is defined as the base learning rate, α is defined as the learning rate in the first batch, and β is defined as the learning rate in the last batch. λ=4·e -4may be used as the best overall learning rate for all structures. For the warm-up phase, α = 0.6λ may be used, reaching λ in epoch = 20. For the decay phase, β = 0.6λ may be used, with decay ending after epoch = 100.
[0137] Returning to FIG. 2, in step 240, the four predicted segmentation images 14c1-14c4 are aggregated (eg, by an aggregation function 150a) to generate an aggregated segmentation image 150b.
[0138] An aggregated segmentation image 150b is generated (by an aggregation function 150a) by combining the four predicted segmentation images 14c1-14c4 at the outputs of the intrinsic neural network and the multi-layered neural network.
[0139] Combining the four segmentation images 14c1-14c4 may be performed such that the voxels in the aggregated segmentation image 150b are: - if corresponding voxels in the first, second and third segmentation images are detected as representing separate anatomical structures, the corresponding voxels in the fourth segmentation image 14c4 generated by the multi-structure segmentation network; in an alternative embodiment, when corresponding voxels in the first, second and third segmentation images represent separate anatomical structures, the voxel in the aggregated image is the label corresponding to the anatomical structure with the highest probability score among the separate anatomical structures, which probability score is calculated based on the softmax output of the multi-structure segmentation network 140b4 and the intrinsic segmentation networks 140b1 to 140b3 ax described above for the training phase with respect to Figure 1. a corresponding voxel of the first, second or third segmentation image 14c1, 14c2 or 14c3 generated by one of the unique segmentation networks in which an anatomical structure is detected in the corresponding voxel, if only one anatomical structure is detected in only one of the first, second and third segmentation images in the corresponding voxel; - the corresponding voxels of the preprocessed image 130d at the input of the segmentation subsystem 140 if no anatomical structure is detected in the corresponding voxels in the first, second and third segmentation images.
[0140] Once the machine learning algorithm is sufficiently trained (e.g., after 200 training epochs for single-organ networks, 500 training epochs for multi-organ networks, or when the results on the validation dataset improve by less than 0.01%), the trained algorithm can be applied to newly acquired images for automatic and fast segmentation of new patient images: see, for example, the method steps described with reference to FIG. 8. Computational segmentation times of approximately 30 seconds per image can be achieved.
[0141] In step 236, an automatic refinement process may be performed (e.g., by refinement subsystem 160) on one or more predicted segmentation images 14c1-14c4 at the output of intrinsic neural networks 140b1-140b3 and multi-layered neural network 140b4. For each patch extracted by patch extraction algorithm 141, one or more refinement cycles may be iteratively performed, with each refinement cycle producing a new segmentation mask and a new interaction image.
[0142] During the refinement cycle, user input is simulated: for each training patch, a binary segmentation mask is calculated using the output predicted segmentation images 14c1-14c4 generated by the associated neural network 140b1-140b4. One or more false positive voxels and / or one or more false negative voxels simulating the user input may be randomly selected and used to shrink or expand the binary segmentation mask.
[0143] One or more interaction maps are generated based on the false-positive and / or false-negative voxels. The interaction map can be a binary mask (or a single-channel image) of the same size as the input image patch. Two interaction maps can be created: one for voxels selected as false-positives and one for voxels selected as false-negatives. In each interaction map, a binarized sphere can be generated around each selected voxel: for example, a sphere of radius r = 3 voxels is centered on the associated voxel. The binary segmentation mask and the two interaction maps can be combined to form a three-channel 3D image (e.g., an RGB image) to form a three-channel 3D interaction image.
[0144] This interaction image is then used as input to an additional convolutional layer (also called a refinement block) added to the associated segmentation neural network (either the intrinsic neural network 140b1-140b3 or the multi-layered neural network 140b4) that generated the predicted segmentation images 14c1-14c4.
[0145] The outputs of these refinement convolutional layers are then merged with the inputs of the encoder blocks of the U-net architecture of the associated segmentation neural network (e.g., the output images are added voxel-by-voxel, such that corresponding voxels are added). Figure 5 illustrates an example of a U-net architecture merged with additional convolutional layers used as refinement blocks via the encoder blocks, and is described in detail below.
[0146] The loss function is calculated based on the segmentation image at the output of the neural network after correction with false positive and / or false negative voxels generated during the refinement cycles, such that the backpropagation algorithm automatically takes into account the corrections brought about by the refinement cycles.
[0147] 4 shows a block diagram of a bounding box detection subsystem for implementing the bounding box detection algorithm 130c. The system receives an RGB image 410 as input and has the objective of providing a minimal bounding box 470 that contains the structure of interest as output. The bounding box detection subsystem may be based on the Faster R-CNN architecture as described herein. The architecture may be a 2D Faster R-CNN architecture that uses a ResNet-101 backbone model trained on the ImageNet
[11] dataset.
[0148] According to the example illustrated in FIG. 4, the architecture includes a feature extraction network 420, a region proposal network (RPN) 430 with an ROI (region of interest) pooling module 440, a detection network 450, and a 3D merge module 460.
[0149] The feature extraction network 420 is configured to extract and represent low-level information of the initial image using a series of convolutional layers. The feature extraction network 420 used can be a ResNet-101 model. All parameters of the feature extraction network 420 can be calculated using the available ImageNet
[11] dataset as a target. The RPN module generates region proposals, which means selecting regions of the image that are likely to contain foreground objects extracted by the feature extraction network 420. These regions are then evaluated as false positives or true positives. Because the proposals can be of variable size, which is not suitable for processing by a neural network, the ROI pooling module 440 extracts and adapts the features extracted by the feature extraction network 420 with the coordinates provided by the RPN module (e.g., using a "max-pool" operator) to obtain a fixed-size feature map. The region proposal network 430 and the detection network 450 are specifically trained on a training dataset to localize the exact location of bones from 2D RGB coronal MRI slices. To create a final 3D bounding box that contains the bone, the 3D merge module 460 is used to merge together all the 2D bounding boxes predicted by the bounding box detection subsystem for different MRI slices to generate a 3D bounding box. During the merging operation, it ensures that all 2D predictions are included in the 3D result.
[0150] The feature extraction network 420 can be implemented as a convolutional neural network and is configured to extract low-level information about the image. The resulting feature maps are used to generate a set of anchors (image regions of a predetermined size) that are utilized by the region proposal network 430 to obtain a list of possible regions containing structures of interest. The list of proposed regions is provided to the ROI pooling module 440, which generates a list of feature maps all of the same size at the locations indicated by the regions. A max-pooling operator can be used by the ROI pooling module 440 to obtain a fixed-size feature map. This fixed-size feature map is then fed to the detection network 450, which is composed of fully connected layers, and the detection network outputs a set of bounding boxes associated with a probability score indicating whether the object is actually contained in the bounding box. The set of bounding boxes with a probability score higher than 95% is received as the final prediction. The selected bounding boxes are merged together in three-dimensional space at 460. The output of the bounding box detection subsystem is a smaller 3D bounding box 470 that contains all the 2D bounding boxes selected by the 3D merge algorithm 460.
[0151] For the region proposal network 430, we use binary cross-entropy as the classification loss (the probability that an anchor contains a foreground object) and mean absolute error as the regression loss (the probability that an anchor contains a foreground object). The loss function can be used to classify the boxes into foreground and background using a categorical cross-entropy loss function.
[0152] FIG. 5 shows a block diagram of an exemplary convolutional neural network that may be used for the segmentation network 140b (140b1, 140b2, 140b3, 140b4).
[0153] Each of the convolutional neural networks 140b used for segmentation may have a U-net architecture. Figure 5, described below, shows an example embodiment. These convolutional neural networks may have an encoder-decoder structure. The network is composed of an encoder made up of a series of convolution / batch normalization / downsampling operators MP and a decoder formed by a series of convolution / batch normalization / upsampling operations US.
[0154] As known in the art, the downsampling operation MP may be a max-pooling operation over a matrix divided into submatrices (which may have different sizes and may be vectors or scalars) that extracts the maximum value from each submatrix to generate a downsampled matrix. After each max-pooling operator, the size of the input is reduced to take into account details at different resolution levels. After each upsampling operator US, the size of the input is increased to restore the initial resolution level. All convolutions use a three-dimensional kernel. An activation layer (e.g., softmax activation) is added to the final convolution of the entire network for label prediction.
[0155] The base architecture is an encoder-decoder convolutional neural network (540). 510 is a patch extracted from the preprocessed image 130a. The encoder is made up of blocks 540a, 540b, and 540c, and the decoder is made up of blocks 540e, 540f, and 540g.
[0156] Each encoder-decoder block 540a, 540b, 540c, 540e, 540f, 540g is configured to execute a processing pipeline including a 3D convolution, an activation function (e.g., ReLu activation), and a normalization function (e.g., batch normalization). Each encoder block 540a, 540b, 540c at level n in the encoder may receive as input the output of the previous block (level n+1) after application of a downsampling operator MP (e.g., a max pooling operator represented by a vertical arrow), as represented by FIG. 5, and the output of one of the refinement blocks 550a, 550b, 550c.
[0157] Each block of the decoder receives as input the concatenation between the output of the previous decoder block after application of the upsampling operator US and the output of the encoder block at the respective level. For example, decoder block 540f receives as input the concatenation between the output of the previous decoder block 540e after application of the upsampling operator US and the output of encoder block 540b at the corresponding level (by stacking the block outputs along the outermost dimension). The first encoder block 540a of the network may receive the initial image 510 as input and the output of refinement block 550a as input. In the last layer, the output of decoder block 540g is connected to an activation function AF (e.g., a softmax operator) to produce a final segmented image 570 showing segmented anatomical structures 570a, 570b, 570c, and 570d.
[0158] Block 540d is defined as the bottleneck of the structure. The bottleneck block 540d is defined using the same configuration of convolutions from the previous encoder layer, but starting from this point on the max-pooling operator, it is replaced by an upsampling operator. The bottleneck block 540d receives as input the output of the last max-pooling operation MP. Like the other encoder blocks, the bottleneck block 540d is configured to execute a processing pipeline including a 3D convolution, an activation function (e.g., ReLu activation), and a normalization function (e.g., batch normalization). The output of the bottleneck block 540d is provided to the first decoder block 540e after application of the upsampling operator US.
[0159] Pyramid layers 520 and / or depth look-up 560 may be used.
[0160] For organ and bone segmentation, a pyramid layer 520 may be added at each level of the encoder (e.g., after every max pooling operation). The first convolution block 540b after the max pooling operator receives as input the concatenation between the result of the previous layer from block 540a and a downsampled version of the original image from block 520a.
[0161] For vessel segmentation, deep supervision block 560 can be used, adding an activation layer (e.g., softmax activation) to the decoder convolution. The loss function is then computed as a weighted sum of the losses at different levels of the decoder, with higher weights for the outermost levels, which correspond to the results produced by blocks 540a and 540g.
[0162] As represented by FIG. 5, block 520 represents an optional pyramid input function block that includes one or more pyramid blocks. Each pyramid block 520a, 520b, and 520c performs a downsampling operation DS, for example, using trilinear interpolation. The first block 520a receives the original image 510 as input. The downsampling operation DS can reduce the size of the image by half. The output of each pyramid block 520a, 520b, and 520c is optionally connected to the output of a respective encoder block 540a, 540b, and 540c of the base network 540 and provided as an input to the next encoder block. For example, as represented by FIG. 5, encoder block 540b receives the outputs of pyramid block 520a and encoder block 540a as inputs, and can also receive the output of refinement block 550b as input. All inputs of an encoder block or decoder block are connected in the outermost dimension.
[0163] Block 560 is a deep supervision module that includes one or more operator blocks. Each of operator blocks 560a, 560b, and 560c applies an activation function AF (e.g., a softmax operator) to the output of a respective decoder block 540f, 540e, and 540d of 540. A loss function may be calculated based on a comparison of the output image of the outermost operator block 560a with a corresponding reference segmentation image (i.e., manually generated ground truth). The loss function may optionally be calculated as a weighted sum of losses between each deep supervision block (560a, 560b, and 560c) and the corresponding reference segmentation image.
[0164] Blocks 550 and 530 are one example of a refinement subsystem (see FIGS. 1 and 7, refinement subsystems 160, 760) that includes one or more refinement blocks 550a, 550b, 550c and an error detection module 530. False positives or false negatives identified by the user during the user-interactive refinement process (e.g., step 83) are identified as false positives or false negatives. An interaction image encoding erroneous voxels (which may be encoded as spheres), selected by user interaction (see FIG. 8 in 5) or automatically by a random algorithm (see FIG. 2 in step 236), is provided by error detection module 530 and fed to a series of refinement blocks 550a, 550b, 550c. As described herein with respect to steps 236 and 835 of FIG. 2 or FIG. 8, error detection module 530 may generate, during the refinement step (using either user interaction or a random algorithm), an interaction image that includes three channels: a segmentation mask and two binary interaction maps for false positive and false negative voxels, respectively.
[0165] Each refinement block 550a, 550b, 550c, like the encoder block, is configured to execute a processing pipeline including a 3D convolution, an activation function (e.g., ReLu activation, rectified linear unit activation function), and a normalization function (e.g., batch normalization). The convolutions are calculated, for example, over blocks of 3x3x3 voxels using a stride equal to 2, effectively reducing the size of the input in the process. The output of each refinement block is added to the output of the respective encoder block of 540. This mechanism allows for consideration of user input in the segmentation network 540.
[0166] Block 570 represents the final anatomical structure segmentation resulting from the last layer of base network 540 after application of an activation function (e.g., a softmax operator) to block 540g.
[0167] To reduce the required computational time and resources, we distill the set of ensembles (teacher models) from the expert neural network into a student neural network (or student model), as described in detail with respect to Figure 6. We provide the student with patches, labels, and ensemble predictions.
[0168] FIG. 6 shows a block diagram of a subsystem for ensemble learning and generation of trained neural networks that can be used to obtain any of the trained segmentation networks 140b1-140b4.
[0169] By aggregating the predictions obtained by the specific neural networks, it is possible to reduce variance and improve the accuracy of the results. In addition, for each type of anatomical structure, one specific neural network can be trained per fold and per type of anatomical structure, and the predicted segmentation images obtained by the specific neural networks are aggregated through averaging, thus creating one ensemble per type of anatomical structure, resulting in N ensembles (N=3, bone, blood vessel, organ).
[0170] Prior to aggregation, the predicted segmentation images 14c1-14c3 of the outputs of the intrinsic segmentation neural networks are compared to detect inconsistent voxels and determine the segmentation uncertainty. If there is a discrepancy for a given voxel (e.g., two different intrinsic neural networks assign the same voxel two different labels), the label provided for this given voxel by the multi-layered neural networks is used to estimate the segmentation uncertainty via a maximum softmax probability μ.
number
[0171] An uncertainty criterion can be used to establish the final label of a voxel where two or more voxels disagree. The class label i with the lower uncertainty or highest probability score is assigned to the voxel. As described above for the training phase, when corresponding voxels in the first, second, and third segmentation images represent distinct anatomical structures, the voxel in the aggregated image is the label corresponding to the anatomical structure with the highest probability score among the distinct anatomical structures, which probability score is calculated based on the voxel-wise softmax outputs of the multi-structure segmentation network 740b4 and the unique segmentation networks 740b1-740b3.
[0172] 6, for each image of the training dataset in the form of a patch (610) and for each type of anatomical structure network (bone, vessel, organ), block 620 calculates the output of an expert model (i.e., neural network) 620 as the average of the outputs of the trained models (i.e., neural networks) obtained for folds 620a-620e (one trained model is obtained for each fold, and a single expert model is obtained from the trained models obtained for all folds). This result is calculated as the average of the last layer of the trained model, which corresponds to a sigmoid activation function.
[0173] To reduce the computational load, the knowledge of these trained models is distilled into a student model (630), which has the same architecture as the expert model 620, but instead of being trained on ground truth labels, it is trained on the initial raw images and ensemble predictions provided by the output of block 620. The final segmentation result (640) for each new anatomical structure to be segmented is then predicted solely by this trained student network as obtained at the end of training stage A1.
[0174] For this training, a binomial loss function may be used. The first term (structural loss) is calculated as the Soft Dice score between the predicted output of the student model and the ground truth (as the average of the trained model's outputs obtained for different folds). The second term (distillation loss) is calculated as the Kullback-Leibler divergence between the outputs of the innermost convolutional layers. We propose to use this final distilled model for fully automatic semantic segmentation of OAR.
[0175] For each type of anatomical structure (bone, blood vessel, organ), a unique student neural network is trained to generate trained segmentation networks 740b1-740b3, and all of these trained neural networks 740b1-740b3 (see the discussion of FIG. 7 relating to the bone-specific neural network, organ-specific neural network, and blood vessel-specific neural network) are used for each new image to be segmented during inference stage A2 to generate three segmentation images 74c1, 74c2, 74c3 (instead of using a corresponding more complex expert neural network obtained for this type of anatomical structure, as described with reference to FIG. 6). Similarly, the multi-structure neural network 140b4 described herein can be trained in the same manner and used during inference stage A2 to generate a fourth segmentation image 74c4.
[0176] FIG. 7 illustrates a trained image segmentation system 700 configured to perform segmentation of medical images using a trained supervised machine learning algorithm. 1 shows a block diagram of a supervised machine learning algorithm (hereinafter referred to as a trained system). Training of the supervised machine learning algorithm may be performed as described herein. Training of the supervised machine learning algorithm may be performed, for example, as described with reference to FIGS. 1 and 2.
[0177] The trained system 700 enables accurate computer-aided 3D modeling of major organs at risk from preoperative pediatric MRI (or CT) images of children who also have tumors and malformations. The system automatically adapts functions and parameters based on the structures of interest. The system allows for the creation of patient-specific 3D models in a matter of minutes, significantly reducing the amount of user interaction required. The 3D models can be exported for visualization, interaction, and 3D printing.
[0178] The trained system 700 may include one or more subsystems: a preprocessing subsystem 720, a bounding box detection subsystem 730, a segmentation subsystem 740, an aggregation subsystem 750, and a refinement subsystem 760.
[0179] The pre-processing subsystem 720 may be, for example, identical to the pre-processing system 120 described with reference to Figures 1 and 2. The pre-processing subsystem 720 is configured to generate a pre-processed image 73a from the input image 71a.
[0180] The bounding box detection subsystem 730 may implement a bounding box detection algorithm 730c corresponding to the trained bounding box detection algorithm 130c, e.g., as obtained at the end of the training stage A1 and described with reference to Figures 1, 2, and 4. The bounding box detection subsystem 730 is configured to generate a bounding box 73d from the preprocessed image 73a. The bounding box detection subsystem 730 is configured to generate a reduced preprocessed image 74a from the preprocessed image 73a and the corresponding bounding box 73d generated by the preprocessing subsystem 720.
[0181] The segmentation subsystem 740 may implement a segmentation algorithm 740b corresponding to the trained segmentation algorithm 140b, e.g., as obtained at the end of the training stage A1, and described with reference to Figures 1, 2, 5, and 6. Each of the segmentation networks 140b1, 140b2, 140b3, 140b4 obtained after the training stage A1 corresponds to a respective trained neural network 740b1, 740b2, 740b3, or 740b4 of the segmentation subsystem 740.
[0182] The segmentation subsystem 740 is configured to receive a reduced preprocessed image 74a generated from the preprocessed image 73a and the corresponding bounding box 73d. The segmentation subsystem 140 can apply a patch extraction algorithm 741 to each reduced preprocessed image 74a received as input to extract one or more patches. The extracted patches are fed to individual segmentation neural networks 740b1-740b3 and a multi-structure segmentation neural network 740b4.
[0183] The aggregation subsystem 750 includes an aggregation function 750a, which, similar to the aggregation function 150a described with reference to Figures 1 and 2, is configured to generate an aggregated segmentation image 75b from four segmentation images 74c1-74c4 obtained at the output of the unique segmentation neural networks 740b1-740b3 and the multi-structure segmentation neural network 740b4.
[0184] The refinement subsystem 760 may be identical to the refinement subsystem 160 described with reference to Figures 1, 2, and 5, for example, except that it uses true user input instead of a random algorithm to identify false positive and / or false negative voxels. The refinement subsystem 760 is configured to receive the segmentation images 74c1-74c4 as inputs and generate an interaction map that is fed to one or more additional convolutional layers of one or more associated neural networks 740b1-740b4.
[0185] FIG. 8 illustrates a schematic flowchart of a method for generating a segmentation image, according to one or more exemplary embodiments.
[0186] An input image 71a is obtained in step 800. The input image 71a may be, for example, an original image such as that acquired by an MRI system.
[0187] In step 810, preprocessing is performed by preprocessing subsystem 720. Input image 71a is preprocessed (e.g., by preprocessing subsystem 720) to generate preprocessed image 73a. The preprocessing operation may be identical to the preprocessing operation described for step 210.
[0188] Bounding box detection is performed by the bounding box detection subsystem in step 820. The preprocessed image 73a is provided as an input image to a trained bounding box detection algorithm 730c to detect a bounding box 73d corresponding to a region of interest that is used as a 3D search space and to generate a reduced preprocessed image 74a.
[0189] In step 830, the reduced preprocessed image 74a is segmented using the segmentation subsystem 740. The reduced preprocessed image 74a is first provided as input to a patch extraction algorithm 741, and then the extracted patches are provided to the four trained segmentation neural networks obtained in step 830. For each reduced preprocessed image 74a, four segmentation images 74c1 to 74c4 are obtained: a first segmentation image 74c1 showing one or more detected bones is generated by the bone segmentation neural network 740b1; a second segmentation image 74c2 showing one or more detected blood vessels is generated by the blood vessel segmentation neural network 740b2; a third segmentation image 74c3 showing one or more detected organs is generated by the organ segmentation neural network 740b3; A fourth segmentation image 74c4 showing one or more detected anatomical structures (including blood vessels, bones, or organs, and / or specific regions of such anatomical structures) is generated by the multi-structure segmentation neural network 740b4.
[0190] In step 840, aggregated segmentation image 75b is generated (e.g., by aggregation function 750a) by combining the four segmentation images 74c1-74c4 obtained in step 823. The combination of the four segmentation images 74c1-74c4 is performed such that the voxels in aggregated segmentation image 75b are: - if corresponding voxels in the first, second, and third segmentation images are detected as representing distinct anatomical structures, then the corresponding voxels in the fourth segmentation image 74c4 generated by the multi-structure neural network 740b4, alternatively, if corresponding voxels in the first, second, and third segmentation images are detected as representing distinct anatomical structures, represents the label of the voxel in the aggregated image corresponding to the anatomical structure with the highest probability score among the distinct anatomical structures, where this probability score is calculated based on the voxel-wise softmax output of the multi-structure segmentation network 740b4 and the intrinsic segmentation networks 740b1-740b3; the corresponding voxel of the first, second, or third segmentation image 74c1, 74c2, or 74c3 generated by one of the unique neural networks 740b1-740b3 in which an anatomical structure is detected in the corresponding voxel, if only one anatomical structure is detected in only one of the first, second, and third segmentation images in the corresponding voxel; - the corresponding voxels of the preprocessed image 73d at the input of the segmentation subsystem 740 if no anatomical structure is detected in the corresponding voxels in the first, second and third segmentation images.
[0191] In step 850, the aggregated segmentation image 75b can optionally be refined through smoothing using a Laplacian operator (e.g., lambda=10, where lambda is the displacement of the vertex along the curvature flow, and higher lambda corresponds to stronger smoothing) and exported for clinical use.
[0192] In step 835, an optional interactive refinement process may be performed by the refinement subsystem 760 on one or more of the segmentation images 74c1-74c4 obtained in step 830. This interactive refinement process may be performed before (or after) the aggregation of the segmentation images during step 840. The results from the automatic segmentation method, while mostly correct, may still be locally erroneous, reducing their clinical potential. For each patch extracted by the patch extraction algorithm 741, one or more refinement cycles may be iteratively performed, with each refinement cycle producing a new segmentation mask and a new interaction image.
[0193] During the interactive refinement process, binary segmentation masks of the associated segmentation images 74c1-74c4 obtained at the output of the intrinsic neural networks 740b1-740b3 or the multi-layered neural network 740b4 are automatically calculated. Additionally, the associated segmentation images 74c1-74c4 obtained in step 830 are displayed (e.g., together with the input image 71a), and a user interface is provided to allow the user to select one or more erroneous voxels encoded as false positives or false negatives, respectively. A corrected segmentation image 76a is generated from the associated segmentation images 74c1-74c4 by correcting the selected erroneous voxels based on the anatomical structure identifications provided by the user for the erroneous voxels.
[0194] Step 835 may be repeated iteratively to allow the user to iteratively correct more voxels and generate second, third, etc. corrected segmentation images. At the end of the refinement process, a final corrected segmentation image 76a is obtained.
[0195] The final corrected segmentation image is used to refine the associated segmentation neural network, e.g., at least the specific segmentation neural network that generated the erroneous voxel. For example, if the erroneous voxel is part of a bone and a blood vessel is detected instead, it may be necessary to refine both the bone segmentation neural network and the blood vessel segmentation neural network. For example, if the erroneous voxel is part of a bone and a different bone is detected instead, it may be necessary to refine both the bone segmentation neural network and the blood vessel segmentation neural network. In this case, only the bone segmentation neural network needs to be improved. In all cases, the multi-structure segmentation neural network can also be improved.
[0196] As previously described herein, an interaction map is generated from the erroneous voxels. The interaction map can be a binary mask of the same size as the input image patch. Two interaction maps can be created: one for voxels selected as false positives and one for voxels selected as false negatives. For each interaction map, a binarized sphere can be generated around each selected voxel: for example, a sphere of voxels with radius r=3 is centered on the associated voxel. The binary segmentation mask and the two interaction maps can be combined to form a three-channel 3D image (e.g., an RGB image) to be used as the interaction.
[0197] The interaction image is used to refine the associated segmentation neural network, e.g., at least the inherent segmentation neural network that generated the erroneous voxel. This interaction image is fed to an additional convolutional layer, the output of which is merged with the input of the corresponding encoder block of one or more associated segmentation networks 740b1-740b4, respectively, as further described with reference to FIG.
[0198] One or more additional convolutional layers (see refinement blocks 550a, 550b, and 550c in FIG. 5) can be added to the base segmentation architecture of the refined intrinsic segmentation neural network to process the interaction image and generate outputs to feed into the corresponding encoder blocks of the U-Net architecture. The additional convolutional layers are stacked together to form parallel branches to the encoder of the U-Net, as illustrated in FIG. 5. For each refinement block, a 3D convolution is followed by a linear activation unit and batch normalization. The resulting output of the refinement block is added to the respective output of the base convolution just before the max pooling operation from the previous encoder block.
[0199] During the training phase A1, the learning rates of these additional convolutional layers can be set higher (e.g., 10 times, 3 times, 5 times higher) than the learning rates of the base layers (encoder and decoder blocks) so that the outputs of the refinement blocks have more weight in the learning phase.
[0200] The interaction images are provided as input training data to these one or more additional convolutional layers of the associated segmentation neural network. For example, to align the network output with the user annotations (e.g., the output voxel labels must belong to the same class as indicated by the user), a specific loss function can be used to train the additional convolutional layers and update the coefficients of the refinement blocks during the inference stage A2 (see blocks 550a-550c in FIG. 5). This loss function can be defined as:
number
number
number
[0201] Steps 810-840 may be repeated for each new input image. When the additional convolutional layers are sufficiently trained, step 835 does not need to be performed, and the output segmentation image 750a generated at the end of step 830 may be automatically corrected using the trained additional convolutional layers (trained refinement blocks) to reduce or avoid the time required for the user to correct erroneous voxels in the output segmentation image from the four segmentation images obtained in step 830 in step 835. Thus, during step 830, the additional convolutional layers are used to automatically correct the output segmentation image obtained in step 830 to generate a corrected output segmentation image.
[0202] FIG. 9 shows a flowchart of a method for generating a 3D anatomical model.
[0203] In step 910, a training stage A1 is performed using the training data set 91a to generate a trained neural network 91b. This training stage A1 may include, for example, one or more or all of the steps of the method for training a supervised machine learning algorithm described with reference to FIG. 2.
[0204] In step 920, an inference stage A2 is performed using one or more new images 92a to generate corresponding segmentation images 92b representing the vessels, organs, and bones (and / or particular regions of such anatomical structures). This inference stage A2 may include, for example, one or more or all of the steps of the method for generating segmentation images described with reference to Figure 8. An interactive refinement process is part of this inference stage A2.
[0205] In step 930, a nerve detection phase B is performed using the nerve definition 93a and segmentation image 92b to generate a 3D anatomical model 93b that includes nerves, blood vessels, organs, and bones (and / or specific regions of such anatomical structures). This nerve detection phase B is described in detail below.
[0206] Neural Detection and Modeling
[0207] The proposed method combines the segmentation of anatomical structures from MRI and / or CT imaging with diffusion MRI image processing for nerve fiber reconstruction by exploiting their spatial properties. The tractography algorithm (deterministic tracking) , applied to diffusion images without any constraints on fiber start / end conditions.
[0208] Due to the huge number of false positives encountered in whole-body tractograms, a filtering algorithm is used that exploits the spatial relationship between nerve bundles and segmented anatomical structures.
[0209] Due to resolution, noise, and artifacts present in diffusion images acquired in clinical settings, it is often difficult to obtain a complete and accurate tractogram. Instead of using traditional ROI-based approaches, whole-body tractogram analysis is used. Seed points are placed within each voxel of the image, regardless of their anatomical significance. Alternatively, they may be randomly placed within a mask. A selected segmentation algorithm is then used to remove all irrelevant streamlines. We extend this approach to extract a complete pelvic tractogram from the numerous streamlines. The end result is visualization of smaller fascicles (e.g., the S4 sacral nerve root) that are typically impossible to obtain using ROI-based tractography. For example, a tractogram of the entire pelvis can be calculated using the following parameters: a DTI deterministic algorithm with the following termination criteria: fractional anisotropy (FA) < 0.05, seed criteria: FA > 0.10, minimum fiber length: 20 mm, and maximum fiber length: 500 mm.
[0210] Approaches based on simple relationships between organ boxes are not accurate enough given the complexity of the pelvic structure. The recognition of each nerve bundle is based on natural language descriptions that define spatial information, including directional information (left, right, anterior, posterior, superior, inferior), path information (proximity, between), and / or connectivity information (crossing, endpoint in). For example, a nerve may be described as "passing through the S4 sacral foramen and crossing the sacral canal." Similar descriptions can be developed for all relevant nerve bundles. To translate these natural language descriptions into a manipulation algorithm, each spatial relationship must be modeled by an anatomical fuzzy 3D area that represents the nerve's defined direction, path, and connectivity information.
[0211] Nerve detection relies on a segmentation image, obtained, for example, as described herein or using another segmentation method, a fuzzy definition of the spatial relationship of nerves with respect to segmented anatomical structures (bones, vessels, and / or organs), and a representation of fibers in the anatomical region of interest (e.g., by a tractogram). The spatial relationship of nerves can be represented with respect to specific anatomical structures (bones, vessels, organs, and / or specific regions of such anatomical structures).
[0212] The nerve detection logic is based on spatial relationships with respect to known anatomical structures. For each nerve, a query is generated from the nerve definition in natural language. The query can be viewed as a combination of spatial relationships (e.g., using Boolean and / or logical operators). To determine whether a fiber meets the nerve definition, the query is converted into a fuzzy 3D anatomical area that represents the search space for searching for fibers.
[0213] The nerve definition can be an English-like text syntax of fiber tracts based on the expertise of neuroanatomists, or detailed descriptions in human anatomy atlases, or medical textbooks, and the English-like text syntax is a natural language translation to facilitate mathematical correspondence for performing anatomy-aware nerve detection. For example, the natural language expressions "going through" or "passing through" are related in the English-like text syntax. Translated as "crossing." The inherent ambiguity of the anatomical definition of a nerve, and of the translation from natural language to an English-like text syntax, is expressed through the use of fuzzy set theory.
[0214] For example, the neurodefinition is "anterior with respect to the colon and not posterior with respect to The term "the piriformis muscle or between the artery and the vein" may refer to the nerve being anterior to the colon but not posterior to the piriformis muscle, or between a given artery and vein.
[0215] Each spatial relationship used in the neural definition is expressed as a human-readable keyword (e.g., "lateral of"), "anterior" or "lateral" depending on the geometric parameters. The geometric parameters (e.g., "in front of," "posterior of," etc.) are converted into 3D fuzzy structuring elements. The geometric parameters may include, for example, one or several angles (α1 and α2). Other types of spatial relationships can be modeled in a similar manner.
[0216] 10 shows a flowchart of a method for nerve detection, according to one or more exemplary embodiments. The steps of the method may be implemented by an apparatus comprising means for performing one or more or all steps of the method.
[0217] Although these steps are described in a progressive manner, one skilled in the art will understand that some steps may be omitted, combined, performed in a different order, and / or in parallel.
[0218] The method allows for detecting, in one or more input images, at least one fiber representing a nerve based on one or more segmentation images comprising an anatomical structure and a nerve definition.
[0219] In step 1090, one or more input images (i.e., diffusion MRI images) are processed to obtain a 3D representation of the fibers present in the anatomical region of interest for which a segmentation image has been obtained. The segmentation image used for nerve detection may be, for example, aggregated segmentation image 75b obtained using the method disclosed with reference to Figure 8, optionally in combination with the embodiments of Figures 7 and 9.
[0220] A 3D representation of fibers can be obtained by applying a tractography algorithm to one or more diffusion-weighted MR images within a convex hull corresponding to a region of interest that encompasses at least some of the anatomical structures being segmented in the aggregated segmentation image. The tractography algorithm can be applied to one or more diffusion images using the convex hull to define a tractogram creation area. The tractography algorithm can be applied without constraints on fiber start / end conditions. The tractography algorithm can be applied by using (e.g., random) seed points for the tractography algorithm, which are placed within voxels of the diffusion images that belong to the convex hull. From these seed points, paths followed by fibers can be detected.
[0221] In step 1091, a nerve definition is obtained for a given nerve. The nerve definition expresses, for example in natural language, one or more spatial relationships between the nerve and one or more corresponding anatomical structures detected in the segmentation image; the anatomical structures may include bones, blood vessels, organs. The nerve definition may also describe specific regions (e.g., For example, it may be based on the sacral foramen.
[0222] A nerve definition may include a single definition for the entire nerve, or a chain of definitions (or chains of queries) of successive portions of the nerve. The definitions of portions of the nerve are referred to herein as sub-definitions. For each definition (or each sub-definition), a search space is obtained in which fibers may be found.
[0223] Multiple queries can be chained together to define different fuzzy 3D anatomical areas that the nerve must pass through in a defined order. Chaining multiple queries can be done in various ways: i) using the nerve streamlines (i.e., streamlines are curves in the tractogram that trace the estimated course of a bundle of nerve fibers, each streamline in the tractogram represents the reconstructed path of a specific fiber) obtained from the previous query as new input instead of the entire tractogram, together with the refined aggregated segmentation image and a new anatomical description of the nerve tract, to give results; ii) using the term "then" to separate and chain multiple queries, thereby defining different 3D search spaces, each associated with the portion of the nerve to be detected.
[0224] For example, a nerve definition using multiple queries can be "crossing sacral canal then crossing sacral hole S1 then anterior with respect to the piriformis muscle and posterior with respect to the obturator muscle," indicating that the nerve passes anteriorly inside the sacral canal, then passes through the sacral foramen S1, and finally crosses the search space anterior to the piriformis muscle and posterior to the obturator muscle. In this case, each fuzzy 3D anatomical area between the "then" operators is combined into a local fuzzy 3D anatomical area (e.g., anterior to the The combination of these is a fuzzy 3D anatomical area.
[0225] Steps 1092-1097 are performed for the nerve definition, or nerve portion definition, of a given nerve, respectively, and may be repeated iteratively for each portion in the sequence of portions defined for the nerve.
[0226] When steps 1092-1097 are applied to a nerve defined by its successive portions, an order for traversing the 3D search areas obtained for each nerve portion can be established (e.g., based on the term "then" used in the nerve definition). The first fibers may be detected based on a threshold applied within the first 3D search area obtained in the first nerve portion, and these first fibers may be provided as input (instead of the entire tractogram within the anatomical region of interest) for the next execution of steps 1092-1097 for the second nerve portion, according to the order of the nerve portions in the nerve definition. Finally, the fibers detected within the last 3D search area corresponding to the last portion of the nerve provide the representation of the nerve. Searching portion-by-portion according to the nerve portion sequence allows for an accurate nerve representation to be obtained.
[0227] In step 1092, the nerve definition (or, respectively, the nerve part definition) is converted into a set of at least one spatial relationship expressed in terms of one or more anatomical structures.
[0228] In step 1093, each spatial relationship with respect to the anatomical structure of the nerve definition (or, respectively, nerve part definition) is converted into a 3D fuzzy structuring element. This 3D fuzzy structuring element is used as the structuring element in the morphological expansion. This 3D fuzzy structuring element In a structuring element, the value of a voxel represents how well the relationship to the origin of space is satisfied for the associated voxel. The value of a voxel in the extension of an anatomical structure with this structuring element represents how well the relationship to the anatomical structure is satisfied for the associated voxel.
[0229] The 3D fuzzy structuring element encodes at least one of directional, path, and connectivity information of the associated spatial relationship of a nerve definition (or nerve part definition, respectively). For example, the directional information of the spatial relationship with a given structure may be modeled as a fuzzy cone originating from the structure and oriented in a desired direction, with fibers within the fuzzy cone satisfying the directional relationship with the structure.
[0230] In step 1094, a segmentation mask is calculated for each anatomical structure used in the nerve definition (or nerve part definition, respectively).
[0231] The segmentation mask may be a binary segmentation mask, calculated from at least one segmentation image that includes relevant anatomical structures.
[0232] In an embodiment, to improve the representation of areas where nerves and anatomical structures are in contact, while at the same time avoiding considering the structures themselves as anatomical areas of possible nerve passages, the segmentation mask generated in step 1094 and used in the next step 1095 may be a fuzzy segmentation mask. This fuzzy segmentation mask may be obtained by dilating an initial segmentation mask (corresponding to voxels representing the relevant anatomical structures) based on a structuring element, then extracting voxels corresponding to the initial segmentation mask from the dilated segmentation mask, and finally normalizing the selected voxels, for example with 0 and 1. The fuzzy segmentation mask may be used instead of a binary segmentation mask to obtain a fuzzy anatomical 3D area in step 1095 below.
[0233] In step 1095, for each spatial relationship expressed in the nerve definition (or nerve part definition, respectively) for an anatomical structure, a fuzzy 3D anatomical area is generated. The fuzzy 3D anatomical area is obtained for a spatial relationship by fuzzy expansion of the segmentation mask of the associated anatomical structure based on a structuring element expressing the associated spatial relationship. The fuzzy anatomical 3D area defines an expansion score for each voxel within the fuzzy anatomical 3D area. The structuring element used to generate the fuzzy segmentation mask may be the same as the structuring element used to expand the segmentation mask.
[0234] More precisely, for each considered spatial relationship S with an anatomical structure, a segmentation mask R of the relevant anatomical structure is extracted from the segmentation image, and this segmentation mask is then scaled by a 3D fuzzy structuring element μ that models the spatial relationship. s It is extended using the 3D fuzzy structuring element μ s can be defined by a mathematical function (e.g., using one or more geometric parameters) that assigns a score to the voxels. The result of the fuzzy expansion is D(R,μ s where D denotes the fuzzy extension and the value D(R, μ s )(P) is the degree to which voxel P satisfies relation S with respect to anatomical structure R, σ s (P). This degree is used as the dilation score for voxel P.
[0235] In embodiments where a fuzzy segmentation mask is generated in step 1094, for each considered spatial relationship S with an anatomical structure, the segmentation of the associated anatomical structure A mask R is extracted from the segmentation image, and this segmentation mask is first converted into a fuzzy segmentation mask μR, which is then converted into a 3D fuzzy structuring element μ that models the spatial relationships. s It is extended using the 3D fuzzy structuring element μ scan be defined by a mathematical function (e.g., using one or more geometric parameters) that assigns a score to the voxels. The result of the fuzzy expansion is D(μR,μ s ), where D denotes the fuzzy extension, and the value D(μR,μ s )(P) is the degree to which voxel P satisfies relation S with respect to anatomical structure R, σ s (P). This degree is used as the dilation score for voxel P.
[0236] Fuzzy 3D structuring element μ s The fuzzy expansion of an object R (or, respectively, a fuzzy object μR) with, for example, at each voxel P, D(R, μ s )(P)=sup{μ s (P-P'),P'∈R}, where μ s (P-P') is the translated μ at voxel P s at voxel P'. Any other fuzzy expansion function or algorithm may be used.
[0237] Thus, spatial relationships regarding anatomical structures are defined by 3D fuzzy anatomical areas derived from structuring elements defined based on geometric parameters.
[0238] In step 1096, if several fuzzy 3D anatomical areas were obtained in step 1095 for a nerve (or for nerve parts, respectively), the fuzzy 3D anatomical areas are combined. A fuzzy anatomical 3D area may be combined by computing it as a voxel-wise combination of one or more fuzzy anatomical 3D areas obtained in step 1095. The combined fuzzy 3D anatomical area obtained in this step defines a search space for the associated nerve (or, respectively, the associated nerve part) in the 3D representation of the fibers (i.e., the tractogram in the anatomical region of interest).
[0239] For example, when a combination of spatial relations is used in the definition of a nerve (or, respectively, the definition of a nerve part), the expanded score σ obtained at voxel P for the fuzzy anatomical 3D area represents the spatial relations S1, S2, etc. used in the definition. s1 (P), σ s2 A combined score is calculated based on (P), etc.
[0240] The Boolean function used to calculate the combined score depends on how the spatial relations S1, S2, etc. are combined in the nerve definition (or nerve part definition), which also defines how the fuzzy anatomical 3D areas are combined.
[0241] For example, a logical combination of fuzzy anatomical 3D areas may include several 3D areas combined with the logical operator "AND": in this case, a nerve can be located within a search space defined by the combination of these fuzzy anatomical 3D areas. A logical combination of fuzzy anatomical 3D areas may include several 3D areas combined with the logical operator "OR": in this case, a nerve must be located in at least one of the 3D areas. A logical combination of fuzzy anatomical 3D areas A1 and A2 may include several 3D areas combined with the logical operator "AND NOT": for example, "A1 AND NOT A2", in this case, a nerve must be located in A1 but not in A2. Any other complex definition may be used with any logical operator and any number of fuzzy anatomical 3D areas.
[0242] The extended scores obtained for fuzzy anatomical 3D areas expressing spatial relationships are are combined in a corresponding manner.
[0243] For example, when two relations are combined by the "AND" Boolean operator, i.e., "S1 AND S2", the combined score at voxel P is the expanded score σ at voxel P. s1 (P), σ s2 The minimum value MIN(σ) for each voxel of (P) s1 (P),σ s2 (P)).
[0244] For example, when two relations are combined by the “OR” Boolean operator, i.e., “S1 OR S2”, the combined score at voxel P is the expanded score σ at voxel P. s1 (P), σ s2 The maximum value of each voxel in (P) is MAX(σ s1 (P),σ s2 (P)).
[0245] For example, if a nerve cannot exist within a given anatomical area (spatial relationship NOT(S0)), then the degree σ S0 (P) approaches zero in a given anatomical area.
[0246] In step 1097, for each fiber in the 3D representation of the fibers (i.e., the tractogram in the anatomical region of interest), a fiber score is calculated based on the fuzzy 3D anatomical areas representing the search space (either the single fuzzy 3D anatomical area obtained in step 1095 or the combined fuzzy 3D anatomical areas obtained in step 1096). The fiber score may be calculated for a fiber as the sum of the dilation scores of voxels representing associated fibers that fall within at least one fuzzy 3D anatomical area.
[0247] One or more detection criteria may be applied from those listed below: The detection criteria may be based on a fiber score.
[0248] For example, according to a first detection criterion, a detected nerve corresponds to a collection of one or more fibers that are completely or partially within the fuzzy anatomical 3D area.
[0249] For example, according to the second detection criterion, the percentage of fiber points that are inside this fuzzy anatomical 3D area is calculated for a given fiber, and this percentage is compared with a threshold to determine whether the fiber corresponds to the search nerve.
[0250] For example, according to a third detection criterion based on fiber score, a detected nerve corresponds to a set of one or more fibers having a fiber score above a threshold.
[0251] For example, the nerve detection method may retain all fibers with a fiber score higher than a threshold, calculated, for example, through the expanded scores (e.g., as a weighted average, or sum, or other aggregation function) for fiber points passing through the fuzzy anatomical 3D area. The threshold may be set, for example, to 0.5, 0.2, 0.1, or 0.8.
[0252] Further details, examples, and explanations relating to methods for neural detection are provided below.
[0253] Figure 11A shows the relation "lateral of R in direction α" (the direction vector is
number
[0254] Since the score of the directional relationship assumes cone symmetry along the direction axis, selecting the cone angle β is equivalent to selecting the tightness of the relationship. Thus, the bounding box can be represented by a cone having an angle equal to π. The user can specify another cone angle to change the tightness of the relationship. The viewing angle can be normalized between 0 and 1 for intuitiveness, where 1 indicates a π aperture and 0 indicates a linear projection.
[0255] Figure 11B shows the corresponding extended fuzzy 3D anatomical area obtained by applying such a 3D fuzzy structured element to an anatomical structure.
[0256] Another example of a spatial relationship is <<in proximity of R (close to R)>>. In this case, μ s can be expressed as the distance to the original O. The Euclidean distance ED can be used as the distance: ED(O,P) = ||P - O||2. Then, μ s can be defined as a decreasing function of the distance from the origin O of voxel P. Then, the anatomical structure R is μ sto define a fuzzy 3D anatomical area corresponding to the relation "in proximity of R". A threshold on the distance may be applied. For interpretability reasons, proximity may be converted from voxel space to millimeters, and it may also be possible to consider anisotropic voxel sizes.
[0257] For a "between" relationship, a fuzzy 3D anatomical area is defined by two anatomical structures. A direction vector can be calculated as the vector connecting the structure centroids. For 3D anatomical areas that contain multiple disjoint areas, the calculation of the 3D anatomical area is performed for each pair of connected anatomical structures.
[0258] "Connectivity" can be modeled using the definitions "crossing" and "endpoints in" to specify the regions of the image that fibers cross and the regions of the image that fibers connect.
[0259] The complete list of natural language relations can be represented as an abstract syntax tree, which can be reduced in hierarchical order to finally compute fuzzy 3D anatomical areas that satisfy the combination of relations.
[0260] FIG. 11C illustrates an example of the spatial definition of a nerve.
[0261] Element 1010 in FIG. 11C represents various 3D geometric areas 1010a, 1010b, 1010c that can be used as directional relationship definitions for simple points with varying rigor. 3D area 1010a includes two cones with small openings (angles α1 and α2), each defined from an origin P1, P2, respectively. 3D area 1010b includes two cones with large openings (angles α1 and α2), each defined from an origin P3, P4, respectively. ) 3D zone 1010c is a 3D parallelepiped (i.e., a bounding box with aperture equal to π) defined from origin P5. Cone 1010a has an aperture value of 0.125. Cone 1010b has an aperture value of 0.5. Parallelepiped 1010c has an aperture value of 1.
[0262] Elements 1020 represent different types of spatial relationships for anatomical structures obtained by automatic segmentation of MRI images: 1020a represents the definition of the anterior coccygeus muscle; 1020b represents the definition of the lateral obturator muscle; 1020c represents the definition of the proximity of the ovaries; and 1020d represents the definition between the piriformis muscle and the ureter.
[0263] Element 1030 represents an example of a logical combination of multiple spatial relationships, each defining a 3D fuzzy area, forming a complex 3D fuzzy area definition defined by several 3D fuzzy areas. The logical combination of 3D fuzzy areas defines a search space by logical combinations of 3D fuzzy areas where a nerve may and / or must not be located. 1030a represents a definition of the posterolateral side of the coccygeus muscle. 1030b represents a definition of the lateral superior side of the bladder. These complex areas allow for a more precise definition of the possible locations of the nerve (or nerve segments, respectively), while excluding some impossible locations.
[0264] 11D and 11E illustrate an example of the spatial definition of a nerve.
[0265] Figure 11D shows an example of a connectivity definition. Figure 11D shows a "crossing" definition for an anatomical area. Fibers that do not have any points inside the highlighted region 1040a are discarded.
[0266] 11E shows an example of a connectivity definition: "endpoints in" definition for an anatomical area. Fibers that do not have a pole point (first or last point of the fiber) inside the highlighted region 1040b are discarded.
[0267] 12 illustrates an aspect of a method for neural recognition, according to one embodiment. A T2w MRI image 1110a is provided as input to a convolutional neural network 1120 for segmentation, e.g., obtained as described herein or using another segmentation method. For each detected nerve, an anatomical definition 1130 is generated using natural language: the anatomical definition includes a definition of the nerve's possible locations, thereby enabling the determination of one or more 3D fuzzy anatomical areas that define a 3D search space corresponding to the spatial definition for the nerve. A list of anatomical-spatial relationships is generated 1130 for each nerve of interest using natural language and logical operators.
[0268] The anatomical spatial relationships 1130 of the nerves are given in natural language with respect to the segmented anatomical structures 1120. A 3D search space 1150 is then determined using the anatomical definitions and corresponding geometric definitions of the structuring elements based on mathematical functions. As explained, the search space can be either a single fuzzy 3D anatomical area obtained in step 1095 or a combined fuzzy 3D anatomical area obtained in step 1096. A combined 3D fuzzy anatomical area is determined based on a logical combination of the spatial relationships of the nerve definitions. The combined 3D fuzzy anatomical area can be computed as a voxel-wise combination of the fuzzy anatomical 3D areas that each represent these spatial relationships.
[0269] The diffusion image 1110b is used to generate a whole-body tractogram 1140 using a DTI deterministic tractography algorithm, such as that disclosed in
[10] . Fibers of the whole-body tractogram 1140 that lie partially or completely within the defined 3D search space 1150 are detected as corresponding to one of the searched nerves 1160a-1160e, allowing for the creation of an accurate nerve-based representation of the peripheral nerve network 1160. The detected nerves are added 1170 to the 3D anatomical model to create a final 3D representation 1170 of the patient's anatomical and functional features.
[0270] Following the detailed description of specific embodiments, it appears clear that the invention can be understood in a more general way, which generalization refers to the possibility of implementing the invention for other body regions, both for children and adults, in addition to the pelvis.
[0271] The present invention therefore relates to a general method for generating a segmentation image representing an anatomical structure, comprising at least one bone, at least one organ or at least one blood vessel, said general method comprising: - performing a segmentation of the input image to generate a first segmented image by detecting voxels representing bones using a first 3D convolutional neural network CNN applied to the input image, the first 3D CNN being a bone-specific neural network trained to detect one or more bones in the image; - performing a segmentation of the input image to generate a second segmented image by detecting voxels representing organs using a second 3D CNN applied to the input image, the second 3D CNN being an organ-specific neural network trained to detect one or more organs in the image; performing segmentation of the input image to generate a third segmented image by detecting voxels representing blood vessels using a third 3D CNN applied to the input image, the third 3D CNN being a blood vessel-specific neural network trained to detect one or more blood vessels in the image; performing a segmentation of the input image to generate a fourth segmented image by detecting voxels representing bones, organs, or blood vessels using a fourth 3D CNN applied to the input image, the third 3D CNN being a multi-layered neural network trained to detect bones, blood vessels, and at least one organ in the image; generating an aggregated segmentation image by combining the first, second, third, and fourth segmentation images.
[0272] According to a variation of this general method, if corresponding voxels in the first, second, and third segmentation images are detected as representing distinct anatomical structures, then a voxel in the aggregated image is the corresponding voxel in the fourth segmentation image.
[0273] A variation of this general method involves obtaining a 3D representation of the fibers present in the input image; - Detecting at least one fiber representing a nerve in the input image using an output segmentation image and a nerve definition, where the nerve definition is expressed in the form of one or more spatial relationships between the nerve and one or more anatomical structures detected in the output segmentation image, and using a fuzzy 3D anatomical area corresponding to the nerve definition.
[0274] According to a variation of this general method, a score may be associated with each voxel within the fuzzy anatomical 3D area, and the general method may include calculating, for each of one or more fibers, a fiber score as the sum of the scores of voxels representing the associated fibers that fall within the fuzzy anatomical 3D area, and the detected nerves are included within the fuzzy anatomical 3D area. It is represented by at least one fiber, completely or partially, and with a score higher than the threshold.
[0275] According to a variation of this general method, the general method is: - Obtaining a segmentation mask of the anatomical structures used in the nerve definition, - Obtaining a fuzzy anatomical 3D area by expanding a segmentation mask based on a structuring element that encodes at least one of directional information, path information, and connectivity information of neural-defined spatial relationships.
[0276] According to a variation of this general method, one or more or each of the first, second, third and fourth 3D CNNs is based on a U-Net architecture.
[0277] According to a variation of this general method, the general method is: obtaining a 3D representation of the fibers present in the input image; - detecting at least one fiber representing a nerve in the input image using an output segmentation image and a nerve definition, where the nerve definition is expressed in the form of one or more spatial relationships between the nerve and one or more anatomical structures detected in the output segmentation image, and using a fuzzy 3D anatomical area corresponding to the nerve definition.
[0278] According to a variation of this general method, one or more or each of the first, second, third, and fourth 3D CNNs may be trained using a loss function defined as a weighted sum of specific loss functions, where the weighting coefficients of the weighted sum are specific to the 3D CNN under consideration and may vary dynamically during training of the 3D CNN.
[0279] According to a variation of this general method, a first specific loss function measures accuracy based on global and local information over a 3D volume in the reference and current segmentation images.
[0280] According to a variation of this general method, the second specific loss function measures the local consistency between voxels of the reference segmentation image and voxels of the current segmentation image.
[0281] According to a variant of this general method, a third specific loss function measures the distance between the boundary points of the reference segmentation image and the boundary points of the current segmentation image.
[0282] According to a variation of this general method, one or more or each of the first, second, third and fourth 3D CNNs are trained using image patches extracted from an input image, the centers of the image patches having coordinates determined using a sampling algorithm adapted to the type of anatomical structure to be detected by the 3D CNN considered.
[0283] According to a variation of this general method, the general method is: For an input image, determining a bounding box of an object of interest using a deep regression network; applying one or more or each of the first, second, third, and fourth 3D CNNs only to the bounding box of the object of interest.
[0284] According to a variation of this general method, the general method is: -CNN, wherein the associated CNN has a corresponding encoder block whose output is associated with the associated CNN. Displaying the segmentation image generated by one of the CNNs, each of which includes one or more additional convolutional layers, each of which is merged with the input of the lock; generating a segmentation mask of an anatomical structure in the segmentation image; allowing the user to select one or more voxels in the displayed segmentation image that are encoded as false positives or false negatives; providing the segmentation mask and the selected voxels as input images to one or more additional convolutional layers.
[0285] Disclosed herein is a system for modeling 3D anatomical structures based on preoperative MRI scans. This system enables the creation of patient-specific digital twins in a matter of minutes, significantly reducing the amount of user interaction required. This 3D model can be exported for visualization, interaction, and 3D printing.
[0286] It should be understood by those skilled in the art that any functions, engines, block diagrams, flow diagrams, state diagrams, flowcharts, and / or data structures described herein represent conceptual views of illustrative circuitry embodying the principles of the invention. Similarly, it will be understood that any flowcharts, flow diagrams, state diagrams, pseudocode, etc. are substantially represented on a computer-readable medium and, therefore, represent various processes that may be performed by such computers or processing devices, whether or not a computer or processing device is explicitly shown.
[0287] Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel, concurrently, or simultaneously. Also, some operations may be omitted, combined, or performed in a different order. A process may terminate when its operations are completed, but may have additional steps not disclosed in the illustration or description. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to a calling function or a main function.
[0288] Each illustrated function, engine, block, or step described herein may be implemented by hardware, software, firmware, middleware, microcode, or any suitable combination thereof.
[0289] As used herein, "means configured to perform a function" or "means for performing a function" may correspond to one or more functional blocks comprising circuitry adapted or configured to perform this function. A block may perform this function by itself or may cooperate and / or communicate with one or more other blocks to perform this function. A "means" may correspond to or be implemented as "one or more modules," "one or more devices," "one or more units," etc.
[0290] When implemented in software, firmware, middleware, or microcode, the instructions to perform the necessary tasks may be stored on a computer-readable medium that may or may not be included in the host device or host system. The instructions may be transmitted via the computer-readable medium and loaded into the host device or host system. The instructions are configured to cause the host device / host system to perform one or more functions disclosed herein. For example, as described above, according to one or more embodiments, at least one memory may contain or store instructions, and at least one memory The memory and instructions may be configured to cause the host device / host system, using at least one processor, to perform one or more functions. Additionally, the processor, memory, and instructions act as a means for providing or causing the host device / host system to perform one or more functions disclosed herein.
[0291] A host device or host system may be a general-purpose computer and / or computing system, a special-purpose computer and / or computing system, a programmable processing device and / or system, a machine, etc. A host device or host system may be or include a user equipment, a client device, a cell phone, a laptop, a computer, a network element, a data server, a network resource controller, a network appliance, a router, a gateway, a network node, a computer, a cloud-based server, a web server, an application server, a proxy server, etc.
[0292] The instructions may correspond to computer program instructions, computer program code, or may comprise one or more code segments. A code segment may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program state. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable techniques including memory sharing, message passing, token passing, network transmission, etc.
[0293] When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors, some of which may be shared. The term "processor" should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include one or more processing circuits, whether programmable or not. Processing circuits may include digital signal processors (DSPs), network processors, application specific integrated circuits (ASICs), and the like. ASIC, field programmable gate array The processor may correspond to a FPGA, ...
[0294] A computer-readable medium or computer-readable storage medium may be any storage medium suitable for storing instructions readable by a computer or processor. A computer-readable medium may more generally be any storage medium capable of storing and / or containing and / or carrying instructions and / or data. A computer-readable medium may be a portable storage medium or a fixed storage medium. A computer-readable medium may be a permanent, mass storage medium. It may include one or more storage devices such as a storage device, magnetic storage medium, optical storage medium, digital storage disk (CD-ROM, DVD, Blue Ray, etc.), USB key or dongle or peripheral, memory card, Random Access Memory (RAM), Read Only Memory (ROM), core memory, flash memory, or any other non-volatile storage device.
[0295] Suitable memory for storing instructions may be random access memory (RAM), read-only memory (ROM), and / or a permanent mass storage device such as a disk drive, a memory card, random access memory (RAM), read-only memory (ROM), core memory, flash memory, or any other non-volatile storage device.
[0296] Terms such as first, second, etc. may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as the first element, without departing from the scope of the present disclosure. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0297] When an element is referred to as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present. Other words used to describe relationships between elements should be interpreted similarly (e.g., "between" vs. "directly between," "adjacent," etc.). (e.g., "adjacent" vs. "directly adjacent").
[0298] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0299] References [1]K.He, [2]Ron Kikinis, Steve D.Pieper, and Kirby G. Vosburgh.”3D Slicer:A Platform for Subject-Specific Image Analysis,Visualization,and Clinical Support',pages 277-289. Springer New York,New York,NY,2014. [3]Shaoqing Ren,Kaiming He,Ross Girshick,and Jian Sun.Faster r-cnn:Towards realtime object detection with region proposal networks.In C.Cortes,N.Lawrence,D.Lee,M.Sugiyama,and R.Garnett,editors,Advances in Neural Information Processing Systems,volume 28.Curran Associates,Inc.,2015. [4]Olaf Ronneberger,Philipp Fischer,and Thomas Brox.U-net:Convolutional networks for biomedical image segmentation.In Nassir Navab,Joachim Hornegger,William M.Wells,and Alejandro F.Frangi,editors,Medical Image Computing and Computer-Assisted Intervention MICCAI 2015,pages 234-241,Cham,2015.Springer International Publishing. [5]Konstantin Sofiiuk,Ilia A.Petrov,and Anton Konushin.Reviving Iterative Training with Mask Guidance for Interactive Segmentation.arXiv e-prints,page arXiv:2102.06583,Feb.2021. [6]N.J.Tustison,B.B.Avants,P.A.Cook,Y.Zheng,A.Egan,P.A.Yushkevich,and J.C.Gee.N4itk:Improved n3 bias correction.IEEE Transactions on Medical Imaging,29(6):1310-1320,2010. [7]S.Ren,K.He,R.Girshick,and J.Sun,”Faster R-CNN:Towards Real-Time Object Detection with Region Proposal Networks,” [8]F.Isensee et al.,”nnU-Net:Self-adapting Framework for U-Net-Based Medical Image Segmentation” [9]Sofiiuk K.,Petrov I.,Konushin A.,”Reviving Iterative Training with Mask Guidance for Interactive Segmentation”
[10] Mori,S.;Crain,B.J.;Chacko,V.P.& van Zijl,P.C.M.Three-dimensional tracking of axonal projections in the brain by magnetic resonance imaging.Annals of Neurology,1999,45,265-269
[11] Deng J,Dong W,Socher R,Li L-J,Li K,Fei-Fei L.Imagenet:A large-scale hierarchical image database.In:2009 IEEE conference on computer vision and pattern recognition.2009.p.248-55.
Claims
1. 1. A method for generating a 3D anatomical model (1170) of an anatomical region of interest, the method comprising: obtaining an input image of the anatomical region of interest and a 3D representation (1140) of fibers within the anatomical region of interest; generating an aggregated segmentation image (75b) representing an anatomical structure, the aggregated segmentation image (75b) including at least one bone, at least one organ, or at least one blood vessel; detecting at least one fiber representing a nerve in the 3D representation of the fibers within the anatomical region of interest; generating the 3D anatomical model using the aggregated segmentation image and at least one fiber representing the detected nerve; generating the aggregated segmentation image (75b) - performing (830) a segmentation of said input image (71a) to generate a first segmented image (74c1) by detecting voxels representing bones using a first segmentation neural network applied to said input image, said first segmentation neural network being a bone-specific neural network trained to detect one or more bones in the image; - performing (830) a segmentation of said input image (71a) to generate a second segmented image (74c2) by detecting voxels representing organs using a second segmentation neural network applied to said input image, said second segmentation neural network being an organ-specific neural network trained to detect one or more organs in the image; - performing (830) a segmentation of said input image (71a) to generate a third segmented image (74c3) by detecting voxels representing blood vessels using a third segmentation neural network applied to said input image, said third segmentation neural network being a blood vessel-specific neural network trained to detect one or more blood vessels in the image; - performing (830) a segmentation of said input image (71a) to generate a fourth segmented image (74c4) by detecting voxels representing bones, organs or blood vessels using a fourth segmentation neural network applied to said input image, said fourth segmentation neural network being a multi-layered neural network trained to detect bones, blood vessels and at least one organ in the image; combining said first, second, third and fourth segmentation images; detecting the at least one fiber representing the nerve; - obtaining a nerve definition expressed in the form of one or more spatial relationships between said nerve and one or more anatomical structures segmented in said aggregated segmentation image (75b); detecting (930) that the at least one fiber satisfies the nerve definition (93a) based on at least one fuzzy 3D anatomical area corresponding to the one or more spatial relationships expressed in the nerve definition.
2. 2. The method of claim 1, wherein detecting at least one fiber that satisfies the nerve definition comprises searching for fibers within a 3D search space (1150) defined by the at least one fuzzy 3D anatomical area.
3. The method comprises: - Each anatomical structure used in the nerve definition is included in the aggregated segmentation image. Extracting from the image (75b), - obtaining a segmentation mask from each extracted anatomical structure (1094); 3. The method of claim 1, further comprising: obtaining (1095) a fuzzy anatomical 3D area representing the given spatial relationship of the nerve definition by expanding the segmentation mask based on a structuring element, the structuring element encoding at least one of directional information, path information, and connectivity information of the given spatial relationship.
4. an expansion score is associated with each voxel in the at least one fuzzy anatomical 3D area, and the method comprises: - calculating (1097) for each of said at least one fiber that falls within said at least one fuzzy anatomical 3D area a fiber score based on the dilation scores of the voxels that represent the fiber under consideration; - detecting that a nerve is represented by at least one fiber when said fiber is completely or partly within at least one fuzzy anatomical area and has a fiber score higher than a threshold.
5. the nerve definition comprises a sequence of nerve sub-definitions, each nerve sub-definition being associated with a portion of the nerve; The method further comprises: for a first portion of the nerve having an associated first sub-definition expressing at least one first spatial relationship between the first portion of the nerve and at least one first anatomical structure segmented in the aggregated segmentation image; a) obtaining a segmentation mask of the at least one first anatomical structure (1094); b) obtaining at least one first local fuzzy anatomical 3D area representing the at least one first spatial relationship by expanding the segmentation mask based on a structuring element, the structuring element encoding at least one of direction information, path information, and connectivity information of the given spatial relationship (1095); c) calculating (1097) for each of a plurality of fibers a fiber score based on the expansion scores of voxels representing said considered fiber that fall within a first fuzzy anatomical 3D area defined by the at least one first local fuzzy 3D anatomical area; d) detecting, for at least one of the plurality of fibers, that the first portion of the nerve is represented by the considered fiber when the considered fiber is completely or partially within the first fuzzy anatomical 3D area and has a fiber score higher than a threshold; 5. The method of claim 1, further comprising repeating steps a) to d) for a second portion of the nerve and by using the at least one fiber detected for the first portion to narrow down the plurality of fibers for which the fiber score is calculated.
6. the segmentation mask of the at least one anatomical structure is a fuzzy segmentation mask of the at least one anatomical structure, and the method comprises: obtaining an initial segmentation mask by extraction of the at least one anatomical structure from the aggregated segmentation image; dilating the initial segmentation mask based on a structuring element to obtain a dilated segmentation mask; The expanded segmentation mask is compared to the initial segmentation mask. extracting the corresponding voxels and normalizing the extracted voxels; The method of claim 3 or 5, comprising defining the fuzzy segmentation mask as the extracted voxels normalized.
7. 7. The method of claim 1, wherein each voxel in the aggregated segmentation image (75b) is a corresponding voxel in the fourth segmentation image (74c4) if corresponding voxels in the first (74c1), second (74c2), and third (74c3) segmentation images are detected as representing distinct anatomical structures.
8. 8. The method of claim 1, wherein one or more of, or each of, the first (740b1), second (740b2), third (740b3), and fourth (740b4) segmentation neural networks is trained using a loss function defined as a weighted sum of specific loss functions, the weighting coefficients of the weighted sum being specific to the segmentation neural network under consideration.
9. The method of claim 8 , wherein the first particular loss function measures accuracy based on global and local information over a 3D volume in the reference segmentation image and the current segmentation image.
10. The method according to claim 8 or 9, wherein the second specific loss function measures the local consistency between voxels of the reference segmentation image and voxels of the current segmentation image.
11. The method according to any one of claims 8 to 10, wherein the third specific loss function measures the distance between the boundary points of the reference segmentation image and the boundary points of the current segmentation image.
12. The method according to any one of claims 8 to 11, wherein the weighting coefficients are dynamically varied during the training of the segmentation neural network.
13. 13. The method of any one of claims 1 to 12, wherein one or more of the or each of the first, second, third and fourth segmentation neural networks are trained using image patches extracted (230) from the input image, the centers of the image patches having coordinates determined using a sampling algorithm adapted to the type of anatomical structure detected by the considered segmentation neural network.
14. The method comprises: - determining a bounding box of an object of interest for said input image using a deep recurrent network; - applying one or more or each of the first, second, third and fourth segmentation neural networks only to the bounding box of the object of interest.
15. The method comprises: one of the segmentation neural networks, wherein the associated segmentation neural network includes one or more additional convolutional layers whose outputs are each merged with the input of a corresponding encoder block of the associated segmentation neural network; Displaying the segmentation images (74c1 to 74c4) generated by - generating a segmentation mask of the anatomical structure in said segmentation image; - allowing the user to select one or more voxels in the displayed segmentation image that are encoded as false positives or false negatives; - providing the segmentation mask and the selected voxels as input images to the one or more further convolutional layers.
16. generating the aggregated segmentation image; 16. The method of any one of claims 1 to 15, further comprising, prior to combining the first, second, third, and fourth segmentation images (74c1-74c4), refining (835) at least one of the first, second, third, and fourth segmentation images (74c1-74c4) by using an iterative or automatic refinement process.
17. 17. The method of any one of claims 1 to 16, wherein each of the first, second, third and fourth segmentation neural networks is a 3D convolutional neural network CNN.
18. 18. The method of claim 17, wherein one or more of, or each of, the first, second, third, and fourth 3D CNNs is based on a U-Net architecture.
19. 19. The method of claim 1, comprising obtaining a 3D representation (1140) of the fibers by applying a tractography algorithm to a diffusion-weighted MR image (1110b) inside a convex hull corresponding to the anatomical region of interest that encloses the anatomical structure being segmented in the aggregated segmentation image, and by using random seed points for the tractography algorithm on the convex hull.
20. Apparatus comprising means for carrying out the steps of the method according to any one of claims 1 to 19.