Neural network-based reconstruction of three-dimensional bone models from MRI data
Patent Information
- Application Number
- PCT/NL2025/050463
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-24
- Filing Date
- 2025-09-18
- Publication Date
- 2026-08-27
Smart Images

Figure NL2025050463_27082026_PF_FP_ABST
Abstract
Description
[0001] Title: Neural network-based reconstruction of three-dimensional bone models from MRI data
[0002] TECHNICAL FIELD AND BACKGROUND
[0003] The present disclosure relates to medical imaging, and in particular to methods, systems, and computer-implemented techniques for automatically reconstructing three-dimensional bone models from magnetic resonance imaging (MRI) data using neural networks. The disclosure further relates to methods for training such neural networks, to the use of trained networks for bone segmentation from MRI data, and to systems and software products configured to perform such reconstruction processes. Applications may include virtual surgical planning, the design of patient-specific implants, and image-guided manufacturing of surgical guides.
[0004] 3D virtual surgical planning (VSP) is a current standard of care for maxillofacial tumor surgery for several years already. Therewith the tumor margins and other soft tissue features can be accurately determined from MRL based tumor segmentations, as MRI In a subsequent fibula free flap reconstruction the affected body parts are replaced by parts that are sculpted from a portion of the fibula of the patient. If a fibula free flap reconstruction is not possible, a patient specific implant may be milled from a bio compatible material, such as titanium, PEEK or a ceramic. A template is required to enable a proper milling of the replacement parts in these reconstruction approaches, which in turn relies on accurate bone segmentation image data. At present this segmentation data is obtained by additional CT-scans. This workflow requires rigid matching between MRI and CT, which introduces an error due to anatomical changes between scans.
[0005] The requirement for multiple imaging modalities further places additional burdens on patients, including increased time and radiation exposure from CT scans. Additionally, the need for extra scans causes strain on hospital resources, contributing to longer waiting times, higher healthcare costs and energy consumption.An article by Dong Nie et al. (DOI: 10.1007 / 978-3-319-67389-9_31) describes segmentation of craniomaxillofacial (CMF) bony structures from MRI using a three-dimensional deep-learning-based cascade framework. In this framework, an initial coarse segmentation is generated by a 3D U-net model. Since these initial segmentations may be relatively thick, they are subsequently refined by a second network, referred to as a Context-CNN, which produces a more detailed and anatomically consistent representation. The second network is trained using a Prior-guided Patch Sampling strategy, in which training patches are densely extracted around voxels exhibiting ambiguous probabilities in the output of the first stage. The framework is trained on a dataset consisting of 16 pairs of MRI and CT scans, with the CT scans segmented to obtain ground-truth bone structures for the MRI data. During preprocessing, each CT image is linearly aligned to its corresponding MR image. Afterward, the bony structures are segmented from each aligned CT image, and these CT segmentation results provide training labels for the network.
[0006] However, the process of aligning data from different imaging modalities can still present challenges. The known registration techniques may not fully account for all anatomical variations between scans acquired at different times. Such discrepancies can result in registration errors in the aligned data, which may reduce the anatomical accuracy of the ground truth used for training.
[0007] Accordingly, there remains a need for improved techniques that can more robustly account for anatomical discrepancies when generating ground-truth data from paired imaging modalities, which may enable more accurate MRI-based bone segmentation and reconstruction of three-dimensional models suitable for virtual surgical planning and the manufacture of patient-specific implants or surgical guides.
[0008] SUMMARY
[0009] It is an object of the present disclosure to at least mitigate the above-mentioned disadvantages. In accordance therewith a method for training a neural network is provided to automatically reconstruct a three-dimensional bone model from MRI-scan data. The trained neural network obtained therewithrenders it possible to determine the 3D virtual surgical planning (VSP) with MRI as the single imaging modality.
[0010] Embodiments of the method for training comprise preparing a respective pair of an MRI-scan comprising MRI-scan data, i.e. a 3D sequence, and a CT-scan with CT-scan data. Then a ground truth bone segmentation data is obtained for each region of interest in the CT-scan corresponding to a respective affected bone.
[0011] In case there are two or more bones affected that are movable relative to each other, respective ground truth bone segmentation data is obtained for the respective regions of interest in the CT-scan corresponding to the two or more bones. The ground truth bone segmentation data is linked with the CT-scan data to obtain enriched CT-scan data. Then the MRI-scan data is fused with the enriched CT-scan data to obtain enriched MRI-scan data including ground truth bone segmentation data.
[0012] By performing the step of fusing the MRI-scan data with the enriched CT-scan data separately for each region of interest, improved alignment of groundtruth data to the MRI anatomy can be achieved. By carrying out a first fusion with the region of interest set on the first bone part, a first aligned ground-truth dataset is obtained that corresponds specifically to the anatomical position of the first bone part in the MRI scan. By subsequently performing a second, separate fusion with the region of interest set on the second bone part, a second aligned ground-truth dataset is generated that corresponds to the anatomical position of the second bone part in the MRI scan.
[0013] By treating each movable bone part independently during the fusion process, discrepancies that may arise due to different relative positions of the bones between the CT and MRI acquisitions can be more effectively addressed. By isolating the motion and alignment of each region of interest, localized registration can be optimized for each bone part rather than relying on a single global alignment. This may reduce the propagation of registration errors at joint interfaces or overlapping structures and may result in more anatomically consistent ground-truth segmentation data for training the neural network.
[0014] By generating individually aligned datasets for each movable bone part, the training data may better reflect the true MRI anatomy, which can improvethe robustness and accuracy of the trained neural network when later used to process MRI scans without a corresponding CT scan. This in turn can facilitate the generation of bone models suitable for use in workflows where MRI may serve as the sole imaging modality for virtual surgical planning.
[0015] The neural network to be trained is used to estimate tentative enriched MRI-scan data from the MRI-scan data. The tentative enriched MRI-scan data comprises the MRI-scan data and tentative bone segmentation data. Estimating tentative enriched MRI-scan data can take place at any point in time once the MRI-scan data is available. It can be performed in parallel to obtaining enriched MRI-scan data. A training step of the neural network is then performed by backpropagating a loss that is indicative for a deviation between the tentative bone segmentation data and the ground truth bone segmentation data. These steps are repeated for a plurality of human body part samples.
[0016] In some embodiments the neural network is trained to process Tl-MRI-scan data. Alternatively, another MRI imaging modality, such as T2-MRI, may be used. Tl-imaging is however preferred as this scanning protocol can be performed faster than the T2-MRI protocol, therewith also reducing the risk of movement artifacts.
[0017] The neural network to be trained is typically of an encoder-decoder structure. Neural networks are available, such as SegNet, Transformer-based architectures, such as SegFormer and Swin-Unet and GAN-based models like Pix2Pix. It has been found that a U-type network is of a particular advantage for the present application in that it provides for a pixel-level segmentation accuracy in a relatively simple and efficient manner. A U-type network comprises a contraction path and an expansion path (encoder -decoder), which gives it the U-shaped architecture. The contracting path is a typical convolutional network that consists of repeated application of convolutions, each followed by a rectified linear unit (ReLU) and a max pooling operation. During the contraction, the spatial information is reduced while feature information is increased. The expansion pathway combines the feature and spatial information through a sequence of up-convolutions and concatenations with high-resolution features from the contracting path. An example of a U-type network is nnU-Net.In embodiments of the method the ground truth bone segmentation data is provided as STL-data. Alternatively, another data format may be used, such as OBJ (Wavefront Object Format), PLY (Polygon File Format), 3MF (3D Manufacturing Format), AMF (Additive Manufacturing File Format) and FBX (Autodesk Format). Nevertheless STL files are preferred as the format is widely adopted and it includes all essential information about the bone geometry with a minimum of overhead.
[0018] In embodiments of the method the segmentation data is wrapped.
[0019] Therewith a continuous surface around the segmented bones is created.
[0020] Wrapping smooths and refines the model while filling in small gaps or imperfections in the mesh. By creating a continuous surface around the segmented bones through a wrapping process, small gaps and irregularities in the raw segmentation can be closed. This may help to smooth out rough areas and connect fragmented parts of the mesh into a coherent structure. By filling these imperfections, the resulting model may become easier to interpret and more representative of the actual anatomy. Providing the neural network with such refined and consistent ground-truth data can lead to better training outcomes. Cleaner meshes may allow the network to focus on learning relevant anatomical features rather than noise or artefacts. This, in turn, may result in a trained network that produces more reliable and clinically useful segmentations when applied to new MRI data. In addition, by reducing defects in the groundtruth data, the wrapping process may improve compatibility with downstream applications, such as generating surgical guides or patient-specific implants. The smoother and more continuous surfaces may be easier to handle in design and manufacturing workflows, supporting efficient integration into virtual surgical planning systems.
[0021] The inventors find the following settings may provide particularly advantageous result:
[0022] - Gap Closing Distance: 3 mm. This determines the maximum distance for closing small gaps in the model's surface. Gaps within this distance will be closed to create a smooth and continuous mesh. By this setting, gaps caused by minor segmentation inconsistencies can be closed without over-smoothing largeranatomical openings, resulting in a more natural representation of the bone surface.
[0023] - Detail Level: 0.5 mm. This controls the level of detail retained during the wrapping process. A smaller value usually means higher fidelity to the original surface, preserving finer details. By this setting, subtle structures, such as ridges or small curvature changes, can be maintained, while still removing noise and irregularities that could interfere with training or downstream use.
[0024] By combining these parameter values, a balance can be achieved between producing a continuous, clean mesh and preserving essential anatomical detail. This may improve both the quality of the ground-truth data for training and the usability of the resulting models for subsequent virtual surgical planning and manufacturing workflows.
[0025] Of course, it will be understood that while the above-indicated values were observed to provide optimal results, deviations from these exact values may also be suitable, depending on factors such as the quality of the input data, the anatomical region being processed, and the requirements of downstream applications. For example, the gap closing distance may in some cases be set lower to preserve natural openings in thinner bone structures, or higher to ensure the closure of larger discontinuities, for instance within a range of about 1 mm to 5 mm, preferably 2 to 4 mm, most preferably 2.5 mm to 3.5 mm. Similarly, the detail level may be varied to either emphasize fine anatomical detail or to simplify the model for computational efficiency, for instance within a range of about 0.2 mm to 0.8 mm, preferably 0.3 mm to 0.7 mm, most preferably 0.4 to 0.6 mm. By allowing such flexibility, the wrapping process can be adapted to different clinical scenarios and imaging characteristics while still achieving a balance between smoothness, continuity, and preservation of anatomical fidelity.
[0026] Embodiments of the method may comprise resampling the MRI-scan data or the CT-scan data in order to match the resolution. For example, the MRI scans may be resampled to match the CT voxel size. Resampling may involve interpolation, e.g. linear interpolation.For processing efficiency, the MRI-scan data and the enriched MRI-scan data may be transformed into image data according to a common image data format, e.g. be encoded as an nii.gz file.
[0027] In accordance with the above-mentioned object, the present disclosure further includes the use of a trained neural network obtained with embodiments of the method as presented above for computing bone segmentation data from MRI-scan data. The trained neural network in response to the MRI-scan data at its input generates the bone segmentation data, therewith enabling a 3D virtual surgical planning (VSP) without a separate CT-scan.
[0028] In another or further aspect, the present disclosure relates to a computer-implemented method for automatically reconstructing a three-dimensional bone model of a region of a patient's body as described herein. The region may comprise at least a first bone part and a second bone part that are movable relative to each other, such as a mandible and a cranium in a maxillofacial application.
[0029] MRI-scan data of the region of the patient's body is received, for example, by an input interface of a computing device. The received MRI-scan data is processed with a trained neural network to compute bone segmentation data. The computed bone segmentation data comprises a first segmentation dataset corresponding to the first bone part and a second, separate segmentation dataset corresponding to the second bone part. By providing separate segmentation datasets for each movable bone part, potential overlap or ambiguity between the bones can be avoided, and each structure can be handled individually in later processing stages.
[0030] From the first and second segmentation datasets, a three-dimensional bone model is reconstructed, for example using a reconstruction module or suitable software. By reconstructing the three-dimensional bone model in this manner, the first and second bone parts are represented as distinct, individually manipulable elements. This separation allows subsequent planning or design steps to be performed on each bone independently, such as simulating surgical movements, designing osteotomy guides, or preparing for reconstruction procedures.In some embodiments, the reconstructed three-dimensional bone model is used to manufacture a patient-specific implant or a surgical guide with a manufacturing device, such as a 3D printer or milling machine. By basing the manufacturing process directly on the reconstructed model, patient-specific devices can be produced with high precision, reducing the need for intraoperative adjustments and supporting accurate surgical outcomes.
[0031] In other or further embodiments, the reconstructed three-dimensional bone model is provided for use in virtual surgical planning (VSP). By using a model that represents each movable bone part separately, clinicians may interact with the anatomy in planning software, evaluate different surgical scenarios, and optimize the surgical workflow before entering the operating room.
[0032] The steps of any computer-implemented method described herein may be executed by a suitable device running software instructions stored on a non-transitory computer-readable medium.
[0033] In another or further aspect, a system is provided for carrying out a reconstruction process, such as described herein. The system includes an input interface configured to receive MRI-scan data, a trained neural network configured to process the MRI-scan data and output the first and second segmentation datasets, and a reconstruction module configured to generate the three-dimensional bone model from these datasets. By combining these components into a single system, an end-to-end workflow can be realized in which MRI data is transformed into an accurate, digital anatomical model ready for surgical planning or device manufacturing without requiring additional imaging modalities such as CT. This integration may streamline clinical processes, reduce patient exposure to ionizing radiation, and improve the overall efficiency of treatment planning and preparation.
[0034] BRIEF DESCRIPTION OF THE DRAWINGS
[0035] These and other aspects are described in more detail with reference to the drawing. Therein:FIG 1 schematically shows an embodiment of a method for training a neural network to automatically reconstruct a three-dimensional bone model from MRI-scan data.
[0036] FIG 2A - 2D shows an overview of evaluation metrics for each anatomical structure on a test dataset.
[0037] FIGs 3A, 3B and FIG 4A - 4D provide a visual comparison between the MRI models and CT models in a first exemplary case.
[0038] FIGs 5A, 5B and FIG 6A - 6D provided a visual comparison between the MRI models and CT models in a second exemplary case.
[0039] FIG 7 shows the median deviation of MRI bone models compared to CT bone model.
[0040] FIG 8 schematically shows a use of the trained neural network.
[0041] FIG 9 shows an exemplary U-type network in embodiments of the method disclosed herein.
[0042] FIG 9A shows optional further elements to be combined with the neural network of FIG 9.
[0043] FIG 9B shows an exemplary stage of the neural network.
[0044] FIG 9C shows another exemplary stage of the neural network.
[0045] DETAILED DESCRIPTION OF EMBODIMENTS
[0046] Like reference symbols in the various drawings indicate like elements unless otherwise indicated.
[0047] In the following detailed description numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be understood by one skilled in the art that the present invention may be practiced without these specific details. In other instances, well known methods, procedures, and components have not been described in detail so as not to obscure aspects of the present invention.
[0048] FIG 1 schematically shows a method for training a neural network, such as a U-type network to automatically reconstruct a three-dimensional bone model from MRI-scan data. The steps as described in more detail below are performed for a plurality of human body part samples.In step SIM, SIC a respective pair of an MRI-scan M and a CT-scan C is obtained. Herein, and in the sequel the suffixes M, C respectively indicate an operation that is specifically related to the MRI-scan and an operation that is specifically related to the CT-scan. The MRI-scan M and the CT-scan C respectively comprise MRI-scan data and CT-scan data.
[0049] By way of example in step S IM the MRI scans M are acquired in headfirst supine position with a 3.0 Tesla MRI scanner (SIEMENS Magnetom Prisma, Siemens Healthineers, Forchheim, Germany), with a 64-channel head coil. The scans M are obtained with the suprahyoid protocol in the Tl-weigthed gradient echo (Dixon VIBE) sequence with contrast agent DOTAREM. In-phase scans are used within this study. The parameters of the MRI are: repetition time / echo time, 5.5 / 2.46 ms. The obtained MRI data M, provided in DICOM format, is a matrix of D=256 x W=264 x H=256 voxels, wherein the voxel size is 0.9 x 0.9 x 0.9 mm.
[0050] The CT scans C in step SIC may be acquired in headfirst supine position with either a Siemens SOMATOM Definition AS scanner, or a Siemens SOMATOM Force scanner. The resulting CT scan data C, provided in DICOM format, is a matrix of D=512 x W=580 x H=512 voxels, wherein the median voxel size is 0.6 x 0.6 x 0.6 mm.
[0051] In step S2C ground truth, GT, bone segmentation data Cs is obtained. In this example, the GT bone segmentation data Cs is obtained for two distinct regions of interest, ROI, in the CT-scan obtained in step SIC, i.e. the ROI of the mandible and the ROI of the cranium. In other examples the GT bone segmentation data is obtained for only one ROI or for more than two ROI’s. As specified above, this depends on the number of mutually independently movable bone parts that are affected. Different segmentation tools may be used dependent on the selected regions of interest. For example, the ground truth bone segmentation data is obtained by a manual segmentation of the CT scans C that is performed using Mimics 26.0 research software (Materialise, Leuven, Belgium). In this example, the mandible and cranium are segmented. For the orbita, the thin bone segment tool is utilized. Optionally for the NAI, a spline of 2 mm diameter is drawn through the mandibular canal until the mandibular foramen. After calculating the parts, all segmentations are exported to anstandard tessellation language (STL) format Cs aligned within the coordinate system of the CT scan data C.
[0052] In the embodiment shown, a wrapping is applied in step S3C. Wrapping is performed with a gap closing distance of 3.0 mm and a detail level of 0.5 mm. In an example, step S3C is performed with the 3-Matic 18.0 research software (Materialise, Leuven, Belgium). The resulting wrapped segmentation models Csw are also exported as STL data aligned within the coordinate system of the CT scan data C.
[0053] In step S4C the wrapped ground truth bone segmentation data Csw obtained in step S3C is then linked with the original CT-scan data C to obtain enriched CT-scan data Cl. For this purpose Brainlab software (Brainlab, Miinchen, Germany) is applicable. The enriched CT scan data Cl, provided in DICOM+STL format, is a matrix of D=512 x W=580 x H=512 voxels, wherein the median voxel size is 0.6 x 0.6 x 0.6 mm.
[0054] In step S5 MRLscan data is fused with the enriched CT-scan data to obtain enriched MRLscan data that includes ground truth bone segmentation data. To enable this fusion one of the CT-scan data or the MRLscan data may be resampled in case their image resolution differs. In the example shown in FIG 1, the MRI scans are resampled in step S2M to 0.6x0.6x0.6 mm using linear interpolation. The resulting resampled MRI data Mr, provided in DICOM format, is a matrix of D=384 x W=395 x H=383 voxels, wherein the voxel size is 0.6 x 0.6 x 0.6 mm.
[0055] In the example shown, the resampled MRLscan data Mr is transformed in step S3M into a compressed format Mrt. In the example shown, the resampled MRI scans Mr are compressed into the file format .nii.gz, which is a compressed version of the NIfTI file format. NIfTI, or Neuroimaging Informatics Technology Initiative, is a widely used format for storing neuroimaging data, such as MRI or CT scans. The .nii.gz format is preferred for its smaller file size while retaining all the data integrity of the original .nii file. In principle also other compressed image formats like jpeg, png, tiff, bmp may be used, but care should be taken that data integrity is sufficiently maintained. It may further be contemplated to directly use the DICOM image format as this is the general standard for medicalimage data. At present no ready to use interface is available for processing DICOM image data by the nnU neural network as used in the present investigation. However there are no fundamental reasons that would prevent training and using the nnU neural network to automatically reconstruct bone segmentation data on the basis of MRI-scan data provided in DICOM format.
[0056] In the example shown the fusion of the MRI-scan data Mrt with the enriched CT-scan data Cl in step S5 is performed twice.
[0057] Once the fusion is performed with the region of interest (ROI) set on the mandible during fusion, after which the mandible STL files Fa are exported and once the fusion is performed with the ROI set on the cranium, after which the cranium STL file is exported Fb. The exported STL data is aligned within the coordinate system of the MRI scan data Mrt.
[0058] In subsequent step S6 these STL files are transformed into nii.gz files Fm to serve as a mask in training. The resulting transformed data Fm comprises the wrapped ground truth segmentation data, aligned within the coordinate system of the MRI scan data Mrt, for all ROI’s, in this case the ROI of the mandible and the ROI of the cranium.
[0059] In step S7, as part of the training procedure, tentative enriched MRI-scan data Ft comprising the MRI-scan data and tentative bone segmentation data is estimated with the neural network to be trained. The tentative enriched MRI-scan data Ft, provided in nii.gz format, is a matrix of D=384 x W=395 x H=383 voxels, wherein the voxel size is 0.6 x 0.6 x 0.6 mm. During the training procedure the tentative enriched MRI-scan data Ft of the neural network will more and more accurate approach the wrapped ground truth segmentation as specified in the transformed data Fm.
[0060] In order to compute the tentative enriched MRI-scan data Ft, the MRI-scan data Mrt is provided at the input of the neural network to be trained and the tentative enriched MRI-scan data Ft is obtained at the output of the neural network. A training step is then performed by computing a loss L in step S8 that is indicative for a deviation between the tentative bone segmentation data as comprised in the tentative enriched MRI-scan data Ft and the ground truth bonesegmentation data comprised in the mask Fm and the loss is backpropagated in step S9.
[0061] As mentioned above, the steps of FIG 1 are performed for a plurality of human body part samples.
[0062] In an example, training is performed with a dataset comprising 100 paired T1 MRI scans and CT scans of oncology patients from the department of Oral and Maxillofacial Surgery of the University Medical Center Groningen (UMCG), The Netherlands. The included patients are diagnosed with a tumor in the oral cavity region with a planned mandibulectomy or maxillectomy. Patients with implants in their mouth are excluded since this causes major artifacts in the MRI. The dataset is split into 80 scans for training and 20 scans for testing.
[0063] In one example the nnU-Net convolutional neural network is selected as the framework for automatic segmentation. nnU-Net has a contraction path and an expansion path (encoder-decoder), each with 5 stages as specified in more detail below with reference to FIG 9, 9B, 9C.
[0064] The model is implemented on an Intel® Core™ i9-10980XE CPU with 64 GB RAM and a graphic card of NVIDIA RTX A4000 (NVIDIA Corporation, USA) with 16 GB GDDR6 memory. 80 / 20 split for training and validation is used.
[0065] Training is done using the 3D full resolution configuration with 5-fold cross validation and 1000 epochs per fold using the cross-entropy dice loss function.
[0066] The following classification metrics are used to evaluate the performance of the model on the test dataset:
[0067] Dice:
[0068]
[0069] ixnyi
[0070] Intersection over Union: IoU= , , , ,
[0071] |x|u|r|
[0072] True positives
[0073] Precision: Precision = True positives+False positives
[0074] True positives
[0075] Recall: Recall = True positives+False negativesThe MRI derived bone models are matched to the original CT bone model using the global registration tool in 3-Matic. To compare the CT and MRI -based bone models, median deviation is calculated and distance maps generated using part-to-part comparison in 3-Matic. Only overlapping bone regions on both CT and MRI scans are considered. A deviation of <1 mm between the CT and MRI is deemed acceptable for clinical use.
[0076] It was found that the trained model achieves overall mean Dice similarity coefficient of 0.86 (SD ±0.03, range 0.78-0.91), Intersection over Union of 0.76 (SD ±0.05, range 0.64-0.84), precision of 0.86 (SD ±0.05, range 0.72-0.92) and recall of 0.86 (SD ±0.05, range 0.71-0.93) on the test dataset. A summary of the evaluation metrics for each anatomical structure on the test dataset is presented in FIGs 2A - 2D and the subsequent table.
[0077] Therein FIG 2A - 2D respectively shows the Median Dice scores, the median precision scores, the median intersection over union and the median recall scores per structure (Cranium CR, Mandible MD) and all structures combined (ALL).
[0078] The median deviation between the CT and MRI based bone models is 0.21 mm (IQR 0.05) for the mandible and is 0.30 mm (IQR 0.05) for the cranium
[0079]
[0080] Visual comparison between the MRI models and CT models is presented in FIGs 3 A, 3B and FIG 4A - 4D.
[0081] FIG 3A, 3B show for an example of an edentulous patient a deviation of an MRI based and a CT-based bone model. The color scale indicates that thedeviation is generally below 1 mm. Only exceptionally deviations up to 2 mm occur.
[0082] FIG 4A and 4C respectively show a front view and a side view of the bone model of the edentulous patient generated on the basis of a CT image. FIG 4B and 4D show a front view and a side view of the bone model of the edentulous patient generated on the basis of an MRI image using the trained neural network.
[0083] Likewise FIGs 5 A, 5B and FIG 6 A - 6D present a visual comparison between the MRI models and CT models for a dentulous patient.
[0084] FIG 5A, 5B shows the deviation of an MRI based and a CT-based bone model for the dentulous patient. The color scale indicates that the deviation is generally below 1 mm. Only exceptionally deviations up to 2 mm occur.
[0085] The most significant deviations between the CT and MRI-derived bone models occur in the dorsal part of the cranium, the mandibular condyles, and the teeth. The missing part of the cranium is likely due to differences in scan coverage. It could potentially be improved by manually adding this segmentation to the MRI or building a dataset where the CT fully covers the same area as the MRI.
[0086] Deviation in the condyles could be explained by the fact that it is difficult to differentiate the condyle from the glenoid fossa. The deviation in the teeth can be explained by several aspects. First, the segmentation of the teeth from the CT contains artefacts due to the presence of tooth fillings. Besides, teeth are relatively small and the original MRI scan has a voxel size of 0.9 mm, thus not capturing the entire tooth. Resampling the scan does not add this data to the scan. All three of these factors do not affect the clinical usability of the model, as these areas are not essential for designing surgical guides or patient-specific implants.
[0087] FIG 6A and 6C respectively show a front view and a side view of the bone model of the dentulous patient generated on the basis of a CT image. FIG 6B and 6D show a front view and a side view of the bone model of the dentulous patient generated on the basis of an MRI image using the trained neural network.FIG 7 further shows the median deviation of MRI bone models compared to CT bone model.
[0088] The inventive approach disclosed herein, although having been specifically demonstrated for generating bone models of the cranium and the mandible, can be analogously extended to other anatomical regions, provided paired CT and MRI scans are available. For example, the method can be applied to the knee joint, where the femur and tibia are distinct bones that can move relative to one another, or to the elbow joint, where the humerus and ulna may occupy slightly different positions between scans. Other examples may include the shoulder region, such as the humerus and scapula, or the hip joint, including the femur and pelvic bones. By applying the same principles to these regions, separate ground-truth segmentation datasets can be created for each bone, which may improve the accuracy of the neural network training and subsequent reconstruction for those anatomically distinct structures.
[0089] Of course, it will be understood that the principles described herein are not limited to pairs of bones but may also be applied where three or more bones are present and movable relative to each other. In such cases, a separate fusion step can be performed for each bone individually to generate a distinct, MRI-aligned ground-truth dataset for every relevant anatomical structure. For instance, in the wrist, bones such as the radius, ulna, and selected carpal bones may each be treated as separate regions of interest. Similarly, in the ankle and foot, individual fusions may be performed for the tibia, fibula, and talus. In the spine, vertebrae within a localized segment, such as the cervical vertebrae, can be independently aligned and segmented to provide accurate ground-truth data for each vertebral body.
[0090] By carrying out a separate fusion for each bone, whether two or more bones are involved, localized motion or positional changes between the scans can be more precisely accounted for. This may reduce registration errors that would otherwise occur if a single, global alignment were applied to a complex, articulated region. By generating distinct datasets for each movable bone, the resulting training data can better represent the true anatomy captured in the MRI, thereby supporting more accurate neural network predictions during laterMRI-only inference. This flexibility allows the inventive method to be used across a wide range of anatomical areas and clinical applications beyond the cranium and mandible, including joints, extremities, and other multi-bone structures.
[0091] FIG 8 schematically shows a use of the trained neural network 20. Therein an MRI-image Mn is obtained with an MRI imaging tool 10 and provided at the input of the trained neural network 20. In response thereto, the trained neural network 20 computes wrapped bone segmentation data Bn.
[0092] These in turn may be used as input for a 3D-printer 30 with which cutting guides may be printed to be used for milling replacements for affected body parts from the patient’s bone, e.g. from the fibula and / or for milling a patient specific implant from a bio compatible material. The MRI-image Mn is further provided to a contour analysis tool 40, with which the tumor margins and other soft tissue features can be accurately determined. In the example shown the 3D-printer 30 and the contour analysis tool 40 have a respective interface 35, 45 for facilitating human control.
[0093] It is noted that in case the MRI-image Mn is provided in a compressed format, such as Nii.gz, then the network output Bn is subjected to a decompression step. In the embodiment shown, this this decompression is performed by a decompression unit 25 that provides decompressed wrapped bone segmentation data Sn in a format suitable for driving the 3D printer or milling device 30, e.g. in STL format. Alternatively, the wrapped bone segmentation data serves as input for a virtual chirurgical planning wherein a medical specialist prepares the input for the milling device or 3D printer 30.
[0094] FIG 9 shows an exemplary U-type network as used in the present investigation. The U-type network shown in FIG 9 is configured to process input image data configured as 32 channels at a resolution of 128x128x128 (DxWxH). As noted above, the available image data Mrt is provided as D=384 x W=395 x H=383 voxels. As shown in FIG 9A, to overcome the difference in resolution, sections at the proper resolution of 128x128x128 are sampled by multiplexing element 20M to be provided to the neural network 20 and the output Ft is reconstructed from the output of the neural network 20 by a demultiplexing element 20D. Typically the sections to be processed are mutually partiallyoverlapping, e.g. by sampling the sections with a stride of 64x64x64. A cosine window or the like may be applied to avoid edge effects.
[0095] The exemplary network comprises an input stage CI, contraction stages C1-C5, expansion stages E1-E5 and an output stage CO.
[0096] As shown in FIG 9B, the contraction stages C1-C5 each subsequently comprise a first convolutional layer CL1, a first instance normalization layer INI, a first leaky ReLu layer RL1, a second convolutional layer CL2, a second instance normalization layer IN2, and a second leaky ReLu layer RL2.
[0097] In the contraction stages, the first and the second convolutional layer CL1, CL2 respectively have a stride of 2x2x2 and a stride of Ixlxl. Both convolutional layers have a 3x3x3 kernel. The first convolutional layer CL1 of each contraction stage provides for a down sampling as shown below.
[0098]
[0099] In the embodiment shown in FIG 1, the neural network also comprises an input stage CI that extracts features of the input data without downsampling. In one example the input stage CI is a single convolutional layer. In another embodiment as used in the present investigations the input stage CI has the same architecture as the contraction stages, as shown in FIG 9B, but contrary to the contraction stages its first convolutional layer CL1 has a stride of Ixlxl. The dimension of the data throughout the input stage CI is 32(Ch) x 128(D) x 128(W) x 128(H).
[0100] As shown in FIG 9C the expansion stages El to E5 each comprise a transverse convolutional layer T that spatially expands data from a first input, a concatenation layer CCAT that concatenates the output of the transverse convolutional layer and a convolutional unit C that processes the output of the concatenation layer. The transverse convolutional layer T in the embodiment ofthe neural network as used in the present investigation has a kernel size of 2x2x2 and a stride of 2x2x2. The transverse convolutional layer T provides for a doubling of spatial resolution and a reduction of the number of channels. In the stages El to E5 the dimensions of input and output of the transverse convolutional layer T is specified as follows:
[0101]
[0102] The convolutional unit C provides for a reduction of the number of channels. In the embodiment of the neural network in the present investigation the convolutional unit C has the same architecture as that of the contraction stages as shown in FIG 9B. However contrary to the contraction stages its first convolutional layer CL1 has a stride of Ixlxl. The number of channels is reduced by a factor of two, which may be said to compensate the effect of doubling the number of channels occurring in the concatenation.
[0103] As shown in FIG 9, the inputs of the expansion stages are coupled as follows.
[0104]
[0105] The neural network as used herein further has an output stage CO, herein provided as a convolutional layer with kernel size Ixlxl and stride Ixlxl. The output stage outputs the final bone segmentation map.Whereas in the examples presented above specific dimensions are mentioned, it will be clear that these are not limiting. Dependent on the case data dimensions may be selected differently. For example for applications wherein the region of interest is small it would suffice to have smaller image dimensions. On the other hand when more computation power comes available and more powerful medical imaging tools are developed it is likely that higher resolution image data can be processed. Some steps in the examples presented above are optional. For example, resampling the MRI-scan data or the CT-scan data is not necessary if they are obtained with the same resolution.
[0106] Also various modification of the neural network as present in FIG 9, 9B and 9C are possible depending on the level of accuracy that is required and the available computational facilities. For example it may be contemplated to configure the neural network to process data at an increased resolution, e.g. to 256x256x256 or 512x512x512. This may be combined with an increased number of stages. For example for each doubling of the resolution an additional contraction stage as shown in FIG 9B and an additional expansion stage as shown in FIG 90 may be added. Also the number of channels may be changed. Instead of two convolutional layers per stage C, one deeper convolution per block (using dilated convolutions or residual connections) could reduce complexity while preserving spatial awareness. As an alternative for the transpose convolution T in the expansion stages E, Using simple bilinear upsampling + convolution may simplify the network while maintaining accuracy. The instance normalization in the stages C is useful for particularly for small batch sizes. If training with larger batches, batch normalization or even skipping normalization might suffice. It may further be contemplated to replace the double set of layers in the contraction stages C with a single set of layers (one convolutional layer, optionally one instance normalization layer and one ReLu layer). Still further for the purpose of reducing computation cost standard convolutions could be replaced with depthwise separable convolutions (as in MobileNet). Various measures may be combined. For example it is an option to increase the resolution with which the neural net can process the input data to that of the image data Mrt, toobviate the segmentwise processing while still keeping the computational cost modest by one or more simplifying measures specified above.
[0107] In the claims the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. A single component or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.
Claims
CLAIMS1. A method for training a neural network (20) to automatically reconstruct a three-dimensional bone model from MRI-scan data (M), the model comprising at least a first bone part and a second bone part that are movable relative to each other, the method comprising for a plurality of human body part samples:preparing (SIM, SIC) a respective pair of an MRI-scan comprising MRI-scan data (M) and a CT-scan with CT-scan data (C);providing (S2C) ground truth bone segmentation data (Cs) for a first region of interest corresponding to the first bone part and a second region of interest corresponding to the second bone part in the CT-scan;linking (S4C) the ground truth bone segmentation data (Cs) with the CT-scan data to obtain enriched CT-scan data (Cl);fusing (S5) the MRI-scan data with the enriched CT-scan data (Cl) to obtain enriched MRI-scan data (F a, Fb) including the ground truth bone segmentation data;with the neural network to be trained, estimating (S7) from the MRI-scan data, tentative enriched MRI-scan data (Ft) comprising the MRI-scan data and tentative bone segmentation data;computing (S8) a loss (L) indicative for a deviation between the tentative bone segmentation data and the ground truth bone segmentation data; and backpropagating (S9) the loss (L);wherein the step of fusing (S5) the MRI-scan data with the enriched CT-scan data is performed separately for each region of interest, comprising:performing a first fusion with the region of interest set on the first bone part to obtain a first ground truth bone segmentation dataset (Fa) aligned with the coordinate system of the MRI-scan data; and performing a second fusion with the region of interest set on the second bone part to obtain a second ground truth bone segmentation dataset (Fb) aligned with the coordinate system of the MRI-scan data.
2. The method according to claim 1, wherein the MRI-scan data (M) is Tl- MRI-scan data.
3. The method according to claim 1 or 2, wherein the neural network is a U-type network, such as nnU-Net.
4. The method according to claim 1, 2 or 3, wherein the ground truth bone segmentation data (Cs) is provided (S2C) as STL-data.
5. The method according to one of the preceding claims, further comprising the step of wrapping (S3C) the segmentation data (Cs) to create a continuous surface and fill imperfections in a mesh of the segmentation data.
6. The method according to claim 5, wherein the step of wrapping is performed with a gap closing distance of substantially 3 mm and a detail level of substantially 0.5 mm.
7. The method according to one of the preceding claims, further comprising the step of resampling (S2M) the MRI-scan data (M) or the CT-scan data to match a common voxel size.
8. The method according to one of the preceding claims, further comprising the step of transforming (S3M, S6) the MRI-scan data and the enriched CT-scan data into image data according to a common image data format.
9. The method according to any one of the preceding claims, wherein the loss (L) is computed using a cross-entropy dice loss function.
10. Use of a trained neural network (20) obtained with the method according to one of the preceding claims, for computing bone segmentation data from MRI-scan data.
11. A computer-implemented method for automatically reconstructing a three-dimensional bone model of a region of a patient's body, said region comprising atleast a first bone part and a second bone part that are movable relative to each other, the method comprising:receiving MRI-scan data (Mn) of the region of the patient's body; processing, with a trained neural network (20), the received MRI-scan data to compute bone segmentation data, wherein the computed bone segmentation data comprises a first segmentation dataset corresponding to the first bone part and a second, separate segmentation dataset corresponding to the second bone part; andreconstructing the three-dimensional bone model from the first and second segmentation datasets, wherein the first and second bone parts are represented as distinct parts.
12. The method of claim 11, further comprising manufacturing, with a manufacturing device (30), a patient-specific implant or a surgical guide based on the reconstructed three-dimensional bone model.
13. The method of claim 11 or 12, wherein the reconstructed three-dimensional bone model is provided for use in virtual surgical planning (VSP).
14. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause a device to perform the method according to any of the preceding claims.
15. A system for automatically reconstructing a three-dimensional bone model of a region of a patient's body, said region comprising at least a first bone part and a second bone part that are movable relative to each other, the system comprising:an input interface (10) configured to receive MRI-scan data (Mn);a trained neural network (20) configured to process the received MRI-scan data and compute bone segmentation data comprising a first segmentation dataset corresponding to the first bone part and a second, separate segmentation dataset corresponding to the second bone part; anda reconstruction module configured to generate the three-dimensional bone model from the first and second segmentation datasets, wherein the first and second bone parts are represented as distinct parts.