Medical image conversion method and associated medical image 3D model personalization method

The method transforms X-ray images into local DRRs using GANs and CNNs to separate anatomical structures, addressing the issue of overlapping organs in DRRs, ensuring accurate 3D model personalization and improved image alignment.

JP7837885B2Active Publication Date: 2026-03-31EOS IMAGING SA
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-05-13
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing medical image conversion methods using generative adversarial networks (GANs) and convolutional neural networks (CNNs) fail to effectively separate and distinguish multiple anatomical structures in digitally reconstructed radiographic (DRR) images due to overlapping organ structures, leading to loss of useful information and unsatisfactory results.

Method used

A method using a single generative adversarial network (GAN) or a group of CNNs to transform actual X-ray images into local DRRs, each focusing on a unique anatomical structure, optimizing both conversion and structural separation simultaneously, enabling separate representation of different organs or parts of organs without losing useful information.

Benefits of technology

The method achieves effective separation of overlapping anatomical structures in DRRs, preserving image quality and enabling accurate 3D model personalization by generating patient-specific models from generic models, improving alignment and reducing complexity in image registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837885000025
    Figure 0007837885000025
  • Figure 0007837885000026
    Figure 0007837885000026
  • Figure 0007837885000027
    Figure 0007837885000027
Patent Text Reader

Abstract

The present invention relates to a medical image conversion method for automatically converting at least one or more actual X-ray images (16) of a patient, including at least a first anatomical structure (21) of the patient and a second anatomical structure (22) of the patient, into at least one digitally reconstructed radiographic image (23) of the patient, the at least one X-ray image (16) representing the first anatomical structure (24) without representing the second anatomical structure (26), by a single operation using either one convolutional neural network (CNN) or a group of convolutional neural networks (CNN) (27) pre-trained to both, or simultaneously, distinguish the first anatomical structure (21) from the second anatomical structure (22) and convert the actual X-ray image (16) into at least one digitally reconstructed radiographic image (23).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a medical image conversion method.

[0002] The present invention also relates to an associated medical image 3D model personalization method that uses such a medical image conversion method.

Background Art

[0003] According to the prior art disclosed in US Patent Application Publication No. 2019 / 0259153, regarding a medical image conversion method, a method of transforming a patient's actual X-ray image into a digital reconstructed radiography (DRR) image by a generative adversarial network (GAN) is known.

[0004] However, the obtained digital reconstructed radiography (DRR) image has the following two characteristics. ● First, it includes all the organs that were present in the starting X-ray image before transformation by the generative adversarial network (GAN). ● Second, when pixel segmentation of a specific organ is required, this is performed from the obtained digital reconstructed radiography (DRR) image in a subsequent segmentation step by another dedicated convolutional neural network (CNN) [see page 1, section 5, and page 4, section 51 of US Patent Application Publication No. 2019 / 0259153].

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0006]

Non-Patent Document 1

[0007] According to the present invention, this is considered to be quite complex and not sufficiently effective because there are several, often many, different anatomical structures within organs or areas of a patient's body.

[0008] Different anatomical structures, such as organs, parts of organs, or groups of organs, are at risk of overlap in digitally reconstructed radiographic images (DRRs) because they represent a planar projection view of the 3D structure. [Means for solving the problem]

[0009] The object of the present invention is to at least partially mitigate the aforementioned drawbacks.

[0010] In contrast to the cited prior art, according to the spirit of the present invention, in a dual DRR image matching process, it is considered beneficial for each DRR to represent a unique organ or anatomical part of an organ in order to avoid mismatches between adjacent or overlapping organs or organ structures.

[0011] However, in conventional techniques, pixel classification resulting from segmentation of globally acquired DRRs that include all organs is insufficient to distinguish the respective amounts of DRR image signals belonging to each organ or organ structure, and the deformation process does not enable this structural separation of overlapping organ regions. Therefore, subsequent extraction of local DRRs, each representing only one organ, from globally acquired DRRs is impossible without at least the loss of useful signals.

[0012] Therefore, according to the present invention, a method is proposed for transforming an actual X-ray image of a patient into either a single local digitally reconstructed radiographic image (DRR) or a set of local digitally reconstructed radiographic images (DRRs) using a single generative adversarial network (GAN). ●Through a single, unique operation of a GAN, ●Trained from DRR generated by directly separating the structure in the original 3D volume, ● The process involves simultaneously generating multiple sets of DRRs, each of which focuses on only one anatomical structure of the subject, or generating one DRR that excludes other anatomical structures and focuses only on one anatomical structure of the subject. Both are performed, and therefore both the conversion function and the structural separation function are optimized simultaneously.

[0013] This objective is achieved using a medical image transformation method that automatically transforms at least one or more actual X-ray images of a patient, including at least one first anatomical structure and the second anatomical structure of the patient, into at least one digitally reconstructed radiographic image (DRR) of the patient that represents the first anatomical structure without representing the second anatomical structure, by a single operation using either one convolutional neural network (CNN) or a group of convolutional neural networks (CNNs) pre-trained to both, or simultaneously, distinguish the first anatomical structure from the second anatomical structure and transform the actual X-ray images into at least one digitally reconstructed radiographic image (DRR) of the patient.

[0014] Preferably, the medical image transformation method according to the present invention also automatically transforms the actual X-ray image of the patient into at least another digitally reconstructed radiographic image (DRR) of the patient that represents the second anatomical structure without representing the first anatomical structure, and by the same single operation, one convolutional neural network (CNN), or a group of convolutional neural networks (CNNs), is pre-trained to both, or simultaneously, distinguish the first anatomical structure from the second anatomical structure and transform the actual X-ray image into at least two digitally reconstructed radiographic images (DRR).

[0015] This means that from a single global X-ray image representing several different organs, or several different parts of an organ, or several groups of organs, several local DRR images corresponding to and representing each of these different organs, or several different parts of an organ, or several groups of organs, are obtained by a single convolutional neural network, or a group of convolutional neural networks linked to it, through a single operation of transformation that includes both transformation and structural separation steps.

[0016] This enables each different organ, or each different part of an organ, or each group of organs, to be represented separately from all other different organs, or all other different parts of an organ, or all other groups of organs, without losing useful information in areas where there is some overlap between some of these different organs, or some different parts of an organ, or some groups of organs, or where there is some overlap in the actual X-ray images, or where there is some overlap in the converted global DRRs.

[0017] In fact, when a global DRR is obtained from one or more actual X-ray images, on the DRR, i.e., on the image in this converted DRR domain, several different overlapping organs, or several different parts of an organ, or several groups of organs cannot be separated and are all mixed together such that any given pixel is a mixture of different contributions from these different overlapping organs, or different parts of an organ, or different groups of organs, and there is no way, or at least it is very difficult, to distinguish these different contributions later, which in any case leads to unsatisfactory results. Therefore, the overlap between several different organs, or several different parts of an organ, or several groups of organs cannot be restored without losing useful information, and these several different overlapping organs, or several different parts of an organ, or several groups of organs cannot be separated from each other without losing useful information.

[0018] Thus, this implementation of the present invention not only converts an image from a first domain, e.g., an X-ray image, to a second domain, e.g., a DRR, but also performs a unique, simple, and more effective method to separate several different overlapping organs, or several different parts of an organ, or several groups of organs without losing useful information, i.e., without significantly degrading the original image quality.

[0019] Some of these different organs, or some different parts of an organ, or some groups of organs are examples of different anatomical structures.

[0020] This objective is also achieved by using a single operation that performs both, or simultaneously, differentiating the first anatomical structure from the second anatomical structure and converting an actual X-ray image into at least two digital reconstructed radiography (DRR) images, using either a single pre-trained convolutional neural network (CNN), or a group of convolutional neural networks (CNNs), to automatically convert at least one or more actual X-ray images of the patient, including at least the first anatomical structure of the patient and the second anatomical structure of the patient, into at least the first and second digital reconstructed radiography (DRR) images of the patient, where the first digital reconstructed radiography (DRR) image represents the first anatomical structure without representing the second anatomical structure, and the second digital reconstructed radiography (DRR) image represents the second anatomical structure without representing the first anatomical structure, using a medical image conversion method.

[0021] Another similar objective is achieved by using a single operation that performs both, or simultaneously, differentiating the first anatomical structure from the second anatomical structure and converting an image in a first domain into at least one image in a second domain, using either a single pre-trained convolutional neural network (CNN), or a group of convolutional neural networks (CNNs), to automatically convert at least one or more images of the patient in a first domain, including at least the first anatomical structure of the patient and the second anatomical structure of the patient, into at least an image of the patient in a second domain that represents the first anatomical structure without representing the second anatomical structure, using a medical image conversion method.

[0022] Another similar objective is achieved by a medical image transformation method that automatically transforms at least one or more global images of a patient in a first domain, which include at least several different anatomical structures of the patient, into several local images of the patient in a second domain, each representing one of the different anatomical structures, by a single operation using either one or a group of convolutional neural networks (CNNs) pre-trained to both distinguish the anatomical structures from one another and transform images in a first domain into at least one image in a second domain, or to do so simultaneously.

[0023] Another complementary object of the present invention is related to the object of the present invention cited above, as it is a preferred application of the present invention. In fact, the present invention also relates to an associated medical image 3D model personalization method that uses and utilizes a medical image transformation method, which is the primary object of the present invention, the method particularly utilizes the ability to separate different anatomical structures from one another without losing useful information because there may be overlap between these different anatomical structures.

[0024] A complementary object of the present invention is a medical image 3D model personalization method comprising a medical image conversion method according to the present invention, wherein a 3D generic model is used to generate at least one or more digitally reconstructed radiographic images (DRRs) of a front view of the patient, each representing one or more anatomical structures of the patient, and at least one or more digitally reconstructed radiographic images (DRRs) of a side view of the patient, each representing one or more anatomical structures of the patient, wherein an actual frontal X-ray image is converted by the medical image conversion method to at least one or more digitally reconstructed radiographic images (DRRs) of a front view of the patient, each representing one or more anatomical structures of the patient, and the actual side X-ray image is converted by the medical image conversion method to at least one or more digitally reconstructed radiographic images (DRRs) of a side view of the patient, each representing one or more anatomical structures of the patient, and the 3D generic model is used to generate at least one or more digitally reconstructed radiographic images (DRRs) of a front view of the patient, each representing one or more anatomical structures of the patient, and the 3D generic model is used to generate front view of the patient, each representing one or more anatomical structures of the patient, and the 3D generic model is used to generate at least one digitally reconstructed radiographic images (DRRs) of a front view of the patient, each representing one or more anatomical structures of the patient, and the 3D generic model is used to generate at least one digitally reconstructed radiographic A method for generating a 3D patient-specific model from a generic model, wherein at least one digitally reconstructed radiographic image (DRR) of the front view of the patient, each representing one or more anatomical structures of the patient and obtained from the 3D generic model, is mapped to at least one digitally reconstructed radiographic image (DRR) of the front view of the patient, each representing one or more anatomical structures of the patient and obtained from the actual front X-ray image, and at least one digitally reconstructed radiographic image (DRR) of the side view of the patient, each representing one or more anatomical structures of the patient and obtained from the 3D generic model, is mapped to at least one digitally reconstructed radiographic image (DRR) of the side view of the patient, each representing one or more anatomical structures of the patient and obtained from the actual side X-ray image.

[0025] Mapping can be performed via elastic mapping or via elastic registering.

[0026] Preferably, the 3D generic model is a surface or volume representing a previous shape that can be deformed using a deformation algorithm.

[0027] Preferably, the process of modifying a generic model to obtain a personalized model tailored to a patient is called a modification algorithm. The relationship between the generic model and the modification algorithm is a modifiable model.

[0028] Preferably, the deformable model is a statistical shape model (SSM).

[0029] A statistical shape model is a deformable model based on statistics extracted from a training database to capture deformation patterns of shapes. Statistical shape models can directly infer plausible shape instances from a reduced set of parameters.

[0030] Therefore, the reconstructed patient-specific 3D model Because it was obtained from a 3D generic model or deformable model and two simple orthogonal real X-ray images, ●It is simpler than any model obtained from computed tomography (CT) or magnetic resonance imaging (MRI), In some cases, this is not only related to 3D generic models, or deformable models, or statistical shape models, but also implicitly and initially includes volume information from the beginning for the front and side starting views that are being transformed. ● More specific, more accurate, Different anatomical structures can be separated from one another without losing useful information, in contrast to the conversion methods used in the cited prior art, which are also more complex to implement.

[0031] The prior art previously cited in U.S. Patent Application Publication No. 2019 / 0259153 is in a completely different field of application and is based on performing "post-" segmentation using a dense inter-image network of task-driven generative adversarial networks (GANs) that pre-convert global X-ray images encompassing several anatomical structures into DRRs encompassing the same multiple anatomical structures. Therefore, it does not disclose complementary objects such as 3D / 2D elastic registration as present in the invention, and in particular does not disclose cross-domain image similarity of cross-multi-structure 3D / 2D elastic registration.

[0032] The robustness of intensity-based 3D / 2D registration of a 3D model to a set of 2D X-ray images also depends on the quality of the image correspondence between the actual images and the digitally reconstructed radiographic images (DRRs) generated from the 3D models. In this multimodal registration situation, the tendency to improve the level of similarity between both images is to produce DRRs that are as realistic as possible, i.e., as close as possible to the actual X-ray images. This involves two important aspects: firstly, having 3D models of soft tissues and bones enhanced with accurate density information of all relevant structures; and secondly, a sophisticated projection process that takes into account the physical phenomena of interaction with matter.

[0033] Conversely, in the method proposed by embodiments of the present invention, the opposite approach of bringing actual X-ray images into the DRR image domain is performed using cross-modality image-to-image conversion. By adding a step prior to image-to-image conversion based on a GAN pix-to-pix model, a simple and fast DRR projection process can be used without simulating complex phenomena. In fact, even with simple metrics, the measurement of similarity becomes efficient because both matching images belong to the same domain and contain essentially the same kind of information. The proposed medical image conversion method also addresses the well-known problem of aligning objects in a scene composed of multiple objects. By using separate convolutional neural network (CNN) output channels for each structure in the XRAY to DRR converter, it becomes possible to separate overlapping adjacent objects from each other and avoid mismatches of similar structures in alignment. The method proposed by embodiments of the present invention is applied to the challenging 3D / 2D elastic alignment of a vertebral 3D model in two-planar radiographic images of the spine. Using the XRAY to DRR conversion step improves alignment results and reduces reliance on the choice of similarity measurement, as it converts multimodal alignment to monomodal alignment.

[0034] Preferred embodiments include one or more of the following features, which can be employed separately or together, in part or in whole, with any of the objects of the invention cited above.

[0035] Preferably, the one convolutional neural network (CNN) or group of convolutional neural networks (CNNs) is a single generative adversarial network (GAN).

[0036] Therefore, the simplicity of the proposed transformation method can be further improved while maintaining its effectiveness.

[0037] Preferably, the single generative adversarial network (GAN) is a U-Net GAN or a Residual-Net GAN.

[0038] Preferably, the actual X-ray image of the patient is a direct capture of the patient by an X-ray imaging device.

[0039] Therefore, the images in the first domain being converted are completely easy to obtain and completely patient-specific.

[0040] Preferably, the first and second anatomical structures of the patient are anatomical structures that are adjacent to each other on the actual X-ray image, or even adjacent to each other on the actual X-ray image, or even in contact with each other on the actual X-ray image, or at least partially superimposed on the actual X-ray image.

[0041] Therefore, the proposed transformation method is even more interesting than the fact that the various anatomical structures that are separated from each other are actually in close proximity to each other, because it increases the amount of mixed information that can no longer be separated from each other.

[0042] Preferably, only one of the at least two digitally reconstructed radiographic images (DRRs) can be used for further processing.

[0043] Therefore, the proposed conversion method can be used even if the practitioner actually only requires one DRR.

[0044] Preferably, all of the at least two digitally reconstructed radiographic images (DRRs) can be used for further processing.

[0045] Therefore, the proposed conversion method can be used, especially when all DRRs are actually useful to the executor.

[0046] Preferably, the actual X-ray image of the patient includes at least three anatomical structures of the patient, preferably only three anatomical structures of the patient, and the actual X-ray image is converted by the single operation into at least three separate digitally reconstructed radiographic images (DRRs), each representing at least three anatomical structures, and each of the digitally reconstructed radiographic images (DRRs) represents only one of the anatomical structures, without representing any other of the anatomical structures, preferably into three separate digitally reconstructed radiographic images (DRRs), each representing only the three anatomical structures.

[0047] Therefore, the proposed transformation method is optimized in the following two ways. ●Firstly, using the two closest adjacent structures (especially linear structures like the patient's spine) helps to better isolate and individualize the specific anatomical structure of the subject. ●Secondly, using only the two closest neighbors still helps in doing so using a relatively simple implementation form.

[0048] Preferably, one convolutional neural network (CNN) or group of convolutional neural networks (CNNs) is pre-trained by a set of training groups of at least one or more corresponding digitally reconstructed radiographic images (DRRs), each representing only one of the anatomical structures but not other anatomical structures of the patient.

[0049] Therefore, the proposed transformation method can be trained using pairs of images, one in the first domain for transformation and the other in the second domain for transformation.

[0050] Preferably, one convolutional neural network (CNN) or group of convolutional neural networks (CNNs) is pre-trained with actual X-ray images and a set of several subsets of at least one or more digitally reconstructed radiographic images (DRRs), each of which represents only one of the anatomical structures but not other anatomical structures of the patient.

[0051] Therefore, the proposed transformation method can be trained using unpaired images in the first domain for transformation and in the second domain for transformation.

[0052] Preferably, one convolutional neural network (CNN) or group of convolutional neural networks (CNNs) is pre-trained by a set of training groups of at least one subset of both one frontal and one lateral actual radiographic image, as well as the corresponding frontal and lateral digitally reconstructed radiographic images (DRRs), where each subset represents only one of the anatomical structures but not the other anatomical structures of the patient.

[0053] Therefore, the proposed transformation method can be trained using groups of both frontal and lateral image pairs in a first domain for transformation and in a second domain for transformation, where each frontal-to-lateral image pair is transformed into several pairs of frontal-to-lateral images, and each transformed pair of frontal-to-lateral images corresponds to one anatomical part separated from other anatomical parts.

[0054] Preferably, the digitally reconstructed radiographic image (DRR) of the training group is obtained from a patient-specific 3D model through adaptation to two actual X-ray images taken along two orthogonal directions.

[0055] Therefore, the compromise between the simplicity of the model used to create the training DRR is optimized, on the one hand, and the effectiveness of the model through precise dedication to the specific patients being targeted.

[0056] Preferably, the different anatomical structure of the patient is the adjacent vertebra of the patient.

[0057] Therefore, the proposed transformation method is even more interesting than the fact that the various anatomical structures that are separated from each other are actually in close proximity to each other, because it increases the amount of mixed information that can no longer be separated from each other.

[0058] Preferably, the different adjacent vertebrae of the patient are located within the same single region of the patient's spine, either in the region of the patient's upper thoracic spinal segment, or the region of the patient's lower thoracic spinal segment, or the region of the patient's lumbar spine segment, or the region of the patient's cervical spine segment, or the region of the patient's pelvis.

[0059] Preferably, the different anatomical structures of the patient are located within a single, identical region of the patient, such as the lumbar region, the lower limb region of the patient (including the femur or tibia), the knee region of the patient, the shoulder region of the patient, or the rib cage region of the patient.

[0060] Preferably, each of the different digitally reconstructed radiographic images (DRRs) representing the different anatomical structures of the patient simultaneously includes an image having pixels representing different gray levels and at least one tag representing anatomical information related to the anatomical structure represented by the tag.

[0061] Therefore, this tag is used to distinguish between different anatomical structures and to separate these different anatomical structures from one another. This tag is a simple and effective method when it is still possible to separate the information corresponding to each of the different anatomical structures, namely, in a 3D volume, it is preferable that these images in the first domain, here in the case of X-ray images, preferably frontal and lateral X-ray images, are transformed in the second domain, here in the case of DRR, and if such distinction has not been made previously, it would no longer be possible to distinguish between these different anatomical structures and to separate these different anatomical structures from one another.

[0062] Preferably, the image is a square image with dimensions of 256 x 256 pixels.

[0063] Therefore, this represents a suitable compromise between image quality on the one hand and processing complexity and required storage capacity on the other.

[0064] Preferably, one of the convolutional neural networks (CNNs) or a group of convolutional neural networks (CNNs) is pre-trained on actual X-ray images, i.e., both X-ray images and deformations of actual X-ray images, of a diverse range of patients, ranging from 100 to 1000, preferably from 300 to 700, and more preferably about 500.

[0065] Therefore, this is a simple and effective way to provide a very large number of well-distinguished training images for training, while at the same time being able to freely use a very limited amount of data to construct these training images.

[0066] There is an optimal range for the number of training images; in fact, using a reasonable range of training images yields nearly optimal results at a very reasonable cost.

[0067] Preferably, at least both the actual frontal X-ray image of the patient and the actual lateral X-ray image of the patient are converted, and both X-ray images each include the same anatomical structures of the patient.

[0068] Therefore, for example, with the help of a 3D generic model, or a deformable model, or possibly a statistical shape model, anatomical structures can be more accurately understood from two orthogonal views, and in some cases even reconstructed in three-dimensional space (3D). By using both frontal and lateral actual X-ray images as the images to be transformed, it becomes possible to obtain frontal and lateral DRRs and to improve the accuracy of the resulting DRRs.

[0069] Further features and advantages of the present invention will become apparent from the following description of embodiments of the invention, given as non-limiting examples, with reference to the accompanying drawings listed below. [Brief explanation of the drawing]

[0070] [Figure 1A] This figure shows the first step in creating a global DRR for training the GAN during the training phase, using conventional technology. [Figure 1B] This figure shows the second step in conventional techniques, in which segmentation is performed to separate anatomical structures from one another during the training phase. [Figure 1C] This figure shows a conventional method for converting X-ray images into segmented images of different anatomical structures. [Figure 2A] This figure shows a first step according to one embodiment of the present invention, in which a local DRR of a first anatomical structure is created for training a GAN during the training phase. [Figure 2B] This figure shows a second step according to one embodiment of the present invention, in which, during the training phase, another local DRR of a second anatomical structure is created for training a GAN. [Figure 2C] This figure shows a method according to one embodiment of the present invention for converting an X-ray image into several different local DRRs, each representing several different anatomical structures. [Figure 3A] This figure shows some examples of 3D / 2D alignment using conventional techniques. [Figure 3B] This figure shows an example of 3D / 2D alignment according to an embodiment of the present invention. [Figure 4A] This figure shows some examples of 3D / 2D alignment using conventional techniques. [Figure 4B] This figure shows an example of 3D / 2D alignment according to an embodiment of the present invention. [Figure 5] This figure shows a global flowchart illustrating an example of a 3D / 2D alignment method according to an embodiment of the present invention. [Figure 6] This figure shows a more detailed flowchart of an example of a training stage before performing a 3D / 2D alignment method according to an embodiment of the present invention. [Figure 7] This figure shows a more detailed flowchart of an example of performing a 3D / 2D alignment method according to an embodiment of the present invention. [Figure 8A] This figure shows an example of the principle of creating the inner layer from the outer layer using the previously defined cortical thickness and normal vector, according to an embodiment of the present invention. [Figure 8B] This figure shows an example of an interpolated thickness map of the entire mesh of the L1 vertebra according to an embodiment of the present invention. [Figure 9] This figure shows an example of the principle of raycasting through a two-layer mesh. [Figure 10] This figure shows an example of a GAN network architecture and training from XRAY to DRR according to an embodiment of the present invention. [Figure 11A] This figure shows an example of GAN-based XRAY to DRR conversion, where an actual X-ray image is converted to a 3-channel DRR image corresponding to the adjacent vertebral structure. [Figure 11B]This figure shows the transformation, also known as the conversion from X-ray domains to DRR domains, under adverse conditions, to demonstrate the robustness of the method proposed by the embodiments of the present invention. [Figure 12A] This figure shows the target alignment error as a function of the initial pose shift on the vertical axis. [Figure 12B] This figure shows an example where a vertebra has been moved vertically upward. [Figure 13A] This figure shows the anatomical regions used to calculate the statistical distance from the node to the surface. [Figure 13B] This figure shows the distance map of the maximum error calculated for the L3 vertebra. [Figure 14A] This figure shows the cost values ​​of PA views and LAT views using GAN DRR compared to actual X-ray images. [Figure 14B] This figure shows the similarity between PA views and LAT views using GAN DRR compared to actual X-ray images. [Figure 14C] This figure shows the better fit observed when using DRR generated by a GAN compared to Figure 14D, which uses actual X-ray images. [Figure 15] This figure shows similar results using the GNCC (Gradient Normalized Cross-Correlation) metric, similar to Figure 14B. [Figure 16] This figure shows the results of the target alignment error (TRE) test in the Z direction (vertical image direction). [Modes for carrying out the invention]

[0071] All subsequent explanations will refer to bodily representations that have several different organs, but the same approach can be applied to bodily representations that have several different groups of organs, or even organ representations that have several different parts of these organs.

[0072] The rear-front (PA) view is the same as the front view (FRT), and the side view follows a direction perpendicular to the common direction of the rear-front and front views, so both are symmetrical to the side view (LAT).

[0073] Figure 1A shows the first step in the conventional technique of creating a global DRR for training the GAN during the training phase.

[0074] Raycasting is performed from DRR X-ray source 1 through 3D volume 2, which contains two different organs, the first organ 3 and the second organ 4.

[0075] The first organ 3 and the second organ 4 have planar projections on the global DRR 5, which are the contributions of the DRR image signal 6 from the first organ 3 and the DRR image signal 7 from the second organ 4, respectively.

[0076] Image regions 6 and 7 have a crossover zone 8 where such signals from organs 3 and 4 are superimposed.

[0077] In this crossover zone 8, it is impossible to know which part of the signal originates from organ 3 and which part originates from organ 4.

[0078] Figure 1B shows a second step in which segmentation is performed to separate anatomical structures from one another during the conventional training phase.

[0079] Segmentation of the first organ 3 and the second organ 4 is performed so that they are separated from each other by the DRR X-ray source 1.

[0080] The segmentation of the first organ 3 and the second organ 4 is given by the first label image 12 and the second label image 13, respectively, with the first trace 14 of the first organ 3 and the second trace 15 of the second organ 4 on top of them, respectively.

[0081] However, even if these segmentation traces 14 and 15 are later applied to the global DRR 5, the mixing of signals within the crossover zone 8 will not allow for separation between their respective original contributions, namely the contribution from the first organ 3 and the contribution from the second organ 4, and therefore they will not become the first DRR of the first organ 3 and the second DRR of the second organ 4.

[0082] Therefore, in all crossover zones 8, it becomes impossible to separate the mixed signals from several different organs into their original contributions, and applying segmentation to the global DRR does not result in several different local DRRs from different organs. Useful signals that are irrecoverable or at least very difficult to recover are lost.

[0083] Figure 1C shows a conventional method for converting X-ray images into segmented images of different anatomical structures.

[0084] The X-ray image 16 is transformed into a transformed global DRR5 by GAN 17, and then segmented by DI2I CNN 18 as follows: ●The first binary mask 12 of this transformed global DRR5 corresponds to the segmentation 14 of the first organ 3, ●The second binary mask 13 of this transformed global DRR5, corresponding to segmentation 15 of the second organ 4. Separately acquiring local DRR images corresponding to the first organ 3 or the second organ 4 from the first binary mask 12 or the second binary mask 13 would inevitably result in a loss of useful signals in the crossover zone 8.

[0085] Figure 2A shows a first step according to one embodiment of the present invention, in which a local DRR of a first anatomical structure is created for training a GAN during the training phase.

[0086] Raycasting is performed from DRR X-ray source 1 through a 3D volume 20 that contains only the first organ 21, which has already been separated from the second organ 22, which is different from the first organ 21.

[0087] The first organ 21 has a planar projection on the local DRR, and each DRR image 23 is dedicated to the first organ 21, representing a DRR image 24 that accurately corresponds to the planar projection of the first organ 21 without loss of useful signals.

[0088] In the DRR image 24, all signals originate from the first organ 21, and no signals originate from the second organ 22. Useful signals related to the planar projection of the first organ 21 are not lost.

[0089] Figure 2B shows a second step according to one embodiment of the present invention, in which, during the training phase, another local DRR of a second anatomical structure is created for training the GAN.

[0090] Raycasting is performed from DRR X-ray source 1 through a 3D volume 20 that contains only the second organ 22, which has already been separated from the first organ 21 and is distinct from the second organ 22.

[0091] The second organ 22 has a planar projection on the local DRR, and each DRR image 25 is dedicated to the second organ 22, representing a DRR image 26 that accurately corresponds to the planar projection of the second organ 22 without loss of useful signals.

[0092] In the DRR image 26, all signals originate from the second organ 22, and no signals originate from the first organ 21. Useful signals related to the planar projection of the second organ 22 are not lost.

[0093] Figure 2C shows a method according to one embodiment of the present invention to convert an X-ray image into several different local DRRs, each representing several different anatomical structures.

[0094] The X-ray image 16 is transformed by GAN27 into two separate transformed local DRRs 23 and 25, which correspond to the following: ●In contrast to image 12 in Figure 1C, the first local DRR image accurately corresponds to the planar projection of the first organ 21 without losing any useful signals related to the first organ 21. ●In contrast to image 13 in Figure 1C, a second local DRR image that accurately corresponds to the planar projection of the second organ 22 without losing useful signals related to the second organ 22.

[0095] Strength-based elastic registration of a 3D model to a 2D planar image is one of the methods used as a crucial step in the 3D reconstruction process. This type of registration is similar to multimodal registration and relies on maximizing the similarity between the actual X-ray image and the digitally reconstructed radiographic image (DRR) generated from the 3D model. Typically, 3D models do not contain all the information, for example, density is not always included, and the two images, such as the actual X-ray and the DRR, can differ significantly from each other, making the optimization process very complex and often unreliable. A standard solution is to bring the DRR generation process as close as possible to the image formation process by adding additional information and simulating complex physical phenomena. In the algorithm proposed by embodiments of the present invention, the reverse approach is used to facilitate image matching by transforming the actual X-ray image into a DRR-like image.

[0096] An image transformation step has been proposed that transforms X-ray images into DRR-like images, enabling the removal of background and noise from the original image and creating a DRR domain. Because the images have similar properties, matching between the two images becomes simpler and more efficient, improving the optimization process even when using standard similarity metrics.

[0097] This is achieved by using pix-to-pix deep neural network training based on a U-Net convolutional neural network (CNN), which enables the transformation of images from one domain to another. Using X-ray image deformation such as DRR facilitates mesh 3D / 2D alignment. Furthermore, the proposed CNN-based converter can separate adjacent bone structures by outputting one transformed image per bone structure, in order to better handle articulated bone structures.

[0098] Embodiments of the present invention are applied to 3D, which challenges the reconstruction of vertebral structures from two-plane X-ray imaging modalities. Vertebral structures are periodic multistructures that can exhibit anatomical deformation. A pix-to-pix network converts actual X-ray images into virtual X-ray images. This makes the following possible: ●Firstly, improve the correspondence and alignment performance between images. Secondly, in order to avoid inconsistencies, specific structures are identified and separated from similar adjacent structures.

[0099] The principle of this 3D / 2D alignment is presented by comparing the prior art shown in Figure 3A with an embodiment of the present invention shown in Figure 3B. Figure 3A shows several examples of 3D / 2D alignment according to the prior art. Figure 3B shows an example of 3D / 2D alignment according to an embodiment of the present invention.

[0100] Figure 3A shows several classical methods using prior art for 3D / 2D registration that utilize the similarity between DRR generated from a mesh and X-ray images. Unfortunately, it is not easy to distinguish the target vertebra 31 from the rest of the vertebra 30.

[0101] In Figure 3B, the alignment process proposed by the embodiment of the present invention uses a prior image-to-image conversion to transform the target image in order to facilitate image correspondence in alignment-like functions and to avoid discrepancies in adjacent structures. Here, the target vertebra 31 is separated from the rest of the vertebra 30.

[0102] The application of this principle of 3D / 2D alignment is also demonstrated by a comparison between the prior art shown in Figure 4A and embodiments of the present invention shown in Figure 4B, with respect to the front and side planar projections of a 3D model. Figure 4A shows some examples of 3D / 2D alignment using the prior art. Figure 4B shows an example of 3D / 2D alignment according to embodiments of the present invention.

[0103] Figure 4A illustrates one classic method for 3D / 2D alignment of a 3D model using the similarity between DRR projections generated from a 3D mesh and two planar X-ray images, based on several prior art techniques. The 3D model 40 is positioned between a first frontal X-ray source 41 and a planar X-ray image 43 in which the frontal projection of the 3D model 40 is imaged. Unfortunately, the frontal image of the vertebra 47 in question is confused with the rest of the vertebra 45. The 3D model 40 is also positioned between a second lateral X-ray source 42 and a planar X-ray image 44 in which the lateral projection of the 3D model 40 is imaged. Unfortunately, the lateral image of the vertebra 48 in question is confused with the rest of the vertebra 46.

[0104] Intensity-based methods, or icon registration, aim to find the optimal model parameters that maximize the similarity between an X-ray image and a digitally reconstructed radiographic image (DRR) generated from a 3D model. They are more robust and flexible, requiring no image extraction. However, they involve using complex models and algorithms to generate DRRs that are as realistic as possible—i.e., as close as possible to actual X-ray images—and to perform efficient similarity measurements. Because they are optimized criteria in the registration process, they must reflect the actual degree of structural agreement in both the changing image and the target image, and be robust to perturbations to avoid trapping at minimums. The main perturbation is that the modalities of both matching images differ, even if the DRR generation appears fairly realistic. To compare an image from domain A (e.g., an actual X-ray image) with an image from domain B (e.g., a DRR image), sophisticated similarity metrics are required. These domain differences limit the performance of model registration on real-world clinical data. Alignment in planar X-rays has another additional peculiarity: the presence of overlapping adjacent structures in the environment, which leads to mismatches, especially when the initial 3D model is not close enough, as seen in Figure 4A.

[0105] Figure 4B shows the method proposed according to an embodiment of the present invention. In the step prior to the proposed XRAY-to-DRR conversion, the X-ray target image is converted to a DRR having a single structure to improve the level of similarity and prevent mismatches of adjacent structures in the alignment process. The 3D model 50 is positioned between the first frontal X-ray source 41 and the planar DRR converted image 51 in which the frontal projection of the 3D model 40 is imaged. Here, it is much simpler to perform the superposition of the frontal projection 55 of the 3D model 40 of the vertebra of interest, so that the exact same vertebra of interest 53 remains intact and is separated from the original X-ray image converted to DRR (the rest of the spine, i.e., other adjacent vertebrae, are removed). The 3D model 50 is also positioned between the second lateral X-ray source 42 and the planar DRR converted image 52 in which the lateral projection of the 3D model 40 is imaged. Here again, it is far simpler to perform the superposition of a lateral projection 56 of the 3D model 40 of the target vertebra, leaving the exact same target vertebra 54 intact and separated from the original X-ray image converted to DRR (the rest of the spine, i.e., other adjacent vertebrae, are removed).

[0106] The method proposed by the embodiment of the present invention shown in Figure 4B performs the opposite approach, thereby changing the paradigm. In fact, instead of improving either the similarity metric or the DRR X-ray simulation, we explore a method to convert X-ray images to DRR-like images using a cross-modality image-to-image conversion model based on a pix-to-pix generative adversarial network (GAN). The XRAY-to-DRR conversion simplifies the entire process and allows the use of both simpler DRR generation and similarity metrics, as shown in Figure 4B.

[0107] The experiment applies to 3D, challenging the reconstruction of multiple objects and periodic structures of the vertebrae from a two-plane X-ray imaging modality. The XRAY-to-DRR conversion step is added before elastic 3D / 2D registration to convert the actual X-ray image into a DRR-like image with the anatomical structures separated. Using this conversion step facilitates mesh 3D / 2D registration because image matching is more efficient, even when using standard similarity metrics, since both images have similar properties and similarity measurements are performed only on the separated anatomical structures of interest.

[0108] Figure 5 shows a global flowchart of an example of performing a 3D / 2D alignment method according to an embodiment of the present invention. This flowchart 60 converts X-ray images to DRR images. A frontal X-ray image 61 is converted to a frontal DRR image 63 by a GAN 62 having a frontal weight set. A lateral X-ray image 64 is converted to a lateral DRR image 66 by a GAN 65 having a lateral weight set. Starting from both the initial 3D model 67 and the initial DRR superposition 68, both the modified aligned 3D model 74 and the modified DRR superposition 75 are obtained by iterative optimization 69. This iterative optimization comprises a cycle including the following consecutive steps: firstly, step 70 of the 3D model instance, then step 71 of the 3D model DRR projection, then step 72 of the similarity score, then step 73 of the optimization, and then back to step 70 of the 3D model instance with updated parameters.

[0109] To efficiently address DRR image matching in radiographic images with overlapping objects, improve robustness and alignment accuracy, and reduce computational complexity, it is proposed to introduce a pre-image-to-image conversion step using a pre-trained pix-to-pix GAN network that converts X-ray images to DRR images, as shown in Figure 5. The resulting XRAY-to-DRR conversion network allows for both efficient cross-image similarity measurement and the use of a simple DRR generation algorithm, as the various DRRs generated from both matching images, the 3D mesh (DRR3DM), and the target DRR belong to the same domain. The proposed XRAY-to-DRR conversion step converts the X-ray image to a DRR3DM-like image with a uniform background and noise, and all soft tissues and adjacent structures removed, as shown in Figure 4B. In fact, the image converter is trained to separate adjacent bone structures by outputting one converted DRR image for each bone structure. In this way, the problem of aligning a single-structure 3D model in a scene composed of multiple objects, such as a skeletal structure with joints, is addressed.

[0110] Figure 6 shows a more detailed flowchart of an example of a training stage before performing a 3D / 2D alignment method according to an embodiment of the present invention. The method proposed by the embodiment of the present invention requires a training stage before providing a U-Net converter to deform X-ray patches into DRR-like patches in order to facilitate 3D / 2D alignment.

[0111] The training database 80 includes a set of two-plane X-ray images 81, both frontal (also called posterior-anterior) and lateral, and a 3D reconstructed model 82 of the vertebrae.

[0112] The training data extraction process 83 is performed using, for example, a process similar to that described in patent application WO2009 / 056970, which is an operation 84 of DRR calculation using a raycasting model and an operation 85 of region of interest extraction and resampling. The training phase 83 is required to train the U-Net XRAY to DRR converter. Both the DRR calculation operation 84 and the ROI extraction and resampling operation 85 use both a set of two-planar X-ray images 81 and a 3D reconstruction model of the vertebrae 82 as inputs. The DRR calculation operation 84 produces a set of two-planar DRRs of the model 86 as output, while the ROI extraction and resampling operation 85 uses this set 86 of two-planar DRRs of the model as input.

[0113] The training data 87 includes a set of X-ray patches 88 and a set of DRR patches 89, where these patches represent different anatomical structures. This set of X-ray patches 88 and this set of DRR patches 89 are generated by the ROI extraction and resampling operation 85. The data 87 required for this training is a pair of patch sets 88 and 89, which are all square images with a size of 256 x 256 pixels.

[0114] Both this set of X-ray patches 88 and this set of DRR patches 89 are used as inputs during operation 90 to train the U-Net converter using a GAN to generate a set of weight parameters 91 for the U-Net converter.

[0115] Patch 88, indicated by X, is extracted from the two-plane X-ray image 81. Patch 89, indicated by Y, is extracted from the two-plane DRR 86 generated from the reconstructed 3D model.

[0116] The training aims to fit these data to a U-Net model, where Y = predict(X, W) + ε, where ε is the residual of the prediction and W is the neural network parameter, resulting in the trained weights.

[0117] X and Y are 4D tensors of the following size: X:[N, 256, 256, M]:N is the total number of patches, and M is the number of input channels in the following example. M=1: The patch belongs to the front (FRT) or side (LAT) view. M=2: Train the joint model FRT+LAT Y:[N, 256, 256, K×M]:N is the total number of patches, and K is the number of anatomical structures (DRR images) in the output channel of the following example. K=1: There is only one anatomical structure. K=3: Adjacent vertebrae above / middle and below, what is the proposed model for 3D / 2D alignment of the vertebrae in the embodiment? K=4: For example, a model with left and right femurs and tibias in the LAT view. The extraction of training data, i.e., patches 88 and 89, is performed as follows: ●M patients will be used for training. Each patient has the following data. FRT and LAT calibrated radiographic images, DICOM images, 3D modeling of the spine. ●M patients are divided into two distinct sets: a training set used to modify the weights of the neural network, and a test set used to examine the convergence of training, the training curve, and select the best model using optimal generalized inference / prediction. ● For each patient, a set of P patches is extracted such that N = M × P. Expert 3D models are used to generate both PA (pre- and post-posterior) and LAT DRR images. A set of P patches, each 256x256 pixels in size, is extracted from both the DRR (Y) and the actual image (X). The patch locations are randomly shifted from the center of the 3D vertebral body within the range of [25, 25, 25] (relative to the unified method) to artificially increase the size of the dataset by shifting the structures on the image. Another data augmentation is performed by using: • Global rotation centered on 2D VBC: [-20 20 degrees] • Global scaling [0.9, 1.1] • Localized image deformation with fluctuations of [-5, +5 degrees] near the center of the end plate.

[0118] Figure 7 shows a more detailed flowchart than that in Figure 5, illustrating an example of performing a 3D / 2D alignment method according to an embodiment of the present invention. The represented iterative alignment process aims to find model parameters that maximize the similarity between the transformed target DRR image and the moved DRR image calculated from the 3D model (DRR3DM) in both rear-front (PA) (also called frontal) and lateral (LAT) views. This more detailed flowchart includes the process of transforming the X-ray image into a DRR image 102.

[0119] The spine statistical shape model (SSM) 112 and the vertebral deformation model 113 (each consisting of a generic model, a dictionary of labeled 3D points, and a set of parametric deformation handles controlling the translation least-squares deformation) are used as input by an initialization process 101 that performs an automated global fit operation 103 to give a first estimate 115 of the 3D model. This first estimate 115 of the 3D model is then inserted into a set of initial parameters 116 for use by an iterative optimization loop 100.

[0120] The first estimate 115 of this 3D model is also used as input to a target generation process 102, along with a set of two-planar X-ray images 114, both frontal and lateral. This target generation process 102 performs a pixel-to-pixel (pix-to-pix) image-to-image conversion. This target generation process 102 includes the following steps: Firstly, a step 104 of ROI extraction and resampling, starting with the first estimate 115 of this 3D model and this set of two-planar X-ray images 114 as input to generate a set of X-ray patches 105, denoted by X; from there, a step 106 of DRR inference using a U-Net converter, weighting parameters 91, 117, thereby generating a set of predicted DRR patches 107, denoted by Y.

[0121] This iterative optimization loop 100 of the 3D / 2D alignment process comprises the following cycle of consecutive steps. ●The first step of mesh instance generation 108 uses both the initial parameter set of 116 and the local vertebral statistical shape model (SSM) as input. ● Second step 109 of calculating the DRR of right vertebral ROI and patch size using ray casting. ●The third step of similarity scoring 110 also uses the set of predicted DRR patches 107 indicated by Y, ● A fourth step in which the model is optimized using a Covariance Matrix Adaptive Evolutionary Strategy (CMA-ES), and then the parameters are updated, returning to the first step 108, and at the end of the iterative cycle, the optimized parameters 119 corresponding to the aligned model are produced as the final output.

[0122] This novel method for resolving the elastic or even rigid 3D / 2D alignment of a 3D model on a planar radiographic image, as proposed by embodiments of the present invention, includes training an XRAY to DRR converter network and a fast DRR generation algorithm calculated from a bone bilayer mesh model that does not require accurate tissue density. In step 110, the similarity between the variable DRR generated from the 3D mesh (DRR3DM) in step 108 in step 109 and the converted DRR image (DRRGAN) used as the target in step 107 is maximized, firstly, because the correspondence between images is facilitated by being in one unique domain, and secondly, because mismatches between adjacent structures are avoided thanks to structural separation in the XRAY to DRR conversion, so that the optimal model parameters can be easily found in step 119. The iterative alignment process 100 aims to maximize the similarity between DRR3DM and DRRGAN simultaneously in all views.

[0123] Formally, this is the optimal parameter for controlling the pose and shape of a 3D model.

[0124]

number

[0125] To find 119, we can consider the following cost function (Equation 1) as the way to maximize it.

[0126]

number

[0127] In the above equation,

[0128]

number

[0129] This is a variable image of view υ.

[0130]

number

[0131] and target image

[0132]

number

[0133] An arbitrary similarity function calculated between (for a total of L views), where both images are regions of interest (ROIs) of the original images around the target structure to be aligned.

[0134]

number

[0135] This depends on the model's parameter vector p and the DRR projection function P of the view υ. v It is calculated as (Equation 3).

[0136]

number

[0137] In the above equation, Ψ(p) (Equation 4) generates a 3D model instance controlled by the parameter vector p. Target image

[0138]

number

[0139] This is defined as the image transformed using U-Net prediction (Equation 5).

[0140]

number

[0141] In the above equation,

[0142]

number

[0143] υ is the original X-ray ROI patch of view υ, and f(I, W) represents the feedforward inference of a GAN-based trained U-Net with CNN parameters W and input image I. In this context, a two-planar radiographic system with L=2 views (PA and LAT) is used. The cost function (Equation 1) can be maximized in the iterative process using an optimizer, as shown in the flowchart of the proposed method shown in Figure 7.

[0144] The following paragraphs present a method for DRR projection from a two-layer mesh surface and a CNN architecture according to embodiments of the present invention, as well as training of an XRAY to DRR converter. Although the description of this method is directed toward a vertebral structure, it can be applied to 3D / 2D alignment of any (multi) structural 3D model in a calibrated planar view.

[0145] Here, we describe the DRR projection that forms a two-layer mesh. This is preferably implemented by using a process similar to that described in patent application WO2009 / 056970. This involves the projection function P v This is the process of generating a DRR from a 3D surface mesh S on a single view υ related to (S)(Equation 3). The virtual radiographic image (DRR) is calculated using ray casting intersecting a 3D surface model consisting of two layers separating two media: cortical and spongy bone media. The media separation plane is used to account for two different factors of attenuation corresponding to the material properties of both. The cortical structure has the highest X-ray energy absorption and appears brighter in radiographic images. The 3D model has 3D vertices with (x, y, z) coordinates.

[0146]

number

[0147] A set of vertices and a set of faces having vertex indices that define the triangular faces.

[0148]

number

[0149] The mesh surface is represented by S = {V, F}, which is defined by and .

[0150] Figure 8A shows an example of the principle of creating an inner layer from an outer layer using the previous cortical thickness and normal vector, according to an embodiment of the present invention. It can be seen that both the previous cortical thickness ti121 and the normal vector Ni122 have been added to the mesh surface 120.

[0151] A two-layer mesh surface is created by adding an inner layer to mesh 120. Surface vertex V i For each, the internal vertices are calculated using Equation 6.

[0152]

number

[0153] In the above equation,

[0154]

number

[0155] is vertex V i This is the surface normal in , and as can be seen in Figure 8A, t i is vertex V iThis represents the cortical thickness at a given point. The normal to each vertex is calculated as the normalized vector of the sum of surface normals belonging to the vertex ring. Values ​​for the cortical thickness of specific anatomical landmarks can be found in literature studies, for example, for vertebral endplates and pedicles.

[0156] Figure 8B shows an example of an interpolated thickness map of the entire mesh of the L1 vertebra according to an embodiment of the present invention. Different regions of map 124, typically represented by zones of different colors in map 124, correspond to different values ​​ranging from 0.5 to 1.8 (top-down), with the intermediate value being 1.2 on scale 125, which is shown to the right of map 124.

[0157] These values ​​for the cortical thickness of specific anatomical landmarks are interpolated across the vertices of the mesh using a thin-plate spline 3D interpolation technique, which enables the calculation of the cortical thickness map 124, as seen in Figure 8B.

[0158] Figure 9 illustrates the principle of raycasting through a two-layer mesh. A ray connecting the X-ray source 126 to pixel 129 on the DRR image 128 alternately traverses the cortical and spongy media of vertebra 127. Raycasting replicates the principle of X-ray imaging and calculates the cumulative thickness traversing the two bone media of vertebra 127.

[0159] Figure 10 shows an example of a GAN network architecture and training from XRAY to DRR according to an embodiment of the present invention.

[0160] The U-Net generator 130 is trained to convert the input XRAY patch 143 into a DRR image 144 having three channels 146, 147, and 148 represented by red, green, and blue (RGB) images, i.e., one channel per vertebral DRR 146, 147, and 148.

[0161] The generator 130 has several convolutional layers, each layer containing one or more slices, and each slice represents an action. Several types of slices or actions are available. ●Slice or operation 131: A 2D convolutional layer with a size of 4x4, a stride of 2, and using the Leaky Rectified Linear Unit activation function. ● Slicing or Operation 132: Batch Normalization ● Slicing or operation 133: Upsampling 2D layer of size 2 ●Slicing or operation 134: Dropout layer ● Slice or Operation 135: A 3-channel output convolutional 2D layer with a size of 4x4, a stride of 1, and using a hyperbolic tangent activation function. ● Slice or operation 137: Size 4x4, stride 1, output convolution 2D layer

[0162] From left to right, from XRAY patch 143 to DRR image 144, the configuration of the consecutive convolutional layers in generator 130, where the input is XRAY patch 143 and the output is DRR image 144, is as follows: ●Layer 1: Slice 131 ●Layer 2: Slice 131, followed by slice 132 ●Layer 3: Slice 131, followed by slice 132 ●Layer 4: Slice 131, followed by slice 132 ●Layer 5: Slice 131, followed by slice 132 ●Layer 6: Slice 131, followed by slice 132 ●Layer 7: Slice 131, then slice 132, then slice 133 ●Layer 8: Slice 131, then slice 134, then slice 132 ●Layer 9: Slice 133, then slice 131, then slice 134, then slice 132 ●Layer 10: Slice 133, then slice 131, then slice 134, then slice 132 ●Layer 11: Slice 133, then slice 131, then slice 134, then slice 132 ●Layer 12: Slice 133, then slice 131, then slice 132 ●Layer 13: Slice 133, then slice 131, then slice 132 ●Layer 14: Slice 133, then slice 131, then slice 132, then slice 135

[0163] The layers can be connected as shown by horizontal line 136. ● Layers 1 and 14 are connected. ● Layers 2 and 13 are connected. ● Layers 3 and 12 are connected. ● Layers 4 and 11 are connected. ● Layers 5 and 10 are connected. ● Layers 6 and 9 are connected. ● Layers 7 and 8 are connected.

[0164] The discriminator 140 has five consecutive layers, from left to right, from input to output. ●Layer 1: Slice 131 ●Layer 2: Slice 131, followed by slice 132 ●Layer 3: Slice 131, followed by slice 132 ●Layer 4: Slice 131, followed by slice 132 ●Layer 5: Slice 137

[0165] The output of the discriminator 140 has a loss function 139. Following this loss function 139 is a part 142 of an adversarial switch having two switchable positions 0 and 1.

[0166] There is a loss function 138 between the generated DRR image 144 (actually three images 146, 147, and 148, one for each vertebra) and the corresponding ground truth DRR image 145 (actually three images 146, 147, and 148, one for each vertebra). From the generated DRR image 144, there is feedback toward a switchable "fake" position of portion 141 of the adversarial switch. From the corresponding ground truth DRR image 145, there is feedback toward another switchable "actual" position of this portion 141 of the adversarial switch. The ends of this portion 141 of the adversarial switch opposite the fake and actual positions are then fed back into one of the inputs to discriminator 140, the other input to discriminator 140 being an XRAY patch 143.

[0167] Next, we will discuss the XRAY to DRR converter. The XRAY to DRR image converter (Equation 5) is a U-Net network trained using a Generative Adversarial Network (GAN). U-Net training creates a high-density pixel-by-pixel mapping between paired input and output images, enabling the generation of realistic DRR images with a consistent overall appearance. To solve this problem, we need to construct a deep abstraction. The Generative Adversarial Network (GAN) consists of a U-Net generator (G) 130 and a CNN discriminator (D) 140. Thus, two loss functions are defined: loss function 138 for the generator 130 and loss function 139 for the discriminator 140. Firstly, the residual between the converted 144 DRR images from the training database and the actual 145 DRR images is calculated by the following mean absolute loss function 138 error.

[0168]

number

[0169] In the above equation, (n, w, h, c) is the dimension of the 4D tensor of the training data output (DRR image) which has n 3D images of size width (w) × height (h) × channel (c).

[0170]

number

[0171] This is the j-th pixel of the i-th sample image, and f(XRAY i , W) is the U-Net prediction for the i-th input XRAY, with the current CNN parameters W. The input data is also a 4D tensor of size (n, w, h, k). Sizes k and c are the number of channels in the input (XRAY) image and the output (DRR) image, respectively. The discriminator (D) 140 aims to classify whether the XRAY / DRR pair of images presented to its input is a DRR-generating image (false class) or an actual DRR image (real class). Thus, the loss function ξD 139 is defined by the binary cross-entropy loss.

[0172] The training data is divided into small minibatches (e.g., size 32). For each minibatch, the generator (G) 130 predicts a fake DRR image. The discriminator 140 is then first trained independently to predict a binary output using two images assigned to the generator output: one for the fake image (output set to zero: fake class) and one for the real image (output set to 1). Finally, gradient backpropagation is performed to update the generator weights (W) using the combined G130 and D140 networks and the total loss L = ξD + ξG (however, the discriminator weights are fixed during this step of generator training).

[0173] The architecture of the U-Net generator G130 consists of 15 convolutional layers for encoding / decoding image features, as described in the article [P. Isola, J.-Y. Zhu, T. Zhou, and AAEfros, “Image-to-Image Translation with Conditional Adversarial Networks,” CoRR, vol.abs / 1611.0, 2016]. Dropout slices 134 are used in the first three decoding layers to facilitate good generalization of the image-to-image translation model. The discriminator output neurons have a sigmoid activation function. In the current example, the input size of the CNN was 256×256×1 (XRAY patch in one view), and the output size of the DRR was 256×256×3. One output channel is assigned to one different anatomical structure to separate them in training: upper 146, middle 147, and lower 148, the vertebrae.

[0174] The pix-to-pix network uses the training procedure of a Generative Adversarial Network (GAN) to train the generator parameters (weights). The generator has a U-Net architecture that allows the image-to-image mapping to transform actual X-ray patches 143 into DRR patches 144 (from domain A to domain B). The entire network consists of a U-Net generator 130 and a CNN discriminator 140. In this example, the GAN input size is 256 × 256 × 1 (X-ray patch 143), and the DRR 144 output size is 256 × 256 × k, where k is defined as the number of anatomical structures in the output, which here are three vertebrae 146, 147, and 148. Therefore, to separate them in training, one output channel is assigned to each different anatomical structure (different vertebrae in this case). Output channels for k=3 are defined for visible bone structures with k=3, specifically for the target vertebra 147 (Ki=1) and its adjacent vertebrae 146 (Ki=0) and 148 (Ki=2).

[0175] Once trained, the network can transform an input image 143 (X-ray) belonging to a complex domain A into a simplified virtual X-ray (like DRR) 144 belonging to domain B. Thanks to the image transformation, the generator (U-Net shape) 130 network can suppress noise, adjacent structures, and background from the original image.

[0176] Using the converted image, the following becomes possible: ●Firstly, to facilitate the matching of images (similarity calculation), Secondly, to avoid inconsistencies between them, adjacent structures are separated.

[0177] When used in a 3D / 2D alignment process scheme, these bone multistructure pix-to-pix networks assist in the similarity calculation stage between the target DRR image generated using the GAN U-Net and various images generated from the 3D model. Since the similarity value is used as the cost to find the optimal model parameters that maximize the similarity between the model and the images, a simpler similarity metric can be used because both images belong to the same domain diversity, resulting in fewer maxima induced during optimization due to mismatches between adjacent structures.

[0178] Next, the experimental set is described. The dataset for training the XRAY to DRR converter includes two-plane acquisitions (PA and LAT views) from 463 patients performed using a geometrically calibrated system and a known 3D environment. Reconstructed 3D spines are used to have ground truth in the 3D models, integrating a semi-automated 3D reconstruction method. For evaluation of the method proposed by embodiments of the present invention, a separate clinical dataset of 40 adolescent idiopathic scoliosis patients (mean age 14 years, mean principal Cobb angle 56°) was used. A bronze standard was constructed for each 3D model to define ground truth as the mean of 3D reconstructions by three experts.

[0179] During training of the XRAY to DRR converter, PA and LAT DRR images were generated individually for each patient in the training dataset, using the previously presented algorithm, for each reconstructed 3D model.

[0180] Figure 11A shows an example of GAN-based XRAY to DRR conversion from an actual X-ray image to a 3-channel DRR image corresponding to an adjacent vertebral structure.

[0181] Image 151a is a converted frontal X-ray image, showing vertebra L4, vertebra L5, and sacral plate S1 from top to bottom.

[0182] Image 152a is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in an embodiment of the present invention. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, blue = sacral plate S1.

[0183] Image 153a is the original frontal DRR image corresponding to frontal X-ray image 151a, and frontal DRR image 152a is a reference image to which the algorithm of the method proposed by embodiments of the present invention is compared.

[0184] Image 151b is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in an embodiment of the present invention, and corresponds only to vertebra L4. It also corresponds to the upper part of image 152a.

[0185] Image 152b is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in an embodiment of the present invention, and corresponds only to vertebra L5. It also corresponds to the middle portion of image 152a.

[0186] Image 153b is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in an embodiment of the present invention, and corresponds only to the sacral plate S1. It also corresponds to the lower part of image 152a.

[0187] Image 151c is the original frontal DRR image corresponding to frontal X-ray image 151a, and is a reference image to be compared with frontal DRR image 151b, corresponding only to vertebra L4. It also corresponds to the upper part of image 153a.

[0188] Image 152c is the original frontal DRR image corresponding to frontal X-ray image 151a, and is a reference image to be compared with frontal DRR image 152b, corresponding only to vertebra L5. It also corresponds to the middle section of image 153a.

[0189] Image 153c is the original frontal DRR image corresponding to frontal X-ray image 151a, and is a reference image to be compared with frontal DRR image 153b, corresponding only to the sacral plate S1. It also corresponds to the lower part of image 153a.

[0190] Image 151d is a converted lateral X-ray image, showing vertebra L4, vertebra L5, and sacral plate S1 from top to bottom.

[0191] Image 152d is a lateral DRR image generated from lateral X-ray image 151d by the method proposed in an embodiment of the present invention. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, blue = sacral plate S1.

[0192] Image 153d is the original lateral DRR image corresponding to lateral X-ray image 151d, and lateral DRR image 152d is a reference image compared to it for training the algorithm of the method proposed by embodiments of the present invention.

[0193] Following morphological changes in the vertebrae along the spine, a view-by-view converter is trained for each vertebral segment: T1 to T5, T6 to T12, and L1 to L5. Twenty patches per vertebra are extracted around the vertebral body center (VBC) with random displacement, rotation, and scale. The region of interest (ROI) is defined with a dynamic scaling factor such that the target vertebra and at least adjacent levels of the vertebral endplate remain visible in a 256x256 pixel patch, the scaling factor depending on the dimensions of the vertebra in the image. The training dataset was split into two sets: training (70%) and test (30%). Training was run for 100 epochs. In each epoch, the mean squared error (MSE), which is the error of the predicted patches (test set), was calculated to select the best model for the entire epoch. The lumbar converter (LAT view) training consists of c=1 or c=3 (Equation 7), i.e., the output image has either only one channel or a layered image with three channels containing the DRR of three vertebrae. The MSE of the DRR of the intermediate vertebrae is 0.0112 and 0.0108 for one channel and three channels, respectively, thereby demonstrating that the additional output information of adjacent vertebrae is useful for GAN training. Therefore, even when the application to 3D / 2D alignment of a 3D model of a single vertebra uses the central channel as the target, three output channels were used per training. Once trained, the network can convert X-ray images into layered images containing DRR images for each structure, as seen in Figure 11A. Qualitative results of the converter are shown in both Figure 11A and Figure 11B.

[0194] Figure 11B shows a transformation, also known as the conversion from X-ray domains to DRR domains, under adverse conditions, to demonstrate the robustness of the method proposed by the embodiments of the present invention.

[0195] Image 154a is a converted frontal X-ray image, showing vertebra L4, vertebra L5, and sacral plate S1 from top to bottom. The X-ray image shows poor visibility.

[0196] Image 155a is a frontal DRR image generated from frontal X-ray image 154a by the method proposed in an embodiment of the present invention, corresponding to the "conversion" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, blue = sacral plate S1.

[0197] Image 156a is the original frontal DRR image corresponding to frontal X-ray image 154a, and frontal DRR image 155a is a reference image to which the effectiveness and performance of the algorithm of the method proposed in the embodiment of the present invention are compared.

[0198] Image 154b is the transformed frontal X-ray image, showing vertebra L4, vertebra L5, and sacral plate S1 from top to bottom. The X-ray images are superimposed.

[0199] Image 155b is a frontal DRR image generated from frontal X-ray image 154b by the method proposed in an embodiment of the present invention, corresponding to the "conversion" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, blue = sacral plate S1.

[0200] Image 156b is the original frontal DRR image corresponding to frontal X-ray image 154b, and frontal DRR image 155b is a reference image to which the effectiveness and performance of the algorithm of the method proposed in the embodiment of the present invention are compared.

[0201] Image 154c is a converted frontal X-ray image, showing vertebra L4, vertebra L5, and sacral plate S1 from top to bottom. The X-ray image shows a circular metal component 157.

[0202] Image 155c is a frontal DRR image generated from frontal X-ray image 154c by the method proposed in an embodiment of the present invention, corresponding to the "conversion" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, blue = sacral plate S1.

[0203] Image 156c is the original frontal DRR image corresponding to frontal X-ray image 154c, and frontal DRR image 155c is a reference image to which the effectiveness and performance of the algorithm of the method proposed in the embodiment of the present invention are compared.

[0204] Image 154d is a converted frontal X-ray image, showing vertebra L4, vertebra L5, and sacral plate S1 from top to bottom. The X-ray image shows a metal screw 158.

[0205] Image 155d is a frontal DRR image generated from frontal X-ray image 154d by the method proposed in an embodiment of the present invention, corresponding to the "conversion" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, blue = sacral plate S1.

[0206] Image 156d is the original frontal DRR image corresponding to frontal X-ray image 154d, and frontal DRR image 155d is a reference image to which the effectiveness and performance of the algorithm of the method proposed in the embodiments of the present invention are compared.

[0207] Next, we will discuss the similarity function. Five methods were implemented to evaluate the similarity of the images used in (Equation 2). Thanks to X-ray image transformation using a U-Net neural network, similarity can be calculated between two DRR-like images, and common unimodal similarity functions such as normalized cross-correlation (NCC) and sum of squared differences (SSD) can be used. Measurements based on image gradients were normalized gradient information (NGI) and NCC calculated on the gradient image (NCCGRAD). Normalized mutual information (NMI) measurements were also included, where a joint histogram is calculated from binarized images in 8 bits.

[0208] Next, we will discuss the spine statistical shape model. Deformable models are used when alignment is elastic, such as when the shape of the object is optimized in addition to the model's pose. The deformation technique uses mesh translation least squares (MLS) deformation controlled by a set of regularized geometric parameters using a PCA model. Pre-definition of a simplified parametric model encompassing the object's shape allows for a compact representation of the geometry and directly provides subject-specific landmarks and geometric primitives for use in biomedical applications.

[0209] The resulting PCA model captures the major fluctuations.

[0210]

number

[0211] This provides a linear generative model of the form, where B is PCA-based,

[0212]

number

[0213] is the average model, and m is the deformation mode vector. The generation of mesh instances using the function Ψ(p) (Equation 4) is controlled by the parameter vector p={Tx, Ty, Tz, Rx, Ry, Rz, Sx, Sy, Sz, m}, and in the above equation,

[0214]

number

[0215] It consists of nine parameters of the affine part (translation, rotation, and scaling) for deforming and scaling the 3D model into the correct pose in the X-ray calibrated 3D environment, and a shape vector m with |m| PCA modes. When the vector m is given, the MLS deformation processes the parameter s that is calculated and used to deform the mesh vertices.

[0216] Next, the optimizer will be described. The cost function (Equation 1) is optimized by minimizing the following equation (Equation 8) to solve the 3D / 2D alignment of the two planar PA and LAT radiographic images.

[0217]

Equation

[0218] In the above equation, φ(I1, I2) is a similarity measure restricted to

[0001] in the cases of NGI, NCC, NCCGRAD, NMI, and SSD. To define the cost, the sum of the similarity scores of both PA and LAT was calculated. To minimize the cost function (Equation 8), a derivative-free exploration CMA-ES optimizer was used. Each CMA-ES iteration evaluated the cost function 100 times to construct the covariance matrix. Upper / lower limits can be defined.

[0219] Next, the method proposed by the embodiments of the present invention is evaluated. Three types of experiments were conducted to evaluate the proposed method in the context of 3D / 2D alignment of vertebral 3D models in two-plane X-rays. First, the target registration error (TRE) of the anatomical landmarks belonging to the 3D model was reported to study the behavior of the similarity metric with and without using the GAN transformation step. Additional tests were conducted to show the advantage of structural separation in the conversion from XRAY to DRR, making it less affected by the initial pose. Finally, the results of the accuracy of fully automated 3D reconstruction of the spine are presented.

[0220] Next, the target registration error (TRE) is described. In this experiment, 17 vertebral 3D models from level T1 to L5 were first fitted to the previously reconstructed ground truth 3D model. The registration is configured to solve rigid body deformations with six degrees of freedom. For each vertebra of each patient, a random deformation was applied to the 3D model. For translations Tx, Ty, Tz, the shift of the model was performed in the range of ±3 mm using a random uniform method. For in-plane rotations Rx and Ry, a range of ±5° was defined, and for out-of-plane (axial vertebral) rotation Rz, ±15° was used. The upper and lower limits are also defined as these values in the optimizer. The range of random deformations is limited so that the registration with the original X-ray target can converge. In fact, the 3D / 2D registration is evaluated using both DRRGAN and the original XRAY target to quantify the improvement brought by the GAN-based converter. The U-Net output channel corresponding to the central vertebra was used as the target for the DRRGAN image. TRE was calculated as the root mean square error of the 3D position error of the anatomical landmarks after registration. The landmarks extracted from the registered 3D model were the centers of the upper and lower endplates and the centers of the left and right pedicles. A total of 40 (patients) × 17 (vertebral levels) × 4 (landmarks) measurements were made for each similarity / target pair.

[0221] Table 1 shows the quantitative TRE for each target / similarity pair.

[0222] [Table 1]

[0223] Table 1 compares the TRE results for five similarity measurements and two targets, reporting the error rate of landmarks with TRE > 2.4 mm, the average cost function delta, i.e., amplitude cost Δ = max(costs) - min(costs), and the average number of iterations to converge. In the above formula, the stopping tolerance was fixed at 1 for the SSD metric and 0.001 for the remaining metrics. When the image transformation step is used with DRRGAN as the target, the TRE results are nearly similar for all similarity metrics, from 1.99 to 2.12, except for a higher error of 2.52 for NCC GRAD. When the target used is XRAY, the SSD, NCC, and NMI metrics drop the alignment results because the DRR3DM image has a limited level of similarity to the XRAY image. NMI has a quasi-null variation in cost, indicating convergence issues. As seen in Table 1, only gradient-based metrics reach TRE of 3.6 mm and 3.3 mm for NGI and NCCGRAD, respectively. When using DRRGAN, the evolution of the similarity level during optimization becomes more important. For example, with the NGI metric, the mean cost delta is 0.29 for DRRGAN compared to 0.05 for XRAY, and as seen in Table 1, this means that using a GAN-based converter makes the evaluated cost less sensitive, which helps improve convergence.

[0224] Figure 12A shows the target alignment error as a function of the initial pose shift on the vertical axis. This target alignment error (TRE) is plotted on the y-axis as a function of various initial pose shifts along the vertical axis z on the y-axis. For DRRGAN containing three vertebral structures (all channels flattened in one image), the error value often exceeds 5, 10, or even 20. This error value is significant, and there is also considerable uncertainty in this error value. For a single DRRGAN channel, the error value is less than 5 and in most cases does not exceed 2 or 3. This error value is small, and there is also very little uncertainty in this error value. This indicates that the separation of structures is very important to avoid mismatches between adjacent vertebrae. Therefore, thanks to the method proposed by embodiments of the present invention, the vertebrae are separated from each other, so the error value becomes much smaller and much more well known.

[0225] Figure 12B shows an example where the vertebrae are moved vertically upward. The vertebrae are shifted 10 mm vertically upward. This means that the vertebrae are 10 mm higher vertically than in the ground truth model.

[0226] In this experiment, a ground truth 3D model of the L2 vertebra was shifted in 5mm increments within a range of ±30mm along the z-axis in the proximal-distal direction, as seen in Figures 12A and 12B, thereby giving 13 initial poses. Rigid body registration was then performed using either a single channel corresponding to the central vertebra or an image composed of all channels flattened, as the DRRGAN target image, thereby blending the superior, middle, and inferior vertebral structures. The metric used was SSD, with upper and lower limits defined as ±5mm for Tx and Ty, ±32mm for Tz, and ±5° for rotation. Figure 12A shows a box plot of the transition residuals for each initial pose shift for 40 patients. As seen in Figure 12A, the residual deformation after registration has significant errors behind an absolute shift of 10mm without structural separation. Outliers also occur when the absolute shift is less than 10mm because the registration is perturbed by adjacent levels L1 and L3. Using GANDRR-transformed images separates the target structure, allowing the alignment process to capture a wider range of initial poses and become more accurate.

[0227] Next, sensitivity to the initial pose is studied. For the alignment of single-structure 3D models in scenes composed of periodic multi-structures, such as vertebrae along the spine, sensitivity to the initial model pose resulting from automated coarse alignment is studied, for example, because there is a risk that, beyond a limit, the 3D model alignment will incorrectly converge to adjacent structures.

[0228] Next, the accuracy of the vertebral pose and shape is investigated. In this experiment, fully automated 3D reconstruction of the spine is evaluated. Initial solutions for the shape and pose of the 3D vertebral models are provided by a CNN-based automated method described in the paper ["Towards automated 3D Spine reconstruction from biplanar radiographs using CNN for statistical spine model fitting" by B. Aubert, C. Vazquez, T. Cresson, S. Parent, and J. De Guise, IEEE Trans.Med.Imaging, p. 1, 2019]. This method provides pre-personalized statistical spine models (SSMs) for levels C7 through L5 using endplates and pedicle centers detected in biplanar radiographs. To fine-tune these acquired 3D models, 3D / 2D alignment is applied in elastic mode, with |m|=20PCA mode bound to ±3. Alignment was performed individually for each level using NCC metrics. Next, the resulting aligned model is used to update the SSM regularization in order to obtain the final model after local 3D / 2D fitting.

[0229] Table 2 shows the quantitative landmarks (TREs) for 3D / 2D alignment of the vertebrae.

[0230] [Table 2]

[0231] As shown in Table 2, the 3D positioning of landmarks at the center and corners of the endplate and the center of the pedicle was improved by the proposed 3D / 2D alignment step. As seen in Table 2, the most significant improvements were observed at the center of the pedicle using 3 mm and 2.2 mm TREs before and after fine 3D / 2D alignment. Error rates >2.4 mm dropped by 26.4%, 15.4%, and 19.1% at the pedicle, center and corners of the endplate, respectively.

[0232] Table 3 shows the quantitative errors regarding the position and orientation of the vertebrae.

[0233] [Table 3]

[0234] An axial coordinate system is defined for each vertebra using the pedicle and the center of the endplate. As shown in Table 3, the 3D position error for each vertebral object is calculated for each vertebral segment in terms of the ground truth 3D model for positional and directional agreement. The average error for all conversions is less than 0.5 mm, indicating a low systematic bias in this method. As seen in Table 3, the standard deviation was more significant for X-conversion, i.e., position in the lateral view, particularly at the thoracic level.

[0235] Finally, shape accuracy was estimated by calculating the distance error from nodes to surfaces using a reference 3D model reconstructed from CT scan images of four patients for whom both two-plane acquisitions and CT scan images were available simultaneously. The volume resolution was 0.56 × 0.56 × 1 mm, and segmentation was performed using 3D slicer software.

[0236] Figure 13A shows the anatomical regions used to calculate the statistical distance from the node to the surface. In vertebra 164, the posterior arch 165 can be distinguished from the structures 166 representing the vertebral body and pedicle.

[0237] Figure 13B shows the distance map of the maximum error calculated for the L3 vertebra. Different regions of map 167, which are usually represented by different colored zones of map 167, correspond to different values in the range from 0.0 mm to 6.0 mm (from top to bottom), and the median value is at 3.0 mm on scale 168 shown on the right side of map 167.

[0238] Table 4 shows the results of the shape accuracy for the CT scan model.

[0239]

Table 4

[0240] The objects were precisely aligned and the distance from the nodes to the surface was calculated. The mesh density of the model, i.e., the number of nodes, is specified in Table 4. As shown in Table 4, error statistics are reported for different anatomical regions, which are the mesh of the entire vertebra 164 as shown in Figure 13A, or the mesh of the posterior arch 165, or the mesh of structure 166 that includes the vertebral body and pedicle. As seen in Table 4, the average error was in the range of 1.1 mm to 1.2 mm. According to the error distance map 167 shown in Figure 13B, the maximum error is localized in the posterior arch region 165.

[0241] The method proposed by embodiments of the present invention adds a step prior to the image-to-image conversion of the target image in the intensity-based 3D / 2D alignment process, so that both images have robust dual image matching belonging to different domains, X-ray, and DRR, and so that they are displayed in a number of different environments and structures. As a first advantage, the XRAY to DRR converter improves the level of similarity between the two images by moving the target image to the same domain of various images, as seen in Table 1. The conversion step reduces reliance on the selection of similarity metrics, allowing not only the use of simplified DRR generation, as seen in Table 1, but even the use of a common unimodal metric. This is an interesting characteristic, as conventional experimental methods consisted of selecting a better metric than a set of metrics. However, this often resulted in a trade-off, as some of these metrics yielded good or bad results depending heavily on the particular case.

[0242] A second advantage is that, as seen in Figure 13A, inconsistencies between adjacent overlapping structures are avoided by selecting the structure of interest in the converter output. In fact, to generate local DRRs for each object, the structures were directly separated in the original 3D volume, and each DRR was assigned to a separate layer in the multi-channel output image of the XRAY-to-DRR converter. In the cited prior art, U.S. Patent Application Publication 2019 / 0259153, previous studies using XRAY-to-DRR converters generated only global DRRs that allowed for the recovery of segmentation masks for each organ, but were not able to generate local DRRs for each object.

[0243] When applied to fully automated 3D reconstruction of the spine from two-planar radiographic images, the improved 3D / 2D alignment step, with added enhancements, improves both object localization and shape compared to the CNN-based automated method used for initialization, as seen in Tables 2 and 3. The mean error in 3D localization of landmarks is 2 ± 1 mm, which is better than the 2.7 ± 1.7 mm detected by the CNN-based displacement regression method for pedicle detection, or better than the 2.3 ± 1.2 mm detected by the nonlinear spine model fit using a 3D / 2D Markov random field. Compared to the "quasi" automated 3D reconstruction method, which requires the user to input both spine curves in both views and then manually fine-tune them after the model is fitted on two-planar X-rays, the 3D / 2D alignment algorithm proposed by the embodiments of the present invention achieved better results for all pose parameters in populations with more severe scoliosis.

[0244] A reference method integrated into the SterEOS® software, used to perform 3D spinal reconstruction in clinical routines, was studied for the shape accuracy of ground truth objects obtained from CT scans relative to the reconstructed object. More precisely, this method requires time-consuming manual elastic 3D / 2D alignment (more than 10 minutes) to fit the vertebral contours projected onto the X-ray information. The automated local alignment step proposed herein, requiring less than 1 minute of computation time, favorably reduces operator dependence and yields similar accuracy results, such as 0.9±2.2 in the vertebral body and pedicle regions and 1.3±3.2 in the posterior arch region (mean ±2SD), compared to 0.9±2.2 and 1.2±3mm in some previous studies.

[0245] By employing these resulting generated images instead of actual X-rays, robust unimodal image matching without structural inconsistencies becomes possible. This solution, integrated into a 3D / 2D non-rigid alignment process aimed at adjusting 3D vertebral models from two-plane X-rays, improves accuracy results as previously observed.

[0246] Next, to demonstrate the improvements brought about by using the conversion from previous images, we include several other results, both of which relate to better similarity (Figures 14A-14B-14C-14D-15) and prevention of anatomical inconsistencies (Figure 16).

[0247] Figure 14A shows the cost values ​​of PA and LAT views using GAN DRR compared to actual X-ray images. The cost value is expressed on the y-axis as a function of the number of iterations on the x-axis. Curve 171 shows a higher cost function for X-ray images, and curve 172 shows a lower, and therefore better, cost function for GAN DRR images.

[0248] Figure 14B shows the similarity values ​​of PA and LAT views using GAN DRR compared to actual X-ray images. The similarity value is expressed on the y-axis as a function of the number of iterations on the x-axis. For the front view, curve 173 indicates a low similarity function for the X-ray image, while curve 175 indicates a high similarity function for the GAN DRR image, and therefore better. For the side view, curve 174 indicates a low similarity function for the X-ray image, while curve 176 indicates a high similarity function for the GAN DRR image, and therefore better. The similarity value of the NGI (Normalized Gradient Information) metric reaches a higher value when using the generated DRR GAN than when using the actual X-ray image.

[0249] Figure 14C shows the better fit observed when using GAN-generated DRR compared to Figure 14D, which uses actual X-ray images. In fact, in Figure 14C, the left portion represents the front view, and the fit between the generated image 181 and the target image 180 first transformed in the DRR domain is better than the fit between the generated image 185 and the target image 184 held in the X-ray domain, in Figure 14D, where the left portion represents the front view. Also, in Figure 14C, the right portion represents the side view, and the fit between the generated image 183 and the target image 182 first transformed in the DRR domain is better than the fit between the generated image 187 and the target image 186 held in the X-ray domain, in Figure 14D, where the right portion represents the side view.

[0250] Figure 15, like Figure 14B, shows similar results using the GNCC (Gradient Normalized Cross-Correlation) metric. In the front view, curve 190 using GAN DRR is higher than curve 188 without GAN DRR, indicating good similarity. In the side view, curve 191 using GAN DRR is higher than curve 189 without GAN DRR, indicating better similarity.

[0251] Figure 16 shows the results of the Target Error Revision (TRE) test in the Z direction (vertical image direction). In section A, various alignment errors expressed in mm on the vertical axis are plotted as a function of vertical shift, also expressed in mm on the horizontal axis. The NGI DRR curve 194 shows a much lower error than the NGI X-ray mask curve 192, and the NGI X-ray mask curve 192 shows a lower error than the NGI X-ray curve 193. In section B, the relative position between point 196 and spinal structure 195 indicates the Z shift from -10 mm. In section C, the relative position between point 198 and spinal structure 197 indicates the neutral shift. In section D, the relative position between point 200 and spinal structure 199 indicates the Z shift from 10 mm.

[0252] The objective of the TRE test was to deform a 3D model using known theoretical deformations and then analyze the residual error in the alignment between the theoretical deformation and the restored deformation. Using the generated GAN DRR, it was found to be less dependent on the initial position. Furthermore, the target alignment error (TRE) was found to increase significantly after a Z shift of ±4 mm, which indicates a significant structural mismatch between adjacent objects (vertebral endplates) that would occur if the generated GAN DRR image were not used, corresponding to curves 192 and 193 in section A of Figure 16.

[0253] The present invention has been described with reference to preferred embodiments. However, many modifications are possible within the scope of the invention. [Explanation of Symbols]

[0254] 1 DRR X-ray source 2 3D Volume 3. The first organ 4. The second organ 5 Global DRR 6 DRR image signal 6 Image area 7 DRR image signal 7 Image area 8 Crossing Zones 12. First label image 12. First Binary Mask 13. Second label image 13. Second Binary Mask 14 First Trace 14. Segmentation 15. Second Trace 15 Segmentation 16 X-ray image 17 GAN 18 DI2I CNN 20 3D Volumes 21 The first organ 22 The second organ 23 Local DRR 25 Local DRR 27 GAN 30 spine 31 vertebrae 40 3D models 41. First frontal X-ray source 42 Second lateral X-ray source 43 Planar X-ray image 44 Planar X-ray image 45 Spine 46 Spine 47 vertebrae 48 vertebrae 50 3D models 51 Planar DRR Transformed Image 52 Planar DRR Transformed Images 53 vertebrae 54 vertebrae 55 Front projection 56 Side projection 60 Flowcharts 61 Frontal X-ray image 62 GAN 63 Front DRR image 64 Lateral X-ray image 65 GAN 66 Side DRR image 67 Initial 3D Models 68 Initial DRR superposition 69. Iterative Optimization 74 Aligned 3D Models 75 DRR superposition 80 Training Databases 81 Set of 2-plane X-ray images 81 X-ray 2-plane image 81 Set of 2-plane X-ray images 3D reconstruction model of 82 vertebrae 86 2-plane DRR 87 Training Data Set of 88 X-ray patches Set of 89 DRR patches 91 parameters 100 Iterative Optimization Loops 100 Iterative Alignment Process 101 Initialization Process 102 Target Generation Process 102 DRR images 103 Automated Global Fit Operation 104 Steps for ROI extraction and resampling Set of 105 X-ray patches 106 Steps in DRR inference using the U-Net converter Set of 107 Predictive DRR Patches 112 Statistical Shape Model (SSM) 113 Vertebral Deformity Model 114 Set of 2-plane X-ray images 115 First estimate of 3D model 116 Initial Parameters 117 parameters 119 Optimized Parameters 120 mesh surface 120 mesh 121 Cortical thickness ti 122 Normal vector Ni 124 Maps 124 Cortical Thickness Map 1 / 125 scale 126 X-ray source 127 vertebrae 128 DRR images 129 pixels 130 U-Net Generator 131 slices or actions 132 slices or operations 133 slices or actions 134 slices or operation 135 slices or motion 136 Horizontal line 137 slices or actions 138 Loss Function 139 Loss Function 140 discriminators 140 CNN discriminator 141 The part of the hostile switch 142 The part about the hostile switch 143 XRAY Patches 144 DRR images 145 Ground Truth DRR Image 146 channels 146 Vertebral DRR 146 images 147 channels 147 Vertebral DRR 147 images 148 channels 148 Vertebral DRR 148 images 157 Circular metal parts 158 metal screws 164 vertebrae 165 Back bow 166 Structures encompassing the vertebral body and pedicle 167 Error Distance Map 1 / 68 scale 171 Curve 172 Curve 173 Curve 174 Curve 175 Curve 176 Curve 180 Target Images 181 Generated images 182 Target image 183 The generated images 184 Target Images 185 generated images 186 Target Images 187 generated images 188 Curve 189 Curve 190 curve 191 Curve 192 NGI X-ray mask curve 193 NGI X-ray curve 194 NGI DRR curve 195 Spinal structure 196 points

Claims

1. To simultaneously generate multiple sets of DRRs, each set of DRRs focusing solely on a first anatomical structure (21) that is distinguished from a second anatomical structure (22) by directly separating the structure in the original 3D volume of the subject, or By generating one DRR, but excluding other anatomical structures, and focusing only on the first anatomical structure (21) which is distinguished from the second anatomical structure (22) by directly separating the structure in the original 3D volume of the subject, By a single operation using either one convolutional neural network (CNN) or a group of convolutional neural networks (CNNs) (27) pre-trained to convert an actual X-ray image (16) into at least one digitally reconstructed radiographic image (DRR) (23), At least one or more actual X-ray images (16) of the patient including at least the first anatomical structure (21) and the second anatomical structure (22) of the patient, A medical image conversion method that automatically converts to at least one digitally reconstructed radiographic image (DRR) (23) of the patient that represents the first anatomical structure (24) and does not represent the second anatomical structure (26), The aforementioned training, The steps include preparing a training database which includes a set of two-plane X-ray images, both frontal and lateral, obtained from one or more people, and a 3D reconstructed model (82) of the first anatomical structure (21), A step of obtaining training data (87) including a set of X-ray patches (88) and a set of DRR patches (89) using training data from the aforementioned training database, wherein each X-ray patch and each DRR patch represents a different anatomical structure, and a step of A medical image transformation method comprising the step of training one CNN or a group of CNNs (27) using the training data (87).

2. The actual X-ray image (16) of the aforementioned patient, Automatically convert to at least another digitally reconstructed radiographic image (DRR) (25) of the patient that represents the second anatomical structure (26) without representing the first anatomical structure (24), By the same single operation described above, one convolutional neural network (CNN), or a group of convolutional neural networks (CNNs) (27), By directly separating the structures in the original 3D volume, the aforementioned first anatomical structure (21) can be distinguished from the aforementioned second anatomical structure (22), Each of the sets of DRRs focuses on only one anatomical structure of interest, generating multiple sets of DRRs simultaneously, or By generating a single DRR, but excluding other anatomical structures and focusing only on the single anatomical structure of interest, Converting the actual X-ray image (16) into at least two digitally reconstructed radiographic images (DRRs) (23, 25) A medical image transformation method according to claim 1, which is pre-trained to perform both or simultaneously.

3. To simultaneously generate multiple sets of DRRs, each set of DRRs focusing only on a first anatomical structure (21) that is distinguished from a second anatomical structure (22) by directly separating the structure in the original 3D volume of the subject, or By generating one DRR, but excluding other anatomical structures, and focusing only on the first anatomical structure (21) which is distinguished from the second anatomical structure (22) by directly separating the structure in the original 3D volume of the subject, By a single operation using either one convolutional neural network (CNN) or a group of convolutional neural networks (CNNs) (27) pre-trained to convert an actual X-ray image (16) into at least two digitally reconstructed radiographic images (DRRs) (23, 25), At least one or more actual X-ray images (16) of the patient including at least the first anatomical structure (21) and the second anatomical structure (22) of the patient, Both are automatically converted into at least a first digitally reconstructed radiographic image (DRR)(23) and a second digitally reconstructed radiographic image (DRR)(25) of the patient. The first digitally reconstructed radiographic image (DRR) (23) represents the first anatomical structure (24) without representing the second anatomical structure (26), A medical image transformation method wherein the second digitally reconstructed radiographic image (DRR) (25) represents the second anatomical structure (26) without representing the first anatomical structure (25), The aforementioned training, The steps include preparing a training database which includes a set of two-plane X-ray images, both frontal and lateral, obtained from one or more people, and a 3D reconstructed model (82) of the first anatomical structure (21), A step of obtaining training data (87) including a set of X-ray patches (88) and a set of DRR patches (89) using training data from the aforementioned training database, wherein each X-ray patch and each DRR patch represents a different anatomical structure, and a step of A medical image transformation method comprising the step of training one CNN or a group of CNNs (27) using the training data (87).

4. The medical image transformation method according to any one of claims 1 to 3, wherein the one convolutional neural network (CNN) or group of convolutional neural networks (CNNs) is a single generative adversarial network (GAN) (27).

5. The medical image transformation method according to claim 4, wherein the single generative adversarial network (GAN) (27) is a U-Net GAN or a Residual-Net GAN.

6. The medical image conversion method according to any one of claims 1 to 5, wherein the actual X-ray image (16) of the patient is a direct capture of the patient by an X-ray imaging device.

7. The first anatomical structure (21) and the second anatomical structure (22) of the aforementioned patient are In the actual X-ray image (16) mentioned above, adjacent to each other, Alternatively, they may even be adjacent to each other in the actual X-ray image (16), Alternatively, they are even in contact with each other in the actual X-ray image (16), The medical image transformation method according to any one of claims 1 to 6, wherein the anatomical structure is at least partially superimposed on the actual X-ray image (16).

8. A medical image conversion method according to any one of claims 1 to 7, wherein only one (23) of at least two digitally reconstructed radiographic images (DRRs) (23, 25) can be used for further processing.

9. A medical image transformation method according to any one of claims 1 to 7, wherein at least two digitally reconstructed radiographic images (DRRs) (23, 25) can all be used for further processing.

10. The actual X-ray image of the patient (151a) includes at least three anatomical structures of the patient, preferably only three anatomical structures of the patient. A medical image conversion method according to any one of claims 1 to 9, wherein the actual X-ray image (151a) is converted by the single operation into at least three separate digitally reconstructed radiographic images (DRRs) (151b, 152b, 153b) each representing at least three anatomical structures, and each of the digitally reconstructed radiographic images (DRRs) represents only one of the anatomical structures without representing any other of the anatomical structures, preferably converted into only three separate digitally reconstructed radiographic images (DRRs) each representing only the three anatomical structures.

11. The medical image transformation method according to any one of claims 1 to 10, wherein the different anatomical structures of the patient are the adjacent vertebrae (146, 147, 148) of the patient.

12. The different adjacent vertebrae (146, 147, 148) of the aforementioned patient, Region of the upper thoracic spine segment of the patient, Alternatively, in the area of ​​the patient's lower thoracic spine segment, Alternatively, in the area of ​​the patient's lumbar spine segment, Alternatively, in the area of ​​the patient's cervical vertebral segment, Alternatively, the area of ​​the patient's pelvis. The medical image transformation method according to claim 11, located within a single, identical region of the spine of any of the patients.

13. The different anatomical structures (21, 22) of the aforementioned patients, The patient's lumbar region, Alternatively, in the area of ​​the patient's lower limb, such as the femur or tibia, Or, the area of ​​the patient's knee, Alternatively, the patient's shoulder area, Alternatively, the area of ​​the patient's thoracic cavity A medical image transformation method according to any one of claims 1 to 12, located within a single, same region of any of the patients.

14. Each of the different digitally reconstructed radiographic images (DRRs) (23, 25) representing different anatomical structures (21, 22) of the patient, respectively, is simultaneously, An image with pixels representing different gray levels, At least one tag that represents anatomical information related to the anatomical structure represented by the tag and A medical image conversion method according to any one of claims 1 to 13, including the following:

15. The medical image conversion method according to claim 14, wherein the image is a square image with 256 x 256 pixels.

16. A medical image transformation method according to any one of claims 1 to 15, wherein one of the convolutional neural networks (CNNs) or groups of convolutional neural networks (CNNs) (27) is pre-trained on X-ray images (16) of a large number of diverse patients ranging from 100 to 1000, preferably from 300 to 700, and more preferably about 500, i.e., both actual X-ray images and deformations of actual X-ray images.

17. A medical image conversion method according to any one of claims 1 to 16, wherein at least both an actual frontal X-ray image of the patient (151a) and an actual lateral X-ray image of the patient (151d) are converted, and both of the X-ray images each include the same anatomical structures of the patient (146, 147, 148).

18. A method for personalizing a medical image 3D model, comprising the medical image conversion method described in claim 17, 3D generic model (120) At least one digitally reconstructed radiographic image (DRR) (145) of the patient's front view, each representing one or more of the patient's anatomical structures (146, 147, 148), At least one digitally reconstructed radiographic image (DRR) (145) of a lateral view of the patient, each representing one or more of the anatomical structures (146, 147, 148) of the patient, and Used to generate, The actual frontal X-ray image (143) is converted by the medical image conversion method described above. The images are converted into at least one digitally reconstructed radiographic image (DRR) (144) of the patient's front view, each representing one or more of the patient's anatomical structures (146, 147, 148), The actual lateral X-ray image (143) is converted by the medical image conversion method described above. The patient is converted into at least one or more digitally reconstructed radiographic images (DRRs) (144) of a lateral view of the patient, each representing one or more of the anatomical structures (146, 147, 148) of the patient, In order to generate a 3D patient-specific model from the aforementioned 3D generic model, Each of the one or more anatomical structures (146, 147, 148) of the patient is represented by at least one or more digitally reconstructed radiographic images (DRRs) (145) of the patient's front view, which are obtained from the 3D generic model, and each of the one or more anatomical structures (146, 147, 148) of the patient is represented by at least one or more digitally reconstructed radiographic images (DRRs) (144) of the patient's front view, which are obtained from the actual frontal X-ray image (143), A method wherein at least one or more digitally reconstructed radiographic images (DRRs) (145) of a lateral view of the patient, each representing one or more anatomical structures (146, 147, 148) of the patient and obtained from the 3D generic model, are mapped to at least one or more digitally reconstructed radiographic images (DRRs) (144) of a lateral view of the patient, each representing one or more anatomical structures (146, 147, 148) of the patient and obtained from actual lateral X-ray images (143).

19. The method for personalizing a medical image 3D model according to claim 18, wherein the 3D generic model is a deformable model.

20. The method for personalizing a 3D medical image model according to claim 19, wherein the deformable model is a statistical shape model.

Citation Information

Patent Citations

  • Medical imaging conversion method and associated medical imaging 3D model personalization method

    EP4150564A1

  • Cross domain medical image segmentation

    US20190259153A1

  • Cone-beam CT image enhancement using generative adversarial networks

    US20190333219A1

  • Medical imaging conversion method and associated medical imaging 3D model personalization method

    US20230177748A1

  • Method for reconstruction of a three-dimensional model of an osteo-articular structure

    WO2009056970A2