Medical imaging conversion methods and related methods for personalizing 3D medical imaging models

By using a combination of generative adversarial networks and convolutional neural networks, the problem of overlapping anatomical structures in existing technologies is solved, achieving efficient conversion and separation from real X-ray images to regional DRR, thus improving image quality and separation efficiency.

CN115996670BActive Publication Date: 2025-10-31EOS IMAGING SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080102445.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-13
Publication Date
2025-10-31
Estimated Expiration
2040-05-13

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively isolate and differentiate different anatomical structures when converting a patient's real X-ray image into a digitally reconstructed radiograph (DRR). This results in insufficient pixel classification of overlapping organ regions, making it impossible to extract the regional DRR of each organ without losing effective signals.

Method used

By employing a combination of a single generative adversarial network (GAN) and a convolutional neural network (CNN), a single operation is used to convert real X-ray images into a single or a set of regional DRRs, thereby achieving the isolation and differentiation of anatomical structures and avoiding overlap and mismatch of adjacent organs.

Benefits of technology

It achieves effective isolation and differentiation of DRR for different anatomical structures without loss of useful information, simplifies the conversion process, and improves image quality and separation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115996670B_ABST
    Figure CN115996670B_ABST
Patent Text Reader

Abstract

The present invention relates to a medical imaging conversion method that automatically performs the following conversion: using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) (27) to convert at least one or more real X-ray images (16) of the patient, which at least contain a first anatomical structure (21) and a second anatomical structure (22) of the patient, into at least one digitally reconstructed radiograph (DRR) (23) of the patient representing the first anatomical structure (24) but not the second anatomical structure (26), wherein the CNN is pre-trained to perform, or simultaneously perform, the following: distinguishing the first anatomical structure (21) from the second anatomical structure (22), and converting the real X-ray image (16) into at least one digitally reconstructed radiograph (DRR) (23).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a medical imaging conversion method.

[0002] The present invention also relates to a method for personalizing associated medical imaging 3D models using these medical imaging conversion methods. Background Technology

[0003] According to the prior art disclosed in US 2019 / 0259153 regarding medical imaging transformation methods, a method is known for transforming a patient's real X-ray image into a digitally reconstructed radiograph (DRR) using a generative adversarial network (GAN).

[0004] However, the obtained digital reconstructed radiographs (DRR) exhibit two characteristics:

[0005] First, it includes all organs present in the initial X-ray image prior to transformation by the generative adversarial network (GAN).

[0006] Second, if pixel segmentation of a specific organ is required, this is done in a subsequent segmentation step using another dedicated convolutional neural network (CNN) [see US 2019 / 0259153, page 1 §5 and page 4 §51] from the obtained digitally reconstructed radiograph (DRR).

[0007] According to the present invention, this is considered rather complex and inefficient because there are several, often even many, different anatomical structures within an organ or region of the patient's body.

[0008] The different anatomical structures, whether organs, organ parts, or multiple organ groups, are at risk of overlapping in digital reconstructed radiographs (DRRs) because they represent planar projected views of 3D structures. Summary of the Invention

[0009] The objective of this invention is to at least partially alleviate the disadvantages mentioned above.

[0010] In contrast to the listed prior art, according to the spirit of the present invention, it is considered advantageous in the process of dual DRR image matching that each DRR represents a unique organ or a unique anatomical part of an organ to avoid mismatches in adjacent or superimposed organs or organ structures.

[0011] However, in existing technologies, the transformation process does not allow for structural isolation of overlapping organ regions because pixel classification generated by segmenting the globally acquired DRR, which includes all organs, is insufficient to distinguish the corresponding amount of DRR image signal belonging to each organ or to each organ structure. Therefore, it is impossible, at least not without loss of effective signal, to subsequently extract regional DRRs representing only one organ from the globally acquired DRR.

[0012] Therefore, according to the present invention, a method is proposed to transform a patient's real X-ray image into a single regional digital reconstructed radiograph (DRR) or a set of regional digital reconstructed radiographs (DRRs) using a single generative adversarial network (GAN), wherein the GAN simultaneously achieves the following:

[0013] DRR training was conducted by directly isolating the structure within the original 3D volume.

[0014] Simultaneously, a set of several DRRs are generated, each focusing on only one anatomical structure of interest; or a single DRR is generated, but it focuses only on the one anatomical structure of interest, excluding other anatomical structures.

[0015] It is performed through a single operation of a single and unique GAN, thus simultaneously optimizing both the transformation function and the structural separation function.

[0016] This objective is achieved using a medical imaging conversion method that automatically performs the following conversion: using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) to convert at least one or more real X-ray images of a patient, containing at least a first anatomical structure and a second anatomical structure of the patient, into at least one digitally reconstructed radiograph (DRR) of the patient representing the first anatomical structure but not the second anatomical structure, through a single operation. The CNN is initially trained to perform the following, either simultaneously: distinguishing between the first anatomical structure and the second anatomical structure, and converting the real X-ray image into at least one digitally reconstructed radiograph (DRR).

[0017] Preferably, the medical imaging conversion method according to the invention also automatically performs the following conversion: converting the patient's real X-ray image into at least one other digitally reconstructed radiograph (DRR) of the patient representing the second anatomical structure but not the first anatomical structure through the same single operation, wherein one or more convolutional neural networks (CNNs) are pre-trained to perform, or simultaneously perform, the following: distinguishing between the first and second anatomical structures, and converting the real X-ray image into at least two digitally reconstructed radiographs (DRRs).

[0018] This means that, through a single transformation operation that includes two steps—transformation and structural separation—several regional DRR images are obtained from a global X-ray image representing several different organs or several different parts of an organ or several groups of organs, through a single convolutional neural network or a group of convolutional neural networks linked together thereon, respectively corresponding to and representing these several different organs or several different parts of an organ or several groups of organs.

[0019] This allows each different organ or each different part of an organ or group of organs to be represented separately from all other different organs or all other different parts of an organ or group of organs, without losing effective information in areas where there is some overlap between these different organs or different parts of an organ or group of organs, or where there may be some overlap on the real X-ray image or in areas where there will be some overlap after transformation of the global DRR.

[0020] In fact, once the global DRR is obtained from one or more real X-ray images, the overlap between several different organs or several different parts of one organ or several groups of organs can no longer be undone without loss of effective information. These overlapping several different organs or several different parts of one organ or several groups of organs can no longer be separated from each other without loss of effective information because on the DRR, that is, on the image in this transformed DRR domain, these overlapping several different organs or several different parts of one or several groups of organs are all inseparably mixed together. Any particular pixel is a mixture of different contributions from these overlapping several organs or several different parts of one organ or several groups of organs. There is no way to subsequently distinguish between these different contributions, or at least it is extremely difficult to subsequently distinguish between these different contributions, which will inevitably lead to unsatisfactory results.

[0021] Therefore, this embodiment of the present invention performs a unique, simple, and more efficient method that not only transforms images from a first domain (e.g., X-ray images) to a second domain (e.g., DRR), but also separates several different organs that were originally overlapping, or several different parts of an organ or several groups of organs, without losing effective information, i.e., without significantly degrading the quality of the original image.

[0022] These different organs, or different parts of one organ or several groups of organs, are examples of different anatomical structures.

[0023] This objective is also achieved using a medical imaging conversion method that automatically performs the following conversions: using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) to convert at least one or more real X-ray images of a patient, containing at least a first anatomical structure and a second anatomical structure of the patient, into at least first and second digitally reconstructed radiographs (DRRs) of the patient through a single operation: the first DRR represents the first anatomical structure but not the second anatomical structure, and the second DRR represents the second anatomical structure but not the first anatomical structure, wherein the CNN is initially trained to perform, or simultaneously perform, the following: distinguishing between the first and second anatomical structures, and converting real X-ray images into at least two digitally reconstructed radiographs (DRRs).

[0024] Another similar objective is achieved using a medical imaging conversion method, which automatically performs the following conversion: using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) to convert at least one or more images of the patient in a first domain containing at least a first anatomical structure and a second anatomical structure of the patient into at least one image of the patient in a second domain representing the first anatomical structure but not the second anatomical structure, by a single operation, wherein the CNN is initially trained to perform the following: distinguishing between the first anatomical structure and the second anatomical structure, and converting an image in the first domain into at least one image in the second domain.

[0025] Another similar objective is achieved using a medical imaging conversion method, which automatically performs the following conversion: using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) to convert at least one or more global images of the patient in a first domain containing at least several different anatomical structures of the patient into several regional images of the patient in a second domain, each representing a different anatomical structure, through a single operation. The CNN is initially trained to perform the following, either simultaneously: distinguishing the anatomical structures from each other, and converting images in the first domain into at least one image in the second domain.

[0026] Another supplementary objective of the invention will be related to the objectives previously listed in the invention, as it is a preferred application of the invention. In fact, the invention also relates to a related method for personalizing 3D models of medical imaging, which uses and leverages the medical imaging transformation method that is a primary objective of the invention, and particularly utilizes its ability to separate different anatomical structures from each other without losing effective information due to possible overlap between these different anatomical structures.

[0027] This supplementary objective of the present invention is a method for personalizing 3D medical imaging models, comprising a medical imaging conversion method according to the present invention, wherein: a 3D universal model is used to generate: at least one or more digitally reconstructed radiographs (DRRs) of the patient's frontal view representing one or more anatomical structures, and at least one or more digitally reconstructed radiographs (DRRs) of the patient's side view representing one or more anatomical structures, wherein a true frontal X-ray image is converted by the medical imaging conversion method into: at least one or more digitally reconstructed radiographs (DRRs) of the patient's frontal view representing one or more anatomical structures, and a true lateral X-ray image is converted by the medical imaging conversion method into: at least one or more digitally reconstructed radiographs of the patient's side view representing one or more anatomical structures. The at least one or more digitally reconstructed radiographs (DRRs) representing the patient's one or more anatomical structures and obtained from the 3D universal model are mapped to the at least one or more DRRs representing the patient's one or more anatomical structures and obtained from the frontal real X-ray image, respectively; and the at least one or more DRRs representing the patient's one or more anatomical structures and obtained from the 3D universal model are mapped to the at least one or more DRRs representing the patient's one or more anatomical structures and obtained from the side real X-ray image, respectively, in order to generate a 3D patient-specific model from the 3D universal model.

[0028] Mapping can be performed via flexible mapping or flexible alignment.

[0029] Preferably, the 3D general model is a surface or volume representing a previous shape that can be deformed using a deformation algorithm.

[0030] Preferably, the process of deforming a general model to obtain a personalized model for the patient is called a deformation algorithm. The association between the general model and the deformation algorithm is a deformable model.

[0031] Preferably, the deformable model is a statistical shape model (SSM).

[0032] Statistical shape models are deformable models that capture shape deformation patterns based on statistical data extracted from a training database. They can directly infer plausible shape instances from a simplified set of parameters.

[0033] Therefore, the reconstructed 3D patient-specific model:

[0034] It is simpler than any model derived from computed tomography (CT) or magnetic resonance imaging (MRI).

[0035] ○ Because it is derived from a 3D general or deformable model and two simple orthogonal real X-ray images,

[0036] More specific and more precise

[0037] ○ Because from the outset, it implicitly and initially contains volume information attributable to the front and side initial views to be transformed, as well as related to the 3D general model, the deformable model, or the statistical shape model (depending on the specific case).

[0038] Different anatomical structures can be separated from each other without losing useful information, which contrasts with the conversion methods used in the listed prior art, which are also more complex to implement.

[0039] The prior art previously listed in US 2019 / 0259153 does not disclose this supplementary objective as in this invention, such as 3D / 2D elastic alignment, and in particular does not disclose cross-domain image similarity, such as for multi-structure 3D / 2D elastic alignment, because the application domain of this invention is completely different. It is based on performing "cause-and-effect" segmentation on dense image-to-image networks using task-driven generative adversarial networks (GANs), which has already pre-converted global X-ray images covering several anatomical structures into DRRs covering the same multiple anatomical structures.

[0040] The robustness of intensity-based 3D / 2D alignment of a 3D model over a set of 2D X-ray images also depends on the image correspondence quality between the actual image and the digitally reconstructed radiograph (DRR) generated from the 3D model. In this multimodal alignment scenario, the trend towards improving the similarity level between two images is to generate the most realistic DRR possible, i.e., as close as possible to the actual X-ray image. This involves two key aspects: first, 3D models of soft tissue and bone rich in accurate density information of all involved structures; and second, complex projection processes that consider the physical phenomena of interaction with matter.

[0041] Conversely, in the method proposed in embodiments of the present invention, the relative approach of introducing actual X-ray images into the DRR image domain uses a cross-modal image-to-image transformation. The prior step of adding an image-to-image transformation based on a GAN pixel-to-pixel model allows for a simple and fast DRR projection process without the need for complex phenomenon simulation. In fact, even with simple metrics, similarity measures become effective because the two images to be matched belong to the same domain and contain substantially the same kind of information. The proposed medical imaging transformation method also addresses the well-known problem of aligning targets in scenes composed of multiple targets. Using a separate convolutional neural network (CNN) output channel for each structure in the XRAY-to-DRR converter allows superimposed neighboring targets to be separated from each other and avoids mismatch of similar structures during the alignment process. The method proposed in embodiments of the present invention is applied to the challenging 3D / 2D elastic alignment of vertebral 3D models in biplane radiographs of the spine. The step of using XRAY-to-DRR transformation enhances the alignment results and reduces the dependence on the selection of similarity measures, as multimodal alignment becomes unimodal.

[0042] Preferred embodiments include one or more of the following features, which may be viewed individually or together, and may be partially or completely combined with any of the previously listed objectives of the invention.

[0043] Preferably, the convolutional neural network (CNN) or a group of convolutional neural networks (CNN) is a single generative adversarial network (GAN).

[0044] Therefore, the simplicity of the proposed conversion method is further improved while maintaining its effectiveness.

[0045] Preferably, the single generative adversarial network (GAN) is a U-Net GAN or a residual-Net GAN.

[0046] Preferably, the actual X-ray image of the patient is a direct capture of the patient by an X-ray imaging device.

[0047] Therefore, the images in the first domain to be converted are generally readily available and completely patient-specific.

[0048] Preferably, the patient's first and second anatomical structures are anatomical structures that are adjacent to each other on the real X-ray image, or even close to each other on the real X-ray image, or even in contact with each other on the real X-ray image, or even at least partially superimposed on the real X-ray image.

[0049] Therefore, the proposed conversion method is of more interest than the fact that the different anatomical structures to be separated are actually close to each other, because the amount of mixed information that is no longer separable between them will be higher.

[0050] Preferably, only one of the at least two digitally reconstructed radiographs (DRRs) can be used for further processing.

[0051] Therefore, the proposed conversion method can be used even if a practicing physician actually needs only one DRR.

[0052] Preferably, all of the at least two digitally reconstructed radiographs (DRRs) can be used for further processing.

[0053] Therefore, the proposed conversion method can be used, especially in cases where all DRRs are actually available to practicing physicians.

[0054] Preferably, the patient's actual X-ray image contains at least three anatomical structures of the patient, and preferably only three anatomical structures of the patient, the actual X-ray image being converted by the single operation into at least three separate digital reconstructed radiographs (DRRs) representing the at least three anatomical structures respectively, each of the DRRs representing only one of the anatomical structures and not any other anatomical structures; and preferably converted into only three separate DRRs representing only the three anatomical structures respectively.

[0055] Therefore, the proposed conversion method is optimized in two ways:

[0056] First, using its two closest neighbors (especially in linear structures such as the patient's spine) helps to better isolate and individualize the specific anatomical structures of interest.

[0057] However, secondly, using only its two closest neighbors helps to perform the operation in a still relatively simple implementation.

[0058] Preferably, the convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) has been initially trained using a set of training groups including: a real X-ray image and at least one or more corresponding digitally reconstructed radiographs (DRRs) representing only one of the anatomical structures but not other anatomical structures of the patient.

[0059] Therefore, the proposed transformation method can be trained using paired images in the first domain to be transformed and the second domain to be transformed into.

[0060] Preferably, the convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) has been initially trained using a set of real X-ray images and several subsets of at least one or more digitally reconstructed radiographs (DRRs), each of which represents only one of the anatomical structures but not the other anatomical structures of the patient.

[0061] Therefore, the proposed transformation method can be trained using unpaired images in the first domain to be transformed and the second domain to be transformed into.

[0062] Preferably, the convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) has been initially trained with a set of the following training groups: both a frontal real X-ray image and a side real X-ray image, and at least one or more subsets of the corresponding frontal and side digital reconstructed radiographs (DRRs), each of the subsets representing only one of the anatomical structures but not the other anatomical structures of the patient.

[0063] Therefore, the proposed transformation method can be trained using paired groups of frontal and lateral images in the first domain to be transformed and the second domain to be transformed into, with each pair of frontal and lateral images being transformed in several pairs of frontal and lateral images, and a pair of transformed frontal and lateral images corresponding to an anatomical part separated from other anatomical parts.

[0064] Preferably, the digital reconstructed radiographs (DRRs) of the training group are derived from a 3D model specifically tailored to the patient via adaptation of the model to two real X-ray images acquired along two orthogonal directions.

[0065] Therefore, because it is precisely designed for specific patients under careful consideration, the trade-off between the simplicity of creating a model to train DRR and the effectiveness of the model is optimized.

[0066] Preferably, the different anatomical structures of the patient are the patient's adjacent vertebrae.

[0067] Therefore, the proposed conversion method is of more interest than the fact that the different anatomical structures to be separated are actually close to each other, because the amount of mixed information that is no longer separable between them will be higher.

[0068] Preferably, the different and adjacent vertebrae of the patient are located in a single and identical region of the patient's spine, either the upper thoracic region, the lower thoracic region, the lumbar region, the cervical region, or the pelvic region.

[0069] Preferably, the different anatomical structures of the patient are located in a single and identical area of ​​the patient among the following: the patient's hip area, or the patient's lower limb area (e.g., femur or tibia), or the patient's knee area, or the patient's shoulder area, or the patient's ribcage area.

[0070] Preferably, each of the different digital reconstructed radiographs (DRRs) representing the different anatomical structures of the patient simultaneously includes: an image with pixels presenting different gray levels, and at least one label representing anatomical information related to the anatomical structure it represents.

[0071] Therefore, this label will be used to distinguish different anatomical structures and separate them from each other. This label is a simple and efficient way to separate information corresponding to different anatomical structures in the first domain (here, X-ray images, preferably frontal and side X-ray images) before they are converted in the second domain (here, DRR) while it is still possible to do so (i.e., in 3D volume), in which it will no longer be possible to distinguish these different anatomical structures and separate them from each other if this distinction has not been made before.

[0072] Preferably, the image is a 256×256 pixel square image.

[0073] Therefore, this represents a good trade-off between image quality and processing complexity and required storage capacity.

[0074] Preferably, the convolutional neural network (CNN) or a group of convolutional neural networks (CNN) has been initially trained on 100 to 1000, and more preferably 300 to 700, and more preferably about 500 x-ray images (both real x-ray images and transformed forms of real x-ray images) of several different patients.

[0075] Therefore, this is a simple and effective way to train a very large number of training images that are sufficiently different from each other, while having a very limited amount of data available for disposal to build these training images.

[0076] The range of training images is also optimized; in fact, when using a reasonable range of training images, near-optimal efficiency is achieved at a very reasonable cost.

[0077] Preferably, at least a frontal true X-ray image of the patient and a side true X-ray image of the patient are converted, each of the two X-ray images containing the same anatomical structures of the patient.

[0078] Therefore, using both frontal and lateral real X-ray images as the images to be converted allows for obtaining frontal and lateral DRRs and improves the accuracy of the created DRRs, because the anatomical structures can be known more accurately from two orthogonal views, and may even be reconstructed in three-dimensional space (3D) by means of, for example, a 3D universal model, a deformable model, or a statistical shape model (depending on the case).

[0079] Further features and advantages of the invention will become apparent from the following description of embodiments of the invention given as non-limiting examples, with reference to the accompanying drawings. Attached Figure Description

[0080] Figure 1A This demonstrates the first step in forming a global DRR for training a GAN during the training phase, based on existing techniques.

[0081] Figure 1B This demonstrates the second step of performing segmentation during the training phase, based on existing technology, to separate anatomical structures from each other.

[0082] Figure 1C illustrates a method for converting X-ray images into segmented images of different anatomical structures using existing techniques.

[0083] Figure 2A This demonstrates the first step of forming a regional DRR for training a GAN during the training phase, according to an embodiment of the invention.

[0084] Figure 2B This demonstrates a second step in forming another regional DRR during the training phase of a second anatomical structure for training a GAN, according to an embodiment of the invention.

[0085] Figure 2C This invention demonstrates a method for converting X-ray images into several different regional DRRs representing several different anatomical structures.

[0086] Figure 3A This demonstrates an example of 3D / 2D alignment based on a certain existing technology.

[0087] Figure 3B An example of 3D / 2D alignment according to an embodiment of the present invention is shown.

[0088] Figure 4A This demonstrates an example of 3D / 2D alignment based on a certain existing technology.

[0089] Figure 4B An example of 3D / 2D alignment according to an embodiment of the present invention is shown.

[0090] Figure 5A global flowchart illustrating an example of performing a 3D / 2D alignment method according to an embodiment of the present invention is provided.

[0091] Figure 6 A more detailed flowchart illustrating an example of a training phase prior to implementing a 3D / 2D alignment method according to an embodiment of the present invention is provided.

[0092] Figure 7 A more detailed flowchart illustrating an example of performing a 3D / 2D alignment method according to an embodiment of the present invention is provided.

[0093] Figure 8A An example illustrating the principle of forming an inner layer from an outer layer using prior cortical thickness and normal vector, according to an embodiment of the invention.

[0094] Figure 8B An example of interpolation thickness mapping for the entire mesh structure of the L1 vertebra is shown according to an embodiment of the present invention.

[0095] Figure 9 An example demonstrating the principle of ray projection through a double-layered mesh structure.

[0096] Figure 10 An example of an XRAY-to-DRR GAN network architecture and training according to an embodiment of the present invention is shown.

[0097] Figure 11A An example of GaN-based XRAY to DRR conversion is shown, from an actual X-ray image to a three-channel DRR image corresponding to a nearby vertebral structure.

[0098] Figure 11B The transition (also referred to as the shift) from the X-ray domain to the DRR domain under adverse conditions is demonstrated to illustrate the robustness of the proposed method according to embodiments of the present invention.

[0099] Figure 12A This demonstrates the target alignment error that varies depending on the initial pose shift along the vertical axis.

[0100] Figure 12B An example showing vertical displacement toward the top of the vertebra.

[0101] Figure 13A This section shows the anatomical region used to calculate the statistical distance from nodes to the surface.

[0102] Figure 13B This displays the distance mapping for the maximum error calculated for the L3 vertebra.

[0103] Figure 14A Shows the cost values ​​of PA and LAT views using GAN DRR compared to actual X-ray images.

[0104] Figure 14B Show the similarity values ​​of PA and LAT views using GAN DRR compared to the actual X-ray image.

[0105] Figure 14C Showing compared to using actual X-ray images Figure 14D When using DRR generated by GAN, a better fit was observed for the alignment results.

[0106] Figure 15 The display is as follows Figure 14B Similar results were obtained from studies that use the GNCC (gradient normalized cross-correlation) metric.

[0107] Figure 16 The results of the target alignment error (TRE) test in the Z direction (vertical image direction) are displayed. Detailed Implementation

[0108] All subsequent descriptions will refer to a body representation having several different organs, but may also be applied to a body representation having several different groups of organs or even an organ representation having several different parts of that organ.

[0109] The posterior-anterior (PA) view is equivalent to the front (FRT) view, and both contrast with the side (LAT) view because the side view is along a direction orthogonal to the common direction of the posterior-anterior and front views.

[0110] Figure 1A This demonstrates the first step in forming a global DRR for training a GAN during the training phase, based on existing techniques.

[0111] The X-ray projection is performed from DRR X-ray source 1 through a 3D volume 2 containing two different organs (first organ 3 and second organ 4).

[0112] The first organ 3 and the second organ 4 have planar projections on the global DRR 5, which are respectively contributed by the DRR image signal 6 from the first organ 3 and the DRR image signal 7 from the second organ 4.

[0113] Image regions 6 and 7 have an intersection region 8 in which the two signals from organs 3 and 4 are superimposed.

[0114] In this intersection zone 8, it is impossible to know which part of the signal comes from organ 3 and which part of the signal comes from organ 4.

[0115] Figure 1B This demonstrates the second step of performing segmentation during the training phase, based on existing technology, to separate anatomical structures from each other.

[0116] From DRR x-ray source 1, segmentation of the first organ 3 and the second organ 4 is performed so that the first organ 3 and the second organ 4 are separated from each other.

[0117] The segments of the first organ 3 and the second organ 4 are respectively given as the first label image 12 and the second label image 13, on which the first trace 14 of the first organ 3 and the second trace 15 of the second organ 4 are respectively present.

[0118] However, even when these segmented first traces 14 and second traces 15 are subsequently applied to the global DRR 5, they do not produce a first DRR for the first organ 3 and a second DRR for the second organ 4, because the mixture of signals within the intersection region 8 can no longer be separated between their respective original contributions, i.e., between the respective contributions from the first organ 3 (on the one hand) and from the second organ 4 (on the other hand).

[0119] Therefore, applying segmentation to the global DRR will not produce several different local DRRs for different organs, because in all intersecting regions 8, the mixed signals from several different organs can no longer be separated into their original contributions. There is already a loss of effective signal, which cannot be recovered or will at least become very difficult to recover.

[0120] Figure 1C illustrates a method for converting X-ray images into segmented images of different anatomical structures using existing techniques.

[0121] X-ray image 16 is converted by GAN 17 into a transformed global DRR 5, which is then segmented by DI2ICNN 18 as follows:

[0122] The first binary mask 12 of this transformed global DRR 5, corresponding to segment 14 of the first organ 3,

[0123] The second binary mask 13 of this transformed global DRR 5 corresponds to segment 15 of the second organ 4.

[0124] Obtaining a regional DRR image corresponding to the first organ 3 or the second organ 4 separately from the first binary mask 12 or the second binary mask 13 will be accompanied by the loss of effective signals in the intersection region 8.

[0125] Figure 2A This demonstrates the first step of forming a regional DRR for training a GAN during the training phase, according to an embodiment of the invention.

[0126] From DRR x-ray source 1, a ray projection is performed through a 3D volume 20 containing only the first organ 21, which has been separated from the second organ 22, which is different from the first organ 21.

[0127] The first organ 21 has a planar projection on a local DRR, namely a DRR image 23 dedicated to the first organ 21 and a DRR image 24 representing a planar projection that precisely corresponds to the first organ 21, without loss of effective signal.

[0128] In DRR image 24, all signals originate from the first organ 21, and no signals originate from the second organ 22. No valid signals are lost regarding the planar projection of the first organ 21.

[0129] Figure 2B This demonstrates a second step in forming another regional DRR during the training phase of a second anatomical structure for training a GAN, according to an embodiment of the invention.

[0130] From DRR x-ray source 1, a ray projection is performed through a 3D volume 20 containing only the second organ 22, which has been separated from the first organ 21, which is different from the second organ 22.

[0131] The second organ 22 has a planar projection on a local DRR, namely a DRR image 25 dedicated to the second organ 22 and a DRR image 26 representing a planar projection that precisely corresponds to the second organ 22, without loss of effective signal.

[0132] In DRR image 26, all signals originate from the second organ 22, and no signals originate from the first organ 21. No valid signals are lost regarding the planar projection of the second organ 22.

[0133] Figure 2C This invention demonstrates a method for converting X-ray images into several different regional DRRs representing several different anatomical structures.

[0134] X-ray image 16 was converted by GAN 27 to correspond to the following two separate converted regional DRRs 23 and 25, respectively:

[0135] The first local DRR precisely corresponds to the planar projection of the first organ 21 without losing the effective signal associated with the first organ 21, which contrasts with image 12 in FIG1C.

[0136] The second local DRR precisely corresponds to the planar projection of the second organ 22 without losing the effective signal associated with the second organ 22, which contrasts with image 13 in FIG1C.

[0137] Intensity-based elastic alignment of 3D models with 2D planar images is one of the methods used as a key step in the 3D reconstruction process. This alignment is similar to multimodal alignment and relies on maximizing the similarity between the actual X-ray image and the digitally reconstructed radiograph (DRR) generated from the 3D model. Given that a 3D model typically does not contain all the information (e.g., it does not always contain density), two images, such as the actual X-ray and the DRR, are significantly different from each other, making the optimization process very complex and often unreliable. The standard solution is to make the DRR generation process as close as possible to the image formation process by adding additional information and simulating complex physics. In the algorithm proposed in embodiments of the present invention, the opposite approach is used by transforming the actual X-ray image into a DRR-like image to facilitate image matching.

[0138] A proposed image transformation step converts X-ray images into DRR-like images, allowing background and noise removal from the original image and bringing it to the DRR domain. This makes matching between two images simpler and more efficient because the images have similar properties, and it improves upon optimization processes that even use standard similarity metrics.

[0139] This is achieved by training a pixel-to-pixel deep neural network based on a U-Net convolutional neural network (CNN) that allows images to be transformed from one domain to another. The use of X-ray image transformation to DRR-style images facilitates 3D / 2D alignment of the mesh structure. Furthermore, the proposed CNN-based transformer can separate neighboring bone structures by outputting a transformed image for each bone structure to better handle pivoted bone structures.

[0140] Embodiments of the present invention are applied to the challenging 3D reconstruction of spinal structures from a biplane x-ray imaging modality. The spinal structure is a periodic, multi-structured entity capable of anatomical deformation. Pixel-to-pixel networks convert actual x-ray images into virtual x-ray images, which allows:

[0141] First, improve image-to-image correspondence and alignment performance.

[0142] And secondly, identify and isolate specific structures from similar neighbors to avoid mismatch.

[0143] pass Figure 3A The above represents existing technologies and Figure 3B The above comparison between embodiments of the present invention illustrates the principle of this 3D / 2D alignment. Figure 3A This demonstrates an example of 3D / 2D alignment based on a certain existing technology. Figure 3B An example of 3D / 2D alignment according to an embodiment of the present invention is shown.

[0144] Figure 3A The image above represents a classic method of 3D / 2D alignment based on a prior art, using the similarity between the DRR generated from the mesh structure and the X-ray image. Unfortunately, it is not easy to distinguish the vertebra 31 of interest from the rest of the spine 30.

[0145] Figure 3B The alignment process proposed in the embodiments of the present invention uses a previous image-to-image transformation to convert the target image so that image correspondence in the alignment similarity function is easy and mismatches in adjacent structures are avoided. Here, the vertebra 31 of interest has been separated from the rest of the spine 30.

[0146] It also addresses the front and side planar projections of 3D models through... Figure 4A The above represents existing technologies and Figure 4B The comparison between the embodiments of the invention shown above illustrates the application of this principle of 3D / 2D alignment. Figure 4A This demonstrates an example of 3D / 2D alignment based on a certain existing technology. Figure 4B An example of 3D / 2D alignment according to an embodiment of the present invention is shown.

[0147] Figure 4A Above, according to a certain prior art, a classic method of 3D / 2D alignment of a 3D model is shown, using the similarity between a DRR projection generated from a 3D mesh structure and a biplane x-ray image. The 3D model 40 is located between a first frontal x-ray source 41 and a biplane x-ray image 43 imaged above it by the frontal projection of the 3D model 40. Unfortunately, the frontal image of the vertebra of interest 47 is mixed with the rest of the spine 45. The 3D model 40 is also located between a second lateral x-ray source 42 and a biplane x-ray image 44 imaged above it by the lateral projection of the 3D model 40. Unfortunately, the biplane image of the vertebra of interest 48 is mixed with the rest of the spine 46.

[0148] Intensity-based methods (i.e., icon alignment) aim to find the optimal model parameters that maximize the similarity between an X-ray image and a digitally reconstructed radiograph (DRR) generated from a 3D model. They do not require extraction from the image and are more robust and flexible. However, they involve using sophisticated models and algorithms to generate DRRs that are as realistic as possible—that is, as close as possible to the actual X-ray image—to have an effective similarity metric. Because they are the criteria optimized during the alignment process, they should reflect the actual degree of matching of structures on both the changing image and the target image, and should be robust to perturbations to avoid truncation in local minima. The main perturbation is that the two images to be matched have different modalities, even if the generated DRR behaves fairly realistically. Sophisticated similarity metrics should be used to compare images in domain A (e.g., real X-ray images) with images in domain B (e.g., DRR images). These domain differences limit the performance of model alignment on real clinical data. Alignment in planar X-rays has another additional specificity: the presence of overlap between adjacent structures and the environment, which leads to mismatches where the initial 3D model is not close enough (e.g., ...). Figure 4A This is especially true in cases where it can be seen from the above.

[0149] Figure 4B Above, a proposed method according to an embodiment of the invention is shown. The proposed XRAY-to-DRR conversion process utilizes a single structure to convert an X-ray target image to DRR in order to improve similarity levels and prevent mismatches in adjacent structures during alignment. A 3D model 50 is positioned between a first frontal X-ray source 41 and the frontal projection of the 3D model 40 onto a planar DRR conversion image 51. Here, it is much easier to perform the superposition of the frontal projection 55 of the 3D model 40 for the vertebra of interest, where the same vertebra 53 remains independent and separated from the original X-ray image converted to DRR (the rest of the vertebra, i.e., other adjacent vertebrae, has been removed). The 3D model 50 is also positioned between a second lateral X-ray source 42 and the lateral projection of the 3D model 40 onto a planar DRR conversion image 52. Here, similarly, it is much easier to perform the superposition of the lateral projection 56 of the 3D model 40 for the vertebra of interest, with the same vertebra 54 of interest remaining independent and separated from the original X-ray image converted to DRR (the rest of the spine, i.e., other adjacent vertebrae, has been removed).

[0150] Figure 4BThe proposed method in the embodiments of the present invention described above takes the opposite approach, thereby changing the paradigm. In fact, instead of improving similarity metrics or DRR x-ray simulation, it explores a way to convert x-ray images into DRR-like images using a cross-modal image-to-image transformation model based on pixel-to-pixel generative adversarial networks (GANs). XRAY-to-DRR transformation simplifies the entire process and allows for the use of simpler DRR generation and similarity metrics, such as... Figure 4B As can be seen above.

[0151] This application experiment utilizes challenging 3D reconstructions of multi-target and periodic structures of the vertebral skeleton from a biplane X-ray imaging modality. An XRAY-to-DRR transformation step is added prior to elastic 3D / 2D alignment to convert the actual X-ray image into a DRR-like image with isolated anatomical structures. This transformation step facilitates 3D / 2D alignment of reticular structures because image matching is more efficient, even when using standard similarity metrics, as the two images have similar properties and similarity measurements are performed only on the isolated anatomical structures of interest.

[0152] Figure 5 A global flowchart illustrating an example of performing a 3D / 2D alignment method according to an embodiment of the present invention is provided. This flowchart 60 converts an X-ray image into a DRR image. A frontal X-ray image 61 is converted into a frontal DRR image 63 by a GAN 62 with a frontal weight set. A side X-ray image 64 is converted into a side DRR image 66 by a GAN 65 with a side weight set. Starting from both an initial 3D model 67 and an initial DRR overlay 68, both a modified aligned 3D model 74 and a modified DRR overlay 75 are obtained through iterative optimization 69. This iterative optimization includes a loop comprising the following consecutive steps: first, a 3D model instance step 70; then, a 3D model DRR projection step 71; then, a similarity scoring step 72; then, an optimization step 73; and finally, a return to the 3D model instance step 70 with updated parameters.

[0153] To effectively solve DRR image matching on X-ray photographs with overlapping targets, and to improve robustness, alignment accuracy, and reduce computational complexity, a proposed method is to introduce a previous image-to-image transformation step using a pre-trained pixel-to-pixel GAN ​​network. This step converts the X-ray image into a DRR image, as shown below. Figure 5As can be seen above, the resulting XRAY to DRR transformation network allows for efficient measurement of cross-image similarity and uses a simple DRR generation algorithm because the two images to be matched—the different DRRs (DRR3DM) generated from the 3D mesh structure and the target DRR—belong to the same domain. The proposed XRAY to DRR transformation step converts the X-ray image into a DRR3DM-like image with a uniform background and noise, and all soft tissue and adjacent structures are removed, such as... Figure 4B As can be seen above. In fact, the image converter is trained to separate neighboring bone structures by outputting a converted DRR image for each bone structure. The problem of aligning single-structure 3D models in scenes consisting of multiple targets (e.g., pivoted bone structures) is thus solved.

[0154] Figure 6 A more detailed flowchart illustrating an example of a training phase prior to implementing a 3D / 2D alignment method according to an embodiment of the invention is provided. The proposed method in this embodiment requires a U-Net converter provided in the previous training phase to transform x-ray specks into DRR specks to facilitate 3D / 2D alignment.

[0155] The training database 80 contains a collection 81 of biplane X-ray images of both the frontal (also known as posteroanterior) and lateral views, as well as a 3D reconstruction model 82 of the vertebrae.

[0156] The learning data extraction process 83 performs DRR calculation operations 84 via a ray projection model, for example using a process similar to that described in patent application WO 2009 / 056970, as well as region of interest (ROI) extraction and resampling operations 85. A learning phase 83 is required to train the U-Net XRAYS to DRR converter. Both the DRR calculation operation 84 and the ROI extraction and resampling operation 85 use a set 81 of biplane X-ray images and a 3D reconstruction model 82 of the vertebrae as input. The DRR calculation operation 84 produces a set 86 of biplane DRRs of the model as output, while the ROI extraction and resampling operation 85 uses this set 86 of biplane DRRs of the model as input.

[0157] The training data 87 contains a set of x-ray spots 88 and a set of DRR spots 89, which represent different anatomical structures. This set of x-ray spots 88 and this set of DRR spots 89 are generated by the ROI extraction and resampling operation 85. The data 87 required for this training is a pair of sets of spots 88 and 89, both of which are square images of size 256×256 pixels.

[0158] This set of X-ray specks (88) and this set of DRR specks (89) are both used as inputs during the training operation 90 of the U-Net converter using GAN to generate a set of U-Net converter weight parameters (91).

[0159] Spot 88, labeled X, was extracted from the x-ray biplane image 81. Spot 89, labeled Y, was extracted from the biplane DRR 86, which was generated from the reconstructed 3D model.

[0160] The training aims to fit these data to the U-Net model: Y = predict(X,W) + ε, where ε is the predicted residual and W is the generated training weights, which are the neural network parameters.

[0161] X and Y are 4D tensors, with the following size constraints:

[0162] X: [N, 256, 256, M]: where N is the total number of spots, and M is the number of input channels, for example, among the following:

[0163] M=1: The spot belongs to either the frontal (FRT) or side (LAT) view.

[0164] M=2: Used to train the joint model FRT+LAT

[0165] Y: [N,256,256,K×M]: where N is the total number of spots, and K is the number of anatomical structures (DRR images) in the output channels, for example, between the following:

[0166] K=1: Only one anatomical structure.

[0167] K=3: The upper / middle and bottom adjacent vertebrae are models of 3D / 2D alignment of the vertebrae proposed for use in the embodiments.

[0168] K=4: Models of the left and right femurs and tibias, for example, in a LAT view.

[0169] The training data (i.e., spots 88 and 89) were extracted as follows:

[0170] M patients were used for training. Each patient had the following data:

[0171] ○ FRT and LAT calibration X-ray photographs DICOM images,

[0172] ○ 3D modeling of the spine.

[0173] The M patients are separated into two different sets: a training set for modifying the neural network weights, and a test set for investigating training convergence, training curves, and selecting the best model with the best generalized inference / prediction.

[0174] For each patient, extract a set of P spots such that N = M × P:

[0175] ○ Generate both PA (post-anterior) and LAT DRR images using expert 3D models.

[0176] ○ Extract a set of P spots of size 256×256 from both the DRR(Y) and the actual image (X).

[0177] The location of the spots within the range [25, 25, 25] (relative to the uniform method) was randomly shifted from the center of the 3D cone to artificially increase the dataset size by shifting the structure on the image. Further data augmentation was performed using the following method:

[0178] ■ Global rotation around the 2D VBC: [-20 20 degrees]

[0179] ■ Globally adjusted proportionally [0.9, 1.1]

[0180] ■ Local deformation of the image with variations around the center of the end plate in the range of [-5, +5 degrees].

[0181] Figure 7 Demonstrating embodiments of the present invention compared to Figure 5 A more detailed flowchart is provided, illustrating an example of performing a 3D / 2D alignment method. The iterative alignment process depicted aims to find model parameters that maximize the similarity between the transformed target DRR image and the moving DRR image (DRR3DM) calculated from the 3D model, in both the back-anterior (PA) (also known as the frontal) and side (LAT) views. This more detailed flowchart includes the process of converting an X-ray image into a DRR image 102.

[0182] The vertebral statistical shape model (SSM) 112 and the vertebral deformable model 113 (each consisting of a free general model, a dictionary of labeled 3D points, and a set of parametric deformation handles for controlling the least squares deformation) are used as input by the initialization process 101 to perform an automatic global fitting operation 103 to provide a first estimate 115 of the 3D model. This first estimate 115 of the 3D model is inserted into the initial parameter set 116 for use by the iterative optimization loop 100.

[0183] The first estimate 115 of this 3D model is also used, along with a set 114 of biplane x-ray images (front and side views), as input to the target generation process 102. This target generation process 102 performs a pixel-to-pixel / pixel-to-pixel transformation. The target generation process 102 comprises the following sequential steps: first, a ROI extraction and resampling step 104, starting from the first estimate 115 of the 3D model and using the set 114 of biplane x-ray images as input, to generate a set 105 of x-ray spots labeled X; and subsequently, a DRR inference step 106 using U-Net converter weight parameters 91, 117, thereby generating a set 107 of predicted DRR spots labeled Y.

[0184] The iterative optimization loop 100 of this 3D / 2D alignment process includes a loop of the following consecutive steps:

[0185] The first step 108 in generating the mesh structure instance uses both the initial parameter set 116 and the local vertebral statistical shape model (SSM) as input.

[0186] The second step 109 involves calculating the DRR (Diagnosis Related Ratio) for the correct vertebral ROI and spot size using ray projection.

[0187] The third step 110 of the similarity scoring also uses a set 107 of predicted DRR spots labeled Y.

[0188] The optimization is performed using the Covariance Matrix Adaptive Evolution Strategy (CMA-ES), and then the updated parameters are returned to the fourth step of the first step 108, thereby producing the optimized parameters 119 corresponding to the aligned model as the final output at the end of the iteration loop.

[0189] The novel approach proposed in embodiments of the present invention for resilient and even rigid 3D / 2D alignment of 3D models on planar ray photographs comprises training an XRAY-to-DRR converter network and a fast DRR generation algorithm computed from a skeletal two-layer mesh model that does not require accurate tissue density. Maximizing the similarity (in step 110) between the different DRRs generated from the 3D mesh structure (in step 108) (DRR3DM) and the converted DRR image (DRRGAN) used as the target (in step 107) allows for easier finding of the parameters of the optimal model (in step 119) because: firstly, image-image correspondence is in a unique domain and is facilitated; and secondly, mismatches on neighboring structures are avoided due to structural separation in the XRAY-to-DRR conversion. The iterative alignment process 100 aims to maximize the similarity between DRR3DM and DRRGAN simultaneously in all views.

[0190] Formally, this can be viewed as how to maximize the following cost function (Equation 1) to find the optimal parameters controlling the pose and shape of the 3D model. 119:

[0191]

[0192] in Equation 2 is the image of the changes calculated for view υ (across a total of L views). With target image Any similarity function between the two images represents a region of interest (ROI) in the original image surrounding the target structure to be aligned. (The image is then modified.) It depends on the model's parameter vector p, and is calculated as the DRR projection function P on the view υ. v ()(Equation 3):

[0193]

[0194] Wherein Ψ(p) (Equation 4) generates a 3D model instance controlled by the parameter vector p. Target image Limited to transformed images predicted using U-Net (Equation 5):

[0195]

[0196] in Let f(I, W) be the ROI blob of the original x-rays of view υ, and let f(I, W) represent the feedforward inference of a GaN-based trained U-Net with CNN parameters W and input image I. In this context, a biplane ray-photograph system is used, where L = 2 views (PA and LAT). During the iteration process, methods such as... Figure 7The optimizer presented in the flowchart of the proposed method above maximizes the cost function (Equation 1).

[0197] The following paragraphs provide a method for DRR projection from a two-layer mesh surface according to embodiments of the present invention, as well as a CNN architecture and training for an XRAY-to-DRR converter. While the method is illustrated for a spinal structure, it can be applied to 3D / 2D alignment of any (multi-)structure 3D model on a calibrated planar view.

[0198] The DRR projection forming a double-layer mesh structure is now described. It is preferably implemented using a process similar to that described in patent application WO2009 / 056970. This is from the projection function P v (S) The process of generating DRR on a 3D surface mesh structure S on a related view υ (Equation 3). A virtual X-ray image (DRR) is calculated using ray projection intersecting with a 3D surface model consisting of two layers that define two media: cortical and cancellous bone media. The media-separating surface is used to account for two different factors corresponding to attenuation of the two material properties. The cortical structure corresponds to the highest X-ray energy absorption and appears brighter in the radiograph. The 3D model is represented by a mesh surface S = {V, F}, which consists of a set of 3D vertices with (x, y, z) coordinates. and the set of faces with vertex indices of the defined triangular faces. limited.

[0199] Figure 8A An example illustrating the principle of forming an inner layer from an outer layer using a prior cortical thickness and a normal vector according to an embodiment of the invention is shown. Visible on the mesh surface 120 is the sum of the prior cortical thickness ti 121 and the normal vector Ni 122.

[0200] A double-layered mesh surface is created by adding an inner layer to the mesh structure 120. For each surface vertex V i Use Equation 6 to calculate the internal vertices:

[0201]

[0202] in For vertex V i The surface normal at point t, and t i For vertex V i Cortical thickness at the site, such as Figure 8A As can be seen above, the normal to each vertex is calculated as a normalized vector of the summed surface normals belonging to the vertex loop. Cortical thickness values ​​for specific anatomical landmarks can be found in literature studies, for example, concerning vertebral endplates and pedicles.

[0203] Figure 8BAn example of interpolation thickness mapping for the entire mesh structure of the L1 vertebra is shown according to an embodiment of the invention. Different regions of mapping 124, typically represented by different color areas, correspond to different values ​​(from top to bottom) in the range of 0.5 to 1.8, with the median value located at 1.2 on the scale bar 125 represented to the right of mapping 124.

[0204] These cortical thickness values ​​for specific anatomical landmarks were interpolated across the entire reticular structure vertices using plate-spline 3D interpolation, thus allowing for the computation of cortical thickness mappings 124, such as... Figure 8B As can be seen above.

[0205] Figure 9 This demonstrates the principle of ray projection through a double-layered mesh structure. Rays from pixel 129 on the combined X-ray source 126 and DRR image 128 alternately traverse the cortical and cancellous media of vertebra 127. The ray projection reproduces the principle of X-ray image formation, and it calculates the sum of the thicknesses of the two bone media traversed through vertebra 127.

[0206] Figure 10 An example of an XRAY-to-DRR GAN network architecture and training according to an embodiment of the present invention is shown.

[0207] The U-Net generator 130 is trained to convert the input XRAY speckle 143 into a DRR image 144 with three channels 146, 147 and 148, i.e., one channel DRR 146, 147 and 148 for each vertebra represented by a red-green-blue (RGB) image.

[0208] Generator 130 includes several convolutional layers, each containing one or more slices, each slice representing an operation. Several types of slices or operations are available:

[0209] Slice or operation 131: A 4×4 convolutional 2D layer with a stride of 2 and using a leak-corrected linear unit activation function.

[0210] Slicing or Operation 132: Batch Normalization

[0211] Slicing or Operation 133: Upsampling a 2D layer of size 2

[0212] Slicing or Operation 134: Hidden Layer

[0213] Slice or Operation 135: A 4×4 three-channel output convolutional 2D layer with a stride of 1 and utilizing the hyperbolic tangent activation function.

[0214] Slice or operation 137: Output convolutional 2D layer of size 4×4 with stride 1

[0215] From left to right, from XRAY spot 143 to DRR image 144, the composition of the consecutive convolutional layers within generator 130 (the generator's input is XRAY spot 143, and its output is DRR image 144) is as follows:

[0216] Layer 1: Slice 131

[0217] Layer 2: Slice 131, followed by slice 132.

[0218] Layer 3: Slice 131, followed by slice 132.

[0219] Layer 4: Slice 131, followed by slice 132.

[0220] Layer 5: Slice 131, followed by slice 132.

[0221] Layer 6: Slice 131, followed by slice 132.

[0222] Layer 7: Slice 131, followed by slice 132, followed by slice 133.

[0223] Layer 8: Slice 131, followed by slice 134, followed by slice 132.

[0224] Layer 9: Slice 133, followed by slice 131, then slice 134, and then slice 132.

[0225] Layer 10: Slice 133, followed by slice 131, then slice 134, and then slice 132.

[0226] Layer 11: Slice 133, followed by slice 131, followed by slice 134, followed by slice 132.

[0227] Layer 12: Slice 133, followed by slice 131, followed by slice 132.

[0228] Layer 13: Slice 133, followed by slice 131, followed by slice 132.

[0229] Layer 14: Slice 133, followed by slice 131, then slice 132, and then slice 135. These layers can be chained together, as shown by horizontal line 136.

[0230] Layers 1 and 14 are connected in series.

[0231] Layers 2 and 13 are connected in series.

[0232] Layers 3 and 12 are connected in series.

[0233] Layers 4 and 11 are connected in series.

[0234] Layers 5 and 10 are connected in series.

[0235] Layers 6 and 9 are connected in series.

[0236] Layers 7 and 8 are connected in series.

[0237] From left to right, from input to output, the discriminator 140 consists of 5 consecutive layers:

[0238] Layer 1: Slice 131

[0239] Layer 2: Slice 131, followed by slice 132.

[0240] Layer 3: Slice 131, followed by slice 132.

[0241] Layer 4: Slice 131, followed by slice 132.

[0242] Layer 5: Slice 137

[0243] At the output of discriminator 140, there is a loss function 139. Following this loss function 139, there is a portion 142 of an adversarial switcher with two switchable positions 0 and 1.

[0244] A loss function 138 exists between the generated DRR image 144 (actually, three images 146, 147, and 148, one for each vertebra) and the corresponding baseline truth DRR image 145 (actually, three images 146, 147, and 148, one for each vertebra). There is feedback from the generated DRR image 144 toward a switchable “false” position of the portion 141 of the adversarial switcher. There is also feedback from the corresponding baseline truth DRR image 145 toward another switchable “true” position of this portion 141 of the adversarial switcher. The ends opposite the false and true positions of this portion 141 of the adversarial switcher then return to one of the inputs to the discriminator 140, the other input of which is the XRAY speckle 143.

[0245] The XRAY to DRR converter is described below. The XRAY to DRR image converter (Equation 5) is a U-Net network trained using a Generative Adversarial Network (GAN). U-Net training creates a dense pixel-wise mapping between paired input and output images, allowing the generation of realistic DRR images with constant panorama. This involves constructing a deep abstract representation to solve this problem. The Generative Adversarial Network (GAN) consists of a U-Net generator (G) 130 and a CNN discriminator (D) 140. Therefore, two loss functions are defined, namely, loss function 138 for the generator 130 and loss function 139 for the discriminator 140. First, the residual between the converted 144 and the actual 145 DRR images from the training database is calculated using the following mean absolute loss function 138 error:

[0246]

[0247] Where (n, w, h, c) is the dimension of the 4D tensor of the training data output (DRR image) of n 3D images with size width (w) x height (h) x channels (c). For the j-th pixel of the i-th sample image, f(XRAY) i Let ,W) be the U-Net prediction of the i-th input XRAY with the current CNN parameters W. The input data is also a 4D tensor of size (n,w,h,k). Sizes k and c are the number of channels in the input (XRAY) and output (DRR) images, respectively. The discriminator (D) 140 aims to classify whether the XRAY / DRR pair of the image presented on its input generates an image for DRR (false class) or an actual DRR image (true class). Therefore, the loss function ξD 139 is constrained by the binary cross-entropy loss.

[0248] The training data is split into micro-batches (e.g., size 32). For each micro-batch, the generator (G) 130 predicts fake DRR images. Next, the discriminator 140 is first trained separately to predict a binary output given two images assigned to the generator output: when it's a fake image (output set to zero: fake class) or when it's a real image (output set to one). Finally, gradient backpropagation is performed using the combined G 130 and D 140 networks and the total loss L = ξD + ξG to update the generator weights (W) (however, the discriminator weights are frozen during this step of generator training).

[0249] The architecture for the U-Net generator G130 consists of 15 convolutional layers for image feature encoding-decoding, as illustrated in the paper [P. Isola, J.-Y. Zhu, T. Zhou, and AAEfros, “Image-to-Image Translation with Conditional Adversarial Networks”, CoRR, Vol. abs / 1611.0, 2016.]. Hidden slice 134 is used in the three first decoding layers to facilitate good generalization of the image-to-image translation model. The discriminator output neurons have a sigmoid activation function. In the current example, the CNN input size is 256×256×1 (an XRAY blob on a view), and the DRR output size is 256×256×3. An output channel is assigned to a different anatomical structure to separate it during training: upper 146, middle 147, and lower 148 vertebrae.

[0250] The pixel-to-pixel network uses a Generative Adversarial Network (GAN) training procedure to train the generator's parameters (weights). The generator has a U-Net architecture, allowing image-to-image mapping to transform the actual x-ray spot 143 into a DRR spot 144 (domain A to domain B). The entire network consists of a U-Net generator 130 and a CNN discriminator 140. In the current example, the GAN input size is 256 × 256 × 1 (x-ray spot 143), and the DRR 144 output size is 256 × 256 × k, where k is constrained to the number of anatomical structures in the output, here 3 vertebrae 146, 147, and 148. Therefore, an output channel is assigned to each distinct anatomical structure (here, each distinct vertebra) to separate them during training. For K = 3 visible skeletal structures, k = 3 output channels are constrained: the vertebra of interest 147 (Ki = 1), and its upper 146 (Ki = 0) and lower 148 (Ki = 2) adjacent vertebrae.

[0251] Once trained, the network can transform an input image 143 (x-ray) belonging to the complex domain into a simplified virtual x-ray (DRR type) 144 belonging to domain B. Due to the image transformation, the generator (U-Net type) 130 network suppresses noise, neighboring structures, and background from the original image.

[0252] Using the converted image is permitted:

[0253] First, it facilitates image-image correspondence (similarity calculation).

[0254] And secondly, separate adjacent structures to avoid mismatch between them.

[0255] In 3D / 2D alignment schemes, these skeletal multi-structure pixel-to-pixel networks facilitate the similarity calculation phase between the target DRR image generated using GANU-Net and the modified image generated from the 3D model. Because the similarity value is used as cost to find the optimal model parameters that maximize the similarity between the model and the image, classifying the two images as belonging to the same domain category allows for a simpler similarity metric and introduces fewer local maxima attributable to mismatches between neighboring structures during optimization.

[0256] The experimental ensemble is now described. The dataset used for training the XRAY-to-DRR converter comprises biplane acquisitions (PA and LAT views) of 463 patients performed using a geometrically calibrated system and a known 3D environment. For the reconstructed 3D spine, a semi-automatic 3D reconstruction method is integrated to provide a baseline truth for the 3D model. To evaluate the proposed method of embodiments of the invention, another clinical dataset of 40 adolescent idiopathic scoliosis patients (mean age 14 years, mean principal Cobb angle 56°) is used. Bronze standards are constructed for each 3D model to define the baseline truth as the mean of 3D reconstructions from three experts.

[0257] During the training of the XRAY to DRR converter, PA and LAT DRR images are generated individually for each reconstructed 3D model for each patient in the training dataset using a previously existing algorithm.

[0258] Figure 11A An example of GaN-based XRAY to DRR conversion is shown, from an actual X-ray image to a three-channel DRR image corresponding to a nearby vertebral structure.

[0259] Image 151a is a frontal X-ray image to be converted, which, from top to bottom, represents: vertebral L4, vertebral L5, and sacral plate S1.

[0260] Image 152a is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in the embodiments of the present invention. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, and blue = sacral plate S1.

[0261] Image 153a is the original frontal DRR image corresponding to frontal X-ray image 151a, and is a reference image to which frontal DRR image 152a will be compared in order to train the algorithm of the method proposed in the embodiments of the present invention.

[0262] Image 151b is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in the embodiments of the present invention, which corresponds only to vertebra L4. It also corresponds to the upper portion of image 152a.

[0263] Image 152b is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in the embodiments of the present invention, which corresponds only to vertebra L5. It also corresponds to the middle portion of image 152a.

[0264] Image 153b is a frontal DRR image generated from frontal X-ray image 151a by the method proposed in the embodiments of the present invention, which corresponds only to the sacral plate S1. It also corresponds to the lower portion of image 152a.

[0265] Image 151c is the original frontal DRR image corresponding to frontal X-ray image 151a, and is the reference image to which frontal DRR image 151b will be compared. It corresponds only to vertebra L4. It also corresponds to the upper portion of image 153a.

[0266] Image 152c is the original frontal DRR image corresponding to frontal X-ray image 151a, and is the reference image to which frontal DRR image 152b will be compared. It corresponds only to vertebra L5. It also corresponds to the middle portion of image 153a.

[0267] Image 153c is the original frontal DRR image corresponding to frontal X-ray image 151a, and is a reference image to which frontal DRR image 153b will be compared. It corresponds only to the sacral plate S1. It also corresponds to the lower portion of image 153a.

[0268] Image 151d is a lateral X-ray image to be converted, which, from top to bottom, represents: vertebral L4, vertebral L5, and sacral plate S1.

[0269] Image 152d is a lateral DRR image generated from lateral X-ray image 151d by the method proposed in the embodiments of the present invention. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, and blue = sacral plate S1.

[0270] Image 153d is the original side DRR image corresponding to side X-ray image 151d and is a reference image to which side DRR image 152d will be compared in order to train the algorithm of the method proposed in the embodiments of the present invention.

[0271] Following the morphological changes of vertebrae along the spine, per-view converters were trained for each vertebral segment: T1 to T5, T6 to T12, and L1 to L5. Twenty spots were extracted for each vertebra around the vertebral body center (VBC) with random shifts, rotations, and scaling. Regions of interest (ROIs) were defined with dynamic reduction factors such that the vertebra of interest and at least the endplates of adjacent layers remained visible in 256×256 pixel spots, the reduction factor depending on the vertebral dimension in the image. The training dataset was split into two sets: training (70%) and testing (30%). Training was run for 100 time periods. At each time period, the mean squared error (MSE), i.e., the error over the predicted spots (test set), was calculated to select the best model for that time period. The lumbar vertebra converter (LAT view) was trained in a configuration of c=1 or c=3 (Equation 7), i.e., with an output image having only one channel, or with a layered image having three channels with three vertebral DRRs. The MSE of the DRR for the intermediate vertebra is 0.0112 for one channel and 0.0108 for three channels, revealing that the additional output information from neighboring vertebrae is helpful for GAN training. Therefore, all three output channels are used for each training iteration, even though a vertebral 3D model for 3D / 2D alignment applications will use the intermediate channel as the target. Once trained, the network can convert X-ray images into hierarchical images with one DRR image per structure, such as... Figure 11A As can be seen above. The qualitative results of the converter are... Figure 11A and 11B Presented in both.

[0272] Figure 11B The transition (also referred to as the shift) from the X-ray domain to the DRR domain under adverse conditions is demonstrated to illustrate the robustness of the proposed method according to embodiments of the present invention.

[0273] Image 154a is a frontal X-ray image to be converted, which, from top to bottom, represents: vertebral L4, vertebral L5, and sacral lamina S1. The X-ray image shows poor visibility.

[0274] Image 155a is a frontal DRR image generated from frontal X-ray image 154a by the method proposed in the embodiments of the present invention, corresponding to the "transition" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, and blue = sacral plate S1.

[0275] Image 156a is the original frontal DRR image corresponding to frontal X-ray image 154a and is a reference image to which frontal DRR image 155a will be compared in order to evaluate the effectiveness and performance of the algorithm of the method proposed in the embodiments of the present invention.

[0276] Image 154b is the frontal X-ray image to be converted, which, from top to bottom, represents: vertebral L4, vertebral L5, and sacral lamina S1. The X-ray images are displayed in an overlay.

[0277] Image 155b is a frontal DRR image generated from frontal X-ray image 154b by the method proposed in the embodiments of the present invention, corresponding to the "transition" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, and blue = sacral plate S1.

[0278] Image 156b is the original frontal DRR image corresponding to frontal X-ray image 154b and is a reference image to which frontal DRR image 155b will be compared in order to evaluate the effectiveness and performance of the algorithm of the method proposed in the embodiments of the present invention.

[0279] Image 154c is a frontal X-ray image to be converted, which, from top to bottom, represents: vertebral L4, vertebral L5, and sacral plate S1. The X-ray image shows the circular metal portion 157.

[0280] Image 155c is a frontal DRR image generated from frontal X-ray image 154c by the method proposed in the embodiments of the present invention, which corresponds to the "transition" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, and blue = sacral plate S1.

[0281] Image 156c is the original frontal DRR image corresponding to frontal X-ray image 154c and is a reference image to which frontal DRR image 155c will be compared in order to evaluate the effectiveness and performance of the algorithm of the method proposed in the embodiments of the present invention.

[0282] Image 154d is a frontal X-ray image to be converted, which, from top to bottom, represents: vertebral L4, vertebral L5, and sacral plate S1. The X-ray image shows the metal screw 158.

[0283] Image 155d is a frontal DRR image generated from frontal X-ray image 154d by the method proposed in the embodiments of the present invention, corresponding to the "transition" from the X-ray domain to the DRR domain. Each color represents an anatomical structure: red = vertebra L4, green = vertebra L5, and blue = sacral plate S1.

[0284] Image 156d is the original frontal DRR image corresponding to frontal X-ray image 154d and is a reference image to which frontal DRR image 155d will be compared in order to evaluate the effectiveness and performance of the algorithm of the method proposed in the embodiments of the present invention.

[0285] The similarity function is now described. Five measures are implemented to evaluate the image similarity used in (Equation 2). The similarity between two DRR-like images is computed due to the x-ray image transformation using a U-Net neural network, allowing the use of common unimodal similarity functions such as Regularized Cross-Correlation (NCC) and Sum of Squared Differences (SSD). The gradient-based measure is Regularized Gradient Information (NGI), along with the NCC computed on the gradient image (NCCGRAD). A Regularized Interaction Information (NMI) measure is also included, where a joint histogram is computed from the images in 8-bit binary representation.

[0286] This section describes a statistical shape model for vertebrae. A deformable model is used when alignment is flexible, such as when the target shape is optimized in addition to the model pose. The deformation technique uses mesh-structured least-squares (MLS) deformation controlled by a set of geometric parameters regularized using a PCA model. The prior constraints on the simplified parametric model covering the target shape allow for a compact representation of the geometry and directly provide subject-specific signage and geometric primitives for biomedical applications.

[0287] The resulting PCA model, which captures the main changes, provides a linear generative model in the following form: Where B is the foundation of PCA. The model is a mean-mean model, and m is a vector of deformation modes. The generation of mesh structure instances with the function Ψ(p) (Equation 4) is controlled by the parameter vector p = {Tx, Ty, Tz, Rx, Ry, Rz, Sx, Sy, Sz, m}, where... It consists of nine parameters (translation, rotation, and scaling) for the affine portion used to transform and scale the 3D model to a right pose in an X-ray calibrated 3D environment, and a shape vector m with a |m|PCA mode. Given vector m, the MLS deformation treatment is calculated and used to deform the vertices of the mesh structure using parameters s.

[0288] The optimizer is now described. The cost function (Equation 1) is optimized by minimizing the following equation (Equation 8) to solve for the 3D / 2D alignment on the biplane PA and LAT ray images:

[0289]

[0290] Where Φ(I1, I2) is the similarity metric defined to

[01] for NGI, NCC, NCCGRAD, and NMI. For SSD, the sum of the PA and LAT similarity scores is calculated to limit the cost. To minimize the cost function (Equation 8), a derivative-free CMA-ES optimizer is used. Each CMA-ES iteration evaluates the cost function by 100 times to construct the covariance matrix. Upper and lower limits can be specified.

[0291] The proposed method of embodiments of the present invention is now evaluated. Three experiments were conducted to evaluate the proposed method in the context of 3D / 2D alignment of a vertebral 3D model in biplane X-rays. First, the target alignment error (TRE) of anatomical landmarks belonging to the 3D model was reported to investigate the behavior of the similarity metric with and without a GAN transformation step. Additional tests were conducted to demonstrate the advantage that structural separation in XRAY to DRR transformation is less sensitive to the initial pose. Finally, accurate results of fully automated 3D reconstruction of the spine are presented.

[0292] The target alignment error (TRE) is described below. In this experiment, seventeen vertebrae 3D models from level T1 to L5 were first fitted onto a previously reconstructed baseline truth 3D model. Alignment was configured to solve for rigid body transformations in six degrees of freedom. For each vertebra of each patient, a randomized transformation was applied to the 3D model. For translations Tx, Ty, and Tz, the model was shifted within ±3 mm using a randomized uniform method. For in-plane rotations Rx and Ry, a range of ±5° was defined, while for out-of-plane axial spinal rotations Rz, ±15° was used. Upper and lower limits were also defined for these values ​​in the optimizer. The range of the randomized transformation was constrained so that alignment with the original X-ray target was convergent. In practice, 3D / 2D alignment was evaluated against both DRRGAN and the original X-ray target to quantify the improvement introduced by the GaN-based converter. The U-Net output channel corresponding to the intermediate vertebra was used as the target for the DRRGAN image. TRE was calculated as the root mean square error of the 3D positional error of the anatomical landmark after alignment. The landmarks extracted from the aligned 3D model were the centers of the superior and inferior endplates and the centers of the left and right pedicles. A total of 40 (patient) × 17 (vertebral level) × 4 (landmark) measures were performed for each similarity / target pair.

[0293] Table 1 shows the quantitative TRE for each target / similarity pair.

[0294] Table 1. Quantitative TRE for each target / similarity pair

[0295]

[0296] Table 1 compares the TRE results for five similarity metrics and two targets, reporting the error rate for TRE > 2.4 mm, the mean of the cost function difference (i.e., amplitude cost Δ = max(costs) - min(costs), and the average number of iterations used for convergence, where the stopping tolerance criterion is fixed to 1 for the SSD metric and to 0.001 for the others. When using an image transformation step with DRRGAN as the target, the TRE results are nearly similar for all similarity metrics from 1.99 to 2.12, except for a higher error of 2.52 for NCC GRAD. When XRAY is used as the target, the SSD, NCC, and NMI metrics degrade the alignment results because the DRR3DM image has limited similarity to the XRAY image. NMI exhibits a near-zero cost variation, indicating a convergence problem. Gradient-only metrics achieve TREs of 3.6 and 3.3 mm for NGI and NCCGRAD, respectively, as shown in Table 1. The evolution of similarity during optimization is more important when using DRRGAN. For example, using the NGI metric, the average cost difference is 0.29 for DRRGAN, compared to 0.05 for XRAY. This means that the assessed cost is less sensitive when using GaN-based converters, as can be seen in Table 1, which can contribute to better convergence.

[0297] Figure 12A This illustrates the target alignment error (TRE) varying based on the initial pose shift along the vertical axis. This TRE is plotted on the vertical axis based on different initial pose shifts along the vertical axis z on the horizontal axis. For DRRGAN containing three vertebral structures (all channels flattened in one image), the error value often exceeds 5, even reaching 10 or 20. This error value is significant and exhibits considerable uncertainty. For a single DRRGAN channel, the error value is less than 5, often not exceeding 2 or 3. This error value is smaller and exhibits minimal uncertainty. This demonstrates the critical importance of structural separation for avoiding mismatch between adjacent vertebrae. Therefore, the method proposed by embodiments of the present invention, due to the separation of vertebrae from each other, results in a much smaller and better known error value.

[0298] Figure 12B This example demonstrates a vertical displacement towards the top of the vertebra. The vertebra is vertically displaced upwards by +10mm. This means the vertebra is vertically 10mm above the baseline truth model.

[0299] In this experiment, the baseline true 3D model of the L2 vertebra was shifted in the proximal-distal direction (that is, along the z-axis) within a range of ±30 mm in steps of 5 mm, as shown below. Figure 12A and 12BAs seen in the case study, this provides 13 initial poses. Then, rigid subject alignment is performed using an image corresponding to a single channel of the intermediate vertebra or an image constructed by flattening all channels as the DRRGAN target image, thereby blending the upper, middle, and lower vertebral structures. The metric used is SSD, with upper and lower limits defined to ±5mm for Tx and Ty, ±32mm for Tz, and ±5° for rotation. Figure 12A Box plots are shown for translational residuals for each initial pose displacement in 40 patients. The residual transformation after alignment shows a significant error behind a 10mm absolute displacement, without structural separation, as... Figure 12A As can be seen above, even for absolute displacements <10mm, some outliers occur because the alignment is perturbed by neighboring layers L1 and L3. When using GANDRR to transform the image, the structures of interest are isolated, and the alignment process captures a larger range of initial poses and is more accurate.

[0300] This study investigates the sensitivity to the initial pose. For the alignment of a single-structure 3D model (e.g., a vertebra along the spine) in a scene composed of periodic multi-structures, the sensitivity to the initial model pose, for example, from an automated coarse alignment, is investigated because, behind constraints, 3D model alignment will risk erroneous convergence on neighboring structures.

[0301] This study investigates the accuracy of vertebral pose and shape. In this experiment, fully automated 3D reconstruction of the spine is evaluated. Initial solutions for the vertebral 3D model shape and pose are provided by a CNN-based automated method, as described in the article [B. Aubert, C. Vazquez, T. Cresson, S. Parent, and J. De Guise, “Towards automated 3D Spine reconstruction from biplanar radiographs using CNN for statistical spine model fitting,” IEEE Transactions on Medical Imaging, p. 1, 2019]. This method provides pre-personalized statistical spine models (SSMs) for levels C7 to L5, detected at the endplates and pedicle centers on biplanar X-rays. To fine-tune these obtained 3D models, 3D / 2D alignment is applied in an elastic mode, where |m| = 20 PCA mode is limited to ±3. Alignment is performed individually level by level using the NCC metric. Next, the resulting aligned model is used to update the SSM regularization in order to obtain the final model after local 3D / 2D fitting.

[0302] Table 2 shows the quantitative markers TREs used for 3D / 2D alignment of vertebrae.

[0303] Table 2. Quantitative markers of TRE for 3D / 2D alignment of vertebrae

[0304]

[0305] The 3D locations of the endplate center and corner, as well as the pedicle center, were improved through the proposed 3D / 2D alignment steps, as shown in Table 2. The most significant optimizations were observed for the pedicle center with TREs of 3 and 2.2 mm before and after fine 3D / 2D alignment, respectively, as shown in Table 2. The error rate >2.4 mm decreased by 26.4%, 15.4%, and 19.1% for the pedicle center and corner, respectively.

[0306] Table 3 shows the quantitative error regarding vertebral position and orientation.

[0307] Table 3. Quantitative error (mean ± SD) regarding vertebral position and orientation

[0308] district X(mm) Y(mm) Z(mm) L(°) S(°) A(°) neck -0.1±0.6 -0.5±0.6 0.1±0.8 1.9±2.6 0.2±3.8 -0.5±3.5 Chest 0.0±1.5 0.2±1.0 0.3±0.7 0.5±2.8 1.0±2.3 0.8±4.2 waist -0.2±0.7 -0.0±0.8 0.4±0.5 -1.0±2.0 0.2±2.1 1.3±2.6 all -0.0±1.3 0.1±0.9 0.3±0.7 0.0±2.7 0.8±2.3 0.9±3.8

[0309] The axial coordinate system for each vertebra is defined using the pedicle and endplate centers. The 3D positional error of each vertebral target is calculated for each vertebral segment based on its position and orientation relative to the baseline truth 3D model, as shown in Table 3. All translations have a mean error of less than or equal to 0.5 mm, demonstrating the low systematic bias of this method. The standard deviation is more important for X-translations (i.e., position in the side view), especially for the thoracic level, as shown in Table 3.

[0310] Finally, shape accuracy was estimated by calculating node-to-surface distance errors using a reference 3D model reconstructed from CT scan images of four patients with simultaneously acquired dual-plane and CT scan images. The volumetric resolution was 0.56 x 0.56 x 1 mm, and segmentation was performed using 3D slicer software.

[0311] Figure 13A The anatomical region used to calculate the statistical distance from nodes to surfaces is shown. In vertebra 164, the posterior arch 165 can be distinguished from the structure 166 representing the vertebral body and pedicle.

[0312] Figure 13B This displays the distance mapping for the maximum error calculated for the L3 vertebra. The different areas of mapping 167, typically represented by different colors, correspond to different values ​​ranging from 0.0 mm to 6.0 mm (from top to bottom), with the median value located at 3.0 mm on the scale 168 shown to the right of mapping 167.

[0313] Table 4 shows the shape accuracy results relative to the CT scan model.

[0314] Table 4. Results of shape accuracy relative to CT scan models

[0315]

[0316] The target is rigidly aligned, and the distance from the nodes to the surface is calculated. The mesh density of the model, i.e., the number of its nodes, is specified in Table 4. Error statistics are reported in Table 4 for different anatomical regions, which are mesh structures of the entire vertebra 164, the posterior arch 165, or structures 166 encompassing the vertebral body and pedicles, as shown below. Figure 13A As shown above, the average error range is 1.1 to 1.2 mm, as can be seen in Table 4. According to... Figure 13B The above represents the error distance mapping 167, with the maximum error value located in the rear bow area 165.

[0317] The proposed method in embodiments of the present invention adds a prior image-to-image conversion step of the target image to the intensity-based 3D / 2D alignment process to achieve robust bi-image matching where the two images belong to different domains (X-ray and DRR), and varies with different environments and structures. As a first benefit, the XRAY-to-DRR converter improves the similarity between two images by bringing the target image into the same domain of the varying images, as can be seen in Table 1. The conversion step reduces the dependence on the choice of similarity metric, and it can even use common single-mode attitude metrics (as can be seen in Table 1) and simplified DRR generation. This is a feature of interest because conventional experimental methods involve selecting a better metric within a set of metrics. However, this often leads to trade-offs, as some of these metrics will give good or bad results depending heavily on the specific context.

[0318] As a second benefit, mismatches between adjacent and superimposed structures are avoided by selecting the structure of interest at the converter output, such as... Figure 13A As can be seen above. In practice, the structure directly isolates the original 3D volume to generate a regional DRR per target, and each DRR is assigned to a separate layer in the multi-channel output image of the XRAY-to-DRR converter. In the cited prior art US 2019 / 0259153, previous work using an XRAY-to-DRR converter only produced a global DRR that allowed for the recovery of segmented masks for each organ, but did not generate a regional DRR per target.

[0319] Applied to fully automated 3D reconstruction of the spine from biplane radiographs, the added 3D / 2D alignment optimization step improves both target localization and shape compared to CNN-based automated methods used for initialization, as can be seen in Tables 2 and 3. The average error for 3D landmark localization is 2 ± 1 mm, which is better than the 2.7 ± 1.7 mm found for pedicle detection using a CNN-based shift regression method, or better than the 2.3 ± 1.2 mm found by fitting a nonlinear spine model using a 3D / 2D Markov random field. Compared to “quasi” automated 3D reconstruction methods that require user input of two spine curves in two views and rigid manual adjustments once the model is fitted on biplane X-rays, the proposed 3D / 2D alignment algorithm of this invention achieves better results for all pose parameters for populations with severe scoliosis.

[0320] Integrated with clinical procedures for performing 3D spinal reconstruction The reference method in the software investigated the shape accuracy of the reconstructed target relative to a baseline truth target derived from a CT scan. For greater accuracy, this method requires time-consuming manual elastic 3D / 2D alignment (over 10 minutes) to fit the contours of the vertebral projection to the X-ray information. The proposed automated local fitting step, costing less than one minute of computation time, advantageously suppresses operator dependence and achieves similar accuracy results, for example, 0.9 ± 2.2 mm (mean ± 2 SD) for the vertebral body and pedicle region, and 1.3 ± 3.2 mm for the posterior arch region, contrasting with 0.9 ± 2.2 mm and 1.2 ± 3 mm in a previous study.

[0321] These generated images are used to replace actual X-rays to achieve robust single-modal image correspondence without structural mismatch. This solution, integrated into a non-rigid 3D / 2D alignment process aimed at adjusting a vertebral 3D model relative to biplane X-rays, improves accuracy results, as previously seen.

[0322] Here are some other results to illustrate the difference between better similarity scores ( Figures 14A-14B -14C-14D-15) and prevention of anatomical mismatch ( Figure 16 Both utilize the improvements resulting from previous image-to-image transformations.

[0323] Figure 14A This section displays the cost values ​​of the PA and LAT views using GAN DRR compared to the actual X-ray image. The cost values ​​are expressed on the y-axis based on the number of iterations on the x-axis. Curve 171 shows the higher cost function of the X-ray image, while curve 172 shows the lower and therefore better cost function of the GAN DRR image.

[0324] Figure 14BThis section displays the similarity values ​​of the PA and LAT views using GAN DRR compared to the actual X-ray image. The similarity values ​​are expressed on the y-axis based on the number of iterations on the x-axis. For the front view, curve 173 shows a lower similarity function for the X-ray image, while curve 175 shows a higher and therefore better similarity function for the GAN DRR image. For the side view, curve 174 shows a lower similarity function for the X-ray image, while curve 176 shows a higher and therefore better similarity function for the GAN DRR image. The similarity value measured by NGI (Normalized Gradient Information) reaches a higher value when using the generated DRR GAN than when using the actual X-ray image.

[0325] Figure 14C Showing compared to using actual X-ray images Figure 14D A better fit was observed for the alignment results when using the DRR generated by GAN. In fact, in Figure 14C In the left portion representing the front view, the generated image 181 fits the target image 180, which was first transformed in the DRR domain, better than... Figure 14D The image 185 generated on the left portion of the front view is fitted to the target image 184 held in the x-ray domain. Furthermore, in Figure 14C In the right portion representing the side view, the generated image 183 fits the target image 182, which was first transformed in the DRR domain, better than... Figure 14D The generated image 187, which represents the right portion of the side view, is fitted to the target image 186 held in the x-ray domain.

[0326] Figure 15 The display is as follows Figure 14B Similar results were obtained with the GNCC (Gradient Regularized Cross-Correlation) metric. For the front view, curve 190 with GAN DRR showed higher and better similarity than curve 188 without GAN DRR. For the side view, curve 191 with GAN DRR showed higher and better similarity than curve 189 without GAN DRR.

[0327] Figure 16The results of the Target Alignment Error (TRE) test in the Z-direction (vertical image direction) are presented. In Part A, different alignment errors, expressed in mm on the vertical axis, are plotted as a function of vertical translational displacement, also expressed in mm on the horizontal axis. NGI DRR curve 194 shows a much lower error than NGI x-ray mask curve 192, which in turn shows a lower error than NGI x-ray curve 193. In Part B, Z-translation displacement from -10 mm is shown by the relative position between point 196 and spine structure 195. In Part C, midline translation is shown by the relative position between point 198 and spine structure 197. In Part D, Z-translation displacement from 10 mm is shown by the relative position between point 200 and spine structure 199.

[0328] The TRE test aims to transform a 3D model using a known theoretical transformation and then analyze the alignment residuals between the theoretical transformation and the recovered transformation. It can be seen that the generated GAN DRR has less dependence on the initial position. It can also be seen that the target alignment error (TRE) increases significantly after a Z-shift of ±4 mm, thus demonstrating that without using the corresponding... Figure 16 In the case of the generated GAN DRR images of curves 192 and 193 on part A, a significant structural mismatch occurs between adjacent structures (vertebral endplates).

[0329] The invention has been described with reference to preferred embodiments. However, many variations are possible within the scope of the invention.

Claims

1. A medical imaging conversion method that automatically performs the following conversion: converting at least one or more real X-ray images (16) of the patient, which contain at least a first anatomical structure (21) and a second anatomical structure (22) of the patient, into at least one digitally reconstructed radiograph (DRR) (23) of the patient representing the first anatomical structure (24) but not the second anatomical structure (26), the conversion being performed by a single operation using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) (27). The convolutional neural network (CNN) is initially trained to perform, or simultaneously perform, the following: distinguishing the first anatomical structure (21) from the second anatomical structure (22) by directly isolating the structure in the original 3D volume; and converting the real X-ray image (16) into at least one digital reconstructed radiograph (DRR) (23) by: simultaneously generating a set of several DRRs, each of which focuses on only one anatomical structure of interest; or generating a single DRR that focuses only on one anatomical structure of interest, excluding other anatomical structures.

2. The medical imaging conversion method according to claim 1, characterized in that, It also automatically performs the following conversion: the patient's actual X-ray image (16) is converted into at least another digitally reconstructed radiograph (DRR) (25) of the patient representing the second anatomical structure (26) but not the first anatomical structure (24), which is performed through the single operation. The convolutional neural network (CNN) or the set of convolutional neural networks (CNNs) (27) is initially trained to perform, or simultaneously perform, the following: distinguishing the first anatomical structure (21) from the second anatomical structure (22) by directly isolating the structure in the original 3D volume; and converting the real X-ray image (16) into at least two digitally reconstructed radiographs (DRRs) (23, 25) by simultaneously generating a set of several DRRs, each of which focuses on only one anatomical structure of interest; or generating a single DRR that focuses only on one anatomical structure of interest, excluding other anatomical structures.

3. A medical imaging conversion method that automatically performs the following conversion: converting at least one or more real X-ray images (16) of the patient, containing at least a first anatomical structure (21) and a second anatomical structure (22) of the patient, into at least a first (23) and a second (25) digitally reconstructed radiographs (DRR) of the patient, which is performed by a single operation using a convolutional neural network (CNN) or a set of convolutional neural networks (CNN) (27). The first digital reconstructed radiograph (DRR) (23) represents the first anatomical structure (24) but not the second anatomical structure (26). The second digital reconstructed radiograph (DRR) (25) represents the second anatomical structure (26) but not the first anatomical structure (25). The convolutional neural network (CNN) is initially trained to perform, or simultaneously perform, the following: distinguishing the first anatomical structure (21) from the second anatomical structure (22) by directly isolating the structure in the original 3D volume; and converting the real X-ray image (16) into at least two digitally reconstructed radiographs (DRRs) (23, 25) by: simultaneously generating a set of several DRRs, each of which focuses on only one anatomical structure of interest; or generating a single DRR that focuses only on one anatomical structure of interest, excluding other anatomical structures.

4. The medical imaging conversion method according to claim 3, characterized in that, The convolutional neural network (CNN) or the set of convolutional neural networks (CNN) is a single generative adversarial network (GAN) (27).

5. The medical imaging conversion method according to claim 4, characterized in that, The single generative adversarial network (GAN) (27) is either U-Net GAN or residual-Net GAN.

6. The medical imaging conversion method according to claim 3, characterized in that, The actual X-ray image (16) of the patient is a direct capture of the patient by the X-ray imaging device.

7. The medical imaging conversion method according to claim 3, characterized in that: The first (21) and second (22) anatomical structures of the patient are as follows: Adjacent to each other in the real X-ray image (16), Or they may be in contact with each other on the actual X-ray image (16). Or at least partially superimposed on the real X-ray image (16).

8. The medical imaging conversion method according to claim 3, characterized in that, Only one of the at least two digitally reconstructed radiographs (DRR) (23, 25) (23) can be used for further processing.

9. The medical imaging conversion method according to claim 3, characterized in that, All of the at least two digitally reconstructed radiographs (DRR) (23, 25) can be used for further processing.

10. The medical imaging conversion method according to claim 3, characterized in that: The actual X-ray image of the patient (151a) contains at least three anatomical structures of the patient, or only three anatomical structures of the patient. The real X-ray image (151a) is converted by the single operation into at least three separate digitally reconstructed radiographs (DRRs) (151b, 152b, 153b) representing the at least three anatomical structures respectively, each of the DRRs representing only one of the anatomical structures and not any other anatomical structures; or converted into only three separate DRRs representing only the three anatomical structures respectively.

11. The medical imaging conversion method according to claim 3, characterized in that: The single convolutional neural network (CNN) or the set of convolutional neural networks (CNNs) (27) has been initially trained using a set of the following training groups: A real X-ray image (16). And at least one or more corresponding digital reconstructed radiographs (DRRs) (23, 25) representing only one of the anatomical structures (21, 22) but not other anatomical structures of the patient.

12. The medical imaging conversion method according to claim 3, characterized in that: The single convolutional neural network (CNN) or the set of convolutional neural networks (CNNs) (27) has been initially trained using the following set: Real X-ray image (16). And several subsets of at least one or more digitally reconstructed radiographs (DRRs) (23, 25), each of which represents only one of the anatomical structures (21, 22) but does not represent other anatomical structures of the patient.

13. The medical imaging conversion method according to claim 3, characterized in that: The single convolutional neural network (CNN) or the set of convolutional neural networks (CNNs) (27) has been initially trained using a set of the following training groups: A true frontal X-ray image and a true side X-ray image (16). And at least one or more subsets of corresponding frontal and lateral digital reconstructed radiographs (DRR) (23, 25), each of the subsets representing only one of the anatomical structures (21, 22) but not the other anatomical structures of the patient.

14. The medical imaging conversion method according to any one of claims 11 to 13, characterized in that, The digital reconstructed radiographs (DRRs) (16) of the training group are derived from a 3D model specifically tailored to the patient via adaptation of the model to two real X-ray images acquired along two orthogonal directions.

15. The medical imaging conversion method according to claim 3, characterized in that, The patient has different anatomical structures, and the different anatomical structures of the patient are the patient's adjacent vertebrae (146, 147, 148).

16. The medical imaging conversion method according to claim 15, characterized in that: The patient's adjacent vertebrae (146, 147, 148) are located in a single and identical region of the patient's spine, among the following: The patient's upper thoracic spine region, Or the patient's lower thoracic spine region, Or the patient's lumbar spine area, Or the patient's cervical spine area, Or the patient's pelvic area.

17. The medical imaging conversion method according to claim 3, characterized in that: The patient has different anatomical structures, and the different anatomical structures (21, 22) of the patient are located in a single and identical area of ​​the patient among the following: Patient's buttock area, Or the patient's lower limb area, Or the patient's knee area, Or the patient's shoulder area, Or the patient's ribcage area.

18. The medical imaging conversion method according to claim 3, characterized in that: The patient has different anatomical structures, and each of the at least two digitally reconstructed radiographs (DRRs) (23, 25) representing the different anatomical structures (21, 22) of the patient simultaneously includes: An image containing pixels that display different shades of gray. At least one label representing anatomical information associated with the anatomical structure it represents.

19. The medical imaging conversion method according to claim 18, characterized in that, The image is a 256×256 pixel square image.

20. The medical imaging conversion method according to claim 3, characterized in that, The convolutional neural network (CNN) or the set of convolutional neural networks (CNN) (27) has been initially trained on 100 to 1000, or 300 to 700, or 500 x-ray images (16) of several different patients, the x-ray images being both real x-ray images and transformations of real x-ray images.

21. The medical imaging conversion method according to claim 3, characterized in that, At least two true X-ray images of the patient's frontal side (151a) and the patient's true X-ray side (151d) are converted, each of the two X-ray images containing the same anatomical structures of the patient (146, 147, 148).

22. A method for personalizing a 3D medical imaging model, comprising the medical imaging conversion method according to claim 21, wherein: Generate using a 3D generic model (120): At least one or more digitally reconstructed radiographs (DRRs) (145) of the patient’s frontal view, representing one or more of the patient’s anatomical structures (146, 147, 148). And at least one or more digitally reconstructed radiographs (DRRs) (145) of the patient’s side view, representing the patient’s one or more anatomical structures (146, 147, 148) respectively. The frontal true X-ray image (143) is converted into the following by the medical imaging conversion method: At least one or more digitally reconstructed radiographs (DRRs) (144) of the patient’s frontal view, representing one or more of the patient’s anatomical structures (146, 147, 148) respectively. The true lateral X-ray image (143) is converted into the following by the medical imaging conversion method: At least one or more digitally reconstructed radiographs (DRRs) (144) of the patient’s side view, representing one or more of the patient’s anatomical structures (146, 147, 148). And among them: The at least one or more digitally reconstructed radiographs (DRRs) (145) representing the patient's one or more anatomical structures (146, 147, 148) and obtained from the 3D universal model, are respectively mapped to the at least one or more digitally reconstructed radiographs (DRRs) (144) representing the patient's one or more anatomical structures (146, 147, 148) and obtained from the frontal true X-ray image (143). And the at least one or more digitally reconstructed radiographs (DRRs) (145) representing the patient's one or more anatomical structures (146, 147, 148) and obtained from the 3D universal model, are respectively mapped to the at least one or more digitally reconstructed radiographs (DRRs) (144) representing the patient's one or more anatomical structures (146, 147, 148) and obtained from the true lateral X-ray image (143). In order to generate a 3D patient-specific model from the general 3D model.

23. The method for personalizing medical imaging 3D models according to claim 22, characterized in that: The 3D general model is a deformable model.

24. The method for personalizing medical imaging 3D models according to claim 23, characterized in that: The deformable model is a statistical shape model.

25. A medical imaging conversion method that automatically performs the following conversion: converting at least one or more images (16) of the patient in a first domain containing at least a first anatomical structure (21) and a second anatomical structure (22) of the patient into at least one image (23) of the patient in a second domain representing the first anatomical structure (24) but not the second anatomical structure (26), wherein the conversion is performed by a single operation using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) (27). The convolutional neural network (CNN) is initially trained to perform, or simultaneously perform, the following: distinguishing the first anatomical structure (21) from the second anatomical structure (22) by directly isolating the structure in the original 3D volume, and converting an image (16) in the first domain into at least one image (23) in the second domain by: simultaneously generating a set of several digital reconstructed ray images (DRRs), each of which focuses on only one anatomical structure of interest; or generating a digital reconstructed ray image (DRR) that focuses only on one anatomical structure of interest, excluding other anatomical structures.

26. A medical imaging conversion method that automatically performs the following conversion: converting at least one or more global images (16) of the patient in a first domain containing at least several different anatomical structures (21, 22) of the patient into several regional images (23, 25) of the patient in a second domain representing the different anatomical structures (24, 26), wherein the conversion is performed by a single operation using a convolutional neural network (CNN) or a set of convolutional neural networks (CNNs) (27). The convolutional neural network (CNN) is initially trained to perform, or simultaneously perform, the following: distinguishing the anatomical structures (21, 22) from one another by directly isolating the structures in the original 3D volume; and converting an image (16) in the first domain into at least one image (23, 25) in the second domain by: simultaneously generating a set of several digitally reconstructed ray images (DRRs), each of which focuses on only one anatomical structure of interest; or generating a single digitally reconstructed ray image (DRR) that focuses only on one anatomical structure of interest, excluding other anatomical structures.

Citation Information

Patent Citations

  • Method for reconstruction of a three-dimensional model of an osteo-articular structure

    WO2009056970A2

  • Cross domain medical image segmentation

    US20190259153A1