Structure-aware sparse-view x-ray 3D reconstruction

The line-segment-density machine learning system with self-attention and masked local-global sampling effectively addresses the inefficiencies of existing methods, achieving superior 3D reconstruction accuracy and efficiency in X-ray imaging.

WO2026090053A1PCT designated stage Publication Date: 2026-04-30JOHNS HOPKINS UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/051667
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-22
Filing Date
2025-10-20
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing deep-learning-based methods for reconstructing 3D representations from 2D X-ray projections require a large number of projection-CT pairs for training, which are tedious and harmful to collect, and fail to generalize across different CT datasets due to domain discrepancies, while traditional methods struggle with sparse-view cases and inefficient computations.

Method used

A line-segment-density machine learning system that models internal dependencies using self-attention within X-ray line segments, combined with a masked local-global sampling strategy to efficiently capture contextual and geometric structures, reducing computational complexity to linear in the number of positions.

Benefits of technology

The system achieves significantly improved 3D reconstruction accuracy and efficiency, outperforming state-of-the-art methods by up to 13.76 dB in PSNR and 0.0282 in SSIM across various applications, including medicine, biology, and industry, while requiring fewer training projections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025051667_30042026_PF_FP_ABST
    Figure US2025051667_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for using machine learning to derive radiodensity values for locations within a three-dimensional target object are presented. The techniques include: obtaining a plurality of X-ray projections, each X-ray projection including pixels that correspond to X-ray beam intensity values after traversing the target object along a line segment through the target object; sampling from the X-ray projections; inputting the sample from the plurality of X-ray projections to a trained line-segment-density machine learning system, where the trained line-segment-density machine learning system is trained, using training X-ray projection data, to derive radiodensity values, where the trained line-segment-density machine learning system outputs radiodensity values for positions within the target object along line segments through the target object; forming, from the radiodensity values, image data characterizing an internal structure of the target object; and providing the image data.
Need to check novelty before this filing date? Find Prior Art

Description

STRUCTURE-AWARE SPARSE-VIEW X-RAY 3D RECONSTRUCTIONRelated Application

[0001] This application claims the benefit of U. S. Provisional Patent Application No. 63 / 710,219 entitled “STRUCTURE-AWARE SPARSE-VIEW X-RAY 3D RECONSTRUCTION,” filed October 22, 2024.Field

[0002] This disclosure relates generally to constructing three-dimensional (3D) representations from two-dimensional (2D) projections, e.g., according to systems such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI).Background

[0003] X-ray, known for its ability to reveal internal structures of objects, provides much richer information for 3D reconstruction than visible light. For example, compared with natural light, X-ray has stronger penetrating power to reveal more internal structures of imaged objects. Hence, X-ray is widely used for prospective imaging in medicine, biology, security, industry, etc.

[0004] Two image processing tasks in the context of X-ray imaging include novel view synthesis (NVS) and computed tomography (CT) reconstruction. NVS aims to create new projections of a scene from viewpoints not originally captured. CT reconstruction retrieves the 3D CT volume of the scanned object from multi-view X-ray projections. These two tasks are complementary with an overall objective to reconstruct 3D representations from 2D projections.

[0005] A majority of existing deep-learning-based methods for reconstructing 3D representations from 2D projections employ a powerful model such as convolutional neural network (CNN) to learn a brute-force mapping from 2D X-ray projections to 3D CT volumes. These methods require a large number of projection-CT pairs for training. Yet, CT volumes are not accessible in practice. Collecting even a small projection-CT dataset is tedious, labor-intensive, and harmful to health. In addition, these paired learning-based techniques fail to generalize from one application to another due to the large domain discrepancy of different CT datasets.Summary

[0006] According to various embodiments, a method of using machine learning to derive radiodensity values for locations within a three-dimensional target object is presented. The method includes: obtaining, for a plurality of arrangements of an X-ray source and an X-ray detector relative to the target object, a plurality of X-ray projections, wherein a respective X-ray projection comprises a plurality of pixels, and wherein a respective pixel corresponds to an X-ray beam intensity value after traversing the target object along a line segment through the target object; sampling from the plurality of X-ray projections, from which a sample from the plurality of X-ray projections is obtained; inputting the sample from the plurality of X-ray projections to a trained line-segment-density machine learning system, wherein the trained line-segment-density machine learning system is trained, using training data comprising training X-ray projection data, to derive radiodensity values, wherein the trained line-segment-density machine learning system outputs radiodensity values for a plurality of positions within the target object along line segments through the target object; forming, from the radiodensity values for the plurality of positions within the targetobject along line segments through the target object, image data characterizing an internal structure of the target object; and providing the image data.

[0007] Various optional features of the above method embodiments include the following. The target object may include a portion of a patient’s anatomy, the image data may include computed tomography image data, and the method may further include: displaying a computed tomography image of the internal structure of the portion of the patient’s anatomy based on the computed tomography image data. The sampling may include sampling a plurality of individual pixels and a plurality of patches. The sampling may include masking a respective background of a respective X-ray projection of the plurality of X-ray projections, wherein the sample from the plurality of X-ray projections excludes data from backgrounds in the plurality of X-ray projections. The trained line-segment-density machine learning system may compute selfattention within respective line segments through the target object. The trained line-segment-density machine learning system may model dependencies within respective line segments through the target object. The trained line-segment-density machine learning system may include a line-segment-based multi-head attention block. The training data may further include training view direction data corresponding to the training X-ray projection data. The sampling may include sampling less than 2% of a total area of the plurality of X-ray projections. The trained line-segment-density machine learning system may have a computational complexity that is linear in a number of positions within the target object along line segments through the target object.

[0008] According to various embodiments, a system that uses machine learning to derive radiodensity values for locations within a three-dimensional target object is presented. The system includes: a non-transitory computer readable mediumcomprising instructions; and at least one electronic processor that executes the instructions, which configure the electronic processor to perform operations comprising: obtaining, for a plurality of arrangements of an X-ray source and an X-ray detector relative to the target object, a plurality of X-ray projections, wherein a respective X-ray projection comprises a plurality of pixels, and wherein a respective pixel corresponds to an X-ray beam intensity value after traversing the target object along a line segment through the target object; sampling from the plurality of X-ray projections, from which a sample from the plurality of X-ray projections is obtained; inputting the sample from the plurality of X-ray projections to a trained line-segment-density machine learning system, wherein the trained line-segment-density machine learning system is trained, using training data comprising training X-ray projection data, to derive radiodensity values, wherein the trained line-segment-density machine learning system outputs radiodensity values for a plurality of positions within the target object along line segments through the target object; forming, from the radiodensity values for the plurality of positions within the target object along line segments through the target object, image data characterizing an internal structure of the target object; and providing the image data.

[0009] Various optional features of the above system embodiments include the following. The target object may include a portion of a patient’s anatomy, the image data may include computed tomography image data, and the system may further include: a display that shows a computed tomography image of the internal structure of the portion of the patient’s anatomy based on the computed tomography image data. The sampling may include sampling a plurality of individual pixels and a plurality of patches. The sampling may include masking a respective background of a respective X-ray projection of the plurality of X-ray projections, wherein the samplefrom the plurality of X-ray projections excludes data from backgrounds in the plurality of X-ray projections. The trained line-segment-density machine learning system may compute self-attention within respective line segments through the target object. The trained line-segment-density machine learning system may model dependencies within respective line segments through the target object. The trained line-segment-density machine learning system may include a line-segment-based multi-head attention block. The training data may further include training view direction data corresponding to the training X-ray projection data. The sampling may include sampling less than 2% of a total area of the plurality of X-ray projections. The trained line-segment-density machine learning system may have a computational complexity that is linear in a number of positions within the target object along line segments through the target object.

[0010] Combinations, (including multiple dependent combinations) of the above-described elements and those within the specification have been contemplated by the inventors and may be made, except where otherwise indicated or where contradictory.Brief Description of the Drawings

[0011] Various features of the examples can be more fully appreciated, as the same become better understood with reference to the following detailed description of the examples when considered in connection with the accompanying figures, in which:

[0012] Fig. 1 illustrates a major difference between visible light imaging and X-ray imaging;

[0013] Fig. 2 provides comparisons of X-ray novel view synthesis, depicting images provided by an example embodiment and images provided using prior art techniques;

[0014] Fig. 3 is a schematic diagram of a system for using machine learning to derive radiodensity values for locations within a three-dimensional target object, according to various embodiments;

[0015] Fig. 4 is a schematic diagram of a line-segment based attention block of a trained line-segment-density machine learning system, according to various embodiments;

[0016] Fig. 5 is a schematic diagram of a line segment-based multi-head selfattention neural network, according to various embodiments;

[0017] Fig. 6 illustrates X-ray sampling according to the prior art and X-ray sampling according to various embodiments;

[0018] Fig. 7 depicts qualitative results of a comparison of an example embodiment to various prior art techniques for a novel view synthesis task;

[0019] Fig. 8 depicts qualitative results of a comparison of the example embodiment to various prior art techniques for the CT task in four application scenarios;

[0020] Fig. 9 illustrates a robustness analysis regarding the number of training projections, comparing the performance of the example embodiment to prior art techniques when given fewer X-ray projection views; and

[0021] Fig. 10 is a flow chart for a method for using machine learning to derive radiodensity values for locations within a three-dimensional target object, according to various embodiments.Description of the Examples

[0022] Reference will now be made in detail to example implementations, illustrated in the accompanying drawings. Wherever convenient, the same reference numbers will be used throughout the drawings to refer to the same or like parts. In the following description, reference is made to the accompanying drawings that form a part thereof, and in which is shown by way of illustration specific exemplary examples in which the invention may be practiced. These examples are described in sufficient detail to enable those skilled in the art to practice the invention and it is to be understood that other examples may be utilized and that changes may be made without departing from the scope of the invention. The following description is, therefore, merely exemplary.

[0023] The NeRF technique for 3D image reconstruction from 2D projections presented in Mildenhall, et al., NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, ECCV, 2020, may be trained using projections of just one scene. Nevertheless, directly applying NeRF for X-ray scenes may achieve suboptimal results due to the fundamental differences between visible light and X-ray imaging.

[0024] Fig. 1 illustrates a major difference between visible light imaging 102 and X-ray imaging 104. As illustrated in Fig. 1, visible light imaging 102 relies on the reflection off the surface of an object. It mainly captures external features. By contrast, in X-ray imaging 104, X-rays penetrate the object and attenuate, thereby forming an image. X-ray imaging 104 primarily reveals internal structures, which provide key clues for X-ray 3D reconstruction.

[0025] Nonetheless, current NeRF-based methods overlook this critical property of X-ray imaging. First, they learn NeRF by a simple multilayer perceptron (MLP). X-ray attenuates differently when penetrating different structures. Applyingexisting RGB NeRF methods for X-ray rendering produces suboptimal results due to such differences between visible light and X-ray imaging. Further, MLP treats each point on an X-ray equally, showing limitations in modeling 3D structures of objects. For example, NAF (see Zhu et al., End-To-End Object Detection With Transformers, ECCV, 2020) follows NeRF to employ an MLP model for medical X-ray neural rendering, showing limitations in capturing complex structures of imaged objects in 3D space and 2D projection.

[0026] Second, prior art techniques mainly use a naive pixel-level ray sampling strategy in the training phase. They randomly sample X-rays corresponding to scattered pixels on the whole image coordinate system. As a result, the contextual information and geometric structures in 2D projection are not extracted well. In addition, X-ray projections are spatially sparse. Sampling X-rays on uninformative regions may lead to inaccurate results. In addition, existing methods mainly consider X-ray 3D reconstruction in limited medical scenes, while their performance on other applications is still under-explored.

[0027] Various embodiments solve the above, and other, issues in capturing 3D structures in X-ray imaging. First, some embodiments pass sampled X-ray projections to a special trained line-segment-density machine learning system, which outputs radiodensity values for positions within the target object along line segments through the target object. By computing self-attention within every line segment (projection of the X-ray), the trained line-segment-density machine learning system models internal dependencies and learns complex 3D structures of different parts penetrated by the X-ray. The trained line-segment-density machine learning system thus solves at least one problem that prior MLP systems have, namely, treating each point on an X-ray equally, which generates inaccurate results. Further, the trained line-segment-densitymachine learning system is highly efficient, enjoying linear computational complexity. (For comparison, neural network vision transformers have quadratic complexity. Thus, a direct application of vanilla vision transformers to the task of X-ray neural rendering will suffer from expensive - and potentially intractable - computational cost with respect to the number of input points.)

[0028] Second, some embodiments utilize a specialized sampling strategy, which samples not only individual pixels, but also patches of pixels. This local-global sampling strategy may also use a binary mask to segment informative foreground regions on the projection. Some embodiments crop nonoverlapping patches in these informative regions and then sample X-rays that land on the pixels inside these patches to help the trained line-segment-density machine learning system perceive local contextual information and 2D structures. For the informative regions outside the patches, some embodiments randomly sample X-rays to help the trained line-segment-density machine learning system perceive the scene’s 2D global shape and geometry. The use of the specialized sampling strategy in concert with the trained line-segment-density machine learning system solves problems of inaccurate results and inefficient computations that arise out of the prior art’s random X-ray sampling, which is deficient in extracting contextual information and geometric structures and overly focuses on uninformative regions in spatially sparse X-ray images.

[0029] Additional problems of the prior art are described as follows. Traditional cone beam CT Reconstruction algorithms are mainly divided into two categories: analytical methods and optimization-based methods. Analytical methods predict the CT volume by solving the Radon transformation and its inverse. These techniques fail in handling sparse-view cases. Optimization based algorithms treat the reconstruction as a maximum a posteriori (MAP) problem based on hand-crafted image priors andsolve it by iteratively minimizing the energy function, which takes a long time. CNNs and diffusion models have been applied to CT reconstruction require a number of data pairs for training. Some embodiments represent 3D objects using an implicit function for radiodensity along projections, which solves the aforementioned problems.

[0030] These and other features and advantages are shown and described herein in reference to the figures.

[0031] Note that the disclosed invention may be used to reconstruct 3D representations from 2D projections in any context, not limited to X-ray CT. For example, the disclosed techniques may be used for MRI, or even reconstructing NVS from surveillance or other cameras. Nevertheless, by way of illustration rather than limitation, this disclosure proceeds to described example embodiments of the invention in reference to the low-dose X-ray CT 3D reconstruction problem, which has the advantage of using less radiation than other X-ray CT techniques. In particular, this disclosure describes an example embodiment in the context of decreasing X-ray imaging projections in the circular cone beam X-ray scanning CT scenario.

[0032] Fig. 2 provides comparisons 200 of X-ray novel view synthesis, depicting images provided by an example embodiment and images provided using prior art techniques. The inventors established a larger-scale benchmark, referred to herein as “X3D,” for X-ray 3D reconstruction. As shown in Fig 2, on the X3D dataset, an embodiment surpassed state-of-the-art algorithms including, InTomo (see Zang, etal., Intratomo: Self-Supervised Learning Based Tomography Via Sinogram Synthesis And Prediction, CVPR, 2021), NeRF, NeAT (see Ruckert, et al., Neat: Neural Adaptive Tomography, TOG, 2022), NAF, and TeRF, a.k.a., TensoRF (see Chen et al., TensoRF: Tensorial Radiance Fields, ECCV, 2022) by 10.91, 15.03, 5.13, and 13.76 dB in PSNR on the scenes of medicine, biology, security, and industry, respectively.The average gains over the cited prior art techniques were over 12 dB. The visual comparisons 200 of the embodiment and the second-best algorithms on four scenes (pelvis, bonsai, box, and engine) show that the embodiment yielded more perceptually pleasing results.

[0033] Fig. 3 is a schematic diagram of a system 300 for using machine learning to derive radiodensity values for locations within a three-dimensional target object, according to various embodiments. The left part of Fig. 3 depicts the non-limiting scenario of circular cone beam X-ray scanning, where a scanner emits cone-shaped X-ray beams and captures sparse-view projections from locations at equal angular intervals. According to a non-limiting example use case, the scanner may capture about 1024 projections at each of about 50 different X-ray source locations. The system 300 includes use of a masked local-global sampling strategy 302 to sample an X-ray batch 52. In particular, the / V point positions P = {p1,...,pN] ∈ ℝN×3on each X-ray r e 52 are sampled according to the strategy 302 and input into the trained line-segment-density machine learning system 304 to produce the radiodensity D. The trained line-segment-density machine learning system 304 includes a Line Segment based Attention Block (LSAB), which is shown and described herein in reference to Fig. 4.

[0034] By contrast, RGB NeRF, for example, models colors at positions on the surface of an object, representing visible light of specific wavelengths reflecting off the surface of the object. Here, instead, X-rays penetrate the object, thereby not reflecting color information. Accordingly, the trained line-segment-density machine learning system 304 models the radiodensity property that denotes the degree to which a substance blocks or attenuates the passage of X-rays or other ionizing radiation.Because the radiodensity only depends on the point position, the neural radiodensity fields may be modeled as follows, by way of non-limiting example.FQL(x,y,z) -> p (1)

[0035] In Equation (1), FQLrepresents the mapping function of the trained line-segment-density machine learning system 304 with weights θLand p e K denotes the radiodensity. According to the Beer-Lambert law, the intensity of an X-ray is reduced by the exponential integration of the traversed object’s radiodensity. Hence, the ground-truth intensity Zst(r) e K of the X-ray r(t) = o + td e IR3with the near and far bounds tnand tf e IR can be formulated as follows, by way of non-limiting example.

[0036] In Equation (2), Iois the initial intensity. By discretizing Equation (2), the predicted projection intensity Zpred(r) e K may be derived as follows, by way of nonlimiting example.

[0037] In Equation (3), ptrepresents the predicted radiodensity of the / -th sampled point and δi= ||pi+1— pi|| is the distance between adjacent points. The training objective is to minimize the total squared error £ between the predicted and ground-truth intensities in the training X-ray batch. as follows, by way of non-limiting example.

[0038] In Equation (4), the term Igt(r) represents the pixel value of the projection. The loss £ is depicted on the right side of Fig. 3.

[0039] As described above, X-ray imaging reveals internal structures of imaged objects, which provide key clues for 3D reconstruction. Yet, prior art techniquesoverlook this important imaging property. For example, similar to RGB NeRF algorithms, existing X-ray NeRF methods mainly adopt a simple MLP model to learn the implicit neural representations. X-ray attenuates differently when penetrating different structural contents. Yet, the MLP model treats each sampled point on an X-ray equally, showing limitations in modeling the 3D structures penetrated by the X-ray.

[0040] Some embodiments solve this problem through the use of a trained line-segment-density machine learning system 304 as shown in Fig. 3. The point position P is firstly fed into a hash encoding module J-C to produce point feature F e IRWXCas F = (P). Then F undergoes four LSABs with a skip connection and two fclayers to derive the point radiodensity D e IRW.

[0041] Fig. 4 is a schematic diagram of an LSAB 400 of a trained line-segment-density machine learning system, e.g., 304, according to various embodiments. As shown in Fig. 4, according to various embodiments, an LSAB includes a fully connected (fc) layer, two layer normalization (LN), a feed-forward network (FFN), and a Line Segment-based Multi-head Self-Attention (LS-MSA) neural network 500. Details of the LS-MSA 500 are shown and described in reference to Fig. 5.

[0042] LSAB 400 is the basic unit of a trained line-segment-density machine learning system, e.g., 304. A component of LSAB 400 is the LS-MSA neural network 500, which captures internal structural dependencies by computing self-attention within each line segment of an X-ray.

[0043] Fig. 5 is a schematic diagram of an LS-MSA neural network 500, according to various embodiments. As illustrated in Fig. 5, an input point feature X e IRWXCis partitioned into M segments along the point dimension as follows, by way of non-limiting example.X = [X1, X2. XM]T(5)

[0044] In Equation (5), Xfe and / = 1, 2,..., M. Then each Xfis linearly projected into query Qfex,by three fully connected layers as follows, by way of non-limiting example.

[0045] In Equations (6), WQi, WKi, and WViG IRCXCare learnable parameters for the fully connected layers; biases may also be included. Subsequently, WQi, WKi, and WViare uniformly split into k heads along the channel dimension as follows, by way of non-limiting example.Qt = [Qi, Q. Qi]Ki = [K K?. K^] (7)Vi = [V, V?. VH

[0046] The dimension of each head is dh=c / ^. Fig. 4 illustrates the situation with k = 1 by way of non-limiting example for illustrative purposes. Then the selfattention within each head H / may be computed as follows, by way of non-limiting example.H- = Attn(

[0047] In Equation (8), al e IR is a learnable parameter that adaptively scales the inner product before the softmax function. Successively, k heads are concatenated in channel dimension to pass through a fully connected layer and then plus a positional embeddingto derive the / -th output Yj eas follows, by way of non-limiting example.

[0048] In Equation (9), Wj G IRCXCare learnable parameters of the fully connected layer. Finally, we group the outputs of M segments in point dimension to obtain the output feature Y e IRWXCas follows, by way of non-limiting example.Y = [YI, Y2. YM]T(10)

[0049] By capturing the interactions of points within each line segment, the trained line-segment-density machine learning system 304 is more capable of perceiving the complex internal 3D structures of different parts penetrated by the X-ray and therefore modeling the implicit neural radiodensity fields in Equation (1) more accurately than prior art techniques.

[0050] Note that the computational complexity of the LS-MSA neural network 500, and therefore that of the trained line-segment-density machine learning system 304, is linear in the number of input points ( / V). In big-0 notation, O(LS-MSA) = O(N)By comparison, the complexity of a vision transformer neural network is quadratic in the number of input points. This heavy computational burden impedes the application of vision transformer neural networks for X-ray 3D reconstruction. The significantly reduced computational complexity of LS-MSA allows for the integration of LS-MSA neural network 500 into each basic unit LSAB 400 of the trained line-segment-density machine learning system 304, while maintaining practical computational efficiency.

[0051] Fig. 6 illustrates prior art X-ray sampling 602 and X-ray sampling 604 according to various embodiments. The naive prior art X-ray sampling strategy 602 samples X-rays that land on scattered pixels. By contrast, the masked local-global X-ray sampling 604 according to various embodiments performs pixel-level and patchlevel sampling on foreground regions.

[0052] Existing NeRF algorithms mainly adopt a naive pixel-level X-ray sampling strategy, as 602. They randomly sample X-rays corresponding to scattered pixels on the whole image coordinate system for training. This naive strategy has two drawbacks. First, it shows limitations in extracting local contextual and geometric representations in 2D projection because the semantic information from neighbor pixels is not captured. Second, X-ray images are spatially sparse. Some randomly sampled X-rays may land on the background dark regions of the projection, such as the pixel pbg shown in Fig. 6. These X-rays do not penetrate the object and thus are not imaged on the projection. In other words, these X-rays are uninformative because they do not characterize the radiodensity property of the object being imaged. Learning with these X-rays, as occurs in some prior art systems, degrades accuracy.

[0053] Various embodiments solve the aforementioned problem through the use of a masked local-global X-ray sampling strategy (MLG) 604. MLG first uses a mask M ∈ ℝH×Wto segment the imaged foreground regions. M is derived by binarizing the projection I ∈ ℝH×Wwith a threshold T e K as M = ퟙ[I>τ]. Subsequently,’S ’S HW to avoid redundant sampling, M is partitioned into a set W ∈ ℝ2of — non-overlapping windows with size S * S. Letdenote the set of windows that are entirely contained in the foreground regions as follows, by way of non-limiting example.W / = {W G W | W = 1SXS} (11)

[0054] To capture local semantic information of the object, patch-level sampling is performed. Specifically,windowsare randomly selected from TT, as illustrated in Fig. 6. Then the X-ray setcorresponding to the pixels within W;can be formulated as follows, by way of non-limiting example.Ray(p) (12)

[0055] In Equation (12), Ray(p) is a function that maps from a pixel p to its corresponding X-ray. Furthermore, to assist embodiments in better capturing global contextual representations and perceiving the overall geometric shape of the imaged object, pixel-level sampling is performed. In particular, Ngpixels P are randomly selected from the foreground regions excluding the area ofto avoid repeated ray sampling, as depicted in Fig. 6. The value for P can be formulated as follows, by way of non-limiting example.

[0056] Then the X-ray set Rgcorresponding to P may be obtained as follows, by way of non-limiting example.Rg= ∪p∈PRay(p) (14)

[0057] Finally, the training X-ray batch 7? may be selected as the union of Rsand Rg, which may be expressed as R = Rs∪ Rg.

[0058] Using MLG ray sampling strategy, embodiments can more effectively capture the contextual information and model the geometric structures of the imaged object on 2D projection. According to some embodiments, MLG may sample less than 2% of a total area of the plurality of X-ray projections, e.g., about 1024 projections, in the form of individual pixels and pixels from patches, from each 256x256 detection grid.

[0059] This disclosure proceeds to describe construction and testing of a nonlimiting example embodiment.

[0060] In order to rigorously test an embodiment, the inventors developed a diverse data set. The prior art mainly conducts X-ray 3D reconstruction testing on a limited number of strictly medical applications. For example, NAF was evaluated ononly five medical scenes. The performance of NeRF-based methods on other X-ray applications is under-explored. To fill this gap, the inventors assembled a large-scale dataset, referred to as “X3D,” containing 15 scenes and covering 4 applications, i.e., medicine, biology, security, and industry. The CT volumes were collected from public datasets. The inventors used the tomographic method TIGRE to generate projections by scanning CT volumes with 3% noise in the range of 0° ~ 180°.

[0061] The example embodiment was implemented using PyTorch. The embodiment was trained with the Adam optimizer (β₁ = 0.9 and β₂ = 0.999) for 3000 iterations. The learning rate was initially set to 1×10-4and halved every 1500 iterations during the training procedure. The batch size of X-rays was set to 2048, 1024 of which were from patch level sampling and the other 1024 from pixel-level sampling. The embodiment uniformly sampled 320 points along each X-ray. For each scene, 50 projections were used to train, and another 50 projections were used to test the performance of NVS, and its CT volume to evaluate the results of CT reconstruction. All experiments were conducted on an RTX 8000 GPU. The peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) were used as evaluation metrics.

[0062] Table 1 presents quantitative results of PSNR and SSIM of the example embodiment on the NVS task, as compared with five prior art techniques: InTomo, NeRF, NeAT, TensoRF, and NAF.Table 1

[0063] In Table 1, the best results are in bold and the second best are underlined. The input and output of all methods are set the same as Equation (1 ) for fair comparison. It can be observed that the example embodiment significantly outperforms the state-of-the-art techniques on all scenes. Specifically, when compared with the recent best general RGB NeRF algorithm TensoRF, the embodiment is 13.70 dB (51.37 - 37.67) and 0.0282 (0.9994 - 0.9712) higher in PSNR and SSIM. When compared with the recent best medical NeRF method NAF, the embodiment surpasses it by 12.56 dB in PSNR and 0.0209 in SSIM. The average improvements of the embodiment on the scenes of medicine, biology, security, and industry are 10.91, 15.03, 5.13, and 13.76 dB, as shown in the bar charts of Fig. 2.

[0064] Fig. 7 depicts qualitative results 700 of the comparison of the example embodiment to various prior art techniques for the NVS task. As can be observed from the zoomed-in patches, the prior art techniques are less effective in synthesizing novel projections. They either produced blurry images or failed to reconstruct structural contents. By contrast, the example embodiment yielded more visually pleasing results with clearer textures and more fine grained details while preserving more complete geometric structures. Additional visual comparisons are shown in Fig. 2.

[0065] Table 2 presents the quantitative results of PSNR and SSIM of the example embodiment on the CT task, as compared with five prior art techniques: InTomo, NeRF, NeAT, TensoRF, and NAF.Table 2

[0066] In Table 2, the best results are in bold and the second best are underlined. For fairness, projection-CT paired learning-based algorithms are not compared, but instead the focus is on comparing techniques that only require X-ray projections of single scenes for training or direct processing. In addition to the five SOTA NeRF-based algorithms, the example embodiment was also compared with an analytical method (FDK, see Feldkamp, et al., Practical Cone-Beam Algorithm, Josa A, 1984) and two optimization-based algorithms (ASDPOCS, see Sidky, et al., Image Reconstruction In Circular Cone-Beam Computed Tomography By Constrained, Total- Variation Minimization, Physics in Medicine & Biology, 2008 and SART, see Andersen, et al., Simultaneous Algebraic Reconstruction Technique (SART): A Superior Implementation Of The Art Algorithm, Ultrasonic imaging, 1984). The example embodiment produced the best results on all scenes. In particular, the example embodiment dramatically outperformed previous NeRF-based, optimization-based, and analytical algorithms by over 2.49, 4.92, and 12.13 dB.

[0067] Fig. 8 depicts qualitative results 800 of the comparison of the example embodiment to various prior art techniques for the CT task in four application scenarios, including medicine (head), biology (carp), security (box), and industry (teapot). The prior art techniques either produce over-smooth images, blurring the structural contents, or introduce distracting artifacts. By contrast, the example embodiment is more favorable to reconstruct vivid, high-frequency details, such as sharp edges, while maintaining spatial smoothness of homogeneous regions within complex structures.

[0068] Fig. 9 illustrates a robustness analysis regarding the number of training projections, comparing the performance of the example embodiment to prior art1techniques when given fewer X-ray projection views. The results are plotted as two line charts 902, 904, where the vertical axis is PSNR (in dB performance) and the horizontal axis is the number of training projections. The example embodiment reliably surpasses the prior art techniques by large margins when given different numbers of training projections on both NVS (chart 902) and CT reconstruction (chart 904) tasks. Surprisingly, when using even only 60% of training projections, the example embodiment still outperforms the prior art on the NVS task. These results clearly exhibit the superiority and robustness of the example embodiment.

[0069] The inventors also compared the example embodiment to a system that utilizes the global multi-head self-attention (G-MSA) mechanism of a vanilla transformer neural network, by replacing the LS-MSA module in the example embodiment with a G-MSA module. For fair comparison, the system parameters were kept the same by fixing the number of channels and heads. The results showed that the trained line-segment-density machine learning system of the example embodiment significantly outperformed the vanilla transformer example by 5.30 and 1.28 dB on the NVS and CT reconstruction tasks, respectively, while only requiring 3.41% of vanilla transformer’s computational complexity. This evidence suggests the efficiency advantage of various embodiments.

[0070] Fig. 10 is a flow chart for a method 1000 for using machine learning to derive radiodensity values for locations within a three-dimensional target object, according to various embodiments. The method 1000 may be implemented using a system as shown and described herein in reference to Figs. 2-9. The method 1000 may be used for any technique that converts 2D projections to 3D images, including, by way of non-limiting example, MRI, CT, security cameras, etc. In general, the method 1000 may be used for tomography reconstructions or NVS reconstructions.

[0071] At 1002, the method 1000 includes obtaining, for a plurality of arrangements of an X-ray source and an X-ray detector relative to the target object, a plurality of X-ray projections. Each X-ray projection may include a plurality of pixel values, and each pixel value may correspond to an X-ray beam intensity value after traversing the target object along a line segment through the target object.

[0072] At 1004, the method 1000 includes sampling from the plurality of X-ray projections, from which a sample from the plurality of X-ray projections is obtained. The sampling may include sampling a plurality of individual pixels and a plurality of patches. The sampling may include masking a respective background of a respective X-ray projection of the plurality of X-ray projections. The sample from the plurality of X-ray projections may exclude data from backgrounds in the plurality of X-ray projections. The sampling may sample less than 2% of a total area of the plurality of X-ray projections. According to some embodiments, a MLG sampling strategy may be used.

[0073] At 1006, the method 1000 includes inputting the sample from the plurality of X-ray projections to a trained line-segment-density machine learning system. The trained line-segment-density machine learning system may include a line-segment-based multi-head attention block. The trained line-segment-density machine learning system may compute self-attention within respective line segments through the target object. The trained line-segment-density machine learning system may model structural dependencies within respective line segments through the target object.

[0074] The trained line-segment-density machine learning system may be trained, using training data that includes training X-ray projection data, to derive radiodensity values. The training data may include training view direction data (e.g., one or more angle values) corresponding to the training X-ray projection data.

[0075] The trained line-segment-density machine learning system may have a computational complexity that is linear in a number of positions within the target object along line segments through the target object.

[0076] In response to the inputting, the trained line-segment-density machine learning system outputs radiodensity values for a plurality of positions within the target object along line segments through the target object.

[0077] At 1008, the method 1000 includes forming, from the radiodensity values for the plurality of positions within the target object along line segments through the target object, image data characterizing an internal structure of the target object.

[0078] At 1010, the method 1000 includes providing the image data. The image data may be provided by displaying, e.g., on a monitor, over a network, or by storing in a medical records system, by way of non-limiting examples.

[0079] According to some embodiments, the target object may include a portion of a patient’s anatomy, and the image data may include computed tomography image data. Such embodiments may include displaying a computed tomography image of the internal structure of the patient’s anatomy based on the computed tomography image data.

[0080] Thus, embodiments may be used to solve a fundamental problem in sparse-view X-ray 3D reconstruction, namely, how to effectively capture the various and complex structures penetrated by X-rays. Example embodiments are shown and described in detail. To model 3D structural dependencies in space, embodiments may include a trained line-segment-density machine learning backbone. According to some embodiments, a trained line-segment-density machine learning system partitions an X-ray into different line segments and then computes self-attention within each piece of the X-ray. In addition, to extract 2D geometry and contextualrepresentations in projection, some embodiments utilize a MLG ray sampling strategy that contains pixel-level and patch level sampling on the informative foreground regions. An example embodiment is tested using a large-scale dataset, X3D, covering wider X-ray application scenarios. Comprehensive experiments on X3D show that the example embodiment significantly surpasses state-of-the-art techniques on the NVS and CT reconstruction tasks.

[0081] Certain examples can be performed using a computer program or set of programs. The computer programs can exist in a variety of forms both active and inactive. For example, the computer programs can exist as software program(s) comprised of program instructions in source code, object code, executable code or other formats; firmware program(s), or hardware description language (HDL) files. Any of the above can be embodied on a transitory or non-transitory computer readable medium, which include storage devices and signals, in compressed or uncompressed form. Exemplary computer readable storage devices include conventional computer system RAM (random access memory), ROM (read-only memory), EPROM (erasable, programmable ROM), EEPROM (electrically erasable, programmable ROM), flash memory, and magnetic or optical disks or tapes.

[0082] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented using computer readable program instructions that are executed by an electronic processor.

[0083] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the electronic processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0084] In embodiments, the computer readable program instructions may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the C programming language or similar programming languages. The computer readable program instructions may execute entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.

[0085] As used herein, the terms “A or B” and “A and / or B” are intended to encompass A, B, or {A and B}. Further, the terms “A, B, or C” and “A, B, and / or C” areintended to encompass single items, pairs of items, or all items, that is, all of: A, B, C, {A and B}, {A and C}, {B and C}, and {A and B and C}. The term “or” as used herein means “and / or.”

[0086] As used herein, language such as “at least one of X, Y, and Z,” “at least one of X, Y, or Z,” “at least one or more of X, Y, and Z,” “at least one or more of X, Y, or Z,” “at least one or more of X, Y, and / or Z,” or “at least one of X, Y, and / or Z,” is intended to be inclusive of both a single item (e.g., just X, or just Y, or just Z) and multiple items (e.g., {X and Y}, {X and Z}, {Y and Z}, or {X, Y, and Z}). The phrase “at least one of” and similar phrases are not intended to convey a requirement that each possible item must be present, although each possible item may be present.

[0087] The techniques presented and claimed herein are referenced and applied to material objects and concrete examples of a practical nature that demonstrably improve the present technical field and, as such, are not abstract, intangible or purely theoretical. Further, if any claims appended to the end of this specification contain one or more elements designated as “means for [perform]ing [a function]...” or “step for [performing [a function]...”, it is intended that such elements are to be interpreted under 35 U. S. C. § 112(f). However, for any claims containing elements designated in any other manner, it is intended that such elements are not to be interpreted under 35 U. S. C. § 112(f).

[0088] While the invention has been described with reference to the exemplary examples thereof, those skilled in the art will be able to make various modifications to the described examples without departing from the true spirit and scope. The terms and descriptions used herein are set forth by way of illustration only and are not meant as limitations. In particular, although the method has been described by examples, the steps of the method can be performed in a different order than illustrated orsimultaneously. Those skilled in the art will recognize that these and other variations are possible within the spirit and scope as defined in the following claims and their equivalents.

Claims

What is claimed is:

1. A method of using machine learning to derive radiodensity values for locations within a three-dimensional target object, the method comprising:obtaining, for a plurality of arrangements of an X-ray source and an X-ray detector relative to the target object, a plurality of X-ray projections, wherein a respective X-ray projection comprises a plurality of pixels, and wherein a respective pixel corresponds to an X-ray beam intensity value after traversing the target object along a line segment through the target object;sampling from the plurality of X-ray projections, from which a sample from the plurality of X-ray projections is obtained;inputting the sample from the plurality of X-ray projections to a trained line-segment-density machine learning system, wherein the trained line-segment-density machine learning system is trained, using training data comprising training X-ray projection data, to derive radiodensity values, wherein the trained line-segment-density machine learning system outputs radiodensity values for a plurality of positions within the target object along line segments through the target object;forming, from the radiodensity values for the plurality of positions within the target object along line segments through the target object, image data characterizing an internal structure of the target object; andproviding the image data.

2. The method of claim 1, wherein the target object comprises a portion of a patient’s anatomy, and wherein the image data comprises computed tomography image data, the method further comprising:displaying a computed tomography image of the internal structure of the portion of the patient’s anatomy based on the computed tomography image data.

3. The method of claim 1, wherein the sampling comprises sampling a plurality of individual pixels and a plurality of patches.

4. The method of claim 1, wherein the sampling comprises masking a respective background of a respective X-ray projection of the plurality of X-ray projections, wherein the sample from the plurality of X-ray projections excludes data from backgrounds in the plurality of X-ray projections.

5. The method of claim 1, wherein the trained line-segment-density machine learning system computes self-attention within respective line segments through the target object.

6. The method of claim 5, wherein the trained line-segment-density machine learning system models dependencies within respective line segments through the target object.

7. The method of claim 1, wherein the trained line-segment-density machine learning system comprises a line-segment-based multi-head attention block.

8. The method of claim 1, wherein the training data further comprises training view direction data corresponding to the training X-ray projection data.

9. The method of claim 1, wherein the sampling comprises sampling less than 2% of a total area of the plurality of X-ray projections.

10. The method of claim 1, wherein the trained line-segment-density machine learning system has a computational complexity that is linear in a number of positions within the target object along line segments through the target object.

11. A system that uses machine learning to derive radiodensity values for locations within a three-dimensional target object, the system comprising: a non-transitory computer readable medium comprising instructions; and at least one electronic processor that executes the instructions, which configure the electronic processor to perform operations comprising:obtaining, for a plurality of arrangements of an X-ray source and an X-ray detector relative to the target object, a plurality of X-ray projections, wherein a respective X-ray projection comprises a plurality of pixels, and wherein a respective pixel corresponds to an X-ray beam intensity value after traversing the target object along a line segment through the target object;sampling from the plurality of X-ray projections, from which a sample from the plurality of X-ray projections is obtained;inputting the sample from the plurality of X-ray projections to a trained line-segment-density machine learning system, wherein the trained line-segment-density machine learning system is trained, using training data comprising training X-ray projection data, to derive radiodensity values, wherein the trained line-segment-density machine learning system outputs radiodensity values for a plurality of positions within the target object along line segments through the target object;forming, from the radiodensity values for the plurality of positions within the target object along line segments through the target object, image data characterizing an internal structure of the target object; andproviding the image data.

12. The system of claim 11, wherein the target object comprises a portion of a patient’s anatomy, and wherein the image data comprises computed tomography image data, wherein the system further comprises:a display that shows a computed tomography image of the internal structure of the portion of the patient’s anatomy based on the computed tomography image data.

13. The system of claim 11, wherein the sampling comprises sampling a plurality of individual pixels and a plurality of patches.

14. The system of claim 11, wherein the sampling comprises masking a respective background of a respective X-ray projection of the plurality of X-ray projections, wherein the sample from the plurality of X-ray projections excludes data from backgrounds in the plurality of X-ray projections.

15. The system of claim 11, wherein the trained line-segment-density machine learning system computes self-attention within respective line segments through the target object.

16. The system of claim 15, wherein the trained line-segment-density machine learning system models dependencies within respective line segments through the target object.

17. The system of claim 11, wherein the trained line-segment-density machine learning system comprises a line-segment-based multi-head attention block.

18. The system of claim 11, wherein the training data further comprises training view direction data corresponding to the training X-ray projection data.

19. The system of claim 11, wherein the sampling comprises sampling less than 2% of a total area of the plurality of X-ray projections.

20. The system of claim 11, wherein the trained line-segment-density machine learning system has a computational complexity that is linear in a number of positions within the target object along line segments through the target object.