Cross-region and cross-view learning for sparse view cone beam computed tomography reconstruction
By using cross-regional and cross-view learning within the C2RV framework, the computational cost and radiation dose issues in sparse view reconstruction of CBCT are resolved, enabling efficient and low-dose reconstruction of complex anatomical structures and generating high-quality 3D CT volumes.
Patent Information
- Application Number
- CN202510442936.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-04-10
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies for sparse view reconstruction in cone-beam computed tomography (CBCT) suffer from high computational costs, large radiation doses, and difficulty in handling complex anatomical structures, especially lacking effective cross-view feature learning methods.
We employ a cross-region and cross-view learning (C2RV) framework, which processes a single projected view through a learnable encoder-decoder model to generate multi-view and multi-scale feature maps. We then leverage scale-view cross attention (SVC-Att) to aggregate features and enhance the representational power of points.
It improves the efficiency and quality of CBCT reconstruction, enabling the generation of high-resolution three-dimensional CT volumes with fewer views, reducing radiation dose, and adapting to complex anatomical structures.
Smart Images

Figure CN120823274A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 632,519, filed April 11, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present application generally relates to cone beam computed tomography (CBCT). More specifically, the present application relates to a sparse view CBCT reconstruction framework, namely cross-region and cross-view learning (C 2 RV) framework, which utilizes cross-region and cross-view feature learning to enhance point-wise representation. Background Art
[0004] Computed tomography (CT) has become an indispensable technology for medical diagnosis to accurately and noninvasively visualize internal anatomical structures. Compared with conventional CT (fan / parallel beam), CBCT has many advantages, including faster acquisition speed and higher spatial resolution [see reference 28 listed at the end of this specification]. Figure 1 A schematic diagram of a typical CBCT imaging apparatus 100 is shown, which has a scan source 110 for emitting a cone-shaped X-ray beam 115 and a 2D detector array 120 for measuring the power level of the received X-ray beam. The received X-ray beam forms an image 140, also called a projection, on the 2D array detector 120. Typically, hundreds of projections are required to produce a high-quality CT scan, which involves a high radiation dose from the X-rays. However, in clinical practice, exposing patients to high radiation doses can be a problem, limiting its use in settings such as interventional radiology. Therefore, reducing the number of projections may be one way to reduce radiation dose, which is also known as sparse view reconstruction.
[0005] In the past few decades, there have been many research works to study the sparse view problem for conventional CT by formulating the reconstruction as a mapping from 1D projections to 2D CT slices, where generation-based techniques [Refs 6, 7, 10, 13, 20, 20, 35, 37, 45] have been proposed to operate on the image or projection domain. However, the measurement results of cone-beam CT are 2D projections (e.g. Figure 1 This results in an increase in dimensionality compared to conventional CT. This means that extending previous conventional CT reconstruction methods to CBCT encounters problems such as high computational cost [Ref. 18].
[0006] Implicit neural representations (INRs) have recently been widely applied to 3D reconstruction, including novel view synthesis and object reconstruction. To handle sparse views or even single view cases, geometric priors (e.g., surface points [Ref. 40] and normals [Ref. 41]) or parametric shape models [Refs. 11, 38, 39, 46] (e.g., SMPL [Ref. 19] and SMPL-X [Ref. 24]) have been introduced to improve robustness and generalization. However, unlike visible light, X-rays have a higher frequency and pass through the surfaces of many materials. Therefore, depth or surface information cannot be measured in the projections. In addition, since the internal anatomy of the human body is more complex than the surface model, it is difficult to establish a CT-specific parametric model.
[0007] Although INR has been introduced into CBCT reconstruction in recent years, self-supervised schemes based on neural radiance fields (NeRFs) [Refs. 3, 31, 44] still require dozens of views (i.e., 20 to 50) due to the lack of prior knowledge. On the other hand, current data-driven schemes (e.g., DIF-Net [Ref. 18]) may perform poorly in cases with complex anatomical structures due to the following two reasons: (1) local features queried from projections may have difficulty in identifying different organs with low contrast in projections; and (2) projections of different views are treated equally, while some views do present more information about specific organs than other views. Figure 2 For example, a left-right view 260 and an anterior-posterior view 270 are shown, both of which show the components of the knee joint: femur 211, tibia 212, patella 214, and fibula 213. The left-right view 260 clearly shows the patella 214, while in the anterior-posterior view 270, the patella 214 overlaps with the femur 211.
[0008] There is a need in the art for an improved technique for reconstructing CBCT images that addresses the limitations of the prior work described above. Summary of the Invention
[0009] One aspect of the present disclosure is to provide a computer-implemented method for reconstructing a three-dimensional CT volume from a plurality of projection views generated in CBCT imaging.
[0010] The method comprises the following steps: determining a plurality of points in a three-dimensional CT volume such that the three-dimensional CT volume is reconstructed by estimating an attenuation coefficient of a single point; processing a single projection view using a learnable encoder-decoder model to thereby generate a decoder output feature map and an encoder output feature map for the single projection view, wherein the learnable encoder-decoder model is shared by the multiple projection views when processing the single projection view; querying a plurality of multi-view pixel-aligned features for the single point using corresponding decoder output feature maps generated for the multiple projection views; generating a plurality of multi-view feature maps at different scales, the different scales comprising a highest resolution and one or more reduced resolutions, wherein by grouping together the corresponding encoder output feature maps generated for the multiple projection views, A first multi-view feature map generated at the highest resolution is obtained, and a corresponding multi-view feature map generated at a single reduced resolution is obtained by downsampling the first multi-view feature map; the multiple multi-view feature maps at the different scales are respectively back-projected into corresponding three-dimensional spaces voxelized according to the different scales to thereby form multiple multi-scale three-dimensional volume representations; the multiple multi-scale three-dimensional volume representations are used to query the multiple multi-scale voxel alignment features for the single point; and the multiple multi-view pixel alignment features and the multiple multi-scale voxel alignment features are aggregated according to scale-view cross-attention to estimate the attenuation coefficient of the single point, so as to advantageously utilize cross-region and cross-view feature learning to enhance the representation of the single point before estimating the attenuation coefficient.
[0011] In some embodiments, the plurality of multi-view pixel-aligned features for the single point are obtained from corresponding decoder output feature maps by querying the view-specific pixel-aligned features for the single point under the single projective view using the decoder output feature map. The corresponding view-specific pixel-aligned features generated for the multiple projective views are regarded as the plurality of multi-view pixel-aligned features for the single point.
[0012] In some embodiments, the view-specific pixel-aligned features for the single point under the single projective view are obtained by interpolating the decoder output feature map. In some embodiments, the decoder output feature map is interpolated using k-linear interpolation, where k is an integer greater than one.
[0013] In some embodiments, the step of querying the plurality of multiscale voxel alignment features for the single point using the plurality of multiscale 3D volume representations comprises: interpolating the plurality of multiscale 3D volume representations to generate a plurality of scale-specific voxel alignment features for the single point; concatenating the plurality of scale-specific voxel alignment features to generate a concatenated voxel alignment feature for the single point; and aggregating the concatenated voxel alignment features to generate the plurality of multiscale voxel alignment features, such that channel sizes of the plurality of multiscale voxel alignment features are consistent with channel sizes of the multi-view pixel alignment features. In some embodiments, each of the plurality of multiscale 3D volume representations is interpolated using k-linear interpolation, where k is an integer greater than one.
[0014] When aggregating the concatenated voxel-aligned features to generate the plurality of multi-scale voxel-aligned features, a multi-layer perceptron (MLP) may be used to map channel sizes of the plurality of multi-scale voxel-aligned features to be consistent with the channel size of the multi-view pixel-aligned features.
[0015] In certain embodiments, the step of aggregating the multiple multi-view pixel-aligned features and the multiple multi-scale voxel-aligned features according to scale-view cross-attention to generate an attenuation coefficient for the single point includes: applying self-attention to the multiple multi-view pixel-aligned features to perform cross-view attention across the multiple multi-view pixel-aligned features to generate a plurality of attention-weighted pixel-aligned features; applying cross-attention between the multiple multi-scale voxel-aligned features and the multiple attention-weighted pixel-aligned features to thereby generate a plurality of cross-region cross-view features for the single point; and estimating the attenuation coefficient from the cross-region cross-view features.
[0016] In some embodiments, the attenuation coefficient is estimated from the cross-region cross-view features by using a linear layer to process the cross-region cross-view features.
[0017] The method may also include a step of aggregating the multiple multi-view pixel-aligned features and the multiple multi-scale voxel-aligned features using a learnable aggregation and estimation model to estimate the attenuation coefficient, wherein the learnable aggregation and estimation model includes: multiple scale view cross-attention (SVC-Att) modules stacked together, which are used to apply self-attention to the multiple multi-view pixel-aligned features and apply cross-attention between the multiple multi-scale voxel-aligned features and the multiple attention-weighted pixel-aligned features, wherein the multiple SVC-Att modules output the multiple cross-region cross-view features for the single point; and a linear layer located after the multiple SVC-Att modules, which is used to estimate the attenuation coefficient from the cross-region cross-view features.
[0018] In some embodiments, the learnable encoder-decoder model is implemented in the form of a U-Net.
[0019] The method may further include training the learnable encoder-decoder model before using the learnable encoder-decoder model to process the single projected view.
[0020] Similarly, the method may further include training the learnable aggregation and estimation model before aggregating the plurality of multi-view pixel-aligned features and the plurality of multi-scale voxel-aligned features according to scale-view cross-attention to estimate the attenuation coefficient.
[0021] Other aspects of the present disclosure are shown in the embodiments described later. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A schematic diagram of a typical CBCT imaging setup is shown, in which a cone-shaped X-ray beam is emitted from a scanned source and a 2D detector array measures the power level of the received X-ray beam.
[0023] Figure 2 Shown are left-right and anterior-posterior views of the knee joint formed by the femur, tibia, patella, and fibula, showing that the patella and femur overlap in the anteroposterior view, but not in the left-right view.
[0024] Figure 3 The disclosed sparse view reconstruction framework C is shown in FIG. 2 A conceptual diagram of an RV overview.
[0025] Figure 4 A schematic diagram of the learnable aggregation and estimation model is provided, which includes the SVC-Att module.
[0026] Figure 5 The visualization of chest CT with 6-view reconstruction is provided (from top to bottom: axial slice; coronal slice; and sagittal slice), and the PSNR / SSIM (dB / ×10 -2 )value.
[0027] Figure 6 Visual examples of reconstruction from different numbers of projection views (i.e., 6, 8, and 10) are provided, where the highlighted areas are magnified to show that the reconstruction results of the present application have richer details than those of other methods.
[0028] Figure 7 Provides visualization of lung segmentation on 6-view reconstructed chest CT.
[0029] Figure 8A workflow is shown to illustrate exemplary steps of the computer-implemented method disclosed herein for reconstructing a 3D CT volume from multiple projection views generated in CBCT imaging.
[0030] Figure 9 A workflow is shown to illustrate exemplary steps of the method disclosed herein involving querying multi-scale voxel-aligned features for a single point using multiple multi-scale 3D volume representations.
[0031] Figure 10 A workflow is shown to illustrate exemplary steps of the method disclosed herein involving aggregating multiple multi-view pixel-aligned features and multiple multi-scale voxel-aligned features according to scale-view cross-attention to estimate the attenuation coefficient of a single point.
[0032] Those skilled in the art will appreciate that elements in the figures are illustrated for simplicity and clarity and are not necessarily drawn to scale. DETAILED DESCRIPTION
[0033] Abbreviations
[0034] 1D One-dimensional
[0035] 2D
[0036] 3D three-dimensional
[0037] ART Algebraic Reconstruction Technology
[0038] ASD Average Surface Distance
[0039] C 2 RV cross-region and cross-view learning
[0040] C-Att Cross Attention
[0041] CBCT cone-beam computed tomography
[0042] CNN Convolutional Neural Network
[0043] CT computed tomography
[0044] DRR Digitally Reconstructed Radiograph
[0045] Distance from DSO source to origin
[0046] FBP filtered back projection
[0047] INR implicit neural representation
[0048] MLP Multilayer Perceptron
[0049] MS-3DV Multi-scale 3D volume representation
[0050] MSE mean square error
[0051] PSNR Peak signal-to-noise ratio
[0052] S-Att Self-Attention
[0053] SMPL Skinned Multi-Person Linear Model
[0054] SSIM structural similarity
[0055] SVC-Att Scale-View Cross-Attention
[0056] As used herein, "projection" in the context of CT imaging (including CBCT imaging) refers to the image formed on the X-ray detector by the resulting X-ray beam obtained from the original X-ray beam after the original X-ray beam passes through the imaged object. The object is usually a human body. In this specification and the appended claims, "projection" and "projection view" in the context of CT imaging can be used interchangeably. Figure 1 The CBCT imaging apparatus 100 in FIG. 1 is used as an example. After the cone-shaped X-ray beam 115 passes through the human body object 15, the X-ray beam 115 creates a 2D image 140 formed on the 2D detector array 120. The 2D image 140 is a projection or projection view.
[0057] To address the limitations of prior work, this paper presents a novel sparse view CBCT reconstruction framework, referred to as C 2 RV, which utilizes cross-region and cross-view feature learning to enhance point-by-point representation. 2 After the RV framework, the embodiments of the present disclosure will be explained based on the disclosed details, examples, applications, etc. of the framework.
[0058] 1.C 2 RV Framework Overview
[0059] To explain C more specifically 2 RV framework, this disclosure first introduces MS-3DV, in which each feature is obtained by back-projecting multi-view features of different scales into 3D space. Explicit MS-3DV realizes cross-region learning in 3D space, providing richer information that helps to better identify different organs. Therefore, the features of the query points can be queried in a hybrid manner, namely, multi-scale voxel aligned features from MS-3DV and multi-view pixel aligned features from projection. Then, SVC-Att is proposed to adaptively learn aggregation weights through self-attention and cross-attention, instead of considering the queried features equally. Finally, the multi-scale and multi-view features are aggregated to estimate the attenuation coefficient. C is then used on two CT datasets (i.e., chest and knee joints).2 RV is quantitatively and qualitatively evaluated. Extensive experiments demonstrate that the proposed C 2 RV consistently outperforms previous state-of-the-art methods by a significant margin.
[0060] 2. Used to develop C 2 Related Work on RV Framework
[0061] In explaining C 2 Before the RV framework, we must first mention the framework used to develop C 2 RV related work.
[0062] In the field of computer vision, especially in 3D vision, reconstruction has attracted much attention in recent years. Below, we will review the related work on sparse view reconstruction based on traditional parallel / fan-beam CT, cone-beam CT, and general 3D.
[0063] 2.1. Sparse View CT Reconstruction
[0064] Conventional parallel / fan-beam CT reconstruction can be viewed as reconstructing 2D CT slices from 1D projections. Existing learning-based methods mainly include image domain, projection domain, and dual-domain methods. Specifically, image domain methods [References 6, 10, 13, 20, 35, 45] apply FBP to reconstruct coarse CT slices with streak artifacts and utilize convolutional neural networks (CNNs) such as U-Net [Reference 25] and DenseNet [Reference 9] for denoising and detail refinement. When these methods are extended to CBCT reconstruction, the network should be modified to a 3D CNN, resulting in a significant increase in computational cost. Another approach is to use these methods for slice-by-slice (2D) denoising [Reference 15], but 3D spatial consistency cannot be guaranteed.
[0065] Projection domain methods directly perform computation on sparse visual images by mapping projections to CT slices [Ref. 7] or recovering full view projections [Ref. 37]. Figure 1 [Reference 32] utilized a generative model based on scoring and proposed a sampling method to reconstruct an image that is consistent with both the measurement process and the observed measurement results (i.e., projections). Chung et al. [Reference 2] further introduced a 2D diffusion model into iterative reconstruction. Dual-domain methods operate on both the projection domain and the image domain by combining the denoising processes of the two domains [References 17, 20] or establishing a dual-domain consistency model [Reference 34]. However, projection-based operations cannot be extended to CBCT reconstruction because the measurement process (cone beam compared to parallel / fan beam) is different.
[0066] 2.2. Sparse View CBCT Reconstruction
[0067] Unlike conventional parallel / fan-beam CT, the measurements of cone-beam CT are 2D projections, which means that reconstruction should be formulated as reconstructing a 3D CT volume from multiple 2D projections. Conventional filtered back projection (FDK [Ref. 4]) and ART-based iterative methods [Refs. 1, 5, 22] often suffer from severe streak artifacts and poor image quality when the number of projections is drastically reduced. Recently, learning-based methods for single-view / orthogonal-view CBCT reconstruction have been proposed [Refs. 12, 14, 30, 42], but these methods are specifically designed for single-view / orthogonal-view reconstruction [Refs. 12, 14, 42] or for patient-specific data [Ref. 30], making it difficult to extend them to general sparse-view reconstruction.
[0068] On the other hand, implicit neural representations [Refs. 21, 26] have been introduced to represent CBCT as an attenuation field [Refs. 3, 44] or an intensity field [Ref. 18]. Self-supervised methods including NAF [Ref. 44] and NeRP [Ref. 31] can simulate the measurement process and minimize the error between the real projection and the synthesized projection. However, these methods require a long time to optimize for each sample and are only applicable to reconstructions from tens of views (i.e., 20 to 50) due to the lack of prior knowledge. DIF-Net [Ref. 18] is a data-driven method that formulates the above problem as learning a mapping from sparse projections to an intensity field. However, DIF-Net treats different projections equally and only queries local semantic features for each sample point, resulting in limited reconstruction quality when dealing with complex anatomical structures such as the chest.
[0069] 2.3. Sparse Vision Figure 3 D reconstruction
[0070] In the field of 3D computer vision, implicit representations have been widely used for novel view synthesis [References 21, 40, 41, 43] and object reconstruction [References 11, 23, 27, 38, 39, 46]. In novel view synthesis, geometric priors, such as surface points [Reference 40] and normals [Reference 41], have been introduced to extend NeRF [Reference 21] to sparse view scenes to improve generalization and efficiency. For object reconstruction, particularly digital human reconstruction, previous work [References 11, 38, 39, 46] has utilized explicit parametric SMPL(-X) models [References 19, 24] to constrain surface reconstruction and improve robustness. However, in the attenuation field of CBCT, no depth or surface information is available because X-rays directly penetrate many common materials, such as flesh. SMPL(-X) is a 3D parametric shape model designed specifically for the human surface, while the internal anatomy is too complex to design a CT-specific parametric model. Therefore, parametric shape models cannot be used for sparse view CBCT reconstruction. Furthermore, cross-view relations are rarely considered in surface-based reconstruction, as one or two views are more practical and often sufficient to learn sparse fields with the aforementioned prior knowledge.
[0071] 3. Used to develop C 2 RV Method
[0072] We first review the problem formulation and benchmark DIF-Net for sparse view CBCT reconstruction proposed in [Reference 18]. Then we formally introduce the C++ framework composed of MS-3DV and SVC-Att for cross-region and cross-view learning. 2 RV.
[0073] 3.1. Review of DIF-Net [Reference 18]
[0074] Following previous work [Refs 18, 44], we formulate the CT image as a continuous implicit function It defines a point in 3D space Attenuation coefficient (same as “intensity” in [Ref. 18]) That is, v = g(p). Therefore, in the measurement process, given N view projections (W and H are width and height respectively) and known scanning parameters (such as viewing angle, distance from source to origin), the reconstruction problem is formulated as a conditional implicit function g(·) such that
[0075] In practice, a 2D encoder-decoder (shared across different views) is used to project from N views Extract multi-view feature maps Where C is the output channel size of the decoder. For the i-th view, the projection function is expressed as It maps the 3D point p to the 2D plane where the detector is located, so that p′ i =π i (p). Then, the view-specific pixel alignment feature of p in the i-th view is defined as:
[0076]
[0077] in, Specifically, k=2, and Interp(·) in the above formula is bilinear interpolation.
[0078] The multi-view pixel-aligned features of p are expressed as Then the attenuation coefficient of p is given by:
[0079]
[0080] Where σ(·) is the aggregation function implemented using MLP (or Max-Pooling+MLP) in DIF-Net [Reference 18]. Although the above formula and implementation can provide efficient training for high-resolution sparse view reconstruction, it only considers the local pixel alignment features queried from the projection and treats different views equally, resulting in poor performance on complex anatomical structures; see the above analysis and the results in Table 1. To this end, this application proposes the following C 2 RV.
[0081] 3.2.C 2 RV frame
[0082] To address the above limitations, a C 2 RV frame. Figure 3 C is shown in 2 Overview of the RV framework 300. Given multi-view projections 305, a 2D encoder-decoder 310 is applied to extract per-view feature maps Used to query pixel alignment features In addition, the output feature map F of the encoder 311 1 Downsampling is performed to obtain a set of multi-scale multi-view feature maps 330. For each scale s, the multi-view features are back-projected into 3D space and collected to form a 3D volume representation Used to query voxel alignment features Finally, the multi-scale voxel-aligned features 350 and the multi-view pixel-aligned features 320 are adaptively aggregated via scale-view cross-attention 360 to estimate the attenuation coefficient 390 .
[0083] Low-resolution 3D volume representation. The 3D volume space is defined by voxelizing the 3D space at a low resolution r≤16 set up is the intermediate feature map of the encoder-decoder for the projection of the given i-th view. The volume feature space defined above It is achieved by back-projecting the multi-view feature map onto produced by, namely:
[0084]
[0085] in, The voxel q in is characterized by:
[0086]
[0087] in,
[0088] F i (q)=Interp(F i ,π i (q)),
[0089] and is an aggregation function, which is implemented in practice using Max-Pooling. Therefore, a 3D convolutional layer (denoted by φ) can be used to perform efficient cross-region feature learning, namely:
[0090]
[0091] MS-3DV. In order to further improve the robustness of reconstructing different anatomical structures, this application proposes to use a multi-scale 3D volume representation. Specifically, given the projection of the i-th view, the output feature map of the encoder is represented as Then a series of downsampling operators ρ are applied to generate multi-scale feature maps in s∈{2,…,S}, S is the total number of scales. Then, we define 1 ,…,r s Multi-scale 3D voxelized space And the multi-view feature maps of each scale are back-projected (Formulas 3 and 5) to obtain in,
[0092]
[0093] S∈{1,…,S}. Therefore, in addition to querying the multi-view pixel-aligned features directly from the view-specific feature map, this application also introduces the multi-scale voxel-aligned features for point p into the estimation of the attenuation coefficient, as given by the following formula:
[0094]
[0095] in, Concat[.] represents concatenation, and MLP maps the channel size of the concatenated voxel-aligned features to be consistent with the pixel-aligned features (Formula 1), i.e., C.
[0096] SVC-Att. First, we review the given reference features and source features When C-Att is defined as [Ref 33],
[0097]
[0098] Where Q = F r M q , K=F s M k , V=F s M v ,in and S-Att can be viewed as a special case of cross attention, where F s =F r .
[0099] For point p, let represents the multi-view pixel-aligned features queried from the projection (Formula 1), and Denotes the multi-scale voxel alignment feature queried from MS-3DV (Formula 7). This application proposes the SVC-Att module to adaptively aggregate the above features.
[0100] Figure 4 A schematic diagram of an SVC-Att module 410 for implementing scale-view cross-attention 360 and a learnable aggregation and estimation model 400 for performing aggregation and estimating attenuation coefficients 390 is provided. In the SVC-Att module 410, self-attention 411 is first applied to the multi-view features 320, and then cross-attention 412 is performed between the multi-scale features 350 and the multi-view features 320. Figure 4As shown, M SVC-Att modules 410 are stacked, and finally the attenuation coefficient 390 is estimated by a linear layer 430. The linear layer 430 is also called a fully connected layer. Note that the learnable aggregation and estimation model 400 includes the M SVC-Att modules 410 and the linear layer 430.
[0101] In a practical implementation, the self-attention module 411 is first applied to the multi-view features 320 performs cross-view attention. Then, the cross-attention module 412 uses the multi-scale features 350 as a reference and the output of the self-attention module 411 as a source to perform attention between the multi-scale features 350 and the multi-view features 320. For formulation, the following formula is given:
[0102]
[0103] in
[0104] In practice, M SVC-Att modules 410 are stacked and then the attenuation coefficient 390 is estimated by a linear layer 430 .
[0105] 3.3. Network training
[0106] Following [Ref 18], the reconstruction network is trained on a CT dataset, where projections are simulated from CT via DRR. Specifically, the volume CT is represented as And the projection is expressed as In continuous 3D space The reference truth attenuation field defined above is:
[0107]
[0108] Where Interp(·) is the interpolation operator (Formula 1). 2 The attenuation field of the RV estimate is given by:
[0109]
[0110] Therefore, MSE is used as the objective function to calculate the point-by-point estimation error and is given by the following formula:
[0111]
[0112] During each training iteration, 10,000 points are randomly sampled from P for loss calculation (Formula 12) to reduce memory requirements and thus achieve efficient network optimization. During inference, the 3D space is transformed to a specified resolution (e.g., 256 3 ) is voxelized, where the attenuation coefficient of the voxel is defined as the attenuation coefficient of the voxel by C2 The attenuation coefficient of the RV's estimate of its centroid. This means that the resolution can be chosen based on the desired trade-off between image quality and reconstruction speed.
[0113] Implementation plan. In practice, this application selects S=3, r 1 =16, and r s =r s-1 / 2(s≥2). Following [Reference 18], this application uses U-Net [Reference 25] with C=128 output feature channels as a 2D encoder-decoder, where the encoder output F 1 The size is The function φ(·) in Equation 5 is implemented using three layers of 3D residual convolution, which converts The channel size is mapped to C. For the aggregation method, M = 3 SVC-Att modules are stacked, and the attention module is implemented as a multi-head attention with 8 heads. During training, the stochastic gradient descent method is used to optimize C 2 The learnable parameters of RV are momentum 0.98 and initial learning rate 0.01. 2 RV was trained for 400 epochs with a batch size of 4. The learning rate of each epoch was increased by a factor of (10 -3 ) 1 / 400 was lowered.
[0114] 4. Experiment
[0115] In order to verify the C 2 To evaluate the effectiveness of RV, this paper conducts experiments on two CT datasets with different anatomical structures, including chest and knee joints. In addition to quantitative and qualitative evaluations, automatic segmentation of sparse view reconstruction results is performed to show the 2 Practical potential of CT-based RV reconstruction in downstream applications.
[0116] Experimental setup
[0117] Datasets. Experiments are conducted on two CT datasets, including a public chest CT dataset (LUNA16 [Reference 29]) and a private knee CBCT dataset collected by Lin et al. [Reference 18]. Specifically, LUNA16 [Reference 29] consists of 888 chest CT scans with a resolution of 145 × 145 × 108 mm. 3 to 375×375×509mm 3They are divided into 738 for training, 50 for validation, and 100 for testing; the knee joint dataset [Reference 18] contains 614 knee joint CBCT scan data with a resolution of 236×236×167mm 3 Up to 500×500×416mm 3 They are split into 464 for training, 50 for validation, and 100 for testing. Following the data preprocessing method of [Ref 18], each CT is resampled and cropped (or padded) to have isotropic spacing (i.e., 1.6 mm for the chest and 0.8 mm for the knee) and 256 3 Size. Multi-view Figure 2 The D projection is composed of a resolution of 256 2 DRR simulation, and the viewing angle is uniformly selected within the range of 180° (half a circle of rotation).
[0118] Evaluation Metrics. Following previous work [Refs 18, 31, 44], two quantitative metrics including PSNR and SSIM [Ref 36] are used to evaluate the reconstruction performance, where higher values indicate better image quality.
[0119] Results
[0120] Quantitative evaluation. 2 RV is compared with self-supervised schemes including FDK [Ref 4], SART [Ref 1], NAF [Ref 44], and NeRP [Ref 31] without additional training data. Data-driven schemes including 2D denoising-based methods (i.e., FBPConvNet [Ref 13], FreeSeed [Ref 20], and BBDM [Ref 16]) and implicit neural representation (INR)-based methods (i.e., PixelNeRF [Ref 43] and DIFNet [Ref 18]) are also compared. The experiments are performed with different numbers of projection views (i.e., 6 to 10) and the reconstruction resolution is 256 3 The results are shown in Table 1, which compares the performance of different methods on two CT datasets with different numbers of projection views, one is a CT dataset of the chest and the other is a CT dataset of the knee joint. The resolution of the reconstructed CT is 256 3 The reconstruction results are expressed as PSNR (dB) and SSIM (×10 -2) is used for evaluation, where a higher PSNR / SSIM value indicates better performance. Although DIF-Net [Reference 18] can achieve satisfactory performance on knee CT, its performance drops sharply when adapted to more complex anatomical structures (such as the chest), and the C 2 RV consistently performs well on different datasets. In addition, our C 2 RV outperformed the previous state-of-the-art by a significant margin, achieving PSNR / SSIM (dB / ×10 -2 ), while on knee CT, they were 2.6 / 4.5, 2.4 / 4.2 and 2.2 / 3.0 PSNR / SSIM (dB / ×10 -2 ). More importantly, even with only 6 views, C 2 RV can also reconstruct higher quality CT than other methods with 4 or more views (ie, 10 views).
[0121] Table 1. Comparison of different methods on two CT datasets (i.e., chest and knee) with various numbers of projection views. The best values are in bold and the second best values are in Underline express.
[0122] (a) LUNA16 [Ref. 29] (Chest CT)
[0123]
[0124] (b) Lin et al. [Reference 18] (CBCT of the knee joint)
[0125]
[0126] Visual comparison. Figure 5 An example of 6-view reconstruction for chest CT is visualized in Figure 2 (axial slices, coronal slices, and sagittal slices from top to bottom), where the PSNR / SSIM (dB / ×10 -2) values are shown above for qualitative comparison. Due to the lack of sufficient projection views, the reconstruction results of FDK [Reference 4] are full of streak artifacts, and NeRP [Reference 31] can only reconstruct satisfactory body and lung contours. As for FBPConvNet [Reference 13] and FreeSeed [Reference 20], since they are 2D methods that reconstruct CT slice by slice, jittering occurs near the boundaries of the body and lungs. As for Pixel-NeRF [Reference 43] and DIF-Net [Reference 18], although the reconstructed details are better than those of other methods, there are still some streak artifacts and unclear contours. C 2 The reconstruction results of RV have clearer shape contours, better internal details, and almost no streak artifacts. In addition, Figure 6 Visual examples of reconstructions from different numbers of projection views (i.e., 6, 8, and 10) are provided, showing consistent conclusions with the above. The highlighted areas are magnified, showing that the reconstruction results of our application are richer in details than those of other methods.
[0127] Downstream evaluation. In addition to quantitative and qualitative evaluation, this application also verifies the reconstructed CT in downstream tasks (i.e., segmentation). Specifically, the LungMask toolkit [Reference 8] is used to perform left / right lung segmentation on CT reconstructed by different methods. The segmentation results are shown in Tables 2 and Figure 7 As shown in Figure 7 Provides visualization of lung segmentation on 6-view reconstructed chest CT. Compared with other methods, 2 The segmentation mask on the reconstructed CT of RV is more consistent with the segmentation on the reference truth CT. This means that the C 2 RV has the potential to reconstruct high-quality CT, which can be further applied in downstream scenarios.
[0128] Table 2.6 Lung segmentation of chest CT reconstructed from 6 views. Dice coefficient (%, the larger the better) and mean surface distance (ASD, mm, the smaller the better) were evaluated. The best value is in bold and the second best value is in Underline express.
[0129]
[0130] 5. Ablation studies
[0131] This paper conducts ablation studies to explore the effectiveness of the proposed MS-3DV and SVC-Att, as well as different designs for MS-3DV. In addition, the applicant further analyzes the C 2RV is robust to different viewing angles and noise scanning parameters. All the following ablation experiments are performed at a resolution of 256 3 The 6-view chest CT reconstruction was performed.
[0132] 5.1. Proposed MS-3DV and SVC-Att
[0133] Ablation on MS-3DV and SVC-Att. This application uses DIFNet [Reference 18] as a baseline model and compares the reconstruction performance of the introduced MS-3DV and SVC-Att. In DIF-Net, multi-view features are aggregated by MLP or Max-Pooling+MLP (σ in Formula 2). The comparison between different aggregation methods is shown in Table 3. In (+MS-3DV), multi-scale voxel alignment features are spliced with maximum pooling multi-scale features. In (+SVC), this application randomly initializes the learnable vector before training as a substitute for the reference feature (i.e., in Formula 9). See also Figure 4 Both MS-3DV and SVC-Att can improve the reconstruction performance, and by jointly using the two methods, the proposed framework achieves new state-of-the-art performance.
[0134] Table 3. Ablation studies of different pooling methods (M.: MLPs [Ref. 18]; Max-M.: Max-Pooling + MLPs [Ref. 18]; SVC: scale-view cross-attention proposed in this application) and MS-3DV. PSNR (dB) and SSIM (×10 -2 ).
[0135]
[0136] Table 4 provides information about the number of scales, initial feature maps F 1 and the initial resolution r 1 Ablation study of F 1 The choice of can be the final layer feature map of the encoder or decoder. PSNR and SSIM values are evaluated on 6-view reconstruction of chest CT. As shown in Table 4, this application compares the performance of using different numbers of scales and the initial feature map F 1 and resolution r 1 The introduction of multi-scale features is very important because, compared with single-scale features, multi-scale features can provide richer information for identifying different anatomical structures, such as organs (e.g., lungs) and bones (e.g., spine). Since the size of the feature map at the third scale is too small (i.e., 4×4), this application does not further increase the number of scales (e.g., 4). For F 1The output of the encoder is better because it contains more high-level features than the decoder. According to experience, an initial resolution of 16 is the best choice for the trade-off between global (high-level) features and local (detailed) features.
[0137] Table 4. Regarding the number of scales and initial feature maps F 1 and the initial resolution r 1 Ablation study of F 1 The choice of can be the final layer feature map of the encoder or decoder. PSNR and SSIM are evaluated on 6-view reconstruction of chest CT.
[0138]
[0139] Robustness Analysis
[0140] set up represents the viewing angle in the original evaluation. The first experiment is conducted by choosing different viewing angles, i.e. Where Δα is the angle offset. Table 5 lists the robustness analysis results for different angles and noise scanning parameters. As shown in Table 5, with the change of angle, C 2 The performance of RV is stable. The second study is about the noisy scanning parameters. Taking the view angle as an example, this application assumes that the measurement process is noisy, which means that the multi-view projection is from Measured, where η i is the noise that obeys the uniform distribution U(-∈,+∈). In this case, the projection function π is still based on the original view (i.e. ) is defined because noise is unobservable. In Table 5, this application considers two scanning parameters, including the viewing angle and the distance from the source to the origin, which are the main factors related to the formula of the projection function (see the appendix in [Reference 18]). Experiments show that the C 2 RV is robust to slight variations in scanning parameters.
[0141] Table 5. Robustness analysis for different angles and noise scan parameters, including view angle and distance from the origin (DSO). For the noise scan parameters, the noise offset follows a uniform distribution, i.e., U(-∈, +ε). PSNR and SSIM were evaluated on a six-view reconstruction of chest CT.
[0142]
[0143] 6. Embodiments of the present disclosure
[0144] Based on the above disclosure about C 2Details, examples, applications, etc. of the RV framework, embodiments of the present disclosure may be developed as follows (which may be generalized).
[0145] One aspect of the present disclosure is to provide a computer-implemented method for reconstructing a 3D CT volume from multiple projection views generated in CBCT imaging. The method utilizes cross-region and cross-view feature learning to enhance pixel representation in the reconstructed 3D CT volume.
[0146] Although the disclosed method is particularly advantageous for sparse-view CBCT reconstruction, the present disclosure is not limited to sparse-view CBCT reconstruction; the disclosed method can be used for CBCT reconstruction with any number of views.
[0147] With the help of Figure 3 and Figure 8 The disclosed method is exemplified. Figure 3 Provided for illustrating C 2 Conceptual diagram of the RV framework. Figure 8 A workflow 800 is shown to illustrate exemplary steps of the disclosed method. The method includes steps 820, 830, 840, 850, 860, 870, and 880.
[0148] In step 820, as an initialization step, a plurality of points in the 3D CT volume are determined. The 3D CT volume is reconstructed by estimating the attenuation coefficient 390 of a single point from the plurality of points. A person skilled in the art will be able to determine an appropriate plurality of points for attenuation coefficient estimation based on the specific CBCT imaging situation under consideration.
[0149] In step 830, a single projection view from the multiple projection views is processed using a learnable encoder-decoder model 310, which is a machine learning model having an encoder-decoder architecture, to thereby generate a decoder output feature map and an encoder output feature map for the single projection view. Specifically, when processing a single projection view, the learnable encoder-decoder model is shared by the multiple projection views. Having an encoder-decoder architecture, the learnable encoder-decoder model 310 is formed by an encoder 311 and a decoder 312. The decoder output feature map is a feature map obtained at the output of the decoder 312. Note that the decoder output feature map is also the final feature map obtained by the learnable encoder-decoder model 310. The encoder output feature map is a feature map obtained at the output of the encoder 311. The encoder output feature map is an intermediate feature map and is not the final feature map obtained by the learnable encoder-decoder model 310. Note also that the encoder 311 and decoder 312 are 2D encoders and decoders for processing a single projection view, and the single projection view is a 2D image. In certain embodiments, the learnable encoder-decoder model 310 is implemented in the form of a U-Net.
[0150] After obtaining the corresponding decoder output feature maps (i.e. Then, in step 840, the feature map is output from the corresponding decoder A plurality of multi-view pixel alignment features for the single point are queried 320 .
[0151] Typically, a decoder output feature map obtained for a single projection view is used to query the view-specific pixel-aligned features for a single point under the single projection view, and a plurality of multi-view pixel-aligned features 320 for the single point are obtained from the corresponding decoder output feature map. The corresponding view-specific pixel-aligned features generated for the plurality of projection views 305 are regarded as the plurality of multi-view pixel-aligned features 320 for the single point. In some embodiments, the view-specific pixel-aligned features for the single point under the single projection view are obtained by interpolating the decoder output feature map according to Formula 1. In some embodiments, the decoder output feature map is interpolated using k-linear interpolation, where k is an integer greater than one.
[0152] In step 850, a plurality of multi-view feature maps 330 at different scales are generated. The different scales correspond to different resolutions of the plurality of multi-view feature maps 330. Thus, the different scales are defined as consisting of one highest resolution and one or more lower resolutions. In the plurality of multi-view feature maps 330, the corresponding encoder output feature maps (i.e., ) are grouped together, and the first multi-view feature map generated with the highest resolution is obtained (i.e. ). For a single reduced resolution, the first multi-view feature map is downsampled to obtain a corresponding multi-view feature map generated at the single reduced resolution (i.e. u∈{2,…,N}).
[0153] In step 860, the multiple multi-view feature maps 330 at different scales obtained in step 850 are each back-projected into corresponding 3D spaces voxelized at different scales to form multiple multi-scale 3D volume representations 340. As described above, because scale corresponds to resolution, the 3D space voxelized at the scale refers to a 3D space voxelized at the corresponding resolution. Equation 6 can be used to generate each of the multiple multi-scale 3D volume representations 340.
[0154] In step 870 , a plurality of multi-scale voxel-aligned features 350 for a single point are queried from the plurality of multi-scale 3D volume representations 340 .
[0155] By employing Equation 7, multiple multi-scale 3D volume representations 340 can be used to query multiple multi-scale voxel-aligned features 350 for a single point. Figure 9 The flowchart using Formula 7 when implementing step 870 is shown. Step 870 may include steps 910, 920, and 930.
[0156] In step 910, the plurality of multi-scale 3D volume representations 340 are interpolated to generate a plurality of scale-specific voxel alignment features for a single point (ie, s=1, ..., S). In some embodiments, each of the plurality of multi-scale 3D volume representations 340 is interpolated using k-linear interpolation, where k is an integer greater than one.
[0157] In step 920, the multiple scale-specific voxel alignment features are stitched together to generate a stitched voxel alignment feature for a single point (i.e., ], as shown in Formula 7).
[0158] In step 930, the concatenated voxel-aligned features are aggregated to generate a plurality of multi-scale voxel-aligned features 350, such that the channel sizes of the plurality of multi-scale voxel-aligned features 350 are consistent with the channel sizes of the multi-view pixel-aligned features 320. In some embodiments of step 930, an MLP is used to map the channel sizes of the plurality of multi-scale voxel-aligned features 350 to be consistent with the channel sizes of the multi-view pixel-aligned features 320.
[0159] Reference Figure 8In step 880, the plurality of multi-view pixel-aligned features 320 and the plurality of multi-scale voxel-aligned features 350 are aggregated to estimate the attenuation coefficient 390 of the single point. In particular, the aggregation is performed based on the scale-view cross-attention 360 to advantageously utilize cross-region and cross-view feature learning to enhance the representation of the single point before estimating the attenuation coefficient 390.
[0160] The scale-view cross-attention 360 can be implemented by adopting Equation 9. Figure 10 A flowchart illustrating exemplary steps for implementing step 880 by employing Equation 9 is shown. In some embodiments, step 880 includes steps 1010, 1020, and 1030. In step 1010, self-attention is applied to the plurality of multi-view pixel-aligned features 320 to perform cross-view attention across the plurality of multi-view pixel-aligned features 320, thereby generating a plurality of attention-weighted pixel-aligned features (i.e., In step 1020, cross attention is applied between the plurality of multi-scale voxel-aligned features and the plurality of attention-weighted pixel-aligned features to thereby generate a plurality of cross-region cross-view features for a single point (i.e. In step 1030 , the attenuation coefficient 390 is estimated from the plurality of cross-region and cross-view features by processing the plurality of cross-region and cross-view features using, for example, a linear layer 430 .
[0161] Step 880 may be performed by the learnable aggregation and estimation model 400, which Figure 4 In some embodiments, the disclosed method further includes aggregating the plurality of multi-view pixel-aligned features 320 and the plurality of multi-scale voxel-aligned features 350 using a learnable aggregation and estimation model 400 to estimate the attenuation coefficient 390. The learnable aggregation and estimation model 400 includes a plurality of SVC-Att modules 410 stacked together for applying self-attention to the plurality of multi-view pixel-aligned features 320 and for applying cross-attention between the plurality of multi-scale voxel-aligned features 350 and the plurality of attention-weighted pixel-aligned features. The plurality of SVC-Att modules 410 output a plurality of cross-region and cross-view features for a single point. The learnable aggregation and estimation model 400 further includes a linear layer 430 following the plurality of SVC-Att modules 410. The linear layer 430 is used to estimate the attenuation coefficient 390 from the cross-region and cross-view features.
[0162] Note that prior to performing steps 830 and 880, the learnable encoder-decoder model 310 and the learnable aggregation and estimation model 400 need to be trained, respectively. In one alternative, each of these learnable models 310, 400 is a pre-trained model loaded with predetermined model parameters. In another alternative, the disclosed method further includes a training step 810. This step 810 may include training the learnable encoder-decoder model before using it in step 830. Step 810 may also include training the learnable aggregation and estimation model 400 before using it in step 880.
[0163] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The present embodiments should therefore be considered in all respects as illustrative and not restrictive. The scope of the present invention is indicated by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.
[0164] References
[0165] The following lists references that are occasionally cited in the specification. The disclosures of each of these references are incorporated herein by reference in their entirety.
[0166] [1] Anders H Andersen and Avinash C Kak, “Synchronous Algebraic Reconstruction Technique (SART): An Advanced Implementation of the Algorithm in the Field,” Ultrasonic Imaging, Vol. 6, No. (1): 81–94, 1984.
[0167] [2] Hyungjin Chung, Dohoon Ryu, Michael T McCann, Marc L Klasky, and JongChul Ye, “Using pretrained 2d diffusion models to solve 3d inverse problems,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 22542–22551, 2023.
[0168] [3] Yu Fang, Lanzhuju Mei, Changjian Li, Yuan Liu, Wenping Wang, Zhiming Cui, and Dinggang Shen, “SNAF: Sparse-view CBCT reconstruction using neural attenuation fields,” arXiv preprint arXiv:2211.17048, 2022.
[0169] [4]Lee A Feldkamp, Lloyd C Davis, and James W Kress, “Practical Cone-Beam Algorithms,” in J.O.S.A., vol. 1, no. (6): 612–619, 1984.
[0170] [5] Richard Gordon, Robert Bender, and Gabor T Herman, “Algebraic reconstruction techniques for three-dimensional electron microscopy and X-ray photography,” Journal of Theoretical Biology, vol. 29, no. (3): pp. 471–481, 1970.
[0171] [6] Yo Seob Han, Jaejun Yoo, and Jong Chul Ye, “Deep residual learning for compressed sensing CT reconstruction via persistent homology analysis,” arXiv preprint arXiv:1611.06391, 2016.
[0172] [7] Ji He, Yongbo Wang, and Jianhua Ma, “Inverse Radon transform via deep learning,” IEEE Transactions on Medical Imaging, vol. 39, no. (6): pp. 2076–2087, 2020.
[0173] [8]Johannes Hofmanninger, Forian Prayer, Jeanny Pan, Sebastian Helmut Prosch and Georg Langs, “Automatic lung segmentation in routine imaging is primarily a data diversity problem, not a methodological one,” European Radiology Experimental, vol. 4, no. (1): pp. 1–13, 2020.
[0174] [9] Gao Huang, Zhuang Liu, Laurens Van DerMaaten, and Kilian Q Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700–4708, 2017.
[0175]
[10] Xia Huang, Jian Wang, Fan Tang, Tao Zhong, and Yu Zhang, “Reducing metal artifacts on cervical spine CT images via deep residual learning,” Biomedical Engineering Online, vol. 17: 1–15, 2018.
[0176]
[11] Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li, and Tony Tung, “Arch: Animated reconstruction of clothed human figures,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 3093–3102, 2020.
[0177]
[12] Yixiang Jiang, “Mfct-gan: Multi-information network reconstruction of ct volume for safety inspection”, Journal of Intelligent Manufacturing and Special Equipment, 2022.
[0178]
[13] Kyong Hwan Jin, Michael T McCann, Emmanuel Froustey, and Michael Unser, “Deep Convolutional Neural Networks for Inverse Problems in Imaging,” IEEE Transactions on Image Processing, vol. 26, no. (9): pp. 4509–4522, 2017.
[0179]
[14] Daeun Kyung, Kyungmin Jo, Jaegul Choo, Joonseok Lee, and Edward Choi, “Perspective projection-based 3D CT reconstruction from biplane X-rays,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5, IEEE, 2023.
[0180]
[15] Anish Lahiri, Marc Klasky, Jeffrey A Fessler, and Saiprasad Ravishankar, “Sparse view cone-beam CT reconstruction using data-consistent supervision and adversarial learning from scarce training data,” ArXiv preprint arXiv:2201.09318, 2022.
[0181]
[16] Bo Li, Kaitao Xue, Bin Liu, and Yu-Kun Lai, “BBDM: Image-to-image translation using the Brownian Bridge Diffusion Model,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 1952–1961, 2023.
[0182]
[17] Wei-An Lin, Haofu Liao, Cheng Peng, Xiaohang Sun, Jingdan Zhang, Jiebo Luo, Rama Chellappa, and Shaohua Kevin Zhou, “Dudonet: A dual-domain network for metal artifact reduction in CT,” in Proceedings of the EEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 10512–10521, 2019.
[0183]
[18] Yiqun Lin, Zhongjin Luo, Wei Zhao, and Xiaomeng Li, “Deep intensity field learning for extremely sparse view cbct reconstruction,” in Medical Image Computing and Computer-Assisted Intervention - MICCAI 2023, pp. 13–23, Cham, 2023, Springer Nature, Switzerland.
[0184]
[19] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black, “Smpl: Skinned Multi-Person Linear Models,” ACM Transactions on Graphics, vol. 34, no. (6), 2015.
[0185]
[20] Chenglong Ma, Zilong Li, Junping Zhang, Yi Zhang, and Hongming Shan, “Freeseed: a band-aware and self-steering network for sparse view CT reconstruction,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 250–259, Springer-Verlag, 2023.
[0186]
[21] Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM, vol. 65, no. (1): 99–106, 2021.
[0187]
[22] Jinxiao Pan, Tie Zhou, Yan Han, and MingJ iang, “Variable weighted ordered subset image reconstruction algorithm,” International Journal of Biomedical Imaging, 2006, pp.
[0188]
[23] Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove, “Deepsdf: Learning a continuous signed distance function for shape representation,” in Proceedings of the EEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 165–174, 2019.
[0189]
[24] Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black, “Expressive human capture: Capturing 3d hands, faces, and bodies from a single image,” in Proceedings of the EEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 10975–10985, 2019.
[0190]
[25] Olaf Ronneberger, Philipp Fischer, and Thomas Brox, “Unet: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015: 18th International Conference, Munich, Germany, October 5–9, 2015, Proceedings, Part III 18, pp. 234–241, Springer-Verlag, 2015.
[0191]
[26] Darius Rückert, Yuanhao Wang, Rui Li, Ramzi Idoughi, and Wolfgang Heidrich, “Neat: Neural adaptive tomography,” ACM Transactions on Graphics (TOG), vol. 41, no. (4): pp. 1–13, 2022.
[0192]
[27] Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li, “Pifu: Pixel-aligned implicit functions for high-resolution digitization of clothed humans,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 2304–2314, 2019.
[0193]
[28] William C Scarfe, Allan G Farman, Predag Sukovic, et al., “Clinical applications of cone-beam computed tomography in dental practice,” Journal of the Canadian Dental Association, vol. 72, no. (1): p. 75, 2006.
[0194]
[29] Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas DeBel, Moira SN Berens, Cas Van Den Bogaard, Piergiorgio Cerello, Hao Chen, Qi Dou, Maria Evelina Fantacci, Bram Geurts et al., “Validation, comparison and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images,” in Luna16 Challenge, Medical Image Analysis, 42: 1–13, 2017.
[0195]
[30] Liyue Shen, Wei Zhao, and Lei Xing, “Reconstructing patient-specific volumetric computed tomography images from a single projection view via deep learning,” Nature biomedical engineering, vol. 3, no. (11): pp. 880–888, 2019.
[0196]
[31] Liyue Shen, John Pauly, and Lei Xing, “Nerp: Implicit Neural Representation Learning with Prior Embeddings for Sparsely Sampled Image Reconstruction,” IEEE Transactions on Neural Networks and Learning Systems, 2022.
[0197]
[32] Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon, “Solving the inverse problem in medical imaging with score-based generative models,” arXiv preprint arXiv:2111.08005, 2021.
[0198]
[33] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, LlionJones, Aidan N Gomez, Kaiser and Illia Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.
[0199]
[34] Ce Wang, Kun Shang, Haimiao Zhang, Qian Li, Yuan Hui, and S Kevin Zhou, “Dudotrans: Dual-domain transformer with more attention for sinusoidal recovery in sparse view CT reconstruction,” ArXiv preprint arXiv:2111.10790, 2021.
[0200]
[35] Jianing Wang, Yiyuan Zhao, Jack H Noble, and Benoit M Dawant, “Conditional generative adversarial networks for metal artifact reduction in ear CT images,” in Medical Imaging Computing and Computer-Assisted Intervention - MICCAI 2018: 21st International Conference on, Granada, Spain, September 16–20, 2018.
[0201]
[36] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. (4): pp. 600–612, 2004.
[0202]
[37] Weiwen Wu, Dianlin Hu, Chuang Niu, Hengyong Yu, Varut Vardhanabhuti, and Ge Wang, “Drone: A residual-based dual-domain optimization network for sparse view CT reconstruction,” IEEE Transactions on Medical Imaging, vol. 40, no. (11): pp. 3002–3014, 2021.
[0203]
[38] Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J Black, “Icon: Implicitly Clothed Humans from Normal People,” in 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13286–13296. IEEE, 2022.
[0204]
[39] Yuliang Xiu, Jinlong Yang, Xu Cao, Dimitrios Tzionas, and Michael J. Black, “Econ: Explicitly Dressed Humans via Normal Integration Optimization,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 512–523, 2023.
[0205]
[40] Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, KalyanSunkavalli, and Ulrich Neumann, “Point-nerf: Point-based neural radiance fields,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 5438–5448, 2022.
[0206]
[41] Fukun Yin, Wen Liu, Zilong Huang, Pei Cheng, Tao Chen, and Gang Yu, “Coordinates are not alone - codebook priors help implicit neural 3d representations,” Advances in Neural Information Processing Systems, 35: 12705–12717, 2022.
[0207]
[42] Xingde Ying, Heng Guo, Kai Ma, Jian Wu, Zhengxin Weng, and Yefeng Zheng, “X2ct-gan: CT reconstruction from dual-plane x-rays using generative adversarial networks,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 10619–10628, 2019.
[0208]
[43] Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa, “Pixelnerf: Neural Radiance Fields from One or Few Images,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 4578–4587, 2021.
[0209]
[44] Ruyi Zha, Yanhao Zhang, and Hongdong Li, “Naf: Neural Attenuation Field for Sparse View CBCT Reconstruction,” in Medical Imaging Computing and Computer-Assisted Intervention - MICCAI 2022: 25th International Conference, Singapore, September 18-22, 2022, Proceedings, Part VI, pp. 442-452, Springer-Verlag, 2022.
[0210]
[45] Zhicheng Zhang, Xiaokun Liang, Xu Dong, Yaoqin Xie, and Guohua Cao, “Sparse view CT reconstruction method based on densenet and deconvolution,” IEEE Transactions on Medical Imaging, vol. 37, no. (6): pp. 1407–1417, 2018.
[0211]
[46] Zerong Zheng, Tao Yu, Yebin Liu, and Qionghai Dai, “Pamir: Parametric model-conditional implicit representation for image-based human body reconstruction,” IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. (6): pp. 3170–3184, 2021.
Claims
1. A computer-implemented method for reconstructing a three-dimensional (3D) computed tomography (CT) volume from a plurality of projection views generated in cone-beam computed tomography (CBCT) imaging, the method comprising: determining a plurality of points in a three-dimensional CT volume such that the three-dimensional CT volume is reconstructed by estimating attenuation coefficients at the individual points; processing a single projected view using a learnable encoder-decoder model to thereby generate a decoder output feature map and an encoder output feature map for the single projected view, wherein the learnable encoder-decoder model is shared by the multiple projected views when processing the single projected view; querying a plurality of multi-view pixel-aligned features for the single point using corresponding decoder output feature maps generated for the plurality of projected views; generating a plurality of multi-view feature maps at different scales, the different scales comprising a highest resolution and one or more reduced resolutions, wherein a first multi-view feature map generated at the highest resolution is obtained by grouping together corresponding encoder output feature maps generated for the plurality of projected views, and wherein corresponding multi-view feature maps generated at each reduced resolution are obtained by downsampling the first multi-view feature map; Back-projecting the plurality of multi-view feature maps at the different scales into corresponding three-dimensional spaces voxelized according to the different scales to thereby form a plurality of multi-scale three-dimensional volume representations; querying a plurality of multi-scale voxel-aligned features for the single point using the plurality of multi-scale three-dimensional volume representations; and The multiple multi-view pixel-aligned features and the multiple multi-scale voxel-aligned features are aggregated according to scale-view cross-attention to estimate an attenuation coefficient of the single point, so as to enhance the representation of the single point using cross-region and cross-view feature learning before estimating the attenuation coefficient.
2. The method according to claim 1, wherein The view-specific pixel alignment features for the single point under a single projection view are queried using the decoder output feature map, and the multiple multi-view pixel alignment features for the single point are obtained from the corresponding decoder output feature map, wherein the corresponding view-specific pixel alignment features generated for the multiple projection views are regarded as the multiple multi-view pixel alignment features for the single point.
3. The method according to claim 2, wherein: The view-specific pixel-aligned features for the single point under the single projection view are obtained by interpolating the decoder output feature map.
4. The method according to claim 3, wherein: The decoder output feature map is interpolated using k linear interpolation, where k is an integer greater than one.
5. The method according to claim 1, wherein Querying the plurality of multi-scale voxel-aligned features for the single point using the plurality of multi-scale three-dimensional volume representations comprises: interpolating the plurality of multi-scale three-dimensional volume representations respectively to generate a plurality of scale-specific voxel-aligned features for the single point; stitching the plurality of scale-specific voxel-aligned features to generate a stitched voxel-aligned feature for the single point; and The concatenated voxel-aligned features are aggregated to generate the plurality of multi-scale voxel-aligned features, such that channel sizes of the plurality of multi-scale voxel-aligned features are consistent with the channel size of the multi-view pixel-aligned features.
6. The method according to claim 5, wherein: Each of the plurality of multi-scale three-dimensional volume representations is interpolated using k-linear interpolation, where k is an integer greater than one.
7. The method according to claim 5, wherein: When aggregating the concatenated voxel-aligned features to generate the plurality of multi-scale voxel-aligned features, a multi-layer perceptron (MLP) is used to map channel sizes of the plurality of multi-scale voxel-aligned features to be consistent with the channel size of the multi-view pixel-aligned features.
8. The method according to claim 1, wherein Aggregating the plurality of multi-view pixel-aligned features and the plurality of multi-scale voxel-aligned features according to scale-view cross-attention to generate an attenuation coefficient of the single point includes: applying self-attention to the plurality of multi-view pixel-aligned features to perform cross-view attention across the plurality of multi-view pixel-aligned features, thereby generating a plurality of attention-weighted pixel-aligned features; applying cross-attention between the plurality of multi-scale voxel-aligned features and the plurality of attention-weighted pixel-aligned features to thereby generate a plurality of cross-region and cross-view features for the single point; and The attenuation coefficient is estimated from the cross-region and cross-view features.
9. The method according to claim 8, wherein The attenuation coefficient is estimated from the cross-region cross-view features by using a linear layer to process the cross-region cross-view features.
10. The method of claim 8, further comprising aggregating the plurality of multi-view pixel-aligned features and the plurality of multi-scale voxel-aligned features using a learnable aggregation and estimation model to estimate the attenuation coefficient, wherein the learnable aggregation and estimation model comprises: a plurality of scale-view cross-attention SVC-Att modules stacked together, the plurality of SVC-Att modules being configured to apply self-attention to the plurality of multi-view pixel-aligned features and to apply cross-attention between the plurality of multi-scale voxel-aligned features and the plurality of attention-weighted pixel-aligned features, wherein the plurality of SVC-Att modules output the plurality of cross-region cross-view features for the single point; as well as A linear layer is located after the multiple SVC-Att modules, and is used to estimate the attenuation coefficient from the cross-region and cross-view features.
11. The method according to claim 1, wherein The learnable encoder-decoder model is implemented in the form of U-Net.
12. The method of claim 1, further comprising training the learnable encoder-decoder model before using the learnable encoder-decoder model to process the single projected view.
13. The method according to claim 10 further includes training the learnable aggregation and estimation model before aggregating the multiple multi-view pixel-aligned features and the multiple multi-scale voxel-aligned features according to scale-view cross-attention to estimate the attenuation coefficient.