Deep-learning-based noise reduction for ultra-low-dose computed tomography
By training a machine learning model with realistic noise maps and cleaner target images, the method effectively addresses the issue of hallucinated structures in ultra-low-dose CT images, improving image quality and diagnostic performance.
Patent Information
- Application Number
- PCT/US2025/051875
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-21
- Filing Date
- 2025-10-21
- Publication Date
- 2026-04-30
AI Technical Summary
Conventional deep learning reconstruction methods for ultra-low-dose computed tomography (CT) images tend to hallucinate or generate artificial structures due to the limited quality of training data, particularly when dealing with high noise levels.
A method is employed to train a machine learning model using routine-dose data to generate low-dose CT images by reconstructing thin-slice and thick-slice images, inserting noise maps, and assembling training data pairs to reduce noise in ultra-low-dose CT images, utilizing realistic noise maps and cleaner target images for improved denoising.
The method achieves higher image quality and reduces hallucinated structures in ultra-low-dose CT images, outperforming conventional methods by enhancing spatial resolution and diagnostic confidence.
Smart Images

Figure US2025051875_30042026_PF_FP_ABST
Abstract
Description
DEEP-LEARNING-BASED NOISE REDUCTION FOR ULTRA-LOW-DOSE COMPUTED TOMOGRAPHY CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 709,865, filed on October 21, 2024, and entitled “DEEP-LEARNING-BASED NOISE REDUCTION FOR ULTRA-LOW-DOSE COMPUTED TOMOGRAPHY ’ which is herein incorporated by reference in its entirety.BACKGROUND
[0002] Deep learning reconstruction and noise reduction (DLR) have been actively used in computed tomography (CT) for improving image quality and diagnostic performance. Various training strategies have been proposed to reduce image noise, either applied on projection data, during image reconstruction, or directly on images after reconstruction. Some of these methods have become commercially available either on CT scanners or as third-party products.
[0003] In general, the performance of these methods is limited by the quality of routinedose and low-dose data, either paired or un-paired, for training. When the input images contain high noise (e.g.. acquired at ultra-low dose), existing DLR methods tend to hallucinate or generate faked structures.SUMMARY OF THE DISCLOSURE
[0004] It is an aspect of the present disclosure to provide a method for training a machine learning model to denoise computed tomography (CT) image data. The method includes accessing routine-dose data with a computer system, where the routine-dose data have been acquired from a plurality of subjects using one or more CT systems. Low-dose CT image data are generated from the routine-dose data with the computer system by: reconstructing thin-slice images from the routine-dose data, where the thin-slice images have a first slice thickness; estimating noise maps from the thin-slice images; reconstructing thick-slice images from the routine-dose data, where the thick-slice images have a second slice thickness that is thicker than the first slice thickness; and generating low-dose thick-slice images by inserting noise from the noise maps to the thick-slice images, where the low-dose thick-slice images comprise the low-dose CT image data. Routine-dose CT image data are generated from the routine-dosedata with the computer system by reconstructing routine-dose thick-slice images from the routine-dose data, where the routine-dose thick-slice images have the same slice thickness as the low-dose thick-slice images. Training data are assembled with the computer system as pairs of low-dose CT image data as training inputs and routine-dose CT image data as training targets. A machine learning model is then trained on the training data using the computer system and the trained machine learning model is stored using the computer system.
[0005] It is another aspect of the present disclosure to provide a method for generating denoised CT image data. Ultra-low-dose (ULD) CT image data are accessed with a computer system, where the ULD CT image data were acquired from a subject using a CT system at a dose level less than 1.8 mGy. A machine learning model is also accessed with the computer system. The machine learning model has been trained on training data to denoise CT image data, where the training data include low-dose CT image data as training inputs and routinedose CT image data as targets. The low-dose CT image data include CT images having noise inserted from noise maps having a thinner slice thickness than the CT images. Denoised CT image data are then generated with the computer system by inputting the CT image data to the machine learning model, generating the denoised CT image data as an output. The denoised CT image data are then output with the computer system.
[0006] It is yet another aspect of the present disclosure to provide a method for processing CT image data using a trained machine learning model. The method includes accessing CT image data with a computer system, where the CT image data include at least one of projection data or reconstructed images. A trained machine learning model is also accessed with the computer system. The trained machine learning model has been trained using training data including synthesized low-dose CT image data as training inputs and routine-dose CT image data as training targets. The synthesized low-dose CT image data were generated by inserting noise maps derived from thin-slice images into thick-slice images. The CT image data are processed with the trained machine learning model to generate processed CT image data having reduced noise compared to the CT image data. The processed CT image data are then outputted with the computer system.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 is a flowchart of an example method for denoising computed tomography (CT) image data, including ultra-low-dose CT image data.
[0008] FIG. 2 is a flowchart of an example method for training a neural network, or other suitable machine learning model, to denoise CT image data.
[0009] FIG. 3 is an overview of an example training scheme for a ULD-NET method for ultra-low-dose CT image data noise reduction.
[0010] FIG. 4 shows example CT images of routine dose and quarter dose (i.e., 25% dose) IR images, GARNET, and ULD-NET denoised 25% and 100% of routine dose FBP images. Yellow arrows indicate hallucinated structures generated by GARNET but not by ULD-NET. The display window is (window level, window width) = (-400, 1500).
[0011] FIG. 5 shows example CT images of a pediatric cystic fibrosis patient reconstructed using FBP, IR, GARNET, and ULD-NET from PCD-CT with ultra-high-resolution scan mode. Yellow arrows indicate hallucinated structures generated by GARNET but not by ULD-NET. The display window is (window level, window width) = (-600, 1500).
[0012] FIG. 6 shows example CT images of another pediatric cystic fibrosis patient reconstructed using FBP, IR, GARNET, and ULD-NET from PCD-CT with ultra-high-resolution scan mode. Yellow arrows indicate hallucinated structures generated by GARNET but not by ULD-NET. The display window is (window level, window width) = (-600, 1500).
[0013] FIG. 7 is a block diagram of an example ultra-low-dose CT image data noise reduction system according to some embodiments described in the present disclosure.
[0014] FIG. 8 is a block diagram of example components that can implement the system of FIG. 7.
[0015] FIGS. 9A and 9B illustrate an example CT system that can implement the methods described in the present disclosure.DETAILED DESCRIPTION
[0016] Described here are systems and methods for training and implementing a machine learning model (e.g., a neural network or other suitable machine learning model) for noise reduction in ultra-low-dose (ULD) computed tomography (CT). In some examples, the machine learning model may implement a deep convolutional neural network for ultra-low-dose CT denoising (ULD-NET) in CT.
[0017] In general, the disclosed methods generate thin-slice noise maps by subtracting simulated-lower-dose thin-slice images from routine-dose ones and then applying spatial decoupling and scaling. A neural network, or other suitable machine learning model, is trained by mapping thick-slice filtered back projection (FBP) images with added thin-slice noise mapto routine-dose thick-slice iterative reconstruction (IR) images. Inference is then performed on thin-slice FBP reconstructed images, or other suitable CT image data, acquired at ultra-low-dose.
[0018] Advantageously, the disclosed systems and methods can achieve higher image quality compared to conventional deep learning reconstruction methods. By utilizing realistic noise maps and cleaner target images for training, the disclosed systems and methods reduce hallucinated structures when applied to ultra-low dose CT image data with strongly increased image noise.
[0019] Conventional supervised DLR methods are trained on paired routine-dose and low-dose images that have the same reconstruction slice thickness as the images during inference. When these DLR methods are applied to ultra-low dose CT images (i.e., images with much lower dose than the low-dose images used in the training dataset) with strongly increased image noise, they tend to generate artificial structures without appropriate training. However, it is unrealistic to train a DLIR model with paired routine-dose and ultra-low-dose data because severe artifacts will be introduced at the ultra-low-dose level. The performance of these DLIR methods is also limited by the image quality of ground truth used in these methods during training. To overcome these limitations, the disclosed systems and methods implement a training strategy that utilizes realistic noise maps and cleaner target images, which advantageously results in a reduction of hallucinated structures and achieves higher image quality compared to conventional DLR methods.
[0020] Referring now to FIG. 1, a flowchart is illustrated as setting forth the steps of an example method for generating denoised CT image data using a suitably trained neural network or other machine learning model. As will be described, the neural network or other machine learning model takes CT image data as input data and generates denoised CT image data as output data.
[0021] The method includes accessing CT image data with a computer system, as indicated at step 102. Accessing the CT image data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the CT image data may include acquiring such data with a CT system and transferring or otherwise communicating the data to the computer system, which may be a part of the CT system.
[0022] As described above, the CT image data are generally low-dose CT image data, and in some instances may be ultra-low-dose CT image data. Ultra-low-dose CT image data can include CT image data acquired with CTDIvol of 1.8 mGy or less.
[0023] In some instances, the CT image data may include thin-slice images. As an example, thin-slice images can include images acquired with a slice thickness of 1.0 mm or less. For instance, thin-slice images can include images acquired with a slice thickness of 1.0 mm, 0.9 mm, 0.8 mm, 0.7 mm, 0.6 mm, 0.5 mm, 0.4 mm, 0.3 mm, 0.2 mm, 0.1 mm, or the like. Thin-slice images can also include images acquired with slice thicknesses in ranges between those values stated above, such as slice thicknesses in increments of 0.01 mm, 0.025 mm, 0.05 mm, 0.1 mm, 0.2 mm. 0.4 mm, or the like, between those values stated above. As a non-limiting example, thin-slice images can include images with slice thicknesses such as 0.2 mm, 0.4 mm, 0.6 mm, 0.8 mm, and so on. As another example, thin-slice images may have slice thicknesses such as 0.6 mm, 0.625 mm, 0.65 mm, 0.675 mm, and so on.
[0024] A trained neural network (or other suitable machine learning model) is then accessed with the computer system, as indicated at step 104. In general, the neural network is trained, or has been trained, on training data in order to reduce noise in CT image data, such as thin-slice, ultra-low-dose CT image data. This noise reduction is achieved, in part, by the neural network (or other machine learning model) being trained by mapping synthesized noisy, low-dose CT images to routine-dose thick-slice target CT images.
[0025] The trained neural network can include a neural network with any suitable neural network architecture for generating denoised CT image data. As one non-limiting example, the trained neural network may include a residual neural network, such as a residual U-Net model. The trained neural network may in some instances have multiple inputs (e.g., corresponding to multiple slice inputs).
[0026] Accessing the trained neural network may include accessing network parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the neural network on training data. In some instances, retrieving the neural network can also include retrieving, constructing, or otherwise accessing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0027] An artificial neural network generally includes an input layer, one or more hidden layers (or nodes), and an output layer. Typically, the input layer includes as many nodes as inputs provided to the artificial neural network. The number (and the type) of inputs provided to the artificial neural network may vary based on the particular task for the artificial neural network.
[0028] The input layer connects to one or more hidden layers. The number of hidden layers varies and may depend on the particular task for the artificial neural network. Additionally, each hidden layer may have a different number of nodes and may be connected to the next layer differently. For example, each node of the input layer may be connected to each node of the first hidden layer. The connection between each node of the input layer and each node of the first hidden layer may be assigned a weight parameter. Additionally, each node of the neural network may also be assigned a bias value. In some configurations, each node of the first hidden layer may not be connected to each node of the second hidden layer. That is, there may be some nodes of the first hidden layer that are not connected to all of the nodes of the second hidden layer. The connections between the nodes of the first hidden layers and the second hidden layers are each assigned different weight parameters. Each node of the hidden layer is generally associated with an activation function. The activation function defines how the hidden layer is to process the input received from the input layer or from a previous input or hidden layer. These activation functions may vary and be based on the type of task associated with the artificial neural network and also on the specific type of hidden layer implemented.
[0029] Each hidden layer may perform a different function. For example, some hidden layers can be convolutional hidden layers which can, in some instances, reduce the dimensionality of the inputs. Other hidden layers can perform statistical functions such as max pooling, which may reduce a group of inputs to the maximum value; an averaging layer; batch normalization; and other such functions. In some of the hidden layers each node is connected to each node of the next hidden layer, which may be referred to then as dense layers. Some neural networks including more than, for example, three hidden layers may be considered deep neural networks.
[0030] The last hidden layer in the artificial neural network is connected to the output layer. Similar to the input layer, the output layer typically has the same number of nodes as the possible outputs. In an example in which the artificial neural network is trained to reduce noisein CT image data, the output layer may include, for example, a number of different nodes corresponding to pixels in denoised, or otherw ise noise-reduced, CT image data.
[0031] The CT image data are then input to the trained neural network, generating output as denoised CT image data, as indicated at step 106. For example, the denoised CT image data may include one or more images that have been denoised, or otherwise have had noise in the images reduced.
[0032] As another example, the CT image data may include raw projection data, such that the denoised CT image data may include raw projection data that has been denoised, or otherwise had noise in the raw projection data reduced. From the denoised raw projection data, one or more images can then be reconstructed using any suitable reconstruction technique, such as filtered backproj ection, iterative reconstruction, or the like. In other implementations, the trained neural network may be a deep learning reconstruction (DLR) neural network that collectively performs denoising and image reconstruction. In these instances, the CT image data may be raw projection data and the denoised CT image data may include one or more reconstructed images that have lower noise than the input raw projection data.
[0033] The denoised CT image data generated by inputting the CT image data to the trained neural network(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 108.
[0034] Referring now to FIG. 2. a flowchart is illustrated as setting forth the steps of an example method for training one or more neural networks (or other suitable machine learning models) on training data, such that the one or more neural networks are trained to receive CT image data as input data in order to generate denoised CT image data as output data. An example workflow for training a neural network to denoise CT image data is also illustrated in FIG. 3.
[0035] In general, the neural network(s) can implement any number of different neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, a multi-scale neural network, a multi-resolution neural network, or the like. In some embodiments, the neural network may include attention mechanisms configured to focus on relevant image features during the denoising process. Additionally, ensemble methods may be employed that combine multiple trained models to improve denoising performance and robustness. Alternatively, the neural network(s) could be replaced with other suitable machine learning or artificial intelligence algorithms, such as thosebased on supervised learning, unsupervised learning, deep learning, ensemble learning, dimensionality reduction, and so on.
[0036] The method includes accessing training data with a computer system, as indicated at step 202. In general, the training data can include low-dose CT image data, routinedose CT image data, and one or more noise maps. Additionally or alternatively, the accessed training data can include CT image data from which one or more of low-dose CT image data, routine-dose CT image data, or noise maps are generated. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with a CT system and transferring or otherwise communicating the data to the computer system. In some embodiments, the training data may be acquired from multiple CT scanner manufacturers and models to improve the generalizability of the trained neural network across different imaging systems. Additionally, cross-institutional training data may be utilized, where CT image data are collected from multiple healthcare institutions to enhance the robustness and clinical applicability of the trained model.
[0037] The method can include assembling training data from CT image data using a computer system. This step may include assembling the CT image data into an appropriate data structure on which the neural network or other machine learning model can be trained. Assembling the training data may include assembling CT image data, one or more noise maps, and other relevant data. For instance, assembling the training data may include generating noise maps from the CT image data, generating low-dose CT image data from the CT image data, generating routine-dose CT image data from the CT image data, and the like. In some implementations, the training data may include multiple dose level variations to improve the ability of the neural network to handle a wide range of noise conditions encountered in clinical practice.
[0038] The target images for the training data can include routine dose CT images. As one non-limiting example, target images can be obtained by iterative reconstruction (IR) of routine dose projection data with slice thickness / interval of 2.0 / 0.5 mm and a sharp kernel (Qr68). As described below, the training inputs can be synthesized by superimposing noise maps on the target images, where the noise maps are generated by subtracting the reduced dose (e.g., 25% dose (QD)) filtered back projection (FBP) images from the corresponding routine dose FBP images, both with a thin slice thickness (0.8 mm). The QD or other reduced dose images can be reconstructed from QD or other reduced dose projection data simulated byvarious noise insertion methods, including but not limited to projection domain noise insertion, sinogram domain noise insertion, and image domain noise insertion techniques.
[0039] As a non-limiting example, one or more noise maps can be generated from CT image data as follows. The original projection data in the CT image data, which as described above may be routine-dose projection data, can have noise inserted into the data to generate noise-inserted projection data. For instance, noise insertion can be applied on the original projection data of each patient selected in the training dataset to generate corresponding low-dose projection data. The low-dose projection data can correspond to a percentage of the dose used when acquiring the routine-dose projection data. As one non-limiting example, the percentage dose may include one or more of multiple dose levels such as 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% dose, such that the low-dose projection data may represent one or more reduced-dose conditions. Training with multiple dose levels advantageously improves the robustness and ability of the neural network to handle a broader range of noise conditions encountered in clinical practice. Further still, the percentage dose can also include a percentage in ranges between those values stated above, such as percentages in increments of 0.5%, 1%, 2.5%, or the like, between those values stated above. As a non-limiting example, the percentage dose may be 25% dose, such that the low-dose projection data are quarter-dose proj ection data.
[0040] Images are then reconstructed from the original projection data and from the low-dose projection data. As anon-limiting example, the images may be reconstructed using a filtered backprojection reconstruction. The images may be reconstructed with a thin-slice thickness, such as 1.0 mm and below. One or more noise maps are then generated by subtracting the low-dose thin-slice images from the corresponding routine-dose thin-slice images. In some embodiments, different noise insertion techniques may be employed beyond projection domain noise insertion, including sinogram domain noise insertion, image domain noise addition, and hybrid approaches that combine multiple noise insertion methods to create more realistic and diverse training scenarios.
[0041] Spatial decoupling can be performed to generate synthesized, or augmented noise maps. For example, a random-translation can be applied in the axial plane with the range of -15 to +15 pixels to each noise map, which efficiently generates a large number of noise insertion examples. The spatial decoupled noise maps can additionally or alternatively be scaled to different noise levels. Additional augmentation techniques may include randomrotations, random flipping, and other geometric transformations to further increase the diversity of the training data and improve the generalization capability of the trained neural network.
[0042] As a non-limiting example, low-dose CT image data can be synthesized, or otherwise generated, from CT image data as follows. Images can be reconstructed from the original routine-dose projection data with a thick-slice thickness, but with the same slice interval as the images reconstructed when generating the noise map(s) described above. As a non-limiting example, these low-dose thick-slice images can be reconstructed using a filtered backproj ection with a slice thickness / interval of 2.0 mm / 0.5 mm.
[0043] Then, the noise maps (e.g., noise maps, augmented noise maps, scaled noise maps) and the corresponding low-dose thick-slice images can be cumulatively split into a large number of patches as multiple-slice training inputs (e.g., 128x128x7 pixels). Noise map patches can be inserted into corresponding thick-slice image patches as illustrated in FIG. 3. The low-dose training input patches can be stored as the low-dose CT image data for training.
[0044] As a non-limiting example, routine-dose CT image data can be synthesized, or otherwise generated, from CT image data as follows. Target images can be obtained by reconstructing images from the routine-dose projection data. As one non-limiting example, the target images can be reconstructed using an iterative reconstruction of the original routine-dose projection data with the same thick-slice thickness and interval as used to reconstruct the low-dose thick-slice images above. Routine-dose target patches matched with the low-dose training input patches at the same anatomy location can then be generated and stored as the routinedose CT image data for training. In some embodiments, multi-scale or multi-resolution approaches may be employed where patches of different sizes (e.g., 64x64, 128x128, 256x256 pixels) are generated to train neural networks capable of processing image features at multiple scales simultaneously.
[0045] One or more neural networks (or other suitable machine learning models) are trained on the training data, as indicated at step 204. In general, the neural network can be trained by optimizing network parameters (e.g., weights, biases, or both) based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function. In some embodiments, the neural network may implement a residual U-Net architecture with specific layer configurations including encoder blocks with convolutional layers, batch normalization layers, and activation functions, decoder blocks with transposed convolutional layers for upsampling, and skip connections between corresponding encoder and decoder layers. The residual U-Net may include multiple resolution levels (e.g., 4-6 levels)with feature map channels ranging from 64 to 512 channels depending on the resolution level. Attention mechanisms may be incorporated at various levels of the network to focus on relevant image features and improve denoising performance, particularly in regions with fine anatomical structures.
[0046] Training a neural network may include initializing the neural network, such as by computing, estimating, or otherwise selecting initial network parameters (e.g., weights, biases, or both). During training, an artificial neural network receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding weights. For instance, training data can be input to the initialized neural network, generating output as denoised or otherwise noise-reduced CT image data. The artificial neural network then compares the generated output with the actual output of the training example in order to evaluate the quality of the denoised CT image data. For instance, the denoised CT image data can be passed to a loss function to compute an error. The current neural network can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current neural network can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the loss according to the loss function. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validation iterations being completed, and the like. When the training condition has been met (e.g., by determining whether an error threshold or other stopping criterion has been satisfied), the current neural network and its associated network parameters represent the trained neural network. Different t pes of training processes can be used to adjust the bias values and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent, Newton's method, conjugate gradient, quasi -New ton. Levenberg-Marquardt, among others. In some embodiments, ensemble methods may be employed where multiple neural networks are trained with different initializations, architectures, or training data subsets, and their outputs are combined (e.g., through averaging, weighted averaging, or learned combination strategies) to improve overall denoising performance and reduce the likelihood of hallucinated structures.
[0047] The artificial neural netw ork can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transferlearning, or other suitable learning techniques for neural networks. As an example, supervised learning involves presenting a computer system with example inputs and their actual outputs (e.g., categorizations). In these instances, the artificial neural network is configured to leam a general rule or model that maps the inputs to the outputs based on the provided example inputoutput pairs. Multi-scale or multi-resolution network architectures may be employed where the neural network processes input images at multiple scales simultaneously, allowing the network to capture both fine-grained details and broader contextual information for improved denoising performance.
[0048] The one or more trained neural networks are then stored for later use, as indicated at step 206. Storing the neural network(s) may include storing network parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the neural network(s) on the training data. Storing the trained neural network(s) may also include storing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored. In embodiments utilizing ensemble methods, multiple trained neural networks and their associated combination strategies may be stored for later deployment and inference.
[0049] The performance of ULD-NET was evaluated in an example study including pediatric thoracic CT cases scanned on a photon counting detector (PCD) CT scanner at ULD (CTDIvol < 0.1 mGy). Three conditions were evaluated: (1) iterative reconstruction (IR), (2) GARNET, (3) and ULD-NET in terms of spatial resolution, noise, hallucination in lung parenchyma (1-5, l=widespread fake structures, 5=no fake structures), and overall diagnostic confidence (1-4, 4=high confidence). ULD-NET received the highest ratings in spatial resolution / noise / diagnostic confidence (3.81±0.37 / 3.75±0.46 / 3.75±0.46) compared to GARNET (3.12±0.69 / 2.38±0.35 / 3.13±0.44) and IR (2.63±0.52 / 2.19±0.59 / 3.06±0.62). In terms of hallucination, no significant difference was found between ULD-NET and IR (4.25±0.71 vs. 4.25±1.04, p=l), both higher than GARNET (2.12±0.83, p<0.01), demonstrating a significant reduction of hallucinated structures by ULD-NET.
[0050] FIG. 4 presents example chest CT images with the six dose / reconstruction conditions: (1) Routine dose / IR, (2) 25% dose / IR, (3) 25% dose / GARNET, (4) 25% dose / ULD-NET, (5) routine dose / GARNET, (6) routine dose / ULD-NET. By comparison with the routine dose IR images (ground truth), the previously trained and clinically validated GARNET introduced hallucinated structures in the denoised 25% dose data, which wasalleviated in the denoised routine-dose data. The proposed ULD-NET method did not introduce hallucinations in both denoised 25% and routine-dose data.
[0051] FIG. 5 compares images reconstructed by FBP, IR, an existing model (GARNET) previously trained and clinically validated, and the proposed ULD-NET from a pediatric patient case scanned by PCD-CT with ultra-high-resolution scan mode. All images were reconstructed with a sharp kernel of Qr68. slice thickness / interval of 0.8 / 0.5 mm. As indicated by the arrows, GARNET introduced hallucinations which were not found in the proposed ULD-NET.
[0052] FIG. 6 compares images with the same reconstruction conditions as FIG. 5 for another pediatric patient case scanned by PCD-CT with ultra-high-resolution scan mode. ULD-NET shows images with superior anti-hallucination in comparison with GARNET.
[0053] FIG. 7 shows an example of a system 700 for denoising CT image data in accordance with some embodiments described in the present disclosure. As shown in FIG. 7, a computing device 750 can receive one or more types of data (e.g., CT image data) from data source 702. In some embodiments, computing device 750 can execute at least a portion of an ultra-low-dose CT image datanoise reduction system 704 to denoise CT image data (e.g., ultra-low-dose CT image data) from data received from the data source 702.
[0054] Additionally or alternatively, in some embodiments, the computing device 750 can communicate information about data received from the data source 702 to a server 752 over a communication network 754, which can execute at least a portion of the ultra-low-dose CT image data noise reduction system 704. In such embodiments, the server 752 can return information to the computing device 750 (and / or any other suitable computing device) indicative of an output of the ultra-low-dose CT image data noise reduction system 704.
[0055] In some embodiments, computing device 750 and / or server 752 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a w earable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 750 and / or server 752 can also reconstruct images from the data.
[0056] In some embodiments, data source 702 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data), such as a CT system, another computing device (e.g., a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some embodiments, data source 702 can be local to computing device 750. For example, data source702 can be incorporated with computing device 750 (e.g.. computing device 750 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 702 can be connected to computing device 750 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 702 can be located locally and / or remotely from computing device 750, and can communicate data to computing device 750 (and / or server 752) via a communication network (e.g., communication network 754).
[0057] In some embodiments, communication network 754 can be any suitable communication network or combination of communication networks. For example, communication network 754 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication network 754 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links show n in FIG. 7 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links. Bluetooth links, cellular links, and so on.
[0058] Referring now to FIG. 8. an example of hardware 800 that can be used to implement data source 702, computing device 750, and server 752 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.
[0059] As shown in FIG. 8, in some embodiments, computing device 750 can include a processor 802, a display 804, one or more inputs 806, one or more communication systems 808, and / or memory 810. In some embodiments, processor 802 can be any suitable hardware processor or combination of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and so on. In some embodiments, display 804 can include any suitable display devices, such as a liquid crystal display (LCD) screen, a light-emitting diode (LED) display, an organic LED (OLED) display, an electrophoretic display (e.g., an '‘e-ink” display), a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 806 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0060] In some embodiments, communications systems 808 can include any suitable hardware, firmware, and / or software for communicating information over communication network 754 and / or any other suitable communication networks. For example, communications systems 808 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 808 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0061] In some embodiments, memory 810 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 802 to present content using display 804, to communicate with server 752 via communications system(s) 808, and so on. Memory 810 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 810 can include random-access memory (RAM), read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), other forms of volatile memory, other forms of non-volatile memory, one or more forms of semivolatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 810 can have encoded thereon, or otherw ise stored therein, a computer program for controlling operation of computing device 750. In such embodiments, processor 802 can execute at least a portion of the computer program to present content (e.g.. images, user interfaces, graphics, tables), receive content from server 752, transmit information to server 752, and so on. For example, the processor 802 and the memory 810 can be configured to perform the methods described herein (e.g., the method of FIG. 1, the method of FIG. 2).
[0062] In some embodiments, server 752 can include a processor 812. a display 814, one or more inputs 816, one or more communications systems 818, and / or memory 820. In some embodiments, processor 812 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 814 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 816 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0063] In some embodiments, communications systems 818 can include any suitable hardware, firmware, and / or software for communicating information over communicationnetwork 754 and / or any other suitable communication networks. For example, communications systems 818 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 818 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0064] In some embodiments, memory 820 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 812 to present content using display 814, to communicate with one or more computing devices 750, and so on. Memory 820 can include any suitable volatile memory, non-volatile memory’, storage, or any suitable combination thereof. For example, memory 820 can include RAM, ROM, EPROM, EEPROM, other ty pes of volatile memory, other ty pes of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 820 can have encoded thereon a server program for controlling operation of server 752. In such embodiments, processor 812 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 750, receive information and / or content from one or more computing devices 750, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.
[0065] In some embodiments, the server 752 is configured to perform the methods described in the present disclosure. For example, the processor 812 and memory 820 can be configured to perform the methods described herein (e.g., the method of FIG. 1, the method of FIG. 2).
[0066] In some embodiments, data source 702 can include a processor 822, one or more data acquisition systems 824, one or more communications systems 826, and / or memory7828. In some embodiments, processor 822 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 824 are generally configured to acquire data, images, or both, and can include a CT system. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 824 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of a CT system. In some embodiments, one or more portions of the data acquisition system(s) 824 can be removable and / or replaceable.
[0067] Note that, although not shown, data source 702 can include any suitable inputs and / or outputs. For example, data source 702 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 702 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.
[0068] In some embodiments, communications systems 826 can include any suitable hardware, firmware, and / or software for communicating information to computing device 750 (and, in some embodiments, over communication network 754 and / or any other suitable communication networks). For example, communications systems 826 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 826 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc ), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0069] In some embodiments, memory 828 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 822 to control the one or more data acquisition systems 824, and / or receive data from the one or more data acquisition systems 824; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 750; and so on. Memory 828 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 828 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 828 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 702. In such embodiments, processor 822 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 750, receive information and / or content from one or more computing devices 750, receive instructions from one or more devices (e.g., a personal computer, alaptop computer, a tablet computer, a smartphone, etc.), and so on.
[0070] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory or non-transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM. EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer-readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0071] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).
[0072] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.
[0073] Referring particularly now to FIGS. 9A and 9B, an example of an x-ray CT imaging system 900 is illustrated. The CT system includes a gantry 902, to which at least one x-ray source 904 is coupled. The x-ray source 904 projects an x-ray beam 906, which may be a fan-beam or cone-beam of x-rays, towards a detector array 908 on the opposite side of the gantry 902. The detector array 908 includes a number of x-ray detector elements 910. Together, the x-ray detector elements 910 sense the projected x-rays 906 that pass through a subject 912, such as a medical patient or an object undergoing examination, that is positioned in the CT system 900. Each x-ray detector element 910 produces an electrical signal that may represent the intensity7of an impinging x-ray beam and, hence, the attenuation of the beam as it passes through the subject 912. In some configurations, each x-ray detector 910 is capable of counting the number of x-ray photons that impinge upon the detector 910. During a scan to acquire x-ray projection data, the gantry 902 and the components mounted thereon rotate about a center of rotation 914 located within the CT system 900.
[0074] The CT system 900 also includes an operator workstation 916, which ty pically includes a display 918; one or more input devices 920. such as a keyboard and mouse; and a computer processor 922. The computer processor 922 may include a commercially available programmable machine running a commercially available operating system. The operator workstation 916 provides the operator interface that enables scanning control parameters to be entered into the CT system 900. In general, the operator workstation 916 is in communication with a data store server 924 and an image reconstruction system 926. By way of example, the operator workstation 916, data store server 924, and image reconstruction system 926 may be connected via a communication system 928, which may include any suitable network connection, whether wired, wireless, or a combination of both. As an example, the communication system 928 may include both proprietary or dedicated networks, as well as open networks, such as the internet.
[0075] The operator workstation 916 is also in communication with a control system 930 that controls operation of the CT system 900. The control system 930 generally includes an x-ray controller 932, a table controller 934, a gantry controller 936. and a data acquisition system 938. The x-ray controller 932 provides power and timing signals to the x-ray source 904 and the gantry controller 936 controls the rotational speed and position of the gantry 902. The table controller 934 controls a table 940 to position the subject 912 in the gantry 902 of the CT system 900.
[0076] The DAS 938 samples data from the detector elements 910 and converts the data to digital signals for subsequent processing. For instance, digitized x-ray data is communicated from the DAS 938 to the data store server 924. The image reconstruction system 926 then retrieves the x-ray data from the data store server 924 and reconstructs an image therefrom. The image reconstruction system 926 may include a commercially available computer processor, or may be a highly parallel computer architecture, such as a system that includes multiple-core processors and massively parallel, high-density computing devices. Optionally, image reconstruction can also be performed on the processor 922 in the operator workstation 916. Reconstructed images can then be communicated back to the data store server 924 for storage or to the operator workstation 916 to be displayed to the operator or clinician.
[0077] The CT system 900 may also include one or more networked workstations 942. By way of example, a networked workstation 942 may include a display 944; one or more input devices 946, such as a keyboard and mouse; and a processor 948. The networked workstation 942 may be located within the same facility as the operator workstation 916, or in a different facility, such as a different healthcare institution or clinic.
[0078] The networked workstation 942, whether within the same facility or in a different facility as the operator workstation 916, may gain remote access to the data store server 924 and / or the image reconstruction system 926 via the communication system 928. Accordingly, multiple networked workstations 942 may have access to the data store server 924 and / or image reconstruction system 926. In this manner, x-ray data, reconstructed images, or other data may be exchanged betw een the data store server 924, the image reconstruction system 926, and the networked w orkstations 942, such that the data or images may be remotely processed by a networked workstation 942. This data may be exchanged in any suitable format, such as in accordance with the transmission control protocol (TCP), the internet protocol (IP), or other known or suitable protocols.
[0079] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention.
Claims
CLAIMS1. A method for training a machine learning model to denoise computed tomography (CT) image data, the method comprising:accessing routine-dose data with a computer system, wherein the routine-dose data have been acquired from a plurality7of subjects using one or more CT systems; generating low-dose CT image data from the routine-dose data with the computer system by:reconstructing thin-slice images from the routine-dose data, wherein the thin-slice images have a first slice thickness;estimating noise maps from the thin-slice images;reconstructing thick-slice images from the routine-dose data, wherein the thick-slice images have a second slice thickness that is thicker than the first slice thickness;generating low-dose thick-slice images by inserting noise from the noise maps to the thick-slice images, wherein the low-dose thick-slice images comprise the low-dose CT image data;generating routine-dose CT image data from the routine-dose data with the computer system by reconstructing routine-dose thick-slice images from the routine-dose data, wherein the routine-dose thick-slice images have the same slice thickness as the low-dose thick-slice images;assembling training data with the computer system as pairs of low-dose CT image data as training inputs and routine-dose CT image data as training targets; training a machine learning model on the training data using the computer system;andstoring the trained machine learning model using the computer system.
2. The method of claim 1, wherein reconstructing thin-slice images from the routine-dose data comprises:reconstructing first thin-slice images from simulated lower dose data at a first dose level; andreconstructing second thin-slice images from the routine-dose data at a second dose level that is higher than the first dose level.
3. The method of claim 1. wherein reconstructing thin-slice images from the routine-dose data comprises:simulating first thin-slice images from lower dose data at a first dose level; and reconstructing second thin-slice images from the routine-dose data at a second dose level that is higher than the first dose level.
4. The method of claim 2, wherein the first thin-slice images and the second thin-slice images are reconstructed using a filtered backproj ection reconstruction.
5. The method of claim 2, wherein estimating the noise maps from the thin-slice images comprises generating the noise maps by computing a difference between the first thin-slice images and the second thin-slice images.
6. The method of claim 5. wherein estimating the noise maps comprises applying a spatial decoupling to the noise maps to generate augmented noise maps.
7. The method of claim 6, wherein applying the spatial decoupling comprises applying at least one of a random translation, a random rotation, or a random flipping to each noise map.
8. The method of claim 7, wherein the random translation is applied within an imaging plane of the noise maps.
9. The method of claim 7, wherein the random translation is selected from a range of -15 to +15 pixels.
10. The method of claim 2. wherein the first dose level is a percentage of the second dose level.
11. The method of claim 10, wherein the percentage is 25%.
12. The method of claim 2. wherein the second dose level corresponds to a dose level at which the routine-dose data were acquired.
13. The method of claim 1, wherein the first slice thickness is 1.0 mm or less.
14. The method of claim 1. wherein the machine learning model comprises a neural network.
15. The method of claim 14, wherein the neural network comprises a residual neural network.
16. The method of claim 14, wherein the neural network comprises a convolutional neural netw ork.
17. The method of claim 1. wherein the thin-slice images are reconstructed using a filtered backproj ection reconstruction and the routine-dose thick-slice images are reconstructed using an iterative reconstruction.
18. A method for generating denoised computed tomography (CT) image data, the method comprising:accessing ultra-low-dose (ULD) CT image data with a computer system, wherein the ULD CT image data were acquired from a subject using a CT system at a dose level less than 1.8 mGy;accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to denoise CT image data, wherein the training data comprise low-dose CT image data as training inputs and routine-dose CT image data as targets, wherein the low-dose CT image data comprise CT images having noise inserted from noise maps having a thinner slice thickness than the CT images;generating denoised CT image data with the computer system by inputting the ULD CT image data to the machine learning model, generating the denoised CT image data as an output; andoutputting the denoised CT image data with the computer system.
19. The method of claim 18, wherein the ULD CT image data comprise raw projection data and the denoised CT image data comprise denoised raw projection data.
20. The method of claim 18, wherein the ULD CT image data comprise at least one CT image and the denoised CT image data comprise at least one denoised CT image.
21. The method of claim 20, wherein the at least one CT image comprises a thin-slice image and the at least one denoised CT image comprises a denoised thin-slice image.
22. The method of claim 21, wherein the thin-slice image and the denoised thin-slice image have a slice thickness of 1.0 mm or less.
23. A method for processing computed tomography (CT) image data using a trained machine learning model, the method comprising:accessing CT image data with a computer system, wherein the CT image data comprise at least one of projection data or reconstructed images;accessing a trained machine learning model with the computer system, wherein the trained machine learning model has been trained using training data comprising synthesized low-dose CT image data as training inputs and routine-dose CT image data as training targets, wherein the synthesized low-dose CT image data were generated by inserting noise maps derived from thin-slice images into thick-slice images;processing the CT image data with the trained machine learning model to generate processed CT image data having reduced noise compared to the CT image data; andoutputting the processed CT image data with the computer system.
24. The method of claim 23, wherein the noise maps were generated by computing a difference between first thin-slice images reconstructed at a first dose level and second thin-slice images reconstructed at a second dose level higher than the first dose level.
25. The method of claim 23, wherein the thin-slice images have a slice thickness of 1.0 mm or less and the thick-slice images have a slice thickness greater than 1.0 mm.
26. The method of claim 23, wherein the trained machine learning model comprises a convolutional neural network configured to reduce hallucinated structures in ultra-low-dose CT image data.
27. The method of claim 23, wherein the CT image data were acquired at a dose level of 1.8 mGy or less and the processed CT image data exhibit improved spatial resolution and reduced noise compared to conventional iterative reconstruction methods.