Deep learning-based monoenergetic image generation
Deep learning-based methods combine features from multi-energy CT images at different energy levels to generate optimized VMIs, addressing trade-offs in existing CT techniques by reducing artifacts and enhancing image quality and efficiency.
Patent Information
- Application Number
- PCT/US2025/038311
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-14
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing multi-energy computed tomography (CT) techniques face intrinsic trade-offs between virtual monoenergetic images (VMIs) at different keV levels, such as enhanced iodine and bone contrasts versus reduced beam hardening and metal artifacts, limiting their clinical effectiveness.
A deep learning-based method using generative adversarial networks and machine learning models synthesizes enhanced VMIs by combining features from VMIs at different energy levels, integrating advantageous spectral characteristics to overcome these trade-offs.
The method generates optimized VMIs that reduce blooming artifacts, enhance contrast-to-noise ratio, and preserve spatial resolution, improving image quality and reader efficiency while simplifying the imaging process across various CT scanner types.
Smart Images

Figure US2025038311_22012026_PF_FP_ABST
Abstract
Description
Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 DEEP LEARNING-BASED MONOENERGETIC IMAGE GENERATION CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 673,571, filed on July 19, 2024, and entitled “DEEP LEARNING-BASED MONOENERGETIC IMAGE GENERATION,” and U.S. Provisional Patent Application Serial No. 63 / 758,929, filed on February 14, 2025, and entitled “DEEP LEARNING-BASED MONOENERGETIC IMAGE GENERATION,” both of which are herein incorporated by reference in their entirety. STATEMENT OF FEDERALLY SPONSORED RESEARCH
[0002] This invention was made with government support under EB028590 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND
[0003] Multi-energy computed tomography (CT), including conventional dual-energy CT with energy-integrating-detectors (EID) and photon-counting detector CT, enables the generation of virtual monoenergetic images (VMIs) and multi-energy CT images from the acquired multi-energy data. Each multi-energy CT image acquired at a specific beam energy can be equivalent to a VMI at the effective energy level of the corresponding tube potential. These VMIs, obtained at various photon energy levels (i.e., keVs), exhibit distinct image properties suitable for diverse clinical applications. For instance, VMIs at lower keV levels enhance iodine and bone contrasts, albeit with blooming artifacts. On the other hand, VMIs at higher keV levels effectively reduce beam hardening, calcium blooming, and metal artifacts, though they result in reduced contrast. Thus, there are intrinsic trade-offs between VMIs at different keV levels. SUMMARY OF THE DISCLOSURE
[0004] It is an aspect of the present disclosure to provide a method for generating an enhanced virtual monoenergetic image (VMI). The method includes accessing a VMI with a computer system, and accessing a machine learning model with the computer system. The machine learning model has been trained on training data to synthesize an enhanced VMI from 1 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 a single input VMI by synthesizing and combining learned features from VMIs corresponding to different energy levels. The VMI is input to the machine learning model using the computer system, generating an enhanced VMI as an output. The enhanced VMI combines a first VMI feature that is characteristic of a first energy level and a second VMI feature that is characteristic of a second energy level. The enhanced VMI may then be outputted by the computer system. Other embodiments of this aspect include corresponding systems (e.g., computer systems), programs, algorithms, and / or modules, each configured to perform the steps of the methods.
[0005] It is another aspect of the present disclosure to provide a method for training a generative adversarial network to synthesize an enhanced virtual monoenergetic image (VMI) from an input VMI. The method includes accessing training data with a computer system, where the training data include first VMI data including VMIs associated with a first energy level and second VMI data including VMIs associated with a second energy level that is different from the first energy level. The method also includes accessing a generative adversarial network with the computer system, where the generative adversarial network includes a generator network and a discriminator network. The generative adversarial network is trained on the training data by: inputting a VMI from the first VMI data to the generator network, inputting the VMI from the first VMI data and a VMI from the second VMI data to the discriminator network, and minimizing a first loss function for the generator network and a second loss function for the discriminator network. The method also includes storing the trained generative adversarial network with the computer system. Other embodiments of this aspect include corresponding systems (e.g., computer systems), programs, algorithms, and / or modules, each configured to perform the steps of the methods.
[0006] It is yet another aspect of the present disclosure to provide a method for generating an enhanced VMI, in which VMI image data are accessed with a computer system. The VMI image data include at least a first VMI and a second VMI. A machine learning model is also accessed with the computer system, where the machine learning model has been trained on training data to synthesize an enhanced VMI from multiple input VMIs by synthesizing and combining learned features from VMIs corresponding to different energy levels. The VMI image data are input to the machine learning model using the computer system, generating an enhanced VMI as an output. The enhanced VMI combines a first VMI feature that is characteristic of the first VMI and a second VMI feature that is characteristic of the second VMI. The enhanced VMI is then output using the computer system. 2 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0007] It is still another aspect of the present disclosure to provide a method for generating an enhanced virtual monoenergetic image (VMI), which includes accessing multi- energy computed tomography (MECT) data with a computer system, where the MECT data include images acquired at least at a first energy level and a second energy level. A machine learning model is also accessed with the computer system, where the machine learning model has been trained on training data to selectively integrate spectral characteristics from different energy levels in MECT data. The MECT data are input to the machine learning model using the computer system, generating an enhanced VMI as an output. The enhanced VMI combines a first feature that is characteristic of the first energy level and a second feature that is characteristic of the second energy level. The enhanced VMI may then be output with the computer system. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG.1 is a schematic overview of an example machine learning framework for generating enhanced virtual monoenergetic images (VMIs) that combine the advantageous features that are characteristic of VMIs attained at different energy levels. The framework, which may be referred to as an OPAL framework, includes a two-stage process. Stage 1 involves the training phase where generators learn background and contrast features from high- energy (e.g., 100 keV) and low-energy (e.g., 70 keV) VMIs, respectively, using a PatchGAN for structure discrimination. Stage 2 depicts the inference phase, in which the trained model is applied to a single VMI as an input.
[0009] FIG.2 illustrates an example of a simplified architecture of a U-Net model used in the generator of the machine learning framework illustrated in FIG.1.
[0010] FIG. 3 is a flowchart of an example method for generating an enhance VMI from an input of a single VMI using a suitably trained machine learning model, where the enhanced VMI combines advantageous features characteristic of VMIs attained at different energy levels.
[0011] FIG. 4 is a flowchart of an example method for training a machine learning model to generate an enhanced VMI from an input of a single VMI, where the enhanced VMI combines advantageous features characteristic of VMIs attained at different energy levels.
[0012] FIG.5 shows examples of patient images with standard resolution multi-energy mode including 70 keV, OPAL, and 100 keV VMIs, according to an example study. The white dashed rectangle highlights the zoomed-in region of interest (ROI), with the display window 3 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 (W / L) set at 1500 / 400 HU. Notably, the white arrow points to the enhanced clarity of calcification boundaries achieved by OPAL compared to the 70 keV input.
[0013] FIG. 6 shows contrast-to-noise ratio (CNR) comparisons evidenced by the improved performance of images generated with the disclosed systems and methods (5.09) relative to the 70 keV (4.84) and 100 keV (2.13) VMIs. CNR calculations were performed in areas foregrounded by the red rectangle and backgrounded by the blue rectangle, as indicated in FIG.5.
[0014] FIG.7 shows a line profile, extracted along the red dashed line in FIG.5, which showcases the ability of the disclosed systems and methods to maintain the spatial resolution of the 70 keV input while preserving the detailed calcification edges observed in the 100 keV images as indicated by the red arrow.
[0015] FIG.8 is a schematic overview of an example machine learning framework for generating enhanced VMIs that combine the advantageous features that are characteristic of VMIs attained at different energy levels. The framework, which may be referred to as a CITRINE framework, includes a two-stage process. Stage 1 involves the training phase where generators learn background and contrast features from high-energy (e.g., 100 keV) and low- energy (e.g., 70 keV) VMIs, respectively, using a PatchGAN for structure discrimination. Stage 2 depicts the inference phase, in which the trained model is applied to two or more VMIs as inputs.
[0016] FIG. 9 is a flowchart of an example method for generating an enhance VMI from an input of a two or more VMIs using a suitably trained machine learning model, where the enhanced VMI combines advantageous features of the input VMIs, which may be features characteristic of VMIs attained at different energy levels.
[0017] FIG. 10 is a flowchart of an example method for training a machine learning model to generate an enhanced VMI from an input of two or more VMIs, where the enhanced VMI combines advantageous features characteristic of the input VMIs.
[0018] FIGS. 11A and 11B illustrates a performance comparison of a representative slice from one patient. FIG. 11A shows images from columns left to right are 70 keV VMI, 100 keV VMI, and CITRINE. The regions of interest marked by the white rectangle are zoomed below, respectively. Image display window (WW / WL): 1000 / 400 HU. FIG. 11B shows line profiles across the coronary vessel as indicated by the yellow dashed line, comparing the performance of 70 keV, 100 keV, and CITRINE images. The results confirm that the calcium 4 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 edges in CITRINE are identical to those at 100 keV, while the contrast remains consistent with the 70 keV images.
[0019] FIG. 12 illustrates percent diameter stenosis quantification using original 70 KeV, 100 keV, and CITRINE images in three patient cases demonstrates that CITRINE lowers percent diameter stenosis compared to original 70 keV, yielding results similar to those from 100 keV VMI.
[0020] FIG. 13 is a block diagram of an example system for generating an enhanced VMI according to some embodiments described in the present disclosure.
[0021] FIG. 14 is a block diagram of example components that can implement the system of FIG.13.
[0022] FIGS.15A and 15B illustrate an example CT system that can be configured to acquire multi-energy CT data. DETAILED DESCRIPTION
[0023] Described here are systems and methods for generating optimized virtual monoenergetic images (VMIs) from multi-energy computed tomography (CT) data. An optimized VMI combines the benefits of different energy (keV) levels. In this way, an optimized VMI can be representative of a single keV level while incorporating advantageous features, properties, and / or characteristics from other keV levels. For example, an optimized VMI may include the enhanced contrast of a low keV VMI and the reduced artifacts from a high keV VMI. Advantageously, an optimized VMI generated using the disclosed systems and methods can avoid the intrinsic trade-off of existing techniques for generating VMIs, thereby enabling the full benefits of multi-energy CT.
[0024] A machine learning model is used to selectively integrate advantageous spectral characteristics from different keV levels in multi-energy CT data, which may include multi- energy CT data acquired with an energy-integrating detector, a photon-counting detector, or the like. The machine learning model may be a deep learning model. For instance, the machine learning model may be a deep neural network. It is an aspect of the present disclosure to not only improve the image quality of the resulting VMI and enhance reader efficiency, but to simplify the overall imaging process and workflow.
[0025] The disclosed systems and methods can operate directly on CT images without the need for access to raw projection data or other proprietary information (e.g., vendor reconstruction algorithms, etc.). It is thus an advantage of the disclosed systems and methods 5 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 that they are compatible across a broad range of CT scanner types and models, including EID- CT, PCD-CT, and cone-beam CT systems.
[0026] It is an aspect of the present disclosure to provide an optimized monoenergetic image generation framework via adversarial learning between background and contrast, which may be referred to as an OPAL framework. The OPAL framework employs two deep-learning generators to learn low blooming artifacts and high contrast-to-noise (CNR) features from two different keV VMIs. As a non-limiting example, the OPAL framework can learn these features from 100 keV and 70 keV VMIs. In some embodiments, a PatchGAN discriminator can be used to assess efficacy when learning the blooming artifact and high CNR features.
[0027] FIG.1 illustrates an OPAL framework, which may include two stages. Stage 1 involves training a machine learning model in the OPAL framework. In this stage, generators are employed to assimilate the background and contrast from VMIs corresponding to two different energy levels: a low energy level and a high energy level. As illustrated, a first synthetic VMI is generated using a first generator and a second synthetic VMI is generated using a second generator. The background and contrast information from the first and second synthetic VMIs are separated. The contrast information from the first and second synthetic images is fused, or otherwise combined, to create a contrast-only VMI. Likewise, the background information from the first and second synthetic images is fused, or otherwise combined, to create a background-only VMI. The contrast-only and background-only VMIs are then fused, or otherwise combined, to generate the final, optimized VMI.
[0028] As noted above, in a non-limiting example the low energy level can be 70 keV and the high energy level can be 100 keV. A generative adversarial network (GAN) is used to discriminate between authentic and synthesized image patches created by the generator(s). As noted above, in a non-limiting example the GAN may include a PatchGAN discriminator. The PatchGAN discriminator penalizes structure at the scale of local image patches rather than the global image.
[0029] Upon training completion, in the second stage of the OPAL framework, inference is carried out using a single VMI as an input. For example, a single 70 keV VMI may be used as an input to the trained machine learning model in the OPAL framework.
[0030] FIG.2 illustrates an example neural network architecture for implementing the generator(s) in the OPAL framework. The neural network architecture may be implemented as a simplified U-Net architecture with nine modules. Each module involves convolution, batch normalization (BN), and exponential linear unit (eLU) activation operations sequentially. The 6 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 max pooling layer and convolution transpose operator are applied in the network. The concatenation is added to the network to preserve the similarity between the input and output.
[0031] Referring now to FIG. 3, a flowchart is illustrated as setting forth the steps of an example method for generating an optimized or otherwise enhanced VMI using a suitably trained machine learning model. The machine learning model takes a single VMI as input data and generates an optimized or otherwise enhanced VMI as an output, where the optimized or otherwise enhanced VMI combines the beneficial features, properties, and / or characteristics of both low-energy and high-energy VMIs. In this way, the optimized VMI is enhanced relative to the input VMI. As one example, artifacts (e.g., blooming artifacts) in the input VMI may be reduced in the optimized VMI. Additionally or alternatively, CNR in the input VMI may be increased or otherwise enhanced in the optimized VMI.
[0032] The method includes accessing VMI data with a computer system, as indicated at step 302. Accessing the VMI data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the VMI data may include acquiring multi-energy CT data with a CT system or access previously acquired multi-energy CT data, generating one or more VMIs from the multi-energy CT data, and transferring or otherwise communicating the VMI data to the computer system, which may be a part of the CT system. In still other examples, as described above the multi-energy CT data may be accessed with the computer system and retained as inputs for the machine learning models described in the present disclosure.
[0033] The VMI data may include one or more VMIs. As described above, the VMI data may include a single VMI for processing. Additionally or alternatively, the VMI data may include more than one VMI for processing. In these instances, the multiple VMIs may be individually processed. The multiple VMIs may be processed sequentially or in parallel (e.g., using more than one machine learning model).
[0034] A trained machine learning model is then accessed with the computer system, as indicated at step 304. In general, the machine learning model is trained, or has been trained, on training data in order to generate an optimized VMI that combines the advantageous features, properties, and / or characteristics of at least two different energy levels (i.e., VMIs corresponding to those at least two different energy levels, multi-energy CT data acquired at the at least two different energy levels). In some implementations the machine learning model may receive one or more VMIs as an input. As described above, in other examples the machine learning model may receive multi-energy CT data to selectively integrate advantageous 7 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 spectral characteristics from different energy levels in the multi-energy CT data, which may include multi-energy CT data acquired with an energy-integrating detector, a photon-counting detector, or the like.
[0035] Accessing the trained machine learning model may include accessing model parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the machine learning model on training data. In some instances, retrieving the machine learning model can also include retrieving, constructing, or otherwise accessing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0036] As one non-limiting example, the machine learning model may include a generative adversarial network (GAN). As described above, the generator(s) of the machine learning model may be implemented as an artificial neural network, such as the artificial neural network illustrated in FIG. 2. The GAN may implement a PatchGAN for the discriminator portion of the model.
[0037] An artificial neural network generally includes an input layer, one or more hidden layers (or nodes), and an output layer. Typically, the input layer includes as many nodes as inputs provided to the artificial neural network. The number (and the type) of inputs provided to the artificial neural network may vary based on the particular task for the artificial neural network.
[0038] The input layer connects to one or more hidden layers. The number of hidden layers varies and may depend on the particular task for the artificial neural network. Additionally, each hidden layer may have a different number of nodes and may be connected to the next layer differently. For example, each node of the input layer may be connected to each node of the first hidden layer. The connection between each node of the input layer and each node of the first hidden layer may be assigned a weight parameter. Additionally, each node of the neural network may also be assigned a bias value. In some configurations, each node of the first hidden layer may not be connected to each node of the second hidden layer. That is, there may be some nodes of the first hidden layer that are not connected to all of the nodes of the second hidden layer. The connections between the nodes of the first hidden layers and the second hidden layers are each assigned different weight parameters. Each node of the hidden layer is generally associated with an activation function. The activation function defines 8 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 how the hidden layer is to process the input received from the input layer or from a previous input or hidden layer. These activation functions may vary and be based on the type of task associated with the artificial neural network and also on the specific type of hidden layer implemented.
[0039] Each hidden layer may perform a different function. For example, some hidden layers can be convolutional hidden layers which can, in some instances, reduce the dimensionality of the inputs. Other hidden layers can perform statistical functions such as max pooling, which may reduce a group of inputs to the maximum value; an averaging layer; batch normalization; and other such functions. In some of the hidden layers each node is connected to each node of the next hidden layer, which may be referred to then as dense layers. Some neural networks including more than, for example, three hidden layers may be considered deep neural networks.
[0040] The last hidden layer in the artificial neural network is connected to the output layer. Similar to the input layer, the output layer typically has the same number of nodes as the possible outputs.
[0041] The VMI data are then input to the machine learning model, generating optimized or otherwise enhanced VMI data as an output, as indicated at step 306. For instance, a single VMI from the VMI data is input to the machine learning model to generate an optimized VMI. The optimized VMI combines the beneficial features, properties, and / or characteristics of VMIs from two different energy levels (e.g., both a low-energy VMI and a high-energy VMI). In this way, the optimized VMI is enhanced relative to the input VMI. Artifacts (e.g., blooming artifacts) in the input VMI may be reduced in the optimized VMI. Additionally or alternatively, CNR in the input VMI may be increased or otherwise enhanced in the optimized VMI. As described above, in other implementations the multi-energy CT data accessed with the computer system may be used as an input to the machine learning model to selectively integrate advantageous spectral characteristics from different keV levels in the multi-energy CT data.
[0042] The optimized or otherwise enhanced VMI data generated by inputting the VMI data or multi-energy CT data to the trained machine learning model can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 308. For example, one or more optimized or otherwise enhanced VMIs in the optimized or otherwise enhanced VMI data can be displayed to a user using a computer system. 9 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0043] Referring now to FIG. 4, a flowchart is illustrated as setting forth the steps of an example method for training a machine learning model on training data, such that the machine learning model is trained to receive a single VMI as input data in order to generate an optimized or otherwise enhanced VMI as an output.
[0044] In general, the machine learning model can implement any number of different deep learning model architectures. As one example, the deep learning model may be a GAN. As another example, the deep learning model may be a transformer network or model. In still other examples, the deep learning model may otherwise implement one or more neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like.
[0045] The method includes accessing training data with a computer system, as indicated at step 402. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with a CT system and transferring or otherwise communicating the data to the computer system. Additionally or alternatively, accessing the training data may include acquiring multi-energy CT data with a CT system or accessing previously acquired multi-energy CT data, and generating VMIs as training data from the multi-energy CT data, where the VMIs may then be transferred or otherwise communicated to the computer system. In still other examples, as described above the multi-energy CT data may be accessed with the computer system and retained as training inputs for the machine learning models described in the present disclosure.
[0046] In general, the training data can include VMIs corresponding to at least two different energy levels. In a non-limiting example, the training data can include VMIs corresponding to two different energy levels: a low-energy level and a high-energy level. The low-energy level may be 70 keV and the high-energy level may be 100 keV. Additionally or alternatively, the training data may include multi-energy CT data as described above.
[0047] The method can include assembling training data from the VMIs or multi- energy CT data using a computer system. This step may include assembling the VMIs into an appropriate data structure on which the machine learning model can be trained. Assembling the training data may include assembling VMIs, segmented VMIs, and other relevant data. For instance, assembling the training data may include generating labeled data and including the labeled data in the training data. Labeled data may include VMIs, segmented VMIs, or other 10 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 relevant data that have been labeled as belonging to, or otherwise being associated with, one or more different classifications or categories.
[0048] A machine learning model is then trained on the training data, as indicated at step 404. In general, the machine learning model can be trained by optimizing model parameters (e.g., weights, biases, or both) based on minimizing one or more loss functions. When the machine learning model is a GAN-based model, during training the generator and discriminator update their weights one at a time in an adversarial manner. In this process, the discriminator is trained to detect the synthetic images. On the other hand, the generator is trained to minimize a loss (e.g., an L1 loss) between the synthetic images and ground truth images. Training is complete when an equilibrium is reached between the generator and discriminator losses.
[0049] The input to the generator is a single VMI and the output is a single optimized VMI that is synthesized by the generator. The input to the discriminator was the optimized contrast and background paired with the input VMI. The discriminator is trained to classify patches of the generator output as either synthetic or real data.
[0050] The one or more trained machine learning models are then stored for later use, as indicated at step 406. Storing the machine learning model(s) may include storing model parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the machine learning model(s) on the training data. Storing the trained machine learning model(s) may also include storing the particular model architecture and / or neural network architecture(s) to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.
[0051] In an example study implementing the systems and methods described in the present disclosure, patient images from 10 cases acquired with a dual-source PCD-CT system in multi-energy mode were used to generate VMIs for training and inference. Cases were randomly split into training sets (n=6) and testing sets (n=4). In each case, VMIs at 70 keV and 100 keV were reconstructed following the standard clinical protocol, with an iterative reconstruction (IR) algorithm at strength 4, a Bv60 kernel, 0.6 mm slice thickness, 1024×1024 matrix, and 200x200 mm2field of view. An optimized monoenergetic image generation framework via adversarial learning between background and contrast (OPAL) framework such as the one described above was implemented. As described above, the OPAL framework employs two deep learning generators to learn low blooming artifact and high CNR features 11 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 from 100 keV and 70 keV VMIs, respectively, with a PatchGAN discriminator assessing efficacy. The performance of the OPAL framework was evaluated for its ability to reduce blooming artifacts, enhance contrast, and preserve spatial details.
[0052] The efficacy of the OPAL framework was evaluated through a standard resolution multi-energy mode test, as illustrated in FIG. 5. The findings reveal that OPAL generates superior image quality to both 70 keV and 100 keV VMIs. OPAL-generated images showed reduced blooming artifacts as the 100 keV VMI and high contrast of the coronary lumen as the 70 keV VMI. Notably, a detailed examination of the zoomed-in area in FIG. 5 shows the capability of the OPAL framework to significantly reduce blooming artifacts around calcifications compared to the original inputs. Furthermore, FIG. 6 demonstrates that OPAL achieved the highest contrast-to-noise ratio (CNR) of 5.09, surpassing 4.84 for 70 keV and 2.12 for 100 keV VMIs. A line profile analysis, indicated by the red dashed line in FIG. 5 and presented in FIG.7, confirms that OPAL preserves the spatial resolution of the original 70 keV input while maintaining the detailed calcification edges observed in the 100 keV images as indicated by the red arrow.
[0053] In the examples described above, a single VMI was used as an input to a machine learning model to generate an optimized or otherwise enhanced VMI as an output. In other examples, multiple VMIs acquired at different energy levels may be input to a machine learning model to generate the optimized or otherwise enhanced VMI.
[0054] By way of example, optimized or otherwise enhanced VMIs may be generated with the use of a machine learning model (e.g., a deep neural network) that selectively integrates advantageous spectral characteristics from different keV levels in multi-energy and PCD-CT. This advancement not only improves image quality, and enhances reader efficiency, but also simplifies the overall imaging process and workflow. Advantageously, the disclosed systems and methods can work directly on CT images, bypassing the need for projection data, raw data, and / or proprietary inputs, enabling compatibility across a broad range of CT scanner types and models, including EID-CT and PCD-CT systems.
[0055] In these implementations, the machine learning model may be referred to as a contrast-guided virtual monoenergetic image synthesis (CITRINE) framework that utilizes adversarial learning to synthesize images by integrating beneficial spectral characteristics from various keV levels. An example CITRINE model framework is illustrated in FIG. 8. The framework employs two generators to learn from 70 keV VMI. The first generator, G1, extracts contrast directly into G1_Contrast, while the second generator, G2, processes the background 12 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 to generate G2_Background and derives G2_Contrast by subtracting the input image from this background. Here, contrast refers to the difference between the 70 keV VMI and the chosen higher keV VMI (e.g., a 100 keV VMI). The background is characterized as having fewer blooming artifacts and reduced contrast, and its texture quality is maintained by a specific loss function applied to the 100 keV VMI. Examples of such loss functions are described below in more detail.
[0056] The fused image contrast, F_Contrast, is formed by integrating features from both G1_Contrast and G2_Contrast. Similar to the OPAL model described above, a PatchGAN discriminator may be used to differentiate between authentic and synthesized patches, ensuring structural integrity at the patch level. PatchGAN assesses either the real image pair {Real_Contrast, 70 keV VMI} or the fake image pair {F_Contrast, 70 keV VMI}, where Real_Contrast is calculated by subtracting 70 keV VMI from 100 keV VMI. Additionally, a contrast loss function regulates the correlation between F_Contrast and Real_Contrast to ensure that the generated fused contrast maintains the desired level of contrast differentiation.
[0057] Following this, F_Contrast is fused with the original 100 keV VMI by the second fusion mechanism, Fusion_2, to create a final synthesized image that optimally blends detailed features.
[0058] Both the G1 and G2 networks may utilize a simplified U-Net architecture composed of nine modules. Each module may include a sequence of convolution, batch normalization, and exponential linear unit activation operations. The network may also incorporate max pooling layers and convolution transpose operators to manage spatial dimensions, with concatenation used to maintain consistency between the input and output images.
[0059] The first fusion mechanism, Fusion_1, utilizes a concatenation operation followed by a 1x1 convolution to function as a weighted averaging process, with learnable weights provided by the convolution operation. In the second fusion mechanism, Fusion_2, an addition strategy combines the 100 keV and contrast information to produce the final synthesized output image.
[0060] Referring now to FIG. 9, a flowchart is illustrated as setting forth the steps of an example method for generating an optimized or otherwise enhanced VMI using a suitably trained machine learning model. The machine learning model takes two or more VMIs as an input and generates an optimized or otherwise enhanced VMI as an output, where the optimized or otherwise enhanced VMI combines the beneficial features, properties, and / or characteristics 13 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 of both the two or more input VMIs. In this way, the optimized VMI is enhanced relative to the individual VMIs that are input to the machine learning model. As one example, artifacts (e.g., blooming artifacts) in one or more of the input VMIs may be reduced in the optimized VMI. Additionally or alternatively, CNR in one or more of the input VMIs may be increased or otherwise enhanced in the optimized VMI.
[0061] The method includes accessing VMI data with a computer system, as indicated at step 902. Accessing the VMI data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the VMI data may include acquiring multi-energy CT data with a CT system or access previously acquired multi-energy CT data, generating one or more VMIs from the multi-energy CT data, and transferring or otherwise communicating the VMI data to the computer system, which may be a part of the CT system. In still other examples, as described above the multi-energy CT data may be accessed with the computer system and retained as inputs for the machine learning models described in the present disclosure.
[0062] The VMI data may include two or more VMIs. As one non-limiting example, the VMI data may include two VMIs for processing. The multiple VMIs may be processed together by applying them as inputs to a single machine learning model.
[0063] A trained machine learning model is then accessed with the computer system, as indicated at step 904. In general, the machine learning model is trained, or has been trained, on training data in order to generate an optimized VMI that combines the advantageous features, properties, and / or characteristics of multiple VMIs acquired at different energy levels (i.e., VMIs corresponding to two or more different energy levels, multi-energy CT data acquired at the different energy levels). In some implementations the machine learning model may receive one or more VMIs as an input. As described above, in other examples the machine learning model may receive multi-energy CT data to selectively integrate advantageous spectral characteristics from different keV levels in the multi-energy CT data, which may include multi-energy CT data acquired with an energy-integrating detector, a photon-counting detector, or the like.
[0064] Accessing the trained machine learning model may include accessing model parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the machine learning model on training data. In some instances, retrieving the machine learning model can also include retrieving, constructing, or otherwise accessing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural 14 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0065] As one non-limiting example, the machine learning model may include a generative adversarial network (GAN). As described above, the generator(s) of the machine learning model may be implemented as an artificial neural network, such as the artificial neural network illustrated in FIG. 8. The GAN may implement a PatchGAN for the discriminator portion of the model.
[0066] The VMI data are then input to the machine learning model, generating optimized or otherwise enhanced VMI data as an output, as indicated at step 906. For instance, two VMIs from the VMI data are input to the machine learning model to generate an optimized VMI. The two VMIs may include a first VMI acquired at a first energy level and a second VMI acquired at a second energy level. As a non-limiting example, first energy level may be 70 keV and the second energy level may be 100 keV. The optimized VMI combines the beneficial features, properties, and / or characteristics of the inputs VMIs from (e.g., both a low-energy VMI and a high-energy VMI). In this way, the optimized VMI is enhanced relative to the input VMIs. Artifacts (e.g., blooming artifacts) in the input VMIs may be reduced in the optimized VMI. Additionally or alternatively, CNR in the input VMIs may be increased or otherwise enhanced in the optimized VMI. By utilizing both 70 keV and 100 keV images, CITRINE stabilizes network output and enhances image feature synthesis. The 100 keV or other high keV images serve as a conditional input, ensuring that the synthesized image (i.e., the optimized or otherwise enhanced VMI) incorporates features from both inputs, rather than producing a random sample within the 100 keV target distribution. As described above, in other implementations the multi-energy CT data accessed with the computer system may be used as an input to the machine learning model to selectively integrate advantageous spectral characteristics from different keV levels in the multi-energy CT data.
[0067] The optimized or otherwise enhanced VMI data generated by inputting the VMI data or multi-energy CT data to the trained machine learning model can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 908. For example, one or more optimized or otherwise enhanced VMIs in the optimized or otherwise enhanced VMI data can be displayed to a user using a computer system.
[0068] Referring now to FIG.10, a flowchart is illustrated as setting forth the steps of an example method for training a machine learning model on training data, such that the 15 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 machine learning model is trained to receive two or more VMIs as an input in order to generate an optimized or otherwise enhanced VMI as an output.
[0069] In general, the machine learning model can implement any number of different deep learning model architectures. As one example, the deep learning model may be a GAN. As another example, the deep learning model may be a transformer network or model. In still other examples, the deep learning model may otherwise implement one or more neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like.
[0070] The method includes accessing training data with a computer system, as indicated at step 1002. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with a CT system and transferring or otherwise communicating the data to the computer system. Additionally or alternatively, accessing the training data may include acquiring multi-energy CT data with a CT system or accessing previously acquired multi-energy CT data, and generating VMIs as training data from the multi-energy CT data, where the VMIs may then be transferred or otherwise communicated to the computer system. In still other examples, as described above the multi-energy CT data may be accessed with the computer system and retained as inputs for the machine learning models described in the present disclosure.
[0071] In general, the training data can include VMIs corresponding to at least two different energy levels. In a non-limiting example, the training data can include VMIs corresponding to two different energy levels: a low-energy level and a high-energy level. The low-energy level may be 70 keV and the high-energy level may be 100 keV. Additionally or alternatively, the training data may include multi-energy CT data as described above.
[0072] The method can include assembling training data from the VMIs or multi- energy CT data using a computer system. This step may include assembling the VMIs into an appropriate data structure on which the machine learning model can be trained. Assembling the training data may include assembling VMIs, segmented VMIs, and other relevant data. For instance, assembling the training data may include generating labeled data and including the labeled data in the training data. Labeled data may include VMIs, segmented VMIs, or other relevant data that have been labeled as belonging to, or otherwise being associated with, one or more different classifications or categories. 16 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0073] A machine learning model is then trained on the training data, as indicated at step 1004. In general, the machine learning model can be trained by optimizing model parameters (e.g., weights, biases, or both) based on minimizing one or more loss functions. When the machine learning model is a GAN-based model, during training the generator and discriminator update their weights one at a time in an adversarial manner. In this process, the discriminator is trained to detect the synthetic images. On the other hand, the generator is trained to minimize a loss (e.g., an L1 loss) between the synthetic images and ground truth images. Training is complete when an equilibrium is reached between the generator and discriminator losses.
[0074] The input to the generator includes two or more VMIs and the output is a single optimized VMI that is synthesized by the generator. The input to the discriminator may include the optimized contrast and background paired with one or more of the input VMIs. The discriminator is trained to classify patches of the generator output as either synthetic or real data.
[0075] As a non-limiting example, the total loss for the CITRINE model may include a PatchGAN loss, a contrast loss, and a background loss, which may be defined as: ^^^ூ்ோூோ ൌ ^^^^^^^^ீ^^^^^^^^^^ ^^^ீ^ே^^^,^^^ ^ ^^^^^^௧^^^௧ ^ ^^^^^^^^^^^௨^ௗ , (1)
[0076] ^^ீ^ே^^^,^^^is the GAN loss contributed by PatchGAN, described by:^^ீ^ே^^^,^^^ ൌ ^^^^,^^^^^^^^^^^^^^,^^^^ ^ ^^^^,^^^^^^^^^ ^1 െ ^^^^^,^^^^^^^^, (2)
[0077] where E^∙^ denotes the expectation operator, G indicates the generator, D is the discriminator, X is the input image 70 keV VMI, ^^^^^^ denotes the fused contrast F_Contrast, and ^^ is the Real_Contrast. The contrast loss can be defined as: ^^^^^௧^^^௧ ൌ ^^^^,^^^‖^^,^^^^^^‖^^ (3)
[0078] which refers to the mean absolute error between F_contrast and Real_Contrast. As the texture information of the background can be partly characterized by its gradients, the generated image G2_Background has similar gradients to the 100 keV VMI, denoted as ^^^^^^^^^. Specifically, ^^^^^^^^^௨^ௗmay be defined as: ଶ ^^^^^^^^^௨^ௗൌ ^െ ^^^^^^^^^^^ฮ, (4)
[0079] operator. 17 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0080] The one or more trained machine learning models are then stored for later use, as indicated at step 1006. Storing the machine learning model(s) may include storing model parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the machine learning model(s) on the training data. Storing the trained machine learning model(s) may also include storing the particular model architecture and / or neural network architecture(s) to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.
[0081] In an example study, the CITRINE model described in the present disclosure was implemented to generate optimized or otherwise enhanced VMIs using two or more VMIs as inputs.
[0082] This study included coronary CTA exams from seven patients, scanned using a dual-source photon counting detector CT (PCD-CT, NAEOTOM Alpha, Siemens Healthineers) with approval from our Institutional Review Board. The scans were performed using an ECG-gated cardiac scan mode, featuring settings of 120 kV and 96 × 0.2 mm collimation that enable simultaneous ultra-high resolution and multi-energy imaging capabilities. VMIs at 70 and 100 keV were reconstructed with iterative reconstruction (QIR at strength level 4), with a 512 × 512 matrix, a medium sharp vascular (Bv60) kernel, 0.4 mm slice thickness, and 0.3 mm increment.
[0083] We utilized 2,342 VMI slice pairs at 70 and 100 keV from four patients for training and subsequently applied the trained network to 1,514 VMI slice pairs from three patients that were not seen during the training. All networks were implemented using the PyTorch framework. The Adam optimizer was selected to optimize network training, with momentum 1 and 2 settings of 0.5 and 0.999, respectively. The learning rate was set at 0.0002, with a batch size of 3. The weights ^ and λ were optimized to 100 and 500, respectively, as determined by a range of experiments to achieve the best performance. The training was conducted over 150 epochs to ensure model convergence.
[0084] The ideal output of CITRINE combines high contrast from 70 keV VMI with minimal blooming artifacts from 100 keV VMI. The contrast magnitude in the images (70 keV, 100 keV, and CITRINE) was quantified by measuring the mean CT number within regions of interest (ROIs) in the contrast-enhanced aorta, chosen for their approximately uniform intensity. Spatial resolution was assessed by comparing line profiles across calcifications in the images. To evaluate the effect of blooming artifacts on stenosis quantification, coronary artery 18 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 with dense calcification were identified and processed using a commercial software (Syngo.Via, VB70) to quantify the percent diameter stenosis. Results were compared among the 70 keV, 100 keV, and CITRINE image.
[0085] FIGS. 11A and 11B illustrate the performance of the CITRINE model in this example study. As depicted in FIG. 11A, CITRINE synthesizes the high contrast observed in 70 keV VMI and the reduced blooming artifacts characteristic of 100 keV VMI. In the ROI marked by a red circle, CITRINE achieves a CT number of 331 HU, comparable to the 337 HU at 70 keV VMI and significantly higher than the 177 HU at 100 keV VMI, indicating CITRINE's ability to maintain consistent contrast similar to 70 keV. The zoomed-in view, indicated by a red arrow, shows that CITRINE mitigates blooming artifacts close to the level of 100 keV while maintaining the high contrast of 70 keV. Line profiles (shown in FIG.11B) across the coronary vessel and calcifications confirm that CITRINE effectively combines reduced blooming artifacts from 100 keV with the consistent contrast of 70 keV.
[0086] FIG. 12 demonstrates the effectiveness of CITRINE in reducing percent diameter stenosis of the left anterior descending (LAD) artery compared to the original 70 keV VMI, while closely aligning with the results from the 100 keV VMI. Compared to standard 70 keV VMI, CITRINE images demonstrate an approximate 25% reduction in luminal stenosis, decreasing from mean percentage stenosis of 17% to 12.67%. These results underscore CITRINE’s capability to synthesize the high contrast of 70 keV VMIs with the reduced blooming artifacts of 100 keV VMIs, thereby enhancing diagnostic accuracy and efficiency in cCTA exams leveraging multi-energy and PCD-CT technologies.
[0087] FIG. 13 shows an example of a system 1300 for generating enhanced, or otherwise optimized, VMIs in accordance with some embodiments described in the present disclosure. As shown in FIG. 13, a computing device 1350 can receive one or more types of data (e.g., multi-energy CT data, VMIs) from data source 1302. In some embodiments, computing device 1350 can execute at least a portion of an enhanced VMI generation system 1304 to generate enhanced VMIs from data received from the data source 1302.
[0088] Additionally or alternatively, in some embodiments, the computing device 1350 can communicate information about data received from the data source 1302 to a server 1352 over a communication network 1354, which can execute at least a portion of the enhanced VMI generation system 1304. In such embodiments, the server 1352 can return information to the computing device 1350 (and / or any other suitable computing device) indicative of an output of the enhanced VMI generation system 1304. 19 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0089] In some embodiments, computing device 1350 and / or server 1352 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 1350 and / or server 1352 can also reconstruct images from the data.
[0090] In some embodiments, data source 1302 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data), such as a CT system, another computing device (e.g., a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some embodiments, data source 1302 can be local to computing device 1350. For example, data source 1302 can be incorporated with computing device 1350 (e.g., computing device 1350 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 1302 can be connected to computing device 1350 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 1302 can be located locally and / or remotely from computing device 1350, and can communicate data to computing device 1350 (and / or server 1352) via a communication network (e.g., communication network 1354).
[0091] In some embodiments, communication network 1354 can be any suitable communication network or combination of communication networks. For example, communication network 1354 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication network 1354 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG.13 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.
[0092] Referring now to FIG. 14, an example of hardware 1400 that can be used to implement data source 1302, computing device 1350, and server 1352 in accordance with some embodiments of the systems and methods described in the present disclosure is shown. 20 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0093] As shown in FIG.14, in some embodiments, computing device 1350 can include a processor 1402, a display 1404, one or more inputs 1406, one or more communication systems 1408, and / or memory 1410. In some embodiments, processor 1402 can be any suitable hardware processor or combination of processors, such as a central processing unit (“CPU”), a graphics processing unit (“GPU”), and so on. In some embodiments, display 1404 can include any suitable display devices, such as a liquid crystal display (“LCD”) screen, a light-emitting diode (“LED”) display, an organic LED (“OLED”) display, an electrophoretic display (e.g., an “e-ink” display), a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1406 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0094] In some embodiments, communications systems 1408 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1354 and / or any other suitable communication networks. For example, communications systems 1408 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1408 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0095] In some embodiments, memory 1410 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1402 to present content using display 1404, to communicate with server 1352 via communications system(s) 1408, and so on. Memory 1410 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1410 can include random-access memory (“RAM”), read-only memory (“ROM”), electrically programmable ROM (“EPROM”), electrically erasable ROM (“EEPROM”), other forms of volatile memory, other forms of non-volatile memory, one or more forms of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1410 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 1350. In such embodiments, processor 1402 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 1352, transmit information to server 21 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 1352, and so on. For example, the processor 1402 and the memory 1410 can be configured to perform the methods described herein (e.g., implementing the machine learning model illustrated in FIGS.1 and 2, the method of FIG.3, the method of FIG.4).
[0096] In some embodiments, server 1352 can include a processor 1412, a display 1414, one or more inputs 1416, one or more communications systems 1418, and / or memory 1420. In some embodiments, processor 1412 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 1414 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1416 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0097] In some embodiments, communications systems 1418 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1354 and / or any other suitable communication networks. For example, communications systems 1418 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1418 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0098] In some embodiments, memory 1420 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1412 to present content using display 1414, to communicate with one or more computing devices 1350, and so on. Memory 1420 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1420 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1420 can have encoded thereon a server program for controlling operation of server 1352. In such embodiments, processor 1412 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1350, receive information and / or content 22 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 from one or more computing devices 1350, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.
[0099] In some embodiments, the server 1352 is configured to perform the methods described in the present disclosure. For example, the processor 1412 and memory 1420 can be configured to perform the methods described herein (e.g., implementing the machine learning model illustrated in FIGS.1 and 2, the method of FIG.3, the method of FIG.4).
[0100] In some embodiments, data source 1302 can include a processor 1422, one or more data acquisition systems 1424, one or more communications systems 1426, and / or memory 1428. In some embodiments, processor 1422 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 1424 are generally configured to acquire data, images, or both, and can include a CT system. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 1424 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of a CT system. In some embodiments, one or more portions of the data acquisition system(s) 1424 can be removable and / or replaceable.
[0101] Note that, although not shown, data source 1302 can include any suitable inputs and / or outputs. For example, data source 1302 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 1302 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.
[0102] In some embodiments, communications systems 1426 can include any suitable hardware, firmware, and / or software for communicating information to computing device 1350 (and, in some embodiments, over communication network 1354 and / or any other suitable communication networks). For example, communications systems 1426 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1426 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0103] In some embodiments, memory 1428 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for 23 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 example, by processor 1422 to control the one or more data acquisition systems 1424, and / or receive data from the one or more data acquisition systems 1424; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 1350; and so on. Memory 1428 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1428 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1428 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 1302. In such embodiments, processor 1422 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1350, receive information and / or content from one or more computing devices 1350, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.
[0104] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory or non-transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0105] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a 24 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).
[0106] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.
[0107] Referring particularly now to FIGS. 15A and 15B, an example of an x-ray computed tomography (“CT”) imaging system 1500 is illustrated. The CT system includes a gantry 1502, to which at least one x-ray source 1504 is coupled. The x-ray source 1504 projects an x-ray beam 1506, which may be a fan-beam or cone-beam of x-rays, towards a detector array 1508 on the opposite side of the gantry 1502. The detector array 1508 includes a number of x-ray detector elements 1510. Together, the x-ray detector elements 1510 sense the projected x-rays 1506 that pass through a subject 1512, such as a medical patient or an object undergoing examination, that is positioned in the CT system 1500. Each x-ray detector element 1510 produces an electrical signal that may represent the intensity of an impinging x-ray beam and, hence, the attenuation of the beam as it passes through the subject 1512. In some configurations, each x-ray detector 1510 is capable of counting the number of x-ray photons that impinge upon the detector 1510. During a scan to acquire x-ray projection data, the gantry 1502 and the components mounted thereon rotate about a center of rotation 1514 located within the CT system 1500. Although illustrated as a single-source CT system, in other examples the CT system 1500 may be configured as a dual-source CT system, which may implement two x-ray sources and / or two x-ray detectors.
[0108] The CT system 1500 also includes an operator workstation 1516, which typically includes a display 1518; one or more input devices 1520, such as a keyboard and mouse; and a computer processor 1522. The computer processor 1522 may include a 25 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 commercially available programmable machine running a commercially available operating system. The operator workstation 1516 provides the operator interface that enables scanning control parameters to be entered into the CT system 1500. In general, the operator workstation 1516 is in communication with a data store server 1524 and an image reconstruction system 1526. By way of example, the operator workstation 1516, data store sever 1524, and image reconstruction system 1526 may be connected via a communication system 1528, which may include any suitable network connection, whether wired, wireless, or a combination of both. As an example, the communication system 1528 may include both proprietary or dedicated networks, as well as open networks, such as the internet.
[0109] The operator workstation 1516 is also in communication with a control system 1530 that controls operation of the CT system 1500. The control system 1530 generally includes an x-ray controller 1532, a table controller 1534, a gantry controller 1536, and a data acquisition system 1538. The x-ray controller 1532 provides power and timing signals to the x-ray source 1504 and the gantry controller 1536 controls the rotational speed and position of the gantry 1502. The table controller 1534 controls a table 1540 to position the subject 1512 in the gantry 1502 of the CT system 1500.
[0110] The DAS 1538 samples data from the detector elements 1510 and converts the data to digital signals for subsequent processing. For instance, digitized x-ray data is communicated from the DAS 1538 to the data store server 1524. The image reconstruction system 1526 then retrieves the x-ray data from the data store server 1524 and reconstructs an image therefrom. The image reconstruction system 1526 may include a commercially available computer processor, or may be a highly parallel computer architecture, such as a system that includes multiple-core processors and massively parallel, high-density computing devices. Optionally, image reconstruction can also be performed on the processor 1522 in the operator workstation 1516. Reconstructed images can then be communicated back to the data store server 1524 for storage or to the operator workstation 1516 to be displayed to the operator or clinician.
[0111] The CT system 1500 may also include one or more networked workstations 1542. By way of example, a networked workstation 1542 may include a display 1544; one or more input devices 1546, such as a keyboard and mouse; and a processor 1548. The networked workstation 1542 may be located within the same facility as the operator workstation 1516, or in a different facility, such as a different healthcare institution or clinic. 26 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603
[0112] The networked workstation 1542, whether within the same facility or in a different facility as the operator workstation 1516, may gain remote access to the data store server 1524 and / or the image reconstruction system 1526 via the communication system 1528. Accordingly, multiple networked workstations 1542 may have access to the data store server 1524 and / or image reconstruction system 1526. In this manner, x-ray data, reconstructed images, or other data may be exchanged between the data store server 1524, the image reconstruction system 1526, and the networked workstations 1542, such that the data or images may be remotely processed by a networked workstation 1542. This data may be exchanged in any suitable format, such as in accordance with the transmission control protocol (“TCP”), the internet protocol (“IP”), or other known or suitable protocols.
[0113] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention. 27 QB\630666.01603\97396124.2
Claims
Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 CLAIMS 1. A method for generating an enhanced virtual monoenergetic image (VMI), the method comprising: accessing a VMI with a computer system; accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to synthesize an enhanced VMI from a single input VMI by synthesizing and combining learned features from VMIs corresponding to different energy levels; inputting the VMI to the machine learning model using the computer system, generating an enhanced VMI as an output, wherein the enhanced VMI combines a first VMI feature that is characteristic of a first energy level and a second VMI feature that is characteristic of a second energy level; and outputting, with the computer system, the enhanced VMI.
2. The method of claim 1, wherein the machine learning model comprises a generative adversarial network (GAN).
3. The method of claim 2, wherein the GAN comprises a GAN with a PatchGAN discriminator network.
4. The method of claim 1, wherein the first VMI feature comprises a contrast-to- noise ratio (CNR).
5. The method of claim 4, wherein the enhanced VMI has increased CNR relative to the VMI input to the machine learning model.
6. The method of claim 1 or 4, wherein the second VMI feature comprises a blooming artifact.
7. The method of claim 6, wherein the enhanced VMI has reduced blooming artifacts relative to the VMI input to the machine learning model. 28 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 8. The method of claim 1, wherein the different energy levels comprise the first energy level and the second energy level.
9. The method of claim 1 or 8, wherein the first energy level is lower than the second energy level.
10. The method of claim 9, wherein the first energy level is 70 keV and the second energy level is 100 keV.
11. A method for training a generative adversarial network to synthesize an enhanced virtual monoenergetic image (VMI) from an input VMI, the method comprising: accessing training data with a computer system, wherein the training data comprises: first VMI data comprising VMIs associated with a first energy level; second VMI data comprising VMIs associated with a second energy level that is different from the first energy level; accessing a generative adversarial network with the computer system, wherein the generative adversarial network comprises a generator network and a discriminator network; training the generative adversarial network on the training data by: inputting a VMI from the first VMI data to the generator network; inputting the VMI from the first VMI data and a VMI from the second VMI data to the discriminator network; minimizing a first loss function for the generator network and a second loss function for the discriminator network; and storing the trained generative adversarial network with the computer system.
12. The method of claim 11, wherein the discriminator network comprises a PatchGAN discriminator.
13. The method of claim 11, wherein the first energy level is 70 keV.
14. The method of claim 11 or 13, wherein the second energy level is 100 keV. 29 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 15. A method for generating an enhanced virtual monoenergetic image (VMI), the method comprising: accessing VMI image data with a computer system, wherein the VMI image data comprise at least a first VMI and a second VMI; accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to synthesize an enhanced VMI from multiple input VMIs by synthesizing and combining learned features from VMIs corresponding to different energy levels; inputting the VMI image data to the machine learning model using the computer system, generating an enhanced VMI as an output, wherein the enhanced VMI combines a first VMI feature that is characteristic of the first VMI and a second VMI feature that is characteristic of the second VMI; and outputting, with the computer system, the enhanced VMI.
16. The method of claim 15, wherein the first VMI corresponds to a first energy level and the second VMI corresponds to a second energy level that is different from the first energy level.
17. The method of claim 16, wherein the first energy level is 70 keV and the second energy level is 100 keV.
18. The method of claim 15, wherein the machine learning model comprises a generative adversarial network (GAN).
19. The method of claim 18, wherein the GAN comprises a GAN with a PatchGAN discriminator network.
20. The method of claim 15, wherein the first VMI feature comprises a contrast- to-noise ratio (CNR).
21. The method of claim 20, wherein the enhanced VMI has increased CNR relative to at least one of the first VMI or the second VMI input to the machine learning model. 30 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 22. The method of claim 15 or 21, wherein the second VMI feature comprises a blooming artifact.
23. The method of claim 22, wherein the enhanced VMI has reduced blooming artifacts relative to at least one of the first VMI or the second VMI input to the machine learning model.
24. A method for generating an enhanced virtual monoenergetic image (VMI), the method comprising: accessing multi-energy computed tomography (MECT) data with a computer system, wherein the MECT data include images acquired at least at a first energy level and a second energy level; accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to selectively integrate spectral characteristics from different energy levels in MECT data; inputting the MECT data to the machine learning model using the computer system, generating an enhanced VMI as an output, wherein the enhanced VMI combines a first feature that is characteristic of the first energy level and a second feature that is characteristic of the second energy level; and outputting, with the computer system, the enhanced VMI.
25. The method of claim 24, wherein the MECT data are acquired with an energy- integrating detector (EID) CT system.
26. The method of claim 24, wherein the MECT data are acquired with a photon- counting detector (PCD) CT system.
27. The method of claim 24, wherein the machine learning model comprises a generative adversarial network (GAN).
28. The method of claim 27, wherein the GAN comprises a GAN with a PatchGAN discriminator network. 31 QB\630666.01603\97396124.2Mayo 2024-266, 2025-058 Q&B Docket: 630666.01603 29. The method of claim 24, wherein the first feature comprises a contrast-to- noise ratio (CNR).
30. The method of claim 24 or 29, wherein the second feature comprises a blooming artifact.
31. The method of claim 24, wherein the first energy level is lower than the second energy level.
32. The method of claim 31, wherein the first energy level is 70 keV and the second energy level is 100 keV. 32 QB\630666.01603\97396124.2
Citation Information
Patent Citations
Image processing apparatus, image processing method, and computer-readable medium
US20230263492A1