Spectral computed tomography with disentangled representation
Through deep learning technology, multi-spectral images are encoded and decoded to generate anatomical structure and contrast potential spatial information, solving the problems of image noise and artifacts in photon counting CT technology, and achieving the improvement of high-quality image reconstruction and diagnostic performance.
Patent Information
- Application Number
- CN202510141300.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-08
- Publication Date
- 2025-08-22
AI Technical Summary
Existing photon counting CT technology has challenges in image noise and artifacts, making it difficult to achieve high-quality image reconstruction and material decomposition.
Deep learning technology is used to encode multispectral images through convolutional neural networks to generate anatomical structure and contrast potential spatial information, and combine unsupervised and supervised training methods to separate and combine this information to reconstruct the target image.
It improves image reconstruction quality, enhances diagnostic performance, simplifies image processing flow, reduces computing resource consumption, and is suitable for a variety of medical imaging technologies.
Smart Images

Figure CN120525976A_ABST
Abstract
Description
Technical Field
[0001] A system and method for spectral computed tomography with a disentangled representation. Background Art
[0002] Photon counting computed tomography (PCCT) is an emerging tomographic imaging technique that offers improved diagnostic performance through better spatial and energy resolution. It uses multiple energy bins to measure the spectral dependence of X-ray attenuation, similar to dual-energy computed tomography (CT) or any other form of spectral CT. However, high-quality image reconstruction and material resolution are challenging problems due to various sources of image noise and artifacts (e.g., quantum noise, charge sharing, pulse pile-up, Compton scattering, beam hardening effects, and other issues). Summary of the Invention
[0003] In an embodiment, the present disclosure relates to a system for generating a target image, the system comprising: a database of multispectral images; and a processor configured to train a latent space encoder to generate latent space information by: encoding the multispectral image to generate anatomical structure latent space information and contrast latent space information, the anatomical structure latent space information representing compressed anatomical structure information in the multispectral image, and the contrast latent space information representing compressed contrast information in the multispectral image; combining selected features of the anatomical structure latent space information and the contrast latent space information to reproduce the multispectral image; comparing the reproduced multispectral image with the multispectral image; and adjusting the encoding of the multispectral image and repeating the training until the reproduced multispectral image matches the multispectral image. After training the latent space encoder, an output decoder is trained to generate a target image by inputting the anatomical structure latent space information into a prediction model, predicting a target image through the prediction model based on the anatomical structure latent space information, comparing the predicted target image with a reference target image related to the multispectral image, and adjusting the weights of the prediction model based on the comparison, and repeating the training until the image predicted by the prediction model matches the reference target image related to the multispectral image.
[0004] In an embodiment, the processor is configured to generate the latent space information by applying a convolutional neural network to the multispectral image to extract and separate anatomical structure information and contrast information from the multispectral image.
[0005] In an embodiment, the processor is configured to select an anatomical feature as a common anatomical feature between the anatomical latent space information.
[0006] In an embodiment, the processor is configured to combine the anatomical structure latent space information and the contrast latent space information by concatenating the anatomical structure latent space information and the contrast latent space information.
[0007] In an embodiment, the processor is configured to compare the reproduced multispectral image with the multispectral image by calculating a loss function of the reproduced multispectral image compared to the multispectral image and repeating the training until the loss function is less than a loss function threshold.
[0008] In an embodiment, the multispectral image is an image of an anatomical structure captured by a medical imaging device from a common frame of reference relative to the anatomical structure and operating at different spectral frequencies.
[0009] In an embodiment, the processor is configured to average the anatomical latent space information before inputting the anatomical latent space information into the prediction model.
[0010] In an embodiment, the prediction model is a neural network that predicts the target image and compares the predicted target image with a reference target image associated with the multispectral image in a supervised manner to determine corrective measures to adjust the weights of the prediction model, which are neural network weights.
[0011] In an embodiment, the target image is a medical image associated with a human anatomical state and includes one or more of a virtual monoenergetic image, a virtual non-contrast image, a material map, and a downstream segmentation.
[0012] In an embodiment, the processor is configured to generate new latent space information about a new multispectral image using the trained latent space encoder, and input the new latent space information into the trained output decoder to generate a new target image.
[0013] A method for generating a target image, the method comprising: training a latent space encoder by a processor to generate latent space information by encoding a multispectral image to generate anatomical structure latent space information and contrast latent space information, the anatomical structure latent space information representing compressed anatomical structure information in the multispectral image and the contrast latent space information representing compressed contrast information in the multispectral image; combining selected features of the anatomical structure latent space information and the contrast latent space information to reconstruct the multispectral image; comparing the reconstructed multispectral image with the multispectral image; and adjusting the encoding of the multispectral image and repeating the training until the reconstructed multispectral image matches the multispectral image. After training the latent space encoder, training an output decoder by the processor to generate a target image by inputting the anatomical structure latent space information into a prediction model, predicting a target image based on the anatomical structure latent space information using the prediction model, comparing the predicted target image with a reference target image associated with the multispectral image, adjusting weights of the prediction model based on the comparison, and repeating the training until the image predicted by the prediction model matches the reference target image associated with the multispectral image.
[0014] In an embodiment, the method includes generating, by the processor, the latent space information by performing a convolution on the multispectral image to extract and separate anatomical structure information and contrast information from the multispectral image.
[0015] In an embodiment, the method includes selecting, by the processor, an anatomical feature as a common anatomical feature between the anatomical latent space information.
[0016] In an embodiment, the method includes combining, by the processor, the anatomical structure latent space information and the contrast latent space information by concatenating the anatomical structure latent space information and the contrast latent space information.
[0017] In an embodiment, the method includes comparing, by the processor, the reproduced multispectral image with the multispectral image by calculating a loss function of the reproduced multispectral image compared to the multispectral image and repeating the training until the loss function is less than a loss function threshold.
[0018] In an embodiment, the multispectral image is an image of an anatomical structure captured by a medical imaging device from a common frame of reference relative to the anatomical structure and operating at different spectral frequencies.
[0019] In an embodiment, the method includes averaging, by the processor, the anatomical latent space information before inputting the anatomical latent space information into the prediction model.
[0020] In an embodiment, the prediction model is a neural network that predicts the target image and compares the predicted target image with a reference target image associated with the multispectral image in a supervised manner to determine corrective measures to adjust the weights of the prediction model, which are neural network weights.
[0021] In an embodiment, the target image is a medical image associated with a human anatomical state and includes one or more of a virtual monoenergetic image, a virtual non-contrast image, a material map, and a downstream segmentation.
[0022] In an embodiment, the method includes generating, by the processor, new latent space information about a new multispectral image using the trained latent space encoder, and inputting the new latent space information into the trained output decoder to generate a new target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order that the manner in which the above-described features of the present disclosure may be understood in detail, a more particular description of the present disclosure, briefly summarized above, may be made by reference to exemplary embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only exemplary embodiments of the present disclosure and are therefore not to be considered limiting of the scope of the present disclosure, as the disclosure may admit to other equally effective exemplary embodiments.
[0024] Figure 1 A block diagram illustrating a training model for spectral computed tomography with disentangled representation according to an exemplary embodiment of the present disclosure is shown.
[0025] Figure 2 A flowchart of training a disentanglement model according to an exemplary embodiment of the present disclosure is shown.
[0026] Figure 3 A flowchart of training a prediction model according to an exemplary embodiment of the present disclosure is shown.
[0027] Figure 4 A flow chart for executing a training model for spectral computed tomography with a disentangled representation is shown according to an exemplary embodiment of the present disclosure.
[0028] Figure 5 Example images of spectral computed tomography with a disentangled representation are shown according to an exemplary embodiment of the present disclosure.
[0029] Figure 6 A block diagram of a hardware device for training and performing spectral computed tomography with a disentangled representation is shown, according to an exemplary embodiment of the present disclosure.
[0030] Figure 7A block diagram illustrating components of a hardware device for training and performing spectral computed tomography with a disentangled representation according to an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0031] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of the parts and steps, numerical expressions and numerical values set forth in these exemplary embodiments do not limit the scope of the present disclosure, unless otherwise specifically stated. The following description of at least one exemplary embodiment is essentially only illustrative and is in no way intended to limit the present disclosure, its application or its use. The technology, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but are intended to be part of this specification sheet where appropriate. In all examples illustrated and discussed herein, any specific values are to be interpreted as illustrative and non-restrictive. Therefore, other exemplary embodiments may have different values. Note that in the following figures, similar figure numerals and letters refer to similar items, and therefore once an item is defined in one figure, it may not be necessary to further discuss the item for the next figure. Below, exemplary embodiments will be described with reference to the accompanying drawings.
[0032] Spectral computed tomography provides improved diagnostic performance through improved spatial and energy resolution. The present disclosure presents techniques based on deep learning (DL) to produce representative spectral medical target (e.g., computed tomography (CT)) output images, such as virtual monoenergetic images, material maps, and downstream segmentation, from multispectral input images via a simpler disentangled representation that disentangles the CT image into a patient anatomy latent space (same for all energy levels) and a contrast latent space (different for all energy levels) in a simple unsupervised manner. The patient anatomy space is then decoded into the desired output image by supervised image synthesis. The method is typically described in the CT image domain (post-reconstruction), but can also be applied to the projection domain (pre-reconstruction). In a CT image, the raw measurements (also called projection data) can be represented by a sinusoidal graph, which is reconstructed to obtain the patient image. Therefore, pre-reconstruction methods are methods applied to sinusoidal graphs or projection data, while post-reconstruction methods are methods applied to the reconstructed image.
[0033] For example, a sinogram can be a two-dimensional or three-dimensional data set representing the raw measurements captured by a CT scanner as it rotates around the patient, detecting X-rays that have passed through the body at various angles. Pre-reconstruction techniques involve manipulating these raw data sets to correct for physical phenomena, such as beam hardening, scattering, or noise, which can affect the quality of the final image. Algorithms in this stage can also perform tasks such as filtering, normalization, and calibration to prepare the data for the subsequent reconstruction process. The goal of pre-reconstruction processing is to condition the data in a way that enhances the accuracy and quality of the image that will be generated in the next stage.
[0034] Post-reconstruction processing, on the other hand, is the processing of data after it has been converted into cross-sectional images by reconstruction algorithms. These algorithms take pre-processed sinusoids and apply techniques to construct a three-dimensional volume or a series of two-dimensional slices representing the scanned anatomical structure. Post-reconstruction processing can include various image enhancement techniques, such as denoising, contrast enhancement, and edge sharpening, to improve the diagnostic utility of the image. In addition, it can involve advanced data-driven deep learning methods, such as the disentanglement representation technology described in the present disclosure, which separates anatomical information and contrast information to facilitate the generation of specialized images to improve medical assessments. The two pre-reconstruction and post-reconstruction stages are an essential part of the CT imaging process, and each stage plays a different role in converting raw data into clinically useful images. The disclosed disentanglement method is applicable to any of these stages.
[0035] Specifically, in the pre-reconstruction domain, the disclosed technique will be applied to the raw projection data collected by a CT scanner, typically represented as sinograms. These sinograms contain raw measurements of X-ray attenuation at different energy levels and angles around the patient. Application of the technique in this domain can involve the following steps: 1. Disentanglement of the Raw Data: The unsupervised learning component of the technique will be used to separate the raw projection data into an anatomical structure latent space and a contrast latent space. This step will aim to identify and encode consistent anatomical structures present in the data, while also capturing contrast information that varies depending on the X-ray energy spectrum. 2. Noise and Artifact Reduction: By working with the raw data, the technique can address and correct various artifacts and noise sources, such as beam hardening, scatter, or electronic noise, which can degrade image quality if not properly managed prior to reconstruction. 3. Synthesis of Enhanced Sinograms: The disentangled anatomical structure latent space can then be used to generate enhanced sinograms, which are optimized for subsequent reconstruction, potentially resulting in clearer and more accurate images.
[0036] In the post-reconstruction domain, the technique will be applied to images that have been reconstructed from sinusoidal data. Applications in this area may involve: 1. Disentanglement of reconstructed images: The technique will analyze the reconstructed images, separating the reconstructed images into an anatomical structure latent space and a contrast latent space. The anatomical structure latent space will represent the structural information of the tissue, while the contrast latent space will reflect the differences in X-ray attenuation at various energy levels. 2. Image enhancement: The disentangled anatomical structure latent space will be used to synthesize enhanced images, such as virtual monoenergetic images or material maps, which can provide additional diagnostic information beyond what is available in standard reconstructed images. 3. Downstream analysis: The technique can facilitate downstream tasks, such as segmentation or identification of specific tissues, by providing a clearer representation of the anatomical structure without the confounding effects of varying contrast levels.
[0037] The benefits of the disclosed methods, apparatus, and systems described herein include, but are not limited to, the generation of high-quality spectral CT images (including monochrome images, tissue / material maps, color maps, and segmentations) in a computationally efficient manner. Spectral CT has clinical significance in many applications, such as non-invasive diagnosis in obstructive coronary artery disease, characterization of coronary atherosclerosis, and non-invasive diagnosis of urolithiasis, to name a few. The present disclosure is also useful for downstream image analysis tasks (e.g., segmentation of fine vessels). It should be noted that although the disclosed systems / methods are described with respect to the processing of CT images, it should be noted that the systems / methods may also be applied to the processing of other types of medical images, such as magnetic resonance imaging (MRI), ultrasound, X-ray, and mammography, to name a few.
[0038] Figure 1 A block diagram 100 of a training model for spectral computed tomography with disentangled representation is shown. The technique is a semi-supervised DL method that can be divided into two parts: (1) first disentangle the input CT image at different energy levels into an "anatomy" latent space and a "contrast" latent space, which can be combined to accurately reproduce the input CT image; and (2) decode the anatomical latent space information into a target output image.
[0039] In summary, the proposed method takes multiple spectral CT images as input images and encodes them into an anatomical latent representation and a contrast latent representation for each input image. The anatomical latent representation is then converted into a high-fidelity desired output image. The input images can correspond to different X-ray spectra, different detector bins, any type of energy bin, a base material decomposition image, a preliminary monochromatic image, a linear combination of the above, or a set of images including several of the above categories and base material images. The output images can be monochromatic attenuation images, material density images, material volume fraction images, density images, other physical representation images (e.g., electron density, effective Z, stopping power, etc.), anatomical structure segmentation images, and base material images.
[0040] As mentioned above, the solution can be decomposed into two parts. The first part is unsupervised (i.e., there are no labels for the latent space). An important aspect here is that anatomical structure (e.g., shape and boundaries) is common information between input CT images at different energy levels. In spectral imaging, the common information represents an artifact-free representation of the underlying tissue material. In the second part, the disentangled patient anatomical structure latent space can be decoded into target medical image results, such as virtual monoenergetic images (VMIs), virtual non-contrast images (VNCs), material maps, and downstream segmentation. This second part is a supervised image synthesis task. Alternative implementations in the projected domain can also be envisioned, in which case both input and output are in the projected domain.
[0041] Before accurately generating the target image, the model (i.e., encoder) can be trained in a semi-supervised manner. Training can generally occur in two stages. In the first stage, the disentanglement model (i.e., encoder) can be trained in an unsupervised manner. For example, a multispectral image such as low-energy image 102A and high-energy image 102B is input into a contrast encoder 104A / 104B and an anatomical structure encoder 106A / 106B. Encoders 106A / 106B encode the multispectral image to generate anatomical structure latent space information 108A / 108B, while encoders 104A / 104B encode the multispectral image to generate contrast latent space information 110A / 110B. The anatomical structure latent space information 108A / 108B represents compressed anatomical structure information in the multispectral image, while the contrast latent space information 110A / 110B represents compressed contrast information (e.g., a smaller image, matrix, vector, etc.) in the multispectral image. For example, the anatomical structure latent space information can represent a map (anatomical structure information) corresponding to a specific human tissue. The encoder chooses (through training) how it will encode the anatomical information. Similarly, the contrast latent space can provide a lookup table to transform (i.e., map) the anatomical structure back to the target image.
[0042] Essentially, the anatomical latent space information represents a refined, abstract representation of the patient's anatomy as captured by the CT system. This information is generated by an encoder that processes multispectral CT images to isolate and encode underlying anatomical information that is consistent across different energy levels. The anatomical latent space captures geometric and spatial information of tissues, bones, and organs, effectively removing variations in X-ray attenuation due to different material properties and the energy spectrum of the X-rays used during the scan. This produces information that emphasizes the structural integrity and layout of the patient's anatomy, which is a common factor in images taken at various energy levels. The anatomical latent space serves as a stable foundation for subsequent image synthesis tasks, such as creating virtual monoenergetic images or performing tissue segmentation.
[0043] On the other hand, the contrast latent space information encodes variable aspects of the CT image caused by the different energy-dependent X-ray attenuation properties of various materials in the scan volume. This information is also generated by the encoder, but focuses on capturing contrast information that varies with the energy spectrum of the X-rays. This includes distinguishing materials based on their spectral characteristics, such as identifying iodinated contrast agents, calcium deposits, or other artifacts with different attenuation curves at different energy levels. The contrast latent space provides a complementary dataset that, when combined with the anatomical latent space, can be used to reconstruct the original multispectral image.
[0044] The solution can then combine selected features of the anatomical latent space information and the contrast latent space information to reproduce the original multispectral image. An example combination can select common anatomical features between the anatomical latent space information 108A / 108B and the contrast latent space information 110A / 110B at block 111 and then concatenate the common features separately with the contrast latent space information 110A / 110B at blocks 112A / 112B. The combined image is then decoded by decoders 114A / 114B to produce a reproduced image 116A / 116B, which is compared to the original multispectral image 102A / 102B. The system adjusts the parameters of encoders 104A / 104B and 106A / 106B and repeats the training until the reproduced multispectral image 116A / 116B matches the multispectral image 102A / 102B to a certain extent (e.g., with an acceptable difference less than a threshold). For example, this difference can be determined by a loss function that measures the information loss due to image reconstruction and compares this information loss to a loss function threshold. The type of loss function used in training a spectral CT system with a disentangled representation can vary depending on the specific goals of the model and the characteristics of the data. However, some examples may include, but are not limited to, the mean squared error (MSE) loss, which measures the mean squared difference between the estimated value and the actual value. In the context of this system, it can quantify the pixel-by-pixel difference between the reconstructed multispectral image and the original multispectral image. Another example can be the mean absolute error (MAE) loss, which measures the mean absolute difference between the predicted value and the actual value. This MAE loss is less sensitive to outliers than the MSE, which may be beneficial if the training data contains anomalies. This process results in accurately training the encoder 106A / 106B, which acts as a disentangled model. To facilitate this unsupervised training, each iteration can use a different (i.e., new) set of input multispectral images to update the encoder model until the encoder model accurately reproduces the original input image from the latent space information.
[0045] In the second stage, a target image prediction model (i.e., decoder) can be trained in a supervised manner to predict a target (i.e., desired) image. For example, after training the latent space encoder, as described above, the target image prediction model is trained to generate a target image by combining two or more anatomical latent space information 108A / 108B into average anatomical latent space information 118 (e.g., averaging, median filtering, or other means of combining anatomical latent space information), passing these images to a decoder 120 that predicts target images 124A / 124B, and comparing the predicted target images 124A / 124B to known reference target images (not shown) associated with the multispectral images 102A / 102B. As described above, these target images can include VMI images, VNC images, material maps, downstream segmentations, etc. The method adjusts the decoder weights of the prediction model based on the above comparison. This process may be repeated for various labeled pairs of anatomical latent space information 108A / 108B and reference target images associated with the source multispectral images 102A / 102B until the target image 124A / 124B predicted by the decoder 120 matches the reference target image associated with the multispectral image.
[0046] The result of training the disentanglement model in the first phase and the prediction model in the second phase is that the models can be deployed as a single, integrated model that receives a new multispectral image, accurately converts the multispectral image into anatomical latent space information 108A / 108B (compressed image), and then uses the anatomical latent space information to accurately predict the target image 124A / 124B. In other words, the integrated model can perform predictions based on the simplified anatomical latent space information 108A / 108B (compressed image) rather than the more complex source multispectral image 102A / 102B. This not only produces more accurate predictions, but also reduces resource consumption and computation time when making predictions, as the amount of data present in the anatomical latent space information 108A / 108B is significantly reduced compared to the data in the multispectral image 102A / 102B, while maintaining sufficient information for making accurate predictions.
[0047] The disentanglement model, the prediction model, and / or the synthesis model can be a neural network such as a convolutional neural network (CNN). For example, the disentanglement model can be a neural network in which a multispectral image is input to the CNN, disentangled into anatomical structure and contrast latent space information, and then the information is combined to reproduce the multispectral image. During training, the weights of the CNN are then adjusted by comparing the original multispectral image with the reproduced multispectral image. Similarly, the prediction model can be a neural network in which anatomical structure latent space information is input to the CNN and used to predict the target image. During training, the weights of the CNN are then adjusted by comparing the known target image with the generated target image. The synthesis model can be a synthesis neural network that combines both the disentanglement CNN and the prediction CNN into a synthesis CNN, in which the multispectral image is input to the synthesis CNN and disentangled into anatomical structure latent space information for predicting the target image.
[0048] Note that, depending on the accuracy of the trained model, training can be repeated as needed. In other words, the model drift of the trained model over time can be monitored. If the model drift exceeds the acceptable prediction accuracy, the model can be retrained on new multispectral images to increase accuracy before updating and deploying.
[0049] Figure 2 A flowchart 200 for training a disentanglement model is shown. In step 202, a multispectral image 202 is input to the disentanglement model. As described above, the multispectral image 202 can be an image of human anatomy (e.g., a CT image) captured at two or more energy levels (e.g., a low energy level and a high energy level). In steps 204 and 206, the disentanglement model performs encoding of the multispectral image 202. Encoding includes generating anatomical latent space information representing compressed anatomical information in the multispectral image, and generating contrast latent space information representing compressed contrast information in the multispectral image. The encoded image is then input to a selector in step 208. The selector can be a random selector that selects common anatomical features from the anatomical latent space information from any energy level. The selector within the disentanglement model plays a role in identifying and isolating common anatomical features from the anatomical latent space information. These common features are shared structural elements that exist across the multispectral images, regardless of the energy level at which each image was captured. The selector operates by analyzing the anatomical latent space information that has been encoded to represent the core anatomical information, and then it identifies those features that remain constant across different spectral frequencies. This process can be thought of as a filter that sifts through the anatomical information, pinpointing consistent elements that define the patient's anatomy.
[0050] Once the common anatomical features are selected, they serve as a stable foundation upon which contrast information can be overlaid. The selector can employ a variety of techniques to achieve this, such as statistical analysis to determine feature consistency. The output of the selector is a refined representation of the patient's anatomy, removing spectral variation while preserving structural detail. This refined anatomical information is then concatenated with contrast latent space information, which contains energy-related attenuation information, to reconstruct a multispectral image or synthesize a new image that can be used for further analysis and diagnosis. In step 210, the common features output by the selector are then combined with the contrast latent space information. The images can be combined by separately concatenating (e.g., overlaying) the contrast latent space information with the common features output by the selector to produce two or more combined images that are input to corresponding decoders. In step 212, the decoder decodes the concatenated images in an attempt to reproduce the original multispectral image. In step 214, the method determines whether the decoder correctly reproduced the original multispectral image. If the image is accurately reproduced, the process moves to step 216, where it is determined whether the disentanglement model was accurately trained. If not, the process is repeated for a new set of multispectral images. The disentanglement model can be a neural network with weights that are adjusted until the decoder can correctly reproduce the original multispectral image from the latent space information.
[0051] After accurately training the entanglement model, the system then trains a prediction model for predicting the desired target image. Figure 3 A flowchart 300 for training a prediction model is shown. In step 302, the system inputs anatomical latent space information (compressed image) into the target image prediction model. Optionally, the anatomical latent space information 108A / 108B (compressed image) may be preprocessed (e.g., averaged, etc.) before being input into the target image prediction model. In either case, in step 304, the system compares the predictions of the target image prediction model with a known reference target image corresponding to the multispectral image. In other words, the reference target image corresponding to the multispectral image is known and can be compared with the predictions made by the target image prediction model. The target image prediction model may be a neural network with weights that are adjusted in step 306 based on the comparison between the labeled datasets. In step 308, the system determines whether training is complete. If training is not complete (i.e., the predictions do not reach the desired level of accuracy), the process is repeated for a new pair of labeled data. This process is repeated until training is complete, and in step 310, the system deploys the trained disentanglement model and target image prediction model. For ease of deployment, the trained disentanglement model and target image prediction model can be combined into a single comprehensive model that performs disentanglement and then performs target image prediction based on the disentangled image.
[0052] As described above, the anatomical latent space can be learned in an unsupervised manner, significantly reducing the amount of required training data. Furthermore, utilizing a common anatomical latent space as input to the output decoder simplifies the training of different types of predictive decoder algorithms without changing the rest of the network. In other words, anatomical disentanglement can be generalized, but different predictive models can be developed for different types of target images. These predictive decoders can then be used in a "plug-and-play" fashion based on the target task. For example, specialized decoders for specific types of medical prediction / diagnosis can be developed in a similar manner and then deployed as needed.
[0053] In other words, the training process for the spectral computed tomography system with disentangled representations is designed to be modular, allowing the anatomical latent space to be learned in an unsupervised manner without changing other components of the network. This modularity is achieved by first focusing on disentangling anatomical and contrast features into independent latent spaces. The anatomical encoder is trained to capture invariant anatomical structures across different energy levels, thereby creating a universal representation of the patient's anatomy. This is done without supervision, meaning that the system does not require labeled data indicating the correct output of the anatomical latent space. On the other hand, the contrast encoder learns energy-related aspects of the image that vary with the X-ray spectrum.
[0054] Once the anatomical latent space is established, it can be used as a consistent input for a variety of prediction models, each designed for a different target image, such as a virtual monoenergetic image, a material map, or a segmentation map. These prediction models are trained in a supervised manner, where they learn to map the anatomical latent space to the desired output. The advantage of this approach is that the anatomical latent space serves as a common basis for all prediction models, which means that when new prediction tasks arise or new target image types are introduced, existing anatomical encoders do not need to be retrained. Instead, new prediction models are trained to work with the already established anatomical latent space. This "plug-and-play" capability streamlines the training process because it allows new prediction tasks to be added with minimal adjustments to the overall network architecture, saving time and computational resources.
[0055] After the model is appropriately trained and deployed, it can be used to perform target image prediction based on a new set of multispectral anatomical images. Figure 4A flowchart 400 for executing a training model for spectral computed tomography with a disentangled representation is shown. For example, in step 402, the system can capture or retrieve new multispectral medical images. These images can be captured by a medical device (e.g., a CT device) and input to the model or retrieved from the memory of the user device or a database of medical images. In either case, in step 404, the new image can be input to the trained disentangled model to generate anatomical structure latent space information. In step 406, the anatomical structure latent space information is then input to the trained target image prediction model. The trained image prediction model can then predict the target image in step 408.
[0056] Figure 5 An example image 500 is illustrated, which demonstrates the process of spectral computed tomography with a disentangled representation. Image 502 is the original input multispectral image obtained from a CT machine at varying energy levels. These images capture the patient's anatomical details and the energy-dependent attenuation characteristics of the tissue. After the disentanglement process, image 504 represents the resulting anatomical latent space information, which can be a compressed representation of the patient's anatomy, removing energy-specific contrast information. These images highlight anatomical features that are consistent at different energy levels, such as the outlines of organs and bones. The anatomical latent space information is then used to predict target reference images 506, which are shown as material images, but can be synthetic images that can include virtual monoenergetic images or other clinically relevant visualizations. These target images are generated by decoding the anatomical latent space information to produce a detailed and diagnostically useful representation of the patient's internal structures.
[0057] Figure 6 A block diagram 600 of a hardware device for training and performing spectral computed tomography with disentangled representations is shown. It should be understood that Figure 6 The components of system 600 shown in and described herein are examples only, and systems having additional, alternative, or fewer components are considered to be within the scope of the present invention.
[0058] As shown, system 600 includes at least one end-user device 602, a server 604, a database 606, and a medical imaging device 610 interconnected by a network 608. In the illustrated example, server 604 supports the operation (e.g., training, deployment, and execution) of spectral computed tomography with the detangled representation solution described herein. In the illustrated example, user device 602 is a PC, but it can be any device (e.g., a smart phone, tablet, etc.) that provides access to server 604 and database 606 via network 610. User device 602 has a user interface UI that can be used to communicate with the server and database via a browser or via a software application using network 610. For example, user device 602 can allow a user to access a trained model executed on server 604, and server 604 can be used to train and deploy these models. Network 608 can be the Internet and / or other public or private networks or a combination thereof. Therefore, network 608 should be understood to include any type of circuit-switched network, packet-switched network, or a combination thereof. Non-limiting examples of the network 608 may include a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), and the like.
[0059] In one example, an end-user device 602 can communicate with a server 604 via a software application to access the models disclosed herein. The software application can initiate the server 604 to execute the trained model. For example, the server 604 can receive medical images from one or more of the user device 602, the database 606, and the medical device 610. The server 604 can then train a detachment model and a predictive model based on these images. The server 604 can then deploy the trained model and allow the user to access and execute the trained model via the user device 602 to obtain new medical images.
[0060] For ease of illustration, devices 602, 604, 606, and 610 are each depicted as a single device, but one of ordinary skill in the art will understand that for different implementations, devices 602, 604, 606, and 610 may be embodied in different forms. For example, any one or each of the servers may include multiple servers including multiple databases, etc. Alternatively, the operations performed by any one of the servers may be performed on fewer (e.g., one or two) servers. In another example, multiple user devices (not shown) may communicate with the server. In addition, a single user may have multiple user devices (not shown), and / or there may be multiple users (not shown), each with their own corresponding user device (not shown). In any case, Figure 6 The hardware configuration shown in can be a system that supports the functionality of model training and execution disclosed in this article.
[0061] Figure 7 A block diagram of components for training and executing a hardware device for spectral computed tomography with a disentangled representation is shown. System 700 may represent at least a portion of each of PC 602, server 604, database 606, and medical imaging system 610. One or more components of system 700 may communicate electrically with each other using bus 705. System 700 may include a processing unit (CPU or processor) 710 and a system bus 705 that couples various system components, including system memory 715 (such as read-only memory (ROM) 720 and random access memory (RAM) 725), to processor 710. System 700 may include a cache of high-speed memory directly connected to, in close proximity to, or integrated as part of processor 710. System 700 may copy data from memory 715 and / or storage device 730 to cache 712 for rapid access by processor 710. In this way, cache 712 may provide a performance boost by preventing processor 710 from being delayed while waiting for data. These modules and other modules can control or be configured to control the processor 710 to perform various actions. Other system memory 715 may also be available. Memory 715 may include a variety of different types of memory with different performance characteristics. Processor 710 may include any general-purpose processor and hardware modules or software modules (such as service 1 732, service 2 734, and service 3 736 stored in storage device 730 and configured to control processor 710), as well as special-purpose processors in which software instructions are incorporated into the actual processor design. Processor 710 can essentially be a completely independent computing system that contains multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.
[0062] In order to enable user interaction with the computing system 700, the input device 745 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphic input, a keyboard, a mouse, motion input, voice, etc. The output device 735 can also be one or more of a plurality of output mechanisms known to those skilled in the art. In some examples, a multimodal system can enable a user to provide multiple types of input to communicate with the computing system 700. The communication interface 740 can generally govern and manage user input and system output. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features herein can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0063] The storage device 730 may be a non-volatile memory and may be a hard disk or other type of computer-readable medium that can store data that can be accessed by a computer, such as a magnetic tape cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, random access memory (RAM) 725, read-only memory (ROM) 720, and mixtures thereof.
[0064] The storage device 730 may include services 732, 734, and 736 for controlling the processor 710. Other hardware or software modules are contemplated. The storage device 730 may be connected to the system bus 705. In one aspect, a hardware module that performs a particular function may include a software component for performing that function stored in a computer-readable medium connected to the necessary hardware components (such as the processor 710, the bus 705, the output device 735, etc.).
[0065] Although the foregoing relates to the exemplary embodiments described herein, other and further exemplary embodiments may be envisioned without departing from the basic scope of the present invention. For example, various aspects of the present disclosure may be implemented in hardware or software or a combination of hardware and software. An exemplary embodiment described herein may be implemented as a program product for use with a computer system. The program of the program product defines the functions of the exemplary embodiments (including the methods described herein) and may be contained on various computer-readable storage media. Exemplary computer-readable storage media include, but are not limited to: (i) non-writable storage media (e.g., a read-only memory (ROM) device within a computer, such as a CD-ROM disk that can be read by a CD-ROM drive, flash memory, ROM chip, or any type of solid-state non-volatile memory), on which information is permanently stored; and (ii) writable storage media (e.g., a floppy disk within a floppy disk drive or hard drive, or any type of solid-state random access memory), on which variable information is stored. When carrying computer-readable instructions that direct the functions of the disclosed exemplary embodiments, such computer-readable storage media are exemplary embodiments of the present disclosure.
[0066] Those skilled in the art will appreciate that the foregoing examples are illustrative and not restrictive. All permutations, enhancements, equivalents, and improvements to the present disclosure will be apparent to those skilled in the art upon reading the specification and studying the drawings, and are intended to be encompassed within the true spirit and scope of the present disclosure. Therefore, the appended claims are intended to encompass all such modifications, permutations, and equivalents that fall within the true spirit and scope of these teachings.
Claims
1. A system for generating a target image, the system comprising: Database of multispectral images; and A processor configured to: The latent space encoder is trained to generate latent space information by the following operations: encoding the multispectral image to generate anatomical structure latent space information and contrast latent space information, wherein the anatomical structure latent space information represents compressed anatomical structure information in the multispectral image and the contrast latent space information represents compressed contrast information in the multispectral image; combining selected features of the anatomical latent space information and the contrast latent space information to reconstruct the multispectral image, comparing the reproduced multispectral image with the multispectral image, and adjusting the encoding of the multispectral image and repeating the training until the reproduced multispectral image matches the multispectral image, and After training the latent space encoder, the output decoder is trained to generate the target image by the following operations: Inputting the anatomical structure latent space information into the prediction model, Predicting a target image using the prediction model based on the anatomical structure latent space information, comparing the predicted target image with a reference target image associated with the multispectral image, and The weights of the prediction model are adjusted based on the comparison, and the training is repeated until the image predicted by the prediction model matches the reference target image associated with the multispectral image.
2. The system of claim 1 , wherein the processor is configured to generate the latent space information by applying a convolutional neural network to the multispectral image to extract and separate anatomical structure information and contrast information from the multispectral image. 3 . The system of claim 1 , wherein the processor is configured to select an anatomical feature as a common anatomical feature between the anatomical latent space information. 4 . The system of claim 1 , wherein the processor is configured to combine the anatomical structure latent space information and the contrast latent space information by concatenating the anatomical structure latent space information and the contrast latent space information.
5. The system of claim 1 , wherein the processor is configured to compare the reproduced multispectral image with the multispectral image by calculating a loss function of the reproduced multispectral image compared to the multispectral image and repeating the training until the loss function is less than a loss function threshold.
6. The system of claim 1, wherein the multispectral image is an image of an anatomical structure captured by a medical imaging device from a common reference frame relative to the anatomical structure and operating at different spectral frequencies. 7 . The system of claim 1 , wherein the processor is configured to average the anatomical latent space information before inputting it into the prediction model.
8. The system of claim 1 , wherein the prediction model is a neural network that predicts the target image and compares the predicted target image to a reference target image associated with the multispectral image in a supervised manner to determine corrective measures to adjust the weights of the prediction model, the weights being neural network weights.
9. The system of claim 1, wherein the target image is a medical image related to a human anatomical state and comprises one or more of a virtual monoenergetic image, a virtual non-contrast image, a material map, and a downstream segmentation.
10. The system of claim 1, wherein the processor is configured to generate new latent space information for a new multispectral image using the trained latent space encoder, and input the new latent space information into the trained output decoder to generate a new target image.
11. A method for generating a target image, the method comprising: The latent space encoder is trained by the processor to generate latent space information through the following operations: encoding a multispectral image to generate anatomical structure latent space information and contrast latent space information, wherein the anatomical structure latent space information represents compressed anatomical structure information in the multispectral image, and the contrast latent space information represents compressed contrast information in the multispectral image, combining selected features of the anatomical latent space information and the contrast latent space information to reconstruct the multispectral image, comparing the reproduced multispectral image with the multispectral image, and adjusting the encoding of the multispectral image and repeating the training until the reproduced multispectral image matches the multispectral image, and After training the latent space encoder, the processor trains an output decoder to generate a target image by: Inputting the anatomical structure latent space information into the prediction model, Predicting a target image using the prediction model based on the anatomical structure latent space information, comparing the predicted target image with a reference target image associated with the multispectral image, and The weights of the prediction model are adjusted based on the comparison, and the training is repeated until the image predicted by the prediction model matches the reference target image associated with the multispectral image. 12 . The method of claim 11 , generating the latent space information by the processor by performing convolution on the multispectral image to extract and separate anatomical structure information and contrast information from the multispectral image. 13 . The method of claim 11 , wherein the processor selects an anatomical feature as a common anatomical feature among the anatomical latent space information. 14 . The method of claim 11 , combining, by the processor, the anatomical structure latent space information and the contrast latent space information by concatenating the anatomical structure latent space information and the contrast latent space information.
15. The method of claim 11, wherein the processor compares the reproduced multispectral image with the multispectral image by calculating a loss function of the reproduced multispectral image compared to the multispectral image and repeating the training until the loss function is less than a loss function threshold.
16. The method of claim 11, wherein the multispectral image is an image of the anatomical structure captured by a medical imaging device from a common frame of reference relative to the anatomical structure and operating at different spectral frequencies. 17 . The method of claim 11 , wherein the processor averages the anatomical structure latent space information before inputting the anatomical structure latent space information into the prediction model.
18. The method of claim 11, wherein the prediction model is a neural network that predicts the target image and compares the predicted target image with a reference target image associated with the multispectral image in a supervised manner to determine corrective measures to adjust the weights of the prediction model, the weights being neural network weights.
19. The method of claim 11, wherein the target image is a medical image related to a human anatomical state and comprises one or more of a virtual monoenergetic image, a virtual non-contrast image, a material map, and a downstream segmentation.
20. The method of claim 11, wherein the processor utilizes the trained latent space encoder to generate new latent space information for a new multispectral image, and inputs the new latent space information into the trained output decoder to generate a new target image.