Patient-specific 3D medical image reconstruction from 2d medical images using a vision transformer-based machine learning model

A machine learning model with hierarchical vision transformer blocks reconstructs 3D medical images from 2D kV images, addressing setup errors in radiotherapy by providing real-time, accurate patient alignment and adaptive therapy.

WO2025184512A1PCT designated stage Publication Date: 2025-09-04MAYO FOUNDATION FOR MEDICAL EDUCATION & RESEARCH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017852
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Current methods for patient alignment in radiotherapy, such as 3D on-the-board imaging and cone beam computed tomography, suffer from high imaging dose, limited field-of-view, and positional uncertainties, especially when tumors are obscured by high-density structures, leading to setup errors.

Method used

A machine learning model using a dual-model framework with hierarchical vision transformer blocks is employed to reconstruct 3D medical images from 2D kV images, enabling real-time, zero-dose patient alignment by utilizing clinically available images and accounting for anatomical changes.

Benefits of technology

The method achieves high-accuracy, real-time 3D image reconstruction, improving patient alignment accuracy and enabling online adaptive radiation therapy, reducing setup errors and enhancing treatment effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017852_04092025_PF_FP_ABST
    Figure US2025017852_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Three-dimensional (3D) medical images are generated from two-dimensional (2D) medical images using a vision transformer-based machine learning model. 2D image data is accessed with a computer system. The 2D image data may include 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system. A machine learning model trained on training data to reconstruct 3D medical images from 2D medical images is accessed with the computer system. The 2D image data is input to the machine learning model using the computer system, generating 3D image data as an output. The 3D image data may then be output via the computer system.
Need to check novelty before this filing date? Find Prior Art

Description

PATIENT-SPECIFIC 3D MEDICAL IMAGE RECONSTRUCTION FROM 2D MEDICAL IMAGES USING A VISION TRANSFORMER-BASED MACHINE LEARNING MODELCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 559,453, filed on February 29, 2024, and entitled “PATIENT-SPECIFIC 3D MEDICAL IMAGE RECONSTRUCTION FROM 2D MEDICAL IMAGES USING A VISION TRANSFORMER-BASED MACHINE LEARNING MODEL,” which is herein incorporated by reference in its entirety.BACKGROUND

[0002] In radiotherapy, image-guided patient alignment enables accurate treatment plan delivery. Methods used for patient alignment include 3D on-the-board imaging (OBI) or acquiring 2D kV images taken at fixed, oblique angles. The visibility of tumor in 2D kV images is limited since the patient’s anatomy is projected onto a 2D plane, especially when the tumor is behind high-density structures such as bones. This can lead to significant patient setup errors. Using 3D OBI (e.g., cone beam computed tomography (CBCT), CT-on-rails (CToR)) also has its drawbacks. CBCT can impart an unnecessarily high imaging dose that is unfavorable for pediatric patients, its field-of-view is limited, and it can suffer from artifacts that can complicate tumor visualization. CToR has the added drawback that the patient must be transferred from the CT scanner to the treatment position, which may introduce positional uncertainties.SUMMARY OF THE DISCLOSURE

[0003] The present disclosure addresses the aforementioned drawbacks by providing a method for generating three-dimensional (3D) medical images from two-dimensional (2D) medical images. The method includes accessing 2D image data with a computer system. As a non-limiting example, the 2D image data include 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system. A machine learning model is accessed with the computer system. The machine learning model has been trained on training data to reconstruct 3D medical images from 2D medical images. The 2Dimage data are then input to the machine learning model using the computer system, generating 3D image data as an output. The 3D image data may then be output via the computer system.

[0004] It is another aspect of the present disclosure to provide a method for generating 3D medical images from 2D medical images. The method includes accessing 2D image data with a computer system. As a non-limiting example, the 2D image data include 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system. An autoencoder network is accessed with the computer system. The autoencoder network has been trained on training data to reconstruct 3D medical images from 2D medical images. As a non-limiting example, the autoencoder network is an asymmetric autoencoder network that includes vision transformer blocks, where each vision transformer block implements a window-based self-attention. The 2D image data are input to the autoencoder network using the computer system, generating 3D image data as an output. The 3D image data may then be output via the computer system

[0005] It is yet another aspect of the present disclosure to provide a method for generating 3D medical images from 2D medical images. The method includes accessing 2D image data with a computer system. As a non-limiting example, the 2D image data include 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system. First and second machine learning models are accessed with the computer system. The first machine learning model has been trained on first training data to reconstruct an image of a first anatomical region, and the second machine learning model has been trained on second training data to reconstruct an image of a second anatomical region that is different than the first anatomical region. The 2D image data are input to the first machine learning model, generating first output image data that depict the first anatomical region. The 2D image data are also input to the second machine learning model, generating second output image data that depict the second anatomical region. 3D image data are then generated by concatenating the first output image data and the second output image data. The 3D image data may then be output via the computer system.BRIEF DESCRIPTION OF THE DRAWINGS

[0006] FIG. 1A shows an example architecture for a machine learning model framework for reconstructing 3D CT images from 2D kV images.

[0007] FIG. IB shows an example of a hierarchical vision transformer block in the encoder portion of the machine learning model framework of FIG. 1A.

[0008] FIG. 1C shows an example of a hierarchical vision transformer block in the decoder portion of the machine learning model framework of FIG. 1A.

[0009] FIG. ID shows an example of a detailed illustration of a window-based multihead attention (W-MHA). Tokenized patches are first split to nW non-overlapping windows of size wxw and the attention is calculated on the windows instead of the whole inputs.

[0010] FIG. 2 is a flowchart setting forth the steps of an example method for reconstructing 3D CT images from 2D kV images using a suitably trained machine learning model.

[0011] FIG. 3 is a flowchart setting forth the steps of an example method for training a machine learning model to reconstruct 3D CT images from 2D kV images.

[0012] FIG. 4 shows an example of coordinate conversion between planning CT, CToR, and kV images.

[0013] FIG. 5 shows an example layout of an image-guided radiation treatment system showing the geometrical relationship between component of the system.

[0014] FIG. 6 is a histogram of shift error (SE) (in mm) from 10 patients in an example study. The y-axis shows the bin edges for the histogram. It can be observed that the majority of the SEs are less than 0.4 mm, which is far smaller than a clinically acceptable patient alignment tolerance for H&N patients, w hich is often in the range of 2-3 mm.

[0015] FIG. 7 shows the CT number absolute difference volume histogram from an example study. The x-axis represents the HU number difference between the ground -truth CT (gCT) and synthesized CT (sCT) and the y-axis represents the percentage of the volume.

[0016] FIG. 8 shows the comparison of one CT slice between the gCT and the corresponding sCT from an example study.

[0017] FIG. 9 shows a dose profile comparison between doses calculated on the sCT and the doses calculated on the gCT in both R-L and A-P directions.

[0001] FIG. 10 is a block diagram of an example system for reconstructing 3D CT images from 2D kV images.

[0002] FIG. 11 is a block diagram of example components that can implement the system of FIG. 10.DETAILED DESCRIPTION

[0003] Described here are systems and methods for reconstructing a 3D CT image from 2D kV images obtained at the treatment position for a patient. In general, a machine learningmodel is trained on training data to reconstruct 3D CT images from clinical 2D kV x-ray images. Using the disclosed systems and methods. 3D images may be obtained in milliseconds at the position of treatment, enabling better patient alignment, real-time adaptive re-planning, and better visibility of tumors. The real-time CT reconstruction with limited projections enabled by the disclosed systems and methods can also benefit photon therapy in low-income, rural areas by providing an alternative and affordable solution where many photon machines lack 3D OBI capability.

[0004] As described in more detail below, the machine learning model used to reconstruct 3D CT images from 2D kV images implements a dual-model framework built with vision transformer blocks, such as hierarchical vision transformer blocks. This dual-model framework receives 2D kV images as the sole input and can synthesize accurate, full-size 3D CT images in real time (e.g., within milliseconds), which can be used for reflecting the realtime patient position in 3D, thus achieving high-quality but “zero-dose” image-guided patient alignment.

[0005] As an advantage, the systems and methods described in the present disclosure can utilize only clinically available images (e.g., daily kV images) to reconstruct real-time, ready-to-use, full-size 3D CT images, thereby enabling online adaptive radiation therapy (ART).

[0006] The machine learning model described herein adopts and adapts a hierarchical vision transformer to the medical images (i.e., kV images and CT images) with a dual-model setting and a geometry property reserved shifting and sampling (GRSS) data augmentation strategy. The GRSS data augmentation takes the geometrical relation between the treatment couch and kV imaging source and detector into consideration. This data augmentation strategy enables the trained machine learning model to take full advantage of noisy, but sparse 2D kV images to fulfill accurate 3D CT image reconstruction while avoiding model overfitting.

[0007] It is another advantage of the disclosed systems and methods to provide a radiation treatment planning workflow that implements an Al-assisted model to account for inter-fraction anatomical changes, which improves patient alignment accuracy and significantly improves the accuracy, effectiveness, and robustness of online adaptive radiation therapy. As a benefit, a patient’s post-treatment quality7of life can be significantly improved. The more accurate anatomical changes determined using the systems and methods described in the present disclosure can also be used to provide robust plan re-optimization for online ART that takes the real-time anatomical changes into account.

[0008] Although the disclosed machine learning models are described with respect to reconstructing 3D CT images from 2D kV images, it will be appreciated that the models can be adapted to other medical imaging modalities, such as magnetic resonance imaging (MRI).

[0009] The machine learning model framework used by the disclosed systems and methods implements dual models. Advantageously, the disclosed dual-model framework enables vision transformers to be adapted to medical images. The primary model may be dedicated to identifying the positions of structures of interest and the secondary model may be focused on reasoning the 2D-3D relations and reconstructing the voxel-level fine details in 3D CT. As another advantage, the dual-model framework is resource-efficient and can be generalized to other medical imaging modalities, such as MRI, PET, etc.

[0010] Each of the dual models has an asymmetric autoencoder-like architecture, including an encoder and a decoder with hierarchical vision transformer blocks as the basic building blocks. An example overall architecture is shown in FIG. 1A. The kV images are simultaneously input to dual models (i.e., primary model and secondary model) to generate the whole CT and CT that covered only a specific anatomical location or region (e.g., the head region), respectively. The full-size synthesized CT is generated by overlaying and concatenating the outputs from the tw o models according to their spatial relationship.

[0011] Both models in the dual-model framework may include a patch embedding layer (e.g., a convolutional layer), an encoder Ek, a decoder Dr. and a fully connected layer. The patch embedding layer is used for projecting non-overlapping raw kV image patches to initial high-dimensional feature representations that serves as the input for the encoder Ek. Both the encoder network, Ek, (e.g., the encoder network shown in FIG. IB) and the decoder network, Dr, (e.g., the decoder network shown in FIG. 1C) may include multiple hierarchical vision transformer blocks, having a pattern of layer normalization, followed by window-based multihead attention (W-MHA), followed by layer normalization, followed by multilayer perceptron (MLP), and followed by a patch merging layer (in the encoder network) or a patch unmerging layer (in the decoder network). The W-MHA (an example of which is illustrated in FIG. ID) calculates the attention within the windows only instead of the entire image, which greatly reduces the computational complexity. The patch merging layer in the encoder network concatenates nearby patches (e.g., 2 x 2 patches) with a linear merging layer to obtain a hierarchical representation. Likewise, the unmerging layer in the decoder network enlarges each patch by a factor (e.g., by a factor of 2) along each dimension through a fully connectedlayer. The fully connected layer converts from the learned representations to the final output (i. e. , the 3D sCT).

[0012] By way of example, the raw input kV images are first split into non-overlapping patches, each with a size HxH. Then, a patch embedding layer projects each patch to an arbitrary dimension, C. Next, several transformer blocks with localized self-attention (e.g., window-based self-attention) are applied to these tokenized patches. A detailed illustration of an example window-based self-attention is shown in FIG. ID.

[0013] The patches are merged by concatenating nearby 2x2 patches with a linear merging layer after each transformer block to get the hierarchical representations. Finally, the encoder Ek (FIG. IB) with N transformer blocks (e.g., with N = 4 as a non-limiting example) converts the raw kV images to a latent representation of size H'xH', which can be used as the input for the decoder, Dr(FIG. 1C). The decoder also includes transformer blocks with window -based self-attention. The number of vision transformer blocks in the decoder depends on the output 3D image size. Instead of patches merging in the decoder, the patches are enlarged by a factor of two along each dimension through a fully connected layer after each transformer block.

[0014] Lastly, tw o fully connected layers are applied to the output of the decoder to yield the reconstructed CT images Xc.

[0015] The details of the window-based multi-head attention is shown in FIG. ID. The attention is calculated on the windows only with a computational complexity of Q(W - MHA) = 4H2C2+ 2W2HC. where H represents the raw7image size along x and y directions, C is the number of channels and w is the size of each window. In comparison, the computational complexity of global attention is Q(G - MHA) = 4H2C2+2H2C. When the raw image size is large, the computation for global attention is generally unaffordable.

[0016] Referring now to FIG. 2, a flowchart is illustrated as setting forth the steps of an example method for generating classified feature data using a suitably trained neural network or other machine learning algorithm. As will be described, the neural network or other machine learning algorithm takes 2D image data as input data and generates 3D image data as output data.

[0017] The method includes accessing 2D kV image data with a computer system, as indicated at step 202. Accessing the 2D kV image data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the 2D kV image data may include acquiring such data with an x-ray imaging systemand transferring or otherwise communicating the data to the computer system, which may be a part of the x-ray imaging system, radiation therapy system, or the like.

[0018] In general, the 2D kV image data includes 2D kV images acquired from a subject. The 2D kV image data may include at least two images of the subject. These images may be acquired, for example, at two orthogonal angles (i.e., in two orthogonal imaging planes).

[0019] A trained machine learning model is then accessed with the computer system, as indicated at step 204. For example, the machine learning model illustrated in FIG. 1A may be accessed with the computer system. In general, the machine learning model is trained, or has been trained, on training data in order to reconstruct 3D CT images from 2D kV images.

[0020] Accessing the trained machine learning model may include accessing network parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the machine learning model on training data. In some instances, accessing the machine learning model can also include retrieving, constructing, or otherwise accessing the particular machine learning model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.

[0021] An artificial neural network generally includes an input layer, one or more hidden layers (or nodes), and an output layer. Typically, the input layer includes as many nodes as inputs provided to the artificial neural network. The number (and the type) of inputs provided to the artificial neural network may vary based on the particular task for the artificial neural network.

[0022] The input layer connects to one or more hidden layers. The number of hidden layers varies and may depend on the particular task for the artificial neural network. Additionally, each hidden layer may have a different number of nodes and may be connected to the next layer differently. For example, each node of the input layer may be connected to each node of the first hidden layer. The connection between each node of the input layer and each node of the first hidden layer may be assigned a weight parameter. Additionally, each node of the neural network may also be assigned a bias value. In some configurations, each node of the first hidden layer may not be connected to each node of the second hidden layer. That is, there may be some nodes of the first hidden layer that are not connected to all of the nodes of the second hidden layer. The connections between the nodes of the first hidden layersand the second hidden layers are each assigned different weight parameters. Each node of the hidden layer is generally associated with an activation function. The activation function defines how the hidden layer is to process the input received from the input layer or from a previous input or hidden layer. These activation functions may vary’ and be based on the type of task associated with the artificial neural network and also on the specific type of hidden layer implemented.

[0023] Each hidden layer may perform a different function. For example, some hidden layers can be convolutional hidden layers which can, in some instances, reduce the dimensionality of the inputs. Other hidden layers can perform statistical functions such as max pooling, which may reduce a group of inputs to the maximum value; an averaging layer; batch normalization; and other such functions. In some of the hidden layers each node is connected to each node of the next hidden layer, which may be referred to then as dense layers. Some neural networks including more than, for example, three hidden layers may be considered deep neural networks.

[0024] The last hidden layer in the artificial neural network is connected to the output layer. Similar to the input layer, the output layer typically has the same number of nodes as the possible outputs.

[0025] The 2D kV image data are then input to the machine learning model, generating output as 3D CT image data, as indicated at step 206.

[0026] The 3D CT image data generated by inputting the 2D kV image data to the trained machine learning model can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 208.

[0027] Referring now to FIG. 3. a flowchart is illustrated as setting forth the steps of an example method for training a machine learning model on training data, such that the machine learning model is trained to receive 2D image data as input data in order to generate 3D image data as output data.

[0028] In general, the machine learning model can implement any number of different model architectures. For instance, the machine learning model can implement the dual-model architecture illustrated in FIG. 1 A. Alternatively, the machine learning model could be replaced with other suitable machine learning or artificial intelligence algorithms, such as those based on otherwise implementing supervised learning, unsupervised learning, deep learning, ensemble learning, dimensionality reduction, and so on.

[0029] The method includes accessing training data with a computer system, as indicated at step 302. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with an imaging system (e.g., a CT system, an x-ray system) and transferring or otherwise communicating the data to the computer system.

[0030] In general, the training data can include 2D image data (e.g., 2D kV image data) and 3D image data (e.g., 3D CT image data). Advantageously, the machine learning model may be trained on 2D kV images and corresponding 3D images (e.g., CToR images) as the training and testing datasets, without referring to supplementary7images such as digital reconstructed radiography (DRR) images. The training data can include 2D image data and 3D image data acquired from two different anatomical regions. For instance, the training data can include 2D image data and 3D image data acquired from a first anatomical region in addition to 2D image data and 3D image data acquired from a second anatomical region in each of a plurality of subjects. The first anatomical region may correspond to a larger imaging volume than the second anatomical region. For example, the first anatomical region may correspond to a whole-body region whereas the second anatomical region may correspond to a more limited anatomical region (e.g., the head region, the head and neck regions).

[0031] The method can include assembling training data from 2D image data and 3D image data using a computer system. This step may include assembling the 2D and 3D image data into one or more appropriate data structures on which the neural network or other machine learning algorithm can be trained. Assembling the training data may include assembling 2D and / or 3D image data, segmented 2D and / or 3D image data, and other relevant data. For instance, assembling the training data may include generating augmented data and including the augmented data in the training data. Augmented data may include 2D and / or 3D image data that can been augmented using one or more data augmentation techniques, such as those described below.

[0032] In anon-limiting example, training data included full-size planning CT, CToR, and kV images taken from the same day as the CToR and the rigid registration (RGs) files involving the above-mentioned CT images and kV images from 10 head and neck (H&N). Each patient had at least three CToR of size 512 x 512 x N, where N is the length of CT slices, which varies from patient to patient, and the same-day two orthogonal kV images, each of size 1024 x 1024. The raw CT images and kV images were augmented, such as by synthesizing thedataset for training and validation with rigid image registration (RIR), cropping, and geometric property-reserved shifting and sampling augmentation.

[0033] Offline patient alignment is often conducted by registering both CToR and kV images to pre-acquired planning CT images for double verification. The resulting registration matrix can be stored in the RGs files mentioned above. Therefore, with the planning CT as the base and the isocenter of the planning CT as the origin of the coordinates, all pairs of kV images and CToR can be registered to the same coordinate, thereby simplifying the subsequent data processing without considering the conversion among different coordinates. Moreover, in the reference stage, the relative position of any new kV images can be readily derived. FIG. 4 shows an example coordinate conversion of the images.

[0034] When using a machine learning model with a dual-model framework, the images can be cropped to two different sizes for the primary model and the secondary model, respectively. For the primary’ model, the CT images can be cropped to size 384 x 336 x 448 to exclude the excessive background with low entropy (e.g., air) as well as forming a dataset with same-size samples. If the length of the CT slices is less than 448. the CT slices can first be transposed to shape 512 x 512 x N (N is the number of slices), followed by the cropping operation. Similarly, the kV images can be cropped to size 1008 x 1008. For the secondary model, the CT images can be cropped to size M x 224 x 224, where M indicates the number of minimum voxels that cover the head region along the right-left (R-L) direction, which may vary from patient to patient. The corresponding kV images can be cropped to size 1008 x 1008. the same as the size of kV images used for the primary model.

[0035] Additionally or alternatively, a geometric property -reserved shifting and sampling data augmentation strategy can be implemented to yield a sufficient number of samples for model training while preserv ing the geometrical property of the CT images and kV images. Given the layout of the kV imaging system in the treatment room (FIG. 5), it can be observed that the movement in the superior-inferior (S-I) direction of the CT is solely reflected in the same direction of the kV images, and the ratio is 1 : 1.5 given the position of the treatment table, the radiation source, and the receiver (i.e., image panel). Hence, the kV-CT pairs can be further augmented by simultaneously moving them along the S-I direction of the CT images within ±5 mm with a minimum step of 0. 1 mm (0. 15 mm for kV images).

[0036] Then, for the primary' model, the CT images can be resampled to 128 x 112 x 112 voxels with a step of 4 voxels in the S-I direction and a step of 3 voxels in the anterior- posterior A-P) and R-L directions. Correspondingly, the two orthogonal kV images can beresampled to 168 x 168 voxels with a step of 6 in both directions. Likewise, for the secondary model, the CT images can be resampled to M x 112 x 112 with a step of 2 voxels in both S-I and A-P directions, and kV images can be resampled to size 504 x 504 with a step of 2 in both directions. Thus, a single pair of CT images and kV images can yield 36 and 4 different smallsize pairs of data samples for the primary model dataset and secondary model dataset, respectively.

[0037] This GRSS data augmentation method helps to avoid overfitting issues due to the limited number of training samples. In addition, it allows for efficient model training since the size of each sample can be relatively small (e.g., less than 200 voxels) along any direction. Finally, a high-resolution CT of full size (e g., 512 x 512 x N), desirable for clinical applications, can be obtained by spatially stacking the small-size reconstructed CT images generated by both the primary model and secondary model. Moreover, the augmented CT-kV pairs are physically rational, potentially guaranteeing the reconstructed 3D CT images are clinically useable.

[0038] The machine learning model is then trained on the training data, as indicated at step 304. In general, the machine learning model can be trained by optimizing network parameters (e.g., weights, biases, or both) based on minimizing a loss function. As one nonlimiting example, the loss function may be a mean squared error loss function.

[0039] Training a neural network may include initializing the neural network, such as by computing, estimating, or otherwise selecting initial network parameters (e.g.. weights, biases, or both). During training, an artificial neural netw ork receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding weights. For instance, training data can be input to the initialized neural network, generating output as 3D image data. The artificial neural network then compares the generated output with the actual output of the training example in order to evaluate the quality of the 3D image data. For instance, the 3D image data can be passed to a loss function to compute an error. The current neural network can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current neural network can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the loss according to the loss function. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validationiterations being completed, and the like. When the training condition has been met (e.g.. by determining whether an error threshold or other stopping criterion has been satisfied), the current neural network and its associated network parameters represent the trained neural network. Different types of training processes can be used to adjust the bias values and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent, Newton’s method, conjugate gradient, quasi-Newton, Levenberg-Marquardt, among others.

[0040] The machine learning model can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transfer learning, or other suitable learning techniques for neural networks.

[0041] In an example implementation, distributed data parallel (DDP) can be employed to minimize memory7usage and significantly accelerate the training speed. AdamW can be used optimizer with |3i = 0.9 and P2 = 0.999 and a cosine annealing learning rate scheduler with an initial learning rate of e-7 and 20 warm-up epochs.

[0042] The machine learning model is then stored for later use, as indicated at step 306. Storing the machine learning model may include storing network parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the machine learning model on the training data. Storing the machine learning model ) may also include storing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.

[0043] An asymmetric autoencoder network built with vision-transformer blocks was developed. Primary and secondary models were trained, focusing on the reconstruction of the whole CT and head region, respectively. The following data was collected with head and neck cancer patients: 3 pairs of orthogonal kV images and the corresponding 3D CT-on-the-rails (CToRs) (i.e., the ground truth CT (gCT)) from different fractions. A GRSS data augmentation strategy was implemented to create more data samples, where the kV and gCT were simultaneously and arbitrarily shifted within ±5 mm. Moreover, the kV and gCT were geometry7property-reserved resampled to small-size data pairs. For the secondary model, the gCT was cropped to focus on the head region before the resampling. Finally, the full-size synthetic CT (sCT) was formed by concatenating the outputs from two models based on their spatial relationship. The image quality of the sCT was evaluated using mean absolute error(MAE) and per voxel absolute CT number difference volume histogram (CDVH). Gamma analysis of the CT numbers comparing the sCT to the gCT was done. Forward dose calculation was performed on the sCT and gCT using the same plan and compared using Gamma analysis as well. Dose volume histogram indices of targets were compared as well. The robustness of the sCT against random shifts was tested through simulations mimicking patient alignment during treatment, where the model predicted the shifted sCTs given manually shifted kVs as input. The predicted shifted sCTs were compared with all the gCTs with known shifts and the one with the minimum MAE was considered as the closest matched CT (mCT). In some examples, the difference of the associated shifts between the mCT and gCT was calculated as the shift error (SE).

[0044] To evaluate the robustness of the disclosed machine learning model framework to generate sCTs in the face of patient setup uncertainties, a comprehensive analysis of was conducted. Random shifts were applied to the kV images to simulate patient setup uncertainty. Given manually shifted kV images within ±4.5 mm as input, on one hand, the model predicted the shifted sCT (ssCT). On the other hand, the shifted gCT (sgCT) could be calculated based on the geometrical relation between the treatment couch and kV imaging system (see FIG. 5 for the detailed treatment room layout). To obtain the SE in this example, a searching pool S= sgCT±5, 8e[- 1, 1] was created, which included sgCT and its variances (shifting sgCT within ± 1mm with a step of 0.1 mm). The MAE between each candidate sgCT in S and ssCT was also calculated. By linear search, the sgCT with 6m that gave the minimum MAE was identified and the absolute value of 8m was defined as SE. The results of the calculated SEs are illustrated in FIG. 6. The model yielded a mean SE of only 0.40 ± 0.16 mm on average in the sCT robustness test mimicking daily clinic practice.

[0045] The model achieved a MAE of <30HU and the CDVH showed that <5% of the voxels had a per voxel absolute CT number difference larger than 128HU (FIG. 7). The profile of a typical gCT slice and its corresponding sCT slice exhibited a high agreement, indicating the high similarity between the gCT and sCT, as shown in FIG. 8. 3D Gamma passing rates of 0.999 and 0.998 with criteria of 3% / 3mm and 2% / 2mm were obtained in the CT number comparisons, respectively. 3D Gamma passing rates of 0.999 and 0.986 with criteria of 3% / 3mm and 2% / 2mm in the dose comparisons, respectively, were achieved. FIG. 9 depicts an example dose profile comparison between the doses calculated on the sCT and the dose calculated on the gCT in both R-L and A-P directions. It can be observed that the dose calculated on the sCT was very close to that calculated on the gCT for every treatment fieldand all fields accumulated. The difference of CTV D95 between the sCT and gCT was 90 cGy[RBE], The model yielded a mean SE of only 0.2 mm in the sCT robustness test.

[0046] FIG. 10 illustrates an example of a system 1000 for reconstructing 3D medical images (e.g., 3D CT images) from 2D medical images (e.g., 2D kV images) in accordance with some embodiments of the systems and methods described in the present disclosure. As shown in FIG. 10, a computing device 1050 can receive one or more types of data (e.g., 2D kV image data, other 2D medical image data) from data source 1002. In some embodiments, computing device 1050 can execute at least a portion of a 3D medical image reconstruction system 1004 to reconstruct 3D medical images (e.g., 3D CT images) from 2D medical images (e.g., 2D kV images) received from the data source 1002.

[0047] Additionally or alternatively, in some embodiments, the computing device 1050 can communicate information about data received from the data source 1002 to a server 1052 over a communication network 1054, which can execute at least a portion of the 3D CT image reconstruction system 1004. In such embodiments, the server 1052 can return information to the computing device 1050 (and / or any other suitable computing device) indicative of an output of the 3D CT image reconstruction system 1004.

[0048] In some embodiments, computing device 1050 and / or server 1052 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 1050 and / or server 1052 can also reconstruct images from the data.

[0049] In some embodiments, data source 1002 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data), such as an x-ray imaging system, another computing device (e.g.. a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some embodiments, data source 1002 can be local to computing device 1050. For example, data source 1002 can be incorporated with computing device 1050 (e.g., computing device 1050 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 1002 can be connected to computing device 1050 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 1002 can be located locally and / or remotely from computing device 1050, and can communicate data to computing device 1050 (and / or server 1052) via a communication network (e.g., communication network 1054).

[0050] In some embodiments, communication network 1054 can be any suitable communication network or combination of communication networks. For example, communication network 1054 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA. GSM, LTE, LTE Advanced. WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication network 1054 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi -private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG. 10 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.

[0051] Referring now to FIG. 11, an example of hardware 1100 that can be used to implement data source 1002, computing device 1050, and server 1052 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.

[0052] As shown in FIG. 11, in some embodiments, computing device 1050 can include a processor 1102, a display 1104, one or more inputs 1106, one or more communication systems 1108, and / or memory71110. In some embodiments, processor 1102 can be any suitable hardware processor or combination of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and so on. In some embodiments, display 1104 can include any suitable display devices, such as a liquid crystal display (LCD) screen, a light-emitting diode (LED) display, an organic LED (OLED) display, an electrophoretic display (e.g., an “e- ink” display), a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1106 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0053] In some embodiments, communications systems 1108 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1054 and / or any other suitable communication networks. For example, communications systems 1108 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1108 can include hardware, firmware, and / or software that can beused to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0054] In some embodiments, memory 1110 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1102 to present content using display 1104, to communicate with server 1052 via communications system(s) 1108, and so on. Memory 1110 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 11 10 can include random-access memory (RAM), read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), other forms of volatile memory, other forms of non-volatile memoiy, one or more forms of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1 110 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 1050. In such embodiments, processor 1102 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 1052, transmit information to server 1052, and so on. For example, the processor 1102 and the memory 1110 can be configured to perform the methods described herein (e.g., the method of FIG. 2, the method of FIG. 3).

[0055] In some embodiments, server 1052 can include a processor 1112, a display 1114, one or more inputs 1116. one or more communications systems 1118, and / or memory 1120. In some embodiments, processor 1 112 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 1114 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1 116 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0056] In some embodiments, communications systems 1118 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1054 and / or any other suitable communication networks. For example, communications systems 1118 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1118 can include hardware, firmware, and / or software that can beused to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0057] In some embodiments, memory 1120 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1112 to present content using display 1114. to communicate with one or more computing devices 1050, and so on. Memory 1120 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1120 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other ty pes of non-volatile memory7, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1120 can have encoded thereon a server program for controlling operation of server 1052. In such embodiments, processor 1112 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1050, receive information and / or content from one or more computing devices 1050, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.

[0058] In some embodiments, the server 1052 is configured to perform the methods described in the present disclosure. For example, the processor 1112 and memory71120 can be configured to perform the methods described herein (e.g., the method of FIG. 2, the method of FIG. 3).

[0059] In some embodiments, data source 1002 can include a processor 1122, one or more data acquisition systems 1124, one or more communications systems 1126, and / or memory 1128. In some embodiments, processor 1122 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 1124 are generally configured to acquire data, images, or both, and can include an x-ray imaging system, a CT imaging system, or the like. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 1124 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of an x-ray imaging system, a CT imaging system, or the like. In some embodiments, one or more portions of the data acquisition system(s) 1124 can be removable and / or replaceable.

[0060] Note that, although not shown, data source 1002 can include any suitable inputs and / or outputs. For example, data source 1002 can include input devices and / or sensors thatcan be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 1002 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.

[0061] In some embodiments, communications systems 1126 can include any suitable hardware, firmware, and / or software for communicating information to computing device 1050 (and, in some embodiments, over communication network 1054 and / or any other suitable communication networks). For example, communications systems 1126 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1126 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0062] In some embodiments, memory 1128 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1 122 to control the one or more data acquisition systems 1 124, and / or receive data from the one or more data acquisition systems 1124; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 1050; and so on. Memory 1128 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1128 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1128 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 1002. In such embodiments, processor 1122 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1050, receive information and / or content from one or more computing devices 1050, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.

[0063] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory or non-transitory.For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.

[0064] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).

[0065] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.

[0066] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention.

Claims

CLAIMS1. A method for generating three-dimensional (3D) medical images from two- dimensional (2D) medical images, the method comprising: accessing 2D image data with a computer system, wherein the 2D image data compnse 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system; accessing a machine learning model trained on training data to reconstruct 3D medical images from 2D medical images; inputting the 2D image data to the machine learning model using the computer system, generating 3D image data as an output; and outputting the 3D image data via the computer system.

2. The method of claim 1 , wherein the machine learning model comprises an autoencoder network.

3. The method of claim 2, wherein the autoencoder network is an asymmetric autoencoder network comprising vision transformer blocks.

4. The method of claim 3. wherein each vision transformer block implements a window-based self-attention.

5. The method of claim 1, wherein the machine learning model comprises a first machine learning model and a second machine learning model, wherein the first machine learning model is trained on first training data to reconstruct an image of a first anatomical region and the second machine learning model is trained on second training data to reconstruct an image of a second anatomical region that is different than the first anatomical region.

6. The method of claim 5. wherein the first anatomical region comprises a whole body region.

7. The method of any one of claims 5 or 6, wherein the second anatomical region comprises a head and neck region.

8. The method of claim 5, wherein the second machine learning model is trained on second training data that includes cropping ground truth image data acquired from the first anatomical region to create cropped ground truth image data that focuses on the second anatomical region.

9. The method of claim 5, wherein inputting the image data to the machine learning model using the computer system comprises: inputting the 2D image data to the first machine learning model, generating first output image data that depict the first anatomical region; inputting the 2D image data to the second machine learning model, generating second output image data that depict the second anatomical region; and generating the 3D image data by concatenating the first output image data and the second output image data.

10. The method of claim 9, wherein the first output image data and the second output image data are concatenated based on a spatial relationship between the first anatomical region and the second anatomical region.

11. The method of claim 5. wherein each of the first machine learning model and the second machine learning model comprise an autoencoder network.

12. The method of claim 11, wherein the autoencoder network is an asymmetric autoencoder network comprising vision transformer blocks.

13. The method of claim 12, wherein each vision transformer block implements a window-based self-attention.

14. The method of claim 1, wherein the training data comprise pairs of orthogonal kV images and corresponding ground truth images.

15. The method of claim 14, wherein the ground truth images comprise 3D CT images.

16. The method of claim 15, wherein the 3D CT images are obtained with a CT- on-rails imaging system.

17. The method of claim 14, wherein the training data include augmented training data generated using a geometry property-reserved shifting to create the augmented training data from the pairs of orthogonal kV images and corresponding ground truth images.

18. The method of claim 17, wherein the augmented training data are generated using the geometry property -reserved shifting by shifting each image in the training data by a shift value that accounts for a geometrical relation between a treatment couch of the radiation treatment system and a kV imaging source and detector.

19. The method of claim 18, wherein the images in the pairs of orthogonal kV images and corresponding ground truth images are simultaneously shifted by the shift value.

20. The method of claim 19, wherein the shift value is within ±5 mm.

21. The method of claim 14, wherein the training data include augmented training data generated using a geometry property -reserved resampling to create the augmented training data as smaller-sized data pairs from the pairs of orthogonal kV images and corresponding ground truth images.

22. A method for generating three-dimensional (3D) medical images from two- dimensional (2D) medical images, the method comprising: accessing 2D image data with a computer system, wherein the 2D image data comprise 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system; accessing an autoencoder network trained on training data to reconstruct 3D medical images from 2D medical images, wherein the autoencoder network is an asymmetric autoencoder network comprising vision transformer blocks, wherein each vision transformer block implements a window-based self-attention;inputing the 2D image data to the autoencoder network using the computer system, generating 3D image data as an output; and outputing the 3D image data via the computer system.

23. The method of claim 22, wherein the training data comprise pairs of orthogonal kV images and corresponding ground truth images.

24. The method of claim 23, wherein the ground truth images comprise 3D CT images.

25. The method of claim 24, wherein the 3D CT images are obtained with a CT- on-rails imaging system.

26. The method of claim 23, wherein the training data include augmented training data generated using a geometry7property-reserved shifting to create the augmented training data from the pairs of orthogonal kV images and corresponding ground truth images.

27. The method of claim 26, wherein the augmented training data are generated using the geometry' property -reserved shifting by shifting each image in the training data by a shift value that accounts for a geometrical relation between a treatment couch of the radiation treatment system and a kV imaging source and detector.

28. The method of claim 27. wherein the images in the pairs of orthogonal kV images and corresponding ground truth images are simultaneously shifted by' the shift value.

29. The method of claim 28, wherein the shift value is within ±5 mm.

30. The method of claim 23, wherein the training data include augmented training data generated using a geometry property -reserv ed resampling to create the augmented training data as smaller-sized data pairs from the pairs of orthogonal kV images and corresponding ground truth images.

31. A method for generating three-dimensional (3D) medical images from two- dimensional (2D) medical images, the method comprising:accessing 2D image data with a computer system, wherein the 2D image data comprise 2D kV images acquired from a subject while the subject is positioned at a treatment position of a radiation treatment system; accessing a first machine learning model with the computer system, wherein the first machine learning model has been trained on first training data to reconstruct an image of a first anatomical region; accessing a second machine learning model with the computer system, wherein the second machine learning model has been trained on second training data to reconstruct an image of a second anatomical region that is different than the first anatomical region; inputting the 2D image data to the first machine learning model, generating first output image data that depict the first anatomical region; inputting the 2D image data to the second machine learning model, generating second output image data that depict the second anatomical region; and generating 3D image data by concatenating the first output image data and the second output image data; and outputting the 3D image data via the computer system.

32. The method of claim 31, wherein the first anatomical region comprises a whole body region.

33. The method of any one of claims 31 or 32. wherein the second anatomical region comprises a head and neck region.

34. The method of claim 31, wherein the second machine learning model is trained on second training data that includes cropping ground truth image data acquired from the first anatomical region to create cropped ground truth image data that focuses on the second anatomical region.

35. The method of claim 31, wherein the first output image data and the second output image data are concatenated based on a spatial relationship between the first anatomical region and the second anatomical region.

36. The method of claim 31 , wherein at least one of the first machine learning model or the second machine learning model comprises an autoencoder network.

37. The method of claim 36, wherein the autoencoder network is an asymmetric autoencoder network comprising vision transformer blocks.

38. The method of claim 37, wherein each vision transformer block implements a window-based self-attention.

39. The method of claim 31, wherein the training data comprise pairs of orthogonal kV images and corresponding ground truth images.

40. The method of claim 39, wherein the ground truth images comprise 3D CT images.

41. The method of claim 40, wherein the 3D CT images are obtained with a CT- on-rails imaging system.

42. The method of claim 39, wherein the training data include augmented training data generated using a geometry' property-reserved shifting to create the augmented training data from the pairs of orthogonal kV images and corresponding ground truth images.

43. The method of claim 42. wherein the augmented training data are generated using the geometry property -reserved shifting by shifting each image in the training data by a shift value that accounts for a geometrical relation betw een a treatment couch of the radiation treatment system and a kV imaging source and detector.

44. The method of claim 43, wherein the images in the pairs of orthogonal kV images and corresponding ground truth images are simultaneously shifted by the shift value.

45. The method of claim 44, wherein the shift value is w ithin ±5 mm.

Citation Information

Patent Citations

  • Deep learning apparatus and method for segmentation and survival prediction for head and neck tumors

    US20230414189A1