A deep learning model to interpolate temporal series for time resolved medical imaging
A deep learning model enhances temporal resolution in CMR scans by interpolating spatiotemporal data, addressing the inefficiencies and artifacts of existing methods to reduce scan time and improve image quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MAYO FOUNDATION FOR MEDICAL EDUCATION & RESEARCH
- Filing Date
- 2025-12-01
- Publication Date
- 2026-06-04
AI Technical Summary
Current CMR scanning methods require long breath-holds, which are burdensome for patients, and existing acceleration techniques suffer from residual artifacts, computational inefficiencies, and suboptimal image quality, particularly in undersampled datasets.
A deep learning model is applied to spatiotemporal medical images to interpolate temporal series, utilizing 2D convolutional neural networks and cascaded models to enhance temporal resolution while reducing computational resources and artifacts.
The method significantly shortens scan time and improves image quality, maintaining diagnostic accuracy by reconstructing high temporal resolution images from low-resolution acquisitions, compatible with existing acceleration strategies.
Smart Images

Figure US2025057536_04062026_PF_FP_ABST
Abstract
Description
Mayo 2023-546630666.01650A DEEP LEARNING MODEL TO INTERPOLATE TEMPORAL SERIES FOR TIME RESOLVED MEDICAL IMAGINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 726,370, filed on November 29, 2024, and entitled “A DEEP LEARNING MODEL TO INTERPOLATE TEMPORAL SERIES FOR TIME RESOLVED MEDICAL IMAGING,” which is herein incorporated by reference in its entirety.BACKGROUND
[0002] CINE cardiac magnetic resonance imaging (CMR) is routinely used in diagnostic radiology to evaluate anatomical structure and functional parameters of the heart. An adequate image featuring good SNR and high spatial and temporal resolution is essential for accurate diagnosis. Most currently used clinical protocols require the patient to hold their breath during the scan to eliminate artifacts from respiratory motion. However, breath holding can be a serious burden for patients. Therefore, accelerating scan speed is favorable for CINE CMR scanning.
[0003] CMR scans require long breath-holds to achieve adequate spatial and temporal resolutions for accurate diagnosis. The images are used to evaluate the anatomical structure of the heart and functional parameters via quantitative analysis. A typical scan may require a subject to hold their breath frequently, in many examinations exceeding 50 times or more. The duration of each breath hold may typically take between 10 to 20 seconds or so, depending on the exact parameters used for the acquisition. Final image quality is heavily dependent on the subject’s capability to hold their breath during the scan, which can be a burden for some patients.
[0004] Several acceleration methods have been disclosed to increase scan efficiency of CMR. Parallel imaging methods are popular implementations. Typical acceleration factors are 2 to 3, lowering breath holding time to approximately 10 seconds per slice. To cover the full heart in all desired orientations, many breath-holds are required for an exam, frequently 50 or more. Higher acceleration rates can be achieved via regularized reconstruction such as compressed sensing (CS). Regularization weighting is debated and residual image artifacts due to regularization remain an issue. Moreover, the computation time for the regularized1QB\630666.01650\99677203.2Mayo 2023-546630666.01650 reconstructions is orders of magnitude longer compared to other methods and can sometimes amount to tens of minutes, making it less favorable in clinical settings.SUMMARY OF THE DISCLOSURE
[0005] According to an aspect of the present disclosure, a method is provided. The method includes obtaining a spatiotemporal medical image. The method includes processing the spatiotemporal medical image to generate a set of two-dimensional (2D) spatiotemporal data. The method includes applying a 2D deep learning model to the 2D spatiotemporal data to generate a set of transformed 2D spatiotemporal data. The method includes generating a transformed spatiotemporal medical image based on the set of transformed 2D spatiotemporal data.
[0006] According to another aspect of the present disclosure, a method is provided. The method includes obtaining a target spatiotemporal medical image. The method includes processing the target spatiotemporal medical image to generate a simulated spatiotemporal medical image having a lower temporal resolution than the target spatiotemporal medical image. The method includes processing the simulated spatiotemporal medical image to generate a set of training 2D spatiotemporal data. The method includes processing the input spatiotemporal medical image to generate a set of target 2D spatiotemporal data. The method includes applying a 2D deep learning model to generate a set of training temporally upsampled 2D spatiotemporal data based on the training 2D spatiotemporal data. The method includes updating the 2D deep learning model based on comparing the set of training temporally upsampled 2D spatiotemporal data to the set of target 2D spatiotemporal data to generate a trained 2D deep learning model.
[0007] According to another aspect of the present disclosure, a method is provided. The method includes obtaining a spatiotemporal medical image. The method includes processing the spatiotemporal medical image to generate a set of 2D spatiotemporal data. The method includes applying a 2D deep learning model to temporally upsample the 2D spatiotemporal data to generate a set of upsampled 2D spatiotemporal data. The method includes generating a reconstructed temporally upsampled spatiotemporal medical image based on the set of upsampled 2D spatiotemporal data.2QB\630666.01650\99677203.2Mayo 2023-546630666.01650BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is an example of undersampling k-space data to generate simulated low temporal resolution images, including spatiotemporal MRI images.
[0009] FIG. 2 is an example of k-space radial undersampling in spatiotemporal image data.
[0010] FIG. 3 is an example deep learning model for temporal enhancement in spatiotemporal medical images.
[0011] FIG. 4 is another example deep learning model for temporal enhancement in spatiotemporal medical images including a denoising deep learning model portion.
[0012] FIG. 5A is another example deep learning model for temporal enhancement in spatiotemporal medical images including a spatial correlation and 2D spatial image deep learning model portion.
[0013] FIG. 5B is another example deep learning model for temporal enhancement in spatiotemporal medical images, which is configured to remove residual spatial artifacts.
[0014] FIG. 5C shows an example residual DENSE block architecture used in the deep learning model of FIG. 5B.
[0015] FIG. 6 is an example of results of temporal enhancement image data reconstruction.
[0016] FIG. 7A is an example deep learning model for temporal enhancement in spatiotemporal medical image including a deep learning model cascade.
[0017] FIG. 7B is an example deep learning model component of a deep learning model cascade as illustrated in FIG. 7A.
[0018] FIG. 8 is an example U-net neural network to support reconstructing spatiotemporal medical image data.
[0019] FIG. 9 is a flowchart of an example method for reconstructing spatiotemporal MRI image data, including CINE CMR image data.
[0020] FIG. 10 is a flowchart of an example method for training a deep learning model, or other suitable machine learning model, to reconstruct medical image data.
[0021] FIG. 11 is a block diagram of an example MRI image data noise reconstruction system according to some embodiments described in the present disclosure.
[0022] FIG. 12 is a block diagram of example components that can implement the system of FIG. 11.3QB\630666.01650\99677203.2Mayo 2023-546630666.01650DETAILED DESCRIPTION
[0023] As used herein, the term “spatiotemporal data” refers to data including spatial characteristics and temporal characteristics. For example, four-dimensional (4D) spatiotemporal data may comprise (x,y,z,t) data where x, y, z index position data and t indexes time data (e.g., volume image data over time), 3D spatiotemporal data may comprise (x,y,t) data (e.g., flat image data over time). As another example, 2D spatiotemporal data may comprise a ID vector of pixels over time (denoted (y, t) data). For example, the ID vector may comprise a slice, projection, or other mapping of a 2D or 3D spatial image over time. By way of example, implementations may be described with respect to (y, t) data derived from y- coordinate data over time. Of course, implementations are not limited to y-coordinate data over time. For instance, implementations may apply to any spatiotemporal data such as time data along with x-coordinate data (e.g., (x, t) data), z-coordinate data (e.g.. (z, t) data), oblique slice data, projected or otherwise mapped x, y, or z data to ID data, etc.
[0024] Various acceleration strategies such as parallel imaging have been proposed to address the issues in scan efficiency for CINE CMR. These approaches have partially ameliorated the problem of long scan duration but have shown to still be insufficient for the sicker patients in particular, who are unable to sustain long breath-holds. However, higher acceleration in parallel imaging becomes difficult to advance due to the physical and hardware limitations. Some approaches have exploited regularized reconstruction such as kt-blast and compressed sensing to improve scan speed while maintaining image quality. These frameworks usually require empirical tunings of the regularization parameters to balance the image noise, amplified artifacts, and loss of spatial and temporal information. Another drawback of the regularized methods are their long reconstruction times, sometimes taking tens of hours for a larger dataset with sophisticated regularization. While these methods have been investigated for many years and some have been proved useful in concept, there is still lack of consensus regarding the optimal selection of the regularization parameters and how these regularization affects the clinical interpretations.
[0025] Deep learning (DL) models are emergent approaches to facilitate faster CMR scanning, as well as faster scanning with other time-resolved imaging techniques and modalities, such as x-ray computed tomography (x-ray CT), single-photon emission computed tomography (SPECT), positron emission tomography (PET), ultrasound imaging (e.g., echocardiography and others), etc. Although there may be an overhead computational cost in training a model, an advantage of DL approaches is very fast processing after the scan is done.4QB\630666.01650\99677203.2Mayo 2023-546630666.01650DL frameworks often focus on two different approaches. One approach includes image restoration from a lower spatial resolution scan. This approach may provide a scan speed that may be performed a little bit faster than conventional parallel imaging. However, it may be relatively less compatible other parallel imaging methods. Another approach is to simulate an undersampling process from the raw data, whereafter a DL model is built to restore the image. While this framework may be sound in principle, it may require substantial computational resources. For example, inaccurate fitting can be expected if the hardware (e.g., computation resources) or the data do not meet the requirement of the model, potentially leading to suboptimal image quality'.
[0026] DL methods have drawn attention in image reconstruction. However, while typical DL based strategies, may in general, allow for faster image reconstruction, they ty pically require substantial amounts of data and time in advance to train an appropriate model. Despite the overhead computational cost, the application of DL methods can be a useful tool for reconstructing accelerated dataset for diagnosis due to the shorter reconstruction time.
[0027] Many acceleration strategies for imaging modalities such as CINE CMR have been developed. Parallel imaging methods are often popular in clinical practice due to their robustness. A typical acceleration factor via parallel imaging is between 2-4 depending on the coil arrangement. While parallel imaging methods may reduce breath-holding time to around 10 seconds per slice, but it may still require about 15 breath-holds to obtain sufficient information. In some cases, higher acceleration can be achieved by employing spatial-temporal correlation in signal reconstruction, such as via spatial frequency techniques (e.g., k-t SENSE) and compressed sensing (CS) methods. These methods utilize additional regularization information to avoid noise amplification during image reconstruction from raw data. How ever, the introduction of regularization often causes residual artifacts. Moreover, optimal parameters in the regularized reconstruction remain inconclusive and reconstruction can sometimes take tens of minutes to finish. Accordingly, these methods may be insufficient to meet the clinical needs.
[0028] Aspects of the described technology may provide DL spatiotemporal image processing systems, methods, applications, devices, etc. For example, further aspects may provide such technology7to generate training data for reconstructing low temporal resolution images. In contrast to some existing approaches, training data need not necessarily come from raw data. While raw data may be used in some implementations, implementations may use encoded images such as DICOM images. In some examples, undersampling functions along5QB\630666.01650\99677203.2Mayo 2023-546630666.01650 time that resemble a sampling strategy in an accelerated scan may be used to create the low temporal resolution images for training.
[0029] Aspects of the described technology may include a deep learning model to recover diagnostic temporal resolution from a faster undersampled short axis CINE CMR scan. For instance, models implemented as described may recover data with a large variety of undersampling factors, such as from about 2 and beyond.
[0030] As an example, the described technology may support an improved / faster cardiac scan. For instance, MR pulse sequences for CINE CMR may be provided to match supported scan parameters for the faster cardiac scan with a reduction factor in temporal resolution. For example, in some cases acquisition duration may be reduced significantly more than 2 fold (e.g., with respect to the above undersampling factors). Implementations may provide image quality and physiological parameters including the ejection fraction similar or superior to existing clinical protocols with a faster scanning method.
[0031] Implementations of the described technology may improve CINE CMR scan acquisition speed several times over the existing spatial acceleration methods. As such, it may ease a patient’s burden during a CMR exam without loss of diagnostic information. In addition to the application in CINE CMR, the described technology may be implemented with respect to other imaging modalities for time resolved images such as x-ray CT, SPECT, PET, ultrasound imaging (e.g., echocardiography and others), etc. The implementation to other image modalities may reduce exam time or radiation dose in a high temporal resolution scan.
[0032] An example technique to reduce scan time may include reducing temporal resolution, as breath-hold duration is inversely proportional to the temporal resolution. However, decreasing of temporal resolution may introduce temporal blurring and is likely to hinder assessment of cardiac function such as the end systolic volume and the ejection fraction. Therefore, the described technology may provide an efficient algorithm to reconstruct a high temporal resolution time series needed for accurate diagnosis from a low temporal resolution acquisition, which may be beneficial to the patients undergoing a temporally resolved imaging exam.
[0033] One challenge for temporal algorithms in CMR, or other time-resolved imaging techniques, is that undersampling in the time domain may unfavorably alter the image quality in the spatial domain. Deep learning models may be expected to handle a signal in both spatial and temporal dimensions when a temporally undersampled dataset is to be reconstructed to a series of images with high temporal resolution sufficient for diagnosis. Another challenge is6QB\630666.01650\99677203.2Mayo 2023-546630666.01650 that most of the reconstruction methods for CMR need training from the raw data, which are often discarded after an exam (e.g., due to its large size). Raw data formats also vary with vendors. Aspects of the described technology may utilize images (e.g., DICOM images, other existing or to-be-developed formats, etc.) for diagnosis to build a DL model that may address these challenges and can become vendor neutral.
[0034] Temporal interpolation using deep learning may be performed as a computational technique to generate super-resolution data along the temporal dimension. DL reconstruction methods for recovery' of CINE CMR data from low spatial resolution datasets have been applied in clinical research with success. Example implementations may reconstruct a series of low spatial resolution images to at least their original full spatial resolution. Examples may be combined with various other imaging modalities, such as Generalized Autocalibrating Partially Parallel Acquisition (GRAPPA) techniques.
[0035] While temporal undersampling in time resolved medical imaging may reduce the number of time frames, it may also introduce spatial artifacts. An example technique to generate temporally undersampled dataset is to perform simulated subsampling from the raw data followed by application of a DL transform via a DL network with high dimensional convolution to reconstruct the image. This example provides a general framework in the reconstruction for the data with Cartesian sampling. However, it may incur a cost in computation including hardware (e.g., physical computing resources) and time (e.g., computation time). Additionally, this example may introduce residual artifacts, particularly in a 2D scan.
[0036] CMR is a useful tool in diagnostic radiology7. It provides both anatomical information and functional evaluation for the heart. Accurate diagnosis may benefit from good image quality7, such as featuring good SNR and high spatial and temporal resolution. Although reduction in temporal resolution allows shorter scan time for CINE CMR exams, it may induce errors, such as in assessment of wall motion abnormalities, as well as in quantitative estimation of physiological parameters. The disclosed technology may offer a deep learning framework that may perform interpolation on a series of spatiotemporal images (e.g.. CINE CMR images) to improve the temporal resolution so as to restore accurate diagnostic information while keeping acquisition time short.
[0037] Aspects of the described technology may address these challenges, providing relatively higher compatibility to other (e.g.. existing or later-developed) acceleration methods7QB\630666.01650\99677203.2Mayo 2023-546630666.01650 to increase scan speed and / or to restore images with a smaller number of computational resources.
[0038] Implementations may include a principal model that utilizes sparse signal characteristics in a 2D spatiotemporal space (e.g., a (y, t) space) to setup a deep learning model (e.g.,2D convolutional neural network (“CNN”)) to recover temporal information. Implementations may include further models, such as a 2D or a relatively small-sized 3D CNN, cascaded before or after the principal model to remove the residual spatial image artifacts. For example, each cascaded model can be individually trained, allowing lower computational expense in hardware and in time. For example, the entire method can accelerate the image acquisition process by several folds as the framework is compatible with existing parallel imaging methods. In principle, the disclosed method can also be applied to any other image modality used to obtain temporal information including but not limited to, echocardiography, x-ray CT, SPECT, and PET.
[0039] The described technology is not limited to CINE CMR, but may be applied to various other medical imaging modalities to increase temporal resolutions, such as, for example, echocardiography, x-ray CT, SPECT, PET, etc. Additionally, the disclosed technology7, such as the temporal interpolation method may provide an image processing pipeline that is compatible with routine imaging modalities used in clinics. Additionally, some disclosed DL models to improve image quality from low temporal resolution scans may be established using training data based on simulated undersampling from DICOM data without the need for raw data.
[0040] Aspects of the described technology may provide a framework for reconstructing temporally undersampled data. An example aspect of utilizing temporal undersampling is that full resolution frames in the spatial domain are used. The described technology may be compatible with a variety of existing acceleration strategies such as SENSE, GRAPPA, and CS. Accordingly, the disclosed temporal interpolation strategy7has the potential to shorten the scan time even with existing acceleration strategies being applied. In some cases, the disclosed technology may be employed without raw data. For instance, in some examples, DICOM images may be sufficient for model training and image generation. Of course, the disclosed technology7may be employed on raw data as well as encoded data. Various aspects may use relatively reduced computational resources, such as a relatively small memory7space of the hardware, which may support faster computation during the DL training phase and application of techniques on shared computational resources (e.g., in a clinical setting on a8QB\630666.01650\99677203.2Mayo 2023-546630666.01650 computer hosting other applications). Further, implementations may be compatible with different types of spatial sampling schemes used in MRIs. such as radial and spiral trajectories. Accordingly, the disclosed technology may be applied to the reconstruction of different applications. Furthermore, it can also be extended to the implementation for ultrasonography, dynamic x-ray CT, SPECT, PET scans, etc. Implementations may be applied to any spatiotemporal imaging modality and training may be performed when the sampling scheme is known for simulating the undersampled data to train the model. Of course, other implementations may be applied without simulated training data, such as via suitable imaging datasets.
[0041] Sparsity-related techniques employ a strategy to reconstruct undersampled signals robustly. For example, the use of sparsity (phase encoding - cardiac phase) to reconstruct undersampled CMR has been documented in parallel imaging and compressed sensing. Sparse characteristics allow dimensional reduction in recovering high dimensional data. Some implementations of the disclosed technology may include lower dimensional convolution DL network transforms. In some cases, 2D spatiotemporal data may provide sparse features similar to a 3D (e.g., (x, y, t)) or a 4D (e.g., (x, y, z, t)) domain of a given dataset yet requiring much lower computational expense. Accordingly, the aspects of the described technology applying 2D reconstruction (e.g., in the (y, t) domain) may both capture major signal features efficiently and also reduce computational demands for much faster processing speed over prior approaches. Additionally, implementations may be extensible to cascade additional deep learning models (e.g., to polish the image quality by removing residual artifacts). Further aspects of the disclosed technology may include preparation of training data and a DL network structure.
[0042] Temporal undersampling in cardiac imaging may be implemented in various manners. For example, temporal undersampling may be implemented by skipping time frames and, more commonly, by exchanging temporal frames in a period for a greater spatial coverage to shorten the exam time. For instance, FIG. 1 shows an example for reducing temporal resolution by three times while maintaining spatial resolution. For example, the illustrated method may be applied to Cartesian encoding with segmented k-space sampling methods. For instance, such technologies may be used in MRI and other imaging modalities, including analogues in modalities such as ultrasonography. FIG. 2 demonstrates another example data acquisition scheme implementing radial encoding, which may be used in CT. PET, SPECT, MRI modalities, etc. Example implementations of the described technology may simulate these9QB\630666.01650\99677203.2Mayo 2023-546630666.01650 actual undersampling functions on DICOM images (e.g., existing DICOM image datasets) to prepare for deep learning model training data. In some examples, data generated by such implementations may simulate (e.g., resemble) an image that is coarsely sampled in time, for example, as would be generated during a physical scan. This may, for example, provide training data and techniques supporting deployed DL deep learning models in clinical or other environments that may feature artifacts, subsequently enhancing image quality in deployed implementations (e.g., in a clinical environment).
[0043] FIGS. 3-5 illustrate various example implementations illustrating various aspects of the disclosed technology. For example, the illustrated deep learning models may be implemented via hardware logic, such as an Application Specific Integrated Circuit (ASIC), programmed Field-Programmable Gate Array (FPGA), neuromorphic computer, etc. As another example, the illustrated deep learning models may be implemented via instructions stored on a non-transitory computer readable medium and executed by a computer (e.g., general purpose processor, central processing unit (CPU), general-purpose graphic processing unit (GPGPU). neural network processor, Al processor, accelerator, etc. In further examples, the illustrated networks may be implemented via combinations thereof or other manners of providing a physically implementation. Accordingly, FIGS. 3-5 should be understood to illustrate deep learning models embodied in a physical form.
[0044] FIG. 3 illustrates a principal deep learning model 301 that may perform transforms on input image data 310 to recover the temporal information and to remove potential artifacts due to undersampling. In various implementations, input image data 310 may comprise a spatiotemporal medical image, such as, for example clinical data (e.g., data obtained via a patient scan or other clinical source), training data, test data, validation data, etc. In some implementations, deep learning model 301 may process the spatiotemporal medical image to generate a set of two-dimensional (2D) spatiotemporal data. For example, simulated undersampled data 310 may be reformatted and transformed along the (y , t) domain as the input 311 of the deep learning model 301.
[0045] Deep learning model 301 may apply a transformation to the spatiotemporal medical images as an intermediate input 311 to generate a set of transformed 2D spatiotemporal data 314. For example, deep learning model 301 may comprise a series of individual convolutional neural network transformers 302-305. For instance, individual 2D convolutional neural networks (2D CNN) 302-305 with batch normalization 306-308 may be used to produce an output 312 with enhanced image quality for (y, t) data in the output 312 of the model 301.10QB\630666.01650\99677203.2Mayo 2023-546630666.01650The structure of the 2D CNN in this method can be a convolution network, a U-net, a densely connected network, a generative Al model, a fully connected network, a recursive network, a residual net, a transformer, an attention module, generative adversarial model, inception module, or reserv oir computer, a transformer network, and so on, as different model structures may be suitable for different sampling patterns. Once the model is built, the training process may match the undersampled (y. t) data to the reference counterpart.
[0046] Some implementations may include additional image processing operations. For instance, processing operations may be performed to account for spatial / spatiotemporal correlations outside the (y, t) domain. For example, spatial correlations may be present in the (x,y). (y,z) or (x,y,z) domains or spatiotemporal correlations in the (x.y.t), etc. domains (e.g., a y value at t=l might correlate to an x value at t=2, etc.). For example, operations for the treatment of information in the spatial domain may be included in the aforementioned principal model 301. For example, FIG. 4 illustrates a denoise model 404 cascaded in front of the principal model 301 to reduce the variations across the space as shown in FIG. 4. In this example, a noisy input spatiotemporal image 401 (represented by image 402 combined with noise 403) is input to a denoising model 404 (e.g., a denoising CNN) to produce a denoised output spatiotemporal image 405. As illustrated, the denoised image 405 may serve as input 310 in the above-described model 301.
[0047] Some implementations may include additional image processing operations following an application of a pnncipal model 301. For example, FIG. 5 A illustrates an example implementation where an output 312 of model 301 is reformatted at process block 501 to an intermediate higher-dimensional spatiotemporal image 502 (e.g., output (y,t) data 314 may be combined to form a 3D spatiotemporal image (e.g., an image with (x,y,t) data) or a 4D spatiotemporal image (e.g., (x.y.z.t) data). As illustrated, spatiotemporal image 502 may be an input to further processing operations, such as a second deep learning model 503. For instance, model 503 may restore spatial correlation information that may have become decorrelated due to the spatiotemporal image 502 reconstructed from the (y, t) data that are the output 312 of the model 301. As an example, deep learning model 503 may comprise a CNN, such as a 2D or 3D CNN, to restore spatial correlation. In some implementations, deep learning model 503 may comprise a cascade of models 504-506. For instance, deep learning model 503 may comprise another 2D convolutional network along the x-y space or a high dimensional convolutional network in x-y-t or in x-y-z-t space depending on the signal characteristics. In some such implementations, the 3D network may be relatively smaller than might be used in an image11QB\630666.01650\99677203.2Mayo 2023-546630666.01650 processing model lacking the (y, t) space model 301. In some examples, each cascade 504, 505, 506 may be trained individually, for example, to reduce memory space requirements.
[0048] FIG. 5B illustrates an alternative model architecture that is configured to remove residual spatial artifacts. This architecture makes use of residual densely connected convolutional (DENSE) blocks to remove spatial and temporal blurring and skip connections (e.g., between the first 2D convolution block and the summation block). In the illustrated example, the 1st cascade contains residual DENSE blocks, of which the structure is shown in FIG. 7B. The 2nd cascade is a 2D U-Net to reduce decorrelation artifact in the x-y domain.
[0049] FIG. 7A illustrates another example implementation of a 2D spatiotemporal deep learning model 701. For example, deep learning model 701 may illustrate an example implementation of a deep learning model 301. In this example, model 701 may comprise a 2D generative residual DENSE network to recover the temporal resolution of input data 702. In this example, input spatiotemporal 702 may be parsed into 2D spatiotemporal image data 703 (e.g., (y, t) data). Data 703 may be transformed via a cascade of DENSE networks 704-707. FIG. 7B illustrates an example implementation of a DENSE network 704-707. In the illustrated example, a DENSE net may comprise a plurality of 2D convolutional layers 722-727, 737-742, and residual connections 729-736 to extract high-level features from the data.
[0050] In some implementations, other processing operations may be performed, such as batch normalizations (BN) 708-711 (e.g., to assist convergence during a training process. In this example, the input data 702 may be convolved 712 with the output of the cascade (e.g., output of BN 711) to produce output data 713, which may be combined to generate an output spatiotemporal image 714.
[0051] FIG. 8 illustrates an example additional processing model 800. For instance, the illustrated model may comprise a denoising model 800. For instance, model 800 may be included as an additional cascade in front of the generative model for resolution enhancement. For example, in implementations, medical images may be obtained with different signal-to- noise ratios (SNR), which may result from differences between scans including, for instance, different voxel size, coil arrangement and vendors of magnet. In some examples, denoiser model 800 may constrain the SNR within a specific range before the temporal treatment (e.g., so that unexpected noise will not be amplified in the generative model).
[0052] As an example, denoising model 800 may be implemented via U-net structure as illustrated. For instance, model 800 may receive an input image 801 and process it to produce an output image 811 via one or more downsampling 802 convolutional layers 803, 804, 805, a12QB\630666.01650\99677203.2Mayo 2023-546630666.01650 lowest layer 806, and one or more upsampling 810 layers, 807, 808, 809 with skip connections as illustrated. In some examples, additional simulated noise within specific SNR ranges may be added to obtained data. Such datasets with added noise and datasets without added noise may be used to fit the model to stabilize the SNR within the range of the obtained dataset.
[0053] In some examples, patient data used to build / train models may be obtained from existing scans in a patient database or other preexisting scan (e.g., as DICOM images). In some cases, data may be preprocessed, such as interpolated to a common number of frames. Continuing the example, each subject’s data may be undersampled to simulate the acquisition with different lower temporal resolutions, as discussed above. Model parameters may be tuned and fitted, e.g., based on the undersampled data and the original data, using deep learning toolkits, such as, for example high performance GPU workstations and software packages.
[0054] In some examples, testing data or validation data may be obtained from such a database or from other sources. For example, subjects (e.g., volunteers) may be scanned during model development to provide testing data to confirm a model built from simulated data is compatible with real scans. As another example, preclinical validation may be performed, for instance by including further newly acquired datasets (e.g., from volunteer scans) obtained from lower temporal resolution settings than standard imaging. Image quality such as artifact power and contrasts as well as quantitative physiological parameters such as ejection fraction may be evaluated, e.g.. by a radiologist, to verify performance of a trained model.
[0055] Referring now to FIG. 9. a flowchart is illustrated as setting forth the steps of an example method for generating reconstructed spatiotemporal medical image data, and as an example, reconstructed CINE CMR data, using a suitably trained deep learning model or other machine learning model. As will be described, the deep learning model takes medical image data as input data and generates reconstructed medical image data as output data.
[0056] The method includes accessing medical image data with a computer system, as indicated at step 902. Accessing the medical image data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the medical image data may include acquiring such data with a medical imaging system, such as MRI, x-ray CT, ultrasonography (e.g., echocardiography, etc.), PET, SPECT and so on and transferring or otherwise communicating the data to the computer system, which may be a part of the medical imaging system. As described above, the medical image data are generally time resolved image data in DICOM format, and in some instances may be CINE DICOM image data.13QB\630666.01650\99677203.2Mayo 2023-546630666.01650
[0057] A trained deep learning model (such as a neural network, tree-based transformer, reservoir computer, or other deep learning model) is then accessed with the computer system, as indicated at step 904. In general, the deep learning model is trained, or has been trained, on training data in order to reconstruct medical image data, such as temporally upsampled CINE MRI image data. This reconstruction is achieved, in part, by the deep learning model (or other machine learning model) being trained by mapping synthesized low-temporal resolution medical images to higher-temporal resolution medical images.
[0058] The trained deep learning model can include a deep learning model with any suitable deep learning model architecture for generating reconstructed medical image data. As one non-limiting example, the trained deep learning model may include a convolution network, a U-net, a densely connected network, a generative Al model, a fully connected network, a recursive network, a residual net, a transformer, an attention module, generative adversarial model, inception module, reservoir computer, etc. The trained deep learning model may in some instances have multiple inputs (e.g., corresponding to multiple slice inputs).
[0059] Accessing the trained deep learning model may include accessing network parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the deep learning model on training data. In some instances, retrieving the deep learning model can also include retrieving, constructing, or otherwise accessing the particular deep learning model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0060] As an example deep learning model, a neural netw ork generally includes an input layer, one or more hidden layers (or nodes), and an output layer. Typically, the input layer includes as many nodes as inputs provided to the neural network. The number (and the type) of inputs provided to the neural network may vary based on the particular task for the neural network.
[0061] The input layer connects to one or more hidden layers. The number of hidden layers varies and may depend on the particular task for the neural network. Additionally, each hidden layer may have a different number of nodes and may be connected to the next layer differently. For example, each node of the input layer may be connected to each node of the first hidden layer. The connection between each node of the input layer and each node of the first hidden layer may be assigned a weight parameter. Additionally, each node of the neural14QB\630666.01650\99677203.2Mayo 2023-546630666.01650 network may also be assigned a bias value. In some configurations, each node of the first hidden layer may not be connected to each node of the second hidden layer. That is. there may be some nodes of the first hidden layer that are not connected to all of the nodes of the second hidden layer. The connections between the nodes of the first hidden layers and the second hidden layers are each assigned different weight parameters. Each node of the hidden layer is generally associated with an activation function. The activation function defines how the hidden layer is to process the input received from the input layer or from a previous input or hidden layer. These activation functions may vary and be based on the type of task associated with the neural network and also on the specific type of hidden layer implemented.
[0062] Each hidden layer may perform a different function. For example, some hidden layers can be convolutional hidden layers which can, in some instances, reduce the dimensionality of the inputs. Other hidden layers can perform statistical functions such as max pooling, which may reduce a group of inputs to the maximum value; an averaging layer; batch normalization; and other such functions. In some of the hidden layers each node is connected to each node of the next hidden layer, which may be referred to then as dense layers. Some neural networks including more than, for example, three hidden layers may be considered deep neural networks.
[0063] The last hidden layer in the neural network is connected to the output layer. Similar to the input layer, the output layer typically has the same number of nodes as the possible outputs. In an example in which the neural network is trained to reconstruct medical image data, the output layer may include, for example, a number of different nodes corresponding to pixels in reconstructed medical image data.
[0064] The medical image data are then input to the trained deep learning model, generating output as reconstructed medical image data, as indicated at step 906. For example, the reconstructed medical image data may include one or more spatiotemporal images that have been temporally upsampled (e.g., had their temporal resolution increased), or otherwise have had a temporal aspect of the images improved. For example, reconstructed medical image data may include DICOM files that encode / contain the reconstructed medical image data. In other implementations, the trained deep learning model may be a deep learning reconstruction (DLR) deep learning model that collectively performs spatiotemporal correlation reconstruction and image reconstruction. In some such instances, the medical image data may be DICOM rawprojection data and the reconstructed medical image data may include one or more reconstructed images that have high temporal resolution than the input image data.15QB\630666.01650\99677203.2Mayo 2023-546630666.01650
[0065] As another example, the medical image data may include raw medical data, such that the reconstructed medical image data may include raw medical data that has been temporally upsampled (e.g., had its temporal resolution increased, for example, via generation of additional simulated raw frame data for interpolated times), or otherwise have had a temporal aspect of the images improved. From the reconstructed raw data, one or more images can then be reconstructed using any suitable reconstruction technique, such as filtered backproj ection, iterative reconstruction, parallel imaging reconstruction or the like.
[0066] The reconstructed medical image data generated by inputting the medical image data to the trained deep learning model(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 908.
[0067] Referring now to FIG. 10, a flowchart is illustrated as setting forth the steps of an example method for training one or more deep learning models (such as a neural network, tree-based classifier, reservoir computer, etc.) on training data, such that the one or more deep learning models are trained to receive medical image data as input data in order to generate reconstructed medical image data as output data. An example workflow for training a deep learning model to reconstruct MRI image data is also illustrated in FIG. 3.
[0068] In general, the deep learning model(s) can implement any number of different deep learning model architectures. For instance, the deep learning model(s) could implement a convolutional neural network, a residual neural network, a transformer, or the like. As another example, the deep learning model could implement other suitable machine learning or artificial intelligence algorithms, such as those based on supervised learning, unsupervised learning, deep learning, ensemble learning, reinforcement learning, dimensionality7reduction, and so on.
[0069] The method includes accessing training data with a computer system, as indicated at step 1002. In general, the training data can include temporally downsampled training data generated from input medical image data. Additionally or alternatively, the accessed training data can include medical image data received from an example database or a scan from a medical imaging system. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with a medical imaging system and transferring or otherwise communicating the data to the computer system.
[0070] The method can include assembling training data from medical image data using a computer system. This step may include assembling the medical image data into an appropriate data structure on which the deep learning model or other machine learning model16QB\630666.01650\99677203.2Mayo 2023-546630666.01650 can be trained. Assembling the training data may include assembling low-temporal resolution data. For instance, assembling the training data may include generating temporally downsampled medical image data from the medical image data, denoising medical image data, and the like (e.g., such as described with respect to FIGS. 3-5, 7).
[0071] One or more deep learning models are trained on the training data, as indicated at step 1004. In general, the deep learning model can be trained by optimizing parameters (e.g., neural network parameters (e.g., weights, biases, or both)) based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function. As another example, the loss function may comprise a minimax loss function, Wasserstein loss function, etc. For instance, such loss functions may be used for a generative adversarial network or other generative deep learning model.
[0072] Training a deep learning model may include initializing the deep learning model, such as by computing, estimating, or otherwise selecting initial parameters (e.g., parameters such as weights, biases, or both). During training, a deep learning model receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding weights. For instance, training data can be input to the initialized deep learning model, generating output as temporally reduced spatiotemporal medical image data. The deep learning model then compares the generated output with the actual output of the training example in order to evaluate the quality of the reconstructed medical image data. For instance, the reconstructed medical image data can be passed to a loss function to compute an error. The current deep learning model can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current deep learning model can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the loss according to a loss function, such as a mean-squared error, etc, or a combined loss function in a generative adversarial network (GAN), such as the combination of a generator loss and a discriminator loss, etc. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validation iterations being completed, and the like. When the training condition has been met (e.g., by determining whether an error threshold or other stopping criterion has been satisfied), the current deep learning model and its associated parameters represent the trained deep learning model. Different types of training processes can be used to adjust the bias values17QB\630666.01650\99677203.2Mayo 2023-546630666.01650 and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent. Newton’s method, conjugate gradient. quasiNewton, Levenberg-Marquardt, among others.
[0073] The deep learning model can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transfer learning, or other suitable learning techniques for deep learning models. As an example, supervised learning involves presenting a computer system with example inputs and their actual outputs (e.g., categorizations). In these instances, the deep learning model is configured to leam a general rule or model that maps the inputs to the outputs based on the provided example input-output pairs.
[0074] The one or more trained deep learning models are then stored for later use, as indicated at step 1006. Storing the deep learning model(s) may include storing parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the deep learning model(s) on the training data. Storing the trained deep learning model(s) may also include storing the particular deep learning model architecture to be implemented. For instance, data pertaining to layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.
[0075] A test has been performed, described as follows: The test comprised an implementation using TensorFlow / Keras on the CINE CMR DICOM images in the Mayo Clinic’s research database obtained from GE, Siemens, and Philips MRI scanners. Patient data for which the cardiac cycle (e.g., RR-interval duration) was between 700ms to 1200ms were included in the training data. Each subject’s data was normalized to a mean of 1024. After training was completed, testing data were used to check the reconstruction quality. Testing data was acquired on a Philips 1.5T scanner from a healthy volunteer. Six different acquisitions were performed with different temporal resolutions, ranging between 40ms-190ms. FIG. 6. are example images from the testing dataset. Spatial and temporal blurring can be noticed in the dataset acquired with temporal resolution of 125ms. These blurring effects were improved when DL was applied. The DL reconstruction time for the given dataset only took around 20 seconds for the entire volume.
[0076] FIG. 11 shows an example of a system 1100 for reconstructing medical image data in accordance with some embodiments described in the present disclosure. As shown in18QB\630666.01650\99677203.2Mayo 2023-546630666.01650FIG. 11, a computing device 1150 can receive one or more ty pes of data (e.g., MRI image data) from data source 1102. In some embodiments, computing device 1150 can execute at least a portion of a low-temporal resolution image temporal enhancement system 1104 to generate medical image data (e.g., reconstructed MRI image data) from data received from the data source 1102.
[0077] Additionally or alternatively, in some embodiments, the computing device 1150 can communicate information about data received from the data source 1102 to a server 1152 over a communication network 1154, which can execute at least a portion of the low-temporal resolution image temporal enhancement system 1104. In such embodiments, the server 1152 can return information to the computing device 1150 (and / or any other suitable computing device) indicative of an output of the low-temporal resolution image temporal enhancement system 1104.
[0078] In some embodiments, computing device 1150 and / or server 1152 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 1150 and / or server 1152 can also reconstruct images from the data.
[0079] In some embodiments, data source 1102 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data), such as a medical imaging system (e.g., a MRI system, an x-ray CT system, a PET system, etc.), another computing device (e g., a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some embodiments, data source 1102 can be local to computing device 1150. For example, data source 1102 can be incorporated with computing device 1150 (e.g., computing device 1150 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 1102 can be connected to computing device 1150 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 1102 can be located locally and / or remotely from computing device 1150, and can communicate data to computing device 1150 (and / or server 1152) via a communication network (e.g., communication network 1154).
[0080] In some embodiments, communication network 1154 can be any suitable communication network or combination of communication networks. For example, communication network 1 154 can include a Wi-Fi network (which can include one or more19QB\630666.01650\99677203.2Mayo 2023-546630666.01650 wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication network 1154 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG. 1 1 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.
[0081] Referring now to FIG. 12, an example of hardware 1200 that can be used to implement data source 1102, computing device 1150, and server 1152 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.
[0082] As shown in FIG. 12, in some embodiments, computing device 1150 can include a processor 1202, a display 1204, one or more inputs 1206, one or more communication systems 1208, and / or memory 1210. In some embodiments, processor 1202 can be any suitable hardware processor or combination of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and so on. In some embodiments, display 1204 can include any suitable display devices, such as a liquid cr stal display (LCD) screen, a light-emitting diode (LED) display, an organic LED (OLED) display, an electrophoretic display (e.g., an “e- ink" display), a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1206 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0083] In some embodiments, communications systems 1208 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1154 and / or any other suitable communication networks. For example, communications systems 1208 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1208 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0084] In some embodiments, memory’ 1210 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for20QB\630666.01650\99677203.2Mayo 2023-546630666.01650 example, by processor 1202 to present content using display 1204, to communicate with server 1152 via communications system(s) 1208, and so on. Memory 1210 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory71210 can include random-access memory (RAM), read-only memory7(ROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), other forms of volatile memory, other forms of non-volatile memory, one or more forms of semi -volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1210 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 1150. In such embodiments, processor 1202 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 1152, transmit information to server 1152, and so on. For example, the processor 1202 and the memory71210 can be configured to perform the methods described herein (e.g., the method of FIG. 9, the method of FIG. 10).
[0085] In some embodiments, server 1152 can include a processor 1212, a display 1214, one or more inputs 1216, one or more communications systems 1218, and / or memory 1220. In some embodiments, processor 1212 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 1214 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1216 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0086] In some embodiments, communications systems 1218 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1154 and / or any other suitable communication networks. For example, communications systems 1218 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1218 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0087] In some embodiments, memory71220 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for21QB\630666.01650\99677203.2Mayo 2023-546630666.01650 example, by processor 1212 to present content using display 1214, to communicate with one or more computing devices 1150. and so on. Memory 1220 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1220 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1220 can have encoded thereon a server program for controlling operation of server 1 152. In such embodiments, processor 1212 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1150, receive information and / or content from one or more computing devices 1150, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.
[0088] In some embodiments, the server 1152 is configured to perform the methods described in the present disclosure. For example, the processor 1212 and memory' 1220 can be configured to perform the methods described herein (e.g., the method of FIG. 9, the method of FIG. 10).
[0089] In some embodiments, data source 1102 can include a processor 1222, one or more data acquisition systems 1224, one or more communications systems 1226, and / or memory 1228. In some embodiments, processor 1222 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 1224 are generally7configured to acquire data, images, or both, and can include a medical imaging system. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 1224 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of an MRI system. In some embodiments, one or more portions of the data acquisition system(s) 1224 can be removable and / or replaceable.
[0090] Note that, although not show n, data source 1102 can include any suitable inputs and / or outputs. For example, data source 1102 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, datasource 1102 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display , an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.22QB\630666.01650\99677203.2Mayo 2023-546630666.01650
[0091] In some embodiments, communications systems 1226 can include any suitable hardware, firmware, and / or software for communicating information to computing device 1150 (and, in some embodiments, over communication network 1154 and / or any other suitable communication networks). For example, communications systems 1226 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1226 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0092] In some embodiments, memory’ 1228 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1222 to control the one or more data acquisition systems 1224, and / or receive data from the one or more data acquisition systems 1224; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 1150; and so on. Memory 1228 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1228 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other ty pes of non-volatile memory7, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1228 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 1102. In such embodiments, processor 1222 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1150, receive information and / or content from one or more computing devices 1150, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.
[0093] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory or non-transitory. For example, non-transitory7computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory7, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during23QB\630666.01650\99677203.2Mayo 2023-546630666.01650 transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0094] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms 'component. ” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).
[0095] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.
[0096] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the disclosure.24QB\630666.01650\99677203.2
Claims
Mayo 2023-546630666.01650CLAIMS1. A method, comprising: obtaining a spatiotemporal medical image: processing the spatiotemporal medical image to generate a set of two-dimensional (2D) spatiotemporal data; applying a 2D deep learning model to the 2D spatiotemporal data to generate a set of transformed 2D spatiotemporal data; and generating a transformed spatiotemporal medical image based on the set of transformed 2D spatiotemporal data.
2. The method of claim 1 , wherein the 2D deep learning model comprises a convolutional neural network.
3. The method of claim 1, wherein the 2D deep learning model comprises a cascade of individual deep learning models.
4. The method of claim 3. wherein at least one individual deep learning model of the cascade of individual deep learning models comprises a convolution network, a U-net, a densely connected network, a generative Al model, a fully connected network, a recursive network, a residual net, a transformer, an attention module, generative adversarial model, inception module, or reservoir computer.
5. The method of claim 1. further comprising: processing the transformed spatiotemporal medical image to generate a set of 2D spatial image slices; applying a second 2D deep learning model to the 2D spatial image slices to generate a set of transformed 2D spatial image slices; and generating a second transformed spatiotemporal data based on the set of transformed 2D spatial image slices.
6. The method of claim 5, wherein the second 2D deep learning model comprises a convolutional neural network.25QB\630666.01650\99677203.2Mayo 2023-546630666.016507. The method of claim 5. wherein the 2D deep learning model comprises a cascade of individual deep learning models.
8. The method of claim 7, wherein at least one individual deep learning model of the cascade of individual deep learning models comprises a convolution network, a U-net, a densely connected network, a generative Al model, a fully connected network, a recursive network, a residual net, a transformer, an attention module, generative adversarial model, inception module, or reservoir computer.
9. The method of claim 1, wherein the spatiotemporal medical image is a second spatiotemporal medical image obtained by denoising a first spatiotemporal medical image.
10. The method of claim 1, wherein the spatiotemporal medical image comprises a DICOM (Digital Imaging and Communications in Medicine) standard image.
11. The method of claim 1. wherein the transformed spatiotemporal medical image has a higher temporal resolution than the spatiotemporal medical image.
12. A method, comprising: obtaining a target spatiotemporal medical image; processing the target spatiotemporal medical image to generate a simulated spatiotemporal medical image having a lower temporal resolution than the target spatiotemporal medical image; processing the simulated spatiotemporal medical image to generate a set of training two-dimensional (2D) spatiotemporal data; processing the target spatiotemporal medical image to generate a set of target 2D spatiotemporal data; applying a 2D deep learning model to generate a set of training temporally upsampled 2D spatiotemporal data based on the training 2D spatiotemporal data; and updating the 2D deep learning model based on comparing the set of training temporally upsampled 2D spatiotemporal data to the set of target 2D spatiotemporal data to generate a trained 2D deep learning model.26QB\630666.01650\99677203.2Mayo 2023-546630666.0165013. The method of claim 12, further comprising: processing the training temporally upsampled 2D spatiotemporal data to generate a set of training 2D spatial image slices; processing the target spatiotemporal medical image to generate a set of target 2D spatial image slices; applying a second 2D deep learning model to the training 2D spatial image slices to generate a set of spatially correlated 2D spatial image slices; and updating the second trained 2D deep learning model based on comparing the set of spatially correlated 2D spatial image slices to the set of target 2D spatial image slices to generate a second trained 2D deep learning model.
14. The method of claim 12, further comprising: processing the training temporally upsampled 2D spatiotemporal data to generate a training three-dimensional (3D) spatiotemporal data; applying a 3D deep learning model to the training 3D spatiotemporal data to generate a spatially correlated 3D spatiotemporal data; and updating the 3D deep learning model based on comparing the training 3D spatiotemporal data to the target spatiotemporal medical image to generate a trained 3D deep learning model.
15. The method of claim 12, further comprising: processing the training temporally upsampled 2D spatiotemporal data to generate a training four-dimensional (4D) spatiotemporal data; applying a 4D deep learning model to the training 4D spatiotemporal data to generate a spatially correlated 4D spatiotemporal data; and updating the 4D deep learning model based on comparing the training 4D spatiotemporal data to the target spatiotemporal medical image to generate a trained 4D deep learning model.
16. The method of claim 12, further comprising: obtaining an input spatiotemporal medical image; processing the input spatiotemporal medical image to generate a set of 2D spatiotemporal data;27QB\630666.01650\99677203.2Mayo 2023-546630666.01650 applying the trained 2D deep learning model to the set of 2D spatiotemporal data to generate a set of upsampled 2D spatiotemporal data: and generating a reconstructed temporally upsampled spatiotemporal medical image for the input spatiotemporal medical image based on the set of upsampled 2D spatiotemporal data.
17. The method of claim 12, wherein the 2D deep learning model comprises a convolutional neural network.
18. The method of claim 12, wherein the 2D deep learning model comprises a cascade of individual deep learning models.
19. The method of claim 12, wherein the target spatiotemporal medical image is a second target spatiotemporal medical image obtained by denoising a first target spatiotemporal medical image.
20. The method of claim 12, wherein the spatiotemporal medical image comprises a DICOM (Digital Imaging and Communications in Medicine) standard image.
21. A method, comprising: obtaining a spatiotemporal medical image: processing the spatiotemporal medical image to generate a set of two-dimensional (2D) spatiotemporal data; applying a 2D deep learning model to temporally upsample the 2D spatiotemporal data to generate a set of upsampled 2D spatiotemporal data; and generating a reconstructed temporally upsampled spatiotemporal medical image based on the set of upsampled 2D spatiotemporal data.
22. The method of claim 21, wherein the 2D deep learning model comprises a convolutional neural network.
23. The method of claim 21, wherein the 2D deep learning model comprises a cascade of individual deep learning models.28QB\630666.01650\99677203.2Mayo 2023-546630666.0165024. The method of claim 23, wherein at least one individual deep learning model of the cascade of individual deep learning models comprises a convolution network, a U-net, a densely connected network, a generative Al model, a fully connected network, a recursive network, a residual net, a transformer, an attention module, generative adversarial model, inception module, or reservoir computer.
25. The method of claim 21, further comprising: combining the set of upsampled 2D spatiotemporal data to generate a set of 2D spatial image slices; applying a second 2D deep learning model to the 2D spatial data to generate a set of spatially correlated 2D spatial image slices; and generating the reconstructed spatiotemporal image based on the set of spatially correlated 2D spatial image slices.
26. The method of claim 25, wherein the second 2D deep learning model comprises a convolutional neural network.
27. The method of claim 26, wherein the convolutional neural network comprises a cascade of individual 2D deep learning models.
28. The method of claim 27, wherein at least one individual deep learning model of the cascade of individual deep learning models comprises a convolution network, a U-net, a densely connected network, a generative Al model, a fully connected network, a recursive network, a residual net, a transformer, an attention module, generative adversarial model, inception module, or reservoir computer.
29. The method of claim 21, wherein the spatiotemporal medical image is a second spatiotemporal medical image obtained by denoising a first spatiotemporal medical image.
30. The method of claim 21, wherein the spatiotemporal medical image comprises a DICOM (Digital Imaging and Communications in Medicine) standard image.29QB\630666.01650\99677203.2