Systems and methods for unsupervised learning for inverse problems using compressibility

The described method trains PD-DL models using a sparsity and consistency term-based loss function on clinical images, addressing the lack of raw k-space data availability in MRI reconstruction, achieving high-quality results across diverse clinical environments.

WO2025222143A1PCT designated stage Publication Date: 2025-10-23REGENTS OF THE UNIVERSITY OF MINNESOTA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025406
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-18
Filing Date
2025-04-18
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing deep learning methods for MRI reconstruction, such as physics-driven deep learning (PD-DL), require access to raw k-space data, which is often not available outside specialized centers, limiting their application in clinical settings.

Method used

A system and method for training PD-DL models using a loss function that combines a sparsity term with a consistency term, allowing training on routine clinical images without raw k-space data, utilizing a reweighted ℓ1 norm and parallel imaging fidelity to ensure non-zero output sparsity and data fidelity.

Benefits of technology

Enables high-quality MRI reconstruction in various clinical settings without raw k-space data, outperforming conventional methods and maintaining image quality comparable to supervised and unsupervised techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025406_23102025_PF_FP_ABST
    Figure US2025025406_23102025_PF_FP_ABST
Patent Text Reader

Abstract

In some aspects, the present disclosure provides a computer-implemented method for training a machine learning model to reconstruct an image from subsampled data. The method includes accessing training data, which includes subsampled medical imaging data, with a computer system. The method further includes training a machine learning model on the training data. Training the model includes minimizing a loss function that includes a sparsity term and a consistency term. The sparsity term includes a reweighted norm in a sparsifying transform domain, and the consistency term ensures a non-zero output when the sparsity is minimized. The method further includes storing the machine learning model in the computer system for later use.
Need to check novelty before this filing date? Find Prior Art

Description

UMN 2024-231 920171.00646 SYSTEMS AND METHODS FOR UNSUPERVISED LEARNING FOR INVERSE PROBLEMS USING COMPRESSIBILITY CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 635,935, filed on April 18, 2024, and entitled “SYSTEMS AND METHODS FOR UNSUPERVISED LEARNING FOR INVERSE PROBLEMS USING COMPRESSIBILITY,” which is herein incorporated by reference in its entirety. STATEMENT OF FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under HL153146, EB027061, and EB032830 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND

[0003] Several deep learning methods have been proposed for reconstructing subsampled data, such as undersampled magnetic resonance imaging (MRI) data. For example, physics-driven deep learning (PD-DL) methods have gained popularity for improved reconstruction of fast MRI scans. While supervised learning has been used in early works, there remains a need to developing unsupervised learning methods for training PD-DL. Moreover, flexible methods that do not require the use of raw data are desirable. SUMMARY OF THE DISCLOSURE

[0004] The present disclosure addresses the aforementioned drawbacks by providing a system and method for training and applying a machine learning model for solving inverse problems. A loss function is described that can be flexibly applied in a variety of learning scenarios, including supervised learning, unsupervised learning, zero-shot learning, and when raw data are not available.

[0005] In some aspects, the present disclosure provides a computer-implemented method for training a machine learning model to reconstruct an image from subsampled data. The method includes accessing training data, which includes subsampled medical imaging data, with a computer system. The method further includes training a machine learning model on the training data. Training the model includes minimizing a loss function that includes a 1 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 sparsity term and a consistency term. The sparsity term includes a reweighted norm in a sparsifying transform domain, and the consistency term ensures a non-zero output when the sparsity is minimized. The method further includes storing the machine learning model in the computer system for later use.

[0006] In other aspects, a computer-implemented method for training a machine learning model to solve inverse problems is presented. The method includes accessing training data using a computer system and training a machine learning model using the training data. Training includes minimizing a loss function, which includes a consistency term and a reweighted norm in a sparsifying domain. The method further includes storing the trained machine learning model in the computer system for later use.

[0007] In still other aspects of the present disclosure, a method for reconstructing an image from undersampled k-space data is provided. The method includes using a computer system to access a pre-trained neural network that has been trained on subsampled k-space data using a loss function. The loss function includes a consistency term and a sparsity term that includes a sum of a reweighted norm in a sparsifying domain over ^^ transform domain coefficients. The method further includes accessing undersampled k-space data obtained from a subject using a magnetic resonance imaging (MRI) system. The method further includes inputting the undersampled k-space data to the pre-trained neural network and generating output as a reconstructed image that depicts the subject. The method also includes displaying the image to a user using the computer system. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIGS.1A and 1B illustrate an example workflow for a Compressibility-inspired Unsupervised Learning via Parallel Imaging Fidelity (CUPID) method for training physics- driven deep learning (PD-DL) models in an unsupervised and / or zero-shot manner without requiring any raw k-space data. The network is unrolled for T units, with each unit including aregularizer (R) and data fidelity (DF). The loss function includes two terms: a reweighted ^ 1component that assesses the compressibility of the network’s output (FIG.1A); and a fidelity term that ensures the output stays consistent with parallel imaging reconstructions via carefully designed perturbations, thereby avoiding the network from producing a sparse all-zeros output (FIG.1B).

[0009] FIGS.2A–2H illustrate example perturbation images. Added perturbations may include precisely positioned letters that have the same intensity in within each shape (FIG.2A), 2 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 randomly positioned letters (FIG. 2B), circles or other shapes that have different intensities (FIG.2C), or randomly rotated card suits (FIG.2D. Furthermore, when accelerated by R = 4, even without ACS data, whose corresponding zero-filled image is shown in FIG. 2E, these perturbations can be resolved via parallel imaging methods such as CG-SENSE: 20 iterations (FIG.2F), 40 iterations (FIG.2G), 80 iterations (FIG.2H).

[0010] FIG.3A shows a schematic summarizing a training procedure that may be used in accordance with some embodiments described in the present disclosure.

[0011] FIG. 3B demonstrates an example image pattern than can be described as a perturbation image in accordance with some embodiments described in the present disclosure.

[0012] FIG.4 provides a flowchart setting forth steps of an example process that may be used to train a machine learning model in accordance with the present disclosure.

[0013] FIG.5 provides a flowchart setting forth steps of an example process that may be used to apply a trained machine learning model in accordance with the present disclosure.

[0014] FIG. 6 shows example images reconstructed in accordance with the present disclosure for varying levels of tuning parameter λ.

[0015] FIG. 7 shows example images reconstructed using known methods and in accordance with the present disclosure.

[0016] FIG. 8 shows example data comparing image metrics achieved using the disclosed method with those of known methods.

[0017] FIG.9 shows example images in which the disclosed method was applied in a zero-shot scenario in comparison with alternative methods.

[0018] FIG.10 shows representative slices reconstructed at an acceleration factor of R = 4 using equidistant undersampling from coronal PD and coronal PD-FS knee MRI, as well as axial T2-weighted brain MRI. The baseline CG-SENSE, CS, EI-trained PD-DL, and DDS suffer from residual artifacts highlighted by red arrows. PD-DL trained with CUPID improves upon them while delivering reconstruction quality comparable to supervised and SSDU-trained PD-DL.

[0019] FIG. 11 shows representative subject-specific / zero-shot learning results for various algorithms on coronal PD and coronal PD-FS knee MRI for retrospective R = 4 equidistant undersampling, along with results from database-trained diffusion models on knee data. Baseline CGSENSE, CS, ScoreMRI and DDS suffer from residual artifacts (red arrows) and blurring. PD-DL with CUPID loss successfully removes these artifacts, and functions in a similar manner to ZS-SSDU. 3 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0020] FIGS. 12A–12D show example prospective acceleration results for various methods that operate on parallel imaging-reconstructed magnitude-only DICOM images exported from the scanner. Vendor-provided PI DICOM (R = 4) (FIG.12A), CS (FIG.12B), DDS (FIG.12C), and CUPID (FIG.12D). CS reconstruction suffers from visible blurring. DDS reduces artifacts and noise, but both are still present. CUPID achieves effective artifact and noise mitigation with the highest-quality reconstruction.

[0021] FIG. 13 shows example results from SSDU and supervised PD-DL models trained on 3T coronal PD data, which show noticeable artifacts when applied to 1.5T data with a lower SNR. DDS, trained on 1.5T data as well as 3T data, better mitigates these artifacts, although some performance degradation remains. CUPID, zero-shot fine-tuned for the target population without raw data access, achieves the best performance with high image quality and significantly reduced artifacts.

[0022] FIG. 14 is a block diagram of an example training and image reconstruction system that can implement the methods of the present disclosure.

[0023] FIG. 15 is a block diagram of example components that can implement the system of FIG.14. DETAILED DESCRIPTION

[0024] Before any aspects of the present disclosure are explained in detail, it is to be understood that the invention is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings. The invention is capable of other embodiments and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings.

[0025] The present disclosure relates to systems and methods for training machine learning models to solve inverse problems, such as physics-driven deep learning (PD-DL) models for magnetic resonance imaging (MRI) reconstruction. In particular, the disclosure 4 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 describes techniques for training PD-DL networks using routine clinical reconstructed MR images, without requiring access to raw k-space data.

[0026] In some cases, PD-DL approaches may be used to improve reconstruction of fast MRI scans. However, conventional PD-DL training methods typically require access to raw k-space measurements, which may not be available outside of specialized MRI centers. While all MRI scanners inherently sample k-space data, the majority of scanners outside specialized academic or research institutions lack the ability to export this data due to vendor restrictions. As a result, most DL-based reconstruction methods, which typically rely on raw k-space during training or fine-tuning, cannot be applied in these settings.

[0027] The techniques described herein enable training of PD-DL reconstruction models using only the final reconstructed images that are routinely exported from MRI scanners, such as Digital Imaging and Communications in Medicine (DICOM) format images. By eliminating the need for raw k-space data during training, the disclosed techniques may allow PD-DL reconstruction to be applied more broadly, including at facilities that do not have access to raw k-space data for training. This may enable improved fast MRI reconstruction in a wider range of clinical settings.

[0028] The present disclosure provides systems and methods for solving inverse problems using a machine learning model or model that can be trained using an unsupervised approach. Such methods can be applied for several inverse problems, including MRI image reconstruction (e.g., undersampling, parallel imaging, denoising, super resolution, and so forth). For example, the described systems and methods can be used to advantageously achieve higher acceleration or undersampling rates than standard parallel imaging techniques. Advantageously, described systems and methods can be applied even when the raw data (e.g., k-space) are not available. In some aspects, the described methods can also be applied to solve other inverse problems, such as denoising, deblurring, inpainting, and others. In other aspects, the present disclosure provides methods for training physics-driven deep learning models.

[0029] As will be described, a machine learning model can be trained using a loss function that is defined using a sparsity-inspired loss. For example, the loss can be based on a reweighted ℓ1norm in a sparsifying transform domain. The loss may further include a consistency term that can be used to ensure a non-zero output when the sparsity is minimized. In this way, a novel convex loss function is proposed as an alternative learning strategy. For example, in the context of MRI data reconstruction, the loss function can evaluate the compressibility of the output image while ensuring data fidelity to assess the quality of 5 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 reconstruction in versatile settings, including supervised, unsupervised, and zero-shot scenarios, including when raw k-space data are not available. In particular, the described methods leverage the reweighted ℓ1norm, which has been shown to approximate the ℓ0norm for quality evaluation. In an example application described below, results show that the PD- DL networks trained with the proposed loss formulation outperform conventional methods, while maintaining similar quality to PD-DL models trained using existing supervised and unsupervised techniques.

[0030] Another advantage of the described approach is the ability to train a machine learning model without access to raw data. For example, the network can be trained based on an image estimate (e.g., DICOM images available from an MRI system), even when the image estimate is sub-optimal. This allows for a flexible method that can be applied in a wider range of settings. For example, such DICOM images can be accessed from standard clinical MRI systems, hospital databases, and historical data, whereas raw data (e.g., k-space) are often not stored or accessible. This can provide a larger and more complete training database.

[0031] An example for MRI is described below, but it will be appreciated that the systems and methods described in the present disclosure can be adapted for training machine learning models for solving inverse problems in other applications, including image reconstruction in other imaging modalities, data denoising, and so forth.

[0032] PD-DL reconstruction has gained importance for MRI reconstruction. Typical methods are trained in a supervised manner, which required the availability of fully-sampled k-space data to be used as reference labels. Noting the difficulty or impossibility of obtaining fully-sampled data in various practical MRI scenarios, several alternative unsupervised learning approaches have been proposed, including self-supervised learning or generative modeling. In the former, a common approach, as in self-supervision via data undersampling (SSDU), masks parts of the acquired k-space, and trains the network to learn these from the remainder of acquired k-space points. As one example, SSDU is described further in co- pending US Patent Application Publication No.2021 / 0118200, the entire contents of which is incorporated herein by reference. In the latter, a generative model is employed to estimate the prior distribution of images. This evaluation is then utilized in combination with a log- likelihood data term during the testing phase.

[0033] The present disclosure describes an alternative method for training PD-DL networks in an unsupervised manner. The approach combines concepts from statistical image processing and compressed sensing (CS). Specifically, the described methods continue to 6 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 leverage the effectiveness of non-linear low-dimensional representations provided by PD-DL networks for image reconstruction while also assessing the quality of these reconstructions using compressibility. The disclosed methods can be applied in supervised, unsupervised, and zero-shot settings.

[0034] CS seeks to find the sparsest / most compressible image in a fixed linear transform domain that is consistent with the acquired measurements. Conventional CS solves a regularized ℓ1 minimization problem, which is the convex relaxation of the constrained ℓ0 minimization problem:

[0035] where W is a sparsifying transformation, such as a wavelet transform, finitedifferences, or others; y ^ is the acquired k-space data; E ^ is the multi-coil encoding operatorthat samples k-space locations, ^ ; x is the image of interest; and ^ is a (e.g., empirically)tunable parameter used to control the sparsity requirement. Equation (1) may be solved using an iterative optimization algorithm; however, at high acceleration rates, especially with uniform undersampling, CS reconstruction has difficulties that lead to residual artifacts.

[0036] A PD-DL reconstruction can be used as an alternative to the CS-based approach described by Equation (1). The inverse problem for MRI reconstruction can be given by a regularized least squares objective function as:

[0037] where the first quadratic term imposes data fidelity with sampledmeasurements, while R ^^ ^ is a regularizer.

[0038] A commonly used approach for PD-DL involves unrolling an iterative algorithm, such as proximal gradient descent or variable splitting, for a fixed number of steps.Then through end-to-end training, the network learns both the proximal operator forRandthe associated weights for data fidelity. In the supervised learning setup, this is performed by comparing the network outputs with a fully-sampled reference data. In the aforementioned self-supervised learning with SSDU, this is performed without fully-sampled data by dividing ^into two disjoint subsets: ^ and ^ . In these cases, y ^ is used in the data fidelity portion ofthe unrolled network, while y ^ , which is hidden from the network, is used to calculate the lossin k-space. 7 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0039] Among different PD-DL methods, unrolled networks remain the highest performer in reconstruction challenges. These methods unroll iterative algorithms for solving the regularized least squares objective in Equation (2), such as proximal gradient descent or variable splitting with quadratic penalty (VS-QP), over a fixed number of steps. Unrolled networks are conventionally trained using supervised learning over a database, where thereference raw k-space measurements are first retrospectively undersampled to form y ^ .Subsequently, the network is trained to map to the original full reference k-space or the corresponding reference image by minimizing:

[0040] where ^ are the network parameters, f ^y^, E ^ ; ^ ^ denotes the networkoutput for inputs y ^ and E ^ , E full is the fully sampled encoding operator, y ref is the fully-sampled reference k-space data, and L ^^, ^ ^ is a loss function.

[0041] Obtaining fully-sampled reference data in MRI can be infeasible due to prolonged scan durations, organ movement in acquisitions such as real-time cardiac imaging or myocardial perfusion, or signal decay in acquisitions like diffusion MRI with EPI. To enable training of PD-DL networks without fully sampled raw MRI data, a variety of unsupervised learning can be used, including self-supervised learning techniques and generative modeling approaches. Self-supervised methods generate supervisory labels from the undersampled data itself. One example of a self-supervised learning approach is the SSDU technique describedabove. As noted above, SSDU involves partitioning the acquired measurement indices ^ intotwo disjoint subsets ^^ ^ ^^ ^ ^ to train the network in a self-supervised manner:m(4).

[0042] Even though these self-supervision based approaches demonstrate exceptional performance across various tasks, they lack the ability to train the model without access to undersampled raw data, as they cannot operate solely using images that are exported from the scanner.

[0043] It is an advantage of the systems and methods described in the present disclosure that PD-DL models can be trained utilizing only routinely available clinical images exported directly from MRI scanners. In some cases, a compressibility inspired loss is adapted and augmented with a parallel imaging fidelity to successfully reconstruct clinical images in 8 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 DICOM format without needing any raw k-space data. Such a framework may be referred to as Compressibility-inspired Unsupervised Learning via Parallel Imaging Fidelity (CUPID) for high-quality PD-DL training using only routine clinical reconstructed images exported from an MRI scanner.

[0044] Compressibility and / or sparsity in the output of the PD-DL network can beenforced by utilizing a weighted1 norm, which can provide a closer approximation to the ^ 0norm compared to the standardnorm. Thus, this compressibility of the output image inCUPID may be achieved by the loss term:

[0045] where x PI denotes the DICOM input acquired using parallel imaging, Wrepresents the wavelet transform or another suitable sparsifying transform, N is the totalnumber of wavelet coefficients (or other sparsifying transform coefficients), andx^m^signifiesthe signal estimate following the training during the m th reweighting step. As a non-limitingexample, the initial weights may be chosen from a CS reconstruction that has a largeregularization and ^ may be added for numerical stability. Note that f ^ ^, ^ ^ is defined withoutthe network parameters, ^ , and x is ^PI used as the network input instead of y to simplifynotation.

[0046] Relying solely on Equation (5) may result in inaccurate training as the neural network learns to produce an all-zeros image in an effort to drive the wavelet coefficients in the numerator to zero, which minimizes the loss function in Equation (5). To address this, a fidelity operator that stabilizes the training of the reconstruction algorithm, building on ideas from parallel imaging, can be introduced.

[0047] Specifically, the network outputs are ensured to be consistent with any clinicalparallel imaging reconstruction through carefully crafted perturbations, ^pk ^ . Theseperturbations for R-fold acceleration are designed in such a way that R-fold aliasing do not create overlaps in the field-of-view, indicating that they could be resolved by parallel imaging reconstruction. The idea behind this design choice is to ensure that the network, when appliedto the unperturbed x PI , yields an accurate estimate of x , and when applied to x PI ^ p ,9 QB\920171.00646\95951430.1UMN 2024-231 920171.00646similarly recovers x ^ p , as the perturbation p must be resolvable within the framework ofany parallel imaging approach. Both processes are visualized in FIGS.1A and 1B. By doing so, the consistency term ensures a non-zero output when the sparsity is minimized. Thus, the second loss term that enforces parallel imaging fidelity may be given as:

[0048] whereFrom an implementationperspective, the expectation over p is calculated over K such perturbations, ^pk ^ . The fold-over constraint for each ^pk ^ may be achieved by picking the perturbations as randomlyrotated and positioned alphanumeric characters (e.g., letters, numbers), card suits, other symbols, or other shapes that have different intensity values. These choices also ensure that high-frequency information, such as edges, are accurately reconstructed by the regularization process. Some example perturbation images that can be used in the parallel imaging fidelity loss are illustrated in FIG.1B. Additional example perturbation images are illustrated in FIGS.2A–2H. Using a single pattern (K ^ 1 ) does not capture the true mean and exhibits artifacts.As more perturbations are introduced, artifacts and noise amplification are reduced. At a certain point, increasing the number of perturbations becomes counterproductive, yielding only marginal gains while significantly increasing the computation time. It was observed that usingsix distinct p k patterns (K ^ 6 ) offered a reasonable trade-off between computation time andreducing artifacts and noise amplification.

[0049] Combining the compressibility and parallel imaging fidelity loss terms results in the final loss function for CUPID, which may be expressed as:

[0050] where ^ is a trade-off parameter between two terms. The effect of the trade-offparameter, ^ , in CUPID by training was evaluated in an example study by training five distinctPD-DL networks using λ ∈ {0, 50, 100, 200,∞}. Using λ= 0 corresponds to using only the compressibility term (Equation (5)), whereas using λ → ∞ translates to using solely the parallelimaging fidelity term (Equation (6)). Using only the compressibility loss,leads tooverly-smooth reconstructions due to network forcing the wavelet coefficients towards zero without maintaining consistency with the data. On the other hand, solely using the parallel 10 QB\920171.00646\95951430.1UMN 2024-231 920171.00646imaging fidelity loss, L pif , results in reconstructions where the network overfits the datawithout any regularization, resulting in noise amplification. CUPID with λ ∈ {50, 100, 200} combines both loss terms to attain high-fidelity reconstructions. It was observed that CUPID demonstrates robust performance across a wide range of λ values, provided that λ is chosen within a reasonable range that avoids weighting substantially only the compressibility or parallel imaging fidelity loss terms.

[0051] In resource-limited or underserved settings, it may be more practical to fine- tune the method using only a few subjects, or even a single subject, to significantly reduce computational costs. As Equation (5) does not solely focus on the subtraction between two entities, it lacks an inherent mechanism to drive the loss to zero through overfitting. Therefore, CUPID can be tailored to suit a scan-specific context without any modification to the loss given in Equation (7).

[0052] In still other cases, other loss functions may be used within the framework described in the present disclosure. These additional loss functions will be described herein in the context of several learning scenarios, including supervised learning, unsupervised learning, learning without access to raw (e.g., k-space) data, and scan-specific or zero-shot learning. In general the loss function includes a sparsity term that includes a reweighted norm and promotes sparsity of the output in a transform domain. Such sparsity term can include a sum of the reweighted norm in the sparsifying transform domain over some number (N) of transform domain coefficients. The reweighted norm can be weighted based on a ratio of output imaging data and baseline imaging data. As non-limiting examples, the baseline imaging data may include fully-sampled imaging data, which may be referred to as reference imaging data. As another non-limiting example, the baseline imaging data may include an imaging estimate that can be updated for a number of epochs. Such imaging estimate may include an image reconstructed from undersampled data (e.g., using compressed sensing, parallel imaging, or another reconstruction technique). If access to k-space is not available, the imaging estimate can be formed based on a linear reconstruction of the undersampled k-space data.

[0053] The loss function also includes a consistency term that can be used in addition to the data fidelity that is enforced using a physics driven approach (e.g., via unrolled neural network). The consistency term ensures a non-zero output of the machine learning model while the sparsity term is minimized during training. Such consistency term may include a norm (e.g., ℓ1, ℓ1, or other) of a difference between acquired raw data and output data. For example, the acquired data may include raw data (e.g., k-space) that can be reconstructed using multi-coil 11 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 sensitivity data. In other implementations, the consistency term can be configured to generally ensure consistency in parallel imaging for DICOM imaging data or other reconstructed image data, as will be further described below.

[0054] The present disclosure introduces a loss function aimed at promoting sparsity of the reconstruction in a transform domain under which MR images are compressible. The approach can be first motivated in the supervised scenario, where the reference or fully- sampled imaging data is available. To this end, a weighted ℓ1 norm can be used, which has been shown to approximate the ℓ0 norm better than the ℓ1 norm alone:

[0055] where x୰^^corresponds to baseline imaging data, which may be a fully-sampled reference image in the supervised setting, x୭^^corresponds to the output imaging data in the form of a reconstructed image, ^^ is the sparsifying transformation, and ^^^x^^denotes the ^^thtransform-domain coefficient. ^^ is the total number of transform-domain coefficients, which is known based on the image size. ^^ is added to ensure numerical stability for the division. In some implementations, ^^ can be chosen empirically. As non-limiting examples, ^^ may be set between 1 / 100thand 1 / 10thof the average pixel intensity in the imaging data. λ is a tunable hyper-parameter that can control a tradeoff between blurring and noise. λ can be tuned empirically and may be set between 20-200 as a non-limiting example. In this way, the first term of Equation (8) provides a sparsity term that includes a reweighted norm in a sparsifying transform domain. Note that both terms in Equation (8) are normalized to ensure that the loss function is unitless.

[0056] The present disclosure recognizes that even though the PD-DL network has data fidelity in its structure, the second term in Equation (8), which may be referred to as a consistency term, can be used to prevent the network output from learning to output an all-zeros image. This is because if the network is trained purely on the reweighted ^ 1 term (e.g.,the first term in Equation (8)), then the loss will be minimized by an all-zeros image. Thus, the proximal operators in the PD-DL network will learn to output zero, while simultaneously pushing data fidelity weights for the regularization term to infinity, to force the final network output to be the all-zeros image. The second term in Equation (8) can ensure this does not happen and is in addition to the data fidelity already built into the unrolled network.

[0057] The supervised formulation in Equation (8) suggests that the reweighted ℓ1 norm minimization can be modified to an unsupervised learning setup. In the disclosed 12 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 approach, the training can start from an initial estimate x^^^of the signal that is used as the denominator weights ൫^^x^^^൯^. The network can be trained with these weights for a fixed number of epochs. The number of epochs may be tuned empirically. As a non-limiting example, the number of epochs may be set to 100 or between 5 and 1000. Then the network outputs, x^^^, can be used to generate the new weights in the loss function, similar to reweightedℓ1 norm minimization. This leads to the following reweighted loss function for the ^^^ ^ 1^threweighting:

[0058] where x^୩^represents the signal estimate after the training in the ^^threweighting step. Initial weights x^^^can be obtained using a CS reconstruction with a large regularization term. In this way, the baseline imaging data does not require fully sampled data. Rather, the baseline imaging data can include an image estimate reconstructed from undersampled data, which may be updated for a fixed number off epochs.

[0059] The proposed reweighted training procedure is summarized in FIG. 1A. The first term is a reweighted ℓ1component designed to evaluate the compressibility of the network output, while the second term is a consistency term that prevents the network to output 0.

[0060] Similar to CS reconstruction with reweighted ℓ1 norm minimization, the unsupervised learning loss described with respect to Equation (9) can be adapted to a scan- specific setup. Since the first term in Equation (9) does not merely work on the subtraction between two entities, there is no inherent mechanism to drive the loss to 0 via overfitting. Thus, this loss function can also be used for scan-specific or zero-shot learning, which may be of interest in scenarios where it is difficult to curate a large database or where generalizability problems are observed with out-of-distribution samples. In contrast, previous work on zero- shot learning using mask-based self-supervised learning employed an explicit self-validation mechanism to avoid overfitting.

[0061] Equation (9) can further be modified for the setting in which raw data yஐis not available. Instead a reconstructed image can be used. For example, a parallel imaging solutioncan be used, which can be described as x^ூwhere^∙^ுis the Hermitianoperator. Without access to yஐ, the second consistency term of Equation (9) needs to be redefined. To do so, it is useful to think of an example in which parallel imaging will be successful. For example, consider recovering the image provided in FIG. 1B using an acceleration of R = 3. The true image (as shown in FIG.1B) can be recovered because the R = 13 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 3 fold-overs (shown as dotted lines) will not create overlap where the pattern folds over onto itself. Many patterns that satisfy the same conditions can be randomly or pseudo-randomly generated for any acceleration rate R. Let p be one such pattern and let qஐbe the raw datacorresponding to the pattern p. Reasonably, the pattern can be described as p ^In this way, p defines a perturbation image that can be undersampled with some acceleration rate R without overlap of the image content.

[0062] The present disclosure recognizes that a deep learning model should accomplish the same. For example, if yஐis input and the DL model provides output x୭^^, then input yஐ^q should le ୭^^ஐ ad to output x ^ p. With this insight, a new loss can be defined asarg m^൫^^^ ^^ ,^^ ^൯ ^ఏin^∑ே^ ಐ ಈ ಈ ^^ ^^^ ^ ^ λ ∙ฮ൫^ಐ^^ಈା୯ಈ, ^^ಈ^ି^ಐ^^ಈ, ^^ಈ^൯ି୮ฮభே^ୀ^ ൫^ౡ^൯^^ାఢ‖୮‖భ(10);

[0063] or arg^^^ ^^ ^ ^^ ^^ ^ ^^ಈ, ^^ ^మ(11)

[0064] Thus, the consistency term includes a norm (e.g., ℓ1or ℓ2) of a differencebetween an estimate of a perturbation image (e.g., ൫^^^^yஐ ^ qஐ, ^^ஐ^ െ ^^^^yஐ, ^^ஐ^൯) and theperturbation image (e.g., p) that was defined. The estimate of the perturbation image compares the output function of an image representing the raw k-space data (e.g., yஐ) and the outputfunction of an image representing the perturbed raw k-space data (e.g., yஐ ^ qஐ). However,such image representation does not require direct access to the raw k-space data, as will be described. The perturbation term is the perturbation image, p. Because p should be fully recoverable using parallel imaging reconstruction techniques, the reconstruction term should be equivalent to the perturbation image itself, thereby minimizing the consistency term, when the reconstruction output function (^^^) is optimized.

[0065] It is also contemplated that, in addition to ℓ1 (e.g., as in Equation (10)) and ℓ2 (e.g., as in Equation (11)) norms, L-infinity (ℓ∞), p-norms (Lp), perceptual loss, other norms or variations on such norms can be used.

[0066] Equations (10) and (11) and variations thereof can also be modified in order to average over several different patterns of p, which can be described as ^p^^^ୀ^. For example, several patterns that do not contain overlap of imaging content in the undersampled image can be randomly generated. Such averaging can provide a more robust solution. Moreover, the example image, p, can include additional structure or shapes. As a non-limiting example, p 14 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 could include randomly rotated and positioned letters, numbers, or other shapes, which provide sharper images for training the reconstruction.

[0067] Advantageously, the loss defined in Equations (10) and (11) and variations thereof can flexibly be used both when the raw data is available and when raw data is not available. For example, if raw k-space data (yஐ) is available, the inputs of ^^^can be defined asthe undersampled or zero-filled k-space data (yஐ ^ qஐ or yஐ) transformed to the image domainby encoding matrix (^^ஐ), which incorporates multi-coil sensitivity. However, if the k-space data is not available, the inputs of ^^^can be defined as the corresponding zero-filled images (e.g., ^^ஐுyுஐ, and ^^ஐyஐ ^ ^^ஐுqஐ).

[0068] Even more advantageously, the training can be performed when only reconstructed images are available. For example, when DICOM images are available from the imaging system, the inputs of ^^^can be defined based on such reconstructed images. Forexample, ^^^^yஐ ^ qஐ, ^^ஐ^ can be defined as ^^^^^^^ஐு^^ஐ^ି^^^ுஐyஐ ^ ^^^ஐு^^ஐ^ି^^^ஐுqஐ^, where^^^ஐு^^ஐ^ି^^^ஐுyஐcan be represented by the linearly reconstructed DICOM image and ^^^ஐு^^ஐ^ି^^^ுஐqஐ ൌ p, and ^^^^yஐ, ^^ஐ^ can be defined as ^^^^^^^ஐு^^ஐ^ି^^^ஐுyஐ^. Regarding the sparsity term, x^^^can also be estimated without access to raw data. As one non-limiting example, x^^^can be defined using a non-linear reconstruction (e.g., compressed sensing) using a data fidelity unit that is based on the reconstructed images as described in co-pending PCT Patent Application No. PCT / US2025 / 011819, the entire contents of which is incorporated herein by reference.

[0069] Referring now to FIG.4, a flowchart is illustrated as setting forth the steps of an example method for training one or more neural networks (or other suitable machine learning models) on training data, such that the one or more neural networks are trained to receive imaging data as input data in order to generate reconstructed imaging data as output data. In some implementations, such reconstructed imaging data have improved image quality (e.g., reduced noise or aliasing artifacts) with respect to the input data.

[0070] In general, the neural network(s) can implement any number of different neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like. The machine learning model may in some instances be an unrolled machine learning model. Alternatively, the neural network(s) could be replaced with other suitable machine learning or artificial intelligence models, such as those based on supervised learning, unsupervised learning, deep learning, ensemble learning, dimensionality reduction, and so on. 15 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0071] The method includes accessing training data with a computer system, as indicated at step 402. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium (e.g., hospital imaging database, DICOM storage). Alternatively, accessing the training data may include acquiring such data with a medical imaging system (e.g., MRI system) and transferring or otherwise communicating the data to the computer system.

[0072] In general, the training data can include medical imaging data, which may be fully sampled (e.g., in the supervised approach) or subsampled or undersampled imaging data (e.g., in the unsupervised approach). Such training data may include raw data (e.g., k-space data) or may advantageously include processed imaging data (e.g., DICOM images reconstructed with parallel imaging techniques or other image reconstruction). Additionally, the training data may include other data, such as multi-coil sensitivity data.

[0073] The method can include assembling training data from an imaging database using a computer system. This step may include assembling the imaging data into an appropriate data structure on which the neural network or other machine learning model can be trained.

[0074] One or more neural networks (or other suitable machine learning models) are trained on the training data, as indicated at step 404. In general, the neural network can be trained by optimizing network parameters (e.g., weights, biases, or both) based on minimizing a loss function, as previously described. As non-limiting examples, the loss function may be described by one of Equations (7)–(11), or a variation thereof (e.g., other norm).

[0075] Training a neural network may include initializing the neural network, such as by computing, estimating, or otherwise selecting initial network parameters (e.g., weights, biases, or both). During training, an artificial neural network receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding weights. For instance, training data can be input to the initialized neural network, generating output as reconstructed imaging data. The artificial neural network then compares the generated output with the actual output of the training example in order to evaluate the quality of the reconstructed imaging data. For instance, the reconstructed imaging data can be passed to a loss function to compute an error. The current neural network can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current neural network can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the 16 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 loss according to the loss function. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validation iterations being completed, and the like. When the training condition has been met (e.g., by determining whether an error threshold or other stopping criterion has been satisfied), the current neural network and its associated network parameters represent the trained neural network. Different types of training processes can be used to adjust the bias values and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent, Newton's method, conjugate gradient, quasi-Newton, Levenberg-Marquardt, among others.

[0076] The artificial neural network can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transfer learning, scan-specific or zero-shot learning, or other suitable learning techniques for neural networks. As an example, supervised learning involves presenting a computer system with example inputs and their actual outputs (e.g., fully-sampled imaging data). In these instances, the artificial neural network is configured to learn a general rule or model that maps the inputs to the outputs based on the provided example input–output pairs. As another example, unsupervised learning can involve presenting a computer system with example inputs and initial estimates of their actual outputs (e.g., imaging data using compressed sensing or other parallel imaging technique).

[0077] The one or more trained neural networks are then stored for later use, as indicated at step 406. Storing the neural network(s) may include storing network parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the neural network(s) on the training data. Storing the trained neural network(s) may also include storing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.

[0078] Referring now to FIG.5, a flowchart is illustrated as setting forth the steps of an example method for generating classified feature data using a suitably trained neural network or other machine learning model. As will be described, the neural network or other machine learning model takes subsampled or undersampled imaging data as input data and generates reconstructed imaging data as output data. 17 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0079] The method includes accessing imaging data with a computer system, as indicated at step 502. Accessing the imaging data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the imaging data may include acquiring such data with a medical imaging system (e.g., MRI system) and transferring or otherwise communicating the data to the computer system, which may be a part of the medical imaging system.

[0080] A trained neural network (or other suitable machine learning model) is then accessed with the computer system, as indicated at step 504. In general, the neural network is trained, or has been trained, using the techniques described above in order to reconstruct images from subsampled medical image data, such as undersampled k-space data.

[0081] Accessing the trained neural network may include accessing network parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the neural network on training data. In some instances, retrieving the neural network can also include retrieving, constructing, or otherwise accessing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.

[0082] The imaging data (e.g., subsampled k-space data) are then input to the one or more trained neural networks, generating output as reconstructed imaging data, as indicated at step 506. For example, the reconstructed imaging data may include fully-sampled k-space data, reconstructed images, or denoised imaging data.

[0083] The reconstructed imaging data generated by inputting the medical image data to the trained neural network(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 508.

[0084] Experiments were performed using fully-sampled coronal proton density (PD) knee data from the NYU-fastMRI database. The relevant imaging parameters were: matrix size = 320×368, in-plane resolution = 0.49×0.44 mm2, slice thickness = 3 mm. These datasets were retrospectively undersampled using a uniform / equidistant undersampling pattern of R = 4 and 24 central k-space (ACS) lines, noting that the recovery of images from uniform undersampling patterns has been shown to be more challenging than from random undersampling. End-to-end database training was performed on 300 knee slices taken from 10 subjects. 18 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0085] The unrolled network utilized variable splitting, with data fidelity based on conjugate-gradient (CG) with 15 iterations, while the proximal operator for the regularizer was implemented via ResNet structure. Testing was performed on 380 knee slices from 10 different subjects. For subject-specific or zero-shot learning, only one knee slice was used.

[0086] For the proposed loss, biorthogonal 1.5 wavelets were selected as the transform ^^. λ in Equations (8) and (9) is a tunable hyperparameter, the effect of which was also studiedfor λ ∈ ^25, 50, 100, 150, 250^. For the unsupervised and zero-shot learning, 4 reweightingswere used, with initial x^^^estimated using conventional ℓ1-wavelet-regularized least-squares minimization. ^ for numerical stability was chosen as 10−3.

[0087] Comparisons were made to supervised learning with a normalized ℓ1-ℓ2loss, as well as to SSDU and EI. Additionally, conventional CS reconstruction with hand-tuned parameters was implemented to emphasize the distinctions between PD-DL with the described loss and convex CS reconstruction with linear representations. In subject-specific or zero-shot learning, the proposed loss was compared to zero-shot (ZS)-SSDU and CS. Results were quantitatively evaluated using structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR).

[0088] FIG.6 shows the supervised training outputs achieved using the loss described by Equation (8) for various values of λ. This parameter controls the trade-off between blurring (low values) and noise (high values). Visual and quantitative results indicate λ = 100 has the best performance, which was used for the rest of the study.

[0089] FIG.7 depicts a representative coronal PD knee slice reconstructed with various approaches for R = 4 uniform undersampling. Comparisons are made among CS, EI, SSDU, disclosed unsupervised, conventional supervised, and disclosed supervised PD-DL strategies. The disclosed approach visibly improves on CS, as expected, as well as on EI, with both these methods suffering from aliasing artifacts (green arrows). Overall, PD-DL reconstructions using the disclosed loss formulation provides similar reconstruction quality to both supervised PD- DL and SSDU-trained PD-DL methods.

[0090] FIG.8 summarizes the mean and standard deviation of SSIM and PSNR metrics for these reconstructions for 380 test slices. The SSIM and PSNR metrics are aligned with the visual observations, shown in FIG. 7. The disclosed loss formulation improves upon CS and EI, while performing similarly to existing unsupervised and supervised PD-DL methods that utilized k-space-based loss functions. 19 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0091] FIG.9 illustrates representative zero-shot or scan-specific reconstructions. CS shows residual artifacts (green arrows), while both ZS-SSDU and the proposed zero-shot loss perform similarly with no artifacts.

[0092] The present disclosure introduced a novel loss formulation, which uses the compressibility of the output image for supervised, unsupervised and subject-specific learning in PD-DL networks. Using concepts from statistical image processing, the capabilities of deep neural networks can be employed to achieve high-quality signal reconstructions.

[0093] In particular, non-linear low-dimensional representations afforded by neural networks can be used to solving the inverse imaging problem. However, such representations are used to evaluate the goodness of the reconstructions, thus differentiating the disclosed approach from conventional CS, as confirmed by the results provided herein. The results demonstrate comparable performance to existing unsupervised and supervised learning methods and present an alternative method that eliminates the need for masking the k-space in self-supervised learning.

[0094] The supervised loss in Equation (8) used a data fidelity term with acquired data. In the supervised setting, this term can be replaced with consistency across the whole k-space since the reference data is available, which leads to performance surpassing the conventional ℓ1-ℓ2 loss (results not shown). The fidelity term with respect to yஐcan be extended to unsupervised and zero-shot setups, while also allowing for a controlled tuning of the λ parameter.

[0095] In the provided example, a biorthogonal wavelet was used for ^^. Further gains may be achieved by using a combination of wavelets, overcomplete representations, or cycle- spinning approaches.

[0096] In another example study, the CUPID framework described in the present disclosure was used to train a PD-DL image reconstruction. The performance of the method was assessed through both qualitative and quantitative analyses, and focused on uniform / equidistant patterns, which produce coherent artifacts that are more difficult to remove compared to the incoherent artifacts from random undersampling.

[0097] In retrospective studies, fully-sampled multi-coil knee and brain MRI data from the fastMRI database were used. Knee dataset included fully-sampled coronal proton density- weighted (coronal PD) and PD with fat suppression (coronal PD-FS) data, both with matrix sizes of 320 × 320. For brain MRI, axial T2-weighted (ax T2) and axial FLAIR (ax FLAIR) datasets with matrix size of 320 × 320 were used. Both knee datasets include data collected 20 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 from 15 receiver coils, while the brain datasets include data collected from 16 and 20 receiver coils for the ax T2 and FLAIR, respectively. Every dataset were retrospectively undersampled using a uniform / equidistant pattern at R = 4. 24 lines of autocalibration signal (ACS) from center of the raw k-space data were kept. DICOM images to train the proposed model wereeach dataset, models were trained using 300 slices, and testing was performed using 380 slices for knee MRI and 300 slices for brain MRI, from distinct subjects.

[0098] In prospective studies, the practical pipeline for CUPID was replicated, wheredata were acquired at the desired high acceleration rate, and reconstructed to x PI with noiseand aliasing artifacts, using parallel imaging. To this end, a 3D MP2RAGE sequence was acquired on a 7T MRI scanner, with matrix size= 320 × 300 × 224, isotropic resolution = 0.75mm, prospective undersampling R = 4 (in ky only). This was the desired target acceleration, which was higher than the standard clinical acquisition. Low-resolution images were acquired in the same orientation for sensitivity estimation. Training and reconstruction with CUPID was done in a zero-shot subject-specific manner.

[0099] The exported DICOM images used for finetuning the PD-DL model with CUPID were magnitude-only, meaning phase information was not retained. This is commonin MRI scans, as most DICOMs only include magnitude of x PI . In applications where phaseis important (e.g., flow imaging), complex-valued DICOMs may be exported, and CUPID naturally works in this scenario as well. Nonetheless, CUPID does not make assumptions about the final image phase, and thus treats the zero phase of the magnitude-only DICOM as theglobal phase of x PI to reconstruct the image, making it compatible with all scans.

[0100] The results of the CUPID-based reconstructions were compared with several database training methods that have access to raw k-space data, including supervised PD-DL, SSDU, and equivariant imaging (EI). All PD-DL methods used the same unrolled network and components to ensure that only the training process differed for fair comparisons. CUPID was also compared with recent diffusion model based techniques, ScoreMRI and DDS. These train a time-dependent score function using denoising score matching on a large dataset of reference fully sampled DICOM images or raw data, and uses this score function during inference to sample from the conditional distribution given the measurements. Conventional reconstructionmethods were also included, such as CG-SENSE, which was used to generate the original x PI21 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 as the clinical baseline comparison that is typically not clinically usable at high acceleration rates, as well as CS.

[0101] The aforementioned methods use E H^ y ^ for data fidelity during inference,since this is what is needed to run gradient descent or conjugate gradient on ^ 2 -based fidelityterms including yH ^and E ^ x . E^ y ^ can be formed without accessing y ^ by multiplyingx PI with EH ^E ^ . Note that E ^ includes information about the undersampling pattern ^ ,which is completely known from the acquisition parameters, and coil sensitivities, which can be estimated from separate calibration scans in DICOM format. What sets CUPID apart fromother PD-DL strategies is that it can train the unrolled network without using y ^ . Thus,without loss of generality, E ^ is known both in training and in testing for all methods.

[0102] In the zero-shot setup, zero-shot results were compared with zero-shot SSDU (ZS-SSDU) as well as CS, and a baseline method, CG-SENSE, all of which are compatible with zero-shot inference. DDS and ScoreMRI were also included for inference in this setup, while noting that these diffusion priors have been pre-trained on very large databases. This was done to highlight the performance and versatility of CUPID, which is trained in a zero-shot manner on the given slice. All quantitative evaluations used structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR). CUPID does not use raw k-space data, and as a result, it does not benefit from the redundancies across multiple coils due to the lack of access to multi-coil raw k-space data.

[0103] Representative results in FIG. 10 show that baseline CG-SENSE, CS, and EI reconstructions exhibit residual artifacts, whereas DDS suffers from minor noise amplification. In contrast, CUPID successfully eliminates these artifacts from the CG-SENSE image using a well-trained PD-DL network, achieving a state-of-the-art reconstruction quality comparable tosupervised PD-DL and SSDU, despite only having access to x PI for training, and not to rawk-space data unlike these other methods. Parallel imaging reconstruction is not clinically usable at higher acceleration rates, but it is improved using a CUPID-trained PD-DL reconstruction. CUPID also outperformed ScoreMRI and DDS across various scenarios in the majority of cases, while achieving performance comparable to supervised PD-DL and SSDU, both of which have access to raw data.

[0104] FIG.11 shows zero-shot reconstruction results. Both CG-SENSE and CS suffer from noise amplification and persistent residual artifacts, with CG-SENSE displaying a more 22 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 pronounced degradation in quality. On the other hand, ScoreMRI and DDS exhibit residual artifacts in coronal PD, while ScoreMRI also introduces blurring in coronal PD-FS. This blurring artificially inflates the distortion metrics, as they favor smoother reconstructions by penalizing high-frequency details, aligning with the perception-distortion trade-off. Again, CUPID demonstrates superior artifact and noise reduction, closely matching the quality of ZS- SSDU, despite not having access to raw data and an explicit self-validation mechanism to prevent overfitting.

[0105] As discussed above, brain data were acquired at a target acceleration rate, reconstructed via parallel imaging, and exported in DICOM format to perform zero-shot fine- tuning. FIG.12 shows reconstruction results for the vendor parallel imaging reconstruction, as well as CS, DDS and CUPID. Vendor-provided R = 4 reconstruction on the scanner exhibits residual artifacts in FIG. 12A (shown in zoomed insets). CS effectively reduces noise but results in an overly-smooth reconstruction. DDS slightly mitigates the artifact and reduces noise, but both remain present. CUPID delivered the most effective artifact and noise mitigation in the DICOM image without requiring any raw k-space data, attesting to the effectiveness of CUPID in real-world scenarios. ZS-SSDU cannot be applied here due to the unavailability of raw data. The vendor-provided DICOM was generated using k-space interpolation instead of an image domain formulation. Furthermore, vendor-provided DICOMs may also include additional zero-padding, which the physics-based approach handled naturally, as well as additional filtering / processing.

[0106] Generalizability was evaluated with a representative out-of-distribution example. Specifically, in FIG.13, it is demonstrated how SSDU and supervised PD-DL models trained on 3T coronal PD data struggle when applied to 1.5T data with a different SNR, leading to noticeable artifacts. DDS, trained on both 3T and 1.5T data, does a better job of mitigating this issue, though some performance degradation persists. CUPID, zero-shot trained on the given slice without the need for raw data, performs the best, effectively reducing artifacts and maintaining high image quality.

[0107] FIG. 14 shows an example of a system 1400 that may be used in accordance with some aspects of the systems and methods described in the present disclosure. As shown in FIG.14, a computing device 1450 can receive one or more types of data (e.g., training data, DICOM images, signal evolution data, k-space data, receiver coil or multi-coil sensitivity data) from data source 1402. In some configurations, computing device 1450 can execute at least a portion of a training and image reconstruction system 1404 to train a machine learning model 23 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 or reconstruct images from medical imaging data (e.g., magnetic resonance data acquired using an undersampling technique). In some configurations, the training and image reconstruction system 1404 can implement an automated pipeline to provide a trained machine learning model or reconstructed imaging data (e.g., MRI images, medical images, reconstructed k-space data).

[0108] Additionally or alternatively, in some configurations, the computing device 1450 can communicate information about data received from the data source 1402 to a server 1452 over a communication network 1454, which can execute at least a portion of the training and image reconstruction system 1404. In such configurations, the server 1452 can return information to the computing device 1450 (and / or any other suitable computing device) indicative of an output of the training and image reconstruction system 1404.

[0109] In some configurations, computing device 1450 and / or server 1452 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 1450 and / or server 1452 can also reconstruct images from the data.

[0110] In some configurations, data source 1402 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data, DICOM images, training data), such as an MRI or other medical imaging system, another computing device (e.g., a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some configurations, data source 1402 can be local to computing device 1450. For example, data source 1402 can be incorporated with computing device 1450 (e.g., computing device 1450 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 1402 can be connected to computing device 1450 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some configurations, data source 1402 can be located locally and / or remotely from computing device 1450, and can communicate data to computing device 1450 (and / or server 1452) via a communication network (e.g., communication network 1454).

[0111] In some configurations, communication network 1454 can be any suitable communication network or combination of communication networks. For example, communication network 1454 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any 24 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some configurations, communication network 1454 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG.14 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.

[0112] Referring now to FIG. 5, an example of hardware 1500 that can be used to implement data source 1402, computing device 1450, and server 1452 in accordance with some configurations of the systems and methods described in the present disclosure is shown.

[0113] As shown in FIG. 15, in some configurations, computing device 1450 can include a processor 1502, a display 1504, one or more inputs 1506, one or more communication systems 1508, and / or memory 1510. In some configurations, processor 1502 can be any suitable hardware processor or combination of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and so on. In some configurations, display 1504 can include any suitable display devices, such as a liquid crystal display (LCD) screen, a light- emitting diode (LED) display, an organic LED (OLED) display, an electrophoretic display (e.g., an “e-ink” display), a computer monitor, a touchscreen, a television, and so on. In some configurations, inputs 1506 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0114] In some configurations, communications systems 1508 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1454 and / or any other suitable communication networks. For example, communications systems 1508 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1508 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0115] In some configurations, memory 1510 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1502 to present content using display 1504, to communicate with server 1452 via communications system(s) 1508, and so on. Memory 1510 can include any suitable 25 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1510 can include random-access memory (RAM), read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), other forms of volatile memory, other forms of non-volatile memory, one or more forms of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some configurations, memory 1510 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 1450. In such configurations, processor 1502 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 1452, transmit information to server 1452, and so on. For example, the processor 1502 and the memory 1510 can be configured to perform the methods described herein (e.g., the method illustrated in the workflow of FIGS.1A and 1B, the method illustrated in the workflow of FIG.3A, the method of FIG.4, the method of FIG.5).

[0116] In some configurations, server 1452 can include a processor 1512, a display 1514, one or more inputs 1516, one or more communications systems 1518, and / or memory 1520. In some configurations, processor 1512 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some configurations, display 1514 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some configurations, inputs 1516 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0117] In some configurations, communications systems 1518 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1454 and / or any other suitable communication networks. For example, communications systems 1518 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1518 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0118] In some configurations, memory 1520 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1512 to present content using display 1514, to communicate with one 26 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 or more computing devices 1450, and so on. Memory 1520 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1520 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some configurations, memory 1520 can have encoded thereon a server program for controlling operation of server 1452. In such configurations, processor 1512 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1450, receive information and / or content from one or more computing devices 1450, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.

[0119] In some configurations, the server 1452 is configured to perform the methods described in the present disclosure. For example, the processor 1512 and memory 1520 can be configured to perform the methods described herein (e.g., the method illustrated in the workflow of FIGS.1A and 1B, the method illustrated in the workflow of FIG.3A, the method of FIG.4, the method of FIG.5).

[0120] In some configurations, data source 1402 can include a processor 1522, one or more data acquisition systems 1524, one or more communications systems 1526, and / or memory 1528. In some configurations, processor 1522 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some configurations, the one or more data acquisition systems 1524 are generally configured to acquire data, images, or both, and can include an MRI or other medical imaging system. Additionally or alternatively, in some configurations, the one or more data acquisition systems 1524 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of an MRI or other medical imaging system. In some configurations, one or more portions of the data acquisition system(s) 1524 can be removable and / or replaceable.

[0121] Note that, although not shown, data source 1402 can include any suitable inputs and / or outputs. For example, data source 1402 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 1402 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on. 27 QB\920171.00646\95951430.1UMN 2024-231 920171.00646

[0122] In some configurations, communications systems 1526 can include any suitable hardware, firmware, and / or software for communicating information to computing device 1450 (and, in some configurations, over communication network 1454 and / or any other suitable communication networks). For example, communications systems 1526 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1526 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0123] In some configurations, memory 1528 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1522 to control the one or more data acquisition systems 1524, and / or receive data from the one or more data acquisition systems 1524; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 1450; and so on. Memory 1528 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1528 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some configurations, memory 1528 can have encoded thereon, or otherwise stored therein, a program for controlling operation of medical image data source 1402. In such configurations, processor 1522 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1450, receive information and / or content from one or more computing devices 1450, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.

[0124] In some configurations, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some configurations, computer-readable media can be transitory or non-transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during 28 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.

[0125] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “controller,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).

[0126] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.

[0127] As used herein, the phrase “at least one of A, B, and C” means at least one of A, at least one of B, and / or at least one of C, or any one of A, B, or C or combination of A, B, or C. A, B, and C are elements of a list, and A, B, and C may be anything contained in the Specification.

[0128] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention. 29 QB\920171.00646\95951430.1

Claims

UMN 2024-231 920171.00646 CLAIMS 1. A computer-implemented method for training a machine learning model to reconstruct an image from subsampled data, the method comprising: accessing training data with a computer system, the training data comprising subsampled medical imaging data; training a machine learning model on the training data, wherein training the machine learning algorithm includes minimizing a loss function comprising a sparsity term comprising a reweighted norm in a sparsifying transform domain and a consistency term that ensures a non-zero output when sparsity is minimized; storing the trained machine learning model in the computer system for later use.

2. The method of claim 1, wherein the sparsifying transform domain is a wavelet domain.

3. The method of claim 1, wherein the sparsity term comprises a sum of the reweighted norm in the sparsifying transform domain over ^^ transform domain coefficients.

4. The method of claim 3, wherein the reweighted norm is weighted based on a ratio of output imaging data and baseline imaging data.

5. The method of claim 4, wherein the baseline imaging data is imaging data acquired with full sampling.

6. The method of claim 4, wherein the baseline imaging data is an imaging estimate that acts as a surrogate for reference data updated for a fixed number of epochs.

7. The method of claim 6, wherein the imaging estimate is data reconstructed from accelerated data using compressed sensing for a first epoch, and the imaging estimate is the output imaging data for subsequent epochs.

8. The method of claim 6, wherein the imaging estimate is an image reconstructed with a non-linear reconstruction using a linearly reconstructed image to enforce 30 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 data fidelity for a first epoch, and the imaging estimate is the output imaging data for subsequent epochs.

9. The method of claim 1, wherein the subsampled medical imaging data is undersampled k-space data acquired using a magnetic resonance imaging system with an acceleration rate.

10. The method of claim 9, wherein the undersampled k-space data is acquired with uniform undersampling.

11. The method of claim 9, wherein the consistency term comprises a norm of a difference between an estimate of a perturbation image and the perturbation image, wherein the perturbation image comprises an image reconstructed from undersampled data that depicts a pattern without the pattern folding over on itself.

12. The method of claim 11, wherein the estimate of the perturbation image comprises a difference between a reconstructed image and a perturbation of the reconstructed image, wherein the perturbation of the reconstructed image is perturbed using the pattern depicted in the perturbation image.

13. The method of claim 12, wherein the reconstructed image is a linearly reconstructed DICOM image.

14. The method of claim 11, wherein the norm is an ℓ1 norm or an ℓ2 norm.

15. The method of claim 11, wherein the perturbation image comprises a randomly generated image that can be undersampled without overlap at the acceleration rate.

16. The method of claim 11, wherein the pattern depicted in the perturbation image comprises one of a pattern of shapes, a pattern of symbols, or a pattern of alphanumeric characters. 31 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 17. The method of claim 1, wherein the machine learning model comprises a neural network.

18. The method of claim 17, wherein the neural network is implemented with an unrolled neural network architecture comprising a plurality of steps, each step comprising reducing the loss function.

19. The method of claim 1, further comprising reconstructing an image by retrieving the trained machine learning model with the computer system and inputting the subsampled data to the trained machine learning model, generating output as a reconstructed image.

20. A computer-implemented method for training a machine learning model to solve inverse problems, the method comprising: accessing training data with a computer system; training a machine learning model using the training data by minimizing a loss function, the loss function comprising a consistency term and a reweighted norm in a sparsifying domain; storing the trained machine learning model in the computer system for later use.

21. The computer-implemented method of claim 20, wherein the reweighted norm comprises a ratio of an output solution to the inverse problem and baseline data in the sparsifying transform domain.

22. The computer-implemented method of claim 21, wherein the baseline data comprise reference data and the reweighted norm is defined as:wherein x୰^^is the reference data, ^^ is a transform into the sparsifying domain, ^^ denotes an ^^thtransform domain coefficient, ^^ is a total number of transform-domain coefficients, ^^ is added to ensure numerical stability, and x୭^^is the solution to the inverse problem. 32 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 23. The computer-implemented method of claim 21, wherein the baseline data comprise a signal estimate and the reweighted norm is defined as:wherein x^୩^is the signal estimate after training in a ^^threweighting step, ^^ is a transform into the sparsifying domain, ^^ denotes an ^^thtransform domain coefficient, ^^ is a total number of transform-domain coefficients, ^^ is added to ensure numerical stability, and x୭^^is the solution to the inverse problem.

24. The computer-implemented method of claim 20, wherein the sparsifying domain is a wavelet transform domain.

25. The computer-implemented method of claim 20, wherein the consistency termis defined as a normalized norm of ൫^^^^yஐ ^ qஐ, ^^ஐ^ െ ^^^^yஐ, ^^ஐ^൯ െ p, wherein ^^^represents an output reconstruction function, ^^^^yஐ, ^^ஐ^ represents a function of an imagerepresentative of the training data, ^^^^^ಈା୯ಈ, ^^ಈ^represents a function of the image representative of the training data perturbed by p, and p represents a perturbation image.

26. A method for reconstructing an image from undersampled k-space data, the method comprising: accessing a pre-trained neural network with a computer system, wherein the pre- trained neural network has been trained on magnetic resonance images using a loss function comprising a consistency term and a sparsity term comprising a sum of a reweighted norm in a sparsifying domain over ^^ transform domain coefficients, wherein the magnetic resonance images were reconstructed from subsampled k-space data; accessing undersampled k-space data with the computer system, wherein the undersampled k-space data were obtained from a subject using a magnetic resonance imaging (MRI) system; inputting the undersampled k-space data to the pre-trained neural network, generating output as a reconstructed image that depicts the subject; and displaying the image to a user using the computer system. 33 QB\920171.00646\95951430.1UMN 2024-231 920171.00646 27. The method of claim 26, wherein the reweighted norm comprises a ratio of an output reconstructed image and a baseline reconstructed image in the sparsifying transform domain.

28. The method of claim 26, wherein the sparsity term comprises a compressibility loss term.

29. The method of any one of claims 26 or 28, wherein the consistency term comprises a parallel imaging fidelity loss term. 34 QB\920171.00646\95951430.1

Citation Information

Patent Citations

  • Accelerated pseudo-random data magnetic resonance imaging system and method

    US20110241676A1

  • Unsupervised learning-based magnetic resonance reconstruction

    US20210150783A1

  • Using sparsity metadata to reduce systolic array power consumption

    US20220413924A1

  • Image processing method, system, device and storage medium

    US20240029204A1