Context-aware optimal transport learning for retinal fundus photograph enhancement

The context-aware OT learning framework addresses the challenge of preserving contextual information and minimizing artifacts in retinal fundus image enhancement, achieving superior image quality and diagnostic accuracy through deep feature extraction and adversarial training.

WO2026011185A1PCT designated stage Publication Date: 2026-01-08THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/036679
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-07-07
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing image processing systems for retinal fundus images struggle with preserving contextual information and minimizing unwanted artifacts, particularly in unsupervised learning methods that rely on quadratic or SSIM costs, leading to suboptimal image enhancement and compromised diagnostic accuracy.

Method used

A context-aware optimal transport (OT) learning framework that leverages deep contextual features using the earth mover's distance (EMD) to minimize unwanted artifacts and preserve structural information by shifting computation from image space to embedding space, employing a U-Net generator and VGG-19 CNN for feature extraction and adversarial training.

Benefits of technology

The framework effectively enhances retinal fundus images, improving signal-to-noise ratio and structural similarity, and significantly outperforms existing methods in downstream tasks such as vessel and lesion segmentation, while maintaining computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025036679_08012026_PF_FP_ABST
    Figure US2025036679_08012026_PF_FP_ABST
Patent Text Reader

Abstract

An image processing system to improve the quality of images via a context-aware OT framework. The system combines a representation of each one of a set of unlabeled low-quality images and corresponding high-quality images as a collection of feature vectors in a deep layer of a neural network, then obtains contextual information, abstracted in the deep layer of the neural network, associated with the representation of each one of low-quality images and the high-quality images. The system then derives a context-aware optimal transport method based on the obtained contextual information, and trains the OT framework, during its training phase, to learn, via application of the context-aware optimal transport method, a mapping between each one of the low-quality images and high-quality images. Thereafter, a low-quality image can be translated, via the trained OT framework, during its inference phase, to a corresponding high-quality image.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) CONTEXT-AWARE OPTIMAL TRANSPORT LEARNING FOR RETINAL FUNDUS PHOTOGRAPH ENHANCEMENT CLAIM OF PRIORITY

[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 668,056, filed July 5, 2024, entitled “CONTEXT-AWARE OPTIMAL TRANSPORT LEARNING FOR RETINAL FUNDUS PHOTOGRAPH ENHANCEMENT”, the disclosure of which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0002] The present disclosure relates to imaging and, in particular, to techniques for enhancing images such as color retinal fundus images through a context-aware optimal transport (OT) machine learning framework. BACKGROUND

[0003] Prior image processing systems and methods suffer from various deficiencies. Accordingly, improved image processing systems and methods are desirable. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] With reference to the following description and accompanying drawings:

[0005] FIG.1 illustrates a prior approach for image processing

[0006] FIG. 2 illustrates an exemplary approach for image processing, in accordance with an exemplary embodiment;

[0007] FIG. 3 shows an adversarial training scheme for a contextual OT framework in accordance with an exemplary embodiment;

[0008] FIG.4 shows the performance with respect to different values of the multiplication factor λ to regulate the importance of contextual loss, according to an exemplary embodiment;Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0009] FIG.5 illustrates qualitative results of the disclosed embodiments on degraded images;

[0010] FIG.6 illustrates visual comparison of the disclosed embodiments with baseline methods;

[0011] FIG. 7 illustrates lesion segmentation performance in accordance with the disclosed embodiments over the enhanced images obtained from different methods;

[0012] FIG. 8 illustrates vessel segmentation on the DRIVE dataset showing that the disclosed embodiments provide the best qualitative results on downstream vessel segmentation;

[0013] FIG.9 presents Table 1 which lists train-test splits, and metrics used in experiments involving an exemplary embodiment; and

[0014] FIG. 10 presents Table 2 which compares the performance of experiments involving an exemplary embodiment with certain prior approaches. DETAILED DESCRIPTION

[0015] Retinal fundus photography offers a non-invasive way to diagnose and monitor a variety of retinal diseases but is prone to inherent quality glitches, for example, arising from systemic imperfections or operator / patient-related factors. However, high-quality retinal images are crucial for carrying out accurate diagnoses and analyses, whether manual or automated analyses. Retinal fundus image enhancement is typically formulated as a distribution alignment problem, by finding a one-to- one mapping between a low-quality image and its high-quality counterpart. The disclosed embodiments provide a context-informed, or context-aware, optimal transport (OT) learning framework for tackling unpaired retinal fundus image enhancement. In contrast to prior generative image enhancement methods, which struggle with handling contextual information (e.g., over- tampered local structures and unwanted artifacts), the context-aware OT learning paradigm according to this disclosure better preserves local structures and minimizes unwanted artifacts.

[0016] Leveraging deep contextual features, the context-aware OT learning framework described herein is derived using the earth mover’s distance (EMD, or EM distance). Over probabilityAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) distributions, the earth mover's distance is also known as the Wasserstein metric W1, Kantorovich– Rubinstein metric, or Mallows' distance. It is a solution for an optimal transport (OT) problem, which in turn is also known as the Monge-Kantorovich problem, or sometimes the Hitchcock–Koopmans transportation problem. When the measures are uniform over a set of discrete elements, the same optimization problem is known as minimum weight bipartite matching.

[0017] Experimental results of the context-informed optimal transport (OT) learning framework on a large-scale dataset demonstrate the superiority of the disclosed embodiments over several other supervised and unsupervised machine learning methods in terms of signal-to-noise ratio, structural similarity index, as well as two downstream tasks. 1. Introduction

[0018] Doctors can examine the retinal fundus using cameras to accurately diagnose and treat eye diseases based on the symptoms shown in the photos taken by these cameras. A retinal fundus camera is a type of retinal camera that allows ophthalmologists to view the retina in greater detail. The results can be stored for further research and comparison. More importantly, both doctors and patients can view clearer images, facilitating more in-depth disease analysis and further treatment.

[0019] Under intense light, the pupil shrinks, making it difficult for a retinal fundus camera to capture a clear and comprehensive picture of the retinal fundus. Therefore, eye drops may be applied to a patient’s eye before a doctor takes photos with a mydriatic fundus camera to dilate the pupil. This provides a clearer view of the inner surface of the eye for examination. The use of a non-mydriatic fundus camera means that high-definition pictures of the optic disc, retina, and lens can also be achieved by the particular low power microscope of the instrument without increasing the size of the pupil.

[0020] Fundus photography involves taking continuous photographs of the eye through the pupil. In ophthalmic medicine, fundus cameras can be categorized into non-mydriatic and mydriatic fundusAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) cameras. Retinal color fundus photography (CFP) is vital for diagnosing ocular diseases, with non- mydriatic CFP being increasingly used for point-of-care diagnosis. Retinal CFP also plays an indispensable role in screening neurodegenerative disorders, such as Alzheimer’s disease, and systemic conditions like diabetes mellitus. However, the quality of non-mydriatic retinal CFPs can be impacted by various factors making it harder for accurate diagnosis in some cases.

[0021] Early optical models could handle the quality degradation due to opaque internal cataractous media and yielding high-quality counterparts. However, CFP is highly influenced by the manual errors arising from the imaging equipment and environmental conditions (e.g., dim surroundings without sufficient lights). The quality of fundus photography exhibits significant variability attributed to multiple factors, including varying operator expertise, fluctuations in illumination during image capture, lens contamination, and abrupt adjustments to focus settings resulting in blurred images. These noises compromise image quality and obscure crucial details such as blood vessels and lesions. Developing a single technique to robustly improve low-quality retinal CFPs suffering from the aforementioned factors would aid in disease (e.g., diabetic retinopathy) diagnosis as well as in developing automated tools for screening and population studies of neurological disorders.

[0022] Recently, deep learning-based methods have achieved state-of-the-art performance in enhancing the quality of fundus images. Early work in retinal fundus image enhancement revolved around supervised machine learning models which required noisy-clean pairs of images. However, the collection of paired noisy-clean retinal training samples proved arduous and costly in practice. To mitigate this challenge, unsupervised machine learning methods such as Generative Adversarial Networks (GANs) have drawn significant attention in recent years by modeling fundus image enhancement as an image translation task. One notable work is the OT-based generative models for fundus image enhancement, which leverages the fact that low-quality images and their high-qualityAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) counterparts share similar underlying structures. This helps reduce the search space for unpaired image-to-image tasks without a computationally expensive cycle consistency.

[0023] Another major challenge towards robust enhancement is understanding the contextual differences between the high-quality and poor-quality domains. Most of the latest enhancement techniques rely on a quadratic cost or an SSIM cost for the preservation of delicate details such as lesions and blood vessels. While this helps preserve structural information, it also distills contextually unwanted artifacts (e.g., light spots) to the enhanced images, due to the fundamental limitation of SSIM in overlooking contextual information in image enhancement.

[0024] Drawing on the understanding that contextual information is often embedded in the deep layers of pre-trained neural networks, the disclosed embodiments shift the computation of the OT cost from image space to embedding space. The innovative context-aware OT framework leverages deep feature spaces for more accurate fundus image enhancement, supported by the theoretical foundations of earth mover’s distance and OT theory, thus providing robust theoretical underpinnings. The contributions of this disclosure are threefold: (i) introducing a novel OT retinal image enhancement learning paradigm based on the deep layer feature space, aiming to minimize undue excessive tampering to lesions and structures while effectively removing noise; (ii) offering a strong theoretical foundation for general image enhancement tasks by ensuring that the transport cost reflects the intrinsic geometrical and contextual properties of the data in the deep feature space; and (iii) a comprehensive evaluation across three large publicly available retinal imaging datasets demonstrating the superiority of the proposed method over competing unsupervised and supervised methods. 2. Related Work

[0025] Recent advancement of deep learning has achieved state-of-the-art performance on the fundus image enhancement task. Prior deep learning-based methods for fundus image enhancement can be roughly divided into three categories:Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0026] (i) supervised methods; (ii) self-supervised methods; and (iii) unsupervised methods. The supervised methods have a hard requirement on paired noisy-clean images, while the unsupervised methods are typically trained on an unpaired dataset. The self-supervised methods rely on the information of the training dataset itself to formulate a supervised learning scheme.

[0027] Focusing on the supervised methods, a previous approach introduces a clinically-focused fundus enhancement network (cofe-Net) that learns direct mapping from degraded noisy images to high-quality clean images in a supervised fashion. Specifically, this approach leverages early-stage low-quality region activation and continuous retinal structure injection via a cascaded encoder-decoder network at multiple scales with shared weights. Recently, another approach, the proposed PCE-Net, decomposes low-quality images into Laplacian pyramid features for multi-resolution based enhancement. The added feature, pyramid constraint for the sequence, guides the PCE-Net to be degradation-invariant. However, these methods have limitations in real practice due to their reliance on paired noisy-clean images.

[0028] To relax the requirement of paired training samples, one approach introduces a fundus image enhancement network boosted by frequency self-supervised representation learning with structure- aware enhancement. This approach combines frequency self-supervision and synthesized data to train GFE-Net. Following this vein, SCR-Net introduces an enhancement network, which involves synthesizing multiple cataract-affected images from a clear fundus image followed by aligning and restoring high-frequency components. SCR-Net comprises an encoder for capturing high-frequency components and a decoder for enforcing high-frequency alignment to facilitate structure alignment and fundus image enhancement through feature sharing.

[0029] Due to the difficulty of collecting noisy-clean retinal image pairs, unsupervised methods such as GANs have drawn significant attention in recent years by modeling fundus image enhancement as an image-to-image translation task. In particular, CycleGAN serves as the most common imageAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) translation-based method for fundus image enhancement by training on unpaired low-quality and high- quality fundus images. Arguably, it suffers from a large search space with a huge computational overhead as well as the failure to preserve vessel and lesion structures. To address this limitation, ArcNet utilizes multiple quadratic loss functions to enforce the alignment of high-frequency components between low-high-quality images. I-SECRET introduces a dual-stage approach, including a supervised learning stage trained on degraded-clean image pairs and an unsupervised learning stage that focuses on generalizing enhancements using GAN. NAGAN introduces an approach that focuses on learning speckle noise patterns for OCT image-to-image translation. This is achieved by using a generator that takes images from the source domain as input and produces output images with noise resembling that of the target domain. Two discriminators are then utilized: one ensures that the generated images replicate the noise patterns of the target domain, while the other ensures that the structures from the source domain remain intact. However, this method is specifically designed for OCT image translation, where the primary distinction between the source and target domains lies in the speckle noise patterns.

[0030] In addition, GANs that leverage optimal transport (OT) theory to reduce search space have also been explored. These methods are contingent on the fact that low-quality images and their high- quality counterparts should share the same underlying structures. One approach proposed an OT guided GAN (OTTGAN) for unsupervised image denoising with a single generator and discriminator. Although it achieved significant results in natural image de-noising, its adoption of a quadratic OT cost led to the destruction or over-tampering of the vessel and lesion structures. To address these challenges, another approach introduced OT-GAN and OTEGAN, which leverage a structural similarity index (SSIM) cost to preserve structural information (e.g., lesions, vessel structures, and optical discs) between enhanced and low-quality images. It is worth noting that OTEGAN is an extension of OTGAN with an additional post-preprocessing step termed regularization by enhancing. While this helpsAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) preserve structural information, SSIM also distills contextually unwanted artifacts (e.g., light spots) to the enhanced images. It arises due to the fundamental limitation of SSIM in overlooking contextual information in image enhancement, as it operates on image space. 3. Method

[0031] Monge’s formulation. Image-to-image translation is a class of vision problems where the goal is to learn the mapping between an input image and an output image using a training set of aligned image pairs. However, for many tasks, paired training data is not available. For example, the fundus image enhancement task is formulated as an unpaired image-to-image translation task. Specifically, given a source domain Y (low-quality domain) and a target domain X (high-quality domain), the goal is to find a direct transformation f : Y → X. The idea of applying OT to this problem is natural, assuming that there is a one-to-one mapping between low-quality images and high-quality images. The image enhancement problem can be defined by Monge’s OT formulation as in Zhu, W., Qiu, P., Dumitrascu, O.M., Sobczak, J.M., Farazi, M., Yang, Z., Nandakumar, K., Wang, Y.: Otre: Where optimal transport guided unpaired image-to-image translation meets regularization by enhancing, International Conference on Information Processing in Medical Imaging, pp. 415–427, Springer (2023). Specifically, for two probability measures ν ∼ P(X) and µ ∼ P(Y ), this is given aswhere c(· , ·) is a cost function. Commonly used cost functions are linear cost and quadratic cost. Moregenerally, c(· , · ) can be defined as c(|x − y|) for some convex function c that measures the discrepancybetween x and y. It is worth noting that each image is treated as a point in its support, i.e., x ∼ X and y ∼ Y. The resulting OT learning scheme for image enhancement is shown in FIG.1. FIG.1 depicts a plurality of low-quality images such as images 100, 105, 110 and 115. These images are numericallyAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) represented in corresponding embedding vectors Y1, Y2, Y3 and Y4 in a manifold Y 120 of a low-quality image embedding space. Similarly, FIG. 1 depicts a plurality of high-quality images such as images 130, 135, 140 and 145. These images are numerically represented in corresponding embedding vectors X1, X2, X3 and X4 in a manifold X 150 of a high-quality image embedding space. Each embedding vector can be thought of as a point in a multidimensional space, where a vector’s location carries information about the represented data. This transformation allows machine learning models to perform various tasks more effectively by understanding the semantic relationship between data points. The manifolds X 120 and Y 150 are abstract collections of the data points in the respective embedding spaces, for example, the manifolds may be respective slices, or two-dimensional planes, within the low- and high- quality, image embedding spaces.

[0032] Prior art learning schemes work with these two manifolds, using the embedding vectors to determine the distance between the manifold of the low-quality image embedding space and the manifold of the high-quality image embedding space, according to the cost functions c(Y1, X1), c(Y2, X2), c(Y3, X3) and c(Y4, X4). The challenge with doing so is that when the plurality of low-quality images is converted to embedding vectors Y1, Y2, Y3and Y4in the manifold Y 120 of the low-quality image embedding space, noise artifacts, such as blurs, spots and light, are included in or represented in the embedding vectors Y1, Y2, Y3and Y4. Further, when the embedding vectors Y1, Y2, Y3and Y4in the manifold Y 120 are transported to the corresponding embedding vectors X1, X2, X3 and X4 in the manifold X 150 of the high-quality image embedding space, these noise artifacts also get transported to the corresponding embedding vectors X1, X2, X3 and X4.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0033] Lagrangian relaxation. With a Lagrangian multiplier, Eq. (1) can be reformulated as an unconstrained optimization problem:where d(· , ·) measures the divergence between probability distribution pXˆ and pX . Here, Xˆ= f (Y ) is used to define the enhanced domain. The first term in Eq. (2) minimizes the transport cost from the low-quality domain to the high-quality domain, facilitating maximal preservation of information from low-quality images in the enhanced images. The second term aligns the enhanced domain with the high-quality domain distribution. Notably, this formulation does not require cycle consistency, ensuring computational efficiency. It also leverages prior knowledge that enhanced / high-quality images should share underlying structures (e.g., optic disc, lesions, vessels) with low-quality images but are degraded by factors like illumination pollution, retinal artifacts, and blurring. This approach can reduce the search space in cycle consistency and mitigate the introduction of unrealistic components from CycleGAN. 3.1 Context-Aware OT Learning Framework for Image Enhancement

[0034] The problem defined in Eq. (2) is suboptimal for a retinal fundus image enhancement task. This is largely due to the fact that the common choice of the transport cost function c (e.g., linear cost c(Y, f (Y)) = ||Y − f (Y )||, quadratic cost c(Y, f (Y )) = ||Y − f (Y )||2in Wang, W., Wen, F., Yan, Z., Liu, P.: Optimal transport for unsupervised denoising learning, IEEE PAMI pp.1–1 (2022), and SSIM cost in Zhu, W., Qiu, P., Dumitrascu, O.M., Sobczak, J.M., Farazi, M., Yang, Z., Nandakumar, K., Wang, Y.: Otre: Where optimal transport guided unpaired image-to-image translation meets regularization by enhancing, International Conference on Information Processing in Medical Imaging, pp. 415–427, Springer (2023)) cannot effectively handle the image contextual information. Specifically, theAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) linear / quadratic cost only encourages the enhanced images to have the same arithmetic median / mean as low-quality images. However, the linear / quadratic cost treats each pixel independently but cannot effectively preserve structural information. The SSIM cost is only locally quasi-convex, which causes the solution to Eq. (2) to be inherently suboptimal. Although the SSIM cost helps preserve local information, it also retains contextually unwanted retinal artifacts in the enhanced images, as retinal artifacts also exhibit meaningful local structures.

[0035] In short, the aforementioned cost functions cannot effectively handle the image context, as they operate in the image space. Instead, the contextual information is typically abstracted in the deeper layers of a neural network.

[0036] Thus, the disclosed embodim ^en^ts incorporate contextual information into the OT problems defined in Eq. (1) and (2). An exemplary embodiment derives the context-aware OT using the earth mover’s distance (EMD), considering discrete sets instead of a continuous set in the previous derivation. An exemplary embodiment represents each image X and Y as a collection of feature vectors in the deep layer of a neural network Φ (e.g., VGG in Mechrez, R., Talmi, I., Zelnik-Manor, L.: The contextual loss for image transformation with non-aligned data (2018)): Y = {Φ(yi)} and f (Y ) = { Φ(f (yj))} with |X| = |Y | = N . An exemplary embodiment also normalizes the feature vectors to have a unit length ||Φ(yi)||2= 1, ||Φ(f (yj))||2= 1 to make it scale invariant. The EMD is given aswhere F is the flow matrix, and C is the cost matrix. Due to its computational intractability, an exemplary embodiment considers the relaxed EM (REM) distance, which can be easily optimized via gradient descent. Formally, this is given as:Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0037] According to the OT property, the transport map is symmetric from X to Y and Y to X (i.e., with the same cost), which is valid for the image enhancing task, so an exemplary embodiment can remove the maximum operation in Eq. (4). To this end, an exemplary embodiment formally derives the context-aware OT cost as

[0038] The yielded, novel, context-aware OT learning scheme is shown in FIG. 2. Generally speaking, and with comparison to FIG.1, the disclosed embodiments introduce intermediate manifolds 125 and 155 that respectively focus on only contextual information and entirely filter out the noise- based artifacts that otherwise get transported when the embedding vectors Y1, Y2, Y3 and Y4 in the manifold Y 120 are transported to the corresponding embedding vectors X1, X2, X3 and X4 in the manifold X 150 of the high-quality image embedding space. Thus, the disclosed embodiments, as depicted in FIG. 2, generate embedding vectors Φ(Y1), Φ(Y2), Φ(Y3) and Φ(Y4) in intermediate manifold 125 of the low-quality image embedding space that numerically represent only contextual information from the plurality of low-quality images. Similarly, the disclosed embodiments generate embedding vectors Φ(X1), Φ(X2), Φ(X3) and Φ(X4) in intermediate manifold 155 of the high-quality image embedding space that numerically represent only contextual information from the plurality of high-quality images.

[0039] One advantage to doing so is reducing the high-dimensionality of data points representing both wanted (contextual) and unwanted (artifact) information in manifold Y 120 of low-quality image embedding space to a smaller dimensionality of data points representing only contextual information in the intermediate manifold 125 of the low-quality image embedding space. Reducing the imageAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) embedding space in this manner reduces the computational costs associated with translating the low- quality images to high-quality images.

[0040] Then, the disclosed embodiments, working with these two intermediate manifolds 125 and 155, determine the distance between the contextual features in embedding vectors Φ(Y1), Φ(Y2), Φ(Y3) and Φ(Y4) in intermediate manifold 125 of the low-quality image embedding space and the contextual features in embedding vectors Φ(X1), Φ(X2), Φ(X3) and Φ(X4) in intermediate manifold 155 of the high- quality image embedding space, according to the cost functions ccon(Φ(Y1), Φ(X1)), ccon(Φ(Y2), Φ(X2)), ccon(Φ(Y3), Φ(X3)) and ccon(Φ(Y4), Φ(X4)). The contextual information, in the example of retinal fundus images, may be lesions or blood vessels, fundus positioning, fundus color, etc. More generally, contextual information comprises information that is desired to be retained versus information that is considered superfluous, uninteresting, unimportant or unwanted.

[0041] This calculation of the distance between the contextual features in embedding vectors Φ(Y1), Φ(Y2), Φ(Y3) and Φ(Y4) in intermediate manifold 125 of the low-quality image embedding space and the contextual features in embedding vectors Φ(X1), Φ(X2), Φ(X3) and Φ(X4) in intermediate manifold 155 of the high-quality image embedding space allows for improved translation of the low-quality images to high-quality images without bringing along unwanted artifacts.

[0042] An exemplary embodiment considers the convex function with respect to the distance to obtain the cost matrix C (e.g., Euclidean distance). The exemplary embodiment also considers the distance defined in Mechrez, R., Talmi, I., Zelnik-Manor, L.: The contextual loss for image transformation with non-aligned data (2018):where h > 0 is a smoothing band-width parameter; when h = 0.5, the above cost converges to a cosine similarity-based cost.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0043] Context-aware OT learning as a GAN. For simplicity, and with reference to FIG. 3, the optimization problem defined in Eq. (2) can be realized as a Generative Adversarial Network (GAN) with an EM distance (also known as Wasserstein metric W) as the divergence measure (WGAN loss) 325 combined at 330 with an additional transport cost term 320 . FIG. 3 illustrates the adversarial training scheme for a contextual OT framework 300 in accordance with an exemplary embodiment that improves upon the OT GAN learning scheme. A generator (fθ) 305 may be a U-Net convolutional neural network (CNN) generator, or simply, U-Net generator, with residual connection and channel attention as outlined in Wei Wang, et. al., Optimal Transport for Unsupervised Denoising Learning, IEEE PAMI, pages 1–1, 2022, and Wenhui Zhu, et. al., Otre: Where Optimal Transport Guided Unpaired image-to-Image Translation Meets Regularization by Enhancing, in International Conference on Information Processing in Medical Imaging, pages 415–427, Springer, 2023. The U- Net generator 305 receives a plurality of low-quality images y 301 and generates a corresponding, or counterpart, plurality of high-quality images Y 302. The generated plurality of high-quality images Y 302 is fed to a convolutional neural network (CNN), for example, a deep CNN such as a Visual Geometry Group CNN like the Visual Geometry Group (VGG)-19 CNN depicted at 310. The VGG- 19 CNN 310 performs contextual feature extraction on the embedding vectors Y1, Y2, Y3 and Y4 in the manifold Y 120 of the low-quality image embedding space, yielding the embedding vectors Φ(Y1), Φ(Y2), Φ(Y3) and Φ(Y4) in intermediate manifold 125 of the low-quality image embedding space that numerically represent only contextual information from the plurality of low-quality images. Likewise, the VGG-19 CNN 310 performs contextual feature extraction on the generated plurality of high-quality images Y 302, yielding the embedding vectors Φ(X1), Φ(X2), Φ(X3) and Φ(X4) in intermediate manifold 125 of the low-quality image embedding space that numerically represent only contextual information from the generated plurality of high-quality images. The distance between the embedding vectors in the intermediate manifold 125 for the plurality of low-quality images and corresponding embeddingAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) vectors in the intermediate manifold 150 for the generated plurality of high-quality images 302 is then calculated as depicted at 320, as the contextual transport cost.

[0044] The discriminator (Φ) 315 disproves the claims output by the U-Net generator. To do so, it receives as input the embedding vectors of the generated plurality of high-quality images 302 and embedding vectors of a verified, ground-truth or reference set of high-quality images, X, depicted at 303 and bridge the distance between these two sets of embedding vectors. According to one embodiment, the discriminator 315 may be the discriminator documented in Wei Wang, et. al., and Wenhui Zhu, et. al. VGG-19 is used for encoding the contextual information onto feature space and computes the contextual transport cost based on these feature embeddings. Following the convention of the WGAN, we can derive the adversarial training scheme:where fθ, Dψ are generator 305 and discriminator 315 parameterized by θ and ψ, respectively. It is worth noting that the GAN solution is subjected to the constraint that the function fψ and Dψ are both 1-Lipschitz continuous. In one implementation, an exemplary embodiment uses spectral normalization and gradient penalty to impose 1-Lipschitz constraint to fθ and Dψ, respectively, while other alternative solutions (not discussed here as beyond the scope of this disclosure). For a fair comparison, the same generator and discriminator network architectures outlined in Zhu, W., et al. are leveraged. Once the training scheme depicted in FIG.3 is complete, the VGG-19 and Discriminator no longer play a role in the framework, only the trained U-Net generator 305 operates to receive low quality images 301 and generate counterpart high-quality images 302.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) 4. Experiments and results 4.1 Experimental design

[0045] The effectiveness of the disclosed embodiments was validated on three publicly available datasets: theEyeQ, DRIVE, and IDRID. Following Zhu, W., et al., the experiment trained the method according to the disclosed embodiments using the EyeQ dataset and evaluated it on downstream tasks, such as vessel segmentation and diabetic lesion segmentation, using DRIVE and IDRID datasets.

[0046] Degradation Experiment. The degradation experiment was conducted to observe the effectiveness of the proposed method on the synthetically degraded retinal fundus images. The training set consists of a subset of 3560 high-quality images from the EyeQ dataset based on the Grading label provided. The training weights were evaluated on complete drive and images, along with the subset of 1819 high-quality images from the IQ data set.

[0047] PSNR and SSIM were estimated between the enhanced images of low-quality images, which were generated by degrading the high-quality image using the model outlined in Shen, Z., Fu, H., Shen, J., Shao, L.: Modeling and Enhancing Low-Quality Retinal Fundus Images. IEEE Trans Med Imaging 40(3), 996–1006 (2021), and the corresponding high-quality images. The prowess of the disclosed embodiments over the Degradation cases is showcased in the combination of Light Transmission Disturbance, Image Blurring, and Retinal Artifact. Downstream segmentation tasks are conducted to showcase the effective preservation of the intricate details from the low-quality fundus images post- enhancement.

[0048] Downstream Vessel Segmentation. The vessel segmentation task is conducted using the DRIVE dataset, where annotated masks are available. The disclosure uses the official training / testing split, which results in 20 subjects in training and testing set. The Vessel Segmentation task is evaluated based on the Area under ROC, Precision-Recall curve, Sensitivity, and Specificity.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0049] Downstream Diabetic Lesion Segmentation. This disclosure chooses the segmentation masks provided with the IDRID dataset. Since the training and testing of downstream segmentation tasks were based entirely on enhanced images, without adding any preprocessing and additional tricks, only large blocks of lesions that were easy to train, such as EX and HE. The Training set is made up of 54 subjects, and the Testing set is made up of 27 subjects. The performance is quantified using Area under ROC, Precision-recall, and Jaccard Index. The disclosed embodiments leveraged the vanilla UNet model to train from scratch for both segmentation tasks.

[0050] All images underwent center-cropping and resizing to a dimension of 256 × 256. The disclosed embodiments were compared with four recent robust research methods for retinal fundus image enhancement: CycleGAN, OTTGAN, OTEGAN, and PCE-Net.

[0051] Implementation details. For the degradation experiment, the disclosed embodiments trained the model for 100 epochs using an RM-Sprop optimizer. The batch size was set to two. The initial learning rate was set to be 1×10−4for the discriminator and 5×10−5for the generator. The learning rate decayed by a factor of 10 every 50 epochs. To prevent overfitting during training, data augmentation was used for training, such as random horizontal / vertical flips, random crops, and random rotations. The UNet as a backbone is leveraged for the Lesion Segmentation task. The model was trained on the summation of BCELoss and Dice loss. An Adam optimizer was used with the initial learning rate of 2 × 10−4along with the weight decay set as 5 × 10−4. The disclosed embodiments maintained the batch size of 4 and trained for 300 epochs. The following data augmentation traits were incorporated: Horizontal and Vertical Flip, Random Grid Shuffle, and coarse dropout with a probability of 0.5. Similar to Lesion segmentation, the disclosed embodiments used UNet for the Vessel segmentation task on the DRIVE dataset. The disclosed embodiments implemented the Adam optimizer with the criterion as Cross Entropy loss, with the initial learning rate as 5×10−5and batch size set to 64. The best model was saved from 50 epochs. To overcome the limited dataset issues, the following dataAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) augmentation techniques were used: Random Crop, Random flip (Left-Right and Up-Down) with a probability of 0.5, and Random rotation. All the mentioned experiments were performed on an Nvidia RTX3090 GPU. 4.2 Ablation studies

[0052] For the ablation study conducted, Eq. (7) explains the cost function that the disclosed embodiments utilize. To understand the significance of the contextual transport cost, the PSNR and SSIM trends were studied over the varying importance of contextual loss in training the Generator. A multiplication factor λ was introduced to regulate the importance of Contextual loss. The λ covers a wide range from 10 to 100 with a stride of 10. Apart from λ, the remaining Hyperparameters and training loop are set to the values discussed in section 4.1 above for the degradation experiment. The performance with respect to the different values of λ is illustrated in FIG.4. A slow upward movement of PSNR and SSIM was observed with the increasing λ value. The performance peaked when λ = 50 (see FIG. 4), after which, increasing the value of λ leads to inferior performance. It is hypothesized this might be attributed to the instability of GAN training and the inherent trade-offs between distribution alignment and structure preserving. 4.3 Experimental Results

[0053] Enhancement over multiple degradation cases. A robust enhancement technique should be able to handle all sorts of degradation and artifacts present in poor-quality images. FIG.5 illustrates the performance of different degradation cases, that generally occur while capturing the retinal images. The figure provides qualitative results of the disclosed embodiments on degraded images over the combinations of different noise (i.e., spot artifacts, illumination, and blurring). The disclosed embodiments achieve good enhancement performance even on severely degraded images (the third, fourth and fifth column of images in FIG.5). The disclosed embodiments preserve finer and thinner blood vessels very well across the spectrum of noises. In the cases of Blur and / or IlluminationAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) degradation where the thinner blood vessels are vaguely visible (see the images in the first, third, and fifth columns of FIG.5), an exemplary embodiment has improved the visibility of these vessels, aiding the diagnosis. The context-aware optimal transport between the quality domains illustrate complete eradication of the sports artifacts (See the images in columns two, four and five in FIG.5).

[0054] Results in degradation experiments. For the degradation experiment, a qualitative evaluation was conducted, as illustrated in Table 1, presented in FIG.9. Table 2, presented in FIG.10, provides a performance comparison with the SOTA methods. The best performance within each column is highlighted in bold. (∗: p < 0.01; with the paired t-test to the baseline methods.) Pre-trained weights were used to obtain the enhanced images for cofe-Net, GFE-Net, SCR-Net. For the remaining techniques, the models were trained on default settings shared in the official repository. The proposed method according to the disclosed embodiments outperformed all baseline methods in terms of PSNR on all three datasets. Specifically, the proposed method surpassed the supervised methods by a significant margin in PSNR on EyeQ, DRIVE, and IDRID datasets, respectively. A similar trend was observed for the GAN-based methods: the proposed method surpassed the recently introduced OT- based OTEGAN by 1.28, 11.54, and 3.45 in PSNR on three datasets, proving the proposed method is an efficient OT-guided learning method.

[0055] However, a slight SSIM performance drop was observed in the IDRID and DRIVE datasets (which contain more complex lesion structures). It is anticipated that the problem here is the limited subjects in these datasets. It is evident that the disclosed embodiments stood second for both datasets and maintained a very close margin of 0.007 with Cycle- GAN for the SSIM on the IDRID dataset. Thus, the proposed method still demonstrated reasonably good generalizability to out-of-distribution datasets (see Table 1 presented in FIG. 9). It is hypothesized that this was credited to the robust contextual feature encoded in the embedding space that better characterizes image quality. It was also observed that although the supervised method showed a satisfactory enhancement performance, it wasAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) likely to introduce unrealistic structures (see FIG.6). Besides, the other GAN-based methods struggled with removing retinal artifacts, whereas the method according to the disclosed embodiments can remove retinal artifacts.

[0056] Results in downstream segmentation tasks. To assess the efficacy of the disclosed embodiments in diagnostic tasks, the performance of the disclosed embodiments was evaluated on diabetic retinopathy lesion and Blood Vessel Segmentation tasks. As shown in Table 2, presented in FIG. 10, the disclosed method outperformed others in both tasks, achieving the best ROC and PR scores, proving the claim to better preserve the relevant artifacts during enhancement, essential for proper diagnosis.

[0057] Two types of lesions were selected, i.e., Hard exudates (EX) and Hemorrhages (HE). This disclosure shows the qualitative assessment (FIG. 7) supporting the findings for consistent, superior lesion identification, especially with the HE blocks. Despite a performance dip in image enhancement across most methods when untrained on IDRID, the focus remained on lesion preservation. Notably, the other approaches struggled with accurate lesion contouring. Of the two kinds of lesions presented here, EX blocks were easy to locate with the naked eye. But it is another thing to solidly enhance the area. Both cofe-Net and PCE-Net, even being the supervised method trained over paired data, have yielded incomplete masks. In contrast, the GAN-based techniques classified more areas as lesions. However, the OT-based techniques (OTTGAN, OTEGAN, and the method according to the disclosed embodiments) performed with near-accurate prediction. The same is evident from the Quantitative analysis (see Table 2 presented in FIG.10).

[0058] Hemorrhages are highly indistinguishable when compared to the fundus background. The same is highlighted in FIG.7 with box 700. Considering its delicate nature, most of the methods failed to localize it properly. For instance, I-SECRET often over-generated the lesions or distorted structures in areas with hemorrhages. However, the context-aware mechanism has overcome this hindrance byAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) outperforming the other methods and achieved top scores for Area under ROC, Precision-Recall, and Jaccard Index.

[0059] The visualization of the Blood Vessel Segmentation task conducted over the DRIVE dataset is presented in FIG.8. The disclosed embodiments outperformed the benchmark methods in the Area under ROC, Precision-recall, and Specificity. The cofe-Net yielded the best Sensitivity value by beating the disclosed method by 0.33. Although the added advantage of paired data for training in supervised methods is noted, cofe-Net is missing the thinner blood vessels in the denser region. On the other hand, PCE-Net predicted false vessels (see highlighted region in FIG.8). Although the OT-based techniques (OTTGAN and OTEGAN) precisely predicted the finer details in the mask, the disclosed embodiments beat them quantitatively.

[0060] Thus, the disclosed embodiments provide a computer-implemented method and system for an image processing system to improve quality of images via a context-aware optimal transport (OT) machine learning framework (“the OT framework”), comprising receiving a plurality of unlabeled low- quality images (“the low-quality images”) and a corresponding plurality of high-quality images (“the high-quality images”), obtaining a representation of each one of the low-quality images and the high- quality images as a collection of feature vectors in a deep layer of a neural network, obtaining contextual information, abstracted in the deep layer of the neural network, associated with the collection of feature vectors for each one of low-quality images and the high-quality images, deriving a context-aware optimal transport method based on the obtained contextual information, using an earth mover’s distance measure, training the OT framework, during its training phase, to learn, via application of the context-aware optimal transport method, a mapping or a translation between each one of the low-quality images and each one of the high-quality images, receiving a low-quality image, and translating, via the trained OT framework, during its inference phase, the low-quality image to a corresponding high-quality image.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0061] The disclosed embodiments may normalize the collection of feature vectors.

[0062] The disclosed embodiments may involve training the OT framework, during its training phase, to learn, via application of a Monge’s formulation of the context-aware optimal transport method, the mapping or the translation between each one of the low-quality images and each one of the high-quality images.

[0063] The disclosed embodiments may involve training the OT framework, during its training phase, to learn, via application of a Monge’s formulation of the context-aware optimal transport method, reformulated with a Lagrangian multiplier as an unconstrained optimization problem, the mapping or the translation between each one of the low-quality images and each one of the high-quality images to maximize preservation of relevant information in the low-quality images in the high-quality images.

[0064] The disclosed embodiments may receive the low-quality images from a retinal color fundus photography camera. The disclosed embodiments may receive the low-quality images from a non- mydriatic retinal color fundus photography camera.

[0065] Accordingly, The disclosed embodiments translate a plurality of low-quality images to a plurality of high-quality images, by receiving the plurality of low-quality images at a generator of a Generative Adversarial Network (GAN) framework, generating, via the generator 305, a counterpart plurality of high-quality images from the plurality of low-quality images, receiving the plurality of low-quality images and the generated counterpart plurality of high-quality images at a convolutional neural network (CNN) of the GAN framework, extracting, via the CNN, contextual information from the plurality of low-quality images and the generated counterpart plurality of high-quality images and creating therefrom respective low-quality image embedding vectors and high-quality image embedding vectors, calculating, via the CNN, a distance between the respective low-quality image embedding vectors and high-quality image embedding vectors, receiving at a discriminator of the GANAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) framework a reference set of high-quality image embedding vectors and the high-quality image embedding vectors created from the generated counterpart plurality of high-quality images, comparing, via a discriminator of the GAN framework, a similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors, and training the GAN framework to translate the low-quality images to a plurality of high-quality images based on the distance between the respective low-quality image embedding vectors and high-quality image embedding vectors and the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

[0066] The disclosed embodiments may receive the plurality of low-quality images at a generator 305 of a GAN optimal transport (OT) framework.

[0067] The disclosed embodiments may receive the plurality of low-quality images at a U-NET CNN generator 305 of a GAN framework.

[0068] The disclosed embodiments may receive the plurality of low-quality images and the generated counterpart plurality of high-quality images at a VGG-19 CNN 310 of the GAN framework.

[0069] The disclosed embodiments calculating, via the CNN, the distance between the respective low-quality image embedding vectors and high-quality image embedding vectors, may involve calculating, via the CNN, a contextual transport cost 320 between the respective low-quality image embedding vectors and high-quality image embedding vectors.

[0070] The disclosed embodiments comparing, via the discriminator 315 of the GAN framework, the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors, may involve calculating, via the discriminator of the GAN framework, a Wasserstein GAN loss 325 between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO)

[0071] The disclosed embodiments may train the GAN framework to translate the low-quality images to the plurality of high-quality images based on the distance between the respective low-quality image embedding vectors and high-quality image embedding vectors. The similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors, may involve training the GAN framework to translate the low-quality images to the plurality of high- quality images based on the contextual transport cost between the respective low-quality image embedding vectors and high-quality image embedding vectors and the Wasserstein GAN loss 325 between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

[0072] The disclosed embodiments may further involve the generator 305, during an inference workflow, receiving and translating a plurality of low-quality images to a plurality of high-quality images. 5. Conclusion

[0073] The disclosed embodiments provide for a Context-Aware Optimal Transport (OT) learning scheme for enhancing retinal fundus images. For this purpose, an exemplary embodiment leverages the earth mover’s distance within the context domain to characterize quality features over extraneous image information. Experimental findings demonstrated a notable enhancement over existing benchmarks across three datasets, evidencing the potential of context-aware OT-guided learning in image enhancement. The results of the disclosed embodiments were compared with recent state-of- the-art supervised and unsupervised learning methods and techniques. While the experiments are limited to non-severely damaged datasets, the methodology has the potential for broader application in medical image enhancement, including optical coherence tomography and endoscopy images.

[0074] Embodiments of the invention contemplate a machine or system within which embodiments may operate, be installed, integrated, or configured. In accordance with one embodiment, the systemAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) includes at least a processor and a memory therein to execute instructions including implementing any application code to perform any one or more of the methodologies discussed herein. Such a system may communicatively interface with and cooperatively execute with the benefit of remote systems, such as a user device sending instructions and data, a user device to receive output from the system.

[0075] A bus interfaces various components of the system amongst each other, with any other peripheral(s) of the system, and with external components such as external network elements, other machines, client devices, cloud computing services, etc. Communications may further include communicating with external devices via a network interface over a LAN, WAN, or the public Internet.

[0076] In alternative embodiments, the system may be connected (e.g., networked) to other machines in a Local Area Network (LAN), an intranet, an extranet, or the public Internet. The machine may operate in the capacity of a server or a client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or series of servers within an on-demand service environment. Certain embodiments of the machine may be in the form of a personal computer (PC), a tablet PC, a set top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, computing system, or any machine capable of executing a set of instructions (sequential or otherwise) that specify and mandate the specifically configured actions to be taken by that machine pursuant to stored instructions. Further, the term “machine” shall also be taken to include any collection of machines (e.g., computers) that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0077] An exemplary computer system includes a processor, a graphics processor, such as the Nvidia RTX3090 GPU, a main memory (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc., static memory such as flash memory, static random access memory (SRAM), volatile but high-dataAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) rate RAM, etc.), and a secondary memory (e.g., a persistent storage device including hard disk drives and a persistent database and / or a multi-tenant database implementation), which communicate with each other via a bus. Main memory includes code that implements the three branches of the SSL framework described herein, namely, the localizability branch, the composability branch, and the decomposability branch.

[0078] The processor represents one or more specialized and specifically configured processing devices such as a microprocessor, central processing unit, graphics processor, or the like. More particularly, the processor may be a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, processor implementing other instruction sets, or processors implementing a combination of instruction sets. The processor may also be one or more special-purpose processing devices such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. The processor is configured to execute processing logic for performing the operations and functionality discussed herein.

[0079] The system may further include a network interface card. The system also may include a user interface (such as a video display unit, a liquid crystal display, etc.), an alphanumeric input device (e.g., a keyboard), a cursor control device (e.g., a mouse), and a signal generation device (e.g., an integrated speaker). According to an embodiment of the system, the user interface communicably interfaces with a user client device remote from the system and communicatively interfaces with the system via public Internet.

[0080] The system may further include peripheral device (e.g., wireless or wired communication devices, memory devices, storage devices, audio processing devices, video processing devices, etc.).

[0081] A secondary memory may include a non-transitory machine-readable storage medium or a non-transitory computer readable storage medium or a non-transitory machine-accessible storageAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) medium on which is stored one or more sets of instructions (e.g., software) embodying any one or more of the methodologies or functions described herein. The software may also reside, completely or at least partially, within the main memory and / or within the processor during execution thereof by the system, the main memory and the processor also constituting machine-readable storage media. The software may further be transmitted or received over a network via the network interface card.

[0082] In addition to various hardware components depicted in the figures and described herein, embodiments further include various operations which are described herein. The operations described in accordance with such embodiments may be performed by hardware components or may be embodied in machine-executable instructions, which may be used to cause a specialized and special- purpose processor having been programmed with the instructions to perform the operations described herein. Alternatively, the operations may be performed by a combination of hardware and software. In such a way, the embodiments of the invention provide a technical solution to a technical problem.

[0083] Embodiments also relate to an apparatus for performing the operations disclosed herein. This apparatus may be specially constructed for the required purposes, or it may be a special purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as, but not limited to, any type of disk, CD-ROMs, and magnetic-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.

[0084] While the algorithms and displays presented herein are not inherently related to any particular computer or other apparatus, they are specially configured and implemented via customized and specialized computing hardware which is specifically adapted to more effectively execute the novel algorithms and displays which are described in greater detail herein. Various customizable and special purpose systems may be utilized in conjunction with specially configured programs in accordance withAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) the teachings herein, or it may prove convenient, in certain instances, to construct a more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will appear as set forth in the description. In addition, embodiments are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the embodiments as described herein.

[0085] Embodiments may be provided as a computer program product, or software, that may include a machine-readable medium having stored thereon instructions, which may be used to program a computer system (or other electronic devices) to perform a process according to the disclosed embodiments. A machine-readable medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., a computer) readable storage medium (e.g., read only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory devices, etc.), a machine (e.g., computer) readable transmission medium (electrical, optical, acoustical), etc.

[0086] Any of the disclosed embodiments may be used alone or together with one another in any combination. Although various embodiments may have been partially motivated by deficiencies with conventional techniques and approaches, some of which are described or alluded to within the specification, the embodiments need not necessarily address or solve any of these deficiencies, but rather, may address only some of the deficiencies, address none of the deficiencies, or be directed toward different deficiencies and problems which are not directly discussed.

[0087] While the subject matter disclosed herein has been described by way of example and in terms of the specific embodiments, it is to be understood that the claimed embodiments are not limited to the explicitly enumerated embodiments disclosed. To the contrary, the disclosure is intended to cover various modifications and similar arrangements as are apparent to those skilled in the art. Therefore,Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) the scope of the appended claims is to be accorded the broadest interpretation to encompass all such modifications and similar arrangements. It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other embodiments will be apparent to those of skill in the art upon reading and understanding the above description. The scope of the disclosed subject matter is therefore to be determined in reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) CLAIMS 1. A computer-implemented method for an image processing system to improve quality of images via a context-aware optimal transport (OT) machine learning framework (“the OT framework”), comprising: receiving a plurality of unlabeled low-quality images (“the low-quality images”) and a corresponding plurality of high-quality images (“the high-quality images”); obtaining a representation of each one of the low-quality images and the high-quality images as a collection of feature vectors in a deep layer of a neural network; obtaining contextual information, abstracted in the deep layer of the neural network, associated with the collection of feature vectors for each one of low-quality images and the high-quality images; deriving a context-aware optimal transport method based on the obtained contextual information, using an earth mover’s distance measure; training the OT framework, during its training phase, to learn, via application of the context-aware optimal transport method, a mapping or a translation between each one of the low-quality images and each one of the high-quality images; receiving a low-quality image; and translating, via the trained OT framework, during its inference phase, the low-quality image to a corresponding high-quality image.

2. The computer-implemented method of claim 1, further comprising normalizing the collection of feature vectors.

3. The computer-implemented method of claim 1, wherein training the OT framework, during its training phase, to learn, via application of the context-aware optimal transport method, theAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) mapping or the translation between each one of the low-quality images and each one of the high-quality images, comprises training the OT framework, during its training phase, to learn, via application of a Monge’s formulation of the context-aware optimal transport method, the mapping or the translation between each one of the low-quality images and each one of the high-quality images.

4. The computer-implemented method of claim 1, wherein training the OT framework, during its training phase, to learn, via application of the context-aware optimal transport method, the mapping or the translation between each one of the low-quality images and each one of the high-quality images, comprises training the OT framework, during its training phase, to learn, via application of a Monge’s formulation of the context-aware optimal transport method, reformulated with a Lagrangian multiplier as an unconstrained optimization problem, the mapping or the translation between each one of the low-quality images and each one of the high-quality images to maximize preservation of relevant information in the low-quality images in the high-quality images.

5. The computer-implemented method of claim 1, wherein receiving the low-quality images, comprises receiving the low-quality images from a retinal color fundus photography camera.

6. The computer-implemented method of claim 5, wherein receiving the low-quality images from the retinal color fundus photography camera, comprises receiving the low-quality images from a non-mydriatic retinal color fundus photography camera.

7. A computer-implemented method for translating a plurality of low-quality images to a plurality of high-quality images, comprising: receiving the plurality of low-quality images at a generator of a Generative Adversarial NetworkAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) (GAN) framework; generating, via the generator, a counterpart plurality of high-quality images from the plurality of low- quality images; receiving the plurality of low-quality images and the generated counterpart plurality of high-quality images at a convolutional neural network (CNN) of the GAN framework; extracting, via the CNN, contextual information from the plurality of low-quality images and the generated counterpart plurality of high-quality images and creating therefrom respective low- quality image embedding vectors and high-quality image embedding vectors; calculating, via the CNN, a distance between the respective low-quality image embedding vectors and high-quality image embedding vectors; receiving at a discriminator of the GAN framework a reference set of high-quality image embedding vectors and the high-quality image embedding vectors created from the generated counterpart plurality of high-quality images; comparing, via a discriminator of the GAN framework, a similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors; and training the GAN framework to translate the low-quality images to a plurality of high-quality images based on the distance between the respective low-quality image embedding vectors and high- quality image embedding vectors and the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

8. The computer-implemented method of claim 7, wherein receiving the plurality of low-quality images at the generator of the GAN framework, comprises receiving the plurality of low- quality images at a generator of a GAN optimal transport (OT) framework.Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) 9. The computer-implemented method of claim 7, wherein receiving the plurality of low-quality images at the generator of the GAN framework, comprises receiving the plurality of low- quality images at a U-NET CNN generator of a GAN framework.

10. The computer-implemented method of claim 7, wherein receiving the plurality of low-quality images and the generated counterpart plurality of high-quality images at the CNN of the GAN framework, comprises receiving the plurality of low-quality images and the generated counterpart plurality of high-quality images at a VGG-19 CNN of the GAN framework.

11. The computer-implemented method of claim 7, wherein calculating, via the CNN, the distance between the respective low-quality image embedding vectors and high-quality image embedding vectors, comprises calculating, via the CNN, a contextual transport cost between the respective low-quality image embedding vectors and high-quality image embedding vectors.

12. The computer-implemented method of claim 11, wherein comparing, via the discriminator of the GAN framework, the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors, comprises calculating, via the discriminator of the GAN framework, a Wasserstein GAN loss between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

13. The computer-implemented method of claim 12, wherein training the GAN framework to translate the low-quality images to the plurality of high-quality images based on the distance between the respective low-quality image embedding vectors and high-quality image embedding vectors and the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors, comprises training theAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) GAN framework to translate the low-quality images to the plurality of high-quality images based on the contextual transport cost between the respective low-quality image embedding vectors and high-quality image embedding vectors and the Wasserstein GAN loss between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

14. The computer-implemented method of claim 7, further comprising the generator, during an inference workflow, receiving and translating a plurality of low-quality images to a plurality of high-quality images.

15. A system comprising: A memory to store instructions; a processor to execute the instructions stored in the memory; wherein the system is configured to translate a plurality of low-quality images to a plurality of high- quality images by executing the instructions via the processor, comprising: receiving the plurality of low-quality images at a generator of a Generative Adversarial Network (GAN) framework; generating, via the generator, a counterpart plurality of high-quality images from the plurality of low- quality images; receiving the plurality of low-quality images and the generated counterpart plurality of high-quality images at a convolutional neural network (CNN) of the GAN framework; extracting, via the CNN, contextual information from the plurality of low-quality images and the generated counterpart plurality of high-quality images and creating therefrom respective low- quality image embedding vectors and high-quality image embedding vectors;Attorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) calculating, via the CNN, a distance between the respective low-quality image embedding vectors and high-quality image embedding vectors; receiving at a discriminator of the GAN framework a reference set of high-quality image embedding vectors and the high-quality image embedding vectors created from the generated counterpart plurality of high-quality images; comparing, via a discriminator of the GAN framework, a similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors; and training the GAN framework to translate the low-quality images to a plurality of high-quality images based on the distance between the respective low-quality image embedding vectors and high- quality image embedding vectors and the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

16. The system of claim 15, wherein calculating, via the CNN, the distance between the respective low-quality image embedding vectors and high-quality image embedding vectors, comprises calculating, via the CNN, a contextual transport cost between the respective low- quality image embedding vectors and high-quality image embedding vectors.

17. The system of claim 16, wherein comparing, via the discriminator of the GAN framework, the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors, comprises calculating, via the discriminator of the GAN framework, a Wasserstein GAN loss between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

18. The system of claim 17, wherein training the GAN framework to translate the low-quality images to the plurality of high-quality images based on the distance between the respectiveAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) low-quality image embedding vectors and high-quality image embedding vectors and the similarity between the reference set of high-quality image embedding vectors and the high- quality image embedding vectors, comprises training the GAN framework to translate the low-quality images to the plurality of high-quality images based on the contextual transport cost between the respective low-quality image embedding vectors and high-quality image embedding vectors and the Wasserstein GAN loss between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

19. A non-transitory computer-readable storage media having instructions stored thereupon that, when executed by a system having at least a processor and a memory therein, translate a plurality of low-quality images to a plurality of high-quality images by executing the instructions via the processor, comprising: receiving the plurality of low-quality images at a generator of a Generative Adversarial Network (GAN) framework; generating, via the generator, a counterpart plurality of high-quality images from the plurality of low- quality images; receiving the plurality of low-quality images and the generated counterpart plurality of high-quality images at a convolutional neural network (CNN) of the GAN framework; extracting, via the CNN, contextual information from the plurality of low-quality images and the generated counterpart plurality of high-quality images and creating therefrom respective low- quality image embedding vectors and high-quality image embedding vectors; calculating, via the CNN, a distance between the respective low-quality image embedding vectors and high-quality image embedding vectors; receiving at a discriminator of the GAN framework a reference set of high-quality image embeddingAttorney Docket No. 37684.6116WO (M24-235L-WO1-g / Mayo 2024-240 / WashU 021152 / WO) vectors and the high-quality image embedding vectors created from the generated counterpart plurality of high-quality images; comparing, via a discriminator of the GAN framework, a similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors; and training the GAN framework to translate the low-quality images to a plurality of high-quality images based on the distance between the respective low-quality image embedding vectors and high- quality image embedding vectors and the similarity between the reference set of high-quality image embedding vectors and the high-quality image embedding vectors.

20. The non-transitory computer-readable storage media of claim 19, wherein calculating, via the CNN, the distance between the respective low-quality image embedding vectors and high- quality image embedding vectors, comprises calculating, via the CNN, a contextual transport cost between the respective low-quality image embedding vectors and high-quality image embedding vectors.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method and computer-readable medium

    US20210224997A1

  • Unsupervised learning method for general inverse problem and apparatus therefor

    US20220027741A1