Motion artifact reduction in optical coherence tomography angiography images
The OSAT framework, combining HFormer and RCDM, addresses displacement, duplicated scanning, and white line artifacts in OCTA images, enhancing image quality and accuracy by leveraging hierarchical feature representation and contextual reconstruction.
Patent Information
- Application Number
- PCT/CN2025/074132
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-31
AI Technical Summary
Existing deep learning methods for optical coherence tomography angiography (OCTA) images primarily focus on white line artifacts, neglecting displacement and duplicated scanning artifacts, leading to incomplete motion artifact removal and potential misinterpretation of retinal vascular structures.
A two-stage framework, OSAT, utilizing a hierarchical transformer model (HFormer) to address displacement and duplicated scanning artifacts, followed by a residual conditional diffusion model (RCDM) to remove white line artifacts, enhancing feature representation and reconstruction precision.
OSAT effectively reduces all three types of motion artifacts in OCTA images, preserving anatomical details and improving image quality, outperforming existing methods in accuracy and clinical applicability.
Smart Images

Figure CN2025074132_31072025_PF_FP_ABST
Abstract
Description
Motion Artifact Reduction in Optical Coherence Tomography Angiography ImagesCROSS-REFERENCE TO RELATED APPLICATIONThis application claims priority to, and the benefit of, US Provisional Patent Application Serial No. 63 / 624,295 filed on January 24, 2024, the disclosure of which is incorporated by reference herein in its entirety.ABBREVIATIONSA-MSA axial-wise multi-head self-attentionC-MSA channel-wise multi-head self-attentionCNN convolutional neural networkCV computer visionFFN feed-forward networkHTM hierarchical transformer modelIA-MSA intra-axial multi-head self-attentionLFFN local feed-forward networkMAE mean absolute errorMLP multilayer perceptronMSA multi-head self-attentionOCT optical coherence tomographyOCTA optical coherence tomography angiographyOSAT One Step at A TimePSNR peak signal-to-noise ratioRCDM residual conditional diffusion modelSSIM structural similarity indexTECHNICAL FIELDThe present disclosure generally relates to image processing of medical images. Particularly, the present disclosure relates to a two-stage, deep-learning-based image processing technique for mitigating eye motion artifacts in OCTA images.BACKGROUNDOCTA is a non-invasive imaging technique that is widely used for retinal vascular imaging [1] . This technique leverages motion contrast imaging to generate angiographic images from high-resolution volumetric blood flow data rapidly. Specifically, OCTA is obtained by extracting the variance within multiple repeated B-scans at a cross-sectional location [2] , [3] . The application of OCTA has demonstrated considerable promise in the diagnosis of various retinal disorders, including diabetic retinopathy [4] , age related macular degeneration [5] , and glaucoma [6] .Nonetheless, the constrained raster scanning speed of B-scans makes them vulnerable to motion artifacts caused by a range of movements, including eye motion and blinking [7] . These motion artifacts, which are noted in OCTA images with a high occurrence rate of 93.1% [8] , can significantly affect the quality of en-face OCTA images, leading to potential disruptions in their visualization, interpretation, and analysis, ultimately influencing the accuracy of clinical diagnoses [9] .Based on the differences in imaging principles, the motion artifacts can be categorized into displacement artifacts, duplicate scanning artifacts, and white line artifacts
[0010] . Displacement artifacts occur when eye movements occur during raster scans, resulting in disconnected vertical image segments originating from various retinal regions
[0011] . Duplicate scanning artifacts stem from repetitive scanning of the same retinal areas, causing the appearance of duplicated vertical image segments [8] . White line artifacts manifest themselves as vertical lines within the OCTA image, potentially leading to the misinterpretation of underlying vascular structures
[0012] ,
[0013] .Recently, considerable research efforts have been devoted to addressing the issues caused by various artifacts via improving the scanning stability and speed of OCTA, including developing high-speed cameras
[0014] , advanced scanning mirrors
[0015] , and eye-tracking devices
[0016] . However, these hardware-based solutions are often associated with implementation complexity and high acquisition costs, which have limited their widespread adoption
[0017] . By contrast, software-based methods provide a cost-effective and competitive alternative. Classical software-based methods utilize statistical analysis
[0018] -
[0020] and registration techniques
[0021] ,
[0022] to improve the OCTA image quality. However, these classical methods often exhibit limited generalization capabilities to the motion artifacts and, in some cases, may even introduce additional artifacts
[0021] ,
[0023] . Recent progress in deep learning has led to the emergence of innovative strategies aimed at mitigating the white line artifacts in OCTA images
[0012] ,
[0024] -
[0026] . Although these deep learning methods surpass classical approaches in performance, some of these methods necessitate an additional vessel segmentation mask to obtain clear OCTA images
[0025] ,
[0026] . Notably, all these deep learning methods primarily target the white line artifacts
[0012] ,
[0024] -
[0026] so far. This specific concentration is a notable limitation, as existing deep learning techniques often neglect other essential motion artifacts in OCTA images, such as displacement and duplicated scanning artifacts. Consequently, there is a pressing need to create holistic solutions that tackle all three types of motion artifacts to generate high-quality OCTA images.A potential solution to address displacement and duplicated scanning artifacts is employing advanced Unet
[0027] ,
[0028] or transformer models
[0029] -
[0031] in computer vision. However, as shown in Table I, these methods do not perform well because they ignore the unique nature of these artifacts, such as direction misalignment and discontinuity. To mitigate these issues, we develop HFormer, a novel model specifically designed to better represent directional and local features in OCTA images. This makes it highly effective in tackling displacement and duplicated scanning artifacts.For white line artifact removal, well-designed generative models such as transformer
[0032] and diffusion model
[0033] in computer vision have shown potential. Nonetheless, as shown in Table I, which includes experimental results obtained for some existing deep learning models, effectiveness of the aforementioned generative models is constrained when applied directly to OCTA images. These methods often struggle to capture the critical fine details necessary for precise white line artifact removal, resulting in over-smoothing.There is a need in the art for a unique deep-learning technique applicable to mitigating displacement artifacts, duplicate scanning artifacts and white line artifacts in OCTA images to generate high-quality OCTA images.SUMMARYMathematical equations referenced in this Summary can be found in Detailed Description.An aspect of the present disclosure is to provide a computer-implemented method for reducing eye motion artifacts in a degraded OCTA image to yield a quality-enhanced OCTA image. The eye motion artifacts for reduction include displacement artifacts, duplicate scanning artifacts and white line artifacts.The method comprises: setting up a HTM configured to remove displacement artifacts and duplicated scanning artifacts from a first input OCTA image to thereby yield a first output OCTA image; setting up a RCDM configured to remove white line artifacts from a second input OCTA image to thereby yield a second output OCTA image; using the HTM to remove any displacement artifact and any duplicated scanning artifacts from the degraded OCTA image to yield an intermediate OCTA image; and using the RCDM to remove any white line artifact from the intermediate OCTA image to yield the quality-enhanced OCTA image.The HTM comprises a plurality of transformer blocks. The plurality of transformer blocks is configured to employ self-attention to capture a plurality of features utilizable for removing the displacement artifacts and duplicated scanning artifacts from the first input OCTA image. Preferably and advantageously, the plurality of features includes global features, vertical-line features and intra-vertical-line features of the first input OCTA image.In generalization, the HTM is configured to utilize the plurality of features, which includes the global features, vertical-line features and intra-vertical-line features of the first input OCTA image, to remove the displacement artifacts and duplicated scanning artifacts from the first input OCTA image.In certain embodiments, the HTM has a U-shaped hierarchical network structure incorporated with skip connections between encoder and decoder blocks of the HTM.In certain embodiments, an individual transformer block for processing an input feature map X to yield an output feature mapcomprises a MSA module and LFFN such that the individual transformer block computesfrom X according to EQNS. (6) and (7) .In certain embodiments, the LFFN is formed with two simultaneous convolutional blocks, and wherein each of the two simultaneous convolutional blocks contains a 1×1 point-wise convolution layer and a 3×3 depth-wise convolution layer.In certain embodiments, the plurality of transformer blocks includes: a first plurality of transformer blocks for capturing the global features; a second plurality of transformer blocks for capturing the vertical-line features; and a third plurality of transformer blocks for capturing the intra-vertical-line features.In certain embodiments, respective transformer blocks in the first plurality of transformer blocks are configured to employ channel-wise MSA for capturing the global features.In certain embodiments, respective transformer blocks in the second plurality of transformer blocks are configured to employ axial-wise MSA for capturing the vertical-line features.In certain embodiments, respective transformer blocks in the third plurality of transformer blocks are configured to employ intra-axial MSA for capturing the intra-vertical-line features.In certain embodiments, the HTM is trained to minimize a loss function given by EQN (1) .In certain embodiments, the setting up of the HTM includes training the HTM before the HTM is used to process the degraded OCTA image.Preferably, the HTM is realized as HFormer.In certain embodiments, the RCDM is a diffusion model configured to optimize a likelihood of correctly reconstructing a first region of the second input OCTA image overlapping with the white line artifacts given that the reconstructing of the first region is conditioned on image data of a second region of the second input OCTA not overlapping with the white line artifacts.In generalization, the RCDM is configured to reconstruct the first region while avoiding modifying the second region for removing the white line artifacts from the second input OCTA image.In certain embodiments, the setting up of the RCDM includes training the RCDM before the RCDM is used to process the intermediate OCTA image.A computing system for reducing eye motion artifacts in a degraded OCTA image to yield a quality-enhanced OCTA image is realizable based on the disclosed method. In particular, the computing system comprises one or more computers configured to execute a process of reducing the eye motion artifacts in the degraded OCTA image to yield the quality-enhanced OCTA image according to any of the embodiments of the disclosed method.Other aspects of the present disclosure are disclosed as illustrated by the embodiments hereinafter.BRIEF DESCRIPTION OF THE DRAWINGSFIG. 1 depicts a schematic diagram illustrating a framework of OSAT, where: the left side of the diagram shows an architecture of HFormer in Stage 1, a specific global transformer block and an axial transformer block; and the right side shows an architecture of RCDM in Stage 2 and the Residual block and self-attention block.FIG. 2, which includes subplots (a) - (d) , depicts schematic diagrams showing exemplary architecture of self-attention, where: subplot (a) is related to pixel-wise self-attention; subplot (b) is related to channel-wise self-attention; subplot (c) is related to axial-wise self-attention; and subplot (d) is related to intra-axial self-attention.FIG. 3 illustrates a forward diffusion process q (left to right) that gradually adds Gaussian noise to white line regions, and a reverse inference process p (right to left) that iteratively denoises the white line regions conditioned on the xc.FIG. 4 depicts exemplary schematic diagrams of FFN and LFFN as disclosed herein.FIG. 5 graphically displays results of different motion artifact removal methods on the real clinical corrupted OCTA images.FIG. 6 depicts a flowchart showing exemplary steps of a computer-implemented method as disclosed herein for reducing eye motion artifacts in a degraded OCTA image to yield a quality-enhanced OCTA image.Skilled artisans will appreciate that elements in the figures are illustrated for simplicity and clarity and have not necessarily been depicted to scale.DETAILED DESCRIPTIONAs used herein, the term “avoid” or “avoiding” refers to any method to partially or completely preclude, avert, obviate, forestall, stop, hinder or delay the consequence or phenomenon following the term “avoid” or “avoiding” from happening. The term “avoid” or “avoiding” does not mean that it is necessarily absolute, but rather effective for providing some degree of avoidance or prevention or amelioration of consequence or phenomenon following the term “avoid” or “avoiding” .To address various challenges as mentioned above in removing displacement artifacts, duplicated scanning artifacts and white line artifacts from an OCTA image, the present discloses develops a two-stage framework for removing these three types of eye motion artifacts sequentially. The two-stage framework is named as OSAT (viz., One Step at A Time) . The disclosed OSAT includes several novel designs, including a HFormer for removing the displacement artifacts and the duplicated scanning artifacts, and a RCDM for removing the white line artifacts.The HFormer is tailored for the OCTA image by enhancing the holistic feature representation across the global, local, and vertical ranges. This design enables the HFormer to capture fine details and complex textures, integrating information from multiple perspectives and detail levels for more effective artifact removal.As mentioned above, existing methods for white-line artifact removal often struggle to capture the critical fine details necessary for precise white-line artifact removal, resulting in over-smoothing. To address this challenge, a novel approach, the RCDM, which integrates contextual information to enhance the focus on the reconstruction of details within the white line artifacts in the OCTA image, is developed. The RCDM effectively leverages contextual information in the OCTA image during iterative removal of the white line artifacts, ensuring the preservation of crucial pathological details while efficiently targeting and removing the artifacts.The advancement provided by the present disclosure to the art includes the following three aspects. First, a two-stage framework, OSAT, for removing displacement artifacts, duplicated scanning artifacts, and white line artifacts sequentially is proposed. Second, a new transformer HFormer is proposed. HFormer is used as the first stage method in OSAT to remove the displacement and duplicated scanning artifacts. The hierarchical self-attention modules in HFormer enhance the holistic feature representation across the global, local, and vertical ranges. Third, in the second stage of OSAT, a new diffusion model known as RCDM is proposed. RCDM is used to remove white line artifacts. RCDM strategically leverages context information as a condition, focusing its attention on the white line areas and bolstering its capability to fill in corrupt regions.1. Related Works Useful for Developing OSAT1.1. Deep Learning for Motion Artifact Removal in OCTADeep learning methods have garnered significant attention for motion artifact removal in OCTA images. Gao et al.
[0024] employed a revised U-Net architecture and incorporated a low-rank matrix constraint as well as a total variational norm constraint into the loss function, effectively mitigating white line artifacts. A similar approach that utilizes low-rank approximation and total variation was also proposed for denoising OCT images
[0034] and OCTA images
[0035] . Under the ComNet framework
[0025] , a two-stage convolutional neural network is designed for both white line artifact removal and vessel segmentation in OCTA images. Ren et al.
[0026] proposed a self-supervised approach called CABR, in which a reconstruction model is trained on clean regions, and vessel segmentation annotations are utilized during inference to remove white line artifacts efficiently. Meanwhile, SR-Net
[0012] utilizes a straightforward U-net to directly address white line artifacts and employs an additional refinement network to further enhance image quality. These methods mainly target white line artifacts but neglect displacement and scanning duplication artifacts. Additionally, these methods often rely on complex vessel segmentation annotations, making them less practical for real-world applications. In contrast, the present disclosure strives to address all three types of motion artifacts without using vessel segmentation annotations, marking a pioneering effort in this direction.1.2. Transformer for Image ReconstructionTransformers have demonstrated superior performance over CNNs due to their ability to capture non-local information. As a result, many efforts have been devoted to exploring the application of Transformers to image reconstruction
[0029] ,
[0031] ,
[0036] -
[0040] . Swin Transformer
[0041] uses local windows to focus attention on regions and employs shift operations to enhance the interactions within these windows. This approach, exemplified by SwinIR
[0037] , is effective for image reconstruction. DaViT
[0042] incorporates the dual self-attention mechanism to efficiently capture global context with linear complexity, enhancing image reconstruction. SFormer
[0031] integrates self-attention into the U-Net framework, combining both paradigms for strong performance. Restormer
[0029] removes spatial self-attention, maintaining performance and showcasing Transformer adaptability. Uformer
[0030] addresses high-resolution maps by using shifting windows and adapting Transformers to computational constraints without compromising results in reconstruction. In the present disclosure, a hierarchical transformer is proposed to enhance the feature representation across the global, local, and vertical ranges for better motion artifact removal.1.3. Diffusion Model for Image ReconstructionDiffusion models have demonstrated remarkable proficiency in generating high-quality images. Their versatility and effectiveness has recently been expanded into the realm of image restoration tasks
[0043] ,
[0044] . These tasks encompass a broad spectrum of challenges, including denoising, super-resolution, and inpainting
[0033] ,
[0045] -
[0047] . Notably, SR3
[0047] and cascaded diffusion model
[0048] have exemplified the significant potential of diffusion models in super-resolution tasks. Additionally, Palette
[0049] has drawn inspiration from conditional generation models to introduce a conditional diffusion model tailored for various image reconstruction tasks. Latent diffusion models
[0050] have greatly enhanced the efficiency of image restoration processes in latent spaces. RePaint
[0033] has made improvements in inpainting techniques through the reassessment of iterations within diffusion models. However, reconstructing the entire image in RePaint may introduce additional artifacts due to the diffusion model’s inherent output diversity. In contrast, the present disclosure takes the clear regions in OCTA images as a condition and pays more attention to the completion in the white line regions.2. Development of OSATIn what follows, a two-stage approach OSAT designed to remove three types of motion artifacts is presented. In the first stage, a U-shaped hierarchical Transformer model, HFormer, is proposed in Section 2.1 to tackle the displacement and duplicate scanning artifacts. In the second stage, a conditional diffusion model, RCDM, is proposed in Section 2.2 to alleviate the negative impacts of white line artifacts and the additional white regions introduced after removing the displacement and duplicate scanning artifacts.2.1. Stage 1: Hierarchical TransformerAs depicted in FIG. 1, the overall architecture of the proposed HFormer follows a U-shaped hierarchical network structure, incorporating skip connections between the encoder and decoder blocks. To elucidate, commencing with an input degraded OCTA image denoted aswe initiate the process with a convolution layer to derive the low-level feature representationSubsequently, X0 undergoes four sequential encoder stages. Each stage comprises a down-sampling layer, which reduces the size of the feature maps by half while doubling the feature channels, along with a stack of HFormer blocks. The number of blocks gradually increases from the top to the bottom levels. For the feature reconstruction stage, the decoder consists of an up-sampling layer, which decreases the feature channels by half while doubling the size of the feature maps, and a stack of Transformer blocks akin to those employed in the encoder. Additionally, the encoder features are fused with the decoder features through skip connections. Subsequently, a 3×3 convolution layer is applied to obtain a residual image denoted asUltimately, the restored image is acquired as I′=I+R. The HFormer is trained by employing the subsequent loss function:where: represents the first-stage ground-truth image with white line artifacts; is given by
[0051] andis defined asIn EQNS. (2) and (3) , ε is empirically set to 1×10-3 for all the experiments, and Δ denotes the Laplacian operator. The Laplacian operator Δ (x, y) of an image with pixel density value I (x, y) is given by:This can be calculated by a Laplacian of Gaussian kernel on the OCTA images, and the kernel function isThe edge-based loss function ensures that the reconstructed OCTA images maintain the original anatomical boundaries of the vessels. By focusing on the edges, we can preserve fine details and prevent any structural deformation that might arise from the refinement process. Compared to methods that require detailed vessel segmentation masks, our edge-preserving loss function does not necessitate pretraining another segmentation model. This advantage simplifies the process and enhances the efficiency of our approach, making it more practical for real-world applications. The parameter λ in EQN. (1) controls the relative importance of the two loss terms. A value of λ may be set to 0.05. In EQN. (5) , δ may be set to be 0.5. Other values of λ and δ may be used depending on practical situations.2.1.1. Hierarchical self-attentionApplying a Transformer for motion artifact removal presents three primary challenges. Firstly, as shown in subplot (a) of FIG. 2, traditional Transformers calculate spatial self-attention, leading to a quadratic increase in computational cost proportional to the feature map size, making it impractical for high-resolution feature maps. Secondly, preserving local context information is crucial for artifact removal, as the surrounding area of a displaced or duplicated scanning pixel can help restore its original state. However, Transformers have limitations in capturing these local dependencies effectively. Lastly, both vertical feature representation and correlation play a significant role in removing displacement and duplicate scanning artifacts, but self-attention struggles to capture these aspects adequately. To address the three mentioned issues, we introduce a hierarchical Transformer block in Stage 1. This block utilizes self-attention to capture dependencies across long-range and vertical-range elements and integrates depth-wise convolution into both self-attention modules and FFNs to improve the capture of crucial local-range context. When dealing with the feature map X, the HFormer block consists of three types of Transformer blocks with MSA modules and a LFFN. The computation for each Transformer block is as follows:X′=MSA (X) +X (6)andwhere X′ andare the outputs of the MSA module and of the LFFN module, respectively. In the following, we elaborate on three types of Transformer blocks with MSA modules and LFFNs.Channel-wise Multi-head Self-Attention (C-MSA) . Instead of employing pixel-wise self-attention to capture global features, we conduct self-attention operations across the channels. As depicted in subplot (b) of FIG. 2, consider an input tensorWe denote the number of heads as h, and the head dimension is D=C / h. For the i-th head input feature Xi, C-MSA initiates the following operations sequentially: layer normalization φ (. ) , a 1×1 point-wise convolution layer Fl (. ) , and a 3×3 lightweight depth-wise convolution layer Fd (. ) , to obtain the projectionsand respectively. By utilizing the 3×3 depth-wise convolution, the resultant projectionsandare enriched with local contextual information. Subsequently, the channel-wise attention for the i-th head can be expressed as follows:whererepresents the single-head channel-wise self-attention and it captures the relationship across the channel dimension. αi serves as an adaptive and trainable scaling parameter, allowing us to modulate the dot product’s magnitude before applying the softmax function. The channel-wise multi-head self-attention can be expressed as follows:whererepresents the operation of 1×1 convolution.Axial-wise Multi-head Self-Attention (A-MSA) . To remove displacement and duplicate scanning artifacts in OCTA images, it is crucial to capture the vertical features. Thus, we introduce axial-wise self-attention, where each vertical line feature is treated as a token, allowing the model to learn relationships among these vertical lines. As depicted in subplot (c) of FIG. 2, consider the input tensorWe partition X into W vertical segments, where each vertical segment is denoted aswith j=1, …, W. Let the number of heads be d and the head dimension be D=W / d. For X, we denote the i-th head feature as Xi, similar to the projection operation in C-MSA, A-MSA also applies the operations φ (. ) , Fl (·) and Fd (·) to compute the projections: andSubsequently, after feature reshaping, the resulting projections areandTherefore, the formulation for the i-th single-head axial-wise attention can be expressed as follows:whereis the single-head channel-wise self-attention, βi is the learnable scaling parameter. The axial-wise multi-head self-attention is depicted as follows:whererepresents the operation of 1×1 convolution.Intra-Axial Multi-head Self-Attention (IA-MSA) . AMSA captures the relationship among various vertical lines but overlooks the correlation among pixels within each individual vertical line, which is important in displacement and duplicate artifacts removal. As a result, we introduced an additional self-attention module to improve the feature representation within each vertical line. As depicted in subplot (d) of FIG. 2, considering the input tensorwe continue to partition X into W vertical lines. Each vertical line feature, denoted as (where j=1, …, W) , consists of H tokens with C dimensions. Assuming there are d heads and each head has a dimension of D=H / d, we represent the i-th head feature of Xj as Xji. Similar to the projection and reshape operations seen in C-MSA and A-MSA, IA-MSA also employs the operations φ (. ) , Fl (·) and Fd (·) to yield the projectionsandConsequently, the i-th single-head intra-axial self-attention mechanism for Xj can be defined as follows:whererepresents the single-head intra-axial self-attention in the j-th vertical line, γji is the learnable scaling parameter. The formulation for intra-axial multi-head self-attention is given as follows:whererepresents the operation of 1×1 convolution.2.1.2. LFFNIn the vanilla Transformer, the FFN (shown in subplot (a) of FIG. 4) uses a MLP combined with a non-linearity to produce the outputwhere X represents the input tensor. However, MLPs possess a notably lower image-specific inductive bias than CNNs and are less adept at leveraging local context. To address this, we propose a LFFN, illustrated in subplot (b) of FIG. 4. Specifically, we replace the standard MLP in the FFN with two simultaneous convolutional blocks and each block contains a 1×1 point-wise convolution layer and 3×3 depth-wise convolution layer to replace the plain MLP in the FFN. This modification results in two projections, X1=Fd1 (φ (X) ) and X2=Fd2 (φ (X) ) , where Fd1 and Fd2 denote the convolution operations. The final output is then derived from the dot product of these two projections, which can be represented as follows:LFFN (X) =Fp (X1⊙g (X2) ) +X (14)where Fp (·) represents the 1×1 convolution layer, ⊙ represents the dot product, and g (·) represent the GeLU activation function.2.2. Stage 2: RCDMWhile the HFormer in Stage 1 effectively addresses displacement and duplicate artifacts, it is important to note that the images processed after Stage 1 may still exhibit white blocks along the border of the image together with the original white line artifacts. As such, RCDM is designed to remove them in the second stage. In the context of removing white line artifacts, we define the following terms for Stage 2: the input image as x, the target clean image as x0, the white line regions as m, the regions within the white lines as xw=m⊙x, and the known clean regions as xc= (1-m) ⊙x. The classical denoising diffusion probabilistic model
[0052] utilizes a forward process that sequentially transforms the data distribution q (x0) intousing Markov diffusion kernels, defined by a fixed variance schedulewhere:This formulation allows for closed-form sampling of xt from x0 at any timestep t:Subsequently, a parameterized Markov chain is trained to reverse this forward process, effectively denoising arbitrary Gaussian noise into a data sample. This reversal is described by the equation:The training process involves maximizing the model loglikelihood with appropriate parameterization and simplification:However, it is noticed that regions beyond the white-line areas do not require reconstruction, thus obviating the necessity of applying the diffusion model to the entire image. Therefore, we design RCDM to optimize the likelihood pθ (xw|xc) under the constraint that the conditional data distribution followsAs the forward process is determined by a Markov Chain involving the addition of Gaussian noise, we have the ability to sample the intermediate imagegiven the known region xc at any time point t. This enables us to consistently retain the known regions xc preventing the introduction of external noise and enabling us to focus more on the white line regions at each time step t. Within the framework of RCDM, we define both the reverse process and the forward process (shown in FIG. 3) . The reverse process, denoted asis characterized as a Markov chain with Gaussian transitions learned from an initial distributionand is formulated as follows:andThe forward process, denoted asis characterized by its gradual introduction of Gaussian noise into the data, following a variance schedule β1, …, βT, which is expressed as follows:With the notationwe can express the following equation:We taketo be σtI, and the optimization objective can be seen as a denoising operation specifically targeting the white line regions, expressed as follows:3. Experiments and Analysis3.1. Implementation DetailsImplementation Details of HFormer in OSAT. For the HFormer architecture, there are four-level encoder-decoder Transformer blocks. The numbers of Transformer blocks range from level-1 to level-4 as follows: [4, 6, 6, 8] . The initial channel number is 48. The numbers of attention heads for C-MSA, A-MSA, and IA-MSA at each level are specified as [1, 2, 4, 8] , [2, 2, 1, 1] , and [2, 2, 1, 1] , respectively. We train the model with a batch size of 8. We utilize the AdamW optimizer with specific parameters: β1=0.9, β2=0.999, and a weight decay of 1×10-4. Our training process spans 300, 000 iterations and comprises two stages. Initially, we train the model for 10, 000 iterations with a starting learning rate of 3×10-4. Subsequently, we employ cosine annealing to gradually reduce the learning rate to 1×10-6 throughout the remaining 290, 000 iterations.Implementation Details of RCDM in OSAT. We employ a U-Net architecture
[0052] as our foundation, incorporating various modifications influenced by recent research
[0047] . Specifically, we adopt the residual blocks introduced in
[0049] , and utilize adaptive group normalization techniques
[0054] . We set the mini-batch size to 8 and the settings of the experiment follow the previous works
[0047] ,
[0052] . We set T = 2000 in training and T = 1000 for inference.Evaluation Metrics. To evaluate the efficacy of different methods in removing motion artifacts, we computed MAE (viz., mean absolute error) , PSNR (viz., peak signal-to-noise ratio) , and SSIM (viz., structural similarity index) scores on the synthetic dataset. For the real clinical data, we invited the doctor to rate the results of the different methods.3.2. DatasetsIn the present study, we have curated two comprehensive synthetic datasets OCTA-6MM and OCTA-3MM for training and testing the motion artifact removal models in OCTA images. In addition, we have incorporated a real clinical image dataset to support visual representation and qualitative analysis of different methods.Synthetic data. We introduce various motion artifacts into the clear OCTA images to generate the data pairs. To create displacement artifacts, we start by dividing the input image I into a variable number of vertical blocks, denoted as M. The specific value of M is randomly chosen from a range of two to ten. Then, each of these vertical blocks, labelled as Im, undergoes one of three operations: staying in place, moving upwards along the vertical axis, or moving downwards along the vertical axis. These operations can mimic displacement artifacts that might occur during image acquisition. According to the recommendation of doctors, we have set the maximum value for the movement to 25. To create displacement artifacts, we randomly divide the input image I into a variable number of vertical blocks, denoted as S with S being randomly chosen from a range of two to ten. Then, we duplicate Is to the right of itself to generate the desired scanning artifact. For the white line artifacts, we randomly selected strips within the images and replaced them with strips containing white lines. Therefore, OCTA-6MM first collects 4, 151 clear 6mm×6mm OCTA images from COINPS. These images are then augmented to synthesize 67, 660 image pairs for training and 768 image pairs for testing. Similarly, OCTA-3MM first collects 4891 clear 3mm×3mm OCTA images from COINPS and then synthesizes 63, 582 image pairs for training and 824 image pairs for testing.Real clinical data. A comprehensive examination was conducted on a cohort of 70 patients at the ophthalmology department of the Sixth Affiliated Hospital of South China University of Technology. The images are obtained by employing a high-speed spectral domain OCTA machine with a scanning frequency of 120kHz, provided by Weiren Meditech (Model C3000) . It encompassed a variety of scanning ranges, including a 6mm×6mm fovea scan and a more expansive 12mm×12mm wide-range scan. The dataset derived from these scans comprises images of retinas under both normal and diseased conditions, including ailments such as diabetic retinopathy, age-related macular degeneration, and others.3.3. Comparison MethodsTo the best of our knowledge, there are currently no existing methods that directly tackle the challenge of eliminating three distinct types of motion artifacts in OCTA images. Therefore, we have undertaken a comprehensive evaluation of two OCTA artifact removal methods, including SR-NET
[0012] and RUNet
[0024] , six SOTA image reconstruction methods in the computer vision field, including DeblurGAN
[0053] , MIMOUnet
[0027] , MPRNet
[0028] , SFormer
[0031] , Uformer
[0030] , and Restormer
[0029] to compare with the proposed OSAT. SRNET
[0012] is designed to remove the white line noise in OCTA images and further enhance the OCTA image quality by using a refinement network. RU-Net
[0024] modified the typical UNet and utilized low-rank regularization during the model training to remove the white line artifacts. DeblurGAN
[0053] incorporated a feature pyramid network within its generator and leveraged a multi-scale discriminator to enhance image restoration. MIMO
[0027] employed a coarse-to-fine approach and utilized a multi-input multi-output U-net for superior feature representation. MPRNet
[0028] adopted a multi-stage architecture for progressive learning of restoration functions. SFormer
[0031] integrated efficient self-attention modules into the U-net to process high-level features for image restoration. Uformer
[0030] incorporated the shift window into self-attention to enhance local features in the transformer block. Restormer
[0030] utilized a single efficient self-attention module to enhance global features. Notably, among these six methods, DeblurGAN, MIMO-Unet, and MPRNet are based on CNN, whereas SFormer, Uformer, and Restormer are based on the transformer. In Stage 2, we extend our comparison to include two SOTA methods for mitigating white line artifacts: the transformer-based approach T-Former
[0032] and the diffusion method RePaint
[0033] .Table I: Comparison between OSAT and various methods on OCTA-6MM and OCTA-3MM, where SR-Net
[0012] and RU-Net
[0024] are specifically designed for OCTA images.Table II: User study on real clinical data.Table III: Comparison of different stage 1 methods in OSAT.Table IV: Comparison of different stage 2 methods in OSAT.Table V: Comparison of OSAT artifact removal results with different settings of self-attention in HFormer.Table VI: Comparison of OSAT artifacts removal with different FFNs in HFormer.Table VII: Effect of loss function design of HFormer.Effect of the loss function in HFormer. We compare the effectiveness ofandwithandin Table VII. The results show thatoutperforms theand loss. Additionally, when we removed thewe noticed a decrease in artifact removal performance, highlighting the importance and effectiveness ofin Stage 1.3.4. Motion Artifacts Removal ResultsResults on the synthetic datasets. The artifact removal results on OCTA-6MM and OCTA-3MM are summarized in Table I. It first showcases a comparative analysis between OSAT and eight one-stage methods, involving two recent OCTA-specific models SR-NET
[0012] and RU-Net
[0024] , and six SOTA image restoration methods in the CV field, including DeblurGAN
[0053] , MIMO-Unet
[0027] , MPRNet 28] , SFormer
[0031] , Uformer
[0030] , and Restormer
[0029] . Furthermore, we also compared our OSAT with different two-stage methods. As illustrated in Table I, while the six CV methods outperform SR-NET and RU-Net, these one-stage approaches still struggle to concurrently address displacement, duplication scanning, and white line artifacts and achieve unsatisfied performance in MAE, PSNR, and SSIM metrics. Besides, Restormer and Uformer outperform the other CNN-based methods, demonstrating the superiority of the transformer methods. MPRNet achieves better performance than SFormer owing to the multi-scale aggregation scheme and the additional refinement networks, and these complex operations boost the capability of MPRNet. Furthermore, SFormer still utilizes the CNN backbone and the self-attention blocks are employed as the additional attention module, which limits the global representation of SFormer. The proposed OSAT removes three types of motion artifacts sequentially and surpasses all one-stage methods, demonstrating the superiority of the two-stage framework. Within the two-stage framework, we evaluate six CV methods alongside our proposed HFormer during the first stage. In the second stage, we compare the proposed RCDM with the other two distinct inpainting techniques, T-former
[0032] and Repaint
[0033] . The results of the two-stage methods are summarized in Table I, demonstrating the superiority of OSAT on artifact removal with much better MAE, PSNR, and SSIM results for both the OCTA-6MM and OCTA-3MM datasets. The HFormer in OSAT captures feature representations across local, global, and vertical ranges, significantly bolstering its potential for streamlined motion artifact removal. Furthermore, the guidance of the clear condition regions makes RCDM effective in white line artifact removal. To further analyze the impact of the HFormer and RCDM, we replace them with other alternative SOTA methods, and the results are summarized in Table III and Table IV, respectively.Results on real clinical data. We conducted a comprehensive assessment of our OSAT using a dataset consisting of 70 real clinical OCTA images sourced from a hospital setting. To ensure the robustness and clinical applicability of our approach, we sought the input of experienced medical professionals who meticulously evaluated and assigned scores to the outputs generated by various techniques when applied to this real-world dataset. We compare the OSAT with other two-stage methods, in which the DeblurGAN, MIMO, MPRNet, SFormer, UFormer, and Restormer are involved in the stage 1 and the RCDM as the stage 2 method. The results of this evaluation are summarized in Table II, which unmistakably highlights the superior average quality score achieved by OSAT in comparison to other two-stage methods (p-value < 0.05) . We can observe the performance of various combinations in FIG. 5, where areas containing motion artifacts are distinctly outlined with red rectangles, while critical points are indicated by yellow arrows. These visualizations serve as confirmation of the remarkable ability of HFormer and RCDM in OSAT to generalize effectively when presented with the intricate motion artifacts within the OCTA image.3.5. Ablation StudyEffect of HFormer in OSAT. We evaluate the effect of the proposed HFormer in comparison to six other methods including DeblurGAN
[0053] , MIMO-Unet
[0027] , MPRNet
[0028] , SFormer
[0031] , Uformer
[0030] , and Restormer
[0029] in stage 1, and all these two-stage frameworks using the RCDM as the stage 2 model. As shown in Table III, HFormer outperforms other methods in all the MAE, PSNR, and SSIM metrics on both OCTA-6MM and OCTA-3MM datasets. By capturing the long-dependence relation in OCTA images, transformers can effectively remove the motion artifacts and achieve better performance on motion artifact removal than the CNN-based methods. Furthermore, the proposed HFormer employed hierarchical self-attention modules to capture a comprehensive range of feature representations, thereby boosting the capability of the model.Effect of RCDM in OSAT. We also evaluate the effect of RCDM in comparison to T-Former and RePaint in stage 2, T-Former is the SOTA transformer-based image completion method and RePaint is a diffusion-model-based method. In this case, all the frameworks utilize HFormer as the stage 1 method. The results are summarized in Table IV. The proposed RCDM outperformed the compared two methods, while RePaint achieved better performance than T-former. The thousands of feedforward and inverse processes enhance the ability of RePaint in artifact removal and RCDM outperforms RePaint due to more focus on the white line regions.Effect of different settings of self-attention. We conducted an ablation study to assess the effectiveness of different self-attention modules within the HFormer architecture. We maintained the foundational self-attention module, C-MSA, throughout the study. Additionally, we examined the impact of two other modules, A-MSA and IA-MSA, by removing them individually. To ensure fairness in comparisons, we replaced any removed module with an instance of C-MSA to maintain consistent model size. The results in Table V indicate that removing IA-MSA led to a decrease in PSNR by 0.53 and 0.65 for OCTA-6MM and OCTA-3MM, respectively. In contrast, excluding A-MSA resulted in a more substantial drop in PSNR, with reductions of 1.00 and 0.98 for OCTA-6MM and OCTA-3MM, respectively. This underscores the greater importance of A-MSA compared to IA-MSA, emphasizing that inter-strip relationships are more significant than relationships within individual vertical strips. We also conducted additional experiments to explore the impact of the sequence order of IA-MSA and A-MSA. When applying IA-MSA followed by A-MSA the results, as presented in Table V, demonstrated a superior representation compared to the reverse sequenceEffect of FFN. We compared the LFFN with the traditional MLP FFN by using an expansion rate of 4. The results are summarized in Table VI, clearly demonstrating that the LFFN outperforms the plain MLP in all metrics, including MAE, PSNR, and SSIM. This provides strong evidence for the superior performance of LFFN in artifact removal, emphasizing the effectiveness of local representations.Effect of the loss function in HFormer. We compared the effectiveness ofandwithandin Table VII. The results show thatoutperforms theand loss. Additionally, when we removed thewe noticed a decrease in artifact removal performance, highlighting the importance and effectiveness ofin Stage 1.4. Embodiments of Present DisclosureEmbodiments of the present disclosure are developed as follows based on the details, examples, applications, etc. regarding the OSAT as disclosed above possibly with generalization and extension.An aspect of the present disclosure is to provide a computer-implemented method for reducing eye motion artifacts in a degraded OCTA image to yield a quality-enhanced OCTA image.The disclosed method is illustrated with the aid of FIG. 6, which depicts a workflow 600 showing exemplary steps of the disclosed method. Exemplarily, the disclosed method comprises steps 610, 620, 630 and 640. The steps 610 and 620 are initialization steps. The steps 630 and 640 collectively realize the OSAT.In the step 610, a HTM configured to remove displacement artifacts and duplicated scanning artifacts from a first input OCTA image is set up. The HTM is arranged to receive the first input OCTA image and as a result, yield a first output OCTA image as an output of the HTM. Preferably and advantageously, the HTM is realized as any one of the embodiments of HFormer as disclosed above.The HTM comprises a plurality of transformer blocks. The plurality of transformer blocks is configured to employ self-attention to capture a plurality of features utilizable for removing the displacement artifacts and duplicated scanning artifacts from the first input OCTA image. In setting up the HTM, as demonstrated above for HFormer, it is preferable and advantageous that the plurality of features includes global features, vertical-line features and intra-vertical-line features of the first input OCTA image. As a result, the plurality of transformer blocks includes a first plurality of transformer blocks for capturing the global features, a second plurality of transformer blocks for capturing the vertical-line features, and a third plurality of transformer blocks for capturing the intra-vertical-line features.In generalization, the HTM is configured to utilize the plurality of features, which includes the global features, vertical-line features and intra-vertical-line features, to remove the displacement artifacts and duplicated scanning artifacts from the first input OCTA image.In certain embodiments, the first, second and third pluralities of transformer blocks are configured as follows. A first transformer block in the first plurality of transformer blocks is configured to conduct a first plurality of self-attention operations (which are channel-wise self-attention operations) across channels of a first input feature map received by the first transformer block. A second transformer block in the second plurality of transformer blocks is configured to conduct a second plurality of self-attention operations (which are axis-wise self-attention operations) across vertical lines of a second input feature map received by the second transformer block such that relationships among the vertical lines are learnt. A third transformer block in the third plurality of transformer blocks is configured to conduct a third plurality of self-attention operations (which are intra-axial self-attention operations) across different pixels within an individual vertical line of a third input feature map received by the third transformer block such that information on correlation among the pixels within the individual vertical line is extracted. More details on the first, second and third pluralities of self-attention operations are provided in Section 2.1.1 and in subplots (b) -(d) of FIG. 2.In certain embodiments of realizing the first, second and third transformer blocks, an individual transformer block in the plurality of transformer blocks is formed by including a MSA module and a LFFN module. Based on the MSA module and the LFFN module, the computation performed by the individual transformer block is governed by EQNS. (6) and (7) .Note that channel-wise MSA is employed by respective first transformer blocks in the first plurality of transformer blocks for capturing the global features. Similarly, axial-wise MSA is employed by respective second transformer blocks in the second plurality of transformer blocks for capturing the vertical-line features. Also similarly, intra-axial MSA is employed by respective third transformer blocks in the third plurality of transformer blocks for capturing the intra-vertical-line features.In certain embodiments, the LFFN module is formed with two simultaneous convolutional blocks. Each of the two simultaneous convolutional blocks contains a 1×1 point-wise convolution layer and a 3×3 depth-wise convolution layer. As mentioned above, using the LFFN module has advantages over using a MLP in that the MLP possesses a notably lower image-specific inductive bias than a CNN, which is used in the LFFN module, and are less adept at leveraging local context.The HTM adopts a U-shaped hierarchical network following an encoder-decoder architecture. As a result, respective transformer blocks in the plurality of transformer blocks can generally be classified into encoder blocks and decoder blocks. In certain embodiments, the U-shaped hierarchical network includes skip connections between the encoder blocks and the decoder blocks. Alternatively, it is also possible that the skipped connections are absent in the U-shaped hierarchical network adopted by HTM.The HTM may be a pre-trained model such that model parameters of the HTM are pre-stored and the pre-stored model parameters are loaded into the HTM in the step 610. Alternatively, the HTM may be trained during setting up the HTM in the step 610. In certain embodiments, the step 610 comprises training the HTM. In certain embodiments, the HTM is trained according to a loss function given by EQN. (1) .In the step 620, a RCDM configured to remove white line artifacts from a second input OCTA image is set up. The RCDM is arranged to receive the second input OCTA image and thereby yield a second output OCTA image as an output of the RDCM.The RCDM is a diffusion model for reconstructing a first region of the second input OCTA image overlapping with the white line artifacts while avoiding modifying a second region of the second input OCTA image not overlapping with the white line artifacts. The first and second regions may be named as an artifact region and a residual region, respectively. It is preferable and advantageous that this diffusion model is configured to optimize a likelihood of correctly reconstructing the first region given that the reconstructing of the first region is conditioned on image data of the second region.In generalization, the RCDM is configured to reconstruct the first region while avoiding modifying the second region to remove the white line artifacts from the second input OCTA image.Similar to the HTM, the RCDM may be a pre-trained model such that model parameters of the RCDM are pre-stored and the pre-stored model parameters are loaded into the RCDM in the step 620. Alternatively, the RCDM may be trained during setting up the RCDM in the step 620. In certain embodiments, the step 620 comprises training the RCDM.The step 630 reflects realization of the first stage of OSAT. In the step 630, the HTM is employed to remove any displacement artifact and any duplicated scanning artifacts from the degraded OCTA image to yield an intermediate OCTA image. Note that the HTM is used to process the degraded OCTA image in the step 630 after the HTM is trained.The step 640 reflects realization of the second stage of OSAT. In the step 640, the RCDM is employed to remove any white line artifact from the intermediate OCTA image obtained in the step 630 to thereby yield the quality-enhanced OCTA image. Note that the RCDM is used to process the intermediate OCTA image in the step 640 after the RCDM is trained.A computing system for reducing eye motion artifacts in a degraded optical OCTA image to yield a quality-enhanced OCTA image is realizable by including one or more computers, where the one or more computers are configured to execute a process of reducing the eye motion artifacts in the degraded OCTA image to yield the quality-enhanced OCTA image according to any of the embodiments of the disclosed method. An individual computer may be a general-purpose computer, a special-purpose computer such as the one implemented with artificial intelligence processor (s) or graphics processing unit (s) , a desktop computer, a physical computing server, a distributed computing server, or a mobile computing device such as a smartphone and a tablet computer.The present disclosure may be embodied in other specific forms without departing 1from the spirit or essential characteristics thereof. The present embodiment is therefore to be considered in all respects as illustrative and not restrictive. The scope of the invention is indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.REFERENCES There follows a list of references that are occasionally cited in the specification. Each of the disclosures of these references is incorporated by reference herein in its entirety.[1] C. Wang, H. Ning, X. Chen, and S. Li, “Db-unet: Mlp based dual branch unet for accurate vessel segmentation in octa images, ” in ICASSP, 2023.[2] R.F. Spaide, J.G. Fujimoto, N.K. Waheed, S.R. Sadda, and G. Staurenghi, “Optical coherence tomography angiography, ” Prog. Retin. Eye Res., 2018.[3] W. Yang, Y. Shao, and Y. Xu, “Guidelines on clinical research evaluation of artificial intelligence in ophthalmology, ” in Int J Ophthalmol, 2023.[4] C. -H. Hua, K. Kim, T. Huynh-The, J.I. You, S. -Y. Yu, T. Le-Tien, S. -H. Bae, and S. Lee, “Convolutional network with twofold feature augmentation for diabetic retinopathy recognition from multi-modal images, ” IEEE J. Biomed. Health., 2020.[5] E.W. Schneider and S.C. Fowler, “Optical coherence tomography angiography in the management of age-related macular degeneration, ” Curr Opin Ophthalmol, 2018.[6] Y. Jia, J.M. Simonett, J. Wang, X. Hua, L. Liu, T. S. Hwang, and D. Huang, “Wide-field oct angiography investigation of the relationship between radial peripapillary capillary plexus density and nerve fiber layer thickness, ” Investig. Ophthalmol. Vis. Sci., 2017.[7] T.T. Hormel, D. Huang, and Y. Jia, “Artifacts and artifact removal in optical coherence tomographic angiography, ” Quant. Imaging. Med. Surg., 2021.[8] I.C. Holmen, S.M. Konda, J.W. Pak, K.W. McDaniel, B. Blodi, K. E. Stepien, and A. Domalpally, “Prevalence and severity of artifacts in optical coherence tomographic angiograms, ” JAMA Ophthalmol., 2020.[9] Y. Zhou, K. Yu, M. Wang, Y. Ma, Y. Peng, Z. Chen, W. Zhu, F. Shi, and X. Chen, “Speckle noise reduction for oct images based on image style transfer and conditional gan, ” IEEE J. Biomed. Health., 2021.
[0010] R.F. Spaide, J.G. Fujimoto, and N.K. Waheed, “Image artifacts in optical coherence angiography, ” Retina, 2015.
[0011] P. Anvari, M. Ashrafkhorasani, A. Habibi, and K. G. Falavarjani, “Artifacts in optical coherence tomography angiography, ” J. Ophthalmic Vis. Res., 2021.
[0012] J. Cao, Z. Xu, M. Xu, Y. Ma, and Y. Zhao, “Atwo-stage framework for optical coherence tomography angiography image quality improvement, ” Front. Med., 2023.
[0013] K. Hu, S. Jiang, Y. Zhang, X. Li, and X. Gao, “Joint-seg: Treat foveal avascular zone and retinal vessel segmentation in octa images as a joint task, ” IEEE Trans. Instrum. Meas., 2022.
[0014] S. Chen, B. Potsaid, Y. Li, J. Lin, Y. Hwang, E.M. Moult, J. Zhang, D. Huang, and J.G. Fujimoto, “High speed, long range, deep penetration swept source oct for structural and angiographic imaging of the anterior eye, ” Sci. Rep., 2022.
[0015] A. Yasin Alibhai, C. Or, and A.J. Witkin, “Swept source optical coherence tomography: a review, ” Curr. Ophthalmol. Rep., 2018.
[0016] M.F. Kraus, B. Potsaid, M.A. Mayer, R. Bock, B. Baumann, J.J. Liu, J. Hornegger, and J.G. Fujimoto, “Motion correction in optical coherence tomography volumes on a per a-scan basis using orthogonal scan patterns, ” Biomed. Opt. Express, 2012.
[0017] A. Wolf, K. Tripanpitak, S. Umeda, and M. Otake-Matsuura, “Eyetracking paradigms for the assessment of mild cognitive impairment: a systematic review, ” Front. Psychol., 2023.
[0018] A. Uji, S. Balasubramanian, J. Lei, E. Baghdasaryan, M. Al-Sheikh, and S.R. Sadda, “Choriocapillaris imaging using multiple en face optical coherence tomography angiography image averaging, ” JAMA Ophthalmol., 2017.
[0019] X. Wei, A. Camino, S. Pi, W. Cepurna, D. Huang, J.C. Morrison, and Y. Jia, “Fast and robust standard-deviation-based method for bulk motion compensation in phase-based functional oct, ” Opt. Lett., 2018.
[0020] Y. Zhang, W. Gao, and C. Xie, “Fourier spatial transform-based method of suppressing motion noises in octa, ” Opt. Lett., 2022.
[0021] A. Camino, Y. Jia, G. Liu, J. Wang, and D. Huang, “Regression-based algorithm for bulk motion subtraction in optical coherence tomography angiography, ” Biomed. Opt. Express, 2017.
[0022] P. Zang, G. Liu, M. Zhang, C. Dongye, J. Wang, A.D. Pechauer, T.S. Hwang, D.J. Wilson, D. Huang, D. Li et al., “Automated motion correction using parallel-strip registration for wide-field en face oct angiogram, ” Biomed. Opt. Express, 2016.
[0023] A. Camino, M. Zhang, C. Dongye, A.D. Pechauer, T.S. Hwang, S.T. Bailey, B. Lujan, D.J. Wilson, D. Huang, and Y. Jia, “Automated registration and enhanced processing of clinical optical coherence tomography angiography, ” Quant. Imaging. Med. Surg., 2016.
[0024] D. Gao, N. Celik, X. Wu, B.M. Williams, A. Stylianides, and Y. Zheng, “A novel deep learning based octa de-striping method, ” in MIUA, 2020.
[0025] A. Li, C. Du, and Y. Pan, “Deep-learning-based motion correction in optical coherence tomography angiography, ” J. Biophotonics, 2021.
[0026] J. Ren, K. Park, Y. Pan, and H. Ling, “Self-supervised bulk motion artifact removal in optical coherence tomography angiography, ” in CVPR, 2022.
[0027] S. -J. Cho, S. -W. Ji, J. -P. Hong, S. -W. Jung, and S. -J. Ko, “Rethinking coarse-to-fine approach in single image deblurring, ” in ICCV, 2021.
[0028] S.W. Zamir, A. Arora, S. Khan, M. Hayat, F.S. Khan, M. -H. Yang, and L. Shao, “Multi-stage progressive image restoration, ” in CVPR, 2021.
[0029] S.W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. -H. Yang, “Restormer: Efficient transformer for high-resolution image restoration, ” in CVPR, 2022.
[0030] Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li, “Uformer: A general u-shaped transformer for image restoration, ” in CVPR, 2022.
[0031] F. -J. Tsai, Y. -T. Peng, Y. -Y. Lin, C. -C. Tsai, and C. -W. Lin, “Stripformer: Strip transformer for fast image deblurring, ” in ECCV, 2022.
[0032] Y. Deng, S. Hui, S. Zhou, D. Meng, and J. Wang, “T-former: An efficient transformer for image inpainting, ” in ACM MM, 2022.
[0033] A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, “Repaint: Inpainting using denoising diffusion probabilistic models, ” in CVPR, 2022.
[0034] P.G. Daneshmand, A. Mehridehnavi, and H. Rabbani, “Reconstruction of optical coherence tomography images using mixed low rank approximation and second order tensor based total variation method, ” IEEE Trans. Med. Imaging, 2020.
[0035] X. Wu, D. Gao, D. Borroni, S. Madhusudhan, Z. Jin, and Y. Zheng, “Cooperative low-rank models for removing stripe noise from octa images, ” IEEE J. Biomed. Health., 2020.
[0036] F. Yang, H. Yang, J. Fu, H. Lu, and B. Guo, “Learning texture transformer network for image super-resolution, ” in CVPR, 2020.
[0037] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “SwinIR: Image restoration using swin transformer, ” in ICCV Workshops, 2021.
[0038] M. Kumar, D. Weissenborn, and N. Kalchbrenner, “Colorization transformer, ” in ICLR, 2021.
[0039] H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer, ” in CVPR, 2021.
[0040] Z. Wang, X. Cun, J. Bao, and J. Liu, “Uformer: A general u-shaped transformer for image restoration, ” arXiv: 2106.03106, 2021.
[0041] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows, ” arXiv: 2103.14030, 2021.
[0042] M. Ding, B. Xiao, N. Codella, P. Luo, J. Wang, and L. Yuan, “Davit: Dual attention vision transformers, ” in ECCV, 2022.
[0043] Z. -X. Cui, C. Cao, S. Liu, Q. Zhu, J. Cheng, H. Wang, Y. Zhu, and D. Liang, “Self-score: Self-supervised learning on score-based models for mri reconstruction, ” arXiv: 2209.00835, 2022.
[0044] C. Cao, Z. -X. Cui, S. Liu, H. Zheng, D. Liang, and Y. Zhu, “High-frequency space diffusion models for accelerated mri, ” arXiv: 2208.05481, 2022.
[0045] W. Xia, Q. Lyu, and G. Wang, “Low-dose ct using denoising diffusion probabilistic model for 20x speedup, ” arXiv: 2209.15136, 2022.
[0046] Y. Shi and G. Wang, “Conversion of the mayo ldct data to synthetic equivalent through the diffusion model for training denoising networks with a theoretically perfect privacy, ” arXiv: 2301.06604, 2023.
[0047] C. Saharia, J. Ho, W. Chan, T. Salimans, D.J. Fleet, and M. Norouzi, “Image super-resolution via iterative refinement, ” IEEE Trans. Pattern Anal. Mach. Intell., 2022.
[0048] J. Ho, C. Saharia, W. Chan, D.J. Fleet, M. Norouzi, and T. Salimans, “Cascaded diffusion models for high fidelity image generation, ” J. Mach. Learn. Res., 2022.
[0049] C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi, “Palette: Image-to-image diffusion models, ” in SIGGRAPH, 2022.
[0050] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models, ” in CVPR, 2022.
[0051] P. Charbonnier, L. Blanc-Feraud, G. Aubert, and M. Barlaud, “Two deterministic half-quadratic regularization algorithms for computed imaging, ” in ICIP, 1994.
[0052] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models, ” NeurIPS, 2020.
[0053] O. Kupyn, T. Martyniuk, J. Wu, and Z. Wang, “Deblurgan-v2: Deblurring (orders-of-magnitude) faster and better, ” in ICCV, 2019.
[0054] P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis, ” NeurIPS, 2021.
Claims
1.A computer-implemented method for reducing eye motion artifacts in a degraded optical coherence tomography angiography (OCTA) image to yield a quality-enhanced OCTA image, the method comprising:setting up a hierarchical transformer model (HTM) configured to remove displacement artifacts and duplicated scanning artifacts from a first input OCTA image to thereby yield a first output OCTA image;setting up a residual conditional diffusion model (RCDM) configured to remove white line artifacts from a second input OCTA image to thereby yield a second output OCTA image;using the HTM to remove any displacement artifact and any duplicated scanning artifacts from the degraded OCTA image to yield an intermediate OCTA image; andusing the RCDM to remove any white line artifact from the intermediate OCTA image to yield the quality-enhanced OCTA image.2.The method of claim 1, wherein the HTM is further configured to utilize a plurality of features to remove the displacement artifacts and duplicated scanning artifacts from the first input OCTA image, and wherein the plurality of features includes global features, vertical-line features and intra-vertical-line features of the first input OCTA image.3.The method of claim 1, wherein the HTM comprises a plurality of transformer blocks, the plurality of transformer blocks being configured to employ self-attention to capture a plurality of features utilizable for removing the displacement artifacts and duplicated scanning artifacts from the first input OCTA image, and wherein the plurality of features includes global features, vertical-line features and intra-vertical-line features of the first input OCTA image.4.The method of claim 1, wherein the HTM has a U-shaped hierarchical network structure incorporated with skip connections between encoder and decoder blocks of the HTM.5.The method of claim 3, wherein an individual transformer block for processing an input feature map X to yield an output feature map comprises a multi-head self-attention (MSA) module and a local feed-forward network (LFFN) such that the individual transformer block computes X′=MSA (X) +Xandwhere MSA (·) and LFFN (·) are mathematical functions performed by the MSA module and by the LFFN, respectively.6.The method of claim 5, wherein the LFFN is formed with two simultaneous convolutional blocks, and wherein each of the two simultaneous convolutional blocks contains a 1×1 point-wise convolution layer and a 3×3 depth-wise convolution layer.7.The method of claim 5, wherein the plurality of transformer blocks includes:a first plurality of transformer blocks for capturing the global features;a second plurality of transformer blocks for capturing the vertical-line features; anda third plurality of transformer blocks for capturing the intra-vertical-line features.8.The method of claim 7, wherein respective transformer blocks in the first plurality of transformer blocks are configured to employ channel-wise MSA for capturing the global features.9.The method of claim 7, wherein respective transformer blocks in the second plurality of transformer blocks are configured to employ axial-wise MSA for capturing the vertical-line features.10.The method of claim 7, wherein respective transformer blocks in the third plurality of transformer blocks are configured to employ intra-axial MSA for capturing the intra-vertical-line features.11.The method of claim 1, wherein the HTM is trained to minimize a loss function, given by where: I′ represents a restored OCTA image due to using the HTM; represents a ground-truth OCTA image corrupted with white line artifacts; is given byε being a predetermined offset; is given byΔ denoting a Laplacian operator; and λ is a predetermined parameter controlling relative importance betweenand12.The method of claim 1, wherein the setting up of the HTM includes training the HTM before the HTM is used to process the degraded OCTA image.13.The method of claim 1, wherein the HTM is realized as HFormer.14.The method of claim 1, wherein the RCDM is configured to reconstruct a first region of the second input OCTA image overlapping with the white line artifacts while avoiding modifying a second region of the second input OCTA not overlapping with the white line artifacts for removing the white line artifacts from the second input OCTA image.15.The method of claim 1, wherein the RCDM is a diffusion model configured to optimize a likelihood of correctly reconstructing a first region of the second input OCTA image overlapping with the white line artifacts given that the reconstructing of the first region is conditioned on image data of a second region of the second input OCTA not overlapping with the white line artifacts.16.The method of claim 1, wherein the setting up of the RCDM includes training the RCDM before the RCDM is used to process the intermediate OCTA image.17.A computing system for reducing eye motion artifacts in a degraded optical coherence tomography angiography (OCTA) image to yield a quality-enhanced OCTA image, the computing system comprising one or more computers configured to execute a process of reducing the eye motion artifacts in the degraded OCTA image to yield the quality-enhanced OCTA image according to the method of any one of the preceding claims.
Citation Information
Patent Citations
OCTA image motion correction method based on deep learning
CN114332278A
Medical image segmentation method based on conditional diffusion model
CN116596949A
Eye fundus blood vessel image segmentation method and system based on hierarchical diffusion model
CN117237640A
Method and system for enhancing and segmenting low-quality eye fundus image based on diffusion model
CN117314935A
Method and apparatus for motion correction and image enhancement for optical coherence tomography
US20110267340A1
Cited By
Diabetic retinopathy early screening method based on diffusion model
CN122066711A