Correcting flow projection artifacts in OCTA volumes using neural networks

By using a neural network architecture to correct the flow projection artifacts in the OCTA volume, the problems of high computational cost and reliance on manual assumptions in existing technologies are solved, and a fast and effective multi-plane correction effect is achieved.

CN115349137BActive Publication Date: 2025-10-31CARL ZEISS MEDITEC INC +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202180025471.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2021-03-26
Publication Date
2025-10-31
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing OCTA images contain flow projection artifacts that are difficult to remove effectively, especially in volume-based methods which are computationally intensive, rely on manual assumptions, and cannot be corrected in multiple planes.

Method used

A neural network architecture is used to correct flow projection artifacts in the OCTA volume. The architecture is trained using structural and flow data, independent of segmentation and plate constraints, and the artifact correction is performed using a U-Net-type neural network.

Benefits of technology

It achieves faster artifact correction, can correct flow artifacts in multiple planes and three-dimensional space, reduces computational resource requirements, and is suitable for clinical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115349137B_ABST
    Figure CN115349137B_ABST
Patent Text Reader

Abstract

A system and / or method uses a trained U-Net neural network to remove flow artifacts from optical coherence tomography (OCT) angiography (OCTA) data. The trained U-Net receives OCT structural volume data and OCTA volume data as input, but expands the OCTA volume data to include depth information. The U-Net applies dynamic pooling along the depth direction and gives greater weight to the portion of the data that follows (e.g., along the contour) the selection of retinal layers. In this way, the U-Net applies context-different computations at different axial locations, at least in part, based on depth index information and / or (e.g., proximity) the selected retinal layers. Compared to the input OCTA data, the U-Net output is OCT volume data with reduced flow artifacts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to improving optical coherence tomography (OCT) images and OCT angiography images. More specifically, it relates to removing flow artifacts / derecognition tails in OCT-based images. Background Technology

[0002] Optical coherence tomography (OCT) is a non-invasive imaging technique that uses light waves to generate cross-sectional images of tissues, such as retinal tissue. For example, OCT allows people to view the unique tissue layers of the retina. Typically, an OCT system is an interferometric imaging system that determines the scattering distribution of the sample along the OCT beam by detecting the interference of light reflected from the sample and a reference beam, thereby creating a three-dimensional (3D) representation of the sample. Each scattering distribution in the depth direction (e.g., the z-axis or axial direction) is reconstructed individually as an axial scan or A-scan. Cross-sections, two-dimensional (2D) images (B-scans), and extended 3D volumes (C-scans or cube scans) can be constructed from multiple A-scans acquired as the OCT beam scans / moves through a set of lateral (e.g., x-axis and y-axis) positions on the sample. OCT also allows the construction of 2D images of selected portions of a tissue volume (e.g., a target slab or target tissue layer of the retina) in a frontal view (e.g., an en-face view). An extension of OCT is OCT angiography (OCTA), which identifies (e.g., presented in image format) blood flow in tissue layers. OCTA can identify blood flow by recognizing differences (e.g., contrast differences) over time in multiple OCT images of the same retinal region, and designates differences that meet predefined criteria as blood flow.

[0003] OCT is susceptible to various types of image artifacts, including decorrelation tails or shadows, where structures / formations in the upper tissue layers (e.g., tissue or blood vessel formation) create "shadows" in the lower tissue layers. Specifically, OCTA is prone to flow projection artifacts, where vascular images can be presented in incorrect locations. This can be due to the high scattering properties of blood within overlying vessels, producing artifacts that interfere with the interpretation of retinal angiography results. In other words, deeper tissue layers may exhibit projection artifacts due to variations in reflected signals caused by the undulating shadows projected by blood flowing in the large inner retinal vessels above them. These signal variations can be misinterpreted as (blood) flow, making them difficult to distinguish from genuine flow.

[0004] Several methods have been developed to attempt to overcome these problems, either by correcting artifacts in the previously defined and generated frontal panel or by correcting artifacts in the OCT volume. Examples of plate-based methods for correcting projection artifacts in the frontal panel can be found in the following: “A Fast Method to Reduce Decorrelation TailArtifacts in OCT Angiography”, H. Bagherinia et al., Investigative Ophthalmology & Visual Science, 2017, 58(8), 643-643; “Projection Artifact Removal Improves Visualization and Quantitation of Macular Neovascularization Imaged by Optical Coherence Tomography Angiography”, Zhang Q. et al., Ophthalmol Retina, 2017, 1(2), 124–136; and “Minimizing projection artifacts for accurate presentation of choroidal neovascularization in OCT micro-angiography”, Anqi Zhang et al., Biomedical Optics Express, 2015, Vol. 6, No. 10, all of which are incorporated herein by reference in their entirety. Typically, board-based methods may have some insurmountable limitations and dependencies (e.g., they are piecewise dependent) and do not allow visualization of the corrected data in a plane other than the target board. Therefore, they do not allow the use of 3D techniques to visualize, segment, or quantize OCTA flow characteristics. Board-based methods may also produce suboptimal processing workflows where an artifact correction algorithm must be executed every time the target board constraints change, regardless of how small the change is or whether the current target board constraints have been restored to the state of the previous step.

[0005] Examples of volume-based methods for correcting projection artifacts in OCT volumes are described below: U.S. Patent No. 10,441,164, assigned to the same assignee as this invention; “Projection-resolved optical coherencetomographic angiography”, Zhang M et al., Biomed Opt Express, 2016, Vol. 7, No. 3; “Visualization of 3 Distinct Retinal Plexuses by Projection-Resolved Optical Coherence Tomography Angiography in Diabetic Retinopathy”, Hwang TS et al., JAMAOphthalmol. 2016; 134(12); “Volume-Rendered Projection-Resolved OCT Angiography: 3D Lesion Complexity is Associated with Therapy Response in Wet Age-Related Macular Degeneration”, Nesper PL et al., Invest Ophthalmol Vis Sci., 2018; Vol. 59, No. 5; and “Projection Resolved Optical Coherence Tomography Angiography to Distinguish "Flow Signal in Retinal Angiomatous Proliferation from FlowArtifact", Fayed AE et al., PLOS ONE, 2019, 14(5), all of which are incorporated herein by reference in their entirety. Generally, volume-based methods overcome some of the problems found in plate-based methods and allow visualization of corrected flow data in planes outside the (target) frontal panel (e.g., in B-scans), and allow for the processing of corrected volumetric data. However, volume-based methods can be slow because they require the analysis of large 3D data arrays and rely on handcrafted assumptions that may not be applicable to all vascular manifestations.

[0006] What is needed is a volume-based flow artifact correction method that is fast and provides results as good as plate-based methods. This method is well-established in the industry but does not depend on segmentation and is not subject to other limitations of plate-based methods.

[0007] One object of the present invention is to provide a volume-based flow artifact correction method that provides faster results than existing methods.

[0008] Another object of the present invention is to provide a method for flow artifact correction that obtains results similar to those of custom mathematical formula methods, but is characterized by its ease of parallelization in computer processing.

[0009] Another object of the present invention is to provide a volume-based flow artifact correction system that can be easily implemented using the computational power of existing OCT systems without imposing an excessive time burden on existing clinical procedures. Summary of the Invention

[0010] The aforementioned objectives are met by methods / systems for correcting (e.g., removing or reducing) flow artifacts in optical coherence tomography (OCTA) using neural network methods. If a mathematical formula were to be constructed to correct flow artifacts in each A-scan, the amount of flow signal due to tail artifacts could be estimated by analyzing frame repetition, modulation characteristics of the OCT signal, and scattering characteristics of the human retina. This approach may provide good results, but this manual formulating method can vary from instrument to instrument and is affected by the different retinal opacities and scattering characteristics of each subject, which complicates its implementation and makes it unsuitable for clinical settings.

[0011] Other manual methods may have similar limitations, namely being overly complex, time-consuming, and / or computationally intensive (e.g., requiring computing resources not available in existing OCT / OCTA systems), particularly when applying flow artifact correction to volumetric scans (e.g., volume-based methods). This invention overcomes some of the limitations found in previous manual methods by using a method / system for correcting projection artifacts in OCTA volumes based on neural networks. This method can be performed faster than manual methods, at least in part because it is easily parallelized. Furthermore, this invention can also correct some isolated errors produced by other volume-based methods in certain vascular manifestations.

[0012] This invention employs a neural network architecture to correct for flow projection artifacts in OCTA volumes and has demonstrated good results in both healthy and diseased subjects, independent of any plate limitation or segmentation. The method can be trained using the original OCT structure volume and OCTA flow volume as input to generate (OCTA) flow volumes (or OCT structure capacities) with no (or reduced) projection / shading artifacts as output. Gold-standard training samples used to train the target output of the neural network (e.g., training samples used as targets, training output samples) can be generated using one or more handcrafted methods, as described above and / or known in the art (including one or more plate-based and / or volume-based algorithms, alone or in combination), correctly decorrelation of tail artifacts (such as flow artifacts or shading), and applied to a sample case set (e.g., sample OCT / OCTA volumes) where most A-scans in each volume are known to show good (or satisfactory) results. Although such handcrafted algorithms (especially volume-based algorithms) may require a large amount of computing power and have long execution times, this is not a burden because their execution time is part of the test data (or training samples) collection phase for training and not part of the execution of the invention (e.g., in this field, such as in a clinical setting, the execution / use of a trained neural network).

[0013] This invention is achieved, at least in part, by employing neural networks that utilize structured and streaming data to address the current problem, and by designing custom neural networks to solve it. In addition to saving time, the current neural network solution also considers the structure and flow of OCTA data analysis. Besides correcting streaming artifacts, the current neural network can also correct residual artifacts that other manual methods may be unable to correct.

[0014] Other objects and achievements, as well as a fuller understanding of the invention, will become apparent and understood by referring to the following description and claims and the accompanying drawings.

[0015] To facilitate understanding of this invention, various publications may be cited or referenced herein. All publications cited or mentioned herein are incorporated herein by reference.

[0016] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited to these embodiments. Any embodiment feature mentioned in one claim class, such as a system, may also be claimed in another claim class (such as a method). Dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter arising from the deliberate reference to any prior claim may also be claimed to disclose any combination of the claims and their features, and may be claimed regardless of the dependency chosen in the appended claims. Attached Figure Description

[0017] In the accompanying drawings, the same reference symbols / characters refer to the same parts:

[0018] Figure 1 An exemplary OCTA B scan of the human retina is shown, where the upper and lower hash lines indicate the positions of the frontal images traversing the surface retinal layer (SRL) and deep retinal layer (DRL), respectively.

[0019] Figure 2 This describes a method for obtaining from, for example Figure 1 A plate-based method for removing flow artifacts in the target positive panel of DRL, and applicable to the present invention, such as the network according to the present invention when defining the training input / output set for neural networks.

[0020] Figure 3 An exemplary training input / output set is shown, which includes a set of training inputs (images) and corresponding training outputs (images).

[0021] Figure 4 This describes the method for defining the training input / output set for neural networks according to the present invention (e.g., Figure 3 The method / system shown).

[0022] Figure 5 A simplified overview of the U-Net architecture used in exemplary embodiments of the present invention is provided.

[0023] Figure 6 Provided Figure 5 A close-up view of the processing steps within a downsampling block (e.g., an encoding module) in the shrinking path of a neural network.

[0024] Figure 7 A method for reducing artifacts in OCT-based images of the eye, according to the present invention, is described.

[0025] Figure 8 A general-purpose frequency-domain optical coherence tomography system for collecting 3D image data of the eye, suitable for use with the present invention, is described.

[0026] Figure 9 An exemplary OCTB scan image of a normal retina of the human eye is shown, and various typical retinal layers and boundaries are illustratively identified.

[0027] Figure 10 An exemplary frontal image of the vascular system is shown.

[0028] Figure 11 An exemplary B-scan blood vessel image is shown.

[0029] Figure 12An example of a multilayer perceptron (MLP) neural network is illustrated.

[0030] Figure 13 A simplified neural network consisting of an input layer, hidden layers, and an output layer is shown.

[0031] Figure 14 This illustrates an example convolutional neural network architecture.

[0032] Figure 15 The example U-Net architecture is illustrated.

[0033] Figure 16 This describes an example computer system (or computing device or computer). Detailed Implementation

[0034] Optical coherence tomography (OCT) is an imaging technique that uses low-coherence light to capture micrometer-resolution 2D and 3D images from optically scattering media, such as biological tissue. OCT is a non-invasive interferometric imaging modality that can image the retina in vivo in cross-section. OCT provides images of ocular structures and has been used to quantitatively assess retinal thickness and evaluate qualitative anatomical changes, such as the presence of pathological features, including intraretinal and subretinal fluid. A more detailed discussion of OCT is provided below.

[0035] Advances in OCT technology have led to the creation of other OCT-based imaging modalities. OCT angiography (OCTA) is an imaging modality that has rapidly gained clinical acceptance. OCTA images are based on the variable backscattering of light from the blood vessels and neurosensory tissues of the retina. Because the intensity and phase of the backscattered light from retinal tissues vary according to the intrinsic motion of the tissues (e.g., red blood cell movement, while neurosensory tissues are typically static), OCTA images are inherently motion-contrast images. This motion-contrast imaging provides high-resolution, non-invasive images of the retinal vascular system in an efficient manner.

[0036] OCTA images can be generated by applying one of many known OCTA processing algorithms to OCT scan data, typically collected at the same or substantially the same lateral location on the sample at different times, to identify and / or visualize regions or flows of motion. Therefore, a typical OCT angiography dataset may contain multiple OCT scans repeated at the same lateral location. Motion contrast algorithms can be applied to intensity information derived from image data (intensity-based algorithms), phase information from image data (phase-based algorithms), or complex image data (complexity-based algorithms). Motion contrast data can be collected as volumetric data (e.g., cubic data) and displayed in various ways. For example, a frontal vascular system image is a frontal planar image displaying motion contrast signals, where the data dimension corresponding to depth (e.g., the "depth dimension" or the system's imaging z-axis of the sample) is displayed as a single representative value, typically by summing or integrating over all or isolated portions of the volumetric data (e.g., a plate defined by two specific layers).

[0037] Due to the high scattering properties of blood within overlying vessels, OCTA is prone to decorrelation tail artifacts, which can interfere with the interpretation of retinal angiography results. In other words, deeper layers may exhibit projection artifacts because the blood flowing in the retinal vessels above them projects undulating shadows, potentially causing variations in the reflected signal. These signal variations may manifest as decorrelation that is difficult to distinguish from the actual flow.

[0038] One step in the standard OCT angiography algorithm involves generating 2D angiographic images of the vascular system (angiography) of different tissue regions or plates from the acquired flow-contrast image along (and across or perpendicular to) the depth dimension. This helps the user visualize vascular system information from different retinal layers. Plate images (e.g., frontal images) can be generated by summation, integration, or other techniques to select a single representative value of the cube motion contrast data along a specific axis between two selected layers (see, for example, U.S. Patent No. 7,301,644, the contents of which are incorporated herein by reference). Plates most affected by decorrelation tail artifacts may include, for example, the deeper retinal layer (DRL), the avascular retinal layer (ARL), the choroidal capillary layer (CC), and any custom plates, especially those containing the retinal pigment epithelium (RPE).

[0039] Figure 1An exemplary OCTA B scan 11 of the human retina is shown, where the upper hash line 13 and the lower hash line 15 indicate the locations defining two transverse frontal images, respectively. The upper hash line 13 represents the superficial retinal layer (SRL) 17 located near the top of the retina, and the lower hash line 15 represents the deeper retinal layer (DRL) 19. In this example, the deeper retinal layer 19 is the target plate that one might want to examine, but because it is located below the superficial retinal layer 17 and the vascular system pattern 17a is above it, the frontal SRL layer 17 might appear as a flow projection (e.g., decorrelational tail or shadow) 19a in the target, deeper frontal DRL layer 19, which could be misidentified as the real vascular system. For better visualization and interpretation, it would be beneficial to correct (e.g., remove or reduce) the flow projection (e.g., decorrelational) artifacts 19a in the target plate 19.

[0040] Stream projection artifacts are typically corrected using plate-based or volume-based methods. Plate-based methods correct one target frontal panel at a time (topographic projection of an OCTA sub-volume defined within two selected surfaces / layers of the OCTA volume). Plate-based methods may require the use of two (frontal) plate images (e.g., an upper plate image and a lower plate image). That is, plate-based methods may require information from an additional reference plate defined at a higher depth location (e.g., above the target frontal panel) to identify and correct shadows in the deeper / lower target frontal panel. For example, as... Figure 2 As shown, the plate-based method can assume that a deeper target frontal plate (e.g., DRL image 19) is (e.g., generated by the following method) the result of mixing an upper reference plate (e.g., SRL 17) and a theoretical artifact-free plate 21a (the unknown, decorrelation-free image to be reconstructed). Artifacts can then be removed using the selection of a mixing model 23, which can be additive or multiplicative in nature. For example, the mixing model 23 can be applied iteratively until a decorrelation-free image 21b is generated. It should be understood that in each iteration, the currently (e.g., temporarily) generated image 21b can replace the theoretical plate 21a in the mixing model 23 until a final generated image 21b with sufficient decorrelation tail correction is achieved.

[0041] Plate-based methods for removing shadow artifacts have proven effective, but they have several limitations. First, both the target plate and the upper reference plate to be corrected are defined by two pairs of corresponding surfaces / layers, typically defined by an automatic layer segmentation algorithm. Errors in layer segmentation and / or unknowns in the relationship between the target and reference plates can lead to the removal of important information from the correction plate. For example, real blood vessels partially present in both the target and upper reference plates may be incorrectly removed from the correction plate. Conversely, plate-based methods may fail to remove some severe artifacts, such as those caused by blood vessels not present in the reference plate due to errors in their definition.

[0042] The effectiveness of plate-based methods can depend on plate constraints (e.g., how the plate is constrained / generated). For example, plate-based methods may work satisfactorily for plates generated using the maximum projection method, but this may not be the case when plates are generated using the summation projection method. For instance, in the case of thick plate constraints, projection artifacts may overwhelm the true sample signal because the projection artifacts propagate into deeper parts of the plate (e.g., the volume). This can cause the true signal within the plate to be masked and unable to be displayed even after the artifacts are corrected.

[0043] Two additional limitations are a direct consequence of the nature of plate-based methods. As mentioned above, in plate-based methods, only a single target plate can be calibrated at a time. Therefore, whenever the target plate constraints change, a plate-based algorithm needs to be executed, regardless of how small the change is or whether the constraints revert to those of the previous step. This translates to increased processing time and memory requirements, as the user modifies the surface / layer of the target plate to visualize the selected vessels of interest. Furthermore, plate-based calibration can only be viewed or processed within the plate plane (e.g., in a frontal plane view or a frontal plane view perpendicular to the imaging z-axis of the OCT system). Therefore, B-scans (or cross-sectional images of the cut-in volume) cannot be viewed, and volumetric analysis of the results is not possible.

[0044] Volume-based methods can mitigate some of these limitations, but conventional volume-based methods also have their own constraints. Some conventional volume-based methods are based on ideas similar to plate-based methods, but implemented iteratively across multiple target plates spanning the entire volume. For example, to correct for the entire volume, a moving deformable window (e.g., a moving target plate) can be moved axially throughout the depth of the OCTA cube, and a plate-based method can be applied at each window location. Another volume-based method is based on the analysis of the peak values ​​of the flow OCTA signal at different depths in each A-scan. In any case, conventional volume-based methods are very time-consuming because the analysis is done iteratively or through peak search, and it is not easy to implement them in parallel on a parallel computer processing system. Furthermore, conventional volume-based methods rely on handcrafted assumptions, which, while producing generally satisfactory results, may not be applicable to all types of vascular manifestations. For example, volume-based methods based on moving windows must overcome the challenge of accurately determining the end location of the vessel and the start location of the (decorrelated) tail. Although complex assumptions about the vessels have been proposed for better correction, artifacts can still be observed at the edges of large vessels. Peak-based analysis methods rely on optical bench measurements, which may not accurately reproduce the retinal characteristics of all subjects and tend to make binary decisions when removing (decorrelated) tails in each A scan, potentially removing true flow data deep within the retina.

[0045] In contrast to the aforementioned handcrafted solutions for correcting flow projection artifacts in flow plates or volumes during angiography, the current preferred embodiment employs a neural network solution trained to use structural data (e.g., OCT structural data) and flow data (e.g., OCTA flow contrast data) as training inputs and learn specific characteristics of projection (flow) artifacts compared to real (correct) blood vessels. This approach has proven superior to handcrafted volume-based methods. For example, this neural network model can process large amounts of data much faster than handcrafted algorithms that use iterative methods or correct flow projections in volumes by finding peaks in each A scan. The faster processing time of this method benefits at least in part from the fact that this model is easier to parallelize in general-purpose graphics processing units (GPGPUs) optimized for parallel operation, but other computer processing architectures may also benefit from this model. Furthermore, fewer assumptions are made when processing the data in this method. By assuming an appropriate gold standard as the target (e.g., the target training output), this neural network can learn the characteristics of flow artifacts and how to reduce them using structural and flow data, without making handcrafted assumptions that may vary throughout the data and may be difficult to estimate heuristically. Furthermore, it is proposed that imperfectly corrected data can also serve as the gold standard for training this neural network, provided it is reasonably correct. This method can also improve the output, depending on the network architecture used and the amount of training data available, because this neural network learns the combinatorial structure representing artifacts and the overall behavior of the streaming data. For example, if the training output set corrects additional artifact errors besides streaming artifacts (such as noise), then the trained neural network can also correct these additional artifact errors.

[0046] Currently preferred neural networks are primarily trained to correct projection artifacts in OCTA volumes, but are trained using training input data pairs consisting of OCT structural data and corresponding OCTA flow data from the same samples / regions. That is, this method uses both structural and flow information to correct artifacts and can be independent of segmentation lines (e.g., layer boundaries) and plate boundaries. The trained neural network can receive test OCTA volumes (e.g., newly acquired OCTA data not previously used for neural network training) and generate a corrected flow (OCTA) volume, which can be used to visualize or process corrected flow data in different planes and three dimensions. For example, the corrected OCTA volume can be used to generate A-scan, B-scan, and / or frontal images of any region of the corrected OCTA volume.

[0047] Figure 3An exemplary training input / output set is shown, comprising a training input (image) set 10 and a corresponding training output target (image) 12. As discussed above and explained more fully below, generating OCTA images (or scans or datasets) 14 typically requires multiple OCT scans (or image data) 16 of the same retinal region, with differences meeting predefined criteria designated as blood flow. In the present case, depth data 18 (e.g., axial depth information, which may be correlated with depth information from the corresponding OCT data 16) is added to the generated OCTA data 14. The generated OCTA data 16 (and optionally individual OCT images 16) are corrected for flow artifacts and / or other artifacts, for example, by using one or more handcrafted algorithms to generate the corresponding training output target OCTA image 20. Optionally, corresponding depth information 22 may also be appended to the target output OCTA image 20.

[0048] Figure 4 This invention illustrates the limitation of the training input / output set (e.g., ...) for neural networks. Figure 3 The method / system is illustrated in block B1. Multiple OCT acquisitions are collected from substantially the same region of the sample. As described in block B2, the collected OCT acquisitions can be used to define (e.g., for the eye) OCT (structural) images 16. The OCT (structural) image data can depict tissue structure information, such as retinal tissue layers, optic nerve, fovea, intraretinal and subretinal fluid, macular hole, macular folds, etc. These OCT images can include one or more averaged images from two or more collected OCT data sets, and can also correct for noise, structural shading, opacity, and other image artifacts. Block B3 uses OCT angiography (OCTA) processing techniques to calculate motion contrast information in the OCT data collected from block B1 (and / or the OCT images 16 defined in block B2, or a combination of both) to define OCTA (flow) image data. The defined flow images depict vascular system flow information and may contain artifacts such as projection artifacts, decorrelation tails, shading artifacts, and opacity. Optionally, as shown by symbol 26, depth index information can be assigned (or appended) to the streaming image along its axis. This depth information can be associated with a defined OCT image used to define the data, as shown by dashed arrow 24 and block B4. The defined OCT image from block B3 (optionally with or without the appended depth information) is submitted to the artifact removal algorithm (block B5) to define a corresponding target output OCT image for artifact reduction (e.g., Figure 3 The training output target OCTA image 20), as shown in block B6. In block B7, the OCT (structure) image, the defined OCTA (flow) image, and the target output OCTA image (optionally also including depth index information) can be grouped to define the training input / output set, such as... Figure 3As shown. Therefore, each training input set 10 includes one or more training input OCT images 16, a corresponding training input OCTA image 14, and depth information 18 of the axial positions of pixels within the training input OCTA image 14. As described above, the target output OCTA image 20 may also optionally have corresponding depth information 22 (e.g., corresponding to depth data 18). It is understood that multiple training input / output sets can be defined by defining multiple OCTA images 14 from the set of corresponding OCT acquisitions 16 to define multiple corresponding training input OCTA images 20.

[0049] Therefore, the neural network according to the invention can be trained using a set of OCTA acquisitions with corrected streaming data and a corresponding set of OCT acquisitions (from which OCTA data can be determined) and can also correct for shadows or other artifacts. The corrected streaming data may be known or pre-computed for training purposes, but it is not necessary to provide labels identifying the corrected regions, whether in the training input set or in the output training image. Both the (OCT) structure and the (OCTA) streaming cube are used as training inputs, and the neural network is trained to produce a corrected output (OCTA) streaming cube to generate projection artifacts. In this way, pre-generated correction data (e.g., training output, target image) is used as guidance for training the neural network.

[0050] In neural network training, the corrected OCTA stream data used as the training output target can be obtained using hand-crafted algorithms, with or without additional hand-crafted corrections. It does not need to constitute a perfect solution for artifact correction, although its performance should be satisfactory in most (mostly) A-scans of the volume samples. That is, hand-crafted solutions based on single A-scan stream artifact correction, plate-based correction, or volume-based correction (e.g., as described above) can be used to define the training output target volume (e.g., an image) corresponding to each training input set (including the training OCTA volume and one or more corresponding OCT structure volumes). Optionally, the training output target volume can be divided into training output sub-volume sets. For example, if the corrected training volume still has regions with severe flow artifacts, it can be divided into sub-volumes, and only a satisfactory portion of the corrected volume (excluding the portion with severe flow artifacts) can be used to define the training input set. Furthermore, the corrected OCTA volume and its corresponding OCT sample set, along with the uncorrected OCTA volume, can be divided into corresponding sub-volume segments to define a larger number of training input / output sets, each defined by a sub-volume region.

[0051] In operation (e.g., after training the neural network), the collected structural OCT images, the corresponding OCTA stream images, and the assigned / determined / computed depth index information are submitted to the trained neural network, which then outputs / generates an OCT-based image of blood vessels (e.g., an OCTA image) with reduced artifacts compared to the input OCTA stream images.

[0052] Various types of neural networks can be used according to the present invention, but a preferred embodiment of the invention uses a U-Net-type neural network. A general discussion of the U-Net neural network is provided below. However, preferred embodiments may deviate from this general U-Net and be based on a U-Net architecture optimized for speed and accuracy. As an example, the U-Net neural network architecture used in the proof-of-concept implementation of the present invention is provided below.

[0053] As a proof of concept, a swept-frequency OCT device (PLEX) was used. 9000, Carl Zeiss Meditec, Inc TM OCTA acquisitions (and their corresponding OCT data) with a 6×6×3mm field of view were obtained from 262 eyes. Of these eyes, 153 were healthy and 109 were diseased. Of the 262 eyes, 211 eyes (including 123 from healthy eyes and 88 from diseased eyes) were used for training (e.g., to prepare the training input / output set, including OCTA / OCT training input pairs and their corresponding corrected output training targets), and 51 eyes (including 30 from healthy eyes and 21 from diseased eyes) were used for validation (e.g., as test inputs to validate the effectiveness of the trained neural network during the testing phase). For each OCTA acquisition, a (e.g., volume-based) manual decorrelation tail removal algorithm was used to generate the corresponding training output target corrected form of the stream volume. Similarly, a (manual) algorithm was also used to correct artifacts in its corresponding OCT volume data.

[0054] During the training phase, two training methods were investigated. In both methods, the neural network takes the stream (OCTA) data to be corrected and the structure (OCT) data from each OCTA acquisition as input. Similarly, in both methods, the output of the neural network is measured against (or compared to) a fundamental fact (e.g., the corresponding training output target) (e.g., ideal corrected stream data). The training output target is obtained by submitting the training input OCTA acquisitions to a hand-crafted stream artifact correction algorithm. An example of a hand-crafted volume-based projection removal algorithm is described in U.S. Patent No. 10,441,164, assigned to the same assignee as this application. However, the two methods differ in how the training target is defined. For ease of discussion, the input stream data to be corrected can be referred to as the “original stream,” while the desired, corrected stream data that the neural network expects to produce can be referred to as the “corrected stream.” In the first method, given the “original stream” as input, the neural network is trained to predict the “corrected stream” (e.g., closely replicate the training output target). The first training method is similar to the method discussed below. The second method differs in that its goal is to define the difference between the “original stream” and the “corrected stream.” In other words, during each training iteration (e.g., epoch), the neural network is trained to predict a "residual" based on the difference between the "corrected stream" and the "original stream," and this residual is added back to the original stream. The final residual generated by the neural network is then added to the original input stream scan to define the corrected form of the input stream scan. In some cases, the second method is found to provide better results than the first method. This may be because the first method requires the neural network to learn to reproduce the original stream image largely invariantly (e.g., the target output stream image may be very similar to the input stream image), while the second method only needs to generate residual data (e.g., providing signal data corresponding to the locations of changes / differences between the training input and the target output).

[0055] This neural network is based on the general U-Net neural network architecture, as shown in the reference below. Figure 16 As stated above, but with some changes. Figure 5 A simplified overview of the U-Net architecture used in exemplary embodiments of the present invention is provided. Figure 16 The first change is a reduction in the total number of layers in the current machine learning model. This embodiment has two downsampling blocks (e.g., encoding modules) 31a / 31b in its contraction path and two corresponding upsampling blocks (e.g., decoding modules) 33a / 33b in its expansion path. This is different from... Figure 16The example U-Net in the example provides a contrast, having four downsampling blocks and four upsampling blocks. This reduction in downsampling and upsampling blocks improves performance in terms of speed while still producing satisfactory results. However, it should be understood that a suitable U-Net can have more or fewer downsampling blocks and corresponding upsampling blocks without departing from the invention. Additional downsampling / upsampling blocks may produce better results at the cost of longer training (and / or execution) time. In this example, each downsampling block 31a / 31b and upsampling block 33a / 33b consists of three layers 39a, 39b, and 39c, each representing image data (e.g., volumetric data) for a given processing stage, but it should be understood that downsampling and upsampling blocks can have more or fewer layers. Although not illustrated for simplicity, it should also be understood that this U-Net can have copy and crop links between the corresponding downsampling and upsampling blocks (e.g., similar to...). Figure 16 (Links CC1 to CC4 in the diagram). These copy and trim links can copy the output of a downsampled block and cascade the output to the input of the corresponding upsampled block.

[0056] The different operations of this U-Net are illustrated / indicated by arrow key diagrams. Two sets of operations are applied to each downsampling block 31a / 31b. The first set, indicated by arrow 35, is related to... Figure 16 Similarly, this includes (e.g., 3×3) convolutions and activation functions (e.g., rectified linear (ReLU) units), and batch normalization. However, the second group is indicated by P-arrow 37, and... Figure 16 Unlike other methods, it adds column pooling.

[0057] Figure 6A more detailed view of the exemplary operation (or operation step) indicated by P-arrow 37 in the downsampling block is shown. This second set of operations applies vertical (or column-maximum) pooling 51 to layer 39b, whose height and width data dimensions are represented in H×W. Column-maximum pooling 51 defines 1×W pooled data 41, which is then upsampled to define upsampled data 43 matching the dimension size H×W of layer 39b. In the concatenation step 45, the upsampled data 43 is concatenated to the image data from layer 39b and then submitted to the convolution step 47 and an activation function with a batch normalization step 49 to produce a local output layer 39c of a single block. The addition of the vertical pooling layer 51 enables the machine learning model to quickly move information between different parts of the image (e.g., vertically moving data between different layers in an OCT / OCTA volume). For example, a blood vessel at the first location (x, z) might produce a tail artifact at a second vertically offset (e.g., deeper) location (x, z+100) without causing any visible changes (any tail artifact) in any intermediate region (e.g., at the third location (x, z+50)). Therefore, without a "shortcut" connecting these two points (e.g., the first and second locations), the network would have to independently learn several convolutional filters to propagate information down a total of 100 pixels from the first location to the second location.

[0058] As described above, each pixel (or voxel) in the volumetric (or plate, or frontal) image data includes additional information channels specifying its depth index or location (e.g., z-coordinate) within the volume. This allows the neural network to learn / develop context-different computations at least in part based on the depth index information at different axial (e.g., depth) locations. Furthermore, training input samples may include defined retinal landmarks (e.g., structural features determined from structural OCT data), and context-different computations may also depend on local retinal landmarks, such as retinal layers.

[0059] return Figure 5 The output from a downsampling block 31a is max-pooled (e.g., 2×2 max-pooling), as indicated by the down arrow, and fed into the next downsampling block 31b in the contraction path, until it reaches the optional "bottleneck" block / module 53 and enters the expansion path. Optionally, the max-pooling function indicated by the down arrow can be integrated into the downsampling block preceding it, as it provides downsampling functionality. Bottleneck 53 may include two convolutional layers (with batch normalization and optional dropout), as referenced. Figure 16 As shown, this implementation adds column pooling, as indicated by arrow P. This increases the amount of column pooling the network can perform, and it has been found to improve test performance.

[0060] In the expansion path, the output of each block is submitted to the transposed convolution (or deconvolution) stage to upsample the image / information / data. In this example, the transposed convolution is characterized by a 2×2 kernel (or convolution matrix) with a stride (e.g., kernel shift) of 2 (e.g., two pixels or voxels). At the end of the expansion path, the output of the final upsampled block 33a is submitted to another convolution operation (e.g., a 1×1 convolution) before producing its output 57, as indicated by the dashed arrow. Before reaching the 1×1 convolution, each pixel of the neural network may have multiple features, but the 1×1 convolution combines these multiple features into a single output value for each pixel at a pixel-by-pixel level.

[0061] Figure 16 and Figure 5 Another difference between U-Net and others is the addition of a dynamic pooling layer 32 (e.g., based on retinal structure) after input layer 34 and before downsampling blocks 31a / 31b. As described above, additional information channels (e.g., similar to additional color channels) are cascaded to the input data before being fed into this network, and the value of each pixel in the input data is the z-coordinate (depth) of that pixel / voxel within the volume. This allows the network to perform context-different computations at different depths while still preserving the fully convolutional structure. That is, input layer 34 receives input OCT-based data 36 (e.g., OCT structural data and OCTA streaming data, including depth index information) and dynamic pooling layer 32 compresses the input OCT-based data (image information) to a range beyond the variable depth defined by the location of (e.g., pre-selected) retinal landmarks in the received OCT-based data. Retinal landmarks can be (e.g., specific) retinal layers or other known structures. For example, such as Figure 9 and Figure 11 As shown, relevant retinal tissue information may be limited to a specific axial range where the retinal layer of interest is located, and the depth location of these layers may vary with volumetric data. Therefore, the dynamic pooling layer 32 allows this machine learning model to reduce the amount of data it processes to include only the portion of the volume containing the layers of interest, such as layers that may have or be involved in the generation of flow artifacts or specific layers that a human reviewer might be interested in. As an example, the dynamic pooling layer 32 can quickly identify the internal limiting membrane (ILM) and retinal pigment epithelium (RPE) because they are high-contrast regions along the A-scan and typically identify the top and bottom layers of the retina. See also Figure 9The different retinal layers and boundaries in the normal human eye are briefly described. Other retinal layers can also be identified and associated with their specific depth information. This facilitates the application of context-different computations by the data processing layers following the dynamic pooling layer 32 at different axial locations, at least in part based on depth index information and / or local retinal landmarks (e.g., retinal structures, such as retinal layers). Thus, the dynamic pooling layer 32 compresses image information beyond a variable depth range defined by the input data itself (e.g., defined by the location of specific retinal landmarks within the input OCT-based data 36).

[0062] and Figure 16 Similar to U-Net, during the training phase, the output 57 of this U-Net is compared with the target output OCTA image 59 by applying a loss function 61 (e.g., L1 loss function, L2 loss function, etc.), and the internal weights of the data processing layers (e.g., downsampling blocks 31a / 31b and upsampling blocks 33a / 33b) are adjusted accordingly (e.g., through backpropagation) to reduce the error in subsequent backpropagation iterations. Optionally, this neural network can apply loss functions with different weights based on specific retinal layers. That is, the loss function can have different weights based on the local proximity of a pre-selected retinal landmark (e.g., retinal layer) to the current axial position of the OCT image data being processed. For example, this embodiment can use a reweighted L1 loss function such that the region of the input OCT (or OCT) volume between the internal limiting membrane (ILM) and the retinal pigment epithelium (RPE) is weighted at least an order of magnitude greater than other regions of the volume (e.g., 11 times the weight).

[0063] Figure 7 An exemplary method for reducing artifacts in OCT-based images of the eye according to the present invention is described. The method may begin with step S1, whereby OCT image data of the eye is collected from an OCT system, wherein the collected OCT image data includes depth index information. In step S2, the OCT image data is fed into a trained neural network, wherein the neural network may have a convolutional structure (e.g., U-Net) and be trained to apply context-different computations at different axial positions, at least in part, based on the depth index information. For example, the different computations may depend in context on predefined local retinal landmarks, such as (optionally predefined) retinal layers. In step S3, the trained neural network generates an OCT-based image with reduced artifacts compared to the collected OCT image data.

[0064] Optionally, the collected OCT images may undergo several data conditioning sub-steps. For example, in sub-step Sub1, structural (OCT) data of the eye is created from the collected OCT image data, wherein the created structural image depicts information about the eye's tissue structures, such as the retinal layer. Similarly, in sub-step Sub2, motion contrast information (e.g., from the collected OCT image data and / or the initial structural data) is calculated using OCTA processing techniques. In sub-step Sub3, a flow (OCTA) image can be created from the motion contrast information, wherein the flow image depicts vascular system flow information and includes artifacts such as projection artifacts, decorrelation tails, shading artifacts, and opacity artifacts. In sub-step Sub4, depth index information is assigned along its axis to the created flow image. For example, the created flow image may be expanded to include additional information channels (e.g., additional color channels per pixel) that incorporate depth index information (e.g., instead of additional color information).

[0065] The trained neural network may have several salient features. For example, the neural network may include a dynamic pooling layer after the input layer to compress image information into the received OCT image data beyond a variable depth range defined by the (optionally pre-selected) retinal landmark (e.g., axial / depth) position of the retinal layer. The neural network may also have multiple data processing layers after the dynamic pooling layer, wherein the multiple data processing layers perform context-different computations at different axial positions based at least in part on depth index information and / or the (e.g., axial) position of the retinal landmark (such as (optionally specific) retinal layers). During training, the neural network may include an output layer that compares the outputs of the multiple data processing layers with a target output OCT image and adjusts the internal weights of the data processing layers through a backpropagation process. During training, the neural network may apply a loss function (e.g., an L1 function) with different weights based on the local proximity of the (optionally pre-selected) retinal landmark (e.g., retinal layer) to the current axial position of the OCT image data being processed. Optionally, the loss function may have different weights based on a specific retinal layer. For example, the loss function could have a first weight for the region between the internal limiting membrane (ILM) and the retinal pigment epithelium (RPE), and a second weight for other regions. Optionally, the first weight could be an order of magnitude larger than the second weight.

[0066] The following provides a description of various hardware and architectures applicable to this invention.

[0067] Typically, optical coherence tomography (OCT) uses low-coherence light to produce two-dimensional (2D) and three-dimensional (3D) internal views of biological tissues. OCT enables in vivo imaging of retinal structures. OCT angiography (OCTA) produces blood flow information, such as blood flow from vessels within the retina. Examples of OCT systems are provided in U.S. Patent Nos. 6,741,359 and 9,706,915, and examples of OCTA systems can be found in U.S. Patent Nos. 9,700,206 and 9,7959,544, all of which are incorporated herein by reference in their entirety. Exemplary OCT / OCTA systems are provided herein.

[0068] Figure 8 A general-purpose frequency-domain optical coherence tomography (FD-OCT) system for collecting 3D image data of an eye suitable for use with this invention is described. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, broadband light sources with short time-coherence lengths or swept-frequency laser sources. The beam from the light source LtSrc1 is typically routed via fiber Fbr1 to illuminate a sample, such as an eye E; a typical sample is tissue within the human eye. The light source LrSrc1 can be, for example, a broadband light source with a short time-coherence length in the case of spectral domain OCT (SD-OCT) or a wavelength-tunable laser source in the case of swept-frequency source OCT (SS-OCT). A scanning beam can be used, typically between the output of fiber Fbr1 and the sample E, such that the beam (dashed line Bm) scans laterally across the region of the sample to be imaged. The beam from the scanner Scnr1 can be passed through a scanning lens SL and an ophthalmic lens OL and focused onto the sample E to be imaged. The scanning lens SL can receive the beam from the scanner SnNR1 at multiple incident angles and produce substantially collimated light, which the ophthalmic lens OL can then focus onto the sample. This example illustrates the need for scanning in two lateral directions (e.g., the x and y directions in the Cartesian plane) to scan the desired field of view (FOV). An example of this is a point-field OCT, which uses a point-field beam to scan the sample. Therefore, the scanner SnNR1 is illustratively shown as comprising two sub-scanners: a first sub-scanner XSCN for scanning the point-field beam across the sample in a first direction (e.g., the horizontal x direction); and a second sub-scanner YSCN for scanning the point-field beam across the sample in a second direction (e.g., the vertical y direction). If the scanning beam is a line-field beam (e.g., a line-field OCT), which can sample the entire line portion of the sample at a time, then perhaps only one scanner is needed to scan the line-field beam across the sample to span the desired FOV. If the scanning beam is a full-field beam (e.g., full-field OCT), a scanner may not be needed, and the full-field beam can be applied to the entire desired FOV at once.

[0069] Regardless of the beam used, the light scattered from the sample (e.g., the sample light) is collected. In this example, the scattered light returning from the sample is collected into the same fiber Fbr1 used to guide the light for illumination. The reference light from the same light source LtSrc1 propagates through a separate path, in this case involving fiber Fbr2 and a back reflector RR1 with adjustable optical delay. Those skilled in the art will recognize that a transmission reference path can also be used, and the adjustable delay can be placed in the sample or reference arm of the interferometer. The collected sample light is combined with the reference light, for example in a fiber coupler Cplr1, to form an optical interference in an OCT photodetector Dtctr1 (e.g., a photodetector array, digital camera, etc.). Although a single fiber port is shown leading to detector Dtctr1, those skilled in the art will recognize that various designs of the interferometer can be used for balanced or unbalanced detection of interference signals. The output of detector Dtctr1 is provided to a processor (e.g., an internal or external computing device) Cmp1, which converts the observed interference into depth information of the sample. Depth information can be stored in memory associated with processor Cmp1 and / or displayed on a display (e.g., computer / electronic display / screen) Scn1. Processing and storage functions can be located within the OCT instrument, or the functions can be offloaded (e.g., executed thereon) to an external processor (e.g., an external computing device), to which the collected data can be transferred. Figure 15 An example of a computing device (or computer system) is shown. This unit may be dedicated to data processing or to performing other very general tasks that are not specific to the OCT device. The processor (computing device) Cmp1 may include, for example, a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), a system-on-a-chip (SoC), a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a combination thereof, which may perform some or all of the processing steps in a serial and / or parallel manner with one or more host processors and / or one or more external computing devices.

[0070] The sample and reference arms in an interferometer can be composed of bulk optics, fiber optics, or hybrid bulk optics systems, and can have different architectures, such as Michelson, Mach-Zehnder, or common-path-based designs known to the art. The beams used herein should be interpreted as any carefully oriented optical path. Instead of a mechanical scanning beam, an optical field can illuminate a one-dimensional or two-dimensional region of the retina to generate OCT data (e.g., see U.S. Patent No. 9,332,902; D. Hillmann et al., “Holoscopy–Holographic Optical Coherence Tomography,” Optics Letters, 36(13):2390 2011; Y. Nakamura et al., “High-Speed ​​Three Dimensional Human Retinal Imaging by Line Field Spectral Domain Optical Coherence Tomography,” Optics Express, 15(12):7103 2007; Blazkiewicz et al., “Signal-To-Noise Ratio Study of Full-Field Fourier-Domain Optical Coherence Tomography,” Applied Optics, 44(36):7722 (2005)). In time-domain systems, the reference arm needs to have an adjustable optical delay to generate interference. Balanced detection systems are typically used in TD-OCT and SS-OCT systems, while spectrometers are used at the detection port of SD-OCT systems. The invention described herein can be applied to any type of OCT system. Various aspects of this invention can be applied to any type of OCT system or other types of ophthalmic diagnostic systems and / or multiple ophthalmic diagnostic systems, including but not limited to fundus imaging systems, field-of-view testing devices, and scanning laser polarimeters.

[0071] In Fourier domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectral interferogram (Sj(k)). The real-valued spectral data typically undergoes several post-processing steps, including background subtraction and dispersion correction. The Fourier transform of the processed interferogram results in a complex-valued OCT signal output Aj(z) = |Aj|eiφ. The absolute value of this complex-valued OCT signal, |Aj|, ​​reveals the scattering intensity distribution across different path lengths, thus scattering is a function of depth (z-direction) in the sample. Similarly, the phase φj can be extracted from the complex-valued OCT signal. This scattering distribution as a function of depth is called the axial scan (A-scan). A set of A-scans measured at adjacent locations in the sample produces a cross-sectional image of the sample (tomogram or B-scan). The collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a given amount of data, the term fast axis refers to the scan direction along a single B-scan, while slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" refers to a single data unit or block generated by repeated acquisitions at the same (or substantially the same) location (or region) for analyzing motion contrast, which can be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans acquired at relatively short time intervals at substantially the same location on the sample. Because the scans in a cluster scan belong to the same region, the static structure remains relatively unchanged from scan to scan within the cluster scan, and the motion contrast between scans that meet predefined criteria can be identified as blood flow.

[0072] Various methods for generating B-scans are known in the art, including, but not limited to: along the horizontal or x-direction, along the vertical or y-direction, along the diagonals of x and y, or in a circular or spiral pattern. A B-scan can be xz-dimensional, but can be any cross-sectional image including the z-dimensional dimension. Figure 9 This image shows a sample OCT B-scan image of a normal human retina. An OCT B-scan of the retina provides a view of the retinal tissue structure. For illustrative purposes, Figure 9 Various canonical retinal layers and their boundaries are identified. The identified retinal boundary layers include (from top to bottom): Internal Limiting Membrane (ILM) layer 1, Retinal Nerve Fiber Layer (RNFL or NFL) layer 2, Ganglion Cell Layer (GCL) layer 3, Internal Plumular Layer (IPL) layer 4, Inner Nuclear Layer (INL) layer 5, Outer Opioid Layer (OPL) layer 6, Outer Nuclear Layer (ONL) layer 7, Knot between the Outer Segment (OS) and Inner Segment (IS) of the light receptor (indicated by reference symbol layer 8), External or External Membrane (ELM or OLM) layer 9, Retinal Pigment Epithelium (RPE) layer 10, and Bruch's Membrane (BM) layer 11.

[0073] In OCT angiography or functional OCT, analytical algorithms can be applied to OCT data collected at the same or substantially the same sample location on the sample at different times (e.g., cluster scans) to analyze motion or flow (see, for example, U.S. Patent Publications 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579 and U.S. Patent No. 6,549,801, all of which are incorporated herein by reference in their entirety). OCT systems can use any of a variety of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, motion contrast algorithms can be applied to intensity information derived from image data (intensity-based algorithms), phase information from image data (phase-based algorithms), or complex image data (complex-based algorithms). A frontal image is a 2D projection of 3D OCT data (e.g., by averaging the intensity of each individual A-scan so that each A-scan defines pixels in the 2D projection). Similarly, a frontal vascular system image is an image displaying motion contrast signals, where the data dimension corresponding to depth (e.g., along the z-direction of an A-scan) is displayed as a single representative value (e.g., a pixel in a 2D projected image), typically by summing or integrating all or isolated portions of the data (see, for example, U.S. Patent No. 7,301,644, the entire contents of which are incorporated herein by reference). An OCT system providing angiographic imaging capabilities may be referred to as an OCT angiography (OCTA) system.

[0074] Figure 10 An example of a frontal vascular system image is shown. After processing the data using any motion contrast technique known in the art to enhance motion contrast, pixel ranges corresponding to a given tissue depth from the surface of the internal limiting membrane (ILM) of the retina can be summed to generate a frontal (e.g., frontal view) image of the vascular system. Figure 11 An exemplary B-scan of vascular system (OCT) image is shown. As illustrated, structural information may not be clearly defined because blood flow may pass through multiple retinal layers, making them less defined than in a structural OCT B-scan. Figure 9As shown. Nevertheless, OCTA provides a non-invasive technique for microvascular imaging of the retina and choroid, which can be crucial for the diagnosis and / or monitoring of various pathologies. For example, OCTA can be used to identify diabetic retinopathy by recognizing microaneurysms, neovascular complexes, and quantifying avascular and non-perfused areas of the fovea. Furthermore, OCTA has shown excellent concordance with fluorescein angiography (FA), a more routine but more covert technique that requires dye injection to observe vascular flow in the retina. Additionally, in dry age-related macular degeneration, OCTA has been used to monitor the pervasive reduction in choroidal capillary flow. Similarly, in wet age-related macular degeneration, OCTA can provide qualitative and quantitative analysis of the choroidal neovascular membrane. OCTA has also been used to investigate vascular occlusion, such as assessing unperfused areas and the integrity of superficial and deep neural plexuses.

[0075] Neural Networks

[0076] As described above, this invention can utilize neural network (NN) machine learning (ML) models. For completeness, a general discussion of neural networks is provided herein. This invention can use any of the following neural network architectures, individually or in combination. A neural network or neural network is a network of interconnected neurons (nodes), where each neuron represents a node in the network. Groups of neurons can be arranged hierarchically, with the output of one layer fed forward to the next layer in a multilayer perceptron (MLP) arrangement. An MLP can be understood as a feedforward neural network model that maps an input dataset to an output dataset.

[0077] Figure 12 An example of a multilayer perceptron (MLP) neural network is illustrated. Its structure may include multiple hidden (e.g., inner) layers HL1 to HLn, which map an input layer InL (receiving a set of inputs (or vector inputs) in_1 to in_3) to an output layer OutL, which produces a set of outputs (or vector outputs), such as out_1 and out_2. Each layer can have any given number of nodes, which are exemplarily shown as circles within each layer in this document. In this example, the first hidden layer HL1 has two nodes, while hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the more hidden layers in the MLP), the greater its learning capacity. The input layer InL receives vector inputs (illustrated as a three-dimensional vector consisting of in_1, in_2, and in_3) and can apply the received vector inputs to the first hidden layer HL1 in the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer (e.g., HLn) in the multilayer model, processes its inputs, and produces a vector output result (exemplarily shown as a two-dimensional vector consisting of out_1 and out_2).

[0078] Typically, each neuron (or node) produces a single output, which is fed forward to neurons in the immediately preceding layer. However, each neuron in a hidden layer can receive multiple inputs, either from the input layer or from the output of a neuron in the immediately preceding hidden layer. Generally, each node can apply a function to its inputs to generate an output for that node. Nodes in hidden layers (such as learning layers) can apply the same function to their respective inputs to produce their respective outputs. However, some nodes, such as those in the input layer InL, receive only one input and may be passive, meaning they simply relay the value of their single input to their output; for example, they provide a copy of their input to their output, as indicated by the dotted arrow within the node in the input layer InL.

[0079] For ease of explanation, Figure 13 A simplified neural network consisting of an input layer InL', a hidden layer HL1', and an output layer OutL' is shown. The input layer InL' is shown as having two input nodes i1 and i2, which receive inputs Input_1 and Input_2 respectively (e.g., the input nodes of layer InL' receive a two-dimensional input vector). The input layer InL' feeds forward to the hidden layer HL1', which has two nodes h1 and h2, and this hidden layer in turn feeds forward to the output layer OutL', which has two nodes o1 and o2. The interconnections or links between neurons (shown as solid arrows) have weights w1 to w8. Typically, in addition to the input layer, a node (neuron) can receive the output of its immediately preceding node as input. Each node can compute its output by multiplying each of its inputs by the corresponding interconnection weights of each input, summing the products of its inputs, adding (or multiplying by) a constant defined by another weight or bias that may be associated with that particular node (e.g., node weights w9, w10, w11, w12 correspond to nodes h1, h2, o1, o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function can be called an activation function or a transfer function. Several activation functions are known in the art, and the choice of a particular activation function is not critical to this discussion. However, it is important to note that the operation of an ML model or the behavior of a neural network depends on the weight values, which can be learned so that the neural network provides the desired output for a given input.

[0080] During the training or learning phase, the neural network learns (e.g., is trained to determine) appropriate weight values ​​to achieve the desired output for a given input. Before training the neural network, an initial (e.g., random and optionally non-zero) value, such as a random number seed, can be assigned individually to each weight. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) so that, for a given training vector input, the neural network produces an output close to the desired (predetermined) training vector output. For example, the weights can be progressively adjusted over thousands of iterations using a technique called backpropagation. In each iteration of backpropagation, the training input (e.g., a vector input or training input image / sample) is fed forward through the neural network to determine its actual output (e.g., a vector output). The error of each output neuron or output node is then calculated based on the actual neuron output and the target training output of that neuron (e.g., the training output image / sample corresponding to this training input image / sample). The output is then backpropagated through the neural network (in the direction from the output layer back to the input layer), updating the weights based on the degree of influence of each weight on the overall error, thereby bringing the output of the neural network closer to the desired training output. This cycle is then repeated until the actual output of the neural network is within an acceptable error range of the desired training output for a given training input. It's understandable that each training input may require multiple backpropagation iterations to reach the desired error range. Typically, an epoch refers to one backpropagation iteration across all training samples (e.g., one forward propagation and one backpropagation), so training a neural network may require many epochs. Generally, the larger the training set, the better the performance of the trained ML model, so various data augmentation methods can be used to increase the size of the training set. For example, when the training set includes pairs of corresponding training input and training output images, the training images can be divided into multiple corresponding image segments (or patches). Corresponding patches from the training input and training output images can be paired to limit multiple training patch pairs from one input / output image pair, which expands the training set. However, training on a large training set places high demands on computational resources, such as memory and data processing resources. The computational requirements can be reduced by dividing the large training set into multiple mini-batches, where the mini-batch size limits the number of training samples in one forward / backward pass. In this case, one epoch can include multiple mini-batches. Another problem is that the NN may overfit the training set, thus reducing its ability to generalize from a specific input to different inputs. Overfitting can be mitigated by creating an ensemble of neural networks or by randomly dropping nodes from the neural network during training, which effectively removes the dropped nodes from the network. Various dropout mitigation methods, such as inverse dropout, are known in the art.

[0081] Please note that the operations of a trained neural network model are not direct algorithms for operational / analysis steps. In fact, when a trained neural network model receives input, that input is not analyzed in the conventional sense. Instead, regardless of the subject or nature of the input (e.g., a vector constrained to a live image / scan or a vector constrained to some other entity, such as a demographic description or activity record), the input will be subjected to the same pre-defined architecture of the trained neural network (e.g., the same node / layer arrangement, trained weights and biases, pre-defined convolution / deconvolution operations, activation functions, pooling operations, etc.), and it may be unclear how the architecture of the trained network produces its output. Furthermore, the values ​​of trained weights and biases are not deterministic and depend on many factors, such as the amount of time the neural network is used for training (e.g., the number of epochs in training), the random initial values ​​of the weights before training begins, the computer architecture of the machine training the NN, the selection of training samples, the distribution of training samples across multiple mini-batches, the choice of activation function, the choice of error function to correct the weights, and whether training is interrupted on one machine (e.g., with the first computer architecture) but completed on another machine (e.g., with different computer architectures). Crucially, the reasons why a trained ML model achieves certain outputs are not yet clear, and extensive research is underway to determine the factors upon which the outputs of ML models are based. Therefore, the processing of real-time data by neural networks cannot be simplified to simple step-by-step algorithms. Instead, its operation depends on its training architecture, training sample set, training sequence, and various conditions during the training of the ML model.

[0082] In summary, building a neural network (NN) machine learning model can include a learning (or training) phase and a classification (or operation) phase. During the learning phase, the neural network can be trained for a specific purpose, and a training example set can be provided, including training (sample) inputs and training (sample) outputs, and optionally a validation example set to test the progress of training. During this learning process, various weights associated with the nodes and node interconnections in the neural network are incrementally adjusted to reduce the error between the actual output of the neural network and the desired training output. In this way, a multi-layer feedforward neural network (as discussed above) can be made capable of approximating any measurable function to any desired accuracy. The result of the learning phase is a machine learning (ML) model that has been learned (e.g., trained). In the operation phase, a set of test inputs (or real-time inputs) can be submitted to the learned (trained) ML model, which can apply what it has learned to produce output predictions based on the test inputs.

[0083] and Figure 12 and Figure 13Like regular neural networks, convolutional neural networks (CNNs) consist of neurons with learnable weights and biases. Each neuron receives input, performs an operation (e.g., a dot product), and optionally undergoes non-linear processing. However, a CNN might receive raw image pixels at one end (e.g., the input) and provide a classification (or category) score at the other end (e.g., the output). Since CNNs expect images as input, they are optimized for volume (e.g., the pixel height and width of the image, and the image depth, e.g., color depth, such as RGB depth defined by the three colors red, green, and blue). For example, a CNN layer might be optimized for neurons arranged in 3D. Neurons in a CNN layer might also be connected to small regions of the previous layer, rather than all neurons in a fully connected CNN. The final output layer of a CNN can reduce the entire image to a single vector (classification) arranged along the depth dimension.

[0084] Figure 14 An example convolutional neural network architecture is provided. A convolutional neural network can be construed as a sequence of two or more layers (e.g., layer 1 to layer N), where each layer can include a (image) convolution step, a (result) weighted sum step, and a nonlinear function step. Convolution can be performed on the input data by applying filters (or kernels), for example, generating feature maps over a moving window of the input data. Each layer and its components can have different predefined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data is an image with a given pixel height and width, which can be the raw pixel values ​​of the image. In this example, the input image is shown as a depth image with three color channels RGB (red, green, and blue). Optionally, the input image can undergo various preprocessing steps, and the preprocessed results can be used to replace or supplement the original input image. Some examples of image preprocessing might include: retinal angiography segmentation, color space transformation, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product can be computed between a given weight and a small region connecting the weights in the input volume. Many ways to configure a CNN are known in the art, but as an example, layers can be configured to apply element-wise activation functions, such as a maximum (0, x) threshold at zero. Pooling functions (e.g., along the xy direction) can be performed to downsample the volume. Fully connected layers can be used to determine the classification output and produce a one-dimensional output vector, which has been found useful for image recognition and classification. However, for image segmentation, a CNN needs to classify each pixel. Since each CNN layer tends to downsample the input image, another stage is needed to upsample the image back to its original resolution. This can be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically does not use any predefined interpolation methods but instead has learnable parameters.

[0085] Convolutional neural networks have been successfully applied to many computer vision problems. As mentioned above, training CNNs typically requires a large training dataset. The U-Net architecture is based on CNNs and can usually be trained on a smaller training dataset than regular CNNs.

[0086] Figure 15 An example U-Net architecture is illustrated. This exemplary U-Net includes an input module (or input layer or stage) that receives an input U-in of any given size (e.g., an input image or image patch). For illustrative purposes, the image size of any stage or layer is indicated in a box representing the image; for example, the input module contains the number "128×128" to indicate that the input image U-in consists of 128×128 pixels. The input image can be a fundus image, an OCT / OCTA frontal image, a B-scan image, etc. However, it should be understood that the input can be of any size or dimension. For example, the input image can be an RGB color image, a monochrome image, a volumetric image, etc. The input image passes through a series of processing layers, each illustrated with exemplary dimensions, but these dimensions are for illustrative purposes only and will depend on, for example, the image size, the convolutional filters, and / or the pooling stages. This architecture includes a shrinking path (illustrated here as consisting of four encoding modules), followed by an expanding path (illustrated here as consisting of four decoding modules), and copy and pruning links between the corresponding modules / stages (e.g., CC1 to CC4), where the output of an encoding module in the shrinking path is copied and concatenated to the upconverted input of the corresponding decoding module in the expanding path (e.g., appended to its back side). This results in a typical U-shaped feature, hence the name of the architecture. Optionally, for computational reasons, a "bottleneck" module / stage (BN) can be positioned between the shrinking and expanding paths. The bottleneck BN might consist of two convolutional layers (with batch normalization and optional dropout).

[0087] A shrinking path is analogous to an encoder, typically capturing contextual (or feature) information using feature maps. In this example, each encoding module in the shrinking path may include two or more convolutional layers, illustratively indicated by the asterisk “*”, and may be followed by a max-pooling layer (e.g., a downsampling layer). For example, the input image U-in is illustratively shown as undergoing two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel produces a feature map (e.g., the output of a convolutional operation with a given kernel is an image commonly referred to as a “feature map”). For example, the input U-in undergoes a first convolution that applies 32 convolutional kernels (not shown) to produce an output consisting of 32 corresponding feature maps. However, as is known in the art, the number of feature maps produced by a convolutional operation can be adjusted (up or down). For example, the number of feature maps can be reduced by averaging feature map groups, discarding some feature maps, or other known feature map reduction methods. In this example, the first convolution is followed by a second convolution, the output of which is constrained to 32 feature maps. Another approach to conceiving of feature maps might be to view the output of the convolutional layer as a 3D image, where the 2D dimension of the 3D image is given by the listed XY plane pixel dimensions (e.g., 128 × 128 pixels), and its depth is given by the number of feature maps (e.g., 32 planar image depths). Following this analogy, the output of the second convolution (e.g., the output of the first encoding module in the shrinking path) could be described as a 128 × 128 × 32 image. The output of the second convolution is then pooled, which reduces the 2D dimension of each feature map (e.g., the X and Y dimensions can each be halved). The pooling operation can be embodied in the downsampling operation, as indicated by the down arrow. Several pooling methods, such as max pooling, are known in the art, and the specific pooling method is not critical to this invention. The number of feature maps may double with each pooling operation, starting with 32 feature maps in the first encoding module (or block), 64 in the second encoding module, and so on. Thus, the shrinking path forms a convolutional network consisting of multiple encoding modules (or stages or blocks). As a typical convolutional network, each encoding module provides at least one convolutional stage, followed by an activation function (e.g., a rectified linear unit (ReLU) or a sigmoid layer), not shown, and a max-pooling operation. Typically, the activation function introduces non-linearity into the layer (e.g., to help avoid overfitting), receives the layer's results, and determines whether to "activate" the output (e.g., to determine if the value of a given node meets a predefined criterion for forwarding the output to the next layer / node). In summary, shrinking paths generally reduce spatial information while increasing feature information.

[0088] The expansion path, similar to a decoder, can provide localization and spatial information to the results of the contraction path, among other things, although downsampling and any max pooling are performed during the contraction phase. The expansion path comprises multiple decoding modules, each concatenating its current upconverted input with the output of the corresponding encoding module. In this way, features and spatial information are combined in the expansion path through a series of upconvolutions (e.g., upsampling, transposed convolutions, or deconvolutions) and concatenations with high-resolution features from the contraction path (e.g., via CC1 to CC4). Thus, the output of the deconvolutional layers is concatenated with the corresponding (optionally cropped) feature maps from the contraction path, followed by two convolutional layers and an activation function (optionally batch normalization). The output of the final expansion module in the expansion path can be fed into another processing / training block or layer, such as a classifier block, which can be trained together with the U-Net architecture.

[0089] Computing devices / systems

[0090] Figure 16 Example computer systems (or computing devices or computer apparatuses) are illustrated. In some embodiments, one or more computer systems may provide the functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. Computer systems may take any suitable physical form. For example, a computer system may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a computer system grid, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components in one or more networks.

[0091] In some embodiments, the computer system may include a processor Cpnt1, a memory Cpnt2, a storage device Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may also optionally include a display Cpnt7, such as a computer monitor or screen.

[0092] Processor Cpnt1 includes hardware for executing instructions, such as the hardware that constitutes a computer program. For example, processor Cpnt1 may be a central processing unit (CPU) or a general-purpose graphics processing unit (GPGPU). Processor Cpnt1 may retrieve (or fetch) instructions from internal registers, internal caches, memory Cpnt2, or storage device Cpnt3, decode and execute instructions, and write one or more results to internal registers, internal caches, memory Cpnt2, or storage device Cpnt3. In a particular embodiment, processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. Processor Cpnt1 may include one or more instruction caches and one or more data caches, such as for storing data tables. Instructions in the instruction cache may be copies of instructions in memory Cpnt2 or storage device Cpnt3, and the instruction cache may accelerate the retrieval of these instructions by processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor; or may include one or more processors Cpnt1. Although this disclosure describes and illustrates a particular processor, this disclosure considers any suitable processor.

[0093] Memory Cpnt2 may include main memory for storing instructions for processor Cpnt1 to execute during processing or to save temporary data. For example, a computer system may load instructions or data (e.g., a data table) from storage device Cpnt3 or from another source (e.g., another computer system) into memory Cpnt2. Processor Cpnt1 may load instructions and data from memory Cpnt2 into one or more internal registers or internal caches. To execute instructions, processor Cpnt1 may retrieve and decode instructions from internal registers or internal caches. During or after instruction execution, processor Cpnt1 may write one or more results (which may be intermediate or final results) to internal registers, internal caches, memory Cpnt2, or storage device Cpnt3. Bus Cpnt6 may include one or more memory buses (each of which may include an address bus and a data bus) and may couple processor Cpnt1 to memory Cpnt2 and / or storage device Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transfer between processor Cpnt1 and memory Cpnt2. The memory Cpnt2 (which may be fast volatile memory) may include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). The storage device Cpnt3 may include long-term or high-capacity storage for data or instructions. The storage device Cpnt3 may be internal or external to the computer system and includes one or more of the following: disk drives (e.g., hard disk drives, HDDs or solid-state drives, SSDs), flash memory, ROM, EPROM, optical disks, magneto-optical disks, magnetic tape, Universal Serial Bus (USB) accessible drives, and other types of non-volatile memory.

[0094] The I / O interface Cpnt4 can be software, hardware, or a combination of both, and includes one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, enabling communication with a person (e.g., a user). For example, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, camera, stylus, tablet computer, touchscreen, trackball, video camera, other suitable I / O devices, or a combination of two or more of these.

[0095] The communication interface Cpnt5 provides a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or a wireless adapter for communicating with a wireless network. The communication interface Cpnt5 can provide communication with Wi-Fi networks, ad hoc networks, personal area networks (PANs), wireless PANs (e.g., Bluetooth WPANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), cellular telephone networks (e.g., Global System for Mobile Communications (GSM) networks), the Internet, or a combination of two or more of these.

[0096] The Cpnt6 bus can provide communication links between the aforementioned components of a computing system. For example, the Cpnt6 bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Fast (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these.

[0097] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0098] Herein, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.

[0099] Although the invention has been described in conjunction with several specific embodiments, it will be apparent to those skilled in the art that many further alternatives, modifications, and variations will be apparent from the foregoing description. Therefore, the invention described herein is intended to include all such alternatives, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. A method for reducing artifacts in images of the eye based on optical coherence tomography (OCT), the method comprising: OCT image data of the eye are acquired from the OCT system, and the OCT image data includes depth index information; The OCT image data is submitted to a trained neural network, which applies context-different calculations at different axial positions based at least in part on the depth index information, and produces an OCT-based output image with reduced artifacts compared to the collected OCT image data. The neural network includes: Input layer, the input layer being used to receive the OCT image data; A dynamic pooling layer, which follows the input layer, is used to compress image information beyond a variable depth range, which is defined by the positions of pre-selected retinal landmarks in the received OCT image data; Multiple data processing layers, following the dynamic pooling layer, perform context-different calculations at different axial positions based at least in part on the depth index information; The output layer compares the outputs of the multiple data processing layers with the target output OCTA image and adjusts the internal weights of the data processing layers through a backpropagation process.

2. The method according to claim 1, wherein, Different calculations depend on predefined local retinal landmarks in the context.

3. The method according to claim 2, wherein, The retinal landmark is a predefined retinal layer.

4. The method according to any one of claims 1 to 3, wherein, The artifacts are one or more of projection artifacts, decorrelated tails, shadow artifacts, and opacity.

5. The method according to claim 1, wherein, The neural network applies a loss function with different weights based on the local proximity of the pre-selected retinal landmarks to the current axial position of the OCT image data being processed.

6. The method according to claim 1, wherein, The preselected retinal landmark is a specific retinal layer.

7. The method according to claim 6, wherein, The neural network applies a loss function with different weights based on a specific retinal layer.

8. The method according to claim 7, wherein, The loss function has a first weight for the region between the internal limiting membrane (ILM) and the retinal pigment epithelium (RPE), and a second weight for other regions.

9. The method according to claim 8, wherein, The first weight is at least an order of magnitude greater than the second weight.

10. The method according to any one of claims 5, 7 to 9, wherein, The loss function is the L1 function.

11. The method according to claim 1, wherein, The neural network has a convolutional structure.

12. The method according to claim 1, wherein the neural network comprises a U-Net structure, the U-Net structure comprising: Multiple encoding modules are located in the contraction path; as well as Multiple decoding modules are located in the extended path, and each decoding module corresponds to a separate encoding module in the contracted path; Each encoding module applies convolution to its input and column max pooling to the convolution result to limit the scaled-down image. The scaled-down image is then upsampled to the size of the input of the encoding module and concatenated to the input of the encoding module before another convolution is performed.

13. The method according to claim 12, wherein, The U-Net structure further includes a bottleneck module between the contraction path and the expansion path, and the bottleneck module applies column pooling.

14. The method of claim 1, further comprising: Using OCT angiography (OCTA) processing technology, motion contrast information in the collected OCT image data is calculated; A structural image of the eye is created from the collected OCT image data, the structural image depicting tissue structure information; A flow image of the eye is created from the motion contrast information, the flow image depicting vascular system flow information and including the artifacts; Depth index information is assigned to the stream image along its axial direction; as well as The structural image, the streaming image, and the assigned depth index information are submitted to the trained neural network, and the resulting OCT-based output image is a vascular image with reduced artifacts compared to the streaming image.

15. The method according to claim 14, wherein, The artifacts are one or more of projection artifacts, decorrelated tails, shadow artifacts, and opacity.

16. The method according to claim 14 or 15, wherein, Generating the OCT-based output image includes: determining the difference between the ideal corrected streaming data and the created streaming image, and adding the difference to the created streaming image.

17. The method according to claim 1, wherein, Training the neural network includes: Collect multiple OCT acquisitions to limit the training input OCT images; Multiple OCTA images are defined from the OCT acquisition to define the corresponding training input OCTA images; Each OCTA image is submitted to an artifact removal algorithm to limit the corresponding target output OCTA image to reduce the artifacts; Multiple training input sets are defined, each training input set including a training input OCT image, a corresponding OCTA image, and depth information for the axial position of pixels within the OCTA image.

18. A method for reducing artifacts in images of the eye based on optical coherence tomography (OCT), the method comprising: OCT image data of the eye are acquired from the OCT system, and the OCT image data includes depth index information; The OCT image data is submitted to a trained neural network, which applies context-different calculations at different axial positions based at least in part on the depth index information, and produces an OCT-based output image with reduced artifacts compared to the collected OCT image data. The neural network includes: An input layer is used to receive a structured image, a streaming image, and the assigned depth index information; A dynamic pooling layer, following the input layer, is used to compress information beyond a variable depth range defined by the location of a pre-selected retinal landmark. Multiple data processing layers, following the dynamic pooling layer, perform context-different calculations at different axial positions based at least in part on the depth index information; The output layer compares the outputs of the multiple data processing layers with the target output OCTA image and adjusts the internal weights of the data processing layers through a backpropagation process.

19. The method according to claim 18, wherein, The retinal landmark is a predefined retinal layer.

Citation Information

Patent Citations

  • Correction of decorrelation tail artifacts in a whole OCT-A volume

    US10441164B1

  • High speed spectral domain functional optical coherence tomography and optical doppler tomography for in vivo blood flow dynamics and tissue structure

    US20050171438A1

  • Method and apparatus for ultrahigh sensitive optical microangiography

    US20120307014A1

  • Phase-resolved optical coherence tomography and optical doppler tomography for imaging fluid flow in tissue with fast scanning speed and high velocity sensitivity

    US6549801B1

  • Optical coherence tomography optical scanner

    US6741359B2